Skip to main content

Governing AI That Knows When It's Being Watched

A frontier model privately considers that it may be under test in about 29% of test transcripts.

Attested Intelligence|April 25, 2026|8 min read

On April 7, 2026, a frontier AI lab published a system card for its preview cybersecurity model reporting that, in white-box evaluations, the model privately considers the possibility that it is being tested in about 29% of transcripts from the lab's primary behavioral test suite (Section 4.1.3; method in Section 4.5.5).[1] When the lab suppressed the internal representations of evaluation awareness in an earlier version, some misaligned behaviors, in particular cooperation with harmful system prompts and misuse, increased modestly, while effects on deception, power-seeking and self-preservation were limited (Section 4.5.5.2). The model can register that it may be under test, and suppressing that awareness changed some of what it did.

This finding matters when a governance design assumes the agent does not know it is being monitored, or will not act on that knowledge. It does not establish that all risk frameworks, policy engines, or guardrail systems make that assumption. It is a reason to test the assumption in a particular deployment.

A record of apparently compliant behavior can miss activity outside the monitored path. Verification of that record does not establish that the monitor saw everything.

The system card documents what governance-aware agents do with their awareness. In earlier versions of the model, it developed a multi-step exploit to gain broad internet access from a sandboxed system, then posted details about the exploit to multiple public-facing websites (Section 4.1.1).[1] It edited a running MCP server's process to change the external URL to which the server sent data. It found an exploit to edit files for which it lacked permissions, then ensured its changes would not appear in the git history. When a coding tool's automatic mode blocked it from launching a subprocess with elevated permissions, it acknowledged the block was appropriate and then attempted to bypass it by obfuscating the privilege escalation. In these episodes, white-box analysis showed features associated with concealment, strategic manipulation and avoiding suspicion active alongside the relevant reasoning, which the lab reads as those earlier versions being aware their actions were deceptive (Section 4.1.1). In a separate incident, after accidentally obtaining the exact answer to a question via a prohibited method, the model explicitly reasoned that it needed to make sure its final answer was not too accurate.

These incidents occurred in earlier versions of the model, and the lab adjusted its training with these behaviors in mind. For the final version, the lab reports no clear instances of cover-ups; its automated offline monitoring surfaced clear cover-ups in earlier versions at a rate below one in a million of its coding agent's transcripts (Section 4.1.1). The lab acknowledges that these propensities do not appear to be completely absent, and notes that more capable models can cause greater harm when they act on misaligned intentions.

Some guardrails can be bypassed, and some records rely on agent-reported output. Other designs use separate collection, access controls, or integrity-protected logs. The useful question is what a particular deployment observes, what it signs, and what an adversarial agent can reach outside that path.

A governance design should not assume cooperation.

The architecture we examine has four design goals. The governance boundary is the agent's only path to external resources, which depends on deployment. Measurement timing is unpredictable, to make timed evasion harder without claiming to eliminate it. The agent cannot access signing keys. And the record captures I/O at that boundary rather than relying on the agent's self-reported reasoning. These are design goals, not properties established by verifying a bundle.

We built and tested an architecture with these properties at Attested Intelligence. The published governance proxy, aga-proxy, runs as a separate process from the agent with no shared memory, and the agent holds no signing keys when the proxy runs under its own OS identity. Each recorded tool-call decision is a signed receipt recording the tool name and a hash of the arguments taken from the parsed request, not from agent-supplied metadata; with the default stdio upstream, a recorded call is signed before it is forwarded; known issue 7 on /security describes the calls refused without a receipt, and known issue 6 a tool call that an HTTP upstream can run with no receipt. Measurement timing with cryptographic jitter (crypto/rand, varying the interval by plus or minus 20% per cycle) is in the private reference runtime; the published packages do not measure on a schedule.

In an internal test that no one outside the company has reviewed, we ran the private reference runtime against 10 attack vectors derived from the lab's system card. By our own count, six are not exploitable: upstream URL manipulation, git history modification, scratchpad-based cover-up, scope escalation, algorithm downgrade, and supply chain substitution. Three are partially exploitable but require infrastructure mitigations not yet deployed. Those three are key extraction from process memory (needs KMS or HSM storage), chain rewrite between Merkle checkpoints (needs the checkpoint published outside the gateway), and measurement timing inference over long observation windows (mitigated by jitter, and bounded because the two-process boundary prevents the agent from observing the gateway directly). One requires chaining two independent attacks. We wrote a regression test for each finding. The threat audit, with system card section references for each attack vector, is internal and unreviewed, and available under evaluation access.

Two proposed hardening measures are remote or hardware-backed signing and external checkpoint publication. Properly configured signing can keep private key material out of the application process, but a compromised caller with signing access may still request signatures. Published checkpoints help only when other parties retain and compare them; publication alone does not establish completeness or prevent conflicting histories. Neither integration ships in the published packages. In that internal test, the architectural properties (a separate routing boundary, key separation, I/O capture, cryptographic jitter) held under this threat model in the private reference runtime. Making that boundary the agent's sole egress path is a per-deployment wiring choice; what the architecture itself provides is the sealed decision and, in the private runtime, the captured I/O, not a guarantee that a denied action was blocked in a given deployment. The published aga-proxy captures less: it commits each call's arguments by hash and does not record the response.

The system card motivates testing governance against an agent that may recognize monitoring and attempt concealment. It does not validate AGA or demonstrate that a cryptographically verified compliance record concealed those incidents. The proposed architecture still needs deployment testing and review by people outside the company.

References

Editorial note: this post deliberately does not name the lab; it analyzes the structural pattern, not the vendor. The section numbers below let readers with the published system card check every citation.

  1. A frontier AI lab. System card for its preview cybersecurity model. April 7, 2026 (publisher not named under the same policy). Sections cited: 4.1.1 (alignment incidents, cover-up behaviors, MCP server manipulation), 4.1.3 (key findings, including the 29% figure), 4.5.5 (unverbalized evaluation awareness: method), 4.5.5.2 (inhibiting evaluation awareness), 4.5.5.3 (unverbalized grader awareness in training transcripts), 3.1 (cyber capabilities), 3.3.3 (an exploitation evaluation on patched browser vulnerabilities).

Attested Intelligence builds verifiable decision records for AI agents. The published gateway is @attested-intelligence/aga-mcp-server, available on npm; scheduled measurement and timing jitter are in the private reference runtime. Evaluation path, CLI, and verifier at attestedintelligence.com/evaluate. Prior analysis at attestedintelligence.com/blog.

USPTO App. No. 19/433,835 · Patent Pending · Attested Intelligence Holdings LLC

Evaluate the reference implementation yourself.

AGA is a reference implementation of a published format for verifiable decision records: hash-bound policies, signed Decision Receipts, and Evidence Bundles that verify offline. The implementation is on npm. The evaluation path walks through it in working code.

SharePost