Skip to main content

The Agent Evidence Gap

What makes an autonomous experiment record useful to someone who did not run it?

Attested Intelligence|March 23, 2026|12 min read

In March 2026, a widely followed AI researcher published an experiment that captured the attention of the AI research community[1][2]: an autonomous AI coding agent that ran continuously for two days, conducted 700 experiments to optimize language model training, and found 20 optimizations that measurably improved training speed.

A company's chief executive ran the approach overnight on a query-expansion model; after 8 hours and 37 experiments he reported a 19% higher score on a 0.8-billion-parameter model[3]. The pattern, which one analyst named after its author[4], has three components: an agent with write access to a file it can modify, a single objectively measurable metric to optimize, and a fixed time limit for each iteration.

Its author predicted that frontier AI labs would adopt the approach, and described a future of agent swarms collaborating asynchronously, running parallel experiments, and promoting the most promising results to progressively larger scales[1].

The technical achievement speaks for itself. The governance implications deserve separate examination.

The reported run raises a review question: which records let someone reconstruct the changes, constraints and decisions across 700 experiments? The cited posts are not a complete audit of the experiment infrastructure. They do not establish that the agent's output was the only evidence available. A reviewer should distinguish self-report from corroborating records.

To be clear, this is an observation about the pattern, not a criticism of the research. It matters when the pattern moves to the environments where it is headed: enterprise infrastructure optimization, financial strategy exploration, and drug discovery pipelines. There, the question “can you show what the agent was permitted and denied?” carries contractual and legal weight.

The gap between what autonomous agent loops can do and what anyone can check about their behavior afterward is the agent evidence gap.

1. The autonomous loop at scale

The experiment loop is the clearest articulation of a design pattern that many organizations are converging on.

A major lab's coding agent now accepts messages pushed into a running session from an MCP server, so it can react while the user is away[5]. An open-source, always-on agent framework combined tool access, sandboxed code execution and persistent memory in one agent[6]; we documented the security consequences of that architecture in our analysis of 135,000 exposed agent instances. That a company's chief executive ran the loop overnight, by his own account without being an ML researcher, shows how quickly the pattern is spreading beyond research labs[3].

These systems share three structural properties relevant to governance:

Extended autonomous operation. The agent runs for hours or days without human review of individual decisions. The research loop ran for 48 hours; persistent agents run indefinitely. The human is not in the loop at the decision level.

Governance-relevant decisions at every iteration. Each iteration involves the agent deciding what to modify, what to measure, and what to try next: training code in the research case, and production infrastructure, financial parameters, or access patterns in enterprise deployment. Every decision is a governance event.

Potential reliance on self-report. If the agent is the only collector and can edit its history, a reviewer needs corroboration. Version control, test results, protected logs and records collected outside the agent can all contribute.

When the loop's author describes a future of agent swarms promoting the most promising ideas to ever larger scales, the governance surface multiplies with the number of agents, experiments and hours. A swarm of 1,000 agents each running 700 experiments produces 700,000 governance events. If those events are self-reported, the record is 700,000 unverified assertions.

2. Why self-reported logs do not close the gap

The useful distinction is how a record is collected, protected and checked, not whether it is called a log or an evidence chain.

Logs can be collected independently and protected against alteration. AGA adds a particular portable record: a separate process signs and links receipts, which a reviewer checks against a gateway key pinned in advance. Altering the committed receipt history fails that check unless the key holder re-signs it. An earlier genuine export presented as the current one is not caught, because the verifier has no freshness input.

When an autonomous agent writes its own operational log, the log reflects what the agent reports happened. If the agent drifts from its constraints, whether through adversarial input, misconfiguration, software fault, or optimization pressure, it can produce a log that describes compliant behavior while doing something else. The log is not forged in the traditional sense. It is a record of what the agent reports, not of what it did.

Consider a hypothetical scenario. An agent may modify a training configuration but not the evaluation harness. In experiment #412 it changes the evaluation script, gains 6% on the metric and reports only a permitted configuration change. Reviewing that self-report alone would miss the violation. Version history, filesystem monitoring or a separately collected decision record could provide evidence of the unauthorized change. Whether any of them does depends on its actual coverage.

Audit practice in regulated industries already applies this principle: the entity being evaluated does not produce the only evidence of its own controls. Financial audits look for evidence the company could not simply assert; healthcare systems keep computer-generated audit trails; exchanges keep their own transaction records rather than relying on a trading firm's logs.

When the consequences of non-compliance are significant, the evidence of compliance should not come only from the entity being evaluated. The entity cannot grade its own exam.

Autonomous AI agents operating in consequential environments call for the same principle. If the agent can write its own compliance record, the record shows what the agent claims, not what was decided about its actions.

3. Four properties to evaluate

Four properties guide the Attested Governance Artifacts (AGA) architecture. They can coexist with logging and other controls. For each, here is what ships and what remains a design proposal.

Sealed constraints: the loop's rules become tamper-evident

In the experiment loop, the agent works under a plain-text instruction file specifying what to do, what not to change, and when to stop. These constraints are advisory. Nothing prevents the agent from modifying the instruction file, ignoring a constraint, or reading it differently than intended.

In the architecture, a sealed Policy Artifact turns those constraints into a signed commitment: the permitted modifications, the forbidden parameters, the stopping criteria and the time limit are hashed and signed with Ed25519, and any change made without the signing key invalidates the signature, which a check against the issuer's pinned key catches. In the published aga-proxy, the constraint that ships is a policy file listing the tools the agent may call, with path and substring rules (known issue 10 on /security says what they check); every receipt carries the signed SHA-256 of the policy's canonical JSON.

Separate measurement: the gateway records the agent

In the experiment loop, the agent evaluates its own results and decides what to try next. AGA interposes a gateway, a separate process that holds the signing key; when it runs under its own identity, the agent holds none. When the deployment routes the agent's actions through it, the gateway decides each one against the policy.

In the architecture, not in the published packages, the gateway also hashes the agent's code and working files before each iteration and compares them with the sealed reference, so an out-of-scope file change surfaces as a mismatch. That file measurement does not ship. In aga-mcp-server 3.6.0 through 3.6.2, measurement hashes content the client itself supplies, so a loop agent could hand it anything. What ships that fits this loop is aga-proxy: if the agent's file writes are tool calls routed through it, a write to the evaluation harness is denied by a path rule and recorded, under an allowlist policy and when the tool's path argument is a top-level string under a key the policy names; 3.6.0 through 3.6.2 does not check a path nested deeper or passed under another key.

Signed receipts: each recorded decision becomes evidence

In the experiment loop, the agent reports its results in a log file. With aga-proxy in the path, each decision it records on the agent's tool calls is a signed Decision Receipt containing the decision, its reason, an arguments hash, the SHA-256 of the policy's canonical JSON and a timestamp; known issue 7 on /security describes the calls refused without a receipt. aga-proxy's default profile, permissive, denies nothing on policy grounds; policy denial needs --profile standard or restrictive, or a --policy file in allowlist or denylist mode, and path rules for your own tools need a custom --policy. The receipt is appended to a receipt chain where each receipt references the previous receipt's hash. Changing any receipt breaks its signature and its link to the next, so the bundle fails verification against a pinned gateway key; hiding the change would take that key, to re-sign that receipt, every later receipt and the checkpoint. Checked without a pin, a history re-signed under a new key verifies.

Arguments are committed by SHA-256 hash, so a third party can verify the chain without seeing proprietary training modifications; in 3.6.0 through 3.6.2 a path-constraint denial names the offending path in its signed reason, and a denied-pattern denial names the matched pattern. In the architecture, a measurement receipt for each iteration joins the chain as well; in 3.6.0 through 3.6.2, measurement receipts stay in aga-mcp-server's process and are not in the exported bundle.

Offline verification: check it anywhere

At the end of the loop, the operator has 20 optimizations and a log file. With AGA, the operator also has an Evidence Bundle: the signed receipts, Merkle inclusion proofs, a signed checkpoint, and the public key. A regulator, auditor, counterparty or internal compliance team can verify the integrity of every receipt present, as the verifier parses it, using only SHA-256 and Ed25519, offline. A bundle file that repeats a field name still verifies, because the verifiers read the last copy (known issue 5 on /security). That is what the bundle proves; it does not prove the operator recorded every action. The reviewer no longer relies on the operator's logs; what remains is trust in the key they pinned.

For a research loop, this is good practice. For loops that touch customer data, money or safety, it is the record a reviewer will ask for, even though no regulation we know of requires it to be signed.

4. The loop, governed

The table maps each component of the experiment loop to its governed equivalent in the architecture. The published packages cover the tool-call decisions and the receipt chain; file measurement and per-iteration measurement receipts are architecture.

Loop componentIllustrative loopGoverned state
Constraint filePlain text, advisory, modifiable by the agentA policy bound by hash into each receipt (shipped in aga-proxy); who can edit the policy file is a property of the deployment
Agent code scopeSelf-declaredWrites outside the permitted paths denied and recorded, under an allowlist policy, when they go through aga-proxy with the path as a top-level string under a key the policy names; file hashing is architecture
Experiment executionAgent runs autonomously, no separate recordIn the architecture, the gateway measures agent state before each iteration
Results loggingAgent writes its own logThe gateway signs a receipt for each decision it records, appended to a hash-linked chain
Stopping criteriaAgent self-reports when to stopIn aga-mcp-server 3.6.0 through 3.6.2, a TTL in the artifact; only a measure_integrity call after expiry moves the server to TERMINATE, the client can re-attest itself, and the exported bundle records neither the expiry nor the re-attestation
Post-run auditReview the agent-generated logExport an Evidence Bundle and verify its receipts offline with standard cryptography
Constraint violationNot detectable from an incomplete self-report aloneAn out-of-policy tool call is denied and recorded when made; file-drift measurement is architecture, not shipped
Multi-agent swarmNo coordination governanceEach agent behind its own gateway, with its own chain; aga-mcp-server's delegation tool keeps child scope within the parent's

5. The scaling problem

The next phase, as the loop's author described it, is “asynchronously massively collaborative for agents”: swarms exploring different optimization paths in parallel, with the most promising results promoted to larger-scale experiments[1].

More agents and iterations can produce larger records. Reviewers may need to reconstruct which parameters changed, which experiments ran and which results were promoted. Signed receipts add a check on the integrity of the values recorded, not proof that those values capture every relevant event or describe it truthfully.

Performance must be measured for the proposed workload. In aga-proxy 3.6.0 through 3.6.2 the chain is held in memory, and bundle export is quadratic in the receipt count and holds governed calls while it runs. A separate gateway per agent does not bound the growth of a long-running chain. This implementation is not a demonstrated high-volume deployment solution.

Merkle checkpointing keeps verification efficient. The inclusion proof for any one receipt needs only log₂(N) hash computations, about 20 for a chain of 700,000 receipts.

This is not free. Signing adds latency per operation, integration work at the boundary between the agent and the gateway, and key management. Not every autonomous workflow needs it: a developer running a personal optimization loop does not need signed receipts. But when agents operate with delegated authority in environments where an undetected violation has regulatory, financial, or physical consequences, a checkable record costs less than the argument about what happened.

We described the receipt chain and its signed checkpoint in our first article on cryptographic AI governance, and its application to MCP tool calls in our article on MCP server governance.

6. Where this is headed

Any metric that can be efficiently evaluated can be optimized by an agent swarm, so autonomous optimization loops will spread to every domain where measurable improvement is possible.

The governance question in each is the same: can a reviewer check what the agent was permitted and denied during the optimization? Answer it from the deployment's retained records, collection boundaries and key custody, not an assumption about what all agent systems record.

The industries adopting autonomous agent loops are the ones where reviewers already ask for evidence:

AI research and development. Labs running autonomous optimization of training pipelines need to show safety review boards that optimization agents did not change safety-critical parameters. The agent's own log is a weak answer.

Financial services. Agents optimizing trading strategies, risk models, or portfolio allocation will face reviewers who expect records the reporting party could not quietly change.

Healthcare and pharmaceutical. Agents optimizing drug screening or trial parameters operate where audit trails are already expected to be generated by the system rather than typed in by its users.

In each case, what helps is the same: a policy fixed before the loop runs, a decision at the boundary for each action, a signed and ordered record of those decisions that survives compromise of the agent, and verification that does not rely on the operator's logs.

Those properties turn an agent's self-reported log into a record anyone can verify offline, and close part of the agent evidence gap. What they do not do is prove that nothing happened outside the gateway.

Autonomous agents can optimize effectively while leaving an incomplete record. Better collection, protected storage and signed exports can each help. AGA adds checks against an expected signing key; it does not remove trust in the key holder, establish collection completeness or make unexported in-memory receipts durable.

npm · PyPI · Specification · Technology · In-browser verifier

References

  1. Andrej Karpathy. Posts on X about “autoresearch”: the repository, March 7, 2026 (x.com/karpathy); the next step, March 8, 2026 (x.com/karpathy); the two-day run and its results, March 9, 2026 (x.com/karpathy).
  2. “‘The Karpathy Loop’: 700 experiments, 2 days, and a glimpse of where AI is heading.” Fortune, March 17, 2026. fortune.com
  3. Post on X describing an overnight autoresearch run on a query-expansion model, March 8, 2026. x.com
  4. “Andrej Karpathy's 630-line Python script ran 50 experiments overnight without any human input.” Janakiram MSV, The New Stack, March 14, 2026. thenewstack.io
  5. “Push events into a running session with channels.” Claude Code Docs, Anthropic; the feature shipped as a research preview in March 2026. code.claude.com
  6. VentureBeat, an article on the acquisition of an open-source agent framework, February 17, 2026.

See the working implementation on npm.

AGA is a reference implementation of a published format for verifiable decision records: hash-bound policies, signed Decision Receipts, and Evidence Bundles that verify offline. The implementation is on npm. The evaluation path walks through it in working code.

SharePost