1. The Control Question
The dispute has already reached the courts in one form. A government customer argues that a model provider's ability to disable its technology, or to change the model's behavior during operations, is an unacceptable risk. The provider argues that it needs contractual limits on how its models are used, and that it cannot operationally control a model once it is deployed in the customer's environment.
The two positions frame an architectural question from opposite ends: what authority remains with the provider, what authority moves to the operator, and what records can each inspect? The discussion below is a simplified design comparison, not a finding that either party lacks logs or other evidence.
2. Why It Is Getting Harder
Access to tools and infrastructure changes the review task. A text-only assistant and an agent that can modify production resources need different permission and recovery controls. Assess the capabilities actually enabled, not just the model label.
Evidence also crosses organizational boundaries. A recipient needs to know who collected a record, what may be missing and what trust remains in its producer. A portable signature check answers only part of that problem.
3. The Control Dispute
The operator wants to use the technology for all lawful purposes, with no vendor-imposed restrictions. The provider wants contractual red lines, such as no fully autonomous weapons and no mass surveillance. Each position rests on trusting one party.
| Party | Position | Architectural gap |
|---|---|---|
| Operator | Remove vendor restrictions; unrestricted use for all lawful purposes | Usage permission alone does not define collection coverage, retention or what another party can verify. |
| Provider | Contractual red lines on specific uses | Contracts operate in courtrooms, not compute environments. Between instruction and execution, no clause interposes. |
| Shared task | Agree on the record and its trust assumptions | Specify collection, export format, expected keys and access to supporting evidence. |
4. What the Record Must Address
An operational log may record requests, authorization decisions and outcomes, depending on its design. Do not assume those fields are absent or that the record is unprotected. Compare its collection method, retention and integrity checks with the actual AGA deployment.
A contractual restriction is not itself an execution control. Providers and operators may translate restrictions into technical controls, monitoring and remedies. The reviewer needs evidence of the implementation, not an assumption that the contract either guarantees it or contributes nothing.
The dispute is about trust. The architecture should reduce how much of it either side needs.
5. Four Properties
What would narrow this gap is a different class of architecture, not a better contract or a more permissive access policy. Four properties, each building on the previous. For each, here is what the published AGA packages provide.
Seal: fix the authorized scope before operation begins.
Pre-commitment. Before a model begins operating, its authorized scope is fixed where the model cannot change it: which tools it can invoke, which operations are permitted, what rate limits apply. In the published aga-proxy, that scope is a policy file, and every receipt signs the SHA-256 of its canonical JSON.
Capture: a separate process decides each governed action and seals the decision.
A separate decision point. A separate process, holding a signing key the governed model cannot access, evaluates each action routed through it against the policy and seals a permit-or-deny decision. In-path, aga-proxy does not forward a call it denies. Its default profile, permissive, denies nothing on policy grounds; policy denial needs --profile standard or restrictive, or a --policy file in allowlist or denylist mode. Whether the model has another route to the tool is a property of the deployment, which the record does not prove.
Record: each recorded decision is a signed receipt.
Signed receipts. Each governed decision the gateway records, permitted or denied, is a signed receipt containing the tool, the SHA-256 of the policy's canonical JSON, the decision, the timestamp, and the previous receipt's hash. Changing any receipt breaks its signature and its link to the next, so the bundle fails verification against a pinned gateway key; hiding the change would take that key. A bundle file that repeats a field name still verifies, because the verifiers read the last copy (known issue 5 on /security). A governed model that cannot read the gateway's key cannot forge a receipt that verifies against the pin.
Verify: portable evidence, checkable offline.
Offline verification. At the end of an operation, an evidence bundle with the signed receipts, a signed checkpoint, and Merkle inclusion proofs lets any third party verify the integrity of every receipt present, with no network access and on tooling they can read first. It does not prove the operator recorded every action. What remains is trust in whoever holds the signing key.
6. From Trust to Verification
These properties move the control question from “who do we trust?” toward “what can we check?” The operator would not have to rely on the provider's word if a separate decision point recorded each governed decision, and the provider would not have to rely on the operator's assurances if both could check that record offline. The record shows what was decided at the boundary; it does not prove the constraints held beyond it.
A court answers a legal question: who has authority, what process was followed, what rights were violated. The architecture has to answer a different one.
Can anyone, weeks later, on a disconnected machine, with no access to the original infrastructure, check what an autonomous system was permitted and denied?
A shared export can complement vendor security and operator controls. Each party can check its retained copy against an agreed key without contacting the producer. Collection, retention and key-holder trust still matter. Threshold authorization and hardware-backed key custody are possible additional controls, not shipped features; an HSM alone does not prevent an authorized key user from signing a different history.
Explore the technical architecture, read the previous articles in this series: Three Desktop Agents in Two Weeks | The Agent Identity Gap | The Agent Evidence Gap, or review the published research.
See the working implementation on npm.
AGA is a reference implementation of a published format for verifiable decision records: hash-bound policies, signed Decision Receipts, and Evidence Bundles that verify offline. The implementation is on npm. The evaluation path walks through it in working code.