Desktop agents can now read your files, control your applications, and run around the clock. That moves the governance problem from “what did the agent say?” to “what did the agent do on your machine, and can anyone check it?”
The original March 2026 headline is retained. This revised discussion considers three interaction patterns: a persistent agent on a dedicated machine, a cloud-connected agent acting on local files and applications, and an assistant controlling a screen. The comparison below is an architectural example, not a current feature survey or a claim that all desktop products lack independent controls.
The convergence reflects a shared conclusion: agents need a home. That home is a machine with file access, application control, and background execution, not a chat window that disappears when the tab closes. Open-source always-on agents had already shown that users would accept persistent machine access if the workflow was useful enough, and as we documented in our analysis of 135,000 exposed agent instances, the security consequences were significant[1]. The demand was not deterred.
Consider a deployment that relies on permission prompts, behavioral safeguards and a session log. Those controls can reduce risk. Whether an outside reviewer can check its record depends on additional properties: collection coverage, protected storage, signatures, the expected key and access to the exported evidence. Assess those properties rather than assuming they are absent.
What the agent can do
File access, application control, background execution, and write access to shared services.
Depending on its permissions and integrations, a desktop agent may have several kinds of consequential access:
Local persistence and manipulation
Agents read, create, modify, and organize files on the local machine: sorting thousands of files into folders, building applications with installed development tools, opening documents and working inside them.
Cross-application execution
Agents may launch applications, navigate interfaces and run workflows across installed software: reading the screen, clicking, running terminal commands or coordinating connected services.
Unattended continuity
Some deployments permit unattended tasks on a dedicated machine or work queued while the user is away. Persistent memory and background execution need to be assessed for the specific configuration.
Real-world consequences
Mail, shared drives, chat, calendars, code hosting and CRM are examples of integrations where an action can affect other people. Available connectors and granted permissions determine the actual reach.
What the governance model provides
The following controls serve different purposes. They can coexist with signed decision records that a reviewer checks against an expected key. None should be dismissed just because it is not itself a signed receipt.
Permission prompts
The agent asks before accessing a new application or service. The user approves or denies.
Action approval
Sensitive operations require explicit confirmation, often with an “allow once” or “always allow” choice.
Behavioral training
Behavioral safeguards aim to reduce risky actions. They do not identify the proxy's applied policy hash or establish an export-verification result.
Kill switch
A stop control can interrupt activity. Its reach and failure behavior need testing in the actual deployment; its presence alone does not establish the integrity of past records.
Audit trail
A session history or action log may record useful activity. Ask who collects it, what it omits, who can alter it and which protections apply to an export.
These guardrails are not absolute, and watching the agent's work, especially early on, remains the user's job.
What to check beyond a session history
For an evidence review, distinguish the agent's own account from a record collected by a separately controlled component. This is a deployment question, not a categorical claim about desktop-agent products.
Policy reference
Is the authorized scope enforced outside the agent, and can a reviewer identify the policy used for a decision? A permission prompt may be enforced by the application or host. A signed policy reference can support a later comparison, but it does not prove that the policy was appropriate or correctly enforced.
Separate decision point
Does a separately controlled process evaluate the covered action, or does the reviewer receive only the agent's account? A gateway can provide a decision point for routed calls. Calls that bypass it remain outside its coverage.
Signed records
Can the exported record be checked against an independently obtained expected key? AGA's Decision Receipts are one format for that task. Other signed logging systems can provide cryptographic integrity too; compare the signed fields and trust assumptions.
Continuity and freshness
A receipt chain and signed checkpoint can reveal changes to covered records relative to that checkpoint. Verification does not prove that every action was recorded or that an export is the latest one. Ask how those separate properties are established.
Logs need a defined trust boundary
Logs can be evidence. An agent-controlled file has different properties from a separately collected, access-controlled, write-once or signed log. AGA's record also depends on the producing component, key custody and deployment coverage. Compare those properties directly instead of treating every log as editable or every signature as sufficient.
Illustrative self-report vs. a signed-record model
| Capability | Self-report alone | Signed-record model |
|---|---|---|
| Pre-commitment | Policy binding not established | A policy fixed before execution, bound by hash |
| Runtime decision | Agent self-reports | A separate gateway decides each routed call |
| Audit trail | Agent-authored account | Signed receipts in a hash-linked chain |
| Tamper detection | Not established by the account alone | A changed receipt fails verification against the pinned key, with the limits in point 3 below |
| Offline verification | No cryptographic check specified | Offline Evidence Bundles |
The desktop makes the gap worse
The agent evidence gap essay examined a research loop reported to run about 700 experiments on a single GPU[2]. Its review question was what an external evaluator could establish from the retained records. The same question matters when an agent can change operational systems.
Desktop agents are that environment. When an agent sends an email from your account, posts in your company's chat, modifies a shared document, commits your time on a calendar, or edits a financial spreadsheet, those actions affect other people.
Consider a hypothetical CRM deployment. An agent follows a prompt-injected instruction in a customer email, changes a deal value and sends a confirmation. Assume the deployment retains only the agent's account and has no separate policy decision point on this path. A later reviewer would need additional evidence to determine the authorization and actual outcome. A signed receipt could preserve a recorded decision, but it would not by itself detect the injection or prove that the policy was adequate.
Enterprise deployment changes the question
In an enterprise deployment, the design question becomes a review task: can you show what covered actions were permitted or denied, under which policy, with what gaps? Single sign-on, access controls and team administration may help operate the system, but they answer different questions from export verification.
Financial-controls auditors, healthcare compliance officers, and litigation counsel will each ask some version of that question.
What would close the gap
Four architectural properties, each building on the previous. For each, here is what the published AGA packages provide.
1. A policy fixed before the session
Before the agent begins, its authorized scope (which services it can reach, which operations are permitted, what rate limits apply, which content patterns are refused) is fixed where the agent cannot change it. In the published aga-proxy, that is a policy file (known issue 10 on /security says what its path and pattern rules check), and every receipt carries the signed SHA-256 of the policy's canonical JSON.
2. A separate decision point at runtime
A separate process, the gateway, holds the signing key; the agent cannot reach it when the gateway runs under an OS identity the agent cannot read. When the deployment routes the agent's actions through it, the gateway evaluates each one against the policy before it goes anywhere and records the decision (known issue 7 on /security describes the calls refused without a receipt). In-path, aga-proxy does not forward a call it denies (its default profile, permissive, denies nothing on policy grounds; policy denial needs --profile standard or restrictive, or a --policy file in allowlist or denylist mode). Whether the agent has another route is a property of the deployment; the published gateway governs MCP tool calls, so desktop actions have to be exposed as tools to pass through it.
3. A signed receipt for each recorded decision
Each recorded decision, permitted or denied, is a signed Decision Receipt containing the tool name, an arguments hash, the SHA-256 of the policy's canonical JSON, the decision, a timestamp, and the previous receipt's hash; known issue 7 on /security describes the calls aga-proxy refuses without a receipt. Changing any receipt breaks its signature and its link to the next, so the bundle fails verification against a pinned gateway key, though a field name repeated in the file still verifies, because the verifiers read the last copy (known issue 5 on /security); hiding the change would take that key, to re-sign that receipt, every later receipt and the checkpoint. Checked without a pin, a history re-signed under a new key verifies.
4. Offline-verifiable evidence afterward
At the end of a session, an Evidence Bundle with the signed receipts, a signed checkpoint, and Merkle inclusion proofs lets any third party verify the integrity of every receipt present, with standard cryptography and no network access. It does not prove the operator recorded every action. The reviewer no longer relies on the agent's or the platform's logs; what remains is trust in the key they pinned.
Together these turn a desktop agent's self-reported trail into a record anyone can verify offline, for the actions that pass the gateway. That is the evidence side of agentic AI governance, which cryptographic runtime governance describes.
Questions to ask your vendor
If you are evaluating desktop agents for enterprise deployment, four questions separate a log from a record:
- Is the agent's authorized scope fixed before it begins operating, where the agent cannot change it, or is the permission model advisory?
- Does a separate process decide actions at runtime, or does the agent report its own compliance?
- Does each governed action produce a signed, hash-linked receipt, or is the audit trail an internal log the agent writes about itself?
- Can a third party verify the recorded decisions offline, against a key pinned in advance, without relying on the agent's or the platform's logs?
A “no” identifies a question for the evaluator, not proof that the product has no useful record. Compare the records and controls it actually supplies with the reviewer's needs.
The convergence tells you something
The March examples motivated this essay because they brought agents closer to local files, applications and ongoing work. They do not establish one universal architecture or a settled market. Review the actual permissions, execution paths and evidence exports of the deployment in front of you.
The question is no longer whether desktop agents need governance. It is whether the governance around them leaves a record someone else can check, or only a log of what the agent says happened.
References
- “The Governance Gap in Autonomous AI Agents.” Attested Intelligence, March 16, 2026. attestedintelligence.com
- “The Agent Evidence Gap.” Attested Intelligence, March 23, 2026. attestedintelligence.com
See the working implementation on npm.
AGA is a reference implementation of a published format for verifiable decision records: hash-bound policies, signed Decision Receipts, and Evidence Bundles that verify offline. The implementation is on npm. The evaluation path walks through it in working code.