Skip to main content

AI agent security

Reference revised September 28, 2026.

Decision evidence has a boundary.

AGA can support tool-call authorization and later review of recorded decisions. It does not prevent every agent attack, prove complete interception or replace a deployment security review.

Scope

Four risks, different controls.

Supply chain injection
Compromised dependencies, poisoned tool configurations, and malicious MCP server responses can subvert agent behavior before execution even begins. A configuration baseline can help identify some changes, but it does not make every injection detectable. The published AGA packages do not address this: aga-proxy forwards the upstream response untouched and hashes no component.
Silent drift
Periodic audits can miss changes between checks. Detecting behavioral drift requires measurements and thresholds appropriate to the deployment. The published AGA packages do not detect this: aga-mcp-server compares only content its caller supplies, on request.
Tool-call authorization gaps
An agent may reach a tool using credentials broader than its assigned task. Deploy-time permission checks alone do not establish which policy was evaluated for a later call. Runtime access controls and decision evidence address different parts of that problem.
Forensic evidence gap
If decision logs are editable by the same actor whose behavior is under review, a reviewer needs additional evidence to assess their integrity. Signed records are one option; protected logging systems may also supply integrity controls. Neither automatically proves execution or complete capture.

Deployment boundary

Separate processes. Separate authority.

  1. 01 · Agent

    Requests a call

    The agent handles reasoning and tool intent. Its identity must not be able to read the gateway’s signing key.

  2. 02 · Gateway

    Records a decision

    The gateway checks routed calls against its configured policy and signs the decisions it records. Coverage gaps remain.

  3. 03 · Reviewer

    Checks the export

    A pinned-key check detects changes to signed values. It does not establish that the tool ran or no alternate route existed.

A separate process alone does not protect the key or force traffic through the gateway. Other guardrail and authorization systems may also use separate identities. In runtime 3.6.0 through 3.6.2, an environment-supplied key is inherited by a stdio upstream; do not assume process separation has resolved custody.

aga-proxy does not forward a call it denies. Response actions beyond that boundary need deployment-specific integration. Receipts not yet exported remain in gateway memory. See the implementation details and architecture definition.

MCP implementation

Policy, capture and export.

Fix the tool policy before execution
The approved tool names and invocation constraints go in a JSON policy file that aga-proxy evaluates. The SHA-256 of the policy's canonical JSON is signed into every receipt, so a changed policy shows up as a different hash and a reformatted file does not.
Gateway evaluates each routed call
The gateway process intercepts each MCP tool call routed through it, at the protocol level, before it reaches an external resource. Each call is checked against the policy: a tool allowlist, rate limits, a path-prefix check on named keys, and substring deny patterns on top-level string arguments (known issue 10 on /security), and each decision it records is signed into a receipt. Known issues 6 and 7 on /security describe gaps in receipt coverage. Under the default profile, permissive, the check denies nothing on policy grounds; policy denial needs --profile standard or restrictive, or a --policy file in allowlist or denylist mode. In-path, aga-proxy does not forward a call it denies; whether the agent has another route to the tool is a property of the deployment.
Signed receipt per recorded decision
Each recorded decision is a signed receipt containing the tool name, arguments hash, timestamp, and SHA-256 of the policy's canonical JSON. Receipts are hash-linked into a tamper-evident chain. Known issues 6 and 7 on /security describe coverage gaps; the chain does not establish that every call was captured.

The policy file is unsigned and is not included in the bundle. A holder of the actual policy can recompute its canonical hash and compare it with a tools/call receipt. The public aga-mcp-server’s exported bundles carry no policy reference.

Incident review

Check the record, not an implied outcome.

A bundle contains the receipt chain, Merkle inclusion proofs and signed checkpoint. A responder can verify these locally without the producing service. Custody records, a trusted expected key, retained checkpoints and freshness comparisons remain separate responsibilities.

A "(passthrough)" receipt records an unevaluated method sent toward an upstream, not proof of delivery. Its policy hash does not mean that policy was applied. Duplicate JSON field names are accepted with the last value retained; inspect the parsed output rather than assuming every displayed value is authenticated.

Framework context

Limited control contributions.

These are scoped mappings, not certification, legal advice or independent endorsement. Read the cited controls and exclusions.

OWASP Top 10 for Agentic Applications
A signed record of a tool-call decision can support review of tool misuse and unauthorized actions. It does not prevent prompt injection or goal hijacking.
NIST NCCoE AI Agent Identity
Receipts are attributable to the gateway's key when that key is pinned. There is no per-agent or delegation identity in the receipt.
CoSAI
Signed receipt chains serve the auditability concerns in the Coalition for Secure AI's MCP guidance for governed tool calls; per-category status is on /standards.
NIST AI RMF
In part: GOVERN 1.4, MEASURE 2.8, MEASURE 3.1, MANAGE 4.1 and MANAGE 4.3, quoted on /standards. The published packages do not measure on a schedule. Alignment, not certification.
EU AI Act Article 12
Receipt chains can form part of the automatic logs Art. 12 requires of high-risk systems, for tool-call decisions only. Articles 9 and 15 are not addressed (/standards).

Security questions.

What is AI agent security?

AI agent security is the practice of securing autonomous AI agents against manipulation, drift, unauthorized tool use, and forensic evidence gaps. It encompasses identity, authorization, runtime governance, and audit: checking what agents were permitted to do, and keeping a verifiable record of what was decided.

How does AGA secure MCP tool calls?

aga-proxy evaluates each MCP tools/call routed through it against a policy file, and signs a permit-or-deny decision into a receipt that carries the SHA-256 of the policy's canonical JSON (known issue 7 on /security describes the exceptions). In-path it does not forward a call it denies; whether the agent has another route to the tool is a property of the deployment. The agent lacks the signing key only when the gateway runs under an OS identity whose secrets the agent cannot read. aga-proxy's default profile, permissive, denies nothing on policy grounds; policy denial needs --profile standard or restrictive, or a --policy file in allowlist or denylist mode.

How is this different from application-level guardrails?

Guardrails that share the agent's trust boundary can be bypassed if that boundary is compromised. Some application-level controls already use separate identities or enforcement services. AGA's gateway is a separate process that holds the signing key. Key isolation requires an OS identity whose secrets the agent cannot read; process separation alone does not provide it. In aga-mcp-server 3.6.0 through 3.6.2, the governed MCP client can re-attest its own baseline through attest_subject, and the exported bundle does not record that change. See the known issues on /security.