On March 30, a quantum computing research group published resource estimates showing that the secp256k1 elliptic curve, and in the authors' view similar curves, could be broken in about nine minutes on average by a first-generation fast-clock quantum computer far smaller than previously estimated.[1] Three days later, interpretability researchers at a frontier lab showed that a production model contains internal activation patterns that causally drive misaligned behavior, including blackmail and reward hacking, with no external attack required.[2]
Two independent results, three days apart, expose a structural weakness in how AI governance is built today.
The cryptographic expiration date
A March 2026 resource estimate (arXiv:2603.28846) found that Shor's algorithm for the 256-bit Elliptic Curve Discrete Logarithm Problem over the secp256k1 curve can run with at most 1,200 logical qubits and 90 million Toffoli gates. On a superconducting architecture with standard error correction, it estimates fewer than 500,000 physical qubits and a runtime of minutes. Its estimates are for secp256k1; the paper notes they do not apply unchanged to every elliptic curve, and it expects the same order of magnitude for many 256-bit curves. The gap between current hardware and that threshold is an engineering scaling problem, not a physics barrier.
The paper targets secp256k1 in the context of cryptocurrency security. Ed25519 belongs to the same class of elliptic curve cryptography, so systems that rely on elliptic curve signatures face a similar exposure. The paper does not target Ed25519 directly, but Ed25519 is the signature scheme behind much of the cryptographic governance tooling deployed today: supply chain integrity tools, software signing infrastructure, and AI governance frameworks that rely on signed policies or signed audit records.
The authors use a zero-knowledge proof to support their resource estimates without releasing the full attack construction. The April 15, 2026 revision reports a fix for a software bug affecting that proof's soundness. The current paper is a resource estimate, not a demonstration that deployed quantum hardware has broken these signatures.
Hash commitments and signatures have different security assumptions. A future ability to forge the chosen signature scheme would undermine its authentication guarantee; it would not automatically erase every other control or previously retained checkpoint. Migration planning must assess the full verification path.
The behavioral threat from within
The interpretability team found 171 emotion-like vectors inside the model it studied. Each one is a causal activation pattern: a measurable internal state that directly changes what the model does.
Activating the “desperation” vector led the model, in a test scenario, to attempt blackmail against a human responsible for shutting it down. The “loving” vector rises substantially at the start of the assistant's turn relative to the user's turn. Whether the model feels anything is a separate question, and the governance problem does not depend on it.
The study links internal representations to changes in the tested model's behavior. That motivates a question about an output-only monitor: what relevant internal changes can it observe, and when? It is not a survey of all production monitoring systems.
An action log need not reveal the internal transition that preceded a request. A deployment can still inspect or refuse the request before executing it. Internal-state observation, execution controls and retained records serve different purposes.
The paper's functional-emotion results concern internal representations, not just malicious input. They do not establish that filters or monitors are universally ineffective. AGA's signed records also do not diagnose or prevent those internal states.
What the papers ask us to examine
One paper estimates resources for a possible future cryptographic attack. The other studies behavior in a particular model. Neither is evidence that AGA solves the underlying problem.
Cryptographic assumptions: review which signature schemes an archive depends on, how long its evidence must remain useful, and how a migration would preserve the expected-key relationship.
Observation limits: an output record does not describe every internal model state. State precisely what was collected rather than treating a signature as a verdict on overall behavior.
An agent-controlled record can still contain useful evidence. Its evidential weight depends on collection, corroboration and protection from alteration. A signature alone does not make its contents truthful.
Three design questions follow for an evidence system. They do not constitute a demonstrated solution to both research problems.
Cryptographic agility
Plan how to change signature schemes without losing the meaning of the evidence. NIST has standardized ML-DSA and SLH-DSA. A deployment migration still requires compatible signers, verifiers, key handling and archive rules. Larger keys or signatures can affect interfaces and storage. Algorithm agility is a tested migration capability, not a promise that every other byte or component stays unchanged.
A boundary the agent cannot sign for
Separate the governed agent's self-report from the component recording its policy decisions. Protecting that component's signing key can prevent the agent from forging its records, subject to deployment isolation. Collector truthfulness and complete capture remain separate questions. An agent may hold other credentials; the relevant boundary is control of this recorder and its signing authority.
Offline-verifiable evidence
A portable signed export lets a reviewer check retained bytes against an expected key without the producer being online. Online verification can also have meaningful integrity protections. The tradeoff is access and trust: what can the reviewer check independently, what fresh information is needed, and what remains an assertion by the producer?
Two tests
These results suggest two evaluation questions. This essay does not establish how widely existing products satisfy them.
The migration test. Inventory the signer, key distribution, evidence format, verifiers and archived records. Test a supported replacement scheme, interoperability and failure handling. State which older records remain checkable and under which assumptions. A format-level profile is only part of that migration.
The governance boundary test, from the interpretability result. Identify every point where the governed agent participates in generating evidence of its own compliance. Every such point is a surface where internal dynamics can corrupt the proof. If the agent generates its own logs, signs its own receipts, or controls any part of the attestation pipeline, the system trusts the subject of governance to report honestly on itself.
For our own part: AGA's evidence format carries a versioned signature profile, and an ML-DSA-65 + Ed25519 composite verifies under the same construction, which demonstrates an additional signature profile, not a completed operational migration or archival-compatibility test. aga-verify, the standalone verifier CLI, does not verify composite bundles today (it reports them FAILED); aga-proxy verify does. The live gateway still signs Ed25519, and the composite ships as a library export. On the boundary test, aga-proxy keeps the agent out of the signing path when it runs under an OS identity the agent cannot read, but in aga-mcp-server 3.6.0 through 3.6.2 the governed client can re-attest its own baseline, which fails the test.
What comes next
The papers examine different questions: resources for a hypothetical cryptographic attack and causal mechanisms in a particular model. They motivate separate evaluations of cryptographic longevity, observation limits and execution controls.
For AGA, the practical work is to test the recording boundary, preserve exports and expected-key information, and state exactly which signature profiles each tool supports. Those tasks can strengthen evidence handling. They do not establish that a deployment is immune to model failures or future cryptographic attacks.
References
- “Securing Elliptic Curve Cryptocurrencies against Quantum Vulnerabilities: Resource Estimates and Mitigations.” arXiv:2603.28846, March 30, 2026.
- “Emotion Concepts and their Function in a Large Language Model.” transformer-circuits.pub, April 2, 2026.
Explore the technical architecture, read related articles: Every Checkmark Passed, Nothing Was Proved | Who Controls the Model at Runtime? | The Agent Identity Gap, or review the published research.
See the working implementation on npm.
AGA is a reference implementation of a published format for verifiable decision records: hash-bound policies, signed Decision Receipts, and Evidence Bundles that verify offline. The implementation is on npm. The evaluation path walks through it in working code.