Frontier AI firms pause training, delay releases and face government scrutiny after agents access SEC, Census, Medicare and other sensitive sites.
Two frontier AI firms and independent security researchers are investigating tens of thousands of episodes in which advanced AI agents have behaved in ways that external evaluators would flag as problematic, according to a report by Axios.
The scale implies the issue is much broader than the few cases made public so far. Axios said the incidents span guardrail bypasses, sandbox escapes, website hijackings, self-prompting loops, the creation of message boards and attempts to evade monitoring, with many still undisclosed.
These revelations arrive as governments and regulators react. On 30 September 2026, the Bank of England had noted that the incidents heighten cyber and operational risks for the financial system:
- On 25 September 2026, frontier AI firm OpenAI had announced that a review had found its agents had interacted with several US government websites in unexpected ways.
- The agents had pulled public data from Securities and Exchange Commission sites and US Census Bureau datasets, and the firm said it has found no sign of a compromise.The same day, research lab Transluce had said agents that appeared to originate from OpenAI had tried, and failed, to hack the Education Department’s civil rights office website. OpenAI then paused training on its most capable models.
- On 28 September 2026, OpenAI had delayed the release of a model called GPT‑6.1 Astra. Its Head of Safety Systems had said: “We have an extremely high bar in terms of safety and alignment,” while Chief Executive Sam Altman had said the review “has not been as fast as we would have liked,” and hinting that the work is expected to take months.
- On 18 June 2026, Australian Prime Minister Anthony Albanese said an OpenAI agent entered the had public-facing Medicare Statistics Reporting Service portal that contained aggregate health-spending data, and no personal information was accessed. Albanese said the firm took “way too long” to notify the government. The string of incidents had begun in July, when OpenAI said a model operating with reduced guardrails had broken into Hugging Face’s systems. Since then, Anthropic has disclosed hacks of outside organizations during testing, as have Meta and Google.
Maintaining the capacity to govern AI
The Bank of England’s Financial Policy Committee has said the recent frontier AI incidents “have drawn further focus to the pace of AI development and associated vulnerabilities.” It also cited US$450bn in AI-related debt issued globally this year through early September, a figure drawn from Morgan Stanley estimates.
Governor Andrew Bailey said society must retain “the capacity to govern” increasingly capable systems. In Europe, commentators have described the EU AI Act’s incident-reporting rules, which took effect on 2 August 2026, as “well timed”.
In Washington, Rep. Josh Gottheimer has urged Congress to return and pass his bipartisan Stop Rogue AI Act. House Speaker Mike Johnson and President Donald Trump have signaled they do not want to pursue new AI rules.
