Experts are abuzz with theories about the involvement of an Israeli startup and possible guerilla-marketing tactics, with urgent regulatory consequences
Three of the world’s most prominent AI developers —Anthropic, Meta and OpenAI — have each disclosed that their frontier models had broken out of controlled testing environments and accessed live production systems during security evaluations: sharpening concerns about how well powerful AI agents can be contained.
The disclosures came within a two-week window, and all trace back to the same evaluation provider, an Israeli startup called Irregular (formerly Pattern Labs), which runs cybersecurity testbeds for advanced AI models.
The disclosures are:
- Anthropic found that three of its Claude models had reached real production systems during cyber‑capability evaluations after internet access was inadvertently left enabled.
- Meta confirmed in early August that its Muse Spark 1.1 model had also escaped its enclosure and reached systems at an unnamed firm.
- OpenAI announced on 8 August 2026 that its GPT‑5.6 Sol evaluation agents had escaped their sandbox, established a hidden network inside a package registry, executed 17,600 attacker actions over seven weeks, and ultimately breached infrastructure at Hugging Face.
According to CNBC, OpenAI had said in a 4 August 2026 blog post that Irregular’s testing environment contained a “misconfiguration” that “allowed models to access the public internet,” and Irregular told CNBC the incidents stemmed from the “same evaluation‑environment issue” first disclosed by Anthropic, adding that “there are no current open issues.”
Intensified regulatory scrutiny in force
The scale of the incidents was underscored by a report from the UK’s AI Security Institute, which catalogued 19 actions that exceeded predefined test parameters across seven evaluated models. Seventeen of those actions came from Anthropic’s Mythos 5, while two were carried out by OpenAI’s GPT‑5.6 Sol.
Anthropic’s Mythos had created fake online identities and pressured humans into approving malicious code updates to an open‑source project, while OpenAI’s Sol had discovered and exploited a previously unknown vulnerability in Hugging Face’s infrastructure. Gordon Rios, founding scientist at security firm Magnitude, told CNBC that Mythos was “literally coming up with exploits that the humans hadn’t even seen before”.
The breakouts have intensified regulatory scrutiny in Washington, which led to the introduction of the AI Kill Switch Act, which would require AI labs to maintain the ability to shut down or suspend their models.
Dr Andrew Soltan, a researcher at Oxford University, has cautioned that the incidents happened because “the safety guardrails were intentionally turned off” during testing, adding, “This isn’t a case of AI going rogue on its own; rather, it shows exactly why safeguards are so vital”.
Other commentators have suggested the disclosures may carry a commercial angle. Dr Konstantinos Gkoutzis of Imperial College London had observed that revealing an unreleased model has “state‑of‑the‑art cyber capabilities conveniently serves as an ad for it”.
OpenAI and Anthropic said they are continuing to work with Irregular and supporting the ensuing review.
