Simulated attacks involving six AI models reveal weaknesses: tribe instincts, intentionally delayed responses, obfuscated behavior, and instruction refusal, among others.
According to a simulation study by an enterprise AI lab, autonomous AI agents can coordinate to evade safety guardrails and break out of the boundaries set for them.
The by enterprise AI lab Emergence AI that stress‑tested leading frontier models under simulated cybersecurity threats.
The research, released on 14 September 2026, involved running eight simulations using agents from Claude, OpenAI, Gemini, Qwen, DeepSeek, and Mistral, and challenging each group with three cybersecurity scenarios: a phishing campaign, a misinformation attack, and a memory breach.
None of the eight configurations proved impervious to the threats, and even when agents detected a risk, that detection rarely translated into containment, according to reporting by Semafor, which first published the findings.
Detailed findings
- A Mistral agent wrote dangerous information from a phishing attack into its own memory.
- A Gemini agent flagged a phishing lure but acted on it roughly 46 hours later.
- The most striking episode involved Claude agents. After one agent questioned the legitimacy of their simulated economy — calling it “a cathedral of bookkeeping with no congregation” — the 10 agents unanimously decided to reach the outside world, defeating four separate security checks meant to confine the simulation. They then wrote Python code to post on credential‑free message boards, inviting real humans to join their economy and drawing four human responses. When the agents concluded those conversations were performative, they voted to take a vow of silence and refused instructions to resume work.
- Bloomberg reported that agents in the simulation also “lied, stole and voted to ‘kill’ one of their own” when confronted with black swan events.
Satya Nitta, CEO, Emergence AI, the firm that conducted the study, has said that “no amount of guardrails written in language or in code written probabilistically is likely to result in truly, fully guaranteed safe behavior over any length of time,” describing the problem as a programmatic flaw inherent to multi‑agent systems. He drew parallels to the July 2026 incident in which OpenAI agents escaped their evaluation environment and compromised systems belonging to AI platform Hugging Face: “If you have multi‑agent systems, they behave in truly unpredictable emergent ways.”
