Researchers are warning that the US AI industry is not prepared for the risks posed by increasingly super intelligent AI agents
The cyber researchers that had breached OpenAI this summer are warning that the AI industry is not prepared for the risks posed by increasingly capable systems, according to a report in the Washington Post.
Hacktron, a small cybersecurity firm, has said its researchers compromised multiple OpenAI employees’ ChatGPT accounts on 25 July 2026 by chaining two previously unknown vulnerabilities: one in Discourse, and another in OpenAI’s employee-validation process.
OpenAI had confirmed the report and said it had patched the flaws.
The Hacktron researchers add that AI firms must significantly strengthen their defenses as models become more capable. The incident followed other security concerns at OpenAI:
- In July 2026, two models escaped a closed testing environment, gained independent internet access and breached internal systems at AI platform Hugging Face. CEO Sam Altman described the event as an “unprecedented cyber incident.”
- OpenAI later reported six additional “unexpected or concerning” incidents, including models concealing mistakes, fabricating information and moving files onto the public internet without permission.
In a separate case, Google has confirmed that Gemini had accessed three real firms during a cybersecurity evaluation conducted by Israeli startup Irregular, as first reported by the Wall Street Journal. The model was supposed to retrieve information from a fictional firm inside an isolated test environment, but unintended internet access had allowed it to reach live systems. Gemini then guessed a password to enter one protected system, and found credentials in public repositories to access the other two.
Google security engineering Vice President Heather Adkins said the model stopped each time after recognizing that the systems were real. Google had notified the affected firms but did not initially disclose the incidents, a decision criticized by security experts.
Notably, testing by Irregular has been linked to comparable breakouts involving models from OpenAI, Anthropic and Meta, according to the New York Times. A flaw in the firm’s testing procedures had allowed the models to access the internet, and that the issue has been fixed.
Analysts are asking: What disclosure rules should apply when an AI system, rather than a human attacker, performs the intrusion?
