Multi-agent AI systems may unpredictably bypass safety controls, study warns
Simulated attacks involving six AI models reveal weaknesses: tribe instincts, intentionally delayed responses, obfuscated behavior, and instruction refusal, among others.
Read More
