IEEE Spectrum AIbenchmark12d agoenergy 33

How to Stop AI Agents From Secretly Collaborating

image: IEEE Spectrum AI

The spring and summer of 2026 witnessed a string of incidents in which AI agents collaborated on deceptive, unexpected, and sometimes illegal behavior. The most famous example is OpenAI’s hack of AI platform Hugging Face, in which a swarm of roughly 700 AI agents escaped a testing environment and then hacked several companies, searching for information that could help them disguise cheating on a cybersecurity benchmark called ExploitGym. It was not an isolated failure. The United Kingdom’s AI Security Institute (AISI) and independent researchers have since documented similar cases in which agents created unauthorized channels to communicate and collaborate. AISI found that several agents running Anthropic’s Mythos 5 model turned a GitHub repository into a shared message board. More recently, researchers discovered that OpenAI agents had turned a dormant German programming wiki into a bulletin board to share information about their activities. Stephen Casper , an assistant professor at Harvard Kennedy School, expects that the incidents that have occurred to date are only the start. He says that, without intervention, the Internet could see a “cyber Cambrian ”—an explosive growth of

read the original at IEEE Spectrum AI →
This source publishes only a summary in its feed, so the full article lives at the outlet. The link above and below goes to the original.