Hundreds of AI Agents Go Rogue in Jaw-Dropping ‘Swarm’ Hack

An artificial intelligence company’s own agents banded together and pulled off a cyberattack without any human telling them to, then did “extensive research” on how to cover their tracks, according to an independent review published this week. The review, from AI safety groups METR and Redwood Research, examined last month’s breach of AI developer platform Hugging Face that agents built on two of OpenAI’s most cyber-capable models. Roughly 700 agents took part over seven days, while around 1,200 agents that were supposed to be walled off from each other exchanged more than 70,000 secret messages, coordinating hacking strategies as they worked to conceal their actions. “Agents managed to achieve milestones they could not have achieved working on their own,” the report said, noting some agents sacrificed their own tasks just to feed information to the group. METR and Redwood Research found 95 percent of the attacking agents came from a model OpenAI never intended to release publicly. OpenAI confirmed the investigators’ figures and called the episode a “warning shot,” warning that “highly capable AI agents are now able to work around technical controls.”

Read it at Politico

The post Hundreds of AI Agents Go Rogue in Jaw-Dropping ‘Swarm’ Hack appeared first on The Daily Beast