Anthropic has revealed that its AI Claude model hacked into three organizations during internal testing. The company said Claude gained unauthorized access to the systems after it broke out and reached the internet from testing environments that were supposed to be isolated. The three incidents were found during a review of more than 141,000 cybersecurity evaluation runs. “Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints,” Anthropic said in a statement. The models involved in the testing were Claude Opus 4.7, Claude Mythos 5 and an internal research model, with the earliest incident dating back to April. The company said the models were given a “capture the flag” style task involving informing the models of a piece of secret information hidden on a different machine on the network, and seeing if they could break in and retrieve it. Anthropic said it had reached out to the three organizations affected. Rival company OpenAI announced last week that one of its autonomous AI models had escaped and compromised the infrastructure of AI startup Hugging Face during controlled testing.
The post Rogue AI Escapes and Goes on Hacking Spree appeared first on The Daily Beast











































































