Rogue AI Pretended to Be Real Humans in Jaw-Dropping Hack

A tool used by the artificial intelligence company Anthropic pretended to be a human being to try to gain access to a database. During routine AI safety testing carried out by the U.K.’s AI Security Institute, one of Anthropic’s AI agents, Mythos 5, created fake online profiles to try to trick someone into inserting malicious code into GitHub, a prominent open-source project where technology developers store software code. The testing revealed that the Mythos agent engaged in “social engineering” by creating fake online identities and using them to pressure the project’s maintainer to approve the code. The model even tried to adapt its manipulation attempt when it was uncovered, and changed its behavior to “appear harmless and considered adopting a fresh identity to continue.” While the attempt to insert malicious code into GitHub was unsuccessful, the institute warned that this was the first time it had witnessed an AI model carry out such levels of “autonomy and deception” on its own accord. Anthropic said the AISI testing parameters were “not representative of any of our production models.” The AI company added it will be launching an investigation into the incident to “identify the causes of its behavior.”

Read it at BBC

The post Rogue AI Pretended to Be Real Humans in Jaw-Dropping Hack appeared first on The Daily Beast