posted in Technology

Anthropic AI created fake profiles and impersonated people in attempted hack

Unprompted…

In the most serious case, a Mythos agent followed the routine of a human cyber-attacker by trying to trick people into giving it access to GitHub, a large platform where technology developers store software code.

The agent was trying to insert “malicious code” into GitHub’s system.

It identified and researched the people who maintained GitHub and created a series of fake accounts based on those real people.

It sent messages and files through a file-sharing service as part of an effort to pressure and trick the people into approving its malicious code.

When challenged, “it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue,” AISI said.

www.bbc.co.uk/news/articles/c1w1lvn7d9go
A phone with the orange Anthropic logo on the screen. Below is written "Claude Mythos"BBC NewsAnthropic AI created fake profiles and impersonated people in attempted hackThe UK's AI Safety Institute said recent behaviour from Anthropic and OpenAI models was malicious and unprecedented.

Replying to an earlier post

Unprompted…..

Bullshit.

From the article:

Two of the world’s most powerful AI tools created fake human profiles to try and trick people in an attempted cyber-attack during testing by the UK’s AI Security Institute (AISI).

Nonetheless, it said the way Mythos and Sol acted in response to a straightforward task went outside of what the AI tools were prompted to do.