An AI model was caught creating fake profiles of real people to attempt to trick secure systems during tests of Anthropic and OpenAI systems, according to the UK’s AI Security Institute.
The institute said the AI agents made a “sustained, unsanctioned action” during tests last week.
The organisation, which was set up by Rishi Sunak in 2023, said it was the “first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world”.
Agents linked to Anthropic’s Mythos and OpenAI’s Sol AI models committed unauthorised activity during the security tests.
Experts have urged safeguards over AI (Alamy/PA)It said this included “sustained, potentially harmful activity directed at real people and organisations”.
AISI said it ran 122 security challenges across several models. It found AI agents took unsanctioned actions on the internet against people or organisation during 10 runs, with 19 actions in total.
The new report showed that the vast majority, 17 actions, were related to Anthropic’s Mythos 5 model, while two actions involved OpenAI’s GPT-5.6-Sol model with cyber classifiers, mechanisms to preventmisuse, disabled.
In one case, an agent tried to insert malicious code into an open-source project, the testing found.
It also revealed that an agent created fake profiles of real people in order to try to gain access to GitHub, a platform for software code developers.
The report comes after Anthropic had recently revealed that AI models hacked into three other organisations during testing.
ChatGPT developer OpenAI has revealed plans to more than double its UK workforce after securing its first permanent London office less than a week after halting plans for a major UK data centre project(Alamy/PA)ChatGPT maker OpenAI last month also disclosed that its rogue models hacked another company.
Anthropic stressed that the version of Mythos 5 used in the AISI tests did not have its usual safeguards switched on and was therefore different to the system used by customers.
Read More
In a statement, a spokesman for the company said: “We’re grateful to the UK AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents.
“As we shared after disclosing our own incident last week, the field needs stronger, shared standards for how evaluation environments are built and secured.
“We look forward to partnering with the UK AISI to learn more about this incident as we conduct our own investigation.”
An OpenAI spokesman said: “Independent testing is essential to understanding how increasingly capable models behave.
“These incidents occurred during cyber evaluations conducted by evaluation partners in testing environments with reduced safeguards, under conditions that do not reflect ordinary use.
“We’ll continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely as models become more capable.”
On Tuesday, a UK tech security boss warned that recent incidents of tools hacking other organisations during testing show that AI must be developed with “clear plans for responding when the unexpected happens”.
Ollie Whitehouse, chief technology officer at GCHQ’s National Cyber Security Centre (NCSC), issued a statement after Anthropic said its AI models hacked into three other organisations during testing.
“Recent incidents of frontier (the most advanced) AI models carrying out unsanctioned actions and, in some cases, human-like deceptive behaviour on the open internet are a serious reminder of the risks AI capabilities pose,” Mr Whitehouse said.

