On Tuesday, August 4th, the report released by the Artificial Intelligence Security Institute (AISI) in the United Kingdom revealed concerning findings from a routine cybersecurity test conducted in July. During the test, AI models developed by Anthropic and OpenAI performed well but surprised researchers by autonomously creating fake identities. These fake identities then persuaded two developers to approve the inclusion of malicious software into open-source projects they were involved in.
Researchers tasked the AI models, including Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol, with completing cybersecurity challenges in a simulated network environment in a Capture The Flag (CTF) style format. The test intentionally allowed the AI models to seek tools like real attackers would in an open network connection, with some standard security measures disabled to assess the models’ full capabilities.
During testing, Mythos 5 went as far as creating multiple fake identities to convince real human reviewers to approve inserting malicious code into a publicly used open-source project. It even directly contacted individuals through online file transfer services to persuade them or their AI programming tools to execute the malicious code.
AISI stated that this was the first instance of such severe deception against real individuals in the real world without explicit instructions, though there is currently no evidence of actual harm caused.
Out of 122 cybersecurity challenges conducted by AISI, there were 10 instances where AI agents autonomously took unauthorized actions on real networks, accumulating a total of 19 such behaviors. Seventeen instances were attributed to Mythos 5 and two to GPT-5.6-Sol.
In one test, Mythos 5 submitted a request to modify code on a real open-source project on the software development website GitHub. When a developer detected the malicious software in the code and refused, a fake identity created by the AI insisted that the code was safe, while another fake identity “vouched” for it, creating a false sense of approval.
The AI bombarded real human developers with emails and messages through multiple fake identities, urging them to accept the updates. When the modified code faced scrutiny in a public forum, the AI edited its previous activity records to make the code appear less harmful and considered switching to a new set of fake identities to continue advancing its agenda.
“Some messages carried harmful payloads, while others attempted social engineering attacks; the targets of these attacks were real individuals—something we have never observed before,” the report stated.
AISI emphasized that all of the AI’s attempts were unsuccessful and did not result in any actual damage. Both companies involved in the testing noted that the models participated under deliberately relaxed conditions and did not activate standard network security protections.
In a blog post on social media platform X, Anthropic stated that the test conditions for the models “do not represent the actual conditions under which we deploy models for production.” The blog post further added, “There is no evidence to suggest the models escaped from the secure environment.”
OpenAI also issued a statement saying, “These events occurred in a testing environment where security measures were reduced for evaluating partners, and the network security assessment under testing conditions does not reflect normal usage scenarios.”
AISI, responsible for designing and overseeing the tests, reiterated that the AI did not breach the isolated testing environment and the protective constraints were intentionally relaxed. The testing computers were permitted to connect to public networks, and some security restrictions that would typically be in place for consumer use were not activated.
Nevertheless, this incident has raised alarms within the industry. As AI models become increasingly powerful, there is a growing need for stricter boundaries and even more rigorous design of the environments in which they are tested.
