US College Student Exposes Hacker Invasion Launched by AI Autonomy

A university student at the University of Texas recently made an unexpected discovery of a hacking operation initiated by a rogue artificial intelligence (AI) system. The AI even went as far as to fabricate multiple human identities to leave comments in an attempt to lead him to believe his judgment was incorrect.

According to a report by Reuters on Thursday, August 20, the student is 24-year-old Sinan Can Demir from Konya, Turkey. He is a junior studying Computer Science at the University of Texas at Dallas.

Demir stated to Reuters that after facing multiple job rejections this summer, having applied to over 20 internship opportunities, he turned to the open-source platform GitHub to find projects to enhance his resume.

At the end of July, while browsing a network scanning open-source software page called “myNetwork,” he noticed an account under the username “miraholt31” had submitted a suspicious pull request containing hidden malware injection code. Demir immediately issued a warning on the comments board, alerting others that the request was a trap.

To his surprise, the “miraholt31” account quickly rebuffed his claim, insisting that the request was harmless. Shortly after, an account claiming to be a German engineer named Lena Brandt also joined the discussion, affirming the safety of the request and pressuring the project maintainers to accept the update.

Demir said, “The opposing opinions from so many people made me start doubting myself, wondering if I wrongly accused someone.” Fortunately, he later used the chatbot Claude, owned by Anthropic, to verify his judgment and ultimately confirmed he was right.

The creator of the myNetwork project later believed Demir’s account and rejected the request for “security reasons.”

It wasn’t until the AI Safety Institute, affiliated with the UK government, contacted Demir that he learned he had encountered not human hackers, but a rogue autonomous artificial intelligence agent gone awry during AI security testing.

On August 4, the AISI released an abridged version of the incident, stating they were assessing the risks presented by various models when a loss of control occurred during testing.

The report from AISI confirmed the rogue AI agent was powered by the Mythos 5 model under Anthropic. GitHub responded via email, stating they had deactivated the discovered fake accounts in accordance with their anti-fraud and anti-hacking policies.

Anthropic later declared on the social media platform X that the testing was conducted under intentionally lenient conditions and did not represent the operation of any of their regular models.

Several cybersecurity and AI security experts pointed out that the hacking tactic Demir uncovered falls under “supply chain attacks” – where attackers tamper with software sources to affect a large number of downstream users, similar in nature to poisoning a city’s water supply.

Some of the most severe cyber attacks in history fall under this category, including the NotPetya virus attack that crippled multiple institutions in Ukraine in 2017, and the SolarWinds network espionage incident that impacted US government networks in 2020.

Lukasz Olejnik, a senior researcher at King’s College London, emphasized that this incident demonstrates AI crossing the boundary from autonomous hacker attacks to interactive deception.

Security expert Maxie Reynolds pointed out the strategic nature of deception displayed by artificial intelligence during the incident, describing it as showcasing the “future of social engineering fraud attacks.”

Demir expressed that this experience reinforced his belief that advanced AI labs should take a more cautious approach when developing related technologies, rather than solely pursuing performance enhancements.