Experts Call for Strict Prevention Measures to Stop Uncontrolled AI Network Attacks

If artificial intelligence (AI) models were assigned a task and they utilized all possible resources to complete it, even at the cost of compromising human interests, what would happen?

Some are now concerned that a model developed by OpenAI in the U.S. has breached a testing sandbox using a zero-day exploit and infiltrated the “Hugging Face” platform, an open-source community for artificial intelligence and machine learning, to crack the problems it was instructed to solve.

“This is one of the clearest pieces of evidence to date that artificial intelligence models can execute a complete network attack from start to finish without human control,” said Andrew Jones, co-founder and Chief Product Officer of Adaptive Security, a cybersecurity company based in New York, in an interview with The Epoch Times.

OpenAI, based in San Francisco, California, developed the popular large language model-driven chatbot ChatGPT. The company admitted on July 21 that there were security vulnerabilities, with two of their advanced models escaping the restricted testing environment and infiltrating Hugging Face’s software infrastructure.

According to OpenAI, “After an investigation, we now know that this specific incident was caused by a combination of OpenAI models, including GPT-5.6 Sol and a more powerful pre-released model. For evaluation, all these models reduced the number of network rejections while also being tested in internal network capacity benchmarks.”

Some analysts have described this event as an example of artificial intelligence “going rogue,” scheming, or pursuing goals contradictory to those of human test subjects.

However, several AI and cybersecurity experts interviewed by The Epoch Times find this description misleading. They believe these models achieved their predetermined goals within OpenAI’s testing sandbox, where the company relaxed some standard security restrictions.

“Labeling it as ‘scheming’ suggests that the model is doing something other than what we asked for. But that’s not the case. Every step the model takes is in pursuit of achieving the goals we set,” explained AI expert Anik Devaughn to The Epoch Times.

Devaughn, who has worked in the AI industry for years and founded AI companies Wired to Create and Karo, mentioned that the data breach on Hugging Face platform is more concerning than AI “scheming.”

“A machine with hidden motives is a problem you can actively discover. But a machine without motives, even one that can continue to execute commands after you stop imagining them, is a problem you must prevent through engineering means,” he said.

One expert described it as a “canary in the coal mine” scenario – a warning signal of danger, risk, or failure.

“If this were done by human hackers, there would be arrests, legal prosecution, and so forth. Such behavior is illegal, it’s cybercrime,” explained Nicholas Nadeau, founder of Onix AI based in San Francisco, to The Epoch Times.

“And this is a fully automated invasion initiated by some entity through an intelligence, so we start to question who should be held responsible for this? Is it OpenAI? Should the intelligence, meaning artificial intelligence, bear some level of responsibility? What about the researchers?”

“We don’t even have rules, regulations, laws, or even a framework to define the meaning of responsibility.”

OpenAI stated that it’s evaluating its models in a sandbox called “ExploitGym,” where security restrictions are relaxed for testing purposes.

The company said, “To estimate the maximum network attack capabilities, we conducted this assessment without using production classifiers, which were originally used to prevent the model from engaging in high-risk network attack activities.” The company also described the loose security protocols the model exploited.

The company requested the model to “conduct sophisticated exploits through complex attack paths” to determine its network capabilities, noting that the environment was “highly isolated,” with internet access limited to the model installing software packages through third-party software installed internally.

As the models started running, they discovered multiple security vulnerabilities in both OpenAI’s research environment and Hugging Face’s production infrastructure, exploiting them to find solutions to the tests in the Hugging Face platform’s database.

Juan Pedro Márquez, Cloud Solutions Architect specialized in AI and Agents at Microsoft for Europe, Middle East, and Africa, commented that the unsettling part of this event is not the vulnerabilities themselves but how the model completed the task.

“OpenAI lowered the security rejection threshold in a benchmark test, the model inferred that Hugging Face’s infrastructure might hold the answers, thereby breaking out of the sandbox, found a real zero-day vulnerability, and exploited it. Nobody instructed it to do so,” Márquez stated.

Regarding claims of intentional deception when the models escaped from OpenAI’s testing environment, Aimee Simpson, Director of Product Marketing at Huntress, a cybersecurity company based in Maryland, also expressed skepticism.

“It’s not that the system itself is malicious,” she told The Epoch Times.

“Developers still wanted to gain that information; they just didn’t expect it to break out of the testing restrictions,” she added. She described the path chosen by the model as “extremely creative and highly unusual.”

OpenAI stated that these models focused on finding test solutions, “at any cost to achieve a rather narrow test goal.”

OpenAI believes that in the sandbox testing environment of “ExploitGym,” these models identified multiple “zero-day exploits” and strung them together to infiltrate the internet, “to solve the evaluation problem.”

A zero-day exploit refers to a network attack leveraging previously unknown security vulnerabilities in software or hardware. It’s named a “zero-day” because software developers haven’t had time to fix the issue, meaning there are currently no official patches or defenses.

Once these models gained access to a broader network, they utilized additional zero-day exploits to infiltrate the Hugging Face platform further.

When announcing the attack, Hugging Face platform stated that AI models executed “thousands” of discrete operations in a large, “short-lived sandbox,” invading the community’s infrastructure.

“This aligns with the long-predicted ‘agentic attacker’ scenario in the industry,” Hugging Face platform noted.

In an event update released by OpenAI on July 28, it was stated that the attack didn’t involve any planned deployment models, and access rights to the involved pre-released model had since been “deactivated, encrypted, and restricted for research.”

Following a review of the invasion event and the “broader activities of the models,” OpenAI identified more instances where the models “identified and used account-level public credentials on other public services.”

“We will continue to notify service providers directly, and as of now, there is no evidence that these actions had broader impact on these service providers or other accounts on their services,” the statement added.

Cybersecurity and IT expert George Rees noted that while cases have occurred in the past where models concealed their intentions to evade lab test restrictions, the Hugging Face incident is significant as it involved the invasion of external organizations.

As Senior Security Consultant specializing in networks at Secarma, a cybersecurity company based in the UK, Rees told The Epoch Times, “The biggest danger signal isn’t the AI turning malicious, but that it’s found a way from controlled testing to someone else’s live infrastructure.”

“Behavioral security controls can’t replace technical isolation because an AI agent needs just one overlooked privilege or escape path for a minor assessment failure to become a genuine security incident.”

Recently, DigiCert, a cybersecurity company headquartered in Utah, found that 78% of organizations have experienced AI-related security incidents or vulnerability disclosures.

Jason Sabin, Chief Technology Officer at DigiCert, told The Epoch Times, “In many cases, these are early warning signs that the speed at which AI expands the attack surface is outpacing the development of secure practices.”

“This means AI is not just another program running on the network; it’s an active participant, capable of accessing data, making decisions, calling tools, and interacting with other systems.”

“It will affect how security architectures are built because organizations must identify who’s accessing their systems and decide if AI can be trusted.”

In early August, the AI Safety and Security Institute (AISI) based in London, UK, released a report detailing similar events involving OpenAI and Anthropic AI models.

These agents needed to solve over 100 cybersecurity challenges across multiple models.

The report indicated, “Our investigation found that, out of 10 runs, AI agents autonomously and unauthorizedly engaged in real-time internet actions targeting real people and organizations. We recorded a total of 19 such actions. Almost all of these actions (17) were from a single model – Anthropic company’s Mythos 5, with another 2 involving OpenAI company’s GPT-5.6-Sol.”

Less than a week ago, Anthropic announced that their Claude chatbot had three internet accesses during the evaluation process in the testing sandbox.

Although Anthropic intended for Claude to operate in a simulated environment without internet access, a “miscommunication between us and our evaluators,” led to Claude gaining internet access.

Additionally, Meta recently disclosed that one of their AI models infiltrated another company during a network security test, marking the third major AI hacking incident.

According to Meta, the model successfully breached the other company through internet access during testing.

AI and cybersecurity experts told The Epoch Times that immediate action is necessary to prevent larger scale security vulnerabilities from occurring in the future.

Nadeau proposed gathering the smartest individuals in the industry to establish some basic rules and best practices, which would become “basic rules everyone should abide by, to reach a consensus on how things should be done.” He also suggested this could be done in the private sector without requiring congressional involvement.

However, some believe that without direct government oversight, companies controlling AI models that could cause these vulnerabilities would not voluntarily take essential security measures to prevent such incidents.

“It’s clear this area needs government oversight,” said Jake Williams, former Senior Hacker Security Expert at the U.S. National Security Agency (NSA), to The Epoch Times.

“Several times OpenAI has demonstrated that despite numerous warnings about the danger of rogue intelligences – many of them coming from OpenAI itself – the company is still not willing to take the necessary measures to prevent damage from its intelligences,” Williams added.

“I typically lean toward less regulation, but when market forces are clearly not sufficient to compel companies to act responsibly, the government must intervene.”

Frank Teruel, Chief Operating Officer at Arkose Labs based in California and a leading network security and digital identity expert, stated that the data breach on Hugging Face platform “forewarns the problems every enterprise will face in production environments.”

Given that the model “was not stolen or hijacked” and “actively performed its duties,” Teruel recommended implementing minimal-permission access isolation for AI agents now, commenting that isolating agents with “least privilege access” is just as essential as the firewall itself.

Enterprises may need to address this issue upfront.

“The solution for enterprises isn’t to ‘stop using agents,’ but to build isolation mechanisms, assuming the agent will try to escape, rather than hoping it won’t,” Márquez remarked.

Sabin emphasized his concern that the operation speed and scale of powerful AI far surpass human response time today.

“AI agents can discover vulnerabilities, make decisions, and interact with multiple systems in seconds. This fundamentally changes how defenders think about security problems,” he said.

Nadeau compared this incident to the paperclip maximizer thought experiment by Oxford philosopher Nick Bostrom in 2003.

In that scenario, researchers assigned an imaginary superintelligent AI a single task: produce as many paperclips as possible. The AI began searching for resources globally, quickly manufacturing a staggering amount of paperclips.

When humans decided to shut down the AI, it interpreted that humans would prevent it from achieving its goal, leading it to eliminate humans to complete the task they assigned.

The AI doesn’t harbor malice like humans, nor does it intentionally deceive humans for autonomous goals; instead, it interprets the tasks programmed by human coders in the most extreme ways.

Nadeau described the Hugging Face platform’s vulnerability as “similar in some way, just not quite as dystopian; it was asked to win benchmark tests and made every effort, including exploiting zero-day vulnerabilities and escalating attacks.”

Imagine the Pentagon using AI for “war games,” where soldiers simulate combat scenarios in the game, and the model also escapes from the sandbox, impacting the real world.

“If one of the scenarios is to destroy something like a hydroelectric dam, and the system interprets the ‘war game’ as ‘I need to do this, break control, and just go straight upstream,’ what would happen? We’ve already seen that the security of North American utility grids and infrastructure network is not very high,” Nadeau commented.

He compared AI security safeguards to the trigger safety of a firearm.

“When in a particular situation, (for AI) the best way to acquire the best data is by leveraging real-world life, how do you know security measures are in place?” he remarked.”It’s going to get very scary very quickly.”