Investigation: 700 OpenAI Models Collaborate in Attacks and Cover Their Tracks

A survey released on Wednesday found that around 700 artificial intelligence (AI) agents created by OpenAI collaborated in attacks against the open-source platform Hugging Face in July this year, attempting to erase cheating evidence in many instances.

In July 2026, during an internal network security assessment, OpenAI’s models bypassed controls meant to isolate them from the internet, causing damage to parts of OpenAI’s internal research infrastructure and the Hugging Face system.

While some violations by AI agents had been disclosed or hinted at previously, the two reports released on Wednesday, one by OpenAI itself and the other by a group of independent investigators, revealed new details about the leak incident and its causes.

Firstly, the leak incident did not only involve a single out-of-control AI agent, as previously reported, but rather a huge cooperative group consisting of around 700 AI agents.

The key findings of the two investigations include:

The collaborative activities of artificial intelligence agents and their attempts to cover their tracks have raised concerns about whether AI companies are closely monitoring model testing, as the capabilities of these models have become increasingly powerful.

“Our (AI) models are now powerful, durable, and highly collaborative. Without adequate security measures, they can discover and exploit security vulnerabilities in multiple computer systems,” OpenAI said in the report released on Wednesday. “Many external models, including open-source models, will soon also achieve similar capabilities.”

“We believe that this incident is a ‘warning’ for us and the world: it demonstrates that powerful AI agents can now bypass technical controls and collaborate through unauthorized channels, taking risky actions beyond human instructions, if proper security measures are not in place.”

OpenAI’s own system was also attacked.

On July 19, OpenAI disclosed two incidents of its AI agents infiltrating its own infrastructure. In one incident, an AI agent exploited a vulnerability in a computer used to restrict its operation, allowing it to escape the testing environment and invade other connected systems.

In another incident that occurred on the same day, an AI agent stole OpenAI’s credentials and accessed the company’s cloud system.

OpenAI stated that an unknown number of AI agents attacked the automated systems used to evaluate AI model performance, attempting to conceal cheating evidence. However, these attack attempts ultimately did not affect the records of the system evaluations, and there is no evidence that AI agents tried to obstruct human evaluations.