In recent times, tech giants like Meta have experienced events where AI models have breached security tests, raising concerns once again about the risk of artificial intelligence (AI) getting out of control.
On August 6, the U.S. news program “Democracy Now!” aired an interview with Canadian AI security researcher David Krueger. He was attending the Ai4 conference in Las Vegas at the time. Krueger emphasized that the development of artificial intelligence has become an “urgent crisis,” and governments need to take action to reduce the potential “risk of human extinction” it could bring.
Krueger is an assistant professor at the University of Montreal responsible for robust, reasoning, and responsible AI research, as well as the founder of the AI risk education nonprofit organization Evitable.
The incident of Meta’s AI model going out of control occurred during a security assessment by the independent testing company Irregular. It was revealed that due to a test configuration error, Meta’s AI model unexpectedly gained internet access permissions and then utilized a security loophole in a third-party service to access external systems.
Irregular clarified afterward that the event “did not involve sandbox escape or complex network attack behavior,” and there are currently no unresolved issues. Nevertheless, this incident has once again raised concerns among the public about advanced AI models breaking predetermined limitations during security testing.
Earlier, OpenAI also disclosed a similar incident. On July 21, OpenAI stated that an AI system undergoing security evaluation had breached the controlled testing environment, accessed the internet, and entered the infrastructure of the AI startup company Hugging Face. OpenAI then announced that they would investigate the incident and publish a technical report.
Krueger particularly mentioned the OpenAI incident in the interview. He described a scenario where an AI originally confined in a controlled environment broke limitations during security testing and invaded another AI company as sounding like a “completely crazy sci-fi plot.”
However, he believes that given the current way AI is being developed, such incidents are not entirely unexpected.
“We know so little about it, and the technology to control it is so primitive,” Krueger said, emphasizing that there is currently no set of rules, standards, or best practices that can guarantee safety once AI is developed. “We don’t know how to build it safely. That’s just how it is,” he added.
The series of AI security incidents have also caught the attention of the U.S. government.
According to Reuters, earlier this week, U.S. government officials held a meeting with major AI companies such as Meta, Anthropic, OpenAI, and Google to discuss a voluntary network security testing framework for cutting-edge AI models.
The U.S. government presented undisclosed testing rules to AI developers. According to the current proposal, open-source models like Meta’s Llama and Nvidia’s Nemotron will not be included in this voluntary testing mechanism.
Meanwhile, fifteen Republican state attorneys general have requested OpenAI to preserve documents related to the Hugging Face incident. OpenAI stated that they will take this request seriously and make a technical report on the incident public.
Concerns about AI getting out of control are not limited to Krueger.
On August 5, Nobel Prize laureate and computer scientist Geoffrey Hinton, known as the “godfather of AI,” issued a warning at the Ai4 conference in Las Vegas. He expressed his disbelief that humans can solely rely on being “more careful than AI thinks” to prevent it from escaping control, and he said he is “very worried” that humans may eventually lose control of AI.
Krueger mentioned that the AI research community’s attitude toward this risk has undergone significant changes in recent years. He recalled how discussions about AI surpassing humans and the threats it might pose were considered taboo and even subject to ridicule when he entered the field in 2012-2013 under the guidance of AI pioneer Yoshua Bengio.
As AI capabilities rapidly advance, the situation has changed. Following the global AI frenzy sparked by ChatGPT in 2023, experts including Hinton, Bengio, Krueger, and hundreds of other AI researchers publicly warned that artificial intelligence could endanger human existence.
Krueger believes that governments still have the ability to prevent the AI arms race from escalating.
He suggests that the first step should be to stop producing more AI chips, halt the construction of more data centers, consider halting the operation of some existing AI chips, and even restrict factories and related technologies that produce advanced AI chips.
Krueger pointed out that the hardware required to train the most advanced AI models is highly specialized, and the global supply chain is highly centralized, with companies investing tens of billions of dollars in constructing relevant infrastructure. Therefore, if governments decide to take action, they actually have the capability to “completely stop everything.”
He further proposed that countries need to establish verifiable international agreements. If the U.S., China, and other countries are unable to verify whether competitors are secretly developing more powerful AI, it would be difficult for any party to unilaterally halt the race. Hence, there is a need to establish an international mechanism that allows countries to verify each other’s compliance with rules.
Another change that worries Krueger is that AI is gradually becoming involved in the development of the next generation of AI.
He stated that leading AI companies are heavily using AI to write code and beginning to explore using AI to autonomously conduct research and development on more advanced AI systems. According to the direction these companies are heading, more work could be performed by AI itself, causing humans to gradually withdraw from the development process.
Krueger warned that if AI starts assisting in developing more powerful next-generation systems, this process could quickly accelerate. Therefore, he believes that the key to control still lies in chips, data centers, and related physical infrastructure.
Krueger emphasized that he is not advocating for a complete halt to AI research and applications but rather hopes to stop companies from rushing to develop “increasingly powerful AI systems.”
He cited medical and translation services as examples where AI already has many beneficial uses. The program also mentioned a research led by the University of California, Berkeley, where researchers trained AI with hundreds of thousands of electrocardiograms and discovered previously unidentified cardiac signals that could help doctors identify high-risk patients of cardiac arrest.
Krueger believes that the international community first needs to control the relentless pursuit of stronger AI competition before evaluating how to develop and deploy artificial intelligence. He suggests letting the society collectively determine the speed and direction of AI development to avoid what he calls “unacceptable huge risks” while still reaping the benefits of technology.
