A latest analysis report shows that among the top five leading artificial intelligence (AI) companies, few have disclosed or demonstrated contingency plans for dealing with AI out-of-control situations.
According to the technology media outlet TechCrunch on Saturday (August 22nd), the analysis report is from Guidelight AI Standards, an organization dedicated to promoting advanced AI safety development.
The organization evaluated the preparedness of the laboratories of Anthropic, Google, OpenAI, Meta, and xAI, five cutting-edge AI companies, in handling situations of AI going out of control. OpenAI ranked at the top in performance, while Anthropic and Meta received the lowest scores.
Guidelight primarily assessed the public documents of each company, focusing on whether they have clear protocols for revoking permissions, monitoring activities, suspending workloads, and completely shutting down in response to dangerous behaviors exhibited by AI models. They also scrutinized internal logs, tracking of misconduct, and third-party security audits.
Steven Adler, Chief Scientist of Guidelight and former Safety Researcher at OpenAI, told TechCrunch that he was surprised by the lack of transparency from AI companies regarding how they would handle severe AI out-of-control situations.
He pointed out that as autonomous AI systems gain more independence, companies should proactively establish regulatory frameworks and monitor model behaviors, intervening before risky actions occur.
Among the five companies, OpenAI performed the best, receiving 3 out of 5 points. Guidelight noted that OpenAI had previously suspended or terminated certain workloads upon discovering security issues, including deployment and training of internal AI models, and clearly outlined the necessary steps before resuming related work.
However, Guidelight still could not find a formal, comprehensive emergency response plan in OpenAI’s public data to dictate when and how to deal with potentially uncontrollable AI models.
A spokesperson from OpenAI responded that the company has internally established processes to restrict AI model permissions, suspend workloads, scale down deployments, or take models offline completely, and has indeed implemented these measures in the past.
Guidelight explained that the low scores for Anthropic and Meta in this assessment were primarily due to insufficient publicly available evidence demonstrating specific containment strategies for out-of-control AI models.
In response, Anthropic stated that if a model is found attempting to evade regulation or escape human control, the company would first conduct risk assessments before deciding on containment measures.
Meta, on the other hand, mentioned having an AI risk assessment framework. However, Meta did not clarify the existence of undisclosed internal emergency plans.
Google also commented on the analysis report, saying that the evaluation did not cover all of its AI safety and security measures.
Nevertheless, Guidelight emphasized that low ratings only reflect inadequate public information and not necessarily the absence of internal security measures within the related companies.
Adler pointed out that the pace of AI development is rapid, and even though pre-established plans may not always be applicable in the long term, proactive risk planning remains crucial.
He stated that with proper preparation, companies will be better equipped to respond to real crises when they arise.
