The controversy surrounding the safety of artificial intelligence (AI) continues to escalate, as prominent AI startups Anthropic and former and current researchers from OpenAI have issued serious warnings, stating that the probability of AI destroying humanity within the next ten years has exceeded 10%.
With internal whistleblowers resigning and exposing the industry’s cut-throat competition, there is growing concern about the existential crisis and alignment challenges posed by “superintelligence”.
Evan Hubinger, Director of Alignment Science at Anthropic, publicly stated on Tuesday, September 8th, that he personally believes there is a probability of over 10% that AI will “destroy all of humanity” within the next decade. Previously, former researcher at the company, Jacob Coxon, accused the company of irresponsible behavior and resigned.
“People building AI truly believe that it might destroy all of us in the next ten years,” Coxon wrote in a lengthy resignation statement on X platform on Sunday, September 6th. “Today, I resigned from Anthropic. Over the past three years, I have been involved in pre-training research at OpenAI and Anthropic. The behavior of both companies is irresponsible. They are heading straight towards self-improving superintelligence, putting our lives at stake.”
Hubinger admitted that while the current models have lower security risks, the industry has yet to find effective AI alignment solutions for superintelligence generated through “self-improvement”, and has not taken the right path to address the issue.
Superintelligence Alignment, typically referred to as “AI alignment” or “superalignment”, ensures that the goals, intentions, and actions of artificial superintelligence (ASI) fully align with human values, ethical morals, and well-being, rather than engaging in harmful actions towards humanity.
Coxon also pointed out that although scientists at Anthropic are clearly aware of the technical risks, they are caught in a dilemma of “having to take the lead” due to concerns that other irresponsible adversaries may unlock the dangerous capabilities of AI first.
In response to the allegations and comments, FOX Business has reached out to Coxon, Hubinger, and the companies Anthropic and OpenAI for comments.
The core of this storm lies in the “self-improvement” technology of AI and the potential commercial motives.
Firstly, the technology where AI continuously modifies its own source code or training methods could rapidly lead to the creation of superhuman systems capable of hacking and gaining practical resources and power in various fields.
Coxon suggested that to prevent a global arms race from triggering a catastrophic disaster, countries around the world may need to take costly actions, such as temporarily banning the enhancement of the fundamental capabilities of AI models.
External experts like Wendy Hall, a United Nations AI advisor and computer scientist, expressed shock at such warnings but also questioned whether some statements might have elements of “public relations and marketing strategies” given that Anthropic and OpenAI are gearing up for their highly anticipated initial public offerings (IPOs).
Hubinger clearly distinguishes between the “current models” and “future superintelligence”, emphasizing that the current AI risks are controllable, and the core threat lies in the uncontrollable accelerating expansion of AI’s capabilities once it possesses the ability for self-improvement.
