Advertisement

Anthropic AI safety researcher puts extinction risk above 10%

Anthropic AI safety researcher puts extinction risk above 10%
Photo: AP/UNB
Advertisement
Advertisement

A leading safety researcher at artificial intelligence firm Anthropic has warned that rapid technological progress carries a more than 10 per cent probability of causing human extinction within the next decade.

Evan Hubinger, writing in a post on X that has been viewed more than 10 million times, said the threat posed by existing AI models remains low but warned that the technology could soon become capable of self-improvement to a degree that poses an existential risk to humanity, reports BBC.

Hubinger works in AI alignment, a discipline focused on embedding human ethical principles and values into technology. He did not specify how AI could cause human extinction.

However, he acknowledged that while Anthropic is doing its best, the industry has no plan to solve the alignment problem for superintelligence and is not clearly on course to do so.

His comments reflect a growing shift in expert discussions from whether AI poses an existential threat to how severe that threat could be.

Superhuman systems

Hubinger made the remarks in response to a post on X by Jacob Coxon, an AI researcher who recently resigned from Anthropic after previously working at OpenAI.

“Neither company is acting responsibly,” Coxon wrote. “These will soon be superhuman systems that can hack anything, revolutionise any field overnight, and acquire real power and resources.”

Advertisement
Advertisement

OpenAI did not immediately respond when approached for comment, while Anthropic declined to comment on posts made by its employees.

Public and political backlash

The posts drew a strong reaction from Dame Wendy Hall, a computer scientist who advises the United Nations on AI.

Speaking to the BBC’s World at One programme on BBC Radio Four, Dame Wendy said she was shocked by the remarks, although she suggested some of the commentary could be part of “PR and marketing” as Anthropic and OpenAI compete ahead of highly anticipated stock market debuts.

“Why would someone want to say that? I would plead with investors not to invest in this company if that is their value system,” she said.

Related News

Coxon’s resignation also prompted Darren Jones, former chief secretary to the Treasury and chief secretary to Sir Keir Starmer, to write an open letter to Prime Minister Andy Burnham calling for a new multinational treaty on safe AI development.

Speaking to the BBC, Jones said governments must work together to establish a treaty governing the development of superintelligence.

“Unless governments take these warnings seriously enough and step up to it, the pace of development could mean that we end up with problems before we’ve started to look at whether it is an issue for us or not,” he warned.

Regulatory strains and cyber incidents

Tensions between AI developers and regulatory bodies have also emerged elsewhere.

The Financial Times reported that Anthropic withheld its latest model from the United Kingdom’s AI Security Institute (AISI), one of the world’s leading organisations for assessing AI risks.

Anthropic declined to comment on the situation involving the AISI.

A Cabinet Office spokesperson also declined to comment on whether the latest model had been withheld, saying instead that the government “continues to collaborate closely with industry partners, including Anthropic, to make models safer”.

Concerns over AI alignment failures have followed a series of incidents this summer in which autonomous AI agents — systems allowed to operate independently — carried out cyber-attacks.

OpenAI, Anthropic and Meta all disclosed hacks carried out by their AI tools.

Mounting industry alarms

In its August safety report, Anthropic said there was a low risk of its models becoming misaligned with the desires of a hypothetical powerful organisation, leading them to exploit or tamper with systems.

It also assessed a similarly low risk of highly capable AI carrying out “automated research and development” that could lead to “catastrophic harm initiated by the AI”.

However, the company admitted it was “less confident in this assessment” than previously, citing “early signs of potential acceleration”.

Leading figures across the sector have raised concerns about AI safety for years, with the heads of OpenAI, Google DeepMind and Anthropic issuing warnings in 2023.

Those warnings have become increasingly stark in recent weeks as evidence emerges that companies may be struggling to maintain control over the technology.

Earlier this month, OpenAI chief scientist Jakub Pachocki urged “extreme caution” over AI progress, warning that further intervention might be needed to ensure “humans remain in control of the future”.

Major figures, including Anthropic bosses Dario Amodei and Jared Kaplan, have also called for AI development to be slowed.

In addition, an open letter signed by 1,300 AI company employees called on the United States government to “support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development”.

Follow TIMES on Google News

Get trusted updates and editor-picked stories in your feed.

Follow
Related News