Anthropic Insiders Warn AI Race Could Outrun Safety Promises

Written by Published

Three researchers from Anthropic, one of the most prominent firms in the artificial intelligence race, have stepped forward to warn that the very systems they are helping to build could bring about human extinction before the decade ends.

According to Breitbart, Jacob Coxon resigned from his position as an AI researcher at Anthropic on Tuesday for the express purpose of issuing a public warning about the existential risks posed by advanced AI. He immediately took to X after his departure, declaring: The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible but I hear the same people express fear privately. No other human activity poses this level of danger, Coxon wrote.

Coxons alarm was swiftly reinforced by two of his former Anthropic colleagues, who joined the same X thread to corroborate his claims. Evan Hubinger, who leads alignment science at Anthropic, stated: Jacob is correct here we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.

Samuel Marks, Anthropics scalable-oversight lead, added that the concern is not confined to a fringe minority within the company but is widespread among senior staff. AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years. In general, the more senior the employee, the more concerned they are.

These warnings echo similar concerns raised at rival firm OpenAI, where top scientists have urged the industry to proceed with extreme caution. Breitbart News reported this week that OpenAIs leadership is openly contemplating voluntary slowdowns in AI development until robust safety standards can be agreed upon across the sector.

In a recent essay, OpenAIs Jakub Pachocki argued that the industry is approaching a point where self-restraint will be necessary to avert catastrophe. This is a time that calls for extreme caution, he wrote, explaining that OpenAI would be prepared to unilaterally hold back further scaling when needed, while insisting that broader action from industry and government is also required.

Pachockis essay describes AI as moving toward recursive self-improvement, a phase in which systems could begin advancing their own capabilities without direct human guidance. He warned that the pace of progress is unlikely to slow on its own, writing: I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence, a statement that underscores how even insiders feel the technology is outpacing human institutions.

He further criticized one of the industrys preferred safety tools, chain-of-thought monitoring, which involves scrutinizing a models step-by-step reasoning to detect harmful or erroneous outputs. According to Pachocki, its effectiveness is deteriorating because the boundary between a models internal reasoning and its supervised communication with people is dissolving, models are increasingly shaping their own reasoning processes, and dangerous capability gains can now emerge even when there is no explicit verbal reasoning to inspect.

Skeptics, however, question whether these dire pronouncements are entirely altruistic, noting that both Anthropic and OpenAI stand to benefit from portraying their products as uniquely powerful and dangerous. By amplifying the specter of extinction, critics argue, these corporations can inflate their valuations and pressure lawmakers into adopting regulatory frameworks that entrench incumbent giants while stifling smaller competitors and open-source alternatives.

The spectacle of AI titans simultaneously unleashing ever more capable systems and demanding government regulation has not gone unnoticed on the right, where there is deep suspicion of handing technological power to unaccountable elites in Silicon Valley or Beijing. Breitbart News social media director Wynton Hall has addressed this tension in his instant bestseller, which he describes as a roadmap for how the MAGA movement can shape AI policy that benefit humanity without handing control of our nation to the leftists of Silicon Valley or allowing the Chinese to take over the world, emphasizing that conservatives must engage this issue before progressive bureaucrats and foreign adversaries define the future of machine intelligence.