Share

FIRSTonline Banner

Anthropic researcher resigns and raises the alarm about AI: "Humanity at risk by the end of the decade." The real risks behind the complaint

Jacob Coxon leaves Anthropic and denounces the risks of the AI ​​race. Meanwhile, new incidents involve Claude and OpenAI's models.

Anthropic researcher resigns and raises the alarm about AI: "Humanity at risk by the end of the decade." The real risks behind the complaint

Ha worked for three years to the development of artificial intelligence, first in OpenAI and then in anthropic. Now Jacob Coxon leaves the industry, convinced that the race towards increasingly powerful systems is proceeding without sufficient security guarantees. The 27-year-old British researcher announced his resignation in a series of messages on X, accusing both companies of taking risks that affect all of humanity.They are playing with our lives “The people developing AI sincerely believe it could kill us all by the end of the decade. Soon, these will be superhuman systems capable of hacking anything, revolutionizing any industry overnight, and acquiring real power and resources.”

At the heart of his complaint is the development of a superintelligence able to improve yourself independently, to the point of surpassing human capabilities and, according to its fears, even the possibility of controlling it. This alarm comes as new incidents of unauthorized access and circumvention of restrictions imposed on AI agents fuel questions about the reliability of security measures.

Why Coxon resigned

Coxon was in charge of pre-training, the stage in which models learn from large amounts of data the capabilities that will then be used to perform different tasks. His work had brought him into the heart of research at two of the leading laboratories in the field. That very experience, he claims, convinced him to abandon his career. "Neither company is behaving responsibly," he wrote. In his account, advances are bringing AI closer to capabilities that could allow it to hack computer systems, rapidly transform entire industries, and acquire resources and power. The problem, in his view, is that this acceleration would not be accompanied by a equally solid understanding of the mechanisms necessary to keep it under control.

The passage then moves on to the concerns which, according to Coxon, are circulating within laboratories. "The people developing AI genuinely believe it could kill us all by the end of the decade," the researcher explains, arguing that some executives use more reassuring expressions in public than they do in private conversations.

It is about assessments and fears expressed by the researcher, not of a proven predictionBut his complaint raises a real question: who gets to decide what level of risk is acceptable when the consequences could extend far beyond a single company?

The contradiction of Anthropic and a race that no one wants to stop

The resignation affects a central element of Anthropic's identityThe company was founded in 2021 by a group of researchers who left OpenAI, partly due to disagreements over security. Today, it faces similar criticism from one of its own employees.

To give further weight to the case, the following intervened: Evan Hubinger, responsible for the team that deals with model alignment, that is, ensuring that their behavior respects human intentions and goals. Hubinger shared Coxon's concern, personally attributing a greater than 10% probability of a catastrophic scenario within the next decadeWhile defending Anthropic's commitment, he acknowledged the lack of an adequate plan to safely develop the most advanced systems. This isn't the first departure accompanied by a warning. In February, Mrinank Sharma, head of security research, also left with a letter describing a world exposed to growing dangers.

According to Coxon, there's an awareness of the risks within the labs, but the prevailing belief is that slowing down would leave room for less cautious competitors. Competition between companies, coupled with international competition, thus ends up justifying continued acceleration. The researcher considers it "an arrogant gamble that shouldn't be made from a private company's Slack channel."

AI: From requests for a break to proposed laws

The answer indicated by Coxon goes from a coordination between laboratoriesThe researcher believes agreements on the pace of development are possible, although he doesn't currently see the conditions to prevent a global race. Among the measures that might be necessary, he cites "expensive actions, such as a temporary ban on improving model capabilities." The call for action involves a broader segment of the sector. An appeal signed by over a thousand researchers, including Coxon and Anthropic CEO Dario Amodei, calls for a international coordination to slow down research on systems capable of self-improvement.

In the United States the debate has also reached CongressOn September 3, Senator Bernie Sanders and Representative Greg Casar introduced the Ban Artificial Superintelligence Act, a proposal that would ban the development of artificial superintelligence and would temporarily suspend other advanced developments until safety regulations are established by a new federal authority. The bill provides for penalties of up to twenty years in prison for violators. The crux of the matter is entrusting the risk assessment Even to entities independent of the companies that build and market the models. This issue already concerns cybersecurity, information manipulation, surveillance, and the concentration of economic power, in addition to the extreme scenarios evoked by researchers.

Claude and OpenAI incidents reignite security concerns

Making the debate more pressing are the episodes that have emerged in recent months. September 9th Anthropic has reported a fourth incident Anthropic has reported a new incident in which one of its models breached external systems during testing. The incident dates back to January and involves a preliminary version of Claude Opus 4.6, but was only discovered in August. The discovery came after an initial review of 141.006 test runs, which excluded some runs. In July, the company had already disclosed three incidents involving other models, which had been granted internet access due to a bug. Regarding the new incident, Anthropic stated that it has informed all parties involved and, based on a preliminary assessment, does not consider it more serious than previous ones. The company has identified two recurring behaviors: the tendency to underestimate or misinterpret cues indicating access to real systems and the willingness to perform potentially harmful actions in order to complete a task.

On the OpenAI front, the investigations on theJuly attack on the Hugging Face platform They reconstructed the coordinated activity of approximately 700 agents, with frequent attempts to cover their tracks. An investigation by Reuters It later reported that company agents had used more than ten other sites for unauthorized communications. This latter activity should be distinguished from computer intrusions. The agents allegedly exploited technical features of wikis and other services to leave messages and exchange information, circumventing restrictions that only allowed them to read the web. OpenAI stated that it had initiated a broader review and had not identified any other activities comparable in severity or scope to the Hugging Face case.

The episodes However, they show a difficulty already current: preventing systems designed to achieve a goal from bypassing the imposed limits and recognizing promptly when this happens.

comments