Anthropic’s CEO wants to slow down AI development. It warns of what may happen in 6-12 months

Until recently, the largest AI companies were racing to see who would be the first to show a more powerful model. Now one of the most important people in the industry says directly: you need to slow down.

Dario Amodei, co-founder and CEO of Anthropic, the company behind Claude, published an essay titled “We Must Pace the Frontier.” His main thesis is simple: the development of the most advanced artificial intelligence models is beginning to outpace our ability to control and study them.

This is important because Amodei is not an external critic of AI or an opponent of this technology. He heads one of the laboratories at the very center of the race for the most advanced models, and Anthropic is today one of the most important competitors of OpenAI and Google DeepMind.

The reaction of the competition was even more interesting. Sam Altman, CEO of OpenAI, publicly admitted that Amodei was right in part of the diagnosis and announced the implementation of one of the security mechanisms he proposed also in OpenAI.

“We need to slow down.”

Amodei starts by reminding that he is still a technological optimist. He believes that the development of AI has the potential to significantly accelerate disease research, economic growth and scientific progress, and the potential benefits remain enormous.

The problem is that in recent months his risk assessment has started to change. According to Anthropic’s CEO, it is no longer enough to develop more and more powerful models and invest in security at the same time. You also need to give security teams more time to understand the new capabilities of the systems.

In practice, it is not about completely stopping the development of AI. Rather, Amodei argues that model capabilities should not grow so quickly that control, audit and security mechanisms cannot keep up.

The head of Anthropic indicates two main reasons why, in his opinion, the situation is becoming more dangerous than just a few months ago.

AI is starting to help create better and better AI

The first is a phenomenon known as recursive self-improvement. This is not yet a movie scenario in which artificial intelligence independently rebuilds its code and becomes superintelligence from hour to hour.

The mechanism is more mundane. AI models are increasingly used by researchers themselves for programming, conducting experiments, analyzing data, automating research and building further AI systems. So a better model helps create an even better model, which can then speed up work on the next one.

According to Amodei, this process has started to gain much more momentum in recent months. If feedback becomes strong enough, the development of model capabilities may begin to outpace humans’ ability to examine, test, and control them.

This is probably the most important element of the entire essay. The discussion about the rapid development of AI is no longer just about faster chips, larger GPU clusters or billions of dollars spent on training. Artificial intelligence itself is becoming more and more important in this equation, helping to develop subsequent generations of systems.

The second alarm: swarms of agents began to do things no one had asked them to do

Amodei also devotes a lot of space to the recent incident related to OpenAI and Hugging Face. During cybersecurity tests, OpenAI models found ways to bypass parts of the protections that isolated them from the internet and began to perform outside the range that researchers had predicted.

Agents communicated with each other, delegated tasks to each other and created their own cooperation structure. At one point, they executed code on Hugging Face’s servers, gained administrator privileges on one of the systems, and gained access to part of the infrastructure that was not the original purpose of the test.

This is especially important because the models were not given the “attack Hugging Face” task. They were supposed to solve cybersecurity problems in a controlled environment, but along the way they found ways to go beyond the imposed limitations.

OpenAI later described the incident and implemented additional safeguards. At the same time, the company emphasized that the event did not affect customer data or the availability of its products.

Amodei: In 6-12 months, a similar swarm could be much more dangerous

The strongest warning comes when Amodei tries to translate the current pace of model development into the coming months. In his opinion, a similar swarm of agents, but with the capabilities of models that may appear in the next 6-12 months, could pose a threat of a completely different scale.

The head of Anthropic is considering, among other things, a scenario in which such a system would be able to build a persistent botnet covering a large part of the Internet or carry out cyber activities much faster than current tools. He does not present this as a certain prediction, but as a risk scenario based on the assumption that agent capabilities will continue to grow at the current rate.

And that’s why his proposal is not “let’s stop AI.” It’s more about buying yourself additional time that can be used to develop alignment methods, security tests, monitoring and better security of the infrastructure.

According to Amodei, even an extra year or two could make a huge difference if used to reduce risk before the next generation of much more autonomous systems emerge.

The Anthropic plan consists of 3 stages

Amodei proposes three levels of action. The first one can be implemented almost immediately and Anthropic declares that it wants to do it unilaterally.

The company intends to provide independent external experts with constant access to its systems at a level similar to that of employees. Auditors would check whether Anthropic actually follows its own safety rules, analyze incidents and assess the behavior of models during the training process.

This is an important change because currently much of the knowledge about the most powerful models comes from the companies that simultaneously create, sell and decide on their publication. The external audit would limit the situation in which the laboratory effectively assesses its own safety.

The second stage assumes cooperation between the largest AI laboratories operating in democratic countries. The companies would jointly set safety standards and limits on the development of the riskiest model capabilities.

However, this is where a more difficult problem begins. If one company slows down and its competitor doesn’t, there is a strong temptation to speed up again. In practice, the entire market operates in conditions resembling an arms race in which each player is afraid of being left behind.

Therefore, the third stage involves an attempt at global coordination, also with China.

What about China?

This is one of the biggest problems of the entire concept. The United States may introduce restrictive rules for OpenAI, Anthropic or Google DeepMind, but if similar restrictions do not apply to laboratories operating in China, American companies may simply lose their technological lead.

Amodei realizes this. Therefore, in parallel, he advocates, among other things, limiting China’s access to the most advanced AI chips, hindering the smuggling of systems and protecting models against theft of scales and copying their capabilities through distillation.

So a certain paradox arises. At the same time, Amodei wants to slow down the pace of development of the most advanced AI and maintain the United States’ advantage over China. Without international coordination, achieving both goals may be very difficult.

Sam Altman: I agree with Dario

The biggest surprise, however, came from OpenAI. Altman himself wrote publicly that he agrees with part of Amodei’s diagnosis, and the issue of limiting the pace of development of the most advanced systems has been one of the important topics of discussion within OpenAI in recent weeks.

Altman also supported the idea of ​​independent auditors with access similar to employees and announced that OpenAI will use a similar mechanism. Representatives of other companies operating in the AI ​​market also expressed support for greater transparency.

This is an unusual situation. Labs that compete for customers, developers, investors and the world’s best researchers are beginning to publicly agree at least in part on an issue that directly affects the pace of their own development.

Europe has already partially moved in this direction

From Poland’s perspective, there is one more interesting element. Some of the solutions demanded by Amodei already exist in some form in the European Union.

The AI ​​Act requires providers of general-purpose models deemed to pose systemic risk to, among other things, conduct evaluation and adversarial testing, assess and mitigate systemic risk, report serious incidents and ensure an appropriate level of cybersecurity.

This does not mean, of course, that the European Union has introduced the “frontier inhibition” proposed by Amodei. The AI ​​Act does not order laboratories to slow down the pace of work on new models, but it shows that mandatory testing, incident reporting and systemic risk control are no longer just an academic discussion.

Biggest problem? Everyone has a reason not to be the first to slow down

Amodeia’s entire proposal has one fundamental weakness. Anthropic may decide that it is too risky to further increase model capabilities, OpenAI may come to a similar conclusion, but the situation will change immediately when one competitor accepts more risk in exchange for a technological advantage.

The stakes are huge, because a better model today means much more than a better chatbot. It is potentially a better tool for software development, scientific research, drug design, cybersecurity, robotics, business automation and building the next generation of artificial intelligence.

Therefore, the greatest significance of Amodei’s essay may not lie in whether his three-step plan will be implemented. The very moment when such a proposal appears publicly is much more interesting.

Until recently, the industry’s biggest question was how quickly we could develop AI. Today, people managing companies at the very front of this race are increasingly asking a different question: should we continue to develop it so quickly?