Stuart Russell, a computer science professor at UC Berkeley and president of the International Association for Safe and Ethical AI, published a column in The Guardian criticizing Dario Amodei, CEO of Anthropic, for proposing to develop frontier models in a 'paced' mode — at a reduced speed under common standards. Russell calls this approach deeply flawed and proposes an aviation certification analogy: a model with a new capability should only be released with a certificate for the corresponding safety properties. The response to this argument will determine how the market entry of frontier models will look in the future.

image

What happened

The trigger for the column was a letter of about 3,800 words, 'We Must Pace the Frontier,' written by Dario Amodei and supported by Sam Altman, Elon Musk, Demis Hassabis, and Satya Nadella. In the letter, Amodei uses the metaphor of a safety car in Formula 1: companies continue development at a reduced speed, buying time for safety checks, with third-party auditors inside each lab and common standards for 'democratic countries.' Russell rejects this logic in his Guardian column and offers a contrasting analogy: just as Boeing does not release an aircraft until it has passed all tests and received an airworthiness certificate, models with capability X should only be released with certification of the corresponding properties. He separately notes that Amodei, in his own letter, effectively agreed with the logic of red lines by using the phrase 'certifications of alignment properties Y and Z.'

Context

The dispute between Russell and Amodei is not about the speed of development as such, but about the point of market entry for a model: an F1 'pace car' versus a 'red flag' with pre-certification. According to the figures cited in the column, AI lab leaders themselves assess the risk of catastrophe as 1 in 10 to 1 in 5, while the generally accepted acceptable level of loss of control is about 1 in 100 million per year; this gap of several orders of magnitude makes current release-decision practices unjustified even without references to regulation. At the same time, this is currently a public debate, not a regulatory requirement: the column and the letter are arguments, not standards, and they do not change any rules or procurement. The article mentions a possible 'window of change' around the Trump and Xi summit as a scenario in which the discussion could evolve into specific standards; this is a supposition, not an established event.

Why this matters for the industry

The column changes the frame of the regulatory debate: instead of 'slow down to buy time,' it proposes a strict market-entry model where safety is a precondition for releasing a model, not a research program running alongside a race. The fact that a proponent of the 'paced' approach himself allowed the formulation of certifying alignment properties gives regulators and labs a ready-made anchor for demands to pause releases that do not meet standards. If the certification scenario is at least partially realized, releasing a model will become more expensive and slower for everyone, but the burden will fall disproportionately on startups: they will have to maintain compliance teams and undergo external audits with a shorter runway to a cash-flow gap. In parallel, a framework for future demand for compliance tools is forming — eval sets for specific properties, audit trails, and reporting for third-party auditors, similar to how regulatory data requirements spawned the privacy-tools industry.

Why this matters for users

The debate does not yet bring direct product changes: this is a regulatory debate, not a feature release, so serving, latency, and inference costs are not affected by it. For ML researchers working with evals and alignment, the vector itself is important: if the certification framework reaches practice, it will determine which measurements become mandatory and who will recognize their results, and safety evals will turn from a research practice into a certification artifact with legal force and external auditors. A reasonable action now is cheap: conduct an inventory of eval coverage, logging, and release reproducibility, and include an optional audit-readiness module in the roadmap — logging of agent decisions and reproducible eval reports — without making it a mandatory layer. The 'certification instead of paced' framework is already useful as an argument in negotiations with enterprise clients and in positioning safety products.

What is still unknown / limitations

The main weakness of both positions is the lack of methodology: neither the letter nor the column contains a method for measuring properties Y and Z, and the source provides no evidence that standardized, reproducible, and gaming-resistant evals for 'alignment properties Y and Z' even exist. The objection also concerns the aviation analogy itself: aircraft certification works because it relies on a set of standardized, physical, reproducible tests, and transferring this model to alignment properties without such methodologies remains rhetoric. Operational questions also remain open: which specific properties to certify, which thresholds to consider sufficient, who will audit, and how to ensure auditor independence. Finally, the resonance of the discussion at the time of publication was low — the Hacker News discussion received 2 points and 0 comments — so the current public weight of the debate should not be overestimated.

Sources

Author

Look at AI, editorial team