At the All-In Summit in Los Angeles on September 14, Elon Musk proposed turning AI lab competition into a safety mechanism: before public release, developers would gain access to a shared test harness and test competitors' models. According to him, xAI, OpenAI, Anthropic, Google, Meta, and three to four leading Chinese companies should participate. The proposal currently exists only as a statement: no competing lab has agreed to participate, and it came amid rare unity among industry leaders around slowing frontier AI development.


What happened
Speaking at the All-In Summit in Los Angeles on Monday, September 14, Elon Musk proposed that the largest developers of frontier AI models share access to a shared test harness — a standard testing platform — and test each other's models before public release. In his design, xAI, whose business was merged with SpaceX in February, OpenAI, Anthropic, Google, Meta, and "three to four leading Chinese companies" should participate in the scheme. Musk acknowledged that the method is not perfect, but stated that "the chances of finding problems will become significantly greater," and expressed confidence that Chinese companies would likely agree to such testing.
Context
The proposal came amid a notable shift in the public debate on frontier AI safety. Dario Amodei published an essay on a pause in frontier model development, supported by Sam Altman and Musk himself — figures who usually compete in the same market; such unity among industry leaders has been rarely noted. The debate was intensified by the departure of researcher Jacob Coxon from Anthropic, who stated that labs are "putting our lives at risk," and Evan Hubinger from Anthropic supported the assessment that the risk of human extinction from AI within a decade exceeds 10 percent. Political pressure is going in the opposite direction: Trump called AI safety concerns a "hoax" and a "scam," and China's Ministry of Foreign Affairs called them "fear mongering." Against this backdrop, the idea of mutual testing looks like an attempt to find a safety mechanism that requires neither government regulation nor a voluntary halt to development.
Why this matters for the industry
For the industry, the essence of the proposal is that competitors become independent red-teamers: a rival has a direct incentive to find a vulnerability that the model's creators missed, and external adversarial testing is usually more effective than internal evals, where the same team makes the tests and the model. Such a mechanism would be an alternative to both government regulation and self-restraint, which the geopolitical race with China makes politically difficult. The main barrier is economic: providing the best model to competitors before release means revealing a competitive advantage, so labs whose business is built on closed frontier models are more likely to participate formally or refuse altogether. If the scheme takes root even in a reduced form, a category of neutral testing infrastructure could grow on its basis: common eval protocols, red-teaming, machine-readable evaluation reports, and for frontier model developers — a permanent expense on safety audits.
Why this matters for users
For the reader, the main point is that the abstract debate "AI is dangerous, we need to slow down" has for the first time been formulated as a specific working mechanism involving companies whose models people use every day. If mutual pre-release testing becomes the norm, new models will be released less frequently, but with fewer untested vulnerabilities, meaning a choice will have to be made between the speed of feature appearance and their reliability. Specific indicators to watch: official responses from OpenAI, Anthropic, Google, Meta, and xAI to Musk's proposal, the appearance of pilot external evaluations, and the publication of evaluation reports before model releases. If even one major lab agrees to external testing, it will be a signal of a real change in safety policy, not just rhetoric.
What is still unknown / limitations
The proposal currently exists as a public statement: there is no methodology, no test harness specification, and no criteria by which a model is considered to have passed testing. It is unknown which version of the model access is given to — a snapshot, API, or weights — and how found problems will be verified and published. No competing lab has agreed to participate yet. Musk's statement that China would likely agree is his personal assessment: consent from Chinese companies has not been provided, and China's Ministry of Foreign Affairs previously called AI safety concerns "fear mongering."
Sources
- Musk urges top AI labs, Chinese companies to test each other's models amid calls for slowdown — CNBC
- Elon Musk calls on rival AI labs and Chinese companies to test each other's models amid calls for AI slowdown — Tech Startups
- Hacker News discussion: Musk urges top AI labs, Chinese companies to test each other's models
Author
Look at AI, editorial team
