According to The Information, Google, OpenAI, and Anthropic are close to creating an independent standards body for frontier AI with the working name Standards Authority for Frontier AI (SAFA), which is intended to operate autonomously and fill gaps in government regulation. In parallel, OpenAI and Anthropic are discussing a mutual testing agreement for their commercial models. The speeches by the heads of OpenAI and Anthropic at the UN Security Council confirm that rules for advanced AI are now being determined not only in laboratories but also at the political level.

image

What happened

According to The Information, cited by Proactive Investors, the three largest developers of frontier models — Google, OpenAI, and Anthropic — are close to creating an independent standards body with the working name Standards Authority for Frontier AI (SAFA). The body is designed to be autonomous and is intended to fill gaps where government regulation does not exist. In parallel, OpenAI and Anthropic are discussing a mutual testing agreement: the companies will gain access to each other's commercial models for independent stress testing for vulnerabilities and unexpected behavior, with restrictions on storing the data obtained. On September 23, 2026, Sam Altman and Dario Amodei spoke at the UN Security Council: Altman stated that key decisions on AI should be made through democratic institutions and governments, while Amodei called for international agreements and global standards for testing new models.

Context

The initiative is based on a regulatory vacuum: there are no mandatory government standards for testing frontier models, and vendor safety assessments are mostly self-reported and methodologically conflicting. The Trump administration is against new restrictions for the industry, so flagship laboratories are demonstrating their readiness to fill regulatory gaps themselves. The OpenAI and Anthropic agreement under discussion is essentially cross-lab red-teaming: competitors are giving each other access to commercial models, which moves assessment beyond the developer's internal reporting and reduces the conflict of interest of "developer = assessor." At the same time, the topic of frontier AI is definitively moving from the engineering plane to the political one, and rules may be determined faster than expected.

Why this matters for the industry

If SAFA is formed, the three largest developers of frontier models will for the first time have a common independent mechanism for standards and mutual testing — a de facto industry regulator where there is no government oversight, and the mutual stress testing agreement will become a rare precedent of such transparency between direct competitors. For startups, this is a double signal: the market for eval audits and compliance services is growing, but at the same time there is a risk that standards will be written to suit the interests of the three founding participants. In a successful scenario, safety testing transforms from a PR activity into an industry function and over time may become a de facto condition for purchasing models: model verifiability as a feature, reliability metrics in the user experience, mandatory evals in CI for agentic systems. For now, however, this is a governance level, not an API or SDK, and there is nothing to incorporate SAFA into architectural decisions.

Why this matters for users

For readers, this is currently a signal, not a product: it does not yet affect models, benchmarks, or access. However, if the results of checking other companies' commercial models become at least partially public, independent benchmarks for the safety of the models you use will appear — a comparison of behavior outside vendor marketing reports. The practical step now is to monitor SAFA documentation and the terms of mutual testing and not to take negotiations for an already launched body. If the agreements are signed, closed pilot tests with NDAs and limited disclosure are likely; the main marker of real benefit will be whether public protocols and statistics on found vulnerabilities appear, in which case dashboards and bots tracking model comparisons will quickly emerge around the reports.

What is still unknown / limitations

SAFA is a working name: the body does not exist as a documented structure, the mutual testing agreement has not been signed, and results are not published. The sources contain no test protocols, metrics, criteria for "unexpected behavior," or rules for disclosing results; without a public methodology, any future conclusions are indistinguishable from marketing. Restrictions on storing data obtained during testing are logical from the perspective of protecting intellectual property and preventing leaks, but they may technically limit the depth of analysis: serious evals require accumulating statistics and long-term work with the model. OpenAI's call for international work on technical standards for systems with recursive self-improvement runs into the absence of an established methodology for such assessments — standardizing "forward" risks cementing weak practices. Finally, the outcome depends on negotiations and the political background: a scenario is possible in which the body remains a declaration and its standards do not receive public disclosure.

Sources

Author

Look at AI, editorial team