On September 18, 2026, a group of researchers led by Martin Jaggi and Robert West (EPFL), Philip Torr (Oxford), and Anna Hedström (ETH Zurich) published an open appeal titled “A Call for Open Science in AI Safety.” The authors ask developers of frontier, i.e., the most powerful, AI models to publish safety methods: evaluations, safety training recipes, code, data, and evidence of desired and undesired model behavior, so that independent researchers can verify, reproduce, and improve these measures. By the time of verification, the appeal had been signed by about 208 people, including employees of Google DeepMind, Hugging Face, and EleutherAI, with all signatures declared in a personal capacity.

image

What happened

The appeal is hosted on a dedicated page at make-safety-open.github.io, where a list of signatories and a signing form are also maintained. The authors demand that labs disclose verifiable safety artifacts: conducted evaluations, safety training recipes, corresponding code and data, as well as evidence of how models behave in desired and undesired scenarios. The text includes a caveat: misuse-sensitive information may be withheld, but each such exception must be targeted and proportionate. Among the signatories are Fabian Pedregosa and Alexander Ram (Google DeepMind), Lewis Tunstall (Hugging Face), and Stella Biderman (EleutherAI), as well as researchers from MIT, Princeton, Stanford, Oxford, and the University of Zurich. The authors promise to update the document and the list of signatures.

Context

The discussion about AI safety has so far been largely built around closedness: labs test their own models themselves and publish only brief reports, while external researchers are left with assurances that cannot be verified. The standard objection to disclosure is well known: training recipes and data may reveal trade secrets or help malicious actors. The authors of the appeal respond to this by choosing the unit of disclosure: they request methods and evidence, not model weights, so scientific verifiability is separated from the most commercially and misuse-sensitive part of development. The composition of signatories shows that the idea is finding resonance within the labs themselves, although none of them as an organization has made any statements.

Why this matters for the industry

For the industry, the appeal shifts the focus of the debate on AI safety from closedness to verifiability: labs will likely have to prove safety with artifacts, not formulations. If academic community pressure works, frontier labs may expect requirements to publish safety recipes and evaluation reports in a verifiable form, which will affect model card practices and the order of safety reporting before releases. For startups, this is an early market signal: around independent verification and safety infrastructure — repositories of eval datasets, harnesses, reporting templates — a separate market layer may form. For now, this is community pressure without obligations on the part of the addressees, and the most honest indicator is one: whether at least one frontier lab will produce a reproducible artifact in response to the appeal.

Why this matters for users

For the reader, the appeal is useful as an explanation of what AI safety consists of in practice: it is not abstract assurances, but a set of specific verifiable documents and artifacts that can be read, requested from vendors, and compared with each other. The text of the appeal and the list of about 208 signatories are open on the website, the appeal can be signed through a form, and the document is promised to be updated, so it is convenient to track the list as a map of the community. Engineers and team leaders can apply the framework of the appeal right now: transfer it to internal requirements for vendors and contractors, requesting methods and evidence of conducted checks, not just final statements about safety.

What is still unknown / limitations

The appeal does not set any metrics, reporting formats, or a threshold of sufficient transparency: it is unknown what to consider sufficient publication of evals, who acts as an arbiter, and how to verify the completeness of disclosure, so the risk of declarativity is high. This is a meta-scientific declaration, not a technical result: it has no API, releases, or benchmarks, and there is nothing to deploy yet. The practical effect today is zero, the discussion on Hacker News at the time of verification was limited to one point without comments, and there are no institutional obligations from any lab, so there is no evidence base for a change in practice. The number of signatories is fixed at the time of verification and will change, and any predictions about the reaction of the labs remain interpretations until specific published artifacts appear.

Sources

Author

Look at AI, editorial team