NSA Deputy Director Tim Kosiba stated at a panel discussion in Bethesda: “We want access to all models, and we will use it.” The agency is in negotiations with frontier developers and is claiming a key role in the voluntary pre-release testing program created under the June White House executive order. The program gives the government up to 30 days to work with closed frontier models before their public release.

image

What Happened

The mechanics of the program under the June White House executive order are structured as follows: closed models that qualify as covered frontier models are voluntarily handed over to the government for up to 30 days before public release and undergo classified tests of the “hacking” capabilities of AI systems. The covered threshold is determined by the NSA Director jointly with the Office of the National Cyber Director and CISA, meaning the model selection criteria are centralized and not publicly fixed. The decision on who is included in the program is de facto made by the same trio of agencies. The program only applies to the closed systems of OpenAI, Anthropic, and Google, while open-weight models are not included. Mandatory model licensing is explicitly prohibited by the executive order, so there is no coercive part to the procedure — company participation is formally voluntary.

Context

The Anthropic precedent shows the real cost of refusal: the Department of Defense designated the company a supply-chain risk, the dispute went to court, and a federal judge temporarily blocked part of the government's actions. Formally the format is voluntary, but in practice it is backed by pressure on federal procurement. The NSA has already secured a special status in this structure: after the dispute with the Department of Defense, the agency obtained carveouts and is experimentally using the closed Anthropic Mythos model — some models like the first Mythos do not go into public access at all due to zero-day capabilities that the state keeps for itself. From an engineering perspective, this event is not about the models, but about the evaluation infrastructure: 30 days of pre-release access to closed weights is effectively external red-teaming with much deeper access than any public API. Such a check is not reproducible for external researchers, because the classified format provides neither methodology, nor metrics, nor datasets. The irony is that the only class of models available for independent audit by the community — open-weight — remains outside the government audit.

Why This Matters for the Industry

For the OpenAI, Anthropic, and Google labs, these are specific terms for negotiations on pre-release access: one-off agreements are turning into a structured norm with a 30-day window. If major players sign voluntary agreements, this window will become part of the release cycle for closed models: versions will be held back longer before rollout, and migration and deprecation schedules will become less predictable. The key fork in the road for the next six months is how the NSA Director, together with ONCD and CISA, will draw the boundary of a covered frontier model: the number of releases that will go through the 30-day cycle depends on this threshold. A new vendor risk question is emerging in the vendor assessment of closed models — whether the provider is participating in negotiations with the NSA and whether its release could have passed the check. A scenario worth tracking, but not to be taken as fact: the list of signing labs becomes a de facto procurement checklist, and the market splits into two tracks — closed systems with government checks versus open-weight models outside this framework. For those selling to the US government sector, the federal distribution channel is already familiar: GPT-5.6 is available to federal customers through a FedRAMP-authorized service.

Why This Matters for Users

For current users, there is no direct effect: the news introduces neither a new API nor a new UX pattern, and endpoints work as they did before. If you use open-weight models, they are outside the NSA's pre-release check — the government audit only covers closed systems, while open-weight remains the only class of models that can be independently checked by yourself. Practically, it is worth keeping an eye on two things: which labs will sign the voluntary agreement and how the NSA Director will determine the covered frontier model threshold — this will determine which future releases will go through the 30-day cycle and with what delay they will appear. If you are targeting the US government sector, the check channel is already effectively working through FedRAMP-authorized services. The material also explains why some models, like the first Anthropic Mythos, do not go into public access at all: their zero-day capabilities remain inside the state.

What Is Still Unknown / Limitations

Classified tests without disclosed methodology, metrics, datasets, and the ability to reproduce do not provide scientific evidence of either the safety or the quality of a model, so the status of “checked by the government before release” cannot be transferred to risk registers as a positive signal. Kosiba's statement is a declaration of intent and a negotiating position, not technical evidence; the results of the tests will not be visible to an external observer. It is still unknown which labs will actually sign voluntary agreements, what the covered frontier model threshold will be, and how many releases will go through the 30-day cycle. The two-track market scenario — closed models with government audit for critical infrastructure versus open-weight outside it — is an interpretation that needs to be tracked, not an established fact.

Sources

Author

Look at AI, editorial team