On September 17, 2026, OpenAI introduced Astra for Law — a vertical configuration of the GPT-6 Astra model for legal work, not a new model. Instead of training a separate network, the company connected it to a Legal Search Index for US law with over 230 million URLs and supplemented it with instructions for legal analysis and writing. On 200 questions from a private validation sample of the Vals AI Legal Research Bench, the configuration scored 54.0% compared to 38.7% for the base GPT-6 Astra with web search. The product cannot be tried yet: access is open to select firms, and the API is promised later.

image
image

What happened

Astra for Law is built on the same GPT-6 Astra weights: what changed is not the language itself, but what the model turns to for facts. Instead of general web search, the configuration uses a Legal Search Index — over 230 million URLs for US law: case law, statutes, regulations, and agency decisions. The foundation of the index was a partnership with the non-profit Free Law Project/CourtListener, which covers more than 99.9% of published US case law. On top of the index, OpenAI supplemented the model with instructions for legal analysis and writing. OpenAI's key measurement was conducted on 200 questions from a private validation sample of the Vals AI Legal Research Bench: the configuration scored 54.0% compared to 38.7% for the base GPT-6 Astra with web search — this is a 40% increase in relative terms. According to OpenAI, the configuration extracted up to 54% more relevant fragments from correct court decisions. At the same time, the company opened 26 partner plugins — among the named partners are Thomson Reuters, Harvey, Legora, iManage, Relativity, Clio, Intapp, and DeepJudge — and 47 community plugins, and also made a ChatGPT add-on for Microsoft Word publicly available. The product itself remains closed: it is available to select firms through Trusted Access in ChatGPT and Codex, and the gpt-6-astra-law API endpoint is promised later.

Context

The main point of the release is not a new model capability, but an architectural conclusion: the accuracy increase was achieved by the same GPT-6 Astra, where only the retrieval layer and instructions changed, but not the weights. It follows that the main contribution to the result is made by specialized search over a curated legal index, not a new ability that appeared in the model. The same pattern is already used by Claude's domain configurations, so vertical packaging of frontier models as a combination of "model + index + instructions" is ceasing to be an experiment and becoming a working industry recipe. The second contextual element is the measurement methodology. The Vals AI Legal Research Bench has an open public leaderboard, where Muse Spark 1.3 Max and Claude Opus 5 currently lead with a result of 55.29% on the public set. OpenAI's measurement was performed on a private validation sample of 200 questions, so the claimed 54.0% is not directly comparable to the public list: these are two different samples of the same benchmark. The third element is the openness of the index foundation. CourtListener is a non-profit project, its database is potentially reproducible by other teams, which distinguishes the legal domain from closed corporate data. The limitation is that the coverage is limited to US law, and the release does not show how the pattern behaves in other jurisdictions and in other languages.

Why this matters for the industry

For the industry, this is the first case where a frontier model has been packaged as a vertical product for lawyers: the combination of "model + specialized search index + industry instructions" provided an accuracy increase on the benchmark without any fine-tuning. For product builders, this is a dual-action signal: legal research and citation extraction cease to be a protected niche of thin wrappers over web search, and value shifts to curated indexes, connectors, and integration into firm workflows. Connectors to iManage, Relativity, Clio, and Thomson Reuters, plus integrations with Harvey and Legora, change OpenAI's position in the market: the company becomes an intelligence layer under existing legal software, not a replacement for it. Vendor competition shifts from chats to workflows and firm data management — ethical walls and access rights were developed by OpenAI together with Latham & Watkins, meaning the stake is not only the model, but also who controls client data. If the promised gpt-6-astra-law API is released, a wave of legal-tech builds on top of it is expected: automation of contract analysis, due diligence, and IPO documentation preparation, as well as a fight for connector quality. Within a couple of years, vertical configurations of "model + index + instructions" will likely become standard packaging for frontier models in regulated industries, and benchmarks will begin to measure the "model + retrieval" combination — with independent reproduction and citation faithfulness metrics.

Why this matters for users

Astra for Law cannot be tried now: access is open only to select firms through Trusted Access in ChatGPT and Codex, the API endpoint is not yet open, and prices and latency are not disclosed. Of the announced set, only the ChatGPT add-on for Microsoft Word is publicly available. It is practical to watch two things: the open Vals AI Legal Research Bench leaderboard, where the results of all participants on the public set are visible, and the future gpt-6-astra-law API — its release will be the first moment when OpenAI's configuration can be compared with competing solutions on the same sample and US legal search can be integrated into one's own pipelines, if the promised Zero Data Retention is confirmed. Lawyers and everyone working with US law should remember the caveats: the claimed result is OpenAI's own measurement on a private sample chosen by it without independent reproduction, and almost half of the questions in this sample were still failed. Mandatory verification of citations by a lawyer is not yet canceled.

What is still unknown / limitations

The main unknowns are related to methodology: the measurement on 200 questions was conducted by OpenAI itself on a private validation sample chosen by it and has not been reproduced by independent parties; there is no direct comparison of the configuration with Muse Spark 1.3 Max and Claude Opus 5 on the same sample yet. The report does not contain metrics for incorrect or fabricated citations — there is no data on citation faithfulness, so the reliability of references for real legal practice cannot be assessed. Commercial parameters are not disclosed: prices and latency are unknown, and the API endpoint is not open. Index coverage is limited to US law, and transferring the "model + index + instructions" pattern to other jurisdictions, languages, and industries is a hypothesis, not a proven result: one vendor self-assessment in one domain does not by itself establish transferability.

Sources

Author

Look at AI, editorial team