The US Department of Justice publicly supported OpenAI in The New York Times' copyright lawsuit: on September 2, 2026, the department filed a statement of interest in the federal court in Manhattan, calling the training of large language models on copyrighted texts fair use. This is the first direct intervention by the US federal government in the central copyright case of the generative AI era. There is no decision on the case itself yet.

image
image

What happened

On September 2, 2026, the US Department of Justice filed a statement of interest in the federal court in Manhattan, meaning the department's opinion on the case, in support of OpenAI in The New York Times' lawsuit filed in December 2023 against OpenAI and Microsoft. The department's position is that training large language models on copyrighted texts is fair use; the department warned that "the US has a substantial interest in the court rejecting any arguments that LLM training on copyrighted texts violates copyright." Among the arguments are national security, as foreign rivals "will not be limited by copyright," and competition: according to the department, only the largest companies could afford licenses, and the "creative possibilities and public benefit" of LLM training "outweigh any competitive harm." In response, The New York Times stated that the administration is siding with a handful of trillion-dollar AI companies.

Context

The filed document is a procedural statement of the department's position, not a court decision, so it does not change the rules for the market today. The stakes in the process are high: its outcome will determine the framework for all disputes over model training on others' data, from copyright lawsuits to licensing negotiations with publishers. Within the administration itself, the position is contradictory: expert Evan Schwartztuber notes that the administration's March AI framework assumed collective licenses for media, meaning a model in which publishers would be paid for the use of their materials in training. Technically, fair use disputes usually come down to memorization: the question is whether the model reproduces a copyrighted expression or learns from non-copyrighted statistical patterns.

Why this matters for the industry

For the industry, the effect is currently signaling, not technical: training pipelines, datasets, and existing contracts continue to operate as before. The negotiation landscape is changing: the thesis that "training is fair use" is now supported at the federal department level, which strengthens the hand of AI companies in current licensing negotiations with publishers and reduces the perceived legal risk of training on open data. The DOJ's competition argument is an explicit acknowledgment that the cost of access to data concentrates training among giants: the department's position is de facto against licensing data moats and indirectly in favor of broad access to data. If the Manhattan court takes the statement into account, the legal risk premium for public data will decrease, the pace of mandatory licensing will slow down, and some products that were delayed due to data risks will be released faster; in the opposite outcome, uncertainty will remain and the licensing model for data access will be strengthened.

Why this matters for users

For product users, nothing changes today: existing AI services, subscriptions, and publishing deals work as before. The practical significance of the signal is for ML teams and laboratories planning training on web corpora: the legal assumptions of their data strategy, whether it is training on open corpora, buying licenses, or a RAG approach, now depend on how the Manhattan court evaluates the department's position, so the internal risk assessment of their own data corpora should be updated now. For readers of news publications, the dispute determines how their favorite materials will end up in AI products: with a sustained recognition of fair use, the data market may be divided into "free for training" and "premium by license" layers, and value will shift to products that create benefit on top of any data — agentic scenarios, model memory, and verified sources with clear attribution.

What is still unknown / limitations

It is unknown whether the Manhattan court will take the department's statement into account, to what extent, and when decisions on the merits of the lawsuit will follow. The national security argument that "foreign rivals will not be limited by copyright" assumes a direct connection between access to training data and model capabilities, but there are no public measurements in the sources of how much limiting licensed texts actually slows progress. The DOJ's position relies on legal argumentation, not on measurable model behavior: in available publications, there are no studies of memorization or extraction of protected expression. Long-term scenarios of restructuring data licensing markets remain speculative until procedural outcomes.

Sources

Author

Look at AI, editorial team