In the case of The New York Times v. OpenAI and Microsoft, 92 pages of internal documents have been unsealed, showing that both companies were aware in advance of the consequences of their AI strategy for the web. Microsoft's Director of Applied Sciences, Brent Hecht, called the collection of training data an "unprecedented, staggering theft," while an internal Microsoft document acknowledged that the AI strategy had triggered a "doom loop" undermining the economics of models and the entire web. The outcome of the lawsuit will determine whether training models on others' texts constitutes a copyright violation, and this will determine how readers will access news.

image

What Happened

As part of the proceedings initiated by The New York Times, 92 pages of internal materials from OpenAI and Microsoft, along with testimony from executives of both companies, have become publicly available. Satya Nadella confirmed under oath that chatbots "replace" visits to original websites, while ChatGPT head Nick Turley called AI products an "existential threat" to publishers who have "no compelling reason to click" on links. An internal OpenAI note from June 2022 records the expectation that GPT-4 would "memorize a ton of data and be insanely good at regurgitation"; meanwhile, company representatives were unaware of attempts to find or remove paywalled content from the datasets. Among the unsealed materials are cases of verbatim reproduction of articles from The New York Times, Mercury News, and The Denver Post.

Context

The central issue of the proceedings is the applicability of the fair use doctrine to the training of language models on copyrighted texts. The key precedent here is considered to be the 2023 Andy Warhol Foundation v. Goldsmith decision, where the replacement of the original by a derivative work became the main argument against fair use, so testimony about traffic replacement has direct legal weight. The background of the case is the "Google Zero" phenomenon: the collapse of publishers' referral traffic, where chatbot and AI search answers replace link clicks. Notably, this is the first case where the companies themselves have documented in writing that LLMs are "destroying their own content supply chain": without clicks, publishers lose revenue and produce less new text for training future models. At the same time, this is a court document as reported by The Verge, not a scientific work: these are documented statements by specific people and companies, not peer-reviewed evidence.

Why This Matters for the Industry

The closed loop described in the documents means for the industry a shift in competitive advantage toward products based on licensed data and explicit attribution: the "free scraping" paradigm has been recognized as self-destructive by its own participants. Companies will have to demonstrate the cleaning of datasets from paywalled content and auditable provenance processes — the absence of such processes at OpenAI became a legal, not just technical, defect. Negotiations over content licensing and data prices are becoming more complex in favor of rights holders: for the first time, the argument is based on the written admissions of the companies themselves, not on critics' assessments. If your product is fine-tuned on scraped data or generates summaries of news sites, it is reasonable to audit the datasets for paywalled content and measure the share of verbatim reproduction in the output. As the case progresses, an expansion of licensing deals with publishers and increased demand for memorization suppression methods are expected: deduplication, unlearning, and retrieval instead of memorization. If fair use for training is limited, licensed corpora will become a mandatory part of training, the open web as a data source will shrink, and the share of synthetic data will grow with all the risks of quality degradation; the web may also split into a layer open to agents and a layer of closed licensed access.

Why This Matters for Users

For readers, the documents are a rare opportunity to see that top executives at OpenAI and Microsoft knew in advance about the side effects of their products, and from the primary source, not from critics' retellings. The practical meaning is direct: when a chatbot summarizes a news story instead of linking to the original source, the publisher loses a click and revenue, which is what journalism exists on. In the long term, this threatens readers with tighter paywalls, licensing walls, and the division of the web into parts open and closed to agents. It is also worth watching how the AI products used attribute sources: this determines whether readers will have a path to the original. The full filings file and The Verge's analysis are publicly available, so the conclusions can be verified independently.

What Is Still Unknown / Limitations

The June 2022 note on memorization does not prove a measured model capability: the documents contain no methodology — no memorization levels, no deduplication statistics, no eval protocol, and a documented intention is not the same as a measured result. The material is a collection of quotes from court documents as reported by The Verge: these are documented statements by specific people, but not peer-reviewed evidence or benchmarks. It is unknown how the court will ultimately rule on fair use, what materials will be unsealed next, and what the actual data cleaning pipelines at the companies look like.

Sources

Author

Look at AI, editorial team