Court-ordered unsealed materials from the New York Times lawsuit against Microsoft and OpenAI showed that both companies had documented the replacement effect of their products and recognized harm to publishers before the public launch of ChatGPT and Copilot. Microsoft Applied Science Director Brent Hecht called news scraping for AI training “astonishing theft of unprecedented proportions,” ChatGPT head Nick Turley acknowledged chatbots as an “existential threat” to publishers, and Microsoft’s own telemetry recorded an 83–93% drop in click-through rates to publisher websites in Copilot compared to regular Bing search. The outcome of the case will determine whether all AI companies must license news content.


What happened
By court order in the New York Times v. Microsoft and OpenAI case, previously sealed internal documents from both companies have been unsealed. In a memorandum dated January 2023, Microsoft Applied Science Director Brent Hecht called news scraping for AI training “astonishing theft of unprecedented proportions” and possibly “the largest theft of labor in human history.” ChatGPT head Nick Turley acknowledged during the same period that chatbots represent an “existential threat” to publishers because the products “largely replace” news. According to Microsoft’s own data, click-through rates to plaintiffs’ websites in Copilot fell by 83–93% compared to regular Bing search. In OpenAI datasets, plaintiffs found more than 91,692 copies of NYT, Daily News, and Center for Investigative Reporting works and over 2 million documents from nytimes.com, while OpenAI President Greg Brockman replied “Ah, nice” to an engineer’s message about a “hack” to bypass the paywall. Plaintiffs are asking the court to rule on summary judgment only for articles with extensive verbatim overlap, where the drop in click-through rates reached maximum values.
Context
The New York Times lawsuit was filed in December 2023 and became a central test of whether training and using generative AI falls under fair use protection. A key parameter of this dispute is replacement: fair use protection assumes that use should not replace the original, so measurable reader shift from publishers to AI products strikes at the foundation of the defense. The timeline in the case matters: internal harm assessments are dated January 2023 and relate to the period before the public launch of ChatGPT and Copilot, which, according to plaintiffs’ logic, indicates awareness of the actions. Plaintiffs proved replacement with prompts like “what's the next line?” and “rate the bias,” which are methodologically equivalent to known academic tests for verbatim memorization and extraction, but for the first time work as judicial evidence rather than a research protocol. The rarity of the unsealed materials also lies in the fact that they provide an external look inside OpenAI’s closed training corpora, which the company had not previously disclosed.
Why this matters for the industry
For the AI industry, the unsealed documents turn replacement of source content from an abstract legal risk into a measurable quantity: Microsoft’s own telemetry shows that an answer-based interface takes the overwhelming majority of transitions away from the source. This is a direct signal that the “summarize news without attribution” pattern is not viable, and the advantage shifts to citation-first products, licensed feeds, and data provenance infrastructure. Legally, the materials undermine the fair use defense: internal assessments from both companies document both replacement and awareness of harm before product launch, so any datasets with news scraping now carry documented awareness of harm. Plaintiffs made a “two-fisted blow” bet — replacement plus verbatim excerpts — and are asking for summary judgment only for articles with extensive verbatim overlap; if the court accepts replacement telemetry and prompt probes as evidence, internal product evals risk becoming standard judicial evidence. Companies whose products rely on news summarization should already conduct an audit of the origin of training and retrieval data, add a licensing payment scenario to their cost model, and start measuring CTR of outgoing links and the share of answers that replace the source. If publishers win, licensing news content could become an industry norm for all AI companies — exactly this scenario is recorded in the documents themselves as a way out of the “prisoner’s dilemma,” where it is beneficial for each company to use free content while others pay.
Why this matters for users
If the court sides with publishers, ChatGPT and Copilot answers on news topics may become less complete, with mandatory links to sources and licensing payments — or services will have to conclude paid deals with media. Nothing in the products breaks immediately: these are unsealed court documents, not a court decision, so today’s answers look as before. Notably, Satya Nadella’s acknowledgment of paywalls: even the main beneficiaries of AI consider licensing paywalled content mandatory, meaning some answers may already rely on licensed rather than freely collected data. Users should get used to more explicit attribution: in a scenario where publishers win, a link to the original source will cease to be decorative and become a mandatory element of a news answer.
What is still unknown / limitations
The unsealed materials do not include the methodology for measuring the 83–93% drop in click-through rates: it is not disclosed which queries and over what period were analyzed, how the denominator for comparison with Bing was calculated, and how uniform the effect was, so it is premature to interpret this figure as a ready-made industry replacement metric or product benchmark. It is also not disclosed what exactly was considered a “copy” among the 91,692 works found and how deduplication was performed when counting documents from nytimes.com. Finally, these are unsealed court documents, not facts established by the court: neither replacement nor the obligation to license content has been legally recognized yet, and the outcome of the requested summary judgment is unknown.
Sources
- The New York Times — Inside Microsoft and OpenAI, Worry About Damaging the Publishing Industry
- TechCrunch — Microsoft exec called AI scraping 'the largest theft of labor in human history,' new unredacted filings reveal
- Ars Technica — Microsoft exec called AI scraping the "largest theft of labor in human history"
- Hacker News — discussion of the material
Author
Look at AI, editorial team
