Microsoft Executive Admits AI Training Is 'Theft' as OpenAI Leaders Warn of 'Existential Threat' to Publishers in Unredacted NYT Lawsuit
By admin | Sep 17, 2026 | 4 min read
Newly unredacted documents from the copyright lawsuit filed by The New York Times against OpenAI and Microsoft three years ago contain striking admissions: AI scraping was effectively treated as theft, and AI products represent a serious danger to publishers. According to the lawsuit, a senior Microsoft executive privately characterized the companies' AI training methods as "theft," while OpenAI's own leadership acknowledged that its AI models represented an "existential threat" to the publishers and journalists whose work was used to train them.
The unsealed material further describes how the companies allegedly accessed and used that content — slipping past paywalls without being detected, assembling training datasets through mass scraping, and intentionally removing copyright notices from training data. It's important to note that a significant portion of this new information originates from The Times' own brief rather than the underlying exhibits, which are still sealed. The quotes provided below appear without their original context.
The unredacted filing marks the most recent development in the three-year-old lawsuit, in which The New York Times originally claimed the firms broke copyright law by training generative AI models on its content. Whether AI companies can legally use copyrighted material for AI training remains unresolved, though judges have generally favored AI companies' arguments that training qualifies as "fair use." This legal doctrine permits the use of copyrighted work without permission in specific circumstances, such as parody, news reporting, or criticism. Earlier this month, the Trump administration filed a brief supporting OpenAI's unlicensed use of copyrighted material to train its LLMs.
However, several of the new admissions work against OpenAI's fair use defense — especially the doctrine's requirement that the use must not substitute for or damage the market of the original work. For instance, Microsoft's own data reveals that its Copilot "answer engine" drove click-through rates for The New York Times' domain down by as much as 93% compared to traditional Bing search. An internal Microsoft presentation authored by Microsoft's Director of Applied Science, Brent Hecht, in January 2024 characterizes this decline as a "doom loop" that would "hurt the performance of our models and the entire web at the same time."
"It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain,'" states the Microsoft document, as quoted in the filing. Microsoft CEO Satya Nadella also testified in a deposition earlier this year that "anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training," and made clear that if he "had been made aware that OpenAI had scraped and trained on information that was behind a paywall," he would have "invoked [Microsoft's right to] require OpenAI to retrain its models."
Additional admissions undermine other pillars of the fair-use test. OpenAI's Head of ChatGPT, Nick Turley, wrote in internal communications that publishers confront an "existential threat" from products like the chatbot, which are "largely substitutive" and "will get more and more substitutive as they get better." OpenAI President Greg Brockman described the models as "excellent at news." Nadella concurred under oath earlier this year that conversing with chatbots "has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source."
Such language speaks to how the technology could directly compete with — rather than transform — the original work. A Microsoft document notes a "real risk" that generative AI could "significantly disrupt the employment of the very people who generated the data on which the foundation model was trained."
The sheer magnitude of the copying is remarkable. The documents reveal for the first time that OpenAI's mid-training datasets alone contain more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting. A Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone. In a January 2023 internal memo, Hecht called it "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history."
The filing lays out in fresh detail how OpenAI and Microsoft went about obtaining the plaintiffs' content, including scraping it from the Bing Index. "OpenAI delivered the entire GPT-3 training dataset to Microsoft, which Microsoft used to evaluate how to implement OpenAI's models within its own commercial products," the filing reads. "Microsoft similarly provided training data to OpenAI through initiatives called Project Taxi and Project Mango."
The companies allegedly compiled the Project Mango data into a training dataset containing copies of at least 160,903 unique works from the news publishers. To maximize the value of their scraping, OpenAI employees allegedly devised a plan to circumvent paywalls undetected. The filings show that when OpenAI researcher Nick Ryder informed Brockman about a "hack to get around nytimes paywall," Brockman responded: "ah nice."
OpenAI employees also allegedly constructed training datasets such as WebText and WebText2 that leaned disproportionately on scraped news content. They additionally pulled millions of articles from Common Crawl, a free, open repository of web crawl data. The findings also describe deliberate efforts to strip copyright notices from training data before it reached the model, since researchers "wouldn't want model outputting" "copyright notices" to users. OpenAI and Microsoft did not return requests for comment.
Comments
Please log in to leave a comment.
No comments yet. Be the first to comment!