Tech Giant Outrage: Microsoft Executive Declares AI Scraping 'Largest Labor Theft' in History
Unredacted documents in The New York Times' copyright lawsuit against OpenAI and Microsoft reveal executive admissions of 'theft' and an 'existential threat' to publishers from AI training practices. The filings detail alleged paywall circumvention, mass scraping, and stripping of copyright notices, significantly challenging the companies' fair use defense and impacting news organizations.
New unredacted information from The New York Times' three-year-old copyright lawsuit against OpenAI and Microsoft has unveiled significant admissions, revealing that top executives privately acknowledged their AI training practices as 'theft' and recognized the 'existential threat' their AI products pose to publications and journalists. The lawsuit's filings detail how the companies allegedly acquired and utilized content by circumventing paywalls, mass scraping to build training datasets, and deliberately removing copyright notices from the data.
A top Microsoft executive reportedly described the companies’ AI training methods as 'theft,' while OpenAI’s own leadership expressed concern that its AI models presented an 'existential threat' to the very publishers and journalists whose work constituted their training data. Furthermore, the unsealed material describes how these companies allegedly obtained content by bypassing paywalls undetected, created training datasets through extensive scraping, and intentionally stripped copyright notices from the acquired material. Much of this new information stems from The Times’ own brief, rather than the underlying sealed exhibits, and specific quotes are presented without their original context.
The lawsuit addresses the complex question of whether AI firms can legally use copyrighted material for training. While judges have often favored AI companies' arguments that such training falls under 'fair use'—a legal rule permitting copyrighted work use without permission for certain purposes like parody or news reporting—many of the new admissions contradict OpenAI’s fair use defense. Notably, the fair use doctrine requires that the use does not substitute or harm the market for the original work. Microsoft's internal data, for instance, indicated that its Copilot 'answer engine' caused click-through rates for The New York Times' domain to plummet by as much as 93% compared to traditional Bing search. An internal Microsoft presentation from January 2024, authored by Brent Hecht, the director of Applied Science, characterized this decline as a 'doom loop' that would 'hurt the performance of our models and the entire web at the same time.'
A Microsoft document, as quoted in the filing, candidly states, 'It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’' Microsoft CEO Satya Nadella testified in a deposition that 'anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training.' He added that had he known OpenAI had scraped and trained on paywalled information, he would have 'invoked [Microsoft’s right to] require OpenAI to retrain its models.'
Other admissions challenge additional pillars of the fair-use test. Nick Turley, OpenAI's head of ChatGPT, internally communicated that publishers face an 'existential threat' from products like the chatbot, which are 'largely substitutive' and 'will get more and more substitutive as they get better.' OpenAI President Greg Brockman described the models as 'excellent at news.' Nadella, under oath, agreed that conversing with chatbots 'has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source.' This language suggests the technology directly competes with, rather than transforms, original works. A Microsoft document also states there is a 'real risk' that generative AI could 'significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.'
The sheer scale of the copying is remarkable. The documents reveal that OpenAI’s mid-training datasets alone contain over 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting. A Common Crawl-derived dataset contained more than 2 million documents from nytimes.com alone. In a January 2023 internal memo, Hecht labeled this 'an astonishing theft of unprecedented proportions' and 'the largest theft of labor in human history.' The filing details how OpenAI and Microsoft allegedly acquired content, including scraping from the Bing Index. 'OpenAI delivered the entire GPT-3 training dataset to Microsoft, which Microsoft used to evaluate how to implement OpenAI’s models within its own commercial products,' the filing states. 'Microsoft similarly provided training data to OpenAI through initiatives called Project Taxi and Project Mango.' These projects allegedly resulted in a training dataset for Project Mango containing copies of at least 160,903 unique works from news publishers.
To maximize their scraping efficiency, OpenAI employees reportedly devised a plan to bypass paywalls undetected. Filings show that when OpenAI researcher Nick Ryder informed Brockman of a 'hack to get around nytimes paywall,' Brockman responded with 'ah nice.' OpenAI employees also allegedly constructed training datasets like WebText and WebText2, which disproportionately relied on scraped news content, and pulled millions of articles from Common Crawl. The findings also describe deliberate efforts to remove copyright notices from training data before it reached the model, as researchers 'wouldn’t want model outputting' 'copyright notices' to users. Steven Lieberman, counsel for the New York Daily News, stated, 'The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong.' OpenAI and Microsoft have not yet returned requests for comment.