AI companies relied on enormous collections of published writing to train their large language models, and executives at those firms knew that doing so amounted to taking copyrighted material, according to remarks quoted by news publishers in a court filing that was unredacted this week.
Brent Hecht, Microsoft’s director of applied science, described the practice as “an astonishing theft of unprecedented proportions” and possibly the “largest theft of labor in human history,” according to the publishers’ filing. The complete exhibits that provide context for the quotations remain sealed.
The remarks cited by the publishers appear to challenge the position taken by AI companies that using copyrighted material was legally protected as fair use. Under that doctrine, whether a use of copyrighted material is protected depends on factors such as how it is used, the nature of the original work, the amount taken and the impact on the market for that work.
Microsoft CEO Satya Nadella testified under oath that conversations with chatbots provided information “right there on the website on the AI platform” instead of requiring users to visit the original source, such as a publisher’s website that had reported the information, according to the filing. Likewise, an OpenAI executive wrote that publishers faced an “existential threat” from products such as the company’s chatbot.
The publishers’ brief also outlines measures taken by AI developers to bypass paywalls, including The New York Times’ paywall. It says an OpenAI employee told company president Greg Brockman about “a hack to get around nytimes paywall,” to which Brockman replied, “ah nice.”

The filing also states that Nadella testified anything behind a paywall “should be licensed by anyone who wants to use it” for AI development and that, had he known OpenAI had trained on paywalled content, he would have required the company to retrain its models.
Representatives for OpenAI and Ziff Davis did not immediately respond to requests for comment.
Publishers also argue that OpenAI and Microsoft employees knew their chatbots’ ability to locate, copy, summarize and sometimes repeat publisher content was keeping readers from visiting news organizations’ websites.
“Defendants’ own experts acknowledge that grounded LLMs exploit their sources rather than promote them like search uses that have been deemed fair use,” the publishers’ filing states.

