Court documents unsealed in the New York Times' lawsuit against OpenAI and Microsoft reveal that both companies' internal documentation explicitly warned about the damage their AI products would cause to the web. The companies described their own data scraping practices in terms that would embarrass a defendant in any other context, and acknowledged the economic destruction they were setting in motion before carrying forward anyway.

The most striking language comes from Microsoft's Director of Applied Science, Brent Hecht, who characterized the harvesting of web data for AI training as "the largest theft of labor in human history" and said Microsoft's legal defense makes "a complete mockery of the idea of fair use." Microsoft has attempted to distance itself from those comments, calling them "one employee's individual perspective" that "are not a legal analysis and do not represent the company's views."

A separate filing from Jordan Usdan, GM for Data Strategy and Ops at Microsoft AI, went further, describing Hecht's role as deliberately adversarial. Usdan characterized Hecht as someone "employed at Microsoft to bring asymmetrical, futuristic, and academic points of view" and said he "is not someone who speaks for Microsoft specifically as to his theoretical views on AI's potential effect on content creators."

The Doom Loop They Documented Themselves

Among the unsealed materials is an internal Microsoft document that lays out the problem with unusual clarity. The document states that Microsoft's "AI content strategy has started a 'doom loop' that will hurt the performance of our models and the entire web at the same time." It continues: "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain.'"

In other words, Microsoft's own analysts identified that their AI products were consuming the output of the publishers, journalists, and creators whose work made those products valuable, and that this consumption would degrade the quality of the products themselves. The document frames this as a structural problem, not an unforeseen consequence.

Satya Nadella, in his own testimony, acknowledged that chatbots have essentially replaced search and removed the need for users to visit source websites directly. OpenAI's own media and economic experts attributed the decline in referral traffic to sites like the Times directly to AI summaries, speculating that search referrals may have dropped as much as 60 percent.

Knowing and Proceeding

The filings reveal a pattern of internal awareness coupled with external denial. Nadella testified that "anything that is paywalled should be licensed," yet an OpenAI representative admitted being "unaware" of any effort to detect or remove paywalled content from training data. The company built a product that consumes paywalled journalism while claiming no knowledge of how that journalism was obtained for training.

Internally, OpenAI recognized the memorization problem. Employees acknowledged that "prevention of memorization" was important to "minimize copyright violations," but admitted that GPT-4 "memorized a ton of data and therefore will be insanely good at regurgitation." The filing cites multiple examples of ChatGPT reproducing long passages verbatim from the Times, Mercury News, The Denver Post, LifeHacker, and Eurogamer in response to user queries.

Microsoft's internal documents acknowledged the same dynamic from the supply side. The company admitted that "almost no one intended for the content they created to be used in this fashion, nor are they compensated for its use," and described its scraping as "hoovering up" creators' work.

The Substitute Problem

The most damaging admissions concern the substitution effect. OpenAI Policy Director Jack Clark described the company as "creating systems that substitute for the labor of the people that define the 'culture' of society." Internal documents characterized ChatGPT as "the modern newsstand." Nick Turley, OpenAI's Head of ChatGPT, said that once a user gets an answer from the chatbot, there is "no good reason to click" on a source link.

Microsoft's own analysis reached a similar conclusion in blunter terms: "LLMs are a product that destroys its own supply chain." The statement acknowledges that the AI's output replaces the very training data that made it possible, creating a feedback loop where the product degrades the quality of the inputs it depends on.

OpenAI cofounder Greg Brockman's internal comments provide the motivational context. His focus was on "gazillions" of potential revenue from commercial AI, not on the economic consequences for the publishing industry that employs millions of people.

What This Means for the Lawsuit and the Industry

The unsealed documents strengthen the Times' case by establishing that both companies understood the harm their products would cause and proceeded regardless. Copyright law's fair use defense requires consideration of the effect on the market for the original work. Internal documentation admitting that AI products are substitutes for that work, that they destroy referral traffic, and that they threaten the economic foundations of content creators directly undermines that defense.

For the publishing industry, the filings confirm what many suspected: the companies building AI products were aware from the beginning that those products would cannibalize the content ecosystem. The question now is whether courts will hold them accountable for proceeding despite that knowledge, and whether the "doom loop" Microsoft identified will continue to play out as the lawsuit proceeds.

For developers and teams building AI applications that rely on web-scraped data, the filings establish a record of documented awareness. The legal and ethical landscape around training data is shifting, and the internal communications from these companies will serve as reference points in future cases involving data scraping, fair use, and the economic impact of AI on content industries.