The Seattle Times and Newsday have filed suit against OpenAI and Microsoft, alleging that the companies used their journalism as training data for AI models without permission and regularly reproduce passages from their reporting in response to user queries. The lawsuits join a growing wave of legal actions from news organizations against the companies behind frontier AI systems.
What the publishers are alleging
The core claim is straightforward. OpenAI used copyrighted journalism from the Seattle Times and Newsday to train its language models. The models, when prompted by users, reproduce passages from that journalism. The publishers argue this constitutes copyright infringement, and they are seeking the destruction of any copies of their works the companies are holding, along with training datasets and AI models that incorporate them.
The request for destruction of trained models is significant. It goes beyond monetary damages and asks the court to order the companies to rebuild their systems from scratch, excluding the copyrighted material. If granted, such an order would force OpenAI and Microsoft to retrain their models without the news content that was included without permission, a process that would be costly and time-consuming.
Microsoft is named as a defendant because Copilot, its AI coding and productivity tool, is built on OpenAI's technology. The publishers argue that Microsoft benefits from the infringement by using the same trained models in products that reach millions of users.
Part of a larger pattern
The Seattle Times and Newsday are not the first publishers to take this action. The New York Times filed a similar suit against OpenAI and Microsoft in late 2023, alleging that ChatGPT reproduces passages from its reporting. Ziff Davis, Merriam-Webster, and Encyclopedia Britannica have all filed their own suits with comparable claims.
More recently, nearly 400 local newspapers joined together to sue the two companies. Their argument adds an economic dimension to the copyright claim: chatbots reduce the need for users to visit news sites directly for reporting and answers, which costs the publishers subscription revenue. When a user asks ChatGPT about a local event and gets a summary drawn from a newspaper's reporting, the newspaper gets nothing while the user never visits the site.
The pattern is clear. News organizations across the industry, from national outlets to local papers, are converging on the same legal theory: their copyrighted content was used without permission to train AI models, and those models now compete with them for the audience and revenue that journalism depends on.
What destruction of models would mean
The publishers are asking for something unusual in copyright litigation: the destruction of trained AI models. Monetary damages are standard in infringement cases. Destruction of the infringing work is reserved for cases where the infringement is ongoing and the harm cannot be remedied by payment alone.
In the context of AI training, this argument has merit. If a model was trained on copyrighted journalism and can reproduce passages from that journalism, the model itself is the infringing work. Damages compensate for past harm, but they do not prevent the model from continuing to reproduce copyrighted material in future interactions. Destruction, or retraining without the copyrighted data, addresses the ongoing nature of the infringement.
Whether courts will agree is another question. The fair use doctrine has been invoked by AI companies to justify training on copyrighted material, and several courts have yet to rule on whether training constitutes fair use. The outcome of these cases will determine not just the financial liability of AI companies but the legal framework that governs how AI models are trained going forward.
What this means for developers
For developers building on language models, these lawsuits create uncertainty about the training data behind the models they use. If courts order the destruction or retraining of models that incorporate copyrighted journalism, the models currently in use may change. Features that depend on the model reproducing or summarizing news content may break. The legal risk is not theoretical. It is active, with dozens of cases pending in courts across the country.
The practical response for developers is to be aware that the models they depend on are the subject of ongoing litigation, and that the outcome could affect model availability, capability, and cost. Building systems that assume current models will remain unchanged is a risk. The legal landscape is shifting, and the models may shift with it.
Neither OpenAI nor Microsoft responded to requests for comment on the Seattle Times and Newsday lawsuits. The cases are in their early stages, and the companies have not yet filed their defenses. The outcome will shape the relationship between AI companies and the news organizations whose content trained their models.