The legal storm surrounding AI training data just got bigger. The Seattle Times and Newsday have filed new lawsuits against OpenAI and Microsoft, accusing the companies of using their copyrighted journalism to train AI models without authorization. Their complaints echo a growing chorus from publishers who argue that AI companies are building billion‑dollar products on the back of newsrooms that were never asked — and never compensated.
Image Courtesy : blogs.microsoft.com
The lawsuits claim that OpenAI’s models can reproduce, summarize, or paraphrase their articles in ways that directly rely on copyrighted material. Both organizations argue that this goes beyond fair use, especially because the AI outputs can compete with their own subscription‑based reporting. They also point to examples where AI systems allegedly generated text that closely mirrors their original articles, suggesting the models were trained on full copies of their content.
For The Seattle Times and Newsday, the issue is existential. Local and regional newsrooms have already been strained by declining ad revenue, shifting reader habits, and competition from national outlets. Now they face AI systems capable of producing news‑like content without paying for the journalism that informs it. Their lawsuits argue that this dynamic threatens the economic foundation of local reporting — and that AI companies should be held accountable for using their work without permission.
These filings join a growing list of legal actions from publishers including The New York Times, The Intercept, Raw Story, and others. Together, they represent a broader push to force AI companies to negotiate licensing deals rather than scraping content freely. Some publishers have already struck agreements — Axel Springer and The Financial Times among them — but many argue that the industry needs standardized protections, not one‑off deals.
OpenAI and Microsoft maintain that their training practices fall under fair use and that their models do not store or reproduce copyrighted articles verbatim. They also argue that AI systems transform information rather than replicate it. But courts have not yet delivered a definitive ruling on how copyright law applies to large‑scale AI training, leaving the industry in a legal gray zone.
The new lawsuits underscore how quickly the conflict is escalating. What began as a dispute between a handful of publishers and AI labs has grown into a nationwide battle over the future of journalism, intellectual property, and the economics of information. As more news organizations join the fight, the pressure increases for courts — or lawmakers — to define clear rules for how AI systems can use copyrighted content.
The stakes are enormous. For newsrooms, the outcome could determine whether AI becomes a partner that pays for content or a competitor that drains their audience. For AI companies, it could reshape how models are trained, what data they can access, and how much they must pay to use it. And for readers, it will influence the future of trustworthy reporting in an era where AI‑generated text is everywhere.
