Two major U.S. newspapers filed copyright infringement lawsuits against OpenAI and Microsoft, joining a growing roster of media organizations challenging the AI companies over unauthorized use of their journalism.
The Seattle Times and Newsday argue that OpenAI scraped their articles to train large language models without permission or compensation. The complaint centers on a core grievance: when users query ChatGPT and other OpenAI systems, the AI frequently reproduces passages from these outlets' reporting verbatim or near-verbatim, effectively serving copyrighted content without attribution or licensing fees.
This lawsuit follows a predictable pattern. The New York Times filed a similar suit in December 2023, claiming OpenAI and Microsoft systematically used millions of articles from its archives to build competitive products. Authors and publishers have pursued comparable claims. What distinguishes this wave of litigation is scale. Major newsrooms now recognize that AI training on their work threatens their business model while generating zero revenue for journalism production.
The mechanics are straightforward. OpenAI trained models like GPT-4 largely on internet-scraped data harvested before 2021. Common Crawl, a public repository, provided bulk data. Newsrooms have traced specific articles appearing in responses to queries, proving direct reproduction. The company has defended training on public data as fair use. Courts have not yet settled whether AI model training qualifies as transformative use or straightforward infringement.
For publishers, the stakes extend beyond copyright. News organizations operate on thin margins. Ad revenue has collapsed over two decades. Subscriptions remain modest. If ChatGPT delivers news summaries directly without driving readers to newspaper sites, traffic and subscription growth suffer. OpenAI's refusal to license content or negotiate revenue-sharing agreements intensifies the threat.
Microsoft compounds this problem. The company owns a substantial stake in OpenAI and has integrated ChatGPT functionality into search products and Office tools. Microsoft's distribution muscle amplifies the potential harm to news outlets. Every time a user asks Copilot for news summaries instead of visiting a newspaper website, that represents lost traffic and potential subscription revenue.
OpenAI has attempted to preempt litigation through selective partnerships. The company signed content licensing deals with outlets including News Corp and Financial Times properties, establishing a potential precedent. Yet most major publishers refused initial terms, viewing OpenAI's offers as inadequate. The Seattle Times and Newsday presumably exhausted negotiation options before litigation.
The legal theory rests on established copyright doctrine. Publishers own the copyright to their journalists' work. Reproduction, even partial, without license constitutes infringement. Fair use claims typically require transformative use. Courts have questioned whether an AI system that reproduces copyrighted passages wholesale qualifies as transformative. The question remains unsettled in AI litigation.
Regulators have begun responding. The European Union's Digital Services Act includes provisions protecting publishers. Some jurisdictions explore AI-specific copyright frameworks. U.S. lawmakers from both parties have indicated interest in protecting news organizations, though comprehensive legislation remains unlikely before 2025.
These lawsuits will take years to resolve. Discovery will likely reveal detailed information about OpenAI's training data sourcing practices. Settlements or judgments could force AI companies to negotiate licensing agreements with publishers, fundamentally altering the economics of large language model development. For now, OpenAI and Microsoft face mounting legal pressure from the institutions whose work powered their systems.
