# OpenAI and Microsoft's Fair Use Defense Crumbles Under Internal Scrutiny

OpenAI and Microsoft face a severe credibility crisis in ongoing litigation over AI training practices. Internal communications obtained through discovery reveal company executives openly describing their data practices as theft, directly undermining their public fair use arguments.

A Microsoft director characterized the company's approach to gathering training data as the "largest theft of labor in human history," according to sworn testimony. Meanwhile, OpenAI's head of ChatGPT acknowledged in writing that their products "are largely substitutive, period." These statements directly contradict the companies' legal position that using copyrighted material without permission constitutes fair use under copyright law.

Fair use doctrine allows limited use of copyrighted material without permission under specific conditions. Courts examine four factors: the purpose of use, the nature of the copyrighted work, the amount used, and the effect on the original work's market value. Tech companies have argued their use of published content for training qualifies as transformative fair use. That argument faces severe damage when executives explicitly state their products substitute for the original works.

The substitutive nature claim matters legally because courts weigh whether the new product harms the market for the original. When OpenAI's own leadership acknowledges their tools directly replace human labor and creative work, they eliminate one of the cornerstone arguments for fair use protection. A court can now point to internal admissions that demonstrate market harm.

The Microsoft characterization as "astonishing theft" adds another layer of legal jeopardy. If discovery produces emails where company leaders knew their practices constituted theft, litigation teams can pursue claims of willful infringement. Willful infringement allows courts to award treble damages, multiplying financial penalties by three, rather than standard damages.

These revelations emerge from multiple lawsuits. Authors including Michael Chabon and John Grisham filed suit against OpenAI in September 2023. The New York Times sued both OpenAI and Microsoft in January 2024, claiming billions in damages. Getty Images also filed suit over unauthorized use of its photographs. As discovery proceeds, more internal communications likely surface.

The stakes extend beyond OpenAI and Microsoft. If courts rule against fair use arguments based on these internal admissions, the entire AI industry's training methodology faces legal vulnerability. Anthropic, Google, Meta, and other companies building large language models used similar approaches. They assumed fair use provided legal cover. These trials will clarify whether that assumption holds.

Company executives may have been candid in internal contexts assuming confidentiality. That assumption backfired. Discovery rules require companies to produce relevant internal communications. Once produced, opposing counsel uses them strategically in court filings and public statements, exactly as happened here.

The legal question now shifts from abstract fair use doctrine to factual findings about what these companies knew and intended. When leaders call their own practices theft, courts gain documentary evidence of knowledge. Intentionality becomes harder to deny. Damages calculations become larger.

These cases will likely reach appeals, possibly the Supreme Court. The court system has not yet issued comprehensive guidance on AI training and fair use. These internal admissions may prove decisive in establishing legal precedent about what constitutes acceptable use of copyrighted material for model training. The technology industry waited years for clarity on this question. These trials may finally provide it, though the answer will likely prove costly for the companies involved.