Sony Music Group and Warner Music Group filed a federal lawsuit against Anthropic on Tuesday, alleging the AI company systematically scraped copyrighted music metadata and lyrics from publicly available sources to train Claude, its AI assistant. The lawsuit centers on what the music labels describe as a "brazen campaign" of intellectual property theft that violates the Digital Millennium Copyright Act and common copyright law.

The complaint, filed in federal court, accuses Anthropic of circumventing technical protections designed to prevent unauthorized collection of copyrighted material. Lawyers for Sony and Warner argue that Anthropic deliberately harvested song titles, artist names, album information, and lyrical content without licensing agreements or permission from the rights holders. The labels claim this data training pipeline directly enabled Claude to reproduce copyrighted works verbatim or near-verbatim when prompted by users.

This case arrives as the music industry escalates legal action against generative AI companies. Record labels have long contended that AI developers must license content before using it for training, establishing a market precedent similar to how search engines and news aggregators historically paid for content rights. Anthropic's Claude, however, generates text, poetry, and song lyrics based on patterns learned during training, creating a direct conflict with music copyright holders who view this capability as derivative of their protected works.

The lawsuit targets both Anthropic's training methodology and the technical architecture of Claude itself. Sony and Warner allege that Anthropic used web scrapers and data collection tools to extract content from multiple sources, including lyrics databases, music metadata repositories, and streaming platforms. The labels contend that even though this material appeared publicly online, copyright protections still applied, and Anthropic's collection violated those restrictions.

Anthropic has not yet responded to the lawsuit. The company previously faced similar copyright challenges from The New York Times and other media organizations in late 2023, which resulted in disputes over whether AI training on publicly available content constitutes fair use under copyright law. That case remains unresolved.

The music industry's aggressive stance reflects a broader shift in how rights holders approach AI development. Unlike some tech companies that argue training on public data qualifies as fair use, music labels demand explicit licensing before their content trains AI systems. This position differs from academic and research arguments that view data ingestion as analogous to human learning.

Sony and Warner seek damages, injunctive relief to prevent further scraping, and destruction of any datasets derived from their copyrighted works. The outcome will likely establish precedent for whether AI companies must license music metadata and lyrics during model training, a distinction that could reshape how generative AI developers source training data going forward. Courts have not yet issued definitive rulings on this specific question, making this case a pivotal test of copyright enforcement in the AI era.