A federal judge ruled that training artificial intelligence models on copyrighted books is legal, even as he ordered AI company Anthropic to pay $1.5 billion to writers whose works were used without permission, according to an August 23 analysis by TechCrunch. Judge William Alsup's landmark decision last year distinguished between the act of AI training itself—which he found lawful—and Anthropic's method of obtaining the books from illegal online shadow libraries, which constituted piracy. The ruling highlights how copyright law written in 1976 struggles to address questions that could shape the future of the AI industry.

The AI models that power ChatGPT, Gemini, Claude, and similar chatbots train on databases containing hundreds of millions of books, online articles, academic papers, and essentially anything available on the internet, TechCrunch reports. Most published authors have unknowingly contributed to developing the same AI tools that may threaten their livelihoods. In a separate case, Thomson Reuters successfully sued research firm Ross Intelligence after the court determined Ross copied Reuters' content to build a competing AI-based legal platform. Anthropic projects roughly $200 billion in annual revenue by 2028, making the $1.5 billion penalty a relatively small fraction of future earnings.

Judge Alsup compared how large language models ingest trillions of words to a writer studying literature, writing that Anthropic's models "trained upon works not to race ahead and replicate or supplant them—but to turn a hard corner and create something different." Attorney Cathy Gellis, who specializes in intellectual property and technology law, told TechCrunch the ruling generally favors AI companies because it treated AI training as analogous to reading rather than copying. "Copyright law hinges on copying, but it doesn't hinge on using the work or experiencing the work, consuming the work, reading the work," Gellis explained. Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International, noted that judges are "all over the place" in their reasoning because "the law has not really caught up to that question."

These legal disputes frequently turn on fair use law—specifically, whether use of copyrighted material is "transformative" enough to be legally permissible, according to the TechCrunch analysis. Fair use carves out exceptions to copyright law for criticism, parody, education, and other purposes, with judges weighing factors including the work's purpose, the amount used, and market impact. Henderson observed that courts tend to disapprove when companies train on copyrighted material to build direct competitors, as in the Thomson Reuters case where Judge Stephanos Bibas ruled Ross Intelligence's use "not transformative" because it lacked a "further purpose or different character." However, while authors could argue chatbots compete by using their works to generate synthetic books, that argument hasn't prevailed in court. Separately, courts have ruled that fully AI-generated works can't be copyrighted, raising thorny questions about how to prove whether content was created with AI assistance and in what proportion.

Most AI companies remain locked in ongoing litigation over these issues, meaning definitive answers won't arrive soon. Gellis noted that early rulings are influential but "could be undone if other courts decide different things," requiring later stages of litigation to determine which precedents prevail. The current legal uncertainty forces AI companies to pay attention to emerging decisions even as the broader framework remains unsettled. For executives navigating partnerships or content licensing, the distinction between lawful training and unlawful sourcing creates a narrow but critical compliance path. Whether that path widens or closes entirely depends on judges interpreting half-century-old statutes through a lens those lawmakers could never have imagined.