AI's Legal Labyrinth: The Complicated Truth of Training Models on Copyrighted Books
The intersection of AI training and copyright law presents a highly complex legal challenge, as AI models use vast databases of copyrighted works. Landmark cases illustrate the struggle to apply outdated laws to new technologies, distinguishing between lawful training and illegal content sourcing. Ongoing litigation and the nuances of fair use will continue to shape the future of AI and intellectual property rights.
The burgeoning field of artificial intelligence (AI) has sparked a complex legal debate, particularly concerning the vast databases of copyrighted works used to train models like ChatGPT, Gemini, and Claude. Many published authors find their creations have contributed to these AI tools without their knowledge or consent, raising significant questions about intellectual property rights and fair compensation.
A critical examination of this issue reveals a legal landscape fraught with complexity, largely due to the outdated nature of copyright law, which has not seen a substantial update since 1976. This forces judges to interpret decades-old guidelines in the context of advanced technologies that were unimaginable at the time. Cathy Gellis, an attorney specializing in intellectual property, copyright, and technology, emphasizes the intricacy of the situation, noting, "It’s very complex and there are a lot of raw feelings about what is happening, both for and against."
One landmark ruling involved Anthropic, an AI company, ordered to pay a $1.5 billion copyright settlement to a group of writers. While seemingly a victory for authors, Judge William Alsup's decision in fact affirmed the lawfulness of Anthropic's AI training. The penalty stemmed from Anthropic's practice of pirating books from illegal online shadow libraries, rather than the act of training itself. Judge Alsup likened an LLM's ingestion of vast textual data to a writer's study of literature, stating, "Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different." Gellis interprets this ruling as largely advantageous for AI companies, underscoring that copyright law primarily concerns "copying" a work, not merely "using," "experiencing," or "consuming" it.
Central to these legal deliberations is the doctrine of fair use, a carve-out in copyright law that permits the use of copyrighted materials without explicit permission for purposes like criticism, parody, or education. For a use to be considered fair, judges assess several factors: the purpose and character of the use, the nature of the copyrighted work, the amount and substantiality of the portion used, and the effect of the use upon the potential market for or value of the copyrighted work. The crucial question often revolves around whether the use is sufficiently "transformative."
Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International, observes that courts are currently divided in their reasoning. However, a pattern is emerging: courts tend to disapprove of AI training when its purpose is to directly compete with the original copyrighted work. Conversely, if the AI's output is not intended to compete, courts are more likely to find the use permissible. Henderson cited the case of Thomson Reuters suing Ross Intelligence for copying its content to build a competing AI-based legal platform. Judge Stephanos Bibas ruled that Ross's use was not transformative because it lacked a "further purpose or different character" than Thomson Reuters's, thus deeming it not fair use.
While authors could argue that chatbots, by generating new, synthetic books, directly compete with them, this argument has not yet prevailed in court. The legal framework also distinguishes between AI training and the copyrightability of AI-generated content itself. The Thaler v. Perlmutter case, for instance, ruled that a 100% AI-generated work is not copyrightable. This decision introduces a new set of challenges, including how to definitively prove the extent of AI involvement in a creative work.
With most AI companies currently embroiled in ongoing litigation, a definitive resolution to these complex legal questions remains distant. Initial court decisions, while not necessarily final, are significantly shaping the industry. As Gellis concludes, these initial legal "volleys" are influential, though their impact could be undone by future rulings. Nevertheless, AI companies would be "foolish to ignore them" as the legal landscape continues to evolve.