It’s widely acknowledged that AI systems like ChatGPT, Gemini, and Claude are trained using vast collections of publicly available materials, including books, articles, and academic papers found online. Many authors have unknowingly contributed to these AI tools, which could potentially impact their careers. But is this practice illegal?
The situation is more nuanced than it appears.
Cathy Gellis, an attorney specializing in intellectual property, copyright, and technology, explained to JS, “One of the challenges with this field of law and technology is its complexity and the mixed emotions surrounding it.”
In a landmark ruling, Judge William Alsup required Anthropic to pay $1.5 billion in a copyright settlement to a group of writers whose works were used to train the company’s AI models. Although this seemed like a win for authors, Judge Alsup deemed Anthropic’s AI training lawful. The penalty was for sourcing these works from illegal online shadow libraries.
The judge remarked, “Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different,” likening the process to a writer studying literature.
Gellis believes this ruling benefits AI companies. A $1.5 billion fine is insignificant for a company aiming for $200 billion in annual revenue by 2028.
“The decision is generally positive for AI training as it views the process as analogous to reading a copyrighted work rather than copying it,” Gellis noted. “Copyright law revolves around copying, not merely using or experiencing the work.”
Copyright law hasn’t seen updates since 1976, requiring judges to interpret decades-old guidelines on modern AI-related legal questions.
Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International, told JS, “There’s a lot of concern because the law hasn’t kept pace with the volume of material AI models consume.”
These concerns often center around fair use law, particularly whether the use of copyrighted material is transformative enough to be legally permissible.
Fair use allows for the use of copyrighted materials without explicit permission, supporting commentary, parody, education, and more. Judges assess factors like the work’s purpose, nature, amount used, and market impact to determine fair use.
Henderson noted, “Copyright aims to protect and expand the market. Courts have varied reasoning in AI cases. If training on someone’s work is meant to compete directly, courts disapprove. If it’s not, they find ways to allow it.”
He referenced a case where Thomson Reuters sued Ross Intelligence for using its content to create a rival AI-based legal platform.
Judge Stephanos Bibas wrote, “Ross’s use is not transformative because it does not have a ‘further purpose or different character’ than Thomson Reuters’s.”
In that case, the court ruled against fair use when training aimed to create a competing platform. While authors might argue that chatbots compete by using their works to generate new content, this claim hasn’t succeeded in court.
Gellis emphasized the distinction between copyright in AI training and AI-generated content.
In a case, Thaler v. Perlmutter, the court ruled that a 100% AI-generated work isn’t copyrightable, raising questions about identifying AI-generated content and its extent.
“If you write your novel in [Microsoft] Word and run spell check, we kind of feel comfortable with the idea of saying that Word does not own your novel,” Gellis explained. “AI challenges us to reconsider decisions we previously overlooked.”
Most AI companies are currently embroiled in ongoing litigation over these issues, delaying any immediate resolution.
“Initial legal decisions are influential, but could be overturned if courts rule differently later. These rulings are shaping the industry, and AI companies would be wise to heed them,” Gellis stated.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

