Courts split on AI training and book copyrights

Recent rulings in cases involving Anthropic and Ross Intelligence show judges still disagree on when AI training counts as fair use.
AI companies built chatbots like ChatGPT, Gemini, and Claude on vast collections of books, articles, and papers, and that has pushed a long-running copyright fight into sharper focus. The core question is whether training a model on published works is more like reading them or illegally copying them. As IT-PUB News notes, recent court rulings suggest the answer still depends heavily on how judges apply older copyright law to a very new technology.
The stakes go well beyond the courtroom. Authors want to know whether their books were used in training data without their knowledge, while AI companies need clearer rules for building products that depend on large-scale text collections. More broadly, the dispute cuts to a basic question for the digital economy: how much of the internet and published culture can be used to train AI systems, and under what terms?
The Anthropic ruling focused on pirated books, not training alone
One of the clearest recent decisions came in a case involving Anthropic. Judge William Alsup ordered the company to pay a $1.5 billion copyright settlement to a group of writers whose works were used in training its AI models.
At first glance, that looked like a major victory for authors. The ruling, though, was more specific than that. Alsup did not say AI training itself was unlawful. He penalized Anthropic for pirating books from illegal online shadow libraries.
In the judge’s view, the way an LLM processes text could be compared to how a person studies books in order to write something new. He wrote that Anthropic’s models were not trained to copy or replace the original works, but to create something different.
That distinction now sits at the center of many AI copyright fights. The same training process can be presented either as transformative use of text or as unauthorized copying, depending on the facts and on how a court reads the law.
Courts are applying a 1976 law to modern AI systems
Part of the uncertainty comes from the age of the law. Copyright rules in the United States have not been updated since 1976, leaving judges to apply older standards to technologies that were never explicitly addressed.
Cathy Gellis, an attorney with expertise in intellectual property, copyright, and technology, said the area is unusually complicated. She described it as a field where a great deal is happening at once, with strong views on both sides.
Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International, said many people are concerned because the law has not kept pace with the scale of AI training. In his view, the core problem is simple enough: models are trained on enormous amounts of material, while the legal framework has not fully caught up.
That gap has left courts to decide how familiar copyright principles should apply to systems that learn from massive datasets.
Fair use is emerging as the main legal test
Many of these cases turn on fair use, the copyright doctrine that allows some use of protected works without permission. It can cover criticism, parody, and education, among other things, but judges weigh several factors before deciding whether a particular use qualifies.
One of the biggest questions is whether the use is “transformative” enough. Put simply, that means whether the new use has a different purpose or character from the original work.
Henderson said courts appear more likely to side with AI companies when the training does not directly compete with the original creator’s market. If the purpose is to build a product that competes with the source material, judges tend to view it less favorably.
He pointed to a case involving Thomson Reuters and Ross Intelligence. In that dispute, Thomson Reuters sued the research firm for copying its content to build a competing AI-based legal platform. Judge Stephanos Bibas ruled that Ross’s use was not transformative because it did not serve a different purpose from Thomson Reuters’s work.
That ruling matters because it shows where at least one court drew the line. Training on copyrighted material for a product that directly competes with the source may be treated very differently from using text in a way that does not.
So far, authors have not succeeded in making the broader argument that chatbots themselves compete with them simply by generating synthetic books or other text. But that remains part of the wider legal fight.
AI-generated content brings a different copyright dispute
The copyright battle is not only about training data. It also extends to what AI systems produce.
Gellis said it is important to separate the question of training from the question of whether AI-generated content can itself be copyrighted. In the case Thaler v. Perlmutter, the court ruled that a work created entirely by AI is not copyrightable.
That opens another difficult issue: how can anyone prove whether a work was generated with AI, and if so, how much human input was involved? Gellis compared the situation to using Microsoft Word spell check. Most people are comfortable saying that software assistance does not make the software the owner of a novel. AI tools, though, are pushing courts and creators to think more carefully about where assistance ends and authorship begins.
The distinction matters for writers, publishers, and businesses that want to know whether AI-assisted content can be protected in the same way as fully human-created work.
The legal picture remains unsettled
Even with several important rulings already on the books, there is still no final answer. Most AI companies remain involved in pending litigation, and different courts may still reach different conclusions.
Gellis warned that early decisions are already influencing the field, but those outcomes could still be changed by later rulings. The first wave of cases is shaping the debate. It is not settling it.
That uncertainty helps explain why the issue is drawing so much attention. Authors want to know whether their work can be used to train AI systems without permission. AI companies want clearer rules for building products at scale. Judges, meanwhile, are being asked to apply decades-old copyright law to a technology that was never part of the original framework.
For now, the legal line between training and copying is still blurred. Courts have started to sketch it out, but the boundaries remain far from settled.