Insights / Blog

AI Copyright and Training on Books: What Courts Have Actually Decided So Far

Courts issued conflicting AI copyright rulings in 2025. The Anthropic case, Ross Intelligence, and Thaler decisions each point in different directions…
L
Lam Nguyen - Founder
Share
AI Copyright and Training on Books: What Courts Have Actually Decided So Far
ON THIS PAGE

Table of Contents

The AI copyright question around book training has no clean answer yet. Courts issued conflicting rulings in 2025: one found that training AI models on books is lawful in principle but penalized how the data was obtained; another ruled that training on copyrighted content to build a direct competitor crosses a legal line. No single decision covers all cases, and U.S. copyright law has not been updated since 1976.

$1.5 billionAnthropic copyright settlement ordered by Judge Alsup (training ruled lawful; penalty for pirating from shadow libraries)TechCrunch, August 2026
$200 billionAnthropic projected annual revenue by 2028 (context for scale of $1.5B fine)TechCrunch, August 2026
1976Last year U.S. copyright law was updated, the framework judges must apply to AI cases todayTechCrunch, August 2026

What Did the Anthropic Ruling Actually Decide?

One of the first major rulings on AI copyright came in 2025, when Judge William Alsup ordered Anthropic to pay a $1.5 billion settlement to a group of writers whose works were used to train its AI models, according to TechCrunch. The headline number obscures what the judge actually found. Alsup ruled that Anthropic’s AI training itself was lawful. The penalty arose from how Anthropic obtained the books: by downloading them from illegal online shadow libraries (unauthorized digital repositories where pirated copies of books circulate freely), not from the act of training on them.

Judge Alsup compared the process to a writer studying literature, writing that Anthropic’s large language models (LLMs, the core AI systems behind chatbots like Claude and ChatGPT) trained on text “not to race ahead and replicate or supplant” the works “but to turn a hard corner and create something different.”

Attorney Cathy Gellis, who specializes in intellectual property and technology law, told TechCrunch she views the ruling as broadly favorable for AI companies. “Copyright law hinges on copying, but it doesn’t hinge on using the work or experiencing the work, consuming the work, reading the work,” she said. For context on scale: Anthropic is projecting roughly $200 billion in annual revenue by 2028, per TechCrunch, which puts the $1.5 billion penalty in perspective.

Why “Transformative Use” Is the Pivotal AI Copyright Test

Fair use is a carve-out in U.S. copyright law that permits using copyrighted material without permission when the use is “transformative” enough to serve a genuinely different purpose. Judges weigh the nature and purpose of the use, how much of the original work was taken, and whether the use harms the market for the original.

That final factor is where courts have been most decisive. In a separate 2025 ruling, Thomson Reuters sued legal research firm Ross Intelligence for training an AI-powered platform on Reuters content to build a product that would directly compete with Reuters. Judge Stephanos Bibas ruled against Ross, writing that “Ross’s use is not transformative because it does not have a ‘further purpose or different character’ than Thomson Reuters’s,” per TechCrunch.

Jason Henderson, Senior Attorney and Founder of the IP and Media Practice at JWL International, told TechCrunch the pattern forming across cases is: “What’s tending to win is if what you’re doing is you’re training on somebody’s property because your purpose is to directly compete, then the courts will frown on it. If what you’re doing is not going to compete, then the courts are tending to find ways that it will be okay.”

Authors have argued that AI chatbots compete with them directly by generating synthetic books using their writing. That argument has not yet prevailed in court, according to TechCrunch.

Can AI-Generated Content Be Copyrighted?

The legal complexity extends to the output side. In Thaler v. Perlmutter, a court ruled that content that is 100% AI-generated cannot receive copyright protection at all, according to TechCrunch. This creates an immediate practical problem: there is no definitive way to prove what percentage of any given work was created with or without AI assistance.

Gellis illustrated the difficulty with an analogy: “If you write your novel in [Microsoft] Word and run spell check, we kind of feel comfortable with the idea of saying that Word does not own your novel.” She told TechCrunch that AI is forcing courts to confront questions legislators have long deferred.

Henderson summed up the broader gap: “They know that the AI model has been trained on so much stuff, and the law has not really caught up to that question.” With copyright law last updated in 1976, judges are applying a 50-year-old framework to questions that will shape the AI industry’s future. Most major AI companies remain lodged in pending litigation, and Gellis told TechCrunch that early rulings “could be undone if other courts decide different things.”

What This Means for AI-Search Visibility

Which sources AI systems are legally allowed to train on will determine which sources those AI systems quote and surface. Right now, courts are in the middle of deciding that question, and the outcome matters directly to any brand that wants to show up in AI-generated answers.

Think of it this way: if courts eventually narrow lawful AI training to licensed, clearly attributed content, publishers and websites with formal data agreements will gain a structural advantage inside AI-generated responses. An AI model trained on a curated, licensed corpus (a structured, legally cleared collection of text) will cite different sources than one trained on everything scraped from the open web.

The Ross Intelligence ruling makes the competitive angle concrete. Courts appear most protective of content owners when AI training directly enables a product that competes with the original content creator. A brand whose articles train an AI tool that then replaces that brand in search results sits in a legally and strategically different position than a brand contributing content to a general-purpose assistant that answers questions in other domains.

One angle the source reporting does not address: publisher behavior is already shifting in anticipation of future rulings, regardless of how those rulings land. Those shifts will change the composition of what future AI training datasets look like, which changes which sources AI models will be able to cite reliably.

In Hingewise’s assessment, the practical implication is this: AI systems trained on legally licensed, clearly authored, well-attributed content are better positioned to surface and cite that content going forward than systems trained on legally contested corpora (collections of text). Waiting for a definitive court ruling before taking any action on content attribution and licensing is effectively a choice to let others shape the training data landscape first.

What to Check Before Assuming the AI Copyright Landscape Is Settled

  • Verify your website’s robots.txt file and whether it distinguishes between standard search crawlers and AI training crawlers.
  • Check whether your content management platform or syndication partners have signed AI data-licensing agreements, and what terms those agreements include.
  • Identify whether your content appears in shadow library-type repositories. That is the specific exposure the Anthropic ruling penalized, separate from the training question itself.
  • Review whether your content licensing agreements include explicit clauses about AI training use.
  • If you publish AI-assisted content, document the human authorship involved. The Thaler ruling raises questions about copyright protection for fully AI-generated output.
  • Track which AI platforms are citing competitors in your category and investigate whether those platforms hold data-licensing agreements for that content.

The 2025 rulings signal that purpose and competitive impact matter more than the volume of content used when courts evaluate AI copyright claims. What counts as “transformative enough” will be refined through years of further litigation. The next round of appellate decisions will carry more weight than any single district court ruling so far, and that is the part of this story still worth watching closely.

Lam Nguyen, Hingewise

Sources

Keep reading

All articles →
Free · no strings

See what AI says about you, today.

Get the report showing how ChatGPT, Gemini & Perplexity answer about your brand.

Get free report →
Reply within 48 hours.