A federal court called training AI on copyrighted books “quintessentially” transformative fair use. Anthropic still ended up on the hook for a $1.5 billion copyright settlement. Both of those things are true, and if you can hold them in your head at the same time, you already understand more about AI copyright law than most people selling you AI tools.
I review toolkits for a living. That means I spend a lot of time reading marketing pages that say things like “ethically trained” or “licensed data” with zero substantiation behind them. The recent court decisions out of the Northern District of California gave me a much sharper lens for that kind of claim, because they drew a line that most vendors are still pretending doesn’t exist.
The line the courts actually drew
Two cases matter here: Kadrey v. Meta Platforms, Inc. and Bartz v. Anthropic PBC. Both came out of the Northern District of California, and both point the same direction. Training a model on copyrighted books can be fair use, because the training itself is transformative. You are not producing a substitute for the book. You are producing something structurally different.
That is the part AI companies quote in press releases. The part they quote less often is what happened to Anthropic’s pirated library. The court denied summary judgment on Anthropic’s use of pirated copies to assemble a central library of copyrighted works. Acquisition and use got treated as two separate questions. The use looked legal. The acquisition did not.
Judge William Alsup ordered Anthropic to pay a $1.5 billion settlement to a group of writers whose works were used in training. One of the first rulings of its kind, and the number is not a rounding error for anyone.
Why this changes how I read a tool’s data page
Before these rulings, “we trained on public data” was a sentence I mostly skipped past. Now I read it as an unanswered question. Public where? Downloaded from where? A model can be trained in a way courts consider transformative and still sit on top of a library that was assembled illegally. The transformation defense does not launder the download.
So when I evaluate an AI writing tool, a code assistant, or one of the dozens of research summarizers that ship every month, the questions I want answered look different than they did a year ago:
- Where did the training corpus come from, specifically, and can the vendor name sources?
- Did the company purchase or license the material, or did it scrape whatever a torrent index served up?
- Is there any indemnification for users if the vendor’s data sourcing gets challenged?
- Has the vendor updated its public statements since the 2026 fair use developments, or is the page frozen in 2023?
Most tools fail that fourth question immediately, which tells you something about how closely they are tracking their own legal exposure.
What this means if you build on these tools
I want to be careful not to overstate the risk to end users. The litigation so far has targeted the companies doing the training, not the freelancer using a chatbot to draft blog posts. But indemnification clauses exist for a reason, and vendors that quietly lack them are making a bet with your money as the stake.
The practical read: if your business depends on a single AI provider, the sourcing of that provider’s training data is now a supply chain question, not a philosophy question. A $1.5 billion settlement is survivable for a company with Anthropic’s funding. It is not survivable for a smaller vendor with a similar data history, and small vendors with sketchy data histories are exactly the ones that disappear without warning.
The honest summary
Training on copyrighted books can be legal. Two conditions have to hold up: the works were acquired legally, and the use is transformative. Miss the first one and the second one will not save you. Pirating books is still illegal, and courts are treating that as its own separate violation with its own separate price tag.
This is also a moving target. A publication dated March 2026 describing a federal court finding on transformative fair use is not the last word, and anyone telling you the question is settled is selling something. Check current rulings before you make a decision that depends on the answer.
What I would push back on is the framing that this is all too complicated to have an opinion about. It is not that complicated. Buy the books or license the data. The companies that did that are in a defensible position. The ones that grabbed a pirate library and hoped the transformation argument would cover both halves of the problem are finding out that it does not.
đź•’ Published: