\n\n\n\n Fair Use Won, But Your Bookshelf Receipt Still Matters - AgntBox Fair Use Won, But Your Bookshelf Receipt Still Matters - AgntBox \n

Fair Use Won, But Your Bookshelf Receipt Still Matters

📖 5 min read•802 words•Updated Aug 24, 2026

Do you actually know what your AI writing tool was trained on? Not the marketing page version. The real answer. Because I’ve spent a lot of time testing these tools for this site, and I can tell you that almost none of them will give you a straight answer, and now the courts have made that silence a lot more interesting.

Here’s where things stand. A United States District Court ruled that training large language models on copyrighted books counts as fair use. Judge Alsup, in the case authors brought against Anthropic, granted summary judgment on fair use for the training itself. A separate decision from Judge Chhabria in Kadrey v. Meta Platforms landed in similar territory. Both came out of the Northern District of California, and a later federal decision described AI training as “quintessentially” transformative fair use.

If you stopped reading there, you’d think the question was settled. It isn’t, and the gap between those two ideas is the whole story.

The split that actually matters

The rulings drew a line that most coverage flattened out. Training on copyrighted books that were legally acquired? Fair use. Training on pirated copies? That’s a different question, and in the Anthropic case, the claims about pirated copies were sent toward trial rather than dismissed.

So the legal victory here isn’t “AI companies can train on anything.” It’s closer to “AI companies can train on books they actually got their hands on legitimately.” The transformation argument covers what you do with the text. It doesn’t cover how the text arrived on your server.

That’s a meaningful distinction for anyone evaluating tools, because it turns a copyright question into a supply chain question. And supply chain questions have paper trails.

Why I’m bringing this up in a tool review context

I review AI toolkits. Writing assistants, code helpers, research tools, the whole shelf. And the honest assessment is that data provenance has never been part of anyone’s spec sheet. You get context window size, you get pricing tiers, you get benchmark scores. You almost never get a clear statement about training data sourcing.

These rulings change the calculus on that silence. Before, “we don’t discuss our training data” was standard industry posture. Now there’s a legal framework where one category of sourcing is defensible and another is heading to trial. A vendor who won’t distinguish between the two is telling you something, even if they aren’t saying it.

What I’d want to see on a product page, and what I’m now going to start asking about:

  • Whether training data was licensed, purchased, or scraped
  • Whether the vendor can describe its acquisition process at all, even in general terms
  • Whether there’s any indemnification for users if a training data claim goes against the vendor
  • Whether the vendor’s public statements have changed since these decisions came down

None of those are unreasonable asks. Some vendors already answer a few of them. Most answer none.

The risk you’re actually carrying

Let me be careful about scope. I’m a tool reviewer, not a lawyer, and nothing here is legal advice. But you don’t need a law degree to notice that the unresolved piece of these cases sits with the model builders, not the people typing prompts into them.

That’s a real distinction, and it should lower your anxiety somewhat. The exposure in these cases is about how models were built. If you’re using a commercial AI tool to draft emails or refactor functions, your legal position isn’t the one being litigated.

What you are carrying is dependency risk. If you build a workflow around a tool, integrate it into your stack, train your team on it, and that tool’s underlying model is tangled in unresolved claims about pirated training material, an adverse outcome could mean pricing changes, feature changes, or a model getting pulled and replaced with something that behaves differently. That’s not a copyright problem for you. It’s a continuity problem, and continuity problems are exactly what tool selection is supposed to manage.

How I’d handle it right now

Don’t panic-migrate. The training-is-fair-use side of these decisions is genuinely favorable for the companies building these tools, and treating every AI product as legally radioactive would be an overreaction to rulings that mostly went the vendors’ way.

Do add provenance to your evaluation checklist. Ask the question in sales calls. Note who answers clearly and who deflects. Keep a fallback option identified for any AI tool that’s load-bearing in your workflow, which is decent practice regardless of what courts decide.

And treat vendor transparency as a quality signal in its own right. In my experience testing these products, the teams that are straight with you about how the thing works tend to be straight with you about everything else too. That correlation has held up better than most benchmarks.

🕒 Published:

🧰
Written by Jake Chen

Software reviewer and AI tool expert. Independently tests and benchmarks AI products. No sponsored reviews — ever.

Learn more →
Browse Topics: AI & Automation | Comparisons | Dev Tools | Infrastructure | Security & Monitoring
Scroll to Top