When a settlement lands over books used to train an AI model, who gets the money — the person who wrote the book, or the company that printed it? If your instinct says “the author, obviously,” you’re about to learn how publishing contracts actually work.
Authors are disputing claims made by publishers and agents on payments from the Anthropic settlement. Some writers say claims are being filed against older works in ways they consider unfair. Mystery and thriller author April Henry posted publicly asking, “WTF is Harpe” — a fragment that says plenty about the tone of the conversation happening right now on author social media. The broader complaint: publishers appear to be claiming more than their fair share of certain payments.
I review AI tools for a living. I’m the guy who tells you whether a coding assistant actually saves time or just generates confident garbage. So why am I writing about book contracts? Because this dispute is the clearest look yet at something every AI tool user should understand — the gap between who produces the input and who collects when that input turns into money.
Training data has a paper trail, and it’s messy
Every AI product I test was built on somebody’s work. Text, code, images, audio. For years that was an abstraction — a line in a model card, a vague reference to “publicly available data.” A settlement changes the abstraction into a wire transfer, and wire transfers require someone to decide where the money goes.
That’s where it falls apart. The rights to a book published in 2004 are governed by a contract signed in 2003, drafted by lawyers who had no concept of a language model. Nobody wrote a clause about machine training because machine training wasn’t a thing you could write a clause about. So now both sides read the same document and reach opposite conclusions, and the party with the in-house legal department tends to read it more confidently.
Agents complicate it further. Standard agency agreements take a percentage of income derived from a work. Is a legal settlement over unauthorized use “income derived from the work”? Reasonable people disagree. Authors who thought their agency relationship ended years ago are finding out otherwise.
The scam economy showed up right on schedule
There’s a second story here that deserves its own attention. Fraudulent emails impersonating the United States Copyright Office are circulating, asking authors to verify copyright registrations. This is not a coincidence. Scammers follow confusion, and a settlement process that thousands of writers don’t fully understand is ideal territory.
The pattern is familiar to anyone who’s watched crypto airdrops or data breach settlements. A legitimate process creates a legitimate need for people to submit information. Scammers replicate the request with a different destination. Direct solicitation by email has been a primary victim-recruiting method for years, and it works because the fake message arrives in the same week as the real news.
Practical guidance, since I’d rather be useful than dramatic:
- The Copyright Office does not email you asking to verify your registration. Treat any such message as fraudulent until proven otherwise.
- Never click a link in an unsolicited email about a settlement. Navigate to the official site yourself.
- If a message creates urgency, that urgency is the tell.
- Pull out your original publishing contract before you argue about splits. The answer, or the ambiguity, is in there.
What this means for the rest of us
I test tools built on scraped data every week. Most of them are useful. Some are genuinely good. That doesn’t make the sourcing question go away, and this dispute shows exactly how it resolves in practice: slowly, expensively, and mostly to the benefit of whoever holds the better contract.
If you write, code, design, or record anything that ends up online, this is a preview. When compensation for training data becomes normal rather than exceptional, the fight won’t be about whether you’re owed something. It’ll be about which entity in your history has a claim on it. Check your contracts now, while there’s no money on the table and nobody’s motivated to interpret them creatively.
For AI companies, there’s a lesson in here too. Vague sourcing creates vague liability, and vague liability turns into settlements that nobody can distribute cleanly. The tools I’d bet on long-term are the ones that can tell you where their data came from with a straight face.
Authors got the first real test of what AI training compensation looks like. The result is a mess of competing claims, contracts written for a different era, and scammers working the margins. Everyone else building on top of other people’s work should watch closely — the same questions are coming for you, and the contracts you signed will answer them for you if you don’t read them first.
🕒 Published: