Graphic by Peg Ihinger

Having lost the past week to the Anthropic mess, I am going to indulge myself in a minor rant. Bear in mind that I am neither a lawyer nor a computer programmer, and there are much broader issues than just the ones that affect published writers.

For those not aware, Anthropic is one of the companies that has been researching AI and producing LLMs (Large Language Models, which is what ChatGPT and the other “AI” programs actually are). Anthropic’s LLM is called Claude.

LLMs need to be fed massive amounts of data to be “trained” properly. There’s a lot of data available on the Internet, and back in 2022, when Claude was being developed, Anthropic decided to scrape up some of it and feed it to Claude. Among the things they used were books, both fiction and non-fiction.

Some authors objected to having their work used to train LLMs, and sued Anthropic for copyright violation. The lawsuit quickly became a class-action suit. (Anthropic is currently facing similar lawsuits from the music industry and Reddit, but I am sticking to the book part.) The court decided that using books to train an LLM comes under Fair Use (which a lot of people still strongly dislike, but reading books is a good way for anybody to learn things).

Unfortunately for Anthropic, they weren’t as careful about their sources as they should have been, and some of the books they used were from pirate websites. They also kept copies of the pirated books in their training library, to the tune of seven million pirated titles. Anthropic decided to settle, rather than let the case go to court. They decided to offer $3,000 per title to the rights holders (split between author and publisher). The settlement was accepted last year.

At this point Anthropic began the normal process for handling a class-action lawsuit: first, notifying possible claimants that they can file a claim; second, collecting and processing claims that do get filed; and finally, making the payouts.

And this is where things started going south for the rest of us.

In fairness, I think Anthropic was doing their best to make things easy and accessible. Instead of the usual little postcard saying “You are a possible claimant…” in eye-wateringly small print, they set up a nice database with all the information they could find about all the titles they’d used. The trouble was that they found a lot of data, not all of which was accurate. I suspect, but cannot prove, that a) they didn’t understand the book business at all, and b) they used an AI to look for possible claimants.

I can almost (but not quite) forgive b). It’s a lawsuit, and going over and above in their effort to identify any and all possible claimants is a lot easier to justify than missing some people who may later file more lawsuits. But there is no universe in which my artist sister’s name should appear under “author” for my first novel next to mine. She didn’t even get to do the cover painting. It also makes no sense that several other books list three authors: “Wrede, Patricia C.”, “Patricia C Wrede,” and “Patricia C. Wrede 1953–”  (That last entry has a lot to do with my suspicion that they used their AI to collect “any and all possible rights holders,” because any sane human being would recognize in a second that all three of them are the same person, and no one would add my birth year as part of my name.)

The other issue is not realizing that they should have checked with somebody who understands the book business. Again, I think they tried, but even in the initial list of “possible claimants,” they only recognized authors and publishers. Which doesn’t cover things like work-for-hire contracts, LLCs and S corporations that authors have set up for tax and/or liability protection, or literary trusts established as part of an author’s estate, to name a few. I am also completely unclear about what they did about anthologies (as opposed to single-author collections).

In addition, Anthropic’s initial definition of “publisher” apparently meant “anybody claiming an association to the book,” including literary agents, library rebinding companies, and pirates/scammers (which is especially ironic, since scraping books from pirate web libraries is what got them in trouble in the first place).

Going through all that in order to file a claim was a time-consuming, annoying pain that cost me a good two weeks of writing time and some valuable brain cells. Last week came the second part, where Anthropic presented their “cleaned-up” list of people whom they thought were the actual rights holders.

This time, it only took me another week to figure out what they were actually asking for and how much I could and couldn’t do. (How do you prove you never had a contract with someone claiming to have rights to one of your books? Show them an empty file folder?) Anthropic also keeps sending me notices that a publisher is claiming 100% of “one of my titles” (they don’t say which, but I suspect one or more of the Star Wars novelizations, because those are all work-for-hire and I therefore don’t hold any of the rights to them, even though my name is on the cover). They have provided no information about which title, which publisher, or what I should or could do about having 100% of a title claimed, they just keep sending alarming notices.

My “cleaned up, final list of claimants” included a bunch of duplicates (one of mine still had me listed three times under different versions of my name), two legitimate literary agencies (neither of which submitted a claim, and both of which were dumbfounded to find themselves listed as “rights holders”. They gladly supplied me with disclaimers to that effect, but I still had to spend time chasing them down.), several titles with the same publishing house listed twice under slightly different versions of their corporate names, and a bunch that showed a library rebinding company as the publisher.

I was lucky; I wasn’t targeted by any scammers (and from what I hear, quite a few people were). Also, I keep very good records (it comes from having a business degree). I know at least two different people who took a look at the list, buried their heads in their hands, and clicked “submit” even though they could see there were similar, obvious errors. Because they simply couldn’t face trying to correct the mess again.

Apart from somewhat relieving my feelings on the matter, the point of this post is that the business end of writing is a pain in the neck that sooner or later wants records, which one needs to have kept from the start. Contracts, letters, royalty statements, rights reversions, options…you keep them all, forever. Because stuff like this can happen forty or fifty years later, and you can’t do anything about it if you don’t have the documents and the will to use them.

1 Comment
  1. I’ve been scraped, and I’m still really angry about it. But I won’t rant. (Much.) I will say, though, the more often things like this happen, the less confidence people have in the legal system.

    That’s “fair use”? Really?

    Also, I don’t say justice system, I say legal system.

    Ms. W, my sympathies. I certainly haven’t had to spend two weeks on this mess.

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.