Earlier today I happened upon an article in Futurism describing the latest on using published books to train AI. Basically, AI providers are buying used print books (not ebooks!) in bulk, taking off the covers by machine, and then scanning, OCRing, and training their AIs on the OCRed pages.
The twist is that unlike Anthropic’s $1.5B settlement over training AIs on pirated ebooks, buying physical books to train AIs is protected by something called the doctrine of first sale. What this means is that if you buy a print book but don’t distribute it in violation of copyright, the copyright owner can’t stop you from doing anything to that book. Scanning a legally purchased book and OCRing up a copy for yourself isn’t illegal. You bought the book. You OCRed it and loaded it on your tablet. You didn’t give a copy to anybody or post it in a pirately fashion. Theoretically, that’s legal. There are judicial rulings specifying that such an action is “transformative” and considered fair use.
Now, the weird part: Is feeding a scanned book to an AI a violation of copyright? Many think it is. I thought so at first. Much depends on how the AI is used. Suppose an AI trained on loads of books (including mine) were asked, “Show me the full text of Jeff Duntemann’s novel The Cunning Blood.” If it coughs up substantial parts of the book, that’s copyright infringement. But if it says, “Sorry, Dave. I’m afraid I can’t do that,” things get very fuzzy. Could it fork over a few paragraphs and a synopsis? Or just a synopsis? Without producing substantial chunks of literal text from the work in question, my understanding of the law suggests that it’s not infringement.
Now, suppose you ask an AI trained on my books, “Write a novel about a prison planet protected by nanomachines that corrode electrical conductors and thus make electrical devices impossible.” If it drops big chunks of my novel into its output, yes, that’s infringement. But if it writes an original novel with that core idea but without including substantial text from my book, I don’t think it’s infringement. Sleazy, maybe. But not illegal.
One metaphor here would be reading a lot of books we’ve purchased or borrowed from a library (something many of us do) and then writing material inspired or informed by what we’ve read. We digest a lot of factual or fictional material and then talk or write about that material. Absent literal transcription, that’s not copyright infringement.
The Futurism article is a little too worried about destroying huge numbers of print books and not worried enough about the theoretical legality of training AIs on legally-purchased print books. I suspect some law group may try to put together a lawsuit about it, but I think that won’t pass judicial muster.
There may already be machines that can flip through the pages of a print book and OCR its text without destroying the book. Even if they don’t exist (yet), if this means of training AIs on purchased print books becomes clearly legal, my guess is that those machines will happen.
I’m sure it’s possible to build guardrails into AI software preventing it from delivering literal content from OCRed books in copyright. I don’t know how hard that will be, but I suspect that such guardrails could be added to AI. And then, if I understand the legalities correctly, the whole problem goes away.
Sooner or later the problem will be solved. Stay tuned—and bring some popcorn.











