{"id":5677,"date":"2026-07-29T10:23:42","date_gmt":"2026-07-29T17:23:42","guid":{"rendered":"https:\/\/www.contrapositivediary.com\/?p=5677"},"modified":"2026-07-29T10:24:45","modified_gmt":"2026-07-29T17:24:45","slug":"ai-and-the-doctrine-of-first-sale","status":"publish","type":"post","link":"https:\/\/www.contrapositivediary.com\/?p=5677","title":{"rendered":"AI and the Doctrine of First Sale"},"content":{"rendered":"<p>Earlier today I happened upon <a href=\"https:\/\/futurism.com\/artificial-intelligence\/ai-companies-destroying-rare-books\">an article in Futurism describing the latest on using published books to train AI<\/a>. Basically, AI providers are buying used print books (<em>not<\/em> ebooks!) in bulk, taking off the covers by machine, and then scanning, OCRing, and training their AIs on the OCRed pages.<\/p>\n<p>The twist is that unlike Anthropic\u2019s&#160; $1.5B settlement over training AIs on pirated ebooks, buying physical books to train AIs is protected by something called the doctrine of first sale. What this means is that if you buy a print book but don\u2019t distribute it in violation of copyright, the copyright owner can\u2019t stop you from doing anything to that book. Scanning a legally purchased book and OCRing up a copy for yourself isn\u2019t illegal. You bought the book. You OCRed it and loaded it on your tablet. You didn\u2019t give a copy to anybody or post it in a pirately fashion. Theoretically, that\u2019s legal. There are judicial rulings specifying that such an action is \u201ctransformative\u201d and considered fair use.<\/p>\n<p>Now, the weird part: Is feeding a scanned book to an AI a violation of copyright? Many think it is. I thought so at first. Much depends on how the AI is used. Suppose an AI trained on loads of books (including mine) were asked, \u201cShow me the full text of Jeff Duntemann\u2019s novel <em>The Cunning Blood<\/em>.\u201d If it coughs up substantial parts of the book, that\u2019s copyright infringement. But if it says, \u201cSorry, Dave. I\u2019m afraid I can\u2019t do that,\u201d things get very fuzzy. Could it fork over a few paragraphs and a synopsis? Or just a synopsis? Without producing substantial chunks of literal text from the work in question, my understanding of the law suggests that it\u2019s not infringement.<\/p>\n<p>Now, suppose you ask an AI trained on my books, \u201cWrite a novel about a prison planet protected by nanomachines that corrode electrical conductors and thus make electrical devices impossible.\u201d If it drops big chunks of my novel into its output, yes, that\u2019s infringement. But if it writes an original novel with that core idea but without including substantial text from my book, I don\u2019t think it\u2019s infringement. Sleazy, maybe. But not illegal.<\/p>\n<p>One metaphor here would be reading a lot of books we\u2019ve purchased or borrowed from a library (something many of us do) and then writing material inspired or informed by what we\u2019ve read. We digest a lot of factual or fictional material and then talk or write about that material. Absent literal transcription, that\u2019s not copyright infringement.<\/p>\n<p>The Futurism article is a little too worried about destroying huge numbers of print books and not worried enough about the theoretical legality of training AIs on legally-purchased print books. I suspect some law group may try to put together a lawsuit about it, but I think that won\u2019t pass judicial muster.<\/p>\n<p>There may already be machines that can flip through the pages of a print book and OCR its text without destroying the book. Even if they don\u2019t exist (yet), if this means of training AIs on purchased print books becomes clearly legal, my guess is that those machines will happen.<\/p>\n<p>I\u2019m sure it\u2019s possible to build guardrails into AI software preventing it from delivering literal content from OCRed books in copyright. I don\u2019t know how hard that will be, but I suspect that such guardrails could be added to AI. And then, if I understand the legalities correctly, the whole problem goes away.<\/p>\n<p>Sooner or later the problem will be solved. Stay tuned\u2014and bring some popcorn.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Earlier today I happened upon an article in Futurism describing the latest on using published books to train AI. Basically, AI providers are buying used print books (not ebooks!) in bulk, taking off the covers by machine, and then scanning, OCRing, and training their AIs on the OCRed pages. The twist is that unlike Anthropic\u2019s&#160; [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[23],"tags":[118,156,122,280],"class_list":["post-5677","post","type-post","status-publish","format-standard","hentry","category-ideasandanalysis","tag-ai","tag-copyright","tag-law","tag-pubishing"],"_links":{"self":[{"href":"https:\/\/www.contrapositivediary.com\/index.php?rest_route=\/wp\/v2\/posts\/5677","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.contrapositivediary.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.contrapositivediary.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.contrapositivediary.com\/index.php?rest_route=\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.contrapositivediary.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=5677"}],"version-history":[{"count":1,"href":"https:\/\/www.contrapositivediary.com\/index.php?rest_route=\/wp\/v2\/posts\/5677\/revisions"}],"predecessor-version":[{"id":5678,"href":"https:\/\/www.contrapositivediary.com\/index.php?rest_route=\/wp\/v2\/posts\/5677\/revisions\/5678"}],"wp:attachment":[{"href":"https:\/\/www.contrapositivediary.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=5677"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.contrapositivediary.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=5677"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.contrapositivediary.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=5677"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}