If you’re following the AI copyright wars, you may have run across a recent study on the impact of AI use on the market for self-published books. Analyzing a sample of 14,419 ebooks, the authors examine which ones include AI-generated text and how successful these AI-assisted books are in the marketplace. The authors have been promoting the study actively, and getting coverage in mainstream outlets like The New York Times and The Atlantic. The paper supports plaintiffs’ arguments against fair use in lawsuits against AI companies, which is deeply ironic because the research methods used in the paper, including massive amounts of unlicensed copying, are extremely similar to AI training. The authors even rely indirectly on some of the same “shadow library” datasets that have triggered copyright lawsuits. If not for fair use, the authors and their institutions would be liable for copyright infringement on a massive scale: up to $2 billion in potential damages for a single paper. Here’s how.
To analyze their sample of 14,419 self-published books, they had to get copies of the books. The study says, “We obtained full texts three ways: authors sent us digital files, we borrowed ebooks from online libraries, and we bought ebooks from Amazon.” They don’t say what format the digital files sourced from authors were in, but Amazon allows self-published authors to use Kindle Digital Rights Management (DRM), and most library ebooks are also protected with DRM. To modify and analyze the ebooks’ full text would likely require circumventing any such technical protection measures, a prima facie violation of Section 1201, part of the Digital Millennium Copyright Act (DMCA). Fair use is not necessarily a defense to a Section 1201 claim, and damages can run as high as $2500 per act of circumvention. Depending on how many ebooks were cracked, this could trigger very substantial liability (over $36 million, if every book was cracked), leaving aside the cost and stress of defending a lawsuit.
Thanks to fair use advocates like the Authors Alliance and the Library Copyright Alliance, there is a DMCA exemption that allows scholars at research institutions to crack DRM on ebooks “solely to deploy text and data mining techniques on a corpus of literary works for the purpose of scholarly research and teaching.” If running chapters of books through a third-party AI detector counts as “text and data mining,” this research may qualify for the exemption. Of course, the rule is needlessly complex, and it can be tricky to satisfy all of its requirements. One of those requirements is that any cracked ebook should be “lawfully acquired and owned by the institution, or licensed to the institution without a time limitation on access.” The authors must also “use[] effective security measures to prevent dissemination or downloading of literary works in the corpus” and ensure “all access [to cracked ebooks is] provided only through secure connections and on the condition of authenticated credentials.” Hopefully, if their work involved cracking DRM, the authors were able to structure their research to satisfy these nitpicky provisions.
Once they had access to the full text of thousands of ebooks, the authors made thousands more copies of portions of the books and processed each one using an AI-powered text classifier: “We segment each book into individual chapters after excluding Contents and Acknowledgements, Dedications, Copyright pages, Author bios, Previews and pass each chapter to Pangram’s v3.3 API.” In addition to copying, this process involved distribution, another exclusive right under the Copyright Act, as the researchers sent text files containing the derivative excerpts to Pangram through their API.
For books obtained from authors, it’s possible the researchers obtained adequate permissions to engage in all this copying and distribution. For books purchased from Amazon or borrowed from libraries, however, this copying is unlicensed. Unless fair use applies, the study’s authors may have exposed themselves and their institutions to potentially ruinous levels of copyright liability. Statutory damages for copyright infringement can be up to $30,000 per work infringed. If the author-sourced books did not clearly include a license to make these copies, then damages could apply for all 14,419 books: that’s ~$432 million. It is unlikely a court would find these infringements to be “willful,” since the authors must have believed they were engaged in fair use, but the stakes of such an unlikely event are very high: at $150,000 per work infringed, damages for willful infringement exceed $2 billion. Even a very slight chance of being liable for $2 billion in damages could be enough to deter some scholars or their institutions from pursuing this kind of research.
Unfortunately, our intrepid researchers aren’t out of the woods, yet. Anyone who follows best practices in processing text to prepare it for analysis faces another source of liability: Section 1202 of the DMCA, which bars the removal of so-called “copyright management information” (“CMI”) under certain circumstances. CMI includes things like the “copyright pages” in the front matter of most commercial books. Like AI developers, the authors of the AI book study removed the copyright pages from the ebook copies before they fed the books into the computer. They did this because copyright pages are “noise” in the data (it’s not the actual text of the book) and could corrupt the results. Some of the biggest book publishers are including claims for removal of CMI in their lawsuits against AI developers. If courts accept their broad interpretations of this statute, it would certainly have a chilling effect on all research like this.
Proving the article’s final claim, that successful AI books use “rare existing-book language,” required even more unlicensed copying – dicing all 14k+ books into thousands of phrase “tokens” for computer analysis. It also involved reliance on two massive unlicensed datasets: “An expression counts as rare existing-book language when it appears in at most five Google Books volumes and has count zero in infini-gram, a 4.7-trillion-token web snapshot.” According to its website, infini-gram is based on troves of unlicensed data, including disputed datasets RedPajama and The Pile, which purportedly contain pirated books from the Books3 dataset. And of course Google Books was created largely without licensing and would not exist without fair use. If the plaintiffs suing AI companies get their way in court, these datasets will disappear and research like this will become impossible.
This case illustrates a fundamental irony in critiques of AI grounded in copyright law. Many scholarly critics of AI must themselves rely on the exact same computational methods—bulk copying, DRM bypasses, data cleaning, and massive un-licensed text corpuses—as AI developers. If the maximalist approach to copyright advanced by plaintiffs is applied consistently, it will outlaw powerful research methods that help us understand AI and its impacts. It is impossible to study, understand, and critique AI robustly without the same fair use rights that permit AI training.
This isn’t just a gotcha about hypocrisy by a handful of academics. It’s a warning: weakening fair use and expanding statutory liability doesn’t just check the power of tech giants. It jeopardizes the future of independent research using computer analysis of culture. Without robust legal protections like fair use, the empirical research needed to understand AI’s real-world impact will be priced out of existence, crushing free expression and scientific progress under the weight of billion-dollar liabilities.