A note from Re:Create: A new paper on “market dilution” actually shows why it’s irrelevant to copyright

If you’re following the AI copyright wars (and readers of this blog surely are), you may have run across a recent study purporting to show the impact of authors using AI on the market for self-published books. The title says it all: “Generative AI floods and dilutes the market for books.” The study’s authors, Tuhin Chakrabarty, Xinyue Liu, Jane C. Ginsburg, and Paramveer Dhillon, have been promoting it actively on social media, and getting some uptake in mainstream outlets like The Atlantic. They say they hope to influence ongoing litigation with this work, which may be why they were able to obtain all of their market data from a “proprietary sales panel…provided by one of the ‘Big Five’ publishers.” The findings are interesting: they show that although self-published authors using AI are generally much less successful than those who do not use AI, some AI-assisted books do compete in the market for self-published ebooks and garner a small portion of sales and revenue. They also show that the AI-assisted portion of sales grew over the three years they studied. 

Ironically, the paper’s findings also show why “market dilution” is not a market effect that matters for fair use. 

Fair use is a flexible doctrine that “permits courts to avoid rigid application of the copyright statute when, on occasion, it would stifle the very creativity which that law is designed to foster.” Courts considering whether a use is fair apply a four-factor test. The fourth factor, which can receive substantial weight, is “the effect of the use upon the potential market for or value of the copyrighted work.” “Market dilution” is the novel (and constitutionally dubious) theory that increased competition from authors using AI tools to create new, non-infringing works should count as an adverse “market effect” and weigh against fair use for AI training.

Not all market effects count in the fair use calculus. The most common example of an effect that doesn’t count is a scathing book review, which might quote liberally from its target to demonstrate how bad it is, deflating consumer interest in the process. In that case, the market effect doesn’t count because it results from the critic’s compelling arguments, not the reader’s ability to ‘enjoy’ the unlicensed quotations without paying for the book. To count against fair use, a market effect should be caused by the alleged infringement rather than the user’s ingenuity. 

One way to show empirically that a particular market effect isn’t caused by alleged infringement, and so should not count in the fair use analysis, would be to show that the same effect is felt equally by works that were not used by the alleged infringer. That is what Chakrabarty et al. show in their paper. The market effects they show are market-wide—they touch every book in the market in the same way, without regard to whether a book was used in AI training. Every book in their dataset (including other books by AI-assisted authors) feels the same generic effect of competitive pressure from more new works. And the competition they document is between books released in the same relatively short time period, making it unlikely that the human-authored books were used to train the AI models used by the AI-assisted books. (Such an effect is impossible in many cases—AI-assisted books published in 2023 can’t have benefited from training on books that didn’t exist until 2025.) 

Another way we know that “market dilution” doesn’t result from infringement is that licensing all AI training, if it were possible, would not protect authors from market dilution. Fully-licensed AI tools would still allow AI-assisted authors to compete with no-AI authors in exactly the same way and to the same extent as AI tools trained on unlicensed data are used today. That’s because increased competition from AI is not a copyright injury. It’s not a consequence of infringement; it’s a consequence of the existence of generative AI

To put it another way, increased competition from AI-assisted authors is an effect of the novelty and ingenuity of AI technology—its ability to generate new, non-infringing expression. Dilution is not an effect of the specific raw material in the training data. Like the clever critic, an AI developer impacts the market by adding something new, not merely repackaging infringing copies. That is the hallmark of fair use, not infringement.

Remarkably, the authors acknowledge in footnote 11 that, “[Books that contain AI] outputs may also compete with [entirely] human-written works that the AI system did not copy, but that substitution effect is not within the purview of copyright law: competition without copying is not infringement.” Indeed. And yet all of the market effects documented in their paper, between AI-assisted and non-AI-assisted books published at roughly the same time, are exactly that: competition without copying. These effects, which are “not within the purview of copyright law,” are exactly the same effects alleged by copyright plaintiffs: AI allows authors to create more, which means all authors are competing with more works in the marketplace. When the outputs are non-infringing new works, there is no reason to treat this harm as a copyright harm. It’s just competition.

Copyright is not a collective right to a particular set of market conditions. It did not protect painters from photography, silent film actors from talkies, or composers from player pianos. Copyright has never been the right policy tool for managing disruptions from technological change. The advent of AI is no exception.