Last week, the less-redacted versions of the plaintiffs’ briefs in the OpenAI MDL were released. I was glad – the redacted versions were so heavily blacked out as to be almost unreadable:

On the other hand, it was clear from context that the redactions were mostly statements attributed to people working at OpenAI or Microsoft; it was a game of gotcha. As I’ll explain below, this signaled to me that whatever I was missing, it didn’t really matter—there’s not much a user can say that should change a fair use analysis. But that doesn’t mean it won’t change the vibes outside the courtroom.
Sure enough, when the black bars were removed, we learned that some people inside OpenAI and Microsoft were concerned about how AI models are trained and the effects AI models might have on the internet generally, and on news specifically. News outlets had a field day with quotes about the “largest theft of labor in human history” and a “doom loop” of substitution that would kill news, the internet, and eventually AI itself. Headlines blared that engineers thought sourcing book data from the shadow library LibGen was “sketchy AF.” Gotcha! How seriously should we take these “gotcha” quotes from non-lawyer personnel at OpenAI and Microsoft, when it comes to fair use?
I sometimes tell my clients that they can trust their own instincts when it comes to fair use. Other times, I have to tell them to forget everything they think they know about copyright. Fair use depends on the facts of a particular situation, and similarly, whether copyright law fits a layperson’s intuitions can vary depending on the context (and how they got their intuitions). In this case, I see a few reasons to think we should ignore the still small voices telling engineers that unlicensed AI training is icky.
Most importantly, these stirrings of conscience appear to spring from a fundamentally mistaken understanding of how copyright works and what it protects. For example, Microsoft researcher Brent Hecht’s “biggest theft of labor” line suggests he may be under the mistaken impression that copyright exists to protect labor. It doesn’t.
As Justice Sandra Day O’Connor wrote in the landmark case Feist v. Rural Telephone Co., “all facts—scientific, historical, biographical, and news of the day…are part of the public domain available to every person.” That bedrock principle doesn’t change just because sometimes people have to work to discover facts. That was the issue at the heart of Feist, a case about competing phone books (remember those?). The first publisher argued that they expended substantial effort gathering all those phone numbers, and so copyright should block a second publisher from just copying those numbers for a competing directory. It was known as the “sweat of the brow” theory: the argument that copyright should protect facts that required work to uncover or compile. The Supreme Court rejected that theory categorically in Feist.
Justice O’Connor explains, “The primary objective of copyright is not to reward the labor of authors, but ‘[[t]o promote the Progress of Science and useful Arts.’” That ultimate goal of progress is best served when facts and ideas are free to circulate, even if it means someone who collects and reports those facts loses an opportunity to make a buck. Feist quotes from another landmark case, Baker v. Selden, which explains that “The very object of publishing a book on science or the useful arts is to communicate to the world the useful knowledge which it contains. But this object would be frustrated if the knowledge could not be used without incurring the guilt of piracy of the book.” If Professor Hecht had a crisis of conscience over using the facts revealed in news reporting without first asking permission from the publishers, his conscience is out of sync with copyright law.
Hecht’s “doom loop” theory (that people will prefer to get information from AI chatbots rather than websites, which will undermine interest in other information sources, undercutting those sources and ultimately starving AI of training data) misses the point for the same reason. If AI tools satisfy users’ interest in finding facts better than news websites, for example, that’s a good thing from a copyright law point of view. Copyright protects the news publishers’ investments in compiling and presenting news stories by ensuring that only they can offer consumers the facts in that particular expressive package, but it leaves others free to present the same facts in a different package. No publisher has ever been allowed to prevent competitors from re-reporting the same facts. Finding the most attractive package for public domain information is a business problem, not a copyright one.
Another reason to discount the “gotchas” revealed last week is that fair use doesn’t depend on the user’s subjective knowledge or intent. What matters is what you actually do. That’s why you can’t successfully invoke fair use just by writing “No copyright infringement intended!” in the description of a social media post that is just a copy of a song you like. It’s also why Andy Warhol’s estate wasn’t able to defeat photographer Lynn Goldsmith just by invoking Warhol’s different aesthetic intent. Courts look at what the user actually made and the use’s relationship to the work(s) they used, in culture and in the economy. Whether engineers or even executives speculated about a model’s potential impact on news or books doesn’t matter nearly as much as what a model actually is and does.
We’ve been down this road before. From time to time over the last several years, earnest ‘whistleblower’ engineers at AI labs have gotten a few minutes of fame by coming forward to offer their armchair copyright opinions, or reporters have gotten ‘scoops’ with some leaked memo or email from an executive opining on copyright. I can’t help but notice none of these people are attorneys, and nothing they say about fair use ever adds up. Thankfully, none of it really matters in the courts, either.