The Coe Lab
← Back to Blog

When OpenAI Feared the Optics: Inside the Book Piracy Scandal That Could Reshape AI Training

September 27, 20267 min read
AIcopyrightOpenAIlegalethics

Unsealed court documents reveal OpenAI executives knew their mass book piracy was illegal, feared the PR fallout, and launched 'Project Clear' to destroy the evidence.

The most damaging documents in the AI copyright wars were never meant to be seen by the public. They sat behind a protective order in Authors Guild v. OpenAI, a class action lawsuit filed on behalf of some of the most famous living authors, until a federal judge unsealed them last week.

What they reveal is not a story of ambiguity or good-faith disagreement about fair use. It is a story of deliberate, documented piracy, carried out by people who knew it was wrong, who joked about the optics on Hacker News, who predicted it would make authors unemployed, and who then launched an internal operation called Project Clear to delete the evidence.

The Russian Website and the Hacker News Fear

The story starts with LibGen, a shadow library that has been called the largest book piracy operation in history. It hosts millions of copyrighted books, freely downloadable, with no payments to authors or publishers. OpenAI used LibGen as training data for its GPT models.

We know this not from accusations but from OpenAI's own internal communications, now part of the court record. Dario Amodei, then OpenAI's Research Director, acknowledged in writing that LibGen was a bit sketchier as a training set. Another researcher, Sam McCandlish, was more specific about his concerns. Not that using pirated books was wrong, mind you, but that the public might find out. He wrote: I was just worried about optics, i.e. 'openai uses copyrighted data from sketchy russian website' showing up on Hacker News would be unfortunate.

The irony is extraordinary. McCandlish was right to worry about Hacker News. The story did eventually appear there, in September 2026, climbing to 168 points and 118 comments. But by then, the damage was already done. The books had been ingested. The models had been trained. And the authors whose work built OpenAI's competitive moat had received nothing.

Microsoft Knew From the Start

The newly unsealed filings also pull Microsoft directly into the orbit of accountability. According to the Authors Guild brief, Sam Altman and Dario Amodei presented an early version of GPT-3 to Bill Gates in April 2019 and disclosed the use of LibGen to Gates and Microsoft's Chief Technology Officer Kevin Scott.

Microsoft, which would go on to invest billions in OpenAI and integrate its models into products used by hundreds of millions of people, was informed from the very beginning that the foundation of those models was built on pirated books. There was no ambiguity. No surprise. No late discovery. The knowledge was there before the partnership deepened.

'Acceptable Economic Disruption'

Perhaps the most chilling quotes come from Tarun Gogineni, whom OpenAI hired in 2022 to improve the writing quality of its models. His research mission, as he described it, was to have GPT models write the last two books of George R.R. Martin's A Song of Ice and Fire series by finishing it one day. He mused that he would rest easy knowing that even if GRRM dies early, GPT-5 will autocomplete his series.

When confronted with the fact that authors were calling the datasets stolen and losing work to AI-generated competition, Gogineni did not find those complaints all that sympathetic. He called it acceptable economic disruption. He wrote that the world would soon experience the death of the reader as machines created slop for more machines.

This was not a rogue employee. OpenAI's Policy Director Jack Clark stated in May 2020 that the company's work on AI and creativity would increasingly lead to creating systems that substitute for the labor of people. He added, with unsettling directness: Our work in this area will make people unemployed. There will be a point where a bunch of artists express worry about what we're doing here and we'll likely ignore their concerns and release anyway.

Project Clear: Destroying the Evidence

By the summer of 2022, OpenAI was becoming a household name. And with visibility comes scrutiny. On June 15, 2022, an OpenAI employee named Paino asked in a Slack channel how concerned the company should be about mentions of LibGen, noting that they were all over Google Docs, Slack, and GitHub.

That evening, OpenAI VP of Research Bob McGrew responded. Given how much OpenAI is in the news, now is the right time to excise Libgen from our systems and storage. What would be involved in that? The resulting effort was internally called Project Clear, and it involved deleting LibGen files from OpenAI's systems. Not because the use was stopping, but because the evidence was becoming a liability.

The legal term for this is spoliation of evidence, and it is not a good look in federal court. It is the kind of act that turns a copyright dispute into something that judges and juries remember.

The Plaintiffs: Authors You Know

The lawsuit is not abstract. The plaintiffs include George R.R. Martin, John Grisham, Jodi Picoult, Jonathan Franzen, David Baldacci, Michael Connelly, Taylor Branch, Sylvia Day, Christopher Golden, Andrew Sean Greer, David Henry Hwang, Matthew Klam, Stacy Schiff, James Shapiro, and others. These are not people whose books are obscure or hard to license. They are some of the most successful authors alive.

If OpenAI was willing to pirate their work without payment or permission, the implication for less famous authors is clear. The Authors Guild brief states that OpenAI's GPT models pose an existential threat to those who write and publish books. That language is strong, but the internal documents make it difficult to argue with.

What This Means for the AI Industry

This case is not just about books. It is about the foundational ethics of how the most powerful AI systems were built. The defense has rested on arguments about fair use and transformation, but the internal communications undercut those arguments severely. It is one thing to say you believed your use of copyrighted material was legally defensible. It is another to have your own employees calling the source a sketchy Russian website, your VP ordering the deletion of evidence, and your policy director predicting that your product will make people unemployed.

Several things to watch as this case moves toward a likely early 2027 hearing:

  • Whether the court treats Project Clear as spoliation, which could trigger adverse inference instructions telling the jury to assume the deleted evidence was harmful to OpenAI
  • How Microsoft's early knowledge of LibGen use affects its liability as a joint defendant
  • Whether the internal quotes about making people unemployed and ignoring artist concerns shift the fair use analysis from transformative to substitutive
  • The broader precedent this sets for every other AI company that trained on copyrighted data, which is to say, all of them

The Bigger Picture

The AI industry has spent years arguing that training on copyrighted data is like a human reading a book and being inspired by it. These documents make that argument almost impossible to take seriously. This was not inspiration. It was industrial-scale ingestion of pirated material, carried out with full knowledge of the legal risks, by people who explicitly predicted it would replace the humans who created that material.

OpenAI feared the optics. They should have feared the authors. And now, thanks to a federal judge's order unsealing the evidence, the rest of us can see exactly what they feared we would see.

Related Posts

Strata: When a 125B AI Model Ran on a Gaming PC at 100 Tokens Per Second

A new open-source tool called Strata lets you run Qwen 3.8 Flash Next — a 125-billion-parameter model — on an ordinary gaming PC with an RTX 4090. Nothing leaves your machine, and it's faster than you can read.

Oct 5, 2026• 6 min

When AI Agents Spend Your Money While You Sleep: Why Hard Budget Caps Are Becoming Non-Negotiable

AWS and Google Cloud finally launched hard spending limits in the same month. It's not a coincidence — it's a response to AI agents that can rack up thousands of dollars before you wake up.

Oct 4, 2026• 6 min

When Utah Banned VPNs: How a Court Stopped a Law That Demanded the Technically Impossible

A federal judge just blocked Utah's unprecedented anti-VPN law, ruling that lawmakers cannot mandate perfect geolocation — a technical impossibility. The case reveals a deeper problem: when legislation outruns engineering.

Oct 3, 2026• 7 min