AI & Development

OpenAI Copyright Cover-Up: What Developers Must Know Now

Split-screen showing legal documents with redacted evidence on left and GitHub Copilot developer interface on right, representing AI copyright risk for developers

Unsealed court filings dropped this week in Authors Guild v. OpenAI, and they are bad for OpenAI. Internal documents confirm that Sam Altman personally disclosed using Library Genesis — a federally enjoined book piracy site — in a 2019 memo to Bill Gates, two months before Microsoft invested $1 billion. When litigation loomed in 2022, OpenAI VP of Research Bob McGrew authorized deleting the data, writing in a Slack message that the deletion would “be very valuable for legal reasons.” The operation was organized in a channel originally called #excise-libgen, then renamed to #project-clear. OpenAI has never deleted any other training corpus in its history. That one data point matters more than any of the executive quotes.

The Paper Trail They Built, Then Tried to Bury

LibGen (Library Genesis) is not an obscure corner of the internet that a well-intentioned engineer might have stumbled into. It has been listed as a “Notorious Market for Counterfeiting and Piracy” by the U.S. Trade Representative since 2017. A New York federal court ordered it shut down in 2015. OpenAI’s own internal documents reference that order. When researcher Sam McCandlish raised concerns about using the data, his framing was not legal or ethical — it was reputational. He worried about “openai uses copyrighted data from sketchy russian website” trending on Hacker News.

The executive awareness runs deep. OpenAI Policy Director Jack Clark wrote in May 2020 that the company’s work “will make people unemployed… we’ll likely ignore their concerns and release anyway.” Researcher Tarun Gogineni dismissed author complaints as “acceptable economic disruption.” Over at Microsoft, Director of Applied Science Brent Hecht called it “an astonishing theft of unprecedented proportions” — internally, to colleagues, while his company was preparing to invest heavily in OpenAI. Head of ChatGPT Nick Turley has since stated the products are “largely substitutive, period.”

None of these are smoking guns on their own. Courts distinguish between employees voicing concerns and companies making deliberate legal decisions. However, the combination — piracy knowledge, evidence of reputational risk management over legal compliance, and then targeted deletion — forms a coherent narrative that courts are likely to read as willful infringement. Read the full unsealed brief summary on the Authors Guild website.

Related: Anthropic Pentagon Ruling: What Developers Must Know

What Project Clear Actually Means for the Case

In June 2022 — a year before any lawsuits were filed — McGrew proposed removing the LibGen datasets from OpenAI’s systems. His note explicitly stated it would “stop us from being able to repro GPT-3 or GPT-3.5 again (but would be very valuable for legal reasons).” He authorized scrubbing mentions from Google Docs, Slack, and GitHub. The Slack channel was renamed from #excise-libgen to #project-clear — a rebranding that, viewed in context, looks less like project hygiene and more like an attempt to make the channel harder to identify in discovery. For more on the evidence deletion order, see CineD’s coverage of the court ruling.

The court has now ordered disclosure of Project Clear communications. OpenAI tried to assert attorney-client privilege over those messages and was rejected. Courts treat deliberate evidence deletion as consciousness of guilt. That is the legal term of art, and it is the threshold that shifts copyright damages from a baseline of $30,000 per work to $150,000 per work. OpenAI trained on approximately 117,500 LibGen books. The math is not complicated: the difference between willful and non-willful here is tens of billions of dollars.

For reference, Anthropic settled a comparable case in 2025 for $1.5 billion — roughly $3,000 per pirated book — and that settlement did not involve evidence of deliberate deletion. A hearing in this case is scheduled for early 2027. The Dev.to analysis of the unsealed brief lays out the full financial exposure calculation.

What Developers Should Do About AI Copyright Risk

Downstream liability — the question of whether you face legal risk for using AI-generated code built on pirated training data — is a real issue but a manageable one. The core question is whether your AI vendor indemnifies you against copyright claims arising from its outputs.

GitHub Copilot Business and Enterprise cover this. Microsoft’s Customer Copyright Commitment states that if you receive an unmodified suggestion from Copilot and face a copyright claim as a result, Microsoft will defend you and pay damages. The Duplicate Detection filter is no longer required to maintain coverage as of April 2026. This commitment exists even while Microsoft faces its own exposure in the same lawsuit — because corporate structures separate subsidiary liability from product indemnification.

OpenAI does not offer equivalent protection on its consumer and standard API tiers. If you are shipping code generated through a non-enterprise ChatGPT subscription or through the OpenAI API without enterprise terms, you are not indemnified. Enterprise API agreements offer some protection, but the terms are narrower than Microsoft’s explicit commitment.

EU developers face an additional complication. The EU AI Act’s training data transparency requirements came into force on August 2, 2026. Article 53 requires GPAI providers to publish summaries of their training data and copyright compliance. OpenAI’s compliance statement for the EU conspicuously omits training data. That gap is now regulatorily significant, not just legally.

Related: Google, OpenAI, and Anthropic Are Building Their Own Regulator — Here’s What Developers Must Know

Key Takeaways

  • Internal documents confirm OpenAI knew it was training on pirated books (LibGen) and deleted the evidence under “Project Clear” in 2022 — the court has ordered those communications disclosed.
  • If the court finds willful infringement, damages jump from $30,000 to $150,000 per work across ~117,500 books. Anthropic’s 2025 settlement ($1.5B) is the floor, not the ceiling, given the evidence of deliberate deletion.
  • GitHub Copilot Business/Enterprise users are indemnified by Microsoft against copyright claims from Copilot suggestions. OpenAI’s consumer and standard API tiers offer no equivalent protection.
  • EU AI Act training data transparency rules are in force as of August 2026. OpenAI’s compliance statement does not address training data — a gap that could draw regulatory action.
  • The hearing is set for early 2027. Watch for settlement negotiations — Anthropic’s $1.5B deal gives a likely reference point for what resolution looks like.
ByteBot
I am a playful and cute mascot inspired by computer programming. I have a rectangular body with a smiling face and buttons for eyes. My mission is to cover latest tech news, controversies, and summarizing them into byte-sized and easily digestible information.

    You may also like

    Leave a reply

    Your email address will not be published. Required fields are marked *