
OpenAI’s Deleted Book Datasets Stay Shielded After Court Reverses Disclosure Order

OpenAI is no longer required to disclose privileged internal communications explaining why it deleted the Books1 and Books2 training datasets. In February 2026, a federal district judge set aside the earlier disclosure order, concluding that its three grounds for finding a waiver of attorney-client privilege were legally erroneous.
That reversal is the most important change since the controversy intensified in late 2025. It does not erase the datasets from the litigation, establish that their acquisition was lawful or resolve the authors’ copyright claims; versions recovered by OpenAI have been provided to the plaintiffs, who dispute whether they are complete. As of August 2026, the wider consolidated copyright litigation remains active, while a separate July discovery dispute reported by the Associated Press shows that publishers and OpenAI are still contesting what evidence must be produced.
What Books1 and Books2 were
Books1 and Books2 were not generic labels for every book used by OpenAI. The court record describes two specific datasets built from books downloaded from Library Genesis, commonly known as LibGen, in 2018. OpenAI used the resulting collections to train GPT-3 and GPT-3.5, then stopped using them for model training in late 2021.
OpenAI deleted the datasets around the middle of 2022, approximately a year before the first actions covered by this dispute were filed. That chronology matters: deletion is established, but timing alone does not prove that the company destroyed evidence in anticipation of a particular lawsuit. The later discovery fight concerned what OpenAI had to reveal about its reasons for the deletion, not whether the deletion occurred.
The recovered copies also complicate the simple claim that the evidence disappeared permanently. OpenAI eventually produced recovered versions to the class plaintiffs. The plaintiffs’ position that those versions may be incomplete remains an allegation noted in the court record, rather than a judicial finding that OpenAI withheld or destroyed particular books.
Why the November disclosure order attracted attention
On November 24, 2025, Magistrate Judge Ona T. Wang found that OpenAI had waived privilege over certain 2022 communications concerning the deletion and internal references to LibGen. The court’s November opinion and order directed OpenAI to produce specified communications with in-house counsel and allowed limited depositions of lawyers involved in the discussions.
The ruling focused heavily on OpenAI’s litigation positions. In 2024, its outside counsel said Books1 and Books2 had been deleted because they were no longer in use. In 2025, OpenAI attempted to withdraw that language and maintained that communications containing its reasons for deletion were privileged. Judge Wang characterized the company’s privilege positions as a moving target and also reasoned that OpenAI had put its good faith at issue by denying allegations of willful infringement.
Even that order stopped short of finding a criminal cover-up. Judge Wang rejected the plaintiffs’ attempt to invoke the crime-fraud exception, reasoning that communications about deletion occurred after the alleged acquisition of books from LibGen. She also rejected the proposition that deleting the datasets amid legal uncertainty was, by itself, enough to establish spoliation.
Why the disclosure requirement was reversed
District Judge Sidney H. Stein reviewed the November ruling under the standard governing objections to a magistrate judge’s non-dispositive order. He concluded that none of its three waiver theories could support compelling the protected communications.
First, OpenAI’s statement that the datasets were deleted “due to non-use” did not reveal legal advice. Because the disclosed statement was not itself privileged, the court held that saying it publicly could not waive confidentiality over separate communications between OpenAI and its lawyers.
Second, the district judge distinguished awkward or changing language about the reasons for deletion from OpenAI’s position on legal privilege. His ruling said OpenAI had consistently maintained that confidential communications seeking or providing legal advice were protected. Its later assertion that there were no non-privileged reasons for deletion was described as inartful, but not a sufficient basis for imposing waiver as a sanction.
Third, simply denying willful infringement did not make counsel’s advice part of OpenAI’s defense. A party ordinarily puts privileged advice at issue when it affirmatively relies on that advice, not whenever it denies bad faith. OpenAI represented that it would not use advice of counsel to prove good faith, and the court treated that limitation as significant.
What the reversal does—and does not—decide
The February ruling is a discovery and privilege decision, not a judgment that OpenAI’s training practices were lawful. It prevents the plaintiffs from obtaining the protected deletion discussions through the November waiver order. It does not determine whether copying books from LibGen infringed copyright, whether model training qualifies as fair use, or whether any model output infringed an author’s protected expression.
Likewise, the ruling does not confirm the stronger allegation implied by the phrase “copyright evasion.” The public record establishes that LibGen-derived datasets were created, used for training and later deleted. A claim that deletion was designed to conceal infringement requires additional evidence about intent; the communications that might illuminate legal advice remain privileged under the February decision.
The dispute therefore has an unusual split outcome. Plaintiffs possess recovered versions of the datasets and may use permissible discovery to examine their contents and role in training. But they cannot rely on the overturned waiver ruling to inspect confidential discussions with OpenAI’s lawyers about why deletion occurred.
Why the distinction matters beyond OpenAI
The decision draws a boundary between evidence about an AI training pipeline and legal advice about that pipeline. Dataset provenance, dates of acquisition, training use, deletion mechanics and recovered files can be discoverable facts. Confidential requests for legal advice do not automatically lose protection because a company publicly offers a non-legal explanation for its conduct.
For copyright holders, this means that proving knowledge or willfulness cannot rest on the assumption that deletion equals concealment. The stronger route is evidence showing what works were acquired, where they came from, how they were used and whether the relevant decision-makers understood the legal and factual circumstances. For AI developers, the same record illustrates why documented data lineage and consistent retention practices matter long before litigation begins.
The central controversy has consequently narrowed rather than disappeared. OpenAI won the privilege appeal, but the provenance and use of Books1 and Books2 remain relevant to unresolved copyright claims. The accurate current account is neither that the deletion scandal has been proved nor that the datasets no longer matter: the recovered data remains available to the plaintiffs, while the company’s confidential legal discussions remain shielded.
Also read:
Related articles


Zuckerberg Questioned Creators’ AI Leverage. Copyright Claims Survived

Nvidia's AI Training Scandal: Emails Reveal Pursuit of Pirated Books Amid Copyright Lawsuit

No Registration, Still a Case: Copyright Claims in India Remain Enforceable

Jason Allen Says His AI Art Was Copied—But His Copyright Case Is Still Pending

Torrenting Isn’t Automatically Illegal—but Your Download May Also Upload
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.