© Copyright 2026 by Anderson Kill P.C. ClickySoft - WordPress Development Company
AI Updates with AK
Special Edition: OpenAI’s Discovery Dispute
A major discovery ruling in the ongoing In re OpenAI, Inc. Copyright Infringement Litigation, No. 1:25-md-3143 (S.D.N.Y.), may become the first significant test case addressing the discoverability of AI system logs. This development is particularly relevant to any organization that relies on AI tools internally because courts are beginning to confront when and how AI-generated records must be produced in litigation. The ruling arises out of The New York Times v. Microsoft & OpenAI, No. 23-cv-11195 (S.D.N.Y.), one of the lead cases in the multi-district litigation (MDL) challenging OpenAI’s use of news publishers’ content to train ChatGPT.
Magistrate Judge Wang has ordered OpenAI to produce twenty million anonymized consumer ChatGPT output logs covering December 2022 through November 2024. This dataset represents a very small sample of the billions of logs OpenAI maintains and excludes temporary chats, deleted chats, and Enterprise ChatGPT logs. OpenAI removed all personally identifiable information, though the thoroughness and effectiveness of that de-identification process remains unclear, and the production is governed by a strict confidentiality order. The publishers sought this data, having originally requested 1.4 billion logs, to assess how often ChatGPT generates text that is infringing, near verbatim, or substitutes for their reporting.
OpenAI resisted production on multiple grounds, including responsiveness, privacy concerns, proportionality, and a proposal to run search terms itself rather than turning over the raw data. Judge Wang rejected those arguments and granted the motion to compel, and although OpenAI has moved for reconsideration, the court declined to stay its order. This ruling provides an early indication that courts may be willing to require disclosure of extensive AI-generated logs even when doing so implicates privacy concerns and significant burdens.
The takeaway of the dispute thus far is clear. AI output logs are discoverable when relevant to a disputed issue. While only a one-percent sample of OpenAI’s full dataset was ordered for production here, and temporary and deleted chats were excluded, user-generated content stored by an AI provider may be subject to discovery if litigation targets that provider’s systems or training methods. Users concerned about confidentiality should be aware of how and when their chat histories are stored.
Further developments are expected in the coming weeks, and we will continue to monitor and report on them.


© Copyright 2026 by Anderson Kill P.C. ClickySoft - WordPress Development Company