In December 2024, a federal magistrate judge in the Northern District of California refused to let plaintiffs turn a targeted training-data case into an audit of a company's entire copyright practice, holding that a demand for all copies of copyrighted works was not proportional to the needs of the case. Eleven months later, a different court in a consolidated New York proceeding ordered an AI developer to produce 20 million de-identified ChatGPT conversation logs to news-publisher plaintiffs, a sampling drawn from tens of billions of preserved records. Those two orders, arising from the same wave of generative-AI copyright litigation, bracket the question every litigator and general counsel now faces: when the disputed system is a machine-learning model, what can you actually get in discovery, and where do courts draw the line? This article offers practical guidance on three things: (1) why discovery into an AI model differs from ordinary document discovery, (2) what categories of proof courts have allowed and refused, and (3) how model-inspection protocols and neutrals can keep the exercise proportional.
Why Is Discovery Into an AI Model Different?
A machine-learning model is not a document, and it is not a database in the ordinary sense. It is the product of a pipeline: a training corpus, the code and configurations used to ingest and filter that corpus, the compute run that produces a set of learned parameters (the model weights), and the outputs the finished model generates in response to prompts. Each layer answers a different question. The training corpus speaks to what the model learned from; the code speaks to how the data was processed and whether protective measures were applied; the weights are the model itself; and the outputs show how the system behaves in the wild. Counsel who ask for 'the model' without specifying the layer will get an objection, and they will deserve it.
This layered structure is why AI discovery fights look unfamiliar even to experienced litigators. The proof a plaintiff needs is often spread across sources that are enormous, sensitive, and expensive to collect, and the relevant material frequently sits inside the defendant's most closely guarded trade secrets. The rules, however, have not changed. Relevance under Federal Rule of Civil Procedure 26(b)(1) and proportionality to the needs of the case still govern, as does the court's authority to appoint a neutral under Rule 53 to manage the technical work. What is new is the scale and the medium, not the framework.
What Can Counsel Get?
Courts have been willing, in the right circumstances, to order production of the pieces of the pipeline that map directly to a pleaded claim. In Kadrey v. Meta Platforms, Inc., No. 23-cv-03417 (N.D. Cal.), Magistrate Judge Thomas S. Hixson ordered Meta to search the files of its designated custodians and produce documents and communications regarding the Llama models and the stripping or removal of copyright management information from literary works, reasoning that removal of that information is relevant to willfulness. The court also treated the acquisition of the LibGen dataset by torrenting as relevant to the claims as pleaded. The lesson is that discovery is available where the request is tied to a specific element of the case.
As a practical matter, four categories are most often in play. First, targeted training-data productions — the specific datasets, or the portions of them, alleged to contain the plaintiff's works, rather than the entire corpus. Second, source code and configurations governing how data was collected, filtered, deduplicated, and processed, which bear on both liability and any 'protective measures' defense. Third, model outputs, including logs of prompts and responses, which plaintiffs use to test whether a system reproduces protected material. The New York Times–led publishers pursued exactly this theory in the consolidated OpenAI litigation, and the court ordered production of a 20-million-log de-identified sample under a protective order. Fourth, the ordinary custodial record — email and internal chat — that explains what the engineers and business owners knew and decided.
What Counsel Cannot Get
The same orders that grant targeted discovery are notable for what they refuse. Proportionality is doing the heavy lifting. In Kadrey, the court declined to compel every copy of every copyrighted work Meta possessed, explaining the mismatch between the demand and the claim.
The Court also does not see how all copies of copyrighted works is proportional to the needs of this case, which is about the use of copyrighted materials to train the Llama models, not all copyright infringement committed by Meta.
The court applied the same discipline to custodial scope. Plaintiffs sought to treat work email and internal chat as 'non-custodial' sources so they could search the files of every one of the roughly one to two thousand Meta employees working in AI; the court refused, observing that this would effectively blow up the custodial limitations in the ESI order, and noting that the demand to expand from fifteen custodians to a thousand or more arrived nine days before the close of fact discovery. Timing and proportionality, not the novelty of the technology, decided the question.
Privilege is the other predictable limit, and it cuts both ways. The Kadrey court sustained many of Meta's attorney-client redactions after in camera review, but it also overruled redactions that shielded business decisions dressed up as legal advice, reminding the parties that putting a lawyer in a meeting does not make everything privileged. For counsel producing model-development records, the takeaway is that a decision made for mixed legal and business reasons will not be privileged as to its business components, and courts will look behind a privilege log that claims otherwise.
How Should a Model-Inspection Protocol Be Structured?
Because the sensitive layers of a model implicate trade secrets and, increasingly, third-party privacy, courts lean on protective orders and inspection protocols rather than open-ended production. The 20-million-log order in the consolidated OpenAI matter was expressly conditioned on de-identification and a protective order limiting who could access the data, how it could be used, and where it would be stored. Source code, model weights, and training data are candidates for the same treatment: production into a secure review environment, on non-networked machines, with access limited to named experts and logged. This is the model-inspection analog of the source-code review rooms that have long been standard in patent and trade-secret litigation.
A neutral with the requisite technical and legal expertise can make that structure work. Under Rule 53, a court may appoint a special master or technical neutral to supervise inspection, resolve disputes over what a dataset actually contains, and translate between the engineering reality and the legal question without either side having to expose more than the case requires. A neutral is a fact gatherer, not a problem solver, and that boundary is precisely what makes the role useful here: the neutral can confirm whether a filtering step existed or a dataset included the plaintiff's works while leaving the ultimate infringement and fair-use questions to the court. For more on scoping that role, see our discussion of working with technical neutrals and special masters.
Before you propose model discovery — or resist it — work through a short diagnostic list:
- Which layer of the pipeline does your claim actually require: training data, code, weights, outputs, or the custodial record?
- Can you tie each request to a specific element — infringement, willfulness, a protective-measures defense — the way the Kadrey plaintiffs tied CMI removal to willfulness?
- Is the volume proportional, or are you asking for the entire corpus when a targeted subset or a statistically sound sample will do?
- Does the request respect the existing ESI order's custodial limits, or does it try to rewrite them under a new label?
- What protective-order terms, secure-environment controls, and de-identification steps are needed before sensitive material changes hands?
- Would a Rule 53 technical neutral resolve the dispute faster and at lower cost than motion practice?
Conclusion
Discovery into AI models is expanding quickly, but the cases decided so far send a consistent message to the bench and the bar: the medium is novel, the rules are not. Courts have ordered training data, processing code, and even massive samples of model outputs where the request maps to a pleaded claim and rides inside a workable protective order, and they have refused corpus-wide demands, custodial end-runs, and over-broad privilege claims where it does not. The value of model discovery, in the right circumstances, can be decisive, but only when the request is scoped to the layer, tethered to an element, and protected by the right controls. As these disputes proliferate, we should continue to watch how courts calibrate proportionality against the scale of modern training and inference data — and counsel who bring in the requisite technical and legal expertise early will spend far less time litigating the shape of discovery and far more time using it. Early intervention is key.
The content is intended for general informational purposes only and should not be construed as legal advice.
Sources
- 01Kadrey et al. v. Meta Platforms, Inc., Public Version of Discovery Order (ECF No. 401) — U.S. District Court, N.D. Cal. (via Justia Dockets & Filings)
- 02OpenAI Loses Privacy Gambit: 20 Million ChatGPT Logs Likely Headed to Copyright Plaintiffs — The National Law Review
- 03OpenAI Loses Privacy Gambit: 20 Million ChatGPT Logs Likely Headed to Copyright Plaintiffs — Jones Walker LLP
- 04OpenAI Court Case: 20M ChatGPT Logs Ordered — Preservation and Production Timeline — Terms.Law
Editorial note — This briefing was drafted by an AI system from editor-selected sources and published under the editorial standards set out in our newsroom. It carries no individual byline because no individual wrote it. It is general information, not legal advice.
Who publishes this
Technical Special Master is edited and published by Daniel B. Garrie. Daniel B. Garrie is a court-appointed technical special master, discovery referee, and forensic neutral, and the founder of Law & Forensics LLC. He has served in more than one hundred court-appointed and expert-witness matters involving source code, e-discovery, cybersecurity, and artificial-intelligence systems, and is an adjunct professor at Harvard University.
For counsel: proposing a special master, model appointment orders, who pays a special master, e-discovery references, and the appointment packet.
Related reading
Technical Special Masters and AI: Resolving Algorithmic and Source-Code Disputes
Machine-learning systems break the assumptions behind ordinary discovery. A technical special master gives courts a way to examine models, training data, and outputs that the rules were not built to reach.
Authenticating AI-Generated Evidence Under Rule 901 in the Deepfake Era
How courts are applying Rules 901 and 902 to synthetic audio, video, and documents — and what proposed Rules 901(c) and 707 would change.
When Should a Court Appoint a Technical Special Master?
Rule 53 gives courts the authority to refer technical disputes to a neutral. The harder question is when an appointment is warranted. Five recurring signals tell the bench it is time.