<
Case of the Week Blog Series

The AI Company Is the Evidence Now: What Britannica v. Perplexity Teaches About RAG and UAL Discovery

Share this article

[Editor’s Note: This article has been republished with permission. It was originally published September 9, 2026 on the Minerva26 Blog]

This week’s case teaches litigators and discovery professionals how AI works, and what they need to know to ask for the discovery of AI generated evidence. A Southern District of New York court ordered Perplexity AI to produce snapshots of its own retrieval database and months of its own user activity logs in a copyright case, the AI company’s own infrastructure treated as the evidence of infringement. Reading time approximately 11 minutes.

Podcast | Transcript

How AI Works, and What You Need to Know to Ask for It

This week’s case teaches litigators and discovery professionals how AI works, and what they need to know to ask for the discovery of AI generated evidence.

A court just ordered Perplexity AI to hand over snapshots of its own retrieval database and months of its own record of what its users asked it as evidence in a copyright case. Not a chatbot log a litigant typed into. Not a document review tool counsel used to decide what gets produced. This is Perplexity’s own database, and its own record of what every user asked it and how it answered.

Here’s why this is important, and I mean this for everybody, not just the handful of firms litigating against the Perplexitys and OpenAIs of the world. This order is a plain English education in how a retrieval-based AI system works: what it stores, what it logs, and where the record of its own activity lives. That’s exactly the vocabulary you need the next time any party in any case of yours used an AI tool of any kind, and you need to know what to ask for. More and more companies, even small ones, are building on these tools now, so you’re going to need this vocabulary.

Understanding RAG and UAL

Perplexity runs what it calls an answer engine. A user asks it a question, and instead of generating an answer purely from what the underlying model learned during training, Perplexity goes out and grabs current information to ground its answer. The way it does that is called retrieval-augmented generation, or RAG. Think of RAG as a live lookup step bolted onto the front of a large language model: the system searches an index of web content, called the RAG database, pulls back whatever looks relevant to your question, and hands that retrieved content to the model along with your original prompt. The model then writes its response using that retrieved material as its source.

Britannica and Merriam-Webster, both in the business of dictionaries and encyclopedias, say Perplexity copies their copyrighted content into that RAG index at the input stage, and reproduces or paraphrases it in the answers handed back to users at the output stage. They also allege a trademark claim that Perplexity’s hallucinated content gets falsely attributed to their brands.

The second piece of technology here is the UAL database, short for User Activity Log. Every time someone uses Perplexity, four things get recorded: what the user asked, how the system decided to retrieve and rank content to answer it, the instructions sent to the underlying language model, and the answer the user ultimately got back. That’s Perplexity’s own record of its RAG system in action, at scale, across millions of queries. When Britannica went looking for evidence of infringement, the RAG database and the UAL logs are exactly where that proof lived.

What’s Different About this Fight

This is not the same fight that we saw in the OpenAI copyright litigation, and the difference is in when Perplexity accesses the alleged copyrighted material.

The OpenAI cases are a training-time theory. OpenAI allegedly copied millions of articles into the data used to train its models, baked that content into the model’s own weights, and the trained model can now reproduce it whenever any user, anywhere, asks the right question. That’s why a court ordered OpenAI to produce twenty million ChatGPT logs sampled out of tens of billions it retains, aimed at OpenAI’s fair use defense.

Perplexity’s RAG architecture is a retrieval-time theory. Nobody claims Perplexity trained its model on Britannica’s content. The claim is that Perplexity copies content into its RAG index live, off the open web, and reproduces it whenever a user’s question happens to retrieve it. There’s no training corpus at issue, which is exactly why this discovery is scoped precisely to Britannica’s own copyright registration dates, not a random sample across millions of users. That’s also why a RAG claim is proving easier to litigate than a training claim: you can point to the actual article, the index entry that copied it, and the output that reproduced it, a much more direct line than proving a work influenced a model’s parameters somewhere inside an opaque training run.

The Backstory

This is not Perplexity’s first fight over this exact issue, and the order walks through the history in detail. Dow Jones and the New York Post sued Perplexity in 2024 over the same RAG architecture, in the same District Court. In an August 2025 opinion denying Perplexity’s motion to dismiss, United States District Judge Katherine Polk Failla described the RAG database itself as the mechanism by which the alleged infringement happens, copying protected works in as inputs and reproducing them as outputs. That’s the same framework that United States Magistrate Judge Cave adopts here.

In the Dow Jones case, Perplexity had already produced two RAG snapshots totaling more than 300 terabytes and over 100 terabytes of UAL data, and after losing a motion to compel, seven more months of UAL data on top of that, with Judge Failla calling the burden proportional to the needs of the case. Read those numbers again – a total of more than 400 terabytes of data

Once Britannica sued, Perplexity didn’t start by fighting. It handed over that same baseline data from Dow Jones without a dispute. The fight only started when Britannica came back on May 8, 2026 asking for ten more months of data. Even then, Perplexity’s first move wasn’t a flat refusal, it proposed a cost-sharing arrangement, and only argued undue burden once that didn’t resolve things.

One more note, and this comes from outside the order itself. Perplexity is fighting versions of this same fight against several publishers at once, and a live account of the actual hearing in this case, reported by journalist Matthew Russell Lee for Inner City Press, describes a parallel dispute at the same conference over whether Perplexity has to produce its source code in native format or a converted one. That reporting is paywalled past its opening section, but if you have a subscription to Substack, you can read that here.

The Holding: Relevance, Proportionality, and Costs

In the Court’s own words: while the Additional Data was relevant, requiring Perplexity to produce all of it was not proportional to the needs of the case. Perplexity would produce one additional RAG snapshot and host six more months of additional UAL data for inspection. That’s the ruling in one sentence, decided under Rule 26(b)(1) and Rule 26(b)(2)(C).

On relevance, the relevant window for this kind of technical production tracks the plaintiffs’ copyright registration dates, the same rule that District Judge Failla applied in Dow Jones. Most of Britannica and Merriam-Webster’s registrations became effective between April and December 2025, so that’s where the window starts. Your technical discovery window in a case like this isn’t set by when you filed suit. It’s set by when your copyrights became effective, because this is a retrieval case, not a training case.

On proportionality, Britannica wanted ten months of data. The court granted one additional RAG snapshot and six months of UAL data, not the full ten. Nobody got everything they asked for.

On costs, this is the part I find most interesting. Ordinarily, under the presumption from Oppenheimer Fund v. Sanders, the responding party bears its own production costs. Perplexity asked the court to shift the cost to Britannica under Rule 26(c)(1)(B) and the Zubulake framework, and the court agreed there was good cause, but capped Perplexity’s relief at $6,000 a month, exactly what Britannica had already offered, not the number Perplexity asked for.

Here’s why. Perplexity had told Britannica informally that hosting this data would run $12,000 to $14,000 a month. In its formal declarations, that number jumped to the hundreds of thousands. Given the volumes here, a single day of UAL data alone runs about 20 terabytes, so the bigger number isn’t implausible on its face. But nothing in the record explained why the estimate moved that much, and the court called the gap an “orders of magnitude difference,” viewing the higher figure “with considerable skepticism.” My own read: the $12,000 to $14,000 figure was probably accurate for hosting the data alone, and the cost of extracting it out of deep storage into that hosting environment just wasn’t priced in. That’s a plausible explanation, and Perplexity still lost the argument, because it never put that explanation on the record. You have to document the basis for costs, and counsel should never provide an estimate without knowing everything that goes into it. Likely someone gave a hosting estimate and that was communicated without any requisite follow up. Communication is critical in all of these discovery discussions. You won’t always get it right the first time, and if you aren’t sure you have all the information, say so. 

What to do this week

Learn this vocabulary even if you never litigate against an AI company directly. This order is a working education in how a retrieval-based AI system operates: what it retrieves, what it logs, and where that record lives. You need it any time an AI tool shows up on the other side of a discovery dispute, no matter what the case is actually about.

If you’re litigating against an AI company on a copyright theory, use the registration-date rule to scope your request before you file a motion to compel. Your technical discovery window is set by when the copyrights became effective, not by when you filed suit. Draft to that window, and expect a court to hold you to it if you overreach.

If you’re defending an AI system and fighting proportionality on cost, the lesson isn’t that your number has to stay fixed, it’s that you have to show your work when it changes. Costs change as you learn what a production actually requires. Perplexity’s estimate moved from roughly $13,000 a month to hundreds of thousands without ever explaining why, and that silence is what cost it the bigger number. Put the reason on the record, or a court will assume the worst explanation for you. It’s the same lesson running through James v. Cerebras Systems and Schulte v. LinkedIn: in Cerebras, the parties bought predictability by agreeing to a framework before anyone fought; in Schulte, nobody agreed to anything and lost trying to get the same transparency by motion; here, nobody agreed to anything either, but this time it’s the responding party absorbing the cost of not settling the terms in advance.

What to Watch Next

Perplexity is fighting versions of this fight against multiple publishers at once, in the same District, and the rulings are already citing each other, Dow Jones is informing Britannica, and Britannica likely to inform whatever comes next. The discovery playbook for suing, or defending, a RAG-based AI company on copyright is being built in real time.

This whole line of cases reminds me of the early OpenAI litigation, because it’s the first time we’re seeing this kind of evidence have to be produced at all. My prediction: we’ll see more negotiated protocols like Cerebras in this exact context, parties getting ahead of the fight by agreeing up front on how the real-time costs of pulling this data get shared, rather than litigating it order by order the way Perplexity has now done twice. Watch the Generative AI tag on Minerva26 for what comes next.

Kelly Twigger on EmailKelly Twigger on Linkedin
Kelly Twigger
CEO at Minerva26
Kelly Twigger is a practicing attorney, software developer, consultant, writer, and speaker on issues in electronic discovery, the development and implementation of legal technology, and how to effectively use data in planning for and during litigation.

She is a co-author of Electronic Discovery and Records and Information Management, and host of Case of the Week at Minerva26. As Principal at ESI Attorneys, Kelly manages the boutique eDiscovery and information law firm that acts as operational business partners with its clients to advise law firms, corporations, and municipalities on all areas of electronic information including eDiscovery, privacy, cybersecurity, and information governance.

Kelly is also the CEO of Minerva26. — a SaaS-based practical resource for litigators handling eDiscovery — that curates discovery decisions, rules, and additional content. She is developing an online academy to provide on-demand education for lawyers and legal support professionals to stay abreast of changes in the law and technology that affect litigation and clients’ obligations to respond.

You can reach Kelly at [email protected], join her Facebook community group at Let’s Talk eDiscovery, or connect with her on Twitter @kellytwigger.

Share this article