Technical GUIDE

Recall, Precision and Elusion Testing in Document Review

Recall estimates how many responsive documents a review found; precision estimates how many documents labeled responsive are truly responsive.

  • 3 min read
  • Last updated
On this page3 min read
  1. Overview
  2. Deep Dive
  3. Strategic Impact
  4. The Future of Recall, Precision and Elusion Testing in Document Review
  5. Real-World Implementation
  6. Risks & Guardrails
  7. Implementation Roadmap
  8. Keep Exploring
  9. Frequently asked questions

Overview

Elusion examines responsive documents hidden in the set classified nonresponsive. Sampling supports estimates, not a guarantee of perfect production.

Deep Dive

In technology-assisted document review, recall is the proportion of all truly responsive documents that the process identifies. Precision is the proportion of documents identified as responsive that are truly responsive. A confusion matrix separates true positives, false positives, false negatives, and true negatives against a defined reference classification. These measures answer different questions: high recall can come with lower precision, and vice versa.

Elusion estimates how many responsive documents are present in the population the system classified as nonresponsive. It can be easier to sample than the entire responsive population, but low elusion does not always establish high recall, especially when responsive documents are rare. EDRM’s statistical sampling guide warns that inference depends on prevalence and a valid sample, and that derived recall estimates do not automatically inherit the confidence level of their component estimates.

Statistical random sampling can support quantitative estimates when the population, coding criteria, sample design, and uncertainty are documented. Judgmental review may help find examples or guide training, but it does not provide the same statistical conclusions. The result is an estimate tied to its sample and assumptions, not proof that every responsive document was found. Teams should define responsiveness with counsel, compare reviewers to a defensible reference standard, and explain limitations in any validation report. The specific protocol should fit the matter and any agreements or court orders.

Strategic Impact

Cost and budget

Architecture decisions drive performance and operating cost for years.

Clearer decisions

Technical education helps teams choose the right stack, not just the newest one.

Quality control

Better engineering choices reduce reliability incidents in production.

The Future of Recall, Precision and Elusion Testing in Document Review

As review platforms add analytics and generative features, defensible validation still depends on a clear population, documented criteria, and appropriate sampling. Teams may refine workflows as they learn more about the collection. Measurements help assess risk and workload, but they do not replace legal judgment or matter-specific agreements. Courts and parties may choose different protocols based on collection size, claims, and production needs. As models and search tools change, teams should explain their design and test behavior on representative data. Transparent estimates help discuss risk, but no single threshold resolves every legal dispute.

Real-World Implementation

Counsel estimates whether the responsive set may contain important uncoded material.

A team samples documents coded nonresponsive to estimate elusion.

A reviewer compares machine coding with a documented human-coded reference set.

Parties agree on sampling method, confidence goals, and the review population.

Risks & Guardrails

  • Optimizing one benchmark can hide broader system weaknesses.

  • Infrastructure and maintenance costs are often underestimated.

  • Security and observability gaps can grow as systems become more complex.

Implementation Roadmap

  1. Define latency, quality, and cost targets before implementation.

  2. Benchmark under realistic load and data conditions.

  3. Instrument monitoring for errors, drift, and user impact.

  4. Prepare rollback and incident response paths before scaling.

Keep Exploring

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the Recall, Precision and Elusion Testing in Document Review quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

Start quiz

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

Frequently asked questions

What is Recall, Precision and Elusion Testing in Document Review?

Recall estimates how many responsive documents a review found; precision estimates how many documents labeled responsive are truly responsive. Elusion examines responsive documents hidden in the set classified nonresponsive. Sampling supports estimates, not a guarantee of perfect production.

A sample shows 80 responsive documents out of 100 documents coded responsive. Which measure is 80%?

Precision asks what share of identified responsive documents are actually responsive.

A responsive document was classified nonresponsive. Which cell does it occupy in a confusion matrix?

The item is responsive in the reference coding but missed by the review.

Which sample most directly estimates elusion?

Elusion concerns responsive items among the nonresponsive population.

Why can a low elusion estimate fail to establish high recall when responsiveness is rare?

EDRM cautions that low elusion may be misleading when responsive prevalence is very low.

Which sampling approach supports a statistical claim about a population?

Statistical inference requires an appropriate selection design, not convenience selection.