E-Discovery and Litigation Data
Collection and Processing
Extraction pulls text from native files. Text-based formats yield content directly, while image-based files require optical character recognition. Extraction quality determines what is searchable.
Extracted text is stored alongside the native file and is produced with it in most formats.
Alternative Names:
Text Extraction, Document Text
Why it Matters?
Search operates on extracted text, so extraction failures make documents invisible regardless of their content. This is the mechanism behind most silent gaps in e-discovery: a scanned contract that OCR handled poorly will not surface for any search term it contains. Reviewing extraction quality on a sample, particularly for image-heavy and handwritten material, is a basic quality control step that is frequently skipped.
Frequently Confused with
Related terms
Frequently asked questions
Why does extraction quality matter?
Which files present extraction problems?


