E-Discovery and Litigation Data

Collection and Processing

Extracted Text

Extracted Text

Extracted Text

Extraction pulls text from native files. Text-based formats yield content directly, while image-based files require optical character recognition. Extraction quality determines what is searchable.

Extracted text is stored alongside the native file and is produced with it in most formats.

Alternative Names:

Text Extraction, Document Text

Why it Matters?

Search operates on extracted text, so extraction failures make documents invisible regardless of their content. This is the mechanism behind most silent gaps in e-discovery: a scanned contract that OCR handled poorly will not surface for any search term it contains. Reviewing extraction quality on a sample, particularly for image-heavy and handwritten material, is a basic quality control step that is frequently skipped.

Frequently asked questions

Why does extraction quality matter?

Which files present extraction problems?