Tagged “ocr”
-
Extracting text from real documents
A PDF that extracts as scrambled reading order does not raise an error. It produces text, reports success, and poisons everything downstream of it.
-
Deciding which documents need OCR
Recognition is the most expensive stage in ingestion, and the documents that need it do not announce themselves. Routing has to be inferred per page.