Tagged “extraction”
-
Boilerplate is not content
Navigation, footers and disclaimers extract perfectly and index as fact. Repeated across a corpus, they crowd out the pages they were attached to.
-
Extracting text from real documents
A PDF that extracts as scrambled reading order does not raise an error. It produces text, reports success, and poisons everything downstream of it.