Dmitriy Kononov.
Let’s talkContact

News and analysis

PDF retrieval: preserving meaning in tables and charts

Choose between text and visual PDF retrieval, evaluate source pages and distinguish a correct citation from a correct numerical answer.

AIPublished:

A PDF search can find the right phrase and still produce the wrong answer. A system might select a page mentioning production volume while the actual quantity sits in a table on the next page. It might extract chart labels but lose the relationship between a coloured line and its legend. The business outcome is the ability to verify the evidence behind a decision.

When designing PDF retrieval, separate three outcomes: locating a relevant page, reading its data correctly and producing a justified answer. Success at the first stage does not establish success at the others. This distinction matters when a reported number affects purchasing, a calculation or a production instruction.

What visual retrieval changes

In its September 1, 2026 article, Weaviate describes indexing page images rather than relying exclusively on extracted text. A late-interaction model represents different image regions with multiple vectors and matches parts of a question against parts of a page. The demonstration uses NVIDIA quarterly investor presentations containing charts and financial tables.

This is a supplier demonstration on a particular corpus, rather than an independent test of every kind of PDF. Its useful design insight is that layout contains information. Reducing a page to a sequence of lines can obscure which column a number belongs to, what an axis compares or which qualification appears in a footnote. Image retrieval retains a source that people can inspect. It does not guarantee accurate reading of small text, measurement units or calculations.

Text extraction remains useful. For a contract number, date or exact clause, a text index may be an appropriate starting point. The PDF processing project outline describes a different task: extracting structured fields for another system. Finding a page that discusses a trend and writing every quantity into a database need different acceptance criteria.

A person beside a chart and data cards.
K. Limpitsouni / unDraw · License

Choose by the questions people ask

Collect actual user questions first. Some people search for a specification clause; others compare periods or inspect a product composition. For each question, identify the expected page, relevant area and conditions for a correct answer. A change-over-time question requires the periods being compared, the measurement unit and the direction of change. A material question requires its identifier, quantity and document revision.

Compare text search, page-image retrieval and a combined approach on the same test set. Look at repeated failures alongside compelling examples. Complex tables may benefit from visual representation, while exact codes may benefit from lexical matching. The latter has its own design considerations, covered in hybrid search tokenization.

Avoid collapsing every outcome into one unexplained score. A system can find the correct page but select the wrong row. It can return an accurate number from an outdated document. It can answer for one period when the user requested a comparison. These failures require different repairs. Keep the question, retrieved pages, file revision and generated answer in the evaluation record.

A citation must support inspection

In the Weaviate retrieval-quality overview, the author associates failures in experimental RAG pipelines with irrelevant, truncated or outdated context. These are the author's research findings. The practical implication for a PDF workflow is to evaluate the supplied evidence separately from the fluency of the answer.

For numerical claims, show the document name, revision, page and source area. An answer relying on two pages needs both citations. A citation does not validate the arithmetic. The reviewer should be able to check which values were compared and whether the system confused thousands with millions, financial years with calendar years or forecasts with actual results.

Provide an explicit insufficient-evidence outcome. A document might contain the current value without a previous-period value. Confidently describing growth would then mislead the reader. Return the available material and identify the missing part of the question rather than completing the evidence with an assumption.

A production-document evaluation scenario

Consider an engineer checking whether component quantities changed between two document revisions. This is a proposed scenario, rather than a claim about an implemented service. It relates to generating production orders from engineering documents, but that published case does not establish the use of visual RAG.

Include similar product names, multiple revisions, rotated pages, blurry scans and tables continuing onto another page in the evaluation set. Add questions the corpus cannot answer. Test document permissions separately: a relevant page must not be returned to a user without access. File updates and deletion must affect the index and previously saved references.

A pilot can limit document types and the decision being supported. An operator receives a page and confirms the values before they move downstream. Python data processing addresses that handoff together with cleaning rules and exceptions. Retrieval does not have to create production instructions immediately; establishing an understandable, repeatable verification process is a useful first result.

For each failure, record its practical consequence. Confusing a current revision with a superseded one is more significant than returning a less convenient page containing the same valid information. This helps the team decide where manual review remains necessary and prevents a ranking metric from becoming the only release criterion.

When the added complexity is justified

A visual index requires evaluation of image storage, processing costs, latency and performance across languages. Compare these costs with specific meaning lost by the text-based alternative. If most questions involve concepts and relationships spanning several documents, a different design may be necessary; see knowledge graphs versus vector RAG.

A useful pilot produces a list of questions for which a retrieved page supports a real decision, plus exceptions with an assigned reviewer. That gives the team a reason to expand the system. An attractive answer becomes useful when a person can trace it to the correct document and verify it without repeating an entire archive search.

Sources