ELAI S.r.l.

Multimodal AI in documents: beyond character recognition

Text, tables and images: where multimodal AI can help document analysis and what to verify before using it.

Multimodal AI in documents: beyond character recognition

A business document contains more information than a text copy reveals. The position of a signature, a note below a table, an arrow in a drawing or a selected checkbox can change a page's meaning. This makes multimodal AI relevant to businesses: it attempts to connect words with visual content. The practical question is where that capability adds value and where traditional automated reading remains easier to control.

Reading characters and interpreting a page

OCR recognizes characters in an image. A document system can then reconstruct rows, columns and fields. A multimodal model can receive images and language instructions to answer questions about the document. Both approaches can coexist: extracted text supports search, while the original page preserves elements that a linear conversion would lose. The choice depends on the question and the quality of the material, rather than the technology's label.

DocVQA by Mathew, Karatzas and Jawahar, introduced in 2020 and published at WACV 2021, studies questions and answers on document images. It is a historical reference for this problem, not a ranking of today's best solutions. Correct answers on a benchmark do not demonstrate that a system can approve business documents without oversight.

An example: comparing a specification and a photograph

Imagine an engineering office receiving a photograph of a component and a manufacturing specification. This is a hypothetical scenario. An initial task might identify the part family and show where its code appears in the specification. The system should return the field, page and relevant crop. An elegant answer without visual evidence would offer little help to the person checking the material.

A second, very different task would declare that a component's measurement meets a tolerance. An ordinary photograph may not contain enough information: perspective, scale and image distortion affect measurement. Document identification, interpretation and dimensional verification should therefore be separated. The model must be able to indicate that an instrument measurement is needed, rather than infer conformity from the general appearance.

Design the output before the input

A useful trial starts with the desired result. Each answer can require the extracted value, unit, document, page and verification status. Missing data should remain explicitly unresolved. This format makes it possible to compare models without confusing writing quality with factual accuracy. It also lets a reviewer correct a single field instead of rewriting an entire summary.

The relationship between original and result must also be preserved. If a page is rotated or cropped before analysis, reference coordinates must still match the document shown to the user. If an attachment is replaced, an answer concerning its previous version should not appear to describe the new one. These product details directly affect confidence in the system.

Which errors to look for

The sample should include poor scans, tables spanning pages, annotations, similar abbreviations and documents with the same layout but different values. One useful check deliberately changes a visual element while preserving the main text: does the answer change when it should? Another removes the page containing the information and checks whether the system admits it cannot answer. These tests distinguish evidence use from mere plausibility.

Cost should cover the whole process: capture, preparation, analysis and review. Analyzing every page at high resolution may be unnecessary for an easily readable field; compressing everything may erase a decisive note. One strategy to test is locating relevant pages first, then analyzing only those in detail. Its value must be checked on the organization's own archive.

The connection with EL-AI

EL-AI presents ELAI Nexus as an environment connecting documents with project context. That organization could provide a starting point for exploring multimodal document assistance. However, the public page does not document automatic visual inspection of components: this example is an application prospect, not an announced feature. An experiment would require authorized documents, a narrow question and agreed verification criteria.

The worthwhile outcome is not a longer description of a document. It is helping someone find and verify the information supporting a decision, while recognizing when the available material is insufficient.

Article prepared with AI assistance and verification of the cited sources. Application examples are hypothetical unless stated otherwise. Sources consulted on September 20, 2026.

Illustrative AI-generated cover; it does not depict actual EL-AI people, premises or installations.