OCR vs LLM: why document validation needs deterministic rules
What OCR does, what an LLM does, why a model's self-reported confidence is not a probability, and how to chain both with deterministic validators in a cascade.
By Constaia team6 min read
Also in: Español
A few years ago, "reading a document" meant OCR. Today many teams send the photo straight to a multimodal language model and ask for JSON. It works surprisingly well… until the day the model returns an ID number that is nowhere in the photo.
This article explains what each piece does, why an LLM alone is not enough to validate documents, and how to combine them in a cascade where deterministic rules have the final say.
What OCR does
OCR (optical character recognition) turns pixels into text. Modern OCR also returns:
- coordinates for each word or block (bounding boxes),
- per-word confidence,
- structure: headings, tables, signatures.
For example, Mistral OCR 4 returns bounding boxes, block classification and inline confidence scores, and costs $4 per 1,000 pages ($2 with the batch API), according to its June 2026 announcement.
What OCR does not do is understand. It does not know that 12345678Z is a Spanish ID number, that the date in the bottom right is the expiry date, or whether the document is a medical certificate or a prescription.
What an LLM does
An LLM (or a multimodal model) does understand context. It knows that on a Spanish ID card (DNI) the field next to "APELLIDOS" is the surname, that a sports medical certificate usually ends with a fitness statement, and it can return the data in a JSON Schema.
The problem is that an LLM generates plausible text. When the image is poor, it does not leave a blank: it fills in. This Hacker News thread puts it well: just as an LLM can fix OCR mistakes, it can "fix" things that were correct and hallucinate instead. For a summary that is tolerable. For a document number, an expiry date or an IBAN, it is not.
An LLM's self-reported confidence is not a probability
It is tempting to ask the model: "also return your confidence from 0 to 1". That number is not a calibrated probability:
- In Xiong et al., the authors evaluate how LLMs express uncertainty and find that, when verbalising their confidence, they tend to be overconfident.
- Earlier, Guo et al. showed that modern neural networks are poorly calibrated: their confidence does not reflect the real likelihood of being right.
Calibrated means that across all the cases where you say "0.9", you are right about 90% of the time. A "0.95" written by the model is text, not a measurement. In serious validation, confidence has to be built from signals outside the model.
The cascade architecture
The idea is simple: let each piece do what it is good at, and let hard rules decide.
1. Ingestion and quality
Detect the real file type, convert formats, straighten the image and measure quality (size, blur). If the photo is unreadable, reject it before paying for OCR.
2. OCR with coordinates
Text, blocks and bounding boxes. Every later value can be pointed to on the image.
3. Classification
Cheap signals first (is there an MRZ? do certain keywords appear?) and, if needed, an LLM choosing from a closed list of types plus "other".
4. Structured extraction
An LLM fills a JSON Schema from the OCR text and returns, for each field, which OCR block it came from. If it cannot point to a source, the value is suspect.
5. Deterministic validators
Code, not AI. For example:
- Spanish NIF/NIE check letter: the number modulo 23 gives the letter from the official table; for the NIE (foreigner ID number), X/Y/Z become 0/1/2 (Spanish Ministry of the Interior).
- MRZ: ICAO 9303 check digits with 7-3-1 weights and modulo 10 (Wikipedia), plus consistency between the MRZ and the printed data.
- IBAN: ISO 13616, modulo 97 (Wikipedia).
- Dates: expiry against today or the event date, maximum age of a certificate, age against the category.
- Totals: net + VAT = total on an invoice, expected amount on a receipt.
6. Confidence and verdict
Confidence combines the OCR's confidence, the validator results and, when in doubt, agreement with a second pass. From that you decide: valid, invalid or human review.
The point of the cascade is that an LLM mistake runs into a rule. If the model "invents" a digit of an ID number, the check letter no longer matches. If it mixes up the issue and expiry dates, consistency with the MRZ fails. The error does not disappear, but it gets caught and ends up in review instead of passing as valid.
What about building it yourself?
That is a real option. An August 2026 article on Beri.net suggests combining $1.50-per-1,000-pages OCR with a Gemini Flash-class model for extraction, for about $3.20 per 1,000 pages, and compares it with Textract's bundle, which by its numbers costs twenty times more. It also keeps coordinates so a reviewer can see where each value came from.
With the team, volume and time, that can make sense. What such a pipeline does not include out of the box is the expensive part: validators per document type, business rules (valid on a given date, holder matches the registration), confidence calibrated against labelled data, explainable messages, file deletion and GDPR compliance. Make the decision with that list in front of you, not just the price per page.
Honesty: a valid checksum says nothing about authenticity
Deterministic validators catch misreadings and inconsistencies. They do not catch well-made forgeries. The NIF letter and MRZ algorithms are public, and there are MRZ generators with valid check digits built for test data. A fabricated document can pass every checksum.
So:
- "Checksum passed" means "read correctly and consistent", not "authentic".
- Tampering signals (screen photo, editing) are signals only.
- If you need real identity verification, with biometrics and liveness, you need a KYC provider.
How Constaia does it
Constaia follows this cascade: Mistral OCR in the EU, cascaded classification, LLM extraction on AWS Bedrock eu-central-1 citing the source blocks (source.bbox on every field) and open-source deterministic validators (@constaia/validators: NIF/NIE, MRZ, IBAN, dates). We never expose the LLM's raw self-reported confidence: the confidence in the response combines OCR, validators and, when in doubt, a second pass with another model.
curl https://api.constaia.com/v1/analyze \
-H "Authorization: Bearer $CONSTAIA_API_KEY" \
-F file=@dni_valid.jpg \
-F 'options={"expect":"es_dni","checks":{"not_expired":true},"storage":"none","language":"en"}'In the response, checks lists each deterministic validator with passed and a message; fields gives each value with its confidence, whether it was validated and its position on the page. With a ck_test_… key you can try it for free with files such as dni_valid.jpg or blurry.jpg (test mode).
More in verdicts, checks and the /v1/analyze reference.
In short
- OCR reads pixels and gives coordinates; an LLM understands context but may infer instead of read.
- An LLM's self-reported confidence is not a calibrated probability.
- A cascade of OCR → classification → extraction with sources → deterministic validators → verdict turns silent errors into review cases.
- A valid checksum proves consistency, not authenticity.
To see the cascade with your own documents, create a free account: 250 credits a month and test keys that don't use credits.
Sources
- 01Hacker News — LLMs for OCR is super risky…
- 02Xiong et al. — Can LLMs Express Their Uncertainty? (arXiv 2306.13063)
- 03Guo et al. — On Calibration of Modern Neural Networks (arXiv 1706.04599)
- 04Mistral AI — Mistral OCR 4
- 05Beri.net — Document extraction: Textract vs Azure Document Intelligence vs LLM (2026)
- 06Spanish Ministry of the Interior — NIF/NIE check letter calculation
- 07Wikipedia — Machine-readable passport (MRZ)
- 08Wikipedia — International Bank Account Number
- 09GitHub — mrz-passport-generator