AI document processing with consensus checking
Reference client: Logistics and trading group in Tyrol, Austria · ~100 employees · 4 sites
Customer data in copy and images has been neutralised.
Starting point
Transport orders and purchase orders arrive as PDFs by email: each structured differently, and each used to be typed into the system by hand. Typos in quantities, prices or loading points included.
Classic OCR fails on the variety: every sender formats differently, and a misread price is more expensive than none at all.
The hard part
Language models are confidently wrong. They report no doubt and return no error; they return a number. On a transport order that number is a price, and a silently misread price costs more than all the retyping ever did.
So the decisive problem is not the reading. It is knowing when the reading cannot be trusted, and that question a single model fundamentally cannot answer: it has no independent yardstick for itself.
Then there is the variety of senders. Every company builds its form differently, many documents are scans, some skewed, some photographs of a printout. Classic text recognition breaks on that, and the worse the source, the more convincingly a model invents what is missing.
Solution
The solution reads every document three times: a vision model sees the PDF the way a person does, a text model reads the extracted content, and a deterministic parser checks tables and anchor terms.
A reconciliation merges the three results: if at least two of the paths agree, the field counts as reliable. Arithmetic checks (quantity times unit price equals total) expose reading errors, and uncertain fields are flagged rather than guessed.
The verified result goes straight into the TMS/ERP import. A person reviews the flagged spots instead of retyping everything. It all runs on the company's own GPUs; no document leaves the premises.
How it is built
Every document runs through three deliberately different paths. A vision model sees the page the way a person does: layout, table rules, position. A text model reads the extracted content. A deterministic parser looks for anchor terms and table structures with no AI at all.
That the paths differ is the entire point. Two readers built on the same text recognition can agree and still be wrong together, because they inherit the same fault; only a path that looks at the page as an image can notice an error made by the text recognition in the first place.
A reconciliation merges the three results field by field. If at least two agree, the field counts as reliable. Otherwise it is flagged rather than guessed. Added to that are arithmetic checks: quantity times unit price against the total, line sums against the final amount.
Every document has a hard time budget. A model that talks itself into a loop is aborted and the case handed to manual review. It is the least spectacular safeguard in the whole design, and the most effective.
Day-to-day operation
The human stays in the process, only at a different point. They review the flagged fields instead of typing everything, and with that the work shifts from capture to checking: precisely the activity where a person is genuinely better than the system, and precisely the one that used to get squeezed out by the typing.
New sender formats need no programming. They announce themselves, because the number of flagged fields rises for them. That doubles as the ongoing quality measurement.
The entire process runs on the company's own GPUs, which means no document, no price and no customer address leaves the premises at any point in the chain.
Outcome
The speed-tuned reading path captures a transport order in about half a minute, hard-capped at three minutes per document, with per-field confidence instead of blind trust.
Retyping becomes the exception: the effort per order shrinks to a short review, and reading errors surface in the reconciliation. Not later on the invoice.
What we learned
The stronger model was the worse choice. A model that emits its reasoning does give good answers but it writes that reasoning into the answer and destroys the expected data format. Structured extraction needs a model that answers and not one that thinks. That was our most expensive mistake in this project, and no datasheet would have told us.
The most effective check in the whole system is not AI. It is a multiplication. Quantity times unit price against the total finds more reading errors than any model's confidence score, because it makes a statement about reality rather than about the reader's self-assurance.
The reconciliation needs fields defined beforehand. Three readers writing into different fields produce no majority, just three opinions. The target schema is the precondition of the consensus check and not a by-product of it.
Other projects we run

Enterprise AI stack, built in-house
Four GPU hosts, one model family for every chat lane, automatic failover and an honest look at five weeks of rebuilding.

Company AI portal on its own GPUs
ChatGPT comfort in-house: 5 models, 32 tools, ~90 user accounts. Chats and documents stay in-house.
Similar problem?
Tell us what you're planning, a short call clarifies whether it pays off.
senn-tech