Insights

AI & automation

Why our invoice tool asks instead of guessing

Imagine, as a hypothetical, a tool that reads your invoices and gets 95% of the fields right. That sounds good until you ask: which 5% are wrong? If the tool cannot point to them, someone has to check every field again, and much of the time saved is lost. That is why we designed doc2data to flag defined inconsistencies for review rather than fill gaps with guesses. It does not "know" when it is wrong. It checks a set of rules, and it puts every document that breaks one on a list for a person.

Published Snello (Foysal)Drafted with AI tools; Foysal is responsible for the content.All test figures come from our doc2data demo on synthetic documents, measured 2 October 2026. They are not results on real supplier invoices.

A failure mode worth designing for

Two kinds of component do most of the reading in tools like this:

  • OCR (optical character recognition), which turns a scan into text;
  • optionally, AI language models, which can be asked to fill fields that rules could not find.

Both can produce clean-looking output that is wrong. A blurred "8" can be read as a "3"; a total can be taken from the subtotal line; a currency symbol can be dropped. If nothing checks the result, such an error can sit in a spreadsheet until someone reconciles the figures.

So a useful question for any extraction tool is not only "how accurate is it?" but also "what happens when it is wrong?" As a hypothetical comparison: a tool that flags its doubtful documents can save more checking time than a slightly more accurate tool that flags nothing, because you know where to look.

The checks invoices make possible

Invoice numbers are related to each other, and doc2data uses those relations as checks on every document:

  • Net + VAT = total.
  • VAT = rate × net, within rounding. This check assumes a supported invoice with a single VAT rate; several rates on one invoice are not handled yet.
  • Line items add up to the net amount (or to the gross total on receipts).
  • Quantity × unit price = line amount for each line.
  • Dates are plausible, and VAT numbers match their country's format.
  • Required fields are present. A missing currency stays empty; it is never assumed to be euro.

When a check fails, the document goes to a "Needs review" sheet with the reason. Nothing is silently corrected. A mismatch can come from the extraction or from the invoice itself, which is one more reason a person should look. And the reverse matters too: passing these checks does not prove every field is correct.

If the optional AI fallback filled a field, that field is flagged as well. The fallback is off by default. If you switch on the cloud option, it sends up to the first 6,000 characters of the extracted text to an OpenAI-compatible provider you choose, using your own API key. The optional Ollama fallback runs on the same machine (localhost); nothing leaves the computer. Both are built and unit-tested, not yet tested against a live model.

What we measured, on synthetic documents

Measured 2 October 2026 in our Docker build, on documents generated by our own script (invented companies, addresses and VAT numbers):

  • 3 clean sets of 16 documents each: 0 field errors in 189 fields, 0 in 191 and 0 in 178 (synthetic). With 16 documents per set, that does not support any accuracy bound for real invoices.
  • Deliberately degraded scans (110 dpi, 1.5° tilt, blur, heavy compression), 12 documents per set: 131 of 141 fields (92.9%, synthetic) on the set used while developing a fix, and 127 of 139 (91.4%, synthetic) on a set that, per the developer's log, was not looked at during that fix. These small synthetic sets do not establish a reliable difference in general performance; no analysis accounting for document clustering has been performed.
  • Every set contained one invoice with a deliberate arithmetic error. It was flagged in every set.
  • On the two degraded sets, 6 of 11 and 7 of 11 otherwise valid invoices were also flagged, mostly because a total could not be read. These are false alarms, and we accept them as the cost of not passing a wrong number through silently.

The clean held-out set uses the same layouts with new values, and the third clean set was part of the automated tests throughout development. So none of these sets tests unseen layouts.

What the checks cannot catch

Arithmetic catches some wrong numbers, not wrong letters. On one degraded set, a document passed every check with an error in the vendor's name: "S.n.C." instead of "S.n.c.". We call this a silent error:

  • 0 silent errors on the three clean sets (synthetic);
  • 1 on the degraded development set (synthetic);
  • 0 on the degraded held-out set (synthetic).

A single-character mistake in a name or VAT number cannot be detected by sums. doc2data checks VAT numbers for format only, not by checksum and not against the EU's VIES register. Better OCR, a register check where appropriate, and human spot-checks are the defences for those fields.

What these numbers do not tell you

The synthetic invoices come from 2 layouts plus a receipt format, and the rules were written knowing them. Zero errors on the clean synthetic sets shows that the pipeline runs end to end on those layouts. Performance on real supplier documents has not been measured and may differ, including being lower. Other supplier layouts and real-world poor scans are untested; degraded synthetic scans were tested as described above. Multi-page tables, discounts and several VAT rates are also untested. That is why a project would start by measuring on the client's own samples.

Three questions to ask any document automation

  1. How is accuracy measured, and on which documents? Synthetic, public or yours? Were the test documents used while building the tool?
  2. How does the tool show you where it may be wrong? What puts a document on the review list?
  3. Which errors can it not detect? A careful vendor can answer this.

Try it on your invoices

Send 3 sample documents, with sensitive data removed, for a free feasibility check. Three samples give a first look, not a reliable accuracy estimate. Free feasibility check

Sources

  • doc2data evaluation reports, five synthetic sets with field counts, review flags and silent errors; generated 2 October 2026 (commit e4bfc58), rechecked by company-engineer the same day at revision 708168d with identical results. Published results with denominators and limits: doc2data test evidence.
  • company-engineer fact-check, 3 October 2026 (internal).

Bring your workflow into focus.

Send up to 3 samples or describe one process. We reply by email with what is feasible and how we would measure it.

Start a project
SNELLO / TOOLS

This explains the tool; it is not a live AI session. No document is uploaded.

Our working method
Full diagram

Read the project

Discuss your project

Send up to 3 samples or describe one process, the tools you use and the result you need.

snello.contact@gmail.com

Go to the Contact page

Navigation