Insights

Internal tools & prototypes

Running a document pipeline locally, without per-page API charges

Some document-AI services charge per page and process documents on their own servers. Amazon Textract, for example, publishes per-page pricing. For a small office, that can mean a recurring bill and documents leaving the building. A local pipeline is one alternative worth evaluating. This is how doc2data works in its default local mode: it runs on your own computer, offline, with no per-page API charge. It also covers what changes if you switch on the optional cloud fallback, and its running costs and limits.

Published Snello (Foysal)Drafted with AI tools; Foysal is responsible for the content.Based on our doc2data demo. Figures measured 2 October 2026 on synthetic documents.

The design rule: simplest reliable method first

We build in layers, from the most predictable to the most flexible:

  1. read text already embedded in the file;
  2. use OCR only for image-only pages;
  3. extract fields with rules;
  4. optionally, ask an AI model for fields the rules missed.

Each layer handles only what the previous one could not. That keeps costs and failures easier to locate.

Step 1: read the embedded text layer

Many invoices produced by accounting software are PDFs with an embedded text layer. pdfplumber, an open-source Python library, extracts that text without OCR. It works best on machine-generated PDFs. It is not a guarantee of correct invoice fields: layout parsing can still fail, which is why the later checks exist. How large a share of a given office's invoices have a usable text layer has to be measured on its own documents; we have no data on that.

Step 2: OCR for image-only pages, on your machine

Image-only scans have no text layer, so they need OCR (optical character recognition). Some scanned PDFs already contain an OCR text layer from the scanner. We use RapidOCR on ONNX Runtime, which runs the recognition model on your own processor, with no cloud account and no per-page fee.

Choosing the model mattered. Our first OCR model had no "€" or "£" in its character set and dropped the spaces between words. A Latin-script model (PP-OCRv5 Latin) fixed that; another candidate read "€" as "£", so we did not use it. We also made line rebuilding follow the page's tilt, which recovered most line items on our skewed synthetic scans.

What we have tested is synthetic rendered scans, clean and degraded. Phone photos have not been tested. OCR results on degraded scans also differed slightly between Windows and Linux in our runs. During feasibility work we can benchmark on the machine you intend to use.

Step 3: rules, then checks

Fields are extracted with label rules and patterns for Italian, English, German and French invoices: "Totale", "Total", "Gesamtbetrag", EU number and date formats, and VAT-number patterns. Rules are predictable and easy to test.

Then the checks run: net + VAT = total, lines add up, plausible dates, valid VAT-number format. A failed check puts the document on a "Needs review" list. The checks flag some inconsistencies; passing them does not prove a document is correct. (More in Why our invoice tool asks instead of guessing.)

Step 4: the optional AI fallback, and what it sends

For fields the rules could not find, doc2data has an optional AI fallback. It is off by default. Be clear about what it does when switched on:

  • Cloud option: it asks for the missing fields, but its prompt includes up to the first 6,000 characters of the extracted invoice text, sent to an OpenAI-compatible provider you choose, using your own API key. That provider receives the text, and may charge for it.
  • Local option (Ollama): The optional Ollama fallback runs on the same machine (localhost); nothing leaves the computer.
  • Every AI-filled field is flagged for checking.
  • Both options are built and unit-tested with a simulated backend; neither has been tested against a live model yet.

Requirements and limits

  • No per-page API charge in default mode. Local processing still uses hardware and electricity, plus a one-time setup and someone's time on the review list. The optional cloud fallback may cost money at your provider.
  • Synthetic results (2 October 2026): 0 field errors in 189, 191 and 178 fields on 3 clean sets of 16 documents; 131 of 141 (92.9%) and 127 of 139 (91.4%) fields correct on deliberately degraded scans of 12 documents each. Performance on real supplier documents has not been measured and may differ, including being lower.
  • Not handled yet: several VAT rates on one invoice, line tables across pages, credit notes, descriptions that wrap onto two lines. VAT numbers are format-checked only.

Packaging for repeatability

We package doc2data with Docker. The image bundles Python, the specified library versions and the OCR model, and the 21 automated tests run during the build (all passing on 2 October 2026, build at commit 708168d). Exact reruns require preserving the tested image; rebuilding later can change the base image and system packages. Docker improves repeatability, but it does not guarantee identical behaviour on every machine or forever. So far the image has been built and run only on Windows 11 with Docker Desktop (x86-64). Other platforms need their own check, and the image and dependency versions are kept so a run can be repeated.

When a hosted service may be worth evaluating

A local pipeline is not always the right choice. Hosted document services are candidates to evaluate if you:

  • process very large and varied volumes;
  • need handwriting or difficult photographs read;
  • have no computer to run the pipeline, or nobody to maintain it.

We have not benchmarked any hosted service against doc2data, so we make no accuracy or speed comparison.

Could your documents be processed locally?

Send 3 samples, with sensitive data removed, for a free feasibility check. Emailing samples means they travel by email; see our Privacy page. Free feasibility check

Sources

  • doc2data evaluation reports and code (pipeline and LLM fallback, including the 6,000-character prompt limit), commit e4bfc58, 2 October 2026; rechecked at 708168d. Published results with denominators and limits: doc2data test evidence.
  • company-engineer fact-check, 3 October 2026 (internal).
  • pdfplumber, https://github.com/jsvine/pdfplumber (retrieved 3 October 2026).
  • RapidOCR, https://github.com/RapidAI/RapidOCR (retrieved 3 October 2026).
  • ONNX Runtime, https://onnxruntime.ai/ (retrieved 3 October 2026).
  • Docker documentation, "Multi-platform builds", https://docs.docker.com/build/building/multi-platform/ (retrieved 3 October 2026).
  • Amazon Textract pricing, https://aws.amazon.com/textract/pricing/ (retrieved 3 October 2026), as an example of per-page pricing only.

Bring your workflow into focus.

Send up to 3 samples or describe one process. We reply by email with what is feasible and how we would measure it.

Start a project
SNELLO / TOOLS

This explains the tool; it is not a live AI session. No document is uploaded.

Our working method
Full diagram

Read the project

Discuss your project

Send up to 3 samples or describe one process, the tools you use and the result you need.

snello.contact@gmail.com

Go to the Contact page

Navigation