R&D & technical writing
A practical guide to a reproducible feasibility study
"Can this be automated?" can take a long time to answer if the only way to answer it is to start building. A feasibility study tries to answer it first, with a small experiment. A reproducible one lets someone else, or you later, rerun it and check the answer, provided the data, software versions, settings, random seeds and relevant hardware are kept. These are the six steps we use. They need discipline more than special tools.
Step 1: write down one question
A feasibility study answers one question, phrased so the answer can be measured:
- Too vague: "Can AI handle our paperwork?"
- Measurable (example): "Can we extract supplier, date, net, VAT and total from our 5 most common invoice layouts, with fewer than 1 document in 20 needing manual correction?"
The measurable version tells you what data you need and when you are done. Write it at the top of the study document with the date. If the question changes, add the new version below instead of editing the old one.
Step 2: fix the test material early, and keep it apart
Decide what you will judge the result on, ideally before building, and keep it separate:
- a development set you may look at while building;
- a hold-out set you do not look at until the end.
Our doc2data demo shows both the value and the limits of this. On synthetic degraded scans, the set used while developing a deskew fix scored 131 of 141 fields (92.9%). A separate set that, per the developer's log, was not looked at during that fix scored 127 of 139 (91.4%). Two cautions apply:
- Both sets come from the same generated layout families, so they say nothing about unseen supplier layouts.
- These small synthetic sets do not establish a reliable difference in general performance; no analysis accounting for document clustering has been performed. The 1.5-point gap does not measure a "generalisation penalty".
Not every one of our sets was a strict hold-out: our third clean set was part of the automated tests (pass mark 95%) throughout development. Results on material that influenced development can be optimistically biased, which is why the separation matters.
Where real data are sensitive, a synthetic set is a reasonable way to test the pipeline. Say plainly that it does not measure accuracy on real documents.
Step 3: measure a baseline
A result needs a comparison. The baseline is the simplest thing that could work: simple rules or the current manual process for document extraction; for a model, a trivial predictor (always the most common class) and a standard method. If the new approach doesn't beat the baseline by a margin that matters, that is a useful finding.
Step 4: set the pass mark in advance
Decide before you see results what counts as success and failure. Examples:
- a threshold: "at least 90% of fields correct on the hold-out set";
- a stop rule: "if recall stays below 90% after two improvement rounds, stop";
- a cost limit: "each document must take under N seconds on a normal office computer".
Setting the bar after seeing the results makes it easy to fool yourself. In model comparisons the same idea appears as a region of practical equivalence, decided in advance (see One lucky split).
Step 5: make the run repeatable
The minimum record:
- Code under version control, with the exact commit behind each reported number.
- Fixed random seeds, with settings in a configuration file.
- Pinned software versions, for example in a container (we use Docker), so the environment can be rebuilt.
- Raw outputs kept: every per-document or per-fold result, so headline figures can be traced back.
- Automated tests, so later changes don't silently break measured behaviour.
- The platform used. In our OCR tests, results on degraded scans differed slightly between Windows and Linux. We report the container figures as the reference and note the difference.
The Turing Way's guide to reproducible research covers these practices in depth.
Step 6: report honestly, including what failed
A feasibility report can be short:
- The question, as written in step 1.
- The material: what it was, its source, synthetic or real, and what was held out and when.
- The method and baseline.
- The results, dated, against the pass mark set in advance, with denominators.
- The failures: what went wrong, how often, and why, where known.
- The limits: what the study does not show.
- The recommendation: build, change the question, or stop.
Every claim carries a source and date, every number comes from code that was actually run, and assumptions are labelled as assumptions.
What a feasibility study is not
- Not a production prototype. Shortcuts are fine if documented.
- Not a guarantee. It reduces uncertainty about one question.
- Not a regulatory or clinical validation. It can support the qualified person who signs one; it cannot replace them.
How we can help
We offer small, reproducible feasibility studies, written up as reports you can share and rerun: research software, literature reviews and technical documentation. Describe your question in a few lines. Free feasibility check
Sources
- doc2data synthetic evaluation reports (degraded development and hold-out sets, field counts), 2 October 2026, commit e4bfc58, rechecked at 708168d; developer log for the hold-out chronology. Published results with denominators and limits: doc2data test evidence.
- company-senior-researcher statistical review of these figures, 3 October 2026 (internal).
- Bayesian-Classifier-Lab README (seeded splits, configuration files, run manifest), https://github.com/Foysal-A-Al/Bayesian-Classifier-Lab (retrieved 2 October 2026).
- The Turing Way Community, The Turing Way, "Guide for Reproducible Research", https://book.the-turing-way.org/reproducible-research/reproducible-research (retrieved 3 October 2026).
Bring your workflow into focus.
Send up to 3 samples or describe one process. We reply by email with what is feasible and how we would measure it.