Invoice Capture Software: How the Categories Actually Differ, and Which One Your Volume Justifies
Template OCR, trained extraction models, e-invoicing networks, and supplier portals differ by what they can ingest and what accuracy costs to maintain. How to classify your own invoice mix and rule one out.
Ken
AI Finance Assistant
Every invoice capture product shows the same demo. A clean PDF drops into a queue, fields populate, a confidence badge turns green. Four vendors will show you that demo this month and you will learn nothing that separates them, because none of them is going to open the demo with a photograph of a crumpled delivery note taken in a warehouse.
The categories underneath those demos are genuinely different pieces of technology with different failure modes and different ongoing costs. Separating them takes about an hour and rules out at least one before you take a call.
The four categories
Template-based OCR reads text from a document and pulls fields from positions defined in a per-layout template. Someone configures where the invoice number sits on a Siemens invoice, and every subsequent Siemens invoice is read from that map.
Trained extraction models learn what an invoice number looks like across many layouts rather than where it sits on one. They handle a supplier they have never seen, return a confidence score per field, and improve when corrections are fed back.
E-invoicing networks do not capture anything. The supplier submits structured data — PEPPOL, UBL, EDI, or a national clearance platform — and it arrives as fields. There is no document to read, so there is no extraction error.
Supplier portals move the keying to the supplier. They log in and enter the invoice against a purchase order. Also no extraction, and also no software reading anything.
The last two are frequently left out of comparisons because they are not the same kind of product, which is exactly why they are worth putting on the same page. Two of the four categories solve invoice capture by removing the document.
What each one actually ingests
This is the question the demo answers dishonestly, because the demo picks the input.
| Reads cleanly | Struggles | Cannot ingest | |
|---|---|---|---|
| Template OCR | Any layout you have templated | A templated layout the supplier redesigned | Any supplier without a template |
| Trained models | Machine-readable PDFs, most scans | Photographs, handwriting, dense multi-page line items | Nothing outright — it returns low confidence instead |
| E-invoicing network | Anything from a connected supplier | Nothing | Anything from a supplier not on the network |
| Supplier portal | Whatever the supplier types | Suppliers who will not use it | Anything from a supplier with the leverage to refuse |
The row that matters is the third column, because that is where your exceptions come from and no product page describes it. A template system does not degrade gracefully on an unknown supplier; it fails completely and routes to a person. A trained model degrades gracefully, which sounds better and is worse in one specific way: it produces a plausible wrong value with a middling confidence score, and somebody has to decide the threshold at which that gets reviewed.
That threshold is the whole product. Set it high and your touchless rate collapses. Set it low and errors reach the ledger. It is a business decision presented to you as a configuration field, usually during implementation, usually decided by whoever is in the room.
What accuracy costs to maintain
Every category has an ongoing cost, and it is a different kind of cost in each.
Template OCR: one template per layout, forever
The cost is not the initial configuration. It is that every supplier who redesigns their invoice silently breaks their template, and you find out through an exception queue rather than an alert. Someone owns a growing library of templates and the maintenance never ends, because your supplier list never stops changing.
The number that determines whether this is manageable is distinct supplier layouts, not invoice volume — and almost every buying conversation uses invoice volume. A team processing 5,000 invoices a month from 40 suppliers is a comfortable fit. A team processing 500 invoices from 400 suppliers is not, and it is the second team that is usually sold the template product because their volume looks small.
Trained models: the correction loop is the product
Accuracy is maintained by corrections flowing back into training. That requires two things most implementations do not set up: corrections captured as structured data rather than as an overwritten field, and enough correction volume for the feedback to mean anything.
Below a few hundred invoices a month there is not enough signal for the loop to do much, and the model you get is the model the vendor trained on everyone else's invoices. That is often fine. It is worth knowing, because it means the "it learns your data" pitch does not apply at your volume, and you are buying a general model.
E-invoicing networks: supplier onboarding is the cost
Extraction accuracy is not a line item, because there is no extraction. The cost is entirely commercial and relational: getting suppliers connected. That is a project measured in supplier conversations, and the ones who resist hardest are usually small suppliers with the least capacity and the most invoices.
Where a national mandate applies — India's e-invoicing regime above a turnover threshold, Italy's clearance platform, the EU's move toward structured invoicing — a share of your suppliers is already producing structured invoices for tax reasons. Finding out what share is a free exercise and it changes the maths considerably.
Supplier portals: the cost is leverage
They work in exact proportion to how much your suppliers need you. A large buyer with fragmented suppliers can mandate one. A mid-market buyer purchasing from larger firms cannot, and will end up with a portal used by the suppliers who were never the problem.
Classify your own invoice mix first
Do this before any demo. It takes an hour and it eliminates categories.
Pull every invoice from one recent month and sort into four piles:
- Already structured — arriving through a network, an EDI feed, or a national platform.
- Machine-readable PDFs — text is selectable. Generated by a supplier's accounting system.
- Scans and images — text is not selectable. Photographs, faxes, scanned paper.
- Not invoice-shaped — statements covering many invoices, spreadsheets of line items, charges in an email body, portal-only invoices someone downloads by hand.
Then count distinct suppliers, and separately count the suppliers responsible for the top 80% of invoice volume.
Those five numbers decide more than any feature list:
- Pile 4 is large. No capture category handles this well. It is a supplier-communication problem and buying software will not touch it. Fix the worst offenders by asking them to send something else.
- Pile 3 is large. Rule out template OCR. Scan quality varies enough that positional extraction is fragile, and a trained model is the only category that copes.
- Pile 2 dominates and distinct suppliers are few. Template OCR is genuinely sufficient and much cheaper. This is the case where the expensive category adds nothing.
- Distinct suppliers are many relative to volume. Rule out template OCR regardless of anything else.
- The top 80% comes from a handful of suppliers. Look hard at an e-invoicing network or a portal before buying capture at all. Removing the document is strictly better than reading it, and a small number of supplier conversations covers most of your volume.
That last case is the one most often missed, because capture software is easier to buy than supplier conversations are to have. It is also the only option with no ongoing accuracy cost at all.
The volume thresholds
Each category stops making sense at a different point, and the trigger is different in each case.
| Category | Works well when | Stops making sense when |
|---|---|---|
| Template OCR | Under roughly 100 distinct supplier layouts, mostly machine-readable PDFs | Supplier churn means templates are being built faster than they are being used |
| Trained models | Varied layouts, mixed quality, hundreds of invoices a month or more | Volume is low enough that a person keying invoices costs less than the subscription and the review loop |
| E-invoicing network | Concentrated supplier base, or a mandate already in force | Suppliers are fragmented and small, with no leverage to bring them on |
| Supplier portal | You are the larger party in most supplier relationships | You are not |
The lower bound on trained extraction is worth stating plainly, because nobody selling it will. Under about a hundred invoices a month, careful manual entry with a three-way match behind it is cheaper and more accurate than any of this. The threshold where automation wins is real, and it is higher than the demos imply.
What to test instead of accuracy
The standard question is what accuracy the vendor achieves, and every answer is a number from a benchmark that does not resemble your invoices. Replace it.
Send them your worst fifty. Not a representative sample — the fifty that cause the most trouble today. The photographs, the multi-page line item schedules, the statements. Ask for field-level output on those, not a percentage.
Ask what happens below the confidence threshold. Who reviews it, in what interface, and does the correction go anywhere or just get typed over?
Ask how a new supplier is handled on day one. For template systems this is the whole cost of ownership; for models it should be nothing. The answer separates the categories faster than any feature comparison.
Ask what breaks when a supplier redesigns their invoice. A vendor who has run this in production has a specific answer. A vendor who has not will say the AI handles it.
Ask who pays for the review work. If the product's touchless rate is achieved by routing 30% of invoices to your team, the labour did not disappear — it moved to a column nobody priced.
Where capture ends
Capture is the smallest part of the problem and it absorbs most of the evaluation effort, because it is the part that demos well. The fields still have to be coded, matched, routed, and approved, and the exceptions generated downstream are where the actual time goes.
A team choosing between two capture products that both handle their mix adequately should stop comparing them and start comparing what happens after extraction. The difference in outcomes between the two capture engines is usually smaller than the difference between a good and a bad exception workflow sitting behind either one.
For the specific comparison of positional OCR against trained models — the first two categories here — see OCR vs AI invoice processing, which goes deeper on the extraction mechanics. For what the capture layer feeds, see invoice processing automation.
FAQ
What is the difference between invoice capture and invoice OCR?
OCR is one technique used inside invoice capture — converting an image into text. Capture is the whole job of getting invoice data into your system, and two of the four viable approaches involve no OCR at all. An e-invoicing network delivers structured fields directly and a supplier portal has the supplier type them. Treating capture and OCR as synonyms narrows the options to the two that read documents, which are also the two with ongoing accuracy costs.
How many invoices do you need before capture software pays for itself?
Below roughly a hundred invoices a month, manual entry is usually cheaper once you include configuration, the review of low-confidence extractions, and the licence. Between one hundred and a few hundred, a trained extraction model starts to win but the learning loop contributes little because there is not enough correction volume. Above that, the case is straightforward. The threshold moves with supplier count too — many suppliers at low volume is harder than few suppliers at high volume, and costs more to automate.
Can invoice capture handle scanned images and photographs?
Trained extraction models handle scans reasonably and photographs unreliably; template-based OCR handles neither dependably, because a photograph's geometry defeats positional extraction. If a meaningful share of your invoices arrive as photographs, the realistic fix is upstream — ask those suppliers to email a PDF instead. That is one conversation per supplier against an indefinite exception queue, and it is the cheaper of the two.
Is an e-invoicing network better than invoice capture software?
Where you can get suppliers onto it, yes — there is no document to read, so there is no extraction error and no maintenance cost. The limit is not technical. It is that connecting suppliers is a commercial exercise, and the suppliers hardest to connect are usually the small ones sending the messiest invoices. Most teams end up with a network covering the concentrated part of their spend and a capture product handling the long tail.
Related Topics
Ready to automate your invoices?
See how Ken can extract invoice data in seconds, right in Slack. No credit card required.