Machine Learning Invoice Classification: What the Model Decides, and What It Should Hand Back
Accuracy is the wrong number to ask a vendor for. Invoice classification is four separate decisions with four different error profiles, and the useful model is the one that refuses the invoices it cannot call.
Ken
AI Finance Assistant
Quick Answer: "Invoice classification accuracy" is a single number laid over four separate decisions — which expense category, which cost centre, what tax treatment, and who approves it. Each fails differently, each costs something different when it is wrong, and averaging them hides the two that matter. The number worth asking a vendor for is not accuracy. It is the share of invoices the model will refuse to call, and what happens to those.
One number, four decisions
When a vendor tells you their model codes invoices at 94% accuracy, ask which decision they measured.
Coding an invoice is not one classification. It is at least four, and they are usually produced by the same model in the same pass, which is exactly why they get reported together:
- Expense category — which general ledger account the cost lands in
- Cost centre — which department, project, or entity carries it
- Tax treatment — recoverable or not, which rate, whether withholding applies
- Approver routing — who has to look at it before it is paid
A blended 94% can be 99% on expense category, which is mostly vendor-to-account memory, and 71% on cost centre, which depends on facts that are not written on the invoice at all. The blend is dominated by whichever decision is easiest, because the easy decision is also the one with the most training examples. You are being shown the number the vendor's data made flattering.
Worse, the four failures are not equally expensive. A miscoded expense category is a journal entry. A miscoded approver is a payment that went out without the person who would have recognised it ever seeing it. Both count as one error in the same percentage.
The four decisions, and how each one actually fails
Expense category
This is the decision models genuinely do well, and it is the one every demo shows. The signal is strong: the same vendor, the same line-item language, the same amount range, coded the same way eighty times before.
It fails in two narrow places. The first is a vendor who sells more than one thing — an office supplier who also invoices for a furniture fit-out, a consultancy that bills both advisory and software licences. The model has learned "this vendor equals that account" and applies it to a line item that does not belong there. The second is a genuinely ambiguous chart of accounts, where two codes overlap so completely that your own team disagrees. A model trained on inconsistent history does not resolve that disagreement. It reproduces it, and it reproduces it confidently, because both answers appeared in training.
Cost of being wrong: a correcting journal entry, found at close or not found at all. Low, and self-limiting.
Cost centre
This is where the blended number goes to die, and it is worth understanding why: the correct cost centre is frequently not on the invoice.
A contractor's invoice says what work was done and how much it cost. It does not say which of your three product lines requested it. That fact lives in a purchase order, an email thread, or someone's head. A model can learn a proxy — this vendor usually bills marketing — and the proxy holds until the quarter marketing lends the vendor to sales.
The consequence is that cost centre accuracy is bounded by how much of your spend goes through a purchase order or a requisition. If 40% of your invoices arrive with no upstream record naming a requester, no model reaches high accuracy on that 40%, and no amount of training data will change it. This is not a model quality problem. It is a missing input, and it is the single most useful thing to check before you evaluate any vendor: what fraction of your invoices carry a PO or a named requester? That fraction is roughly the ceiling.
Cost of being wrong: a department's budget looks fine while another one looks blown, and the manager who would have queried the spend never sees it in their numbers. Moderate, and slow to surface.
Tax treatment
Tax is the decision most likely to be quietly wrong, because it depends on facts about the transaction rather than about the vendor: where the service was delivered, whether the vendor is registered, whether the expense is recoverable in your jurisdiction, whether withholding applies to this payment type.
Some of that is on the invoice. Some of it is in the vendor master. Some of it is a judgement about the nature of the expense. A model that has seen the same vendor treated one way ninety times will apply that treatment to the invoice where the underlying facts changed, and there is nothing in the document to signal the change.
Cost of being wrong: an under-claimed recovery, or an under-withheld payment that surfaces in an assessment months later with interest attached. Highest cost per error, lowest visibility.
Approver routing
Routing looks like the simplest decision and it is the one you should let a model touch last.
Routing is usually a rule — amount thresholds, category, cost centre — so a model does not add much over the rule. What it adds is failure. If the model mis-routes, the invoice goes to someone who has no basis for questioning it, and they approve it because it arrived in their queue looking like everything else that arrives in their queue. The control that was supposed to catch an unusual invoice was the fact that a specific person sees invoices of that type. Route around that person and the control is gone, silently, with a complete audit trail showing a correctly approved invoice.
This is the concrete version of the position we take everywhere on this site: extraction and classification should augment the person, not replace their judgement on approvals. Not because models are untrustworthy in general, but because approval is the one step whose entire value is that a particular human with particular context looked at it.
Cost of being wrong: the approval control does not fire, and nothing in the record shows that it did not. Unbounded.
Three ways a classification is wrong
Vendors report errors as a single bucket. Split them into three, because they need different handling:
- Absent — the model produced nothing, or produced a null. This is the good failure. It is visible, it queues, someone deals with it.
- Wrong value — the model produced a plausible but incorrect value in the right field. Cost centre "Marketing" when it should have been "Sales". This is the ordinary failure, and it is what accuracy percentages measure.
- Right value, wrong field — the model read the correct string off the document and put it somewhere it does not belong. The PO number captured as the invoice number. The delivery address's cost centre applied as the billing entity. These are the ones that survive review, because every field looks populated and plausible, and they are the ones almost never reported separately.
A vendor who cannot break their error rate into these three has not looked at their errors closely enough to be useful to you.
A worked sample
Six invoices from a month, showing which single decision went wrong on each and what it cost:
| Invoice | Decision that failed | Proposed | Correct | Kind | What it cost |
|---|---|---|---|---|---|
| Office supplier, 12,400, includes a desk fit-out | Expense category | Office Supplies | Furniture and Fixtures | Wrong value | A journal entry at close |
| Contractor, 18,000, "Q3 platform work" | Cost centre | Engineering | Product | Wrong value | Two budgets misstated for a quarter |
| Overseas SaaS renewal, 9,600 | Tax treatment | Standard recoverable | Reverse charge | Wrong value | An under-declared liability, found in an assessment |
| Utility bill, 2,150 | — | Correct on all four | — | — | Nothing. This is the 80% |
| Legal firm, 46,000, litigation | Approver routing | AP manager (amount rule) | General counsel | Wrong value | An invoice approved by someone with no basis to question it |
| Freight invoice, 3,300, references a PO | Expense category | Freight | Freight | Right value, wrong field | PO number captured into the invoice number field, duplicate check defeated |
Row four is most of your volume. Rows one and two are what an accuracy percentage is measuring. Rows three, five and six are what it is hiding — and rows five and six were both, on the model's own reckoning, high-confidence answers.
The model that refuses is worth more than the model that guesses
A model that is right 94% of the time and silent about which 94% is less useful than a model that is right 88% of the time and tells you which invoices it could not call. The second one converts an unknown error rate into a known queue.
Made concrete, that means two things:
A confidence threshold per decision, not one for the invoice. The threshold for expense category can be low — being wrong costs a journal entry. The threshold for tax treatment and approver routing should be high enough that the model refuses often, because being wrong costs much more. A single invoice-level threshold forces you to price the cheapest decision and the most expensive one identically. Ask whether the system supports per-field thresholds. Many do not, and it is a real limitation rather than a detail.
A refusal queue that is a work queue. "Low confidence" that routes to the same inbox as everything else is not a refusal; it is a relabelled guess, because the reviewer sees a populated field and confirms it. A refusal has to look different from a proposal.
What the refusal queue should look like
Four properties separate a refusal queue from an exceptions folder nobody opens:
- The uncertain field is empty, not pre-filled. If the model's best guess is shown in the field, the reviewer confirms it. Show the guess beside the field as a suggestion the reviewer has to actively take, or do not show it at all.
- It says why it refused. "New vendor, no coding history." "Two candidate cost centres, no PO." "Vendor's tax status changed since the last invoice." A reason turns a queue item into a question with an answer, and the reasons aggregate into a list of what to fix upstream.
- It has an owner and an age. An unowned queue grows until someone bulk-approves it, which is the failure this whole design exists to prevent. Refusals older than a couple of days are a staffing problem, and the queue should make that visible before month end does.
- Resolutions feed back. Every refusal a human resolves is a labelled example on exactly the invoices the model finds hardest. If corrections do not return to training, the refusal rate never falls, and you have bought a permanent manual step.
Sized properly, the queue should be small: on a book of 800 invoices a month with reasonable PO coverage, expect tens of refusals, not hundreds. If the refusal rate does not fall over the first quarter, the model is not learning from resolutions, and that is worth raising with the vendor while you still have leverage.
What to ask a vendor for
Not their accuracy number. Ask for a run on your own invoices, and specify the sample yourself:
- Give them 200 of your own invoices, coded, from the last quarter. Include the awkward ones deliberately — new vendors, split-coded invoices, credit notes, a vendor who sells two unrelated things, anything with a foreign tax treatment.
- Ask for per-decision results. Expense category, cost centre, tax treatment and routing reported separately, each with its own accuracy and its own refusal rate.
- Include invoices where refusing is the correct answer. Salt the sample with five or ten invoices whose correct coding genuinely cannot be determined from the document — a contractor invoice with no PO and two plausible cost centres. A model that confidently codes all of them has told you its confidence score means nothing. This is the single most informative thing in the exercise, and no vendor will suggest it.
- Ask what happens to a refusal. Have them show you the queue, not describe it. Check it against the four properties above.
- Ask how corrections re-enter training, and how long that takes. "Continuously" is not an answer. Per batch, nightly, on a retraining cadence you control — those are answers.
Steps one and three are the whole test. The rest is diligence.
What this does not fix
Classification quality is bounded by the records upstream of it. If your invoices arrive without purchase orders, cost centre accuracy has a ceiling no model clears. If your chart of accounts has fifteen overlapping variants of the same thing, the model learns your ambiguity and returns it with a confidence score attached. If your team codes the same expense two different ways depending on who is on shift, that inconsistency is now the training data.
None of that argues against automating the coding. It argues for fixing the inputs first, because a model is a very fast, very consistent reproduction of whatever your history already contains. The broader sequencing — what to standardise before automating anything — is covered in invoice processing automation.
FAQ
What accuracy should I expect from machine learning invoice classification?
Ask for the number per decision rather than as a blend. Expense category is the easiest and models do genuinely well on vendors with coding history. Cost centre is bounded by how many of your invoices carry a purchase order or a named requester, and no model exceeds that ceiling. Tax treatment depends on facts about the transaction that often are not on the document. A single blended figure is dominated by whichever decision has the most training examples, which is usually the cheapest one to get wrong.
Should a model assign the approver as well as the GL code?
Treat routing as the last thing you automate, and keep a person in the approval itself. Routing is normally a rule, so a model adds little, but a mis-route removes the control silently: the invoice reaches someone with no basis to question it, they approve it, and the audit trail shows a correctly approved invoice. The value of an approval is that a specific person with specific context looked at it, which is precisely the thing a classifier cannot supply.
What is a refusal queue, and how big should it be?
A refusal queue holds the invoices the model declined to code because its confidence fell below the threshold for that field. It differs from an ordinary exception folder in that the uncertain field is left empty rather than pre-filled, each item carries the reason it was refused, the queue has an owner and a visible age, and human resolutions feed back into training. On 800 invoices a month with reasonable PO coverage, expect tens of items rather than hundreds — and expect the count to fall over the first quarter. If it does not, corrections are not reaching the model.
Why does classification accuracy fall for some vendors and not others?
Two reasons dominate. A vendor who sells more than one kind of thing breaks the vendor-to-account association the model relies on, so a fit-out invoice from an office supplier gets coded as office supplies. And a vendor whose invoices arrive without a purchase order gives the model no signal at all for cost centre, leaving it to guess from history — which holds until the pattern changes. Both are input problems rather than model problems, and both are visible before you buy anything.
Related Content
- Invoice Processing Automation — what to standardise upstream before automating the coding
- AP Audit Trail — what the record has to show for an approval to be evidence of anything
- AP Month-End Close Checklist — where miscoded invoices surface, and what they cost by then
- AI Fraud Detection for Invoices — the anomaly side of the same signal
Related Topics
Ready to automate your invoices?
See how Ken can extract invoice data in seconds, right in Slack. No credit card required.