skip to content

Twenty senders carry most of your scanned-invoice volume — what recurring cost decides how long template rules stay the cheapest option?

level: seniorimportance: must knowfreq 56%

answer

  1. cheap to run, expensive to keep
  2. the bill is human, not compute
  3. scales with templates, not documents
  4. layout churn sets the patch rate
  5. the long tail breaks the amortisation

basics

~20 s

Human maintenance time. Per-sender template rules cost almost nothing to run, so the recurring bill is patch work: templates held multiplied by how often each sender changes its layout. It grows with senders covered, not with documents processed.

solid answer

~40 s

The rule baseline is the third option beside training your own extractor and paying per document, and it wins exactly where volume is concentrated and layouts are stable. Twenty templates covering most of the documents cost near-zero compute, run deterministically and can be explained to a finance team field by field. What they cost is upkeep: roughly `templates held * breaks per template per month * hours to patch`. Note what is absent from that product — the document count. Ten times the invoices from the same twenty senders changes nothing; a hundred more senders changes everything, because each tail template is amortised over a handful of documents a month. The baseline stops being cheapest when you must cover the tail, or when layout churn outruns whoever patches it.

code

pseudocode · 14 lines
pseudocode
head = { senders: 20,  docs_per_month: 72000 }   // most of the volume
tail = { senders: 900, docs_per_month: 18000 }   // ~20 docs each

PATCH_HOURS      = 3
BREAKS_PER_MONTH = 0.05        // one layout change per template every ~20 months

for each group in [head, tail]:
    upkeep_hours     = group.senders * BREAKS_PER_MONTH * PATCH_HOURS
    doc_thousands    = group.docs_per_month / 1000
    hours_per_1k     = upkeep_hours / doc_thousands
    report(group, upkeep_hours, hours_per_1k)

// head: 3 upkeep hours/month -> 0.04 hours per 1000 documents
// tail: 135 upkeep hours/month -> 7.5 hours per 1000 documents

go deeper

for a junior

Remember that rule templates are almost free to run and are paid for in people's time instead. The bill arrives when a sender redesigns its invoice and someone has to fix the template.

for a middle

Be able to write the upkeep as templates held times break rate times patch hours, and to point out that the document count is absent from that product.

for a senior

Show the head-versus-tail amortisation with numbers, and budget the silent-failure check as part of upkeep rather than only the patch itself.

for a principal

Argue the split as a portfolio: deterministic rules over the concentrated head, a paid or trained extractor sized for the tail only, and an owner named for the upkeep hours.

## Why the concentrated head is cheap to rule Invoice volume concentrates and invoice layouts do not. If twenty senders account for most of the documents, twenty templates buy most of the coverage. On that head, rules have a cost shape neither of the other two options can match: - **Near-zero marginal cost.** Matching a region of a known layout is trivial work; the compute does not register against the other two options. - **Deterministic and explainable.** A finance team can be shown exactly which region of which layout produced a total. Neither a trained extractor nor a bought one offers that by construction. - **Immediate.** A template for a sender you already receive is hours of work, not a quarter. So the argument against rules is never the price of running them. It is the standing bill for keeping them true. ## The upkeep bill, written out ``` monthly upkeep = templates_held * breaks_per_template_per_month * hours_to_patch * hourly_cost ``` The document count does not appear. That single absence is the whole economics of the baseline, and it cuts both ways: - Volume can grow tenfold on the **same** senders and the upkeep bill does not move — which is why the baseline can beat a per-document price precisely when a per-document price is at its most painful. - Coverage that grows by senders moves the bill linearly, while the documents those senders contribute may be a rounding error. | option | cost per document | what the recurring cost grows with | how it fails | |---|---|---|---| | per-sender template rules | near zero | templates held, times layout churn | silently: one field reads wrong | | your own trained extractor | marginal compute | volume, plus retraining cadence | gradually, across the whole distribution | | hosted per-document extraction | the quoted rate | volume | on the provider's schedule, not yours | ## The long tail is where rules lose Amortisation, not difficulty, is what kills tail coverage. A template for a sender that submits twenty documents a month costs the same to write and patch as one for a sender submitting several thousand. Spread the same upkeep hours over a two-hundredth of the volume and the cost per thousand documents grows by orders of magnitude — the worked figures in the snippet below differ by roughly 180 times between head and tail. That is the honest division of labour between the three options: **rules for the concentrated head, a trained or bought extractor for the tail**, and a per-document contract sized for the tail only, which is a much smaller contract than one covering every document. ## Four things that end the baseline's reign 1. **Coverage has to reach past the head** — a customer commitment to handle every supplier, not most of the volume. 2. **Churn outruns the patch queue.** Upkeep is a rate, not a stock; when breaks arrive faster than they are fixed, accuracy decays continuously and nobody can say by how much. 3. **A field the rules cannot express** — handwritten amendments, a total that is tabular for one sender and free-form for the next. 4. **Nobody owns it.** Upkeep arrives as hours inside somebody else's sprint rather than as an invoice, so it is the easiest cost in the comparison to leave off the sheet and the easiest to stop paying. ## Budget the detection, not just the patch A template that stops matching fails loudly: no field is found and the document lands in review. A template that **still matches the wrong region** fails quietly — a plausible number is written into a ledger and nothing signals it, because a region match produces no confidence score to be low. Any honest upkeep figure therefore includes the standing cost of checking a sample of the head's extractions against the documents, not only the cost of patching once someone notices. That asymmetry is also the reason the head keeps earning its rules after a model arrives for the tail. The senders that carry the money are the ones where a silent wrong value is most expensive, and a deterministic path over them is the cheapest form of certainty on offer.

  • Why does covering the long tail of senders cost so much more per document than the first twenty?
    Because volume concentrates and authoring cost does not. A tail template costs the same hours to write and patch as a head template, but is amortised over a handful of documents a month instead of thousands. The tail is exactly where a trained or bought extractor earns its price.
  • What makes template upkeep so easy to leave off a cost comparison?
    It never arrives as a bill. It is engineer or analyst hours absorbed inside another team's sprint. Count it the way you count a monthly floor: templates held, times the rate each breaks, times the hours to patch one, times a loaded hourly cost.
  • Once an extractor handles the tail, what do the head's rules still earn?
    The cheapest path over the documents that carry the money, with a deterministic explanation of every field, and a much smaller per-document contract — one sized for the tail rather than for every invoice you receive.

Per-sender templates are like keys cut for twenty doors: using them costs nothing, and the bill arrives every time a landlord changes a lock. Holding keys to nine hundred doors you rarely open is what makes it expensive.

saying these in an interview costs you the question

  • Calling the rule baseline free because it needs no accelerator
  • Assuming a template that worked last quarter still parses today's layout
  • Extending per-sender rules into the tail at the same cost per document
  • Believing that rules covering 80% of volume cover 80% of senders
  • Treating a broken template as a loud failure that pages somebody