skip to content

Given a learning curve over training-set size, how do you decide whether 20,000 more labels are worth buying?

level: principalimportance: nice to knowfreq 28%

answer

  1. it is a purchase, not a plot
  2. count doublings, not rows
  3. extrapolate one doubling, not ten
  4. price the drop in hours and money
  5. buy a small tranche and re-plot first

basics

~20 s

Extrapolate the curve one or two doublings at most, price the projected error drop in hours and money, and ask whether it changes a decision the product makes. Buy a small tranche first to test the projection.

solid answer

~50 s

Treat it as a purchase, not a modelling question. Plot the curve with size on a log axis: over a useful range each **doubling** of the data buys a roughly constant absolute drop in error, so gains per row shrink fast — on one document-extraction task, halving the error cost about ten times the rows. Project **one or two doublings ahead only**; anything further is fiction. Then price it: 20,000 expert hours against a projected drop of, say, one and a half points, and ask whether that changes a decision the product makes — if the triage threshold does not move, the accuracy did not matter. If the projection lands near the level the curve is already asymptoting on, stop. De-risk by buying a **small tranche first** and confirming the next point falls where you predicted.

go deeper

for a junior

Recall that gains from extra data shrink as the dataset grows, so a fixed number of new rows is worth much less to a large dataset than to a small one. Think in doublings rather than raw counts.

for a middle

Explain how to read diminishing returns off a log-scaled size axis and why a short projection with a band beats a confident single number extended far past the measured range.

for a senior

Show the operating translation: projected error drop into queue volume, review hours, or automation rate at the actual threshold, plus a staged purchase that validates the projection before the full spend.

for a principal

Own the budget call across options — more rows, better labels, broader coverage, or nothing — and be willing to recommend stopping. Expect to defend the recommendation to people who do not read curves.

## The question is an investment question A learning curve's whole commercial purpose is to answer "should we buy more data?" — and that is a question about **cost per unit of realised value**, not about whether the curve is still descending. It is nearly always still descending a little. The discipline is turning the slope into a number that can sit next to a price. ## Step 1: read the slope on the right axis Put training-set size on a **logarithmic** axis. Empirically, over the practically relevant range, held-out error falls by a roughly constant amount for each doubling of the data and then bends toward its floor. Two consequences follow immediately: - **Gains per row collapse.** Going from 1,000 to 2,000 rows and from 10,000 to 20,000 rows can buy comparable improvements — but the second costs ten times as much labelling. On one scanned-document field-extraction task, halving the error rate required roughly ten times the rows. - **Absolute row counts are the wrong unit.** "We will add 20,000 rows" means something very different at 5,000 existing rows (two doublings) than at 200,000 (a rounding error). Always convert the proposed purchase into doublings before reasoning about it. ## Step 2: extrapolate, but barely Project **one doubling, at most two**, beyond your largest measured point. The curve's shape further out is unknown: it may bend toward the floor much sooner than a straight-line extension suggests, and the floor is exactly where the extension is least reliable. Anyone confidently quoting the error at ten times the current data, from a curve measured over one decade, is quoting an assumption, not a measurement. State the projection as a range and carry the uncertainty band from your replicates into it. ## Step 3: price the projected drop, in the units the business uses Twenty thousand labelled scans is not a number of rows; it is a number of expert hours, a scheduling problem, and an opportunity cost for the people who would otherwise be doing something else. Set that against the projected improvement — say one and a half points of error — and then ask the question that actually decides it: **does one and a half points change any decision the system makes?** - If the model feeds a triage queue with a fixed capacity, the relevant quantity is how many more true cases surface in the top of the queue, not the aggregate error. - If a review threshold is set by a fixed tolerance, the relevant quantity is how much manual review volume falls at that tolerance. - If the improvement is real but no threshold, queue, or automation rate moves, the purchase buys a nicer number in a report. This translation is the part junior analyses skip, and it is the part that makes the recommendation credible to whoever signs off. ## Step 4: check the projection against the floor If the projected error lands close to where the curve appears to be asymptoting — the level set by label ambiguity and by what the current model and features can express — then the purchase is buying the last sliver of an exhausted resource. A projection that comfortably beats the visible plateau, on the other hand, should make you suspicious of the extrapolation rather than optimistic about the purchase. ## Step 5: de-risk with a staged buy The cheapest way to test an extrapolation is to buy a small piece of it. Commission a tranche that adds perhaps a quarter of the proposed volume, re-plot the curve, and check that the new point lands where you predicted. If it does, the remaining purchase is now supported by measurement instead of by a fitted line. If it does not, you have spent a fraction of the budget to avoid a bad decision. Staging also protects against a subtler failure: new data collected under a different protocol, from a different source, or at a different time can fail to behave like the old data at all, and a small tranche exposes that immediately. ## Step 6: compare against other uses of the same budget More rows of the same kind is one option among several with the same price tag: - **Better labels on the cases that matter.** If a meaningful share of remaining errors sit on cases the current labels get wrong, re-examining those may move the achievable floor rather than just walking further down the curve. - **Broader coverage instead of deeper volume.** Rows from a site, device, language, or customer segment the sample barely contains change what the model can generalise to; the existing curve says nothing about them, because it was drawn from the population you already have. - **Nothing at all.** If the projected gain is small, the floor is near, and no decision threshold moves, the honest recommendation is to stop investing in data for this model and say so plainly. ## What a strong answer sounds like "Our curve spans 1k to 20k. Adding 20,000 rows is one doubling; the last doubling bought about two points, and the curve is bending, so I would project one to one and a half points with a wide band. That is roughly this many expert hours. At our current threshold, one and a half points changes the review queue by this much, which is worth about that much per year. I want to buy 5,000 first and confirm the next point before we commit the rest." Numbers, a doubling-based projection, a staged commitment, and an explicit link to a decision the business makes.

  • Why is it dangerous to extrapolate a learning curve several doublings beyond the measured range?
    Because the curve bends toward its floor and you cannot see where. A straight-line extension on a log axis assumes constant gains per doubling, which holds only until the floor starts to bind — precisely the region the extrapolation is being used to predict. The honest form is a short projection with a band, backed by a staged purchase that tests it.
  • The projected error drop is real but small. How do you decide whether it is worth anything?
    Translate it into the decision the system makes. Work out how the queue length, the manual-review volume, or the automation rate changes at the operating threshold, and price that. If no threshold, capacity, or downstream action moves, the improvement is not worth its labelling cost regardless of how statistically real it is.
  • Would 20,000 rows from a new hospital move the curve you drew?
    Not in the way the curve predicts. That curve was drawn by resampling your existing population, so its slope describes more of the same. Rows from a new source change coverage rather than depth: they may help enormously on that source and barely at all on the old one. Treat it as a different question and, if it matters, draw a separate curve.

Drilling a well: each extra hundred metres costs the same to drill but yields less water, and there is a depth below which there is no water at all. You test with a shallow bore before funding the deep one.

saying these in an interview costs you the question

  • Recommends more data because the curve still slopes down
  • Extrapolates many doublings from a short measured range
  • Quotes error improvement without pricing the labelling
  • Ignores that the curve bends toward a floor
  • Assumes new rows from any source behave like the old ones

context