skip to content

How do you choose n for self-consistency on a nightly batch of 3,000 tickets under a fixed budget?

level: principalimportance: should knowfreq 48%

answer

  1. plot the curve before picking n
  2. cost is linear, accuracy is not
  3. gains concentrate in the first few samples
  4. spend the large n only on hard items

basics

~20 s

Measure vote accuracy against n on a labelled slice: cost grows linearly with samples while accuracy saturates. Pick the n where the marginal accuracy per dollar stops being worth the cost of the errors it prevents, and reserve large n for hard items.

solid answer

~50 s

There is no universal n — it is a capacity-planning decision, so I would derive it. First build a labelled slice of the ticket corpus and plot vote accuracy at n = 1, 3, 5, 10, 20, 40. Almost always the curve rises steeply over the first handful of samples and flattens well before 40, while token cost rises strictly linearly, so marginal accuracy per dollar falls fast. Then price the error. On a nightly root-cause scoring job, a wrong label costs a misrouted ticket; if that is cheap, n=5 is likely past the knee already, and jumping to n=40 multiplies the bill eightfold for a point or two. Finally, stop paying the same n for every item: sample a small n first and extend only the items whose answers are still split, which concentrates spend on the hard tail. Re-measure after any model change — the knee moves.

code

python · 11 lines
python
ITEMS = 3000            # tickets per nightly run
OUT_TOKENS = 350        # generated tokens per sampled chain
PRICE_PER_MTOK = 3.0    # illustrative output price, USD per million


def batch_cost(n):
    return ITEMS * n * OUT_TOKENS * PRICE_PER_MTOK / 1_000_000


for n in (1, 5, 10, 20, 40):
    print(n, round(batch_cost(n), 2))

go deeper

for a junior

Know that more sampled chains cost proportionally more tokens and that the accuracy benefit tails off, so n is a tradeoff rather than a number to maximize.

for a middle

Be able to describe the measurement: sweep n on a labelled slice, plot vote accuracy against cost, and find the knee. Mention that the shared prompt can be cached while generated tokens scale linearly.

for a senior

Show operational judgement — confidence intervals on the accuracy deltas, staged sampling so easy items finish cheap, batch endpoints for offline jobs, and re-measuring the knee after any model or prompt change.

for a principal

Own it as capacity planning: allocate n by error cost across the request mix, set floors, ceilings and a total job budget, and be willing to conclude that a stronger model retires self-consistency entirely for this workload.

## The two curves that decide n Self-consistency has one lever and two curves. The lever is n, the number of reasoning chains sampled per input. The cost curve is **linear**: n chains generate roughly n times the output tokens, and output tokens usually dominate the bill for chain-of-thought work. The accuracy curve is **saturating**: gains concentrate in the first few samples and then flatten, because the marginal chain is increasingly likely to duplicate a route already in the ensemble. Choosing n is nothing more than finding where those two curves cross in business terms — and the answer is a policy, not a constant, which is why interviewers ask it as a judgement question. ## Build the curve before you argue about it Take a labelled slice of the real corpus — for a nightly job scoring 3,000 telecom trouble tickets for root cause, a few hundred adjudicated tickets is usually enough to see the shape. Score vote accuracy at n = 1, 3, 5, 10, 20, 40, holding temperature and prompt fixed, and repeat each point a few times because the measurement itself is stochastic. Report a confidence interval: a one-point difference between n=20 and n=40 on 200 items is inside the noise, and teams routinely buy that difference for real money. What you will typically see is that n=1 to n=5 captures most of the available lift, n=5 to n=10 adds a little, and beyond that the curve is flat enough that the error bars overlap. That shape is the whole reason the question has a defensible answer. ## Price the errors, not just the tokens The knee of the curve is not the answer by itself — the answer is where marginal accuracy per dollar stops beating the cost of the mistakes it prevents. For 3,000 tickets a night, work it through concretely: if each extra chain costs a fixed amount per ticket, going from n=5 to n=40 multiplies the generation bill by eight. If that buys 1.5 points of root-cause accuracy, you are buying 45 correctly-routed tickets a night. Whether that is a bargain or a waste depends entirely on what a misrouted ticket costs — an engineer's triage minutes, or a missed outage. This also means the answer changes by error class. Asymmetric costs justify asymmetric n: if one root-cause category triggers an expensive field dispatch when wrong, that class can carry a much larger n than the rest of the batch without moving the total much, because it is a small share of volume. ## Stop paying a flat n A single global n is the beginner's configuration. Two refinements pay for themselves: 1. **Staged sampling.** Draw a small n first — say 3 to 5. If the sampled answers already point overwhelmingly one way, stop; if they are split, draw more up to a cap. Easy inputs finish cheap and the budget flows to genuinely hard items. Set both a floor (never fewer than the minimum that gives a meaningful vote) and a ceiling (so one pathological ticket cannot consume the batch budget). 2. **Routing by difficulty.** Cheap signals — ticket length, presence of logs, historical ambiguity of the product area — can pre-sort items into small-n and large-n lanes before the first call. Both are budget-allocation policies, and both need a hard total-spend cap on the job so an unusually hard night cannot overrun. ## Batch economics Offline nightly work has advantages the interactive path does not. The n chains share a long prefix, so prompt caching (where the provider offers it and the requests land inside the cache window) makes the shared instructions and few-shot exemplars nearly free after the first chain, leaving generated tokens as the real variable. Asynchronous batch endpoints usually price below the interactive tier in exchange for a delivery window, which for a nightly job is free money. Both change the arithmetic enough that you should compute the budget with them in place rather than with list interactive pricing. ## Re-measure on every change The knee moves whenever the model, the prompt, or the corpus mix moves. A stronger model often needs a smaller n to reach the same accuracy — sometimes the honest conclusion is that self-consistency has stopped paying at all and one chain from the better model beats twenty from the old one. Bake the n-sweep into the periodic eval so the constant in the config never becomes folklore, and record the date and model version alongside the chosen n. ## What a weak answer looks like "We use 40 because the paper did", "more samples are always better", or quoting a cost number without ever having priced an error. The strong answer names the two curve shapes, the measurement that finds the knee, the business cost that sets the threshold, and the policy that varies n across the batch instead of flattening it.

  • What does the marginal-accuracy-per-dollar calculation look like as n grows?
    Cost per item rises linearly with n while vote accuracy saturates, so the marginal accuracy bought by chain n+1 shrinks monotonically. Divide the accuracy delta between consecutive n values by the cost delta and you get a strictly decreasing series. Stop at the first n where that ratio falls below what an avoided error is worth to the business — that framing survives model changes, whereas a hard-coded n does not.
  • How does prompt caching change the arithmetic for large n?
    The n chains share a long prefix — instructions, few-shot exemplars, the ticket text — so where the provider caches prefixes and the requests land inside the cache window, that portion is billed once at a discount instead of n times. Generated tokens remain fully linear. Effectively caching flattens the fixed part of the cost and makes larger n cheaper than list pricing suggests, but it never makes fan-out free.
  • How would you set n when errors are asymmetric — one wrong root cause triggers an expensive field dispatch?
    Split the batch by error cost, not by convenience. Give the expensive class a large n and a conservative abstain threshold, and leave the cheap majority at the knee of the curve. Because the expensive class is usually a small share of volume, this barely moves the total bill while concentrating reliability where a mistake actually hurts.
  • Your team upgrades to a stronger model. What happens to the chosen n?
    Re-run the sweep. Stronger models typically need fewer samples for the same vote accuracy, and sometimes the single-chain baseline already exceeds the old ensemble — at which point self-consistency has stopped paying and should be removed rather than kept out of habit. Treat n as a measured parameter with a recorded model version, never an inherited constant.

saying these in an interview costs you the question

  • Picking n from a paper without measuring on the target corpus
  • Assuming accuracy scales linearly with the number of samples
  • Ignoring that cost is strictly linear in generated tokens
  • Applying one flat n to every item in the batch
  • Never pricing what an individual wrong answer actually costs

context