A captioning service widened its accelerator batch window from 20 ms to 40 ms; p95 rose 40 ms while throughput rose 4% - what does that say?
answer
- read marginal gain, not absolute throughput
- past the knee the window charges twice
- wait grows and the pass grows
- flat throughput means saturated, not idle
- the knee moves with the arrival rate
basics
~20 sIt says the service is past the knee of its throughput-versus-latency curve. The fixed per-pass cost is already amortised, so a wider batch grows pass time roughly in step with batch size while buying almost no extra chunks a second. Revert the window.
solid answer
~40 sBelow the knee, each extra millisecond of collection window gathers chunks the accelerator absorbs nearly free, so per-chunk cost drops fast. Past it the device's parallel width is saturated and pass time grows about linearly with batch size, so the wider window costs **twice** - once as queueing delay waiting for the batch to close, once as a longer pass that every chunk in it sits through. A 40 ms p95 rise from a 20 ms window increase is exactly that double charge: roughly 20 ms more wait plus roughly 20 ms more pass. The 4% throughput gain is the flat part of the curve. Go back to 20 ms, and re-sweep the window whenever the arrival rate changes materially, because the knee moves with it.
code
pseudocode · 15 lineson chunk_arrival(chunk):
queue.append(chunk)
if queue.size == 1:
arm_flush_timer(WINDOW_MS) // armed by the batch's FIRST chunk
if queue.size >= MAX_BATCH:
cancel_flush_timer()
submit(queue.take_all()) // size cap closed it early
on flush_timer_fires():
if queue.size > 0:
submit(queue.take_all()) // window closed it
// the earliest chunk's added wait is bounded by WINDOW_MS;
// every chunk in the batch then pays the pass time, which
// grows with queue.size once the device is saturatedgo deeper
Recall that throughput against latency is a curve with a knee, and that after the knee extra batch size buys rate you can barely measure while the tail keeps climbing.
Explain both charges in the tail: more queueing delay while the batch fills, and a longer pass once the device is saturated, paid by every chunk in the batch rather than only the earliest.
Diagnose from a measured sweep - marginal throughput per added millisecond, a device-bound plateau against an arrival-bound one - and revert or re-tune against the deadline share this stage holds.
Own the rule rather than the value: who re-sweeps, on what traffic change, and how much of the end-to-end tail the organisation is willing to spend for a few per cent of rate.
## The curve, measured The only honest way to find a batch window is to sweep it at a fixed arrival rate and record both axes. A live captioning service scoring fixed-length audio chunks might measure: | collection window | mean batch | chunks a second | p95 caption latency | |---|---|---|---| | 10 ms | 9 | 620 | 55 ms | | 20 ms | 17 | 780 | 72 ms | | 40 ms | 31 | 810 | 112 ms | | 80 ms | 65 | 815 | 172 ms | The rows are self-consistent: with the batch-size cap set high enough not to bind, the window is what limits the batch, so mean batch is roughly throughput times window (780 x 0.020 is about 17; 810 x 0.040 is about 31). Read the marginal gains rather than the absolute numbers: - 10 to 20 ms: **+26%** throughput for **+17 ms** p95 - a good trade; - 20 to 40 ms: **+4%** throughput for **+40 ms** p95 - the trade has inverted; - 40 to 80 ms: **+0.6%** throughput for **+60 ms** p95 - pure loss. The **knee** is the window past which marginal throughput per added millisecond collapses. Here it sits at about 20 ms, and the reported move went straight past it. ## Why the far side of the knee charges twice On the near side, the accelerator's fixed per-pass cost - crossing the device boundary, starting the work, reading the model's weights - is being divided among more and more chunks, and the extra chunks ride in lanes that were idle anyway. Pass time barely moves; per-chunk cost falls steeply. On the far side the lanes are full. Adding chunks now adds real arithmetic, so **pass time grows roughly in proportion to batch size**. That is why p95 rose by 40 ms for a 20 ms window change: 1. the batch's earliest chunk now waits up to 40 ms instead of 20 ms - about **20 ms** more queueing delay; 2. the pass it finally rides in grew from roughly 22 ms to roughly 38 ms - about **16 ms** more service time, paid by **every** chunk in the batch, not just the earliest; 3. the longer pass leaves the next batch waiting slightly longer to start, which shows up in the tail. The per-chunk cost meanwhile fell only from about 1.28 ms to about 1.23 ms. That is the 4%. ## The other way a window stops paying Saturation is one ceiling; **arrivals** are the other. If the service only takes 300 chunks a second, a 40 ms window collects about 12 chunks no matter how much room the device has, and widening it further just waits. Distinguishing the two matters because the fixes differ: - **device-bound** (this case): throughput climbed with batch size right up to the plateau, so the plateau is the hardware. A wider window cannot help; a cheaper pass could. - **arrival-bound**: the measured batch stays far below what the device handles happily, and throughput tracks the arrival rate. The window is not the constraint at all, and shortening it costs almost nothing in rate. ## The knee is not a constant Because the window collects roughly `arrival rate x window` chunks, the position of the knee **on the window axis** moves with traffic. Halve the arrival rate and the same window builds half the batch, putting the service back on the steep part of the curve with a lower throughput ceiling. A window tuned at peak and left alone is mistuned off-peak, and vice versa. That is why the window belongs to a sweep that is repeated, with a rule attached: pick the largest window whose p95 still fits the share of the caption deadline this stage holds, and stop earlier than that if marginal throughput has already collapsed. ## Doing the diagnosis in an interview The answer an interviewer is listening for has three moves: 1. **Name the shape** - throughput against latency as batch size grows, with a knee where per-chunk cost stops falling. 2. **Attribute the latency** - split the 40 ms into added wait and a longer pass, rather than calling it all queueing delay. 3. **Act on the deadline, not the curve alone** - 112 ms of p95 may still be fine against a generous caption deadline, but you are paying it for 4%, and that tail is budget you will want back during a spike or a retry. A weak answer stops at "bigger batches are better for throughput" and never notices that this particular move bought almost none.
- At the window you keep, p95 still misses the caption deadline at the required arrival rate. What does the curve tell you about how much cheaper the scorer must get?It converts the deadline into a requirement. Subtract the wait needed to collect the batch that rate implies from the share of the deadline this stage holds, and what remains is the pass time you can afford at that batch size - a per-pass millisecond target. That target is a serving-side number you hand to whoever chooses how to make the scorer cheaper; picking the technique is not this decision.
- Why does the knee move when the arrival rate drops?The window collects roughly arrival rate times window, so at half the rate the same window assembles half the batch. The service slides back onto the steep part of the curve, where per-chunk cost is higher and the achievable throughput ceiling is lower, and the window that was just past the knee at peak now sits before it. Re-sweep after any material traffic change.
- How do you tell a saturated device from a service that is simply arrival-bound?Compare measured batch size against what the device handles without pass time growing. If throughput climbed with batch size up to a plateau and pass time now grows in step with the batch, the device is the ceiling. If the batch stays small and throughput tracks incoming traffic while pass time is flat, arrivals are the ceiling and the window is not the constraint at all.
saying these in an interview costs you the question
- Says a wider batch always raises chunks a second
- Reads the flat part of the curve as the accelerator being idle
- Assumes the window only adds waiting, never a longer pass
- Tunes the window once and treats the knee as fixed across arrival rates
- Picks the window that maximises throughput without checking the caption deadline
- Blames an unrelated regression for a latency rise the window fully explains