skip to content

questions

4

What does batching audio chunks into one accelerator pass buy a live captioning service, and what does it cost?

level: juniorimportance: must knowfreq 64%

answer

  1. fixed cost per pass, not per item
  2. throughput up, earliest arrival waits
  3. the window bounds the first chunk's wait
  4. one pass, weights read once
  5. 12 ms alone against 40 ms for thirty-two

basics

~20 s

Batching amortises the accelerator's large fixed per-pass cost over many chunks, so throughput rises several-fold. It costs latency: the first chunk in a batch waits for the batch to fill or the window to expire before any scoring starts.

solid answer

~40 s

An accelerator charges a fixed cost per *pass* - moving inputs across the device boundary, starting the work, reading the model's weights - and that cost barely changes with how many chunks ride in the pass. Scoring one 500 ms audio chunk alone might take 12 ms; scoring 32 in a single pass might take 40 ms, or 1.25 ms a chunk. Throughput goes from about 83 chunks a second to 800. The bill arrives as queueing delay: the server holds arriving chunks until the batch reaches its cap or a collection window expires, and the chunk that arrived *first* opened that window, so it waits longest. The last chunk to join waits almost nothing. So batching trades a per-request delay, bounded by the window, for a large gain in chunks a second.

go deeper

for a junior

Recall the shape of the trade: one pass over many chunks amortises a fixed cost, so throughput rises, and the chunk that arrived first waits for the batch to close. Be able to say which side each number is on.

for a middle

Explain the mechanics: what the fixed per-pass cost actually covers, why the window is armed by the first arrival, and why per-chunk service time and per-chunk end-to-end latency move in opposite directions.

for a senior

Quote the trade as an exchange rate against a stated deadline - so many chunks a second bought for so many milliseconds of worst-case wait - and say how you measured both rather than asserting that batching is good.

for a principal

Frame batching as spending tail latency, a budget shared with retries, upstream variance and traffic spikes. Decide how much of that budget a throughput gain is allowed to consume and make the exchange rate explicit for the teams downstream.

## What an accelerator charges per pass An accelerator is a device built to run the same arithmetic across a very wide set of lanes at once. Every invocation carries a **fixed cost that does not depend on how many items ride in it**: moving the input across the device boundary, starting the work on the device, reading the model's weights out of device memory, and collecting the output on the way back. That cost is paid once per **pass**, not once per chunk. A live captioning service cuts each session's audio into fixed-length chunks - say 500 ms of audio - and scores every chunk to produce the next span of caption text. One such chunk uses a sliver of the device's lanes. Sent alone, it pays the entire fixed cost to occupy a fraction of the hardware. Sent with thirty-one others, the same fixed cost is divided thirty-two ways and the lanes are actually used. ## The worked numbers | | one chunk a pass | 32 chunks a pass | |---|---|---| | pass time | 12 ms | 40 ms | | cost a chunk | 12 ms | 1.25 ms | | chunks a second | about 83 | 800 | | earliest chunk's added wait | none | up to the collection window | The pass got slower in absolute terms - 12 ms became 40 ms - and *cheaper* per chunk, by roughly ten times. That is the whole mechanism in two rows. ## What the earliest chunk pays The server cannot batch chunks that have not arrived yet, so it holds the ones it has. A batcher closes a batch on whichever comes first: 1. the batch reaches its configured maximum size, or 2. a **collection window**, armed when the first chunk of the batch arrived, expires. That ordering matters for who waits. The window is armed by the **first** arrival, so that chunk's added wait is bounded by the window length, and every later arrival joins an already-open batch and waits only the remainder. The chunk that lands just before the batch closes waits essentially nothing. Under even arrivals the mean added wait is about half the window; the worst case is the full window, and only one chunk a batch pays it. Note what this is **not**: it is not a flat delay added to every chunk, and it is not paid at all when arrivals are dense enough to hit the size cap before the timer fires. ## Which throughput are you quoting Two numbers in this conversation are both called throughput and they are not the same: - **chunks a second out of the service** - the rate captions are produced, the number a capacity conversation means; - **chunks inside one pass** - the batch size, a configuration choice. Saying "we do 32" answers neither question cleanly. Always name the unit: 800 chunks a second, at a batch of 32. Likewise, two latencies live here. **Service time** is the pass itself (40 ms). **Queueing delay** is the wait to be included in a pass. End-to-end caption latency is the sum plus whatever the rest of the path costs, and only the queueing delay is what batching added. ## What batching does not do - It does not reduce total work - the same arithmetic runs, just arranged into fewer, wider passes. - It does not make any individual chunk's scoring faster; per-chunk *cost* falls, per-chunk *latency* rises. - It does not help a stream whose arrival rate is too thin to collect a worthwhile batch inside the wait the caption deadline allows - there the window expires on a near-empty batch and you have bought delay for nothing. - It does not keep paying indefinitely. Once one pass already saturates the device's parallel width, pass time grows about linearly with batch size, so a wider batch stops lowering per-chunk cost while the wait keeps climbing. ## How to talk about it in a design round State the trade as an exchange rate rather than a preference. "A 20 ms collection window buys us roughly an order of magnitude in chunks a second and costs the earliest chunk in each batch at most 20 ms, against a caption deadline of 300 ms" is an answer. "We batch because it is faster" is not - it is also wrong for the request that waited. The interviewer is listening for whether you know which side of the trade each number sits on, and whether you checked the wait against a stated deadline rather than against nothing.

  • Why does the first chunk in a batch wait longer than the last one?
    The collection window is armed by the first arrival, which opens the batch. Every later chunk joins a batch already partway through its window and waits only the remainder, and the one that lands as the batch closes waits almost nothing. Under even arrivals the mean added wait is roughly half the window; the full window is a worst case one chunk a batch pays.
  • Does raising the maximum batch size always raise chunks a second?
    No. While the fixed per-pass cost is still being amortised, per-chunk cost falls steeply. Once a single pass already saturates the device's parallel width, pass time grows roughly in proportion to batch size, so per-chunk cost flattens while the collection wait needed to fill that batch keeps growing. Past that point a wider batch buys tail latency, not rate.
  • If the caption deadline is generous, should the window just be as long as the deadline allows?
    Only up to the point where extra window still collects chunks worth having. Beyond it you are adding wait to every batch's earliest arrival for a rate gain of a percent or two, and you have spent tail latency you may want later for a retry, a slower upstream or a traffic spike. Tune the window against measured rate gain, not against the deadline alone.

A shuttle bus that leaves when it is full or when the clock says go moves far more people per trip than a taxi moves per passenger. The person who boarded first pays for that in waiting.

saying these in an interview costs you the question

  • Says batching lowers latency for every request, not just raising throughput
  • Treats the accelerator's per-chunk cost as fixed regardless of batch size
  • Assumes the last chunk to join a batch waits as long as the first
  • Describes the collection window as a flat delay added to every chunk
  • Claims a wider batch always raises chunks a second with no latency cost
open as a page

A captioning service widened its accelerator batch window from 20 ms to 40 ms; p95 rose 40 ms while throughput rose 4% - what does that say?

level: seniorimportance: must knowfreq 71%

basics

~20 s

It says the service is past the knee of its throughput-versus-latency curve. The fixed per-pass cost is already amortised, so a wider batch grows pass time roughly in step with batch size while buying almost no extra chunks a second. Revert the window.

open as a page

What does Little's Law give a captioning service that takes 2,000 chunks a second at a mean 60 ms from arrival to caption?

level: middleimportance: should knowfreq 52%

basics

~20 s

Little's Law gives the work in flight: concurrency equals arrival rate times mean time in the system, so 2,000 a second at 60 ms means about 120 audio chunks are inside the service at any instant. It is an identity over averages, not a capacity plan.

open as a page

A low-traffic caption language pair takes 12 chunks a second - what decides whether it is scored on an accelerator or on general-purpose cores?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Whether the arrival rate can assemble a worthwhile batch inside the wait the caption deadline allows, and whether one chunk alone already carries enough parallel work to use the device. At 12 chunks a second neither holds, so an accelerator runs at its worst operating point.

open as a page