A low-traffic caption language pair takes 12 chunks a second - what decides whether it is scored on an accelerator or on general-purpose cores?
answer
- can the rate fill a batch in time
- batch of one is the worst operating point
- peak throughput assumes a full batch
- a big enough item needs no batching
- 12 ms against 45 ms under a 300 ms deadline
basics
~20 sWhether the arrival rate can assemble a worthwhile batch inside the wait the caption deadline allows, and whether one chunk alone already carries enough parallel work to use the device. At 12 chunks a second neither holds, so an accelerator runs at its worst operating point.
solid answer
~50 sAn accelerator's advantage is throughput bought by batching, so it needs enough arrivals to fill a batch inside the window the deadline permits. At 12 chunks a second a 30 ms window collects `12 * 0.030 = 0.36` chunks - practically every pass is a batch of one, which pays the whole fixed per-pass cost for a sliver of the device's lanes. General-purpose cores start each chunk on arrival with no collection window at all: at 45 ms a chunk, Little's Law puts `12 * 0.045 = 0.54` cores busy, and against a 300 ms caption deadline that is comfortable. The picture flips two ways: raise the rate toward 800 chunks a second and the cores need `800 * 0.045 = 36` busy cores where one accelerator pass of 32 in 40 ms does it; or lengthen the chunk so one item saturates the device on its own.
go deeper
Recall that an accelerator earns its keep through batching, so a stream too thin to fill a batch inside the allowed wait does not get the benefit at all.
Do the arithmetic: arrival rate times the allowed collection wait gives the batch you can actually assemble, and per-item cost at batch one is what you compare against general-purpose cores.
Locate the crossover and say what moves it - arrival rate, chunk length, the deadline share - and include readiness, since an accelerator-backed replica is not servable until its artifact is resident and warm.
Decide the rule for a fleet of scorers rather than one pair: at what measured rate a scorer graduates to accelerators, who re-checks it, and how much operational surface a long tail of quiet scorers is allowed to add.
## The decision is an arithmetic one A live captioning service typically runs several scorers - a busy language pair and a long tail of quiet ones. "Accelerator or general-purpose cores" is not a preference; it falls out of three measured numbers: 1. **the arrival rate** of audio chunks for that scorer; 2. **the per-item cost on each kind of hardware**, measured at batch one; 3. **the wait the deadline allows** before a chunk must be in a pass. With those, ask the only question that matters: *can this arrival rate assemble, inside that allowed wait, a batch big enough to reach the accelerator's efficient region?* ## Running the numbers on the quiet pair - Arrival rate: 12 chunks a second. - Allowed collection wait: 30 ms out of a 300 ms caption deadline. - Chunks collected in that window: `12 * 0.030 = 0.36`. So nearly every pass is a batch of one. That is the accelerator at its **worst** operating point: the full fixed per-pass cost - crossing the device boundary, starting the work, reading the model's weights out of device memory - amortised over a single item using a fraction of the available lanes. | | accelerator, batch of 1 | general-purpose cores | |---|---|---| | per-chunk time | 12 ms | 45 ms | | collection wait | window, collecting ~0.36 chunks | none; scoring starts on arrival | | work in flight at 12 a second | 0.14 passes | 0.54 cores busy | | against a 300 ms deadline | met with room | met with room | | artifact residency | must be loaded into device memory before traffic | loaded into host memory | The accelerator is genuinely faster per chunk here - 12 ms against 45 ms. It just wins a race nobody is running: both arrangements clear a 300 ms deadline with an order of magnitude to spare, so the 33 ms saved buys nothing, while the device sits reserved and mostly idle and the replica cannot take traffic until the artifact is resident. ## Where the crossover actually is The decision inverts as the arrival rate climbs, and Little's Law - concurrency equals arrival rate times time in the system - puts a number on it. At 800 chunks a second: - general-purpose cores need `800 * 0.045 = 36` cores busy at all times just to keep up, before any allowance for bursts; - the accelerator, now able to collect 32 chunks in about 40 ms, runs one pass of 32 in 40 ms, which is `32 / 0.040 = 800` chunks a second on one device. At that rate the batching the accelerator needs is free - the chunks are already there - and the per-chunk cost falls to about 1.25 ms. The same hardware that was a poor choice at 12 a second is the obvious one at 800. ## Chunk size moves the line too Batching exists to give the device enough parallel work in one pass. A single item can do that by itself if it is big enough. A service that scores 500 ms of audio per chunk has a small item; one that buffers 8 seconds of audio per chunk has an item carrying sixteen times the work, which may saturate the device's lanes alone. Then the accelerator's advantage arrives **without** any collection window, and the arrival-rate argument stops applying. That is why chunk length and hardware choice must be decided together. Lengthening the chunk is not free either: it delays the caption by the buffering time, which comes straight out of the same deadline. ## The checks worth stating out loud - **Measure batch one on both.** If the accelerator is not dramatically better at batch one, its entire advantage lives in the batching your arrival rate cannot do. - **Convert the deadline into an allowed wait first**, then see what that wait collects. Doing it the other way round - picking a batch size and hoping the deadline stretches - is how services ship a window nothing fills. - **Account for readiness.** An accelerator-backed replica must read a possibly large artifact into device memory and warm up before it is fit to serve; a replica that takes traffic before that is done fails the first batches it is handed. - **Do not choose on peak throughput alone.** Peak throughput is the number at a full batch, and a scorer that never fills one never sees it. ## The shape of a good answer "Twelve chunks a second cannot fill a batch inside 30 ms, so the accelerator would run at batch one and we would pay its fixed cost for a fraction of its width; general-purpose cores meet the 300 ms deadline at 45 ms a chunk with under one core busy. I would revisit at roughly 800 chunks a second, where cores need about 36 busy and one accelerator pass covers the whole rate, or sooner if we lengthen the audio chunk enough that one item saturates the device." That is a decision with arithmetic under it, which is the point of the question.
- Why can a freshly started accelerator-backed replica not take caption traffic the moment its process is up?The model artifact still has to be read into device memory, and the first passes run at an atypical cost until buffers and memory layout settle. Readiness must therefore gate on a completed load plus at least one warm pass, not on the process being alive. A replica marked ready too early is handed real batches it either drops or serves far outside the deadline.
- How does lengthening the audio chunk change the choice?A longer chunk carries more parallel work per item, so one item can saturate the device's lanes without any batch at all - which removes the arrival-rate objection entirely. The cost is paid elsewhere: buffering that much audio delays the caption by the buffer length, straight out of the same end-to-end deadline. Chunk length and hardware choice are one decision, not two.
- The accelerator is measured at 12 ms a chunk and the cores at 45 ms. Why is that comparison not enough on its own?Because it compares a single number against a deadline neither breaches. What decides the choice is the rate you must sustain and what it costs to sustain it: at 12 chunks a second both are idle most of the time, while at 800 the cores need about 36 busy and the accelerator needs one pass. Per-item speed only decides when the deadline is actually tight.
saying these in an interview costs you the question
- Assumes an accelerator is always faster than general-purpose cores for the same scorer
- Picks the accelerator on peak throughput without checking the rate fills a batch
- Thinks a batch of one costs a thirty-second of a full pass
- Ignores that the artifact must be resident in device memory before traffic arrives
- Treats chunk length as irrelevant to whether batching is needed at all
- Compares per-item speed against a deadline that neither option breaches