skip to content

Arithmetic and Precision

A GPU spends its time on parallel multiply-accumulates bounded by memory bandwidth, and on the number formats those operations run in. Interviewers start here before any scaling question.

on this pageshow

explore

questions

10

A training run's GPU sits at 30% utilisation while the host CPU is pinned decoding image tiles — what limits step time?

level: juniorimportance: must knowfreq 60%

answer

  1. look at who is busy and who is idle
  2. the model is not the suspect here
  3. decode and resize repeated every epoch
  4. prepare the next batch during this one
  5. time a step on one resident batch

basics

~20 s

The input pipeline, not the model. The GPU idles waiting for batches while the host decodes and resizes images. Fixes are overlapping data preparation with compute, adding host-side parallelism, and pre-processing tiles into a cheaper stored form.

solid answer

~50 s

A saturated host CPU next to an idle accelerator says the device is starved: batches are not being produced as fast as it can consume them, so most of the step is waiting, not computing. On a satellite-imagery run at 512x512 tiles the usual culprit is per-sample compressed-image decode plus a resize, done on the host for every sample of every epoch. Confirm it before changing anything: repeat training steps on a single batch already resident in device memory — if step time collapses and utilisation jumps, the model was never the limit. Then fix the pipeline: prepare batch k+1 while batch k trains, use more parallel worker processes if cores are free, pre-decode the tiles once into a compact stored format at the training resolution, and overlap the host-to-device copy with compute. A faster accelerator buys nothing here; it would simply idle more.

go deeper

for a junior

Be ready to name the input pipeline as the suspect when the accelerator is idle and the host CPU is busy, and to say that the fix is preparing the next batch while the current one trains.

for a middle

Explain the mechanics: per-sample decode, resize and augmentation repeated every epoch, the collate and the host-to-device copy, and which of those can be overlapped, parallelised or precomputed away.

for a senior

Show the diagnostic discipline — a fixed-batch timing test to isolate the model, a measured per-batch production time, and a fix chosen by measured payoff rather than by reflex. Say plainly that faster hardware would only idle more.

for a principal

Own the systemic version: how many host cores and how much storage throughput a device needs to stay fed, whether a one-off pre-processed dataset is worth its storage and its staleness risk, and how utilisation targets are set and monitored across a fleet.

## Reading the symptom Two numbers together tell the story: the accelerator is busy only 30% of the time, and the host CPU is at its limit. That combination is almost never a model problem. It says the training loop spends most of each step waiting for the next batch to be produced and delivered, while the host burns all of its cores producing it. The device's compute ceiling and its bandwidth ceiling are both irrelevant while it has nothing to work on. ## Where the host time goes A typical image pipeline per sample does: read bytes from storage, decode a compressed image, resize or crop to the training resolution, apply augmentations, convert to the numeric layout the model expects, and collate samples into a batch. On 512 x 512 tiles the decode and the resize dominate, and both are CPU-heavy and done *again on every epoch* even though their result is identical each time. Multiply by the batch size and it is entirely ordinary for one host to be unable to feed one accelerator. The delivery leg matters too. Once a batch exists on the host it must be copied to device memory. If that copy is issued synchronously between steps rather than overlapped with the previous step's compute, it adds dead time even when the producers keep up. ## Confirming rather than guessing The cheap decisive test is to remove the pipeline from the measurement: take one batch, keep it resident in device memory, and run training steps on that same batch in a loop. Compare the resulting step time against the real run's step time. - Step time collapses and utilisation rises: the input pipeline is the limit, exactly as the symptom suggested. - Step time barely changes: the device really is busy for that long, and the bottleneck is inside the model — at which point the compute-versus-bandwidth classification of its layers is the next question, not the loader. The second useful measurement is simply how long batch production takes on the host per batch, compared with the model-only step time from the test above. Whichever is larger sets the pace, because the two run in parallel once you overlap them. ## The fixes, roughly in order of payoff 1. **Overlap production with compute.** Produce batch k+1 while batch k is training, keeping a small queue of ready batches. This alone converts a serial `prepare + compute` step into `max(prepare, compute)`. 2. **Parallelise on the host.** Multiple worker processes decoding and augmenting concurrently, sized to the cores actually available. Note the ceiling: if the host is already saturated, more workers only add contention. 3. **Stop paying the decode repeatedly.** Pre-process the dataset once into a compact stored format at the training resolution — already decoded, already resized, written in large sequential records — and read from that. This is usually the single biggest win for image workloads, because it deletes the expensive step rather than parallelising it. 4. **Move work to the device.** Some augmentation and the numeric conversion can run on the accelerator, which has spare capacity by definition in this scenario. 5. **Overlap the transfer.** Issue the host-to-device copy for the next batch while the current step computes, so delivery is not serialised behind compute. 6. **Check storage.** If the tiles come from a remote or slow filesystem in many small files, the host may be blocked on reads rather than on decoding; large sequential reads and local caching fix that variant. ## What not to do Do not start by shrinking the model, lowering its resolution, or reaching for a device with a higher FLOP rating. Every one of those makes the device finish its share sooner and idle *more*, leaving wall-clock time roughly where it was. Do not assume low utilisation always means a small batch either — a bigger batch on a starved run just makes each wait longer. And do not accept a utilisation figure alone as proof: pair it with the fixed-batch timing test, because a device can also show low utilisation when it is blocked on synchronisation or on many tiny operations. ## When lower utilisation is acceptable If the pre-processing genuinely cannot be cached — heavy randomised augmentation that must differ every epoch, or data that is generated on the fly — then some host cost is irreducible, and the engineering question becomes how many host cores per accelerator the workload needs. That is a capacity answer, not a code answer, and stating it that way is what separates a considered response from a guess.

  • How do you prove the model is not the bottleneck in this run?
    Repeat training steps on a single batch that already sits in device memory, so no host work happens between steps. If step time drops sharply and device utilisation climbs, the input pipeline was the limit. If step time is unchanged, the device really was busy that long and the investigation moves inside the model. It is a five-minute test and it settles the argument before anyone rewrites a loader or orders hardware.
  • Would moving this run to an accelerator with double the throughput cut wall-clock time?
    No. The device is already idle 70% of the time, so making its busy fraction shorter leaves the waiting untouched and simply lowers utilisation further. Wall-clock time is set by how fast batches arrive. The money is better spent on host cores, faster storage, or a one-off pre-processing pass that removes the repeated decode entirely.
  • The pipeline already runs many parallel workers and the host is still saturated. What now?
    Stop parallelising the expensive step and delete it. Pre-decode and resize the tiles once into a compact stored format at the training resolution and read that instead, so every epoch pays only a sequential read. Move cheap augmentations and the numeric conversion onto the accelerator, which has idle capacity. If the work genuinely cannot be cached, the honest conclusion is that this workload needs more host cores per device.

A fast oven and one cook chopping vegetables. The oven's temperature is not the problem; the meals come out slowly because nothing is ready to go in.

saying these in an interview costs you the question

  • Blames the model and starts shrinking it first
  • Proposes a faster accelerator for a run that is already idle
  • Assumes low utilisation always means the batch is too small
  • Never times a step with the input pipeline removed
  • Adds more worker processes to an already saturated host

context

open as a page

Why can a gradient that is nonzero in 32-bit floats round to exactly zero in 16-bit?

level: juniorimportance: must knowfreq 64%

basics

~20 s

A 16-bit float's 5 exponent bits bottom out near 6e-5 for normal values and near 6e-8 once subnormals run out. A gradient smaller than that has no representation, so it stores as exactly zero and that weight stops moving.

open as a page

What is mixed-precision training, and where do its speed and memory wins come from?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Mixed-precision training runs the forward and backward passes in a 16-bit float while keeping a single-precision copy of the weights. Wins come from moving half the bytes and from matrix units that multiply 16-bit inputs at much higher throughput.

open as a page

What is arithmetic intensity, and how does it decide whether a GPU operation is compute- or bandwidth-bound?

level: middleimportance: must knowfreq 58%

basics

~20 s

Arithmetic intensity is the floating-point operations an op performs per byte it moves to and from device memory. Compare it with the device's peak-FLOP-rate divided by peak bandwidth: below that ridge point the op is bandwidth-bound, above it compute-bound.

open as a page

Which parts of a mixed-precision training step must stay in single precision, and why?

level: middleimportance: must knowfreq 70%

basics

~20 s

A mixed-precision step keeps the weights the optimizer updates, large reductions, the loss and normalization statistics in 32-bit. Those are the places where a tiny quantity is added to or accumulated with a much larger one, which 16-bit arithmetic destroys.

open as a page

In 16-bit floats, why can adding a 1e-7 update to a weight of 1.0 change nothing?

level: middleimportance: should knowfreq 46%

basics

~20 s

A 16-bit float keeps about 11 significand bits, so the next value above 1.0 is roughly 1.001. A 1e-7 update is far below half that gap, so the sum rounds back to 1.0 and the weight never moves.

open as a page

Why does half-precision training multiply the loss by a large scale factor before backpropagation?

level: middleimportance: should knowfreq 58%

basics

~20 s

Half-precision gradients can be too small to represent and collapse to zero. Multiplying the loss by a large constant scales every gradient up by that factor through the chain rule; the scale is divided back out before the optimizer step.

open as a page

A speech encoder's normalization, activation and residual-add chain dominates step time — what does fusing it into one pass remove?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Fusion removes the trips to device memory between the stages. Unfused, the activation tensor is written and re-read after each operation; fused, it is read once and written once. The arithmetic is identical — only the memory traffic shrinks.

open as a page

Your half-precision run needs a tuned loss scaler; why does bf16 let you delete it?

level: seniorimportance: should knowfreq 46%

basics

~20 s

bf16 carries the same exponent range as 32-bit float, so gradients that underflow in fp16 stay representable and no loss scaler is needed. The price is fewer mantissa bits, so 32-bit master weights and accumulation still matter.

open as a page

Why can a global gradient norm in 16-bit come out infinite when every gradient is finite?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

A global norm sums squares across every parameter. Each square is finite, but the running total crosses the 16-bit ceiling of 65504 and saturates to infinity. The failure lives in the accumulator, not in any individual gradient.

open as a page