skip to content

Compute vs Memory Bound

Dense matrix multiplication reuses each loaded value many times, while elementwise ops read and write far more than they compute. Interviewers use it to explain why some layers refuse to speed up.

on this pageshow

questions

3

A training run's GPU sits at 30% utilisation while the host CPU is pinned decoding image tiles — what limits step time?

level: juniorimportance: must knowfreq 60%

answer

  1. look at who is busy and who is idle
  2. the model is not the suspect here
  3. decode and resize repeated every epoch
  4. prepare the next batch during this one
  5. time a step on one resident batch

basics

~20 s

The input pipeline, not the model. The GPU idles waiting for batches while the host decodes and resizes images. Fixes are overlapping data preparation with compute, adding host-side parallelism, and pre-processing tiles into a cheaper stored form.

solid answer

~50 s

A saturated host CPU next to an idle accelerator says the device is starved: batches are not being produced as fast as it can consume them, so most of the step is waiting, not computing. On a satellite-imagery run at 512x512 tiles the usual culprit is per-sample compressed-image decode plus a resize, done on the host for every sample of every epoch. Confirm it before changing anything: repeat training steps on a single batch already resident in device memory — if step time collapses and utilisation jumps, the model was never the limit. Then fix the pipeline: prepare batch k+1 while batch k trains, use more parallel worker processes if cores are free, pre-decode the tiles once into a compact stored format at the training resolution, and overlap the host-to-device copy with compute. A faster accelerator buys nothing here; it would simply idle more.

go deeper

for a junior

Be ready to name the input pipeline as the suspect when the accelerator is idle and the host CPU is busy, and to say that the fix is preparing the next batch while the current one trains.

for a middle

Explain the mechanics: per-sample decode, resize and augmentation repeated every epoch, the collate and the host-to-device copy, and which of those can be overlapped, parallelised or precomputed away.

for a senior

Show the diagnostic discipline — a fixed-batch timing test to isolate the model, a measured per-batch production time, and a fix chosen by measured payoff rather than by reflex. Say plainly that faster hardware would only idle more.

for a principal

Own the systemic version: how many host cores and how much storage throughput a device needs to stay fed, whether a one-off pre-processed dataset is worth its storage and its staleness risk, and how utilisation targets are set and monitored across a fleet.

## Reading the symptom Two numbers together tell the story: the accelerator is busy only 30% of the time, and the host CPU is at its limit. That combination is almost never a model problem. It says the training loop spends most of each step waiting for the next batch to be produced and delivered, while the host burns all of its cores producing it. The device's compute ceiling and its bandwidth ceiling are both irrelevant while it has nothing to work on. ## Where the host time goes A typical image pipeline per sample does: read bytes from storage, decode a compressed image, resize or crop to the training resolution, apply augmentations, convert to the numeric layout the model expects, and collate samples into a batch. On 512 x 512 tiles the decode and the resize dominate, and both are CPU-heavy and done *again on every epoch* even though their result is identical each time. Multiply by the batch size and it is entirely ordinary for one host to be unable to feed one accelerator. The delivery leg matters too. Once a batch exists on the host it must be copied to device memory. If that copy is issued synchronously between steps rather than overlapped with the previous step's compute, it adds dead time even when the producers keep up. ## Confirming rather than guessing The cheap decisive test is to remove the pipeline from the measurement: take one batch, keep it resident in device memory, and run training steps on that same batch in a loop. Compare the resulting step time against the real run's step time. - Step time collapses and utilisation rises: the input pipeline is the limit, exactly as the symptom suggested. - Step time barely changes: the device really is busy for that long, and the bottleneck is inside the model — at which point the compute-versus-bandwidth classification of its layers is the next question, not the loader. The second useful measurement is simply how long batch production takes on the host per batch, compared with the model-only step time from the test above. Whichever is larger sets the pace, because the two run in parallel once you overlap them. ## The fixes, roughly in order of payoff 1. **Overlap production with compute.** Produce batch k+1 while batch k is training, keeping a small queue of ready batches. This alone converts a serial `prepare + compute` step into `max(prepare, compute)`. 2. **Parallelise on the host.** Multiple worker processes decoding and augmenting concurrently, sized to the cores actually available. Note the ceiling: if the host is already saturated, more workers only add contention. 3. **Stop paying the decode repeatedly.** Pre-process the dataset once into a compact stored format at the training resolution — already decoded, already resized, written in large sequential records — and read from that. This is usually the single biggest win for image workloads, because it deletes the expensive step rather than parallelising it. 4. **Move work to the device.** Some augmentation and the numeric conversion can run on the accelerator, which has spare capacity by definition in this scenario. 5. **Overlap the transfer.** Issue the host-to-device copy for the next batch while the current step computes, so delivery is not serialised behind compute. 6. **Check storage.** If the tiles come from a remote or slow filesystem in many small files, the host may be blocked on reads rather than on decoding; large sequential reads and local caching fix that variant. ## What not to do Do not start by shrinking the model, lowering its resolution, or reaching for a device with a higher FLOP rating. Every one of those makes the device finish its share sooner and idle *more*, leaving wall-clock time roughly where it was. Do not assume low utilisation always means a small batch either — a bigger batch on a starved run just makes each wait longer. And do not accept a utilisation figure alone as proof: pair it with the fixed-batch timing test, because a device can also show low utilisation when it is blocked on synchronisation or on many tiny operations. ## When lower utilisation is acceptable If the pre-processing genuinely cannot be cached — heavy randomised augmentation that must differ every epoch, or data that is generated on the fly — then some host cost is irreducible, and the engineering question becomes how many host cores per accelerator the workload needs. That is a capacity answer, not a code answer, and stating it that way is what separates a considered response from a guess.

  • How do you prove the model is not the bottleneck in this run?
    Repeat training steps on a single batch that already sits in device memory, so no host work happens between steps. If step time drops sharply and device utilisation climbs, the input pipeline was the limit. If step time is unchanged, the device really was busy that long and the investigation moves inside the model. It is a five-minute test and it settles the argument before anyone rewrites a loader or orders hardware.
  • Would moving this run to an accelerator with double the throughput cut wall-clock time?
    No. The device is already idle 70% of the time, so making its busy fraction shorter leaves the waiting untouched and simply lowers utilisation further. Wall-clock time is set by how fast batches arrive. The money is better spent on host cores, faster storage, or a one-off pre-processing pass that removes the repeated decode entirely.
  • The pipeline already runs many parallel workers and the host is still saturated. What now?
    Stop parallelising the expensive step and delete it. Pre-decode and resize the tiles once into a compact stored format at the training resolution and read that instead, so every epoch pays only a sequential read. Move cheap augmentations and the numeric conversion onto the accelerator, which has idle capacity. If the work genuinely cannot be cached, the honest conclusion is that this workload needs more host cores per device.

A fast oven and one cook chopping vegetables. The oven's temperature is not the problem; the meals come out slowly because nothing is ready to go in.

saying these in an interview costs you the question

  • Blames the model and starts shrinking it first
  • Proposes a faster accelerator for a run that is already idle
  • Assumes low utilisation always means the batch is too small
  • Never times a step with the input pipeline removed
  • Adds more worker processes to an already saturated host

context

open as a page

What is arithmetic intensity, and how does it decide whether a GPU operation is compute- or bandwidth-bound?

level: middleimportance: must knowfreq 58%

basics

~20 s

Arithmetic intensity is the floating-point operations an op performs per byte it moves to and from device memory. Compare it with the device's peak-FLOP-rate divided by peak bandwidth: below that ridge point the op is bandwidth-bound, above it compute-bound.

open as a page

A speech encoder's normalization, activation and residual-add chain dominates step time — what does fusing it into one pass remove?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Fusion removes the trips to device memory between the stages. Unfused, the activation tensor is written and re-read after each operation; fused, it is read once and written once. The arithmetic is identical — only the memory traffic shrinks.

open as a page