Your GPU idles between batches in PyTorch training — how do you find and fix the input bottleneck?
answer
- measure before you tune
- run the loader with no model
- overlap, then transfer, then work
- pinned pages enable real async copies
- non_blocking needs pin_memory
basics
~20 sFirst prove it: time a pass that iterates the DataLoader with no model. If that alone matches epoch time, the pipeline is the limit. Then raise num_workers, add persistent_workers, pin_memory with non_blocking copies, and cut per-sample decode cost.
solid answer
~50 sStart with measurement, not tuning. Iterate the loader alone, no forward or backward, and time it; compare with a full epoch. If the data-only pass takes about as long, the input pipeline is the ceiling. `torch.profiler` confirms it by showing the gaps: long CPU time attributed to the dataloader iterator between short bursts of kernel activity. Then work the ladder. Raise `num_workers` toward the CPU quota and set `persistent_workers=True` so they survive epoch boundaries; raise `prefetch_factor` if per-sample cost is spiky. Set `pin_memory=True` and copy with `.to(device, non_blocking=True)` — without pinned memory the copy is effectively synchronous, so the pair only works together. If the CPU is genuinely saturated, reduce the work: cheaper decoding, resize offline, store a packed preprocessed format, or move batch-level normalization and augmentation onto the GPU. And check the floor: raw disk or network throughput, `/dev/shm` size, and thread oversubscription in workers.
code
python · 13 linesimport time
import torch
from torch.utils.data import DataLoader, TensorDataset
if __name__ == "__main__":
ds = TensorDataset(torch.randn(4096, 128))
loader = DataLoader(ds, batch_size=64, num_workers=4, pin_memory=True)
start = time.perf_counter()
for _ in loader: # no model: this times the input pipeline alone
pass
print(f"data-only epoch: {time.perf_counter() - start:.2f}s")go deeper
Know the first-aid settings — num_workers above 0 and pin_memory=True — and that the data pipeline, not the model, is a common reason a GPU looks underused.
Explain the mechanism: workers prefetch batches, pinned memory allows an asynchronous host-to-device copy, and non_blocking=True only helps when the source is pinned.
Lead with measurement — a data-only pass, a synthetic-batch pass, the profiler — then work a ladder and read the counter-signals that say whether you are CPU-bound, I/O-bound or already fine.
Treat it as a data-layout problem, not a flag problem: shard format, offline preprocessing, storage tier and container CPU/shm budgets decide the ceiling for every job on the cluster, and you should know when the right answer is to stop optimizing the pipeline.
## Prove the diagnosis before you tune "GPU utilization is low" is not a diagnosis. Three cheap experiments separate the causes. **Data-only pass.** Iterate the `DataLoader` in a loop that does nothing with the batch and time it. If that takes 90 seconds and a full epoch takes 100, the model is nearly free and the pipeline is the whole problem. If the data-only pass takes 10 seconds, stop tuning the loader. **Synthetic-batch pass.** Train on one preloaded batch repeated N times. That gives the pure compute time. The gap between it and a real epoch is the pipeline's unhidden cost. **Profiler.** `torch.profiler.profile` over a handful of steps shows where wall-clock goes; wrapping regions in `record_function` makes the fetch phase legible. What you are looking for is dead time between kernel bursts, attributed to CPU-side iteration. Utilization sampled from `nvidia-smi` is a weak signal — it reports whether *any* kernel was resident during a sample window, not how much of the SM capacity was used, so a starved run can still read high. ## The fix ladder **More overlap.** `num_workers` toward the CPU quota (not the visible core count — containers lie), `persistent_workers=True` to stop paying respawn cost every epoch, `prefetch_factor` above 2 when per-sample cost is uneven so a slow sample does not stall the consumer. Measure after each change; throughput usually plateaus and then regresses. **Faster transfer.** `pin_memory=True` makes DataLoader copy each collated batch into page-locked host memory using a dedicated thread in the main process. Page-locked pages cannot be swapped, which is exactly what lets the CUDA driver do a true DMA transfer while the GPU keeps computing — but only if you also request it: `batch.to(device, non_blocking=True)`. From ordinary pageable memory, `non_blocking=True` is silently near-synchronous. The pair costs page-locked RAM and one extra host copy, so it is a win when transfers are large and the host has memory to spare, and a mild loss on a memory-tight box. **Less work per sample.** This is where the big multiples live. Decoding full-resolution JPEGs to then crop to 224x224 is the archetypal waste — pre-resize offline. Convert thousands of small files into a handful of sequential shards so reads stop being random 4 KB seeks. Cache decoded tensors in a memory-mapped array. Do normalization and simple geometric augmentation on the GPU over the whole batch rather than per-sample on the CPU. Avoid per-sample Python overhead: a dataset that constructs several intermediate objects per call spends more time in the interpreter than in the decoder. **Fix the floor.** If the storage layer delivers 200 MB/s and your pipeline needs 800, no amount of worker tuning helps — measure raw read throughput separately. In containers, a default 64 MB `/dev/shm` causes worker crashes or throttling once batches get large. And thread oversubscription is a classic own-goal: every worker may launch its own BLAS/OpenMP pool, so `torch.set_num_threads(1)` inside `worker_init_fn` often beats adding workers. ## Reading the counter-signals If raising `num_workers` does nothing and CPU sits near 100%, you are compute-bound on the CPU: reduce work, do not add processes. If CPU is idle and workers are many, you are I/O-bound: change the storage layout. If throughput improves then collapses at high worker counts, you are hitting memory, shared-memory or scheduler limits. If the first batch of every epoch is slow but the rest are fine, that is worker startup — turn on `persistent_workers`. ## Knowing when to stop The goal is not a saturated GPU for its own sake; it is that the pipeline stays comfortably ahead of the model with a modest buffer. Once the data-only pass is meaningfully faster than the compute-only pass, further pipeline work buys nothing, and the next lever is the model side — larger batches, mixed precision, or a compiled graph. Chasing the last 10% of a pipeline that already leads is time spent on the wrong bottleneck.
- What does pin_memory actually cost?An extra host-side copy per batch, performed by a dedicated pinning thread, plus page-locked RAM that the OS cannot swap or reclaim. With large batches and many workers in flight that memory adds up, and on a host already short on RAM or CPU it can make things worse. It pays off when device transfers are large enough to be worth overlapping with compute.
- When does moving augmentation to the GPU help?When the CPU is the bottleneck and the augmentations are batch-friendly tensor ops — normalization, resize, flips, crops — which run far faster on a whole batch on the GPU than per sample in Python. The catch is that this compute now competes with the model's, so it only wins while the GPU has headroom. Decode-heavy work usually stays on the CPU.
- How do you tell an I/O bottleneck from a CPU bottleneck?Watch CPU while raising num_workers. If the cores are pegged and adding workers changes nothing, you are CPU-bound — cut per-sample work rather than adding processes. If cores are mostly idle while the loader still lags, you are waiting on storage or the network: measure raw read throughput independently and fix the layout, typically by converting many small files into a few sequentially-read shards.
saying these in an interview costs you the question
- Adds workers without measuring the pipeline first
- Trusts nvidia-smi utilization as a precise signal
- Uses non_blocking=True without pin_memory
- Believes pin_memory moves data to the GPU
- Tunes the loader when storage throughput is the real limit