skip to content

In pipeline-parallel training, what is the pipeline bubble and how do micro-batches shrink it?

level: middleimportance: must knowfreq 58%

answer

  1. devices wait while the chain fills
  2. count wasted slots at each end
  3. S-1 in front, S-1 behind
  4. split the batch so stages overlap
  5. fraction sets stages against micro-batches

basics

~20 s

The bubble is the idle time while the pipeline fills and drains: with S stages, S-1 stage-slots are wasted at each end. Splitting the batch into M micro-batches makes the bubble fraction (S-1)/(M+S-1), so more micro-batches shrink it.

solid answer

~50 s

Pipeline parallelism puts consecutive groups of layers on different devices, so stage 2 cannot start until stage 1 has produced something for it. Feed the pipeline a single batch and stage `s` sits idle for the `s-1` stage-times it takes the work to reach it, and idle again while the tail drains — with `S` stages that is `S-1` wasted slots at each end. That idle time is the bubble. The fix is to cut the batch into `M` micro-batches and push them back to back so stages overlap: the step takes `(M + S - 1)` stage-times against an ideal of `M`, giving a bubble fraction of `(S-1)/(M + S - 1)`. With 8 stages and 1 micro-batch, 7/8 of the machine is idle; with 32 micro-batches it falls to about 18%. Micro-batches are not free — smaller ones use each device less efficiently, and raising `M` raises the global batch size.

go deeper

for a junior

Recall that stages form a chain, so a device with nothing arriving yet simply waits, and that feeding the batch in smaller pieces lets the stages work at the same time.

for a middle

Be able to derive the timeline: the step costs micro-batch count plus stage count minus one slots against an ideal of the micro-batch count, and state the resulting idle fraction for concrete numbers.

for a senior

Show you know why the micro-batch count cannot simply be raised without limit, and that you would separate fill-and-drain loss from stage imbalance before tuning anything.

for a principal

Own the interaction with the optimization recipe: pipeline depth forces a micro-batch count, which forces a global batch size, which forces a learning-rate schedule. That chain should be a deliberate choice, not a side effect.

## What pipeline parallelism does Pipeline parallelism splits a network *by depth*. Layers 1-8 live on device A, layers 9-16 on device B, and so on; each group is a **stage**. A micro-batch of activations flows forward through the stages, and its gradient flows backward through them in reverse. Unlike splitting a single matrix across devices, no layer is ever divided — each stage owns whole layers, so a stage's own compute needs no collective at all. The only traffic is the activation tensor handed across each stage boundary, and its gradient handed back. ## Where the idle time comes from The stages form a dependency chain. Stage 2 has literally nothing to compute until stage 1 has finished a piece of work and passed it on. If you feed the pipeline the whole batch as one unit, the timeline is brutally simple: stage 1 works while everyone else waits, then stage 2 works while everyone else waits, and so on. With `S` stages, the total work per device is one stage-time and the wall clock is `S` stage-times, so device utilisation is `1/S`. Eight stages, one batch: 87.5% of the fleet is idle. That wasted time is the **bubble**. It has two halves: the *fill*, while the first micro-batch is still working its way down the chain and later stages have nothing yet, and the *drain*, while the last micro-batch finishes and earlier stages have nothing left. Each half is `S-1` stage-slots wide. ## Micro-batches The remedy is to make the unit of work smaller than the batch. Split the batch into `M` **micro-batches** and inject them one after another. As soon as micro-batch 1 leaves stage 1, micro-batch 2 enters it — so after the fill, every stage is working on a different micro-batch simultaneously, and the pipeline is in a steady state. Count the timeline in stage-slots. The last micro-batch enters stage 1 at slot `M`, and it needs `S-1` more slots to reach the final stage. So the whole step occupies `M + S - 1` slots per device, of which `M` are useful: - ideal time = `M` stage-times - actual time = `(M + S - 1)` stage-times - **bubble fraction = `(S - 1) / (M + S - 1)`** Some worked values with `S = 8`, so `S - 1 = 7`: - `M = 1`: 7/8 = 87.5% idle - `M = 4`: 7/11 ≈ 64% idle - `M = 8`: 7/15 ≈ 47% idle - `M = 32`: 7/39 ≈ 18% idle - `M = 128`: 7/135 ≈ 5% idle The useful rule of thumb: you want `M` several times larger than `S`. At `M = 4S` the bubble is roughly 20%; at `M = 16S` it is roughly 6%. ## Why you cannot just make M enormous Three ceilings bite. **Device efficiency.** A micro-batch is the batch dimension of every matmul the stage performs. Shrink it far enough and the matmuls become too small to saturate the device — you spend more time on kernel launch overhead and memory traffic than on arithmetic, so the stage-time itself stops shrinking proportionally. You have traded bubble for poor arithmetic intensity. **Global batch size.** The step's global batch is micro-batch size times `M` times the number of replicas. If you hold the micro-batch size fixed and raise `M`, the global batch grows, which changes the optimization problem: learning-rate schedule, warmup and total step count all have to be re-tuned, and past some point more examples per step stop buying proportionally better progress. **Per-micro-batch overhead.** Every micro-batch pays a fixed cost in scheduling and in the send/receive at each boundary. Enough tiny micro-batches and that fixed cost is the new bottleneck. ## The schedule matters too The formula above describes a schedule that runs all `M` forward passes and then all `M` backward passes. It has a nasty side effect: the activations of every in-flight micro-batch must be kept alive until its backward pass arrives, so peak activation memory grows with `M` — exactly the knob you were turning up. Scheduling one backward pass as soon as it becomes available, interleaved with the remaining forwards, yields the same bubble fraction but bounds the number of in-flight micro-batches per stage to roughly `S`, which decouples the bubble knob from memory. ## A second, different source of idle time The bubble is only the fill-and-drain loss, and it assumes every stage takes the same time. It does not cover imbalance. If one stage is slower than the others, every other stage waits for it on *every* micro-batch, so that loss is multiplied by `M` rather than amortised by it — raising `M` does not help at all. The two problems have different fixes, and diagnosing which one you have is the first step.

  • Why not keep raising the micro-batch count until the bubble is negligible?
    Three limits. Smaller micro-batches make each stage's matmuls too small to saturate the device, so stage-time stops falling proportionally. Holding micro-batch size fixed and raising the count inflates the global batch, which forces the learning-rate schedule to be re-tuned. And each micro-batch pays fixed scheduling and boundary-transfer overhead. The returns are also sharply diminishing: past roughly four times the stage count, further increases buy very little.
  • One stage in a deep encoder-decoder holds twice as many layers as the others. What happens to step time?
    That stage sets the cadence for the whole pipeline: every other stage idles the difference on every micro-batch, so the loss scales with the micro-batch count instead of being amortised by it. More micro-batches will not help. Rebalance by measured stage time rather than layer count — embedding layers and the loss head are often much heavier or lighter than a plain block.
  • How does the backward pass change the picture?
    Each micro-batch must also traverse the stages in reverse, so fill and drain apply to the backward direction as well and the bubble fraction is unchanged. What does change is memory: a schedule that runs all forwards before any backward must retain every in-flight micro-batch's activations. Interleaving a backward pass as soon as one is available keeps the same bubble but caps in-flight micro-batches per stage at roughly the stage count.

saying these in an interview costs you the question

  • Claims more devices always mean proportionally faster steps
  • Confuses micro-batches with per-replica batch splitting
  • Says the bubble disappears entirely with enough micro-batches
  • Ignores that the global batch size grows with micro-batch count
  • Blames the bubble for idle time actually caused by stage imbalance

context