Step time doubled after offloading stored activations to host memory — what do you check?
answer
- a different resource is being spent
- bytes per step against link throughput
- compare a round trip with rerunning it
- issue early, prefetch back, never block
basics
~20 sCheck whether the bytes moved each step exceed what the host link can carry in the time available. Offloading pays only when a round trip is faster than recomputing the values and overlaps with compute.
solid answer
~50 sOffloading trades memory for interconnect bandwidth rather than for arithmetic, so the first measurement is bytes per step against the link's throughput. Compare that transfer time with what recomputing the same activations would cost — if recompute is cheaper, the offload was the wrong lever. Then check overlap: copies out should be issued as soon as a tensor is finished and copies back should be prefetched before the backward pass needs them, so the device is never blocked on a transfer. Watch for synchronous copies that serialise against compute, ordinary pageable host allocations that force an extra staging copy instead of a direct asynchronous transfer, and too many tensors offloaded so the link saturates. Also confirm the host is not itself under memory pressure. The practical outcome is usually a mix: offload a few of the largest, cheapest-to-move tensors and recompute the rest.
go deeper
Know that offloading activations to host memory frees device memory but costs a round trip over a much slower link, so it is not simply extra memory for free.
Explain the comparison that decides it: round-trip transfer time against recompute time, and why the copies have to be asynchronous and overlapped rather than sitting on the critical path.
Diagnose from a profile — device idle aligned with copy activity, step time scaling with bytes moved — and propose the mixed configuration of offload some, recompute some, keep some resident.
Own the conclusion nobody wants: when neither recompute nor offload reaches an acceptable step time, the configuration does not fit the device, and batch size, sequence length or model placement is the real decision.
## Two different currencies Recomputation buys memory with arithmetic. Offloading buys memory with interconnect bandwidth: the stored activation is copied from device memory to host memory after it is produced, and copied back when the backward pass needs it. The device keeps neither the tensor nor the obligation to rebuild it — it keeps a round trip over a link. That link is typically far slower than the device's own memory bandwidth, which is the whole reason this can backfire. A tensor that takes microseconds to read from device memory can take an order of magnitude longer to fetch back across the host connection. ## The arithmetic that decides it For each candidate tensor, two quantities: - Transfer time: bytes divided by the achievable link bandwidth, doubled because the tensor goes out and comes back. - Recompute time: the cost of rerunning the operations that produced it from the nearest retained value. Offload wins only when the transfer is faster than the recompute *and* it can be hidden behind other work. If the round trip is slower than rerunning the segment, you have chosen the worse lever and the step time reflects it directly. Large, cheap-to-produce tensors are bad offload candidates precisely because they are the good recompute candidates: many bytes, few operations. ## Why overlap is the usual culprit Even when the arithmetic favours offloading, the copies must not sit on the critical path. Three checks, in order: 1. **Are the copies asynchronous?** A copy issued synchronously blocks until it finishes, converting a hideable transfer into pure added latency. The tell is a device utilisation trace with regular idle gaps that line up exactly with segment boundaries. 2. **Is the copy-out issued early enough?** A tensor should start moving to the host the moment it is no longer needed forward, so the transfer runs underneath subsequent layers' compute. Deferring all copies to the end of the forward pass creates a burst that nothing overlaps with. 3. **Is the copy-back prefetched?** The backward pass consumes offloaded tensors in reverse order, which is perfectly predictable, so each one should be requested several steps of work before it is needed. Fetching on demand guarantees a stall at every consumption point. A fourth, subtler check: how the host buffer was allocated. A transfer from ordinary pageable host memory cannot be driven directly by the copy hardware, so the runtime stages it through an intermediate page-locked buffer — an extra copy, lower throughput, and no true asynchrony. Page-locked host buffers avoid that, at the cost of host memory that cannot be paged out. ## Saturation and the host side The link is a shared, fixed-capacity resource. Offloading more tensors does not make each transfer faster; past the saturation point every additional offloaded tensor lengthens the queue, and step time grows roughly linearly with bytes moved. If the profile shows transfer time scaling with the number of offloaded tensors while device utilisation falls, the link is saturated and the fix is to offload fewer, larger, higher-value tensors and recompute the rest. The host is not free either. Pinning a large working set removes memory the operating system can reclaim, and if the host starts paging, transfer latency becomes wildly variable. Check host memory headroom before blaming the link. ## Reading the profile The diagnosis is usually visible in a timeline. Device idle time that coincides with copy activity means the transfers are on the critical path. Device busy throughout with a longer step means the extra time is real work, so look at recompute instead. Step time that scales with sequence length or batch size faster than compute does points at bytes moved rather than operations performed. ## What good usually looks like The healthy configuration is rarely all-or-nothing. Rank the retained tensors by bytes and by recompute cost, offload the handful that are large and expensive to reproduce, recompute the ones that are cheap to reproduce, and keep resident anything small enough not to matter. That mix keeps the link inside its capacity, keeps the recompute bill modest, and leaves the device fed. And if neither lever gets the model into memory at an acceptable step time, the honest conclusion is that this configuration does not fit this device, and the batch size, the sequence length or the placement of the model across devices is what has to change.
- How do you decide between offloading a tensor and recomputing it?Compare the round-trip transfer time, bytes over achievable link bandwidth in both directions, with the cost of rerunning the operations that produced it. Recompute wins for large tensors that are cheap to produce; offload wins for tensors that are expensive to reproduce and small enough that the link can carry them under cover of other compute.
- Why does prefetching matter more on the backward pass than the forward pass?Because the backward pass is where offloaded tensors are consumed, and its consumption order is exactly the reverse of production order — fully predictable. That predictability means every fetch can be issued well ahead of need; fetching on demand instead guarantees a stall at each consumption point, which is the most common offload performance bug.
- What limits how much you can offload before it stops helping?The link's bandwidth. It is a fixed shared capacity, so once total bytes per step exceed what it can carry inside the step's compute time, each further offloaded tensor adds queueing delay directly to step time. Past that point the correct move is to offload fewer, higher-value tensors and recompute the rest.
It is like putting rarely used files in off-site storage: fine if the courier is fast and you order them the day before, ruinous if you request each one only when you sit down to read it.
saying these in an interview costs you the question
- Treats host memory as free extra device memory
- Fetches offloaded tensors on demand instead of prefetching
- Ignores whether copies overlap with compute
- Offloads large tensors that are trivially cheap to recompute
- Never measures bytes moved per step against link bandwidth