How do you estimate whether a 3D segmentation run on 256-cubed volumes fits in 40 GB?
answer
- Two piles: state is small here
- Voxels times channels times bytes
- One full-resolution tensor is about a gigabyte
- Skip connections stay resident to the end
- Halving each 3D dimension divides by eight
basics
~20 sBudget two piles. Model state, parameter count times bytes per parameter, is often under a gigabyte here. Activations dominate: one 256-cubed feature map with 32 channels in half precision is about 1 GB, and several stay resident.
solid answer
~50 sDo the arithmetic per pile. Model state is parameter count times bytes per parameter — a 30-million-parameter encoder-decoder trained with Adam in 32-bit is about 0.5 GB, which is noise on a 40 GB device. Then count activations for **one** sample: a 256x256x256 volume is 16.8 million voxels, so a single full-resolution feature map with 32 channels in half precision is `16.8e6 * 32 * 2` bytes, roughly 1 GB. A segmentation network holds several such tensors at or near full resolution, including the skip connections it must keep until the decoder consumes them, so per-sample activations land in the high single-digit gigabytes. That is what sets the ceiling — batch 2 or 3, not 32. If it still does not fit, the levers that actually move the number are the input patch size, which is cubic in the spatial dimension, then the channel width of the full-resolution stages, then depth.
go deeper
Practise turning a tensor shape into bytes: multiply every dimension, then multiply by channels and by the bytes per number. Know that a full-resolution 3D feature map is enormous compared with the weights.
Explain why the deep, narrow stages cost little while the shallow, wide ones cost most, and why skip connections keep early tensors alive. Be able to state the exponent attached to batch, resolution and channel width.
Produce the estimate before the run and defend the resulting batch ceiling. An interviewer expects you to name the order in which you would change the shape and to admit the accuracy cost of each change, rather than asking for a bigger device.
Own the framing that input geometry, not parameter count, sets hardware requirements for volumetric and high-resolution work. Be ready to argue a patch-size and resolution policy across projects, including what modelling quality the organisation is trading for it.
## Budget before you launch The estimate has two independent parts, and mixing them up is the usual reason a plan is wrong by an order of magnitude. **Part one: model state.** Count parameters, multiply by bytes per parameter for your optimizer and precision. A convolutional encoder-decoder for volumetric segmentation is typically in the tens of millions of parameters. At 16 bytes per parameter for Adam in 32-bit, 30 million parameters is about 0.5 GB. On a 40 GB device this is a rounding error, and that is the first surprise for anyone whose mental model of memory is "model size". **Part two: activations for a single sample.** This is where the budget actually goes, and it is pure geometry: ``` bytes_for_one_tensor = D * H * W * channels * bytes_per_number ``` For a 256-cubed input: `256^3 = 16,777,216` voxels. With 32 channels at half precision that is `16.78e6 * 32 * 2 = 1.07e9` bytes — about **1 GB for one tensor**. In 32-bit it is 2 GB. Now count how many such tensors are retained. A typical segmentation network runs several convolutions at full resolution before its first downsampling, each producing a feature map that must live until the backward pass reaches it. The encoder halves the spatial dimensions as it goes, and in 3D each halving divides the voxel count by 8 — so the deeper stages are cheap even when they double the channel count, because 8x fewer voxels beats 2x more channels. But the decoder's skip connections force the full-resolution encoder tensors to stay resident all the way through, and the decoder produces its own full-resolution tensors on the way back up. Sum it and per-sample activations for this shape land somewhere in the high single-digit gigabytes. Against 40 GB minus 0.5 GB of state, that is a batch of two or three. The plan that assumed batch 32 was never physically possible. ## The shape of the scaling Understanding which exponent attaches to which knob is the whole skill here: - **Batch size** — linear. Doubling the batch doubles the activation pile exactly. - **Spatial resolution in 3D** — cubic. Training on 128-cubed patches instead of 256-cubed volumes divides per-sample activations by **8**, not by 2. This is the single largest lever available on a volumetric model, and it is why patch-based training is the norm for this kind of data. - **Spatial resolution in 2D** — quadratic. Going from 384-pixel to 224-pixel inputs on an image encoder cuts per-sample activations by `(384/224)^2`, about 2.9x. Going the other way multiplies them by the same factor, which is why "we just bumped the resolution" so often ends a run that used to fit. - **Sequence or frame count** — linear for convolutional and recurrent sequence models. Doubling a spectrogram's frame count doubles every intermediate tensor while the parameter count is untouched, so an audio model that fit on short clips can fail on long ones with no change to the architecture at all. - **Channel width** — linear in activations at the stage you widen, but the parameter count of a convolution grows with the product of input and output channels. Widening the full-resolution stage is therefore expensive on both piles at once, and widening the deepest stage is nearly free on activations. - **Depth** — roughly linear in the number of retained tensors, but the marginal layer's cost depends on where it sits. Adding blocks at the lowest resolution costs very little memory; adding them at full resolution costs a gigabyte each in the example above. ## Which knob to turn Given a shape that does not fit, the order that follows from the exponents is: **cut the input geometry first** (patch size, or resolution), because it is cubic in 3D and quadratic in 2D; then **narrow the full-resolution stages**, because those tensors are the largest; then **move depth from high-resolution stages to low-resolution ones**, which changes capacity distribution rather than removing it; and treat **batch size** as the fine adjustment, because once you are at two or three there is almost nothing left to give. Each of these is a modelling decision with accuracy consequences, not a free memory trick. Patch-based training limits the spatial context a single prediction can see. Narrowing the full-resolution stage reduces the capacity available for fine boundary detail, which is exactly what a segmentation task cares about. The estimate is what lets you have that conversation with numbers instead of guesses. ## Sanity-checking the estimate The arithmetic gets you to the right order of magnitude, which is all a plan needs; the residual error comes from tensors you did not think to count and from scratch space that individual layers need transiently. Confirm by running a single training step at batch 1 with the target input shape and reading peak memory, then extrapolate the per-sample slope by trying batch 2. Two measurements give you both the intercept (state plus fixed overhead) and the slope (per-sample activations), and from there the largest batch that fits is a division. The conclusion to carry away: for high-resolution and volumetric work the parameter count tells you almost nothing about whether a run fits. **The input geometry does**, and it does so with an exponent equal to the number of spatial dimensions you are scaling.
- Between halving the batch and halving each spatial dimension, which saves more?Halving the spatial dimensions, by a wide margin, on a 3D input. Batch size is linear, so halving it saves a factor of two. Halving depth, height and width divides the voxel count by eight, so per-sample activations fall roughly 8x. On 2D images the same comparison is 2x versus about 4x.
- You cut to batch size one and it still does not fit. What does that tell you?That the floor itself is too high: model state plus one sample's retained activations exceeds the device. No batch-side adjustment can help, because there is no batch left. The shape has to change — smaller input patches, narrower full-resolution stages, or fewer layers operating at full resolution.
- Does doubling the channel width of the first full-resolution block double that block's activations?Yes, activation bytes at that stage are linear in channel count, so a 1 GB tensor becomes 2 GB. The parameter cost grows differently: a convolution's weight count scales with input channels times output channels, so widening squares that term — but on this shape the parameters remain small next to the activations.
- Why does the parameter count mislead people on this kind of model?Because it sizes only the fixed pile. A 30-million-parameter segmentation network needs well under a gigabyte of weights, gradients and optimizer buffers, while one 256-cubed sample can need ten times that in retained feature maps. For high-resolution and volumetric inputs, input geometry rather than model size determines whether the run starts.
saying these in an interview costs you the question
- Sizes the run from the weight count alone
- Treats 3D resolution as scaling linearly
- Forgets skip connections stay resident
- Assumes a bigger device always fixes it
- Cannot convert a tensor shape into bytes