When should sparsity be ramped up gradually during training rather than cut once after training?
answer
- one shock versus many small ones
- recovery needs training left to do
- the schedule moves the knee
- prune fast early, slowly late
- compute budget decides as much as accuracy
basics
~20 sRamp gradually when the target sparsity is high enough that a single cut would remove more than the network can recover from. One-shot pruning of the finished model plus a fine-tune is cheaper and adequate at modest sparsity.
solid answer
~50 sOne-shot means taking the trained model, zeroing the target fraction of smallest weights in a single step, and fine-tuning to recover. Gradual magnitude pruning instead raises the sparsity target across training — commonly on a ramp that prunes quickly early and slowly near the end — recomputing the mask at each pruning step and letting normal training run in between, so the network re-fits after every small cut instead of absorbing one large shock. The tradeoff is compute against reachable sparsity: one-shot costs a fine-tune, gradual costs close to a training run, and their accuracy curves are indistinguishable at low sparsity but diverge sharply past the knee. At around half the weights, one-shot is fine; at 90 percent and beyond it typically lands below what gradual reaches. Budget decides the rest — with only a checkpoint and a short fine-tune, one-shot is the only option on the table.
go deeper
Know the two shapes: a single cut on a finished model followed by fine-tuning, versus sparsity rising in steps across training. Be able to say which one costs more compute.
Explain why interleaved training recovers better than one large shock, describe the ramp shape and why it front-loads the cutting, and note that the mask only ever grows.
Show that you check the recovery conditions before drawing conclusions — learning rate available during fine-tuning, a flat final stretch, and a sweep that locates the knee before you commit a training run to one level.
Frame it as a budget allocation: cheap one-shot sweeps to find where the knee is, one expensive gradual run only for the level you ship, and an explicit decision about how often the team can afford to re-target.
## Two schedules for the same target Both procedures end in the same place — a masked network at a chosen sparsity — but they get there differently. **One-shot.** Train to convergence. Compute the threshold for the full target sparsity. Zero everything below it in one step. Fine-tune under the fixed mask until accuracy recovers or stops improving. One pruning event, one recovery phase. **Gradual.** Define a sparsity *schedule*: a function from training step to target sparsity, rising from zero (or from some initial level) to the final target over a span of training, then held flat for a final stretch. At intervals, recompute the mask so that the current target is met — this always means zeroing more weights, since the target only rises — and otherwise train normally. The commonly used ramp shape prunes aggressively at the start, when the network has the most redundancy and the most remaining training in which to recover, and slows as it approaches the final sparsity, when each additional weight removed is more costly. **Iterative prune-and-retrain** sits between them: prune a slice, retrain to convergence, prune again, repeat. It is the most expensive of the three and generally the strongest at extreme sparsity, because every cut gets a full recovery. ## Why gradual wins where it wins Recovery is a training problem. After a cut, the surviving weights must move to compensate, and how far they can move depends on how much training remains and how large the learning rate still is. A single cut to 95 percent removes so much of the function at once that the fine-tune starts from a badly degraded point — often one from which the optimiser cannot climb back, because the loss surface it now lives on is not a small perturbation of the one it was trained on. A sequence of small cuts, each followed by real training, keeps the model near a good solution the whole way. The empirical signature is a sparsity sweep. Take a text-classification encoder and evaluate at 50, 80, 90, 95 and 99 percent global sparsity under both schedules. At 50 percent the two curves sit on top of each other — the model has that much slack and one-shot plus a short fine-tune reclaims it. At 80 percent a gap opens. At 90 the gap is clear. At 95 and 99 one-shot has typically fallen off the cliff while gradual is still degrading smoothly, until it too finds its knee at a higher sparsity. **The value of the schedule is precisely that it moves the knee.** ## The learning-rate trap The most common way a one-shot cut is judged unfairly: pruning at the very end of a decayed schedule and fine-tuning at the final, tiny learning rate. At that rate the surviving weights barely move, so almost nothing is recovered and the conclusion drawn is that the model cannot tolerate the sparsity. Recovery needs a learning rate large enough to reorganise — restarting at a meaningfully higher rate with its own decay, rather than continuing the exhausted tail of the original schedule. Before concluding that a sparsity level is unreachable, check that the recovery phase had room to work. ## Choosing in practice Drive the decision from three inputs. 1. **Target sparsity relative to the knee.** Below the knee, take the cheap option. Near or beyond it, gradual is what buys the level. 2. **Compute budget.** Gradual pruning costs roughly a training run; if you are pruning a model you did not train and cannot afford to retrain, one-shot plus fine-tuning is the realistic plan and you accept a lower reachable sparsity. 3. **How many times you will do this.** Gradual pruning bakes the schedule into training, so re-targeting a different sparsity means another run. One-shot lets you produce several sparsity points from one trained checkpoint cheaply, which is what you want while you are still exploring where the knee is. A reasonable default: sweep sparsity with cheap one-shot cuts to locate the knee, then commit to a gradual run only for the level you intend to ship. ## Details worth stating - The last stretch of a gradual schedule should hold the final sparsity fixed while training continues, so the model ends on pure recovery rather than on a fresh cut. - The mask is monotone: because targets only rise and pruned weights are held at zero, a weight removed early does not return. - Whichever schedule you use, the mask must be reapplied after every update, or the optimizer will refill the zeros.
- Why does a gradual schedule prune fastest at the beginning?Early on the network is at its most redundant, so the first weights removed are the cheapest, and there is the most training left in which to recover from them. As sparsity climbs, each further weight is more likely to be load-bearing and there is less remaining training to absorb the loss — so the ramp slows, and the final stretch is held flat so the model finishes on recovery rather than on a fresh cut.
- You inherit a trained checkpoint and a small fine-tune budget. What sparsity do you promise?Whatever one-shot plus a short recovery can actually reach, which you find by sweeping rather than by quoting a number. Expect it to be well below what a gradual run would deliver on the same model, and make sure the fine-tune uses a learning rate large enough to reorganise the survivors rather than the exhausted tail of the original schedule, or you will under-report what is reachable.
- How does iterative prune-and-retrain differ from a gradual schedule?Iterative pruning interleaves full recovery phases: cut a slice, retrain to convergence, cut again. Gradual pruning interleaves ordinary training steps instead, with cuts happening on a schedule inside a single run. Iterative is the more expensive and usually the stronger at extreme sparsity because every cut gets a complete recovery; gradual gets most of the benefit for roughly the cost of one training run.
saying these in an interview costs you the question
- Claims gradual pruning is always better, ignoring its compute cost
- Fine-tunes at the final decayed learning rate and calls the sparsity unreachable
- Thinks gradual schedules let pruned weights come back
- Picks a schedule without knowing the target sparsity
- Ends a gradual run on a fresh cut with no recovery left