Model Efficiency and Compression
Shrinking a trained network without wrecking it: teacher-student distillation, magnitude and structured pruning, blocks designed cheap from the start. Interviewers probe the accuracy-for-speed trade.
on this pageshowhide
explore
- Knowledge Distillation7 questions
- Soft Targets and Temperature3 questions
- Feature Distillation and Limits4 questions
- Pruning and Quantized Training17 questions
- Magnitude Pruning5 questions
- Structured Channel Pruning4 questions
- Quantization-Aware Training4 questions
- Layer Sensitivity Budgets4 questions
- Designing Small Networks11 questions
- Factorized Layer Design4 questions
- Scaling and Architecture Search3 questions
- Parameters, FLOPs and Latency4 questions
questions
page 2 of 2Do you factorize a trained model and retrain, or train the factorized shape from scratch?
basics
~20 sFactorize-then-fine-tune when trained weights exist and compute is short: the spectrum picks a per-layer rank and a brief retrain recovers most accuracy. Train the factorized shape from scratch when you want low rank learned, not imposed.
When is running an architecture search worth its compute versus scaling a known design?
basics
~20 sAn architecture search pays when its cost amortizes: unusual target hardware, or one space reused across many deployment targets. For a single model on ordinary hardware, scaling a well-tuned published design is cheaper and often as good.
How do you assign per-layer bit-widths across a network under one total size budget?
basics
~20 sSpend bits where the sensitivity curve is steep and take them from where it is flat, equalizing the accuracy lost per bit saved until the size budget is met. Then validate the joint configuration, because per-layer curves were measured independently.
A researcher proposes iterative magnitude pruning with rewinding to ship a smaller model — how do you evaluate it?
basics
~10 sTreat the lottery-ticket result as a claim about trainability, not a deployment technique. Finding the mask means training the dense network repeatedly, and the artifact is scattered sparsity that ships no faster than dense.
To hit 30 fps on a camera stream, would you channel-prune your trained detector or train a narrower one?
basics
~20 sThe answer turns on what you still have: pruning plus fine-tuning is the cheap path, and the only one available without the original data and recipe. For a large cut, a purpose-built narrow architecture trained to convergence usually wins.
showing 31–35 of 35