skip to content

Model Efficiency and Compression

Shrinking a trained network without wrecking it: teacher-student distillation, magnitude and structured pruning, blocks designed cheap from the start. Interviewers probe the accuracy-for-speed trade.

on this pageshow

explore

questions

page 2 of 2

Do you factorize a trained model and retrain, or train the factorized shape from scratch?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Factorize-then-fine-tune when trained weights exist and compute is short: the spectrum picks a per-layer rank and a brief retrain recovers most accuracy. Train the factorized shape from scratch when you want low rank learned, not imposed.

open as a page

When is running an architecture search worth its compute versus scaling a known design?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

An architecture search pays when its cost amortizes: unusual target hardware, or one space reused across many deployment targets. For a single model on ordinary hardware, scaling a well-tuned published design is cheaper and often as good.

open as a page

How do you assign per-layer bit-widths across a network under one total size budget?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Spend bits where the sensitivity curve is steep and take them from where it is flat, equalizing the accuracy lost per bit saved until the size budget is met. Then validate the joint configuration, because per-layer curves were measured independently.

open as a page

A researcher proposes iterative magnitude pruning with rewinding to ship a smaller model — how do you evaluate it?

level: principalimportance: nice to knowfreq 26%

basics

~10 s

Treat the lottery-ticket result as a claim about trainability, not a deployment technique. Finding the mask means training the dense network repeatedly, and the artifact is scattered sparsity that ships no faster than dense.

open as a page

To hit 30 fps on a camera stream, would you channel-prune your trained detector or train a narrower one?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

The answer turns on what you still have: pruning plus fine-tuning is the cheap path, and the only one available without the original data and recipe. For a large cut, a purpose-built narrow architecture trained to convergence usually wins.

open as a page

showing 31–35 of 35