Do you factorize a trained model and retrain, or train the factorized shape from scratch?
answer
- imposed versus learned low rank
- evidence is cheap, training runs are not
- a well-conditioned layer never yields
- ranks guessed before any evidence
- the product needs its own initialization
basics
~20 sFactorize-then-fine-tune when trained weights exist and compute is short: the spectrum picks a per-layer rank and a brief retrain recovers most accuracy. Train the factorized shape from scratch when you want low rank learned, not imposed.
solid answer
~50 sPost-hoc factorization is evidence-driven and cheap. The trained weights are already there, so you can measure each layer's spectrum and sensitivity, allocate rank non-uniformly, and pay only for a short recovery fine-tune. Its ceiling is how factorizable those weights happen to be: a well-conditioned layer will not yield at any rank that saves parameters, and no fine-tune fixes a bad starting approximation. Training the factorized shape from scratch removes that ceiling, because the network optimises inside the low-rank family from step one instead of losing a dense solution it already learned. The costs are a full training run, ranks fixed before any per-layer evidence exists, and a product parameterization whose initialization scaling needs care. Usually do both in order: a post-hoc study to learn which layers tolerate what rank, then bake that profile into a from-scratch run if the target justifies the budget.
go deeper
Know that a factorized layer can either be created by squeezing an already trained matrix or built that way from the start, and that the second requires training the whole network again.
Explain why post-hoc factorization is capped by how compressible the trained weights happen to be, and why a product of two matrices needs its initialization scaled differently from a single dense layer.
Show you would run the post-hoc study first as measurement, allocate rank per layer from it, and only then decide whether a from-scratch run is worth its compute. Be specific about what the recovery fine-tune needs in data and schedule.
Frame this as a compute-allocation call with an accuracy budget attached, argue from how the target is worded - within X% of today's model versus best model under N parameters - and be ready to say when the size target is not the binding constraint at all.
## The two routes **Route A - factorize afterwards.** Train the dense model normally. Then, layer by layer, replace the trained weight matrix with the best low-rank product at a chosen rank, and fine-tune the whole thing briefly to recover. **Route B - factorize in the architecture.** Decide the ranks up front, define layers as a pair of thin matrices with a narrow waist, and train that network from random initialization. There is no dense model at any point. The two answer different questions. Route A asks "how much of this trained function survives being squeezed?" Route B asks "how good a model can I train inside this budget?" ## What Route A gives you Evidence, and cheaply. Every decision can be measured before you commit: which layers hold the parameters, how each layer's singular values decay, how much validation loss moves when exactly one layer is factorized at a candidate rank. That lets you allocate the compression budget non-uniformly - which is nearly always better than a single global rank - and it lets you stop early if the accuracy curve turns bad. The compute is a fraction of a training run. Its ceiling is the trained weights themselves. Nothing in ordinary training pushes a layer toward low rank, so whether a given layer is compressible is largely luck plus the incidental effects of the regularization and architecture used. When a layer comes out well conditioned, the best rank-r approximation is genuinely far from it, the initial accuracy drop is large, and a short fine-tune starting from a bad point does not rescue it. You can be left with a model that is only mildly smaller because the layers that hold the parameters were the ones that refused. A related asymmetry: fine-tuning after factorization is not free of risk either. It re-touches every weight, so it needs the original training data or a good proxy, the original augmentation and schedule sensibilities, and a low enough learning rate that the model does not drift on distributions the compression was validated against. ## What Route B gives you The network is never asked to give anything up, because it never had it. Optimization happens inside the rank-r family from step one, and the model finds the best solution available there rather than the nearest one to a dense solution. When the target size is far below what post-hoc squeezing achieves, this is usually the only way to hit it at acceptable accuracy. The costs are: a full training run per rank configuration you want to compare, so exploring the design space is expensive; and the ranks must be chosen before you have per-layer evidence, which pushes you toward uniform or heuristic allocation - exactly the thing Route A is good at avoiding. There is also an optimization subtlety. A layer parameterized as a product of two matrices is not the same optimization problem as the dense layer it can express. Initialize both factors as though each were a standalone dense layer and the product's scale is wrong, so activations shrink or blow up through the waist. Gradients to each factor are scaled by the other, so the two can drift out of balance during training - one growing while the other shrinks - which changes the effective step size on the composed map even though the function is unchanged by such a rescaling. The fixes are ordinary but must be deliberate: scale the initialization so that the *product* preserves activation variance, and consider a normalization layer or a weight-decay scheme that treats the pair as a unit. ## How to decide Ask four questions in order. 1. **Does a trained model exist, and can you afford another full run?** If there is no spare training budget, the choice is made for you. 2. **How aggressive is the target?** A 2x parameter cut on over-parameterized layers is usually reachable post-hoc. A 5x cut, or a target on layers that turn out well conditioned, usually is not. 3. **How is the accuracy budget worded?** "Within 1% of the current model" favours Route A, which is measured against exactly that model. "Best model under N parameters" favours Route B, which optimises directly for it. 4. **Will you retrain regularly anyway?** Teams with a routine retraining pipeline pay much less for Route B, because the from-scratch run is a run they were going to do. Teams shipping a one-off model pay full price. The hybrid that usually wins: run the post-hoc study first as *analysis*, even if you intend to train from scratch. It costs little and gives you a defensible per-layer rank allocation to bake into the architecture, replacing a guessed uniform rank with a measured profile. Then train that shape from scratch and compare it against the post-hoc model you already have on the same validation set. ## The framing to own This is a compute-allocation decision as much as a modelling one. The technical difference - imposed low rank versus learned low rank - is real, but the question a lead actually answers is where a fixed engineering and compute budget buys the most accuracy per unit of size, and whether the size target is even the constraint that binds.
- What makes a from-scratch factorized layer harder to optimize than the dense layer it replaces?It is a product of two matrices, so the composed scale depends on both. Initializing each factor as a standalone dense layer gives the product the wrong variance, and activations shrink or grow through the waist. Gradients to each factor are scaled by the other, so the pair can drift out of balance during training. Scale the initialization so the product preserves activation variance, and treat the pair as one unit for weight decay.
- How do you choose ranks for a from-scratch run when there is no trained model to inspect?Work backwards from the size target and the break-even rank per layer, then refine with borrowed evidence: a post-hoc factorization study on any related trained model, or a short pilot run at reduced steps. Even an imperfect per-layer profile beats a single uniform rank, because the parameters and the redundancy are both concentrated in a few wide layers. Budget one comparison run for the profile you are least sure about.
- When is neither route the right answer?When the parameters are not where you assumed - count them per layer before designing anything, because factorizing layers holding a few percent of the model buys nothing. Also when the accuracy budget is effectively zero, or when the deployment constraint that actually binds is not the one factorization moves. Confirm which cost metric your target names before spending a training run on the wrong axis.
saying these in an interview costs you the question
- Assumes training from scratch always beats factorizing afterwards
- Picks a single uniform rank without any per-layer evidence
- Skips the recovery fine-tune after post-hoc factorization
- Initializes both factors as if each were a full dense layer
- Treats it as a pure modelling call and ignores the compute budget
- Expects fine-tuning to rescue a poor low-rank starting approximation