Why does averaging the weights of two independently initialized runs produce a broken model?
answer
- same function, different labelling
- swapping hidden units changes nothing
- the straight line between them matters
- high-loss barrier at the midpoint
- one basin, not merely one run
basics
~20 sThe two runs sit in unrelated regions of a non-convex loss surface, with hidden units learned in different orders and different arrangements. Coordinate-wise averaging mixes unrelated features, and the midpoint has high loss. Averaging only works along one trajectory.
solid answer
~50 sWeight averaging is only meaningful when the points being averaged lie in one low-loss region connected by a straight line. Two runs from different random initializations do not: a network has enormous permutation symmetry - relabelling hidden units together with their incoming and outgoing weights leaves the function identical - so unit 7 in one run has no relationship to unit 7 in the other. Averaging coordinate-wise therefore blends unrelated features, and the loss along the straight line between the two solutions rises to a high barrier in the middle, often near chance. Checkpoints from one run, taken after the trajectory has settled, stay inside one such region, so their mean stays low-loss - which is why stochastic weight averaging harvests its checkpoints from a single run. The cheap test before you commit is to evaluate the loss at a few interpolation points between the two weight vectors and look for a bump.
go deeper
Know the boundary rule: weights may be averaged across checkpoints of the same training run, never across two runs started from different random initializations.
Explain permutation symmetry - relabelling hidden units with their incoming and outgoing weights leaves the function unchanged - and why that makes coordinate-wise averaging meaningless across runs.
Show you would verify rather than assume: interpolate between the two weight vectors, evaluate loss along the path with statistics recomputed, and decide from the curve whether averaging is legal.
Own the distinction when a team proposes 'merging our models': parameter merging needs a shared trajectory and delivers one-model inference cost, while output ensembling needs diversity and multiplies serving cost.
## The claim being tested Averaging parameters assumes something strong and usually unstated: that the straight line between the points you are averaging stays in a low-loss region, so the midpoint is a good model. Along one training trajectory late in a run this tends to hold. Between two independent runs it does not, and the failure is dramatic rather than gradual - the average is typically no better than a random network. ## Permutation symmetry Take any hidden layer and swap two units: exchange their incoming weight rows, their biases and their outgoing weight columns. The network computes **exactly the same function**. A layer of `n` units therefore admits `n!` weight vectors that are functionally identical, and a deep network multiplies these across layers - an astronomically large set of equivalent parameterizations of the same model. Two independent runs, with different initializations and different data orders, land in arbitrary members of these equivalence classes. Even if they learned *the same features*, the run-A vector might place an edge detector at index 7 and the run-B vector at index 214. Coordinate-wise averaging adds the run-A edge detector to whatever run B happened to put at index 7 - a completely unrelated feature - and halves both. Every unit in the averaged network becomes a meaningless blend, and the downstream layers, whose weights were also blended, are reading from features that no longer exist. Symmetry is not the whole story - scale symmetries between layers, and the simple fact that the two runs may have found genuinely different solutions, contribute too - but permutation is the cleanest way to see why coordinate-wise averaging has no reason to work. ## The barrier, and how to measure it The operational version of this is **linear mode connectivity**. Take weight vectors `w_A` and `w_B` and evaluate the training loss at `w(a) = (1 - a) * w_A + a * w_B` for `a` in `0, 0.1, ..., 1.0`. If the curve stays flat or dips, the two points are linearly connected and averaging is safe. If it spikes in the middle, they are in different basins and the average is garbage. This costs a handful of evaluation passes and is worth running once before you build any averaging procedure on an assumption you have not checked. Note that you must recompute any normalization running statistics at each interpolation point, or you will measure a barrier that is really a buffer mismatch. Two important refinements: - **A shared prefix buys connectivity.** If both runs share the first part of training - the same initialization and the same early steps - and only diverge afterwards, the two endpoints often *are* linearly connected, and averaging works. This is the mechanism behind averaging several fine-tunes of one shared pretrained starting point: because they start from a common point and move relatively little, they usually stay in one region, and the average of them can beat any individual fine-tune. The rule is not 'one run' but 'one basin', and a shared trajectory is the practical way to guarantee it. - **Alignment can remove the barrier.** Research on permuting one network's units to match the other's shows that after such an alignment the barrier between two independently trained networks can be greatly reduced. That is a real result, but it requires solving the matching problem explicitly; it is not what happens if you simply add two checkpoints together. ## Contrast with prediction averaging The reason this surprises people is that averaging the **outputs** of two independent runs works fine, and works better the more independent the runs are - diversity is the whole point of an ensemble. Parameter averaging has the opposite requirement: it needs the models to be nearly the same model, differing only by the noise you are trying to cancel. When someone proposes 'averaging the models', the first clarifying question is which of the two operations they mean, because their preconditions are opposites. ## Practical rules 1. Average checkpoints from a single run, harvested after the trajectory has settled. 2. Or average fine-tunes that share one initialization and one early trajectory - and verify with an interpolation check. 3. Never average across different seeds, different architectures, or different layer widths. 4. Recompute normalization statistics after averaging, so a genuine barrier is not confused with a buffer mismatch. 5. If you want the benefit of independent runs, ensemble their predictions instead, and pay the inference cost.
- How would you test cheaply whether two weight vectors can be averaged at all?Interpolate and evaluate. Compute `(1 - a) * w_A + a * w_B` for a few values of `a` between 0 and 1 and measure training loss at each, recomputing normalization statistics at each point. A flat or dipping curve means the two are linearly connected and averaging is safe; a spike in the middle means they are in different basins and the average will be useless. It costs a handful of evaluation passes.
- Are there cases where averaging weights across different runs does work?Yes, when the runs share a trajectory. Several fine-tunes launched from one common pretrained starting point usually stay in the same region, so their average can beat any individual one. The condition is a shared initialization and limited movement, not literally a single run - and it is still worth confirming with an interpolation check rather than assumed.
- Why does averaging the predictions of two independent runs work when averaging their weights does not?Because the two operations want opposite things. Output averaging benefits from diversity: uncorrelated errors partly cancel, so the more independent the runs, the better. Parameter averaging needs the points to be nearly the same model differing only by optimization noise, so independence is exactly what destroys it. The cost differs too - the ensemble pays two forward passes, the average pays one.
Two people each write the same recipe but number the steps in a different order. Averaging the two numbered lists line by line gives you half of 'chop the onion' mixed with half of 'preheat the oven'.
saying these in an interview costs you the question
- Treats weight averaging and prediction averaging as interchangeable
- Blames a learning-rate or batch-size mismatch between runs
- Says it fails only because the data order differed
- Expects averaging two 90-percent models to give about 90 percent
- Proposes averaging models with different layer widths
- Never checks whether the two points are linearly connected