Two tasks' gradients meet a shared trunk at negative cosine similarity — how do you decide what to sacrifice?
answer
- direction disagreement, not size disagreement
- cosine on the shared parameters
- sustained, not one batch
- no positive weight flips a sign
- remove the opposing component by projection
basics
~20 sNegative cosine between two task gradients means an update helping one hurts the other, and no choice of weights removes that. Name the task the product is judged on, fix the degradation you will accept elsewhere, then pick a mechanism to enforce it.
solid answer
~50 sFirst confirm it is real: compute each task's gradient with respect to the shared parameters and track the cosine between them over many batches, not one. A persistently negative cosine says the trunk cannot serve both tasks with the same features — that is a representation problem, not a scaling one, so retuning weights only changes which task loses. My decision starts from the product: name the task whose metric the business is judged on, state in advance how much I will let the other degrade, and make that the acceptance criterion instead of the total loss. The mechanisms are then ordinary: downweight the conflicting term, project one task's gradient onto the plane normal to the other so the directly opposing component is removed, change how often each task is sampled, or drop the term. The obligation is to report the trade-off, not to bury it in a scalar.
go deeper
Know that two tasks trained through one shared network can pull its parameters in opposite directions, so an update that improves one task can measurably worsen the other.
Be able to define the conflict as a negative dot product between per-task gradients on the shared parameters, and explain why a positive scalar weight can rebalance sizes but never reverse a direction.
Show how you would instrument a live run to confirm sustained conflict cheaply, and what you would try first — downweighting, gradient projection, or a sampling change — with a per-task metric to judge whether it worked.
Own the trade-off itself: name the task the product is judged on, state the degradation you will accept elsewhere before the run, and make that exchange rate a reviewed decision rather than a by-product of a weight sweep.
## Two different failures that look alike When a multi-task model underperforms its single-task baselines, there are two distinct causes and they need different responses. - **Magnitude imbalance.** One task's gradient at the shared parameters is far larger than another's, so the small one is effectively ignored. This is a scaling problem and weights fix it. - **Direction conflict.** The tasks' gradients point in substantially opposing directions: their dot product at the shared parameters is negative. The step that reduces one task's loss increases the other's. **No scalar weighting removes this**, because a positive weight cannot change a direction — it only chooses which task the compromise favours. The cosine similarity between per-task gradients on the shared parameters is what separates the two. It is also the quantity to instrument, and confusing the two failures is the single most common mistake in this area. ## Measuring it honestly Computing per-task gradients means one backward pass per task, so measuring conflict continuously roughly multiplies the backward cost by the number of tasks. In practice you sample: every few hundred steps, on a fixed held-out batch or a small window of batches, restricted to the shared parameters (often just the last shared block, which is where the representation is most task-specific and conflict shows up first). Two cautions about reading the number: - **A single mini-batch cosine is noise.** Gradients on one batch are high-variance; conflict means a *sustained* negative average over a window. - **Early conflict is normal.** While the trunk is still learning generic features, tasks disagree about where to go and then stop disagreeing. What matters is conflict that persists once the representation has stabilised, and whether it coincides with a task's validation metric flattening below its single-task baseline. ## The mechanisms available - **Downweight the conflicting term.** Cheap, honest, and it makes the sacrifice explicit: you have decided which task pays. It does not recover both tasks; it chooses. - **Gradient projection.** When two task gradients have a negative dot product, replace one with its projection onto the plane normal to the other: `g_i <- g_i - (g_i . g_j / ||g_j||^2) * g_j`. This is the PCGrad recipe, applied pairwise in random order. It removes the component that directly opposes the other task rather than removing the whole term. It changes the update direction; it does not make the tasks compatible, and it does not deliver single-task performance on both. - **Change the sampling.** Alternating updates or sampling tasks with different frequencies interleaves the conflict in time rather than resolving it: the trunk drifts toward whichever task updated most recently and most often. It is useful when the tasks have very different data volumes, and it makes convergence harder to reason about. - **Remove the term.** If a task conflicts and no one is judged on it, deleting it from the objective is a legitimate and underused answer. ## The decision that is actually being asked for Every one of those mechanisms picks a point on a trade-off surface. The lead-level content of this question is not which mechanism you name, it is whether you make the choice explicitly and in advance: 1. **Name the primary.** Which task's metric is the product judged on? If the answer is "both equally", push until there is a real answer or an explicit exchange rate. 2. **State the budget before the run.** "We will accept up to X worse on the secondary task to hold the primary at Y" is an acceptance criterion. "We picked the weights with the lowest total loss" is not. 3. **Report per task, always.** An aggregate — total loss, or an average of normalised per-task scores — hides which task paid. Two runs with identical aggregates can be very different products. 4. **Re-open the question when data changes.** Conflict is a property of the tasks *and* the data mix. A new data source can turn a conflicting pair into a cooperating one, or the reverse, so the trade-off is a decision with a review date rather than a constant. ## What a strong answer sounds like A strong answer distinguishes the two failure modes, describes measuring sustained cosine rather than a single batch, names at least one mechanism beyond reweighting and is honest that it mitigates rather than resolves, and finishes on the product judgment: who decided which task loses, how much, and where that is written down.
- How would you measure gradient conflict without roughly doubling training cost?Sample instead of streaming: every few hundred steps, take per-task gradients on one fixed held-out batch and restrict the comparison to the last shared block. Average the cosine over a window before acting on it. That turns a per-step cost into a negligible one and removes the mini-batch noise that makes single-step cosines uninterpretable.
- If two tasks conflict, why not simply alternate updates between them?Alternating does not remove the conflict, it interleaves it. The trunk drifts toward whichever task updated most recently, so the final parameters depend on the schedule's phase, and convergence becomes harder to reason about and to reproduce. It is genuinely useful when the tasks have very different data volumes, but as a fix for opposing gradients it hides the trade-off rather than deciding it.
- What should you conclude if the cosine is negative early in training and positive later?Usually nothing. Early in training the trunk is learning generic features and the tasks disagree about where to go; that resolves on its own. Judge on sustained conflict after the representation has stabilised, and only treat it as actionable when it coincides with a task's validation metric flattening below what that task reaches on its own.
- Does gradient projection let both tasks match their single-task results?No. Projection removes the component of one task's gradient that directly opposes another's, which reduces destructive interference, but the tasks still compete for the same shared capacity and the same representation. Expect a better point on the trade-off surface, not the elimination of the trade-off — and validate that claim with per-task metrics against single-task baselines.
Two people carrying one table in different directions: pulling harder does not help. Someone has to decide where the table is going and say so out loud.
saying these in an interview costs you the question
- Assumes retuning loss weights can cancel a direction conflict
- Reads one mini-batch's negative cosine as proof
- Reports only aggregate loss and hides which task got worse
- Believes projection recovers single-task performance on both
- Has no stated exchange rate between the competing tasks