skip to content

Would you serve one shared-trunk model with steering and brake heads, or two separate models?

level: principalimportance: nice to knowfreq 36%

answer

  1. Trunk sums gradients from both heads
  2. One forward pass, one release
  3. Capacity competition between unrelated tasks
  4. Always build the single-task baselines
  5. Task-specific branches as a middle ground

basics

~20 s

Share a trunk when the two tasks need the same perception features and you want one forward pass; keep them separate when they need independent release cadences. Sharing buys compute and transfer, and costs you a coupled artefact where retraining for one task can regress the other.

solid answer

~50 s

For a driving camera the perception features are genuinely common — both a continuous steering-angle regression head and a binary brake/no-brake sigmoid head need lane geometry, obstacle position and distance — so one trunk with two heads is the natural first design. It gives you a single forward pass instead of two, which is decisive on an embedded budget, and each task's data acts as extra signal shaping the shared representation. What you pay for is coupling: the trunk receives the summed gradients from both heads, so the tasks compete for representational capacity, and a retrain motivated by braking can quietly move steering. Operationally the two tasks now share one artefact, one release and one rollback, which is an organisational decision as much as a modelling one. I would share the trunk, keep per-head evaluation gates so neither task can ship a regression in the other, and split into separate models if the tasks' data cadences or accountability diverge enough that the joint release becomes the bottleneck.

go deeper

for a junior

Know what a multi-head model is: one trunk producing shared features, with a small separate output layer per task. Recall that the heads can have different shapes, such as one continuous output and one sigmoid.

for a middle

Explain the mechanics — both heads send gradients back into the shared trunk and those gradients add, so the trunk must find one representation serving both tasks. Name the compute saving and the capacity-competition risk.

for a senior

Show you would build single-task baselines before trusting the joint model, gate releases on per-task metrics rather than an aggregate, and recognise coupled retraining as the operational cost. Be able to describe the freeze-trunk and task-branch middle grounds.

for a principal

Own the decision as an organisational one as much as a modelling one: who owns the shared artefact, whether both tasks can live on one release train, and what evidence would trigger splitting the model back apart later.

## The structure A multi-head model is one trunk — the stack of layers that maps the input to a hidden representation — with two or more small heads reading off the same final hidden vector. Here the trunk consumes a camera frame; one head is a width-1 linear regression head emitting a continuous steering angle, the other is a single sigmoid unit emitting the probability of braking. The heads are tiny relative to the trunk, typically one linear layer each, which is exactly why sharing is attractive: nearly all the compute is in the part both tasks need. During the backward pass, each head produces gradients for its own parameters and also a gradient with respect to the shared hidden vector. Those flow back into the trunk and **sum**. This is the single mechanical fact everything else follows from: the trunk is not trained for one task and then the other, it is trained by both simultaneously, and it can only find one representation that serves both. ## What sharing buys **Compute and latency.** One trunk evaluation per frame instead of two. On an embedded compute budget where the trunk dominates cost, this can be the difference between shipping and not. **Transfer and regularisation.** Braking-labelled frames teach the trunk about obstacles, which is also useful for steering; steering data teaches lane geometry, which is contextual for braking. When tasks are genuinely related, each acts as auxiliary signal for the other and the shared representation generalises better than either task's data alone would support — most visibly when one task has far less labelled data than the other. **One thing to maintain.** A single training pipeline, a single set of input preprocessing assumptions, a single artefact to version and monitor. ## What sharing costs **Capacity competition.** The trunk has a fixed budget of representation. If the tasks want different features — one needs fine-grained texture, the other needs coarse layout — they fight, and both can end up worse than dedicated models. That degradation is the multi-task failure everyone hits eventually, and it is not detectable from a single aggregate number. **Coupled releases.** A retrain to improve braking ships a new trunk, therefore a new steering model, whether or not anybody wanted one. Every release must be gated on **both** tasks' metrics separately, and a rollback rolls back both. If the two tasks are owned by different teams with different risk tolerances, this is where the design starts hurting. **Harder attribution.** When steering degrades, the cause may be a steering data change, a braking data change, a shift in how the two objectives balance, or the trunk itself. Two independent models make regressions trivially attributable; one shared model does not. **Divergent data cadence.** If braking labels arrive continuously and steering labels arrive quarterly, a joint retrain schedule serves neither well. ## Middle grounds The choice is not binary. You can share the lower trunk and give each task its own branch of a few task-specific layers before its head — cheaper than two full models, with room for divergent features. You can freeze the trunk after a joint training phase and retrain only one head, which decouples releases at the cost of the trunk no longer improving. You can start with a shared trunk and split later once one task's requirements outgrow it; going the other direction, merging two mature independent models, is usually harder because their input assumptions have drifted apart. ## How to decide Make it a two-axis judgment. *Task relatedness*: do the tasks plausibly need the same features, and does joint training measurably beat two dedicated baselines on both tasks? Always build the two single-task baselines — a shared trunk that loses to them on both tasks is not worth its coupling, and without the baselines you will never know. *Organisational coupling*: can the two tasks tolerate one release train, one rollback, one on-call rotation, and one owner for the balance between the objectives? If the answer is no, the modelling argument rarely wins — a technically superior joint model that no team can safely retrain is worse in practice than two independent models that ship. My default for tightly related perception tasks under a hard latency budget is: share the trunk, add short task-specific branches, gate every release on per-task metrics, and treat a persistent regression in one task as the trigger to split.

  • What evidence would convince you the shared trunk is hurting one of the tasks?
    A dedicated single-task baseline beating the shared model on that task's own held-out metric, reproduced across seeds. Without that baseline you cannot distinguish task interference from the task simply being hard. Supporting evidence: the regression tracking changes in the other task's data rather than its own.
  • How do you stop a retrain aimed at braking from silently regressing steering?
    Gate the release on per-task metrics evaluated separately, each with its own frozen held-out set and its own acceptance threshold, and block the ship if either fails. An aggregate score hides exactly this failure. Pair it with shadow evaluation of the steering head before the new trunk takes traffic.
  • If you must decouple releases but keep one forward pass, what do you do?
    Freeze the trunk and retrain only the head that needs to change. Releases decouple because the shared computation is fixed, and serving cost is unchanged. The price is that the trunk stops improving, so it works as a stabilising measure between periodic joint retrains rather than forever.

A shared trunk is two products on one factory line: cheaper per unit and the tooling improves both, but you cannot retool for one product without stopping the other.

saying these in an interview costs you the question

  • Assumes sharing a trunk always improves both tasks
  • Ships a joint model without single-task baselines to compare
  • Gates releases on one aggregate score across both tasks
  • Ignores that one retrain now redeploys both tasks
  • Treats the choice as purely technical, ignoring release ownership

context