When would you share one encoder trunk between the actor and critic instead of training two networks?
answer
- who pays for the representation
- two heads write into one encoder
- dense value signal shapes features early
- value loss grows as returns grow
- compare gradient norms, not loss values
basics
~20 sShare a trunk when observations are high-dimensional and the encoder dominates the cost: the value head's dense signal also shapes features early. Keep them separate when the state is small or the value loss swamps the policy gradient.
solid answer
~50 sThe case for sharing is representation cost. If the observation is a large structured state or an image, most parameters and compute sit in the encoder, and training two doubles the bill for nearly identical features. The value head's dense regression signal also shapes those features early, when the policy gradient is still mostly noise. The case against is gradient competition: both heads write into the same trunk, and the value-loss coefficient effectively decides whose features it learns. Value loss scales with the square of the return magnitude, which grows as the agent improves, so a coefficient that was balanced at the start can quietly swamp the policy term later - the symptom is a policy that stops moving while the value curve looks fine. With a low-dimensional state there is little to save and no reason to accept that coupling.
go deeper
Know that the actor and critic can either be two separate networks or one encoder with two heads, and that sharing saves compute because the features are computed once.
Be ready to explain that both heads' gradients add together in a shared encoder, and that the value-loss coefficient decides how much of that encoder serves value prediction rather than the policy.
Show you can diagnose the imbalance in a live run: flat entropy and a stalled return curve while value loss dominates, gradient norms compared at the trunk, and return normalisation tried before coefficient tuning.
Own the tradeoff as an ongoing operational risk, not a one-off architecture pick. Argue where the encoder cost justifies the coupling, what instrumentation must exist before you accept it, and when a team should simply run two networks and spend the tuning budget elsewhere.
## The two designs In the **shared** design, one encoder maps the observation to a hidden representation, and two small heads sit on top: a policy head producing a distribution over actions, and a value head producing a single scalar. In the **separate** design, two independent networks each take the raw observation and produce their own output, and nothing is tied. ## The argument for sharing Consider a datacenter job scheduler whose observation is the current queue - hundreds of pending jobs with resource requests, ages and priorities, plus per-machine utilisation. Encoding that into a useful representation is nearly all the work; deciding where to place the next job, given a good representation, is a small head. Two arguments follow: 1. **Cost.** The encoder dominates parameters, memory and forward-pass time. When you are batching observations from many parallel simulators, one encoder rather than two is a real, immediate saving in both memory and throughput. 2. **Signal quality.** The value head's target is a regression on returns, which is dense and comparatively low-variance. The policy gradient early in training is a high-variance signal built on near-random advantages. The value head therefore teaches the encoder useful structure - which queue states are dangerous, which are comfortable - before the policy gradient has anything coherent to say. ## The argument against sharing Both heads' gradients sum into the same encoder, and there is no guarantee they want the same features. In the joint loss - a policy term, plus a value coefficient times the value loss, minus an entropy coefficient times entropy - that value coefficient is not a cosmetic knob. It sets how much of the encoder serves value regression rather than control. The failure mode to have lived through: value loss grows with the square of the error in predicted return, and return magnitudes grow as the agent improves. On the scheduler, early returns are small and the two terms are comparable; a few hours in, the agent is completing far more jobs, returns are an order of magnitude larger, the squared value error is two orders larger, and the same fixed coefficient now means the encoder is essentially a return predictor with a policy head attached for decoration. The visible symptom is subtle and easy to misread: total loss keeps falling, value loss dominates it, policy entropy goes flat, and the return curve stops climbing. People reach for the learning rate, which is the wrong lever. ## What to do about it In rough order of what to try: 1. **Normalise the return scale.** Keep a running mean and standard deviation of returns and train the value head on the normalised target. This removes the drift at its source and is more robust than chasing the coefficient. 2. **Compare gradient contributions, not loss values.** The number that matters is the norm of each head's gradient measured at the top of the trunk. Two loss values on different scales tell you nothing about who is steering. 3. **Clip the global gradient norm.** This bounds the damage from a value spike but does not fix a persistent imbalance. 4. **Stop the value gradient at the trunk.** The value head then trains on features it does not shape. You keep the compute saving and lose the feature-shaping benefit - a reasonable trade once the encoder is good. 5. **Split the networks.** With a low-dimensional state the encoder is a couple of small layers, sharing saves almost nothing, and separate networks also let you run a larger learning rate on the critic, which usually wants to move faster than the actor. ## How to decide The honest decision rule is about where the cost and the risk actually are. - High-dimensional observations, expensive encoder, many parallel actors: share, and instrument the balance from day one. - Small state vectors, cheap networks: separate, and spend your tuning budget elsewhere. Sharing a trunk to "save parameters" on a ten-dimensional state is a coupling you are buying for nothing. - Anything in between: start separate, because it has fewer failure modes and gives you a clean baseline, then share if profiling says the encoder is the bottleneck. The part worth owning as a lead is that this is not a one-time architecture choice. A shared trunk has a hyperparameter that silently changes meaning as the agent gets better, and unless someone is watching the relative gradient contributions, the run degrades in a way that looks like a policy problem and is actually a plumbing problem.
- How does running many parallel actors change this decision?It pushes toward sharing. With 128 network-routing simulators feeding one synchronous learner, each update sees a large decorrelated batch, so both heads get stable gradients and the coupling is less volatile; and the memory saved by one encoder over hundreds of batched observations is substantial. Asynchronous workers pushing gradients computed under older parameters are the harder case - the trunk absorbs stale updates from both heads at once, and the imbalance is much harder to attribute.
- What single metric tells you the critic is helping the trunk rather than dragging it?The critic's explained variance against observed returns, read together with the ratio of the two heads' gradient norms at the top of the trunk. High explained variance with a balanced gradient ratio means the value head is contributing structure. High explained variance with a lopsided ratio means it has taken over the encoder, and the policy is riding features chosen for prediction rather than control.
- The value loss is swamping the policy loss in a shared-trunk agent. What do you change first?Normalise the return scale before touching the coefficient. The imbalance usually comes from returns growing as the agent improves, which makes squared value error grow faster still; a running normalisation of the value target removes the drift permanently, whereas retuning the coefficient only rebalances for the current stage of training and will drift again.
A shared trunk is two teams sharing one build system. It is a real saving while their needs align, and a permanent negotiation the moment one team's workload grows faster than the other's.
saying these in an interview costs you the question
- Assumes sharing weights is always cheaper and better
- Sets the value-loss coefficient once and never revisits it
- Ignores that return magnitudes grow during training
- Compares raw loss values instead of gradient contributions
- Shares a trunk for a ten-dimensional state to save parameters