skip to content

For a bounded physical plant, would you deploy a deterministic actor with injected exploration noise or a maximum-entropy stochastic policy?

level: principalimportance: should knowfreq 40%

answer

  1. where does exploration come from
  2. who tunes it, and per how many plants
  3. state-blind noise versus state-dependent spread
  4. ship one action, not a sample
  5. bounds are not a safety envelope

basics

~20 s

Prefer the entropy-regularised stochastic learner for training, because its exploration is state-dependent and self-tuning, then ship its median action behind an external rate limiter. Choose the deterministic actor only when a fixed, auditable control law and a reproducible exploration process matter more.

solid answer

~60 s

The real question is where exploration comes from and who pays to tune it. A deterministic actor explores only through noise you inject, so someone must choose a noise scale, a correlation structure and a decay schedule per plant, and re-choose them whenever the action scaling changes. That noise is state-blind: on a datacenter cooling controller it jitters the valve setpoints just as hard when the hall is near its thermal limit as when it is idle. An entropy-regularised stochastic policy derives exploration from its objective and auto-tunes the temperature against a target entropy, so it spreads where the value surface is flat and sharpens where it is not — far less per-task tuning, which matters more than any benchmark difference when you run many plants. My default is therefore the stochastic learner, deployed by acting at the distribution's median rather than sampling, with the squash into physical range, a per-step rate limit and an external supervisory envelope outside the policy. I would take the deterministic actor when the shipped artefact must be a fixed, reproducible function and all exploration happened in simulation.

go deeper

for a junior

Recall the basic contrast: a deterministic controller needs noise added to explore, while a stochastic one is random by construction. Know that whatever explores during training, production usually wants one steady action per state.

for a middle

Be able to explain that injected noise is state-blind and hand-scheduled while entropy-driven spread adapts per state and is priced by a tuned temperature. Say how you would turn a stochastic policy into a single deployed action.

for a senior

Show operational judgment: rate limits, a supervisory envelope and a fallback controller outside the policy, plus a clear account of why bounded outputs are not by themselves a safety argument.

for a principal

Own the decision framing and the cost model. Argue from where exploration comes from, who pays the tuning bill across a fleet, and what artefact must be auditable — and be willing to choose the less fashionable option when reproducibility or an existing pipeline outweighs elegance.

## Frame the decision correctly This is not "which algorithm scores higher". Both families are off-policy continuous-control learners with a critic; both bound actions by squashing into the physical range; both can be made to work. The decision axes that actually matter on a plant are: **where exploration comes from**, **who tunes it**, **what is shipped**, and **what stands between the policy and the actuator**. ## Axis 1 — the source of exploration A deterministic actor emits one action per state, forever. Exploration is therefore an external component: noise added to the action during data collection, with a chosen magnitude, a chosen temporal correlation (independent jitter explores very differently from a slowly drifting offset — on a thermal plant only the latter produces a meaningful excursion), and a chosen decay schedule. An entropy-regularised policy has exploration inside the objective. The agent is paid to keep spread, so it keeps it where spread is cheap and gives it up where reward demands precision. The behaviour is state-dependent for free, and the temperature that prices it is auto-tuned against a target entropy rather than scheduled by hand. On a cooling controller this difference is concrete. Fixed-scale noise applies the same excursion to a valve whether the hall has ample thermal headroom or is minutes from an alarm. State-dependent spread is not a safety mechanism, but it is at least *correlated* with how sharply value varies, which fixed noise is not. ## Axis 2 — the tuning bill The deterministic family is the one that costs engineer-days: noise scale, decay, and a critic that is fragile to it. If you operate one plant with a patient team, that is affordable. If you operate fifty heterogeneous plants, per-site noise scheduling is the dominant lifetime cost, and an auto-tuned temperature is the argument that wins. Be honest, though, that auto-tuning replaces the noise schedule with a **target entropy** — one knob instead of three, and a scale-free one, but not zero knobs. A candidate who claims the temperature removes all tuning has oversold it. ## Axis 3 — what you actually ship A plant operator does not want a jittering setpoint in steady state. Whichever family trains the policy, the deployed controller should be a single action per state: - deterministic actor: ship it as-is, noise off; - stochastic policy: ship the **median** action — the squashed centre of the pre-squash distribution — rather than a sample. Because the squash is monotone it preserves quantiles, so that value is available in closed form and is the natural point estimate. If you truly want live exploration on a production plant, that is a separate and much heavier decision requiring a bounded excursion budget and an interlock, not a side effect of leaving sampling switched on. ## Axis 4 — bounds are not a safety case Both families bound their outputs by construction, and it is tempting to call that safe. It is not. The squash guarantees the action lies in the interval you declared; it guarantees nothing about the *sequence* of actions, the rate of change, or behaviour on states outside the training distribution. On a real plant you want, outside the learned policy: - a per-step rate limit on each setpoint; - a supervisory envelope that overrides the policy on measured plant state; - a fallback controller with a defined handover, and monitoring that triggers it. Adding a movement penalty to the reward is worth doing as well, but it is a preference, not a guarantee — the guarantee has to live outside the network. ## Where the deterministic option genuinely wins - **Auditability and reproducibility.** A fixed function from state to action is easier to certify, diff between releases, and reason about than a policy defined as the centre of a learned distribution. If a regulator or a safety review is in the loop, that is real. - **Exploration confined to simulation.** If all data collection happens in a simulator, the noise schedule is a training-time detail nobody has to defend, and the argument about state-dependent exploration on the plant evaporates. - **An existing, tuned pipeline.** A team that already has a working noise schedule for this class of plant should not throw it away for an argument about elegance. ## The answer to give Default to the entropy-regularised learner for training because it removes per-plant exploration tuning and explores in a state-dependent way; deploy it deterministically at its median action; put the rate limiter, envelope and fallback outside the policy where they belong. Switch to the deterministic actor when the shipped artefact must be a fixed auditable control law, or when all exploration is confined to simulation and a tuned schedule already exists. State the trade explicitly rather than declaring one algorithm better — that is what the question is testing.

  • How would you bound actions on a plant where one bad setpoint is expensive?
    Squash into the physical range inside the policy, then put the real guarantees outside it: a per-step rate limit on each setpoint, a supervisory envelope keyed to measured plant state, and a fallback controller with a defined handover. Add a movement penalty to the reward too, but treat it as a preference, not a guarantee — the guarantee cannot live inside the network.
  • How do you pick the exploration noise scale for a deterministic actor?
    Relative to the actuator's meaningful resolution and the plant's response time, with the temporal correlation chosen to match: independent jitter on a slow thermal system explores almost nothing, while a slowly drifting offset produces a real excursion. Then decay it over training. All of it must be redone whenever the action scaling changes, which is the recurring cost the auto-tuned alternative avoids.
  • What would make you keep the deterministic actor despite the tuning cost?
    A shipped artefact that must be a fixed, auditable function of state — easy to diff between releases and defend in a safety review — combined with exploration confined entirely to simulation, which makes the noise schedule a training-time detail nobody has to justify. An existing tuned pipeline for this class of plant is also a legitimate reason not to switch.
  • Does auto-tuning the temperature really remove the tuning burden?
    It reduces it, it does not remove it. You still choose a target entropy, and that choice sets how exploitative the final policy is. The gain is that one scale-free knob replaces a noise magnitude, a correlation structure and a decay schedule, and that it adapts as the task's precision demands change during training instead of needing a human at that transition.

saying these in an interview costs you the question

  • Claims the entropy temperature eliminates all hyperparameter tuning
  • Keeps sampling actions on the live plant after training
  • Treats the squashed action bound as the safety mechanism
  • Assumes a simulator noise schedule transfers unchanged to hardware
  • Picks an algorithm on benchmark scores with no operational argument

context