Your value agent's Q-values climb without bound - which leg of the deadly triad do you relax?
answer
- first prove it is not a legitimate value
- compare against r_max over one minus gamma
- three levers, each with a real price
- longer returns dilute bootstrapping continuously
- the deployment decides which leg is negotiable
basics
~20 sFirst confirm the growth is divergence, not a legitimately large return, by comparing against the maximum possible discounted value. Then relax whichever leg your problem can afford: on-policy data, longer or full returns instead of bootstrapping, or a simpler representation.
solid answer
~50 sStart by ruling out a non-triad cause. With rewards bounded by r_max, no action value can legitimately exceed r_max / (1 - gamma), so compare against that ceiling before concluding anything; unclipped reward scale, a missing terminal condition and a discount too close to one all mimic divergence. If it really is instability, you have three levers and each costs something specific. Dropping off-policy learning - training on data from the policy being evaluated - restores the convergence result but throws away data reuse and sample efficiency. Dropping bootstrapping, by moving to long n-step or full Monte-Carlo returns, turns the update into honest regression at the price of variance and a need for episodes to end. Dropping or weakening approximation only works when engineered features or a coarse discretisation can carry the task. For a power-grid dispatch agent that must learn a greedy target policy from logged exploratory operations, off-policy is not negotiable, so the realistic move is longer returns plus a lower discount.
go deeper
Know that value estimates have a theoretical maximum set by the reward magnitude and the discount, and that predictions above it mean something is wrong.
Be able to name the three levers and what each costs: sample efficiency, variance, or representational capacity.
Show a diagnostic order - rule out reward scale, missing terminations and discount before touching the algorithm - and justify the lever you pull with evidence.
Own the decision under real constraints: say which leg the deployment has already fixed, what the organisation is willing to spend in samples or variance, and how you will know the fix worked.
### Step one: is it actually divergence? Unbounded-looking value growth has several mundane causes, and relaxing a triad ingredient is an expensive response to a scaling bug. - **Check the ceiling.** In a discounted task with per-step rewards bounded in magnitude by `r_max`, no true action value can exceed `r_max / (1 - gamma)`. With a discount of 0.99 and rewards of at most 1, that ceiling is 100. Predictions passing it are not merely optimistic, they are impossible - that is real evidence, not a hunch. - **Check the reward scale.** A raw reward in the thousands multiplies straight through the target and produces enormous but *correct* values, along with gradients large enough to destabilise anything. - **Check for missing terminations.** If an episode never signals its end, the update keeps bootstrapping past the boundary and accumulates value from a continuation that does not exist. - **Check the discount.** A discount very close to one lengthens the effective horizon and weakens the contraction that keeps the bootstrapped loop in check; it makes genuine divergence far more likely and also inflates legitimate values. Only when the values pass the theoretical ceiling, and the growth persists across step sizes and seeds, should you treat it as triad instability. ### Step two: the three levers and their prices **Relax off-policy.** Learn about the policy you are actually running. Bootstrapped updates with a function approximator converge when the data distribution matches the policy being evaluated, so this genuinely removes the danger. The price is severe in data terms: every transition becomes single-use, stored experience from older policies is no longer valid training material, and sample efficiency collapses. This is the right lever when environment samples are cheap - a fast simulator - and wall-clock stability matters more than sample count. **Relax bootstrapping.** Replace the one-step target with an n-step return for large n, or with the full Monte-Carlo return. The target then rests mostly or entirely on observed rewards, the update becomes ordinary regression on a fixed label, and the self-referential loop that drives divergence weakens in proportion to how much of the target is observed. The price is variance - a full return accumulates the randomness of every step - and a requirement that episodes terminate in reasonable time. This is often the best lever in practice because it is continuous: you can dial n up until the instability stops rather than making an all-or-nothing switch. **Relax approximation.** Use a smaller or better-conditioned function class, engineered features, tile coding, or a coarse discretisation. Less feature overlap means one state's update disturbs fewer others. The price is capacity and generalisation, so this only works for tasks whose structure you understand well enough to encode by hand - and for a genuinely high-dimensional state it is not available at all. ### Step three: what the problem lets you give up The decision is driven by which resource is actually scarce. A power-grid dispatch agent illustrates the constrained case. It learns a greedy target policy from transitions generated by an exploratory or human-operated behaviour policy, and the values climb without bound. Running the target policy to collect on-policy data is not permitted - you cannot explore a live grid to satisfy an algorithm - so the off-policy leg is fixed by the problem, not by preference. That leaves lengthening the returns, lowering the discount to shorten the effective horizon and strengthen the contraction, and normalising rewards so no single update takes a large parameter step. A cheap simulator inverts the calculus. Samples cost almost nothing, so giving up off-policy data reuse costs wall-clock time rather than feasibility, and the on-policy route buys back a convergence guarantee. A well-understood low-dimensional control problem admits the third route: hand-built features with limited overlap can be both stable and adequate, and are much easier to reason about when something goes wrong. ### Step four: keeping all three You can also keep all three ingredients and pay elsewhere. Under linear approximation, the gradient-TD family and emphatic TD are provably convergent off-policy, at the cost of a second parameter vector or high-variance reweighting. In practice most teams instead lean on engineering that damps the loop - bounded rewards, a discount comfortably below one, conservative step sizes, and the standard stabilising machinery of value-based deep RL - accepting that these reduce risk empirically rather than removing it by proof. ### What the interviewer is testing Not a memorised list. They want to see you rule out the boring causes with a quantitative check, name the three levers with their actual costs, and then choose based on a constraint of the deployment rather than on taste. The strongest answers say explicitly which leg the problem has already taken off the table.
- How do you tell an impossibly large Q-value from a correct one?Bound it analytically. With per-step rewards bounded by r_max in magnitude and discount gamma, the largest possible discounted return is r_max / (1 - gamma), so any prediction beyond that cannot correspond to any policy in the environment. Values below the ceiling but growing steadily are weaker evidence, so confirm with multiple seeds and step sizes before concluding it is instability rather than slow, legitimate value propagation.
- When is switching to on-policy data simply not available?Whenever the target policy cannot be executed to collect data. Offline learning from logged operations, safety-critical systems where an exploratory policy is not permitted to touch the plant, and settings where data comes from human operators all fix the off-policy leg. In those cases the negotiable levers are the target construction and the representation, not the data source.
- Why does lowering the discount factor often stabilise a diverging agent?The discount is the contraction factor of the bootstrapped backup: each pass through the loop multiplies existing error by gamma. A discount near one barely shrinks error per iteration, so the projection step's expansion under an off-policy data distribution can dominate and the estimates escalate. Lowering it shortens the effective horizon and strengthens the contraction, buying stability at the cost of caring less about distant reward.
saying these in an interview costs you the question
- Concludes divergence without checking the maximum possible value
- Reaches straight for a smaller learning rate as the fix
- Ignores that offline settings make off-policy non-negotiable
- Claims dropping bootstrapping is free of any cost
- Treats reward rescaling and triad instability as the same problem