skip to content

Value iteration planned an optimal policy from estimated transition probabilities — what can go wrong?

level: seniorimportance: nice to knowfreq 26%

answer

  1. optimal for the model, not the world
  2. the maximum selects favourable errors
  3. reported value is optimistic
  4. errors compound over the horizon
  5. sensitivity sweep the shaky parameters

basics

~20 s

The policy is exactly optimal for the model you supplied, not for the world. Errors in the transition probabilities compound over the discounted horizon, and the planner actively exploits any transition whose probability or reward was estimated too optimistically.

solid answer

~50 s

Planning gives you an exact answer to the question you asked, and the question was posed in terms of your estimated dynamics. Two things bite. First, the optimiser seeks out favourable transitions, so estimation error is not averaged away — it is selected for; the reported optimal value is systematically optimistic. Second, small per-step errors compound: value differences grow with the effective horizon `1/(1-gamma)`, and worst-case bounds scale with its square, so a long-horizon plan built on shaky dynamics can be far off. In a machine replace-or-repair MDP built from reliability tables, the parameters you trust least are usually the rare catastrophic failures with the fewest observations, and those are exactly the ones the policy's safety hinges on. The defences are sensitivity analysis over plausible parameter ranges, checking whether the *structure* of the policy is stable, planning against a worst case within an uncertainty set, and re-solving as new data arrives.

go deeper

for a junior

Know that a planner is only as good as the model it was given, and that transition probabilities estimated from limited records are uncertain inputs rather than facts.

for a middle

Explain that the optimisation step selects favourable estimation errors, which makes the reported optimal value optimistic, and that per-step model error accumulates over the discounted horizon.

for a senior

Demonstrate the working practice: sensitivity sweeps across plausible parameter ranges, checking whether the policy's structure holds, and re-solving as operating data arrives rather than shipping a one-off plan.

for a principal

Own the decision of how much conservatism to buy — worst-case planning over an uncertainty set versus certainty-equivalent re-solving — and set the reporting standard so no single optimal number is presented as a forecast.

## The failure in one sentence Dynamic programming returns a policy that is *provably optimal for the model it was given*. Nothing in the guarantee says anything about the world; the guarantee transfers only as far as the model does. ## Why estimation error does not simply average out The intuition people bring from supervised learning — "the errors are random, they will wash out" — fails here because a maximisation sits between the estimates and the answer. The planner scans all actions and picks the one with the best estimated value. Any action whose transition probabilities happen to have been estimated favourably is more likely to be selected precisely *because* of that error. So the errors are not sampled at random; they are selected for. Two consequences: 1. The value the planner reports for its chosen policy is **optimistically biased** as an estimate of real-world return. Reporting it as the expected outcome is a genuine forecasting error, not a rounding issue. 2. The policy itself may be attracted to poorly-estimated corners of the state space, where the thin data allowed an over-favourable number. ## Why the error compounds A plan is a statement about a long sequence of transitions. Errors in one-step dynamics enter the value function once per step and are discounted, so their total influence scales with the effective horizon `1/(1-gamma)`. Standard results bound the value gap between planning in an approximate model and the true model by a term that grows with the *square* of that effective horizon when the reward scale is fixed — one factor because the error is incurred at every step, another because the policy chosen under the wrong model is itself wrong. The practical reading: at `gamma = 0.99` the effective horizon is a hundred steps and a per-step dynamics error you would call negligible can be amplified enormously in value terms. ## Which parameters actually matter Not all of the model deserves equal suspicion. - **Rare, high-consequence transitions.** A catastrophic failure observed three times in a reliability table has an enormous relative uncertainty, and it is usually the transition the policy is trying to avoid. The parameter you know least is the one carrying the most decision weight. - **Transitions near a decision boundary.** In many maintenance problems the optimal policy has threshold structure — replace once the condition index passes some level. Whether the threshold sits at level seven or level nine may hinge on a handful of transition probabilities; the rest of the model barely moves it. - **The reward or cost scale.** Getting the ratio of downtime cost to replacement cost wrong shifts the policy far more than a small perturbation of the failure rates. Sensitivity analysis tells you which of these you are in: re-solve the MDP with each uncertain parameter moved across its plausible range and look at what changes. ## What to do about it **Look at the policy's structure, not only its value.** If the optimal policy is "replace at condition level eight" and every plausible model in your uncertainty range says between seven and nine, you have a robust recommendation even though the model is uncertain. If the recommended action flips between replace-now and never-replace as you nudge a probability, you do not have an answer yet — you have a signal that more data on that parameter is the highest-value work available. **Plan against an uncertainty set.** A robust MDP replaces each transition distribution with a set of plausible distributions and optimises the worst case within it. You get a more conservative policy with a guarantee that holds for every model in the set. The price is pessimism, which can be severe if the sets are drawn too wide. **Re-solve as data accumulates.** Treat the plan as a standing computation, not a one-off artefact. Certainty-equivalent control — estimate the model, solve it, act, update the estimates, re-solve — is the usual production pattern, and it is fine as long as everyone understands the policy will move as the model does. **Report intervals, not a single optimal value.** The honest deliverable is: here is the recommended policy, here is how it changes across the plausible range of the model, and here is the range of returns it produces — not a single number carrying four decimal places of false precision. **Constrain what the planner may recommend.** Operational constraints (never leave a critical machine unattended past some interval, never order beyond storage capacity) should be encoded as hard restrictions on the action set rather than left to be discovered by an optimiser that is scoring them with uncertain numbers. ## The line to remember "Optimality is relative to the model. The interesting question is not whether the planner solved the MDP correctly — it did — but how much the recommended policy moves when the model moves within its own error bars."

  • Which model parameters deserve the most scrutiny before you trust the plan?
    The ones with thin data and high decision weight — typically rare, high-consequence transitions such as catastrophic failures, and any probability sitting near a decision boundary where the recommended action flips. A sensitivity sweep identifies them: move each uncertain parameter across its plausible range, re-solve, and see which ones change the policy rather than just the value.
  • What is a robust MDP, and when is it worth the pessimism?
    It replaces each transition distribution with a set of plausible distributions and optimises against the worst case in that set, yielding a policy whose guarantee holds for any model inside it. It is worth the conservatism when downside outcomes are expensive and irreversible, and a poor trade when the sets are drawn so wide that the worst case is dominated by scenarios nobody believes.
  • How should the result be presented to a decision maker?
    As a recommended policy plus its stability: what the policy is, how it changes as the uncertain parameters move within their ranges, and what range of outcomes it implies. A single optimal value quoted to several decimal places overstates confidence, because that number is the optimum of an estimated model and is optimistically biased.

saying these in an interview costs you the question

  • Reports the planner's optimal value as the expected real return
  • Treats estimated transition probabilities as ground truth
  • Ignores that rare transitions have the thinnest data
  • Assumes a small model error implies a small value error
  • Never re-solves as new operating data arrives

context