In offline RL on logged ICU sepsis dosing records, what makes the learned policy untrustworthy?
answer
- the log never tried that dose
- taking a maximum finds optimistic errors
- no overlap, no evaluation
- clinicians saw what the record omits
- pessimism costs you upside
basics
~20 sTrained only on logged decisions, an offline policy extrapolates the value of doses clinicians rarely gave, and maximisation selects exactly those over-optimistic estimates. You also cannot test it without acting, and the unrecorded reasons behind each dose confound the data.
solid answer
~50 sOffline RL learns a policy from a fixed log of past decisions and outcomes with no further interaction, which is the only option in an ICU where you cannot explore doses on live patients. Three things make the result untrustworthy. First, distributional shift: the policy will prefer actions the clinicians seldom took, whose value is extrapolated rather than observed, and taking a maximum over estimates systematically picks whichever errors are most optimistic. Second, evaluation: judging a policy that acts differently from the log requires the log to have some real probability of the actions the policy proposes, and where that overlap is thin the estimates have enormous variance. Third, confounding: clinicians chose doses on information never written down, so sicker patients received more aggressive treatment, and a naive model can read the treatment as causing the worse outcome. The mitigations are pessimism, staying near the logged behaviour, and shadow deployment.
go deeper
Know the setting: offline RL means learning from records of past decisions with no chance to try anything new. Be able to say why an ICU forces that setting rather than live experimentation.
Explain distributional shift concretely -- why estimates for rarely taken actions are extrapolated, and why maximising over noisy estimates seeks out the optimistic errors instead of averaging them away.
Show the evaluation discipline: overlap checks before modelling, variance diagnostics beside any off-policy estimate, clinician review of the largest disagreements, and shadow deployment before anything acts.
Own the risk position. Decide what evidence would justify acting on a learned clinical policy at all, who signs off, and whether the honest outcome is a decision-support tool rather than an autonomous policy.
## What offline RL is Offline RL learns a decision policy purely from a fixed dataset of past interactions -- situations, the action taken, and what followed -- with no ability to try anything new during learning. It is the natural framing for the ICU sepsis case: years of records exist showing which fluids and vasopressor doses patients received and how they fared, and nobody is going to let an agent explore doses on live patients to find out what else might work. The appeal is obvious and the difficulty is under-appreciated. ## Problem one: extrapolation plus maximisation RL improves a policy by preferring actions with higher estimated value. Estimates come from the data. For an action-situation combination the clinicians rarely or never chose, the estimate is not measured but extrapolated by the function approximator, and its error can point in either direction. Now take the maximum. Maximising over noisy estimates does not sample errors at random -- it seeks out the largest positive ones. So the policy is actively drawn toward exactly the actions the data cannot support, and it looks best precisely where it is least justified. Online RL self-corrects this: it tries the optimistic action, discovers the truth, and updates. Offline RL never gets that correction, which is why the failure is structural rather than a tuning issue. The standard remedy is **pessimism**: penalise estimated value where data is thin, or constrain the learned policy to stay close to the behaviour that generated the log. The honest cost of pessimism is a ceiling on improvement -- you cannot discover a good treatment that nobody ever tried, because by construction you refuse to trust it. ## Problem two: you cannot evaluate without acting Suppose you have a candidate policy. How good is it? On a supervised model you would score held-out rows. Here the held-out records were produced by the clinicians, not by your policy, so the outcomes you would like to score do not exist. Off-policy evaluation attempts this by reweighting logged trajectories according to how likely your policy was to take the action that was actually taken -- an importance-sampling argument. It works only under **overlap**: every action your policy would choose must have had a non-trivial chance of being chosen in the log for that kind of patient. Where overlap is thin, a handful of trajectories carry gigantic weights, the estimate is dominated by a few patients, and its variance is so large the number is not decision-grade. Worse, the logging behaviour's probabilities are usually unknown for human clinicians and must themselves be estimated, adding another layer of error. Always report an effective-sample-size style diagnostic next to any such estimate, and treat an estimate with none as unreported. ## Problem three: confounding by indication The log was not generated by randomised exploration. Clinicians chose doses using everything they saw, including judgement that never reached a field in the record. Sicker patients systematically received more aggressive treatment and also had worse outcomes. If severity is only partly recorded, the residual severity is absorbed into the dose variable, and the model learns that the aggressive dose causes death. A policy trained on that inference withholds treatment from the patients who need it most -- and its estimated value will look excellent, because it is scored under the same distorted assumptions. No amount of data fixes this; more records of the same biased process sharpen the wrong estimate. It is repaired only by measuring the confounder, finding a source of genuine randomisation, or restricting the policy to comparisons where the confounder is controlled. ## How to proceed anyway - **Check overlap before modelling.** Compare the action distribution your policy proposes against the logged distribution, sliced by patient state. No overlap, no claim. - **Constrain the action space** to a range clinicians actually explored, and to deviations experts consider defensible. - **Prefer pessimistic value estimation** and be explicit that you have bought safety with lost upside. - **Read the disagreements.** Pull the cases where the policy departs most from clinician behaviour and have clinicians review them. This finds confounding and data artefacts faster than any metric. - **Deploy in shadow or advisory mode first**, where the policy's recommendation is recorded but a human acts, which also begins to generate the data an evaluation would need. ## The interview point The weak answer is that a large dataset substitutes for interaction. The strong answer is that offline RL trades the cost of interaction for a much harder evaluation problem, and that in a clinical setting the binding constraint is not the learning algorithm at all -- it is whether anyone can credibly say the proposed policy is better before someone acts on it.
- How would you sanity-check an offline policy before anyone acts on its recommendations?Check overlap first: does the log ever contain the actions this policy proposes, for these patient states, with non-trivial probability? Then report any off-policy value estimate alongside an effective-sample-size diagnostic, never alone. Then pull the cases of largest disagreement with clinicians for expert review. Only after that, run advisory-only for a period and compare recommendations against outcomes.
- Why does constraining the policy to stay near the logged behaviour help, and what does it cost?It keeps value estimates inside the region the data supports, so extrapolation errors cannot be exploited by the maximisation step. The cost is a ceiling on improvement: a genuinely better action that clinicians never tried is unreachable by construction, because the constraint treats unfamiliar and bad as the same thing.
- What is confounding by indication in this dataset, and why does more data not fix it?Sicker patients received more aggressive dosing, so if severity is only partly recorded the dose carries the unmeasured severity with it and appears to cause the worse outcome. More records come from the same biased process, so they sharpen the wrong estimate rather than correcting it. You need the missing measurement or a source of randomisation.
saying these in an interview costs you the question
- Says a large enough dataset removes the need to interact
- Treats logged clinician choices as random exploration
- Reports estimated policy value with no variance diagnostic
- Assumes higher estimated value means a better policy
- Believes more data cures confounding by indication