How do you check that a patient's chart is a Markov state for a discharge decision?
answer
- a property of your state, not the world
- two identical charts, opposite trajectories
- does history change the action?
- lagged features add predictive power
- fold summaries in, pay in dimensions
basics
~20 sAsk whether earlier history would change the action you take. If two patients with identical charts but opposite three-day trends need different decisions, the chart is not a Markov state, so fold trend summaries into it.
solid answer
~50 sThe Markov property is a claim about your **state representation**, not about medicine: it says the next state and reward depend only on the current state and action. The practical test is a counterfactual — take two patients with identical charts today but opposite three-day trajectories, one whose creatinine is climbing and one whose creatinine is falling to the same value. If a clinician would discharge one and keep the other, today's chart is not a Markov state. The fix is to fold the relevant history in: deltas over the last few days, rolling means, time since admission, the last k measurements. You can also test it empirically — if adding lagged features measurably improves prediction of the next day's state on held-out patients, the current state was insufficient. The cost is a larger state space and more data, so you add history that changes the decision, not all of it.
go deeper
Be able to state the property in words: the next state and reward depend only on where you are now and what you do, not on how you got there. Recognise that a trend is not visible in a single snapshot.
Explain the repair as well as the definition. Show how deltas, rolling statistics or a window of the last few observations restore the property, and what each costs in state dimensionality.
Demonstrate that you test the assumption rather than inherit it. Bring the identical-snapshot counterfactual, the held-out lag comparison, and a clear account of how a non-Markov state quietly corrupts value estimates and offline evaluation.
Own the tradeoff between fidelity and learnability. Decide when to enrich the state, when to accept partial observability explicitly, and when the honest call is that the available data cannot support a sequential policy at all.
## What is actually being asked Every reinforcement-learning method assumes the state it is handed is **Markov**: that ``` P(next state, reward | current state, action) ``` is unchanged by conditioning on everything that came before. Interviewers use the clinical framing because it makes the failure vivid, but the question generalises to every applied MDP: *is the thing I am calling a state a sufficient summary of the past?* The crucial move is realising this is **a property of your representation, not of the world**. There is no such thing as a non-Markov environment, only a state description too impoverished to be Markov. That reframing turns an abstract assumption into an engineering task. ## The counterfactual test The cheapest check needs no data. Construct two histories that lead to *identical* current states and ask whether the right action differs. - Patient A: creatinine has risen 0.4 over three days and stands at 1.6 today. - Patient B: creatinine has fallen 0.4 over three days and stands at 1.6 today. If today's chart is the state, these are the *same* state, so any policy must take the same action for both. A clinician would not — one is deteriorating, one is recovering. That single thought experiment proves today's chart is not Markov for the discharge decision. Run the same test on other candidate omissions: time since admission, whether this is a readmission, what treatment was given yesterday, whether a result is still pending. Each is a hypothesis about hidden state. ## The empirical test With data you can go further. The Markov property implies that, given the current state and action, **lagged variables carry no additional predictive information** about the next state or reward. So: 1. Fit a predictor of the next state (or of the outcome) from the current state and action. 2. Fit a second predictor that also receives lagged features — yesterday's and the day before's values, deltas, rolling statistics. 3. Compare on held-out patients. If the second model is materially better, the current state is not Markov and the lags name exactly what is missing. This is a diagnostic, not a proof: passing it says the history you *tested* adds nothing, not that no history could. ## Fixing it Three standard repairs, in increasing cost: **Augment the state.** Add engineered summaries of history — a three-day delta, a rolling mean, a max-so-far, time since admission, a count of prior episodes. This keeps the MDP machinery intact and is almost always the first thing to try. **Stack a window.** Include the last `k` raw observations in the state. Simple and assumption-free, but the state dimension multiplies by `k` and much of it is redundant. **Admit it is partially observed.** When the thing that matters is genuinely unobservable — an underlying disease severity you never measure — the honest model is a partially observable process, and the agent acts on a *belief* over hidden states or on a learned summary of the whole history rather than on a raw observation. This is more powerful and considerably more expensive. ## The cost of adding history History is not free, which is why "add everything" is the wrong answer: - **State-space growth.** Every added dimension spreads the same data more thinly, and estimates for each state get noisier. - **Data requirements.** More state distinctions mean more trajectories are needed before any of them is well estimated. - **Spurious distinctions.** Irrelevant history splits states that should be pooled, so the agent learns separate, worse estimates for what is really one situation. The discipline is to add history that changes the *decision*, checked by the counterfactual test, and to leave out history that merely changes the description. ## What goes wrong if you skip the check Running a learning algorithm on a non-Markov state does not throw an error. It silently breaks things: - **Value estimates become inconsistent.** The same state label mixes genuinely different situations, so the learned value is an average over a mixture that depends on how often each situation occurred in your data. - **Convergence guarantees lapse.** The proofs behind the standard algorithms assume the property; without it, learning may oscillate or settle on a policy that is not optimal for any real situation. - **The best deterministic policy may be worse than a random one.** Under partial observability, committing to a single action per observation can be strictly worse than randomising, which is deeply counter-intuitive to anyone assuming a well-formed MDP. - **Offline evaluation lies.** Estimates computed on logged data inherit the same mixture, so the policy looks fine on paper and misbehaves in deployment. ## How to answer in an interview Say the property is about the representation; give the two-patients counterfactual; name the empirical lag test; then name the fix and its cost. The strongest signal is that you treat "is it Markov?" as a design question you actively test, rather than an assumption you inherit from a textbook.
- Is the Markov property a property of the environment or of your model?Of your state representation. Any process can be made Markov by putting enough of the history into the state — in the limit, the entire trajectory is trivially a Markov state. So the question is never whether the world is Markov, it is whether the summary you chose is rich enough for the decision you are making, and cheap enough to learn from.
- What is the risk of just stacking the last thirty days of observations into the state?You buy the property with dimensionality. The state space explodes, most of the added features are noise for the decision at hand, and the same amount of data now has to support far more distinctions, so every estimate gets noisier. Irrelevant history also splits situations that should be pooled, which makes the learned policy worse rather than better.
- How would you empirically test whether the current state is sufficient?Predict the next state or outcome twice on held-out data: once from the current state and action alone, once with lagged features and trend summaries added. If the second is materially better, the first state was not Markov, and the features that helped tell you exactly what to fold in. Passing only clears the history you actually tested.
A thermometer reading of 38 degrees tells you nothing about whether the fever is breaking or building. The number is the observation; the trend is the state.
saying these in an interview costs you the question
- Says a real environment is either Markov or not, full stop
- Assumes the raw observation is automatically the state
- Adds every available lag without checking relevance
- Thinks a non-Markov state just slows learning down
- Believes more history always improves the policy