What separates model-free from model-based reinforcement learning?
answer
- who knows what happens next
- dynamics versus consequences
- can you simulate without acting
- transition and reward function held or not
- TD needs no transition probabilities
basics
~20 sA model-based agent has (or learns) the environment's transition and reward dynamics and can plan against them. A model-free agent never builds those dynamics; it estimates values or a policy directly from sampled experience. Temporal-difference learning is model-free.
solid answer
~50 sA model of the environment answers the question "if I take action a in state s, what state and reward come next, and with what probability?". Model-based methods either receive that model or estimate it from data, then use it to compute a policy without touching the environment again. Model-free methods, including Monte-Carlo and temporal-difference learning, skip that step entirely: they watch actual transitions `(s, a, r, s')` roll past and nudge value estimates toward what was observed. The practical trigger is availability. On a factory conveyor whose jam-and-breakdown behaviour nobody can write down, there is no transition table to fill in, only logged runs — so the only honest option is to learn from sampled experience. The tradeoff is that a model lets you squeeze many updates out of one piece of data, while model-free learning pays for every update with real interaction.
go deeper
Be ready to state the distinction in one sentence and place temporal-difference learning on the model-free side. Knowing that a model means transition and reward dynamics is the whole ask here.
Explain what an agent can do with a model that it cannot do without one: simulate transitions, reuse each sample many times, answer what-if questions. Name the data-efficiency versus model-error tradeoff.
An interviewer expects you to pick a side for a concrete system and defend it. Argue from whether the dynamics can be estimated honestly, how expensive interaction is, and how badly compounding model error would hurt.
Own the framing decision: whether the team should invest in a dynamics model and a simulator at all, given that an inaccurate model silently corrupts every downstream policy while pure model-free learning simply costs more interaction.
## The object in question: the model In reinforcement learning an agent sits in a state, takes an action, receives a reward and lands in a new state. A **model** of the environment is a description of that last part: the transition dynamics `p(s' | s, a)` (where you end up) and the reward function `r(s, a)` (what you get). Having a model means you can ask "what would happen if…" **without acting** — you can simulate. The model-free / model-based split is simply: does the agent hold such a description, or not? ## Model-based learning A model-based agent either - **is given** the dynamics (a board game's rules, a queueing system written down as equations, a physics simulator), or - **learns** them from data — count how often action `a` in state `s` led to `s'`, average the rewards, and you have an estimated transition and reward model. With a model in hand the agent can compute a policy by reasoning over it, and it can generate as much synthetic experience as it likes. That is why model-based methods tend to be far more **data-efficient**: one observed transition improves the model, and the improved model can be replayed thousands of times. The costs are real. The model has to be accurate, and errors compound: an agent that plans many steps ahead through a slightly wrong model will confidently chase rewards that do not exist — "model exploitation". Learning the dynamics of a rich environment can also be much harder than learning the value of being in a state, because the dynamics carry detail the decision does not need. ## Model-free learning A model-free agent never represents `p(s' | s, a)`. It keeps only what it needs to act — a value function, a policy, or both — and updates it from experience it actually had. Two families dominate: - **Monte-Carlo**: wait until the episode ends, compute the realised return `G` from a state, and move that state's estimate toward `G`. - **Temporal-difference**: after a single step, move the estimate toward `r + gamma * V(s')`, the reward you just got plus your current estimate of where you landed. Both learn the *consequences* of behaviour without ever learning the *mechanics* of the environment. Nothing in a TD update requires knowing why the conveyor jams; it only requires having seen it jam. ## Which one the situation forces The deciding questions are practical: 1. **Can the dynamics be written down or reliably estimated?** For the conveyor, jams depend on part geometry, humidity, wear and operator behaviour. Nobody can enumerate that state space honestly, and a badly estimated model is worse than none. Model-free. 2. **Is interaction cheap?** If every sample costs a production run, the data efficiency of a model becomes very attractive — even an imperfect learned model may pay for itself. 3. **Do you need to answer counterfactuals?** "What if we changed the reward for early completion?" is a model question. A model-free value function is tied to the behaviour it was trained under. ## Why this matters for TD specifically Temporal-difference learning is the canonical model-free prediction method, and it borrows one property from each side of the divide. Like Monte-Carlo, it learns purely from sampled transitions and needs no model. Like model-based reasoning, it **bootstraps** — it updates an estimate using another estimate, which is exactly the kind of self-consistency a model-based computation enforces. That hybrid character is why TD is often described as sampling the experience but reasoning over the estimates, and it is the reason TD can update after one step rather than one episode. A final clarification, because interviewers probe it: *model-free* does not mean *no function approximator*. A model-free agent may hold a large parametric value function. "Model" here means a model of the **environment's dynamics**, not the statistical model doing the estimating.
- Is a learned model of the environment still model-based, or does estimating it make the method model-free?Still model-based. The label refers to whether the agent represents the environment's transition and reward dynamics, not to whether those dynamics were handed over or estimated from data. An agent that counts observed transitions to build `p(s' | s, a)` and then reasons over that estimate is model-based, and it inherits the usual risk: planning far ahead through an inaccurate estimated model produces confident nonsense.
- Why is model-based learning usually more sample-efficient than model-free learning?Because one real transition can be reused many times. It updates the model, and the model can then generate as much simulated experience as compute allows, so many value updates come out of a single interaction. A model-free method extracts roughly one update per transition it actually experienced, so improving the estimate requires proportionally more real interaction.
- Does model-free mean the agent has no learned parameters at all?No. Model-free refers only to the absence of a model of the environment's dynamics. The agent can carry a large parameterised value function or policy with many learned parameters. The distinction is what those parameters describe: the expected value of behaviour, not the probability of the next state given a state and action.
A model-based driver has the road map and can plan a route in their head; a model-free driver has only memories of which turns tended to get them home faster.
saying these in an interview costs you the question
- Says model-free means the agent learns nothing
- Thinks model refers to the neural network or estimator
- Claims model-based always beats model-free
- Believes TD learning estimates transition probabilities internally
- Says a learned dynamics model makes the method model-free