skip to content

Reinforcement Learning

Learning by acting: an agent explores an environment, collects delayed reward and improves its policy with Q-learning or policy gradients. Interviewers check you can tell it from supervised learning.

on this pageshow

explore

questions

page 2 of 2

When is a stochastic policy strictly better than any deterministic policy?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

In a finite, fully observed MDP some deterministic policy is always optimal, so randomising gains nothing. Randomising wins when an opponent can exploit predictability, when distinct states look identical to the agent, or when the policy class is restricted.

open as a page

In tabular Q-learning, when would you prefer a constant learning rate over a decaying one?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Prefer a constant learning rate when the environment drifts, because it keeps weighting recent experience and lets the agent track change. Prefer a decaying one when the problem is stationary and you want the convergence guarantee, which a constant rate forfeits.

open as a page

When is building a simulator to train an RL agent worth the investment, and when is it a trap?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Build a simulator when the dynamics are well understood and the real system is too slow or unsafe to explore. It is a trap when the hardest part to model is the part that matters: the agent optimises your assumptions.

open as a page

showing 31–33 of 33