Reinforcement Learning
Learning by acting: an agent explores an environment, collects delayed reward and improves its policy with Q-learning or policy gradients. Interviewers check you can tell it from supervised learning.
on this pageshowhide
explore
- Agent, Environment, Reward13 questions
- MDPs and Reward Design5 questions
- Policies and Value Functions4 questions
- Exploration vs Exploitation4 questions
- Planning and Learning20 questions
- Temporal-Difference Learning4 questions
- Q-Learning and SARSA3 questions
- Policy Gradient Methods4 questions
- When RL Fits4 questions
- Value and Policy Iteration5 questions
questions
page 2 of 2When is a stochastic policy strictly better than any deterministic policy?
basics
~20 sIn a finite, fully observed MDP some deterministic policy is always optimal, so randomising gains nothing. Randomising wins when an opponent can exploit predictability, when distinct states look identical to the agent, or when the policy class is restricted.
In tabular Q-learning, when would you prefer a constant learning rate over a decaying one?
basics
~20 sPrefer a constant learning rate when the environment drifts, because it keeps weighting recent experience and lets the agent track change. Prefer a decaying one when the problem is stationary and you want the convergence guarantee, which a constant rate forfeits.
When is building a simulator to train an RL agent worth the investment, and when is it a trap?
basics
~20 sBuild a simulator when the dynamics are well understood and the real system is too slow or unsafe to explore. It is a trap when the hardest part to model is the part that matters: the agent optimises your assumptions.
showing 31–33 of 33