Deep Reinforcement Learning
Replacing the Q-table with a network buys generalization and costs stability: replay buffers, target networks, actor-critic and PPO. Interviewers probe where the tabular guarantees break.
on this pageshowhide
explore
- Value Networks11 questions
- Deadly Triad4 questions
- Replay and Target Networks3 questions
- Double and Dueling DQN4 questions
- Policy Optimization13 questions
- Advantage Actor-Critic4 questions
- Clipped Surrogate Objective4 questions
- Continuous Control5 questions
- Sample Cost and Reliability10 questions
- Sample Efficiency3 questions
- Sparse Rewards and Hacking4 questions
- Seed Sensitivity3 questions
questions
page 2 of 2Your simulated robot policy scores by exploiting a physics-engine bug. How do you catch it?
basics
~20 sWatch the trajectories and log physical invariants such as penetration depth, contact impulse and energy, and track a success metric defined independently of the reward. A rising return inside the same buggy simulator is not evidence.
Your RL variant beats a baseline that lacked observation normalisation — is the gain real?
basics
~20 sNot as stated. Observation normalisation, reward scaling and advantage standardisation often move returns more than an algorithmic change does, so the comparison must give both arms the same details and the same tuning budget before any gain can be attributed to the algorithm.
Your value agent's Q-values climb without bound - which leg of the deadly triad do you relax?
basics
~20 sFirst confirm the growth is divergence, not a legitimately large return, by comparing against the maximum possible discounted value. Then relax whichever leg your problem can afford: on-policy data, longer or full returns instead of bootstrapping, or a simpler representation.
When is adding Double, dueling and prioritized replay to a working value-based agent not worth the cost?
basics
~20 sEach extension fixes a specific symptom and adds tuning surface. Add one only when its symptom shows in the diagnostics, one at a time with matched seeds, and skip any whose failure mode your environment lacks.
showing 31–34 of 34