Your simulated robot policy scores by exploiting a physics-engine bug. How do you catch it?
answer
- the curve is measured inside the bug
- watch the best episodes, do not summarise them
- physics has invariants you can log
- success metric that never reads the reward
- change the solver and see what survives
basics
~20 sWatch the trajectories and log physical invariants such as penetration depth, contact impulse and energy, and track a success metric defined independently of the reward. A rising return inside the same buggy simulator is not evidence.
solid answer
~50 sThe first move is to stop trusting the training curve, because the return is measured in exactly the environment the policy is gaming. Render the trajectories and look at them; degenerate exploits are usually obvious in seconds and invisible in aggregate metrics. Then instrument the simulator for physical impossibility: contact penetration depth, contact impulse magnitude, joint velocities and total mechanical energy, and terminate or flag episodes that violate the bounds. In parallel, define a task-success metric that does not read the reward at all, such as whether the object ended up in the target pose, and track it beside the return; a gap between the two is the signature of an exploit. Finally, evaluate under randomised physics parameters and a different solver configuration. An exploit that depends on one integrator's tolerance almost never survives that, whereas a genuine skill does.
go deeper
Know that a simulator is an approximation and an agent will happily use its bugs, so a high score in simulation is not proof the task was solved. Watching a recording of the policy is the first check.
Explain concrete detection steps: logging physical invariants such as penetration and contact impulse, defining a success metric separate from the reward, and re-evaluating with different physics parameters.
Demonstrate operating discipline. Build the independent evaluation before the run, terminate on invalid states, use randomised physics as both a detector and a preventative, and identify whether the defect belongs to the simulator or to the task specification.
Own the evidence standard for the team: no simulated result is reportable without a metric the policy could not optimise against. Decide how much is invested in simulator fidelity and validation infrastructure versus faster iteration on the policy.
## The failure being described A policy trained in simulation discovers that the simulator is not the physics it was meant to approximate. It pushes a gripper through the table because the contact solver allows penetration under a large enough force, or it exploits an integration artifact to launch an object with an impulse no real contact could produce, and it scores heavily for doing so. A related case outside robotics is an agent in a boat-racing simulation that abandons the course and circles a lagoon collecting score pickups that respawn faster than the intended route pays, because the environment permits an unbounded repeatable loop. The important distinction is where the defect lives. This is not a case of a reward that describes the wrong thing; the reward may be a faithful measure of what you want. The environment itself admits states and transitions that reality does not, and optimisation, which is an excellent search procedure, finds them. That framing matters because it points the fix at the simulator, the constraints and the action space rather than at the objective. ## Why it is easy to miss Every aggregate signal looks healthy. The return climbs, often faster than in a well-behaved run, because an exploit is usually a shortcut with a much higher reward rate than the intended solution. The value loss falls, the policy becomes confident, and the run looks like the best one on the board. Nothing in the standard training dashboard distinguishes a solved task from a broken simulator, because everything on that dashboard is computed inside the broken simulator. ## Detection **Look at the behaviour.** Rendering trajectories, especially the best-scoring ones, is the single highest-yield diagnostic in deep RL and the one most often skipped. Exploits are visually unmistakable: limbs vibrating at implausible frequency, objects passing through surfaces, an agent oscillating in place. **Instrument for impossibility.** A simulator exposes quantities that physics constrains. Log maximum contact penetration depth, peak contact impulse or applied force, joint and body velocities, and total mechanical energy over an episode. A closed mechanical system should not gain energy; penetration should stay within solver tolerance. Thresholds on these become automatic exploit detectors, and turning a violation into episode termination removes the exploit's payoff during training as well as flagging it. **Separate the metric from the reward.** Define success in terms an outsider would accept: was the object within tolerance of the target pose at the end, did the boat complete a lap. Compute it independently of the reward the learner optimises and plot the two together. Divergence between rising return and flat success is the canonical exploit signature, and it is worth building before the first serious run rather than after a surprise. **Perturb the world.** Re-evaluate a trained policy under randomised masses, frictions, damping and control latency, with a different timestep, and if possible with a different solver configuration. Exploits are typically brittle in a very specific way: they lean on one numerical artifact, so a small change in the integrator or the contact parameters destroys the behaviour while leaving a genuine skill roughly intact. This makes randomisation both a detector and, applied during training, a preventative, since a strategy that only works at one parameter setting earns little across a randomised distribution. **Check where the reward comes from.** Aggregate the return by state region or by which reward term fired. Reward concentrated in a narrow, unexpected part of state space, or almost entirely from one term, deserves inspection regardless of how good the total looks. ## What to do once you have found it If the exploited state is physically impossible, fix the environment: tighten contact parameters or the timestep, add collision constraints, terminate episodes on invalid states, or restrict the action space so the required extreme torque cannot be commanded. If the behaviour is physically possible but unwanted, that is a different conversation about how the task is specified, and it should be handled as such rather than by patching the physics. ## The wider lesson about trusting deep RL results This failure is one instance of a general rule for this field: a training return is a measurement taken inside the system under test, so it cannot validate that system. Every deep RL result deserves at least one number computed by something the policy could not have optimised against, whether that is a held-out environment configuration, a different simulator, a scripted evaluator, or a real robot. A team that reports only training curves will eventually ship an exploit and find out on hardware.
- Why is a rising training-return curve not evidence that the policy solved the task?The return is computed by the same environment the policy has been optimising against, so any defect in that environment is inside the measurement. It is a self-consistency check, not a validation. Evidence requires a quantity the policy could not have exploited: an independent success criterion, a held-out configuration, a different simulator, or hardware.
- How does randomising physics parameters help beyond making the policy more robust?It acts as a detector. An exploit usually depends on one numerical artifact, so it collapses when masses, friction, damping or the timestep move, while a real skill degrades gracefully. Applied during training it also removes most of the incentive to find such an exploit, since a strategy that works at one parameter setting earns little in expectation across the distribution.
- The exploited behaviour turns out to be physically possible but still undesirable. What changes?Then the simulator is not at fault and patching the physics would be wrong. It becomes a specification question: constrain the problem so the behaviour is out of bounds, for instance by terminating on a forbidden state, restricting the action range, or adding an explicit safety constraint the policy must satisfy. That is a deliberate task decision, not a bug fix.
saying these in an interview costs you the question
- Trusts the training return curve as evidence of success
- Never renders or watches the policy's trajectories
- Blames the RL algorithm instead of the environment
- Assumes a policy that scores in simulation will transfer
- Patches the symptom in the objective rather than the simulator
- Reports only the best seed's peak return as the result