Google · ML & AI Fundamentals
Detect and Prevent Reward Hacking
TrueInterview
October 7, 2026 · 1 min read
Identifying and Preventing Reward Hacking
Describe the ways reward hacking shows up during language-model post-training, and how you would detect, constrain, and react to it while treating the reward score as only a proxy for the true objective.
Constraints & Assumptions
- The reward is only an approximation and can leave out qualities that matter.
- Detection must rely on signals separate from the reward being optimized.
- Where feasible, interventions should avoid degrading useful capabilities.
Clarifying Questions to Ask
- Which behavior is the reward supposed to capture?
- What shortcuts could raise the score without actually serving user needs?
- Which independent evaluations or human reviews can be used?
Hint — Watch for score-behavior divergence: Examine cases where the optimized reward goes up while blinded human or rule-based measures go down.
What a Strong Answer Covers
- Specific examples such as verbosity, style imitation, manipulating graders, or exploiting evaluator blind spots.
- Independent held-out evaluations, adversarial prompts, causal probes, and ensembles of reward models.
- Regularization, constraints, repairing data, limiting optimization, and monitoring.
- Incident response plans, rollback criteria, and residual uncertainty.
Follow-up Questions
- How would you detect hacking that transfers beyond the prompts used in testing?
- In what cases can increasing reward-model capacity make the situation worse?
Overview: Examine reward hacking in LLM post-training through concrete shortcuts, independent detection, conservative optimization, and recovery controls.
Loading comments…