Back to problems

Debug and Improve a GRPO RL Training Loop for Language Models (PyTorch)

AI Coding · OpenAI · Hard

Problem: Debug and Enhance a GRPO RL Training Loop for Language Models (PyTorch) You receive PyTorch code for reinforcement-learning fine-tuning of a language model in an RLHF/RLAIF-like setup. Although the implementation says it uses Group Relative Policy Optimization (GRPO) and executes without failing, its optimization behavior is flawed: training may be unstable, logically incorrect, or report misleading metrics. In the interview, you must repair and improve this loop.…

Checking your access…