Back to problems

RL Fundamentals — GRPO Debug

AI Coding · Anthropic · Hard

Requirements Part 1 — Debug The interviewer provides a compact GRPO trainer written in PyTorch. You must identify three specific defects that candidates have encountered: torch.multinomial receives raw logits because the required softmax is absent. The standard deviation used to normalize advantages lacks an epsilon term, which can produce NaNs when the variance is almost zero. The relationship involving ratio = exp(model_logprob - old_logprob) is subtly wrong: even at the…

Checking your access…