Google · ML & AI Fundamentals
Compare Language-Model Post-Training Methods
TrueInterview
October 7, 2026 · 1 min read
Comparing Post-Training Methods for Language Models
Compare the main post-training methods for language models, describing the signal each one relies on, the model stage or problem it targets, and how they differ in data requirements, stability, control, and evaluation.
Constraints & Assumptions
- Keep supervised adaptation distinct from preference-based or reward-based optimization.
- Do not collapse distinct methods into one another just because they share a label.
- When a method depends on a reference policy, reward model, or preference data, name it explicitly.
Clarifying Questions to Ask
- Is the aim task adaptation, instruction following, style control, or behavior alignment?
- Do you have scalar rewards, only pairwise preferences, or demonstrations?
- Can new on-policy samples be produced?
Hint — Map signal to objective: For each method, state the training data, the optimized loss, and the failure mode it is meant to address.
What a Strong Answer Covers
- Supervised fine-tuning and what it is for.
- Reward-model plus policy-optimization approaches, along with direct preference objectives.
- On-policy versus offline data, KL control, and how much the setup depends on a reward model.
- Evaluation, reward hacking, distribution shift, and criteria for choosing a method.
Follow-up Questions
- When can optimizing against preferences degrade capability?
- How would you decide between gathering better demonstrations and optimizing more aggressively against preferences?
Overview: Compare supervised, reward-based, direct-preference, and group-relative LLM post-training according to signal, objective, data regime, and failure modes.
Loading comments…