Mistral AI · Behavioral
Explain Adam, AdamW, and Decoupled Weight Decay
TrueInterview
September 26, 2026 · 1 min read
Explain Adam and AdamW. Why is adding an L2 penalty to the loss generally different from decoupled weight decay under Adam?
Constraints & Assumptions
Employ the usual first and second moment estimates, bias correction, and a scalar learning rate. Indicate which parameter groups undergo decay; handling of biases and normalization parameters is an independent decision.
Clarifying Questions
Does regularization modify the gradient or update parameters directly? Which learning rate schedule and epsilon convention are in use? Which parameters are subject to decay?
What a Strong Answer Covers
Include the moment update equations, differentiate the two decay methods, and explain how adaptive gradient scaling alters the contribution of an L2 gradient.
Follow-up Questions
In what cases are L2 regularization and weight decay equivalent? Why does the decay coefficient interact with the learning rate schedule? Does AdamW eliminate the need to tune regularization?
Overview: Grasp Adam’s moment estimates and bias correction, then contrast gradient-based L2 regularization with AdamW’s decoupled weight decay and its interaction with the learning rate.
Loading comments…