Google · ML & AI Fundamentals
Interpret Language-Model Perplexity
TrueInterview
October 7, 2026 · 1 min read
Interpreting Language-Model Perplexity
Give the definition of perplexity for an autoregressive language model, show how it relates to the average negative log-likelihood, and discuss what it can and cannot tell you about a model's quality.
Constraints and Assumptions
- Run the evaluation on a fixed tokenized dataset and state which logarithm base is used.
- Mask out padding and any tokens that are not prediction targets.
- Perplexity values from different tokenizers cannot be compared directly.
Clarifying Questions to Ask
- Is the loss averaged per token, per sequence, or per byte?
- Was the evaluation set included in the training data?
- How does the setup handle long contexts and truncation?
Hint: Start from likelihood. Write perplexity as the exponential of the mean token cross-entropy before moving to intuition.
What a Strong Answer Covers
- The formula and its interpretation as a geometric mean branching factor.
- Correct masking, normalization, and the dependence on the tokenizer.
- Its use for held-out predictive fit and training diagnostics.
- Its limits for factuality, safety, usefulness, calibration, and cross-model comparison.
Follow-up Questions
- Why might lower perplexity not lead to better performance on a user's task?
- How would normalizing at the byte level make different tokenizers easier to compare?
Overview: Understand the precise likelihood basis of perplexity, the correct evaluation details, tokenizer caveats, and its limits as a quality metric for LLMs.
Loading comments…