Amazon · ML & AI Fundamentals
Evaluate NLP Classification Models
TrueInterview
October 7, 2026 · 6 min read
Imagine you are in an interview for a Data Scientist intern role at Amazon. The interviewer wants you to explain your approach to an NLP classification task — such as sorting customer messages, search queries, or support tickets into groups — and then tests your grasp of the basics of model evaluation.
Go through the sections that follow. The aim is to show that you can describe key metrics plainly, weigh tradeoffs using business costs, and link evaluation decisions to actual product outcomes.
Boundaries and Premises
- The context is a practical classification system (either binary or multi-class), not a tidy academic dataset.
- Class distributions can be skewed — the category that matters most (like a serious policy breach, fraud, or an escalation) tends to be uncommon.
- The model produces scores or probabilities, and a cutoff value turns those into concrete actions.
- Predictions might trigger an automatic response (such as banning, blocking, or routing) or go to a human review queue.
- Certain responses (Part 1) need to be clear to a lay audience; others (Parts 2–5) require exact definitions and equations.
Questions to Clarify
- Is this a binary or multi-class problem, and how many classes exist?
- How skewed are the classes, and which one is most important for the business?
- What occurs after a prediction — an automated step, or sending it to a person for review?
- In this product, what are the comparative costs of a false positive versus a false negative?
- Are there latency or cost limits that restrict which model you can use?
Part 1 — Describing a confusion matrix to high schoolers
Describe what a confusion matrix is to a class of high school students. Skip technical terms; pick a specific, familiar example.
hint Start with one example Choose a simple yes/no situation they already know (spam or not, sick or healthy) and label the four boxes in everyday language before you mention "true positive."
What This Section Should Include
- Truly understandable for non-experts: a single concrete example, no equations, and the four cells named in simple terms.
- Shows why the matrix is useful — it distinguishes between types of errors, not just their total number.
- Delivers the key point that various mistakes have different consequences in the real world.
Part 2 — Defining precision, recall, F1, and AUC
Provide the exact definition of precision, recall, F1 score, and AUC, with formulas when applicable. In a single sentence, say what each metric reveals.
hint Use the four cells Each of these is a ratio or summary based on the TP/FP/FN/TN numbers from Part 1 — except AUC, which measures ranking without a threshold. Consider which error each ratio places in its denominator.
hint Interpreting AUC ROC AUC has a neat one-sentence probabilistic meaning involving "a randomly selected positive" and "a randomly selected negative" — can you express it? Also consider when PR AUC is preferable to ROC AUC.
What This Section Should Include
- Accurate formulas for precision, recall, and F1, expressed in terms of TP/FP/FN counts.
- A one-sentence intuition for each metric (and which mistake appears in each denominator).
- The probabilistic meaning of ROC AUC and the fact that it does not depend on a threshold.
- Recognition that PR AUC is a superior summary when classes are heavily imbalanced.
Part 3 — Precision versus recall: when to prioritize which
Describe when you would favor precision over recall, and when you would favor recall over precision. Base your argument on a decision rule, not merely examples.
hint The key question Ask: which costs more in this situation, a false positive or a false negative? Then link that to where you place the decision threshold on the model's score — shifting it exchanges one metric for the other. An score allows you to embed the chosen tradeoff into one number.
What This Section Should Include
- One decision rule (the relative cost of FP versus FN), not a set of memorized cases.
- The threshold as a lever: increasing or decreasing it trades precision for recall.
- Understanding that captures a deliberate imbalance, and that 0.5 is not a fixed default.
Part 4 — Assessing a multi-class classification model
How would you assess a multi-class classification model (with one of categories)? Look past a single accuracy figure.
hint One number is not enough A confusion matrix along with per-class precision/recall/F1 uncovers failures that overall accuracy conceals — particularly for rare classes.
hint Averaging choices Micro, macro, and weighted averaging address different questions. Determine which one highlights rare-class performance, and also think about calibration and subgroup breakdowns.
What This Section Should Include
- The confusion matrix and per-class precision/recall/F1, not only overall accuracy.
- A thoughtful selection among micro, macro, or weighted averaging, justified by rare-class concerns.
- Moving beyond single-point metrics: calibration when probabilities trigger actions, and evaluation by subgroups or slices.
Part 5 — Cross-entropy loss
What is cross-entropy loss, and why is it a standard choice for classification? Provide the binary and multi-class versions.
hint What it evaluates Consider what two quantities you are comparing in the formula — and think about how harshly the loss reacts when the model is highly confident yet entirely incorrect. Why does that reaction, along with differentiability, make it suitable for gradient descent?
What This Section Should Include
- Accurate binary and multi-class (softmax) expressions of the loss.
- An explanation: it compares the predicted distribution with the true label distribution, which is the same as maximizing likelihood.
- Why it is the go-to — a steep penalty for confident but incorrect predictions and straightforward differentiability/gradients.
Part 6 — When human evaluation outperforms an automatic metric
In which cases might human evaluation be superior to an automatic objective function or metric? Also mention the expenses and drawbacks of human evaluation.
hint When proxies fail Consider tasks where the automatic metric captures only surface similarity, not meaning (summarization, search relevance, generation). Then assess human evaluation's own flaws (cost, rater variability) and how you would mitigate them.
What This Section Should Include
- Specific tasks where the automatic metric serves only as a superficial stand-in for the quality that truly matters.
- The ways automatic metrics fail (Goodhart's law, noisy or stale references, similarity limited to surface features).
- A candid recognition of human evaluation's own shortcomings and the safeguards that ensure a rigorous study (rubrics, several blind raters, inter-annotator agreement).
Part 7 — Explaining an NLP project you have completed
If you are asked to describe an NLP project you have worked on, which technical details and tradeoffs should you cover? Outline the framework of a strong response.
hint Follow the lifecycle Address problem definition, data and labeling, simple baselines versus complex models, the metric you selected and the reason, error analysis, and deployment/impact. Interviewers appreciate a candidate who can defend a simpler model and demonstrate lessons learned from errors.
What This Section Should Include
- A lifecycle outline: framing → data/labeling → baselines versus complex models → metric selection → error analysis → deployment/impact.
- Rationale for every decision, including the wisdom to advocate for a simpler model.
- Real error analysis and a truthful discussion of tradeoffs and what was gained from errors.
What a Strong Response Includes
These aspects cut across all sections — they are what the interviewer seeks throughout the entire discussion, in addition to the per-section criteria above.
- Audience adjustment: Part 1 is truly understandable to non-experts; Parts 2–5 are exact and formulaically correct — the candidate tailors their language to the listener.
- Cost-driven reasoning: metric and threshold decisions are defended by the cost of each error type, not by memorized examples.
- Imbalance consciousness: acknowledges that accuracy is deceptive for rare classes, and consistently turns to macro F1, PR AUC, or per-class recall.
- Link to decisions: the ideal metric is the one that reflects the product's cost of errors and supports an actual decision, not the most advanced one.
Follow-up Questions
- Your model has an ROC AUC of 0.95, yet business stakeholders say it is "useless" for the rare class. What is probably happening, and what would you measure instead?
- Predictions trigger an automated action. The model's probabilities serve as confidence to decide whether to auto-act or send to human review. How would you verify that the probabilities are reliable?
- Imagine your offline macro F1 improved after a model update, but the online business metric declined. How would you reconcile and diagnose this?
- How would you structure a human evaluation study so that its findings are reproducible and not merely one annotator's viewpoint?
Overview: This question assesses skill in evaluating NLP classification models, including comprehension of confusion matrices, precision/recall/F1/AUC metrics, thresholding and trade-offs, managing class imbalance, and linking metric choices to operational actions and business costs.