Amazon · ML & AI Fundamentals
Explain ML evaluation, sequence models, and optimizers
TrueInterview
October 7, 2026 · 1 min read
Scenario
An interviewer is taking a close look at a machine learning project you built; unless told otherwise, treat it as a supervised model. You will need to defend your modeling choices, evaluation setup, and training decisions.
Part A — Evaluation design and metrics
- Walk through your end-to-end evaluation approach: how you split the data, your validation procedure, and how the test set is used.
- Which metrics do you choose, and what business-versus-ML tradeoffs drive that choice?
- Give the mathematical definitions for each metric you named—for example, accuracy, precision/recall, F1, ROC-AUC, PR-AUC, log loss, MSE/MAE, or calibration metrics.
- Suggest an evaluation workflow that improves on a single holdout set, such as cross-validation, time-based splits, stratification, repeated runs, or confidence intervals.
- If you can bring in human labels or human evaluation, cover:
- what you would label,
- how you would ensure label quality with guidelines and inter-annotator agreement,
- how this improves the evaluation signal.
- If you have no labels at all, what is the simplest way to judge whether two model outputs or answers are similar?
Part B — Transformers vs RNNs on long inputs
- Contrast Transformer architectures with RNN/LSTM/GRU architectures.
- For very long sequences, discuss the strengths and weaknesses of each in terms of training stability, long-range dependency capture, and compute/memory costs.
- Explain why attention is able to capture long-range dependencies while vanilla RNNs often have difficulty doing so.
Part C — Detecting distribution mismatch in images
Given two image sets, Set A and Set B, how would you test whether they seem to be drawn from the same underlying distribution?
Part D — Optimizers
Compare the practical differences and tradeoffs among SGD with and without momentum, RMSProp, Adam, and AdamW. In what situations would AdamW be the better choice?
Overview: This question tests a candidate's grasp of ML model evaluation and metrics, sequence-modeling tradeoffs between transformers and RNNs, detecting distribution shift in images, and comparing optimization algorithms in machine learning.
This question is drawn from a full Amazon Applied Scientist interview experience.