Amazon · ML & AI Fundamentals
Explain Core ML Interview Concepts
TrueInterview
October 7, 2026 · 7 min read
Imagine you are in a phone screen for an applied scientist or machine-learning engineer position and the interviewer asks you to explain several machine-learning fundamentals out loud. For every part, provide an accurate conceptual answer and be prepared to defend the reasoning behind it, not merely the definition. Approach each question as a chance to show depth: first give the central idea, then walk through the logic or intuition that supports it.
Constraints & Assumptions
- The format is a conceptual, whiteboard-style conversation rather than a coding task. You will not be given data, libraries, or executable code.
- Responses should be spoken explanations, with light mathematical notation when it helps, such as loss functions or update rules.
- Unless a part says otherwise, work from standard supervised-learning assumptions.
- Sound reasoning and correctness count more than covering many topics; the interviewer will press on the "why" behind each answer and will challenge vague or hand-wavy statements.
Clarifying Questions to Ask
- For the regression and classification portions, should I emphasize the modeling assumptions, the estimation/optimization perspective, or both?
- When we discuss loss functions, do you want the probabilistic maximum-likelihood justification, or only the optimization properties?
- For the optimizer comparison, are you looking for a specific setting such as large-scale vision, NLP/transformers, or sparse features, or a general comparison?
- For the neural-network portion, should we reason from the classical small-network intuition or the modern overparameterized deep-learning viewpoint?
- How much depth do you expect per part: a one-paragraph summary for each, or a deeper treatment of the one I find most interesting?
Part 1 — Linear Regression
What are the principal assumptions behind linear regression? Why is squared loss the usual choice?
Hint — Where to start: Enumerate the classical assumptions one by one: the model's functional form, the conditional mean of the error term, independence or correlation among errors, error variance, and relationships among the features. Then separate them into those required for unbiased point estimates and those required only for valid inference or standard errors.
Hint — Why squared loss: Ask which probabilistic noise model makes least squares the maximum-likelihood estimator. Also consider convexity, differentiability, and which statistic of squared loss ultimately estimates: the conditional mean rather than, for example, the median.
What This Part Should Cover
- Names the classical assumptions and correctly sorts each into what is needed for unbiasedness or consistency of the point estimate versus what is needed only for valid inference, such as standard errors, confidence intervals, and hypothesis tests.
- Points out which assumption is not required for unbiasedness and explains why it is sometimes included anyway.
- Offers at least two separate justifications for squared loss, one probabilistic and one optimization-based, and identifies which statistic of the minimizer targets.
Part 2 — Logistic Regression
What is logistic regression? Why do logarithms show up in its formulation or in its loss function?
Hint — Where the log enters: Once the model is written out, the logarithm appears in two separate places. Follow the path from a raw probability in to an unconstrained linear score, and separately consider how the parameters are actually fit. Ask what each step would look like without a log and why it would fail.
Hint — The loss: For Bernoulli labels, the likelihood is a product of per-example probabilities. What does taking a do to a product, and why is that useful both mathematically, by turning the objective into a sum, and numerically?
What This Part Should Cover
- Correctly states what logistic regression models and writes out or describes the relationship between the linear score and the output probability.
- Identifies and explains both distinct places where a logarithm appears, one in the model form and one in the fitting criterion, with a clear reason why each is necessary or convenient.
- Explains the practical benefits of the log in the fitting criterion beyond simply saying "it is the MLE."
Part 3 — Random Forest
What is a random forest? During tree construction, how is the set of candidate features chosen at each split?
Hint — Two sources of randomness: A random forest introduces randomness in two independent ways: how the data for each tree is sampled, and how features are considered at each split. Name both, and be ready to say which one the question is really probing.
Hint — Feature selection at a split: Consider whether each split may examine all features or only a restricted random subset, and which tuning knob controls that number. Then push on why deliberately hiding features from a split can make the overall ensemble better rather than worse.
What This Part Should Cover
- Gives a clear definition of the ensemble and how its prediction is formed, whether by voting or averaging.
- Names both sources of randomness accurately, including the relevant hyperparameter for the feature-selection mechanism and common default values.
- Provides a principled explanation grounded in ensemble theory, not just intuition, for why restricting features at each split improves the ensemble's generalization.
Part 4 — Adam vs. SGD
Explain the Adam optimizer. What are its advantages and disadvantages relative to vanilla stochastic gradient descent?
Hint — What state Adam keeps: Adam combines two ideas you have seen elsewhere by maintaining per-parameter running statistics of the gradient stream. Which two quantities about recent gradients would each idea track, and how would the update combine them? Once you name them, write the moving-average updates, the bias-correction step, and the final parameter update.
Hint — Trade-offs to weigh: Be honest about both sides: faster early convergence and per-parameter adaptive step sizes versus extra memory, two states per parameter, and the documented generalization gap compared with well-tuned SGD with momentum. Mention how weight decay interacts with Adam, contrasting L2 regularization with decoupled weight decay or AdamW.
What This Part Should Cover
- Correctly names and describes the two running statistics Adam maintains per parameter and the conceptual ideas each one comes from.
- Explains the bias-correction step, what causes the bias and why correction is needed, and writes or describes the final update formula.
- Provides at least two concrete advantages and at least two concrete disadvantages or caveats relative to SGD, including the weight-decay interaction.
Part 5 — Narrow vs. Wide Networks and Local Minima
Consider two neural networks with the same two-layer structure. One has only a few neurons per layer, while the other has many neurons per layer. Which one is more likely to get trapped in a poor local minimum, and why?
Hint — Frame it as capacity and landscape: Both objectives are non-convex. Reason about how the number of parameters affects the number of low-loss configurations and how connected the good solutions are in the loss landscape, contrasting isolated bad basins with wide connected low-loss regions.
Hint — Don't forget the trade-off: A complete answer names which network is more prone to poor local minima or underfitting, and also flags the cost of the easier-to-optimize one: what does extra capacity risk when data or regularization is limited?
What This Part Should Cover
- Makes a clear, unambiguous choice of which architecture is more susceptible, with a capacity-based justification.
- Explains the loss-landscape argument for why the other architecture is easier to optimize, going beyond "more parameters equals better."
- Acknowledges the countervailing risk that comes with the easier-to-optimize network and names at least one practical mitigation.
What a Strong Answer Covers
The interviewer is listening for these cross-cutting signals across all five parts. This is a checklist of dimensions the interviewer scores, not the answers themselves.
- Explicitly stated assumptions for linear and logistic regression, with awareness of which ones matter for unbiased point estimates versus valid inference.
- Probabilistic grounding: connecting squared loss and log-loss to maximum likelihood under specific noise or label models.
- The mechanism of randomness in ensembles and why it helps through variance reduction via decorrelation, not just "it is a bunch of trees."
- Optimizer internals: what state Adam maintains, the actual update rule, and honest trade-offs versus SGD, including memory, generalization, tuning, and weight decay.
- Non-convex optimization intuition for narrow versus wide networks, including capacity, the structure of the loss landscape, and overfitting risk.
- Calibrated nuance: acknowledging where the textbook answer is incomplete, or where practice diverges from theory.
Follow-up Questions
- For squared loss: how would your answer change if the noise were heavy-tailed, for example Laplacian, instead of Gaussian? What loss would maximum likelihood give you then, and which statistic of would it estimate?
- For random forests: how do
n_estimatorsand the feature-subset size trade off bias, variance, and decorrelation between trees? - For Adam: in what concrete settings have you seen, or would you expect, SGD with momentum to generalize better, and what would you try to close the gap?
- For the narrow-vs-wide question: how does the modern overparameterization view, including loss-landscape connectivity and flat versus sharp minima, reconcile with the classical "more parameters leads to more overfitting" intuition?
Overview: This question evaluates core machine learning fundamentals: statistical modeling assumptions and loss functions for linear and logistic regression, ensemble methods and feature sampling in random forests, optimization algorithms such as Adam versus stochastic gradient descent, and neural network capacity and training dynamics.