Amazon · ML & AI Fundamentals
Derive and compare core ML and RL methods
TrueInterview
October 7, 2026 · 2 min read
Work through the following machine learning fundamentals rigorously: state the assumptions, give the equations, and justify the trade-offs.
-
Derive the update rules for full-batch gradient descent (GD) and stochastic gradient descent (SGD) for . Compare their convergence behavior, gradient variance, and wall-clock efficiency, and explain the conditions under which SGD outperforms GD.
-
Define batch size. Given samples, 5 epochs, and batch size , compute the number of update steps per epoch and in total. If is increased to 2,000, compute the new step counts and propose a learning-rate adjustment based on linear scaling; explain when this rule fails.
-
Classify each of the following algorithms as supervised or unsupervised and give one use case for each: logistic regression, SVM, k-NN, k-means, PCA, t-SNE, and Isolation Forest.
-
Explain how reinforcement learning relates to supervised and unsupervised learning. Write the REINFORCE gradient and show how a baseline keeps the estimator unbiased while lowering variance; express the gradient for a length-3 trajectory with returns and score-function terms using a constant baseline .
-
Explain how neural networks are used in RL, for example in DQN, policy gradient, and actor-critic methods. For DQN, describe why target networks and experience replay make training more stable, and give a failure mode that occurs without them.
-
Compare Transformers and RNNs in terms of parallelism, long-range dependency handling, and complexity. For sequence length and model dimension , estimate the asymptotic time and memory cost of self-attention, and name two techniques that reduce the quadratic scaling.
-
Define embeddings and polysemy. Propose a method to distinguish “King” in chess contexts from monarchy contexts using contextual encoders or multi-sense embeddings; outline an intrinsic evaluation based on word sense disambiguation (WSD) and an extrinsic evaluation based on downstream accuracy.
-
Given a single 24-GB GPU and a 7B-parameter model, design a low-compute fine-tuning plan, such as QLoRA or adapters, 4-bit quantization, gradient checkpointing, and mixed precision. Choose a LoRA rank and specify batch size, sequence length, optimizer, and learning-rate schedule. Provide a back-of-envelope estimate of trainable parameters assuming hidden size and about 32 layers; state any assumptions about which projections you adapt.
Overview: This question tests a candidate's command of core machine learning and reinforcement learning ideas, including optimization (gradient methods and batch-size trade-offs), supervised versus unsupervised algorithms, policy-gradient RL and variance reduction, deep RL stabilization techniques, sequence model complexity (Transformers versus RNNs), embeddings and polysemy, and low-compute fine-tuning strategies. It is often used to assess both theoretical understanding and practical engineering judgment around convergence, variance, computational and memory complexity, representation learning, and resource-constrained model adaptation; the domain is Machine Learning, and the assessment covers conceptual understanding and practical application.
Read the full data scientist interview experience this question came from.