Study
ML Theory Questions for Senior Engineers: 100 Reported (2026)
TrueInterview
October 11, 2026 · 21 min read

As of 2026-10-11, TrueInterview's bank holds 100 candidate-reported ML fundamentals questions from 17 big-tech companies, and 76 of them were reported in phone screens. Amazon files 26 and ByteDance 19. Across the bank, the named rounds run from dropout and random forests to LoRA, RLHF and calibration. Our take: treat machine learning theory as a screen-stage filter on that modern topic set, build your drill list per company and per round, and at senior level prepare to diagnose a failing model rather than recite definitions.
Disclosure: TrueInterview is an interview-preparation product and publishes this article. Facts about other products come from their public pages on the dates listed under Sources.
Which ML theory questions are candidates reporting at big tech?
Eight named rounds carry the most traceable reports, each tied to a single company and each linked from between four and eight candidate write-ups. They range from Pinterest's quick-fire screen to DoorDash's onsite discussion round, so read the list by format first and by topic second: the format decides how long and how deep your answer should be.

| Question | Company | Round | Last reported | Linked candidate write-ups |
|---|---|---|---|---|
| ML Fundamentals & Model Debugging Drill | Apple | Phone screen | 2026-06 | 8 |
| ML Fundamentals Deep Dive (AI/ML & MLE Roles) | Phone screen | 2026-06 | 8 | |
| ML Fundamentals Quick-Fire | Phone screen | 2026-06 | 8 | |
| ML Fundamentals, Transformer & Regularization | Snapchat | Phone screen | 2026-05 | 8 |
| Dropout, Overfitting, Normalization, Loss Functions | ByteDance | Phone screen | 2026-06 | 7 |
| ML Knowledge / Discussion Round | DoorDash | Onsite | 2026-07 | 6 |
| ML Breadth Orals — Linear / Logistic / Random Forest / Optimizers | Amazon | Not stamped | Not stamped | 4 |
| LoRA and PEFT Variants | Amazon | Not stamped | Not stamped | 4 |
The same regularization topic sits inside Pinterest's quick-fire screen, Snapchat's transformer round and ByteDance's oral cluster. In a quick-fire round you owe one precise sentence and a mechanism; in a deep dive you owe the follow-up two layers down; in a discussion round you owe a trade-off tied to a real system. For regularization, practise three depths: one sentence at quick-fire pace, a second-layer follow-up of your own choosing (for example, weight decay under Adam) for a deep dive, and a shipped-model trade-off for a discussion round.
ML Fundamentals & Model Debugging Drill (Apple)
A debugging drill hands you symptoms and grades the order of your diagnosis. The standard order is: compare training and validation curves, rule out data problems such as leakage, duplicated rows across splits, label noise and train-serving skew, check the optimizer for a learning rate that diverges or stalls, and only then change capacity or regularization. The decision an interviewer will push on at senior level is evidence: how you would know the fix worked, which means a held-out slice or an ablation rather than a better training curve. The worked example in the senior section below walks through one.
ML Fundamentals Quick-Fire (Pinterest) and Deep Dive (Google)
Quick-fire rewards a crisp definition followed by the mechanism that explains it. GeeksforGeeks' question list describes L1 regularization as adding the absolute value of weights, which "can shrink some weights to zero and perform feature selection", and L2 as reducing large weights without eliminating them. The mechanism is what turns that into a strong answer: the absolute-value penalty has a corner at zero, so the optimum often lands exactly on it, while the squared penalty is smooth and shrinks every weight proportionally. In Bayesian terms, the lasso penalty is a Laplace prior on the weights and the ridge penalty a Gaussian prior.
A deep dive asks the next question down, so prepare your own follow-ups. Two worth having ready: when you would prefer elastic net (correlated features, where the lasso alone picks one arbitrarily), and why weight decay differs from a squared-weight penalty under Adam (adaptive scaling weakens the penalty on weights with large gradients, which is why AdamW decouples the decay from the gradient step).
ML Fundamentals, Transformer & Regularization (Snapchat) and attention deep-dives
Amazon's Transformer / Attention Deep-Dive was last reported at a phone screen in 2026-08, so attention now belongs to the screen, not only to research orals. Be able to write scaled dot-product attention, softmax(QKᵀ/√d_k)V, and explain the scaling: dot products of random vectors grow in variance with dimension, which pushes the softmax into saturated regions where gradients vanish. Then cover what multiple heads buy, why the causal mask exists, and where dropout and layer normalization sit in the block.
The decisions a senior candidate gets pushed on are cost and stability. Self-attention is quadratic in sequence length, and at inference the KV cache stores past keys and values so each new token computes only its own query, which makes decoding memory-bound. Pre-norm blocks train more stably than post-norm blocks at depth, which is why most large models use them.
Dropout, Overfitting, Normalization, Loss Functions (ByteDance)
This oral cluster tests deep learning mechanics. For dropout, explain inverted dropout, which scales the surviving activations up during training so inference needs no change, and the reading of dropout as an approximate ensemble of thinned networks. For normalization, contrast batch normalization, which normalizes each feature across the batch and switches to running averages at inference, with layer normalization, which normalizes across features within one example and so does not depend on batch size or sequence length; that independence is why transformers use it.
For loss functions, know why cross-entropy beats mean squared error for classification: with a sigmoid or softmax output, the cross-entropy gradient with respect to the logits is the prediction minus the label, while squared error multiplies in the activation's derivative and stalls when the output saturates.
ML Knowledge / Discussion Round (DoorDash)
This is the one onsite round in the table, and it is a conversation rather than a quiz. Anchor each answer in a model you have shipped: the data you had, the constraint that bound you (latency, label delay, cost), and the alternative you rejected and why. A discussion round is where a strong definition with no production context reads as a junior answer.
Are ML theory questions asked in the phone screen or the onsite?
Mostly in the phone screen. Reported rounds for the 100 questions run phone screen 76, onsite 34 and online assessment 1, and one question can count in more than one round. If your screen is inside three weeks, make screen-shaped theory your first practice block and save discussion-round depth for after you clear it.
The subject does not change between the two stages: all 76 phone-screen questions and all 34 onsite questions carry the ML fundamentals subtype. What changes is the format. The yuan-meng.com post on MLE interviews describes the ML fundamentals round as "rapid-fire questions on machine learning foundations", naming architectures, optimization routines, loss functions, activation functions and regularization. That matches the screen: breadth at pace, with no time to build an answer from first principles.
Onsite theory is more applied. One candidate's write-up on Taro (jointaro.com) reports an onsite of four one-hour sessions in which machine learning theory was separate from live coding, statistics and machine learning system design. The same write-up says "The machine learning theory was actually very much applied to a real ranking problem", and lists prompts including "Talk about the bias-variance trade-off." and "What metrics would you use for a ranking algorithm?". For that last prompt, the standard answer separates offline ranking metrics such as NDCG, MAP and recall at k from the online metric the business cares about, and explains why the two can disagree.
Our take: the common advice is to study ML theory in the final week before the onsite. The reported distribution says the opposite: theory is front-loaded in the screen, and the onsite theory session, where it exists, asks you to apply the same concepts to a concrete problem. Drill Pinterest-style recall in week one and DoorDash-style application in week three.
Which companies report the most ML theory questions?
Amazon files the most, 26 of the 100, followed by ByteDance with 19, and no other company reaches ten. Every question filed under Amazon and under ByteDance carries the ML fundamentals subtype. These are filing counts of candidate reports, not how often a company asks, so use them to decide where your practice list has evidence, not to judge a company's bar.
| Company | Questions in our bank | Named round or example in the bank | What to do (our view) |
|---|---|---|---|
| Amazon | 26 | Breadth orals on linear and logistic regression, random forests and optimizers; LoRA and PEFT variants; Transformer / Attention Deep-Dive | Largest share of theory time: the classic floor plus PEFT and attention |
| ByteDance | 19 | Dropout, overfitting, normalization, loss functions | Deep learning mechanics at oral pace |
| Snapchat | 9 | ML fundamentals, transformer and regularization | Attention plus regularization, both directions |
| 8 | ML fundamentals deep dive for AI/ML and MLE roles | Follow-ups two layers below the definition | |
| Microsoft | 7 | No named round in this extract | The classic floor, then the modern pass |
| Meta | 6 | No named round in this extract | The classic floor, then the modern pass |
| Apple | 5 | Model debugging drill | Ordered diagnosis narratives |
| 4 | Quick-fire | Timed recall, one sentence plus a mechanism | |
| NVIDIA | 3 | GPU and Inference Systems Fundamentals, phone screen, 2026-02 | Memory bandwidth, batching, KV cache, quantization |
| OpenAI | 3 | No named round in this extract | The modern pass plus research depth |
Two modern examples come from companies outside that list. Anthropic's "Explain AI Safety and Weigh Advanced AI Benefits and Risks" was last reported at a phone screen in 2026-08, and Netflix's ML Research Orals (Self-Attention / LoRA / Optimizers) at a phone screen in 2026-05. A frontier-lab or research-flavoured loop can put a safety discussion or a research oral in the screen itself.
Our take: generic lists give every candidate the same flat syllabus. If Amazon or ByteDance is on your list, their volume justifies working through their filed questions one by one. If your target files three or four questions, there is too little to memorise; spend that time on the format of its named round and on the modern pass below.
Do the classic fundamentals still get asked?
Yes, as the floor rather than the differentiator. Amazon's breadth orals cover linear and logistic regression, random forests and optimizers, and ByteDance's oral cluster covers dropout, overfitting, normalization and loss functions. Prepare these to a clean, fast standard, then stop: past the floor, more classic theory buys little.
One ML interviewer's Medium post (janiebrooke.medium.com) describes the floor as gradient descent, regularization, the bias-variance trade-off, feature engineering principles "and the properties of the algorithms you claim to have used". The same post reports: "Above that floor, additional theory knowledge is largely table stakes, it does not differentiate you from other candidates who also cleared the floor." That is one interviewer's view, but it matches the level data in the senior section: the drills that skew senior are the ones that grade judgment.
Two classic answers are worth getting exactly right because they recur in reports. For bias and variance, state the decomposition of expected squared error into squared bias, variance and irreducible noise, then explain that regularization trades a little bias for a larger cut in variance. At senior level, add that heavily overparameterized networks can show double descent, where test error falls again past the interpolation point, so the classic U-shaped curve is not the whole story.
For random forest against logistic regression, the Taro write-up above lists the prompt "When would you choose a Random Forest or a Logistic Regression model?". Pick a winner by situation. Logistic regression wins when you need interpretable coefficients, reasonably calibrated probabilities, cheap serving or a monotonic response that extends beyond the training range. A forest wins on tabular data with nonlinear interactions you do not want to engineer by hand, at the cost of larger models, poorer calibration and no extrapolation beyond the training range.
Which modern topics separate strong candidates?
Transformers, parameter-efficient fine-tuning, RLHF, calibration and drift, inference systems and AI safety all appear as named questions in the bank, and two of them, the drift and calibration round and the RLHF round, draw their reported links from senior-or-above candidates. This is the set the classic listicles miss, and the set where a strong candidate stands apart.
The yuan-meng.com post on MLE interviews gives the shift in one contrast: a few years ago many companies asked candidates to implement a simple model such as KNN or logistic regression from scratch, while today "frontier labs may ask you to debug or implement language model training or inference code", naming Transformer encoders and decoders, LoRA, KV cache, beam search and autograd.
LoRA and PEFT variants
Amazon's LoRA round is linked from 4 candidate write-ups, and Netflix's research orals name LoRA in the title. Explain LoRA as a frozen pretrained weight plus a learned low-rank update, the product of two thin matrices, with one of them initialized to zero so training starts from the base model's behaviour. The update can be merged into the weight after training, so it adds no inference latency. Then compare variants: adapters insert small bottleneck layers and add latency, prefix and prompt tuning learn virtual tokens, QLoRA keeps the frozen base quantized to save memory, and DoRA separates magnitude from direction.
The decisions you will be pushed on are rank and placement (which projection matrices get adapters, and what a higher rank buys), when full fine-tuning beats LoRA (a large domain shift that needs more capacity than a low-rank update gives), and how one base model serves many adapters.
RLHF: PPO vs GRPO vs GSPO
This round is about why the newer methods exist. PPO optimizes the policy against a reward model trained on human preference comparisons, with a clipped probability ratio, a learned value function for advantages and a KL penalty that keeps the policy near the reference model and limits reward hacking. GRPO drops the value function: it samples a group of responses per prompt and uses each response's reward relative to the group's mean and spread as its advantage. GSPO moves the importance ratio and the clipping from the token level to the sequence level, which addresses the instability that token-level ratios cause on long responses.
StackScholar's ML fundamentals page lists describing RLHF as "training on human answers" as a mistake, because it trains on human rankings, and lists calling next-token prediction unsupervised as another, because it is self-supervised. Both are quick ways to lose a senior screen.
Optimization, drift and calibration
Calibration means that among predictions made with a given confidence, that fraction turns out correct. Measure it with a reliability diagram and expected calibration error, and fix it after training with temperature scaling, Platt scaling or isotonic regression; temperature scaling changes confidence without changing the predicted class. Modern deep networks are often overconfident, and label smoothing or mixup-style augmentation also changes confidence, so recheck calibration after every training change.
For drift, separate covariate shift (inputs change, the input-to-label relationship holds), label shift (class priors change) and concept drift (the relationship itself changes). Monitor feature and prediction distributions, recheck calibration once delayed labels arrive, and pick the remedy by type: importance weighting for covariate shift, recalibration for label shift, retraining for concept drift.
GPU and inference systems, and AI safety
For NVIDIA's inference round, reason from arithmetic intensity: decoding a single sequence is memory-bandwidth bound, so batching, KV-cache management, quantization and fused kernels that avoid writing the attention matrix to high-bandwidth memory are the levers. For Anthropic's AI safety prompt, take a position: name specific risks such as misuse and reward hacking, the mitigations you would rely on, including evaluations, red-teaming and preference training, and where you think those mitigations fall short.
Our take: spend the floor's time on the classics and everything above it on this section, with dated questions from the bank. A list that stops at SVM kernels and PCA prepares you for the questions that differentiate least.
How recent are the questions?
Recent enough to trust as a live set: 75 of the 100 questions were last reported within the 12 months before 2026-10-11. Every month-stamped example above falls in the current year, from NVIDIA's inference round to Amazon's attention deep-dive. Use the month as a filter: a recent report earns a slot in a three-week plan, while an undated list entry becomes background reading.
Other sites show what a single report looks like. AceOffer's page for one ML fundamentals oral describes a one-hour oral Q&A with no coding that runs from overfitting and underfitting to more advanced topics, and labels it "Reported 1× across candidate reports", last reported November 2025. One report tells you a format exists; it does not tell you it recurs.
What changes at senior and staff level?
The subjects stay the same; the grading moves from definition to diagnosis. Write-ups link these questions 66 times from candidates who reported a senior, staff-plus or manager level and 38 times from candidates who reported junior, new-grad or intern. Levels are self-reported and these are link counts, not distinct write-ups, but the skew by drill is the useful part.
| Drill | Senior-or-above write-ups | Junior, new-grad or intern write-ups | What a senior answer adds (our view) |
|---|---|---|---|
| ML Fundamentals & Model Debugging Drill | 7 | 1 | An ordered diagnosis and proof the fix worked |
| ML Knowledge / Discussion Round | 5 | 1 | Trade-offs tied to a system you shipped |
| ML Fundamentals Deep Dive (AI/ML & MLE Roles) | 5 | 3 | The second follow-up, not just the first |
| ML Fundamentals, Transformer & Regularization | 5 | 3 | Cost and stability, not only the formula |
| Deep Learning Fundamentals: Optimization, Drift, Calibration | 4 | No count reported | Which drift type, and the matching remedy |
| LoRA and PEFT Variants | 3 | 1 | Rank, placement and serving many adapters |
| RLHF: PPO vs GRPO vs GSPO | 3 | No count reported | Why each method replaced the last |
| ML Fundamentals Quick-Fire | 4 | 4 | Precision at pace; the round itself does not change |
Quick-fire is the only even split, and it is the only drill where a definition is the whole answer. Every drill that asks you to diagnose, discuss or compare skews senior. If you are senior or staff, assume your theory round will look like the top of this table.
InterviewNode's article on ML interview mistakes says: "As seniority increases, interviewers become less tolerant of gaps in judgment, even if raw ML knowledge is strong." StackScholar's fundamentals page puts the gap more bluntly: everyone can define overfitting, but "Very few can say how they detected it last time, what they tried first, and why." Our take: for every classic topic, prepare a failure-mode version: the metric that moved, the first thing you tried, why it did not work, and how you confirmed the change that did.
Worked example: a training and validation gap
StackScholar's page lists this prompt: "Your model gets 0.98 training accuracy and 0.71 validation accuracy. Walk me through what you do." A junior answer says overfitting and reaches for dropout. A senior answer checks the cheap explanations first:
- The split. If validation comes from a later period or a different user group, the gap may be distribution shift rather than variance, and regularization will not close it.
- The data. Duplicates within training inflate training accuracy, and label noise caps validation accuracy.
- The metric. Under class imbalance, accuracy can hide a model that ignores the minority class; check precision and recall per class.
- The diagnosis. Plot learning curves. If validation accuracy keeps rising with more data, it is variance: add data or augmentation, weight decay or early stopping, one change at a time.
- The proof. Confirm the winning change on a held-out slice the tuning never touched, and say what result would have made you abandon it.
InterviewNode's article lists the reasoning interviewers listen for: why the current model is failing, what constraints exist in data, latency and cost, what risks added complexity introduces, and how you would know the change helped. The five steps above answer all four.
Which rounds decide whether you are down-leveled?
Our take: the theory screen mostly decides whether you continue, and the level is set elsewhere: in the depth of the debugging and discussion rounds, in the project deep dive or research presentation, and in the scope of your behavioral stories. A clean quick-fire round proves the floor; it rarely proves a level.
The same Medium post (janiebrooke.medium.com) reports that the behavioral round "is the round that senior candidates most frequently underestimate, and it is the round that has the most influence on leveling decisions." The yuan-meng.com post adds that some research engineer roles require "a job talk style presentation on your past work" instead of a verbal project walk-through. Eugene Yan's guide on eugeneyan.com describes science depth interviews as inviting candidates to showcase expertise on a project of their choice.
If you are interviewing at senior or staff level, ask the recruiter which round carries the project or research depth, and prepare it as the level round. Pick a project where you can name the failure you found, the alternatives you rejected and the result you measured.
How much of your prep should ML theory get?
It depends on your stage. If your screen is next and inside three weeks, theory should take most of your hours, aimed at the company with the most reports on your list. If you are past the screen and into onsite preparation, keep theory at maintenance level and move the hours to ML system design and behavioral.
The Medium interviewer's post reports: "Most candidates over-prepare for the theory round and under-prepare for the system design and behavioral rounds." Interview101's Amazon MLE guide says candidates who completed those loops reported that "the ML system design round focuses heavily on cost estimation, operational failure modes, and infrastructure decisions", which is a different skill from fundamentals recall. Our take: if you are senior, split your theory block in half, classics to the floor and a dated modern pass with failure-mode narratives, and stop there.
How we counted
Counts come from TrueInterview's question bank as of 2026-10-11 and cover the ML fundamentals questions filed under big-tech companies. They are candidate-reported questions reconstructed for practice, not any company's official question list. A question can be reported in more than one round, so round tallies overlap; "Not stamped" means our extract gave no round or month for that question; and level splits count write-up links by the level candidates reported for themselves, which is not a verified level or a count of distinct write-ups.
FAQ
Are these the machine learning questions big tech companies officially ask?
No. They are questions candidates reported after real interviews, reconstructed for practice and filed under the company where the candidate met them; no employer published them. Use them to decide what to drill first and to see which round and format a company favours. A topic you can explain at the right depth transfers to a question you have never seen, while a memorised answer fails at the first follow-up.
Is a machine learning interview questions PDF or GitHub list enough?
As background, yes; as a drill list, no. The andrewekhalel/MLQuestions repository on GitHub lists 69 questions with answers across fundamentals, deep learning, vision, NLP and statistics, and has been "Curated and community-maintained since 2018." Lists like that have no company, round or month attached, so they cannot tell you what is current. Use one to fill gaps in the classic floor, then spend your drill time on dated, round-specific questions.
Do freshers and new grads get different ML theory questions than senior candidates?
Mostly the same subjects, with a different emphasis. The level table above shows quick-fire splitting evenly between junior and senior candidates, while debugging and discussion drills skew senior. For a new grad, fresher lists such as InterviewBit's, which include definitions such as classification versus regression and cross-validation, cover the floor. Drill those until they are fast, then add the transformer and regularization material that appears in screens.
Will an ML theory round ask me to write code?
Usually not in the theory round itself, which the sources describe as oral. AceOffer's page on one ML fundamentals oral calls it a one-hour Q&A with no coding. Interview101's Amazon guide says candidates should not expect ML-specific coding such as implementing backpropagation or custom loss functions in the coding round, which uses LeetCode-style problems. Research-flavoured loops are the exception: the yuan-meng.com post's frontier-lab examples above are implementation tasks.
What is the most common mistake in ML theory rounds?
Answering like an exam. InterviewNode's article says one of the most expensive mistakes is that candidates "treat interviews as exams instead of judgment evaluations". InterviewNode's example of a shallow answer to improving a model is "I'd try a more complex architecture, add more features, and tune hyperparameters." Replace that with the diagnosis order from the worked example: what is failing, what you checked first, and what would prove the fix.
Which deep learning topics should I prepare beyond the classics?
Attention and the transformer block, LoRA and its variants, RLHF methods, calibration and drift, and inference cost, in roughly that order for most MLE loops. Each appears above with its company and round. If your target is a frontier lab or an infrastructure team, move inference systems and safety earlier, since those prompts appeared in phone screens rather than onsites.
Your practice plan
The target is narrower than all of machine learning: a set of named rounds, the companies that file them, and a month stamp on each. This plan assumes four or five evenings a week for three weeks while you work full time.
- Week one, first evenings: pick your target company from the company table. If it is Amazon or ByteDance, work through its filed questions in TrueInterview's question bank, starting with the classic floor: Dropout, Overfitting, Normalization, Loss Functions and the bias-variance and random forest answers above.
- Week one, remaining evenings: run ML Fundamentals Quick-Fire on a timer, one precise sentence plus a mechanism per topic, until no answer needs a pause.
- Week two: do the modern pass with Transformer / Attention Deep-Dive and ML Fundamentals, Transformer & Regularization, then LoRA, RLHF and calibration from the sections above. Write each answer with the decision a senior interviewer would push on.
- Week two, if you are senior or staff: work ML Fundamentals & Model Debugging Drill and write a failure-mode narrative for five classic topics, using the worked example's five steps.
- Week three: run ML Fundamentals Deep Dive (AI/ML & MLE Roles) as a mock and record yourself, checking that every answer survives a second follow-up.
- Week three, if your loop includes a discussion or research round: rehearse ML Knowledge / Discussion Round or ML Research Orals around one project you can defend end to end. For infrastructure or frontier-lab targets, add GPU and Inference Systems Fundamentals or Explain AI Safety and Weigh Advanced AI Benefits and Risks.
- After the screen: drop theory to one evening a week and move the rest to ML system design and behavioral stories, using the down-leveling section to pick which round to over-prepare.
Sources
- TrueInterview question bank — ML fundamentals at big-tech companies — counted 2026-10-11
- Machine Learning Interview Questions and Answers - GeeksforGeeks — checked 2026-10-11
- MLE Interview 2.0: Research Engineering and Scary Rounds | Yuan Meng — checked 2026-10-11
- Expedia Group Machine Learning Scientist III Interview Experience - London, United Kingdom — checked 2026-10-11
- Medium — checked 2026-10-11
- Machine Learning Fundamentals Interview Questions (with Diagnostics) | StackScholar — checked 2026-10-11
- ML Fundamentals (Oral) — OpenAI Interview Question | AceOffer — checked 2026-10-11
- Mistakes That Cost You ML Interview Offers (and How to Fix Them) - Interview Node Blog — checked 2026-10-11
- How to Interview and Hire ML/AI Engineers — checked 2026-10-11
- Amazon's MLE Interview Separates ML Theory From Production Systems — And Most Candidates Prepare for the Wrong One — checked 2026-10-11
- GitHub - andrewekhalel/MLQuestions: Machine Learning and Computer Vision Engineer - Technical Interview Questions · GitHub — checked 2026-10-11
- Top Machine Learning Interview Questions & Answers (2025) - InterviewBit — checked 2026-10-11
Last reviewed: 2026-10-11.