Study
ML Coding Questions for Senior Engineers: 81 Reported (2026)
TrueInterview
October 11, 2026 · 18 min read

As of 2026-10-11, TrueInterview's bank holds 81 candidate-reported ML coding questions from 19 big-tech companies, and 34 of its 55 algorithm-format questions carry the math tag. The named examples are from-scratch numerics: Hand-Written K-Means at Microsoft, Implement Scaled Dot-Product Attention at Meta and Hand-Code Self-Attention and Cross-Entropy at ByteDance, each reported at a phone screen between 2026-06 and 2026-07. Our take: the ML coding round is not LeetCode in disguise. Move most of your hard-DSA hours into writing attention, softmax, k-means and backpropagation in NumPy from a blank file, then weight the rest by company and by stage.
Disclosure: TrueInterview is an interview-preparation product and publishes this article. Facts about other products come from their public pages on the dates listed under Sources.
Which ML coding questions do candidates report most?
Seven of the eight most-linked questions sit at OpenAI or Anthropic. Six are ML implementation or debugging tasks, and the other two are an agent build and Coinbase's ML Concepts and Coding Assessment. The table ranks them by how many candidate write-ups link to each, with the split by the candidate's self-reported level. Work it top-down if a frontier lab is on your list.

| Rank | Question | Company | Write-ups linking it | Senior-or-above vs junior links | Practice page |
|---|---|---|---|---|---|
| 1 | Debug a Transformer and Convert It to a Classifier | OpenAI | 14 | 10 vs 4 | Debug a transformer |
| 2 | Agents / Coding with LLMs | Anthropic | 12 | 7 vs 5 | Agents and LLM tool use |
| 3 | Classifier with Noisy Annotators | OpenAI | 11 | 8 vs 3 | Not linked in this guide |
| 4 | Vectorized 1-NN and Neural Network Forward Pass | OpenAI | 11 | 9 vs 2 | Not linked in this guide |
| 5 | Prefix Matrix Products and Backpropagation | OpenAI | 7 | 6 vs 1 | Not linked in this guide |
| 6 | Online Softmax Entropy | OpenAI | 6 | 5 vs 1 | Not linked in this guide |
| 7 | RL Fundamentals — GRPO Debug | Anthropic | 4 | 4 senior-or-above | Not linked in this guide |
| 8 | ML Concepts and Coding Assessment | Coinbase | 4 | No level split reported | Coinbase ML assessment |
A write-up count measures what candidates reported and wrote up, not how often a company asks a question. Treat the ranking as evidence of where preparation has paid off for other people, not as a forecast of your loop. The sections below give the technique that cracks each of the top OpenAI questions, phrased as standard practice rather than as a claim about any interviewer.
Our take: generic ML question lists spend most of their length on definitions such as bias and variance. None of the top eight is a pure definition question. Almost every one asks you to produce working code, or to find the bug in someone else's, under time pressure.
Debug a Transformer and Convert It to a Classifier (OpenAI)
The title names two jobs: find the bugs in a transformer, then repurpose it for classification. In hand-written transformer code, bugs cluster in a few places: softmax taken over the query axis instead of the key axis, the square-root scaling missing, a causal mask that is inverted or applied after the softmax, a head split that reshapes without transposing, and a loss that counts padding tokens. For the conversion, pool the sequence (the first token, or a mean over non-padding positions using the mask), add a linear head that outputs class logits, and train with cross-entropy over classes.
The senior-level decision is how you localise a bug, not how fast you spot one. Check that the loss at initialisation is close to the log of the number of classes, then confirm the model can overfit a single small batch. If it cannot, the bug is in the model or the loss; if it can, look at the data pipeline and the optimiser.
Vectorized 1-NN and Neural Network Forward Pass (OpenAI)
This question rewards broadcasting over loops. For nearest-neighbour search, expand the squared distance as the squared norm of each query, minus twice the dot product, plus the squared norm of each training point, then take the argmin over the training axis. For the forward pass, write each layer as a matrix product plus a broadcast bias, and annotate every array's shape in a comment as you go.
Expect the push on memory. The distance matrix is queries times training points, so say when you would process queries in chunks, and clamp the expanded distances at zero because round-off can make them slightly negative.
Prefix Matrix Products and Backpropagation (OpenAI)
For a product of matrices, the gradient with respect to one factor is the transpose of the product to its left, times the upstream gradient, times the transpose of the product to its right. Compute prefix products on the forward pass and suffix products on the backward pass, and keep the order straight, because matrix products do not commute.
The trade-off worth naming at senior level is memory against compute. Storing every prefix is fast but costs memory proportional to the chain length; recomputing them from a few saved checkpoints is the same idea as activation checkpointing in large-model training.
Online Softmax Entropy (OpenAI)
Softmax entropy can be computed in one pass with constant memory. Keep a running maximum, a running sum of exponentials shifted by that maximum, and a running sum of those exponentials weighted by the logit. When a new maximum arrives, rescale both sums by the exponential of the old maximum minus the new one. The entropy is the log-sum-exp minus the probability-weighted mean logit, which the code below computes. The rescaling step is the same trick streaming attention kernels use, which is a useful point to raise if the conversation turns to inference efficiency.
Classifier with Noisy Annotators (OpenAI)
Start from a majority-vote baseline, then show you know when it fails. Two standard upgrades are training on soft labels, where the target is each item's distribution of annotator votes, and estimating a confusion matrix per annotator with expectation-maximisation (the Dawid-Skene approach) so that reliable annotators count for more. The question a senior candidate should raise unprompted is validation: without a clean held-out set, you cannot tell whether the noise model helped or simply fitted the annotators' shared mistakes.
Is the ML coding round LeetCode or from-scratch ML?
In this bank it is mostly from-scratch ML. Math is the leading tag on the algorithm-format questions, and matrix and sorting have four each, while strings, arrays, greedy and hashing appear once each. That profile describes vectorised numerics and linear algebra, so a plan built on dynamic programming and graph sets prepares you for a different round.
Outside our data, one interviewer's account on Medium, from someone who has interviewed more than 100 ML candidates, says "The coding round in ML interviews is not LeetCode, and it is not Kaggle." The same account describes the round at production-ML companies as a data manipulation problem, an algorithm implementation problem or a debugging exercise on broken code. The yuan-meng.com post on MLE interviews says that a few years ago many companies asked candidates to implement a simple model from scratch or fit a scikit-learn model on toy data. The same yuan-meng.com post says "Today, frontier labs may ask you to debug or implement language model training or inference code", and names Transformer encoders and decoders, LoRA, KV cache, beam search and autograd as examples.
Amazon is the documented exception. Interview101's Amazon MLE guide says "Amazon's MLE coding round uses the same LeetCode-style data structures and algorithms problems as SDE interviews." The same guide says candidates should not expect to implement backpropagation or custom loss functions, which it describes as more common at research-focused companies. Yet the Amazon example in our bank is Write and explain gradient descent pseudocode, reported at an onsite in 2026-08.
Our take: the LeetCode-or-ML question has a company-specific answer. For the ML screens at OpenAI, Anthropic, Meta and ByteDance, from-scratch numerics should take most of your coding hours. For Amazon, split the time between standard DSA and one gradient-descent walkthrough you can explain line by line, because both have been reported.
Which companies and rounds should you weight first?
OpenAI holds 22 of the 81 questions, Amazon 10 and Anthropic 7, so the top three hold 39 between them. A flat topic list spends the same hour on a company with three reports as on one with twenty-two. The table pairs each company with its named example in our bank and the first thing we would implement for it.
| Company | ML coding questions in bank | Named example | Round, last reported | What to implement first (our view) |
|---|---|---|---|---|
| OpenAI | 22 | Debug a Transformer and Convert It to a Classifier | Phone screen, 2026-09 | The top OpenAI rows of the ranked table, in order |
| Amazon | 10 | Write and explain gradient descent pseudocode | Onsite, 2026-08 | Batch, mini-batch and stochastic updates with a convergence test |
| Anthropic | 7 | Agents / Coding with LLMs | Phone screen, 2026-08 | A tool-calling loop with a step limit and error handling |
| ByteDance | 6 | Hand-Code Self-Attention and Cross-Entropy | Phone screen, 2026-06 | Masked attention plus cross-entropy from logits |
| Apple | 5 | Transformer Attention Mask and Heads Coding | Phone screen, 2026-06 | Head split and merge, causal and padding masks |
| Microsoft | 5 | Hand-Written K-Means | Phone screen, 2026-07 | Vectorised assignment, empty clusters, a stopping rule |
| Meta | 4 | Implement Scaled Dot-Product Attention | Phone screen, 2026-07 | Scaling, masking and a stable softmax |
| 4 | MCQ + NN Forward + Coding + ML Implementations | OA, 2026-06 | A forward pass and short ML implementations at speed | |
| Uber | 3 | ML Coding from Scratch (Regression / Markov / Facility) | Phone screen, 2026-04 | Regression by closed form and gradient descent, then the Markov and facility parts |
| Coinbase | 3 | ML Concepts and Coding Assessment | OA, 2026-02 | Concept questions plus short coding under a timer |
Our take: if a frontier lab is your target, the OpenAI rows of the ranked table are the syllabus, and the other companies' examples are warm-ups. If you are interviewing for a broad MLE loop at Amazon, Apple, Microsoft, Meta or ByteDance, cover that company's named example first, because most are single, well-defined implementations you can finish in an evening.
How do phone screens differ from onsites?
The stage changes the format more than the company does. Reported rounds for the 81 questions are phone screen 48, onsite 39, OA 8 and take-home 2, and one question can count in more than one round. Among the 38 algorithm-format questions reported for phone screens, 23 are tagged math, 4 matrix and 3 sorting.
The onsite splits in two. Its algorithm-format questions are still math-led, with 13 of 22 tagged math. Its other-format questions are almost all builds: 16 of the 17 carry the AI/ML build subtype. Across every other-format question in the bank, AI/ML build accounts for 22 of 26 and ML coding for 4.
Our take: prepare two different skills. For the phone screen, drill timed numerics from a blank file until attention or k-means takes well under the length of the call. For the onsite, rehearse one end-to-end build you can narrate, from data loading through training to an evaluation you trust, because a build is judged on decisions as much as on code.
What does a senior-grade from-scratch implementation look like?
It is short, vectorised, numerically stable and checked. At senior level, a correct loop-based answer can still read as mid-level if you cannot say why it would overflow, how much memory a batch needs, or how you know the gradient is right. The guidance below is standard numerical practice, phrased as what to show the interviewer.
Numerical stability
Never exponentiate raw logits. Subtract the row maximum before the softmax, compute log-probabilities with log-sum-exp, and compute cross-entropy from logits rather than taking the log of a softmax output. For logistic regression, use the logits form of binary cross-entropy, which stays finite for large positive and negative scores, and remember that its gradient with respect to the logits is the predicted probability minus the label.
import numpy as np
def log_softmax(z, axis=-1):
m = z.max(axis=axis, keepdims=True)
return z - m - np.log(np.exp(z - m).sum(axis=axis, keepdims=True))
def cross_entropy(logits, y): # logits (N, C), integer labels y (N,)
return -log_softmax(logits)[np.arange(len(y)), y].mean()
def bce_with_logits(z, y): # stable for large |z|; dL/dz = (sigmoid(z) - y) / N
return (np.maximum(z, 0) - z * y + np.log1p(np.exp(-np.abs(z)))).mean()
def online_entropy(stream): # one pass over logits, constant memory
m, s, t = -np.inf, 0.0, 0.0 # running max, sum of e^(z-m), sum of e^(z-m) * z
for z in stream:
if z > m: # new max: rescale the old sums
scale = np.exp(m - z)
s, t, m = s * scale, t * scale, z
w = np.exp(z - m)
s, t = s + w, t + w * z
return m + np.log(s) - t / s # logsumexp(z) - E_p[z]
Masks deserve one sentence out loud. Setting disallowed scores to negative infinity is correct, but a row that is fully masked, such as a padding query, then produces NaN after the softmax. Say how you would handle it: a large finite negative value, or zeroing those rows after the softmax.
Vectorisation and memory
Replace loops over samples with matrix products and broadcasting, and state the shape of every intermediate. For multi-head attention, the common bug is reshaping to heads without the transpose that moves the head axis ahead of the sequence axis, and the same transpose is needed in reverse when you merge heads.
def attention(q, k, v, mask=None): # q (B, H, Tq, d); k, v (B, H, Tk, d)
scores = q @ k.swapaxes(-1, -2) / np.sqrt(q.shape[-1])
if mask is not None: # mask is True where attention is allowed
scores = np.where(mask, scores, -np.inf)
w = np.exp(scores - scores.max(-1, keepdims=True))
return (w / w.sum(-1, keepdims=True)) @ v # softmax over keys
def split_heads(x, h): # (B, T, D) -> (B, h, T, D // h)
B, T, D = x.shape
return x.reshape(B, T, h, D // h).transpose(0, 2, 1, 3)
def causal_mask(t): # query i may attend to keys j <= i
return np.tril(np.ones((t, t), dtype=bool))
def sq_dists(x, c): # x (n, d), c (k, d) -> (n, k), for k-means and nearest neighbour
d = (x ** 2).sum(1)[:, None] - 2 * x @ c.T + (c ** 2).sum(1)[None,:]
return np.maximum(d, 0) # round-off can push values slightly below zero
The memory question follows naturally. A distance matrix or an attention score matrix grows with the product of its two lengths, so name the point at which you would chunk the computation, and what that costs in speed.
Gradient checks
If you write a backward pass, check it. Use central differences in double precision on a small random input with non-square shapes, because square shapes hide transposition bugs. Compare with a relative error rather than an absolute one, since gradients vary in scale across parameters.
def grad_check(f, x, analytic, eps=1e-6): # x must be float64
num = np.zeros_like(x)
for i in np.ndindex(x.shape):
old = x[i]
x[i] = old + eps; fp = f(x)
x[i] = old - eps; fm = f(x)
x[i] = old
num[i] = (fp - fm) / (2 * eps)
rel = np.abs(num - analytic) / np.maximum(1e-12, np.abs(num) + np.abs(analytic))
return rel.max() # around 1e-7 or below is typical for smooth f
Edge cases to raise before you are asked
For k-means, say what happens to an empty cluster (re-seed it from the point farthest from its centroid, or keep the old centroid), how you initialise (k-means++ seeding usually converges to better solutions than uniform random seeding), and when you stop (assignments unchanged, or centroid movement below a tolerance). For standardisation, guard against features with zero variance. For a training loop, state the order: clear gradients, forward, loss, backward, step, and switch to evaluation mode with gradients disabled when you measure validation loss, so dropout and batch statistics behave correctly.
Where do debugging and AI-assisted ML questions fit?
They are a large share of the bank, not a side format. By format, the 81 questions split into algorithm coding 55, AI-assisted coding 22, object-oriented design 2 and behavioral or knowledge 2. The two AI-assisted examples at the top of the ranked table are OpenAI's transformer debug, last reported at a phone screen in 2026-09, and Anthropic's agents question, last reported at a phone screen in 2026-08.
The skills overlap with the numerics above. In our reading, a transformer debug rewards knowing where transformer bugs live, and an agent question rewards a loop that calls the model, parses a tool call, executes it, appends the result and stops on a clear condition or a step limit. The algoengineer.com article on AI-era interviews says that "even in AI-assisted rounds, the grade is on how you frame the problem, steer the tool, and catch its mistakes". Hello Interview's article on Meta's AI-enabled format says it replaced one of Meta's two coding rounds, while candidates still sit a traditional coding interview alongside it, so classic coding has not disappeared.
Our take: rehearse the transformer debug twice, once alone and once with an assistant, and write down every place the assistant's suggestion was wrong. The mechanics of the format itself are covered in TrueInterview's Study hub; this guide stays on the ML content those rounds test.
What changes at senior and staff level?
The data skews senior. Candidate write-ups link these questions 100 times from candidates who reported a senior, staff-plus or manager level and 45 times from those who reported junior, new-grad or intern. Those are link counts, not write-up counts, and the levels are self-reported, but the ratio is roughly two to one.
The skew is steepest on the numerics. The vectorised nearest-neighbour and forward-pass question is 9 senior-or-above links to 2 junior, prefix matrix products and backpropagation 6 to 1, and online softmax entropy 5 to 1. The GRPO debug question and Uber's ML Coding from Scratch appear only in senior-or-above write-ups in this data, with 4 links each. The most level-neutral question is Agents / Coding with LLMs, at 7 senior-or-above links to 5 junior.
| Your level | What the level data points to | What we think is being graded | Where to spend scarce hours |
|---|---|---|---|
| New grad or junior | Junior links still reach the transformer debug (4 links) and the agents question (5) | Correct, readable code with shapes stated | Named company examples first, then the top OpenAI rows |
| Senior (L5/E5) | Senior-or-above links outnumber junior ones about two to one | Stability, vectorisation and a gradient you can verify | The numerics rows of the ranked table, timed |
| Staff (L6+/E6+) | The GRPO debug and Uber's from-scratch question appear only in senior-or-above write-ups | The trade-off behind each choice: memory against compute, what you would validate and why | Debugging and build questions, narrated as design decisions |
Our take: at senior and staff level, the numerics questions carry most of the level differentiation inside the coding round, and the agents question carries the least. If you have one free evening, spend it explaining why your attention implementation is stable and how its memory grows, not polishing an agent loop.
For the GRPO debug, review the standard formulation before you start. Advantages are normalised within each prompt's group of sampled completions, not across the batch; a group with identical rewards needs an epsilon in the denominator; and the policy-ratio clipping follows PPO. Sign errors in the loss and log-probabilities taken over prompt tokens instead of completion tokens are the usual places to look.
Where the level signal sits (down-leveling note)
In our view, the coding round rarely decides the level on its own: it can confirm senior signal or remove it, but a clean implementation alone does not prove staff scope. One interviewer's account on Medium reports that the ML system design round is the most differentiating part of the loop for mid-level and senior roles. The same account reports that the behavioral round has the most influence on leveling decisions and is the one senior candidates most often underestimate.
Down-leveling risk in the coding round shows up as needing hints on stability or memory, or as a correct answer you cannot defend. Ask your recruiter which level the loop is calibrated for, and whether the ML coding round expects NumPy, PyTorch or a plain algorithm problem. If the answer is staff, put as many hours into the system design and behavioral rounds as into ML coding, and in the coding round explain each trade-off before the interviewer asks for it.
How we counted
Counts come from TrueInterview's question database, checked 2026-10-11, covering 81 ML coding questions and their reported rounds. A question can be reported in more than one round, link counts are links from candidate write-ups rather than numbers of write-ups, and levels are what candidates reported about themselves. These are candidate-reported questions reconstructed for practice, not any company's official question list. TrueInterview says engineers who have worked at companies that run these loops review each question by hand before it is filed.
FAQ
Are ML coding interview questions on LeetCode?
Mostly not, for the questions in this bank. Seven of the top eight are numerical implementations, debugging tasks or builds, and the eighth is Coinbase's ML Concepts and Coding Assessment. Amazon is the exception worth knowing: Interview101's Amazon guide says its MLE coding round uses LeetCode-style problems, so for an Amazon loop you should keep some DSA practice. For frontier labs, use LeetCode only to keep your array and hashing fluency warm, and spend most of your coding hours writing models and numerics from scratch.
What should a new grad practise for an ML coding interview?
Start with the named company examples: k-means, scaled dot-product attention, self-attention with cross-entropy, and gradient descent. They are single, well-defined implementations, and each can be finished and tested in an evening. Junior candidates still link the transformer debug and the agents question in their write-ups, so add those after the basics. Write the shape of every array in a comment, and test on a toy input small enough to check by hand. Senior-level depth on memory and trade-offs matters less at entry level than finishing a correct, tested solution within the time.
Where can I find ML coding interview questions with answers on GitHub?
Many widely shared answer lists are concept question-and-answer collections. One widely shared GitHub repository, for example, describes itself as a cheat sheet for machine learning interview questions and answers, with sections on system design and MLOps, probability and statistics, coding, maths, and behavioral and scenario-based questions. Lists like that suit the ML fundamentals round. For the coding round, you need practice that runs your code and shows where it fails, which a static list of answers cannot do.
Do I need PyTorch, or is NumPy enough?
Use NumPy for from-scratch questions and PyTorch for debugging and builds. The yuan-meng.com post on MLE interviews describes ML coding as debugging or building a PyTorch or NumPy model and then training it. NumPy is the better tool for from-scratch questions, because it forces you to write the softmax, the gradient and the masking yourself. PyTorch matters for debugging questions and builds, where you need to read a training loop, a module and a loss quickly and spot what is wrong.
How current is this data?
Recent. Of the 81 questions, 68 were last reported in the 12 months before 2026-10-11, and every named example in the company table carries its last-reported month. Treat the ranking as a snapshot: link counts shift as new write-ups are filed, and a question reported once can drop out of a company's loop. Check the last-reported month on each practice page before you commit a week to it.
Your practice plan
This plan assumes about an hour on weeknights and one longer block at the weekend, for three weeks. Write every implementation from a blank file, run it, and say each design choice out loud as if an interviewer were listening.
- Week one, weeknights: implement Implement Scaled Dot-Product Attention, then Hand-Code Self-Attention and Cross-Entropy, with a causal mask, a padding mask and a stable log-softmax.
- Week one, weekend: do Transformer Attention Mask and Heads Coding, then write the one-pass softmax entropy and a gradient-check harness you can reuse for the rest of the plan.
- Week two, weeknights: implement Hand-Written K-Means with vectorised distances, empty-cluster handling and a stopping rule, then a vectorised nearest-neighbour search and a two-layer forward pass under a forty-minute timer.
- Week two, weekend: work through Write and explain gradient descent pseudocode, then logistic regression from scratch with the gradient-check harness from week one, then backpropagation through a chain of matrix products.
- Week three, weeknights: run Debug a Transformer and Convert It to a Classifier twice, once alone and once with an assistant, and log every wrong suggestion. TrueInterview says coding answers are judged on hidden tests that show you the case you failed.
- Week three, weekend, by target: for Anthropic, do Agents / Coding with LLMs; for Uber, ML Coding from Scratch; for an OA at Coinbase or Pinterest, the Coinbase assessment or the Pinterest OA.
- The last two evenings: pick two implementations from earlier weeks and explain them start to finish without notes, covering stability, memory and how you verified correctness. If you are senior or staff, spend the second evening on one ML system design prompt instead.
Sources
- TrueInterview question bank — ML coding at big-tech companies — counted 2026-10-11
- Medium — checked 2026-10-11
- MLE Interview 2.0: Research Engineering and Scary Rounds | Yuan Meng — checked 2026-10-11
- Amazon's MLE Interview Separates ML Theory From Production Systems — And Most Candidates Prepare for the Wrong One — checked 2026-10-11
- How Coding Interviews Changed in the AI Era — checked 2026-10-11
- Meta's AI-Enabled Coding Interview: How to Prepare | Hello Interview — checked 2026-10-11
- Interview prep compared: your application, end to end · TrueInterview — checked 2026-10-11
- GitHub - amitshekhariitbhu/machine-learning-interview-questions: Your Cheat Sheet for Machine Learning Interview – Questions and Answers. · GitHub — checked 2026-10-11
Last reviewed: 2026-10-11.