Study
ML System Design Questions for Senior MLEs: 125 Reported (2026)
TrueInterview
October 11, 2026 · 20 min read

As of 2026-10-11, TrueInterview's bank holds 125 candidate-reported ML system design questions from 23 big-tech companies, and four problem families hold most of them: ads, recommendation and ranking (41), LLM, RAG and AI infrastructure (30), prediction and classification (16), and trust and safety (15). The round changes the mix: LLM and AI-infrastructure prompts lead the phone screen, while recommendation and ranking leads the onsite. Our take: stop preparing for one generic recommender round. Build a short list scoped by your next round, your target company and your level, and learn the one or two decisions each family's interviewer will push on.
Disclosure: TrueInterview is an interview-preparation product and publishes this article. Facts about other products come from their public pages on the dates listed under Sources.
The tables below cover ML design prompts only. If your loop also has a general distributed-systems round with feeds, chat or payments, use the companion guide to big tech system design questions by company and round for that half.
What ML system design questions do big tech companies ask?
They ask product problems grouped into a few design families, not model-architecture puzzles. Ranking and recommendation is the largest family, LLM and AI infrastructure is second and leads the screen, and trust and safety clusters at a handful of companies. The table gives each family's round split, the seniority of the write-ups behind it, and named prompts from the bank.

| Design family | All questions | Phone screen | Onsite | Senior-or-above vs junior-or-below write-up links | Reported prompts |
|---|---|---|---|---|---|
| Ads, recommendation and ranking | 41 | 12 | 30 | 70 vs 9 | Reels / Short Video Recommendation (Meta); Design A Nearby Restaurant Recommendation System (DoorDash, Uber) |
| LLM, RAG and AI infrastructure | 30 | 14 | 18 | 46 vs 3 | Design GPU Inference Serving System (Anthropic); Route and Batch Inference with Eight GPUs (Anthropic); AI-Enabled System Design — Agent System (Meta) |
| Prediction and classification | 16 | 5 | 10 | Not broken out | Listing Lifetime Value — Estimation (Airbnb); Mining Novel Data from Large Unlabeled Corpus (OpenAI) |
| Trust and safety (fraud, spam, moderation) | 15 | 4 | 11 | 16 vs 2 | Harmful / Weapon-Sales Content Detection (Meta); Design An Account Takeover Detection System (Stripe, Robinhood) |
| Search, autocomplete and crawling | 3 | Not in the top four | Not in the top four | Not broken out | Design App Store Search (Apple); Rank Home Search Results Without a Text Query (Airbnb) |
| Other listed families | Video and media 3, maps and delivery 2, feed, chat and job scheduling 1 each; 10 titles unmatched | Not in the top four | Not in the top four | Not broken out | ML Job Scheduler (Netflix, Apple, Snowflake) |
Read the table as a study order. The two largest families are also the two that lead a round, so one ranking design and one serving design cover the opening move of most loops. Trust and safety sits fourth overall but carries the second-strongest senior tilt after the LLM family, which makes it the best third topic for an experienced hire.
Our take: the systemdesign.academy guide says "A recommendation system is the most common machine learning system design prompt." In the bank, the nearest match, ads, recommendation and ranking, is indeed the largest single family, but it holds only about a third of the questions, and in the phone screen LLM and AI infrastructure outranks it. Prepare ranking as the base case, then spend real hours on serving and safety rather than treating them as optional extras.
Which questions are reported most often, and what does each one test?
Eight named prompts are each linked from five or more candidate write-ups, and Anthropic's GPU inference serving prompt leads with 22, all from one company. Five of the eight are recommendation prompts, two are LLM-infrastructure prompts and one is a safety classifier, which matches the family table: drill the ranking skeleton first, then serving.
| Question | Company | Round and last reported | Write-up links (senior-or-above) | Practice |
|---|---|---|---|---|
| Design GPU Inference Serving System | Anthropic | Phone screen, 2026-07 | 22 (22, of which 2 staff-plus or manager) | Practise GPU inference serving |
| Design A Nearby Restaurant Recommendation System | DoorDash, Uber | DoorDash onsite, 2026-07 | 7 (6, plus 1 junior-or-below) | Practise nearby restaurant recommendation |
| Design A Feed Recommendation System | Onsite, 2026-04 | 6 (6) | Practise feed recommendation | |
| Reels / Short Video Recommendation | Meta | Onsite, 2026-06 | 6 (6) | Practise Reels recommendation |
| Harmful / Weapon-Sales Content Detection | Meta | Not recorded in this snapshot | 6 (5, of which 2 staff-plus or manager, plus 1 junior-or-below) | In the bank, no practice page listed here |
| Nearby Place Recommendation | Meta | Not recorded in this snapshot | 5 (5) | In the bank, no practice page listed here |
| AI-Enabled System Design — Agent System | Meta | Not recorded in this snapshot | 5 (5) | In the bank, no practice page listed here |
| Short Video Recommendation & Ranking | Snapchat | Phone screen, 2026-04 | 5 | Practise short video ranking |
A link count is how many candidates reported a prompt, not how often a company asks it. The GPU serving prompt's lead comes from a single employer, so it is mandatory for an Anthropic loop and a strong second serving drill for OpenAI or Meta, not a universal favourite. Where the level is recorded, nearly every write-up behind these eight came from a candidate who reported senior level or above, which tells you the depth expected: an end-to-end design with trade-offs, not a model choice.
How do you answer recommendation and ranking prompts at senior level?
Make candidate generation plus ranking the one design you can draw from memory. Meta's Reels, Pinterest's feed and DoorDash's nearby restaurants all reuse the same two-stage skeleton with a different item, a different label and a different latency budget, so one well-rehearsed skeleton covers most of the largest family.
The dated variants show how far the skeleton travels. Google's Design a Product or Video Recommendation System was last reported as a phone screen in 2026-08. Uber's ML System Design (Recommendation / Feed Ranking / ETA) was reported onsite in 2026-03, and Amazon's Search / Ranking / Experimentation onsite in 2026-04. The systemdesign.academy guide describes "the two-stage candidate-generation-plus-ranking pattern that narrows millions of items down to a few hundred cheaply, then scores those precisely."
The decisions an interviewer will push on
Where the stages split. Candidate sources such as embedding retrieval, co-engagement, follow graphs and fresh content are tuned for recall under a tight latency budget. The ranker is tuned for the business objective on a short list, and a re-ranking pass handles diversity, freshness and policy. A senior answer says what each stage is allowed to miss and how you would measure retrieval recall separately from ranking quality, because a ranker cannot recover an item that retrieval dropped.
The label, its delay and the offline-online gap. Name the label (click, watch completion, order) and when it becomes final: a purchase or a completed delivery arrives long after the impression, and clicks carry position bias from the old ranker. Grokkingml's guide warns that "A model with strong offline metrics can still underperform online due to distribution shift or user behavior changes." So state the offline gate you would use, the online metric the A/B test decides on, and the guardrail metrics that block a launch.
For the nearby prompts, the location is a hard filter in candidate generation, and open status and delivery time are real-time features. Interview101's guide on Google's MLE loop says "Feature freshness determines whether your recommendation system shows users products they already bought yesterday." For Amazon, Interview101's guide on its MLE loop says candidates are asked to "explain how your system degrades when the feature store goes down". Have a fallback ranking ready: a popularity or recency list served from cache.
How do you answer LLM, RAG and AI-infrastructure prompts?
Treat them as serving and evaluation problems, not as a choice of model. This family leads the phone screen and is the most senior-weighted family in the bank, so a senior candidate with an Anthropic, OpenAI or Meta loop should put it in week one rather than saving it for the last weekend.
Design ChatGPT is filed under Amazon, Anthropic and OpenAI, and Google's ML System Design (Recsys / Chatbot / Image Classifier) mixes a chatbot into a classic prompt. Formation's 2026 blog post says "Generative AI system design is new and has moved from niche to standard in under 18 months." The same Formation post says OpenAI has run extended design rounds around prompts such as Design the OpenAI Playground.
GPU inference serving and batching (Anthropic)
The decisions here are batching and overload. Continuous batching raises throughput but stretches time to first token, so tie the batch policy to a stated latency target per request class. KV-cache memory decides how many sequences fit on one GPU, which drives routing, admission control and when you shed load. The devmatementors prep page asks what should degrade first when an inference service overloads during a burst, and the devmatementors answer is to "Define which requests are essential, bound queue time and concurrency, and reject excess work with an explicit retry contract." Anthropic's Route and Batch Inference with Eight GPUs is the same problem at a smaller scale, so rehearse both with one set of numbers.
LLM features, RAG and agents (Design ChatGPT, Agent System)
The decisions here are the eval set and the fallback. The eval set is a fixed collection of real prompts with graded answers that every model or prompt change must pass before release. The fallback is what the product does when the model is wrong, slow or down: a smaller model, a cached answer, a retrieval-only response or a hand-off to a person. For RAG, measure retrieval separately from generation, because a fluent answer built on the wrong passage still fails. For agents, bound the step count and make tool actions idempotent and permissioned. Formation's post says strong answers "build evaluation and guardrails into the design from the start."
When is trust and safety the main event?
At Meta it is the second cluster after ranking, and at Stripe, Robinhood, Databricks or DoorDash a harmful-content or account-takeover prompt can be the whole round. The family runs 16 senior-or-above write-up links against 2 junior-or-below, so it is a senior-weighted topic that most generic guides treat as a warm-up.
Four of Meta's 20 filed ML design questions are trust and safety, behind 10 for ranking and ahead of 3 for prediction. Design A Harmful Content Detection System is filed under Databricks, DoorDash and Pinterest, and Databricks' version was last reported onsite in 2026-04. Design An Account Takeover Detection System is filed under Stripe, Robinhood and Uber. Meta's Image Copyright Violation Detection sits in the same family. The yuan-meng.com post on MLE interviews lists "Trust and safety (e.g., harmful content detection)" among the other common topics.
The decisions an interviewer will push on
Labels and the operating point. Labels come from human review and user reports, so they arrive late and over-represent what was already flagged. Set thresholds from the cost of each error per policy, with tiered actions: remove, demote, or send to review. For account takeover, score at login in real time and make step-up verification the default action rather than a block.
Adversaries and feedback loops. Enforcement changes the data you train on, and attackers adapt to the model. Keep a random sample out of enforcement so you can still measure precision and recall, and pair the model with a fast rule layer for new attack patterns. Calibrd's senior MLE guide lists "Design the ML monitoring and observability for a fraud-detection system." as a sample prompt, so have alerts and an on-call runbook ready, not only a classifier.
What about prediction, ML infrastructure and the smaller families?
Prepare one prediction prompt and one ML-infrastructure component, then stop. These families are small, but the infrastructure prompts are where staff-level write-ups appear, so they carry level signal out of proportion to their size.
Prediction prompts in the bank include Airbnb's Listing Lifetime Value — Estimation and Anthropic's Estimate FFN Compute, Memory, and Sharding Communication. For a value or demand prediction, the decision is the label horizon: a lifetime value is not observed for months, so you need a proxy target, a time-based split and a plan for censored examples. Grokkingml's guide says to "Split data by time, not randomly, to mimic prediction-on-the-future."
Infrastructure prompts sit beside the product families. ML Job Scheduler is filed under Netflix, Apple and Snowflake. The yuan-meng.com post says "A handful of companies like Netflix, Snap, Reddit, Notion, and DoorDash have an ML infra system design round for MLE candidates". For a feature store, the decision is consistency between training and serving: point-in-time joins for training data and one feature definition for both paths. For a training scheduler, it is gang scheduling of GPUs, preemption and fairness between teams.
Should you study by round or by company first?
Study by round first. Onsite reports outnumber phone-screen reports 87 to 44, and a question can be reported in more than one round, but the family that leads each round differs (see the family table). If your next round is a phone screen, spend week one on LLM and AI-infrastructure prompts and week two on ranking; for an onsite, reverse the order.
Company comes second, as a weighting inside the round. For Meta, the 10: 4: 3 split of its filed questions means roughly half your design time on ranking, then safety, then one prediction prompt. For a company outside Meta, start from the named prompts in the company table below rather than its count.
How does the mix change by company?
Meta files the most ML design questions at 20, followed by OpenAI at 12 and Apple, Amazon and ByteDance at 10 each. Use your target's named prompts as the syllabus rather than its total, because a small row can reorder after a single new report.
| Company | Filed ML design questions | Named prompts in the bank | First design to build |
|---|---|---|---|
| Meta | 20 | Reels / Short Video Recommendation; Nearby Place Recommendation; Harmful / Weapon-Sales Content Detection; AI-Enabled System Design — Agent System | Short-video ranking, then a harmful-content classifier |
| OpenAI | 12 | Design ChatGPT; Mining Novel Data from Large Unlabeled Corpus | A chat-serving path with an eval gate and fallback |
| Apple | 10 | ML Job Scheduler; Design App Store Search | Search ranking, then a training job scheduler |
| Amazon | 10 | Search / Ranking / Experimentation (onsite, 2026-04); Design ChatGPT | Search ranking with an experiment plan and a cost estimate |
| ByteDance | 10 | No named prompt in this snapshot | Follow the round order in the family table |
| Microsoft | 9 | Design a Product Search System | Product search ranking |
| DoorDash | 8 | Design A Nearby Restaurant Recommendation System; Design A Harmful Content Detection System; Design A Personalized Search Ranking System | Nearby recommendation, then harmful content |
| 8 | Design a Product or Video Recommendation System (phone screen, 2026-08); ML System Design (Recsys / Chatbot / Image Classifier) | Two-stage recommendation, then a chatbot variant | |
| 7 | Design A Feed Recommendation System; Design A Harmful Content Detection System | Feed ranking, then harmful content | |
| Uber | 7 | ML System Design (Recommendation / Feed Ranking / ETA); Design An Account Takeover Detection System; Design A Personalized Search Ranking System | Feed ranking, then account takeover |
| Anthropic | Outside the top ten by count | Design GPU Inference Serving System; Route and Batch Inference with Eight GPUs; Estimate FFN Compute, Memory, and Sharding Communication; Design ChatGPT | GPU serving, then batch routing |
Read across a row, never down a column: the rows differ in size, and a row of seven tells you which prompts exist, not their weights. LinkedIn sits outside this table, but its LinkedIn Skills — Data Mining & ML System Design was last reported onsite in 2026-02, so a LinkedIn candidate should prepare an extraction and matching pipeline.
Worked example: a Meta onsite, then an Anthropic phone screen
Take a senior engineer with a Meta onsite in three weeks and an Anthropic phone screen the week after. Meta's mix makes week one a ranking week: Reels / Short Video Recommendation, last reported onsite in 2026-06 with all 6 links from senior-or-above write-ups, and Nearby Place Recommendation. Week two goes to Meta's safety cluster, Harmful / Weapon-Sales Content Detection, plus the Agent System prompt. Week three flips to Design GPU Inference Serving System, the phone-screen prompt behind the most write-up links in the bank, with the batching and overload decisions rehearsed aloud.
Which questions should you practise for several loops at once?
Start with Design A Personalized Search Ranking System, filed under five companies: Instacart, Airbnb, DoorDash, Robinhood and Uber. Then take one from the three-company overlaps, Design ChatGPT, Design A Harmful Content Detection System, Design An Account Takeover Detection System or ML Job Scheduler, matched to your target.
One prepared answer for each covers several loops, which is better value than memorising company trivia. For personalized search, the decision an interviewer pushes on is the blend of query relevance and personal history, and how you debias click labels collected under the old ranking. For the three-company overlaps, reuse the serving, safety and scheduling decisions above: an eval gate and fallback for Design ChatGPT, thresholds and adversarial drift for the two detection prompts, and fairness and preemption for the scheduler.
What changes at senior and staff level?
The write-ups behind these questions come overwhelmingly from experienced candidates: 149 links from candidates who reported senior, staff-plus or manager level against 16 from junior, new-grad or intern candidates. The levels are self-reported, but the tilt is wide enough to set your target depth: production trade-offs, not a model tour.
Staff-plus and manager write-ups link Design GPU Inference Serving System (2), LinkedIn Skills — Data Mining & ML System Design (2), Harmful / Weapon-Sales Content Detection (2), Design A Harmful Content Detection System (1), Distributed Training Data Pipeline (FAR) (1) and Design Linkedin Learning Recommendation System (1). Those are infrastructure and platform components more than model choices. Calibrd's senior MLE guide says "Senior MLE interviews are calibrated against production ownership, not just model quality." The same Calibrd guide puts it more briefly: "The senior signal is the trade-off, not the metric."
Our take: generic guides tell senior candidates to go deeper on models. The bank points the other way. Prepare one component you have owned end to end, such as a feature pipeline, a serving path or a labeling loop, and be ready to defend what you traded away and why. Grokkingml's guide describes the senior opening move as "Clarifies the problem and constraints first": the user, the scale, the latency budget and what counts as success, before any model is named.
Where experienced candidates get down-leveled
In our view, three rounds carry the level signal. The ML design round decides whether you drive the trade-offs or wait to be asked; Interview101's guide on MLE loops says evaluators separate candidates who raise training-serving skew, feature staleness or label feedback loops only when prompted from those who build them in from the start. The project deep dive is the second: Calibrd's guide says at senior level "the deep-dive round becomes a 60-minute walk-through of an ML platform component you've owned for 6+ months." Behavioral scope is the third, where a team-sized story reads as senior and an org-sized one as staff. If you are targeting staff, rehearse the deep dive as hard as the design round.
How current is this list, and is a PDF or GitHub list enough?
It is recent: 105 of the 125 questions were last reported within the 12 months before 2026-10-11. A static PDF or GitHub list cannot tell you that, which is why it works for the answer framework and fails for choosing what to practise.
Free guides age quickly. The patrickhalina.com ML systems design guide, written by an author who reached Staff ML Engineer at Pinterest, shows "Last Updated: Jan 17, 2021". Its advice still holds: the author says "it's very common for consumer big tech companies to ask questions about recommender systems in their system design". What it cannot show is the shift toward serving and LLM prompts that the family table records. TrueInterview says its list is updated as candidates report new questions, that recent and often-reported questions come first, and that reports are dated when the report gives a month. Use a framework for the spine of the answer and a dated, round-labelled inventory to pick the prompts.
How we counted
These are candidate-reported questions reconstructed for practice, not any company's official list. Each question's family is assigned from its title, and the family table shows the ten most frequent families plus the 10 titles that match no family, so the few questions in rarer families are left out and the rows sum to slightly less than the total. Round counts overlap because one question can be reported in more than one round. TrueInterview says an engineer who has worked at this kind of company reviews each question before it is filed. Write-up link counts show how many candidates reported a prompt, and one company can supply all of them, as with the 22 GPU serving links from Anthropic. Levels are what candidates reported about themselves.
FAQ
Are these the official ML system design questions big tech companies ask?
No. They are reconstructions of what candidates reported after their interviews, filed by company, round and month, and no company publishes or endorses them. Treat a prompt as evidence that at least one candidate met it in a named round. The practical use is selection: if a prompt is filed under your target and was last reported within the past year, it earns a timed rehearsal, and if it is only filed elsewhere, the shared design decisions still transfer.
Is a machine learning system design interview PDF or book enough?
A book or PDF gives you the answer framework, which you need. Grokkingml's guide frames it as "One repeatable framework — frame, metric, data, features, model, serve, monitor". The harder skill is pacing. The mockingly.ai guide says "The most common mistake is spending 20 minutes on model architecture and rushing through monitoring in the last 2 minutes." Rehearse each framework step against a clock, and give data, labels and monitoring at least as much time as the model.
What do interviewers expect beyond a standard system design answer?
They expect the ML failure modes inside the architecture. Interview101's guide quotes no-hire feedback that reads "ML knowledge was strong, but system design lacked ML-awareness." The mockingly.ai guide says training-serving skew is one of the most common silent failures of production ML, and that interviewers at Meta and Google probe for it. Draw the data and label path first, then show where features are computed for training and for serving.
Which roles get an ML system design round?
Most applied ML roles do, and the round is spreading beyond them. The patrickhalina.com guide says "If you're applying to be a Data Scientist, ML Engineer or ML Manager at a big tech company, you'll probably face an ML Systems design question." The systemdesign.academy guide says the question appears for machine learning engineers, applied scientists and ML platform roles, and increasingly for senior and staff backend roles that touch ranking or personalization. If your title is backend but your team ships a ranker, prepare the ranking skeleton above.
How long should ML system design prep take while working full time?
Plan design prep around your loop's length rather than a fixed number of prompts. Calibrd's senior MLE guide says "FAANG-level Senior MLE loops typically run 5–7 rounds over 5–7 weeks." Within that window, three focused weeks of evening and weekend rehearsal fit the plan below: one family per week, two or three prompts each, and a final pass on the company you are interviewing with first. Use your current job's systems as material for the deep dive.
Your practice plan
- On day one, write down your target company and your next round, then pick the lead family from the family table: LLM and AI infrastructure for a phone screen, recommendation and ranking for an onsite.
- Week one, ranking: draw the two-stage skeleton on Reels / Short Video Recommendation and Design A Feed Recommendation System. For each, say aloud where the stages split, what the label is, how late it arrives and how an offline win becomes an A/B decision.
- Week one or two, serving: rehearse Design GPU Inference Serving System with a batching policy, a KV-cache budget and an overload plan, then sketch one LLM feature with its eval set and its fallback.
- Week two, safety: work through Design A Harmful Content Detection System, naming the label source, the threshold per policy, the review queue and the holdout you keep to measure drift.
- Week two, overlap: rehearse Design A Nearby Restaurant Recommendation System with location as a hard filter, and sketch personalized search ranking with debiased click labels.
- Week three, company: run your target's prompts from the company table, such as Google's Design a Product or Video Recommendation System, Amazon's Search / Ranking / Experimentation, Uber's ML System Design (Recommendation / Feed Ranking / ETA), Snapchat's Short Video Recommendation & Ranking or LinkedIn's LinkedIn Skills extraction.
- Week three, level: if you are targeting senior or staff, prepare the deep dive on one component you owned, with the trade-off you made, the alternative you rejected and the metric that told you it worked.
- Before each rehearsal, check the prompt's last-reported month and the levels of the write-ups behind it, and drop prompts that only junior candidates reported if you are interviewing at senior level.
Sources
- TrueInterview question bank — ML system design at big-tech companies — counted 2026-10-11
- Design a Recommendation System (ML Design Walkthrough) — checked 2026-10-11
- ML System Design — Grokking the Machine Learning Interview — checked 2026-10-11
- Google MLE Interviews Test System Design Before Model Tuning — Here's What That Means — checked 2026-10-11
- Amazon's MLE Interview Separates ML Theory From Production Systems — And Most Candidates Prepare for the Wrong One — checked 2026-10-11
- 3 Ways AI has Changed FAANG Interviews · Formation Blog — checked 2026-10-11
- Dell Machine Learning Engineer Technical Interview Prep — checked 2026-10-11
- Preparing for ML Infra System Design Interviews | Yuan Meng — checked 2026-10-11
- Senior Machine Learning Engineer Interview Prep — Calibrd — checked 2026-10-11
- MLE Interviews Are Not SWE Interviews With ML Questions Appended — Here Is What Evaluators Are Actually Measuring — checked 2026-10-11
- ML Systems Design Interview Guide · Patrick Halina — checked 2026-10-11
- Real FAANG Interview Questions by Company · TrueInterview — checked 2026-10-11
- Interview prep compared: your application, end to end · TrueInterview — checked 2026-10-11
- Machine Learning System Design Interview: Complete Prep Guide — checked 2026-10-11
Last reviewed: 2026-10-11.