Study
Big Tech System Design Questions by Company and Round (2026)
TrueInterview
October 11, 2026 · 12 min read

Disclosure: TrueInterview is an interview-preparation product and publishes this article. Facts about other products come from their public pages on the dates listed under Sources.
A few classic prompts do travel across big-tech companies, but the problem families each company leans on differ sharply. TrueInterview's bank holds 475 candidate-reported system design questions from 25 big-tech companies. Meta's 75 filed questions lead with ads, recommendation and ranking ML (14) and A/B testing (9), Amazon's 50 with metrics, logging and data pipelines (7), and OpenAI's 44 with LLM, RAG and AI infrastructure (9). Build your plan from your target company's top two families, and treat URL shorteners and chat apps as the backstop rather than the syllabus.
What is in the bank, and how was it filed?
As of 2026-10-11, 399 of the 475 questions were last reported within the previous 12 months. They are candidate-reported prompts reconstructed for practice, not any company's official list. TrueInterview says each one is checked against other reports of the same round and stamped with the month it was asked wherever the report gave one.

For prep, a count tells you where candidates' write-ups cluster. That is the closest public signal of what a loop looks like, but it is not a probability that you will get a given prompt.
How we counted
Each question's problem family is assigned from its title, and 41 titles match no family. Round counts overlap because a question can be reported in more than one round, which is why the onsite (336), phone-screen (162) and OA (8) tallies add up to more than 475. Coverage is uneven: Meta has 75 filed questions and Databricks 25, and the counts measure what candidates reported and chose to write up, not what companies ask. So compare families inside one company's row, never across rows, because rows of such different sizes are not on the same scale.
Which design families dominate, and what should you drill first?
Ads, recommendation and ranking ML leads the bank at 59 of 475, with LLM, RAG and AI infrastructure second at 42, so the two ML-heavy families together account for roughly one in five of the 475 questions.
The rest of the top ten runs metrics/logging/data pipelines (32), A/B testing (30), booking/inventory/ticketing and maps/ride-hailing/delivery (28 each), payments (23), prediction and classification ML (19), chat and messaging (18), and storage/caching (17).
The classics count differently: each is a single prompt filed under several companies, with Design TinyURL under 6, Design Notification System under 8 and Design News Feed under 9. If your next free hour is a choice between a fourth URL-shortener variant and your first ranking or model-serving design, the ranking or serving design covers more of what candidates report.
Ads, recommendation and ranking ML
Reported titles include Netflix's Design an Ads Frequency Cap System and Design the Data Model for an Ads Demand Platform, DoorDash and Uber's Design A Nearby Restaurant Recommendation System, and Meta's Reels / Short Video Recommendation. Whatever the variant, be ready to defend how candidates are generated and fanned out, how fresh the features are, and how a cap is enforced without reading every impression.
LLM, RAG and AI infrastructure
Anthropic's Design GPU Inference Serving System, Databricks and Anthropic's Design GPU Scheduling Platform, and Anthropic's Design AI Prompt Playground sit here. Treat these as capacity and scheduling problems: when to batch versus serve in real time, and what degrades first when the GPU queue backs up.
Metrics, logging and data pipelines
The reported titles include Amazon and DoorDash's Design Metrics System, LinkedIn's Metrics & Monitoring Platform, Oracle's VM Health Monitoring at 1M Scale and DoorDash's alert notification system. Cardinality and cost decide these designs, so be ready to say how many series a new label creates and what each one costs to keep.
A/B testing and experiment design
Uber's Design a switchback and choose block length and Design Pricing Model Experiment, Airbnb's Design an A/B Testing Platform, and LinkedIn's Scale a Distributed Randomized Multiset are the reported examples. Rehearse how interference between units and the choice of block length change the analysis, and how you would size the sample.
Booking, inventory and ticketing
Airbnb and Amazon's Design Hotel Booking System, Pinterest's Design Inventory Management System, Databricks' Design Book Price Aggregator and Salesforce's Coffee Ordering System Design sit here. Defend reservation semantics under concurrency: overselling, holds, expiry and idempotent confirmation.
How does the question mix differ by company?
Each company's leading families are different, so build your syllabus from your target's row rather than from a cross-company average. Filed counts run Meta 75, Amazon 50, OpenAI 44, Google 39, ByteDance 37, Uber 34, DoorDash 31, Microsoft 30, LinkedIn 29 and Databricks 25.
| Company | Filed questions | Leading families | First design to build |
|---|---|---|---|
| Meta | 75 | ads/rec/ranking ML 14, A/B testing 9, news feed/social graph 5 | A short-video or ads ranking design, then an experiment platform |
| Amazon | 50 | metrics/logging/data pipelines 7, maps/ride-hailing/delivery 6, LLM/RAG/AI infrastructure 5 | A metrics pipeline with cardinality and retention costs worked out |
| OpenAI | 44 | LLM/RAG/AI infrastructure 9, payments 5, chat/messaging 4 | A model-serving path, then one payments flow |
| 39 | news feed/social graph 4, chat/messaging 3, A/B testing 3 | A feed, then a chat system | |
| ByteDance | 37 | ads/rec/ranking ML 6, video/media/uploads 4, metrics/logging/data pipelines 4 | A short-video ranking design, then an upload pipeline |
| Uber | 34 | maps/ride-hailing/delivery 7, ads/rec/ranking ML 5, A/B testing 5 | A dispatch or nearby-search design, then a switchback experiment |
Read the table as a syllabus: your target's top two families are your first two designs. Google's row is the flattest, with no family above a handful of questions, so a Google candidate gets more from breadth across feed, chat and experiments than from three passes at one family.
If your target is not in the table, start from the prompts filed under it. Databricks and Anthropic are not in the table, but both are among the three companies, with OpenAI, behind Design GPU Scheduling Platform's 22 write-ups. DoorDash and LinkedIn both have metrics questions in the bank, and both have Design Job Scheduler filed under them. Netflix's two reported ads prompts, the frequency cap and the ads demand-platform data model, sit in the ranking family. Microsoft is one of three companies behind Design Hotel Booking System's 13 write-ups and one of six under which Design TinyURL is filed.
Do phone-screen and onsite questions ask different things?
Yes: the leading families change by round. A/B testing and experiment design leads the phone screen at 24 of 162, with LLM/RAG/AI infrastructure and ads/rec/ranking ML at 16 each. Onsite, ads/rec/ranking ML leads at 45 of 336, followed by LLM/RAG/AI infrastructure (28), booking/inventory/ticketing (24) and metrics pipelines (23), so experiment design drops out of the top four and booking enters it.
For a phone screen, rehearse one narrowly scoped prompt (an experiment, a metrics pipeline or a serving path) and finish the deep dive aloud in one sitting. For the onsite, rehearse one full-scale build from ranking, model serving or booking.
Round labels are not difficulty labels. Anthropic's Design GPU Inference Serving System and Amazon's Design News Feed were both reported as phone screens, each last reported in 2026-07, while Netflix's Ads Frequency Cap System (2026-08) and Databricks' GPU Scheduling Platform (2026-07) were reported onsite. If a serving prompt is plausible for your screen, prepare it to onsite depth.
DesignGurus' guide to Google's process says each round lasts 45 minutes and centres on a single open-ended problem such as Design YouTube or Design Google Maps. The same DesignGurus guide says the interviewer probes 1–2 specific areas, such as sharding strategy, cache eviction logic, consistency trade-offs or fault tolerance. Pick your deep-dive areas before the interviewer does: for a feed, the fan-out path; for a booking system, the hold-and-expire logic.
Is it still mostly distributed systems?
Distributed systems is still the largest subtype tag at 256 questions, but 125 are tagged ML system design and 46 experiment design (data science), with low-level design at 26 and frontend system design at 10. Decide which track your role sits on before choosing material: a candidate for an applied-ML or ML-platform team who rehearses only caching and sharding patterns is preparing for a different test.
A Substack post on senior deep dives refers to "the AI infrastructure question that is now standard at most companies". A Deep Engineering piece on design hiring says one interviewer it quotes opens with a single test: "Does the candidate treat an AI component as an unreliable dependency or a magic box?". The same Deep Engineering piece says weak candidates "draw a box labeled LLM and move on", while strong ones ask what happens when it is wrong, what the fallback is and how they will know it is drifting. Before any AI-infrastructure loop, rehearse one serving path end to end with a named fallback, a drift signal and an evaluation set that must pass before shipping.
Which individual questions recur across companies and write-ups?
Design Job Scheduler travels furthest: it is filed under 12 companies and linked from 21 candidate write-ups across 6 of them.
By write-up count the order changes, with Design Payment System at 27 write-ups from OpenAI and ByteDance candidates, Design GPU Inference Serving System at 22, all from Anthropic, and Design Slack-like Chat System at 19 across Airbnb, OpenAI and Databricks. Company count is a breadth signal and write-up count a depth signal; neither is a rate. Job Scheduler is the best single transfer prompt because it is both wide and well documented, while the inference-serving prompt is essential for an Anthropic loop and optional elsewhere.
The other heavily linked payments prompt, OpenAI's Payment / Coffee-Shop Ordering (read the prompt!), has 16 write-ups, and its title warns that the visible framing may not be the actual ask. Cost is the other trap. One candidate's Stackademic write-up recalls an Amazon interviewer calling cost "a first-class constraint" and pricing the candidate's design at $400K a year against an $80K alternative. Bring rough arithmetic to any payments or serving prompt: requests per second, storage per day and what that storage costs per month.
Worked example: a Meta loop, family by family
Meta's filed questions lead with ads/rec/ranking ML (14), A/B testing (9) and news feed/social graph (5), so three families cover its three most-reported clusters. Week one: a ranking design; practise Design an Ads Frequency Cap System. Week two: an experiment platform plus a switchback with block-length reasoning, the shape of Uber's reported switchback and pricing-experiment prompts. Week three: a feed; practise Design News Feed, filed under 9 companies including Meta.
The opening minutes matter as much as the syllabus. The System Design Interview Roadmap newsletter (systemdrd.com) says candidates who score well choose what they are designing and why in the first 3–5 minutes. For the frequency cap, that means asking whether the cap is per user per campaign or across campaigns, whether it may overshoot slightly or must be exact, and what latency budget the ad server allows for the check. Then name the trade-off you will spend time on, an exact counter on the serving path against an approximate one updated asynchronously, and go deep there: where the counter lives, and whether the ad serves or is suppressed when the counter store is down. A Hashnode post by an engineer who sat on Meta and Microsoft hiring committees describes strong candidates saying "the interesting part here is the ranking, let me spend my time there.".
FAQ
Are these official company question lists?
No. They are candidate-reported questions reconstructed for practice, not any company's official question list. Treat each entry as evidence of what one or more candidates were asked in a specific round, and design against it as a practice prompt. If a company's row is small, read its individual titles rather than trusting the family ranking, because a single extra write-up can reorder a family that only has a few questions.
Which questions are filed under the most companies?
Design Job Scheduler is filed under 12 companies, Design News Feed under 9 and Design Notification System under 8. Being filed under many companies means many different loops produced a report of the prompt, not that every company asks it at the same rate. These are the best prompts for a final timed run when your target's own families are covered, because the design skills they test carry across loops.
Do I need to design production-grade systems, or is drawing boxes enough?
Boxes are the starting point. Google's SRE workbook says designers must turn a whiteboard design into concrete resource estimates at multiple steps, and that in these exercises sound reasoning and assumption making matter more than the final values. In practice, state your traffic and storage assumptions out loud, turn them into machine and cost estimates, and say which component fails first when one assumption doubles.
How deep should I go for my level?
The System Design Interview Roadmap newsletter (systemdrd.com) says you need the right depth for your level, not the maximum possible depth. DesignGurus' guide to Google's process says L5 candidates get one mandatory system design round while L6 and above face two to three of increasing complexity. If you are interviewing at staff level, budget a second full rehearsal per family and practise the failure and operations half of each answer, not only the data flow.
How should experienced candidates prioritise the final week?
Pull your target's top two families from the per-company table, then rehearse one prompt from the family that leads your booked round: A/B testing and experiment design for a phone screen (24 of 162 reports) or ads, recommendation and ranking ML for an onsite (45 of 336). Skip families you already design at work, and spend the saved time on timed run-throughs with a cost estimate.
Your practice plan
- Write your target company's top two families on one card from the per-company table; if its row is small, such as Databricks at 25 filed questions, list its individual prompts instead.
- If your target is Meta, ByteDance or Netflix, start with a ranking design such as Design an Ads Frequency Cap System, then an experiment prompt such as Uber's switchback with block-length choice.
- If your target is OpenAI, Anthropic or Databricks, start with Design GPU Inference Serving System, then Design GPU Scheduling Platform.
- If your target is Amazon, start with metrics, logging and data pipelines; DoorDash and LinkedIn also have metrics questions in the bank. Work out series count and retention cost for every design.
- If your target is Uber, start with maps, ride-hailing and delivery, then a switchback or pricing experiment.
- Whatever your target, do one timed run of Design Job Scheduler, linked from 21 write-ups across six companies, and one payments pass with Design Payment System and Payment / Coffee-Shop Ordering.
- Prefer prompts with a last-reported month inside the past year, as 399 of the 475 have; where no month was recorded you cannot tell when it was asked, so treat that as missing evidence rather than a reason to skip.
- Open every rehearsal with scope questions before you draw, then say aloud the one trade-off you will go deep on.
If you want the dated, round-labelled versions of these prompts in one place, browse TrueInterview's question bank.
Sources
- TrueInterview question bank — system design at big-tech companies — counted 2026-10-11
- Interview prep compared: your application, end to end · TrueInterview — checked 2026-10-11
- Google System Design Interview: What Changed, What They Ask, and How to Pass — checked 2026-10-11
- The Deep-Dive Tape: How I Filled 22 Minutes of a Senior System Design Round With the Decisions That Got the Offer — checked 2026-10-11
- System Design Hiring Is Really a Judgment Test — checked 2026-10-11
- Medium — checked 2026-10-11
- What Interviewers Actually Score in System Design Rounds — checked 2026-10-11
- What 1,000+ System Design Interviews Taught Me — checked 2026-10-11
- Google SRE - System Design: Non-Abstract Large System Design — checked 2026-10-11
Last reviewed: 2026-10-11.