Study
Experiment Design and A/B Test Questions: 46 Reported (2026)
TrueInterview
October 11, 2026 · 19 min read

As of 2026-10-11, TrueInterview's bank holds 46 candidate-reported experiment design questions from 14 big-tech companies; 28 sit in the A/B testing and experiment design family, and 32 were reported from phone screens. The round is a product-judgment screen rather than a statistics quiz. Our take: be able to say, in ten minutes and without a whiteboard, which decision the test serves, which unit you would randomize, which guardrail could veto the launch, and when you would not run the test at all. Sample-size formulas and test selection come after that.
Disclosure: TrueInterview is an interview-preparation product and publishes this article. Facts about other products come from their public pages on the dates listed under Sources.
Which experiment design questions are candidates reporting?
The titles below show the range. Most arrive in a phone screen and read as a product decision with an experiment attached, such as evaluating a detector, pricing a promotion or adding an ad format. Very few ask for a formula, so the right-hand column, our reading of what each title tests, matters more for your preparation than the company column.

| Question | Company | Round | Last reported | What the title tests (our reading) |
|---|---|---|---|---|
| Design a switchback and choose block length | Uber | Phone screen | 2026-01 | Randomization unit when riders share drivers |
| How would you evaluate stolen-post detection? | Meta | Phone screen | 2026-03 | Evaluating a classifier-driven integrity launch |
| Design a free-month experiment | OpenAI | Phone screen | 2026-02 | An outcome that lands after the promotion ends |
| Design experiments for marketplace product changes | DoorDash | Onsite | 2026-02 | Two-sided interference and choice of design |
| How would you evaluate adding video ads? | Amazon | Phone screen | 2025-11 | Revenue against user-experience guardrails |
| Plan DS approach for biker delivery project | ByteDance | Phone screen | 2025-11 | Scoping the question before any test |
| Choose a precise A/B test primary metric | Phone screen | Not recorded | Defining one decision metric | |
| Decide if ad load is optimized | Onsite | Not recorded | Dose-response and long-run effects | |
| Design a network-aware Wi‑Fi badge experiment | Airbnb | Phone screen | Not recorded | Spillover between treated and control listings |
| Design a pricing experiment with network effects | Instacart | Phone screen | Not recorded | Market-level randomization |
Read the table by the decision each title forces. Uber's switchback, Instacart's pricing and Airbnb's Wi‑Fi badge titles are one problem in three costumes: a user-level split leaks, so the unit of randomization has to move. If you prepare that one move well, you have covered a third of the list.
Design a switchback and choose block length (Uber)
In a ride-hailing market, treated and control riders draw on the same drivers, so a rider-level split contaminates the comparison. A switchback randomizes the whole market by time block instead. The decision the interviewer pushes on is block length: short blocks give you many units and tight intervals, but effects carry over into the next block; long blocks contain carryover but leave you with few units and wide intervals. Choose the length from how long the system takes to settle after a switch, discard a burn-in window at each boundary, and analyse at the block level, because the block is now your unit. Uber's engineering blog makes the general point: "the types of units our users need to randomize on can vary according to use case".
How would you evaluate stolen-post detection? (Meta)
A weak answer designs an A/B test in the first minute. A stronger one asks what the evaluation decides: shipping the detector, tuning its threshold, or changing the reporting flow. Offline precision and recall on a labelled set come first, because they tell you whether an online test is even safe to run. If you do test online, name the unit (post, reporter or viewer) and note that viewers who see both variants break a clean user-level split. Then set the primary decision metric, add guardrails such as wrongful takedowns, appeals and support tickets, and say when you would not run it: if the control arm means knowingly leaving stolen content up, a staged rollout with monitoring is the better tool.
Design a pricing experiment with network effects (Instacart)
A price shown to one customer changes demand, and that demand draws on shopper supply that control customers also use. The deciding question is the unit: randomize by market or region, or by time, and accept that fewer units means less power. Recover some of that power with pre-period covariates, and pick the decision metric before launch: revenue per order, order volume and retention can point in different directions, and the interviewer wants to hear which one decides. Guardrails here include cancellations and shopper earnings.
Design a free-month experiment (OpenAI)
The outcome you care about arrives after the free month ends. Randomize eligible users at assignment and analyse by intent to treat, so that users who ignore the offer still count. The decision metric is net revenue per eligible user over a window that runs past the promotion, because the offer also goes to people who would have paid anyway. Ask how long the business can wait for that answer, since an early read on sign-ups will flatter the offer.
How would you evaluate adding video ads? (Amazon)
The bank files this title in the ads, recommendation and ranking ML family, next to Pinterest's Decide if ad load is optimized, ByteDance's Decide launch of downranking suspected bad sellers and Meta's Design a feed ads A/B test with guardrails. The tension is short-run ad revenue against slower losses in engagement and purchase conversion. Treat ad load as a dose with more than one level, keep a long-running holdout to see the slow effects, and name the guardrail that can veto revenue, such as conversion per session or ad-hide rate.
Where do these questions show up: phone screen or onsite?
Mostly in the phone screen. Reported rounds for the 46 questions split into 32 phone-screen and 14 onsite reports, and a single question can be reported in more than one round. Of the 32 phone-screen questions, 23 are in the A/B testing and experiment design family. Build the ten-minute spoken answer first and the hour-long version second.
| Round | Reports in our bank | What to prepare (our view) |
|---|---|---|
| Phone screen | 32, of which 23 in the A/B family | A ten-minute spoken skeleton: decision, unit, metric, guardrails, when not to run, decision rule |
| Onsite | 14 | The same skeleton, defended for an hour: interference, power, trust checks, ramp plan |
Treat the split as a reporting pattern, not a company rule. The practical consequence is pacing. A phone screen leaves no time for a sample-size derivation, so the candidate who states the decision and the unit in the first two minutes is the one who gets to the follow-ups.
Recency is the other limit. Only 10 of the 46 questions were last reported within the 12 months before 2026-10-11. That tail is too thin to memorise as a list, which is another reason to practise the shape of the answer rather than individual prompts.
Our take: most prep lists treat experiment design as an onsite deep dive and spend their pages on test statistics. The reports point the other way. If your loop opens with a phone screen at Meta, Uber, ByteDance, Google or Amazon, spend your first two weeks on the unit of randomization, the decision metric with guardrails, and when not to run the test.
Which companies concentrate the questions?
Meta, by a wide margin. Meta has 16 of the 46 reported questions, and 9 of those 16 are in the A/B testing and experiment design family. ByteDance and Uber follow, then Google and Amazon. If Meta is your target, experiment design deserves your largest single block of evenings; elsewhere the volume falls quickly.
| Company | Questions in our bank | Example title | Where to put extra hours (our view) |
|---|---|---|---|
| Meta | 16 | Design a feed ads A/B test with guardrails | Primary metric, guardrails and a launch call |
| ByteDance | 5 | Plan DS approach for biker delivery project | Scoping, then launch decisions in ranking |
| Uber | 5 | Design Pricing Model Experiment | Switchbacks and marketplace interference |
| 4 | Choose a precise A/B test primary metric | Metric precision and causal edge cases | |
| Amazon | 4 | How would you evaluate adding video ads? | Ads revenue against conversion guardrails |
| Airbnb | 3 | Design a network-aware Wi‑Fi badge experiment | Spillover between listings |
| DoorDash | 2 | Design experiments for marketplace product changes | Two-sided marketplace designs |
| OpenAI, Pinterest, Instacart | 1 each | Design a free-month experiment | One title each: practise it, then return to shapes |
The ten named companies account for 42 of the 46 questions. Read the gap between Meta and the rest as what candidates reported, not as how often each company asks. DataLemur's guide says A/B testing questions appear in about half of data science interviews, especially for product data science roles at consumer-tech companies such as Meta, Airbnb and Uber. That is an estimate of how often the topic comes up across interviews, while our counts are questions collected per company, so the two measure different things and neither forecasts your loop.
What are interviewers grading in an experiment design answer?
A decision. PrepVector's webinar write-up says interviewers at Google, Meta and Microsoft separate reporting results from guiding a decision, and interview101's Meta guide says candidates lose points by jumping to test selection before establishing the causal question. Your answer should end in a ship, ramp or stop call with the threshold that triggers it.
The skeleton below fits every title in the first table. The left column is the order to say it in.
| Step | What you say | Why it is graded (source or standard practice) |
|---|---|---|
| Decision | Ship or no-ship, partial rollout, timing or revenue limits | PrepVector's guide lists asking what decision the test supports as a strong signal |
| Unit | User, session, account, household, market or time block, and why | Interference and shared accounts change the variance and the bias |
| Primary metric | One decision metric that is hard to game | PrepVector's guide gives purchases per visitor as an example |
| Guardrails | Metrics that must not get worse | PrepVector's guide lists AOV, refunds, latency and retention |
| Power and duration | Minimum detectable effect, traffic, whole weeks | Sample size grows with variance and with the inverse square of the effect you need to detect |
| Trust checks | Sample ratio, logging, novelty | A broken split invalidates every metric downstream |
| Decision rule | What result ships, what result stops, what would change your mind | Interviewers want uncertainty turned into a recommendation |
Two weak signals recur across the guides. PrepVector's write-up lists "Run it for two weeks", with clean results assumed and a stop at significance, as a weak answer, and interview101's Meta guide says that "the p-value is below 0.05 so ship it" fails because the causal claim was never interrogated. The same PrepVector write-up says these interviews are "actually a judgment problem and structure is how you demonstrate that judgment under pressure".
Our take: interview101's Meta guide reports that candidates over-index on statistical sophistication and under-prepare the product reasoning that separates hire from no-hire. For an experienced data scientist the risk is sharper, because fluent statistics can fill the whole slot. Spend the first two minutes on the decision and the unit, and let the statistics arrive as support.
How much statistics do you still need?
Enough to defend the design in plain language. An account on sirjohnnymai.com lists the red flags as leaving out confidence intervals, ignoring statistical power and skipping multiple-testing adjustments. Treat the concepts below as follow-up material: know each one well enough to explain when it changes your design, not to derive it.
| Concept | When it comes up | What to say (standard practice) |
|---|---|---|
| Power and minimum detectable effect | Any follow-up on test length | Fix the smallest effect worth shipping, estimate variance from history, and size the test; halving the detectable effect roughly quadruples the sample |
| Duration | Weekly cycles and slow metrics | Run whole weeks, and run past novelty; if the metric moves slowly, use a leading indicator you have validated, or wait |
| Peeking and sequential testing | A request to stop early | Repeated looks at a fixed-horizon test inflate false positives; stop early only under a sequential design planned in advance |
| Sample ratio mismatch | Results look too good, or traffic looks off | Compare assignment counts with the designed split using a chi-squared test; a mismatch points to a bug, so read no metric until it is fixed |
| Variance reduction (CUPED) | Small effects, limited traffic | Adjust each user's metric by their pre-period value of the same metric; the variance falls by the share the covariate explains, with no bias, because the pre-period cannot be affected by treatment |
| Ratio metrics | Randomize by user, measure per session | The unit of analysis differs from the unit of randomization, so use the delta method or a user-level metric |
| Multiple metrics and variants | Many secondaries or arms | Name one primary metric, and correct secondaries for multiple comparisons |
| Novelty and regression to the mean | Effect shrinks over time | Plot the effect by exposure week before calling it |
Two of these deserve a sentence of context. On duration, interview101's Netflix guide says that running a two-week experiment on a metric that responds over a twelve-week cycle produces underpowered, potentially misleading results. On fading effects, interview101's Google guide describes a strong answer that considers novelty, checks for regression to the mean, and asks about spillover when treatment and control users interact.
On trust checks, Uber's engineering blog says its older platform "often produced systematically imbalanced control/treatment cohorts". That is the failure a sample-ratio check exists to catch, and naming the check unprompted is a cheap senior signal.
What question shapes hide behind the A/B testing label?
At least three, and most candidates rehearse only the first. Interview101's Netflix guide lists design an experiment for this feature, critique this experiment brief, and read data from a completed experiment as the three archetypes. In our bank, 28 of the 46 questions sit in the A/B testing family and the rest spread across other design families.
| Shape | Example | First move |
|---|---|---|
| Design a test for a feature | Design a free-month experiment (OpenAI) | State the decision and the unit |
| Critique a brief | interview101's Google guide: a ranking change where users see different results based on past behaviour | Find the broken assumption before fixing details |
| Read a finished test | interview101's Google guide: engagement is higher among users who enable notifications | Ask whether assignment was random |
| Interference | interview101's Google guide: treatment users can interact with control users | Change the unit or the design |
| Metrics and dashboards | Build dashboard; diagnose engagement–purchase gap (Meta) | Define the funnel before the chart |
The non-A/B families behave differently. The bank files 4 questions under ads, recommendation and ranking ML and 3 under metrics, logging and data pipelines, with Meta's How would you evaluate Pixel issue alerts? and Amazon's Design an operations dashboard with justifications in the second group. For those, the deciding question is what the metric or alert will be used to decide, and how you would know it is wrong.
When should you not run the experiment?
When the test cannot answer the decision. Interview101's Netflix guide says that after a design, the interviewer asks "under what conditions would you not run this experiment?". The same guide calls it a probe every candidate gets, aimed at "whether the candidate recognizes that experiment design and experiment appropriateness are separate questions."
Interview101's Netflix guide also says the most common strong-no-hire outcome is a candidate whose statistically valid experiment would mislead at the company's operating scale, with no flag raised. The usual cause is the unit. Interview101's same guide notes that a standard test assigns users and assumes they are independent, and describes a hire-level answer that asks whether household members sharing an account are independent before designing anything else. Interviewnode's Meta guide calls assuming the user is always the right unit one of the most common interview pitfalls.
Say the "not run" conditions out loud. Our view of the standard list:
- The traffic cannot detect the smallest effect worth shipping within the time the business will wait.
- The outcome lands after the decision is due, as with a free month whose revenue arrives later.
- Interference you cannot contain by changing the unit, such as shared supply in a marketplace.
- A control arm that causes harm, such as leaving abuse up or charging some customers more without a reason you could defend.
- The decision is already made, or cannot be reversed, so the result would change nothing.
When the unit must change, name the design. Uber's engineering blog says its older analysis stack supported only user-randomized experiments and that its core abstractions handled only a narrow set of designs correctly. Cluster randomization, geo or market splits and switchbacks are the standard answers. For ranking, interview101's Netflix guide says Netflix documented interleaving because traditional splits lacked statistical power at recommendation timescales.
What changes at senior and staff level?
The prompt usually stays the same; the bar on your answer moves. PrepVector's webinar write-up says that for senior and staff roles, "surface-level experimentation knowledge is now table stakes, not a differentiator." This set carries no level split we can quote, so the table below is our view, built from the cited guides.
| What the interviewer hears | Mid-level | Senior | Staff |
|---|---|---|---|
| Unit of randomization | User, stated | Unit chosen per use case, with the interference named | Unit tied to platform limits and the cost of a wrong unit across many tests |
| Metrics | A sensible primary metric | Primary metric plus guardrails, and which guardrail can veto | Metrics that cannot be gamed, aligned with long-run value |
| Uncertainty | A p-value and a call | Intervals, power and the error rate the business tolerates | Expected loss from both error types, framed as a business trade-off |
| When not to run | Rarely raised | Raised without prompting | Proposes the cheaper evidence instead, such as offline evaluation or a staged ramp |
| Rollout | Ship or not | Ramp plan with stop conditions | Holdouts, ownership and how the result changes the roadmap |
An account on sirjohnnymai.com describes a senior data scientist who asked a candidate to quantify false-positive risk and then what to do if the business tolerates a 10% error rate. The same account on sirjohnnymai.com describes a cost-sensitivity analysis, expected loss from false-positive and false-negative rates, as the correct response. Interview101's Meta guide lists translating statistical uncertainty into product recommendations as a tested competency, and interview101's Netflix guide says Netflix's culture document expects people to push back on approaches, data-driven ones included, when they have well-reasoned grounds.
Our take: do not hunt for a separate list of staff-level prompts. Take the titles in the first table and answer each at the staff column: the unit with its cost, the guardrail that vetoes, the expected-loss framing, and a ramp plan with a stop condition.
Which rounds decide whether you are down-leveled?
Our take: the phone screen mostly decides whether you continue, and the level is set elsewhere. Three places carry the level signal: the onsite case, where interference and rollout come up; the deep dive on experiments you have run; and behavioral questions about decisions you changed with data.
DoorDash's onsite Design experiments for marketplace product changes is the kind of prompt where depth shows. A senior answer picks between a switchback, a geo split and a user split and defends the bias against the variance. For the project deep dive, prepare one experiment you ran end to end: what it decided, what the unit was, what went wrong and what you would change. For behavioral scope, have one story where you argued against running a test, or against shipping a significant result, and say what the evidence was. Ask your recruiter which round carries the case discussion, and treat that round as the level round.
How we counted
Counts come from TrueInterview's question bank as of 2026-10-11 and cover the experiment design questions filed under 14 big-tech companies. They are candidate-reported questions reconstructed for practice, not any company's official list. A question can count in more than one round, families are assigned from question titles, and the "our reading" columns are editorial judgement, not stored tags.
FAQ
Are these the questions big tech companies officially ask?
No. They are candidate-reported questions reconstructed for practice and filed under the company where the candidate met them. Use them to decide what to rehearse first and to hear the shape of a prompt at your target company. Do not expect your exact prompt: the interviewer will change the product, the metric or the constraint, and a memorised answer breaks on the first change. A skeleton you can apply to an unseen product transfers; a script does not.
What are the basics of A/B testing you must know cold?
Random assignment, a pre-registered primary metric, guardrails, a power calculation and a decision rule. Candor's interview guide frames answers in four parts: design and methodology, result measurement, running the experiment, and launch decisions. Beyond that, know how to check the assignment split, why peeking inflates false positives, and when the user is the wrong unit. Experienced candidates are expected to explain each one in a sentence and say when it changes the design.
How are A/B testing questions different for experienced candidates?
The prompt is often the same; the expected answer is wider. The account on sirjohnnymai.com says interviewers expect a discussion of statistical power, prior probability and business impact rather than a significant or not significant label. At senior level, add the error rate the business can tolerate and a ramp plan; at staff level, add what the result changes for the roadmap and which cheaper evidence could replace the test.
How long should an A/B test run?
Long enough to detect the smallest effect worth shipping, in whole weeks, and past any novelty period. DataLemur lists determining test duration as a standard interview question. Work it out from the minimum detectable effect, the metric's variance and the traffic available, then check that the window covers the time your outcome takes to appear. If the answer is longer than the business can wait, say so and propose a validated leading metric or a different design.
Do machine learning engineers get experiment design questions?
Yes, mostly through ranking and ads. Interviewnode's Meta guide says Meta evaluates whether you can prove a model works in the real world, not only whether you can build one. For ranking changes, know interleaving, offline-to-online metric gaps and long-running holdouts, and expect the same decision-first structure as a product data science round. The ads and ranking titles in the company table are the right practice set for this.
Your practice plan
The target is narrower than a statistics curriculum: one answer skeleton, a handful of question shapes and the titles your target company reports. This plan assumes four or five evenings a week for three weeks alongside a full-time job.
- Week one, first evening: write the skeleton from the grading table on one page, then say it aloud against How would you evaluate stolen-post detection? on a ten-minute timer.
- Week one, remaining evenings: run the same skeleton on Choose a precise A/B test primary metric and Design a free-month experiment, and close each answer with a ship, ramp or stop rule. Most phone-screen reports sit in this family, 23 of 32.
- Week two: work the interference set, Design a switchback and choose block length, Design a pricing experiment with network effects and Design a network-aware Wi‑Fi badge experiment, naming the unit and its bias-variance trade-off before anything else.
- Week two, two evenings: drill the follow-up table: power, sample ratio checks, CUPED and peeking, each explained in two sentences without notes.
- Week three: run Design experiments for marketplace product changes and Decide if ad load is optimized as full onsite mocks, including a ramp plan and stop conditions.
- Week three, if you are senior or staff: re-answer How would you evaluate adding video ads? at the staff column of the level table, with an expected-loss framing, and prepare one experiment from your own work for the deep dive.
- Last evening: in TrueInterview's question bank, filter to your target company and say each recent title aloud with the decision, the unit and the condition under which you would not run it.
Sources
- TrueInterview question bank — experiment design at big-tech companies — counted 2026-10-11
- Supercharging A/B Testing at Uber — checked 2026-10-11
- 50 A/B Testing Interview Questions & Answers — checked 2026-10-11
- Cracking A/B Testing Interviews: Webinar Recap — checked 2026-10-11
- Meta Data Scientist Interviews Weight Experiment Design Over Statistical Methods — checked 2026-10-11
- Cracking A/B Testing Interviews — checked 2026-10-11
- P-Value Misinterpretation in Netflix DS Experimentation Interview: A Common Pitfall | Johnny Mai — checked 2026-10-11
- Netflix Data Scientist Interviews Weight Experimentation Design Over Analysis — Here's What That Means for Your Loop — checked 2026-10-11
- Google Data Scientist Interviews Evaluate Statistical Rigor Differently Than They Did in 2021 — checked 2026-10-11
- Meta AI Interview: Large-Scale Experimentation and A/B Testing Questions (2026) - Interview Node Blog — checked 2026-10-11
- A/B Test Interview Questions Asked at FANG | Candor — checked 2026-10-11
Last reviewed: 2026-10-11.