Study
Senior Data Science Product Case Questions: 50 Reported (2026)
TrueInterview
October 11, 2026 · 19 min read

As of 2026-10-11, TrueInterview's bank holds 50 candidate-reported product case questions from 11 big-tech companies, nearly half of them filed under Meta and most of them reported from phone screens. The reported stems are narrower than the phrase product sense suggests: measure a product, check whether a number is real, diagnose why a metric moved, or design a change and say how you would evaluate it. Our take: for a data scientist, product sense is a short list of measure, diagnose and design-and-evaluate prompts at the companies on your loop, so rehearse those stems aloud at phone-screen length before you study any framework taxonomy.
Disclosure: TrueInterview is an interview-preparation product and publishes this article. Facts about other products come from their public pages on the dates listed under Sources.
Which product case questions are candidates reporting at big tech?
Mostly prompts about a number: define it, check that it is trustworthy, or explain why it moved. Ten reported stems show the range, from a Meta phone-screen prompt about repairing a game's revenue reporting to an Uber onsite question that pairs a feature with a safety analysis. The archetype column is our reading of each title, not a tag stored in the bank.

| Question | Company | Round | Last reported | Archetype (our reading) |
|---|---|---|---|---|
| Measure Game Monetization and Validate a Reporting Repair | Meta | Phone screen | 2026-09 | Measure, then validate the number |
| Diagnose Low CTR in an Advertising Campaign Funnel | ByteDance | Phone screen | 2026-04 | Diagnose a funnel metric |
| How would you prevent wrong items in deliveries? | DoorDash | Onsite | 2026-03 | Design and evaluate a prevention plan |
| Design an Uber feature and analyze safety | Uber | Onsite | 2026-01 | Design and evaluate, with a safety guardrail |
| Analyze homepage drop and feed ranking | Phone screen | 2026-01 | Diagnose, with a ranking change in scope | |
| Design and Evaluate a Home Carousel | Phone screen | 2026-01 | Design and evaluate | |
| Evaluate Stripe Capital Loan Performance | Stripe | Phone screen | 2026-01 | Measure a financial product |
| Diagnose sustained drop in executed trades | Robinhood | Phone screen | Not in our extract | Diagnose a sustained decline |
| Drive stakeholder alignment under trade-offs | Amazon | Online assessment | Not in our extract | Behavioral trade-off |
| Lead a zero-to-one initiative effectively | Instacart | Phone screen | Not in our extract | Behavioral, zero-to-one scope |
Half of the ten titles, by our reading, ask you to measure or diagnose a number, three start from a design, and two are behavioral prompts set in a product. If your preparation so far has been feature brainstorming, this table says to move most of your hours to metric definition and diagnosis. The four prompts below are the ones we would rehearse first, with the decision an interviewer is most likely to push on.
Measure Game Monetization and Validate a Reporting Repair (Meta)
The title carries two jobs: define monetization for a game, then prove that a repaired report now tells the truth. Start with the second job, because every product conclusion depends on it. Reconcile the repaired figure against an independent record such as raw purchase events, compare old and new numbers on overlapping days, and say whether the repair changed the definition or only the pipeline. The deciding questions are which table is the source of truth and how you restate history without breaking trend comparisons.
Diagnose Low CTR in an Advertising Campaign Funnel (ByteDance)
Click-through rate is a ratio, so split it before you explain it: did clicks fall, or did impressions grow into lower-intent placements? Then walk the funnel by stage and cut by placement, creative, audience and device. The decision an interviewer will push on is mix shift: a campaign can lose CTR in aggregate while every segment holds steady, and your answer should test for that before naming a cause.
Analyze homepage drop and feed ranking (LinkedIn)
This is a diagnosis prompt with a suspect already named. Scope the drop first (which homepage metric, how large, since when, which surfaces), rule out logging changes, and only then ask whether a ranking change altered what members see. The senior-level push is on the evaluation of ranking itself: which short-term engagement metric the model optimises, and which longer-horizon metric you would hold as a guardrail so a ranking win does not quietly cost return visits.
Evaluate Stripe Capital Loan Performance (Stripe)
A measurement prompt in lending, where the trap is time. Loans mature, so compare repayment and loss by origination cohort rather than as one pooled average that mixes young and old loans. The deciding trade-off is approval volume against loss rate: an interviewer will want to hear which of the two you would hold fixed, and why, before you recommend loosening or tightening anything.
Which companies have enough reported questions to study directly?
By our reading of the counts, only Meta. Of the 50 reported product case questions, the ten companies the bank breaks out hold Meta 24, DoorDash 6, Uber 5, Stripe, ByteDance and Amazon 3 each, Instacart 2, and LinkedIn, Pinterest and Robinhood 1 each, with the remaining question at an eleventh company. If Meta is on your loop, work its set directly; for any other company, drill the archetypes and use that company's titles as a final check.
These counts measure what candidates reported into the bank, not how often a company runs a product case. A company with a single reported title is under-sampled rather than uninterested, and its one question is an archetype example, not a forecast of your interview.
The outside signal points at the same loops. Hacking the Case Interview's data science guide calls product and analytics cases the most common type, "especially at tech companies like Meta, Google, Airbnb, and DoorDash". Two of the four companies the guide names, Meta and DoorDash, are also the top two in our counts, so the sources agree on the heaviest loops but not on the full list.
Our take: big tech is the wrong unit of preparation for this round. Generic lists treat a Meta loop and a Robinhood loop as the same exercise. The reported volume says otherwise: a Meta candidate can rehearse against a real company set, while a LinkedIn, Pinterest, Robinhood or Instacart candidate should stop hunting for a company list that is one or two titles deep and practise the archetype those titles belong to.
Where does the product case show up: phone screen, onsite or assessment?
Mostly in the phone screen. Of the reported rounds, 28 are phone screens, 21 are onsites and 1 is an online assessment. Plan for a spoken case in your first technical conversation, and stop treating the product case as a rare onsite curveball that you can prepare for in the final week.
Within the behavioral/knowledge format, every one of the 25 questions reported from phone screens and every one of the 19 reported from onsites carries the product case subtype. The archetype does not change between rounds; the length and depth do. Hacking the Case Interview's guide says a phone screen may bring a 10 to 15 minute case mixed with technical questions, and that "These are usually simpler and test your basic product sense and analytical thinking." The same guide puts onsite cases at 30 to 45 minutes, and says a case can arrive as a verbal case, a take-home or a presentation.
Our take: rehearse the phone-screen version as a short spoken answer with no slides and no dataset: the archetype, the metric, two or three ranked hypotheses and one recommendation. Save the full-length rehearsals, with experiments and guardrails, for the onsite-style titles from DoorDash and Uber. If your only product case signal is an assessment, as with Amazon's stakeholder trade-off title, prepare a written answer that states the options, picks one and names the metric you would watch.
How do you tell measure, diagnose and design-and-evaluate prompts apart?
By the verb in the stem and the deliverable it implies, and you should say which one you heard within your first minute. Measure prompts want a metric definition, diagnose prompts want ranked causes and the check that separates them, design-and-evaluate prompts want an intervention plus an experiment, and trade-off prompts want a decision you can defend.
| Archetype | Wording in the stem | What you deliver | Titles from the bank | What a senior interviewer pushes on (our view) |
|---|---|---|---|---|
| Measure | Measure, evaluate, define success | Primary metric with numerator, denominator and window, plus guardrails | Meta reporting repair, Stripe loan performance | Whether the number can be trusted, and cohort effects over time |
| Diagnose | Diagnose, analyze a drop, why did it move | Scoped change, ranked hypotheses, the cut that confirms or kills each | ByteDance low CTR, LinkedIn homepage drop, Robinhood executed trades | Mix shift against real behavior change, and logging before product |
| Design and evaluate | Design, prevent, design and evaluate | One intervention, its success metric, a guardrail and an experiment | DoorDash wrong items, Uber feature and safety, Pinterest home carousel | Unit of randomization, and what the change could break |
| Behavioral trade-off | Drive alignment, lead an initiative | A decision, the options you rejected and how you brought people along | Amazon stakeholder alignment, Instacart zero-to-one | Scope of the decision and what you would do differently |
Third-party guides list similar types. DataLemur's product sense guide names four common types of product-focused questions for data scientists: product metrics, diagnosing a metric change, brainstorming features and designing A/B tests. Hacking the Case Interview's guide names product analytics, business strategy, machine learning and experimentation as the four main data science case categories.
Our take: the taxonomies are fine as a map, but the skill being scored is classifying the stem and structuring to it, not reciting a framework. The same Hacking the Case Interview guide says "Candidates who jump straight into analysis without organizing their thoughts rarely do well", while its Microsoft guide warns "Memorized frameworks fail here, so tailor a custom structure to the specific product and objective". The author of intrico.io's list of product design mistakes says being too framework-y was one of the top three reasons they saw interviewers reject candidates at Google. Saying the archetype aloud solves both problems at once: it is a structure, and it is fitted to the prompt in front of you.
For diagnosis prompts, one author's Medium revision guide describes a usable hypothesis checklist: accidental causes such as bugs and logging errors, natural ones such as seasonality, internal ones such as feature or algorithm changes, and external ones such as competitors. Use it as a checklist you run silently, then say only the two or three buckets the stem makes plausible.
What does a strong answer to DoorDash's wrong-items prompt look like?
A metric definition, a stage-by-stage breakdown, ranked fixes that each carry a guardrail, and one experiment. The prompt is reported from a DoorDash onsite, last reported 2026-03, and is linked from 3 candidate write-ups, so you can check your answer against several candidates' reports. What follows is our worked construction from the title, not a reported answer.
Classify it first. It is a prevention prompt, so it belongs to design and evaluate, and the deliverable is an intervention plan with a measurement, not a root-cause hunt. Hacking the Case Interview's Dropbox guide puts the shape of a strong answer simply: "Every strong answer starts with the user problem and ends with a clear success metric".
- Define the metric. Wrong-item rate is deliveries with a reported missing or incorrect item divided by completed deliveries. State the reporting channel and the window in which a complaint counts, because the number only means what its definition says.
- Break it down by stage. Menu and order entry, restaurant picking and sealing, the courier handoff, and the customer's check on arrival. Each stage has a different owner, so each needs its own share of the errors.
- Rank the fixes and attach a guardrail to each. Order hypotheses by volume times fixability. Pair every fix with the metric it could hurt: courier wait time, delivery time or restaurant preparation time, so accuracy is not bought with speed.
- Pick one intervention and test it. An item-count confirmation at handoff is a reasonable first choice. Randomize by store rather than by order, since staff at one store learn the new step and would contaminate an order-level split. Use wrong-item rate as the primary metric and delivery time as the guardrail.
The senior-level push on this prompt is the definition in step one. Customer-reported errors can be inflated by refund abuse, so a strong answer names a second, harder signal, such as restaurant-confirmed errors, and says which one the experiment will be judged on.
Reuse the skeleton. The same four steps carry Uber's feature and safety prompt, where the safety metric becomes the guardrail, and Pinterest's home carousel, where the guardrail is engagement elsewhere on the home surface.
Why do some product cases arrive as system design?
Because a number has to be trustworthy before anyone reasons about the product behind it. Of the 50 product case questions, 45 are filed in the behavioral/knowledge format and 5 as system design. If you rehearse only metric discussion, a system design version of the same case will find you without an answer on instrumentation and data integrity.
Meta's Measure Game Monetization and Validate a Reporting Repair is one of those system design filings, reported from a phone screen. Meta also has 20 questions in the behavioral/knowledge format, all with the product case subtype, so a Meta candidate should expect both shapes. Product manager guides show the same overlap: Aakash Gupta's case interview guide lists "Design a scalable system for video streaming at Spotify." as a technical product system design question.
Our take: treat the reporting-and-instrumentation family as part of product sense, not a separate track. For any prompt that asks you to validate a number, the deciding questions are standard data engineering ones: where events can be dropped or double-counted between client and warehouse, how late-arriving events are handled, and how you would restate history once the bug is fixed. Give a metric tree and one data-integrity hypothesis before you offer any product idea.
How recent are these questions?
Recent enough to schedule against: 19 of the 50 were last reported within the 12 months before 2026-10-11. Put those first in your drill order and treat the older titles as pattern references, because a last-reported date is the most recent report of a question, not proof that older prompts have retired.
Seven of the ten titles in the first table fall inside that window, which is why the practice plan below starts with them. TrueInterview says its questions are written up right after the interview and then sorted by company and round, which is what makes a last-reported month available for each title.
There is also a reason to prefer live prompts that has nothing to do with our bank. The Pragmatic Engineer's survey reports that 58% of Big Tech interviewers have adjusted the kind of questions they ask in response to suspected AI cheating tools, and lists asking candidates why they made each choice as one of the adaptations. Generic sample questions age faster under that pressure than the stems candidates are reporting this year.
What changes at senior and staff level?
The prompt stays the same; what is scored moves from structure to scope. We cannot publish a self-reported level split for this set, so the table below is our view: four reported titles, each answered at three levels. The third-party anchors describe product manager loops at Meta and Google, so read them as the direction of travel rather than a data science rule.
Aced's Meta product sense guide says that for IC roles at L5/P4 to P5 the typical sequence is two phone screens, one product sense and one product execution or analytical, followed by an onsite with another of each and a behavioral or leadership round. Productinterview.com reports that Meta added a fourth final-loop round, Product Sense with AI, for IC6 and M1/M2 roles: 30 minutes of product sense followed by 30 minutes prototyping with a real AI tool. Productinterview.com says evaluators in that prototype segment are not scoring code quality, and adds: "They are scoring judgment about scope". A review of Google PM loops on sirjohnnymai.com quotes one committee comment: "Candidate shows high execution skill but low strategic judgment; risk of plateau at L5."
| Reported title | Below senior | Senior | Staff |
|---|---|---|---|
| Meta reporting repair | Defines monetization and proposes a dashboard | Reconciles the repaired figure against raw purchase events before any product claim | Decides how history is restated and which table becomes the source of truth |
| LinkedIn homepage drop | Lists possible causes | Scopes the drop, rules out logging, then tests the ranking change | Names the longer-horizon guardrail the ranking change should have been held to |
| Stripe loan performance | Reports a pooled repayment rate | Compares repayment and loss by origination cohort | Chooses whether to hold approval volume or loss rate fixed, and the result that would reverse it |
| DoorDash wrong items | Proposes a fix | Defines the metric, randomizes by store, guards delivery time | Adds an error signal that refund abuse cannot inflate, and sequences the rollout |
Our take: below senior, the product case is largely a structure test. At staff level it becomes a scoping test: what you choose not to build, and what you would instrument to learn that you were wrong. Rehearse the staff column on each title in the first table: the cut, the guardrail and the reversal trigger, instead of a longer feature list.
Which rounds decide whether you are down-leveled?
Our take: in a data science loop, the product case phone screen mostly decides whether you continue. The level is set later, in the onsite case follow-ups and in the scope of your behavioral answers. A clean structure gets you to the onsite; your handling of two conflicting metrics, and the size of the decisions in your stories, set the level.
Aced's Meta guide supports the first half: "The initial prompt is the easy part. Meta's evaluation happens in the follow-ups." It describes the approach that works when two metrics conflict as clarifying the metric definition first, listing possible explanations, then prioritising which one to investigate and why. The sirjohnnymai.com review records a senior interviewer's note that one candidate lost sight of the decision-making process and over-engineered the answer, which is the staff-level failure in miniature: more analysis, less decision.
For behavioral scope, the bank's two behavioral titles are the rehearsal material: Amazon's stakeholder alignment under trade-offs and Instacart's zero-to-one initiative. If you are interviewing at senior or staff level, ask your recruiter which onsite rounds carry the product case and which carry leadership, and prepare one story per round where you owned a metric decision across teams.
How we counted
Counts come from TrueInterview's question bank as of 2026-10-11 and cover product case questions filed under 11 big-tech companies. They are candidate-reported questions reconstructed for practice, not any company's official question list. A question can count in more than one round, company and round tallies describe what candidates reported rather than how often a company asks, and the archetype column is our reading of each title. A last-reported date is the most recent report of a question, not its only one.
FAQ
How do you pass a product sense interview as a data scientist?
No preparation guarantees a pass, but the titles in the first table show the deliverables interviewers have asked for. Name the archetype, define the metric with its numerator and denominator, rank two or three hypotheses, and end with one recommendation plus the guardrail you would watch. Hacking the Case Interview's data science guide also stresses communication: "If the interviewer cannot follow your logic, they cannot evaluate it". In practice, say each step before you take it, and confirm the metric definition with the interviewer before you build hypotheses on top of it.
Is a data science product case the same as a PM product sense interview?
They overlap, but the deliverable differs. DataLemur describes the product sense interview as evaluating a candidate's capacity to understand and strategize product development, while Hacking the Case Interview says data science cases test whether you can apply technical analytical skills to real business problems. The titles in our first table lean analytical: validating a reporting repair, diagnosing a funnel, evaluating loan performance. Give feature ideation about a quarter of your hours and spend the rest on metric definition and diagnosis.
Are these the questions big tech companies officially ask?
No. They are candidate-reported questions reconstructed for practice, and no company has published them as its question list. Treat each title as evidence of the kind of prompt a company has used in a given round, not as a promise of what you will be asked. That is also why the plan below drills archetypes on top of titles: if your interviewer changes the product or the metric, a structure you can rebuild will survive where a memorised answer will not.
What if my target company has only one or two reported questions?
Use them to identify the archetype your company tests. Work out which archetype each title belongs to, then practise that archetype on the deeper sets in the first table, especially the Meta and DoorDash titles, which carry more reports. In the last week, return to your company's titles and run each as a timed mock. Also ask your recruiter whether the product case sits in the phone screen or the onsite, because that changes the length you should rehearse.
Should I use AARRR or another framework?
Use one as a private checklist, not as a script you recite. One author's Medium revision guide describes following the AARRR funnel to define success metrics and naming guardrail metrics that should not degrade while you optimise the north star. Both are useful prompts for your own coverage. Hacking the Case Interview's Dropbox guide gives the counterweight: "The interviewer cares less about a memorized framework and more about how clearly you move from a messy prompt to a sharp recommendation."
How many product case questions should I practise?
Start with ten and run each one twice. The ten titles in the first table, rehearsed aloud and then rewritten once at the senior bar, cover all four archetypes used in this article. If Meta is on your loop, add the Meta set. On the second run, change one constraint, such as a fixed budget or a guardrail that now conflicts with the primary metric, because the follow-ups are where the case is judged.
Your practice plan
The target is narrower than a generic product sense list: ten named stems, four archetypes and your own company's set. This plan assumes four or five evenings a week for three weeks while you work full time.
- Week one, first evening: read every title in the first table and say its archetype and deliverable aloud in one sentence each. Then do the same for any titles you find for your target company in TrueInterview's question bank.
- Week one, remaining evenings: run the phone-screen titles as fifteen-minute spoken answers with no slides: Measure Game Monetization and Validate a Reporting Repair, Analyze homepage drop and feed ranking, Evaluate Stripe Capital Loan Performance, Design and Evaluate a Home Carousel and Diagnose sustained drop in executed trades.
- Week two: write out How would you prevent wrong items in deliveries? end to end with the four steps above, then reuse the skeleton on Design an Uber feature and analyze safety and Diagnose Low CTR in an Advertising Campaign Funnel at full onsite length.
- Week two, if Meta is on your loop: filter the question bank to Meta and work through its product case set, alternating measurement prompts with the reporting-repair style.
- Week three: prepare the behavioral edge with Drive stakeholder alignment under trade-offs and Lead a zero-to-one initiative effectively, each answered with one story where you owned a metric decision.
- Week three, if you are senior or staff: rerun your week-one prompts at the staff column of the level table, adding what you would not build and the result that would reverse your recommendation.
- Final evening: record one full mock of the DoorDash prompt and listen back for silent stretches. TrueInterview says it does not sell mock interviews with human engineers and suggests booking one or two elsewhere before an onsite.
Sources
- TrueInterview question bank — product case at big-tech companies — counted 2026-10-11
- Data Science Case Interview: Complete Guide (2026) — checked 2026-10-11
- 13 Product-Sense Interview Questions & Tips for Data Scientists — checked 2026-10-11
- Microsoft Case Interview: Complete Guide (2026) — checked 2026-10-11
- List of common mistakes during product design and product sense case interviews. — intrico.io — checked 2026-10-11
- Medium — checked 2026-10-11
- Dropbox Case Interview: How to Prepare (2026) — checked 2026-10-11
- PM Case Interview Guide: The Rubric and Anti-Patterns — checked 2026-10-11
- Real FAANG Interview Questions by Company · TrueInterview — checked 2026-10-11
- The Pulse #146: How AI is changing tech interviews — checked 2026-10-11
- Meta Product Sense Interview (2026 Guide) - Aced (formerly Exponent) — checked 2026-10-11
- How AI changed what PM interviews test in 2026 · productinterview.com — checked 2026-10-11
- Google PM Interview Process 2026: Data-Driven Teardown of Product Sense and Strategy Rounds | Johnny Mai — checked 2026-10-11
- Interview prep compared: your application, end to end · TrueInterview — checked 2026-10-11
Last reviewed: 2026-10-11.