Google · Statistics & Data Analysis
Compare two stores’ profits rigorously
TrueInterview
October 7, 2026 · 1 min read
Two snack shops are positioned at a school gate and open concurrently. You have a 14-day window to decide which one will earn higher profit over the coming quarter. Create a measurement and analysis plan that produces a defensible recommendation under real-world limitations. Specify:
- The precise primary metric (for example, profit per passerby per open hour) and why it aligns with the decision; state which costs and revenues are included and how you will normalize for foot traffic and operating hours.
- The minimum data to collect: hourly foot traffic counts, transaction counts and average order value (AOV), item-level margins, labor hours, rent and utilities allocation, weather, school calendar, and competitor promotions or price changes.
- Your identification strategy—either an experiment (such as randomized flyer distribution or alternating queueing) or a quasi-experiment (such as hour-level difference-in-differences with fixed effects and weather controls)—that stays valid despite weekday/weekend and exam-week spikes, Store B extending its hours, Store A offering weekend discounts, and one shop changing its price on day 9.
- The model and uncertainty: write the difference-in-differences regression you would fit (defining outcome, treatment, fixed effects, and controls), how you would compute a 95% confidence interval for the profit difference, and a minimal power check indicating that 14 days is enough (order-of-magnitude inputs are acceptable).
- Guardrails for detecting stockouts or cannibalization, plus an explicit decision rule (for example, recommend Shop X if the estimated profit difference exceeds $Y per day and the confidence interval lower bound exceeds $Z).
Overview: This question tests a data scientist's abilities in experimental design, causal inference (including difference-in-differences), profitability measurement and metric construction, uncertainty quantification (confidence intervals and power checks), and operational guardrails for real-world data collection.
This question is drawn from a data scientist interview experience.
Loading comments…