Yelp · Statistics & Data Analysis
Recommend a Decision After a Nonsignificant Experiment
TrueInterview
October 7, 2026 · 2 min read
A primary metric from an A/B test shows an estimated 2.0% relative improvement, with a two-sided 95% confidence interval spanning -0.5% to 4.5% and a p-value of 0.12. Guardrail metrics show no meaningful adverse effects. A stakeholder still finds the change worthwhile.
Describe statistical significance and give a recommendation. Discuss practical significance, uncertainty, power, the pre-specified decision rule, and what further evidence would support shipping, iterating, or halting work.
Constraints & Assumptions
- The experiment and the primary metric were locked in before anyone looked at the data.
- There was no sample-ratio mismatch and no known problem with instrumentation.
- The confidence interval corresponds to the same relative effect as the reported estimate.
- Do not treat “not significant” as proof that there is no effect.
Clarifying Questions to Ask
- Once implementation and maintenance costs are considered, what is the smallest effect that would still be worthwhile?
- Was the experiment powered to detect that effect, and did it reach the intended sample size?
- Can the decision be reversed, and what do a false positive and a false negative cost?
- Were several variants, metrics, or interim peeks part of the analysis?
Hint — Compare the interval with the decision threshold: The interval includes zero, but it may also include effects that are large enough to matter, or too small to justify action.
What a Strong Answer Covers
- Reading the p-value and confidence interval correctly.
- Keeping evidence strength separate from business value.
- Acknowledging that the outcome is inconclusive, not necessarily a negative result.
- Basing the recommendation on the minimum worthwhile effect and the associated risk.
- Choosing a well-reasoned next step instead of post hoc metric shopping.
Follow-up Questions
- How would your recommendation change if the entire interval were below the minimum worthwhile effect but entirely above zero?
- In what situations is a non-inferiority design more suitable?
- How can a Bayesian decision analysis account for implementation cost and reversibility?
Overview: Decide how to handle a product change after a nonsignificant A/B test. Interpret p-values and confidence intervals correctly, separate practical significance from statistical significance, and pick next steps that are guided by the evidence.