Meta · Statistics & Data Analysis
Test two models' proportions for significance
TrueInterview
October 7, 2026 · 1 min read
Two search models, A and B, were each run once by 100 different users, with a single query per user. Success is measured per query using your composite metric, where success is coded as 1 and failure as 0. Model A produced 90 successes and Model B produced 85. For a two-sided test at : 1) Write down and , select the appropriate test (pooled two-proportion z-test), calculate the test statistic and p-value, and decide whether A outperforms B. 2) Construct a 95% confidence interval for and explain its practical significance. 3) What sample size per arm is required to detect a +5 percentage-point uplift over an 85% baseline with 80% power at ? Show the formulas and inputs. 4) If you test the two models across 10 independent intents simultaneously, apply a Bonferroni correction and state whether your conclusion changes. 5) Briefly explain when you would favor an exact test or a Bayesian comparison, and what you would report for each.
Overview: This question assesses skill in statistical inference for proportions, spanning hypothesis tests, confidence intervals, power and sample-size calculations, multiple-testing correction, and frequentist versus Bayesian approaches, in the Statistics & Math domain for data scientist roles.