Microsoft · Statistics & Data Analysis
Test classifier difference with McNemar's test
TrueInterview
October 7, 2026 · 1 min read
Two classifiers, A and B, were evaluated on the same set of 10,000 labeled examples. The paired outcomes are:
- Both correct: ; both wrong: ; A correct and B wrong: ; A wrong and B correct: . Answer the following:
- Apply McNemar's test with a continuity correction to compute the test statistic and p-value for : the error rates are equal. Show the intermediate values , , , and .
- Compute the exact binomial p-value for the same using trials, and explain when you would prefer the exact test.
- Give a 95% confidence interval for the accuracy difference on paired data; state which method you use and why.
- Discuss the assumptions, cases where McNemar's test is inappropriate, and how you would adjust if you compare A against 10 models (multiple testing control). Overview: This question evaluates a data scientist's ability to apply McNemar's test for comparing paired classifiers, a core competency in statistical hypothesis testing and model evaluation. It checks practical knowledge of contingency table analysis, exact versus asymptotic tests, confidence interval construction, and multiple testing correction in machine learning contexts.
Loading comments…