Google · Statistics & Data Analysis
Analyze data duplication effects in linear regression
TrueInterview
October 7, 2026 · 1 min read
Take the ordinary least squares model where has full column rank and the errors are i.i.d. . Imagine you replicate the entire dataset times by stacking and vertically identical copies, then fit OLS again.
- Show algebraically that stays the same, while is multiplied by , so each standard error is multiplied by , and describe what happens to the t-statistics, p-values, and confidence interval widths.
- Numerical check: suppose initially a single coefficient has and . If you incorrectly treat four duplicated copies as independent (), what are the approximate new standard error, t-statistic, and two-sided p-value? Show the calculations.
- Give the correct interpretation of a p-value in this setting, and explain why duplicating data breaks its assumptions. When might analyst choices such as oversampling, data augmentation, or bootstrapping unintentionally produce effects similar to duplication?
- For chi-square tests, explain how a very large can produce very small p-values even when the effect is negligible. Suggest two remedies: (1) report and set thresholds on an effect-size measure such as Cramér’s V or an odds ratio with a confidence interval, and (2) use a penalized or Bayesian approach that shrinks spurious significance, such as penalized likelihood, Firth correction for small cells, or weakly informative priors. Explain when each remedy is preferable and how you would calibrate the thresholds.
Overview: This question tests understanding of OLS estimation, how duplicated observations affect estimator variance and inference, and the relationship among sample size, p-values, and effect-size measures in the Statistics & Math area for a Data Scientist position.
Loading comments…