Instacart · Statistics & Data Analysis
Improve low R² without p‑hacking
TrueInterview
October 7, 2026 · 1 min read
Suppose you fit a linear regression model to predict the contribution amount generated by each order, and the model yields . Address the following parts:
(a) Describe concrete modeling changes that can improve predictive performance without destroying the model's usefulness for inference. Include, at minimum: predictor transformations such as splines for basket size, interactions such as treatment-by-daypart, a more suitable response distribution and link such as Gamma with a log link, and checks for target leakage.
(b) Is it true that simply adding another covariate will reliably increase out-of-sample ? Justify your answer with a cross-validation argument, then propose alternative approaches such as generalized additive models, quantile regression, or gradient boosting while keeping the goal of effect estimation in mind.
(c) Explain how you would use nested cross-validation and target-leakage tests to protect against p-hacking while iterating on the model.
(d) Discuss when a low is acceptable for unbiased average treatment effect estimation but unacceptable for accurate individual-level predictions.
Examples
Example 1
Input:
4 1 2
1 2 3 4
2 4 6 8
Output: 1.0000000000 1.0000000000
Example 2
Input:
6 2 2
1 1 2 1 3 2 4 3 5 5 6 8
0 2 1 0 -4 -11
Output: 1.0000000000 1.0000000000