ByteDance · Statistics & Data Analysis
Interpret and validate regression with interactions
TrueInterview
October 7, 2026 · 1 min read
Consider modeling 7-day retention (retained_7d, 0/1) from user-level data using a linear probability model and logistic regression. The predictors are treated (1 if exposed to the new preloading), watch_time_day1 (minutes), new_user (1/0), and fixed effects for country and signup_date. The model includes a treated × new_user interaction.
(a) Write out both model specifications (LPM and logit) and interpret the coefficients on treated and on treated × new_user. For the logit, translate a coefficient into an odds ratio and then into an approximate marginal effect at the mean.
(b) Discuss the assumptions and diagnostics you would examine: heteroskedasticity, separation, multicollinearity (e.g., VIF), misspecification (e.g., link test), and calibration (e.g., reliability plots). What standard errors would you report and why (HC-robust vs clustered)? Justify the cluster choice.
(c) Suppose watch_time_day1 is endogenous (e.g., influenced by unobserved preferences). Propose two remedies and their assumptions: a control function/IV approach and panel fixed effects with within-user variation. Which instruments or proxies might be plausible here?
(d) You also have count data on daily videos watched. When would Poisson or negative binomial be preferable to OLS? Explain how you would check overdispersion and interpret the exponentiated coefficients.
Overview: This question assesses grasp of regression modeling with interaction terms, binary outcome models and odds/marginal-effect interpretation, causal inference ideas such as endogeneity and instrumental variables, model diagnostics and standard-error choices, and count-data modeling.