Squarepoint · Statistics & Data Analysis
Explain Linear Regression Assumptions, Evaluation, and Feature Addition
TrueInterview
October 7, 2026 · 2 min read
Describe the problem that linear regression addresses, how the model is estimated, which assumptions are important, and how you would assess its quality. Then discuss what happens to when additional features are introduced.
Part 1 — Fitting and Interpreting the Model
Give a linear-regression objective and explain how its coefficients are estimated.
What to Cover in This Part
- The response variable, the predictor matrix, the intercept, and the least-squares objective.
- A numerically suitable fitting approach, plus the role of rank or collinearity.
- How the assumptions differ for a useful conditional-mean model, for inference on coefficients, and for prediction.
Part 2 — Assessing Regression Quality
Select evaluation metrics and diagnostics, and explain what each one reveals.
What to Cover in This Part
- Error metrics and computed on a suitable validation split.
- Residual patterns, outliers, and the limits of depending on a single metric.
- The distinction between fitting the observed data and generalizing to unseen observations.
Part 3 — Adding Predictors
When features are added, does go up, go down, stay the same, or does the answer depend on the setting?
What to Cover in This Part
- The precise fitting and evaluation conditions under which training cannot fall.
- Why the larger model contains the original model as a special case.
- How adjusted and out-of-sample behave differently.
Questions to Clarify
- Are we considering unregularized least squares with an intercept, fitted on the same observations?
- Is the metric computed on training data or held-out data?
- Is the aim prediction, coefficient interpretation, or causal interpretation?
Hint: Compare the two feasible model sets. After a predictor is added, think about whether some choice of coefficients can reproduce every prediction the original model made.
What a Strong Answer Should Cover
A strong answer keeps the fitting objective separate from inference assumptions and evaluation. It states the conditions for the result exactly and does not claim that predictors or the marginal target must be normally distributed.
Follow-Up Questions
- Why can adding a feature raise training while lowering test performance?
- What occurs when two predictor columns are exactly redundant?
- When are normally distributed errors important, and when are they not needed to compute the least-squares fit?
Overview: Explain least-squares regression, inference assumptions, evaluation metrics, and why added features affect training and test differently.