Google · Statistics & Data Analysis
Estimate b when features exceed samples
TrueInterview
October 7, 2026 · 1 min read
Take the linear model , where and an intercept column is included.
a) Work out the OLS estimator , and state both the rank conditions required for identifiability and the sampling distribution of under the classical assumptions.
b) Suppose instead that . Lay out at least three workable strategies (for instance ridge: ; lasso; elastic net; forward selection; PCA/PLS), and for each one explain how you would pick and how you would verify generalization (including the specifics of cross-validation).
c) Under what circumstances does the Moore–Penrose pseudoinverse produce a sensible minimum-norm solution, and what drawbacks come with it?
d) Explain why naively duplicating rows does not fix rank deficiency and can also damage inference.
Overview: This question gauges command of linear regression theory — identifiability and the sampling distribution of OLS — alongside high-dimensional competencies such as regularization, variable selection, dimensionality reduction, the properties of the Moore–Penrose pseudoinverse, and the statistical consequences of naive upsampling.
Read the full Data Scientist interview experience this question came from