ByteDance · ML & AI Fundamentals
Derive L1 vs L2 effects with correlation
TrueInterview
October 7, 2026 · 1 min read
Suppose two standardized predictors and have . The observed Gram matrix is and the cross-product with the response is . (1) For , compute the ridge (L2) coefficient estimate (report the numeric weights to two decimal places), and explain why L2 regularization tends to split the weight between correlated predictors. (2) Without running the full LASSO, explain which coefficient pattern the L1 solution is most likely to yield at — both weights comparable and small, one near zero with the other large, or some other pattern — and justify your answer using the geometry of L1 versus L2 constraint regions when predictors are strongly collinear. (3) Suggest elastic-net penalty settings ( and ) that would make selection more stable while keeping variance in check; explain how you would tune and , and state which validation metric you would choose if the aim is sparse interpretation with the smallest possible drop in PR-AUC.
Overview: This question tests comprehension of ridge/L2, LASSO/L1, and elastic net regularization, the consequences of multicollinearity, geometric intuition about constraint regions, and how hyperparameters affect coefficient patterns within Statistics & Math.