ByteDance · ML & AI Fundamentals
Explain SHAP vs VIF under collinearity
TrueInterview
October 7, 2026 · 1 min read
You have a binary target and two features A and B with , along with other features that are only weakly correlated. You train (i) a logistic regression and (ii) a gradient-boosted tree model. (1) Calculate or estimate the VIF for A and B, and explain what threshold values signal harmful multicollinearity. (2) Describe how SHAP values act when A and B are nearly identical, contrasting interventional and conditional SHAP, and why the attributions can be unstable or distributed unpredictably between A and B. (3) Suggest a sound interpretation workflow: feature clustering or grouped SHAP, permutation importance computed conditionally on the other feature, and refitting after dropping one member of the pair; describe the diagnostics you would expect. (4) Recommend changes to the modeling approach—for example, elastic net for the GLM or feature grouping/penalization for trees—and how you would check that both interpretability and predictive performance remain acceptable.
Overview: The question tests knowledge of multicollinearity diagnostics such as VIF, feature attribution methods such as SHAP, and the consequences of near-duplicate predictors for interpretation and validation in binary classification.