Voleon · ML System Design
Design and diagnose a regression pipeline
TrueInterview
October 7, 2026 · 2 min read
Your task is to predict 90-day customer value, CLV_90, at the user level from features available up to a cutoff date; the target has many zeros and a heavy right tail. The features include continuous marketing-channel spend (Spend_SEM, Spend_Social, Spend_Display), recency/frequency/monetary (RFM) variables, device, region, tenure, and dozens of sparse, high-cardinality categorical campaign IDs. Known problems are strong multicollinearity among the spend channels, heteroskedastic errors, and nonlinear effects. Build a regression pipeline that does the following: (1) select and justify an appropriate loss or distribution (for example, log-link GLM, Tweedie, zero-inflated, quantile regression, or gradient-boosted regression); (2) carry out feature processing (log1p transforms, standardization, rare-category bucketing, and target encoding for high-cardinality features using K-fold out-of-fold encoding to prevent leakage); (3) address multicollinearity and feature selection with ridge/lasso/elastic net: write the elastic net objective in terms of and , and explain when each penalty dominates; (4) create time-based nested cross-validation that avoids leakage from future signals and campaign overlap; (5) check model assumptions: detect and handle heteroskedasticity (for instance, with White-robust standard errors, variance-stabilizing transforms, or modeling the mean-variance relationship in a Tweedie GLM), nonlinearity (splines/interactions), and influential outliers (Huber/Tukey loss); (6) produce calibrated 95% prediction intervals for CLV_90 (such as conformal prediction or bootstrap) and compare them with OLS analytic intervals; (7) interpret effects: compute and interpret standardized coefficients for a regularized linear model versus SHAP values for a tree model; (8) quantify and reduce data leakage risks (for example, target leakage through post-cutoff features, or look-ahead in target encoding); and (9) compare performance and interpretability trade-offs between a regularized linear model and a gradient-boosting model, naming metrics that are robust to heavy tails (such as MAE, quantile loss, MAPE with epsilon). Overview: This question tests a data scientist's ability to design and diagnose an end-to-end regression pipeline for zero-inflated, heavy-tailed targets, with emphasis on feature engineering for high-cardinality variables, handling multicollinearity, choosing an appropriate loss/distribution and regularization, time-aware validation, uncertainty quantification, interpretability, and data leakage mitigation. It is often asked in Machine Learning interviews to assess practical application together with conceptual understanding of regression modeling, distributional assumptions, cross-validation strategies, and performance versus interpretability trade-offs.