Onemain Financial · ML & AI Fundamentals
Handle missing data and outliers robustly
TrueInterview
October 7, 2026 · 1 min read
You are building a customer churn model whose inputs include numeric spend (strongly right-skewed, with about 2% extreme values), count features that are often zero, and categorical plan types; missing values occur both as MAR and MNAR (for instance, high-spend customers sometimes leave income blank). 1) Lay out a preprocessing pipeline suitable for both linear models and tree-based ensembles; it should address imputation options (median, KNN, MICE, model-based), missingness indicator flags, robust scaling, and outlier handling (winsorization, robust estimators, or isolation-based filters). 2) Discuss when each option is beneficial or harmful, and why—for example, how winsorization changes logistic regression versus tree splits, and what leakage risks arise with MICE. 3) Describe how you would empirically evaluate the pipeline’s effect on probability calibration and SHAP explanations while avoiding optimistic bias. 4) If roughly 10% of records are missing not at random on an important feature, what modeling or data-collection approaches would you use to reduce bias?
Overview: This question tests proficiency in ML preprocessing and robustness—particularly handling MAR versus MNAR missingness, treating outliers, choosing feature handling appropriate to linear and tree algorithms, and empirically assessing probability calibration and interpretability.