Onemain Financial · ML & AI Fundamentals
Select and tune XGBoost hyperparameters
TrueInterview
October 7, 2026 · 1 min read
Suppose you are given a binary classification dataset containing 1,000,000 rows and 100 features, of which 20 are numeric and 80 are categorical one-hot encoded, with a 1% positive-class rate. Model training needs to complete within 5 minutes using one 16-core CPU and 32 GB of RAM.
-
Suggest starting XGBoost hyperparameters—eta/learning_rate, max_depth, min_child_weight, subsample, colsample_bytree, lambda, alpha, n_estimators, and max_bin or tree_method—and explain the reasoning for each choice in relation to bias–variance trade-off, class imbalance, and computational limits.
-
Outline a compute-efficient tuning approach, including the search space, early stopping, and a cross-validation design that prevents leakage when the same user appears in more than one fold.
-
Explain precisely how XGBoost treats missing values when splitting trees, and how this behavior interacts with one-hot encoding versus target encoding.
-
With a very small minority class, compare scale_pos_weight, weighted loss, and focal loss; in what situations would each option be most appropriate?
Overview: This question assesses the ability to choose and tune XGBoost hyperparameters, handle severe class imbalance and sparse one-hot encodings, manage missing values, and design compute-efficient training with grouped cross-validation to avoid user-level leakage.