ByteDance · ML & AI Fundamentals
Explain and tune XGBoost; prevent overfitting
TrueInterview
October 7, 2026 · 1 min read
Describe XGBoost's tree booster in sufficient depth to address: (a) Which objective it optimizes, and how the second-order Taylor expansion yields the split "gain" formula? Explain the roles of lambda (L2), alpha (L1), gamma (min_split_loss), and learning rate (eta) in that gain and in pruning. (b) Enumerate the most influential hyperparameters for tabular classification and, for each one, state the expected direction of impact on bias/variance and training time: max_depth, max_leaves, min_child_weight, subsample, colsample_bytree/level, eta, n_estimators, lambda, alpha, gamma, max_delta_step, scale_pos_weight, monotone_constraints. (c) You need to train a model that flags "bad sellers" when the positive rate is 0.5%. Outline a tuning plan that minimizes actual business cost: define a data split strategy that prevents leakage (for example, time- and seller-based splits), the main offline metric (such as PR-AUC), how to select an operating threshold from a cost matrix, how to use early stopping in a robust way, and which diagnostics or plots you would generate to catch overfitting and data leakage. (d) After training, how would you calibrate the predicted probabilities and explain the model to investigators (for example, with SHAP), while keeping attackers from reverse-engineering the rules?
Overview: This question tests knowledge of XGBoost tree-boosting internals, how hyperparameters affect bias/variance and training time, approaches for imbalanced binary classification and leakage-aware validation, post-training calibration, and interpretability for tabular fraud detection.