Capital One · ML & AI Fundamentals
Deep-dive XGBoost handling and overfitting
TrueInterview
October 7, 2026 · 1 min read
Technical / ML Deep Dive
You applied gradient-boosted decision trees (such as XGBoost or LightGBM) to a credit risk or response prediction problem. Respond to the following:
- Missing values: How do boosted trees treat missing values during training and inference? What choices do you have (native support versus imputation), and when would you choose one over the other?
- Overfitting control: What are the main causes of overfitting in boosted trees, and which techniques or hyperparameters would you use to reduce it?
- Evaluation: Which metrics would you use for an imbalanced credit outcome (for example, default), and how would you validate the model to make sure it generalizes? Be ready to discuss practical pitfalls (data leakage, time-based splits, calibration) and how you would debug them. Overview: This question tests proficiency with gradient-boosted decision trees and related competencies: built-in missing-value handling versus imputation, sources and control of overfitting through regularization and hyperparameters, metric selection and validation strategies for imbalanced outcomes, and practical debugging concerns such as data leakage, time-based splits, and calibration for a Data Engineer role. It often appears in Machine Learning interviews to assess both conceptual understanding of how the algorithm behaves and practical application of model evaluation and deployment-ready validation techniques. Read the full Data Engineer interview experience where this question appeared.
Loading comments…