ByteDance · ML & AI Fundamentals
Compare bagging vs boosting on imbalanced data
TrueInterview
October 7, 2026 · 1 min read
You need to identify 0.5% fraud among 10,000,000 time-ordered transactions that have 300 features (100 numeric and 200 one-hot). Pick between Random Forest (bagging) and Gradient Boosting (for example, XGBoost/LightGBM). State: (1) which one you would try first and why, touching on bias–variance trade-offs, margins, and how each algorithm responds to label noise on the minority class; (2) how you will address class imbalance (class weights vs downsampling the majority vs SMOTE/SMOTE-ENN), which primary metric you will optimize (PR-AUC vs ROC-AUC) and why; (3) an initial hyperparameter grid for each method (RF: n_estimators, max_depth, max_features, class_weight; GBM: learning_rate, max_depth, min_child_weight, subsample, colsample_bytree, scale_pos_weight), with concrete starting values and the effects you expect; (4) your validation plan to avoid leakage (for example, time-based blocked CV and grouping by user_id), and how you will use early stopping or OOB estimates; (5) two failure modes where boosting would do worse than bagging on this task and how you would diagnose them with plots or diagnostics.
Overview: This question tests skill in choosing between ensemble models (bagging vs boosting), managing extreme class imbalance, selecting metrics, tuning hyperparameters, validating in a time-aware way, and diagnosing failure modes for large-scale, time-ordered binary classification; it is often asked because it explores bias–variance trade-offs, robustness to label noise on the minority class, and practical evaluation and leakage-avoidance issues in Machine Learning. It evaluates both conceptual understanding of algorithmic trade-offs and statistical robustness, as well as practical application skills such as designing time-blocked cross-validation, choosing appropriate metrics for imbalanced data, specifying hyperparameter grids, and interpreting diagnostic plots.