Amazon · ML & AI Fundamentals
Compare Random Forests vs Gradient Boosting rigorously
TrueInterview
October 7, 2026 · 2 min read
You need to pick either a Random Forest (RF) or a Gradient-Boosted Trees model (GBT, such as LightGBM/XGBoost) for a binary classification task with these properties: 1,000,000 rows; 200 features (70% numeric, 30% categorical, some with high cardinality above 1,000); 20% missing values; class imbalance 1:50; moderate label noise (roughly 5–10% flipped labels); strong feature correlations; a strict online prediction latency limit of 20 ms per example; a training budget of 60 minutes on 16 vCPU and 64 GB RAM; and no deep learning permitted.
Answer every part precisely:
-
Choose RF or GBT for production and defend the choice using bias–variance trade-offs, robustness to label noise and outliers, capacity to model interactions, and stability when features are correlated. State the main risks of that choice.
-
Give concrete starting hyperparameters and tuning ranges for both models (RF: n_estimators, max_depth, max_features, min_samples_leaf, class_weight; GBT: learning_rate, n_estimators, max_depth or num_leaves, subsample, colsample_bytree, min_child_samples, reg_alpha, reg_lambda, scale_pos_weight). Explain the expected impact on bias/variance and latency.
-
Explain how you would encode categorical features (for example, out-of-fold target encoding, one-hot, hashing, or native categorical support) while avoiding leakage and keeping latency low; include your approach for high-cardinality features.
-
Describe your class-imbalance strategy (class weights versus sampling versus loss weighting) and how you will choose the primary metric (such as PR-AUC versus ROC-AUC) and the decision threshold. Include calibration plans (Platt versus isotonic) and how you would validate calibration.
-
Outline a 60-minute experiment plan: data split protocol (time-aware or stratified K-fold), feature preprocessing, tuning schedule (coarse-to-fine with early stopping for GBT and out-of-bag sanity checks for RF), and guardrails to catch leakage. Give a minute-by-minute or staged budget and a fallback path if training runs over.
-
Identify situations where RF would likely beat GBT on this dataset and situations where GBT would likely beat RF. Include how missing-value handling, monotonic constraints, correlated features, and distribution shift influence the decision.
-
Specify how you will generate and validate feature importances (permutation versus gain), partial dependence/ICE checks, and SHAP analyses, noting pitfalls when features are correlated or leakage is present. Finally, explain how you will satisfy the 20 ms inference latency budget (for example, tree depth limits, model compression, batching).
Overview: This question tests a candidate's ability to select and configure tree-based models (Random Forest versus Gradient Boosted Trees), manage high-cardinality categorical features and missing data, reduce the impact of class imbalance and label noise, produce trustworthy feature importance and calibration, and design an experiment and inference approach that satisfies tight latency and resource limits. It is often asked in Machine Learning interviews for Data Scientist positions to assess bias–variance trade-offs, robustness to correlated features and noise, hyperparameter and encoding choices, experiment design and evaluation, and to check both conceptual understanding and practical application for production systems.