Point72 · ML & AI Fundamentals
Explain and tune decision trees robustly
TrueInterview
October 7, 2026 · 1 min read
During an internship you trained a decision tree. Respond to the following with concise formulas and procedures:
-
Describe how a CART tree chooses splits for classification versus regression, covering impurity and variance criteria; give the exact formulas for Gini, entropy, and MSE, and explain how surrogate splits work when features contain missing values.
-
Provide a defensible method for selecting max_depth and min_samples_split: specify a cross-validation design, early stopping or pruning via the cost-complexity alpha path, and the metric you would optimize under severe class imbalance, with justification for PR-AUC versus ROC-AUC versus F1. Explain how you would choose alpha from the CCP path without leakage.
-
Overfitting checks: list at least three diagnostics, for example cross-validated gap relative to training, learning curves, permutation importance stability, and calibration curves. Which patterns specifically indicate overfitting for trees?
-
Given roughly 500k rows and about 300 features, including high-cardinality categoricals and sparse indicators, propose a preprocessing and modeling plan using a single decision tree: encoding choice, treatment of rare categories, monotonic constraints if any, feature binning, and computational cost. Give concrete hyperparameter ranges and the expected order-of-magnitude training time.
-
If you could revisit the project, under what conditions would a random forest or a gradient-boosted tree model such as XGBoost/LightGBM outperform a single tree on this dataset? Name at least three data or target conditions and the associated trade-offs, including variance, interpretability, latency, OOB versus CV, and calibration. How would you compare models fairly, including data splits, nested CV, fixed preprocessing, and an identical evaluation protocol?
Overview: The question tests a candidate's grasp of CART decision tree mechanics, split criteria and surrogate splits for missing values, hyperparameter tuning and pruning, overfitting diagnostics, preprocessing for large feature sets and high-cardinality categoricals, and the conditions for choosing ensemble methods within machine learning.