Citadel · ML & AI Fundamentals
Explain RF optimization and variable-importance pitfalls
TrueInterview
October 7, 2026 · 1 min read
Describe your approach to optimizing and regularizing a Random Forest regressor when working with tabular data. Include: (1) why standard random forests are not pruned after training, and how max_depth, min_samples_leaf, and max_features limit overfitting. (2) two importance measures—mean decrease impurity versus permutation importance—how each is calculated, the situations where they diverge, and their biases (such as favoring high-cardinality categorical variables or correlated features). (3) how to get dependable importance values through out-of-bag estimates, repeated permutations, or conditional permutation approaches that adjust for correlations. (4) practical ways to accelerate training on large datasets, for example subsampling, feature bagging, and warm-starting trees.
Overview: The question tests knowledge of Random Forest regularization and feature-importance diagnostics, including awareness of the biases that separate mean decrease impurity from permutation importance, along with considerations for trustworthy importance estimation and efficient training on large tabular data.