Snapchat · ML & AI Fundamentals
Explain Random Forest randomness and implications
TrueInterview
October 7, 2026 · 1 min read
Random Forest depth: 1) Identify all sources of randomness—bootstrap sampling, feature subsampling at each split, random tie-breaking, randomized split points—and describe how each influences bias and variance. 2) For a dataset containing 100k rows, 100 features, and a 5% positive rate, suggest values for n_estimators, max_depth, and max_features; explain how max_features limits correlation among trees. 3) Contrast out-of-bag (OOB) error with 5-fold cross-validation; when might they produce different results, and why? 4) Why are impurity-based importance scores biased toward continuous or high-cardinality features? Propose and defend a corrected approach (for example, permutation importance with stratified shuffles and repeated runs). 5) Outline class-imbalance strategies (class_weight, threshold moving, balanced subsampling) and discuss the implications for probability calibration and decision thresholds.
Overview: This question tests a candidate's understanding of Random Forest ensemble mechanics—sources of randomness, their impact on bias and variance, hyperparameter effects, evaluation choices (OOB vs cross-validation), feature-importance bias, and class-imbalance strategies—within the Machine Learning domain for binary classification.