Google · ML & AI Fundamentals
Explain logistic regression vs forests and boosting
TrueInterview
October 7, 2026 · 2 min read
Answer every part precisely.
-
Define binary logistic regression and give the model . Work through the negative log-likelihood (log loss) and its gradients with respect to and . Explain what makes the loss convex and what that means for optimization.
-
Compare L1 and L2 regularization for logistic regression with respect to sparsity, multicollinearity, margin geometry, and probability calibration. When would you choose elastic net instead of pure L1 or pure L2?
-
Under what data conditions does logistic regression usually beat a random forest? Cover cases including (a) genuinely linear or nearly linear decision boundaries with few interactions, (b) high-dimensional sparse binary features such as text, (c) small-n, large-p settings where strong regularization is beneficial, and (d) situations where calibrated probabilities and interpretability matter most.
-
Your model is overfitting: list concrete fixes specific to each method—for logistic regression (regularization strength, feature selection, class weighting, calibration, proper cross-validation), for random forests (more trees, limiting depth, max_features, min_samples_* settings, out-of-bag validation), and for boosting (learning rate, number of estimators, max_depth or leaf-wise growth, subsampling, early stopping). Also describe how you would detect overfitting beyond accuracy, such as via calibration curves, PR-AUC versus ROC-AUC, and decision boundary checks.
-
Contrast random forests with gradient boosting on bias–variance behavior, robustness to noisy features, hyperparameter sensitivity, ability to respect monotonic constraints, native handling of missing values, and training/inference cost. Give one real-world scenario where each clearly dominates the other, and justify your choices.
-
Case study: You have 50k rows, 10k sparse binary features, a 1% positive class, and strong temporal drift. Propose an end-to-end pipeline for (a) logistic regression with elastic net and (b) a tree-based method (choose RF or GBDT). Cover feature processing, regularization/hyperparameters, evaluation protocol (time-based cross-validation), threshold selection, probability calibration, and how you would compare the models fairly. State the pitfalls you will avoid, such as leakage through target encoding and improper scaling of sparse inputs.
Overview: This question assesses a candidate's command of supervised learning ideas, including logistic regression's probabilistic formulation and convex optimization, the impact of L1/L2/elastic-net regularization, ensemble methods such as random forests and gradient boosting, and practical skills in diagnostics, probability calibration, and end-to-end pipeline design. It is frequently used in machine learning interviews for data scientist roles because it tests both theoretical foundations—loss functions, gradients, bias–variance trade-offs—and practical application, such as handling high-dimensional sparse or imbalanced data, hyperparameter sensitivity, and time-based evaluation, covering both conceptual understanding and hands-on model selection.