ByteDance · ML & AI Fundamentals
Choose linear regression or decision tree appropriately
TrueInterview
October 7, 2026 · 1 min read
You are given 100,000 independent and identically distributed rows with features (0–100), , , and . The true data-generating process is unknown to you, but it is piecewise linear with a hinge at and an interaction: , with heteroskedastic noise . Outline an analysis to decide between linear regression and a decision tree. Specify: (1) the feature engineering and linearity tests you would run (for example, a spline basis for and an interaction), and how you would check residual diagnostics for heteroskedasticity; (2) a fair comparison protocol (CV split, identical preprocessing) and metrics; (3) how you would enforce monotonicity or interaction constraints in a tree-based model to reflect domain knowledge; (4) which model you expect to generalize better here and why, including bias–variance reasoning and how you would quantify it with learning curves.
Overview: This question evaluates model selection and diagnostic skills in supervised learning, specifically assessing feature engineering, interaction detection, handling heteroskedastic residuals, incorporation of monotonicity or interaction constraints in tree-based models, and fair cross-validation-based comparison between linear and tree approaches.