Thumbtack · ML & AI Fundamentals
Choose clustering vs regression; explain KNN
TrueInterview
October 7, 2026 · 1 min read
In a business problem where only some outcomes are labeled, when is clustering the right choice instead of regression? State the decision criteria: how much labeling is available, the objective, the evaluation metrics, and the cost of errors. List at least four clustering algorithms—K-Means, hierarchical/agglomerative clustering, DBSCAN/HDBSCAN, and Gaussian mixture models—and compare their assumptions, key hyperparameters, scalability, distance metrics, and failure modes (for example, non-spherical clusters, varying density, high-dimensional sparsity, and mixed data types). Give concrete examples of situations where DBSCAN should be preferred over K-Means, and situations where the opposite is true. Finally, explain K-Nearest Neighbors to a non-technical stakeholder using a real-world analogy, then go deeper into selecting k, weighting by distance, the effects of feature scaling, the curse of dimensionality, and how to deploy KNN efficiently with KD-trees, ball-trees, or approximate neighbor methods.
Overview: This question tests the ability to choose between clustering and regression when outcomes are only partially labeled, along with familiarity with clustering algorithms (K-Means, hierarchical/agglomerative clustering, DBSCAN/HDBSCAN, and Gaussian mixture models), how K-nearest neighbors behaves, evaluation metrics, hyperparameters, and deployment considerations for a machine learning data scientist position. It is often used to assess judgment on supervised versus unsupervised approaches, the trade-offs created by label availability, objectives, and error costs, algorithmic assumptions, scalability, and practical deployment techniques; the expected level covers both conceptual understanding and practical application.