Meta · ML & AI Fundamentals
Which clustering algorithm would you use and why
TrueInterview
October 7, 2026 · 1 min read
Question
For a social product such as Meta, your task is to partition users into clusters so that meaningful groupings—communities, interest groups, or usage segments—can be identified. You may have one or both of the following data sources:
- A user feature table — dense numeric/categorical attributes for each user (age bucket, country, activity rate, topics engaged, embeddings, etc.).
- A social network graph — nodes represent users, edges represent friendships / follows / messages / interactions, and may be weighted and directed. Respond to the following:
- Traditional (feature-vector) clustering. Which clustering algorithms would you think about (e.g. k-means, GMM, hierarchical, DBSCAN/HDBSCAN), and how would you decide among them? Cover preprocessing, distance/similarity choices, how you would select the number of clusters, and how you would assess cluster quality.
- Social network / graph clustering. If the main data is a social graph rather than a feature table, which algorithms would you apply for community detection, and how is that fundamentally different from clustering a feature matrix?
- Directed and weighted graphs. How do you account for edge direction and weights when clustering a graph?
- Hybrid. If both graph structure and user features are available, how would you combine them?
- Choosing the number of clusters and evaluating quality. Which metrics and validation strategy would you use in the feature-vector case and in the graph case?
- Scale and operations. What practical issues appear when there are millions of users (compute, dynamic graphs, cold-start, drift), and how would you address them? Overview: A Meta Data Scientist machine learning screen about selecting a clustering algorithm for users of a social product. It compares traditional feature-vector clustering (k-means, GMM, hierarchical, DBSCAN/HDBSCAN) with social-graph community detection (Louvain/Leiden, spectral, SBM, node embeddings), and covers preprocessing, choosing the number of clusters, evaluation, directed/weighted graphs, and scaling to millions of users.
Loading comments…