Microsoft · ML & AI Fundamentals
Explain KNN and PCA and key tradeoffs
TrueInterview
October 7, 2026 · 1 min read
A Data Scientist internship interview includes questions on core machine learning concepts:
- K-Nearest Neighbors (KNN)
- Describe how KNN performs both classification and regression.
- How would you select k? What occurs if k is set too low or too high?
- How do you decide on a distance metric such as Euclidean or cosine?
- Which preprocessing steps matter most, including feature scaling and encoding categorical variables?
- Talk about its computational cost and how you would scale KNN to large datasets.
- What problems appear in high-dimensional feature spaces, often called the curse of dimensionality?
- Principal Component Analysis (PCA)
- Which optimization objective does PCA target? Explain the geometric meaning behind it.
- How is PCA actually calculated: through eigendecomposition of the covariance matrix or via SVD?
- How do you pick the number of components, using explained variance or cross-validation?
- In what situations can PCA reduce performance, such as loss of interpretability, non-linear structure, or data leakage?
- If PCA is applied before KNN, when could it improve results and when could it make them worse?
Give clear, interview-ready responses that include practical tradeoffs.
Overview: It tests knowledge of K-Nearest Neighbors, an instance-based method for classification and regression, and Principal Component Analysis, a linear dimensionality reduction technique, with attention to tradeoffs around distance metrics, preprocessing choices, computational cost, high-dimensional limits, PCA’s optimization and computation strategies, and how reducing dimensions affects downstream nonparametric methods. This type of question is typical for Data Scientist internship interviews in machine learning at a fundamentals-to-intermediate level because it checks both the theory and the practical concerns involved in using non-parametric algorithms and linear feature extraction on real data.
This question is drawn from a Data Scientist interview experience.