ByteDance · ML & AI Fundamentals
How do you choose a classification threshold?
TrueInterview
October 7, 2026 · 1 min read
Context
You have trained a binary sentiment classifier (for example, positive versus negative) and now need to deploy it in a product where downstream actions are driven by the model’s predictions.
Questions
- Describe your ML pipeline from start to finish:
- How you sourced and labeled data, and how you built the dataset (train/validation/test splits).
- Feature engineering or model selection (e.g., TF-IDF plus a linear model versus a transformer).
- Training process and evaluation configuration.
- Main practical obstacles (class imbalance, noisy labels, distribution shift) and what you did about them.
- Modeling decisions:
- Why did you pick method/model X instead of the alternatives?
- What assumptions does it rely on, and what trade-offs does it create (latency, interpretability, cost, robustness)?
- Iterative improvement:
- Explain how you improved the system in cycles (for example, error analysis → new features/data → retrain → re-evaluate).
- What were your most important lessons, and what would you change next time?
- Choosing a threshold for production:
- The model emits a probability score. How do you select the decision threshold?
- Which metrics would you look at (precision, recall, F1, ROC-AUC, PR-AUC, cost-weighted loss), and which would serve as primary vs. diagnostic vs. guardrail?
- How does your answer change with:
- Severe class imbalance
- Different costs for false positives versus false negatives
- A fixed review/operations capacity (for example, only 1,000 items per day can be escalated)
- Defining the metric:
- If a stakeholder suggests treating metric XXX as the definition of “success,” how do you assess whether that definition is suitable?
- What data problems (label leakage, delayed labels, sampling bias) could make the metric misleading? Overview: This question assesses a data scientist’s ability in binary classification threshold selection, end-to-end ML pipeline design, model evaluation, and operational trade-offs such as class imbalance, asymmetric error costs, and limited escalation capacity, in the Machine Learning area of model evaluation and deployment.
Loading comments…