IBM · ML & AI Fundamentals
Describe building statistical vs ML models
TrueInterview
October 7, 2026 · 4 min read
Imagine you are interviewing for a Data Scientist internship with a marketing analytics group. Describe a project in which you built both (a) a statistical model—for example, linear or logistic regression or a GLM—and (b) a machine learning model, such as a tree-based method, boosting, or a neural network. Your answer should address:
- The business problem and the decision the model was meant to inform.
- How you defined the target or label and what counted as “success.”
- Which features you used—behavioral, demographic/firmographic, marketing touchpoints, time-based, text, and so on—and why.
- How you managed stakeholder expectations: did they care only about predictive accuracy, or also about interpretability, meaning which features mattered and why?
- What you would change next time, including data quality, leakage, monitoring, deployment, fairness, and similar concerns.
Summary: This question tests a Data Scientist's ability to build and compare statistical and machine learning models, covering feature engineering, target definition, interpretability, stakeholder communication, and operational issues such as deployment, monitoring, data leakage, and fairness in a marketing analytics setting.
Solution A strong response reads like a short case study and directly compares statistical models with ML models on assumptions, interpretability, and operationalization.
1) Establish the business decision
- Begin with: Who will use the model, and what action will change because of it?
- Example: “Marketing operations uses the score to send leads either to SDRs or to nurture emails; because sales capacity is limited, we need high precision near the top of the ranked list.”
- Specify the unit of analysis—lead, account, or person—and the scoring cadence, such as daily or weekly.
2) Define the label and time window to avoid leakage
- Make the label actionable and bounded in time:
- Example labels:
converted_within_30_days_of_signuporopportunity_created_within_14_days_of_MQL.
- Example labels:
- Confirm that all features are calculated as of the scoring moment, with no signals from after conversion.
- Explain how you dealt with delayed outcomes or right-censoring, for instance by excluding recent leads or applying survival methods.
3) Compare statistical and ML models, and when to use each
Statistical model, such as logistic regression or a GLM:
- Strengths: interpretable coefficients or odds ratios, easier calibration, simpler monitoring, and often more robust with limited data.
- Weaknesses: assumes linearity or additivity, and requires manual feature engineering for interactions and nonlinear relationships. ML model, such as XGBoost or LightGBM:
- Strengths: captures nonlinearities and interactions, delivers strong ranking performance, and handles mixed feature types.
- Weaknesses: less transparent, needs more tuning, can overfit or leak, and may require post-processing for calibration. A useful narrative is: “I began with logistic regression as an interpretable baseline to build stakeholder trust, then moved to gradient boosting to improve lift in the top decile, and used SHAP plus partial dependence plots to explain the main drivers.”
4) Feature examples and why they matter
Include these categories and explain their value:
- Behavioral: sessions, key events, and time since last activity, which signal intent.
- Acquisition: channel or campaign and UTM tags, which reflect marketing efficiency.
- Firmographic: company size, industry, and region, which indicate fit.
- Product usage: feature adoption and activation milestones, which act as product-qualified signals.
- Temporal: day of week and seasonality, plus recency and frequency. Also mention safeguards:
- Remove label proxies that are created only after the outcome, such as “sales contacted” when that contact occurs because the lead already had a high score.
- Handle missing values deliberately, for example with an explicit “unknown” category or missing indicators.
5) Evaluation: match metrics to the use case
For marketing lead scoring, ranking is usually the main concern:
- Offline metrics: AUC-ROC, AUC-PR when conversion is rare, log loss, and Brier score.
- Business or ranking metrics: lift in the top decile, precision@K, recall@K, and expected conversions given SDR capacity.
- Calibration: use reliability curves, and if the score is treated as a probability, make sure it is calibrated. Give a simple example:
- “If SDRs can call 500 leads per week, I would optimize precision@500 and lift compared with random selection.”
6) Interpretability versus “only results” in stakeholder management
Show that you can adjust:
- If stakeholders care only about outcomes, emphasize lift, incremental conversions, and operational limits.
- If they need to understand drivers, provide:
- global importance from gain or SHAP,
- local explanations for an individual lead,
- stable narratives such as “high intent usage plus target industry drives the score.” Important nuance: feature importance is not the same as causal impact; say that you would validate with experiments, for example testing whether contacting leads earlier actually causes more conversions.
7) Deployment, monitoring, and iteration
Show practical data science maturity:
- Keep training and serving consistent, and schedule retraining.
- Monitor feature distribution drift, score drift, and calibration drift.
- Track business KPIs after launch, such as conversion, revenue, and sales efficiency, along with guardrails like spam complaints and fairness concerns.
8) End with “what I would do differently”
High-signal improvements:
- Define the label more tightly around revenue, such as SQL → Opportunity → Closed Won.
- Use stronger validation with time-based splits.
- Address selection bias, since sales touches are not random.
- Add causal tests for interventions, such as call versus email versus holdout.