Citadel · Project Deep Dive
Discuss PhD coursework and research impact
TrueInterview
October 7, 2026 · 5 min read
Talk me through the courses you selected during your PhD and the research you did. Which two classes had the greatest influence on how you do empirical modeling, and why? Tell me about a research project whose first approach did not work — what did you change once you got feedback, and how did you measure the impact (for instance, through ablation, replication, or external validation)? Suppose I asked your advisor and one of your collaborators to name a single area you should improve — what would each of them say, and what steps have you taken in response?
Overview: This question probes how solid a candidate's empirical modeling fundamentals are, whether they can learn from an approach that failed, how well they absorb feedback, and how capably they measure the impact of research in either a research or a product setting.
Solution
Structuring a strong answer (2–3 minutes)
- Begin with a single sentence that states your PhD focus and the modeling areas you care about.
- Coursework: pick two courses and, for each one, say which principle you took from it and how it altered the way you model.
- Research failure: give a brief STAR account (Situation, Task, Action, Result) that includes quantified impact and validation work (ablation, replication, external).
- Feedback: one improvement area from your advisor and one from a collaborator; close with the specific actions you put in place.
Selecting the two courses (and what to say about them)
Choose courses that plausibly build empirical rigor. Some examples and the lessons they carry:
- Bayesian Data Analysis: prior elicitation, posterior predictive checks, and uncertainty quantification; the working habit of examining calibration and coverage rather than only point metrics.
- Causal Inference: identification versus prediction, DAGs, difference-in-differences, instrumental variables; the habit of specifying estimands, protecting against confounding, and choosing validation that matches the question.
- Statistical Learning / Regularization: the bias-variance tradeoff, cross-validation, regularization paths; the habit of nested CV, early stopping, and ablating groups of features.
- Time Series / Panel Methods: avoiding leakage, rolling splits, non-stationarity; the habit of out-of-time validation and stability checks across different regimes.
A phrasing template:
- Course name → principle → change in behavior.
- For example: Causal Inference showed me how to keep identification separate from prediction; these days I build features and targets that match the estimand and validate using out-of-time and subgroup checks.
A research project whose first approach failed (STAR with quantification)
An example you can adapt (swap in your own domain and figures, but keep the structure):
- Situation: I investigated whether alternative app-usage signals could forecast quarterly earnings surprises.
- Task: Build a model that classifies surprise versus no surprise and evaluate its economic value.
- Action (first approach and its failure): I began with an end-to-end LSTM trained on raw daily signals. It performed well under random CV () but fell apart in a rolling time split (). Feedback pointed to leakage (lookahead inside the feature windows), a poorly defined target, and overfitting.
- Action (following the feedback):
- Redefined the target so it is public at , and lagged every feature by at least 7 days to remove lookahead.
- Moved to a transparent baseline (regularized logistic regression plus gradient boosting) using engineered weekly aggregates and seasonality controls.
- Used rolling-window hyperparameter tuning together with out-of-time holdouts.
- Added domain controls (sector, size, prior momentum) to cut down spurious correlations.
- Result (quantified impact):
- Predictive: AUC rose from 0.53 to 0.62 on a 4-quarter holdout; the Brier score fell by 11%; the calibration slope was about 0.97.
- Ablations: dropping the alternative data lowered AUC by 0.06; dropping the seasonality controls lowered AUC by 0.02, which isolated the components that mattered.
- Replication: comparable gains across 3 sectors and an international sample, with .
- External validation: a model trained on 2017–2020 generalized to 2021–2022 with an AUC of 0.60; the long-short backtest IR improved from 0.30 to 0.78 once transaction costs, turnover constraints, and block-bootstrap confidence intervals were included.
What to emphasize:
- The exact failure mode (leakage, a misspecified target, overfitting) and the safeguards you introduced (time-based splits, lagging, calibration checks).
- Quantified impact, and which kinds of validation you relied on (ablation to attribute the gains, replication across subgroups, an external out-of-time sample).
Advisor and collaborator feedback (an area to improve)
Choose two complementary angles along with the actions you took.
- What your advisor would probably say (depth and rigor): a tendency to reach for complex models too early.
- Actions: follow a modeling ladder (dummy/baseline → linear → tree → deep), write pre-analysis plans, and run through a checklist for identification and leakage before any tuning.
- What a collaborator would probably say (engineering and communication): code modularity and reproducibility, or clear communication of uncertainty to people outside the field.
- Actions: unit tests for feature pipelines, versioning of seeds and data, reproducible environments; one-slide experiment templates covering the problem, metric, and decision rule; practice at explaining calibrated probabilities and effect sizes.
Pitfalls and guardrails to call out explicitly
- Avoid leakage: strict temporal splits, lagged features, and no peeking across folds.
- Prevent p-hacking: define metrics and decision thresholds in advance; adjust for multiple comparisons when many features are explored.
- Robust validation: nested or rolling CV; keep one final holdout untouched; report calibration and stability across regimes and subgroups.
- Reproducibility: fixed seeds, data versioning, and code review.
A compact sample answer (adapt it with your own details)
My PhD focused on empirical modeling for economic time series. Two courses shaped how I work. Causal Inference showed me to keep identification apart from prediction; I now define estimands first, build targets and features that avoid confounding, and validate with out-of-time and subgroup checks. Bayesian Data Analysis put uncertainty front and center: I rely on prior elicitation and posterior predictive checks, and I always report calibration rather than accuracy alone. In one project that predicted earnings surprises from app-usage data, my initial LSTM looked strong under random CV but broke down in a rolling split (). Feedback exposed lookahead and a loosely defined target. I lagged every feature by 7+ days, sharpened the target, switched to a regularized logistic baseline and gradient boosting with weekly aggregates and seasonality controls, and tuned within rolling windows. AUC climbed to 0.62 on a 4-quarter holdout, the Brier score improved by 11%, and the calibration slope was 0.97. Ablations showed the alternative data added +0.06 AUC; replication across sectors and an international sample preserved the gains; an out-of-time 2021–2022 test reached an AUC of 0.60, with a transaction-cost-aware backtest IR rising from 0.30 to 0.78. My advisor would say that I sometimes reach for complex models too quickly. I now work through a modeling ladder and write pre-analysis plans. A close collaborator would point to reproducibility. I set up unit tests for feature code, data versioning, and standardized experiment reports so results are easy to audit and re-run.