Citadel · Project Deep Dive
EQR Alpha-Factor Research Deep-Dive + LLM Inference Stability
TrueInterview
July 18, 2026 · 2 min read
Format
This round centered on a resume review and case study and was conducted mainly through verbal discussion. The reported follow-up items were:
Items covered in the assessment
- Q1 — (no title) (non-coding item appearing only in this overview)
- Q2 — (no title) (non-coding item appearing only in this overview)
- Q3 — (no title) (non-coding item appearing only in this overview)
- Q4 — (no title) (non-coding item appearing only in this overview)
- Q5 — (no title) (non-coding item appearing only in this overview)
- Q6 — Case analysis — stability of LLM inference
- Q7 — Final live-coding task
Observations
- Detecting look-ahead bias: standard safeguards include point-in-time dataset snapshots, delayed feature joins enforced by
as_of_timestamp <= signal_timestamp, and adversarial reproducibility tests in which rerunning a strategy for the same as-of date must yield the same signals regardless of the actual execution date. State at least one structural safeguard aloud rather than merely saying that you verify things carefully. - IC and ICIR cutoffs: IC measures the cross-sectional rank correlation between a factor and subsequent returns, while annualized ICIR is
mean(IC) / std(IC). Common working benchmarks are|IC| > 0.02-0.05per period at daily horizons andICIR > 0.5for keeping a factor. Clearly specify the normalization used, including rank versus raw values and period-level versus cumulative measurement. - LLM stability prompt: treat observed labels as noisy estimates of a latent true signal. Given
Var(noise_per_day) = σ^2and true-signal variance ,corr(y_day, ground_truth) = τ / sqrt(τ^2 + σ^2). For two independent daily samples, their correlation is . Taking the mean acrossNindependent days divides the noise variance byN, producingcorr(y_avg, ground_truth)^2 = τ^2 / (τ^2 + σ^2 / N). FindNfrom ; starting from 0.95 gives , or roughly 16 days. - Extending a planned 45-minute session to 75 minutes was a warning sign in the reported result: the candidate was declined even after addressing most of the questions. Evaluation emphasized the rigor of the derivation rather than only the numerical conclusions, so articulate each modeling assumption as you proceed.
- The concluding programming exercise creates a sudden shift away from the discussion-focused tempo. Reserve 15 minutes for it, even when the case-study follow-ups are taking longer than expected.
Preparation
- Develop a five-to-seven-minute, end-to-end explanation of one alpha-research strategy that covers the hypothesis, dataset, factor design, look-ahead protections, IC and ICIR evaluation, hyperparameter selection, overfitting defenses such as time-based cross-validation, regularization, and ensemble averaging, plus attribution of live PnL.
- Practice the noisy-measurement and signal-to-noise setup with small examples until you can mentally determine how many independent observations are needed to shrink noise by a factor of X. The LLM stability problem is one application of this structure, which also appears in many other quantitative case studies.
- Review established methods for limiting overfitting in quantitative research: walk-forward validation, chronological train/test partitioning, the deflated Sharpe ratio from Bailey and López de Prado, and purged k-fold cross-validation. Be prepared to identify at least two by name.
- Prepare the wildcard-matching implementation in advance, beginning with recursion and memoization and then using an iterative two-pointer method for the follow-up, so the final coding task does not consume time intended for the case study.
Loading comments…