Capital One · Statistics & Data Analysis
Use data to resolve an ambiguous problem
TrueInterview
October 7, 2026 · 5 min read
Describe a situation where you tackled an ambiguous business problem from start to finish using data. State the precise hypothesis, the data sources along with their known biases, the statistical approach or model you chose, and how you checked your assumptions. Measure the impact using a credible counterfactual (for instance, difference-in-differences rather than a simple before–after comparison). Also cover how you dealt with data quality problems, how you conveyed uncertainty to stakeholders, and one improvement you would make if you had 10% more data.
Overview: This question assesses a data scientist's ability to perform complete data analysis, identify causal effects, build statistical models, manage data quality and governance, evaluate biases, and communicate uncertainty while working on a vague business issue.
Solution
Example answer (STAR + causal rigor)
Situation (poorly defined problem)
Over several months, delinquencies on a revolving credit product had been gradually rising. Product leadership wanted to know: "Could we encourage customers to sign up for AutoPay to cut down on missed payments without raising risk or hurting the customer experience?" The issue was vague: there were many possible levers (UI, timing, incentives), the target segment was unclear, and macroeconomic seasonality made before–after comparisons unreliable.
Hypothesis (testable)
Displaying an in-app prompt at the moment customers view their bill to those who are eligible will boost AutoPay enrollment within 14 days and lower 60-day delinquency, with no negative impact on charge-offs or complaints. Primary outcome: AutoPay enrollment within 14 days. Secondary outcomes: 60-day delinquency (DPD60), charge-off rate, and customer complaints within 30 days.
Data sources and known biases
- App event logs (screen views, prompt impressions/clicks)
- Biases: users who only use the app differ from web or phone users (selection bias); duplicate session-level records; timezone inconsistencies.
- Payments and statement data (due dates, posted payments, late fees)
- Biases: end-of-month seasonality; partial payments; timing of cutoffs.
- Risk and bureau attributes (internal risk score, external score bands)
- Biases: how often scores refresh; missing values for new accounts; regulatory limits on usage.
- Customer profile and eligibility flags (AutoPay availability, account tenure)
- Biases: survivorship bias because closed accounts are excluded; eligibility can change during the experiment.
- Marketing exposures (email/SMS pushes)
- Biases: interference across channels; incomplete logging on some older campaigns. Mitigations included stratified randomization, logging exposures, normalizing timezones, and pre-registering metrics and windows.
Method and design
- Causal strategy: a randomized controlled experiment for the prompt, and difference-in-differences (DiD) for delinquency to control for time-based shocks.
- Unit of randomization: customer rather than session to prevent spillover across sessions; 50/50 split.
- Stratification variables: risk band, tenure, region, and due-date week to improve balance and statistical power.
- Sample: 200k eligible customers (100k treatment, 100k control) over one billing cycle; power analysis aimed for a minimum detectable effect of 1.0 percentage point on AutoPay with 90% power.
- Estimation:
- AutoPay: difference in proportions using intent-to-treat, with stratification covariates included in a logistic regression for added precision.
- DPD60: two-period DiD on a customer-level panel (pre = prior 2 months, post = experiment month), estimated via OLS with customer and time fixed effects and cluster-robust standard errors by customer.
- Heterogeneous effects (for rollout targeting): uplift modeling using causal forests on treatment-by-feature interactions. Why not a naive before–after? Seasonality and macroeconomic trends (such as tax season) substantially shift delinquency; DiD handles this by using a contemporaneous control trend.
Assumption checks and validation
- Randomization balance: standardized mean differences across more than 20 covariates were all below 0.05.
- SUTVA/interference: holdout users never saw the prompt; cross-device exposures were monitored; contamination was under 1%.
- Parallel trends (DiD): pre-period DPD60 trends over 3 months had slope differences statistically indistinguishable from zero; a placebo DiD on pre-period months showed no effect.
- Overlap: all risk bands were represented in both groups because of stratification.
- Model diagnostics: the logistic model was well-calibrated (Brier score 0.14) with no high VIFs; the causal forest out-of-bag uplift AUC was 0.62.
Data quality issues and fixes
- User ID stitching: app and core systems occasionally had ID mismatches. We fixed this with deterministic keys plus a fuzzy match fallback, then deduplicated; unit tests were added to ensure a 1:1 mapping.
- Timezones/cutoffs: all event timestamps were unified to UTC and aligned with statement cutoff windows; daily snapshots were precomputed to avoid bias from late-arriving data.
- Missing bureau scores: 6% were missing; we imputed a "missing" category for modeling (but not for treatment assignment) and ran a sensitivity analysis excluding these users—the effects were consistent.
- Exposure logging gaps: legacy email campaigns could confound results; we paused overlapping campaigns for the sample and logged any exceptions.
Results with a defensible counterfactual
- AutoPay enrollment (14-day):
- Treatment: 27.1% (27,100/100,000)
- Control: 24.0% (24,000/100,000)
- Uplift: +3.1 percentage points (pp)
- 95% CI: ±0.4 pp (SE ≈ 0.20 pp), p < 1e-10
- DPD60 (difference-in-differences):
- Pre DPD60: both groups ≈ 5.0%
- Post DPD60: control 5.6%, treatment 4.8%
- DiD estimate:
- 95% CI: [−1.1 pp, −0.5 pp] (cluster-robust) Why DiD matters: A naive before–after using only the treatment group would have shown −0.2 pp, overlooking the macroeconomic increase visible in the control group (+0.6 pp). DiD recovers the true causal effect of −0.8 pp. Back-of-the-envelope impact (illustrative, using expected credit loss ECL = PD × LGD × EAD):
- Eligible monthly population at rollout: 2.0M customers.
- Fewer DPD60 accounts: per month.
- Assumptions: conditional charge-off probability 30%; LGD 80%; average EAD $1,500.
- Monthly ECL reduction ≈ ; annualized ≈ $69M.
- Late fee revenue loss from fewer late payments: ≈ $0.7M/month (based on historical fee incidence), net ≈ $5.1M/month. Sensitivity: Changing the charge-off probability by ±5 pp shifts the annualized benefit by about ±$11M; we showed a tornado chart covering the main drivers.
Communication of uncertainty and decision
- Reported point estimates along with 95% confidence intervals and minimum detectable effects; stressed intent-to-treat estimates to mirror real-world adherence.
- Presented the trade-offs: customer benefit (fewer fees), risk reduction, and revenue impact; ensured alignment with compliance and customer fairness.
- Decision: full rollout to all eligible users, with targeted prioritization based on the uplift model for high-impact segments; guardrails were set on complaint rate and any adverse risk shifts.
One improvement with +10% more data
- Improve heterogeneity estimation: With 10% more users, we could estimate segment-level uplift more precisely (especially in sparse strata such as new-to-credit), allowing a stricter treatment policy (for example, targeting only the top decile of uplift), which simulations indicate would add about 10–15% incremental ECL reduction for the same exposure volume.
Key formulas (for clarity)
- Difference-in-Differences:
- Intent-to-Treat uplift on AutoPay:
Pitfalls and how we avoided them
- Naive before–after: we used DiD and checked pre-trends.
- Contamination: we enforced a holdout and monitored cross-channel exposures.
- P-hacking: we pre-registered metrics and windows and did not peek at interim results.
- Overfitting uplift: we used out-of-bag validation and monotonicity checks, and set conservative thresholds for rollout.