Cvs Health · Production Troubleshooting
Lead structured response to accuracy incident
TrueInterview
October 7, 2026 · 9 min read
You lead the on-call rotation for a payment-accuracy team. At 09:00 word arrives that, since 00:00 today, the share of incorrect payments among CA users aged 18–24 has climbed from a long-run 2.0% to 3.1% across roughly 120k transactions; finance is escalating and expects a mitigation inside 24 hours. You can draw on feature logs, model scores, recent code and deployment diffs, and a sampling tool that pulls raw cases for audit.
Lay out a structured, time-boxed plan stating exactly how you would: 1) triage and size the impact (owners, metrics, and the acceptance criteria for an 'all clear'), 2) build testable hypotheses and run the smallest experiments/queries that isolate the root cause (data slices, holdouts, backfills), 3) choose a safe mitigation (a threshold change, a feature hotfix, a rule fallback, and so on), with rollback triggers and predicted side effects (the precision/recall trade-off), 4) report status to executives and partner teams at named times with concrete artifacts (dashboards, runbooks), and 5) stop it recurring (observability, guardrails, pre-deploy checks). Spell out which risks you accept versus which you refuse, and how you would judge success by end of day.
Overview: This question probes a data scientist's incident leadership, operational judgement, and data-driven debugging, with the focus on triage, hypothesis formation, mitigation choice, cross-functional communication, and recurrence prevention.
Solution
Assumptions and Notation
- CA means California — confirm this immediately; if it is Canada, redo every slice keyed on
country_region. - The score is the model's estimated probability that a payment is incorrect.
- Decision rule: (where is the threshold) sends the payment to review/decline, while auto-approves it. Lowering catches more potentially incorrect payments — the incorrect rate drops while friction rises.
- Cohort volume observed since 00:00 is for CA ages 18–24 — validate this; if 120k is total site volume, recompute for the cohort.
- Incorrect rate (IR) is .
0. Timeline Overview (Day-of Plan)
- 09:00 Detect and page. Open an incident channel and ticket, then assign roles.
- 09:05–09:30 Rapid triage: confirm the anomaly and pin down who/what/when. Freeze risky deploys. Stand up a live dashboard by cohort.
- 09:30–10:30 Root-cause probes: config/threshold, model version, feature missingness and drift, label pipeline, traffic mix. Pull 100–200 raw cases for audit.
- 10:30–11:00 Decision checkpoint 1: if a clear cause has surfaced (config/model), roll back. If it has not, ship a gated mitigation (tighten the threshold for CA 18–24 on a 25% canary) with guardrails.
- 11:00–14:00 Keep experimenting (shadow scoring, backfills, merchant/device slices). Adjust the mitigation as evidence arrives.
- 14:00 Decision checkpoint 2: expand or roll back the mitigation based on the metrics.
- 16:00 Draft the postmortem outline; finalize prevention actions and owners.
- 17:30 End-of-day (EOD) report: results, metrics, next steps.
1) Triage and Quantify Impact
Owners (assign in incident ticket)
- Incident commander (IC): you, the on-call DS — decision maker and timekeeper.
- ML engineer: model configs, thresholds, rollbacks.
- Data engineer: feature pipeline and data quality.
- Analyst/BI: dashboards, SQL queries, confidence intervals.
- Risk Ops: manual audit sampling and case categorization.
- Finance liaison: translating impact into dollars and volume, plus stakeholder updates.
Metrics to compute now
IR_18_24_CA(t): rolling 15-minute incorrect rate for CA ages 18–24.- Baselines: the last 4 weeks, matched on day-of-week and time-of-day, giving a mean and bands.
- Volume: transactions per 15 minutes.
- Confusion components (approximated in near real time):
- Block/review rate (BR): the share routed to review/decline.
- Auto-approve rate (AR): the share approved immediately.
- If labels arrive quickly: false approvals (incorrect among auto-approvals) and false reviews (correct among reviewed).
- Feature health: null rate, staleness, and distribution shift for the top-20 features.
- Model health: score distribution per cohort; KS statistic versus baseline; PSI for features.
- Label pipeline health: label latency distribution and completeness.
Quick significance check (example)
- Observed on . The 95% CI is roughly , i.e. (3.0%, 3.2%). Against the historical 2.0% baseline that is highly significant.
- Counts: at 2.0% you would expect 2,400 incorrect payments; you observe 3,720, a gap of about +1,320.
Acceptance criteria for 'all clear'
Any one of:
IR_18_24_CAsettles at or below 2.3% (no more than +15% relative to the 2.0% baseline) across three consecutive 30-minute windows with stable volume; ORIR_18_24_CAsits inside the baseline 95% control limits with no active anomaly flags on score or feature drift; AND- no new anomalies appear in adjacent cohorts (other ages in CA, 18–24 in other states); AND
- the manual review queue stays under 80% capacity at the 95th-percentile SLA.
2) Hypotheses and Minimum Experiments/Queries
Prioritize checks that are fast, discriminative, and reversible.
H1. Threshold/config drift or bad rollout
- Check: diff the current production threshold , rule weights, and
model_versionagainst the last stable release — config files, env vars, feature flags, and change logs since 00:00. - Query:
SELECT distinct model_version, threshold, count(*) FROM events WHERE ts >= today AND cohort = 'CA_18_24' GROUP BY 1,2; - Action if true: roll back immediately to the last known-good config/model.
H2. Feature pipeline issue (missingness/staleness/time-zone)
- Check: for the top features by importance, compare null rate, timestamp skew, and PSI.
- Query (sketch):
- Feature nulls:
SELECT feature_name, avg(is_null(value)) null_rate FROM feature_logs WHERE ts >= today AND cohort='CA_18_24' GROUP BY 1 ORDER BY null_rate DESC LIMIT 20; - Staleness:
SELECT avg(ts_event - ts_feature) FROM feature_logs WHERE cohort='CA_18_24';
- Feature nulls:
- Experiment: re-score a 1k sample from today using yesterday's pipeline artifact (a backfill) and compare the scores; a large divergence signals a pipeline bug.
H3. Labeling pipeline change (definition/latency)
- Check: label arrival latency and label completeness today versus earlier days.
- Query:
SELECT percentile(latency, [50,90,99]), completeness FROM labels WHERE ts >= today AND cohort='CA_18_24'; - If labels have slowed, the observed IR may be inflated or biased. Validate it against a manual audit sample.
H4. Traffic mix shift (merchant, device, payment method, promo)
- Check: contribution analysis by
merchant_id,device_os,payment_method, BIN, referrer. - Query:
SELECT merchant_id, count(*) n, avg(is_incorrect) ir FROM tx WHERE ts>=today AND cohort='CA_18_24' GROUP BY 1 ORDER BY ir DESC LIMIT 20; - Experiment: if 1–2 merchants drive the delta, scope the mitigation to them.
H5. Geography/age parsing bug (CA vs CA-Province; age boundary at midnight)
- Check: validate the geo parser inputs; compare the distribution of state/province codes; examine the 17–18 and 24–25 age edges around midnight.
- Query:
SELECT state, country, count(*) FROM tx WHERE ts>=today AND age BETWEEN 18 AND 24 GROUP BY 1,2;
H6. Score distribution drift/miscalibration
- Check: compare the score histogram with baseline; run a KS test; plot calibration curves if labels are available.
- Experiment: shadow-score today's features with the previous model; if the shadow matches baseline while production drifts, a model or feature change is the likely cause.
H7. Third-party dependency degradation
- Check: timeouts/error rates for external services (device fingerprint, address verification). When fallbacks are applied, feature quality degrades.
- Query:
SELECT service, error_rate FROM svc_metrics WHERE ts>=today AND cohort='CA_18_24';
Minimum manual audit
- Pull 100–200 incorrect cases (and 100 near-threshold approvals) for CA 18–24.
- Tag root-cause categories: feature missing, merchant anomaly, time zone, data mapping, genuinely new behaviour. This steers the mitigation choice.
3) Safe Mitigation Decision, Rollback Triggers, Side Effects
Decide by 10:30–11:00, even if the root cause is not fully confirmed. Use a canary and guardrails.
Candidate mitigations (ordered by safety)
- Roll back config/model to last known-good
- Use if any diff in threshold/model/rules shows up since 00:00.
- Predicted effect: the prior precision/recall is restored. Validate over a 30–60 min window.
- Rollback trigger: N/A; this step is the rollback.
- Targeted threshold tightening for CA 18–24 only (lower )
- Goal: fewer incorrect approvals (better recall) at the cost of a higher review/false-positive rate.
- Deployment: 25% canary of CA 18–24 for 60 minutes; then ramp to 100% if it improves and capacity holds.
- Guardrails:
- Manual review queue length below 80% capacity; 95th-percentile review SLA under target.
- IR improvement of at least 20% relative within 60–90 minutes versus control.
- Side effects (quantify them): simulate using yesterday's ROC/PR. Example:
- At : TPR = 0.72, FPR = 0.04.
- At : TPR = 0.80 (+8 pp), FPR = 0.055 (+1.5 pp).
- With today's cohort volume around 120k and incorrect prevalence around 3.1%:
- Added reviews ≈ extra reviews.
- Incorrect approvals reduced ≈ fewer incorrect approvals.
- Make sure Ops can absorb +1.7k reviews across those hours.
- Rollback triggers:
- Review SLA breach above 2x baseline for 2 consecutive 15-minute windows.
- No IR improvement of at least 10% relative after 90 minutes.
- Feature hotfix: drop degraded features and re-weight via rules fallback
- Use if one or a few features show over 30% missingness or heavy drift at PSI > 0.25.
- Implementation: temporarily exclude the affected features; enable rule-based overrides (e.g. merchant block list, stricter AVS-mismatch handling) for CA 18–24.
- Side effects: loss of model discrimination; expect more false positives; monitor BR and SLA.
- Merchant-targeted policy
- If 1–2 merchants drive the majority of the delta, add temporary stricter rules for those merchants (e.g. lower only for those IDs, or route them to review).
Measuring mitigation effect quickly
- Use an interleaved A/B (canary versus control within CA 18–24) with 15-minute metrics and CIs.
- Primary: IR among auto-approveds. Secondary: total IR, review queue metrics, customer impact (approval rate).
- If labels are delayed, use proxies: the reduction in high-risk approvals near the threshold and audit sampling outcomes.
4) Communication Plan and Artifacts
Cadence
- 09:15 Initial incident notification (execs + finance + partner teams)
- What: anomaly detected; scope (cohort, magnitude), immediate actions (deploy freeze, triage), next update at 09:45.
- 09:45 Triage readout
- Share the dashboard link, the IR time series with CIs, top slices, preliminary hypotheses, and owner assignments.
- 11:00 Decision checkpoint 1
- Announce the chosen mitigation (rollback/threshold canary), guardrails, expected side effects, and when to expect impact (the next 60–90 min).
- 13:00 Midday update
- Results of the canary/experiments, whether to ramp or adjust; any merchant-specific actions.
- 15:00 Decision checkpoint 2
- Finalize the mitigation state (ramp to 100% or revert), outline prevention work.
- 17:30 EOD report
- Metrics versus baseline, incident timeline, root cause (or leading hypothesis), actions taken, next-day items.
Artifacts to produce/share
- Live dashboard (Looker/Superset):
- IR by 15-min for CA 18–24 versus baseline; adjacent cohorts for a regression-to-mean check.
- Score distribution (KS), feature nulls/staleness, label latency.
- Review queue and SLA utilization.
- Runbook page: mitigation levers, thresholds, rollback commands, owners, and escalation paths.
- JIRA/incident ticket: timeline, decisions, diffs, SQL links, shadow/backfill results.
- One-pager for execs: problem, impact ($/volume), mitigation, risk trade-offs, ETA for steady state.
5) Prevent Recurrence
Observability and Guardrails
- Segment monitors: real-time IR for key cohorts (age, geo, merchant), with control charts and anomaly alerts.
- Model monitors: score drift (KS), calibration drift (Brier/NLL), feature PSI and missingness alerts, with auto-fallback to rules once thresholds are exceeded.
- Labeling monitors: label latency/completeness SLOs; alert on shifts.
- Capacity monitors: review queue SLA and auto-throttling of the mitigation to avoid overload.
Pre-deploy checks
- Canary by cohort: 10–20% rollout with guardrail checks (IR, approval rate, SLA) and auto-rollback on breach.
- Config integrity: checksums/signature and diff approvals; no same-day threshold change without a canary.
- Data contracts and schema tests for upstream features; time-zone and geo parsing unit tests.
- Backtest gate: require day-of-week and cohort-sliced validation over the last 2–4 weeks before promoting a model/config.
Process
- Blameless postmortem within 48 hours with concrete action items and due dates.
- Expand runbook coverage; add decision trees for common root causes (feature missingness, merchant surge, label delays).
Risks: Accepted vs Avoided
- Accepted: temporary higher friction (more reviews/false positives) within Ops capacity; localized mitigations (CA 18–24 and/or specific merchants); short-term revenue impact to protect accuracy.
- Avoided: broad global threshold changes without a canary; untested feature code changes; changes that exceed Ops capacity or breach SLAs system-wide; permanent policy shifts without analysis.
Success by End of Day
- Quantitative:
IR_18_24_CAreduced by at least 20% relative from peak and within baseline tolerance bands for 90 minutes; OR explained by a verified label/measurement artifact and mitigated.- No new anomalies in adjacent cohorts; review SLA within target; approval rate impact at or below the agreed guardrail (e.g. a drop of ≤ 2 pp).
- Qualitative:
- Root cause identified, or the top hypothesis validated with evidence; mitigation in place; rollback plan documented.
- Dashboards, runbook updates, and postmortem draft complete; owners assigned for prevention items.
Example SQL/Analysis Snippets (Sketches)
- IR by cohort/time:
SELECT time_bucket('15 min', ts) t, count(*) n, avg(is_incorrect::int) ir FROM tx WHERE ts>=today AND geo='CA' AND age BETW…[truncated]…