Capital One · Product & Business Case
Critique a product interface and propose fixes
TrueInterview
October 7, 2026 · 6 min read
Choose a digital product that you interact with every day. Point to a single interface you appreciate and a single capability you don't like. Drawing on usability heuristics (Nielsen, for instance) and accessibility criteria (WCAG, for example), explain what makes each one work or break down. Then sketch a redesign for the part you dislike, specify success measures and guardrails, and describe an experiment or telemetry approach for validating the change without being fooled by novelty effects.
Overview: This prompt gauges a candidate's product sense and usability evaluation ability — in particular how they apply heuristics and accessibility standards, plus their skill at proposing a redesign, setting success metrics, and planning experiments or telemetry.
Solution
Worked Example (Product: Gmail on Web)
1) Selected Interfaces
- Interface liked: Smart Compose — the inline autocomplete that appears as you compose a message.
- Function disliked: building filters out of the search box (advanced search → Create filter → pick actions).
2) Reasons the liked interface works
- Nielsen heuristics
- Recognition over recall: options show up inline, so nobody has to dredge up whole phrases from memory.
- Flexibility and efficiency: experienced users take the suggestion with Tab, while beginners can simply disregard it.
- User control and freedom: dismissing it is trivial, and no text is locked in until the user accepts.
- Aesthetic and minimalist design: the suggestion is a faint gray, staying out of the way.
- Visibility of system status: feedback arrives the moment you type.
- Accessibility (WCAG)
- 2.1.1 Keyboard: both accepting and dismissing work from the keyboard (Tab/Esc).
- 4.1.3 Status Messages: the arrival of a suggestion ought to be announced to screen readers without interrupting (
aria-live="polite"). - 1.4.1 Use of Color: color alone must not carry the meaning; the styling shift is distinguishable by other means.
- Possible gaps: 1.4.3 Contrast (Minimum) if the gray is too faint for some people; the focus and accept cues need to remain perceivable.
3) Reasons the disliked function breaks down
- Summary of problems: the route to filter creation sits behind a tiny icon; advanced syntax must be remembered; rules can easily end up too broad; and previewing is minimal.
- Nielsen heuristics
- Recognition over recall (breached): people have to keep operators such as from:, has:attachment, older_than: in their heads.
- Visibility of system status (weak): there's little live preview or sense of impact — for instance, how many messages will match?
- Error prevention (weak): it's easy to build sweeping filters that archive far more than intended.
- Match between system and real world: the labels and actions are system-flavored ("skip inbox") rather than stated as plain outcomes.
- Help and documentation: hints are present, yet they aren't woven into the moment of need.
- Accessibility (WCAG)
- 2.5.5 Target Size (2.2): the small hit areas, like the search options icon, are difficult to click.
- 2.1.1 Keyboard: every step has to be reachable by keyboard, and the focus order isn't always intuitive.
- 2.4.7 Focus Visible: in crowded forms the focus ring can be faint.
- 3.3.1 Error Identification / 3.3.3 Error Suggestion: feedback for invalid or overly broad queries is inadequate.
- 1.3.1 Info and Relationships / 4.1.2 Name, Role, Value: form fields and toggle actions need correct labels and roles.
4) Proposed redesign of filter creation
- Goals: the task should be findable, previewable, safe by default, and completely accessible.
- Main changes
- A token-based query builder inside the search bar
- Typing "from: ali" surfaces pickable chips such as From, To, Subject, Has attachment, List-id, and so on.
- Guided, labeled tokens (autocomplete plus examples) take the place of free-form recall.
- A live preview panel
- Display the number of matches plus the top three sample messages, with the matching fields highlighted.
- An impact banner: "This rule would affect ~2% of your new mail (about 15/day)."
- A safe-action wizard
- Step 1: set the conditions. Step 2: pick actions described in everyday words ("Apply label Receipts", "Skip the inbox"), each accompanied by an explanation.
- Start in Test Mode for seven days (labels only, nothing destructive). A clear switch turns on the full actions afterward.
- Guardrails against risk
- Warn when conditions get broad: "Over 1,000 emails match. Consider adding Subject or Sender."
- A confirmation showing a summary and an undo link.
- Accessibility built in from the start
- A keyboard-first flow, sensible tab order, and visible focus states.
- Correct labels and instructions (3.3.2), unambiguous error messages (3.3.1/3.3.3), and status updates through aria-live (4.1.3).
- Adequate contrast (1.4.3) and sufficiently large targets (2.5.5 where it applies).
- A token-based query builder inside the search bar
5) Success measures and guardrails
- Primary measures
- Filter Creation Success Rate: finished filters divided by filter flows begun.
- Time to Create: the median number of seconds between starting the filter flow and confirming.
- Misfilter Rate: messages that filters acted on automatically and that were then undone or returned to the Inbox within 24 hours, divided by all messages auto-acted on by filters.
- Secondary measures
- Adoption: the share of active users who create a first filter within 30 days.
- Precision proxy: the share of filter edits made within 72 hours that narrow the scope by adding conditions.
- Preview Utilization: the share who consult the preview or suggestions before creating.
- Guardrails
- Misfilter Rate may not rise by more than 0.5 per 1,000 auto-acted messages.
- Page performance must not regress (for example, p95 search-to-first-paint no worse than +50ms).
- Accessibility: automated checks (axe/pa11y, say) pass, and manual keyboard and screen-reader smoke tests meet the required WCAG 2.1 AA.
- User support: no meaningful rise in filter-related help tickets per MAU.
- A few small numeric examples
- Baseline success ; target a lift to , with median time dropping from 45s to 30s.
- Baseline misfilter 2 per 1,000; the guardrail ceiling is 2.5 per 1,000.
6) Plan for the experiment and telemetry
- Experiment design
- Randomization unit: A/B at the user level.
- Segmentation/stratification: existing filter users against first-timers; low, medium, or high email volume; desktop against mobile web.
- Ramp plan: 5% → 20% → 50% → 100%, with automated guardrail checks at every stage, and a 10% long-term holdout kept for several weeks.
- Duration: 4–6 weeks, long enough for Test Mode filters to touch enough new mail and to dampen novelty effects.
- Reducing novelty effects
- Learning period: drop the first 2–3 days from the primary analysis, or report stabilized metrics on their own.
- Repeat-exposure metrics: assess outcomes once a user has made at least one filter and had seven days of filter activity.
- Difference-in-differences: compare each user's before/after change across both arms to account for seasonality.
- CUPED/covariate adjustment: prior search usage and email volume serve to shrink variance.
- Telemetry/events to record
- search_opened, filter_builder_opened, token_added/removed, suggestion_clicked, preview_viewed, risk_warning_shown, test_mode_enabled, filter_created, filter_edited, filter_disabled.
- auto_action_applied (label/archive), undo_invoked, message_moved_to_inbox_after_auto_action.
- Timers: start and end timestamps for time-on-task, plus counts of reformulations.
- Power check (approximate)
- With baseline success and target (), a two-proportion test at and power requires roughly ~3,000 filter-start sessions per arm.
- Analysis
- Primary: intent-to-treat averages; medians reported for time-to-create.
- Sensitivity: results by segment; watch for Simpson's paradox across volume cohorts.
- Quality: examine misfilter examples through aggregate signals (no content inspection — only undo and move-back events).
7) Risks and how to mitigate them
- Risk: users quickly build rules that are too broad.
- Mitigation: Test Mode on by default, a prominent impact preview, and warnings.
- Risk: the extra UI raises cognitive load.
- Mitigation: progressive disclosure, sensible defaults, tight copy, and keyboard shortcuts.
- Risk: accessibility regressions.
- Mitigation: a WCAG 2.1 AA checklist in CI, plus manual screen-reader and keyboard QA on the important flows.
The result connects product judgment — heuristics and accessibility — to a redesign that can be measured, backed by a solid experiment plan that handles novelty and shields the user experience through explicit guardrails.