Amazon · Statistics & Data Analysis
Profile and visualize an unfamiliar dataset
TrueInterview
October 7, 2026 · 1 min read
You are handed an undocumented CSV that mixes user events and purchases, with these columns: order_id, user_id, event_ts (UTC), merchant_id, session_id, event_type (view/add_to_cart/purchase), amount_usd, device_type, country. Within 30 minutes, outline precisely how you would make sense of the dataset and build executive-ready visualizations.
Requirements:
- Data understanding: enumerate the first 10 checks you would run (e.g., missing values, duplicate rows, whether timestamps increase monotonically by session, timezone plausibility, cardinality of categorical fields, outliers, consistent units, referential integrity between
event_type='purchase'and amount_usd, weekend versus weekday behavior, coverage across country and device). - Visual plan: propose 3–5 concrete charts (with titles, axes, and grain) that answer “What is happening?” and “So what?”. Explain the reasoning behind each choice and the insight you expect.
- Granularity: pick daily or hourly aggregation for a launch week; justify the trade-offs and describe how you would switch using a parameter.
- Data quality: show how a single day with bad clock skew would show up in your visuals, and how you would annotate or adjust for it.
- Deliverable: describe a one-slide dashboard wireframe (sections, KPIs, filters) and how you would validate it with a stakeholder during a 5-minute readout.
Overview: This question tests a data scientist's skills in exploratory data analysis, data quality assessment, time-series aggregation, visualization design, and concise stakeholder communication when working with an undocumented events-plus-purchases CSV.
Loading comments…