Voleon · ML Coding
Load and visualize large CSV robustly
TrueInterview
October 7, 2026 · 1 min read
You are screen-sharing in a HackerRank environment with Python 3, pandas, numpy, seaborn, and matplotlib available.
You are given a single file, data.csv (about 1.5 GB). Its delimiter is unknown (either , or ;) and its encoding is either UTF-8 or latin-1. The columns are:
id: intdate:YYYY-MM-DDregion: strspend: floatclicks: intsignups: int
Up to 5% of the values may be missing, there can be exact duplicate rows, and some rows have clicks = 0.
Write code to:
- Detect the delimiter and encoding without loading the full file, then load the data in chunks while keeping peak memory under 1 GB.
- Drop exact duplicate rows and enforce the column dtypes.
- Impute missing
spendwith the medianspendwithin the row'sregion. Impute missingclicksandsignupswith 0 only if the row's non-null feature count is at least 4; otherwise drop the row. Justify this rule. - Create
cpc = spend / clicksusing safe division, and winsorizecpcat the 1st and 99th percentiles by region. - Produce and save:
- (a) a scatter plot of
spendvssignupswith a LOWESS smoothed line and a 95% confidence interval; - (b) a boxplot of
signupsbyregion, with regions sorted by median; - (c) a time-series line of daily total
signups.
- (a) a scatter plot of
- Briefly explain your memory and time complexity choices, and how you would test this code.
Provide runnable, end-to-end code, and state any assumptions explicitly.
Hints (3)
Loading comments…