System Design · Amazon · Medium
You are given a tabular dataset for supervised learning with the following columns: F1: a count variable whose values are mostly small integers, with many zeros. F2: a monetary amount in dollars, exhibiting a heavy-tailed distribution. F3: a binary flag. F4 and F5: continuous measurements that are highly correlated with each other. y: the target variable. Your tasks are: Identify precisely which features require standardization or normalization, and justify why. For each…
Checking your access…