Capital One · ML & AI Fundamentals
Apply Leakage-Safe Imputation and Ordinal Encoding
TrueInterview
October 7, 2026 · 2 min read
You are given training and test tables that include numeric columns, an ordered categorical column service_tier with business ordering bronze < silver < gold, and a binary target available only in training. Some feature values are missing, and the test set may contain a tier that never appeared in training. Describe how to create processed_train and processed_test using imputation and ordinal encoding while preventing any information from the test set from leaking into the pipeline. Cover when to call fit_transform versus transform, how to keep feature alignment intact, and how to deal with unknown categories.
Constraints & Assumptions
- The target exists only in the training table and must never be passed into a feature transformer.
- Numeric and categorical columns may be imputed with different strategies.
- The two processed outputs must share the same feature columns in the same order.
- Unknown or missing tier values must not be silently mapped to a higher business tier.
Clarifying Questions to Ask
- Is the tier ordering defined by the business, or should it be learned from the observed labels?
- Should missing tier values become their own level, or be imputed to a known level?
- Which estimator will use the transformed features, and does it support sentinel values?
- Will preprocessing be assessed inside cross-validation, or only on a single train/test split?
Hint — Fit once on training data: All learned statistics, such as medians and category mappings, come from the training fit and are applied unchanged to validation and test data.
What a Strong Answer Covers
- Separate numeric and categorical pipelines that are fit only on training features.
- Explicit ordinal categories plus a deliberate encoding for unknown values.
- Correct use of
fit_transformon training data andtransformon all other data. - Stable column names, ordering, dtypes, and row alignment.
- Cross-validation implications and serialization of the fitted pipeline.
Follow-up Questions
- What kind of leakage happens if the imputer is fit before the cross-validation folds are created?
- When would one-hot encoding be safer than ordinal encoding?
- How would you track a rising rate of unseen service tiers after deployment?
Overview: Describe leakage-safe imputation and ordinal encoding for training and test data. Cover fit versus transform, business-defined category ordering, unseen levels, cross-validation, and stable output schemas.