JPMorgan · ML & AI Fundamentals
Explain Core ML Concepts
TrueInterview
October 7, 2026 · 8 min read
Imagine you are sitting for a senior AI/ML-focused Data Scientist position at a financial institution (J.P. Morgan). This segment is the 'ML fundamentals' portion of a technical screen: a rapid-fire oral Q&A that tests whether you grasp the rationale behind fundamental machine-learning concepts, rather than merely reciting textbook definitions. The interviewer looks for responses anchored in actual production modeling scenarios — credit risk, fraud detection, customer churn, transaction classification — where datasets are tabular, positive labels are rare and arrive late, and observations follow a time order. Address each part below with clarity and sufficient technical depth to meet a senior standard. Throughout your answers, tie definitions to model behavior, validation design, and what happens in deployment.
Constraints & Assumptions
- This is the conceptual ML-fundamentals part of the screen — respond as if in a live discussion. Prepare for the interviewer to press for specifics ("what exactly does L2 do to correlated features?") rather than tolerate vague statements.
- The practical setting is tabular financial data: extremely low positive rate (e.g., fraud under 1%), delayed labels (chargebacks resolve only after weeks), and pronounced drift over time.
- Whenever evaluation or validation comes up, treat the data as time-ordered; a naive random train/test split can leak future information.
- Assume precise terminology matters: bias versus variance, sparsity versus shrinkage, filter versus wrapper versus embedded, self-attention versus recurrence.
Clarifying Questions to Ask
- Which model families already run in production here — gradient boosted trees on tabular data, deep sequence models on transaction streams, or both? The answer changes which of these concepts matter most on a daily basis.
- For the validation and leakage discussion: is the prediction problem point-in-time (must honor an "as-of" timestamp), and do labels arrive late (e.g., chargeback resolution)?
- Is interpretability a hard requirement (for example, adverse-action notices in credit), which would steer me toward sparse or explainable models?
- In any sequence-modeling use case, how long are the sequences — dozens of events, or thousands? That determines whether quadratic attention is even computationally practical.
- For feature selection, is the constraint statistical (generalization) or operational (latency, cost, governance)?
Part 1 — Compare bagging and boosting
Describe the issue each ensemble approach addresses, name representative algorithms for each, and explain the effect each has on bias and variance. Be precise about how the two differ in their mechanics (how base learners are trained and how their outputs are combined).
Hint — Which error term does each target?: The bias-variance decomposition includes two reducible terms. Before you answer, ask yourself: which single term is each method designed primarily to reduce, and which term does it leave essentially untouched? Anchor your answer to that decomposition instead of vague "accuracy." Hint — Mechanics that set them apart: Examine the training procedure of each: are base learners trained in parallel on independently resampled data, or sequentially, each one targeting the mistakes the running ensemble still makes? Are the base trees typically deep or shallow, and why does that pairing make sense given the error term being attacked? Name a flagship algorithm from each family.
What This Part Should Cover
- Correct mechanism, not just names: how each ensemble trains its base learners (parallel/independent bootstrap samples versus sequential/residual-targeting) and how it combines them (averaging/voting versus additive weighted sum).
- Which error term each attacks: correctly identifies which term each method aims to shrink, explains why the typical base-learner depth (deep versus shallow trees) is a deliberate match for that target, and notes how boosting's overfitting risk grows without regularization.
- A flagship algorithm per family (e.g., Random Forest versus XGBoost/LightGBM) and at least one concrete regularization mechanism that curbs boosting's overfitting risk.
Part 2 — Explain the bias-variance tradeoff
Define bias and variance, describe what high bias (underfitting) and high variance (overfitting) look like, and explain how you would diagnose each by comparing training versus validation performance.
Hint — A 2x2 diagnostic: Think about the gap between training and validation error. Examine all four corners: (bad train, bad val), (good train, bad val), (good train, good val), (bad train, good val) — each suggests a different diagnosis. The last corner is suspicious and typically signals a data leakage or sampling issue rather than a genuine fit. Hint — The financial wrinkle: On time-ordered data, ask whether the validation split itself is honest — a random split can make a leaky or drifting model appear healthy.
What This Part Should Cover
- The error decomposition (bias², variance, irreducible noise) and concrete symptoms of underfitting versus overfitting.
- The train-vs-validation gap as a diagnostic tool, including the "too good to be true" quadrant that signals leakage or a non-representative split, and the use of learning curves.
- The time-ordered wrinkle: why a random split is dishonest here and why a time-based or walk-forward split is required.
Part 3 — Methods to reduce model variance, and L1 vs. L2
List the main levers for reducing variance (regularization, more data, cross-validation, ensembling/averaging, early stopping, pruning/complexity limits, feature reduction, dropout). Then go deeper on regularization: write out the L1 and L2 penalty terms, explain their different effects on the weights, and state when L1 is preferred over L2.
Hint — Look at the shape of each penalty: One penalty sums the absolute values of the weights, the other sums their squares. Picture the constraint region each one carves out (think about whether it has sharp corners on the axes or is smoothly rounded). Reason from that geometry to what each does to a typical weight — and decide for yourself which one can pin weights to exactly zero versus only shrink them. Hint — Let your belief about the features pick the penalty: Tie the choice to a prior about the feature set. Ask: do you expect that only a few features genuinely matter, or that many features each contribute a little and several are correlated? Match each belief to whichever penalty's behavior (from the previous hint) is the better fit — and recall there's a hybrid penalty that targets the middle ground.
What This Part Should Cover
- A broad menu of variance-reduction levers beyond regularization, with the financial caveat that "more data" is constrained by rare, delayed positives.
- Regularization geometry: states the L1 and L2 penalty formulas correctly, accurately characterizes the qualitatively different effect each has on individual weights, and gives a geometric or algebraic explanation for why that difference arises — without relying solely on memorized vocabulary.
- A defensible decision rule for L1 versus L2 (sparse/few-relevant-features/governance versus many-small/correlated features), plus where Elastic Net fits.
Part 4 — Explain feature selection
Compare filter, wrapper, and embedded methods (what each does, plus a pro and a con). Explain how to avoid data leakage during feature selection. Finally, explain how feature selection differs for linear, tree-based, and deep learning models.
Hint — Three families, one axis: Organize by how tightly selection is coupled to the model: ranking features by a statistic independently of the model (filter), searching feature subsets by repeatedly training the model (wrapper), or selecting during training (embedded — e.g., an L1 penalty or tree split importance). Hint — Leakage is about when selection happens: The classic trap: selecting features on the full dataset before splitting. Think about where selection must live relative to the train/validation boundary — and, for cross-validation, that it must happen inside each fold, on time-ordered data with a time-based split.
What This Part Should Cover
- The three families correctly distinguished by coupling to the model, each with a representative method and a real pro/con (e.g., filters ignore interactions; wrappers are costly and overfit-prone; embedded importance can be biased toward high-cardinality features).
- Leakage discipline: selection on training data only, inside each CV fold, with point-in-time correctness on time-ordered data.
- Model-type sensitivity: scaling/multicollinearity for linear models; trees' robustness to monotonic transforms but vulnerability to noisy/high-cardinality features (prefer permutation/SHAP over raw impurity importance); deep models relying on representation learning and embeddings for high-cardinality categoricals.
Part 5 — Compare Transformers and RNNs
Explain why Transformers largely displaced RNNs for many sequence tasks. Describe the attention mechanism at a high level and using the query-key-value formulation (including the scaled dot-product form). Then discuss the computational tradeoffs, sequence-length limitations, and interpretability caveats.
Hint — Why Transformers won: Compare how each handles the sequence dimension: an RNN's hidden state is computed step by step (inherently serial, with long-range gradient issues), while self-attention lets every position look at every other position in parallel, directly modeling long-range dependencies. Hint — Frame attention as soft retrieval: Think of it as a lookup: each position issues something like a query, every position advertises a key, and carries a value. Work out how a query and the keys would combine to decide how much of each value to pull in. Then ask two follow-ups for yourself: why might raw similarity scores need to be rescaled before the weighting step, and — if every position attends to every other — how does the cost grow as the sequence gets longer? Hint — The honest caveat: Be ready to push back on a common myth: high attention weight is not a proof of importance/explanation. Mention how you'd actually validate attribution (ablation/counterfactuals), and that attention is permutation-invariant so positional information must be injected.
What This Part Should Cover
- Why Transformers replaced RNNs: training parallelism, direct long-range dependency modeling, and superior scaling — contrasted with the RNN's serial recurrence and vanishing-gradient problems (even with LSTM/GRU gating).
- The attention formulation: correctly identifies the query, key, and value roles; states the scaled dot-product formula accurately (including the scaling term and a valid explanation for why that scaling is needed) — assessed on both correctness of the formula and the reasoning behind it.
- The honest caveats: cost in sequence length and its mitigations, the need for positional encoding, and "attention explanation."
What a Strong Answer Covers
These cross-cutting dimensions span all five parts and are what separate a senior answer from a textbook recitation:
- Leakage awareness throughout — selection inside folds, point-in-time correctness, and time-based splits — framed for the financial setting rather than stated abstractly.
- Connecting back to business cost: under heavy class imbalance, accuracy is the wrong metric; precision/recall, PR-AUC, and expected dollar loss are what matter, and false positives versus false negatives carry very different costs.
- Precise vocabulary and honest tradeoffs: using terms exactly (bias versus variance, sparsity versus shrinkage, filter/wrapper/embedded) and volunteering the failure modes (overfitting, drift, biased importance, attention-as-explanation myth) instead of waiting to be cornered.
Follow-up Questions
- Random Forest and XGBoost both rely on trees — given a noisy, imbalanced, time-drifting tabular dataset, which would you choose first and why?
- Your offline PR-AUC is excellent but the model fails in production. Walk through how you'd determine whether the cause is leakage, drift, or a bad validation split.
- Elastic Net has two hyperparameters. How would you tune the L1/L2 mix without leaking, on time-ordered data?
- For a sequence of 5,000 transactions per customer, full self-attention is too expensive. What concrete options would you consider to make a Transformer-style model tractable? Overview: This question assesses mastery of core machine learning concepts — notably ensemble methods (bagging versus boosting) and the bias-variance tradeoff — and the ability to connect those concepts to model behavior, validation design, and deployment issues in tabular, imbalanced, time-ordered financial data.