Netflix · ML & AI Fundamentals
Compare Losses and Explain LoRA
TrueInterview
October 7, 2026 · 4 min read
ML Fundamentals: Loss Functions and Low-Rank Adaptation
This screen is a rapid-fire pass through ML fundamentals. The expectation is precise reasoning about loss functions and parameter-efficient fine-tuning, not merely repeating definitions. Treat both parts below with the depth an interviewer would push for.
Constraints & Assumptions
- This is a conceptual/whiteboard screen: no coding is required for these two parts, but equations and clear reasoning are expected.
- Assume the LoRA discussion is aimed at a standard transformer, meaning attention plus MLP blocks.
- "Parameter-efficient" is judged by the number of trainable parameters and the optimizer-state memory they demand, not by raw forward-pass FLOPs.
Clarifying Questions to Ask
- Is the Part 1 comparison intended for regression only, or should I also cover what breaks when MSE is applied to a classification problem?
- Do you want full derivations on the board—gradients and parameter counts—or higher-level intuition for why each method behaves as it does?
- For Part 2, should I assume a default LoRA placement, or discuss which transformer modules to target?
- How much practical depth are you looking for—for example, rank/ tuning, merging, and comparisons with other PEFT methods—versus just the core mechanism?
Part 1 — Mean Squared Error vs. Cross-Entropy
Compare the mean squared error (MSE) loss with the cross-entropy (CE) loss. Cover:
- When each one is appropriate and the type of task it targets.
- The probabilistic assumption behind each loss—that is, which likelihood you are implicitly maximizing.
- How their optimization behavior differs, especially why cross-entropy is preferred over MSE for classification with a sigmoid or softmax output.
Hint — Where to start: Each loss is a negative log-likelihood under an assumed output distribution. Ask: when you minimize this loss, what distribution are you implicitly fitting the labels to? One family is continuous, and the other is discrete.
Hint — Optimization behavior: Compute the gradient of the loss with respect to the pre-activation logit for a sigmoid output. Compare with . Check what happens when the model is confidently wrong ( but ): does the gradient vanish or remain strong?
Hint — Convexity / surface: Think about whether each loss is convex in the model weights for a linear classifier, and what happens to that surface when a sigmoid is composed with squared error. This partly explains why MSE-on-sigmoid trains slowly.
What This Part Should Cover
- Task fit: correctly matches each loss to the kind of task and target it suits—continuous regression versus categorical classification.
- Likelihood framing: ties each loss to a maximum-likelihood objective under an assumed output/noise distribution—Gaussian for MSE, Bernoulli/categorical for CE—rather than treating it as an arbitrary distance.
- Gradient argument: derives the logit-gradient for both losses on a sigmoid/softmax head and uses that to explain why CE is preferred, especially in the confidently-wrong or saturation case, instead of merely asserting it.
- Loss-surface intuition: notes the convexity / flat-region difference and why that makes MSE-on-sigmoid slow or unstable.
Part 2 — Low-Rank Adaptation (LoRA)
Explain Low-Rank Adaptation (LoRA). Cover:
- The core idea and the math: how the weight update is parameterized.
- How it is used when fine-tuning a large neural network or LLM—which layers, what is frozen versus trained, and what happens at inference.
- Why it is parameter-efficient, and roughly how the trainable parameter count compares with full fine-tuning.
Hint — Where to start: Do not fine-tune directly. What if you froze the original weights and learned only the update —and what structural constraint could you place on so it has far fewer free parameters than itself?
Hint — The math: A low-rank matrix can be written as the product of two thin factors. Express the update in that form, choose a rank, and reason about how that rank sets the trainable-parameter count relative to the full matrix. Is there a scalar you would add to keep the effective update magnitude stable as the rank changes?
Hint — Initialization & inference: At the start of fine-tuning, the adapted model should behave exactly like the pretrained model—what does that imply about how the two factors should be initialized? At inference, ask whether the learned update can be folded back into the original weights so there is no extra latency.
What This Part Should Cover
- Mechanism: states what is frozen versus trained and how the update is structured—a product of two thin matrices—to be cheap, rather than merely naming the method.
- Math & initialization: writes as a scaled low-rank product, explains the rank choice, and gives the zero-init detail that makes the adapter start as a no-op.
- Efficiency, quantified: quantifies the trainable-parameter reduction, ideally with a concrete count, and identifies optimizer-state / memory savings as the real win rather than forward FLOPs.
- Inference handling: addresses how the adapter is treated at inference—merging versus keeping it separate—and the latency / multi-task serving trade-off.
- Honest limitations: names regimes where a low-rank update falls short of full fine-tuning.
Follow-up Questions
- Derive for both losses on a single sigmoid output and explain precisely why MSE's gradient can stall.
- For a attention projection with , what fraction of that layer's parameters does LoRA train?
- What do the LoRA hyperparameters , , and dropout control, and how would you tune them?
- How does LoRA compare with other PEFT methods such as adapters, prefix-tuning, or QLoRA, and when would you choose each?
Overview: This question tests understanding of loss functions—mean squared error versus cross-entropy—and parameter-efficient fine-tuning through Low-Rank Adaptation (LoRA), with emphasis on probabilistic interpretations, optimization behavior, and low-rank parameterization.