LinkedIn · ML & AI Fundamentals
Explain Logistic Regression, Backprop, and Adam
TrueInterview
October 7, 2026 · 6 min read
Trace the mathematical ideas that link logistic regression to the way deep networks are trained today. You are expected to write out the equations, work through the gradient derivations rather than merely reporting them, and justify the optimizer's design decisions; this is a whiteboard-style 'show your math' session, not a high-level discussion.
Constraints & Assumptions
- Binary classification where the input is and the label is .
- You must be able to obtain results from first principles; simply quoting a final formula without showing the chain-rule steps is not enough.
- Use vector or matrix notation when it is natural, such as a batch design matrix .
- Standard definitions hold: is the logistic sigmoid, log means natural log, and (or ) is a learning rate.
Clarifying Questions to Ask
- Should the derivations start from first principles, or is it acceptable to state the final gradient with a one-line justification?
- Do you prefer scalar notation for a single example, or fully vectorized matrix notation across a batch?
- For the network, should hidden layers use a particular activation such as sigmoid or ReLU, or should remain general?
- For the Part 1 gradient, do you want the single-example version, the batch-averaged version, or both?
- For Adam, do you care about the bias-correction terms and the role of , or only the core moment estimates?
Part 1 — Logistic regression
- Explain how logistic regression represents binary classification.
- Write the model from linear score to sigmoid, the sigmoid function itself, and the binary cross-entropy (BCE) loss for a single example and for a batch of examples.
- Derive the gradient of the loss with respect to the parameters and , and give the gradient-descent update.
Hint — Where to start: Set and , then differentiate the BCE loss. Hold as an intermediate quantity: the chain rule splits into . Hint — The key simplification: The sigmoid has the handy derivative . When this is combined with the BCE loss, look for cancellations: the awkward factors are precisely what allow the result to collapse into a single clean expression. Aim for that cancellation.
What This Part Should Cover
- Probabilistic framing: interpret as , and identify BCE as the Bernoulli negative log-likelihood rather than an arbitrary loss.
- The chain rule carried through the sigmoid-derivative cancellation to arrive at the compact form , derived rather than quoted.
- Both the single-example gradient and the vectorized batch gradient , along with the gradient-descent update and a note that the objective is convex in .
Part 2 — From logistic regression to a neural network
- Show that logistic regression is precisely a one-layer network with no hidden layer, and describe what changes when hidden layers are added.
- For a feedforward network with layers, write the forward pass in mathematical form, then derive the backpropagation equations for the weights and biases.
Hint — Forward pass: For each layer : and , with . Introduce a per-layer error term to structure the recursion. Hint — The backward recursion: The entire derivation depends on expressing in terms of ; once that recurrence is set up, the per-weight gradient follows from the chain rule. While doing so, track which factor propagates the error to the previous layer and which factor accounts for the local activation. Let the matrix shapes tell you where each term belongs.
What This Part Should Cover
- A correct statement that logistic regression is the no-hidden-layer case, and a clear explanation of what hidden layers contribute: learned non-linear features before the final linear-plus-sigmoid stage.
- A correct forward pass and a backprop recursion written in terms of the per-layer error , including the elementwise activation-derivative factor.
- Weight and bias gradients with the correct shapes—outer product per example, matrix product for a batch—and why backprop is because activations are cached.
Part 3 — Mini-batch gradient descent
- Write clear pseudocode for training a model with mini-batch gradient descent, and explain why mini-batches are preferred to full-batch updates or pure SGD.
Hint — Structure: Use two nested loops: an outer loop over epochs that shuffles the data each epoch, and an inner loop over batches of size . Each batch performs forward, loss, backward, and parameter update. Think about where the batch averaging occurs.
What This Part Should Cover
- Pseudocode that shuffles every epoch, splits the data into batches of size , and runs forward, batch-averaged loss, backward, and update, with the averaging placed so the gradient scale is roughly independent of batch size.
- The tradeoff behind mini-batches: better hardware throughput and lower gradient variance than full-batch, while keeping enough stochasticity and bounded memory compared with pure SGD.
Part 4 — Adam optimizer
- Describe Adam's main components and write its update rule: first moment, second moment, bias correction, and parameter step.
- Explain why Adam divides each update by the square root of the gradient's second moment.
Hint — Two ideas combined: Adam combines momentum, an EMA of the gradient , with an adaptive per-parameter step size, an EMA of the squared gradient . Write both EMAs, then reason about why the moving averages need bias correction early in training. Hint — Why the square root: Consider the units. The second moment accumulates , so ask what units it has and what happens to the units of the step after taking the root. From there, reason about the effect on the per-parameter step magnitude.
What This Part Should Cover
- All four ingredients written correctly: the first-moment EMA, the second-moment EMA, bias correction and why it is needed when , and the parameter update with .
- A genuine argument, not just a restatement, for the denominator: dimensional consistency or scale-invariance of the step, and per-parameter adaptive step sizing that protects rare or large-gradient directions.
What a Strong Answer Covers
These dimensions cut across all four parts; the per-part rubrics above cover the part-specific content.
- Notation discipline: one consistent notation reused across parts, such as reusing the Part 1 cancellation, so the parts visibly connect instead of reading as four disconnected facts.
- Derive, don't quote: every gradient is reached through visible chain-rule steps; final formulas are the endpoint of a derivation, not the starting point.
- Intuition alongside algebra: each result is paired with a one-line 'what this means'—residual-times-input gradient, error propagated backward, momentum plus adaptive scaling—and stated assumptions are surfaced up front.
Follow-up Questions
- Why is BCE loss preferred over squared error for classification? What happens to the gradient and optimization landscape if MSE is used with a sigmoid output?
- During backprop with sigmoid or tanh activations, what is the vanishing-gradient problem, and how do ReLU, normalization, or residual connections reduce it?
- How does Adam differ from RMSProp and from SGD with momentum? When might plain SGD with a tuned schedule generalize better than Adam?
- What is weight decay, and why is decoupled weight decay (AdamW) different from adding L2 regularization to the loss when using Adam? Overview: This question assesses supervised learning fundamentals, the mathematical link between logistic regression and neural network training, gradient derivation, and optimizer mechanics in the Machine Learning domain for Data Scientist roles.