Microsoft · ML & AI Fundamentals
Compare CNN, RNN, and LSTM rigorously
TrueInterview
October 7, 2026 · 1 min read
Compare CNNs, RNNs, and LSTMs in a rigorous way for sequence modeling. Answer every part:
- Inductive biases and use cases: For time series, when would you choose a 1D dilated CNN instead of an RNN or LSTM? In what situations does an LSTM clearly beat a vanilla RNN?
- Vanishing and exploding gradients: Give the hidden-state recurrence for a vanilla RNN and explain why gradients can vanish or explode. Then give the LSTM gate equations (input, forget, output, cell) and explain how additive paths and gating reduce the problem.
- Parameter and computation comparison: Given an input of shape (batch=32, time=100, features=64), compute the parameter counts for: (a) a 1D CNN with 128 filters, kernel size 3, stride 1, and no bias-sharing tricks; (b) a single-layer unidirectional GRU with 128 hidden units; (c) a single-layer unidirectional LSTM with 128 hidden units. Show the formulas and totals. Discuss the implications for parallelism and latency.
- Experimental design: You have only 50k labeled sequences and a strict latency budget of less than 5 ms per sample. Propose an ablation plan to choose among the models above, covering regularization, data augmentation, and early stopping criteria. Define the primary metrics and stopping rules. Overview: This question tests sequence modeling skills: comparative understanding of CNN, RNN, and LSTM inductive biases, gradient dynamics and gating mechanisms, parameter and computation trade-offs, and experimental design under constraints, within the Machine Learning area focused on deep learning architectures for time series.
Loading comments…