Mistral AI · Behavioral
Compare Tensor Parallelism and Pipeline Parallelism
TrueInterview
September 26, 2026 · 1 min read
Contrast tensor parallelism with pipeline parallelism when training or serving a transformer that cannot fit on a single device. Describe what gets partitioned, the communication patterns each method demands, and the scenarios where combining them makes sense.
Constraints & Assumptions
No specific model size, device topology, or throughput target is given. Take a transformer as the example and separate the training case from autoregressive inference. Do not presume one technique is always faster.
Clarifying Questions
Is the model exceeding device memory because of weights, optimizer state, or activations? What are the interconnect bandwidth, batch size, sequence length, and latency targets?
What a Strong Answer Covers
Demonstrate a within-layer tensor partition and a between-layer pipeline partition; explain the required collectives and activation forwarding; and discuss memory footprint, utilization, and scheduling.
Follow-up Questions
What causes pipeline bubbles? How do micro-batches reduce them? Why is tensor parallelism often restricted to a high-speed local interconnect? How does a small inference batch size influence the decision?
Overview
Examine transformer tensor and pipeline parallelism by considering parameter sharding, collective communication, activation passing, micro-batch bubbles, memory usage, and inference latency.
Loading comments…