Mistral AI · Behavioral
Explain Mixture of Experts and Expert Parallelism
TrueInterview
September 26, 2026 · 1 min read
Describe what a mixture-of-experts transformer is, and show how expert parallelism spreads its computation across devices. Walk through the path of a token batch as it is routed, processed by experts, and then combined.
Constraints & Assumptions
Take a sparse feed-forward mixture-of-experts layer with top-k routing as the reference case. The expert count, the value of k, the capacity rule, and device layout are not given, so make and state reasonable choices where necessary.
Clarifying Questions
Do all experts live on a single device, or are they spread across devices? Is a token allowed to go to more than one expert? What occurs if an expert is assigned more tokens than its capacity allows? Are we optimizing for training throughput or for serving latency?
What a Strong Answer Covers
Keep the model architecture distinct from how devices are partitioned; explain sparse activation and the communication needed for routing; and call out load imbalance, capacity, and memory overheads.
Follow-up Questions
Why is it possible for a model to have a huge number of total parameters while only a small fraction are active for any one token? What behavior does a load-balancing loss push the model toward? Under what conditions can communication overhead wipe out the expected computational savings?
Overview: Explain sparse mixture-of-experts routing and expert parallelism, covering token dispatch, weighted aggregation, load balancing, overflow handling, and communication costs.
Loading comments…