Amazon · ML & AI Fundamentals
Explain Transformers and MoE in LLMs
TrueInterview
October 7, 2026 · 1 min read
You are interviewing for a position that involves large language models (LLMs). Describe the following concepts and how they connect to building and scaling LLMs:
- Transformer architecture
- What are the main building blocks (for example, self-attention, multi-head attention, positional encodings, feed-forward networks)?
- At a high level, how does self-attention operate?
- Why do Transformers fit language modeling better than RNNs/LSTMs?
- Mixture-of-Experts (MoE) architecture
- What issue does MoE aim to address for LLMs?
- Conceptually, how does expert routing work (for instance, gating networks, top-k experts)?
- What are the primary trade-offs in MoE (compute efficiency versus model complexity, training stability, load balancing)?
- Collective communication and parallelism for LLMs
- Briefly explain the common parallelism strategies used to train and serve large models: data parallelism, tensor/model parallelism, and pipeline parallelism.
- What is collective communication (such as all-reduce, all-gather, broadcast), and why does it matter for large-scale distributed training?
- Provide a simple example of where all-reduce is used during Transformer training. Aim for clear explanations that would help a strong software engineer understand how large language models are structured and scaled. Overview: This question assesses knowledge of large language model architectures and systems-level scaling skills—specifically core Transformer concepts, Mixture-of-Experts routing, and collective communication primitives—in the Machine Learning category.
Loading comments…