ByteDance · ML & AI Fundamentals
How to deploy and tune multimodal models?
TrueInterview
October 7, 2026 · 2 min read
Question
You are interviewing for a new-grad machine learning / data scientist position at ByteDance. Respond to the following related machine-learning and LLM questions.
- Multimodal deployment under constraints: Suppose you must deploy a multimodal model (such as text + image, text + video, or audio + text) with tight limits on GPU memory (VRAM), compute, latency, and cost. How would you rework the model and serving stack to cut memory, latency, and cost while keeping quality acceptable? Cover model-level changes (quantization, distillation, pruning, input/architecture simplification), serving-level changes (batching, caching, efficient kernels, tiered serving), and modality-specific optimizations, and weigh the tradeoffs of each.
- Fast video retrieval with captions and embeddings: In a large corpus, every video already has a caption and one or more precomputed embedding vectors. How would you design a retrieval system that responds to user queries quickly while preserving high recall? Address offline preprocessing, index design, approximate nearest neighbor (ANN) search, lexical versus dense (hybrid) retrieval, reranking, freshness / update trade-offs, and the relevance and serving metrics you would use to evaluate it.
- Overfitting: Define overfitting, explain how you would detect it, and describe the most effective mitigation methods in deep learning systems.
- Dropout: Explain the intuition behind dropout, why it can reduce overfitting, and why the usual "inverted dropout" implementation keeps the expected activation scale consistent between training and inference. Also note situations where dropout may be less effective or even harmful.
- Normalization layers: Compare Batch Normalization, Layer Normalization, Group Normalization, Instance Normalization, and RMSNorm. Which statistics does each use, how do training-time and inference-time behavior differ, and why are certain normalization layers favored in transformers or small-batch settings?
- Reinforcement learning for LLM post-training: Explain how reinforcement learning is applied in LLM post-training, particularly in RLHF. Outline the typical pipeline from supervised fine-tuning to preference modeling and policy optimization, the role of KL regularization, common failure modes, and how you would evaluate it. You may also compare PPO-style RLHF with newer preference-optimization approaches such as DPO. Overview: A six-part ByteDance data scientist onsite ML interview covering deployment and tuning of multimodal models under GPU/VRAM, compute, and latency constraints; design of fast video retrieval over precomputed captions and embeddings with ANN and hybrid reranking; and deep-learning fundamentals — overfitting detection and mitigation, dropout and inverted-dropout scaling, normalization layers (BatchNorm/LayerNorm/GroupNorm/InstanceNorm/RMSNorm), and reinforcement learning for LLM post-training (RLHF). It assesses practical systems tradeoffs alongside theoretical foundations across ML systems, retrieval, and reinforcement learning.
Loading comments…