ByteDance · ML & AI Fundamentals
Explain deployment, retrieval, and regularization
TrueInterview
October 7, 2026 · 1 min read
You are interviewing for a machine-learning position at a large short-video platform. Respond to the conceptual questions below.
- Given strict limits on GPU compute and VRAM, how would you serve a multimodal model for video retrieval or ranking? Cover architecture decisions, compression, batching, caching, and the trade-offs among quality, p99 latency, throughput, and serving cost.
- Assume captions and video embeddings are already precomputed and persisted. How would you speed up online video retrieval? Address indexing, approximate nearest neighbor search, hybrid text-plus-vector retrieval, reranking, memory usage, and freshness.
- Define overfitting, explain how you would detect it, and describe how you would reduce it in deep learning systems.
- Explain the intuition and math behind Dropout, including why inverted Dropout preserves the expected activation scale from training to inference.
- Compare common normalization approaches such as BatchNorm, LayerNorm, GroupNorm, and RMSNorm. When is each suitable, and how are their statistics handled during inference?
- Explain how reinforcement learning is applied in LLM post-training, particularly RLHF. Describe the roles of supervised fine-tuning, preference data, reward modeling, policy optimization, KL regularization, and typical failure modes.
Overview: This question assesses machine learning systems engineering skills, including multimodal model deployment under GPU/VRAM limits, scalable video retrieval and indexing, regularization and normalization techniques, and reinforcement learning–based post-training for language models.
Loading comments…