System Design · ByteDance · Hard
You are a machine learning engineer responsible for productionizing a multimodal video captioning model on a short-video platform. The model receives a video as sampled frames, with audio as an optional input, and outputs a text caption describing that video. Your work has two parts. Part A — Serving the captioning model under tight compute and GPU-memory limits The platform continuously ingests a very high volume of new videos. The captioning system must support both…
Checking your access…