Back to problems

Design video captioning under compute limits

System Design · ByteDance · Hard

You are a machine learning engineer responsible for productionizing a multimodal video captioning model on a short-video platform. The model receives a video as sampled frames, with audio as an optional input, and outputs a text caption describing that video. Your work has two parts. Part A — Serving the captioning model under tight compute and GPU-memory limits The platform continuously ingests a very high volume of new videos. The captioning system must support both…

Checking your access…