Back to problems

Choose Fast or Cheap Models

System Design · Google · Medium

Consider a product backed by machine-learning inference. For every incoming request, the system can call either of two serving paths: Path A: more expensive per generated token, but returns results faster. Path B: less expensive per generated token, but returns results more slowly. How would you choose between these paths for different requests? Discuss the trade-offs related to end-user experience, latency, output quality, reliability, and operating cost. Also describe the…

Checking your access…