Back to problems

Evaluation System with Human + LLM Evaluators

System Design · Waymo · Hard

Requirements Input consists of a flow of model responses that must be assessed using a rubric. Provide two groups of evaluators: Human reviewers: compensated annotators with defined skill levels, capacity constraints, and regional/time-zone availability. LLM evaluators: automated model-based graders that are quicker and lower cost, but less dependable. For every item, return a score (or score distribution), a confidence value, and a traceable audit record. The system must…

Checking your access…