Oracle · Behavioral
Evaluate Subjective, Nondeterministic Agent Outputs
TrueInterview
September 30, 2026 · 1 min read
Requirements
- Describe how to assess an agent whose outputs are both subjective and nondeterministic.
- Respond to three specific objections raised against the evaluation approach:
- A judge built on an LLM can be too confident in its judgments.
- Users might not give explicit feedback.
- Criteria written by developers can embed developer bias.
- Talk through whether online user signals can stand in for absent direct feedback, and say what those signals would actually prove about output quality.
Notes
- The conversation is deliberately adversarial: simply naming user feedback or an LLM judge is the beginning, not a full answer.
- The interviewer presses on how trustworthy and available each evaluator is, so keep the quality signal, its source, and its failure modes distinct.
- The candidate suggested online user signals after the objections about explicit feedback and the LLM judge, but the round continued without settling on a preferred answer.
Preparation
- Rehearse a structured plan for evaluating subjective, nondeterministic outputs, then be ready to defend it against all three objections above.
- Prepare concrete examples of online behavior signals and be able to explain what each one can and cannot establish about agent quality.
- Practice keeping evaluator confidence separate from evaluator correctness when discussing an LLM judge.
Loading comments…