System Design · OpenAI · Hard
Design the execution harness and offline evaluation system for an AI agent that can call external tools. The agent receives a user goal, may perform several model and tool steps, and eventually returns a final result. Your platform must provide reproducible offline runs, a safe isolation boundary, useful failure diagnostics, and controlled comparison of different agent releases. Keep the scope on the harness and evaluator, not on model training or low-level model serving.…
Checking your access…