引导式测验 · 硬 等级

Evaluating AI Agents Quiz

Tests how agent evals score task success, trajectories and tool use, why environment-based grading and repeated trials matter, and grader trade-offs.

相关引导路径Evaluating AI Agents
问题 1 的 8

Why is checking a single final answer usually not enough for agents?