Rehber bağlantılı sınav · Sert Seviye

Evaluating AI Agents Quiz

Tests how agent evals score task success, trajectories and tool use, why environment-based grading and repeated trials matter, and grader trade-offs.

Soru 1 arasında 8

Temsilciler için neden tek bir son cevabı kontrol etmek genellikle yeterli olmuyor?