แบบทดสอบที่เชื่อมโยงกับคำแนะนำ · ยาก ระดับ

Calibrating LLM Judges Against Humans Quiz

Calibrate LLM judges with shared rubrics, independent human labels, confusion matrices, suitable agreement statistics, and repeated held-out checks.

เส้นทางแนะนำที่เกี่ยวข้องCalibrating LLM Judges Against Humans
คำถาม 1 ของ 8

Why is Cohen's kappa preferred over raw percent agreement when calibrating an LLM judge?