Guide-linked quiz · Hard Level
Calibrating LLM Judges Against Humans Quiz
Calibrate LLM judges with shared rubrics, independent human labels, confusion matrices, suitable agreement statistics, and repeated held-out checks.
Question 1 of 8
Why is Cohen's kappa preferred over raw percent agreement when calibrating an LLM judge?
Keep testing yourself
More quizzes picked for your level.