गाइड से जुड़ी प्रश्नोत्तरी · कठिन स्तर

Calibrating LLM Judges Against Humans Quiz

Calibrate LLM judges with shared rubrics, independent human labels, confusion matrices, suitable agreement statistics, and repeated held-out checks.

संबंधित मार्गदर्शक पथCalibrating LLM Judges Against Humans
प्रश्न 1 का 8

Why is Cohen's kappa preferred over raw percent agreement when calibrating an LLM judge?