OkulandelayoUmhlahlandlela olandelayo
The Dual-LLM Pattern Against Prompt Injection
Ubuchwepheshe
UMHLAHLANDLELA Wobuchwepheshe
Calibrating an LLM judge means evaluating how its decisions compare with human judgments under the same rubric and test cases.
Agreement statistics can help characterize a categorical judge, but high agreement alone does not show that the rubric is valid or that the judge is correct.
Using an LLM as a judge is attractive because it scales far beyond what human review can cover, but a judge that has never been checked against human judgment is essentially untested. Calibration starts by collecting a set of examples that have already been rated by humans - ideally by more than one rater, so inter-human agreement provides context for interpreting judge-human agreement - and then running the LLM judge on those same examples using its intended rubric and prompt. The core comparison typically uses Cohen's kappa, a statistic that measures agreement between two raters while correcting for the agreement expected by pure chance, which matters because raw percent agreement can look deceptively high when one category is very common. A confusion matrix, which cross-tabulates human labels against judge labels for each category, then shows exactly where disagreement is concentrated - for example, a judge might agree with humans well on clear passes and clear failures but disagree heavily in a middle 'borderline' category. That pattern has several possible causes, including unclear rubric boundaries, judge-specific errors, or disagreement among human raters. Inspect examples before deciding whether to revise the rubric, add examples, adjust the judge, or adjudicate the reference labels. A common misconception is that a single round of calibration is sufficient forever; judges can drift as the underlying model, prompt, or the distribution of inputs being judged changes over time, so recalibration on fresh human-labeled samples on a regular cadence is standard practice. Another misconception is that perfect agreement is the bar to hit; human-human agreement is a useful reference for the task’s subjectivity, not a universal target or strict ceiling for judge performance.
Izinqumo zezakhiwo ziqhuba ukusebenza kanye nezindleko zokusebenza iminyaka.
Imfundo yobuchwepheshe isiza amaqembu ukuthi akhethe isitaki esifanele, hhayi nje esisha.
Izinketho ezingcono zobunjiniyela zinciphisa izehlakalo ezinokwethenjelwa ekukhiqizeni.
LLM judges are used for more automated evaluation, but their performance can shift with model versions, prompts, answer style, and task mix. Recalibrate when any of those inputs change and keep a human-reviewed holdout. Report agreement with uncertainty and category-level error patterns; do not describe a match rate as accuracy unless the reference labels and evaluation design justify that interpretation. Human review remains essential for contested or high-impact judgments. Repeat checks after dataset or rubric revisions for each important release.
A team building an LLM judge to grade customer-support responses as 'helpful' or 'not helpful' collects 200 human-labeled examples, runs the judge on the same examples, and computes Cohen's kappa to check agreement beyond chance.
An engineer inspects a confusion matrix and discovers the judge systematically rates borderline-acceptable answers as 'excellent,' revealing a rubric that doesn't clearly define the boundary between the two categories.
A company tightens its judge's prompt by adding two concrete example answers for each rating level after finding human-judge disagreement concentrated on mid-range scores rather than clear passes or failures.
A team re-runs its calibration check every quarter on a fresh sample of human-labeled data, since a judge that agreed well with humans six months ago has started drifting after several unrelated prompt updates to the underlying model.
Ukuthuthukisa ibhentshimakhi eyodwa kungafihla ubuthakathaka obubanzi besistimu.
Izindleko zengqalasizinda nezokulungisa zivame ukubukelwa phansi.
Izikhala zokuphepha nokubonakala zingakhula njengoba izinhlelo ziba nzima kakhulu.
Chaza ukubambezeleka, ikhwalithi, nezindleko ezihlosiwe ngaphambi kokuqaliswa.
Ibhentshimakhi ngaphansi komthwalo wangempela nezimo zedatha.
Ukuqapha amathuluzi amaphutha, ukukhukhuleka, nomthelela wabasebenzisi.
Lungiselela izindlela zokuhlehlisa nezigameko ngaphambi kokukala.
Free newsletter
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
Calibrating an LLM judge means evaluating how its decisions compare with human judgments under the same rubric and test cases. Agreement statistics can help characterize a categorical judge, but high agreement alone does not show that the rubric is valid or that the judge is correct.
Raw percent agreement can look high simply because one category dominates; kappa adjusts for that.
A confusion matrix breaks down agreement and disagreement by category, pinpointing where the judge diverges from humans.
A concentrated mismatch can reflect unclear category boundaries, judge-specific behavior, or inconsistent human labels; inspect the cases before choosing a fix.
A new model, prompt, rubric, or task mix can change judge behavior, so calibration must be revisited when conditions change.
Agreement measures similarity under the chosen rubric and reference labels; it does not establish that the rubric captures the intended construct or that its labels are correct.
Qhubeka ufunda
Imihlahlandlela eyengeziwe yalesi sihloko
OkulandelayoUmhlahlandlela olandelayo
The Dual-LLM Pattern Against Prompt Injection
Ubuchwepheshe