VolgendeVolgende gids
The Dual-LLM Pattern Against Prompt Injection
Technisch
Technische GIDS
Calibrating an LLM judge means evaluating how its decisions compare with human judgments under the same rubric and test cases.
Agreement statistics can help characterize a categorical judge, but high agreement alone does not show that the rubric is valid or that the judge is correct.
Using an LLM as a judge is attractive because it scales far beyond what human review can cover, but a judge that has never been checked against human judgment is essentially untested. Calibration starts by collecting a set of examples that have already been rated by humans - ideally by more than one rater, so inter-human agreement provides context for interpreting judge-human agreement - and then running the LLM judge on those same examples using its intended rubric and prompt. The core comparison typically uses Cohen's kappa, a statistic that measures agreement between two raters while correcting for the agreement expected by pure chance, which matters because raw percent agreement can look deceptively high when one category is very common. A confusion matrix, which cross-tabulates human labels against judge labels for each category, then shows exactly where disagreement is concentrated - for example, a judge might agree with humans well on clear passes and clear failures but disagree heavily in a middle 'borderline' category. That pattern has several possible causes, including unclear rubric boundaries, judge-specific errors, or disagreement among human raters. Inspect examples before deciding whether to revise the rubric, add examples, adjust the judge, or adjudicate the reference labels. A common misconception is that a single round of calibration is sufficient forever; judges can drift as the underlying model, prompt, or the distribution of inputs being judged changes over time, so recalibration on fresh human-labeled samples on a regular cadence is standard practice. Another misconception is that perfect agreement is the bar to hit; human-human agreement is a useful reference for the task’s subjectivity, not a universal target or strict ceiling for judge performance.
Architectuurbeslissingen bepalen jarenlang de prestaties en bedrijfskosten.
Technisch onderwijs helpt teams bij het kiezen van de juiste stapel, niet alleen de nieuwste.
Betere technische keuzes verminderen het aantal betrouwbaarheidsincidenten in de productie.
LLM judges are used for more automated evaluation, but their performance can shift with model versions, prompts, answer style, and task mix. Recalibrate when any of those inputs change and keep a human-reviewed holdout. Report agreement with uncertainty and category-level error patterns; do not describe a match rate as accuracy unless the reference labels and evaluation design justify that interpretation. Human review remains essential for contested or high-impact judgments. Repeat checks after dataset or rubric revisions for each important release.
A team building an LLM judge to grade customer-support responses as 'helpful' or 'not helpful' collects 200 human-labeled examples, runs the judge on the same examples, and computes Cohen's kappa to check agreement beyond chance.
An engineer inspects a confusion matrix and discovers the judge systematically rates borderline-acceptable answers as 'excellent,' revealing a rubric that doesn't clearly define the boundary between the two categories.
A company tightens its judge's prompt by adding two concrete example answers for each rating level after finding human-judge disagreement concentrated on mid-range scores rather than clear passes or failures.
A team re-runs its calibration check every quarter on a fresh sample of human-labeled data, since a judge that agreed well with humans six months ago has started drifting after several unrelated prompt updates to the underlying model.
Het optimaliseren van één benchmark kan bredere systeemzwakheden verbergen.
Infrastructuur- en onderhoudskosten worden vaak onderschat.
De lacunes op het gebied van beveiliging en waarneembaarheid kunnen groter worden naarmate systemen complexer worden.
Definieer latentie-, kwaliteits- en kostendoelen vóór implementatie.
Benchmark onder realistische belasting- en gegevensomstandigheden.
Instrumentbewaking op fouten, drift en gebruikersimpact.
Bereid rollback- en incidentresponspaden voor voordat u gaat schalen.
Free newsletter
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
Calibrating an LLM judge means evaluating how its decisions compare with human judgments under the same rubric and test cases. Agreement statistics can help characterize a categorical judge, but high agreement alone does not show that the rubric is valid or that the judge is correct.
Raw percent agreement can look high simply because one category dominates; kappa adjusts for that.
A confusion matrix breaks down agreement and disagreement by category, pinpointing where the judge diverges from humans.
A concentrated mismatch can reflect unclear category boundaries, judge-specific behavior, or inconsistent human labels; inspect the cases before choosing a fix.
A new model, prompt, rubric, or task mix can change judge behavior, so calibration must be revisited when conditions change.
Agreement measures similarity under the chosen rubric and reference labels; it does not establish that the rubric captures the intended construct or that its labels are correct.
Blijf leren
Er zijn meer handleidingen voor dit onderwerp geselecteerd
VolgendeVolgende gids
The Dual-LLM Pattern Against Prompt Injection
Technisch