기술 가이드

Cohen's Kappa

Cohen's kappa measures agreement between two sets of categorical labels after accounting for agreement expected from their label frequencies.

  • 3분 읽기
  • 마지막 업데이트
이 페이지에서3분 읽기
  1. 개요
  2. 심층 분석
  3. 전략적 영향
  4. The Future of Cohen's Kappa
  5. 실제 구현
  6. 위험 및 가드레일
  7. 구현 로드맵
  8. 계속 탐색하세요
  9. 자주 묻는 질문

개요

It helps teams assess annotation consistency, but a high score does not establish that either annotator is factually correct.

심층 분석

Suppose two annotators agree on most examples. That sounds reassuring, but the category frequencies matter. If both almost always choose the same common category, frequent agreement can arise without much ability to distinguish cases. Cohen's kappa compares observed agreement with an expected-agreement calculation based on each annotator's label proportions. The formula subtracts expected agreement from observed agreement, then divides by one minus expected agreement. If observed agreement is 0.8 and expected agreement is 0.5, kappa is 0.6. These values form an illustrative calculation. A kappa of one represents perfect agreement when the denominator is defined; zero means agreement matches the frequency-based expectation. Negative values indicate less agreement than that expectation. The expected term does not prove that either person guessed randomly. It is a mathematical reference constructed from the marginal label frequencies. For example, if one annotator uses a category 60% of the time and the other uses it 50% of the time, that category contributes 0.3 to expected agreement. Add the contributions for every category. Use independently assigned labels on the same cases. If the second annotator copies the first, their agreement says little about independent reliability. Examine disagreement patterns before changing the instructions, then evaluate revised instructions on a fresh sample where practical. Kappa also depends on category prevalence and the annotators' use of labels. There is no universally appropriate cutoff that makes a dataset trustworthy. Report sample size, raw agreement and the agreement table alongside the score. Scikit-learn offers unweighted and weighted calculations. Weighted kappa is suitable when category order and the consequences of different-sized disagreements have been defined clearly.

전략적 영향

비용 및 예산

아키텍처 결정은 수년 동안 성능과 운영 비용을 결정합니다.

더 명확한 결정들

기술 교육은 팀이 최신 스택뿐만 아니라 올바른 스택을 선택하는 데 도움이 됩니다.

품질 관리

더 나은 엔지니어링 선택은 생산 시 신뢰성 사고를 줄입니다.

The Future of Cohen's Kappa

AI-assisted annotation makes it more important to record how each label was produced. Two reviewers who saw the same model suggestion may share an error, even when their agreement is high. Teams can strengthen their process by preserving a sample labeled independently, keeping a record of instruction changes and discussing recurring disagreements with domain experts. Kappa can remain one part of that review, alongside checks against well-supported reference answers. The goal is a dataset whose labels have a clear meaning and a defensible review process, rather than a single impressive agreement number.

실제 구현

Two people independently label the same support tickets as billing or technical. A kappa report measures their agreement and a disagreement review reveals where the labeling instructions need clarification.

In a hypothetical sample, observed agreement is 80% and expected agreement is 50%. Kappa is (0.8 minus 0.5) divided by (1 minus 0.5), or 0.6.

Reviewers assign low, medium or high severity to incidents. A weighted kappa can distinguish a one-level disagreement from a disagreement between the lowest and highest levels.

An analyst uses scikit-learn's cohen_kappa_score for two aligned label arrays, then reports the category counts and agreement table so readers can interpret the summary.

위험 및 가드레일

  • 하나의 벤치마크를 최적화하면 더 광범위한 시스템 약점을 숨길 수 있습니다.

  • 인프라 및 유지 관리 비용은 종종 과소평가됩니다.

  • 시스템이 더욱 복잡해짐에 따라 보안 및 관찰 가능성의 격차가 커질 수 있습니다.

구현 로드맵

  1. 구현하기 전에 지연 시간, 품질, 비용 목표를 정의하세요.

  2. 현실적인 로드 및 데이터 조건에서 벤치마킹합니다.

  3. 오류, 드리프트 및 사용자 영향에 대한 계측기 모니터링.

  4. 확장하기 전에 롤백 및 사고 대응 경로를 준비하세요.

계속 탐색하세요

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the Cohen's Kappa quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

퀴즈 시작

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

자주 묻는 질문

What is Cohen's Kappa?

Cohen's kappa measures agreement between two sets of categorical labels after accounting for agreement expected from their label frequencies. It helps teams assess annotation consistency, but a high score does not establish that either annotator is factually correct.

Two annotators have observed agreement of 0.8 and expected agreement of 0.5. Which kappa value follows?

Kappa is (0.8 minus 0.5) divided by (1 minus 0.5), which equals 0.6.

Why does the expected-agreement term in kappa not establish that annotators were guessing randomly?

The expected term is a mathematical baseline constructed from marginal frequencies; it does not describe the annotators' actual thought process.

For low, medium and high incident severity, why might a team choose weighted kappa?

Weights allow the score to reflect the distance between ordered categories, provided that order and weighting are appropriate.

Both annotators assign every case to the same single category. Why should ordinary kappa not be reported as a straightforward reliability success?

The denominator is one minus expected agreement. When expected agreement is one, the ordinary formula is undefined.

A second reviewer copies the first reviewer's labels and achieves perfect agreement. Which conclusion is justified?

Copied decisions do not provide an independent assessment, even if their numerical agreement is perfect.