Statistics / Agreement

Cohen's Kappa Calculator (Inter-Rater Agreement)

Measure agreement between two raters beyond chance with Cohen's kappa, or weighted kappa with linear or quadratic weights for ordered categories, from a table of how often each pair of ratings occurred.

Cohen's Kappa Calculator (Inter-Rater Agreement): Observed agreement is the share of items on the diagonal; chance agreement is the sum, over categories, of the two raters' shares multiplied. κ = (observed − chance) ÷ (1 − chance). For the example table, observed agreement is 73.8%, chance 34.8%, and κ = 0.597, moderate. Weighted kappa credits near misses on ordered scales. Results match statsmodels. Runs 100% locally in your browser with zero server file uploads.

Runs
In your browser
Cost
Free · no sign-up
Availability
Ready to use
Cohen's kappa calculatorLocal processing

Runs entirely in your browser

Cohen's kappa0.5971moderate
Observed agreement73.8%
Agreement expected by chance34.8%

Kappa measures how much two raters agree beyond what chance alone would give: κ = (observed − expected) ÷ (1 − expected), where expected agreement comes from each rater's overall category shares. For ordered categories, weighted kappa gives partial credit to near misses. Landis and Koch's labels call 0.41–0.60 moderate, 0.61–0.80 substantial, and above 0.80 almost perfect; they are conventions, not tests.

Worked example

Two clinicians rate 80 cases as mild, moderate, or severe and agree on 59 (73.8%). Chance agreement from their category totals is 34.8%, so κ = (0.738 − 0.348) ÷ (1 − 0.348) = 0.597. Linear weights give 0.631 and quadratic 0.671.

To test whether two categorical variables are related rather than how well raters agree, use the chi-square calculator.

Kappa's quirks

Kappa can be low even with high agreement when one category is very rare or very common, the so-called kappa paradox; report the agreement table too.

For more than two raters, Fleiss' kappa is the usual extension.

How to use it

  1. Enter the agreement table: rater A's categories as rows and rater B's as columns, counts in the cells.
  2. Choose unweighted for plain categories, or linear or quadratic weights for ordered ones.
  3. Read kappa, its label, and the observed and chance agreement.

Privacy & limitations

Your data stay in your browser.

Related tools

Frequently asked questions

Why not just report percentage agreement?

Because raters agree some of the time by chance, especially when one category is common; kappa removes that part.

What is a good kappa?

Landis and Koch call 0.41–0.60 moderate, 0.61–0.80 substantial, and 0.81–1 almost perfect; requirements vary by field.

When should I use weighted kappa?

For ordered categories, such as mild, moderate, and severe, where rating moderate instead of severe is a smaller disagreement than mild instead of severe.

Free tool · runs in your browser · no account required