Statistics / Agreement
Cohen's Kappa Calculator (Inter-Rater Agreement)
Measure agreement between two raters beyond chance with Cohen's kappa, or weighted kappa with linear or quadratic weights for ordered categories, from a table of how often each pair of ratings occurred.
Cohen's Kappa Calculator (Inter-Rater Agreement): Observed agreement is the share of items on the diagonal; chance agreement is the sum, over categories, of the two raters' shares multiplied. κ = (observed − chance) ÷ (1 − chance). For the example table, observed agreement is 73.8%, chance 34.8%, and κ = 0.597, moderate. Weighted kappa credits near misses on ordered scales. Results match statsmodels. Runs 100% locally in your browser with zero server file uploads.
- Category
- School & study tools
- Runs
- In your browser
- Cost
- Free · no sign-up
- Availability
- Ready to use
Runs entirely in your browser
Kappa measures how much two raters agree beyond what chance alone would give: κ = (observed − expected) ÷ (1 − expected), where expected agreement comes from each rater's overall category shares. For ordered categories, weighted kappa gives partial credit to near misses. Landis and Koch's labels call 0.41–0.60 moderate, 0.61–0.80 substantial, and above 0.80 almost perfect; they are conventions, not tests.
Worked example
Two clinicians rate 80 cases as mild, moderate, or severe and agree on 59 (73.8%). Chance agreement from their category totals is 34.8%, so κ = (0.738 − 0.348) ÷ (1 − 0.348) = 0.597. Linear weights give 0.631 and quadratic 0.671.
To test whether two categorical variables are related rather than how well raters agree, use the chi-square calculator.
Kappa's quirks
Kappa can be low even with high agreement when one category is very rare or very common, the so-called kappa paradox; report the agreement table too.
For more than two raters, Fleiss' kappa is the usual extension.
How to use it
- Enter the agreement table: rater A's categories as rows and rater B's as columns, counts in the cells.
- Choose unweighted for plain categories, or linear or quadratic weights for ordered ones.
- Read kappa, its label, and the observed and chance agreement.
Privacy & limitations
Your data stay in your browser.
Related tools
Frequently asked questions
Why not just report percentage agreement?
Because raters agree some of the time by chance, especially when one category is common; kappa removes that part.
What is a good kappa?
Landis and Koch call 0.41–0.60 moderate, 0.61–0.80 substantial, and 0.81–1 almost perfect; requirements vary by field.
When should I use weighted kappa?
For ordered categories, such as mild, moderate, and severe, where rating moderate instead of severe is a smaller disagreement than mild instead of severe.
Free tool · runs in your browser · no account required