UK-based online statistics and data analysis support for USA, UK, and international clients. No exams, no impersonation, no fabricated data.
Chi-Square and Categorical Data Tests

Weighted Kappa: Formula, Interpretation, Example, Python, R, SPSS, SAS and Excel Guide

Distance-sensitive agreement for ordered categories Weighted Kappa: Formula, Interpretation, Example, Python, R, SPSS, SAS and Excel Guide Weighted Kappa measures agreement between two ordinal classifications while...

Statistics guide Ethical learning support SPSS/R/Python/Excel friendly
Weighted Kappa: Formula, Interpretation, Example, Python, R, SPSS, SAS and Excel Guide

Distance-sensitive agreement for ordered categories

Weighted Kappa: Formula, Interpretation, Example, Python, R, SPSS, SAS and Excel Guide

Weighted Kappa measures agreement between two ordinal classifications while giving small disagreements less penalty than large disagreements. This worked analysis compares four-level G1 and G3 grade bands for 649 students, reconstructs every linear and quadratic weighting calculation, separates exact agreement from chance agreement, explains the direction of category movement, and provides verified Python, R, SPSS and Excel outputs.

649 valid paired ratingsExact agreement 71.34%Linear κ = 0.6474Quadratic κ = 0.7652648/649 within one category

Weighted Kappa Model Overview

Weighted Kappa, often called weighted Cohen’s kappa, evaluates agreement between two ratings measured on the same ordered scale. It improves on ordinary Cohen’s kappa by recognising that a one-category disagreement is less serious than a two- or three-category disagreement. The present analysis compares G1 and G3 grade classifications using four ordered bands: low, medium, high and very high.

The complete 4 × 4 table contains 649 paired classifications. The diagonal contains 463 exact matches. Of the 186 disagreements, 185 are only one category apart, 1 is two categories apart and 0 are three categories apart. That concentration near the diagonal is the central reason quadratic weighted kappa (0.765160) is higher than linear weighted kappa (0.647398).

Exact Research Question

How strongly do G1 and G3 ordinal grade bands agree beyond the agreement expected from their different marginal distributions, and how does the result change when disagreements are penalised linearly rather than quadratically?

H0: weighted observed disagreement equals weighted disagreement expected by chance
H1: weighted observed disagreement is lower than chance-weighted disagreement
Result componentVerified valueInterpretive role
Unweighted κ0.544050Treats every mismatch equally.
Linear weighted κ0.647398Penalty increases in direct proportion to category distance.
Quadratic weighted κ0.765160Penalty increases with squared category distance; primary result.
Exact agreement463/649 (71.34%)Raw diagonal agreement before chance correction.
Within one category648/649 (99.85%)Shows how rarely the ratings are far apart.
Verified analysis source: the linked Python and R reports independently reproduce linear weighted κ = 0.6473984073 and quadratic weighted κ = 0.7651598550. The SPSS output confirms the same 4 × 4 table and reports ordinary unweighted κ = .544; the weighted values are then calculated from the published matrix and disagreement weights.
Interpretation boundary: weighted kappa measures agreement, not correlation, predictive accuracy, causality or equality of marginal distributions. Two variables can be strongly correlated yet disagree systematically, and weighted kappa can be high even when one rating tends to be slightly higher if most pairs remain close.
AdvertisementGoogle AdSense top placement reserved here

Quick Answer: Weighted Kappa Result

The primary quadratic weighted kappa is 0.765160. Under quadratic disagreement weights, observed disagreement is 0.032357, compared with 0.137785 expected from the row and column margins. The calculation therefore removes 76.52% of the weighted disagreement expected by chance.

Valid paired ratings649
Exact matches463
Linear weighted κ0.6474
Quadratic weighted κ0.7652

Statistical Result

  • Observed quadratic disagreement: 0.032357
  • Expected quadratic disagreement: 0.137785
  • Quadratic weighted κ: 0.765160
  • Supplementary bootstrap 95% CI: 0.7289 to 0.7974
  • Exact agreement: 71.34%

Data Meaning

  • 185 disagreements are adjacent.
  • Only 1 disagreement spans two bands.
  • No pair spans all three possible category steps.
  • 153 pairs move to a higher G3 band; 33 move lower.
  • Among disagreements, 82.26% move upward.
Plain-language conclusion: G1 and G3 grade bands show strong distance-sensitive agreement. Most students remain in the same band, and nearly every mismatch is only one band away. Quadratic weighting appropriately recognises that the observed errors are overwhelmingly minor rather than severe.

Table of Contents

  1. What weighted kappa measures
  2. Research question and hypotheses
  3. When to use weighted kappa
  4. Linear, quadratic and custom weights
  5. Variables and ordinal coding
  6. Formula and manual calculation
  7. Observed, expected and residual tables
  8. Complete weighted kappa results
  9. Assumptions and diagnostic checks
  10. Comparison with related methods
  11. Calculator workflow
  12. Python chart interpretations
  13. Paired R chart interpretations
  14. Python, R, SPSS, SAS and Excel
  15. Expandable code
  16. APA interpretation and reporting
  17. Common mistakes
  18. Practice questions
  19. Reports and worked files
  20. Related guides
  21. Frequently asked questions
  22. Conclusion

What Is Weighted Kappa?

Weighted Kappa is a chance-corrected agreement coefficient for two ratings recorded on the same ordinal scale. It starts with a cross-classification table, assigns each cell a disagreement weight, and compares the observed weighted disagreement with the disagreement expected if the two marginal distributions were independent.

Ordinary kappa treats low-versus-medium and low-versus-very-high disagreements as equally wrong. Weighted kappa does not. In this four-band example, a low-to-medium disagreement receives a linear penalty of 1/3 or a quadratic penalty of 1/9, whereas low-to-very-high receives the maximum penalty of 1 under both systems.

1 — LowGrade below 10
2 — MediumGrade from 10 to 13
3 — HighGrade from 14 to 16
4 — Very highGrade 17 or above
QuantityMeaningWorked value
Observed exact agreementProportion on the diagonal463/649 = 0.713405
Chance exact agreementDiagonal agreement implied by margins0.371433
Unweighted kappaChance-corrected exact agreement0.544050
Linear weighted kappaChance-corrected agreement with proportional penalties0.647398
Quadratic weighted kappaChance-corrected agreement with squared penalties0.765160

Why Weighted Kappa Is Not a Reliability Percentage

The value 0.7652 is not “76.52% of ratings are correct.” It means the observed quadratic disagreement is 76.52% lower than the disagreement expected from the marginal distributions. Raw exact agreement is 71.34%, while the weighted agreement score is 96.76%. These quantities answer different questions and should be reported separately.

Why Marginal Distributions Matter

G1 has 157 low, 340 medium, 128 high and 24 very-high classifications. G3 has 100, 355, 148 and 46, respectively. Because the G3 distribution is shifted upward, raw agreement alone would not fully describe the agreement process. Weighted kappa retains the paired table and corrects for the agreement expected from these margins.

Research Question, Hypotheses and Data Design

First ratingG1 grade band with four ordered levels
Second ratingG3 grade band with the same cut points
Unit of analysis649 students with both classifications

Research Question

Do G1 and G3 place students in the same or nearby ordinal grade bands more often than expected from their separate grade-band distributions?

ComponentFormal statementMeaning in this analysis
Null hypothesisκw = 0Observed weighted disagreement equals chance-weighted disagreement.
Alternative hypothesisκw > 0Observed weighted disagreement is smaller than chance expectation.
Primary weightingQuadraticDistant disagreements receive disproportionately larger penalties.
Secondary weightingLinearEach additional category step adds the same penalty.
Direction analysisG3 − G1Positive cells above the diagonal indicate movement to a higher final-grade band.

Prespecified Analysis Decisions

The grade cut points, category order, pairing rule and weighting scheme should be selected before inspecting the kappa result. Quadratic weights are primary here because a two-band classification error is intended to count four times as much as a one-band error, while a three-band error receives the maximum penalty.

Methods sentence: “Agreement between G1 and G3 four-level ordinal grade bands was evaluated with linear and quadratic weighted Cohen’s kappa. Quadratic weights were primary, with disagreement weights defined as squared category distance divided by nine.”

When to Use Weighted Kappa

Use Weighted Kappa When

  • The same cases receive two ratings or classifications.
  • Both ratings use the same ordered categories.
  • Near disagreements are less serious than distant disagreements.
  • The goal is agreement beyond chance, not merely correlation.
  • Each case contributes one independent pair.
  • The category order and weights are defensible before analysis.
  • The complete cross-classification table can be inspected.

Choose Another Method When

  • Categories are nominal with no defensible order.
  • More than two raters classify the same cases.
  • The ratings are continuous and measurement error is the focus.
  • Repeated ratings create clustering or longitudinal dependence.
  • Different raters use different category systems.
  • Prevalence, bias or marginal homogeneity is the main question.
  • The purpose is prediction rather than agreement.

Decision Flow

Step 1Same cases rated twice?

If no, agreement kappa is not appropriate.

Step 2Are categories ordered?

If no, use ordinary nominal kappa.

Step 3How should distance matter?

Select linear, quadratic or justified custom weights.

For two nominal ratings, use the ordinary kappa statistic. For more than two raters, consider Fleiss’ kappa. For paired binary change rather than agreement, use McNemar’s test; for paired multicategory marginal change, use the McNemar–Bowker test.

Linear, Quadratic and Custom Weighted Kappa

The numerical value of weighted kappa depends on the weight matrix. Weighting is part of the estimand, not a decorative software option.

Linear Disagreement Weights

wij = |i − j| / (k − 1)
G1 \ G3Low (<10)Medium (10–13)High (14–16)Very high (17+)
Low (<10)0.0000000.3333330.6666671.000000
Medium (10–13)0.3333330.0000000.3333330.666667
High (14–16)0.6666670.3333330.0000000.333333
Very high (17+)1.0000000.6666670.3333330.000000

Quadratic Disagreement Weights

wij = (i − j)² / (k − 1)²
G1 \ G3Low (<10)Medium (10–13)High (14–16)Very high (17+)
Low (<10)0.0000000.1111110.4444441.000000
Medium (10–13)0.1111110.0000000.1111110.444444
High (14–16)0.4444440.1111110.0000000.111111
Very high (17+)1.0000000.4444440.1111110.000000
SchemeOne-step penaltyTwo-step penaltyThree-step penaltyWorked κ
Unweighted1110.544050
Linear0.3333330.6666671.0000000.647398
Quadratic0.1111110.4444441.0000000.765160

Why Quadratic Kappa Is Higher Here

There are 185 adjacent disagreements, only 1 two-step disagreement and no three-step disagreement. Quadratic weights sharply discount adjacent errors: a one-step error contributes only 1/9 of a maximum error instead of 1/3. Because almost all mismatches are adjacent, quadratic observed disagreement falls to 0.032357, compared with linear observed disagreement of 0.096045.

Do not choose the larger result after seeing the data. The weighting system must reflect the practical loss attached to category distance. Reporting both linear and quadratic values is useful, but one scheme should be identified as primary in advance.

Variables, Ordinal Coding and Data Dictionary

VariableRoleCodingValid NInterpretation
G1First ordinal classification1 = low; 2 = medium; 3 = high; 4 = very high649Earlier grade-band classification.
G3Second ordinal classificationSame four cut points649Final grade-band classification.
Distance |G3−G1|Disagreement severity0, 1, 2 or 3649Number of category steps separating the pair.
Linear weightDisagreement penaltydistance / 3649Equal penalty added per category step.
Quadratic weightDisagreement penaltydistance² / 9649Large errors penalised disproportionately.

Marginal Grade-Band Distributions

BandG1 countG1 %G3 countG3 %G3 − G1 count
Low (<10)15724.19%10015.41%-57
Medium (10–13)34052.39%35554.70%+15
High (14–16)12819.72%14822.80%+20
Very high (17+)243.70%467.09%+22

Scale-Score Summary

Descriptive quantityValueInterpretation
Mean G1 category score2.029276Average location on the 1–4 ordinal coding.
Mean G3 category score2.215716G3 classifications are somewhat higher overall.
Mean category shift+0.186441Average G3 band exceeds G1 by 0.186 category steps.
Mean absolute category distance0.288136Average pair is 0.288 bands apart.
Root mean squared distance0.539645Gives more weight to the rare larger discrepancy.
Ordinal caution: category scores 1–4 provide a transparent summary of direction and distance, but they do not imply that the underlying grade intervals are psychologically or educationally equal. Weighted kappa relies on the selected category-distance structure.

Category Retention and Movement from Each G1 Band

G1 bandRow NExact NExact %Moved lower NMoved lower %Moved higher NMoved higher %
Low (<10)1578856.05%00.00%6943.95%
Medium (10–13)34026778.53%123.53%6117.94%
High (14–16)1288667.19%1914.84%2317.97%
Very high (17+)242291.67%28.33%00.00%

The very-high band has the highest exact retention at 91.67%, although it contains only 24 cases. The low band has the lowest exact retention at 56.05%; all 69 low-band mismatches move one step upward to medium. Medium is the largest row and retains 267 of 340 students, while 61 move upward and 12 move downward.

Where Each G3 Band Came From

G3 bandFrom Low (<10)From Medium (10–13)From High (14–16)From Very high (17+)Column total
Low (<10)88.00%12.00%0.00%0.00%100.00%
Medium (10–13)19.44%75.21%5.35%0.00%100.00%
High (14–16)0.00%40.54%58.11%1.35%100.00%
Very high (17+)0.00%2.17%50.00%47.83%100.00%

Column percentages reveal a different perspective from row retention. Of the 100 students classified low at G3, 88.00% were already low at G1 and 12.00% came from medium. Among the 46 very-high G3 classifications, 47.83% were already very high, 50.00% moved up from high, and 2.17% moved two bands from medium.

Weighted Kappa Formula and Manual Calculation

This article uses disagreement weights. With cell proportion pij, row marginal pi+, column marginal p+j and disagreement weight wij, weighted kappa is:

κw = 1 − [ΣiΣjwijpij] / [ΣiΣjwijpi+p+j]
SymbolDefinition
pijObserved proportion in row i and column j.
pi+Marginal proportion for G1 category i.
p+jMarginal proportion for G3 category j.
wijDisagreement penalty for cell i,j.
DoObserved weighted disagreement Σwijpij.
DeExpected weighted disagreement Σwijpi+p+j.

Linear Calculation

Observed disagreementDo,L = 0.096045197740
Expected disagreementDe,L = 0.272390141524
Ratio Do/De0.352601592711
Linear κ1 − ratio = 0.647398407289

Quadratic Calculation

Observed disagreementDo,Q = 0.032357473035
Expected disagreementDe,Q = 0.137785100753
Ratio Do/De0.234840144969
Quadratic κ1 − ratio = 0.765159855031
MethodObserved weighted agreementExpected weighted agreementChance-corrected κ
Exact / unweighted71.34%37.14%0.544050
Linear90.40%72.76%0.647398
Quadratic96.76%86.22%0.765160

Manual Check from Disagreement Distances

For the linear system, the 185 one-step disagreements contribute 185 × (1/3) weighted count units and the one two-step disagreement contributes 1 × (2/3). Dividing by 649 gives Do,L = 0.096045. For quadratic weights, the same cells contribute 185 × (1/9) plus 1 × (4/9), producing Do,Q = 0.032357.

Cross-check: the matrix calculation reproduces the verified Python, R and SPSS-integrated values to all displayed digits.

Cell-by-Cell Disagreement Contributions

Observed transitionCountDistanceLinear weightLinear contributionQuadratic weightQuadratic contributionShare of quadratic Do
Low (<10) → Medium (10–13)6910.3333330.0354390.1111110.01181336.51%
Medium (10–13) → Low (<10)1210.3333330.0061630.1111110.0020546.35%
Medium (10–13) → High (14–16)6010.3333330.0308170.1111110.01027231.75%
Medium (10–13) → Very high (17+)120.6666670.0010270.4444440.0006852.12%
High (14–16) → Medium (10–13)1910.3333330.0097590.1111110.00325310.05%
High (14–16) → Very high (17+)2310.3333330.0118130.1111110.00393812.17%
Very high (17+) → High (14–16)210.3333330.0010270.1111110.0003421.06%

The contribution table separates frequency from severity. Low→medium is the largest quadratic contribution because it occurs 69 times. Medium→very-high spans two categories and receives a four-times-larger quadratic weight than an adjacent error, but it occurs once and therefore contributes only 0.000685. This is why a weighted analysis must multiply cell frequency by the selected penalty rather than ranking errors by distance alone.

Expected Disagreement Is Built from the Margins

Expected cell probability is pi+p+j. For example, the expected low-G1/medium-G3 probability is (157/649) × (355/649) = 0.132324. Under quadratic weights, multiplying that probability by 1/9 contributes 0.014703 to expected disagreement. Repeating this for all 16 cells gives De,Q = 0.137785100753.

Chance correction therefore does not assume equal category frequencies. It uses the actual G1 and G3 margins. This matters because medium dominates both distributions and G3 is shifted toward higher categories.

Observed Agreement, Expected Counts and Diagnostic Tables

Observed G1-by-G3 Agreement Matrix

G1 \ G3Low (<10)Medium (10–13)High (14–16)Very high (17+)Row total
Low (<10)886900157
Medium (10–13)12267601340
High (14–16)0198623128
Very high (17+)0022224
Column total10035514846649

The diagonal holds 463 exact matches. Blue cells above the diagonal represent movement to a higher G3 band; amber cells below the diagonal represent movement to a lower G3 band.

Row Percentages

G1 bandLow (<10)Medium (10–13)High (14–16)Very high (17+)Row total
Low (<10)56.05%43.95%0.00%0.00%100.00%
Medium (10–13)3.53%78.53%17.65%0.29%100.00%
High (14–16)0.00%14.84%67.19%17.97%100.00%
Very high (17+)0.00%0.00%8.33%91.67%100.00%

Exact retention is 56.05% for low, 78.53% for medium, 67.19% for high and 91.67% for very-high G1 classifications.

Expected Counts Under Independent Margins

G1 \ G3Low (<10)Medium (10–13)High (14–16)Very high (17+)Row total
Low (<10)24.19185.87835.80311.128157
Medium (10–13)52.388185.97877.53524.099340
High (14–16)19.72370.01529.1909.072128
Very high (17+)3.69813.1285.4731.70124

Expected diagonal counts total 241.060, corresponding to 37.14% chance exact agreement. Observed diagonal agreement is 463, substantially above that expectation.

Pearson Residuals for the Unweighted Table

G1 \ G3Low (<10)Medium (10–13)High (14–16)Very high (17+)
Low (<10)+12.973-1.821-5.984-3.336
Medium (10–13)-5.580+5.941-1.991-4.705
High (14–16)-4.441-6.097+10.515+4.624
Very high (17+)-1.923-3.623-1.485+15.564

Large positive diagonal residuals show that exact matches occur far more often than expected under independence. Negative residuals in distant off-diagonal cells show that severe disagreements are rarer than expected.

Disagreement Distance Distribution

Category distanceCountPercent of all pairsPercent of disagreementsInterpretation
0 — exact match46371.34%Same grade band.
1 — adjacent mismatch18528.51%99.46%One band apart.
2 — two-band mismatch10.15%0.54%Only one observed pair.
3 — maximum mismatch00.00%0.00%No observed pair.

Quadratic Observed Disagreement Contributions

G1 \ G3Low (<10)Medium (10–13)High (14–16)Very high (17+)
Low (<10)0.0000000.0118130.0000000.000000
Medium (10–13)0.0020540.0000000.0102720.000685
High (14–16)0.0000000.0032530.0000000.003938
Very high (17+)0.0000000.0000000.0003420.000000
Largest weighted contributions: low→medium contributes 0.011813, medium→high contributes 0.010272, high→very-high contributes 0.003938, and high→medium contributes 0.003253. Together these adjacent transitions account for nearly all observed quadratic disagreement.

Observed Versus Expected Exact Agreement by Band

BandObserved diagonalExpected diagonalObserved − expectedObserved/expected ratioDiagonal residual
Low (<10)8824.191+63.8093.638×+12.973
Medium (10–13)267185.978+81.0221.436×+5.941
High (14–16)8629.190+56.8102.946×+10.515
Very high (17+)221.701+20.29912.933×+15.564

Every diagonal cell exceeds its chance expectation. The strongest proportional concentration is in the very-high cell: 22 observed versus 1.701 expected. The low cell has 88 observed versus 24.191 expected. These large diagonal excesses are why both unweighted and weighted kappa are well above zero.

Marginal Shift from G1 to G3

BandG1 countG3 countCount changeG1 shareG3 shareShare change
Low (<10)157100-5724.19%15.41%-8.78 pp
Medium (10–13)340355+1552.39%54.70%+2.31 pp
High (14–16)128148+2019.72%22.80%+3.08 pp
Very high (17+)2446+223.70%7.09%+3.39 pp

The margins show a redistribution rather than identical prevalence. Low declines by 57 cases, while high rises by 20 and very high rises by 22. Medium rises by 15. Weighted kappa remains high because most individual transitions are exact or adjacent, but the margin table demonstrates that agreement does not imply an unchanged grade distribution.

Complete Weighted Kappa Results

Exact agreement71.34%

463 diagonal pairs

Unweighted κ0.5441

All mismatches equal

Linear weighted κ0.6474

Proportional distance penalty

Quadratic weighted κ0.7652

Squared distance penalty

Within one category99.85%

648 of 649

Net upward transitions+120

153 upward versus 33 downward

Primary Weighted Kappa Results

MeasureEstimateObserved disagreementExpected disagreementSupplementary bootstrap 95% CI
Linear weighted kappa0.6473984070.0960451980.2723901420.6007 to 0.6909
Quadratic weighted kappa0.7651598550.0323574730.1377851010.7289 to 0.7974

Agreement Decomposition

ComponentCountPercent of all pairsAnalytical meaning
Exact match46371.34%Raw agreement before chance correction.
Adjacent disagreement18528.51%Minor one-band difference.
Two-band disagreement10.15%Rare moderate discrepancy.
Three-band disagreement00.00%No maximum discrepancy.
G3 higher than G115323.57%Upward classification movement.
G3 lower than G1335.08%Downward classification movement.

Bootstrap Precision Check

EstimateBootstrap SE2.5th percentile97.5th percentileResamples
Linear κ0.022900.600730.6908720,000
Quadratic κ0.017540.728940.7974420,000

The bootstrap is a supplementary reproducibility check generated directly from the verified 4 × 4 empirical table with a fixed seed. It is not presented as an output from the linked software reports. Both intervals remain well above zero and preserve the ordering quadratic κ > linear κ.

Directional Transition Detail

TransitionCountPercent of all pairsShare of disagreements
Low → Medium6910.63%37.10%
Medium → High609.24%32.26%
Medium → Very high10.15%0.54%
High → Very high233.54%12.37%
Medium → Low121.85%6.45%
High → Medium192.93%10.22%
Very high → High20.31%1.08%
Data conclusion: weighted agreement is strong because exact matches dominate and severe errors are virtually absent. At the same time, the asymmetry of 153 upward versus 33 downward movements shows a systematic upward shift from G1 to G3 that agreement coefficients alone do not summarise.
AdvertisementGoogle AdSense placement reserved after the Results section

Where the Quadratic Disagreement Comes From

Source groupQuadratic contributionPercent of observed quadratic disagreementInterpretation
Upward adjacent transitions0.02602380.42%Low→medium, medium→high and high→very high.
Downward adjacent transitions0.00565017.46%Medium→low, high→medium and very high→high.
Two-step upward transition0.0006852.12%One medium→very-high pair.
Three-step transitions0.0000000.00%No low↔very-high pair.

Upward adjacent transitions account for 80.42% of observed quadratic disagreement, while downward adjacent transitions account for 17.46%. The coefficient is therefore not only high; its remaining disagreement is directionally concentrated.

Three Complementary Conclusions from the Same Table

QuestionStatistic or evidenceConclusion
How often are classifications identical?463/649 = 71.34%Most pairs match exactly.
How close are classifications beyond chance?Quadratic κ = 0.765160Strong distance-sensitive agreement.
Did the distribution shift over time?153 upward vs 33 downward transitions; mean shift +0.186G3 is systematically higher despite strong agreement.

These conclusions are not contradictory. A student can remain close to the earlier band while still being more likely to move upward than downward. Reporting only weighted kappa would omit the third conclusion; reporting only the margins would omit the strong case-level pairing.

Weighted Kappa Assumptions and Diagnostic Checks

Weighted kappa requires a defensible paired ordinal design. The coefficient can be computed mechanically even when the design is inappropriate, so each condition should be checked explicitly.

Paired observations

Each of the 649 students contributes one G1 and one G3 classification.

Same ordered scale

Both variables use identical low, medium, high and very-high cut points.

Independent pairs

One student’s pair should not determine another student’s pair.

Prespecified weights

Quadratic weighting is chosen for substantive reasons, not because it gives a larger κ.

Complete table

All 16 cells and both margins are inspected, not only the final coefficient.

Adequate category support

The very-high G1 band has only 24 cases, so category-specific percentages need context.

DiagnosticObserved evidenceAssessment
Missing paired values649 valid, 0 missing in SPSS case summarySatisfied for the analysed table.
Scale alignmentIdentical four grade bands for G1 and G3Satisfied.
Distant disagreement support1 two-step and 0 three-step errorsQuadratic result is driven by adjacent-error discounting.
Marginal equalityG1 and G3 margins differNot required for kappa, but direction should be reported.
Sparse cellsSeveral distant cells are zeroSubstantively informative; bootstrap precision is useful.
Rater independenceG1 and G3 are repeated classifications, not independent ratersInterpret as temporal/measurement agreement, not inter-rater reliability.

Prevalence and Marginal Effects

Kappa can change when category prevalence or rater margins change, even with similar raw agreement. Here medium classifications dominate both margins, while G3 contains fewer low and more high/very-high classifications. Reporting the full margins, exact agreement and direction counts protects against overinterpreting one coefficient.

Do not call this “test–retest reliability” without design context. G1 and G3 represent different grading occasions. Their difference may reflect genuine academic change as well as classification consistency.

Sparse-Cell Interpretation

Seven of the 16 cells are zero. In this dataset, zero distant-error cells are substantively favourable because they indicate no low-to-high, low-to-very-high, high-to-low, very-high-to-low or very-high-to-medium movements. They also mean that uncertainty for rare error types cannot be learned directly from many observed cases. The supplementary bootstrap interval therefore resamples the full empirical table rather than applying a normal approximation to individual sparse cells.

Independence and Temporal Pairing

The pair is formed within student, so G1 and G3 are intentionally dependent. The independence assumption applies across students: one student’s grade-band pair should not be duplicated or nested in a way that makes pairs correlated. If students are clustered in classrooms or schools and inference must generalise beyond the observed sample, a multilevel ordinal model may be needed in addition to descriptive weighted kappa.

Weighted Kappa Versus Related Agreement and Change Methods

MethodData structureUses order?Chance correction?What it answers
Weighted kappaTwo paired ordinal ratingsYesYesHow strongly do ratings agree when distance matters?
Ordinary Cohen’s kappaTwo paired nominal ratingsNoYesHow strongly do exact categories agree beyond chance?
Fleiss’ kappaMultiple raters, nominal categoriesNoYesHow strongly do more than two raters agree?
McNemar’s testPaired binary outcomesNoNoDo paired marginal proportions differ?
McNemar–Bowker testPaired multicategory outcomesOptionalNoIs the square table marginally symmetric?
Spearman correlationTwo ordered/continuous variablesYesNoIs there monotonic association?
Intraclass correlationContinuous ratingsContinuousModel-basedHow reliable are continuous measurements?

Weighted Kappa Versus Ordinary Kappa

Ordinary κ is 0.5441; linear weighted κ is 0.6474; quadratic weighted κ is 0.7652. The increase is not evidence that ordinary kappa is wrong. It reflects a different loss function: ordinary kappa treats all 186 disagreements equally, while weighted kappa recognises that 185 of them are only one category apart.

Weighted Kappa Versus Marginal-Homogeneity Tests

Agreement and change are distinct. The 153 upward and 33 downward transitions suggest that G3 tends to be higher than G1. A McNemar–Bowker test would test symmetry of the paired square table, while weighted kappa quantifies closeness. A complete study can report both when both questions matter.

Weighted Kappa Versus Correlation

Correlation is unaffected by certain systematic shifts and does not require identical values. Weighted kappa rewards equality and near equality on a common scale. Use Spearman rank correlation for monotonic ordering and weighted kappa for category agreement.

What Different Analyses Would Say About This Exact Matrix

Analysis applied to the same tableKey worked quantityInformation retainedInformation lost
Exact agreement71.34%Diagonal match rateChance expectation and disagreement distance.
Ordinary kappa0.544050Chance-corrected exact agreementDifference between near and far errors.
Linear weighted kappa0.647398Ordered distance with proportional lossNonlinear severity preferences.
Quadratic weighted kappa0.765160Strong penalty for far errorsSpecific directional asymmetry.
Marginal comparisonLow −57; high +20; very high +22Distributional shiftIndividual pairing closeness.
Transition direction153 upward; 33 downwardDirection of changeChance-corrected agreement.
CorrelationNot the primary estimandMonotonic associationExact category equality.

Weighted Kappa Calculator: Step-by-Step Workflow

A weighted kappa calculator needs the full square count matrix and the intended weight system. Entering only the diagonal percentage is insufficient.

Step 1Enter the 4 × 4 counts

Use rows for G1 and columns for G3 in the same category order.

Step 2Select weights

Linear, quadratic or a justified custom matrix.

Step 3Inspect diagnostics

Margins, exact agreement, distance counts and directional imbalance.

Calculator inputWorked entryValidation rule
Number of categories4Rows and columns must match.
Observed matrix[[88,69,0,0],[12,267,60,1],[0,19,86,23],[0,0,2,22]]All counts nonnegative; total = 649.
Category orderLow, Medium, High, Very highSame order for both ratings.
Primary weightsQuadratic disagreementDiagonal = 0; maximum distance = 1.
Secondary weightsLinear disagreementUsed for sensitivity comparison.
Confidence method20,000-table bootstrapLabel as supplementary if not part of primary software output.

Manual Calculator Audit

Count total649
Diagonal total463
Linear Do0.096045
Quadratic Do0.032357

A trustworthy calculator should reproduce κL = 0.647398407 and κQ = 0.765159855. If the result is instead .544, the calculator is reporting ordinary unweighted kappa.

Common calculator trap: some tools label agreement weights rather than disagreement weights. The matrices are complements of one another, and both conventions can produce the same kappa only when the formula matches the convention.

Python Weighted Kappa Chart Interpretations

These six Python figures are interpreted with the exact matrix values. Each chart supports a different part of the agreement argument: counts, weights, contributions, sensitivity and final reporting.

1. Agreement Matrix Python

Weighted Kappa Agreement Matrix Python chart
Agreement Matrix for the G1-by-G3 four-band weighted kappa analysis.
Exact values

The matrix contains 463 diagonal agreements, 153 cells above the diagonal and 33 below. The largest off-diagonal cells are low→medium (69) and medium→high (60).

Statistical meaning

The dense diagonal and near-diagonal concentration visually support a strong weighted kappa. Exact agreement is 71.34%, and 99.85% of all pairs are no more than one category apart.

Next analytical step: Use the matrix before interpreting any single coefficient; it reveals both severity and direction.

2. Linear Disagreement Weights Python

Weighted Kappa Linear Disagreement Weights Python chart
Linear Disagreement Weights for the G1-by-G3 four-band weighted kappa analysis.
Exact values

Linear penalties are 0, 1/3, 2/3 and 1 for distances 0 through 3.

Statistical meaning

Applying these weights produces observed disagreement 0.096045, expected disagreement 0.272390 and κ = 0.647398.

Next analytical step: Check that category order is identical on both axes before applying the matrix.

3. Quadratic Disagreement Weights Python

Weighted Kappa Quadratic Disagreement Weights Python chart
Quadratic Disagreement Weights for the G1-by-G3 four-band weighted kappa analysis.
Exact values

Quadratic penalties are 0, 1/9, 4/9 and 1.

Statistical meaning

Because 185 of 186 mismatches are adjacent, quadratic observed disagreement falls to 0.032357; κ rises to 0.765160.

Next analytical step: Use quadratic weights only when larger classification errors genuinely deserve disproportionate penalty.

4. Weighted Contributions Python

Weighted Kappa Weighted Contributions Python chart
Weighted Contributions for the G1-by-G3 four-band weighted kappa analysis.
Exact values

The largest quadratic contributions are low→medium (0.011813) and medium→high (0.010272).

Statistical meaning

The single medium→very-high case contributes 0.000685; despite spanning two bands, it remains a small part of total observed disagreement because it occurs once.

Next analytical step: Investigate cells that are both frequent and heavily weighted; frequency alone does not determine contribution.

5. Weighting Comparison Python

Weighted Kappa Weighting Comparison Python chart
Weighting Comparison for the G1-by-G3 four-band weighted kappa analysis.
Exact values

Unweighted κ = 0.5441, linear κ = 0.6474, and quadratic κ = 0.7652.

Statistical meaning

The 0.1178 gap between quadratic and linear results quantifies sensitivity to how adjacent errors are discounted.

Next analytical step: Report the weighting definition with the coefficient so readers know which disagreement loss function was used.

6. Result Summary Python

Weighted Kappa Result Summary Python chart
Result Summary for the G1-by-G3 four-band weighted kappa analysis.
Exact values

The summary combines N = 649, exact agreement 71.34%, linear κ 0.6474 and quadratic κ 0.7652.

Statistical meaning

Observed quadratic agreement is 96.76% compared with 86.22% expected under the margins.

Next analytical step: Pair the summary with the full matrix and direction counts; the coefficient alone does not reveal the upward G3 shift.
AdvertisementGoogle AdSense placement reserved after the Python figures

R Weighted Kappa Charts in Paired Rows

The R figures are displayed in paired rows, matching the fresh permanent sample. Explanations remain separate so each visual has its own numerical interpretation.

R figure pair 1: Agreement Matrix and Linear Weights
Weighted Kappa Agreement Matrix R chart
Agreement Matrix generated for the R weighted kappa analysis.
Weighted Kappa Linear Weights R chart
Linear Weights generated for the R weighted kappa analysis.
Exact values

Agreement Matrix

The R agreement matrix reproduces all 16 counts and both margins. Diagonal agreement is 463/649.

The diagonal is strongest in the medium cell (267), while the most common movement is low→medium (69).

Cross-check against the verified matrix and the reported coefficient.
Exact values

Linear Weights

The linear matrix assigns 0.3333 to adjacent, 0.6667 to two-step and 1.0000 to maximum disagreements.

The resulting κ is 0.647398; observed weighted agreement is 90.40%.

Confirm that the displayed weight convention matches the formula.
R figure pair 2: Quadratic Weights and Weighted Contributions
Weighted Kappa Quadratic Weights R chart
Quadratic Weights generated for the R weighted kappa analysis.
Weighted Kappa Weighted Contributions R chart
Weighted Contributions generated for the R weighted kappa analysis.
Exact values

Quadratic Weights

The quadratic matrix assigns 0.1111, 0.4444 and 1.0000 to one-, two- and three-step errors.

The resulting κ is 0.765160; the higher value reflects the near absence of distant disagreements.

Cross-check against the verified matrix and the reported coefficient.
Exact values

Weighted Contributions

Observed quadratic contributions sum to 0.032357.

Low→medium and medium→high together contribute 0.022085, or 68.25% of observed quadratic disagreement.

Confirm that the displayed weight convention matches the formula.
R figure pair 3: Weighting Comparison and Result Summary
Weighted Kappa Weighting Comparison R chart
Weighting Comparison generated for the R weighted kappa analysis.
Weighted Kappa Result Summary R chart
Result Summary generated for the R weighted kappa analysis.
Exact values

Weighting Comparison

Linear κ = 0.647398; quadratic κ = 0.765160.

Quadratic κ exceeds linear κ by 0.117761 and unweighted κ by 0.221109.

Cross-check against the verified matrix and the reported coefficient.
Exact values

Result Summary

The summary reports the verified coefficients and disagreement terms.

The preferred conclusion is strong ordinal agreement with a separate upward-shift finding: 153 G3 increases versus 33 decreases.

Confirm that the displayed weight convention matches the formula.
AdvertisementGoogle AdSense placement reserved after the R figures

Weighted Kappa in Python, R, SPSS, SAS and Excel

Python

  • Build the 4 × 4 matrix with pandas or NumPy.
  • Use disagreement matrices for linear and quadratic weights.
  • Compute Do, De and κ directly.
  • Cross-check with sklearn.metrics.cohen_kappa_score using weights.

R

  • Create an ordered contingency table.
  • Use irr::kappa2, psych::cohen.kappa or a manual matrix calculation.
  • Confirm category order and weight definition.
  • Export the matrix, weights and result plots.

SPSS

  • CROSSTABS /STATISTICS=KAPPA returns ordinary kappa (.544 here).
  • Weighted kappa requires an extension, embedded Python/R or manual matrix computation.
  • The linked SPSS output verifies the counts, margins and custom weighted results.

SAS

  • Use PROC FREQ with the square table.
  • Request agreement statistics and specify appropriate weights.
  • Verify whether SAS reports agreement or disagreement weighting conventions.

Excel

  • Enter the observed matrix and calculate row/column margins.
  • Create linear and quadratic weight matrices.
  • Use SUMPRODUCT for observed and expected disagreement.
  • Compute 1 - Do/De.

Cross-software target

  • Linear κ = 0.6473984073
  • Quadratic κ = 0.7651598550
  • Quadratic Do = 0.0323574730
  • Quadratic De = 0.1377851008
SoftwareNative weighted kappa?Verified role in this projectExpected result
PythonYes, through libraries or manual formulaPrimary transcript and figuresLinear 0.647398; quadratic 0.765160
RYes, packages or manual formulaIndependent reproduction and figuresLinear 0.647398; quadratic 0.765160
SPSSBuilt-in CROSSTABS is unweightedMatrix verification plus integrated custom calculationUnweighted .544; custom weighted values match
SASAvailable through agreement proceduresReproducible alternative workflowShould match when weights and ordering match
ExcelManual formulasTransparent worked calculationShould match to displayed precision
SPSS distinction: the “Kappa = .544” table in standard CROSSTABS is not the linear or quadratic weighted result. It is ordinary kappa. The weighted coefficients printed later in the linked output come from the explicit weighted formula.

Cross-Software Reconciliation Checklist

CheckpointPython reportR reportSPSS outputRequired agreement
Valid N649649649All software uses the same complete pairs.
Observed matrix4 × 4 countsSame 4 × 4 countsSame crosstabCell order must match.
Linear weights0, .3333, .6667, 1Same matrixCustom calculationDisagreement convention.
Quadratic weights0, .1111, .4444, 1Same matrixCustom calculationDisagreement convention.
Linear κ0.64739840730.64739840730.6473984073 printed in custom blockExact match.
Quadratic κ0.76515985500.76515985500.7651598550 printed in custom blockExact match.
Ordinary κOptionalOptional0.544Do not confuse with weighted κ.

If software values differ, check category ordering first, then whether the program uses agreement weights or disagreement weights, then whether missing values or zero-count categories were dropped. A reversed category on one axis can produce a radically different coefficient even when the same counts are present.

Expandable Weighted Kappa Code

Python: manual linear and quadratic weighted kappa
import numpy as np

O = np.array([
    [88, 69,  0,  0],
    [12,267, 60,  1],
    [ 0, 19, 86, 23],
    [ 0,  0,  2, 22]
], dtype=float)

N = O.sum()
P = O / N
row = P.sum(axis=1)
col = P.sum(axis=0)
E = np.outer(row, col)

idx = np.arange(4)
distance = np.abs(idx[:, None] - idx[None, :])
W_linear = distance / 3
W_quadratic = (distance / 3) ** 2

def weighted_kappa(W):
    observed_disagreement = np.sum(W * P)
    expected_disagreement = np.sum(W * E)
    return 1 - observed_disagreement / expected_disagreement

print(weighted_kappa(W_linear))     # 0.647398407289
print(weighted_kappa(W_quadratic))  # 0.765159855031
Python: scikit-learn cross-check
from sklearn.metrics import cohen_kappa_score

g1 = []
g3 = []
for i in range(4):
    for j in range(4):
        g1.extend([i + 1] * int(O[i, j]))
        g3.extend([j + 1] * int(O[i, j]))

print(cohen_kappa_score(g1, g3, weights="linear"))
print(cohen_kappa_score(g1, g3, weights="quadratic"))
R: manual matrix calculation
O <- matrix(c(
  88,69,0,0,
  12,267,60,1,
  0,19,86,23,
  0,0,2,22
), nrow=4, byrow=TRUE)

P <- O / sum(O)
E <- outer(rowSums(P), colSums(P))
d <- abs(outer(1:4, 1:4, "-"))
W.linear <- d / 3
W.quadratic <- (d / 3)^2

wk <- function(W) {
  Do <- sum(W * P)
  De <- sum(W * E)
  1 - Do / De
}

wk(W.linear)
wk(W.quadratic)
SPSS: verify the table and ordinary kappa
CROSSTABS
 /TABLES=grade1 BY grade3
 /FORMAT=AVALUE TABLES
 /STATISTICS=KAPPA
 /CELLS=COUNT ROW COLUMN
 /COUNT ROUND CELL.

This command verifies the 4 × 4 table and ordinary κ = .544. Use embedded Python/R or an extension for weighted κ.

Excel: observed quadratic disagreement
=SUMPRODUCT(Observed_Proportions, Quadratic_Weights)

For the worked table this returns 0.032357473035.

Excel: expected quadratic disagreement and kappa
=SUMPRODUCT(Expected_Proportions, Quadratic_Weights)
=1-(Observed_Quadratic_Disagreement/Expected_Quadratic_Disagreement)

The final value is 0.765159855031.

How to Interpret and Report Weighted Kappa

Report the weighting scheme, category structure, sample size, coefficient and enough table information to understand disagreement severity. Avoid labels such as “good” or “substantial” without the underlying numbers.

Primary APA-Style Result

Agreement between G1 and G3 four-level grade bands was evaluated using quadratic weighted Cohen’s kappa. Agreement was strong, κw = 0.765, supplementary bootstrap 95% CI [0.729, 0.797], N = 649. Exact agreement was 71.34%, and 648 of 649 pairs (99.85%) were within one category.

Sensitivity Result

Using linear rather than quadratic weights produced κw = 0.647, supplementary bootstrap 95% CI [0.601, 0.691]. The lower linear estimate reflects its larger penalty for adjacent disagreements.

Data-Focused Interpretation

Of 649 paired classifications, 463 matched exactly, 185 differed by one category, one differed by two categories and none differed by three. G3 was higher than G1 for 153 students and lower for 33, indicating that strong agreement coexisted with a systematic upward shift in the final classification.

Reporting Hierarchy

Reporting elementInclude?Worked content
Weighting schemeAlwaysQuadratic primary; linear sensitivity.
Category orderAlwaysLow, medium, high, very high.
Sample sizeAlways649
Weighted kappaAlways0.765160
UncertaintyRecommendedBootstrap 95% CI 0.7289–0.7974.
Raw exact agreementRecommended71.34%
Distance distributionStrongly recommended185 adjacent, 1 two-step, 0 three-step mismatches.
DirectionWhen ratings are ordered in time153 upward vs 33 downward transitions.
Unweighted kappaOptional comparison0.544050.
Reusable templates:Replace the highlighted fields with the values from another weighted kappa analysis.

Quadratic Weighted Kappa

Primary

Method

Agreement between rating 1 and rating 2 was assessed using quadratic weighted Cohen’s kappa across k ordered categories.

Result

Weighted agreement was κw = estimate, 95% CI [lower, upper], N = sample size.

Pattern

exact agreement % matched exactly and within-one-category % were within one category.

Weighting Sensitivity

Secondary

Comparison

Linear weighting produced κw = value, compared with quadratic κw = value.

Meaning

The difference reflected the concentration of disagreements at specified category distances.

How to Discuss the Upward Shift Without Undermining Agreement

A rigorous discussion should state that weighted agreement is strong and that G3 tends to be higher. The agreement coefficient summarises closeness after chance correction; the direction counts summarise change. The recommended wording is: “Ratings were usually identical or adjacent, but disagreements were asymmetric, with 153 upward and 33 downward transitions.” This avoids the misleading implication that high kappa means no systematic movement.

Interpretive Labels Are Secondary

Conventional labels such as moderate, substantial or strong vary across disciplines and should never replace the numerical evidence. In this example, the most defensible interpretation comes from the coefficient together with 71.34% exact agreement, 99.85% within-one-category agreement, and the complete absence of maximum-distance errors.

Common Weighted Kappa Mistakes and Corrections

MistakeWhy it is wrongCorrection for this analysis
Calling .544 the weighted resultSPSS CROSSTABS reports ordinary kappaUse .647398 linear or .765160 quadratic.
Reporting only exact agreement{fmt_pct(exact_pct)} ignores chance and distanceReport raw agreement plus weighted kappa.
Choosing quadratic weights because κ is largerWeight choice becomes outcome-drivenJustify the loss function before analysis.
Treating κ as percent correctKappa is a chance-corrected ratioExplain observed and expected disagreement.
Ignoring table directionAgreement can coexist with systematic changeReport {up} upward and {down} downward transitions.
Using weighted kappa for nominal categoriesDistance would be arbitraryUse ordinary kappa.
Reversing category order on one axisWeights attach to wrong cellsUse the same ordered category list on both axes.
Ignoring zero distant cellsSparse severe-error cells affect uncertaintyShow the full matrix and use a suitable interval method.
Calling the result causalAgreement does not explain why grades changedInterpret as classification agreement and movement.
Using correlation as a substituteCorrelation measures association, not equalityUse weighted kappa for agreement.
Most serious reporting error: stating “quadratic kappa proves G1 and G3 are the same.” The coefficient shows close agreement beyond chance, while the transition counts show G3 is more often higher than lower. Both findings are simultaneously true.

Weighted Kappa Practice Questions

1. Compute exact agreement.

Add the diagonal cells 88 + 267 + 86 + 22 and divide by 649.

Answer: 463/649 = 0.713405 = 71.34%.

2. Count adjacent disagreements.

Add all cells one step above and below the diagonal.

Answer: 69 + 12 + 60 + 19 + 23 + 2 = 185.

3. Why is quadratic κ higher?

Use the distance distribution rather than a verbal benchmark alone.

Answer: 185 of 186 mismatches are adjacent, and quadratic weights give adjacent errors only 1/9 of the maximum penalty.

4. Interpret κQ = 0.7652.

Answer: observed quadratic disagreement is 76.52% lower than chance-weighted disagreement implied by the margins.

5. What does ordinary κ = 0.5441 add?

Answer: it shows agreement when every mismatch receives the same penalty, providing a sensitivity contrast with distance-sensitive coefficients.

6. Describe directional movement.

Answer: G3 is higher for 153 pairs and lower for 33; 82.26% of disagreements move upward.

7. Identify the largest quadratic contribution.

Answer: low→medium contributes 0.011813, slightly more than medium→high at 0.010272.

8. Should weighted kappa replace McNemar–Bowker?

Answer: no. Weighted kappa measures closeness; McNemar–Bowker tests paired marginal symmetry. Use each for its own question.

9. Calculate chance exact agreement.

Multiply corresponding G1 and G3 marginal proportions and sum across the four bands.

Answer: 0.371433, or 37.14%.

10. Calculate ordinary kappa.

Use (Po − Pe)/(1 − Pe).

Answer: (0.713405 − 0.371433)/(1 − 0.371433) = 0.544050.

11. Which band shows the largest upward count?

Answer: low→medium with 69 pairs, followed by medium→high with 60.

12. What percentage of all pairs are within one category?

Answer: (463 + 185)/649 = 99.85%.

Weighted Kappa Reports and Worked Files

Reproduction target: every file should return linear weighted κ = 0.6473984073 and quadratic weighted κ = 0.7651598550 when category ordering and disagreement weights match.

Weighted Kappa Frequently Asked Questions

What is weighted kappa?

Weighted kappa is a chance-corrected agreement coefficient for two ordinal ratings. It assigns smaller penalties to near disagreements and larger penalties to distant disagreements.

What is the weighted kappa formula?

Using disagreement weights, κw = 1 − observed weighted disagreement divided by expected weighted disagreement.

What is the result in this example?

Linear weighted kappa is 0.647398 and quadratic weighted kappa is 0.765160 for 649 paired G1 and G3 classifications.

Why is quadratic kappa higher than linear kappa?

Because 185 of the 186 mismatches are adjacent, one is two categories apart and none are three categories apart. Quadratic weights heavily discount adjacent errors.

Is .765 the same as 76.5% agreement?

No. Exact agreement is 71.34%. The .765 value is the proportionate reduction in quadratic disagreement relative to chance expectation.

What does SPSS kappa .544 mean here?

It is ordinary unweighted Cohen’s kappa from CROSSTABS, not the weighted coefficient.

Can weighted kappa be negative?

Yes. A negative value indicates weighted disagreement greater than expected under the marginal distributions.

Should I use linear or quadratic weights?

Choose based on the practical cost of category distance. Linear weights add equal penalty per step; quadratic weights make distant errors disproportionately worse.

Does weighted kappa test marginal change?

No. Use McNemar’s test for paired binary margins or McNemar–Bowker/Bowker symmetry methods for paired multicategory change.

Can I use weighted kappa for more than two raters?

Standard weighted Cohen’s kappa is for two ratings. Multiple-rater ordinal agreement requires another method or a model designed for several raters.

What should be reported with weighted kappa?

Report category order, weighting scheme, N, κ, uncertainty, raw exact agreement, distance distribution and any directional imbalance.

Why report the full agreement matrix?

The matrix shows that 153 mismatches move upward and 33 downward, information that the single kappa value cannot reveal.

Weighted Kappa Conclusion

Weighted Kappa provides a precise summary of agreement when categories are ordered and disagreement severity matters. In this G1-versus-G3 analysis, 463 of 649 classifications match exactly, 185 differ by one category, one differs by two categories and none differ by three. Linear weighted κ is 0.647398, while quadratic weighted κ is 0.765160.

The quadratic result is higher because almost every error is adjacent. That finding should not hide the transition pattern: 153 classifications move upward from G1 to G3 and 33 move downward. The most complete conclusion is therefore strong distance-sensitive agreement with a clear net upward shift in G3 classifications.

Final reporting sentence: “G1 and G3 four-level grade classifications demonstrated strong quadratic weighted agreement, κw = 0.765, supplementary bootstrap 95% CI [0.729, 0.797], N = 649; 71.34% matched exactly and 99.85% were within one category.”

Back to top

Need help applying this to your own data?

Salar Cafe can help interpret output, clean datasets, review assumptions, build dashboards and explain statistical results ethically.

Need help interpreting your data analysis results?

Contact Salar Cafe
Engr. Muhammad Yar Saqib author profile photo

Engr. Muhammad Yar Saqib

WhatsApp Get Data Analysis Help