UK-based online statistics and data analysis support for USA, UK, and international clients. No exams, no impersonation, no fabricated data.
Agreement among quantitative ratings

Inter Rater Reliability: 7 Essential Steps, ICC Formula and Worked Example

Inter rater reliability describes how consistently two or more raters, graders, observers, judges, clinicians, or measurement occasions assign scores to the same targets. This complete guide explains the ICC(2,1) two-way random-effects absolute-agreement model, assumptions, ANOVA components, confidence intervals, interpretation, and a verified worked analysis of G1, G2, and G3 scores for 649 targets in Python, R, SPSS, and Excel.

Two-way random-effects ICC
Absolute agreement
Single-rater reliability
649 targets × 3 raters
Python + R + SPSS + Excel
Targetsn = 649
ICC(2,1)0.858847
95% confidence interval0.836–0.878
DecisionGood reliability
Quick answer

Single-rater absolute agreement was good and statistically different from zero.

The worked inter rater reliability analysis treated G1, G2, and G3 as three randomly selected quantitative graders or scoring occasions applied to the same 649 targets. The two-way random-effects absolute-agreement coefficient was ICC(2,1) = 0.8588468392, with a 95% CI from 0.836 to 0.878. The corresponding test was F(648, 1296) = 20.246, p < .001. This indicates that one randomly selected rater would provide good absolute agreement for these scores.

Interpret the correct coefficient: the single-measure ICC answers how dependable one rater is. The average-measure ICC of 0.948, 95% CI [0.939, 0.956], answers how dependable the mean of all three ratings is. These values are not interchangeable, so the intended scoring procedure must be stated before interpreting inter rater reliability.
1

What does inter rater reliability measure?

A variance-based answer to an agreement question, not a percentage of identical ratings.

Inter rater reliability is the degree to which different raters produce the same or sufficiently similar judgments when they evaluate the same cases under the same scoring rules. A “target” may be a student, patient, interview, product, photograph, written response, behavior episode, medical image, laboratory specimen, or any other unit being rated. A “rater” may be a human judge, an instrument, a scoring occasion, a trained observer, or a measurement procedure that is treated as exchangeable with comparable raters.

Reliability is about reproducibility

When targets genuinely differ, dependable ratings should preserve those differences. A high-scoring target should generally remain high across raters, and a low-scoring target should generally remain low. However, rank ordering alone is not enough for absolute agreement. Two raters can correlate almost perfectly while one consistently assigns scores two points higher. An absolute-agreement inter rater reliability model counts that systematic offset as disagreement.

This distinction is why the Pearson correlation is not a complete substitute for an intraclass correlation coefficient. Correlation measures association between two variables; ICC measures how much observed score variation is attributable to stable target differences under a specified rating design.

Reliability is not validity

Strong inter rater reliability means raters are consistent with one another, not necessarily that they are correct. Several raters can apply the same flawed rubric and agree closely. Validity asks whether the ratings measure the intended construct and support the intended interpretation. Reliability is necessary for many uses of scores because unstable ratings cannot support precise decisions, but reliability alone does not establish accuracy, fairness, or causal meaning.

The distinction also prevents a common reporting error: a large ICC should not be described as proof that the scoring instrument is valid. It indicates reproducibility under the sampled targets, raters, scoring conditions, model definition, and unit of analysis.

Plain-language definition: inter rater reliability tells you how much confidence to place in a rating when another suitably trained rater could have scored the same target. The answer depends on whether you use one rater or an average, whether raters are fixed or random, and whether consistency or exact agreement matters.

Inter rater reliability in psychology, education, health, and qualitative research

In psychology, inter rater reliability may describe agreement on diagnostic ratings, behavior codes, interview categories, symptom severity, or observational scales. In education it may describe agreement among essay graders, performance assessors, classroom observers, or scoring occasions. In health research it may evaluate repeated measurements by clinicians or readers. In qualitative research it can document the dependability of coding decisions, although categorical coding often calls for Cohen’s kappa, weighted kappa, or Fleiss kappa rather than a continuous-score ICC.

The correct statistic follows the measurement scale and design. Continuous or approximately continuous ratings commonly use an intraclass correlation coefficient. Two nominal raters commonly use Cohen’s kappa. More than two nominal raters may use Fleiss kappa. Ordered categories may use weighted kappa. Selecting the statistic before looking at the result protects the analysis from choosing whichever coefficient appears largest.

2

When should you use inter rater reliability?

Use the decision logic below before selecting an ICC or kappa coefficient.

A formal inter rater reliability study is appropriate when ratings contribute to research variables, diagnoses, grades, eligibility decisions, quality-control classifications, clinical outcomes, audit findings, or machine-learning labels. The analysis is especially important when scoring requires judgment rather than a purely mechanical reading.

Define the target

State exactly what unit is rated and whether each target receives the same number of ratings.

Define the rater population

Decide whether conclusions apply only to named raters or to a wider population of comparable raters.

Choose agreement type

Use absolute agreement when identical scores matter; use consistency only when systematic offsets are acceptable.

Choose the unit

Decide whether decisions use one rating or the mean of several ratings.

Estimate uncertainty

Report the confidence interval, not only the point estimate.

Use ICC when

The outcome is numeric, the same targets are rated repeatedly, and the scientific question concerns reproducibility of individual or averaged scores. ICC is appropriate for interval-like ratings, counts with adequate spread, many-point scales, and measurements where differences between score values are meaningful.

Use kappa when

The outcome is categorical. Cohen’s kappa fits two raters assigning nominal categories; Fleiss kappa extends to multiple raters; weighted kappa recognizes the ordering and distance among categories. These methods answer a different question from continuous-score inter rater reliability.

Use agreement plots when

The practical size and direction of pairwise differences matter. A Bland–Altman plot can reveal bias, changing variability, and limits of agreement that a single ICC may conceal. ICC and agreement plots often complement rather than replace one another.

Examples of questions that require an agreement study

Applied questionRecommended design focusLikely statistic
Would another essay grader assign nearly the same numeric score?Same essays, multiple exchangeable graders, exact scores matterTwo-way random absolute-agreement ICC
Do two clinicians assign the same diagnosis category?Two raters, nominal outcomeCohen’s kappa
Do several observers use an ordered severity scale similarly?Multiple raters, ordinal categoriesWeighted multi-rater approach or ordinal agreement model
Is the mean of three judges dependable?Average of k ratings is the operational scoreAverage-measure ICC
Do two devices track patients similarly but show a fixed offset?Agreement and systematic bias both matterAbsolute-agreement ICC plus Bland–Altman analysis
Do not calculate inter rater reliability only after disagreements become visible. The target sample, rater selection, model, agreement definition, and decision threshold should be planned before scoring. Post hoc model switching can make a weak process look stronger than it is.
3

Inter rater reliability assumptions: conditions to check

The method is model-based and design-dependent, not assumption-free.

The assumptions for inter rater reliability should be evaluated at the design, data, and model levels. A numerically correct ICC can still answer the wrong question if raters are treated as random when they are fixed, if average-measure reliability is reported for a process that uses one rater, or if consistency is reported when exact scores determine decisions.

Design requirements

The same target is rated by all raters included in the complete two-way design.
Targets are independent of one another, even though ratings within a target are related.
Raters use the same score scale, scoring instructions, and operational definitions.
The selected raters represent the population to which the reliability conclusion will generalize when a random-rater model is used.
The unit of analysis is not accidentally transposed; rows represent targets and columns represent raters in this worked example.

Data and model conditions

Scores contain enough between-target variation to distinguish stable target differences from noise.
Extreme outliers, floor effects, ceiling effects, and range restriction are inspected because they can materially alter the ICC.
Missing ratings are handled with a method consistent with the intended design rather than silently changing the target set.
Residual variation is not dominated by a small group of targets or raters.
Confidence intervals are reported because the point estimate alone does not show precision.

Does inter rater reliability require normality?

Classical ICC inference is derived from an ANOVA model and is most straightforward when model residuals are reasonably well behaved. Exact normality of every rater column is not the central requirement. With a large balanced sample such as 649 targets, the point estimate is generally driven by mean-square components, while the confidence interval and F test rely more directly on distributional assumptions. Severe skew, heavy tails, influential outliers, or a bounded scale concentrated at an endpoint should prompt sensitivity checks, robust alternatives, transformations when meaningful, or bootstrap confidence intervals.

Normality should not be used as an automatic gatekeeper. More important questions are whether the scale is suitable for an ICC, whether raters applied it comparably, and whether variance components represent the intended population. For ordinal categories with only a few levels, a weighted kappa may be more defensible than treating category codes as continuous scores.

Why target heterogeneity matters

Inter rater reliability is a ratio of stable between-target variance to total observed variance under the selected model. If a sample includes very diverse targets, between-target variance can be large and the ICC can rise even when absolute discrepancies are clinically or educationally important. If a sample is highly homogeneous, the ICC can be modest despite small raw differences because there is little target variance to preserve. This is not a mathematical defect; it means reliability must be interpreted for the sampled target population and, when practical tolerances matter, supplemented with absolute-error summaries.

Balanced worked data: all 649 targets have valid G1, G2, and G3 ratings, so the SPSS reliability procedure uses 649 cases and excludes none. The absence of missing ratings simplifies the two-way ANOVA decomposition and cross-software verification.
4

Inter rater reliability hypotheses and ICC model choice

State the target, rater population, agreement definition, and measurement unit before calculating the coefficient.

There is no single universal ICC. A complete inter rater reliability report identifies the model, type, and definition. The worked analysis uses a two-way random-effects, absolute-agreement, single-measure coefficient commonly written as ICC(2,1).

Worked model

ICC(2,1)

Two-way random · absolute · single

Targets and raters are modeled as random effects. A fixed increase or decrease in one rater’s scores is treated as disagreement. Reliability refers to one randomly selected rater.

Three decisions behind the notation

Model: two-way randomBoth target effects and rater effects are random. The result is intended to generalize beyond the three observed raters to comparable raters.
Definition: absolute agreementRatings must match in level, not merely preserve rank order. Systematic rater offsets reduce reliability.
Type: single measureThe coefficient describes the dependability of one rater. Use the average-measure coefficient when the operational score is the mean of all k raters.

Null and alternative hypotheses

Null hypothesis

H0: the population single-rater absolute-agreement ICC equals zero. Under this null, ratings do not preserve dependable target differences beyond the rater and residual structure represented by the model.

The SPSS procedure tests the coefficient against a specified value of zero. In this example the observed F statistic is much larger than expected under the null, so the null is rejected.

Alternative hypothesis

H1: the population single-rater absolute-agreement ICC is greater than zero. A significant result shows evidence of nonzero reliability, but statistical significance is not the same as acceptable practical reliability.

With 649 targets, even an ICC too small for operational use could be statistically significant. The point estimate and confidence interval therefore carry more practical information than the p-value alone.

Consistency versus absolute agreement

A consistency ICC removes systematic rater-level differences from the definition of error. It can be appropriate when one rater is known to be stricter but relative rankings are the only concern. An absolute-agreement ICC includes those differences as disagreement. In this worked example the rater means are 11.399 for G1, 11.570 for G2, and 11.906 for G3. Because the scientific interpretation treats those columns as interchangeable ratings, the upward mean shift is relevant and the absolute-agreement definition is appropriate.

DecisionQuestion answeredWorked choice
One-way or two-wayDid every target receive the same raters, and should rater effects be modeled?Two-way
Fixed or random ratersDoes inference apply only to these raters or to comparable raters?Random
Consistency or agreementAre systematic score-level differences acceptable?Absolute agreement
Single or averageWill decisions use one rater or the mean of k raters?Single is primary; average is supplementary
5

Inter rater reliability formula: ICC(2,1) and ANOVA components

The two-way random absolute-agreement coefficient combines target, rater, and residual mean squares.

The ICC(2,1) formula expresses inter rater reliability as the proportion of score variation attributable to persistent differences among targets after accounting for residual disagreement and systematic rater effects.

ICC(2,1) = (MSR − MSE) / [MSR + (k − 1)MSE + k(MSC − MSE)/n]

Here MSR is the target or row mean square, MSC is the rater or column mean square, MSE is the residual mean square, k is the number of raters, and n is the number of targets.

Two-way random-effects model

Yij = μ + Ti + Rj + Eij

Yij is the rating for target i from rater j; μ is the grand mean; Ti is the target effect; Rj is the rater effect; and Eij contains target-by-rater interaction and residual error.

The model does not assume every disagreement has the same cause. It summarizes the observed structure through variance components. Absolute agreement retains the rater-level component in the denominator because a strict or lenient rater changes the score a target receives.

F test for nonzero reliability

F = MSR / MSE = 24.083791 / 1.189560 = 20.245973

The numerator degrees of freedom are n − 1 = 648. The denominator degrees of freedom are (n − 1)(k − 1) = 1296.

The F test asks whether between-target variation exceeds residual variation. The resulting p-value is smaller than .001. Public reporting should use p < .001 rather than p = 0, even though software may display .000 or the numerical calculation may underflow to zero.

Exact ANOVA components from the worked data

SourceSum of squaresdfMean squareRole in ICC(2,1)
Targets (rows)15,606.29686764824.083791Stable differences among targets
Raters (columns)86.330765243.165383Systematic scoring-level differences
Residual / interaction1,541.6692351,2961.189560Unexplained disagreement and target-by-rater interaction
ICC(2,1) = (24.083791 − 1.189560) / [24.083791 + 2(1.189560) + 3(43.165383 − 1.189560)/649] = 0.858846839

The rater mean square is large relative to the residual mean square, reflecting systematic mean differences across G1, G2, and G3. Because the absolute-agreement formula divides the rater component by the large target count, the adjustment is present but does not overwhelm the substantial target signal.

Average-measure inter rater reliability

ICC(2,k) = (MSR − MSE) / [MSR + (MSC − MSE)/n] = 0.948

Averaging three ratings reduces random disagreement. This is why the reliability of the mean of G1, G2, and G3 is higher than the reliability of one randomly selected rating. Report ICC(2,k) only when the average of k ratings is the score actually used.

6

Inter rater reliability example: G1, G2, and G3 ratings

A complete 649-target, three-rater worked analysis with exact values.

The worked inter rater reliability dataset contains 649 rows and three numeric rating columns: G1, G2, and G3. Each row is one target, and each column is treated as one rater or scoring occasion. The original educational labels are retained so every software output can be reconciled exactly.

Variables used

VariableRoleScaleObserved range
subject_idTarget identifierInteger1–649
G1Rater/scoring occasion 1Numeric score0–19
G2Rater/scoring occasion 2Numeric score0–19
G3Rater/scoring occasion 3Numeric score0–19

Raw descriptive statistics

RaterNMeanSDMinimumMaximum
G164911.39912.7453019
G264911.57012.9136019
G364911.90603.2307019

The rater means increase from G1 to G3, so exact agreement is a stricter criterion than consistency. The mean difference is not enormous relative to the score range, but it is systematic. ICC(2,1) appropriately includes this rater-level shift in the denominator rather than treating it as harmless.

Pairwise associations are high but are not the final answer

G1 with G2r = 0.865p < .001
G1 with G3r = 0.826p < .001
G2 with G3r = 0.919p < .001
Target-mean variance8.0279Across 649 targets

All three pairwise Pearson correlations are strong and positive. These correlations show that targets tend to retain their relative positions across scoring occasions. They do not test exact agreement because a fixed shift in all scores leaves a correlation unchanged. The lower single-rater ICC of 0.859 compared with Cronbach’s alpha of 0.951 illustrates why internal consistency and absolute-agreement inter rater reliability should not be treated as synonyms.

Substantive limitation of the demonstration: G1, G2, and G3 are sequential educational grades, not three independent human graders scoring the same response at the same time. The calculation is a valid demonstration of the ICC model, but applied claims about human rater behavior require an actual rater-sampling design.

Seven-step worked process

Arrange the matrix

Place targets in rows and the same three raters in columns.

Check completeness

Confirm 649 complete rows and no excluded cases.

Set the model

Select two-way random, absolute agreement, single measure.

Compute ANOVA

Obtain target, rater, and residual mean squares.

Calculate and report

Report ICC, CI, F, degrees of freedom, and practical meaning.

7

Inter rater reliability statistics, results and interpretation

All software outputs reconcile to the same ICC, confidence interval, F test, and practical conclusion.

The primary inter rater reliability result is the single-measure absolute-agreement ICC because the question concerns one randomly selected rater. The average-measure result is reported separately to show the reliability of averaging all three ratings.

Primary verified result

ICC(2,1) = 0.858847

95% CI [0.836, 0.878] · F(648, 1296) = 20.246 · p < .001

The coefficient indicates good single-rater absolute agreement for the sampled targets and rater design. Both confidence limits remain well above zero and within the range commonly described as good, while the interval also shows that the population value is unlikely to be above 0.90 under this design.

Single measures0.85995% CI 0.836–0.878
Average measures0.94895% CI 0.939–0.956
Cronbach’s alpha0.951Standardized alpha 0.953
Valid targets649Excluded = 0

Primary result table

CoefficientEstimate95% CIFdf1df2p-valueMeaning
ICC(2,1), absolute agreement0.8588468390.836–0.87820.2459736481296< .001Reliability of one randomly selected rater
ICC(2,3), average measures0.9480.939–0.95620.2459736481296< .001Reliability of the average of three ratings

Cross-platform reconciliation

PlatformModel specificationICC(2,1)F statisticVerification
PythonManual two-way ANOVA mean squares and ICC formula0.858846839220.2459730213Pass
RManual two-way random absolute-agreement calculation0.858846839220.2459730213Pass
SPSSRELIABILITY /ICC=MODEL(RANDOM) TYPE(ABSOLUTE)0.85920.246Pass after display rounding
ExcelAuditable sums of squares, mean squares, and direct formula0.858846839220.2459730213Pass

The exact calculations agree to machine precision in Python, R, and Excel. SPSS displays rounded values in the ICC table. Agreement across implementations confirms the row/column orientation, target count, rater count, mean-square decomposition, and absolute-agreement formula.

Do not interpret p < .001 as excellent reliability. The significance test only rejects ICC = 0. The magnitude and confidence interval determine whether the coefficient is adequate for the intended decision. A demanding high-stakes application may require a higher lower confidence bound than a preliminary research study.

How to interpret the inter rater reliability result

Point-estimate interpretation

ICC(2,1) = 0.858847 means that the rating design preserves a large share of the modeled score variation as stable differences among targets, while the remaining variation reflects rater-level disagreement and residual error. On commonly used descriptive bands, values from 0.75 to 0.90 are often called good. These bands are conventions rather than universal decision rules.

The value should not be translated mechanically into “85.9% agreement.” ICC is a variance ratio under a model, not the percentage of identical ratings. Nor does it mean that 85.9% of ratings are correct.

Confidence-interval interpretation

The 95% confidence interval from 0.836 to 0.878 shows the precision of the estimated single-rater reliability. The lower bound is especially important for planning. If an application requires reliability of at least 0.80, the interval supports that requirement in this sample. If an application requires at least 0.90, the interval does not support it.

Interpreting the lower bound rather than the point estimate alone guards against overconfidence. Larger and more representative samples improve precision but do not correct a poorly chosen model or unrepresentative raters.

Common descriptive bands for ICC values

ICC rangeCommon descriptive labelHow to use the label responsibly
Below 0.50PoorUsually inadequate for individual-level decisions; investigate design, training, and scale limitations.
0.50 to 0.75ModerateMay support exploratory or group-level uses but can be insufficient for high-stakes individual decisions.
0.75 to 0.90GoodOften acceptable, subject to the lower CI bound, consequences of error, and discipline-specific standards.
Above 0.90ExcellentOften preferred for high-stakes individual measurement, but still requires validity and bias checks.

These cutoffs are not laws. A reliability requirement should reflect the decision being made, the cost of disagreement, the number of ratings averaged, and available alternatives. A coefficient of 0.86 may be strong for exploratory observational research and insufficient for a clinical threshold that determines irreversible treatment.

Why average-measure reliability is higher

The average of three ratings smooths idiosyncratic disagreement, producing ICC(2,3) = 0.948. This does not prove that each individual rater is excellent. It means a scoring protocol that always averages three comparable raters is more dependable than one that uses a single rater. If the real-world workflow assigns only one rater, reporting only the average-measure coefficient would materially overstate operational inter rater reliability.

Why Cronbach’s alpha is not the same result

Cronbach’s alpha equals 0.951 for the three columns, and standardized alpha is 0.953. Alpha summarizes internal consistency and is closely related to an average-measure consistency concept under particular assumptions. It does not directly answer whether one randomly selected rater produces the same absolute score as another. The single-measure absolute-agreement ICC of 0.859 is the more relevant primary coefficient for this question. The distinction is explained further in the site’s Cronbach’s alpha guide.

Best one-sentence interpretation: ratings showed good single-rater absolute agreement, ICC(2,1) = .859, 95% CI [.836, .878], while averaging all three ratings produced excellent reliability, ICC(2,3) = .948.
8

Inter rater reliability in Python: complete calculation and charts

A reproducible two-way random absolute-agreement analysis with five verified graphics.

The Python workflow calculates inter rater reliability directly rather than relying on a black-box coefficient. This makes the target, rater, and residual mean squares auditable and allows exact comparison with R, SPSS, and Excel.

Python primary metrics chart for inter rater reliability

Python chart 1: primary metrics

The chart places the verified ICC, F statistic, p-value representation, target count, and rater count in one summary view.

The most important statistical values are ICC(2,1) = 0.858846839 and F = 20.245973. The count of 649 targets is numerically much larger than the coefficient, so a common vertical scale visually compresses the ICC bar. The chart should therefore be read from its labels rather than by comparing bar heights across quantities that use different units.

The p-value is shown numerically as zero because floating-point evaluation underflows at an extremely small probability. The public statistical conclusion remains p < .001. The chart also confirms the balanced design: 649 targets and three raters.

Python ICC ANOVA components chart

Python chart 2: ICC ANOVA components

Target, rater, and residual mean squares show the variance structure behind ICC(2,1).

Python rater means chart

Python chart 3: rater means

Mean scores increase from G1 to G2 and G3, documenting a systematic scoring-level difference.

The ANOVA components are MS targets = 24.083791, MS raters = 43.165383, and MS residual = 1.189560. The small residual mean square relative to the target mean square supports dependable target differentiation. The nontrivial rater mean square shows that the three scoring occasions differ systematically in level, which is why an absolute-agreement model is more demanding than a consistency model.

The rater-means chart provides the direction of that systematic effect. G1 averages 11.3991, G2 averages 11.5701, and G3 averages 11.9060. The difference between G1 and G3 is about 0.507 score units. Absolute-agreement inter rater reliability treats this shift as part of disagreement instead of discarding it.

Python target coverage chart for inter rater reliability

Python chart 4: target coverage and record distribution

The target-index histogram verifies broad, even coverage across the complete set of 649 target records.

Python verified result summary chart

Python chart 5: verified result summary

The horizontal summary repeats the exact metrics used for the cross-software check.

The fourth image is best interpreted as a target-coverage check rather than a distribution of the mean ratings themselves. Its bins contain similar numbers of sequential target identifiers, showing that the full target range is represented in the generated workflow. Substantive score interpretation comes from the rater means, pairwise plots, and ANOVA components.

The verified summary reinforces that Python used 649 targets and three raters and obtained the same ICC and F statistic as the other platforms. Like the first chart, its count scale dominates the coefficient visually, so the printed values are the appropriate basis for interpretation.

Python code for ICC(2,1)
Pythonimport numpy as np
import pandas as pd
from scipy.stats import f

ratings = pd.read_csv("dataset.csv")[["G1", "G2", "G3"]].dropna()
Y = ratings.to_numpy(dtype=float)
n, k = Y.shape

grand = Y.mean()
row_means = Y.mean(axis=1)
col_means = Y.mean(axis=0)

ss_rows = k * np.sum((row_means - grand) ** 2)
ss_cols = n * np.sum((col_means - grand) ** 2)
ss_total = np.sum((Y - grand) ** 2)
ss_error = ss_total - ss_rows - ss_cols

ms_rows = ss_rows / (n - 1)
ms_cols = ss_cols / (k - 1)
ms_error = ss_error / ((n - 1) * (k - 1))

icc_2_1 = (ms_rows - ms_error) / (
ms_rows + (k - 1) * ms_error + k * (ms_cols - ms_error) / n
)
f_stat = ms_rows / ms_error
p_value = f.sf(f_stat, n - 1, (n - 1) * (k - 1))

print(f"n={n}, k={k}")
print(f"MSR={ms_rows:.12f}, MSC={ms_cols:.12f}, MSE={ms_error:.12f}")
print(f"ICC(2,1)={icc_2_1:.12f}")
print(f"F={f_stat:.12f}, p={p_value:.6g}")

The direct calculation returns ICC(2,1) = 0.858846839219 and F = 20.245973021325. A production script should also calculate a confidence interval or use a validated reliability package, retain the model specification in the output, and test the matrix orientation before reporting the result.

Open the complete Python inter rater reliability report →

9

Inter rater reliability in R: complete calculation and charts

The R workflow reproduces the Python and SPSS values and the same five result graphics.

The R analysis treats raters as random and systematic rater offsets as disagreement. This specification matches the intended ICC(2,1) definition of inter rater reliability, so its coefficient should match the Python, SPSS, and Excel results apart from display rounding.

R primary metrics chart for inter rater reliability

R chart 1: primary reliability metrics

The primary panel summarizes ICC(2,1), the F ratio, the numerical p-value representation, and the dimensions of the rating matrix.

R confirms ICC(2,1) = 0.858846839 and F = 20.245973 for 649 targets rated in three columns. Because the graphic combines statistics and counts on one scale, the printed labels are more informative than the relative bar heights. The substantive conclusion is good single-rater absolute agreement with strong evidence that the population ICC exceeds zero.

The R report explicitly notes that ICC(2,1) penalizes systematic disagreement among random raters. This is the critical model statement that distinguishes the result from a consistency coefficient.

R ICC ANOVA components chart

R chart 2: ANOVA mean-square components

The target, rater, and residual terms reproduce the exact denominator of the absolute-agreement ICC.

R rater means chart

R chart 3: mean rating by rater

The plot documents the progressive mean increase across G1, G2, and G3.

The target mean square of 24.0838 is about twenty times the residual mean square of 1.1896, producing the F statistic of 20.246. The rater mean square of 43.1654 is larger than either, but its contribution to ICC(2,1) is scaled by the target count. This pattern shows why the formula must be evaluated rather than interpreting any one ANOVA component in isolation.

The R rater-means chart makes the absolute-agreement issue visible. Ratings rise from G1 through G3. A consistency coefficient would focus on whether targets remain ordered similarly; the chosen absolute-agreement coefficient also asks whether the score levels match.

R target coverage chart for inter rater reliability

R chart 4: target coverage

The record-index distribution confirms that all target ranges enter the analysis rather than a truncated subset.

R verified result summary chart

R chart 5: verified result summary

The final panel records the same exact coefficient, test statistic, and matrix dimensions used by the independent checks.

The target coverage image is a workflow diagnostic, not a direct measure of score agreement. It verifies the full 649-row input. The verified summary then provides a compact audit point for matching the R output to Python and Excel.

When R packages are used, model labels can differ. The safest practice is to state the design in words—two-way random effects, absolute agreement, single measure—and verify the package’s documentation rather than relying only on a shorthand label.

R code for ICC(2,1)
Rdat <- read.csv("dataset.csv", stringsAsFactors = FALSE)
Y <- as.matrix(dat[, c("G1", "G2", "G3")])
Y <- Y[complete.cases(Y), , drop = FALSE]

n <- nrow(Y)
k <- ncol(Y)
grand <- mean(Y)
row_means <- rowMeans(Y)
col_means <- colMeans(Y)

ss_rows <- k * sum((row_means - grand)^2)
ss_cols <- n * sum((col_means - grand)^2)
ss_total <- sum((Y - grand)^2)
ss_error <- ss_total - ss_rows - ss_cols

ms_rows <- ss_rows / (n - 1)
ms_cols <- ss_cols / (k - 1)
ms_error <- ss_error / ((n - 1) * (k - 1))

icc_2_1 <- (ms_rows - ms_error) /
(ms_rows + (k - 1) * ms_error + k * (ms_cols - ms_error) / n)

f_stat <- ms_rows / ms_error
p_value <- pf(f_stat, n - 1, (n - 1) * (k - 1), lower.tail = FALSE)

cat(sprintf("ICC(2,1) = %.12f
", icc_2_1))
cat(sprintf("F = %.12f, p = %.6g
", f_stat, p_value))

The R calculation returns the same ICC and F statistic to at least twelve decimal places. This direct approach is useful for auditing; a validated package remains helpful for confidence intervals, unbalanced designs, and alternative ICC definitions.

Open the complete R inter rater reliability report →

10

Inter rater reliability in SPSS: ICC workflow and output

SPSS reports the single- and average-measure coefficients, confidence intervals, and F test.

The SPSS inter rater reliability workflow uses the RELIABILITY procedure with a random-effects ICC model and an absolute-agreement definition. The valid-case table shows 649 included targets and zero exclusions.

SPSS menu path

Choose Analyze → Scale → Reliability Analysis. Move G1, G2, and G3 into the Items box. Open Statistics, request intraclass correlation coefficients, choose a two-way random model, select absolute agreement, request a 95% confidence interval, and retain both single and average measures in the output.

Menu labels vary slightly by SPSS version. The syntax is the clearest permanent record because it preserves the selected model and agreement type.

SPSS output to report

The Intraclass Correlation Coefficient table reports single measures = .859, 95% CI [.836, .878], and average measures = .948, 95% CI [.939, .956]. The F test is 20.246 with 648 and 1296 degrees of freedom, displayed significance .000. Report this as p < .001.

The output also reports alpha = .951 and pairwise correlations from .826 to .919. These values are useful diagnostics but do not replace the selected absolute-agreement ICC.

SPSS syntaxTITLE 'Inter Rater Reliability ICC Two One'.
SUBTITLE 'Random raters, absolute agreement, single measure'.

RELIABILITY
/VARIABLES=G1 G2 G3
/SCALE('Random raters') ALL
/MODEL=ALPHA
/STATISTICS=DESCRIPTIVE SCALE CORR
/SUMMARY=TOTAL
/ICC=MODEL(RANDOM) TYPE(ABSOLUTE) CIN=95 TESTVAL=0.

MEANS TABLES=G1 G2 G3
/CELLS=COUNT MEAN STDDEV MIN MAX.

CORRELATIONS
/VARIABLES=G1 G2 G3
/PRINT=TWOTAIL NOSIG
/MISSING=PAIRWISE.

GRAPH /SCATTERPLOT(MATRIX)=G1 G2 G3.

How to read the SPSS tables

SPSS tableKey valuesInterpretation
Case Processing SummaryValid 649; excluded 0The ICC uses every target with complete G1, G2, and G3 scores.
Reliability StatisticsAlpha .951; standardized .953High internal consistency, not the primary absolute-agreement result.
Item StatisticsMeans 11.40, 11.57, 11.91Systematic mean differences are present across raters.
Intraclass Correlation CoefficientSingle .859; average .948Primary and averaged-score reliability under the selected model.
Inter-Item Correlation Matrix.826 to .919Strong association, but exact agreement still requires the ICC.
Scatterplot MatrixStrong positive linear cloudsTargets retain similar ordering, with some spread and low-score clusters.

The scatterplot matrix visually supports strong association among all three score columns. It does not, by itself, distinguish consistency from absolute agreement. The ICC table supplies that design-specific answer.

Open the complete SPSS inter rater reliability output →

11

Inter rater reliability in Excel: worked ICC calculation

The workbook exposes the targets-by-raters matrix, ANOVA components, formula, diagnostics, and verification ledger.

Excel can calculate inter rater reliability transparently when the rating matrix is balanced and formulas are carefully audited. The downloadable workbook contains Guide, Data_Input, Working, Calculations, Diagnostics, and Reporting sheets.

Workbook structure

Guide: identifies ICC(2,1), the design, exact formula, variables, and expected result.
Data_Input: contains the 649 aligned G1, G2, and G3 rows.
Working: calculates target means, target sums of squares, and within-target residual pieces.
Calculations: derives MS targets, MS raters, MS residual, ICC, F, and p-value.
Diagnostics and Reporting: document assumptions and compare calculated values with the verified reference.

Excel outputs

MS targets24.0837914614
MS raters43.1653826400
MS residual1.1895595947
ICC(2,1)0.8588468392
F statistic20.2459730213

Excel formula pattern

ExcelTarget mean (row 2):
=AVERAGE(B2:D2)

Grand mean:
=AVERAGE(B2:D650)

Target sum of squares:
=3*SUMPRODUCT((E2:E650-$B$654)^2)

Rater sum of squares:
=649*SUMPRODUCT((B655:D655-$B$654)^2)

Residual sum of squares:
=Total_SS-Target_SS-Rater_SS

MS targets:
=Target_SS/(649-1)

MS raters:
=Rater_SS/(3-1)

MS residual:
=Residual_SS/((649-1)*(3-1))

ICC(2,1):
=(MS_Targets-MS_Residual)/(MS_Targets+(3-1)*MS_Residual+3*(MS_Raters-MS_Residual)/649)

F statistic:
=MS_Targets/MS_Residual

P-value:
=F.DIST.RT(F_Statistic,648,1296)

Cell references depend on the workbook layout, but the mathematical sequence should remain visible. Avoid hard-coding verified results into cells that are supposed to calculate them. A reliable workbook should show zero difference between the calculated coefficient and the independent reference.

Excel quality-control checks

Structural checks

Confirm that exactly three rater columns and 649 complete target rows are included. Ensure no header row enters the numeric range, formulas are filled through the final target, and row/column roles are not reversed. The row count should remain constant across every calculation sheet.

Numerical checks

Verify that total sum of squares equals target plus rater plus residual sums of squares within rounding tolerance. Confirm ICC = 0.8588468392 and F = 20.2459730213. Any discrepancy usually signals a range error, a denominator error, or accidental inclusion of missing/text cells.

Open the worked Excel inter rater reliability analysis →

12

Inter rater reliability in MATLAB and SAS

Equivalent workflows for readers who use matrix calculations or mixed-model software.

Inter rater reliability can be reproduced in MATLAB or SAS when the targets-by-raters design and ICC definition are kept explicit. The required model is two-way random effects, absolute agreement, single measure. The same 649 × 3 data matrix must produce ICC(2,1) = 0.858846839 when the ANOVA mean squares and formula are implemented correctly. A valid inter rater reliability calculation must preserve that model definition across platforms.

MATLAB matrix workflow

MATLABX = readmatrix("ratings.csv"); % rows = targets, columns = raters
[n,k] = size(X);
gm = mean(X,"all");
rowMeans = mean(X,2);
colMeans = mean(X,1);

SSR = k * sum((rowMeans-gm).^2);
SSC = n * sum((colMeans-gm).^2);
SSE = sum((X-rowMeans-colMeans+gm).^2,"all");

MSR = SSR/(n-1);
MSC = SSC/(k-1);
MSE = SSE/((n-1)*(k-1));

icc21 = (MSR-MSE) / ...
(MSR+(k-1)*MSE+k*(MSC-MSE)/n);

The calculation should use the complete balanced rating matrix. Verify the output against MS targets = 24.083791, MS raters = 43.165383, and MS residual = 1.189560 before reporting the coefficient.

SAS mixed-model workflow

SASproc mixed data=ratings method=type3;
class target rater;
model score = / solution;
random target rater;
run;

Arrange the ratings in long form with one row per target–rater score. Extract the target, rater, and residual components or corresponding mean squares, then apply the absolute-agreement single-measure ICC formula. State the model in words because software labels alone can conceal whether systematic rater offsets count as disagreement.

Do not report a generic “ICC” from another platform without verifying the model. A one-way ICC, a consistency ICC, or an average-measure ICC answers a different question and will not necessarily reproduce the verified result.
13

Inter rater reliability compared with related statistics

Agreement, consistency, association, internal consistency, and measurement error answer different questions.

Choosing an inter rater reliability statistic requires matching the coefficient to the scale, number of raters, rater sampling, agreement definition, and intended score. The table below prevents common substitutions that can overstate reliability.

Statistic or modelBest useWhat it capturesKey limitation
ICC(2,1)One quantitative rating from a random raterTwo-way random absolute agreementDepends on target heterogeneity and model assumptions
ICC(2,k)Mean of k quantitative ratings from random ratersReliability of the average scoreOverstates operational reliability if only one rater is used
ICC(3,1)Named fixed raters under a two-way mixed designReliability limited to those raters; definition may be consistency or agreementDoes not generalize to a wider rater population
Cohen’s kappaTwo raters, nominal categoriesChance-corrected categorical agreementSensitive to prevalence and marginal distributions
Fleiss kappaMultiple raters, nominal categoriesMulti-rater chance-corrected agreementNot designed for continuous ratings
Weighted kappaOrdered categoriesPartial credit for near agreementResult depends on the weighting scheme
Percent agreementSimple descriptive categorical summaryObserved exact matchesDoes not adjust for chance and ignores disagreement severity
Pearson correlationLinear association between two score setsRank/linear co-movementCan be 1.0 despite a fixed score offset
Cronbach’s alphaInternal consistency of a multi-item scaleShared covariance among itemsNot a direct single-rater absolute-agreement coefficient
Bland–Altman analysisPairwise measurement agreementBias and limits of agreementDoes not summarize multi-rater reliability in one ICC

Inter versus intra rater reliability

Inter rater reliability

Compares different raters evaluating the same targets. It asks whether the measurement process is reproducible across people, devices, judges, or exchangeable scoring occasions. Rater selection and generalization are central design questions.

Intra rater reliability

Compares repeated ratings from the same rater, usually separated by time. It asks whether one rater is stable. Memory effects, learning, fatigue, and genuine target change must be controlled. The same ICC family can be used, but the design and interpretation differ.

Why Pearson correlation can mislead

Imagine Rater B always scores exactly five points higher than Rater A. The Pearson correlation is 1.00 because target ordering is perfectly preserved. Absolute agreement is poor because the score assigned to any target depends on the rater. ICC(2,1) detects this systematic offset through the rater mean square. This is the core reason a correlation coefficient should not be labeled an inter rater reliability coefficient without qualification.

Why the intraclass correlation coefficient has several pages on this site

Readers who need broader model selection can consult the ICC formula and interpretation guide and the concise intraclass correlation coefficient overview. The present article focuses specifically on the random-rater absolute-agreement application to inter rater reliability.

14

Diagnostics, sensitivity checks and common mistakes

A defensible analysis combines the coefficient with design checks, rater patterns, data inspection, and improvement procedures.

Diagnostics for inter rater reliability should examine the rating matrix before and after computing the ICC. The goal is not to search for reasons to discard an inconvenient result; it is to understand what produces agreement and disagreement.

Rater-level diagnostics

Compare means, medians, standard deviations, and score ranges for every rater.
Plot pairwise scores with the identity line, not only a fitted regression line.
Inspect whether one rater is systematically stricter, more variable, or compressed near the center.
Check whether disagreement changes over time, target type, difficulty, or score level.
Review whether raters were independent and blinded during the reliability sample.

Target-level diagnostics

Identify targets with unusually large within-target spreads.
Examine floor and ceiling concentrations that restrict distinguishable scores.
Assess whether the sampled target range matches the population of intended use.
Verify that target identifiers are unique and rows have not been duplicated.
Document missing ratings, reasons for missingness, and the analysis set.

Worked-data diagnostic findings

DiagnosticObserved resultImplication
Complete cases649 of 649; no exclusionsNo listwise-deletion change to the target population
Rater means11.399, 11.570, 11.906Small systematic upward shift across scoring occasions
Rater SDs2.745, 2.914, 3.231G3 is somewhat more dispersed
Pairwise correlations.826 to .919Strong preservation of target ordering
Residual mean square1.189560Residual disagreement is small relative to target signal
Confidence interval.836 to .878Good precision and a lower bound above .80

Negative ICC values

An estimated ICC can be negative when residual disagreement exceeds stable between-target variation. A negative estimate does not mean “negative reliability” in a substantive sense. It signals that the observed data provide no evidence of dependable target differentiation under the chosen model, and the process should be investigated. Some software truncates negative values in summaries; transparent reporting should preserve the estimate and explain the model.

Missing and unbalanced ratings

The direct balanced-ANOVA formulas in this article assume every target has all three ratings. Real studies may have different raters per target or missing cells. Deleting incomplete targets can waste information and introduce bias. Mixed-effects models or specialized reliability methods can estimate variance components from unbalanced data, but the estimand and assumptions must be described. A balanced ICC formula should not be applied to a ragged matrix by replacing missing values with zeros or column means.

Range restriction and transportability

If reliability is estimated only among very similar targets, between-target variance shrinks and the ICC can fall. If the sample deliberately includes extreme targets, the ICC can rise. Therefore, inter rater reliability does not transport automatically from one population to another. Report the target characteristics and score distribution so readers can judge whether the reliability estimate applies to their setting.

Recommended companion checks: report descriptive statistics, confidence intervals, pairwise identity plots, rater mean differences, and—when two methods or raters are central—a Bland–Altman analysis. These diagnostics explain the coefficient rather than competing with it.

How to establish and improve inter rater reliability

Clarify the scoring construct

Define what is and is not being judged. Replace broad labels such as “quality” with observable indicators. Specify boundaries, exceptions, units, and the evidence required for each score. A rubric should reduce avoidable discretion without erasing legitimate expert judgment.

Use anchored examples

Provide representative examples at the low, middle, and high ends of the scale, including borderline cases. Explain why each example receives its score. Anchors make abstract score definitions operational and expose disagreements before production scoring begins.

Calibrate raters

Have raters independently score a common pilot set, compare decisions, discuss reasoning, and repeat until the intended rules are applied consistently. Calibration should focus on recurring disagreement patterns rather than forcing consensus on every unusual case.

Standardize conditions

Control the information available to raters, order of materials, scoring interface, time limits, measurement device, and opportunities for discussion. Blinding can prevent raters from being influenced by another score or an expected outcome.

Monitor drift

Agreement can deteriorate after initial training. Insert periodic duplicate cases, review control charts or rolling agreement estimates, and schedule recalibration when systematic offsets appear. Report whether reliability was checked only at baseline or throughout data collection.

Use adjudication correctly

Adjudication can produce a final consensus score, but consensus after discussion is not evidence of independent inter rater reliability. Estimate reliability from the original independent ratings, then describe adjudication as a separate data-resolution step.

Practical improvement workflow

Observed problemLikely causeCorrective actionDiagnostic to repeat
One rater is consistently higherDifferent threshold or scale calibrationReview anchors, retrain on score levels, verify unitsRater means and absolute-agreement ICC
Disagreement grows for high scoresHeteroscedastic error or unclear upper anchorsAdd upper-range examples and inspect difference plotsBland–Altman plot and residual spread
Only borderline cases disagreeCategory boundaries are ambiguousRefine decision rules and document tie-breaking evidenceCase-level disagreement table
Agreement declines over timeRater drift, fatigue, or changing interpretationPeriodic blinded duplicates and recalibrationReliability by time block
ICC is low in a homogeneous sampleRestricted target varianceSample the intended full target range; add absolute-error metricsTarget variance and limits of agreement
Do not improve inter rater reliability by deleting difficult targets merely because raters disagree. Ambiguous cases may be exactly where the measurement system needs to work. Exclusions require substantive rules established before inspecting their effect on the coefficient.

Adding more raters can improve the reliability of an average, as shown by the increase from single-measure ICC 0.859 to average-measure ICC 0.948. It does not repair systematic bias in every rater. If all raters share the same misconception, averaging can produce a highly reliable but invalid score.

15

How to report inter rater reliability in APA style

Include the ICC model, agreement definition, measurement unit, confidence interval, F test, and practical interpretation.

An APA-style inter rater reliability statement should allow another analyst to reconstruct the intended coefficient. Reporting only “ICC = .86” is incomplete because multiple ICC models can yield different values from the same matrix.

APA-style worked report

Inter rater reliability was evaluated for G1, G2, and G3 scores across 649 targets using a two-way random-effects, absolute-agreement, single-measure intraclass correlation coefficient. Ratings demonstrated good absolute agreement, ICC(2,1) = .859, 95% CI [.836, .878], F(648, 1296) = 20.25, p < .001. The reliability of the mean of all three ratings was higher, ICC(2,3) = .948, 95% CI [.939, .956]. Mean ratings increased from G1 (M = 11.40, SD = 2.75) to G2 (M = 11.57, SD = 2.91) and G3 (M = 11.91, SD = 3.23), indicating a modest systematic scoring-level difference that was counted as disagreement by the absolute-agreement model.

Minimum reporting checklist

Name the targets and raters or scoring occasions.
State the ICC model: one-way, two-way random, or two-way mixed.
State absolute agreement or consistency.
State single-measure or average-measure reliability.
Report the estimate and 95% confidence interval.
Report F, degrees of freedom, and p-value when relevant.
Describe the practical level of reliability and the intended use.

Language to avoid

Do not write “the raters were 85.9% identical,” “85.9% of variance is valid,” or “the scores are accurate because ICC was significant.” Avoid reporting the average-measure coefficient without saying ratings were averaged. Do not call Cronbach’s alpha an inter-rater ICC. Do not use p = .000; write p < .001.

When the rater sample is fixed, do not claim the result generalizes to all raters. When the target sample is narrow, do not assume the coefficient will be the same in a broader population.

Brief reporting versions

Use caseSuggested wording
Methods sectionInter rater reliability was estimated with a two-way random-effects, absolute-agreement, single-measure ICC because the three raters were treated as a sample from a broader rater population and individual ratings were of interest.
Results sectionSingle-rater agreement was good, ICC(2,1) = .859, 95% CI [.836, .878], F(648, 1296) = 20.25, p < .001.
Average-score workflowThe mean of three ratings showed excellent reliability, ICC(2,3) = .948, 95% CI [.939, .956].
Limitations sectionThe worked columns represent sequential scoring occasions rather than independently sampled human raters, so generalization to human-rater behavior is illustrative.

For related reporting foundations, see the site guides on confidence intervals, p-values, null and alternative hypotheses, and effect size.

16

Inter rater reliability PDF, Excel and software downloads

Open the exact Python, R, SPSS, and Excel analysis files.

The downloadable inter rater reliability reports and worked Excel workbook allow readers to verify every result. The inter rater reliability audit is strongest when Python, R, SPSS, and Excel use the same 649 targets, three rating columns, two-way random-effects model, absolute-agreement definition, and single-measure primary coefficient.

17

Inter rater reliability verification sources and documentation

Analysis artifacts and internal method guides used to confirm the model, calculations, and interpretation.

This inter rater reliability guide is grounded in the exact analysis outputs rather than an unverified software screenshot. Each inter rater reliability value is traceable to a named artifact. The SPSS output confirms the confidence intervals and F test, the Excel workbook exposes the formula and data ledger, and the site’s ICC guides explain how model, agreement definition, and measurement unit change the reported coefficient.

Verified SPSS output

The SPSS inter rater reliability report documents 649 valid cases, the rating descriptives, inter-item correlations, ICC(2,1) = .859, 95% CI [.836, .878], average-measure ICC = .948, and F(648, 1296) = 20.246.

Worked Excel audit

The worked Excel analysis contains the targets-by-raters data, target means, ANOVA calculations, ICC formula, diagnostics, reporting sheet, and independent zero-difference verification.

Related ICC methodology

Use the internal intraclass correlation coefficient guide and the detailed ICC formula and interpretation guide to compare one-way, two-way, consistency, absolute-agreement, single-measure, and average-measure choices.

18

Inter rater reliability FAQs

Answers to the questions most often missed in brief reliability explanations.

What is inter rater reliability?

Inter rater reliability is the reproducibility of ratings assigned by different raters to the same targets. It indicates whether another suitably trained rater would produce a similar score or category under the same conditions.

What does inter rater reliability mean in psychology?

In psychology, inter rater reliability describes agreement among observers, coders, interviewers, diagnosticians, or clinicians. It is used for behavior coding, symptom ratings, diagnostic categories, interview scores, and other judgments that may vary across evaluators.

How do you calculate inter rater reliability?

Arrange the same targets in rows and raters in columns, choose a coefficient appropriate for the score scale and design, compute the required agreement statistic, and report a confidence interval. For continuous ratings in a complete two-way design, ICC(2,1) can be calculated from target, rater, and residual ANOVA mean squares.

Which ICC should be used for inter rater reliability?

Use ICC(2,1) when raters are treated as random, exact agreement matters, and one rating is used. Use an average-measure version when the operational score is the mean of several raters. Fixed-rater designs may require a two-way mixed model.

What is a good inter rater reliability score?

Values from 0.75 to 0.90 are often described as good and values above 0.90 as excellent, but requirements depend on the decision, consequences of error, and lower confidence bound. The worked ICC of .859 is good for one rater, with a 95% CI from .836 to .878.

What is an acceptable inter rater reliability?

There is no universal acceptable value. Exploratory research may tolerate moderate reliability, while high-stakes individual decisions often require a coefficient above .90 and a sufficiently high lower confidence bound. The threshold should be justified before analysis.

What is high inter rater reliability?

High inter rater reliability means stable target differences dominate disagreement introduced by raters and residual error. It does not automatically mean the ratings are valid, unbiased, or correct.

What is low inter rater reliability?

Low inter rater reliability means scores depend substantially on which rater performed the assessment. Causes may include an unclear rubric, insufficient training, ambiguous targets, inconsistent conditions, restricted target range, or an unsuitable statistical model.

What is the difference between inter and intra rater reliability?

Inter rater reliability compares different raters. Intra rater reliability evaluates whether the same rater gives stable scores when rating the targets again. The designs address different sources of measurement error.

Is inter rater reliability the same as agreement?

It is a form of agreement analysis, but the exact meaning depends on the coefficient. An absolute-agreement ICC requires score levels to match. A consistency ICC allows systematic rater offsets. Percent agreement and kappa use different definitions for categorical data.

Can Pearson correlation measure inter rater reliability?

Pearson correlation measures linear association, not exact agreement. Two raters can correlate perfectly while one always scores higher. For continuous ratings, an ICC or agreement analysis is usually more appropriate.

Can Cronbach’s alpha be used for inter rater reliability?

Alpha may reflect consistency among columns, but it does not directly estimate single-rater absolute agreement. In this example alpha is .951 while ICC(2,1) is .859. The ICC is the primary result for one random rater and exact agreement.

When should Cohen’s kappa be used?

Use Cohen’s kappa when two raters assign nominal categories to the same targets. Use weighted kappa for ordered categories when near disagreements should receive partial credit.

When should Fleiss kappa be used?

Fleiss kappa is commonly used for nominal ratings from more than two raters. It is not a replacement for an ICC when the ratings are continuous numeric scores.

What is absolute-agreement inter rater reliability?

Absolute-agreement reliability requires raters to assign the same score level. A systematic difference in rater means counts as disagreement. This definition fits grading, clinical measurement, and other settings where the actual score determines the decision.

What is consistency inter rater reliability?

Consistency reliability focuses on whether raters rank targets similarly and ignores a constant strictness or leniency difference. It may be appropriate when scores are later standardized or only relative ordering matters.

Why is average-measure ICC higher than single-measure ICC?

Averaging multiple ratings reduces idiosyncratic error. In the worked example, single-measure ICC is .859 and the reliability of the average of three ratings is .948. Report the average coefficient only when the mean of three ratings is actually used.

What does ICC(2,1) mean?

ICC(2,1) denotes a two-way random-effects, single-measure coefficient. In the absolute-agreement form used here, both targets and raters are random, systematic rater offsets count as disagreement, and reliability refers to one rater.

What does ICC(2,k) mean?

ICC(2,k) estimates the reliability of the mean of k random raters under a two-way random-effects model. In this example k = 3 and the average-measure coefficient is .948.

Can inter rater reliability be negative?

An estimated ICC can be negative when residual disagreement exceeds stable target variation. The result indicates no dependable reliability under the chosen model and should prompt investigation rather than being interpreted as a meaningful negative proportion.

Does inter rater reliability require normal data?

Classical ICC confidence intervals and tests rely on ANOVA assumptions, but exact normality of each rater column is not the only concern. Severe outliers, skew, heteroscedasticity, floor or ceiling effects, and model mismatch should be assessed. Large balanced samples make the point estimate more stable but do not fix design problems.

How many raters are needed?

Two raters can support an inter-rater analysis, but more raters allow the reliability of an average to be evaluated and may better represent the intended rater population. Rater count should follow the operational scoring process and precision requirements.

How many targets are needed?

Required target count depends on expected ICC, desired confidence-interval width, number of raters, and decision threshold. More targets improve precision, but representative target sampling is as important as raw sample size.

How can inter rater reliability be improved?

Clarify the construct and rubric, add anchored examples, train and calibrate raters, standardize conditions, keep ratings independent, monitor drift, and investigate disagreement patterns. Do not inflate reliability by deleting difficult cases without a prespecified rule.

How is inter rater reliability reported?

Report the ICC model, agreement definition, single or average unit, number of targets and raters, coefficient, 95% confidence interval, F statistic, degrees of freedom, p-value, and practical interpretation. Name the exact raters or scoring occasions.

How do I calculate inter rater reliability in SPSS?

Use Analyze → Scale → Reliability Analysis, select the rating variables, request intraclass correlation coefficients, and choose the model, agreement definition, and confidence interval. For this example use a two-way random model, absolute agreement, and single measures.

How do I calculate inter rater reliability in Excel?

Use a balanced targets-by-raters matrix, calculate target, rater, and residual sums of squares and mean squares, then apply the ICC(2,1) formula. Audit ranges carefully and verify the decomposition against an independent implementation.

What did the worked example show?

The worked inter rater reliability analysis found ICC(2,1) = .859, 95% CI [.836, .878], F(648, 1296) = 20.246, p < .001, for 649 targets and three ratings. The average of all three ratings had reliability .948.

+

Related statistical guides

Continue the inter rater reliability workflow with closely connected agreement, reliability, correlation, and measurement-error methods.

↑ Back to top