Intercoder reliability (ICR) — having two or more researchers independently code the same data and measuring how much they agree — is one of the most debated rigor practices in qualitative research. Borrowed from content analysis and quantitative traditions, it’s sometimes treated as a mandatory box to check and sometimes rejected outright as incompatible with interpretive qualitative work. The honest answer sits between those two positions: ICR is a genuinely useful tool in some designs, and a poor fit in others.
What ICR is actually checking
When multiple coders apply the same codebook to the same data independently, then compare results, the level of agreement is a signal — not proof — of how clearly the codebook is defined and how consistently it can be applied by someone other than the person who wrote it. High agreement suggests a codebook that would likely be interpreted similarly by a new coder; low agreement usually means the code definitions themselves need work, not that the coders did something wrong.
It isn’t right for every design
ICR can genuinely improve transparency and — just as importantly — team dialogue, since disagreements force coders to articulate why they interpreted a segment differently, which often sharpens the codebook itself. But it isn’t a universal requirement, and which statistic even makes sense to calculate depends heavily on how the coding units are defined in the first place (O’Connor & Joffe, 2020; Coleman, Ragan, & Dari, 2024; Rau & Shih, 2021). A single-coder interpretive phenomenological study and a five-coder content analysis of thousands of open-ended survey responses are doing fundamentally different kinds of coding — treating both the same way for reliability purposes doesn’t make sense.
The Cohen’s kappa trap
Cohen’s kappa is the most commonly reported ICR statistic, but it isn’t valid for every nominal coding situation it gets applied to (O’Connor & Joffe, 2020; Rau & Shih, 2021). Kappa’s behavior is sensitive to how balanced the code categories are and how the coding units were segmented — in some configurations it can produce a low score even when raw percentage agreement is high, a well-documented statistical quirk that catches researchers off guard when they report kappa without checking whether it’s actually appropriate for their data. Reporting simple percentage agreement alongside kappa, and being explicit about how coding units were defined, avoids presenting a single number as more meaningful than it is.
What to specify before you start
A defensible ICR process is planned in advance, not calculated as an afterthought once two people’s codes don’t match. At minimum, that plan should specify: which statistic will be used and why it fits the coding structure; what threshold counts as acceptable agreement; which subset of the data both coders will code independently before dividing up the rest; the process for resolving disagreements once they’re found; and an audit trail for any codebook revisions made as a result (Halpin, 2024). Without deciding these in advance, ICR calculations risk becoming a number reported to satisfy a reviewer rather than a genuine check on the coding.
The bottom line
ICR is worth using when a project genuinely involves multiple coders applying a shared codebook and the study’s credibility depends on showing that codebook can be applied consistently. It’s worth skipping — or replacing with a different rigor practice like peer debriefing or an audit trail — when the analysis is inherently interpretive and single-coder, where forcing a reliability statistic can misrepresent what the analysis actually is.
Coding agreement in Byleron QDA: when a project has more than one coder, comparing two people’s coding of the same document side by side — segment by segment — makes disagreements visible immediately, rather than only surfacing once someone runs the numbers at the end.
References
- Coleman, M. L., Ragan, M., & Dari, T. (2024). Intercoder Reliability for Use in Qualitative Research and Evaluation. Measurement and Evaluation in Counseling and Development, 57, 136–146. https://doi.org/10.1080/07481756.2024.2303715
- Halpin, S. N. (2024). Inter-Coder Agreement in Qualitative Coding: Considerations for its Use. American Journal of Qualitative Research. https://doi.org/10.29333/ajqr/14887
- O’Connor, C., & Joffe, H. (2020). Intercoder Reliability in Qualitative Research: Debates and Practical Guidelines. International Journal of Qualitative Methods, 19. https://doi.org/10.1177/1609406919899220
- Rau, G., & Shih, Y.-S. (2021). Evaluation of Cohen’s kappa and other measures of inter-rater agreement for genre analysis and other nominal data. Journal of English for Academic Purposes, 53, 101026. https://doi.org/10.1016/j.jeap.2021.101026

