A MAXQDA intercoder reliability check answers the question reviewers ask of any codebook-based study: would a second trained coder, using the same codebook independently, have coded the data the same way? MAXQDA's Intercoder Agreement feature reports percentage agreement and a kappa corrected for chance. Our thematic analysis guide touched on the kappa debate; this post goes further, covering project set-up, the three comparison levels, the output, the arithmetic behind each coefficient and reporting.
When intercoder reliability is appropriate, and when it is not
Intercoder reliability (ICR) suits designs that treat the codebook as a shared instrument: qualitative content analysis, codebook or template thematic analysis, framework analysis, deductive coding and team projects that split the data. High agreement shows that the code definitions are clear enough for someone else to apply. Reflexive thematic analysis sees it differently: themes come from the researcher's interpretation, so agreement is not a quality marker and is often not calculated. Both positions are legitimate; the mistake is mixing them. Decide at the design stage and justify it in your methods chapter.
Preparing a MAXQDA project for an intercoder agreement check
- Settle on a working codebook. Each code needs a definition, inclusion and exclusion criteria and an example extract. Vague definitions cause most disagreements.
- Agree on the coding unit. Sentences, whole answers or free-length passages? The looser this rule, the more disagreements concern boundaries rather than meaning.
- Draw a pilot sample. Double-coding 10–20% of the material is typical; pick documents from different participants or sites, not just the first transcripts.
- Code independently. Each coder works in their own project copy without seeing the other's codings; the files are then combined with Home → Merge Projects.
- Arrange the copies. MAXQDA compares identically named documents placed in two different document groups or sets, such as ‘Coder A’ and ‘Coder B’. The texts must match exactly, as one extra line break can shift segment positions, so check that character counts are equal.
MAXQDA intercoder agreement: three levels of comparison
Open Analysis → Intercoder Agreement, choose each coder's document group or set, select the codes and pick a comparison level:
- Code occurrence in the document: did both coders use the code somewhere in the document? Enough when you only need to know whether a theme appears in an interview. Codes neither coder assigned can be ignored or counted as matches.
- Code frequency in the document: did both coders use the code equally often?
- Segment level with a minimum overlap rate (default 90%): did the other coder give the same code to a passage overlapping each coded segment by at least that percentage? This is the strictest option and the one journals usually expect.
Each level produces a code-specific results table (agreements, disagreements and a percentage for each code, plus a Total row) and a detailed table of documents or segments. At segment level you can list only the disagreements, jump to each passage and copy one coder's coding across once the pair agrees. The Kappa symbol in the results table adds a chance-corrected coefficient.
MAXQDA kappa versus percentage agreement: a worked example
Percentage agreement is agreements ÷ (agreements + disagreements); Turkish theses usually attribute this formula to Miles and Huberman. It is intuitive but ignores agreement that could arise by chance. Kappa removes that share: κ = (Po − Pe) / (1 − Pe), where Po is observed and Pe chance agreement. In the illustrative table below, two coders judged 100 predefined answers against one code.
| Coder A: coded | Coder A: not coded | Total | |
|---|---|---|---|
| Coder B: coded | 55 | 5 | 60 |
| Coder B: not coded | 15 | 25 | 40 |
| Total | 70 | 30 | 100 |
Observed agreement is Po = (55 + 25) / 100 = .80, which clears the usual 80% bar. Cohen's κ estimates chance from how often each coder used the code: Pe = (.70 × .60) + (.30 × .40) = .54, so κ = (.80 − .54) / (1 − .54) = .57, only moderate. Brennan–Prediger κ assumes every category is equally likely by chance, so Pe = 1/k; with k = 2 categories, κ = (.80 − .50) / .50 = .60. MAXQDA applies this at segment level, with k equal to the number of codes in the results table. With 20 codes, chance agreement is only .05, so kappa sits close to the raw percentage. MAXQDA's segment level also counts only passages that at least one coder coded (the ‘neither’ cell is fixed at zero), so the same data would show 55 / 75 = 73%.
Thresholds, disagreement meetings and the codebook loop
No cut-off is universal. The bands below follow the Landis and Koch convention; many supervisors and journals expect at least .60–.70 and treat .80 or above as strong.
| κ value | Landis–Koch label | Practical reading |
|---|---|---|
| ≤ .40 | Poor, slight or fair | Codebook not ready; redefine the codes |
| .41–.60 | Moderate | Revise definitions, add borderline examples, recode |
| .61–.80 | Substantial | Widely accepted minimum for codebook-based studies |
| .81–1.00 | Almost perfect | Strong; check the codes are not trivially easy |
When a code falls short, hold a disagreement resolution meeting: list the disagreements, discuss each passage against the codebook and note the cause (vague definition, missing exclusion rule, boundary problem or slip). Revise the codebook with a version number, date and rationale, have both coders code a fresh sample and rerun the check. Recoding only the disputed passages inflates agreement. Two or three rounds are normal; remaining disagreements are then settled by consensus. If you need an experienced second coder, our qualitative and mixed methods service covers the whole loop.
How to report MAXQDA intercoder reliability in a thesis or article
State the share and selection of double-coded data, the number of codes, the comparison level and overlap threshold, the coefficient and its value before discussion, the lowest single-code value, the number of rounds and how disagreements were resolved. An APA-style example: Two coders independently coded 15% of the transcripts (3 of 20) using the final 24-code codebook. MAXQDA's segment-level intercoder agreement analysis (minimum overlap 90%) showed 86% agreement, Brennan–Prediger κ = .85; agreement for individual codes ranged from 72% to 100%. Disagreements were resolved by consensus. For studies with a quantitative strand, see our guide to mixed methods software.
| Your situation | MAXQDA option / index | Common target |
|---|---|---|
| Reflexive thematic analysis, single analyst | No ICR; reflexivity journal and peer debriefing | Not applicable |
| Codebook thematic or content analysis, coder-defined segments | Segment level (90% overlap); percentage agreement + Brennan–Prediger κ | Agreement ≥ 80%; κ ≥ .60 |
| Only whether each theme appears in each interview | Code occurrence; Kappa (RK) with ≤ 3 codes, shown when unassigned codes count as matches | Agreement ≥ 80% |
| Predefined units (e.g. open-ended survey answers) | Cohen's κ per code in SPSS (Crosstabs), Stata (kap) or Python | κ ≥ .60, ideally ≥ .80 |
| Committee asks for Miles–Huberman | Agreements ÷ (agreements + disagreements), plus κ | ≥ .80 |
Kappa does not validate your themes; it shows whether someone other than you can use your codebook.
Frequently Asked Questions
How do I calculate intercoder reliability in MAXQDA?
Place the two coders' identical document copies in separate document groups or sets, run Analysis → Intercoder Agreement and choose a comparison level. The code-specific results table gives percentage agreement for each code, and at segment level the Kappa symbol adds a Brennan–Prediger coefficient.
What is a good kappa value for qualitative coding?
Under the Landis and Koch convention, .61–.80 is substantial and above .80 almost perfect. Most reviewers accept .60–.70 as a minimum for codebook-based studies, usually alongside percentage agreement of around 80%.
Why is my MAXQDA kappa almost the same as percentage agreement?
At segment level, MAXQDA sets chance agreement to one divided by the number of codes analysed. With many codes that value is tiny, so the correction is small. This follows from the method and is not an error.
Can Celsus run an intercoder reliability check for my study?
Yes. We help build the codebook, provide a trained second coder, run the MAXQDA intercoder agreement analysis and moderate the disagreement rounds. We also write the reliability paragraph for your methods chapter.