How to Analyse Paired Categorical Data: McNemar’s Test

How can we determine whether the prevalence of a symptom changed after treatment when the same patients were assessed before and after the intervention? For two related measurements of a dichotomous variable, the appropriate method is usually McNemar’s test, which accounts for the direction of change in individual patients.

In this article, we explain concordant and discordant pairs, why the standard chi-square test is inappropriate, and how to report the results correctly.

What are paired categorical data?

Categorical data describe the assignment of observations to specific categories. Examples include:

  • presence or absence of a symptom,
  • a positive or negative test result,
  • occurrence or non-occurrence of a complication,
  • response or no response to treatment,
  • a “yes” or “no” answer.

Data are paired when two results can be unequivocally linked to form a pair. Most commonly, these are two assessments of the same person performed at different times or using two different methods.

An example would be the assessment of pain before and after treatment. Each patient contributes two related responses to the analysis: “pain present” or “pain absent”.

Paired categorical data also arise when two diagnostic methods are applied to every patient and their positive and negative results are compared. These are not two independent groups, because both methods were assessed in the same individuals.

If, by contrast, the result of method A was obtained in one group of patients and the result of method B in another group, the observations are independent and McNemar’s test is not appropriate.

What does the table for two paired measurements look like?

Two measurements of a dichotomous variable can be presented in a 2 × 2 table. Suppose we assess the presence of a particular symptom before and after treatment.

Symptom after treatment: presentSymptom after treatment: absent
Symptom before treatment: present3518
Symptom before treatment: absent641

Each patient belongs to one of four groups:

  1. The symptom was present before treatment and remained present after treatment.
  2. The symptom was present before treatment but disappeared after treatment.
  3. The symptom was absent before treatment but appeared after treatment.
  4. The symptom was absent both before and after treatment.

In this example, 35 patients had the symptom at both assessments, while 41 patients did not have it at either assessment. These are concordant pairs because the outcome category did not change.

The symptom disappeared in 18 patients and appeared at the second assessment in 6 patients. These are discordant pairs because the outcome changed between assessments.

Concordant and discordant pairs: which are most important?

In an analysis of change, the discordant pairs provide the most important information:

– patients who changed from “yes” to “no”;
– patients who changed from “no” to “yes”.

Concordant pairs indicate that the result remained unchanged, but they provide no information about the direction of change. McNemar’s test primarily evaluates whether the number of changes in one direction differs from the number of changes in the opposite direction.

In the example above, we compare the 18 patients whose symptom disappeared with the 6 patients whose symptom appeared. These 24 discordant pairs provide the key information about the direction of change.

If the numbers of changes in the two directions are similar, the overall proportion may remain approximately unchanged. If one direction clearly predominates, this suggests a systematic change between the two assessments.

What question does McNemar’s test answer?

McNemar’s test is used to compare two related measurements of a variable with two categories. It evaluates whether changes in one direction occur as frequently as changes in the opposite direction. [1,2]

In practice, it can answer questions such as:

  • Did the proportion of patients reporting a symptom change after treatment?
  • Do two diagnostic methods produce different proportions of positive results?
  • Did the frequency of a “yes/no” response change after an intervention?
  • Does a dichotomous exposure differ between individually matched cases and controls?

The key point for clinicians is that the test uses the direction of change within the same individuals rather than only the two overall proportions.

McNemar’s test and the comparison of two proportions

Before treatment, the symptom may be present in 53% of patients and, after treatment, in 41%. The absolute difference is therefore 12 percentage points.

This descriptive comparison is useful, but it does not account for the fact that both proportions refer to the same individuals. It also does not show how many patients improved and how many changed in the opposite direction.

The same difference between proportions can result from different patterns of within-person change. The analysis should therefore use the complete table of paired outcomes rather than only the proportions observed before and after treatment.

McNemar’s test does not compare two proportions as if they came from independent groups. It compares changes occurring in two opposite directions.

Why is the standard chi-square test inappropriate?

The standard chi-square test for two groups assumes that observations are independent. This assumption is not met when the same patient is assessed twice.

Applying a standard chi-square test to two measurements from the same individuals discards information about which results form pairs. The analysis uses only the numbers or proportions observed at the two assessments, as if they came from different patients.

McNemar’s test is the appropriate method for two related dichotomous measurements. Presenting data in a 2 × 2 table does not automatically mean that a standard chi-square test should be used. The relationship between the observations must also be considered when selecting the statistical method.

When can McNemar’s test be used?

McNemar’s test is appropriate when:

  • the outcome variable has two categories, such as “yes/no” or “positive/negative”,
  • two related measurements or two matched observations are being analysed,
  • each pair is independent of the other pairs,
  • both results are available for every person included in the analysis,
  • the categories have the same meaning at both assessments.

The independence assumption applies to the pairs rather than to the observations within each pair. The two results obtained from the same patient are dependent, whereas the pair contributed by one patient should be independent of the pair contributed by another patient.

McNemar’s test is not limited to before-and-after studies. It can also be used to compare two diagnostic methods applied to the same patients and in individually matched case-control studies.

Standard or exact McNemar’s test?

The standard version of McNemar’s test uses an approximation that performs best when the number of discordant pairs is sufficiently large. Therefore, what matters is not only the total study sample but, more importantly, the number of patients whose outcome actually changed between assessments.

A study may include many participants, but almost all of them may have the same result at both assessments. The number of discordant pairs will then be small, and the standard version of the test may not be the best choice.

In this situation, an exact McNemar’s test or another method intended for a small number of discordant pairs may be used.[3] The publication should specify which version of the test was applied and why.

A mechanical rule based on a single numerical threshold should be avoided. The choice of test should reflect the data distribution, number of discordant pairs, and options available in the statistical software.

Does McNemar’s test assess agreement between two methods?

No. This is an important distinction.

McNemar’s test evaluates whether two related assessments produce the same proportions of positive results. It does not directly determine whether the methods agree for individual patients.

Two methods may produce a similar overall proportion of positive results while frequently disagreeing about which patients have a positive result. McNemar’s test may show no difference in such a situation because the numbers of discordant results in the two directions are similar, even though agreement between the methods is poor.

McNemar’s test and measures of agreement therefore answer different questions:

– McNemar’s test: Do the proportions of positive results differ between the methods?

– Measure of agreement: How often do both methods assign the same category to the same patient?

If the purpose of the study is to evaluate agreement between classifications, the percentage agreement should be reported and an appropriate measure, such as Cohen’s kappa, should be considered. McNemar’s test alone should not be interpreted as a test of agreement.

McNemar’s test when comparing diagnostic methods

McNemar’s test is frequently used to compare two diagnostic tests applied to the same individuals. However, its interpretation depends on whether a reliable reference standard is available.

To compare the sensitivity of two methods, the analysis is conducted among patients whose disease was confirmed using the reference standard. To compare specificity, the analysis is conducted among individuals without the disease. In both cases, the test results are paired because both methods were applied to the same patients.[4]

If no reference standard is available, McNemar’s test can compare the proportions of positive results produced by the two methods. However, it cannot by itself determine which method has greater diagnostic accuracy.

How should the effect size be presented with McNemar’s test?

The p-value does not indicate the magnitude of the change or its clinical relevance. The result of McNemar’s test should therefore be supplemented with information describing the magnitude and direction of the effect.

The results should include:

  1. The proportion of positive results at the first assessment.
  2. The proportion of positive results at the second assessment.
  3. The absolute difference between the proportions in percentage points.
  4. The numbers of changes in both directions.
  5. A relative measure, such as the matched-pairs odds ratio, if relevant to the research question.
  6. A confidence interval and p-value.

The matched-pairs odds ratio is based on the comparison between the numbers of changes in the two directions. The direction in which it is calculated must always be clearly defined, because reversing the coding produces the reciprocal value.

The absolute difference between proportions is often easier to interpret clinically. It should be accompanied by a confidence interval indicating the precision of the estimate.

Missing data and incomplete pairs

A standard McNemar’s test requires two results for every person included in the analysis. A patient with a baseline assessment but no follow-up assessment does not form a complete pair and will not be included in a standard analysis.

The publication should distinguish between:

  • the number of people enrolled in the study,
  • the number assessed at each time point,
  • the number of complete pairs included in McNemar’s test.

Missing data may do more than reduce statistical power. If the absence of the second measurement is related to disease severity, treatment effectiveness, or adverse events, a complete-pair analysis may produce biased results.

A missing response should not be replaced with the most frequent category, nor should it automatically be assumed that the patient’s condition remained unchanged. More complex studies may require models for repeated binary outcomes or an appropriately planned method for handling missing data.

Should a quantitative variable be converted into a categorical variable?

A quantitative variable is sometimes divided into two categories, such as “normal” and “abnormal”. This may be justified if the cut-off has an established clinical meaning and the research question specifically concerns whether that threshold has been crossed.

However, a variable should not be categorised solely to permit the use of McNemar’s test. Dichotomising a quantitative variable results in a loss of information, reduces statistical power, and makes the result dependent on the selected cut-off point.

If exact numerical values are available, the primary analysis should generally use the quantitative data. A categorical analysis may be presented as an additional analysis when it has a clear clinical rationale.

What should be done when there are more than two categories?

The standard McNemar’s test applies to variables with two categories. If the outcome can take three or more values, a different method is required.

Depending on the research question, possible methods include:

  • the Stuart–Maxwell test, which evaluates whether the distribution of categories changed between two measurements;
  • Bowker’s test, which assesses the symmetry of changes between categories. [5,6]

These methods do not evaluate exactly the same hypothesis. The general term “extended McNemar’s test” should therefore not be used without specifying which test was actually performed.

What should be done when there are more than two measurements?

If the same dichotomous variable was assessed at three or more time points, a single McNemar’s test does not account for the complete study structure.

A simple extension is Cochran’s Q test, which evaluates whether the proportion of positive results is the same across all related measurements. [7] If the overall test indicates a difference, appropriately planned pairwise comparisons may be conducted while accounting for multiple testing.

In more complex study designs, particularly when data are missing, assessment times differ, or additional variables need to be included, models for repeated binary outcomes may be more appropriate.

Common errors in the analysis of paired categorical data

The most common problems include:

  1. Using the standard chi-square test for two assessments performed in the same patients.
  2. Reporting only the pre- and post-intervention proportions without showing the directions of change.
  3. Failing to consider the number of discordant pairs when selecting the version of the test.
  4. Reporting a p-value without the magnitude and direction of the effect.
  5. Interpreting McNemar’s test as a test of agreement between two methods.
  6. Using McNemar’s test for a variable with more than two categories.
  7. Failing to account for incomplete pairs.
  8. Dichotomising a quantitative variable without a clinical justification.

How should McNemar’s test be described in a publication?

The methods section should state that the observations were paired, identify the variable that was analysed, and specify which version of the test was used.

Example wording for the Methods section:

“The change in symptom prevalence between the baseline and follow-up assessments was analysed using McNemar’s test for paired dichotomous data. Because of the small number of discordant pairs, the exact version of the test was used.”

Example wording for the Results section:

“The proportion of patients reporting the symptom decreased from 53% before treatment to 41% after treatment, corresponding to an absolute reduction of 12 percentage points. The symptom resolved in 18 patients and developed in 6 patients. The difference between the directions of change was statistically significant in McNemar’s test (p = 0.034; n = 100 complete pairs).”

If a matched-pairs odds ratio or a confidence interval for the difference between proportions is reported, these values should be added to the description. The results should allow the reader to assess both statistical significance and the clinical importance of the change.

Key takeaways for medical researchers

– McNemar’s test is used for two related measurements of a dichotomous variable, such as the presence of a symptom before and after treatment.

– The most important information is provided by discordant pairs-patients whose result changed in one direction or the other.

– The standard chi-square test is inappropriate for two assessments of the same individuals because it assumes that observations are independent.

– The results should include the proportions at both assessments, numbers of changes in both directions, effect size, confidence interval, and p-value.

– McNemar’s test compares paired proportions but is not a test of agreement between diagnostic methods.

– When there are more categories or measurements, or when data are missing, other methods are required, such as the Stuart–Maxwell test, Cochran’s Q test, or models for repeated binary outcomes.

References:

  1. Agresti A. Categorical Data Analysis. 3rd ed. Hoboken: John Wiley & Sons; 2013.
  2. McNemar Q. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika. 1947;12:153–157. doi:10.1007/BF02295996.
  3. Sundjaja JH, Shrestha R, Krishan K. McNemar and Mann-Whitney U Tests. StatPearls. Treasure Island: StatPearls Publishing; 2023.
  4. Fagerland MW, Lydersen S, Laake P. The McNemar test for binary matched-pairs data: mid-p and asymptotic are better than exact conditional. BMC Med Res Methodol. 2013;13:91.
  5. Kim S, Lee W. Does McNemar’s test compare the sensitivities and specificities of two diagnostic tests? Stat Methods Med Res. 2017;26:142–154.
  6. Stuart A. A test for homogeneity of the marginal distributions in a two-way classification. Biometrika. 1955;42:412–416.
  7. Bowker AH. A test for symmetry in contingency tables. J Am Stat Assoc. 1948;43:572–574.
  8. Cochran WG. The comparison of percentages in matched samples. Biometrika. 1950;37:256–266.

Contact us

If you have any questions,
contact us.
We will respond within 24 hours.

+48 601 40 77 71

Show the e-mail address