Meta-analysis, systematic review and umbrella review: what is the difference?
Meta-analyses, systematic reviews and umbrella reviews all synthesise research evidence, but they serve different purposes. Learn…
Read more >How can we determine whether the prevalence of a symptom changed after treatment when the same patients were assessed before and after the intervention? For two related measurements of a dichotomous variable, the appropriate method is usually McNemar’s test, which accounts for the direction of change in individual patients.
In this article, we explain concordant and discordant pairs, why the standard chi-square test is inappropriate, and how to report the results correctly.
Categorical data describe the assignment of observations to specific categories. Examples include:
Data are paired when two results can be unequivocally linked to form a pair. Most commonly, these are two assessments of the same person performed at different times or using two different methods.
An example would be the assessment of pain before and after treatment. Each patient contributes two related responses to the analysis: “pain present” or “pain absent”.
Paired categorical data also arise when two diagnostic methods are applied to every patient and their positive and negative results are compared. These are not two independent groups, because both methods were assessed in the same individuals.
If, by contrast, the result of method A was obtained in one group of patients and the result of method B in another group, the observations are independent and McNemar’s test is not appropriate.
Two measurements of a dichotomous variable can be presented in a 2 × 2 table. Suppose we assess the presence of a particular symptom before and after treatment.
| Symptom after treatment: present | Symptom after treatment: absent | |
| Symptom before treatment: present | 35 | 18 |
| Symptom before treatment: absent | 6 | 41 |
Each patient belongs to one of four groups:
In this example, 35 patients had the symptom at both assessments, while 41 patients did not have it at either assessment. These are concordant pairs because the outcome category did not change.
The symptom disappeared in 18 patients and appeared at the second assessment in 6 patients. These are discordant pairs because the outcome changed between assessments.
– patients who changed from “yes” to “no”;
– patients who changed from “no” to “yes”.
Concordant pairs indicate that the result remained unchanged, but they provide no information about the direction of change. McNemar’s test primarily evaluates whether the number of changes in one direction differs from the number of changes in the opposite direction.
In the example above, we compare the 18 patients whose symptom disappeared with the 6 patients whose symptom appeared. These 24 discordant pairs provide the key information about the direction of change.
If the numbers of changes in the two directions are similar, the overall proportion may remain approximately unchanged. If one direction clearly predominates, this suggests a systematic change between the two assessments.
McNemar’s test is used to compare two related measurements of a variable with two categories. It evaluates whether changes in one direction occur as frequently as changes in the opposite direction. [1,2]
In practice, it can answer questions such as:
The key point for clinicians is that the test uses the direction of change within the same individuals rather than only the two overall proportions.
Before treatment, the symptom may be present in 53% of patients and, after treatment, in 41%. The absolute difference is therefore 12 percentage points.
This descriptive comparison is useful, but it does not account for the fact that both proportions refer to the same individuals. It also does not show how many patients improved and how many changed in the opposite direction.
The same difference between proportions can result from different patterns of within-person change. The analysis should therefore use the complete table of paired outcomes rather than only the proportions observed before and after treatment.
McNemar’s test does not compare two proportions as if they came from independent groups. It compares changes occurring in two opposite directions.
The standard chi-square test for two groups assumes that observations are independent. This assumption is not met when the same patient is assessed twice.
Applying a standard chi-square test to two measurements from the same individuals discards information about which results form pairs. The analysis uses only the numbers or proportions observed at the two assessments, as if they came from different patients.
McNemar’s test is the appropriate method for two related dichotomous measurements. Presenting data in a 2 × 2 table does not automatically mean that a standard chi-square test should be used. The relationship between the observations must also be considered when selecting the statistical method.
McNemar’s test is appropriate when:
The independence assumption applies to the pairs rather than to the observations within each pair. The two results obtained from the same patient are dependent, whereas the pair contributed by one patient should be independent of the pair contributed by another patient.
McNemar’s test is not limited to before-and-after studies. It can also be used to compare two diagnostic methods applied to the same patients and in individually matched case-control studies.
The standard version of McNemar’s test uses an approximation that performs best when the number of discordant pairs is sufficiently large. Therefore, what matters is not only the total study sample but, more importantly, the number of patients whose outcome actually changed between assessments.
A study may include many participants, but almost all of them may have the same result at both assessments. The number of discordant pairs will then be small, and the standard version of the test may not be the best choice.
In this situation, an exact McNemar’s test or another method intended for a small number of discordant pairs may be used.[3] The publication should specify which version of the test was applied and why.
A mechanical rule based on a single numerical threshold should be avoided. The choice of test should reflect the data distribution, number of discordant pairs, and options available in the statistical software.
No. This is an important distinction.
McNemar’s test evaluates whether two related assessments produce the same proportions of positive results. It does not directly determine whether the methods agree for individual patients.
Two methods may produce a similar overall proportion of positive results while frequently disagreeing about which patients have a positive result. McNemar’s test may show no difference in such a situation because the numbers of discordant results in the two directions are similar, even though agreement between the methods is poor.
– McNemar’s test: Do the proportions of positive results differ between the methods?
– Measure of agreement: How often do both methods assign the same category to the same patient?
If the purpose of the study is to evaluate agreement between classifications, the percentage agreement should be reported and an appropriate measure, such as Cohen’s kappa, should be considered. McNemar’s test alone should not be interpreted as a test of agreement.
McNemar’s test is frequently used to compare two diagnostic tests applied to the same individuals. However, its interpretation depends on whether a reliable reference standard is available.
To compare the sensitivity of two methods, the analysis is conducted among patients whose disease was confirmed using the reference standard. To compare specificity, the analysis is conducted among individuals without the disease. In both cases, the test results are paired because both methods were applied to the same patients.[4]
If no reference standard is available, McNemar’s test can compare the proportions of positive results produced by the two methods. However, it cannot by itself determine which method has greater diagnostic accuracy.
The p-value does not indicate the magnitude of the change or its clinical relevance. The result of McNemar’s test should therefore be supplemented with information describing the magnitude and direction of the effect.
The results should include:
The matched-pairs odds ratio is based on the comparison between the numbers of changes in the two directions. The direction in which it is calculated must always be clearly defined, because reversing the coding produces the reciprocal value.
The absolute difference between proportions is often easier to interpret clinically. It should be accompanied by a confidence interval indicating the precision of the estimate.
A standard McNemar’s test requires two results for every person included in the analysis. A patient with a baseline assessment but no follow-up assessment does not form a complete pair and will not be included in a standard analysis.
The publication should distinguish between:
Missing data may do more than reduce statistical power. If the absence of the second measurement is related to disease severity, treatment effectiveness, or adverse events, a complete-pair analysis may produce biased results.
A missing response should not be replaced with the most frequent category, nor should it automatically be assumed that the patient’s condition remained unchanged. More complex studies may require models for repeated binary outcomes or an appropriately planned method for handling missing data.
A quantitative variable is sometimes divided into two categories, such as “normal” and “abnormal”. This may be justified if the cut-off has an established clinical meaning and the research question specifically concerns whether that threshold has been crossed.
However, a variable should not be categorised solely to permit the use of McNemar’s test. Dichotomising a quantitative variable results in a loss of information, reduces statistical power, and makes the result dependent on the selected cut-off point.
If exact numerical values are available, the primary analysis should generally use the quantitative data. A categorical analysis may be presented as an additional analysis when it has a clear clinical rationale.
The standard McNemar’s test applies to variables with two categories. If the outcome can take three or more values, a different method is required.
Depending on the research question, possible methods include:
These methods do not evaluate exactly the same hypothesis. The general term “extended McNemar’s test” should therefore not be used without specifying which test was actually performed.
If the same dichotomous variable was assessed at three or more time points, a single McNemar’s test does not account for the complete study structure.
A simple extension is Cochran’s Q test, which evaluates whether the proportion of positive results is the same across all related measurements. [7] If the overall test indicates a difference, appropriately planned pairwise comparisons may be conducted while accounting for multiple testing.
In more complex study designs, particularly when data are missing, assessment times differ, or additional variables need to be included, models for repeated binary outcomes may be more appropriate.
The most common problems include:
The methods section should state that the observations were paired, identify the variable that was analysed, and specify which version of the test was used.
“The change in symptom prevalence between the baseline and follow-up assessments was analysed using McNemar’s test for paired dichotomous data. Because of the small number of discordant pairs, the exact version of the test was used.”
“The proportion of patients reporting the symptom decreased from 53% before treatment to 41% after treatment, corresponding to an absolute reduction of 12 percentage points. The symptom resolved in 18 patients and developed in 6 patients. The difference between the directions of change was statistically significant in McNemar’s test (p = 0.034; n = 100 complete pairs).”
If a matched-pairs odds ratio or a confidence interval for the difference between proportions is reported, these values should be added to the description. The results should allow the reader to assess both statistical significance and the clinical importance of the change.
– McNemar’s test is used for two related measurements of a dichotomous variable, such as the presence of a symptom before and after treatment.
– The most important information is provided by discordant pairs-patients whose result changed in one direction or the other.
– The standard chi-square test is inappropriate for two assessments of the same individuals because it assumes that observations are independent.
– The results should include the proportions at both assessments, numbers of changes in both directions, effect size, confidence interval, and p-value.
– McNemar’s test compares paired proportions but is not a test of agreement between diagnostic methods.
– When there are more categories or measurements, or when data are missing, other methods are required, such as the Stuart–Maxwell test, Cochran’s Q test, or models for repeated binary outcomes.
Meta-analyses, systematic reviews and umbrella reviews all synthesise research evidence, but they serve different purposes. Learn…
Read more >What are paired categorical data? Categorical data describe the assignment of observations to specific categories. Examples…
Read more >How should two measurements from the same patients be analysed? Learn when to use the paired…
Read more >What is the difference between paired and independent data? Explore examples from medical research and learn…
Read more >Assessing the normality of data distributions is one of the first stages of statistical analysis in…
Read more >The Idea for This Blog Had Been Developing for Years The idea of creating this blog…
Read more >