Meta-analysis, systematic review and umbrella review: what is the difference?
Meta-analyses, systematic reviews and umbrella reviews all synthesise research evidence, but they serve different purposes. Learn…
Read more >How should two measurements obtained from the same patients be compared, and when should the paired t-test or the Wilcoxon signed-rank test be used?
The analysis of paired quantitative data should focus on individual differences between measurements. The choice of method depends, among other factors, on the distribution of these differences, sample size, the presence of outliers, and the research question.
In this article, we explain how to analyse paired quantitative data, avoid the most common errors, and report the results appropriately in a scientific publication.
Measurements obtained before and after treatment in the same patients do not constitute two independent groups. Each baseline result is linked to a specific follow-up result, and this relationship should be taken into account in the analysis. When analysing paired quantitative data, we primarily assess within-pair differences rather than analysing two sets of measurements separately.
The choice between the paired t-test and the Wilcoxon signed-rank test should not be based solely on the result of the Shapiro–Wilk test. The measurement scale, distribution of the differences, presence of outliers, sample size, and research question should all be considered. It is equally important to report the magnitude and precision of the observed change, rather than only the p-value.
Paired data arise when each observation can be assigned to a specific, related observation. Most commonly, they are two measurements obtained from the same person, for example before and after an intervention.
In a study involving 60 patients assessed twice, we obtain 120 measurements but only 60 independent pairs. Observations from the same patient are generally more similar to each other than measurements from two randomly selected individuals. The statistical analysis should make use of this information.
A detailed explanation of the distinction between paired and independent data is available in the previous article. Here, we focus specifically on the analysis of two related measurements of a quantitative variable.
Suppose we measure systolic blood pressure before treatment and after 12 weeks of therapy. For each patient, we can calculate an individual difference: the post-treatment value minus the baseline value. If blood pressure decreases from 145 to 135 mmHg, the difference is −10 mmHg. If it increases from 130 to 134 mmHg in another patient, the difference is +4 mmHg.
A paired test does not compare two means or two distributions as if they came from different individuals. It analyses the set of individual differences. This approach accounts for the fact that each patient serves as their own reference.
This is particularly important in medical research. Patients may differ in age, biological characteristics, disease severity, and baseline values. In a paired analysis, some of this between-person variability is eliminated because the analysis focuses on the change occurring within the same individual.
Test selection should not begin with a normality assessment. First, the structure of the data should be confirmed and the effect to be estimated should be defined. Only then should we determine whether the assumptions of a particular statistical method are sufficiently well met.
The paired t-test, also called the dependent-samples t-test, is used to determine whether the mean difference between two related measurements differs from zero.
A result that is clinically more informative than the p-value alone is the mean difference with its 95% confidence interval. The confidence interval indicates the direction and plausible magnitude of the effect, as well as the precision of the estimate.
One of the most common errors is to assess the normality of the baseline and follow-up measurements separately. This is not the appropriate approach because the paired t-test is based on individual differences.
The distribution that should be assessed is the distribution of the “after minus before” values, not the separate distributions of the baseline and follow-up measurements.
It is possible for both sets of measurements to have skewed distributions while their differences are approximately symmetrical and close to normal. The opposite situation is also possible: both measurements may appear normally distributed, while the distribution of differences is strongly skewed or contains a single extreme value.
I discussed the assessment of normality in greater detail in a separate article. In brief, the distribution of differences should be assessed using a combination of:
The Shapiro–Wilk test should not be used as the sole criterion. In a small sample, it may fail to detect a relevant departure from normality. In a large sample, it may identify a statistically significant but practically negligible deviation. A significant normality test does not automatically mean that the t-test must be rejected.
In larger samples, the t-test is generally relatively robust to moderate departures from normality. However, marked asymmetry, heavy distribution tails, and outliers may remain problematic, particularly when the number of pairs is small.
The paired t-test uses the mean and standard deviation of the differences. Both measures are sensitive to outliers. A single patient with an unusually large change may have a substantial influence on the mean effect, confidence interval, and p-value.
An outlier should not be removed merely because it makes a statistically significant result more difficult to obtain. First, determine whether:
If the value is correct and belongs to the studied population, it should generally remain in the analysis. A sensitivity analysis may also be performed by comparing the results obtained with and without the outlying observation. Any such procedure should be reported transparently.
The Wilcoxon signed-rank test is a non-parametric method for analysing two related measurements. Like the paired t-test, it begins by calculating the differences within individual pairs. It then uses the direction of these differences and the ranks of their absolute values.
In simplified terms, the method considers two pieces of information:
The Wilcoxon test does not require the differences to be normally distributed, but this does not mean that it has no assumptions. For the standard interpretation as a test of a location shift, the distribution of the differences should be approximately symmetrical. If the differences are strongly skewed, the hypothesis evaluated by the test becomes less intuitive.
The decision should not, however, be reduced to the rule: “The Shapiro–Wilk test is significant, therefore use the Wilcoxon test.” The shape of the distribution, sample size, outliers, and whether the mean change is of clinical interest should all be considered.
Scientific publications often state that the t-test compares means whereas the Wilcoxon test compares medians. This is an oversimplification.
The Wilcoxon signed-rank test uses the ranks of positive and negative differences. Under appropriate assumptions, particularly the symmetry of the distribution of differences, it can be interpreted as a test of a central shift in the distribution. It is not, however, a general test of equality between two medians.
Moreover, the median baseline result and the median follow-up result do not directly describe the median within-person change. These are different quantities. The two medians may be identical even when a substantial proportion of patients experienced individual changes.
If the aim is to estimate a typical change, the median of the individual differences may be reported. However, the p-value from the Wilcoxon signed-rank test is not always a direct test of the hypothesis that this median equals zero.
| Criterion | Paired t-test | Wilcoxon signed-rank test |
| Type of comparison | Two related measurements | Two related measurements |
| Basis of calculation | Values of the differences | Ranks and directions of the differences |
| Main effect | Mean difference | Location shift in the distribution of differences |
| Normality of differences | Assumed; particularly important in small samples | Not required |
| Symmetry of differences | Not a separate formal assumption | Important for the standard interpretation |
| Sensitivity to outliers | Higher | Usually lower |
| Preferred reporting | Mean difference and 95% CI | Description of differences and an effect estimate with a CI, if calculated |
If the distribution of differences is approximately symmetrical, contains no extreme observations, and the mean change has a meaningful clinical interpretation, the paired t-test is generally an appropriate choice. If the differences clearly depart from normality or contain influential outliers, the Wilcoxon signed-rank test may be considered.
A non-parametric method is not automatically a “safer” method. It addresses a somewhat different statistical question and may make it more difficult to express the effect in clinically meaningful units. The choice should therefore reflect the objective of the analysis rather than a fear of any departure from perfect normality.
Strong asymmetry in the differences may be problematic both for the paired t-test in a small sample and for the standard interpretation of the Wilcoxon signed-rank test. The available solutions depend on the nature of the data and the research question.
A transformation of the results may be considered. For example, a logarithmic transformation may be suitable for positive variables with a ratio-scale interpretation. Analysing differences between logarithms then corresponds to analysing relative or ratio changes rather than absolute differences. This interpretation may be appropriate for laboratory concentrations but not necessarily for every clinical variable.
Another option is the sign test, which uses only the direction of the changes and does not require the differences to be symmetrical. The price of these weaker assumptions is generally lower statistical power because the magnitude of the observed changes is ignored.
For more complex data, robust methods, bootstrap procedures, or appropriately specified statistical models may be used. A method should not be selected solely on the basis of whether it is labelled “parametric” or “non-parametric.”
A simple paired test requires both results to be available for every person included in the analysis. If a patient has a baseline measurement but no follow-up measurement, their individual change cannot be calculated.
If 100 patients were enrolled but only 82 have both measurements, a simple paired test will be performed using 82 complete pairs. A publication should clearly distinguish between:
The loss of incomplete pairs is not merely a sample-size problem. If a missing follow-up measurement is related to the patient’s condition, adverse events, or treatment failure, a complete-case analysis may produce biased results.
A missing value should not be replaced with the mean or by carrying the baseline measurement forward as the follow-up value. The approach to missing data should depend on the missing-data mechanism, study design, and prespecified analytical strategy.
Separate box plots for the baseline and follow-up measurements show the distributions of the results but conceal the most important feature of the study: the relationship between observations from the same patient.
For paired data, useful graphical presentations include:
A paired dot plot makes it possible to determine whether most patients changed in the same direction or whether the overall mean effect was driven primarily by a small number of large changes. This information cannot be obtained from two separate plots showing only summary valuesdzielnych wykresach prezentujących jedynie wartości zbiorcze.
A statistically significant result means that the observed data are not highly compatible with the null hypothesis under the assumed model and its assumptions. It does not indicate whether the observed change is clinically relevant.
In a large sample, a small difference may be statistically significant despite having little practical importance. In a small sample, a clinically meaningful change may fail to reach statistical significance because the estimate is imprecise.
If a minimal clinically important difference has been defined for the outcome, the estimated effect and its confidence interval should be compared with that threshold., warto porównać z nią oszacowany efekt i jego przedział ufności.
No. The correlation between baseline and follow-up results indicates whether patients with higher values at baseline also tend to have higher values at follow-up. It does not answer whether the outcome systematically increased or decreased after treatment.
A very high correlation may be observed even if the result decreases by exactly 10 units in every patient. The ranking of patients remains almost unchanged, so the correlation remains high despite a clear and consistent change.
Correlation should also not be used as the sole method for assessing agreement between two measurement methods. Two methods may be strongly correlated while producing systematically different results. Agreement requires separate methods, such as Bland–Altman analysis.
A paired test can determine whether a change from baseline occurred within one group. However, it is not sufficient to demonstrate that one treatment is more effective than another.
A common error is to perform a separate paired test in each treatment group and then conclude that treatment A was effective because its result was statistically significant, whereas treatment B was ineffective because its p-value exceeded 0.05.
The difference between a significant and a non-significant result is not, by itself, evidence of a significant difference between treatments. To compare the effectiveness of two treatments, the groups should be compared directly while accounting for the baseline value, for example by using an appropriately specified regression model or ANCOVA.
Depending on the study design, changes may also be compared between groups. However, in randomised trials with baseline measurements, an analysis of the follow-up outcome adjusted for the baseline value is often more efficient.[5]
The paired t-test and Wilcoxon signed-rank test apply to two related measurements. If patients were assessed at baseline, after one month, after three months, and after six months, a series of tests for all possible pairs should not be performed automatically.
Multiple comparisons increase the risk of false-positive results and do not make appropriate use of the complete data structure. Depending on the distribution, study design, and missing-data pattern, the following methods may be considered:
Mixed-effects models are particularly useful when the number or timing of measurements differs between patients or when some observations are incomplete. The analysis of more than two repeated measurements, however, requires a separate discussion.j niż dwóch powtarzanych pomiarów wymaga jednak osobnego omówienia.
The most common problems include:
The statistical methods section should provide more than the name of the test. It should also describe how its assumptions were assessed and identify the quantity being analysed.
Methods:
“The change in the analysed biomarker concentration between baseline and follow-up was assessed using a paired t-test. The normality of the distribution of within-patient differences was evaluated using a Q–Q plot, histogram, skewness and kurtosis values, and the Shapiro–Wilk test.”
Results:
“The biomarker concentration decreased by a mean of 6.20 units (mean difference: −6.20; 95% CI: −9.30 to −3.10; p < 0.001).”
The direction in which the difference was calculated should be defined in advance. If it was calculated as the baseline value minus the follow-up value, a positive result indicates a decrease. If the order was reversed, the interpretation of the sign will also be reversed.
Methods:
“Because of the markedly skewed distribution of within-patient differences and the presence of outliers, the change in the analysed parameter was assessed using the Wilcoxon signed-rank test.”
Results:
“The parameter decreased after treatment. The median individual change was −4 units (IQR: −8 to −1; Wilcoxon signed-rank test, p = 0.003).”
– A paired analysis evaluates the differences between two measurements obtained from the same person, rather than treating the measurements as two independent groups.
– The paired t-test evaluates the mean difference, whereas the Wilcoxon signed-rank test uses the direction and ranks of the differences and is not simply a test comparing two medians.
– Normality should be assessed for the within-pair differences, using not only the Shapiro–Wilk test but also graphical assessment, skewness, kurtosis, sample size, and the presence of outliers.
– The results should include the magnitude of change, a confidence interval, and the number of complete pairs, rather than only the p-value.
– Correlation between baseline and follow-up values does not replace an analysis of change, and separate within-group tests cannot establish that one treatment is more effective than another.
– If patients have more than two measurements or if some data are incomplete, simple paired tests may be insufficient, and repeated-measures methods or mixed-effects models should be considered.
Meta-analyses, systematic reviews and umbrella reviews all synthesise research evidence, but they serve different purposes. Learn…
Read more >What are paired categorical data? Categorical data describe the assignment of observations to specific categories. Examples…
Read more >How should two measurements from the same patients be analysed? Learn when to use the paired…
Read more >What is the difference between paired and independent data? Explore examples from medical research and learn…
Read more >Assessing the normality of data distributions is one of the first stages of statistical analysis in…
Read more >The Idea for This Blog Had Been Developing for Years The idea of creating this blog…
Read more >