Analysis of Paired Quantitative Data – Paired t-test or Wilcoxon test?

How should two measurements obtained from the same patients be compared, and when should the paired t-test or the Wilcoxon signed-rank test be used?

The analysis of paired quantitative data should focus on individual differences between measurements. The choice of method depends, among other factors, on the distribution of these differences, sample size, the presence of outliers, and the research question.

In this article, we explain how to analyse paired quantitative data, avoid the most common errors, and report the results appropriately in a scientific publication.

Measurements obtained before and after treatment in the same patients do not constitute two independent groups. Each baseline result is linked to a specific follow-up result, and this relationship should be taken into account in the analysis. When analysing paired quantitative data, we primarily assess within-pair differences rather than analysing two sets of measurements separately.

The choice between the paired t-test and the Wilcoxon signed-rank test should not be based solely on the result of the Shapiro–Wilk test. The measurement scale, distribution of the differences, presence of outliers, sample size, and research question should all be considered. It is equally important to report the magnitude and precision of the observed change, rather than only the p-value.

What are paired quantitative data?

Paired data arise when each observation can be assigned to a specific, related observation. Most commonly, they are two measurements obtained from the same person, for example before and after an intervention.

Examples of paired quantitative data include:

  • blood glucose concentration before and after treatment,
  • blood pressure at baseline and at a follow-up visit,
  • body weight before and after a dietary intervention,
  • pain intensity measured using a numerical scale before and after a procedure,
  • results obtained using two different measurement methods in the same patients,
  • the value of a parameter measured in the right and left eye of the same person.

In a study involving 60 patients assessed twice, we obtain 120 measurements but only 60 independent pairs. Observations from the same patient are generally more similar to each other than measurements from two randomly selected individuals. The statistical analysis should make use of this information.

A detailed explanation of the distinction between paired and independent data is available in the previous article. Here, we focus specifically on the analysis of two related measurements of a quantitative variable.

Within-pair change is the key element of paired analysis

Suppose we measure systolic blood pressure before treatment and after 12 weeks of therapy. For each patient, we can calculate an individual difference: the post-treatment value minus the baseline value. If blood pressure decreases from 145 to 135 mmHg, the difference is −10 mmHg. If it increases from 130 to 134 mmHg in another patient, the difference is +4 mmHg.

A paired test does not compare two means or two distributions as if they came from different individuals. It analyses the set of individual differences. This approach accounts for the fact that each patient serves as their own reference.

This is particularly important in medical research. Patients may differ in age, biological characteristics, disease severity, and baseline values. In a paired analysis, some of this between-person variability is eliminated because the analysis focuses on the change occurring within the same individual.

What should be checked before selecting a statistical test?

Before analysing paired quantitative data, determine:

  1. Can specific pairs of observations be identified?
  2. Is the analysed variable quantitative?
  3. Are both measurements available for every patient?
  4. What is the distribution of the differences between measurements?
  5. Are there any outliers among the differences?
  6. Are there only two measurements or more than two time points?

Test selection should not begin with a normality assessment. First, the structure of the data should be confirmed and the effect to be estimated should be defined. Only then should we determine whether the assumptions of a particular statistical method are sufficiently well met.

What question does the paired t-test answer?

The paired t-test, also called the dependent-samples t-test, is used to determine whether the mean difference between two related measurements differs from zero.

The paired t-test is appropriate when:

  • two related observations of a quantitative variable are analysed,
  • the individual pairs are independent of one another,
  • the differences are approximately normally distributed, or the sample is sufficiently large and the distribution does not contain major departures from normality or extreme outliers,
  • the mean change is the effect of interest.

A result that is clinically more informative than the p-value alone is the mean difference with its 95% confidence interval. The confidence interval indicates the direction and plausible magnitude of the effect, as well as the precision of the estimate.

Normality applies to the differences, not to the two measurements separately

One of the most common errors is to assess the normality of the baseline and follow-up measurements separately. This is not the appropriate approach because the paired t-test is based on individual differences.

The distribution that should be assessed is the distribution of the “after minus before” values, not the separate distributions of the baseline and follow-up measurements.

It is possible for both sets of measurements to have skewed distributions while their differences are approximately symmetrical and close to normal. The opposite situation is also possible: both measurements may appear normally distributed, while the distribution of differences is strongly skewed or contains a single extreme value.

I discussed the assessment of normality in greater detail in a separate article. In brief, the distribution of differences should be assessed using a combination of:

  • a histogram,
  • a Q–Q plot,
  • skewness and kurtosis,
  • the Shapiro–Wilk test,
  • identification of outliers,
  • the sample size.

The Shapiro–Wilk test should not be used as the sole criterion. In a small sample, it may fail to detect a relevant departure from normality. In a large sample, it may identify a statistically significant but practically negligible deviation. A significant normality test does not automatically mean that the t-test must be rejected.

In larger samples, the t-test is generally relatively robust to moderate departures from normality. However, marked asymmetry, heavy distribution tails, and outliers may remain problematic, particularly when the number of pairs is small.

Why are outliers particularly important?

The paired t-test uses the mean and standard deviation of the differences. Both measures are sensitive to outliers. A single patient with an unusually large change may have a substantial influence on the mean effect, confidence interval, and p-value.

An outlier should not be removed merely because it makes a statistically significant result more difficult to obtain. First, determine whether:

  • the value resulted from a data-entry or coding error,
  • the measurement was performed correctly,
  • a clinical event explains the unusual result,
  • the observation met the predefined criteria for inclusion in the analysis.

If the value is correct and belongs to the studied population, it should generally remain in the analysis. A sensitivity analysis may also be performed by comparing the results obtained with and without the outlying observation. Any such procedure should be reported transparently.

How does the Wilcoxon signed-rank test work?

The Wilcoxon signed-rank test is a non-parametric method for analysing two related measurements. Like the paired t-test, it begins by calculating the differences within individual pairs. It then uses the direction of these differences and the ranks of their absolute values.

In simplified terms, the method considers two pieces of information:

  • whether the result increased or decreased in a given patient,
  • how large that change was relative to the changes observed in the other patients.

The Wilcoxon test does not require the differences to be normally distributed, but this does not mean that it has no assumptions. For the standard interpretation as a test of a location shift, the distribution of the differences should be approximately symmetrical. If the differences are strongly skewed, the hypothesis evaluated by the test becomes less intuitive.

The Wilcoxon signed-rank test may be appropriate when:

  • the distribution of differences clearly departs from normality,
  • outliers have a strong influence on the mean,
  • the sample is small and an assumption of normality would be difficult to justify,
  • the data allow the magnitude of changes to be ranked reliably, but the mean difference is not the most appropriate description of the effect.

The decision should not, however, be reduced to the rule: “The Shapiro–Wilk test is significant, therefore use the Wilcoxon test.” The shape of the distribution, sample size, outliers, and whether the mean change is of clinical interest should all be considered.

The Wilcoxon test is not simply a test of medians

Scientific publications often state that the t-test compares means whereas the Wilcoxon test compares medians. This is an oversimplification.

The Wilcoxon signed-rank test uses the ranks of positive and negative differences. Under appropriate assumptions, particularly the symmetry of the distribution of differences, it can be interpreted as a test of a central shift in the distribution. It is not, however, a general test of equality between two medians.

Moreover, the median baseline result and the median follow-up result do not directly describe the median within-person change. These are different quantities. The two medians may be identical even when a substantial proportion of patients experienced individual changes.

If the aim is to estimate a typical change, the median of the individual differences may be reported. However, the p-value from the Wilcoxon signed-rank test is not always a direct test of the hypothesis that this median equals zero.

Paired t-test or Wilcoxon signed-rank test?

CriterionPaired t-testWilcoxon signed-rank test
Type of comparisonTwo related measurementsTwo related measurements
Basis of calculationValues of the differencesRanks and directions of the differences
Main effectMean differenceLocation shift in the distribution of differences
Normality of differencesAssumed; particularly important in small samplesNot required
Symmetry of differencesNot a separate formal assumptionImportant for the standard interpretation
Sensitivity to outliersHigherUsually lower
Preferred reportingMean difference and 95% CIDescription of differences and an effect estimate with a CI, if calculated

If the distribution of differences is approximately symmetrical, contains no extreme observations, and the mean change has a meaningful clinical interpretation, the paired t-test is generally an appropriate choice. If the differences clearly depart from normality or contain influential outliers, the Wilcoxon signed-rank test may be considered.

A non-parametric method is not automatically a “safer” method. It addresses a somewhat different statistical question and may make it more difficult to express the effect in clinically meaningful units. The choice should therefore reflect the objective of the analysis rather than a fear of any departure from perfect normality.

What should be done when the differences are strongly skewed?

Strong asymmetry in the differences may be problematic both for the paired t-test in a small sample and for the standard interpretation of the Wilcoxon signed-rank test. The available solutions depend on the nature of the data and the research question.

A transformation of the results may be considered. For example, a logarithmic transformation may be suitable for positive variables with a ratio-scale interpretation. Analysing differences between logarithms then corresponds to analysing relative or ratio changes rather than absolute differences. This interpretation may be appropriate for laboratory concentrations but not necessarily for every clinical variable.

Another option is the sign test, which uses only the direction of the changes and does not require the differences to be symmetrical. The price of these weaker assumptions is generally lower statistical power because the magnitude of the observed changes is ignored.

For more complex data, robust methods, bootstrap procedures, or appropriately specified statistical models may be used. A method should not be selected solely on the basis of whether it is labelled “parametric” or “non-parametric.”

Missing data and incomplete pairs

A simple paired test requires both results to be available for every person included in the analysis. If a patient has a baseline measurement but no follow-up measurement, their individual change cannot be calculated.

If 100 patients were enrolled but only 82 have both measurements, a simple paired test will be performed using 82 complete pairs. A publication should clearly distinguish between:

  • the number of participants enrolled,
  • the number of participants with each measurement,
  • the number of complete pairs included in the analysis.

The loss of incomplete pairs is not merely a sample-size problem. If a missing follow-up measurement is related to the patient’s condition, adverse events, or treatment failure, a complete-case analysis may produce biased results.

A missing value should not be replaced with the mean or by carrying the baseline measurement forward as the follow-up value. The approach to missing data should depend on the missing-data mechanism, study design, and prespecified analytical strategy.

How should paired quantitative data be presented graphically?

Separate box plots for the baseline and follow-up measurements show the distributions of the results but conceal the most important feature of the study: the relationship between observations from the same patient.

For paired data, useful graphical presentations include:

  • a paired dot plot in which the two measurements from each patient are connected,
  • a plot of individual changes,
  • a histogram or Q–Q plot of the differences,
  • a plot presenting the differences and their distribution,
  • a plot of the mean change with its 95% confidence interval.

A paired dot plot makes it possible to determine whether most patients changed in the same direction or whether the overall mean effect was driven primarily by a small number of large changes. This information cannot be obtained from two separate plots showing only summary valuesdzielnych wykresach prezentujących jedynie wartości zbiorcze.

A p-value does not indicate the magnitude of change

A statistically significant result means that the observed data are not highly compatible with the null hypothesis under the assumed model and its assumptions. It does not indicate whether the observed change is clinically relevant.

In a large sample, a small difference may be statistically significant despite having little practical importance. In a small sample, a clinically meaningful change may fail to reach statistical significance because the estimate is imprecise.

The reported result should therefore include:

  • the number of complete pairs,
  • descriptive statistics for both measurements,
  • the magnitude and direction of the change,
  • a confidence interval,
  • the p-value,
  • an interpretation of the clinical relevance.

If a minimal clinically important difference has been defined for the outcome, the estimated effect and its confidence interval should be compared with that threshold., warto porównać z nią oszacowany efekt i jego przedział ufności.

Can correlation replace an analysis of change?

No. The correlation between baseline and follow-up results indicates whether patients with higher values at baseline also tend to have higher values at follow-up. It does not answer whether the outcome systematically increased or decreased after treatment.

A very high correlation may be observed even if the result decreases by exactly 10 units in every patient. The ranking of patients remains almost unchanged, so the correlation remains high despite a clear and consistent change.

Correlation should also not be used as the sole method for assessing agreement between two measurement methods. Two methods may be strongly correlated while producing systematically different results. Agreement requires separate methods, such as Bland–Altman analysis.

A one-group pre–post study versus a comparison of two treatments

A paired test can determine whether a change from baseline occurred within one group. However, it is not sufficient to demonstrate that one treatment is more effective than another.

A common error is to perform a separate paired test in each treatment group and then conclude that treatment A was effective because its result was statistically significant, whereas treatment B was ineffective because its p-value exceeded 0.05.

The difference between a significant and a non-significant result is not, by itself, evidence of a significant difference between treatments. To compare the effectiveness of two treatments, the groups should be compared directly while accounting for the baseline value, for example by using an appropriately specified regression model or ANCOVA.

Depending on the study design, changes may also be compared between groups. However, in randomised trials with baseline measurements, an analysis of the follow-up outcome adjusted for the baseline value is often more efficient.[5]

What should be done when there are more than two measurements?

The paired t-test and Wilcoxon signed-rank test apply to two related measurements. If patients were assessed at baseline, after one month, after three months, and after six months, a series of tests for all possible pairs should not be performed automatically.

Multiple comparisons increase the risk of false-positive results and do not make appropriate use of the complete data structure. Depending on the distribution, study design, and missing-data pattern, the following methods may be considered:

  • repeated-measures ANOVA,
  • the Friedman test,
  • mixed-effects models.

Mixed-effects models are particularly useful when the number or timing of measurements differs between patients or when some observations are incomplete. The analysis of more than two repeated measurements, however, requires a separate discussion.j niż dwóch powtarzanych pomiarów wymaga jednak osobnego omówienia.

Common errors in the analysis of paired data

The most common problems include:

  1. Using a test for two independent groups to compare baseline and follow-up results obtained from the same individuals.
  2. Assessing normality separately for the two measurements instead of assessing the distribution of their differences.
  3. Automatically selecting the Wilcoxon signed-rank test solely because the Shapiro–Wilk test is statistically significant.
  4. Reporting only the p-value without the magnitude of change or a confidence interval.
  5. Ignoring incomplete pairs and failing to report the actual number of observations included in the analysis.
  6. Replacing an analysis of change with a correlation coefficient.
  7. Inferring a difference between treatments from tests performed separately in each treatment group.
  8. Performing multiple tests for consecutive time points without accounting for multiple comparisons.

How should a paired-data analysis be described in a publication?

The statistical methods section should provide more than the name of the test. It should also describe how its assumptions were assessed and identify the quantity being analysed.

Example wording for the paired t-test

Methods:
“The change in the analysed biomarker concentration between baseline and follow-up was assessed using a paired t-test. The normality of the distribution of within-patient differences was evaluated using a Q–Q plot, histogram, skewness and kurtosis values, and the Shapiro–Wilk test.”

Results:
“The biomarker concentration decreased by a mean of 6.20 units (mean difference: −6.20; 95% CI: −9.30 to −3.10; p < 0.001).”

The direction in which the difference was calculated should be defined in advance. If it was calculated as the baseline value minus the follow-up value, a positive result indicates a decrease. If the order was reversed, the interpretation of the sign will also be reversed.

Example wording for the Wilcoxon signed-rank test

Methods:
“Because of the markedly skewed distribution of within-patient differences and the presence of outliers, the change in the analysed parameter was assessed using the Wilcoxon signed-rank test.”

Results:
“The parameter decreased after treatment. The median individual change was −4 units (IQR: −8 to −1; Wilcoxon signed-rank test, p = 0.003).”

Practical takeaways for clinicians

– A paired analysis evaluates the differences between two measurements obtained from the same person, rather than treating the measurements as two independent groups.

– The paired t-test evaluates the mean difference, whereas the Wilcoxon signed-rank test uses the direction and ranks of the differences and is not simply a test comparing two medians.

– Normality should be assessed for the within-pair differences, using not only the Shapiro–Wilk test but also graphical assessment, skewness, kurtosis, sample size, and the presence of outliers.

– The results should include the magnitude of change, a confidence interval, and the number of complete pairs, rather than only the p-value.

– Correlation between baseline and follow-up values does not replace an analysis of change, and separate within-group tests cannot establish that one treatment is more effective than another.

– If patients have more than two measurements or if some data are incomplete, simple paired tests may be insufficient, and repeated-measures methods or mixed-effects models should be considered.

References:

  1. Franc JM. Analysis of Paired Data. Prehosp Disaster Med. 2025;40(2):61-63. doi:10.1017/S1049023X25001517.
  2. Bland JM, Altman DG. Correlation, regression, and repeated data. BMJ. 1994;308:896.
  3. Rosner B, Glynn RJ, Lee MLT. The Wilcoxon signed rank test for paired comparisons of clustered data. Biometrics. 2006;62:185–192.
  4. Mishra P, Pandey CM, Singh U, Keshri A, Sabaretnam M. Selection of appropriate statistical methods for data analysis. Ann Card Anaesth. 2019;22:297–301.
  5. Vickers AJ, Altman DG. Analysing controlled trials with baseline and follow up measurements. BMJ. 2001;323:1123–1124.
  6. Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991.
  7. Kirkwood BR, Sterne JAC. Essential Medical Statistics. 2nd ed. Oxford: Blackwell Science; 2003.

Contact us

If you have any questions,
contact us.
We will respond within 24 hours.

+48 601 40 77 71

Show the e-mail address