Assessing Data Normality: How to Do It Correctly

Assessing data normality is one of the first steps in biostatistical data analysis. However, it is frequently reduced to performing the Shapiro–Wilk test. Depending on the resulting p-value, researchers then select either a parametric or a non-parametric test.

This approach is simple, but it is not always correct.

Assessing the normality of data distributions is one of the first stages of statistical analysis in medical research. In many publications, however, the entire procedure is described in a single sentence: “Normality was assessed using the Shapiro–Wilk test.” Depending on the resulting p-value, the authors then automatically choose either a parametric or a non-parametric test.

Although straightforward, this approach is not always appropriate. Normality assessment should not be based solely on the result of a single statistical test. It should also take into account histograms, Q–Q plots, skewness, kurtosis, outliers, sample size and the assumptions of the statistical method that is being considered.

The Shapiro–Wilk test is a useful diagnostic tool, but it should not be treated as an automatic switch between Student’s t-test and the Mann–Whitney U test.

In medical statistics, the most important question is not: “Are the data perfectly normally distributed?” Instead, we should ask: “Could the characteristics of the data adversely affect the statistical method we intend to use?”

What Is a Normal Distribution?

A variable is said to follow a normal distribution, also known as a Gaussian distribution, when most observations are concentrated around the average value and very low or very high values occur less frequently. When presented graphically, such a distribution resembles a symmetrical bell-shaped curve.

In a perfectly normal distribution, the observations are symmetrically distributed on both sides of the centre. The mean, median and mode are equal. Most observations are located close to the mean: approximately 68% fall within one standard deviation of the mean, while approximately 95% fall within two standard deviations.

In real-world medical data, a perfect bell-shaped distribution is uncommon. Minor deviations from this shape occur frequently and do not necessarily create a problem for subsequent statistical analysis.

Why Do We Assess Data Normality?

Assessing the distribution of the data helps determine whether the selected descriptive measures and planned statistical analyses are appropriate for the actual structure of the dataset.

In medical statistics, this is particularly important when choosing between:

  • the mean and the median;
  • parametric and non-parametric tests;
  • conventional statistical methods and methods that are more robust to outliers.

Normality assessment may also help determine whether the assumptions of a planned statistical test or model are satisfied, whether the result may be disproportionately influenced by individual extreme observations and whether a data transformation should be considered.

The objective is therefore not simply to classify a distribution as “normal” or “non-normal”. The real purpose is to establish whether the characteristics of the data could undermine the reliability of the planned statistical analysis.

A non-normal distribution should not automatically be interpreted as an indication that a non-parametric test is required. A more important consideration is whether the type and magnitude of the deviation could materially affect the performance of the selected statistical method.

How Should Data Normality Be Assessed?

Good statistical practice involves combining several diagnostic methods rather than mechanically interpreting a single p-value.

A practical assessment of medical data should include the following steps:

  • checking the dataset for errors, missing values and clinically impossible observations;
  • examining histograms and Q–Q plots;
  • evaluating skewness, kurtosis and outliers;
  • performing a formal normality test, such as the Shapiro–Wilk test;
  • considering the sample size and the robustness of the planned statistical method;
  • identifying exactly which element of the analysis is expected to be normally distributed: raw observations, differences between paired measurements or model residuals.

Normality should therefore be treated as a diagnostic issue rather than as a single formal test that produces a simple “yes” or “no” decision.

Checking the Data Before Statistical Analysis

Before assessing normality, researchers must ensure that the dataset itself is correct. A single data-entry error may substantially alter the appearance of a histogram, the result of a normality test and the conclusions drawn from the analysis.

A simple example is a recorded value of 450 instead of 45. Such an error may be responsible for the apparent asymmetry of the entire distribution. Similarly, a recorded BMI of 285 is more likely to represent 28.5 than a genuine clinical observation.

The dataset should therefore be checked for:

  • data-entry errors;
  • inconsistent measurement units;
  • clinically impossible values;
  • missing values coded as 0, 99 or 999;
  • duplicate records.

Histogram

A histogram is a simple bar chart showing how many observations fall within consecutive ranges of values. It allows researchers to quickly identify whether the data are approximately symmetrical, skewed or multimodal, meaning that the distribution has more than one prominent peak.

A histogram can also help identify extreme observations.

An example of a BMI histogram is presented below.

In practice, a histogram provides an initial assessment of:

  • the symmetry or skewness of the distribution;
  • the number of peaks and the possible presence of several clusters of observations;
  • the length of the distribution tails;
  • extreme values and potential data errors;
  • the presence of clinically relevant subgroups.

Although histograms are extremely useful, they have an important limitation: their appearance depends on the number and width of the selected intervals, commonly referred to as bins.

The same dataset may look different after changing the histogram settings. Bins that are too wide may conceal irregularities, whereas bins that are too narrow may overemphasise random fluctuations.

The histogram below presents the same BMI dataset as the previous figure, but with a different number and width of bins.

Histograms may also be unstable and difficult to interpret in small samples. With only 10 or 15 observations, researchers should not expect to see a perfectly smooth bell-shaped curve, even when the observations genuinely come from a normally distributed population.

Q–Q Plot

A Q–Q plot, or quantile–quantile plot, compares the observed data with the values expected under a normal distribution.

In simplified terms, when the plotted points follow an approximately diagonal straight line, the distribution may be considered reasonably close to normal.

Characteristic deviations from the reference line may indicate:

  • skewness, when the points systematically depart from the line in one direction;
  • heavy tails, indicating more extreme observations than would be expected under a normal distribution;
  • light tails, indicating fewer extreme observations;
  • outliers, when individual points are clearly separated from the remaining observations;
  • a mixture of several distributions, for example when the dataset contains distinct subgroups of patients.

A single deviation at the edge of a Q–Q plot does not necessarily justify rejecting parametric methods. Its relevance depends on the magnitude of the deviation, the sample size and whether the unusual observation represents a data error or a valid clinical result.

Skewness and Kurtosis

Skewness is a descriptive statistic that indicates whether a distribution is symmetrical or whether it has a longer tail on one side.

Positive, or right, skewness means that most observations are concentrated at lower values, while a small number of high observations extend the right tail of the distribution. Negative, or left, skewness describes the opposite pattern.

Right-skewed distributions are very common in medical data. Examples include:

  • length of hospital stay;
  • treatment costs;
  • concentrations of biological markers;
  • the number of complications;
  • time-to-event variables.

In such cases, most patients have low or moderate values, while a small group or individual patients have very high values.

Kurtosis describes certain aspects of the shape of a distribution, particularly the concentration of observations and the behaviour of its tails. High kurtosis may indicate heavy tails and an increased probability of extreme observations.

There are no universal skewness or kurtosis thresholds that can definitively determine whether a parametric analysis is acceptable in every study.

Frequently cited thresholds, such as ±1 or ±2, should be treated as general guidelines rather than absolute decision criteria.

The Shapiro–Wilk Test

The Shapiro–Wilk test evaluates the null hypothesis that the observed data come from a normally distributed population.

It is a useful statistical test, but its result must be interpreted carefully.

The main principles of interpretation are as follows:

  • p < 0.05 indicates that a statistically significant deviation from normality has been detected;
  • p ≥ 0.05 does not prove that the distribution is normal; it only means that there is insufficient evidence to reject the assumption of normality;
  • in small samples, the test may fail to detect genuine deviations from normality;
  • in large samples, the test may detect very small deviations that have little or no practical importance;
  • the result should be interpreted together with the histogram, Q–Q plot, skewness, kurtosis and the characteristics of the data.

Consequently, as the sample size increases, it becomes progressively less reasonable to select a statistical method solely on the basis of the Shapiro–Wilk p-value.

This issue is discussed in more detail in the article: Sample Size and Data Normality: Can the Shapiro–Wilk Test Be Misleading?

Does the Analysed Variable Have to Be Normally Distributed?

This is one of the most common misunderstandings in medical data analysis. The answer is: not always. It depends on the statistical method being used.

For example, in a paired Student’s t-test, the analysis concerns the differences between two paired measurements. It is therefore the distribution of these differences—not the separate distributions of the “before” and “after” measurements—that should be assessed for normality.

In linear regression, the normality assumption primarily concerns the random error component, which is usually evaluated by analysing the model residuals. The independent variables themselves are not required to follow a normal distribution.

Moreover, not all statistical methods require normality assessment.

Logistic regression is one example. Its dependent variable is usually binary and therefore, by definition, cannot follow a normal distribution.

Before performing the Shapiro–Wilk test, researchers should first establish which variable or component of the model is actually expected to follow a normal distribution and therefore requires assessment.

Does p < 0.05 in the Shapiro–Wilk Test Mean That a Non-Parametric Test Is Required?

No. A Shapiro–Wilk result of p < 0.05 indicates that a deviation from a perfectly normal distribution has been detected. However, it does not explain:

  • how large the deviation is;
  • whether it is practically important;
  • whether it results from skewness, kurtosis or a single outlier;
  • whether the planned parametric method is sensitive to this particular deviation.

The result also does not indicate whether a non-parametric method answers the same research question.

The Mann–Whitney U test is not a simple non-parametric substitute for Student’s t-test. Student’s t-test evaluates differences between means, whereas the Mann–Whitney U test compares distributions or ranks.

Interpreting the Mann–Whitney U test as a test of differences between medians requires additional assumptions, including distributions of a similar shape in the compared groups.

The choice of statistical test should therefore be based on:

  • the research question;
  • the type of data;
  • the sample size;
  • the presence of outliers;
  • the nature and magnitude of deviations from normality;
  • the robustness of the planned statistical method.

It should not be determined solely by whether the Shapiro–Wilk p-value falls above or below 0.05.

How Should Normality Assessment Be Reported in a Scientific Paper?

The statement “Normality was assessed using the Shapiro–Wilk test” is technically correct but usually incomplete.

It does not indicate whether the authors examined diagnostic plots, evaluated outliers, considered skewness or took the sample size into account.

A more informative description could be:

The distributions of continuous variables were evaluated using histograms and Q–Q plots, together with skewness and kurtosis statistics. The Shapiro–Wilk test was used as a supplementary formal method for assessing deviations from normality. Group sizes, the presence of outliers and the robustness of the planned statistical tests to violations of the normality assumption were also considered when selecting the statistical methods.

For linear regression, it is useful to specify that the residuals of the model were evaluated rather than all variables separately.

For example: The normality assumption was evaluated by examining the distribution of the model residuals and the Q–Q plot of the residuals.

Common Mistakes in Data Normality Assessment

One of the most common errors is treating p ≥ 0.05 in the Shapiro–Wilk test as proof that a distribution is normal.

Other frequent problems include:

  • automatically selecting a parametric or non-parametric test solely on the basis of a normality test;
  • assessing the distribution of the wrong element of the analysis, such as examining raw variables instead of model residuals in linear regression;
  • performing separate normality tests for dozens of variables and mechanically deciding between the mean and median or between Student’s t-test and the Mann–Whitney U test for each variable.

Although this procedure may appear objective, it does not adequately account for the properties of the statistical methods, the sample size, the research objective or the clinical relevance of the observed deviations.

Key Conclusions for Medical Statistics

Key Conclusions for Medical Statistics

– The Shapiro–Wilk test is a supplementary diagnostic tool, not an automatic criterion for selecting a statistical method.
– A result of p ≥ 0.05 does not prove normality, while p < 0.05 does not automatically indicate that a non-parametric test is required.
– In small samples, a non-significant result may be caused by low statistical power. In large samples, a significant result may reflect a minor and practically irrelevant deviation.
– The normality assumption may concern differences between paired measurements or model residuals rather than the raw values of a variable.
– The choice of statistical method should account for the research question, sample size, outliers, skewness, kurtosis and the robustness of the planned analysis.

Professional statistical analysis of medical data involves understanding the structure of the dataset and selecting a method that answers the appropriate research question while providing reliable results.

When preparing a statistical analysis for a scientific paper, publication or doctoral dissertation, it is advisable to plan the assessment of data distributions before conducting the main analyses. This facilitates the selection of appropriate statistical methods, improves the reporting of results and helps researchers respond to potential comments from reviewers.

Need support with your research? Learn more about our professional medical statistics services and discover how we can help you plan, analyse and report your study.

References

  1. Altman DG, Bland JM. Statistics notes: the normal distribution. BMJ. 1995;310:298.
  2. Ghasemi A, Zahediasl S. Normality tests for statistical analysis: a guide for non-statisticians. International Journal of Endocrinology and Metabolism. 2012;10(2):486–489.
  3. Henderson AR. Testing experimental data for univariate normality. Clinical Chimica Acta. 2006;366(1–2):112–129.
  4. Kim HY. Statistical notes for clinical researchers: assessing normal distribution using skewness and kurtosis. Restorative Dentistry & Endodontics. 2013;38(1):52–54.
  5. Rochon J, Gondan M, Kieser M. To test or not to test: preliminary assessment of normality when comparing two independent samples. BMC Med Res Methodol. 2012;12:81.
  6. Shatz I. Assumption-checking rather than assumption-testing: a terminology and practical framework for improving statistical practice. Biom J. 2024;66(1):e2200356.
  7. Lang TA, Altman DG. Basic statistical reporting for articles published in biomedical journals: the SAMPL Guidelines. International Journal of Nursing Studies. 2015;52(1):5–9.

Contact us

If you have any questions,
contact us.
We will respond within 24 hours.

+48 601 40 77 71

Show the e-mail address