Wilcoxon Signed Rank Test in R: A Practical Guide for Paired Non‑Parametric Data
The Wilcoxon signed rank test is a go‑to method when you need to compare two related samples without assuming normality. In R, implementing this test is straightforward, and the language offers several packages that make the workflow tidy, reproducible, and visually informative. Below you’ll find a step‑by‑step walkthrough, the statistical reasoning behind the test, common pitfalls, and a FAQ section to solidify your understanding Less friction, more output..
Why Choose the Wilcoxon Signed Rank Test?
When you have paired observations—such as measurements taken before and after an intervention on the same subjects—the classic paired t‑test relies on the assumption that the differences are normally distributed. If that assumption is violated (e.g.That's why , skewed data, outliers, or small sample sizes), the Wilcoxon signed rank test provides a solid alternative. It evaluates whether the median of the paired differences differs from zero by ranking the absolute differences and considering their signs Took long enough..
Key points to remember:
- Non‑parametric: No distributional assumptions beyond symmetry of the difference distribution.
- Paired design: Each observation in one group is uniquely linked to an observation in the other.
- Interpretation: Tests the null hypothesis H₀: median difference = 0 against the alternative H₁: median difference ≠ 0 (two‑sided) or a directional version.
Preparing Your Data in R
Before running the test, ensure your data is in a tidy format: one row per subject, with two columns representing the paired measurements. For illustration, we’ll use a built‑in dataset (sleep) that records extra sleep gained by patients under two drugs.
# Load necessary packages
library(tidyverse) # for data manipulation
library(rstatix) # provides a tidy wrapper for wilcox_test
library(ggpubr) # for easy publication‑ready plots
# Inspect the data
head(sleep)
Output:
extra group ID
1 0.7 1 1
2 -1.6 1 2
3 -0.2 1 3
4 -1.2 1 4
5 -0.1 1 5
6 3.4 1 6
The extra column shows the change in sleep hours, while group indicates the drug (1 or 2). To apply the signed rank test we need to reshape the data so each subject has two measurements side‑by‑side Which is the point..
# Widen the data: one row per ID, columns for each drug
sleep_wide <- sleep %>%
pivot_wider(names_from = group, values_from = extra) %>%
rename(drug1 = `1`, drug2 = `2`)
head(sleep_wide)
Output:
# A tibble: 6 × 3
ID drug1 drug2
1 1 0.7 1.9
2 2 -1.6 0.8
3 3 -0.2 1.1
4 4 -1.2 0.1
5 5 -0.1 -0.1
6 6 3.4 4.4
Now each row contains the paired observations for a single participant.
Performing the Wilcoxon Signed Rank Test
Base R Approach
The simplest way is to compute the differences and feed them to wilcox.test() with paired = TRUE Worth keeping that in mind..
# Compute differences
diff <- sleep_wide$drug2 - sleep_wide$drug1
# Wilcoxon signed rank test (two‑sided)
wt_base <- wilcox.test(diff, paired = TRUE, exact = FALSE, correct = FALSE)
wt_base
Typical output:
Wilcoxon signed rank test with continuity correction
data: diff
V = 30, p-value = 0.009
alternative hypothesis: true location shift is not equal to 0
- V is the test statistic (sum of ranks for positive differences).
- The p‑value indicates whether we reject H₀ at a chosen α (commonly 0.05).
- Setting
exact = FALSEuses a normal approximation, which is fine for sample sizes > 20; for smaller samples you may keep the exact calculation.
Tidy Approach with rstatix
If you prefer a tidy workflow that returns a data frame, rstatix::wilcox_test() is handy.
wt_tidy <- sleep_wide %>%
wilcox_test(drug2 ~ drug1, paired = TRUE) %>%
add_significance() # adds stars for p‑value thresholds
wt_tidy
Output:
# A tibble: 1 × 8
.y. group1 group2 n1 n2 statistic p p.adj p.signif
1 drug2 drug1 drug2 10 10 30 0.009 0.009 **
The table includes the statistic, p‑value, and a significance column (** for p < 0.01) Turns out it matters..
One‑Sided Tests
If you have a directional hypothesis (e.g., drug 2 yields more sleep than drug 1), specify alternative = "greater" or "less" Worth keeping that in mind..
wt_greater <- sleep_wide %>%
wilcox_test(drug2 ~ drug1, paired = TRUE, alternative = "greater")
wt_greater
Effect Size and Confidence Intervals
Statistical significance alone doesn’t convey the magnitude of the effect. A common non‑parametric effect size for the Wilcoxon signed rank test is r, calculated as:
[ r = \frac{Z}{\sqrt{N}} ]
where Z is the standardized test statistic and N is the number of pairs Less friction, more output..
# Compute Z from the test statistic (using normal approximation)
z_val <- qnorm(wt_base$p.value/2, lower.tail = FALSE) * sign(wt_base$statistic - (n*(n+1)/4))
r_val <- z_val / sqrt(length(diff))
r_val
Alternatively, rstatix provides an effectsize argument:
wt_eff <- sleep_wide %>%
wilcox_test(drug2 ~ drug1, paired = TRUE) %>%
add_effect_size()
wt_eff
Output includes effsize (the r value) and a 95% confidence interval for the median difference via wilcox., conf.test(...int = TRUE) The details matter here..
wt_ci <- wilcox.test(diff, paired = TRUE, conf.int = TRUE)
wt_ci$conf.int