Contents / Statistics / Nonparametric Methods
Chapter 9
Nonparametric Methods
Distribution-free goodness of fit, including the Kolmogorov-Smirnov statistic and Fisher's exact test; measuring association in tables with Cramer's V and McNemar's test for paired data; the sign, Wilcoxon signed-rank, Mann-Whitney and Kruskal-Wallis rank tests with their null moments derived; Spearman's rho and Kendall's tau; and permutation and bootstrap methods, with the efficiency cost of going distribution-free.
Introduction
Every test so far has assumed a shape. The -tests assumed normal populations, the -tests assumed normal errors with equal variances, and logistic regression assumed a Bernoulli response with a logit-linear mean. Those assumptions buy efficiency: when they hold, the resulting procedures extract about as much from the data as anything can.
When they fail the picture is mixed rather than catastrophic, and it is worth being precise about that, because "the data are not normal, so use a nonparametric test" is a reflex that is often wrong. The -test is robust to moderate non-normality because the central limit theorem is already working on the sample mean. What it is not robust to is heavy tails and outliers, where a single extreme observation can dominate and inflate — and that failure is not repaired by a larger sample, because a larger sample from a heavy-tailed population brings more extremes with it.
The methods in this chapter avoid the problem by throwing away information deliberately. Replace the observations by their ranks and an outlier at becomes simply the largest value, no more influential than one at . The null distribution of a rank statistic can then be computed by counting permutations, with no assumption about the population at all — which is where the name comes from, and why "distribution-free" is the more accurate one.
That freedom costs something, and the chapter quantifies it: against a genuinely normal population the Wilcoxon test needs about more data than the -test to achieve the same power. Against heavy tails it can need dramatically less. The last section covers the computational approach that has largely superseded both — permutation and bootstrap methods, which make almost no assumptions and require almost no distribution theory, only the willingness to resample.
The chi-square tests of Statistical Inference: Hypothesis Testing belong to this family too, being tests about counts that assume nothing about an underlying continuous distribution. They are derived there and used here, with the exact alternatives that chapter did not cover.
9.1Goodness of fit without a parametric model
The chi-square goodness-of-fit test asks whether observed counts match a hypothesised set of probabilities, and it is proved in the testing chapter. Its limitations set up everything else in this section.
Remark (What the chi-square test needs, and when it fails). By the theorem Chi-square goodness-of-fit test, is asymptotically . Two conditions are doing work.
The distribution is asymptotic, requiring expected counts of roughly or more; with small expected counts the statistic is discrete and skewed while is continuous, and the test rejects too often.
And the data must be categorical, or made so. Applying it to a continuous variable requires binning, and the result depends on the bins chosen — a real arbitrariness, since different binnings of the same data give different -values.
The second limitation is removed by comparing distribution functions directly rather than counts.
Definition 9.1 (Kolmogorov-Smirnov statistic). For a sample with empirical distribution function
and a hypothesised continuous distribution function , the Kolmogorov-Smirnov statistic is
Theorem 9.2 (The null distribution of the KS statistic is distribution-free). If the data really come from the continuous distribution , the distribution of does not depend on .
Proof. Apply the probability integral transform: if has continuous distribution function , then is Uniform. Since is non-decreasing, if and only if at continuity points, so the empirical distribution function of the at the point equals .
Writing for the empirical distribution function of the uniforms,
which involves only a uniform sample. No feature of survives.∎
That is the pattern the whole chapter repeats: transform the data into something whose distribution is fixed regardless of the population — here by the probability integral transform, later by ranking — and the null distribution becomes computable once and for all.
Example 9.4 (The KS statistic by hand). A sample of five values has when evaluated at the ordered observations. Compute .
Solution.
- The empirical distribution function jumps to at the -th ordered observation, so at each it takes the value just after the jump and just before.
- Both gaps must be checked at every point. The "after" gaps are
- The "before" gaps are , , , , .
- The largest of all ten is , so .
- Sanity check on why both sides are needed: at the fourth observation the "before" gap is far larger than the "after" gap , so checking only one side would have missed a near-maximum. The supremum of is always attained at a jump, from one side or the other.
- Against the critical value for , about , this is nowhere near significant — five observations can rule out very little.
Pitfall. The theorem assumes is fully specified in advance. Estimating parameters from the same data — testing normality with and plugged in — makes stochastically smaller, so the standard critical values are too large and the test rejects too rarely. The Lilliefors correction exists for exactly this case.
More generally, testing for normality before choosing a test is a two-stage procedure whose overall error rate nobody computes, and it is a poor use of the data besides: with small the normality test has no power to detect anything, and with large it rejects departures too small to matter for the -test that follows.
Definition 9.5 (Fisher's exact test). For a table with fixed margins, the probability of the observed table under independence is given by the hypergeometric distribution,
and the test sums this probability over all tables at least as extreme as the one observed.
Fisher's test is exact for any sample size, which is what recommends it when expected counts are small and the chi-square approximation cannot be trusted.
Example 9.6 (When the approximation cannot be used). A table has counts . Check whether the chi-square test is appropriate.
Solution.
- Margins: rows and , columns and , total .
- Expected counts under independence: , , and the same for the second row.
- Two of the four expected counts are , below the working threshold of , so the chi-square approximation is questionable — and with the statistic is very discrete.
- Use Fisher's exact test. The observed table's probability is
and summing over the more extreme tables ( and ) gives a one-sided of about . 5. Sanity check: the association is stark — of in one row against of in the other — so a very small is expected, and the exact computation confirms it without relying on an approximation the counts cannot support. 6. Note what not to do: pooling categories to force expected counts above changes the hypothesis being tested, and here there is nothing to pool.□
9.2Association in tables
The chi-square test of independence is proved in the testing chapter, with degrees of freedom. What it does not provide is a measure of how strong an association is, and with a large sample it will declare significance for associations of no practical size.
Definition 9.7 (Cramer's V). For an table with chi-square statistic and sample size ,
Proposition 9.8 (Cramer's V is bounded by one). , so , with under exact independence and for a perfect association.
Proof. We use the standard bound , which is attained when each row has all its mass in one column (or each column in one row). Dividing by and taking the square root gives . When the observed counts equal the expected ones exactly, and so is .∎
The important feature is the in the denominator. grows proportionally with the sample size for a fixed pattern of association, so it measures evidence; divides that out and measures strength. Reporting both separates "are we sure?" from "does it matter?", exactly as did beside the ANOVA .
Definition 9.9 (McNemar's test). For paired binary data — the same subjects measured twice — with subjects changing from to and changing from to , the test statistic is
the concordant pairs being ignored entirely.
Proof. Under the null that the two measurements have the same marginal distribution, a subject who changes is equally likely to change in either direction. So conditional on changers, is Binomial, with mean and variance . The standardised value is
and is the stated statistic, asymptotically .∎
That the concordant pairs drop out is the same idea as pairing in the -test: subjects who gave the same answer twice carry no information about a change, so conditioning on the changers removes a nuisance and keeps only the comparison.
Example 9.10 (Association: significance against strength). A table on subjects gives . Test independence at with , and assess the strength of the association.
Solution.
- Degrees of freedom: . Since , reject independence — the association is statistically significant, with .
- Strength: , so
- A Cramér's of is a very weak association — the variables are related, but knowing one tells you almost nothing about the other.
- Sanity check on the role of : the same pattern in a table of would give , nowhere near significance, with the identical . The statistic tracks the sample size; does not.
- The honest report is both numbers. "Significant association () " alone invites a reader to imagine a strong relationship that is not there.
9.3Rank tests
The central idea of the chapter: replace each observation by its rank, and the null distribution becomes a counting problem over permutations, free of any assumption about the population.
Definition 9.11 (Sign test). For paired data, let be the number of pairs with a positive difference, discarding ties. Under the null that the median difference is zero, .
Proof. If the median of the difference distribution is , then each difference is positive with probability independently of the others, so the count of positives is a sum of independent Bernoulli variables. No other feature of the distribution is used — not symmetry, not shape, not variance.∎
The sign test uses only the direction of each difference and is therefore valid under almost nothing. It is also wasteful: a difference of and one of count the same. The Wilcoxon test recovers the discarded magnitude information by ranking it.
Definition 9.12 (Wilcoxon signed-rank test). For paired data with differences , rank the from to and let be the sum of the ranks belonging to positive differences. Under the null that the differences are symmetric about , large or small is evidence against.
Theorem 9.13 (Null moments of the signed-rank statistic). Under the null,
Proof. Write , where indicates that the observation with absolute rank had a positive sign. Under symmetry about zero, each sign is independently or with probability , so the are i.i.d. Bernoulli — and, crucially, independent of the ranks, since symmetry makes sign and magnitude independent.
Hence, using ,
For the variance, independence makes the variance of the sum the sum of the variances, and :
Definition 9.14 (Mann-Whitney U test). For two independent samples of sizes , let
count the pairs in which the second sample's observation is larger. Equivalently, where is the sum of the ranks of the first sample in the combined ranking.
Theorem 9.15 (Null moments of U). Under the null that both samples come from the same continuous distribution,
Proof. Write , a sum of indicators. Under the null, any particular and are two draws from the same continuous distribution, so by symmetry (ties having probability zero). Linearity of expectation — which needs no independence, and the indicators are certainly dependent — gives
The variance requires accounting for the dependence between indicators sharing an index, and we quote the standard result.∎
Example 9.16 (A Mann-Whitney test). Two independent samples of sizes and give a rank sum of for the first sample in the combined ranking. Test at using the normal approximation.
Solution.
- Convert the rank sum to : by the definition, .
- Null moments from the theorem Null moments of U:
so the standard deviation is . 3. Standardise: . 4. Against , do not reject; the two-sided is about . 5. Sanity check on the range of : it must lie between and , and does ✓. The midpoint is the null expectation, and sits below it, indicating the first sample tends to be smaller — but not by enough to rule out chance with these sample sizes. 6. Note the conclusion's form: the evidence does not establish that the distributions are identical, only that it is insufficient to distinguish them — and by the pitfall above, a rejection would not have established a difference in medians either without a shift assumption.□
Pitfall. The Mann-Whitney test is often described as testing whether the medians are equal. That is not what it tests in general. Its null is that the two distributions are identical, and its alternative is a stochastic ordering — .
Two distributions can have equal medians and different shapes, and Mann-Whitney will reject; two can have different medians and the test may not reject. The median interpretation is valid only under the extra assumption that the distributions differ by a shift and are otherwise the same shape. State that assumption if you rely on it.
Definition 9.17 (Kruskal-Wallis test). For independent samples, rank all observations together and let be the rank sum of group . The statistic
is asymptotically under the null that all groups share a distribution.
Kruskal-Wallis is to one-way ANOVA what Mann-Whitney is to the two-sample -test — indeed for it reduces to Mann-Whitney, exactly as the -test reduced to the -test in the proposition With two groups, F is the square of t.
Example 9.18 (A Wilcoxon signed-rank test in full). Ten subjects give differences . Test that the differences are symmetric about zero, using the normal approximation at .
Solution.
- Rank the absolute values as through .
- The negative differences are , and , with ranks , and . So the negative rank sum is , and since all ranks sum to ,
- Null moments from the theorem Null moments of the signed-rank statistic with :
so the standard deviation is . 4. Standardise: . Against , do not reject at the level; the two-sided is about . 5. Sanity check against the sign test, which uses less information: of differences are positive, and for Binomial is , two-sided about — far from significant. The Wilcoxon result is much stronger because the positive differences are also the larger ones, which the sign test cannot see. 6. Note what the data suggest: seven positives, and the three negatives are among the four smallest in magnitude. The evidence points one way but ten observations cannot settle it.□
9.4Rank correlation
Pearson's measures linear association and is sensitive to outliers, both for the reason given in the regression chapter: it is built from products of deviations, so one extreme point dominates. Ranking the data first fixes both problems at once.
Definition 9.19 (Spearman's rank correlation). Rank the 's and the 's separately. Spearman's is Pearson's correlation computed on the two sets of ranks.
Theorem 9.20 (Computing formula for Spearman's rho). When there are no ties, with the difference between the ranks of observation ,
Proof. The ranks of each variable are a permutation of , so both have mean and the same sum of squared deviations,
a standard identity. Now expand with , inserting twice since the two rank means are equal:
so . Therefore
Definition 9.21 (Kendall's tau). For all pairs of observations, a pair is concordant if the two variables order it the same way and discordant otherwise. With and the counts,
Remark (Choosing between the two). Both are distribution-free and both detect any monotone relationship, not merely a linear one — an important gain over Pearson, since gives exactly while is well below .
has the more direct interpretation: it is the difference between the probability that a random pair is ordered the same way by both variables and the probability that it is not. It is also more robust to a few discrepant pairs, since each pair contributes regardless of how far apart the observations are, and its null distribution converges to normality faster at small .
is more familiar and easier to compute by hand. For most data they lead to the same conclusion, with typically smaller than for the same data — roughly for a bivariate normal.
Example 9.23 (One point, two correlations). Nine points lie close to a rising line and a tenth sits far below the trend. Pearson's falls to while Spearman's is . Explain the gap, and say which to report.
Solution.
- Pearson. , and the outlier has a large positive deviation and a large negative deviation. Their product is a large negative number that offsets much of the positive contribution from the other nine points.
- Its influence is quadratic in how far out it lies: doubling its distance from the means roughly doubles its contribution to while the other nine are unchanged.
- Spearman. On the rank scale the outlier is simply the observation with the largest and the smallest — ranks and . Its rank difference is , the largest possible, and it contributes to , but that is bounded no matter how extreme the raw value.
- So reports a strong-but-imperfect monotone association, which describes the ten points fairly, while reports almost no association, which describes none of them.
- Which to report? Neither, alone. The right response is to show the scatterplot and say there is a strong association among nine points and one observation that does not fit. Choosing the statistic that gives the preferred answer is the error; the statistics are flagging a real feature of the data that a single number cannot express.
- Sanity check on the general principle: rank methods bound the influence of any one observation, which is robustness, and that is a different property from being assumption-free.
Example 9.24 (Rank correlation against Pearson). Six observations give ranks and . Compute , and say what Pearson's on the original data would be if grew exponentially with .
Solution.
- Rank differences, taking pair by pair: , that is .
- Every , so .
- By the theorem Computing formula for Spearman's rho with , so :
- On the exponential question. If the underlying relationship were exactly, then is a strictly increasing function of , so the ranks match perfectly and exactly — rank correlation detects any monotone relationship, whatever its shape.
- Pearson's on the raw values would be well below , since the relationship is curved and measures only the linear part. For and , .
- Sanity check on the direction: here reflects a strong but imperfect monotone association — each adjacent pair is swapped, which is a small departure and is penalised accordingly. ✓
9.5Permutation and bootstrap methods
The rank tests replaced the data with ranks so that the null distribution could be computed once and tabulated. Modern computing removes the need for the tabulation, and with it the need to use ranks at all.
Definition 9.25 (Permutation test). To test whether two groups differ, compute the observed statistic — a difference in means, say. Then repeatedly reassign the group labels at random among the pooled observations, recomputing each time. The -value is the proportion of relabellings giving a at least as extreme as .
Theorem 9.26 (Permutation tests are exact). If the null hypothesis is that the two groups have identical distributions, then under the null every assignment of labels is equally likely, and the permutation test has significance level exactly for any sample size — with no asymptotic approximation and no assumption about the population.
Proof. Under the null the pooled observations are exchangeable: every one of the label assignments produces a data set with the same probability. So conditional on the observed multiset of values, is uniformly distributed over the values it takes across those assignments.
Rejecting when falls in the upper fraction of that finite set therefore has probability exactly , by the definition of a quantile of a uniform distribution on a finite set.∎
Two things about that proof deserve emphasis. Exchangeability is the only assumption, and it is implied directly by the null hypothesis rather than added to it. And the statistic is arbitrary — a difference of means, of medians, of trimmed means, of anything — because the argument never used its form. That freedom is what makes permutation tests so general: a statistic with no tractable distribution theory is no harder than one with.
Definition 9.27 (Bootstrap). To estimate the sampling distribution of a statistic , draw resamples of size with replacement from the observed data, compute on each, and use the spread of those values. The bootstrap standard error is the standard deviation of the , and a percentile interval takes the and quantiles.
Intuition. The bootstrap substitutes the sample for the population. The sampling distribution one wants describes how would vary across fresh samples from the population — which is unavailable. So take the empirical distribution of the data as a stand-in and draw fresh samples from that, which is exactly resampling with replacement.
It works because the empirical distribution converges to the true one, so the resampling distribution converges to the sampling distribution. Sampling with replacement is essential: without it, every resample would be the original data in a different order, and nothing would vary.
The striking consequence is that a standard error can be obtained for a statistic whose sampling distribution has never been derived — a median, a trimmed mean, a ratio of quantiles, the output of a fitting procedure. No formula is needed, only the ability to recompute the statistic.
Pitfall. The bootstrap resamples whatever unit was independent in the original sampling. For clustered data, resampling individual observations destroys the clustering and understates the variability — the clusters must be resampled instead. This is the pseudoreplication error of the design chapter, in resampling form.
It also fails for statistics that depend on extreme order statistics, such as the sample maximum. The resample's maximum can never exceed the observed maximum, so the bootstrap distribution is systematically wrong at the boundary — precisely the Uniform situation from the estimation chapter.
Example 9.28 (A bootstrap standard error). A sample of incomes has median . Bootstrap resampling with gives resample medians with standard deviation and and quantiles of and . Report the results and comment on the interval's shape.
Solution.
- Bootstrap standard error of the median: — obtained without any distribution theory for the sample median, which has no simple closed-form standard error.
- Percentile interval: .
- The interval is not symmetric about the estimate: it extends below and above. That asymmetry is informative and would be lost by reporting , which forces symmetry the data do not support.
- Sanity check on the source of the skew: incomes are right-skewed, so resamples that happen to include more of the upper tail move the median further up than unlucky resamples move it down. The percentile method inherits that shape automatically.
- A caution about the sample size: with the median takes only a handful of distinct values across resamples — it is always one of the nine observed incomes — so the bootstrap distribution is coarse and the quantiles are crude. The bootstrap does not manufacture information that nine observations lack.
- For a smoother result at this size, a bootstrap of a trimmed mean would behave better, since it varies continuously with the resample rather than jumping between order statistics.
Theorem 9.29 (Asymptotic relative efficiency of the Wilcoxon test). Against a normal population, the asymptotic relative efficiency of the Wilcoxon signed-rank test compared with the -test is
For heavy-tailed populations the efficiency exceeds , sometimes greatly.
Proof. We quote the value, whose derivation requires the asymptotic distribution theory of rank statistics. Its meaning is what matters: to achieve the same power, the Wilcoxon test needs about times as many observations as the -test when the data really are normal.∎
That number is the honest summary of the trade. Against the population most favourable to it, the -test's advantage over Wilcoxon is under — a remarkably small price for dropping the normality assumption entirely. Against a heavy-tailed population the comparison reverses dramatically: for a Cauchy-like distribution the -test is not merely less efficient but inconsistent, since the sample mean does not converge, while the rank test proceeds unaffected.
Example 9.30 (A permutation test by hand). Two groups give and . Use a permutation test on the difference in means, one-sided against .
Solution.
- Observed: , , so .
- There are ways to split the pooled values into two groups of two. Listing them with :
| group 1 | group 2 | |
|---|---|---|
- The observed is the largest of the six, so exactly one of the six permutations is at least as extreme.
- The one-sided -value is .
- Sanity check on the floor: with the smallest attainable -value is , so no result whatever could reach . The data separate perfectly and the test still cannot reject — a limitation of the sample size, not of the method, and one the exact computation makes unmistakable where an approximation would have hidden it.
- Note the symmetry of the table: the six values come in pairs, which is the exchangeability of the theorem Permutation tests are exact visible in six rows.
- Switching to a nonparametric test because a normality test rejected. That is a two-stage procedure with an uncomputed error rate; with large it flags departures too small to affect the -test, and with small it has no power.
- Describing Mann-Whitney as a test of medians. Its null is that the distributions are identical; the median reading needs the extra assumption of a pure shift.
- Assuming rank tests are assumption-free. The Wilcoxon signed-rank test assumes the differences are symmetric about the null value; without symmetry, use the sign test.
- Using chi-square with small expected counts. Fisher's exact test is exact at any size; pooling categories changes the hypothesis.
- **Reporting a chi-square -value without a measure of strength.** grows with for a fixed association; Cramér's does not.
- Testing goodness of fit with parameters estimated from the same data. The standard KS critical values are then too large and the test rejects too rarely.
- Bootstrapping individual observations from clustered data. Resample the independent unit, which is the cluster.
- Bootstrapping a statistic that depends on the extremes. A resample's maximum never exceeds the sample's, so the distribution is wrong where it matters.
- Believing rank tests cost a great deal of power. Against a normal population the Wilcoxon needs about more data; against heavy tails it needs far less.