Contents / Statistics / Analysis of Variance
Chapter 5
Analysis of Variance
The sum-of-squares decomposition and what each mean square estimates; the F-test and its equivalence to the two-sample t-test; which assumptions matter and why; Tukey, Bonferroni, contrasts and Scheffe for multiple comparisons; two-way designs and interaction; fixed versus random effects and variance components; and ANOVA as a linear model with indicator predictors.
Introduction
Comparing two means is a -test. Comparing several is not simply more -tests, and the reason is the multiplicity arithmetic of the previous chapter: six pairwise comparisons among four groups, each at , reject at least one true null about of the time. A procedure is needed that asks the single question "are these means all equal?" at a single controlled error rate.
Analysis of variance is that procedure, and its name records the idea that makes it work. To decide whether several means differ, do not look at the means directly — look at variation, and split it in two. Some of the spread in the pooled data is variation between the groups, and some is variation within them. If the groups genuinely differ, the between-group part is inflated relative to what the within-group part predicts; if they do not, the two are estimating the same thing and their ratio hovers around one. The whole chapter is that comparison made precise: a theorem that the split is exact, two expectation calculations showing what each piece estimates, and the distribution of their ratio.
Three things follow, and they occupy the rest of the chapter. A significant says some means differ but not which, so a second stage is needed that identifies pairs without giving back the error rate the -test protected. The assumptions deserve scrutiny, because one of them matters much more than the other two. And two-way designs let a second factor enter, which introduces something genuinely new — interaction, the possibility that a factor's effect depends on the level of another.
The last section makes a connection that is usually left implicit: ANOVA is not a separate theory. It is a linear model with indicator predictors, and its -test is the model-comparison test of the next chapter. Seeing that now means the regression chapter has less to prove.
Notation. groups indexed by , with observations in group , total , group means and grand mean . The -test facts used are from *Statistical Inference: Hypothesis Testing*.
5.1Comparing several means: the F-test
Definition 5.1 (The one-way ANOVA model and hypotheses). The one-way model writes each observation as a group mean plus noise,
and tests
Note the alternative. It is not "all differ" but "not all equal", which is why a rejection identifies no particular pair — a point that the post-hoc section exists to address.
The construction rests on an exact algebraic identity, true of any numbers whatever, with no probability in it at all.
Theorem 5.2 (Decomposition of the total sum of squares). With , and ,
Proof. Insert the group mean and expand:
The first term is . In the third, the summand does not depend on , so summing over multiplies it by , giving .
The cross term vanishes. Fix and factor out , which is constant in :
because deviations from a group's own mean sum to zero, by the proposition Deviations from the mean sum to zero. Summing over leaves zero.∎
The identity is exact and unconditional, which is worth pausing on: it holds for any data set, whether or not the model is true. What the probability assumptions determine is not the split but what the two pieces estimate, and that is the content of the next two results.
Definition 5.3 (Mean squares and the F statistic). The mean squares divide each sum of squares by its degrees of freedom,
and the F statistic is their ratio, .
The degrees of freedom count freely varying comparisons, exactly as before. The group means must reproduce the grand mean, one constraint, leaving ; and within the groups, each of the samples loses one degree of freedom to its own mean, leaving . The two add to , the degrees of freedom of , mirroring the sum-of-squares identity.
Theorem 5.4 (What the two mean squares estimate). Under the one-way model,
where . In particular when holds, and otherwise.
Proof. For . Within group , the quantity is times that group's sample variance, whose expectation is by the theorem Unbiasedness of the sample variance. So , and summing over ,
Dividing by gives . Nothing about the entered, which is why this holds whether or not is true.
For . Write with of mean and variance , and . Then , and since the second piece has mean zero,
A computation using gives . Multiplying by and summing,
Dividing by gives the stated expectation. The sum is non-negative and is zero exactly when all are equal.∎
That theorem is the entire logic of the -test in one line. Both mean squares estimate when the null holds, so their ratio should sit near ; when the means differ, only the numerator is inflated, and by an amount measuring exactly how spread out the true means are. A large is evidence against equality, and a small is not evidence of anything in particular.
Theorem 5.5 (Null distribution of F). Under the one-way model with true,
the distribution with and degrees of freedom. The test rejects for large .
Proof. We quote Cochran's theorem, which states that under normality and are independent chi-square variables with and degrees of freedom when holds. Given that, is the ratio of two independent chi-squares each divided by its degrees of freedom,
with cancelling — which is the definition of the distribution. The test is one-sided in the upper tail because, by the theorem What the two mean squares estimate, departures from inflate the numerator only.∎
Proposition 5.6 (With two groups, F is the square of t). When , the ANOVA statistic equals the square of the pooled two-sample statistic, and the two tests reject on exactly the same data.
Proof. With two groups the grand mean is , so and . Hence, writing ,
With we have , and is precisely the pooled variance. Therefore
The critical values match as well, since .∎
So ANOVA does not replace the -test; it extends it, agreeing exactly where both apply. Note that the squaring makes the -test inherently two-sided — the sign of is lost — which is why one-sided alternatives have no place in ANOVA.
Example 5.8 (A complete one-way ANOVA). Three teaching methods are compared on students each. The group means are , and , and each group has sample variance . Test equality at , given .
Solution.
- Grand mean. Equal group sizes, so .
- Between. , on degrees of freedom, so .
- Within. , on degrees of freedom, so .
- Statistic. . Since , reject: the methods do not all have the same mean.
- Check the identity. , on degrees of freedom ✓, which is the theorem Decomposition of the total sum of squares used as an arithmetic check.
- Sanity check on : it estimates , and each group's sample variance was , so exactly as it should — with equal group sizes is just the average of the group variances.
- What has not been established: which methods differ. The alternative was "not all equal", and all the test reports is that. The post-hoc section takes that up.
Definition 5.9 (Effect size: eta-squared). The proportion of total variation attributable to group membership is
Pitfall. A significant and a large are different claims, and with enough data the first arrives without the second. In the example , a large effect — but with per group the same -test would reject for group means differing by a fraction of a standard deviation, where might be .
Always report , or the group means with intervals, alongside the -test. The test answers "is there a difference?"; only the effect size answers "does it matter?".
Example 5.11 (When F does not reject, and why that is not a null result). Four fertilisers are compared on plots each. The yields give and . Test at , given , and report the effect size.
Solution.
- Degrees of freedom: between, within.
- and .
- . Since , do not reject.
- Note . By the theorem What the two mean squares estimate that is entirely ordinary under the null — the ratio of two estimates of the same falls below one about half the time — and it is never evidence for beyond what the test already says.
- Effect size: , so . About of the variation is between fertilisers.
- Sanity check on what can be claimed: with per group and (so ), the standard error of a difference of two group means is . A genuine difference of units — over a standard deviation — would be only two standard errors out, so this design has poor power. "No significant difference" here means the experiment was too small to tell, not that the fertilisers are equivalent.
5.2What the assumptions do, and which one matters
The one-way model made three assumptions, and they are not equally important. Knowing which is which is the difference between a useful diagnostic habit and superstition.
Definition 5.12 (The three ANOVA assumptions).
- Independence — observations are independent within and across groups.
- Normality — the errors are normally distributed.
- Homoscedasticity — every group has the same variance .
Independence is the one that matters. It cannot be checked from the data alone — it is a property of how the data were collected — and violating it invalidates everything, because the whole degrees-of-freedom bookkeeping assumes each observation contributes one independent piece of information. Repeated measures on the same subjects, students clustered in classrooms, measurements taken in time order with drift: each breaks independence, and each requires a different model rather than a correction.
Normality matters least. The statistic is built from sample means, and the central limit theorem is already making those approximately normal; for moderate group sizes the test's actual level stays close to nominal for any population that is not wildly skewed or heavy-tailed. Checking normality of the raw data is in any case the wrong check — the assumption is about the residuals , and the raw data are expected to look multimodal when the means genuinely differ.
Equal variance matters conditionally, and the condition is the group sizes. With equal the -test is remarkably robust to unequal variances; with unequal it is not, and it fails in a specific direction.
Pitfall. The dangerous combination is small groups with large variances. Then , a weighted average of the group variances with weights , is pulled toward the variances of the large groups and so understates the noise affecting the small ones. The result is an that is too large and a test that rejects far more often than .
The reverse pairing — large groups with large variances — makes the test conservative instead. Either way the remedy is Welch's ANOVA, which does not pool, and it is the sensible default whenever the group sizes are unequal.
Example 5.13 (Diagnosing a design before trusting the F). Three groups have and sample standard deviations , and . Should the ordinary -test be used?
Solution.
- The group sizes are badly unequal and so are the variances: the ratio of the largest to smallest variance is , far beyond the informal factor-of-four rule of thumb.
- Worse, the pairing is the dangerous one. weights by , so it is dominated by the large group with the small variance:
- That figure understates the noise in groups 2 and 3, whose variances are and , so any difference involving those groups is divided by too small a denominator and the is inflated.
- Sanity check on the direction: had the large group been the variable one, would have been pulled up and the test made conservative. The asymmetry is entirely about which groups carry the weight.
- Conclusion: use Welch's ANOVA, which estimates each group's variance separately and adjusts the degrees of freedom, exactly as Welch's -test does for two groups.
Remark (On formal tests of equal variance). Bartlett's and Levene's tests test : all variances equal. Bartlett's is powerful under normality and notoriously sensitive to non-normality, so it frequently rejects for the wrong reason; Levene's, which applies ANOVA to the absolute deviations from group medians, is the robust choice.
But the more important point is that using either as a gatekeeper — test variances, then choose the main test accordingly — is itself a two-stage procedure with an error rate nobody computes. With unequal group sizes it is simpler and safer to use Welch's ANOVA unconditionally, which costs almost nothing when the variances are in fact equal.
Example 5.14 (Checking the right thing). A researcher pools all observations from three groups, plots a histogram, finds it clearly bimodal, and concludes the normality assumption fails. Is the reasoning sound?
Solution.
- No — the check is on the wrong quantity. The model assumes the errors are normal, which corresponds to the residuals , not to the raw pooled data.
- If the group means genuinely differ, the pooled data is a mixture of three distributions centred at different points, and a mixture of well-separated normals is bimodal or trimodal by construction. Finding that shape is evidence that the means differ, which is what the analysis is trying to establish.
- The correct procedure: compute the residuals , which recentre every group at zero, and examine those — by a normal quantile plot, ideally, since histograms of points are unstable.
- Sanity check on how much it would matter anyway: normality is the least critical of the three assumptions, since the statistic is built from group means that the central limit theorem is already normalising. With per group and no extreme skew, a modest departure changes the true level very little.
- Contrast with what would be worth worrying about: if those observations were subjects each measured three times, independence fails, and no amount of residual checking would reveal it from the numbers alone.
5.3After a significant F: multiple comparisons
A rejected says the means are not all equal. Finding out which differ means making pairwise comparisons, and doing that naively hands back exactly the error rate the -test was constructed to protect.
Definition 5.15 (Comparisonwise and familywise error rates). The comparisonwise error rate is the probability that a single comparison falsely rejects. The familywise error rate is the probability that at least one of the comparisons in the set does. With groups there are pairs.
For that is six comparisons, and by the proposition Family-wise error rate under multiplicity, running each at gives a familywise rate of about . The whole point of post-hoc procedures is to control the second rate rather than the first.
Theorem 5.16 (Tukey's honestly significant difference). For equal group sizes , all pairwise differences may be compared against
where is the upper- point of the studentised range distribution for groups and degrees of freedom. Declaring whenever holds the familywise error rate at exactly .
Proof. We quote the distribution of the studentised range and show why it is the right reference. A familywise error occurs when any pair is falsely separated, and under the complete null the largest of all the pairwise differences is — the range of the group means. So
since a threshold is exceeded by some pair exactly when it is exceeded by the most extreme pair. The quantity on the left is by definition the studentised range statistic, so choosing as its upper- point makes this probability exactly .∎
That proof explains why Tukey beats Bonferroni here. Bonferroni bounds the familywise rate by the union bound, which ignores the fact that the comparisons are highly dependent — they share the same group means and the same . Tukey uses the exact joint distribution of the most extreme comparison, so it is not conservative: it achieves rather than staying under it, and its intervals are correspondingly narrower.
Example 5.17 (Tukey against Bonferroni on the teaching data). For the three teaching methods (; ; on df), find which pairs differ at familywise . Use and .
Solution.
- Tukey. .
- The three differences are , , . Comparing with : methods 1 and 2 do not differ; method 3 differs from both.
- Bonferroni. With three comparisons the per-comparison level is , two-sided, so the margin is
- Bonferroni's threshold exceeds Tukey's , so Bonferroni is stricter and would reach the same conclusions here with less power to spare.
- Sanity check on the ordering: Bonferroni must be the more conservative of the two, since it bounds a probability that Tukey computes exactly ✓. The gap widens with — at there are comparisons and Bonferroni's penalty becomes severe.
- Note the conclusion is not transitive-looking nonsense but a genuine finding: the data separate method 3 from the other two and cannot separate those two from each other.
Example 5.18 (The cost of more groups). Four treatments are compared, each, with on degrees of freedom and means . Apply Tukey at familywise , with .
Solution.
- .
- The six pairwise differences are , , , , , .
- Against : significant are , and . Treatment 4 differs from all three others; the first three are indistinguishable from one another.
- Sanity check on the multiplicity: had each comparison been run at an uncorrected , the margin would be , which would also declare significant — a fourth finding, bought by abandoning the familywise guarantee.
- The familywise rate of that uncorrected procedure would be about , so roughly a one-in-four chance of at least one false claim even if all four treatments were identical.
Pairwise comparisons are not the only question one might ask, and often not the interesting one.
Definition 5.19 (Contrast). A contrast is a linear combination of the group means
It is estimated by , with standard error .
The constraint is what makes a contrast a comparison rather than a level: it is unaffected by adding a constant to every mean. A pairwise difference is the contrast ; "treatments versus control" is ; a linear trend across ordered doses is .
Example 5.20 (A planned contrast). Test whether teaching method 3 differs from the average of methods 1 and 2, using the same data.
Solution.
- The contrast is , which sums to zero ✓.
- Estimate: .
- Standard error: , so .
- Statistic: on degrees of freedom. Against , this is significant.
- If the contrast was chosen after seeing the data, the critical value is not legitimate — any of infinitely many contrasts could have been picked. Scheffé's method covers all contrasts simultaneously, using the critical value .
- Since , the contrast survives even the fully conservative Scheffé standard ✓ — so the conclusion holds whether it was planned or discovered.
5.4Two factors at once, and interaction
With two factors, a third question appears alongside the two obvious ones, and it is usually the interesting one.
Definition 5.21 (The two-way model with interaction). With factor at levels, factor at levels and observations per cell,
where and are main effects, is the interaction, and the effects are constrained to sum to zero over each index.
Theorem 5.22 (Two-way decomposition). The total sum of squares splits into four orthogonal pieces,
with degrees of freedom , , and summing to . Each effect is tested by its mean square over , giving an with the corresponding degrees of freedom.
Proof. The argument is the one-way proof applied twice. Write each deviation from the grand mean as a telescoping sum
which is an identity — the four right-hand terms telescope to the left-hand side, every intermediate mean cancelling. Squaring and summing, all six cross terms vanish because each is a product in which one factor is constant over an index and the other sums to zero over that index, exactly as in the one-way proof. The degrees of freedom follow from the sum-to-zero constraints on each set of effects.
That the interaction term is the residual from additivity is worth reading off directly: is exactly the amount by which cell departs from what the two main effects alone predict.∎
Intuition. A main effect asks whether a factor moves the response on average. An interaction asks whether how much it moves it depends on the other factor.
On a plot of cell means with one factor on the horizontal axis and one line per level of the other, no interaction means parallel lines — the vertical gap between them is the same everywhere, which is precisely additivity. Interaction means the lines converge, diverge, or cross. Crossing is the extreme case: the factor helps at one level of the other factor and hurts at another, so no single statement about its effect is true.
Example 5.23 (Additive effects, and then an interaction). Two designs, each with factor at levels, at levels, per cell. Compute , and for cell means
(a) and ; (b) and .
Solution.
- (a) Row means and ; column means ; grand mean .
- . .
- Interaction: cell is predicted by , and it is . Every cell matches, so . The design is exactly additive: is above at every level of .
- (b) Row means still and ; but column means are now and the grand mean .
- , unchanged. — factor has no main effect at all.
- Interaction: cell is predicted by but is , a departure of . The six departures are , so
- Sanity check on what this means: in (b) the response rises with under and falls with under , so the two effects cancel exactly in the column means. Reporting "factor has no effect" from would be badly wrong — has a large effect whose direction depends on .
Pitfall. When the interaction is significant, do not interpret the main effects. Design (b) above is the reason: while moves the response by units in opposite directions at the two levels of . A main effect is an average over the other factor's levels, and averaging over a sign change destroys the information.
The correct report when interaction is present is the cell means themselves, or the simple effects — the effect of within , and separately within .
Example 5.25 (A complete two-way ANOVA table). For design (b) — , , — the within-cell sum of squares is , with , , . Build the ANOVA table and test each effect at , given and .
Solution.
- Degrees of freedom: has ; has ; has ; error has . These sum to ✓.
- Mean squares: ; ; ; .
- The three statistics, each over :
- Against the critical values: , not significant; , not significant; , strongly significant.
- So the table reports no main effects and a large interaction — which by the pitfall above is precisely the situation in which the main effects must not be interpreted. Concluding "neither factor matters" from the first two rows would be the exact opposite of the truth.
- Sanity check by looking at the cell means directly: under the response goes , and under it goes . Factor moves the response by units in both cases — in opposite directions. The correct report is the simple effects, not the main effects.
5.5Fixed and random effects
Everything so far treated the groups as the only groups of interest — three specific teaching methods, four specific fertilisers. Sometimes the groups are instead a sample from a population of groups, and then the interesting question changes.
Definition 5.26 (Fixed and random effects). In a fixed-effects model the levels of the factor are the only ones of interest, and the parameters are unknown constants. In a random-effects model the levels are a random sample from a population of levels, and the model is
with the independent of the . The hypothesis of interest becomes .
The distinction is about the inference intended, not about the arithmetic. Testing three specific teaching methods you wish to choose between is fixed. Sampling twenty schools from a district to ask how much schools vary is random: the twenty schools themselves are of no interest, and the estimand is the variance across all schools.
Theorem 5.27 (Expected mean squares under random effects). For the balanced random-effects model with observations in each of groups,
Proof. The within-group calculation is unchanged from the theorem What the two mean squares estimate, since conditional on the the observations within group are i.i.d. normal about ; the argument never used the group centres.
For the between-group part, note , so
the two terms independent. The are i.i.d. across with this variance, and is times their sample variance, which is unbiased for . Hence
The same tests in both models, because under either null both mean squares estimate — which is why the arithmetic is identical and the distinction is easy to miss. What differs is what a rejection means, and what is estimated afterwards.
Corollary 5.28 (Estimating the variance component). Equating the mean squares to their expectations gives the method-of-moments estimator
truncated at zero when the numerator is negative.
Proof. Set and and solve, which is the method of moments of the definition Method of moments estimator applied to the mean squares.∎
Pitfall. That estimator can be negative, and a negative variance estimate is not a computational error — it happens whenever , which under occurs about half the time. Truncating at zero is the usual fix and introduces an upward bias. It is a genuine awkwardness of the moment approach, and the reason maximum likelihood or its restricted variant is preferred in practice for variance components.
Example 5.29 (How much do schools vary?). Twenty schools are sampled and pupils tested in each. The ANOVA gives and . Estimate the between-school variance and the proportion of variation attributable to schools.
Solution.
- By the corollary Estimating the variance component,
- The within-school variance is estimated by .
- The proportion of total variation between schools — the intraclass correlation — is
- So about of the variation in pupil scores is between schools and within them — a finding about the population of schools, not about these twenty.
- Sanity check on the contrast with a fixed-effects reading: on and degrees of freedom is highly significant, so schools certainly differ. But "significantly different" and "an important source of variation" are different claims, and here the second is modest while the first is emphatic. The variance component is what answers the question actually asked.
5.6ANOVA is a linear model
Analysis of variance is often taught as a separate technique with its own vocabulary. It is not: it is a regression with categorical predictors, and seeing that unifies this chapter with the next one.
Definition 5.30 (Indicator coding). For groups, define indicator (dummy) variables
taking group as the reference. The model has and .
Proposition 5.31 (The F-test is a model comparison). Under indicator coding, the hypothesis is exactly , and the one-way ANOVA statistic is the statistic comparing the full model against the intercept-only model:
Proof. Fitting the full model by least squares sets each group's fitted value to its own mean — the theorem The mean minimises total squared deviation applied within each group, since the indicators let every group have a free level. So the residual sum of squares of the full model is .
The intercept-only model fits the grand mean to everything, with residual sum of squares . Therefore by the theorem Decomposition of the total sum of squares, and substituting,
which is the ANOVA statistic. The hypotheses match because , so all vanish exactly when all group means equal .∎
Remark (Why the reference level is arbitrary but the F is not). Choosing group as reference makes each a comparison against group , so the individual coefficients depend on that choice — a different reference gives different coefficients describing the same fit. What does not change is the set of fitted values, hence , hence .
This is worth knowing because a table of regression coefficients for a categorical predictor invites the reader to compare each level against the reference and stop there, which is both an arbitrary set of comparisons and an uncorrected one. The -test asks the question that does not depend on the coding.
Example 5.32 (The same test, twice). For the three teaching methods, write the indicator model, give the least-squares coefficients, and confirm the -test agrees with the ANOVA computed earlier.
Solution.
- With method 1 as reference, and indicate methods 2 and 3. The fitted coefficients are the group means recoded:
- The fitted value for any observation is its own group's mean, so and .
- Model comparison: over , giving ✓ — identical to the ANOVA computation.
- Sanity check on the coefficients: is the difference between methods 3 and 1, and the Tukey analysis found exactly that pair to differ. But note that reading significance off the individual coefficients would be three uncorrected comparisons against the reference, which is why the comes first.
- Had method 3 been the reference instead, the coefficients would be , , — different numbers, identical fit, identical .
- **Running pairwise -tests instead of ANOVA.** Six comparisons at reject at least one true null about of the time. Use the -test first, then a procedure that controls the familywise rate.
- **Concluding which groups differ from a significant . ** The alternative is "not all equal". Identifying pairs requires Tukey, Bonferroni or a planned contrast.
- Interpreting main effects when the interaction is significant. A main effect averages over the other factor, and averaging over a sign change can give exactly zero for a large effect.
- Checking normality of the raw data rather than the residuals. When the means genuinely differ, the pooled raw data *should* look multimodal; the assumption concerns .
- Ignoring unequal variances when the group sizes are unequal. Small groups with large variances inflate and the test rejects far too often. Use Welch's ANOVA.
- Treating independence as checkable from the data. It is a property of the design. Repeated measures and clustered observations need a different model, not a correction.
- Reporting significance without an effect size. With enough data any difference is significant; or the group means with intervals say whether it matters.
- Using a one-sided ANOVA. is a squared quantity and loses the sign; the test is inherently two-sided.