Contents / Statistics / Analysis of Variance

Chapter 5

Analysis of Variance

The sum-of-squares decomposition and what each mean square estimates; the F-test and its equivalence to the two-sample t-test; which assumptions matter and why; Tukey, Bonferroni, contrasts and Scheffe for multiple comparisons; two-way designs and interaction; fixed versus random effects and variance components; and ANOVA as a linear model with indicator predictors.

5.3After a significant F: multiple comparisons

A rejected FF says the means are not all equal. Finding out which differ means making pairwise comparisons, and doing that naively hands back exactly the error rate the FF -test was constructed to protect.

Definition 5.15 (Comparisonwise and familywise error rates). The comparisonwise error rate is the probability that a single comparison falsely rejects. The familywise error rate is the probability that at least one of the comparisons in the set does. With kk groups there are (k2)\binom{k}{2} pairs.

For k=4k=4 that is six comparisons, and by the proposition Family-wise error rate under multiplicity, running each at 5%5\% gives a familywise rate of about 1−(0.95)6=0.2651-(0.95)^6 = 0.265. The whole point of post-hoc procedures is to control the second rate rather than the first.

Theorem 5.16 (Tukey's honestly significant difference). For equal group sizes nn, all pairwise differences may be compared against

HSD=qα, k, N−kMSWn,\text{HSD} = q_{\alpha,\,k,\,N-k}\sqrt{\frac{MSW}{n}},

where qq is the upper-α\alpha point of the studentised range distribution for kk groups and N−kN-k degrees of freedom. Declaring μi≠μj\mu_i \ne \mu_j whenever ∣xˉi−xˉj∣>HSD|\bar x_i - \bar x_j| > \text{HSD} holds the familywise error rate at exactly α\alpha.

Proof. We quote the distribution of the studentised range and show why it is the right reference. A familywise error occurs when any pair is falsely separated, and under the complete null the largest of all the pairwise differences is max⁡iXˉi−min⁡iXˉi\max_i \bar X_i - \min_i \bar X_i — the range of the group means. So

P(any false rejection)=P(max⁡iXˉi−min⁡iXˉiMSW/n>q),P(\text{any false rejection}) = P\left(\frac{\max_i \bar X_i - \min_i \bar X_i}{\sqrt{MSW/n}} > q\right),

since a threshold is exceeded by some pair exactly when it is exceeded by the most extreme pair. The quantity on the left is by definition the studentised range statistic, so choosing qq as its upper-α\alpha point makes this probability exactly α\alpha.∎

That proof explains why Tukey beats Bonferroni here. Bonferroni bounds the familywise rate by the union bound, which ignores the fact that the comparisons are highly dependent — they share the same group means and the same MSWMSW. Tukey uses the exact joint distribution of the most extreme comparison, so it is not conservative: it achieves α\alpha rather than staying under it, and its intervals are correspondingly narrower.

Example 5.17 (Tukey against Bonferroni on the teaching data). For the three teaching methods (xˉ=10,12,17\bar x = 10, 12, 17; n=5n=5; MSW=8MSW = 8 on 1212 df), find which pairs differ at familywise α=0.05\alpha=0.05. Use q0.05,3,12=3.77q_{0.05,3,12}=3.77 and t0.00833,12=2.779t_{0.00833,12}=2.779.

Solution.

  1. Tukey. HSD=3.778/5=3.771.6=3.77×1.26491=4.769\text{HSD} = 3.77\sqrt{8/5} = 3.77\sqrt{1.6} = 3.77\times 1.26491 = 4.769.
  2. The three differences are ∣10−12∣=2|10-12| = 2, ∣10−17∣=7|10-17| = 7, ∣12−17∣=5|12-17| = 5. Comparing with 4.7694.769: methods 1 and 2 do not differ; method 3 differs from both.
  3. Bonferroni. With three comparisons the per-comparison level is 0.05/3=0.01670.05/3 = 0.0167, two-sided, so the margin is
tMSW(1n+1n)=2.7798×0.4=2.7793.2=2.779×1.78885=4.971.t\sqrt{MSW\left(\tfrac1n+\tfrac1n\right)} = 2.779\sqrt{8\times 0.4} = 2.779\sqrt{3.2} = 2.779\times 1.78885 = 4.971 .
  1. Bonferroni's threshold 4.9714.971 exceeds Tukey's 4.7694.769, so Bonferroni is stricter and would reach the same conclusions here with less power to spare.
  2. Sanity check on the ordering: Bonferroni must be the more conservative of the two, since it bounds a probability that Tukey computes exactly ✓. The gap widens with kk — at k=6k=6 there are 1515 comparisons and Bonferroni's penalty becomes severe.
  3. Note the conclusion is not transitive-looking nonsense but a genuine finding: the data separate method 3 from the other two and cannot separate those two from each other.
□

Example 5.18 (The cost of more groups). Four treatments are compared, n=8n=8 each, with MSW=12MSW = 12 on 2828 degrees of freedom and means 20,23,24,2920, 23, 24, 29. Apply Tukey at familywise α=0.05\alpha=0.05, with q0.05,4,28=3.861q_{0.05,4,28}=3.861.

Solution.

  1. HSD=3.86112/8=3.8611.5=3.861×1.224745=4.729\text{HSD} = 3.861\sqrt{12/8} = 3.861\sqrt{1.5} = 3.861\times 1.224745 = 4.729.
  2. The six pairwise differences are ∣20−23∣=3|20-23|=3, ∣20−24∣=4|20-24|=4, ∣20−29∣=9|20-29|=9, ∣23−24∣=1|23-24|=1, ∣23−29∣=6|23-29|=6, ∣24−29∣=5|24-29|=5.
  3. Against 4.7294.729: significant are (20,29)(20,29), (23,29)(23,29) and (24,29)(24,29). Treatment 4 differs from all three others; the first three are indistinguishable from one another.
  4. Sanity check on the multiplicity: had each comparison been run at an uncorrected 5%5\%, the margin would be t0.025,2812(1/8+1/8)=2.048×1.732=3.547t_{0.025,28}\sqrt{12(1/8+1/8)} = 2.048\times 1.732 = 3.547, which would also declare (20,24)(20,24) significant — a fourth finding, bought by abandoning the familywise guarantee.
  5. The familywise rate of that uncorrected procedure would be about 1−(0.95)6=0.2651-(0.95)^6 = 0.265, so roughly a one-in-four chance of at least one false claim even if all four treatments were identical.
□

Pairwise comparisons are not the only question one might ask, and often not the interesting one.

Definition 5.19 (Contrast). A contrast is a linear combination of the group means

L=∑i=1kciμiwith∑i=1kci=0.L = \sum_{i=1}^{k} c_i\mu_i \qquad \text{with} \qquad \sum_{i=1}^{k} c_i = 0 .

It is estimated by L^=∑icixˉi\hat L = \sum_i c_i\bar x_i, with standard error MSW∑ici2/ni\sqrt{MSW\sum_i c_i^{2}/n_i}.

The constraint ∑ci=0\sum c_i = 0 is what makes a contrast a comparison rather than a level: it is unaffected by adding a constant to every mean. A pairwise difference is the contrast (1,−1,0)(1,-1,0); "treatments versus control" is (−1,12,12)(-1, \tfrac12, \tfrac12); a linear trend across ordered doses is (−1,0,1)(-1,0,1).

Example 5.20 (A planned contrast). Test whether teaching method 3 differs from the average of methods 1 and 2, using the same data.

Solution.

  1. The contrast is c=(−12,−12,1)c = (-\tfrac12, -\tfrac12, 1), which sums to zero ✓.
  2. Estimate: L^=−12(10)−12(12)+17=−11+17=6\hat L = -\tfrac12(10) - \tfrac12(12) + 17 = -11 + 17 = 6.
  3. Standard error: ∑ci2/ni=(0.25+0.25+1)/5=0.3\sum c_i^{2}/n_i = (0.25+0.25+1)/5 = 0.3, so SE=8×0.3=2.4=1.5492SE = \sqrt{8\times 0.3} = \sqrt{2.4} = 1.5492.
  4. Statistic: t=6/1.5492=3.873t = 6/1.5492 = 3.873 on 1212 degrees of freedom. Against t0.025,12=2.179t_{0.025,12}=2.179, this is significant.
  5. If the contrast was chosen after seeing the data, the tt critical value is not legitimate — any of infinitely many contrasts could have been picked. Scheffé's method covers all contrasts simultaneously, using the critical value (k−1)Fα,k−1,N−k=2×3.885=7.77=2.787\sqrt{(k-1)F_{\alpha,k-1,N-k}} = \sqrt{2\times 3.885} = \sqrt{7.77} = 2.787.
  6. Since 3.873>2.7873.873 > 2.787, the contrast survives even the fully conservative Scheffé standard ✓ — so the conclusion holds whether it was planned or discovered.
□
About the practice questions. They check that you can carry out this chapter's computations correctly, and each one is graded on a single answer. They are not proof exercises: working through them confirms the mechanics, not that you could prove the results yourself. For that, re-read the statements above and try to reconstruct their proofs with the page closed.
Helpful?