Contents / Statistics / Statistical Inference: Hypothesis Testing
Chapter 4
Statistical Inference: Hypothesis Testing
The logic of a test; Type I and Type II errors, power and sample size; the Neyman-Pearson lemma and uniformly most powerful tests; the p-value, its uniformity under the null and its duality with confidence intervals; the t-test family; chi-square tests for counts; and multiple testing, Bonferroni and the post-test probability of a true effect.
Introduction
The previous chapter asked what values of a parameter the data support. This one asks a narrower question: is one particular value refuted? That sounds like less, and in an important sense it is — a test returns a verdict where an interval returns a range — but the narrowness buys something, because a sharp question can be answered with a controlled error rate.
The logic is a probabilistic version of proof by contradiction, and it is worth stating plainly because almost every misuse of testing comes from forgetting it. Assume the claim you wish to discredit. Work out what the data would look like if it were true. If what you actually observed would be extraordinary under that assumption, the assumption is discredited. The word doing all the work is extraordinary, because unlike in a genuine contradiction, no observation is impossible — so the argument can only ever be probabilistic, and it can always be wrong.
Being wrong comes in two kinds, and they trade against each other. Convicting an innocent hypothesis and acquitting a guilty one cannot both be made rare at a fixed sample size, and the whole technical apparatus of this chapter is machinery for managing that trade: the significance level fixes the first rate by decree, power measures the second, and the Neyman–Pearson lemma identifies the test that does best on the second for a given first. That lemma is the theoretical centre of the chapter and the reason the familiar tests have the form they do.
Then the -value, which is the single most misreported quantity in applied science. We prove the one fact that pins its meaning down — that it is uniformly distributed when the null is true — and derive from that both why the threshold works and why a -value is not the probability that the null is true. The chapter closes with what follows when many tests are run at once, where the arithmetic of multiplicity turns a well-behaved procedure into a machine for generating false findings.
Throughout, is the null value, the significance level, the standard normal distribution function, and the value with . The sampling distributions used are proved in Sampling and Data Distributions, and the pivots and -distribution facts in Statistical Inference: Estimation.
4.1The logic of a test
A test begins by splitting the parameter space into two claims, asymmetrically. The asymmetry is the point and is often mistaken for a defect.
Definition 4.1 (Null and alternative hypotheses). The null hypothesis is the claim to be tested, stated so that it determines a distribution for the data — typically . The alternative is the claim that holds if the null fails, and is two-sided () or one-sided (, or ).
The null must be specific enough to compute with: " " fixes a sampling distribution, while " " does not, which is why the equality is always the null and never the alternative. This is not a statement about which claim is more likely to be true. It is a statement about which claim can be used to generate predictions.
Definition 4.2 (Test statistic, rejection region, significance level). A test statistic is a function of the data whose distribution is known when holds. The rejection region is the set of values of leading to rejection of , and the significance level is
the probability of rejecting a true null.
Theorem 4.3 (The one-sample z-test). Let be i.i.d. with mean and known variance , and suppose is normal (exactly, for a normal population; approximately, by the central limit theorem). For testing , the statistic
is standard normal under . Rejecting when gives a two-sided test of exact level ; rejecting when gives a one-sided test of level against .
Proof. Under the true mean is , so by the theorem Mean and variance of the sample mean, has mean and variance . Standardising, .
For the two-sided region, by symmetry of the normal and the definition of ,
For the one-sided region, directly. In each case the probability is computed under , which is exactly the definition of the significance level.∎
Notice what the theorem does and does not deliver. It guarantees the rate of one specific error — rejecting a true null — and says nothing whatever about the other. A test that never rejects has and is useless; controlling alone is not a criterion for a good test, only a constraint that any admissible test must satisfy.
Intuition. A test is a trial with the null hypothesis as defendant, and the presumption of innocence is built into the asymmetry. We do not ask which hypothesis is more plausible; we ask whether the evidence against the null is strong enough to convict beyond a fixed standard of doubt, and is that standard.
Two consequences follow immediately and are routinely got wrong. An acquittal is not a finding of innocence — failing to reject means the evidence was insufficient, not that is true, exactly as "not guilty" is not "proven innocent". And the standard of doubt is chosen before the evidence is examined, not adjusted afterwards to reach a preferred verdict.
Pitfall. "Fail to reject " and "accept " are different claims, and only the first is licensed. A test with a small sample fails to reject almost everything, because it lacks the power to detect anything; reporting that as evidence for the null converts an absence of evidence into evidence of absence.
If you need to argue that a parameter is close to , a confidence interval does it and a test does not: a narrow interval around is a positive finding, while a large -value is compatible with both a true null and a hopeless experiment.
Example 4.4 (A one-sided z-test, in full). A machine should fill bottles to ml with ml. A regulator suspects underfilling and samples bottles, obtaining ml. Test against at .
Solution.
- The alternative is one-sided and points downward, so the rejection region is .
- The standard error is .
- The test statistic is
- Since , reject at the level. The evidence is consistent with underfilling.
- Sanity check against the interval of the previous chapter: the two-sided interval was , which excludes — the same conclusion by a different route, and that agreement is a theorem proved later in this chapter.
- A caution on the one-sided choice: it was made because the regulator's concern was underfilling specifically, before seeing the data. Choosing the direction after seeing that would double the true error rate while reporting it as .
Example 4.5 (A two-sided test for a proportion). A supplier claims a defect rate of . An inspector finds defects in units. Test against a two-sided alternative at .
Solution.
- Check the approximation is usable: and , both comfortably large, so the theorem Normal approximation to the sample proportion applies.
- The observed proportion is .
- The standard error is computed under the null, using and not — this is a test, not an interval, so the null value is available and should be used:
- The statistic is , and , so reject .
- The two-sided -value is .
- Sanity check on the choice of denominator: had we used in the standard error we would have got and , still significant but noticeably smaller. Using is correct for a test because the null specifies the whole distribution, including the variance.
4.2Two kinds of error, and the power of a test
Definition 4.6 (Type I and Type II errors, power). A Type I error is rejecting when it is true; its probability is . A Type II error is failing to reject when it is false; its probability is written . The power of a test against a particular alternative is
Power is not a single number but a function of the alternative, and that is the useful way to hold it. A test has poor power against alternatives near the null — telling from is nearly impossible at any realistic — and good power against distant ones. Asking "what is the power of this test?" without naming an effect size is asking an incomplete question.
Theorem 4.7 (Power of the one-sided z-test). For the one-sided test of against at level , rejecting when , the power at a true mean is
Proof. Power is the rejection probability computed under the true mean , not under . When , the standardised mean is standard normal. Rewrite the rejection event in terms of it:
obtained by subtracting from both sides and dividing by . The left-hand side is standard normal under , so the probability is of the right-hand side.∎
The quantity is the effect size measured in standard errors, and the formula says power depends on the data only through it. Three levers raise power, and the formula prices each: a larger true effect, a smaller , or a larger — the last only through , so power is bought at the same quadratic rate as precision was in the previous chapter.
Corollary 4.8 (Sample size for a target power). To detect a difference of with power in a one-sided level- -test,
Proof. Require , that is , which rearranges to . Since is increasing and , this holds exactly when
Solving for gives , and squaring — both sides positive — gives the result.∎
Example 4.10 (Sizing a study for power). A treatment is worth adopting if it raises the mean by units, where . How many subjects are needed for a one-sided test to have power?
Solution.
- The critical values are and , so .
- By the corollary Sample size for a target power,
so take . 3. Check with the power formula: at the effect is standard errors, giving power ✓. 4. Sanity check on the cost of a smaller effect: halving to quadruples the requirement to . And raising the power target from to (with on both terms) needs , so — a increase for the last fifteen points of power. 5. The practical moral: a study that does not compute this before collecting data is likely to be underpowered, and an underpowered study that fails to reject has established nothing at all.□
Remark (Why underpowered studies are worse than no study). An underpowered test is usually described as likely to miss a real effect, which is true and is only half the problem. The other half is what happens when it does reject.
To reject, an underpowered study needs an unusually large observed effect — that is what crossing the threshold with little data requires. So among the underpowered studies that reach significance, the estimated effects are systematically inflated, sometimes by several-fold. The literature then contains a set of published findings that are exaggerated conditional on having been published, and that inflation does not wash out as more such studies accumulate.
Example 4.11 (Computing both error probabilities for one test). A test of against uses , and . Find the rejection rule in the original units, then the Type II error probability if the true mean is .
Solution.
- The standard error is , so the rule "reject when " becomes
- A Type II error is failing to reject when , that is observing when the true mean is . Standardise using the true mean:
- So the power against is , and the test misses a real five-unit effect about of the time.
- Sanity check by the power formula: , so power ✓.
- Note the asymmetry deliberately built in: by decree, as a consequence. The two are not treated even-handedly, because rejecting a true null is regarded as the more serious error — and if that judgement does not fit the application, it is that should be reconsidered, not the arithmetic.
4.3Optimal tests: the Neyman-Pearson lemma
The -test rejects for large , which seems natural, but naturalness is not an argument. Among all tests with significance level , is that one best? For a simple alternative the question has a complete answer, and it determines the form of essentially every test in use.
Definition 4.12 (Simple hypothesis and likelihood ratio). A hypothesis is simple if it specifies the distribution of the data completely. For two simple hypotheses and with likelihoods and , the likelihood ratio is
Theorem 4.13 (Neyman-Pearson lemma). Among all tests of a simple against a simple with significance level at most , the test that rejects when
with chosen so the level is exactly , has the greatest possible power.
Proof. Write for the likelihood-ratio test — the indicator of — and let be any other test with level at most , so , where denotes expectation under .
The key is a pointwise inequality. Consider
To see it, take the two cases. Where , the second factor is positive and , so the first factor is non-negative. Where , the second factor is non-positive and , so the first factor is non-positive. In both cases the product of two like-signed quantities is non-negative.
Now integrate the inequality over all . The integral of is , the difference in power; the integral of is , the difference in level. So
the last step because and . Hence : the likelihood-ratio test is at least as powerful as any competitor.∎
Proposition 4.14 (The z-test is the Neyman-Pearson test for a normal mean). For a normal population with known , testing against a simple with , the likelihood-ratio test rejects exactly when exceeds a constant — that is, it is the one-sided -test.
Proof. Write the ratio of normal likelihoods and take logs. The quadratic terms cancel between numerator and denominator, leaving
Since the bracket's coefficient is positive, so is a strictly increasing function of . Therefore is the same event as for a corresponding constant , and choosing to give level gives .∎
That proposition is the justification the -test was missing. It is not merely a reasonable procedure: for a normal population with known variance, no test at the same level has more power against any specific alternative on that side.
Remark (Uniformly most powerful tests, and where they stop existing). The constant in the proposition does not depend on — it dropped out. So the same test is most powerful against every simultaneously, which makes it uniformly most powerful for the composite alternative .
For a two-sided alternative no such test exists, and the reason is visible in the algebra: the optimal test against rejects for large , while against it rejects for small , and no single rejection region can be optimal for both. The usual two-sided test is a compromise, justified by unbiasedness and symmetry rather than by uniform optimality. This is the honest reason two-sided tests have less power than one-sided ones — and the reason the direction must be chosen in advance, since choosing it from the data recovers the two-sided error rate while reporting the one-sided one.
Example 4.15 (Building a most powerful test from scratch). A single observation comes from an exponential distribution with rate . Construct the most powerful level- test of against , and find its power at .
Solution.
- The likelihood ratio is
which is strictly decreasing in . 2. So is the event for a corresponding constant: by the theorem Neyman-Pearson lemma, the most powerful test rejects for small . This is worth noting because "reject for large values of the statistic" is a habit, not a rule — here the larger rate produces smaller observations, so small values are the evidence against . 3. Choose for level . Under , Exp, so , giving . 4. Power: under , Exp, so
- Sanity check: the power exceeds , as any sensible test's must, but only barely — one observation carries very little information about a rate. With observations the ratio becomes , the test rejects for small , and the power rises toward .
4.4The p-value: what it is, and what it is not
Reporting only "reject at " throws information away: a statistic barely past the threshold and one far beyond it lead to the same verdict but are not the same evidence. The -value reports where the observation fell.
Definition 4.16 (p-value). The -value is the probability, computed assuming is true, of obtaining a test statistic at least as extreme as the one observed, where "extreme" is measured in the direction of the alternative:
The decision rule is then simply: reject when . That this reproduces the level- test exactly is the content of the following theorem, which also pins down what a -value is in a way the definition alone does not.
Theorem 4.17 (The p-value is uniform under the null). If the test statistic has a continuous distribution and is true, then . Consequently for every .
Proof. Take a one-sided test rejecting for large , so that where is the distribution function of under . Since is continuous and increasing on the support, for ,
A random variable with on is uniform on . Setting gives the Type I error rate, confirming that the rule "reject when " has level exactly .∎
This is the fact to carry. Under the null, -values are not clustered near — they are spread evenly across , so a -value of from a true null is exactly as likely as one of . Small -values are not rare under the null in any absolute sense; they are rare only at the rate you chose. That is simultaneously why the procedure controls error and why running many tests produces small -values by construction.
Pitfall. Four statements that a -value does not make, each common in print.
- It is not . It is a probability computed assuming , so it conditions in the opposite direction. Converting between the two requires a prior, by Bayes' theorem — see the final section.
- It is not the probability the result was due to chance. That phrase describes again.
- is not the probability the alternative is true, nor the probability the finding will replicate.
- A larger -value is not evidence for . Under the null, is uniform; under a small true effect with a small sample, is also spread widely. A of discriminates between those two situations hardly at all.
The relationship to the previous chapter's intervals is exact, not approximate, and it is the cleanest way to see what a test adds and what it omits.
Theorem 4.19 (Duality of tests and confidence intervals). The two-sided level- -test rejects if and only if lies outside the confidence interval for .
Proof. The test rejects exactly when
multiplying by the positive . The negation of that inequality is precisely
which says lies in the confidence interval of the theorem Confidence interval for a mean, known variance. So the test fails to reject exactly when is inside the interval, which is the claim.∎
Remark (Why the interval is usually the better report). By duality the interval contains the test: it tells you the outcome of the test against every null value at once, since the values it contains are exactly those that would not be rejected. It also reports the effect size and its precision, which the test does not.
So an interval strictly dominates a test as a summary, and the practice of reporting " " alone discards information that was already computed. The case where a test is genuinely the right report is when a decision must be made — accept the shipment or reject it — and there the verdict is the deliverable.
Example 4.20 (From test statistic to p-value, both sides). For the bottle-filling data (, , , ), compute the one-sided and two-sided -values and interpret them.
Solution.
- The test statistic was .
- One-sided against : .
- Two-sided: , double the one-sided value because deviations of either sign count as extreme.
- Both are below , so both reject at the level; the two-sided test is the more conservative and is the right choice unless the direction was fixed in advance.
- The correct reading of : if the machine were correctly calibrated, a sample mean this far from in either direction would occur about of the time. It is not the statement that there is a chance the machine is fine.
- Sanity check via duality: , so should fall outside the interval — and it does ✓.
4.5The t-test family
Replacing the known by changes the reference distribution from normal to , exactly as it did for intervals, and for the same reason.
Theorem 4.21 (One-sample t-test). For i.i.d. with unknown, the statistic
has the distribution under . Rejecting when gives a two-sided test of exact level .
Proof. Under the population mean is , so is the studentised mean of the theorem The studentised mean is pivotal, which is -distributed. The rejection probability under is then by the definition of the critical value and the symmetry of the density.∎
Definition 4.22 (Paired t-test). When observations come in matched pairs , form the differences and apply the one-sample -test to them, testing with where is the number of pairs.
Pairing is not a technicality but usually the single most effective design choice available. Each subject serves as its own control, so any variation between subjects — which is often the dominant source of noise — cancels in the difference and never enters the standard error. The cost is one degree of freedom relative to treating the data as independent observations, and the gain is frequently a several-fold reduction in .
Theorem 4.23 (Two-sample t-test, pooled). For independent samples from and with a common unknown variance, the statistic
has the distribution under , with the pooled standard deviation.
Proof. Under the difference has mean zero, and the construction in the proof of the theorem Confidence interval for a difference of means shows the ratio is . The test is that pivot with set to its null value of .∎
Example 4.24 (Paired against unpaired on the same data). Ten subjects are measured before and after a treatment. The differences have and . Separately, the before and after readings each have standard deviation about . Test at using , and compare with what an unpaired analysis would have given.
Solution.
- Paired. , so
- Since , reject : the treatment has a detectable effect.
- Unpaired, treating the two sets as independent samples of size with : the standard error would be , giving on degrees of freedom — nowhere near significance.
- The same data, the same effect, opposite conclusions. The between-subject standard deviation of swamps the treatment effect of , but it is common to both readings on each subject and cancels entirely in the differences, where the residual noise is only .
- Sanity check on which analysis is correct: the observations are not independent across the two groups — they are the same ten people — so the unpaired analysis violates its own hypotheses. Pairing here is not merely more powerful, it is the only valid option.
Example 4.25 (A two-sample t-test). Two independent groups give , , and , , . Test at , with .
Solution.
- Pooled variance: , so .
- Standard error: .
- Statistic: on degrees of freedom.
- Since , reject at the level.
- Sanity check via duality: the interval for computed in the previous chapter was , which excludes — the same verdict, as the theorem Duality of tests and confidence intervals requires ✓.
- What the interval adds: the test says "not zero", while the interval says the difference is somewhere between half a point and ten points. Reporting only would conceal how imprecisely the effect is pinned down.
4.6Chi-square tests for counts
The tests so far concerned means. For categorical data the natural comparison is between observed counts and the counts a hypothesis predicts.
Theorem 4.26 (Chi-square goodness-of-fit test). Let observations fall into categories with observed counts , and let specify probabilities giving expected counts . Then under , as ,
where is the number of parameters estimated from the data. Large values of are evidence against .
Proof. We give the argument for , where it is exact and transparent, and quote the general case. With two categories, and . Writing , and , note , so
The bracket is exactly the standardised binomial count, which by the theorem Normal approximation to the sample proportion is asymptotically standard normal. The square of a standard normal is , and ✓.
For general the same idea applies to the vector of standardised counts, which is asymptotically multivariate normal; the single linear constraint removes one dimension, and each estimated parameter removes one more, leaving .∎
The degrees of freedom deserve the same reading as in the -test: they count the number of freely varying comparisons. The counts must sum to , so once of the deviations are known the last is determined — one constraint, one degree of freedom lost — and every parameter fitted from the same data costs another.
Proposition 4.27 (Chi-square test for independence). For an contingency table, the hypothesis that the row and column variables are independent gives expected counts
and is asymptotically with degrees of freedom.
Proof. Under independence , and the maximum likelihood estimates of the marginals are the observed row and column proportions, giving , which is the stated formula. For the degrees of freedom, the table has cells, hence free probabilities; estimating the marginals costs parameters. So
Pitfall. The chi-square distribution here is an approximation that requires the expected counts to be reasonably large — the usual working rule is every . With small expected counts the statistic is discrete and skewed while is continuous, and the test rejects far too often. Use an exact test (Fisher's, for a table) instead of pooling categories until the rule is satisfied, since pooling changes the hypothesis being tested.
Note also that the requirement is on the expected counts, not the observed ones. A cell with zero observations is not a problem; a cell with expected count is.
Example 4.28 (A goodness-of-fit test). A die is rolled times with outcomes through observed times. Test fairness at , given .
Solution.
- Under each face has probability , so every expected count is , comfortably above .
- The deviations are , and their squares are , summing to .
- Since every is the same, .
- Degrees of freedom: , no parameters having been estimated. Since , do not reject: the data are consistent with a fair die.
- Sanity check on the scale: under the mean of a is its degrees of freedom, . An observed sits slightly below what a fair die would typically produce, so there is not the faintest suggestion of bias — the -value is about .
- Note the correct conclusion: the die is not shown to be fair, only not shown to be unfair. With rolls the test has little power against a mild bias.
Example 4.29 (A test for independence). A table records treatment against outcome: of treated, improved; of untreated, improved. Test independence at , given .
Solution.
- The observed table has row totals and , column totals improved and not, and .
- Expected counts under independence, by the proposition Chi-square test for independence: , , , . All exceed ✓.
- The deviations all have the same magnitude, , as they must in a table where the margins are fixed. So
- The bracket is , giving .
- Degrees of freedom: . Since , reject — but only just, with .
- Sanity check on how marginal this is: a single patient moving between cells would flip the verdict. Reporting this as "significant, " without noting its fragility would be technically true and substantively misleading — which is the case for reporting the improvement rates, against , with an interval around their difference.
4.7Many tests at once
Every result so far concerned a single test. The moment several are run, the guarantee that provides — a chance of a false rejection — applies to each test separately and not to the collection, and the difference is large.
Proposition 4.30 (Family-wise error rate under multiplicity). If independent tests are each conducted at level and all nulls are true, the probability of at least one false rejection is
which tends to as grows.
Proof. Each test independently fails to reject with probability , so all fail to reject with probability by independence. The complement is the stated probability, and since we have .∎
At this is for , for , and for . Twenty tests on pure noise produce at least one "significant" result about two times in three.
Theorem 4.31 (Bonferroni correction). Testing each of hypotheses at level makes the probability of one or more false rejections at most , whatever the dependence between the tests.
Proof. Let be the event of a false rejection on test , so for each true null. By the union bound — countable subadditivity of probability, requiring no independence whatever —
The proof's indifference to dependence is both its strength and its weakness: it is always valid, and when the tests are strongly correlated it is very conservative, costing power for a guarantee that was not needed. Procedures controlling the false discovery rate — the expected proportion of rejections that are false, rather than the probability of any — are the usual modern alternative when is large.
Theorem 4.32 (Post-test probability of a true effect). Suppose a fraction of the hypotheses tested in some field are genuinely non-null, and tests have level and power . Then the probability that a rejected null is genuinely false is
Proof. By Bayes' theorem, with the two ways of obtaining a rejection in the denominator:
and substituting , , gives the formula.∎
Example 4.33 (Why a significant result is often wrong). In an exploratory field, suppose of tested hypotheses are genuinely true effects, tests run at , and typical power is . What fraction of significant findings are real? What if power is only ?
Solution.
- With and power :
So of significant findings are false positives, despite every test being run correctly at the level. 2. With power instead:
so under of significant findings are real — the majority are false. 3. Sanity check on the mechanism: controls the false positive rate among true nulls, and when true nulls are of everything tested, that small rate applies to a large pool. Low power shrinks the true-positive pool at the same time, and the ratio collapses. 4. The lesson is not that testing is broken but that is not a guarantee, and its meaning depends on two quantities the -value does not contain: the prior plausibility of what is being tested, and the power to detect it.□
Example 4.34 (Correcting for a family of tests). A study measures outcome variables and finds raw -values of , , , , , , and . Which survive a Bonferroni correction at , and what would the uncorrected count have been?
Solution.
- Uncorrected, five of the eight are below and would be reported as significant.
- The Bonferroni threshold is .
- Only the first, , falls below it. The other four "significant" findings do not survive.
- Sanity check on whether the correction is too harsh here: under eight true nulls, the chance of at least one raw below is , so seeing a couple of small values among eight is unremarkable. Getting one below is not.
- The conservative side of Bonferroni: if the eight outcomes are strongly correlated — different scales measuring the same construct, say — then the effective number of independent tests is fewer than eight, and the correction overshoots. A method controlling the false discovery rate would retain more of them, at the cost of a weaker guarantee.
- What must not be done is to report the five uncorrected results and mention the correction in a footnote. The reported error rate has to match the procedure actually used, including the size of the search.
Pitfall. The multiplicity that matters is the number of tests you could have reported, not the number you did. Trying four outcome measures and reporting the one that reached is a family of four tests, and the honest level is . The same applies to trying several subgroups, several cut-points for a continuous variable, or several stopping times for data collection.
This is why pre-registration exists: the correction depends on the size of the search, and only a record made in advance establishes that. A -value computed after an undocumented search has no defined error rate at all.
- **Reading as the probability that is true.** It is computed *assuming* ; the reverse conditional needs a prior and comes out very differently, as the post-test probability calculation shows.
- Accepting the null. Failing to reject means the evidence was insufficient. To argue a parameter is near , report a narrow confidence interval, not a large -value.
- Choosing the direction of a one-sided test after seeing the data. This doubles the true Type I error rate while reporting the nominal one.
- Confusing statistical with practical significance. With large enough, any deviation from , however trivial, becomes significant. The interval reports the size of the effect; the test does not.
- Reporting a test where an interval was available. By duality the interval determines the test's outcome against every null value, and adds the effect size and its precision.
- Ignoring multiplicity. Twenty independent tests at on pure noise give at least one significant result about of the time.
- Treating an underpowered non-significant result as a null finding, or an underpowered significant one as an accurate effect size. Both are wrong, and the second is wrong in a consistently inflationary direction.
- Using chi-square with small expected counts. The approximation requires expected counts of about or more; the condition is on expected counts, not observed ones.
- Running an unpaired test on paired data. It violates independence and usually destroys the power that pairing was designed to provide.