Contents / Statistics / Statistical Inference: Estimation
Chapter 3
Statistical Inference: Estimation
Bias, variance and mean squared error; consistency; the method of moments and maximum likelihood with its invariance property; Fisher information and the Cramer-Rao bound; confidence intervals by the pivotal method for a mean, a variance, a difference of means and a proportion, including the Wilson interval; and sample size.
Introduction
The previous chapter established what a sample mean does across repeated samples: it is unbiased for , its standard error is , and for large it is approximately normal whatever the population looks like. Those are facts about one particular statistic. This chapter asks the general questions behind them.
There are two, and they occupy the two halves of the chapter. First, where do estimators come from, and how do we compare them? Nothing so far explains why rather than the median, or the midrange, or the first observation — all of which are unbiased for a symmetric population. Comparing them needs a criterion, and it turns out the natural one, mean squared error, splits cleanly into bias and variance, which is why an estimator that is deliberately wrong on average can still be the better choice. Then two general recipes — the method of moments and maximum likelihood — manufacture estimators for any model at all, and the Cramér–Rao bound says how good an unbiased estimator is even permitted to be.
Second, how do we report uncertainty rather than hide it? A single number is a claim no data can support: the probability that lands exactly on is zero for any continuous population. A confidence interval reports the estimate together with its precision, and it is constructed by a mechanical device — a pivot — that we set up once and then reuse for means, for means with unknown , and for proportions. The proportion case is the one where the textbook formula is genuinely defective, and we show why and what to use instead.
What is not here: hypothesis tests, -values, and the two kinds of error. Those belong to the next chapter, and separating them is deliberate — an interval answers "what values are plausible?", a test answers "is this specific value refuted?", and treating the second as the primary question is the habit that makes statistics feel like ritual rather than reasoning.
Notation follows the rest of the subject: is an unknown parameter, an estimator of it, and the sample mean and variance, the standard normal distribution function, and the value with . Everything about the sampling distribution of — the theorem Mean and variance of the sample mean, the definition Standard error of the mean, the Central Limit Theorem (Lindeberg–Lévy) — is proved in Sampling and Data Distributions and used here without reproof.
3.1What makes an estimator good: bias, variance, and error
An estimator is a random variable, because it is a function of a random sample. Judging one therefore means judging a whole distribution, not a number, and there are several honest ways to do it that do not always agree.
Definition 3.1 (Estimator, bias, and mean squared error). Let be an estimator of a parameter . Its bias and mean squared error are
The estimator is unbiased if for every value of .
Unbiasedness says the estimator is correct on average over repeated samples. It says nothing about any particular sample, and on its own it is a weak requirement — an estimator can be unbiased and useless, as the example below shows. Mean squared error is the more honest single number, because it penalises an estimator for being far from regardless of the direction, and it decomposes in a way that explains exactly what an estimator can trade.
Theorem 3.2 (Bias-variance decomposition). For any estimator with finite second moment,
Proof. Write and insert it:
The first term is by definition. The middle term vanishes because , the quantity being a constant that comes out of the expectation. The last term is .∎
The decomposition is worth reading as a statement about what you are allowed to buy. Total error is the sum of two non-negative pieces, so an estimator can be improved by shrinking either one — and since the pieces move in opposite directions under most modifications, the question is never "is it unbiased?" but "is the bias worth the variance it bought?".
Example 3.3 (An unbiased estimator that should never be used). A population has mean and variance , and a sample of size is available. Compare two estimators of : , the first observation alone, and .
Solution.
- Both are unbiased: and by the theorem Mean and variance of the sample mean.
- Their variances differ by a factor of : while .
- By the theorem Bias-variance decomposition, with zero bias for both, and . The sample mean is a hundred times better on this criterion.
- The moral is that unbiasedness alone ranks nothing: throws away of the data and remains perfectly unbiased. Bias controls where the estimator points on average; variance controls whether any single estimate is close. Only the two together are informative.
Example 3.4 (A biased estimator that beats every unbiased one). Estimate from a sample of size by shrinking the sample mean toward zero: for a constant . Find the minimising MSE.
Solution.
- , so the bias is , and .
- By the decomposition,
- Differentiate and set to zero: , giving
- Since strictly whenever , the best estimator in this family is biased, and it beats (the case ) on MSE for every .
- Sanity check on the two extremes: as the variance term dies and , recovering — with enough data there is nothing to gain by shrinking. And if is near zero relative to , then and the best estimate is simply , which is right: when the signal is far smaller than the noise, the data are not worth listening to.
- The catch, and it is a real one: depends on the unknown , so this estimator cannot actually be used as stated. It establishes that unbiasedness is not optimal, not that a usable better estimator is in hand. Making the idea practical is the business of shrinkage estimation.
Pitfall. "Unbiased" is not a synonym for "good", and "biased" is not a synonym for "wrong". The maximum likelihood estimator of a normal variance, derived later in this chapter, is biased downward — and it still has smaller MSE than the unbiased for every and every normal population. Reporting an estimator as acceptable because it is unbiased, without looking at its variance, is the most common way to choose a worse estimator on purpose.
A second criterion asks not about a fixed but about what happens as data accumulate.
Definition 3.5 (Consistency). An estimator based on a sample of size is consistent for if in probability: for every ,
Theorem 3.6 (MSE consistency). If as — equivalently, if both the bias and the variance tend to zero — then is consistent.
Proof. Apply Chebyshev's inequality in the form proved in Sampling and Data Distributions, to the non-negative variable at threshold :
The right-hand side tends to by hypothesis, and the equivalence with "bias and variance both vanish" is the theorem Bias-variance decomposition, both terms being non-negative.∎
Corollary 3.7 (The sample mean is consistent). For a population with finite variance, is a consistent estimator of .
Proof. is unbiased, so its bias is , and its variance is . Apply the theorem MSE consistency.∎
That corollary is the weak law of large numbers arriving by a different road, and the agreement is not a coincidence — consistency is the weak law, stated about an estimator rather than about a running average.
Intuition. Bias, variance and consistency answer three different questions, and a good estimator needs all three answers to be satisfactory.
Picture a rifle sighted on a target. Bias is whether the shots centre on the bullseye or systematically off to the left. Variance is how tightly they group, wherever they are centred. Consistency is whether the group tightens onto the bullseye as you fire more rounds. A rifle can be perfectly sighted and hopelessly inaccurate (unbiased, huge variance), or slightly off-centre and very tight — and on a scoring criterion that counts total distance from the centre, the second rifle wins.
3.2Two recipes: the method of moments and maximum likelihood
The estimators so far arrived by common sense. For a model with an unfamiliar parameter, common sense runs out, and a general procedure is needed. There are two, and they are worth having both because they fail in different places.
The older one simply matches sample moments to population moments.
Definition 3.8 (Method of moments estimator). Let be i.i.d. from a distribution with parameters , and write for the population moments. The method of moments estimator solves the equations
equating each population moment to its sample counterpart.
The justification is consistency: each sample moment converges to the corresponding population moment by the law of large numbers, so solving the matched equations returns something converging to the truth. The method is easy, requires no calculus, and can produce estimates that are frankly absurd.
Example 3.9 (The method of moments, working and failing). (a) For a Poisson population with parameter , find the method of moments estimator. (b) For a Uniform population, find it, and evaluate it on the sample .
Solution.
- (a) A Poisson has , so the single equation is . Sensible, and it agrees with maximum likelihood below.
- (b) A Uniform has , so the equation gives .
- On the sample : , so .
- But the sample contains the value , and a Uniform population cannot produce a . The estimate is not merely imprecise — it is logically impossible, assigning probability zero to data that were actually observed.
- Sanity check on why this happens: the method uses only the sample mean and ignores the sample maximum, which for this model carries almost all the information about . A recipe that discards the informative statistic can be expected to misbehave, and here it does so visibly.
The second recipe never makes that mistake, because it is built from the probability of the observed data directly.
Definition 3.10 (Likelihood, log-likelihood, and the MLE). Given observed data from a model with density or mass function , the likelihood and log-likelihood are
The maximum likelihood estimator is the value of maximising over the parameter space.
Two remarks on reading the definition. The likelihood is not a probability distribution over — it does not integrate to in , and treating it as a distribution is the step that turns frequentist inference into Bayesian inference, which the last chapter of this subject takes deliberately. And maximising rather than is legitimate because is strictly increasing, so the two have their maxima at the same ; it converts a product into a sum, which is the difference between a tractable derivative and an intractable one.
Method 3.11 (Finding a maximum likelihood estimate).
- Write the likelihood and take logs.
- Differentiate: , the score, and solve .
- Confirm a maximum, by checking or by inspecting the shape of .
- Check the boundary. If has no interior stationary point, or if the parameter space is restricted by the data, the maximum may sit at an endpoint, where the derivative is not zero.
Step 4 is the one omitted by every treatment that presents maximum likelihood as "differentiate and solve", and the uniform example below exists to show what it catches.
Example 3.12 (Maximum likelihood for a Bernoulli proportion). Let be i.i.d. Bernoulli with successes. Find .
Solution.
- The mass function is , so
- Differentiate: .
- Set to zero and cross-multiply: , so and .
- Confirm a maximum: everywhere in , so is strictly concave and the stationary point is the global maximum.
- Sanity check: the estimate is the observed success fraction, which is what anyone would have guessed — but it is now derived rather than assumed, and the same machinery will work where guessing fails.
Example 3.13 (Maximum likelihood for a normal population). Let be i.i.d. with both parameters unknown. Find the MLEs.
Solution.
- The log-likelihood is
- Differentiating in : , which vanishes when , that is at . This is the theorem The mean minimises total squared deviation of the first chapter arriving from a different direction: maximising the normal likelihood is minimising a sum of squares.
- Differentiating in : . Setting this to zero at gives
- Note the divisor: , not . By the theorem Unbiasedness of the sample variance the unbiased estimator uses , so the MLE is biased downward, with .
- Sanity check on the direction of the bias: measures spread about rather than about the unknown , and is by construction the point closest to the data, so deviations about it are systematically too small. Maximum likelihood does not correct for that, and it does not claim to — it optimises fit, not unbiasedness.
Example 3.14 (Maximum likelihood where the derivative is useless). Let be i.i.d. Uniform. Find , and compare with the method of moments estimate on the sample .
Solution.
- The density is for and otherwise. The likelihood is therefore
because a single observation above makes one factor zero and kills the whole product. 2. On its support, is strictly decreasing, so is never zero. There is no interior stationary point, and "differentiate and solve" returns nothing. 3. The maximum is at the left endpoint of the allowed range: is as large as possible when is as small as it is permitted to be, namely
- On this gives , against the method of moments' impossible . Maximum likelihood cannot produce an impossible estimate, because any contradicting the data has likelihood exactly zero.
- The MLE here is biased downward — it can never exceed , so strictly. In fact , so is unbiased. Sanity check: that correction inflates the estimate slightly, which is right, since the largest of draws systematically falls short of the ceiling.
Theorem 3.15 (Invariance of maximum likelihood). If is a maximum likelihood estimator of and is any function, then is a maximum likelihood estimator of .
Proof. Take one-to-one first. Reparametrise by , so that and the likelihood in the new parameter is . As ranges over the image, ranges over the original parameter space, so the two functions take exactly the same set of values and attain their maxima at corresponding points: is maximised at precisely because is maximised at .
For general , define the induced likelihood , the standard convention. Then
since every lies in exactly one of the sets being maximised over. The outer maximum is therefore attained at .∎
Invariance is a genuine convenience and has no analogue for unbiasedness — if is unbiased for , then is in general biased for , by Jensen's inequality whenever is convex or concave. So the MLE of is the square root of the MLE of , whereas the square root of the unbiased is not unbiased for .
Example 3.16 (Invariance in use). For a Bernoulli sample with , estimate the odds and the log-odds.
Solution.
- By the theorem Invariance of maximum likelihood, the MLE of any function of is that function of ; no new derivation is needed.
- Odds: .
- Log-odds: .
- Sanity check: makes the odds less than and the log-odds negative ✓. Note that — the transformation is convex, so the odds estimator is biased upward by Jensen even though is unbiased. Invariance is a statement about maximisation, not about expectation.
3.3Efficiency, information, and the Cramér-Rao bound
Among unbiased estimators, smaller variance is better, and the natural question is how far that can be pushed. Remarkably, there is a floor — determined by the model alone, before any estimator is proposed — and an estimator attaining it cannot be improved upon.
Definition 3.17 (Score and Fisher information). For a model differentiable in , the score of a single observation is , and the Fisher information is its variance,
Lemma 3.18 (The score has mean zero, and information as curvature). Under regularity conditions permitting differentiation under the integral sign,
Proof. For the first claim, differentiate the identity with respect to and exchange the order:
using .
For the second, differentiate that same identity once more. Since , taking expectations gives
the first integral vanishing because it is the second derivative of the constant .∎
The second form is the one to hold onto: information is the expected curvature of the log-likelihood. A sharply peaked log-likelihood has large curvature, pins down tightly, and carries a lot of information; a flat one is consistent with a wide range of and carries little. Information is a property of the model and the true parameter, not of the data in hand.
Theorem 3.19 (Cramer-Rao lower bound). Let be i.i.d. from satisfying the regularity conditions, and let be any unbiased estimator of . Then
Proof. Write the total score . By the lemma The score has mean zero, and information as curvature applied to each term, , and by independence .
Now compute . Since , it equals , and writing the joint density as and using ,
where the third equality uses unbiasedness, .
Apply the Cauchy-Schwarz inequality to the covariance:
and dividing by gives the bound.∎
Definition 3.20 (Efficiency). An unbiased estimator attaining the Cramer-Rao bound is called efficient. The efficiency of an unbiased is the ratio of the bound to its actual variance, a number in .
Proposition 3.21 (The sample mean is efficient for a normal population). For with known, , and attains the Cramer-Rao bound.
Proof. For a single observation, , so and . By the lemma, .
The bound is therefore , which is exactly by the theorem Mean and variance of the sample mean.∎
That proposition is the precise sense in which the sample mean is not merely reasonable but optimal: for a normal population no unbiased estimator of whatsoever — not the median, not a trimmed mean, not anything yet to be invented — can have smaller variance. It also shows the limits of the claim, since the hypothesis of normality is doing real work. For a heavy-tailed population the median can beat the mean handily, and the Cramér–Rao bound for that model is a different number.
Remark (Why maximum likelihood is the default). Under regularity conditions, the MLE is consistent and asymptotically efficient: as grows,
so its variance approaches the Cramér–Rao bound. The proof needs a Taylor expansion of the score together with the law of large numbers and the central limit theorem, and we quote the result rather than prove it.
This is the whole case for maximum likelihood as a default method. It is not that the MLE is unbiased — it usually is not, as the normal variance and the uniform ceiling both showed. It is that for large samples it is approximately unbiased and approximately as precise as any unbiased estimator is permitted to be, and it can be computed for any model you can write down.
3.4Confidence intervals and the pivotal method
A point estimate answers the wrong question. For a continuous population , so the estimate is certainly wrong, and reporting it alone conceals whether it is wrong by a little or by a lot. An interval estimate reports the precision along with the estimate, and the machinery for producing one is entirely mechanical once the right object is identified.
Definition 3.23 (Pivotal quantity). A pivotal quantity is a function of the data and the parameter whose distribution is completely known — it does not depend on or on any other unknown.
The definition is doing something subtle. is not pivotal, because its distribution involves the unknown . But , measured in units of its own standard error, has a distribution that is the same whatever happens to be, and that is what makes it usable: we can make a probability statement about it before knowing , then rearrange.
Method 3.24 (The pivotal method).
- Find a pivot containing the parameter of interest.
- Find constants with , using the known distribution of .
- Rearrange the inequality into the form , where and involve only the data.
- Report .
Theorem 3.25 (Confidence interval for a mean, known variance). Let be i.i.d. with known. Then
is a confidence interval for : the probability that this random interval contains is exactly .
Proof. By the proposition A normal population gives an exactly normal mean, , so
a distribution free of and — hence a pivot. By the definition of ,
Now rearrange the event, which is a chain of algebraic equivalences and so leaves the probability unchanged. Multiplying through by ,
then subtracting and multiplying by (which reverses both inequalities, exchanging the roles of the two bounds),
The event has not changed, so its probability is still .∎
For a non-normal population the same interval is approximately valid for large , because the Central Limit Theorem (Lindeberg–Lévy) makes approximately standard normal; the coverage is then only in the limit.
Remark (What the confidence level is a statement about). In the probability statement above, is a fixed constant and is the random quantity. The randomness lives entirely in the endpoints: it is the interval that varies from sample to sample, not the parameter.
So " confidence" is a property of the procedure. Repeat the sampling and the construction many times, and about of the intervals produced will contain . Once you have computed from actual data, that particular interval either contains or does not; there is no remaining randomness and therefore no probability to assign. Saying "there is a probability that lies in " treats as random, which is a Bayesian statement requiring a prior — legitimate, but a different framework, and the subject of Bayesian Inference.
Pitfall. A confidence interval quantifies sampling error only — the variability from having observations rather than the whole population. It makes no allowance for a biased sampling design, a leading question, non-response, or a miscalibrated instrument. A interval computed from a badly drawn sample is a precise statement about the wrong population, and increasing narrows it without making it any less wrong. The four failures catalogued in the definition Four failures of a sampling design are not repaired by any amount of data.
Example 3.26 (A z-interval, and what changing the level costs). A machine fills bottles with ml known from long experience. A sample of bottles has ml. Give , and intervals for .
Solution.
- The standard error is ml.
- The three critical values are , , , giving margins , and ml.
- The intervals are , and .
- Note what higher confidence costs: going from to widens the interval by . Confidence and precision are traded against each other at a fixed , and the only way to buy both is more data.
- Sanity check: only the interval contains the nominal ml. So the data are mildly inconsistent with the machine being correctly calibrated — at the level is excluded, at the level it is not. That is the same information a hypothesis test would deliver, which is no accident and is taken up in the next chapter.
3.5Unknown variance: the t-distribution
The -interval assumed known, which it essentially never is. Substituting for seems harmless and is not: it introduces a second source of variability, because itself changes from sample to sample. Ignoring that gives intervals that are systematically too narrow, and the correction is a different distribution.
Definition 3.27 (Student's t-distribution). If and are independent, then
has the -distribution with degrees of freedom. It is symmetric about , bell-shaped, and has heavier tails than the standard normal, converging to it as .
Theorem 3.28 (The studentised mean is pivotal). Let be i.i.d. . Then
a distribution depending on neither nor .
Proof. Two facts about sampling from a normal population are needed, and we quote them: that , and that and are independent. (The independence is special to the normal distribution and is what makes the construction work; proving it requires an orthogonal change of variables beyond this chapter.)
Given those, write
obtained by dividing numerator and denominator by . Here is standard normal, is , and the two are independent because and are. That is precisely the definition of . Both and have cancelled.∎
Corollary 3.29 (Confidence interval for a mean, unknown variance). Under the same hypotheses, a confidence interval for is
Proof. Apply the recipe The pivotal method to the pivot of the theorem The studentised mean is pivotal, exactly as in the proof of the theorem Confidence interval for a mean, known variance, with in place of .∎
The degrees of freedom are rather than for the same reason the unbiased variance divides by : the residuals satisfy one linear constraint, summing to zero by the proposition Deviations from the mean sum to zero, so only of them are free to vary. One degree of freedom was spent estimating .
Example 3.31 (A t-interval, and the error from using z). A sample of has and . Give a interval for , and compare with what the -interval would have given.
Solution.
- is unknown, so the pivot is with . The standard error estimate is .
- , so the interval is .
- Using instead would give — narrower by .
- That narrowness is not a bonus but an error: the -interval's true coverage here is about , not the claimed, because it ignores the variability of . The correction widens the interval by exactly enough to restore the advertised coverage.
- Sanity check on the size of the correction: at the ratio , so the penalty is under . At it would be , and at only . The correction matters for small samples and is negligible for large ones, which is why the two intervals are often described as interchangeable for .
Remark (How much normality is really needed). The theorem assumed a normal population, and real populations are not normal. In practice the -interval is robust: for moderate its coverage stays close to the nominal level for any population that is not strongly skewed or heavy-tailed, because the central limit theorem is already making near-normal while the correction absorbs the remaining uncertainty in .
The failure case is skew, not mere non-normality. For a strongly right-skewed population, and are positively correlated — a sample that happens to catch a large value inflates both — and the interval's coverage can fall well below the nominal level even at . Symmetry is what the procedure really wants, not normality.
3.6Intervals for a variance, and for a difference of means
The pivotal method is not tied to means. Any quantity with a known distribution once the parameter is absorbed will do, and two more cases come up constantly: the spread of a single population, and the gap between two of them.
Theorem 3.32 (Confidence interval for a normal variance). Let be i.i.d. . A confidence interval for is
where denotes the value with .
Proof. The pivot is , quoted in the proof of the theorem The studentised mean is pivotal; its distribution involves neither nor . By definition of the critical values,
Invert the chain. All three quantities are positive, so taking reciprocals reverses both inequalities:
and multiplying through by gives the stated interval.∎
Pitfall. This interval is not symmetric about , and it cannot be written as "estimate margin". The distribution is right-skewed, so the two critical values are not mirror images and the upper arm is longer than the lower one. Reporting something is wrong here, not merely approximate.
It is also the least robust procedure in this chapter. The -interval survives moderate non-normality because the central limit theorem is working on ; nothing correspondingly helpful acts on , whose distribution depends on the population's fourth moment. For a heavy-tailed population the coverage of this interval can be far below its nominal level at any .
Example 3.33 (An interval for a variance, and for a standard deviation). A sample of gives . Construct a interval for , and for . Use and .
Solution.
- The numerator is .
- Lower endpoint: . Upper endpoint: .
- So the interval for is .
- For , take square roots of both endpoints — legitimate because is increasing, so the event is the same event as and the coverage is unchanged. This gives .
- Sanity check on the asymmetry: sits above the lower endpoint and below the upper one — nearly twice as far. With the data are compatible with a standard deviation anywhere from to almost , a reminder that variances are estimated far less precisely than means at the same sample size.
Now two populations. The parameter of interest is , and the construction is the one-sample argument applied to a difference.
Theorem 3.34 (Confidence interval for a difference of means). Let two independent samples of sizes come from and with a common unknown variance. Define the pooled variance
Then a interval for is
Proof. The two sample means are independent and normal, so by the lemma Linearity of expectation and variance of an independent sum their difference is normal with
Standardising gives a . For the denominator, and are independent variables with and degrees of freedom, and a sum of independent chi-squares is chi-square with the degrees of freedom added, so
independently of the two sample means. Forming as in the definition Student's t-distribution gives a pivot, and cancels. Inverting it by the recipe The pivotal method yields the interval.∎
Remark (When the variances are not equal). The pooling step used the common-variance assumption twice, and it is a real restriction. If the two variances differ, use the Welch interval, which estimates each variance separately,
with degrees of freedom given by the Welch-Satterthwaite approximation
This is generally not an integer, and the resulting distribution is an approximation rather than an exact sampling distribution — which is why the pooled version is stated as a theorem and this one as a remark. Welch is nonetheless the safer default: it is barely less precise when the variances are in fact equal, and markedly more accurate when they are not, especially with unequal sample sizes.
Example 3.35 (Comparing two treatments). Two independent groups give , , and , , . Assume a common variance and construct a interval for , using .
Solution.
- The point estimate is .
- Pooled variance: , so .
- Standard error: .
- Margin: . The interval is .
- Sanity check on the pooled variance: it must lie between and , and does, nearer the larger sample's value as the weighting requires ✓.
- Reading the result: the interval excludes , but only just — the data support a real difference, while leaving its size anywhere from half a point to ten points. Reporting "the treatments differ" without the interval would suggest far more precision than observations can deliver.
3.7Intervals for a proportion, and why the textbook formula fails
For a proportion the obvious pivot is the standardised sample proportion, and following the recipe produces the interval found in most introductory texts. It is also, over much of the parameter space, badly wrong — a rare case where the standard formula is not merely crude but delivers far less coverage than it advertises.
Definition 3.36 (The Wald interval for a proportion). From a sample of size with sample proportion , the Wald interval is
obtained by substituting for in the standard error .
Pitfall. The substitution in that definition is the defect. The quantity is an honest pivot, approximately standard normal by the theorem Normal approximation to the sample proportion. Replacing the inside the square root by produces something that is no longer pivotal, and the error is worst exactly where is a poor estimate of — near and .
The degenerate case makes it vivid. If , the Wald interval is , the single point , claiming confidence in a zero-width interval. Observing no successes in trials is entirely consistent with , and the interval excludes it with certainty.
The repair is to not make the substitution, and instead solve the pivotal inequality honestly.
Theorem 3.37 (The Wilson score interval). Solving for gives the interval with centre and half-width
where . The interval lies inside always and has coverage close to across the whole range of .
Proof. Square both sides of the inequality, which is legitimate as both are non-negative, and clear the denominator:
Expand and collect as a quadratic in :
The leading coefficient is positive, so the solution set is the closed interval between the two roots. By the quadratic formula the roots are
Expand the discriminant: the terms cancel, leaving
whose square root is . Dividing numerator and denominator by gives the stated centre and half-width.∎
Remark (Reading the Wilson interval). Three features are worth noticing, and each fixes a specific Wald failure.
The centre is not but is pulled toward — it is the weighted average of and with weights and , which for a interval means adding about successes and failures. That shrinkage is what keeps the interval inside .
The half-width contains the extra term , which never vanishes. So at the interval still has positive width, unlike Wald's.
And as with fixed, and the whole thing collapses back to the Wald interval. Wald is the large-sample limit of Wilson, which is why it is not wrong so much as prematurely applied.
Example 3.38 (Wald against Wilson at a small count). A safety trial reports adverse events in patients. Give the Wald and Wilson intervals for the event rate .
Solution.
- Wald. , so and the interval is — a claim that the rate is exactly zero, with confidence, from patients.
- Wilson. With , and :
- So the Wilson interval is — the lower endpoint lands exactly at zero, as it must when no events were seen, and the upper endpoint says a rate as high as remains consistent with the data.
- Sanity check by direct computation: if , the chance of seeing zero events in patients is — small but not negligible, which is exactly what an endpoint of a interval should look like. At Wald's endpoint of , the chance of zero events is , which is not the boundary of anything.
- The practical consequence is the "rule of three": with no events in trials, the upper bound is about , here , close to the Wilson answer. Reporting instead would be a serious misstatement of what the trial established.
3.8Choosing the sample size
The interval's half-width is the margin of error, and every term in it is known before any data are collected except the estimate itself. That makes the relation invertible: fix the margin you need and solve for .
Proposition 3.39 (Sample size for a target margin of error). To obtain a interval for a mean with margin of error at most , when is known or can be guessed, take
rounding up to the next integer.
Proof. The margin is . Requiring and solving: , and squaring — both sides being positive — gives the bound. Rounding up is necessary because is decreasing in , so any below the threshold gives too wide an interval.∎
The square in that formula is the whole economics of sample size, and it is the square-root law of the previous chapter read backwards. Halving the margin requires four times the data; cutting it to a tenth requires a hundred times. Precision is bought at quadratic cost, which is why very precise estimates are expensive and why beyond a point it is cheaper to reduce — by measuring better or stratifying, per the proposition Stratification cannot hurt, and usually helps — than to collect more observations.
Proposition 3.40 (Worst-case sample size for a proportion). For a proportion, for all , with equality at . Hence
guarantees a margin of at most whatever the true turns out to be.
Proof. The function has , vanishing at , and , so is the global maximum on with . Substituting the worst case into the margin and solving for as in the previous proposition gives the bound.∎
Example 3.41 (Sizing two studies). (a) A interval for a mean is required with margin at most , and is believed to be . (b) A poll needs a margin of at most percentage points, with no prior information about .
Solution.
- (a) , so take .
- Verify: with , the margin is ✓. With it would be , just over — which is why the rounding must go up.
- (b) , so take .
- Sanity check against practice: national polls quote a margin of error of about points and use samples of roughly a thousand. That is not a coincidence — it is this calculation, and it explains why polling a country of million needs about the same sample as polling a city, since depends on the margin and not on the population size.
- Note what halving the margin to points would cost in (b): , four times as many respondents for one extra digit of precision.
- Treating unbiasedness as the goal. MSE is the criterion that counts both ways of being wrong; a biased estimator with small variance frequently wins, and the maximum likelihood estimator of a variance is a standard example.
- Differentiating the likelihood without checking the boundary. For Uniform the score never vanishes and the MLE sits at . Any model whose support depends on the parameter needs step 4 of the recipe.
- **Saying "there is a 95% probability that is in this interval."** After the data are in, nothing is random: the level describes the procedure across repeated samples. The probability statement requires a prior and belongs to Bayesian inference.
- **Using when was estimated.** The interval comes out too narrow and its true coverage falls short of the advertised level. Use unless is genuinely known.
- **Using the Wald interval for a proportion near or . ** Its coverage collapses, and at it returns a zero-width interval. Use Wilson, which costs two extra terms.
- Thinking a wider interval is a worse result. A wide interval is an honest report of limited information. Narrowing it by lowering the confidence level, or by using a formula that under-covers, hides the uncertainty rather than reducing it.
- Forgetting the quadratic cost of precision. Halving the margin of error takes four times the data, not twice.
- Believing a confidence interval accounts for bad sampling. It quantifies sampling variability alone. Non-response and selection bias are invisible to it, and more data makes the interval narrower without making it more correct.