Contents / Statistics / Sampling and Data Distributions
Chapter 2
Sampling and Data Distributions
Population vs. sample, law of large numbers, central limit theorem, and sampling distributions.
Introduction
Every statistical claim you will ever meet has the same shape. Somebody wanted to know something about a large group they could not examine — all voters, all manufactured bearings, all patients with a disease — so they examined a small group instead, computed a number from it, and then asserted something about the large group. The whole of inferential statistics is the study of when that last step is legitimate and how wrong it is likely to be.
This chapter supplies the machinery that makes the step legitimate. The central realisation is deceptively small: a statistic computed from a random sample is itself a random variable. The sample mean is not a number waiting to be measured; it is a quantity with a distribution, a centre and a spread, and that distribution — the sampling distribution — is the object every confidence interval and every hypothesis test is secretly talking about.
Two theorems dominate. The Law of Large Numbers says the sample mean settles down on the population mean as the sample grows. The Central Limit Theorem says that, once you zoom in on the settling by the right amount, what you see is always a normal curve — whatever the population looked like. The first is why estimation works at all; the second is why one small table of normal probabilities suffices for a whole science.
Everything downstream — estimation, confidence intervals, hypothesis testing, regression — rests on the results proved here. The chapter runs in that order: what a population and a sample are and what makes an estimator honest; how samples are actually drawn and what each design costs; the exact mean and variance of ; the Law of Large Numbers, proved from Chebyshev's inequality, itself proved from Markov's; the Central Limit Theorem, with a proof sketch and honest guidance on when to trust it; and finally the sample proportion, which is the whole apparatus specialised to counting.
2.1Populations, samples, parameters, statistics
Before anything can be estimated, two vocabularies must be kept rigidly apart: the language of the thing you want to know about, and the language of the thing you actually have.
Definition 2.1 (Population and parameter). A population is the entire collection of units under study, together with the measurement taken on each. A parameter is a fixed numerical feature of the population — for instance the population mean
or a population proportion . A parameter is a constant. It does not vary, it is not random, and except in a census it is not known.
Definition 2.2 (Sample and statistic). A sample is a subset of the population that is actually observed, written . A statistic is any function of the sample alone — any quantity you could compute knowing only the data, with no knowledge of the parameters. The sample mean
and the sample variance
are statistics; so is the sample maximum, the sample median, and the proportion of observations exceeding .
The clause with no knowledge of the parameters is what makes the definition do work. The quantity is not a statistic, because you cannot evaluate it: is exactly what you do not have. That single obstruction is the reason has in its denominator rather than , as the section on the sampling distribution proves.
Notation is doing real work here, so it is worth fixing. Roman letters for statistics, Greek for parameters: and estimate and , and estimates . A capital denotes the random variable — the observation before it is drawn — and a lower-case the realised number. So is a random variable, and is one of its values.
Definition 2.3 (Simple random sample). A simple random sample (SRS) of size is a sample drawn so that every subset of units of the population is equally likely to be the one selected.
When the population is large relative to — or when sampling is done with replacement — the resulting observations are independent and identically distributed (i.i.d.): each has the population distribution, and knowing some of them tells you nothing about the others. This i.i.d. condition, not the word "random", is the hypothesis every theorem in this chapter actually uses.
Intuition. The population is the whole jar of marbles; the sample is the handful you scoop out. A parameter is a fact about the jar — fixed, hidden, the same whoever asks. A statistic is what you measure in your handful: visible, but jumpy, different for every scoop.
The reason randomising matters is not fairness. It is that random scooping makes the handful's composition predictable in distribution, which is the only thing that lets you say how far off you probably are. Reaching deliberately for the top of the jar might give a perfectly fine handful — you just have no way of knowing, and no theorem will help you.
A statistic is a random variable
This is the hinge of the subject, and it is worth slowing down for. Draw one sample and compute ; you get a number. Draw a second sample from the same population and compute again; you get a different number. Neither is wrong. The variation is not measurement error, it is sampling variability, and it is a feature of the procedure, not a defect in it.
Because takes different values with different probabilities, it has a distribution of its own, and that distribution is what we can reason about. The population parameter sits fixed; the estimator scatters around it. The two questions worth asking are therefore: does the scatter centre on ? (bias), and how wide is the scatter? (standard error).
Definition 2.4 (Estimator, estimate, bias). An estimator of a parameter is a statistic used to approximate it; the number it takes on a particular sample is an estimate. Its bias is
and is unbiased when , that is, when .
Unbiasedness is a statement about the estimator's long-run centre across all possible samples, and about nothing else. An unbiased estimator can be badly wrong on the sample you happen to have. Conversely a slightly biased estimator with a much smaller variance may be preferable in practice; unbiasedness is a desirable property, not a sacred one.
Pitfall. "Unbiased" does not mean "accurate on this sample", and it does not mean "the sampling method was fair". It is a property of the formula, computed as an average over every sample the design could have produced. A perfectly unbiased formula applied to a badly drawn sample estimates the wrong population's parameter, exactly and reliably.
Example 2.5 (Enumerating a sampling distribution by hand). A population of exam scores is . Samples of size are drawn with replacement, so the two draws are i.i.d. List the sampling distribution of , and check that and .
Solution. The population mean is a parameter:
and the population variance, which divides by because this is the whole population and not a sample, is
With replacement there are equally likely ordered samples. Their means run from (draw twice) to , and the possible values occur with frequencies out of — a little triangle centred on . Averaging all sample means gives exactly , so . Averaging the values of gives , so
Two lessons. First, no individual sample mean has to equal : the sample gives , and it is not defective. Second, the spread of is genuinely smaller than the spread of a single observation — against — which is the entire reason for averaging.□
Example 2.6 (An estimator that is biased, and the fix). On the same population and the same with-replacement samples, compare two candidate estimators of : the "natural" one and the sample variance .
Solution. Averaged over all samples, the divide-by- version comes out at and the divide-by- version at .
The first is biased, and severely: its expectation is , half of , so it understates the population variance by a factor of exactly every time on average. The second hits on the nose.
The reason is visible in the sample , where and the squared deviations are both zero. Deviations are measured from , which is pulled toward the data, and never from , which is not. Deviations from the sample's own centre are systematically too small, and dividing by rather than inflates them back by precisely the right amount. The theorem Unbiasedness of the sample variance below proves that "precisely" is exact, not approximate.□
2.2Sampling designs and the biases they control
The theorems in this chapter assume an i.i.d. sample. Real data collection has to earn that assumption, and the design used to draw the sample decides both whether it holds and how much precision you get for a fixed budget.
Definition 2.7 (Four sampling designs). Let the population have units and let be the sample size.
Simple random sampling. Every subset of size is equally likely. Unbiased and assumption-free; requires a complete list of the population (a sampling frame) and can be expensive when units are geographically scattered.
Stratified sampling. Partition the population into strata that are internally homogeneous — age bands, regions, production lines — and draw an independent SRS within each. Under proportional allocation the stratum sample sizes satisfy .
Cluster sampling. Partition the population into clusters — city blocks, schools, shipping crates — that are each meant to resemble the population in miniature, take an SRS of clusters, and measure every unit in the chosen clusters.
Systematic sampling. Order the population, pick a random start between and , then take every th unit thereafter.
Stratifying and clustering look superficially similar — both begin by cutting the population into groups — but they are opposites in intent and in effect. Stratification wants groups that differ from each other and are uniform within, and it samples from every group; it reduces variance. Clustering wants groups that each look like the whole population, and it samples only a few groups; it reduces cost, and it usually increases variance, because units within a cluster tend to resemble one another and therefore carry duplicate information.
Proposition 2.8 (Stratification cannot hurt, and usually helps). Write for the stratum weights, and for the mean and variance within stratum . Then the population variance decomposes as
and under proportional allocation the stratified estimator has variance
against . Stratification therefore deletes the entire between-strata term from the variance.
Proof. The decomposition is the law of total variance applied to the stratum label: conditioning on which stratum a unit falls in, , and the two terms are exactly the within and between sums.
For the estimator, the strata samples are drawn independently, so by the variance of a sum of independent variables,
Proportional allocation sets , so , and the sum collapses to . Comparing with and using the decomposition, the difference is .∎
The inequality is never strict in the wrong direction, which is the practical content: proportional stratification is a free lunch, and the size of the lunch is the spread between stratum means. Strata that differ wildly buy you a lot; strata that all have the same mean buy you nothing.
Example 2.9 (What stratification is worth). A town of adults splits into non-graduates with mean income (thousand) and standard deviation , and graduates with mean and standard deviation . A sample of is to be drawn. Compare the standard error of the mean under simple random sampling and under proportionally allocated stratified sampling.
Solution. The weights are and , and the overall mean is .
Within-strata variance: . Between-strata variance: . So
Proportional allocation takes and , and
The standard error falls from to — a factor of about — for the same interviews. Achieving that by brute force instead would need times the sample, or interviews. Sanity check: almost all the variation in this population is between the two groups, and stratification removes exactly that part, so a dramatic gain is what the proposition predicts.□
The biases a design can introduce
A design controls variance; it can also create bias that no amount of data will remove.
Definition 2.10 (Four failures of a sampling design). Undercoverage (or frame bias): some part of the population cannot be selected at all, because the sampling frame omits it.
Nonresponse bias: selected units do not respond, and the non-responders differ systematically from the responders.
Voluntary response bias: respondents select themselves, which over-represents whoever feels strongly.
Convenience sampling: units are chosen because they are easy to reach, which over-represents whoever is nearby.
All four shift away from by an amount that does not depend on .
That last sentence is the whole point and deserves to be read twice. Sampling variability shrinks like ; sampling bias does not shrink at all. A biased design run at enormous scale produces a very precise estimate of the wrong quantity, with a reassuringly narrow confidence interval around it.
Pitfall. A large sample is not a defence against a bad design. In 1936 the Literary Digest mailed ten million ballots and tabulated about million responses — a sample roughly a thousand times larger than a modern poll — and predicted a landslide for the candidate who then lost of states. Its frame was drawn from telephone and automobile registries, which in 1936 meant the affluent, and its respondents were self-selected on top of that. The margin of error computed from was about , and it was meaningless.
Method 2.11 (Choosing a design).
- Write down the target population explicitly, then the frame you can actually sample from. The gap between them is your undercoverage, and it is best confronted before the data are collected rather than after.
- If you know a variable that is strongly related to the outcome and available for the whole frame, stratify on it. The gain is the between-strata variance, so stratify on something that separates the means.
- If the cost of reaching a unit dominates the cost of measuring it, consider clustering, and then budget for the variance penalty: clusters whose members resemble each other deliver fewer effective observations than their headcount suggests.
- Use systematic sampling only when the ordering of the frame is unrelated to the measurement. If the list has a period that matches — every seventh day, every fourth machine on a four-machine line — systematic sampling locks onto it and the sample is worthless.
- Plan for nonresponse before it happens. A follow-up protocol on a random subsample of non-responders is worth more than a larger initial mailing.
2.3The sampling distribution of the mean
We can now do the central computation of the chapter: find the exact centre and exact spread of , for any population whatsoever, with no assumption of normality anywhere.
Definition 2.12 (Sampling distribution). The sampling distribution of a statistic is its probability distribution, induced by the randomness in the sampling procedure — that is, the distribution of the values the statistic takes across all samples the design could produce.
Two facts about expectation and variance do all the work, so state them first.
Lemma 2.13 (Linearity of expectation and variance of an independent sum). For any random variables and with finite means and any constants ,
If in addition and are independent with finite variances,
Proof. Linearity of expectation follows from linearity of the sum or integral defining , and needs no independence at all — it holds even for variables that are perfectly dependent.
For the variance, write and and expand:
The last expectation is the covariance . Independence gives , hence , and the cross term vanishes.∎
Remark. Independence is used only for the cross term. The correct hypothesis for the variance formula is therefore the weaker one that the variables be uncorrelated. It matters because the cross term is what a bad design reinstates: cluster sampling produces positively correlated observations, so and the true variance of exceeds . Quoting for clustered data understates the uncertainty.
Theorem 2.14 (Mean and variance of the sample mean). Let be i.i.d. with mean and finite variance , and let . Then
No assumption is made about the shape of the population distribution.
Proof. For the mean, apply linearity of expectation to the sum and pull the constant out:
where the identical-distribution hypothesis supplies for every . Independence was not needed.
For the variance, independence is needed. Extending the lemma from two summands to by induction, the variance of an independent sum is the sum of the variances, so
Scaling by the constant multiplies the variance by :
Look closely at where the comes from, because it is the most-quoted and least-understood constant in statistics. Variance is a squared quantity, and the two things that happen to it — summing independent copies, then dividing by — are not symmetric. Summing multiplies variance by ; dividing by the constant divides variance by . The net is . Taking square roots to get back to the scale of the data turns into .
Definition 2.15 (Standard error of the mean). The standard error of is the standard deviation of its sampling distribution:
When is unknown it is estimated by , giving the estimated standard error .
A standard error measures how far the estimate typically lands from the parameter it estimates. It is not a property of the data — is that — but a property of the procedure. Reporting when you meant , or the reverse, is the single most common numerical error in applied work: describes how spread out individual observations are and does not change when you collect more of them, while describes how spread out your estimate is and shrinks as you do.
Corollary 2.17 (The square-root law). To divide the standard error by , the sample size must be multiplied by . In particular, halving the standard error requires quadrupling the sample.
Proof. , so .∎
Intuition. Precision is expensive and gets more expensive. The first hundred observations buy a great deal; the next hundred buy much less, because they are diluted into a bigger average. This is why a national poll of and a national poll of have margins of error of about and — doubling the budget for a third off the error — and why nobody polls a million people.
Why the sample variance divides by
Theorem 2.18 (Unbiasedness of the sample variance). Let be i.i.d. with mean and finite variance , . Then
Proof. The trick is to compare deviations from with deviations from , which is where the expectations are known. Insert and remove inside the square:
Now , so the middle term equals and the last two terms combine:
This identity is worth remembering on its own: the scatter about the sample mean is always less than the scatter about the true mean, by exactly times the squared error of the sample mean. Take expectations of both sides. Each by definition, and by the theorem Mean and variance of the sample mean — which is where independence enters. Therefore
Dividing by gives .∎
The proof explains the rather than merely confirming it. Sum of squared deviations about has expectation , not , because is itself built from the data and sits closer to them than does. The deficit is exactly one — one observation's worth — which is the meaning of the phrase one degree of freedom is spent estimating the mean. The deviations satisfy one linear constraint, , so only of them are free.
Pitfall. is unbiased for , but is not unbiased for . Taking a square root is a strictly concave operation, so Jensen's inequality gives : the sample standard deviation underestimates on average. The bias is small for large and is almost always ignored, but "unbiasedness is preserved under square roots" is false and the belief causes real errors elsewhere.
Example 2.19 (Computing , and the estimated standard error). A sample of bearing diameters (mm) reads . Report , , , and the estimated standard error of the mean.
Solution. The sample mean is
Deviations are , which sum to zero as they must. Their squares are , totalling . Hence
and the estimated standard error of the mean is
Sanity check: dividing by instead would have given , which the theorem says is biased downward by the factor — and indeed . Note also the scale: individual bearings scatter by about mm, but the mean of five scatters by only about mm.□
Finite populations and exact normality
Two refinements complete the picture.
Proposition 2.20 (Finite population correction). When an SRS of size is drawn without replacement from a finite population of size with variance , the draws are no longer independent and
The factor multiplying the standard error is the finite population correction (FPC).
The correction is always less than : sampling without replacement is more precise, because later draws cannot repeat information already collected. At the extreme the factor is — a census has no sampling error at all. The usual convention is to ignore the FPC when , where it exceeds and changes nothing that matters. On the five-score population above, sampling without replacement gives rather than , which direct enumeration of the ten samples confirms.
Proposition 2.21 (A normal population gives an exactly normal mean). If are i.i.d. , then for every , including ,
Proof. The moment generating function of is . For independent variables the MGF of a sum is the product of the MGFs, and , so
which is the MGF of . Since an MGF finite in a neighbourhood of determines the distribution, the claim follows.∎
This matters because it separates two situations that students routinely merge. If the population is normal, no theorem about large is needed — is normal from . The Central Limit Theorem is required only when the population is not normal, and then it delivers normality only in the limit.
2.4The Law of Large Numbers
We now know that sits at on average and that its spread is , which tends to . Intuitively, an estimator centred on the target whose spread collapses must converge to the target. The Law of Large Numbers is that intuition made into a theorem, and the route to it runs through two inequalities that are worth knowing in their own right.
Theorem 2.22 (Markov's inequality). Let be a non-negative random variable with finite mean. Then for every ,
Proof. Define the indicator , equal to when and otherwise. Then pointwise
because when the right side is , and when the right side is , using non-negativity. Expectation is monotone, so taking expectations preserves the inequality:
Dividing by gives the result.∎
Non-negativity is not decoration. Without it the inequality is false: a variable taking and can have a small mean and still exceed with high probability.
Theorem 2.23 (Chebyshev's inequality). Let have finite mean and finite variance . Then for every ,
Equivalently, with : the probability of landing at least standard deviations from the mean is at most .
Proof. Apply Markov's inequality to the non-negative variable with threshold . The events and are the same event, so
Three lines, and it holds for every distribution with a finite variance: at most of any distribution lies two or more standard deviations from its mean, at most lies three or more. The price of that universality is looseness — for a normal distribution the true figures are and — but a bound that needs no assumptions is exactly what is wanted for proving a theorem that needs no assumptions.
Theorem 2.24 (Weak Law of Large Numbers). Let be i.i.d. with mean and finite variance , and set . Then in probability: for every ,
Proof. Fix . By the theorem Mean and variance of the sample mean, has mean and variance . Apply Chebyshev's inequality to with :
With and fixed, the right-hand side is a constant over , which tends to as . The probability on the left is non-negative and bounded above by something tending to , so it tends to .∎
The proof is three lines because all the work was done earlier: the variance calculation is what makes it go through, and the finite-variance hypothesis is inherited from Chebyshev. (The weak law is in fact true assuming only a finite mean, but that proof needs characteristic functions rather than Chebyshev, and the finite-variance version is the one that illuminates.)
Theorem 2.25 (Strong Law of Large Numbers). Let be i.i.d. with finite mean . Then almost surely:
The difference between the two laws is not a technicality, though it is easy to state badly. The weak law says: for each fixed large , it is unlikely that is far from . It permits the sequence to wander back out beyond infinitely often, as long as it does so ever more rarely. The strong law says: the sequence of running means, viewed as a single infinite object, converges — for all but a probability-zero set of possible infinite sample sequences. Excursions beyond eventually stop happening altogether. The strong law implies the weak law; the converse fails.
Pitfall. The Law of Large Numbers is not the "law of averages", and the difference is where the gambler's fallacy lives. After seven heads, the coin does not owe you tails. Each flip is still independent and still , and the running average returns to by dilution, not by compensation: the surplus of seven heads becomes a smaller and smaller fraction of an ever-longer sequence.
The sharp form of the point is that the count does not converge at all. The number of heads minus has standard deviation , which grows without bound: at you should expect to be off by around heads, and at by around . It is the proportion, , that shrinks. Averages settle; totals drift further apart forever.
Example 2.27 (How many observations does Chebyshev demand?). Observations are i.i.d. with unknown mean and known . How large must be to guarantee using only Chebyshev's inequality — that is, with no assumption about the population's shape?
Solution. By the bound used in the proof of the weak law,
Requiring gives .
Compare what the Central Limit Theorem will give for the same guarantee in the next section: it needs standard errors to cover , so , giving and .
Sanity check on the gap: against , a factor of about five. Both are correct. Chebyshev's is a guarantee valid for every distribution with , including maximally awkward ones; the CLT's is an approximation that relies on already being large enough for normality to have set in. The price of assuming nothing is a bound five times too conservative for well-behaved data.□
2.5The Central Limit Theorem
The Law of Large Numbers tells us collapses onto . That is a statement about a point, and a point carries no uncertainty, so on its own the LLN cannot produce a confidence interval. To say anything quantitative we must look at the collapse under a microscope: magnify the deviation by just enough to keep it visible, and ask what shape it has. Since the deviation has standard deviation , the right magnification is by — and the astonishing fact is that what comes into focus is the same curve every time.
Definition 2.28 (Convergence in distribution). A sequence of random variables converges in distribution to , written , if their cumulative distribution functions converge,
This says the probabilities line up in the limit. It says nothing about the values of tracking the values of , which is why it is the weakest of the standard convergence modes.
Theorem 2.29 (Central Limit Theorem (Lindeberg–Lévy)). Let be independent and identically distributed with mean and variance satisfying . Then
Equivalently, for large ,
No assumption whatsoever is made about the shape of the population distribution beyond the existence of a finite, non-zero variance.
Every word of the hypothesis earns its place. Identically distributed and independent give the mean and variance computed earlier. Finite variance is the one that is genuinely restrictive and genuinely fails in practice: for a Cauchy distribution, which has no finite variance (indeed no mean), has exactly the same Cauchy distribution as a single observation, for every . Averaging Cauchy data achieves literally nothing, and no amount of data repairs it. Heavy-tailed financial returns and city-size distributions live uncomfortably close to this regime, which is why "the CLT always applies eventually" is a dangerous thing to believe.
Note also what the conclusion is about. It is a statement about , or equivalently about the sum. It is not a statement about the individual , which remain exactly as skewed and lumpy as they always were, no matter how many of them you collect.
Proof. (Sketch, via moment generating functions.) Assume the MGF exists in a neighbourhood of ; the full theorem replaces MGFs by characteristic functions, which always exist, and the structure of the argument is unchanged.
Standardise first: let , so the are i.i.d. with and , and
Let be the common MGF of the . Differentiating under the expectation gives , and , so the second-order Taylor expansion about is
Independence turns the MGF of a sum into a product, and scaling by evaluates each factor at :
Now use with :
which is the MGF of . By the continuity theorem for MGFs — pointwise convergence of MGFs in a neighbourhood of implies convergence in distribution — .
Two features of the argument explain the theorem's universality. The expansion of keeps only the terms up to , so only the first two moments survive; every higher moment is pushed into the and washed out by the limit. And the normal is the unique distribution whose MGF is , so the limit has nowhere else to land. This is a sketch: the Taylor remainder needs to be controlled uniformly, and MGFs need not exist, both of which the characteristic-function proof handles properly.∎
The figure is the theorem in one picture, and it repays a slow look. The three curves are exact densities, not simulations: for an exponential population with mean , the sample mean of observations has a Gamma distribution with shape , whose density can be written down in closed form. Their standard deviations are , and , matching the theorem Mean and variance of the sample mean exactly. What the CLT adds is the shape: the curve is as far from normal as a distribution gets, and observations are enough to make its average look like a bell.
Example 2.31 (A probability for a sample mean). A population has and . A simple random sample of is drawn. Find and .
Solution. First the sampling distribution. By the theorem Mean and variance of the sample mean, and , so
With the CLT justifies treating as approximately . Standardise:
For the second, , so
Sanity check: an individual observation exceeds with probability , but a mean of exceeds with probability only . That contrast — the same threshold, wildly different probabilities — is the standard error doing its job, and it is the reason a sample mean is informative about when a single observation is not.□
How large is "large enough"?
Note. The familiar advice is . It is a rule of thumb, not a theorem, and it has no mathematical status whatsoever: the CLT is an asymptotic statement, and the rate at which the approximation improves depends on the population. What governs the rate is skewness, and the Berry–Esseen theorem makes this precise — the error in the normal approximation to the CDF of is bounded by a constant times . The third absolute moment is a measure of asymmetry, and it sits in the numerator.
Practical consequences:
- A symmetric, light-tailed population (uniform, triangular) is well approximated by as small as .
- A moderately skewed population needs something in the region of , which is where the rule of thumb comes from.
- A strongly skewed population — exponential, lognormal, insurance claims, income — may need in the hundreds, and the approximation fails first and worst in the tails, which is exactly where -values live.
- Heavy tails without finite variance are not fixed by any .
Example 2.32 (Testing the rule of thumb honestly). For the exponential population with mean of the figure above, the exact probability can be computed from the Gamma distribution. Compare it with the CLT approximation at several sample sizes.
Solution. The CLT approximation uses and reports .
| CLT approximation | exact | relative error | |
|---|---|---|---|
Read the last column carefully, because it contains a lesson that the rule of thumb hides. The absolute error does fall steadily — from at to at — exactly as the CLT promises. But the relative error stops improving after and then gets steadily worse, because by the probability being approximated is itself tiny and sits far out in the tail, where a skewed distribution departs from normality most stubbornly.
Sanity check and moral: at a relative error on a probability near is perfectly usable for a confidence interval, so the rule of thumb is not wrong. But quoting when the truth is — a claim that is half again more extreme than reality — is how a marginal result becomes a published one. Tail probabilities from skewed data deserve an exact method or a bootstrap, not the CLT.□
Corollary 2.33 (The CLT for sums). Under the hypotheses of the Central Limit Theorem, the sum satisfies
so its standard deviation is .
Proof. , and multiplying by the constant multiplies the mean by and the variance by : and . The standardised quantity is the same in both cases, so the limiting normality transfers.∎
Pitfall. Sums and means scale differently and mixing them up is a standing trap. The standard deviation of the sum grows like ; the standard deviation of the mean shrinks like . Both are the law, pointing in opposite directions. When a problem asks about a total load, a total claim amount or a total waiting time, it is asking about , and the spread gets larger, not smaller.
Intuition. Why should averaging manufacture a bell curve out of anything at all? Because an average is a tug-of-war among many small independent pushes, and no single push can dominate once there are enough of them — finite variance is precisely the condition that stops one observation from dominating. The particular quirks of the population get expressed in the third and higher moments, and the algebra in the proof shows those moments being divided away faster than the first two. All that survives the limit is a centre and a spread, and there is exactly one distribution shape determined by a centre and a spread alone.
2.6The sampling distribution of a proportion
Counting is the special case where the whole apparatus becomes concrete, and it is the case that polling, quality control and clinical trials actually use.
Definition 2.34 (Sample proportion). Let each unit be classified as a success or a failure, with population success probability , and let be the number of successes in independent trials, so . The sample proportion is
The key observation is that is not a new kind of statistic at all. Code each observation as for a success and for a failure; then and . Every theorem proved for the sample mean applies verbatim.
Proposition 2.35 (Mean and standard error of the sample proportion). For an i.i.d. Bernoulli sample,
Proof. A Bernoulli variable has and , so
Since , the theorem Mean and variance of the sample mean applies directly with and , giving and . Taking the square root gives the standard error.∎
Remark. The variance is maximised at , where it equals , and shrinks to at either extreme. Proportions near or are therefore estimated more precisely than proportions near a half, for the same . This is also why the conservative planning value for a poll's margin of error always uses : it is the worst case, so the resulting sample size is safe whatever turns out to be.
Theorem 2.36 (Normal approximation to the sample proportion). If and , then
The two conditions are the CLT's "large enough " made checkable. They are needed because a Bernoulli population is maximally skewed when is near or — at , a sample of contains on average a single success and the distribution of is nothing like a bell. Requiring at least about ten expected successes and ten expected failures forces up far enough that the skew has been averaged away. When either condition fails, the normal approximation should be abandoned in favour of exact binomial calculations.
Note. Continuity correction. The binomial is discrete and the normal is continuous, so when approximating a probability about a count, extend each integer to the interval around it: replace by before standardising. For , the exact value is ; with the correction the normal gives , and without it — an error of . The correction is worth the half-unit of arithmetic whenever is small or the probability is needed accurately.
Example 2.38 (A poll, checked and reported). A poll of randomly selected voters finds in favour. Verify the conditions, then construct a confidence interval for the population proportion .
Solution. The point estimate is .
Check the conditions first. The true is unknown, so use : and . Both pass comfortably, so the normal approximation is justified.
The estimated standard error is
At confidence , so the margin of error is
and the interval is , that is .
Sanity check: a margin of about five points is what a poll of a few hundred delivers, and the interval lies entirely below — though only just, so this poll is weak evidence that the true support is under half. No finite population correction is applied, because voters are a negligible fraction of the electorate.□
Method 2.39 (Sample size for a target margin of error). To achieve a margin of error at most at confidence level for a proportion:
- Find the critical value for the level ( for , for , for ).
- Choose a planning value : a prior estimate if you have one, otherwise the conservative .
- Solve for :
- Round up to an integer, and check that and both clear .
For at with : , so . That single calculation is why so many published polls report a sample size of about and a margin of error of about three points.
Proposition 2.40 (Difference of two independent proportions). Let and come from independent samples of sizes and . Then and
Proof. Apply the lemma Linearity of expectation and variance of an independent sum with , : expectations subtract, and variances are multiplied by and , so they add.∎
Pitfall. Variances add for a difference of independent estimators; standard errors do not. The correct step is
Adding the standard errors overstates the uncertainty — for two equal standard errors it inflates by a factor of — and quietly destroys the power of the comparison. The minus sign in does not subtract variance; it is squared away.
Summary. The load-bearing results of the chapter, each with the hypotheses it actually needs.
- A statistic is a random variable. Its distribution across all samples the design could produce is its sampling distribution; a parameter, by contrast, is a fixed constant.
- Unbiasedness. is a statement about the estimator's average over all possible samples, not about accuracy on one sample. It offers no protection against a biased design, whose error does not shrink with .
- Stratification. Under proportional allocation, , which removes the between-strata variance from and can never be larger.
- Mean and variance of . For i.i.d. with mean and finite variance : (needs only identical distribution) and (needs uncorrelatedness). Hence , and dividing by costs a factor in sample size. Without replacement from a finite population, multiply by the FPC .
- Unbiasedness of . For i.i.d. with finite variance, when the divisor is . The identity is why. Note is not unbiased for .
- Markov. and finite give . Non-negativity is essential.
- Chebyshev. Finite mean and variance give , for every distribution — universal, and correspondingly loose.
- Weak Law of Large Numbers. I.i.d. with finite variance: , so in probability. The strong law upgrades this to almost-sure convergence under only a finite mean. Neither says anything about a count, whose deviation from grows like .
- Central Limit Theorem. I.i.d. with : , whatever the population's shape. Finite variance is indispensable — the Cauchy mean never converges. A normal population gives exact normality at every , no CLT required. The adequacy of the approximation is governed by skewness, so is a rule of thumb and tail probabilities from skewed data are the first thing to fail.
- Sums versus means. : the spread of a sum grows like while the spread of a mean shrinks like .
- Proportions. for data, so and , normal when and . For a difference of independent proportions, variances add.
- **Confusing with . ** The population standard deviation describes the spread of individual observations and does not change when you collect more data; the standard error describes the spread of the *estimate* and shrinks like . Using where belongs inflates every interval by a factor of .
- Believing a large sample fixes a biased design. Sampling variability shrinks like ; sampling bias does not shrink at all. A biased design at scale gives a very precise answer to the wrong question.
- Applying the CLT to individual observations. The theorem is about and . The stay exactly as skewed as the population is, forever.
- **Treating as a theorem.** It is a rule of thumb calibrated on moderate skew. Strongly skewed populations need far more, and the approximation degrades first in the tails, which is precisely where -values are read.
- Forgetting that the CLT needs finite variance. For heavy-tailed data with infinite variance, averaging does not help at any : the mean of Cauchy observations has the same distribution as one observation.
- The gambler's fallacy. The LLN says the *average* settles, by dilution. It does not say past outcomes get compensated, and the *count* of heads drifts ever further from as grows.
- **Dividing by in the sample variance.** That estimator has expectation and understates the variance every time on average. Divide by unless you genuinely have the whole population.
- **Assuming is unbiased for . ** is unbiased for ; the square root is concave, so .
- Adding standard errors instead of adding variances. For independent estimators, combine as — never , and never for dependent samples at all.
- **Using the normal approximation for a proportion when or . ** At extreme the Bernoulli population is too skewed for the sample size; use exact binomial methods.
- **Quoting for clustered or otherwise correlated data.** That formula needs uncorrelated observations. Positive within-cluster correlation makes the true variance larger, so the reported uncertainty is too small.