Contents / Statistics / Descriptive Statistics
Chapter 1
Descriptive Statistics
Measures of center, spread, shape, and data visualization.
Introduction
A dataset arrives as a list of numbers, and a list of numbers is not yet
knowledge. Descriptive statistics is the discipline of compressing that list
into a few quantities and a picture, in such a way that the compression is
honest: what survives should be the features of the data that matter, and what
is discarded should be the part a reader would not have used anyway.
Two questions organise the whole subject. Where is the data? — a question
about location, answered by the mean, the median, the mode. *How spread out is
it?* — a question about scale, answered by the variance, the standard
deviation, the interquartile range. A third question, what shape is it?,
decides which of the first two answers you are allowed to trust.
This chapter takes those summaries seriously as mathematical objects rather than
as recipes. The mean is not merely "add up and divide"; it is the unique number
that minimises total squared distance to the data, which is exactly why it turns
up wherever squares are minimised. The median is the corresponding minimiser for
absolute distance, which is exactly why it ignores outliers. The sample variance
divides by and not by for a reason that can be written down and
proved, and we prove it. By the end you should be able to say, of every formula
here, what it optimises and what it assumes.
1.1Populations, Samples and Notation
Almost every confusion in elementary statistics traces back to one blurred
distinction: the difference between the thing you want to know about and the
thing you actually measured.
Definition 1.1 (Population, sample, parameter, statistic). A population is the entire collection of units under study; a sample is a subset of it that we actually observe.
A parameter is a number computed from the whole population. It is a fixed, usually unknown constant — the population mean , the population variance .
A statistic is a number computed from the sample. It is random: draw a different sample and you get a different value. The sample mean and sample variance are statistics.
The parameter is fixed and unknown; the statistic is known and variable. Every
inferential technique in later chapters is a way of managing that trade. Here we
only need the distinction to keep two formulas apart, but keeping them apart is
worth the effort, because the divisor in the variance depends on which of the
two you are computing.
Notation (Parameters versus statistics). Throughout this chapter and the rest of the subject:
| quantity | population (parameter) | sample (statistic) |
|---|---|---|
| size | ||
| mean | ||
| variance | ||
| standard deviation | ||
| proportion |
Greek letters for parameters, Roman letters (often with a bar or a hat) for statistics. A capital denotes the -th observation regarded as a random variable, before it is observed; a lower-case denotes the number it turned out to be. The distinction matters in a statement like , which is a claim about the random variable and is meaningless for the observed number .
We will use one modelling assumption repeatedly, and it deserves a name.
Definition 1.2 (Random sample). is a random sample (equivalently, the are i.i.d.) from a distribution with mean and variance if the are independent of one another and each has that same distribution, so that and for every .
Notice what this does not assume: nothing about normality, nothing about
symmetry, nothing about the shape of the distribution at all. Only independence
and a common distribution with a finite variance. Most of the theorems in this
chapter need exactly that much and no more, and it is worth watching which ones
need more.
Definition 1.3 (Estimator, bias, unbiasedness). An estimator of a parameter is a statistic used to guess it. Its bias is
and is unbiased when the bias is zero, i.e. .
Intuition. Unbiasedness is a statement about the long run, not about your particular sample. It says: if you could repeat the sampling forever and average all the estimates you got, you would land exactly on the truth. It does not say your one estimate is close. An estimator can be unbiased and wildly noisy, or slightly biased and far more reliable. Unbiasedness is one desirable property among several, not a certificate of quality.
1.2Mean
Definition 1.4 (Population mean and sample mean). The population mean of values is
and the sample mean of observed values is
The two formulas are identical; only the set being averaged differs.
The mean uses every data point, with equal weight. That is its great strength —
no observation is thrown away, and the algebra that follows is clean — and its
one great weakness: a single extreme value gets the same vote as every other, so
it can drag the mean far from anything typical.
Intuition. Picture the values as unit weights placed along a rod. The mean is the point where the rod balances. Slide one weight far to the right and the balance point must move right too, even though every other weight stayed put. That is the whole story of the mean's sensitivity to outliers, and it is literally true: the identity below is the statement that the moments about cancel.
The balance property is the first thing to prove, because everything else in the
section leans on it.
Proposition 1.5 (Deviations from the mean sum to zero). For any values with mean ,
Proof. Split the sum and use that does not depend on :
where the middle step is just the definition rearranged.∎
This is why raw deviations are useless as a measure of spread: they always
cancel, for every dataset, tight or scattered. Any workable measure of spread
must first destroy the signs, by squaring (variance) or by taking absolute
values (mean absolute deviation).
The mean as a minimiser
The following theorem is the real definition of the mean. It explains, in one
line, why least-squares regression produces means, why the variance is defined
around rather than around anything else, and why the mean is the
natural centre whenever the cost of being wrong grows with the square of the
error.
Theorem 1.6 (The mean minimises total squared deviation). For fixed data , define . Then is minimised uniquely at , and
Proof. Insert and expand, writing :
The middle term vanishes by Proposition Deviations from the mean sum to zero — this is exactly where that identity earns its keep. What remains is
The second term is non-negative and equals zero only when , so for every .∎
Remark. The same computation done with calculus is shorter but says less: gives , and confirms a minimum. The algebraic version is better because it produces the exact excess incurred by centring at the wrong place, an identity we use again when proving why the sample variance needs .
Linear transformations
Changing units — Celsius to Fahrenheit, dollars to thousands of dollars — is a
linear transformation . Summaries must transform predictably or
they would be useless.
Proposition 1.7 (Mean under a linear transformation). If for constants and all , then .
Proof. Directly from the definition,
So the mean is equivariant: it moves with a shift and scales with a scale.
This is the property the median also has, and it is what lets us standardise
data later without worrying about units.
Weighted and combined means
Definition 1.8 (Weighted mean). Given weights attached to values , the weighted mean is
The ordinary mean is the case for every . Weights are the right
tool when observations carry different amounts of information — a course grade
where assignments are worth different point totals, a survey where respondents
represent different numbers of people.
Proposition 1.9 (Combined mean of two groups). If group has values with mean and group has values with mean , then the mean of the pooled data is
Proof. The pooled sum is the sum of the two sums, and each group's sum is its size times its mean: and . Dividing the total by the total count gives the claim.∎
The pooled mean is therefore a weighted mean of the group means, with the group
sizes as weights. Averaging the two means directly is correct only when the
groups are the same size.
Pitfall. Averaging averages is one of the most common errors in applied work, and it does not merely lose precision — it can reverse a conclusion. If a small group has a high mean and a large group a low one, the naive average of the two means sits far above the true pooled mean. The same mechanism, applied to rates rather than means, produces Simpson's paradox.
Definition 1.10 (Trimmed mean). A trimmed mean discards the lowest fraction and the highest fraction of the sorted values and averages what remains.
Trimming interpolates between the two extremes: gives the ordinary
mean, and pushing toward gives the median. A trimmed mean
keeps most of the mean's efficiency on clean data while refusing to be dragged
by a handful of extreme values.
The sample mean as an estimator
So far has been a descriptive number. Regarded as an estimator of an
unknown , it has two properties that make it the default choice, and both
have short proofs. Note carefully which hypothesis each one needs.
Theorem 1.11 (The sample mean is unbiased). Let be a random sample from a distribution with mean . Then
Only and linearity of expectation are used; independence is not needed.
Proof. By linearity of expectation, which holds whether or not the are independent,
Theorem 1.12 (Variance of the sample mean). If in addition the are independent, each with variance , then
Proof. For a constant , , and for independent summands the variance of a sum is the sum of the variances. Hence
Here independence is doing genuine work: with correlated observations the cross
terms do not vanish, and a positively
correlated sample has a sample mean far noisier than suggests.
This is why cluster sampling needs its own variance formulas, and why treating
measurements on subjects as independent observations
understates the uncertainty.
Intuition. The is the single most consequential fact in applied statistics. Precision improves with the square root of effort: to halve the noise in an estimate you must quadruple the sample. Every cost–benefit argument about sample size is an argument about this square root.
Example 1.13 (Mean, weighted mean, and units). Five service times, in minutes, are . Find the mean. Then find the mean of the same times recorded in seconds.
Solution.
- Sum: , so minutes.
- Seconds are , a linear transformation with , . By Proposition Mean under a linear transformation, seconds.
- Check: lies between the minimum and maximum , as any mean must, and seconds is minutes. ✓
Example 1.14 (Combined mean). Group has values with mean ; group has values with mean . Find the pooled mean.
Solution.
- Total sum .
- Pooled mean .
- Check: the naive average of the means is , but is twice as large, so the truth is pulled toward 's mean of . Indeed . ✓
Example 1.15 (Using the minimising property). A dataset of values has and . Without the raw data, find .
Solution.
- By Theorem The mean minimises total squared deviation, .
- With : .
- Check: , as it must be, since is the unique minimiser. ✓
1.3Median and Mode
Definition 1.16 (Median). Sort the data as . The median is
When is even any number between the two middle values splits the data
equally, so the median is not unique as a splitting point; taking the midpoint
is a convention chosen to make the median a continuous function of the data.
Intuition. Line everyone up from smallest to largest and ask what the person in the middle has. Doubling the richest person's fortune moves the mean but does not move the middle person at all. That insensitivity is not a defect to be corrected — it is the entire point of the median.
The median has its own minimisation property, exactly parallel to the mean's,
with squares replaced by absolute values. This one theorem explains every
robustness claim made about the median.
Theorem 1.17 (The median minimises total absolute deviation). For fixed data , the function is minimised at any median of the data. If is odd the minimiser is unique and equals ; if is even every in is a minimiser.
Proof. is a sum of convex functions, hence convex and continuous, and it is piecewise linear with kinks only at the data values. So it suffices to examine how changes as moves.
Take not equal to any data value, and let be the number of and the number with , so . Increasing by a small that crosses no data value increases each of the terms with by exactly and decreases each of the terms by exactly :
Thus is strictly decreasing while (more points still to the right) and strictly increasing once . The minimum is therefore attained where the counts on the two sides balance, which is precisely the median.
Concretely, for odd, for every and for every , so the unique minimiser is that middle order statistic. For even, throughout the interval , so has slope zero there and every point of the closed interval is a minimiser.∎
Remark (Why the two theorems explain everything). Squared loss punishes one huge error far more than many small ones, so the minimiser must chase an outlier to keep its squared distance down. Absolute loss charges an outlier in proportion to its distance, and moving the centre toward it costs exactly as much on the other side; the trade is a wash, so the minimiser does not move. Robustness is not a property bolted onto the median — it is a consequence of the loss function it minimises.
Definition 1.18 (Mode). The mode is the most frequently occurring value (for discrete or categorical data) or the location of a peak of the density (for continuous data). A distribution is unimodal, bimodal, or multimodal according to the number of peaks.
The mode is the only one of the three centres that makes sense for purely
categorical data: you cannot sort eye colours, so there is no median, and you
certainly cannot average them.
Pitfall. A bimodal distribution usually signals two mixed subpopulations — two machines, two age groups, two experimental conditions. Reporting a single mean for it describes a value that may occur in neither group. Always look at the shape before quoting a centre; this is the one case where no single number is an honest summary.
Example 1.19 (Median, odd and even counts). Find the median of and of .
Solution.
- Sorted: . With odd the median is the rd value, .
- Sorted: . With even the median is .
- Check: with an even count the median need not be an observed value; here is not in the data, which is fine. ✓
Example 1.20 (Mean versus median under an outlier). Twelve help-desk service times (minutes) are , already sorted. Compute both centres, then recompute after deleting the .
Solution.
- Sum , so .
- With even, the median is .
- Deleting the : the sum falls to over values, so — a shift of minutes. The new median is the th of values, — a shift of .
- Check: one observation out of twelve moved the mean by more than five times as much as it moved the median. ✓
1.4Variance and Standard Deviation
The crudest measure of spread is the range, . It uses two
observations out of , throws away everything between them, and is determined
entirely by the two least typical values in the dataset. It is worth quoting and
never worth relying on.
Since deviations from the mean sum to zero, a useful measure must remove the
signs. Squaring is the choice that keeps the algebra differentiable and connects
to the minimisation theorem above.
Definition 1.21 (Variance and standard deviation). For a population of values with mean ,
For a sample of values with mean ,
The divisor is Bessel's correction; the next section proves why it is there. The standard deviation restores the original units of the data, which the variance — measured in units squared — does not.
Intuition. The standard deviation is roughly "the typical distance from the mean". It is not exactly the average distance (that would be the mean absolute deviation), because squaring weights large deviations more heavily, but it is the right order of magnitude and it is the number the empirical rule and every normal-theory method are written in terms of.
For a random variable the same definition reads
, and it has a computational form that
saves a pass through the data and is used constantly in later chapters.
Theorem 1.22 (Computational formula for the variance). For a random variable with finite second moment and ,
In particular , with equality only if is constant.
Proof. Expand the square inside the expectation and use linearity, remembering that is a constant so and :
Since as the expectation of a non-negative quantity, ; equality forces , which for a non-negative random variable means with probability one.∎
Corollary 1.23 (Computational formula for a sample). For any values with mean ,
Proof. Expand: . Substituting turns the middle term into , so the last two terms combine to . The final form follows from .∎
Pitfall. The computational formula is exact in exact arithmetic and treacherous in floating point. When the mean is large compared with the spread — timestamps, years, sensor readings offset from zero — and are two huge nearly equal numbers, and subtracting them destroys most of the significant digits. It can even return a negative variance. Use it for hand computation and for algebra; use the two-pass or Welford formula in code.
Linear transformations of the variance
Proposition 1.24 (Variance under a linear transformation). If for constants , then
In particular a shift () leaves the spread unchanged, and the same holds for and for .
Proof. By Proposition Mean under a linear transformation, . Hence each deviation transforms as
the cancelling exactly. Squaring gives , and summing and dividing by the same gives . Taking square roots, and noting a standard deviation is non-negative, gives .∎
Remark. The absolute value is not decoration. Multiplying data by reflects and stretches it; the spread triples, so , not . A negative standard deviation would be meaningless.
Comparing spreads across scales
Definition 1.25 (Coefficient of variation). For data with , the coefficient of variation is
usually quoted as a percentage. It is dimensionless, since and carry the same units.
Being dimensionless, the CV compares variability across datasets measured in
different units or on wildly different scales. It requires a meaningful zero and
a positive mean — it is a ratio-scale tool, and quoting a CV for temperatures in
Celsius is nonsense, because shifting to Fahrenheit changes it.
Example 1.26 (Sample variance two ways). For (), find and using the definition, then confirm with the computational formula.
Solution.
- Mean: , so .
- Deviations: ; squares summing to .
- and .
- Computational check: and , so . ✓
- Had this been a whole population we would divide by : .
Example 1.27 (Rescaling a summary). A dataset has and . Give the mean and standard deviation of .
Solution.
- Mean: .
- Standard deviation: , and .
- Check: the shift by affected only the mean, and the sign of affected only the mean too — the spread depends on , not on . ✓
Example 1.28 (Coefficient of variation). Heights have cm with cm; weights have kg with kg. Which is more variable?
Solution.
- Heights: .
- Weights: .
- The raw standard deviations are nearly equal, but relative to their own scales weights vary more than twice as much. Comparing cm with kg directly would have been meaningless. ✓
1.5Why the Sample Variance Divides by
This is the single most-asked question in the chapter, and it has a complete
answer. The claim is not that is more cautious, or that it "corrects for
small samples", or that it makes the answer bigger. The claim is exact:
dividing by makes an unbiased estimator of , and dividing
by does not.
The source of the problem is that is defined as a spread about the
true mean , but we can only measure spread about the sample mean
— and by Theorem The mean minimises total squared deviation,
is the one centre that makes that sum of squares as small as it can
possibly be. Measuring spread about the minimiser systematically understates it.
The next lemma computes the shortfall exactly.
Lemma 1.29 (Expected sum of squares about the sample mean). Let be a random sample from a distribution with mean and finite variance , and let be the sample mean. Then
Independence is required: the proof uses .
Proof. Apply the identity of Theorem The mean minimises total squared deviation with , read in the direction , that is
This is an algebraic identity, true sample by sample. Now take expectations of both sides and use linearity.
For the first term, for each , so it contributes .
For the second, by Theorem The sample mean is unbiased, so by Theorem Variance of the sample mean. That term contributes .
Therefore
Theorem 1.30 (Unbiasedness of the sample variance). Under the same hypotheses — a random sample from a distribution with finite variance —
No assumption of normality is needed.
Proof. Pull the constant out of the expectation and apply Lemma Expected sum of squares about the sample mean:
Corollary 1.31 (The n-divisor estimator is biased downward). Let . Then
Proof. , so by Theorem Unbiasedness of the sample variance. Subtracting gives the bias .∎
So the -divisor version underestimates the variance by a factor , on
average, every time. At it is off by half; at by ; as
the bias vanishes, which is why the distinction stops mattering
for large samples and matters enormously for small ones.
Intuition (Degrees of freedom). The deviations are not free numbers. They satisfy one linear constraint, from Proposition Deviations from the mean sum to zero: knowing any of them determines the last. So the deviation vector carries only independent pieces of information about spread, and the honest thing is to average over of them, not .
This is the same bookkeeping that gives degrees of freedom to a one-sample -test, to a simple regression (two parameters estimated), and in general. Every estimated parameter costs one degree of freedom, and the cost is exactly the shrinkage this section quantifies.
Pitfall. Unbiasedness does not pass through a square root. Since is strictly concave, Jensen's inequality gives whenever is genuinely random, so underestimates on average even though is unbiased for . Nobody corrects for this in practice — the correction factor depends on the distribution — but "we use so the standard deviation is unbiased" is a false statement.
Example 1.32 (Bias of the two divisors). A population has . Samples of size are drawn repeatedly and is computed each time. What value does average to, and what is its bias?
Solution.
- By Corollary The n-divisor estimator is biased downward, .
- Bias , which agrees with .
- Check: the bias is negative, as it must be, since the sum of squares is taken about the minimising centre . ✓
Example 1.33 (Recovering the sum of squares). A sample of has . Find and the value the -divisor formula would have reported for the variance.
Solution.
- and .
- The -divisor version gives .
- Check: , matching the factor . ✓
1.6Percentiles, Quartiles and the IQR
Location and spread can also be described positionally, by where a value falls
in the sorted order rather than by how far it is from a centre. Positional
summaries are the natural partners of the median, and they inherit its
robustness.
Definition 1.34 (Percentile and quantile). The -th percentile of a dataset is a value below which approximately of the observations fall. Equivalently, the -quantile with is a value with at least a fraction of the data at or below it and at least a fraction at or above it.
The quartiles are (the th percentile), (the median), and (the th percentile). The five-number summary is .
The word approximately is not laziness. For finite data no value in general
has exactly below it, and different software resolves the ambiguity
differently — there are nine commonly implemented quantile definitions. They
agree to within one observation and disagree in the last decimal place; this
chapter uses the median-of-halves convention throughout.
Method 1.35 (Quartiles by the median-of-halves rule).
- Sort the data.
- Find the median .
- Form the lower half: all values strictly below the median's position. If is odd, exclude the median value itself.
- Form the upper half the same way.
- is the median of the lower half; is the median of the upper half.
Definition 1.36 (Interquartile range). The interquartile range is : the width of the interval containing the middle half of the data.
The IQR is to the median what the standard deviation is to the mean — a measure
of scale matched to its centre. It is determined entirely by the middle of the
data, so moving the largest observation to infinity does not change it at all,
whereas it sends to infinity too.
The box plot
The five-number summary is exactly what a box plot draws.
Read off the figure what the box plot is good at and what it hides. It shows
centre, spread, skew (the median sits left of centre in the box, and the right
whisker is longer) and outliers, in a form compact enough to stack a dozen
groups side by side. What it cannot show is shape: a bimodal dataset and a
uniform one can produce identical box plots, because five numbers do not
determine a distribution.
Where the 1.5 comes from
The outlier rule looks like an arbitrary constant. It is not.
Definition 1.38 (Tukey's fences). A value is a mild outlier if it lies outside
and an extreme outlier if it lies outside the same interval with replaced by .
Proposition 1.39 (What the 1.5 fence flags in normal data). If the data follow a normal distribution with mean and standard deviation , then and , so and the mild-outlier fences sit at
The proportion of a normal population outside those fences is about .
Proof. For the standard normal, to four places, and the distribution is symmetric about , so . An increasing linear map preserves order, so it carries quantiles to quantiles: the -quantile of is , where is the -quantile of . Hence and the upper fence is at above the mean. The tail probability is .∎
So is calibrated: on clean normal data the rule flags roughly seven points
in a thousand — rare enough that a flag means something, common enough that the
rule is not vacuous. The fence sits at , flagging
about two points in a million.
Pitfall. A flagged point is not an error and must not be deleted on sight. The rule is calibrated for a roughly normal, unimodal distribution; on genuinely skewed data — incomes, waiting times, insurance claims — it flags a large fraction of perfectly valid observations, because the long tail is the distribution rather than a contamination of it. Investigate a flagged point; delete it only if you find an actual reason.
Example 1.40 (Five-number summary and fences). For the twelve service times , give the five-number summary, the IQR, and the mild-outlier fences.
Solution.
- Median: is even, so .
- Lower half has median .
- Upper half has median .
- Five-number summary: and .
- Fences: and . Only falls outside, so it is the sole outlier. It is beyond as well, so it is an extreme outlier.
- Check: holds, and the whisker ends at the largest non-outlier, . ✓
Example 1.41 (Locating a percentile). A sorted list has values. Where does the th percentile fall?
Solution.
- The position index is .
- is a whole number, so exactly values lie below the percentile and it is taken as the midpoint of the th and th sorted values.
- Check: of is , matching the count below. ✓
Example 1.42 (IQR versus standard deviation under contamination). In the service-time data, replacing by changes from about to about . What happens to the IQR?
Solution.
- The quartiles depend only on the ordered positions within each half. Replacing the largest value by a larger one changes no quartile.
- So exactly as before, a change of zero.
- Check: this is the robustness claim made concrete — grew by a factor of more than a hundred, the IQR did not move. ✓
1.7Standard Scores, the Empirical Rule and Chebyshev
A raw value carries units and a scale, so " " means nothing until you know
what is typical. Standardising strips both away.
Definition 1.43 (Standard score). The z-score of a value from a distribution with mean and standard deviation is
the number of standard deviations sits above the mean. For sample data, .
Proposition 1.44 (Standardised data has mean 0 and variance 1). If with , then and .
Proof. Write with and . By Proposition Mean under a linear transformation,
and by Proposition Variance under a linear transformation, .∎
Because standardisation is exactly a linear transformation, everything proved
about linear transformations applies, and the two propositions together say that
standardising puts any dataset on a common footing: mean , spread ,
dimensionless. That is what makes two scores on differently-scaled exams
comparable.
Intuition. A z-score is a universal ruler. "Twelve centimetres above average" means nothing without knowing how much heights vary; " standard deviations above average" means the same thing whether the subject is heights, test scores or house prices, and can be compared directly across all three.
Pitfall. A z-score is comparable across distributions but a z-score does not by itself give a percentile. Converting into "the rd percentile" requires the normal distribution; on a skewed distribution that conversion can be badly wrong. The z-score is always a position in standard-deviation units, and only sometimes a probability statement.
The empirical rule
Proposition 1.45 (The empirical rule (68–95–99.7)). If the data are approximately normal with mean and standard deviation , then approximately
- of values lie in ,
- lie in ,
- lie in .
These are not conventions; they are the values of
at , computed from the normal density. The
hypothesis of approximate normality is doing all the work, and the rule fails
badly without it — for a strongly right-skewed distribution, far more than
of the mass can sit above .
Example 1.47 (Applying the empirical rule). Test scores are roughly normal with and . What fraction of scores lie in ? Above ?
Solution.
- , so about .
- . About lies inside , leaving about split between the two tails; by symmetry about lies above .
- Check: the exact normal values are and , so the rule's rounding is what causes the small discrepancy. ✓
Chebyshev's inequality
When normality is not available, a weaker but universal bound is.
Theorem 1.48 (Chebyshev's inequality). Let be any random variable with finite mean and finite variance . For every ,
No assumption whatsoever is made about the shape of the distribution.
Proof. First, Markov's inequality: if and , then . To see it, note that the random variable is non-negative — when it equals , and otherwise it equals . Taking expectations, , which rearranges to Markov's inequality.
Now apply it to , which is non-negative with , and . Since holds exactly when ,
The bound is vacuous for (a probability is always at most ) and is
deliberately loose for — it must hold for every distribution with the
stated variance, including the worst one. That worst case is attainable: the
variable taking values with probability each and
otherwise has variance and achieves equality, so Chebyshev's
inequality cannot be improved without further assumptions.
Example 1.49 (Chebyshev versus the empirical rule). A distribution has and , with unknown shape. What can be guaranteed about the interval ? What would the empirical rule claim?
Solution.
- , so and Chebyshev guarantees at least .
- If the distribution were known to be normal, the empirical rule would give about .
- Check: , as required — the universal guarantee must be weaker than the one that gets to assume a shape. ✓
Example 1.50 (Solving for k). For any distribution, within how many standard deviations of the mean must at least of the values lie?
Solution.
- Set , i.e. .
- So and .
- Check: at the bound is exactly. ✓
1.8Shape and Visualization
Numerical summaries answer questions you thought to ask. A picture answers the
ones you did not, which is why it should always come first.
Histograms bin a quantitative variable and plot the count or density in each
bin. They show shape: symmetric, skewed, unimodal, bimodal, uniform, heavy
tailed. The one choice they force on you is the bin width, and it matters — too
wide hides structure, too narrow turns the picture into noise.
Box plots draw the five-number summary. They compress a distribution into
five numbers, which makes them ideal for comparing many groups at once and
useless for seeing modality.
Scatter plots show two quantitative variables jointly, and are the tool for
relationship rather than distribution.
Definition 1.51 (Skewness). A distribution is right- (or positively) skewed when its right tail is longer, left- (negatively) skewed when its left tail is longer, and symmetric when the two tails mirror each other. Numerically, the sample skewness is
which is positive for right skew and zero for any symmetric distribution.
The cube is what makes this work: it preserves sign, so deviations on the long
side do not cancel those on the short side, and it weights distant points
heavily, so the tail dominates.
Proposition 1.52 (Skew separates the mean from the median). For a right-skewed unimodal distribution one typically has , and the reverse ordering for left skew.
Pitfall. The word typically is honest, not hedging. The ordering is a reliable rule of thumb, not a theorem: discrete and multimodal counterexamples exist in which a right-skewed distribution has mean below median. Use the ordering to form a hypothesis about shape, then confirm it with a plot.
The mechanism behind the ordering is the mean's sensitivity, already proved: by
Theorem The mean minimises total squared deviation the mean chases squared
distance, so a long tail pulls it; by Theorem *The median minimises total
absolute deviation* the median only counts observations on each side, so the
tail's length is irrelevant to it.
Centre is not the whole story
Two distributions can share a mean exactly and behave completely differently,
because the mean says nothing at all about scale. Spread is an independent
dimension of the summary, not a footnote to the centre.
Remark (Anscombe's quartet and the Datasaurus). Anscombe's four datasets share their means, variances, correlation and fitted regression line to two decimal places, yet one is linear, one is curved, one is a perfect line with a single outlier, and one is a vertical stack with a lone influential point. The "Datasaurus dozen" pushes the joke further: a dozen datasets with identical summary statistics, one of which is a dinosaur. The moral is not that summaries are useless — it is that summaries are answers to specific questions, and a plot is what tells you which question to ask.
Example 1.55 (Diagnosing skew from two numbers). A dataset of annual incomes has mean and median . What shape is it, and which summary should be reported?
Solution.
- The mean exceeds the median by , so by Proposition Skew separates the mean from the median the distribution is right-skewed.
- A few very high earners stretch the right tail and drag the mean; the median is the better summary of a typical income, paired with the IQR.
- Check: this matches the known shape of income data, and the direction agrees with the histogram above. ✓
Example 1.56 (Reading a box plot comparison). Two groups have the same median but group 's box is twice as wide as group 's. What do you conclude?
Solution.
- Equal medians mean the two groups have the same typical value.
- 's IQR is twice 's, so 's middle half is twice as spread out; is the more variable group.
- Check: this is exactly the situation the two-normals figure draws, where centre alone cannot distinguish the distributions. ✓
1.9Choosing the Right Summary
Everything above reduces to one decision, made after looking at a plot.
Method 1.57 (Choosing a centre and a spread).
- Plot the data first — a histogram for one variable, box plots for several groups.
- If the distribution is roughly symmetric and unimodal with no extreme values, report and . They use all the data and feed directly into the inferential machinery of later chapters.
- If it is skewed or contains outliers you cannot justify removing, report the median and the IQR, or the five-number summary.
- If it is bimodal, report neither as a single figure. Say so, and summarise the groups separately.
- For categorical data, report the mode and the proportions; a mean is undefined.
- Always report alongside any summary. A mean of from and one from are not the same claim, because the precision differs by a factor of more than thirty.
Summary. Notation. Parameters describe a population and are fixed; statistics are computed from a sample and are random.
Centres. is the unique minimiser of , with the exact identity ; any median minimises , uniquely when is odd. Deviations from the mean always satisfy . The pooled mean of two groups is , never the average of the means unless .
Spread. for a population; for a sample. For a random variable, ; for data, . Under : , , .
Why . For a random sample (independent, identically distributed, finite variance — normality not required), , hence while the -divisor version has expectation and bias . Unbiasedness does not survive the square root: .
Sampling behaviour of . needs only a common mean and linearity of expectation; needs independence and finite variance.
Position. ; Tukey's mild fences are and , which on normal data sit at and flag about of values. The z-score standardises to mean and variance .
Coverage. The empirical rule ( within standard deviations) requires approximate normality. Chebyshev's inequality, , requires only a finite variance and is sharp, hence loose.
- Reporting the mean for heavily skewed data such as incomes or house prices. A right tail drags the mean above anything typical; report the median and IQR.
- Averaging two group means without weighting by group size. The pooled mean is , and the unweighted average is right only when the groups are equally large.
- Using the divisor for a sample variance. It is biased downward by exactly , which is negligible at and catastrophic at .
- Claiming that makes the *standard deviation* unbiased. It does not; Jensen's inequality gives . Only is unbiased.
- Forgetting to sort the data before reading off a median, a quartile or a percentile.
- Quoting the computational formula's result from floating-point code when the mean is large relative to the spread. Catastrophic cancellation can even produce a negative variance.
- Treating the standard deviation as a percentage. It carries the units of the data; the dimensionless comparison is the coefficient of variation, and that one requires a meaningful zero and a positive mean.
- Applying the empirical rule to non-normal data. Without approximate normality the figures have no standing; Chebyshev's bound is what holds universally.
- Expecting Chebyshev's bound to be tight. It is attained by a specific three-point distribution, so it cannot be improved in general, and on normal data it is far weaker than the truth.
- Deleting every point flagged by the rule. On skewed data the rule flags a large fraction of legitimate observations; a flag is a prompt to investigate, not a licence to discard.
- Reading a percentile off a z-score without normality. is the th percentile for a normal distribution and can be anywhere for another shape.
- Using the range as the primary measure of spread. It is a function of the two least typical observations and grows with even when the distribution does not change.
- Reporting a single centre for a bimodal distribution, which describes a value that may occur in neither subpopulation.
- Quoting a summary without . Precision scales as , so the sample size is part of the claim.
- Treating repeated measurements on the same subjects as independent. The formula needs independence, and positive correlation makes the sample mean far noisier than it promises.