Contents / Statistics / Correlation and Regression
Chapter 7
Correlation and Regression
Covariance and correlation with the Cauchy-Schwarz bound; least squares derived from the normal equations; the variance decomposition and R-squared; the LINE conditions and the Gauss-Markov theorem; inference for the slope and the difference between confidence and prediction intervals; multiple regression in matrix form with the hat matrix as an orthogonal projection; multicollinearity and variance inflation; model selection and the nested F-test; and why a coefficient is not a causal effect.
Introduction
Every method so far has compared groups. Regression asks a different question: given a quantitative predictor, how does the response change as the predictor moves? The groups become a continuum, the means become a line, and almost everything from the previous chapter carries over — because, as that chapter's last section showed, ANOVA was already a regression.
The chapter has three movements. The first builds the simple case in scalar algebra, where every formula can be derived by hand: what correlation measures and what it provably cannot exceed, the least-squares line obtained by minimising a sum of squares, and the decomposition of variation that gives its meaning. The second adds inference — the sampling distribution of the slope, why the least-squares estimator is optimal among a large class, and the difference between predicting a mean and predicting an observation, which is the distinction most often collapsed in practice.
The third rewrites everything in matrix form and adds predictors. This is not a new theory but the same one written compactly, and the compression pays immediately: the normal equations become a single matrix equation, the fitted values become an orthogonal projection, and the residual-orthogonality facts that looked like lucky algebra in the scalar case become geometry. With that in hand, the genuinely multivariate problems — predictors that carry the same information, and choosing among models — can be stated properly.
A warning that the chapter returns to at the end: none of this establishes causation, and the mathematics is completely silent on the question. A regression coefficient is a statement about association in the data at hand, and converting it into a claim about intervention requires assumptions that no amount of fitting can supply.
Notation. paired observations , means and , and the shorthand sums
7.1Correlation
Before fitting anything, ask how strongly two variables move together. The raw measure of co-movement is the covariance, and its defect is that it carries units.
Definition 7.1 (Covariance and correlation). The sample covariance and sample correlation of paired data are
The covariance of heights in centimetres against weights in kilograms changes if you switch to inches and pounds; the correlation does not. Dividing by both standard deviations strips the units, and the result is confined to a fixed range — which is the content of the next theorem and the reason is interpretable at all.
Theorem 7.2 (Correlation lies between -1 and 1). For any paired data with ,
with if and only if the points lie exactly on a straight line of non-zero slope.
Proof. Write and . The Cauchy-Schwarz inequality for sums states
that is . Dividing by the positive gives , hence .
For the equality case, recall that Cauchy-Schwarz is an equality exactly when the two vectors are proportional: for all and some constant . That says for every , which is the equation of a line through with slope . The sign of matches the sign of .∎
Proposition 7.3 (Correlation is invariant under increasing linear rescaling). If and with , then . In particular is unchanged when both scales are multiplied by positive constants or shifted.
Proof. Centring removes the additive constants: and similarly for . Hence , and . Substituting,
Pitfall. measures linear association and nothing else. Data lying exactly on the parabola over a symmetric range of has — a perfect, deterministic relationship that the correlation reports as no association at all.
The converse failure is equally common: a single outlying point can move from near zero to near one, since every term in is a product of two deviations and an extreme point contributes an enormous one. Always plot the data. A correlation quoted without a scatterplot is a number whose meaning has not been checked.
Example 7.4 (Computing r by hand). For the data and , compute .
Solution.
- Means: , .
- Deviations: and .
- ; ; .
- Therefore
- Sanity check: ✓ as the theorem requires, and the value is positive, matching the generally upward drift of the data. Note it is not close to despite the upward trend — the point breaks the pattern, and registers that.
Example 7.5 (One point moving the correlation). To the data , (with ) add a sixth point . Recompute .
Solution.
- New means: , .
- The new point's deviations are and , contributing to — against a total of from all five original points.
- Computing the full sums: , , . Hence
- A single point has moved the correlation from to , and it contributes of by itself.
- Sanity check on why: every term of is a product of two deviations, so a point far from both means contributes quadratically. One observation at ten times the spread of the others dominates the sum entirely.
- The reported is not describing the five original points at all — it is describing the fact that one point sits far away in a consistent direction. This is why the pitfall above insists on a scatterplot.
7.2The least-squares line
Correlation summarises association with one number. A regression goes further and produces a function, which can be used to predict.
Definition 7.6 (The simple linear regression model). The model treats the predictor as fixed and the response as random,
so that and each observation deviates from that line by independent noise of constant variance.
Theorem 7.7 (The least-squares estimates). The values of minimising the residual sum of squares are
and this minimum is unique whenever .
Proof. is a quadratic in two variables with positive-definite Hessian, so setting both partial derivatives to zero locates the unique global minimum.
Differentiating in :
which on dividing by gives — the line passes through the point of means.
Differentiating in :
These two are the normal equations. Substitute into the second and rearrange:
and the two sides are precisely and , by the computational identity of the theorem Computational formula for the variance. Hence .
For uniqueness, the Hessian of is , whose determinant is ; with positive diagonal it is positive definite, so the stationary point is a strict minimum.∎
Corollary 7.8 (Slope, correlation and the regression effect). The slope and the correlation are related by
Proof. , and , so . Substituting,
Intuition. That corollary contains the phenomenon that gave regression its name. Measure both variables in standard units and the slope is , which has absolute value at most . So a subject one standard deviation above average in is predicted to be only standard deviations above average in — closer to the mean, in relative terms, than they were.
This is regression to the mean, and it is a property of imperfect correlation, not a force acting on anything. Tall fathers have sons who are tall but on average less exceptionally so; a season of exceptional performance is typically followed by a more ordinary one. Nothing is pulling anyone back: the second measurement simply contains its own independent noise, and an extreme first measurement was probably extreme partly by luck.
Proposition 7.9 (Two properties of least-squares residuals). With the residuals from the least-squares fit,
Proof. These are exactly the two normal equations of the previous proof, rewritten. The first derivative condition says , which is ; the second says , which is .∎
The second identity says the residuals are orthogonal to the predictor: whatever structure is left over carries no remaining linear relationship with . That is a consequence of how the line was chosen, not a property of the data, and it is why a plot of residuals against can never show a linear trend — only curvature or changing spread, which is exactly what makes such a plot a useful diagnostic.
Example 7.10 (Fitting the line). Fit the least-squares line to , , and predict at .
Solution.
- From the correlation example, , , , .
- Slope: . Intercept: .
- The fitted line is .
- Check against the corollary Slope, correlation and the regression effect: , and ✓.
- Fitted values at : . Residuals: , which sum to ✓ as the proposition requires.
- Prediction at : . But is outside the observed range , so this is extrapolation — see the last section for why that number deserves less confidence than the arithmetic suggests.
7.3Residuals, R-squared, and what the fit explains
The same decomposition that drove ANOVA reappears here, and for the same reason: the fitted values play the role the group means played.
Theorem 7.11 (Decomposition of variation in regression). With , and ,
Proof. Write , square and sum:
The cross term vanishes. Since and , we have , and . So the cross term is
by the proposition Two properties of least-squares residuals.∎
Definition 7.12 (Coefficient of determination). The proportion of variation in accounted for by the fit is
Proposition 7.13 (For simple regression, R-squared is r squared). In simple linear regression, , and .
Proof. Using ,
which also equals . Dividing by ,
Definition 7.14 (Residual standard error). The error variance is estimated by
the divisor being because two parameters were estimated from the data.
Pitfall. measures how much variation the fit accounts for, and that is all. It does not say the model is correct, that the relationship is linear, or that causes .
Anscombe's quartet is the standard demonstration: four data sets with identical means, identical variances, identical , identical fitted lines and identical , of which only one is appropriately described by a straight line — the others contain a clear curve, a single high-leverage point driving everything, and an outlier. Every summary statistic agrees; the scatterplots do not.
also rises mechanically with every predictor added, which is what the adjusted version in the model-selection section corrects.
Example 7.15 (Decomposing the variation). For the fitted line on the running example, compute , , and .
Solution.
- By the proposition For simple regression, R-squared is r squared, .
- , so .
- Confirm directly from the residuals : their squares are , summing to ✓.
- , and indeed ✓.
- , so .
- Reading the result: the line accounts for of the variation in , and a typical observation sits about units off it. With , that residual spread is not small relative to the response, which is the honest counterweight to a respectable-sounding .
7.4The conditions, and why least squares is optimal
Definition 7.17 (The LINE conditions). The simple regression model assumes
- Linearity — really is ;
- Independence — the errors are independent of one another;
- Normality — the errors are normally distributed;
- Equal variance — for every (homoscedasticity).
These are not equally important, and the following theorem says exactly which ones the central result needs. Notice what is absent from its hypotheses.
Theorem 7.18 (Gauss-Markov). Assume linearity, independence and equal variance, with — normality is not required. Then among all estimators of that are linear in the and unbiased, the least-squares estimator has the smallest variance. It is the best linear unbiased estimator.
Proof. First write in linear form. Since ,
and these weights satisfy two identities used repeatedly below:
Unbiasedness. .
Variance. By independence and equal variance, .
Optimality. Let be any other linear unbiased estimator. Unbiasedness for every forces and , by the same expansion. Write ; then the two constraints give and . Now
and the cross term vanishes:
Therefore , with equality only when every , that is .∎
The theorem is doing more than it appears. It says least squares is not an arbitrary choice of loss function: under three assumptions that say nothing about the shape of the error distribution, minimising squared error delivers the best estimator in a large and natural class. Normality is needed only later, for the -distributions used in inference — and even there the central limit theorem makes the requirement soft for large .
Corollary 7.19 (Precision of the slope). , estimated by , so the standard error of the slope is .
Proof. Contained in the proof above; the estimate substitutes for .∎
Remark (How to make a slope precise). in the denominator is the design lever. It grows with the sample size and with the spread of the predictor values, so a slope is estimated precisely when the 's are spread widely, and imprecisely when they are bunched together however many there are.
That has a concrete design consequence: if the linear model is trusted and the only goal is estimating the slope, the optimal allocation puts half the observations at each extreme of the feasible range. The reason not to do that in practice is that such a design has no ability whatever to detect curvature — with two distinct values, every fit is a perfect straight line through two means. Spreading buys precision; keeping intermediate values buys the ability to check the assumption.
Pitfall. Of the four conditions, independence is again the one that cannot be repaired after the fact and cannot be checked from the fitted values alone. Time-ordered data with correlated errors produces standard errors that are too small — sometimes by a large factor — so slopes look significant when they are not.
Equal variance is checked by plotting residuals against fitted values and looking for a funnel shape. A widening funnel does not bias , which stays unbiased by Gauss-Markov's independence of the normality assumption, but it invalidates the standard errors and therefore every -value and interval.
Linearity is checked on the same plot, looking for curvature. A residual plot showing a clear U shape means the model is wrong in a way that may not reveal.
Example 7.21 (Reading a residual plot). A regression of sales on advertising spend gives , and the residual plot funnels outward. What is compromised, and what is not?
Solution.
- Still valid. By the theorem Gauss-Markov... almost. Gauss-Markov assumes equal variance, so strictly the optimality claim goes too. But unbiasedness of survives — it needs only and linearity, neither of which a funnel violates. The fitted line is still estimating the right slope on average.
- Compromised. Every standard error, every statistic, every confidence and prediction interval. All of them were computed from a single , and there is no single .
- The direction of the error is not uniform: the intervals are too narrow where the spread is large and too wide where it is small, so the overall coverage can look roughly right while being wrong at every individual .
- Remedies: model the variance (weighted least squares, weighting each observation by ), transform the response (a log transform often stabilises a funnel that widens proportionally), or use standard errors robust to heteroscedasticity.
- Sanity check on the role of : it says nothing about this. is computed from the sums of squares and is entirely blind to how the residuals are distributed across the range — which is exactly the pitfall about Anscombe's quartet, met here in the wild.
7.5Inference for the slope, and two kinds of prediction
Theorem 7.22 (Sampling distribution of the slope). Under the full LINE conditions,
so a confidence interval for is , and is tested by referring to .
Proof. By the theorem Gauss-Markov, is a linear combination of the ; under normality that combination is normal, with mean and variance . So is standard normal.
Independently of it, — quoted, as the analogous fact was in the -test chapter, the degrees of freedom being because two parameters were fitted. Forming the ratio as in the definition Student's t-distribution replaces by and gives .∎
Testing is testing whether the predictor carries any linear information about the response. For simple regression this is the same test as the ANOVA of the previous chapter, and indeed again — the model comparison being between the fitted line and the intercept-only model.
Prediction needs more care, because there are two different questions and they have different answers.
Theorem 7.23 (Confidence and prediction intervals). At a predictor value , with , a interval for the mean response is
while a prediction interval for a single new observation at is
Proof. Write , using . The two terms are independent — and are uncorrelated, since by the weight identity — so
which gives the first interval on replacing by and using .
For a new observation , the prediction error is , and is independent of the data used to fit. Variances therefore add:
which is the second interval.∎
Intuition. The extra under the square root is the whole difference, and it never goes away. Estimating the average response at gets easier with more data, and the interval shrinks toward zero width as . Predicting a single new observation does not, because that observation has its own noise , which no amount of data about other observations can reduce.
So the prediction interval has a floor at roughly , and with infinite data it converges to the interval you would draw around the true line from knowing alone. Confusing the two produces prediction intervals far too narrow, which is one of the more consequential errors in applied work.
Example 7.24 (Both intervals, at the centre and at the edge). For the running example (, , , , , ), give the confidence and prediction intervals at and at .
Solution.
- At . , and .
- Mean response: , margin , interval .
- New observation: , margin , interval — nearly two and a half times as wide.
- At . , and .
- Mean response: , margin , interval .
- Sanity check on the shape: both intervals are narrowest at and widen as moves away, because the term grows. The fitted line pivots about the point of means, so a small error in the slope has little effect near and a growing effect further out — which is the mathematical reason extrapolation is dangerous, quantified.
Example 7.25 (Testing the slope). Test for the running example at .
Solution.
- Standard error: .
- Statistic: on degrees of freedom.
- Against : do not reject. The slope is not distinguishable from zero.
- Sanity check against : the fit accounts for of the variation, which sounds substantial, yet the slope is not significant — because leaves only degrees of freedom and the critical value is correspondingly large. A high on tiny data is not evidence of a real relationship.
- Confirm via : , and ✓, the equivalence noted after the theorem Sampling distribution of the slope.
- The confidence interval says the same thing more usefully: . It contains zero, and it is wide enough to be compatible with a substantial positive slope as well — the experiment has not settled the question either way.
7.6Many predictors: the matrix formulation
With several predictors the scalar formulas become unmanageable, and rewriting in matrix notation does more than save ink: it turns the algebra into geometry, and the geometry explains results that looked accidental.
Definition 7.27 (The multiple regression model in matrix form). With parameters (an intercept and predictors), collect the data into
with the design matrix. The model is
The single equation encodes two of the LINE conditions at once: the off-diagonal zeros are independence, and the constant diagonal is equal variance.
Theorem 7.28 (The normal equations and the least-squares solution). The vector minimising satisfies the normal equations
and if has full column rank then is invertible and
Proof. Expand the objective:
using , both being scalars. Differentiating with respect to ,
and setting it to zero gives the normal equations.
For full rank: if then , so , and full column rank forces . Hence is invertible, and being positive definite by the same computation, the stationary point is the unique minimum.∎
Definition 7.29 (Hat matrix). The hat matrix is , so that the fitted values and residuals are
Theorem 7.30 (Least squares is orthogonal projection). is symmetric and idempotent (), satisfies , and is the orthogonal projection onto the column space of . Consequently : the residual vector is orthogonal to every predictor.
Proof. Symmetry: , since is symmetric and so is its inverse.
Idempotence: the inner factors cancel,
Likewise , so fixes the column space pointwise. A symmetric idempotent matrix fixing a subspace pointwise is the orthogonal projection onto it.
Finally, from the normal equations we get
Intuition. The geometry makes the whole method obvious in hindsight. The vector lives in ; the fitted values , as ranges over all coefficient vectors, sweep out a -dimensional subspace — the column space of , the set of all predictions the model is capable of making.
Least squares asks for the point of that subspace closest to , and the closest point of a subspace is the foot of the perpendicular. So the fit is the orthogonal projection, and the residual is the perpendicular component. Every algebraic identity in this chapter is that picture written in coordinates: residuals summing to zero is orthogonality to the intercept column of ones, is orthogonality to the predictor column, and is Pythagoras' theorem on the right triangle with legs and .
Theorem 7.31 (Mean and variance of the coefficient estimates). Under the model, is unbiased with
and is unbiased for . The standard error of is , and .
Proof. Write , a fixed matrix, so is linear in . Then
For the variance, the rule with gives
The degrees of freedom arise because lies in the orthogonal complement of a -dimensional subspace, which has dimension ; the unbiasedness of and the distribution then follow as in the simple case.∎
Pitfall. In multiple regression, is the change in per unit change in with all other predictors held fixed. That is a different quantity from the simple-regression slope of on alone, and the two can have opposite signs.
The classic instance: regressing income on years of education gives a positive slope; adding occupation to the model can shrink or reverse it, because education's effect operates partly through occupation, and holding occupation fixed removes that pathway. Neither coefficient is wrong; they answer different questions. Reporting a multiple-regression coefficient as "the effect of " without naming what is being held fixed is ambiguous at best.
Example 7.32 (A two-predictor fit read correctly). A model of house price (thousands) on floor area (, in ) and age (, years) gives with standard errors and , from houses. Interpret, and test each coefficient at with .
Solution.
- Interpretation. Among houses of the same age, each additional is associated with more. Among houses of the same size, each additional year of age is associated with less.
- Degrees of freedom: , since parameters were fitted.
- Area: , far beyond — strongly significant.
- Age: , and — significant.
- Confidence interval for the area coefficient: . The association is clearly positive but its magnitude is pinned down only to within about .
- Sanity check on the intercept: is the predicted price of a brand-new house of zero floor area, which does not exist. An intercept is frequently meaningless in isolation and is not evidence of anything wrong with the model — it is a fitted constant, not a prediction at a real design point.
Example 7.33 (The normal equations by hand). For the running data , , write and and solve the normal equations, confirming the scalar answer.
Solution.
- The design matrix has a column of ones and the column, so
using and . 2. The determinant is , so
- Therefore
- This is exactly , from the scalar computation ✓.
- Sanity check on the determinant: ✓, which is the same quantity that appeared in the Hessian in the proof of the theorem The least-squares estimates. It is non-zero exactly when the 's are not all identical — the full-rank condition in this simplest case.
- Standard errors follow from the same inverse: ✓, matching the corollary Precision of the slope.
7.7When predictors carry the same information
The matrix formula warns of a specific failure mode: if is near-singular, its inverse has enormous entries and the coefficients are estimated terribly.
Definition 7.34 (Multicollinearity and the variance inflation factor). Multicollinearity is strong linear association among the predictors. Its effect on predictor is measured by the variance inflation factor
where is the coefficient of determination from regressing on all the other predictors.
Theorem 7.35 (Variance inflation). The variance of the -th coefficient estimate is
where . So the variance is that of a simple regression on alone, multiplied by the inflation factor.
Proof. We quote the standard identity, whose derivation uses the partitioned-inverse formula for . The content is readable directly: regress on the other predictors and let be the residual vector, the part of not explained by the others. Then depends on the data only through , with
the second equality being the definition of applied to that auxiliary regression. Substituting gives the claim.∎
That proof says exactly what is going on, and it is more useful than the formula. A coefficient is estimated from the part of its predictor that the other predictors cannot account for. If is nearly a linear combination of the others, that leftover part is tiny, and the coefficient is being estimated from almost no information.
Remark (What multicollinearity does and does not break). It inflates standard errors, so individual coefficients become insignificant and unstable — dropping one observation can swing them wildly, and two highly correlated predictors often produce large coefficients of opposite sign that nearly cancel.
It does not bias the estimates, and it does not harm prediction within the range of the observed data: the fitted values and are unaffected, because the column space of is unchanged. This is the resolution of a common puzzle — a model with a highly significant and not one significant coefficient. The predictors jointly explain a great deal; the data simply cannot apportion the credit.
So the remedy depends on the goal. For prediction, do nothing. For interpreting individual coefficients, drop one of the redundant predictors, combine them into an index, or collect data that breaks the association.
Example 7.36 (Reading VIFs). A model has for one predictor and for another. Compare their variance inflation factors and the effect on the standard errors.
Solution.
- for the first; for the second.
- Standard errors scale as , so the first coefficient's standard error is times what it would be with an uncorrelated predictor; the second's is times.
- To compensate, the first predictor would need about times as many observations to reach the same precision, since variance falls as .
- Sanity check on the usual thresholds: rules of thumb flag or . Those correspond to and — very strong association among predictors, and the thresholds are conventions rather than results.
- Note what is not implied: the first predictor may still belong in the model. A large VIF is a statement about precision, not about relevance.
7.8Choosing a model
Adding a predictor can never increase — the larger model contains the smaller one as the special case where the new coefficient is zero, so its minimised sum of squares is at least as small. therefore rises with every predictor added, including pure noise, and cannot be used to compare models of different sizes.
Theorem 7.37 (Nested model comparison). Let a full model with parameters contain a reduced model with parameters as a special case. To test whether the extra coefficients are all zero,
rejecting for large .
Proof. Both residual sums of squares are squared lengths of projections onto orthogonal complements of nested subspaces, so and the difference is the squared length of the projection onto the -dimensional part of the full column space orthogonal to the reduced one — by the theorem Least squares is orthogonal projection applied twice.
Under that component consists of noise alone, so , independently of since the two live in orthogonal subspaces. The ratio of independent chi-squares each over its degrees of freedom is , with cancelling.∎
Taking the reduced model to be intercept-only recovers the overall -test of the whole regression, and taking recovers the -test of a single coefficient, since . It is the same machinery as the ANOVA , which the previous chapter showed was this test all along.
Example 7.38 (Do three predictors earn their place?). A model with has with parameters. Dropping three predictors gives with . Test at , given .
Solution.
- The reduction in from the three predictors is , on degrees of freedom, so the numerator is .
- The denominator is .
- . Since , reject: at least one of the three carries information beyond the reduced model.
- Sanity check on what this does not say: the joint test does not identify which of the three matters, exactly as a significant one-way ANOVA does not identify which groups differ. Examining the three individual statistics afterwards is three more comparisons.
- Note the contrast with the individual tests: under multicollinearity all three coefficients could be individually insignificant while this joint test is highly significant, because the three together explain a lot even when no one of them is separable from the others.
Definition 7.39 (Adjusted R-squared).
Proposition 7.40 (Adjusted R-squared penalises useless predictors). Adding a predictor increases only if it reduces by more than the factor by which it reduces the remaining degrees of freedom. A predictor whose statistic has always decreases .
Proof. increases exactly when — the estimate — decreases, since is fixed. Adding a predictor changes the numerator from to and the denominator from to , so falls exactly when
The reduction in from adding one predictor equals for that predictor's statistic, so the condition is essentially .∎
Remark (Information criteria and regularisation). Two other approaches trade fit against complexity more formally. AIC and BIC both add a penalty per parameter, with BIC's penalty larger for , so BIC selects smaller models. Lower is better for both, and neither is a hypothesis test — they rank models rather than accepting or rejecting.
A different strategy shrinks the coefficients instead of dropping predictors. Ridge regression minimises , which has the closed form — note that adding makes the matrix invertible even when is not, which is precisely the multicollinearity cure. Lasso uses instead and sets some coefficients exactly to zero, performing selection and fitting at once.
Both are deliberately biased, and by the theorem Bias-variance decomposition that is the point: they accept bias to buy a larger reduction in variance, and can beat least squares on mean squared error despite Gauss-Markov — which constrains only unbiased estimators.
Pitfall. Selecting predictors by looking at -values, then reporting the final model's -values as though the search had not happened, is the multiplicity error of the testing chapter in its most common form. Stepwise selection over candidate predictors examines far more than hypotheses, and the surviving coefficients are biased away from zero because they survived.
If the model was chosen using the data, its reported standard errors and intervals are not valid for that data. The honest options are to select on one part of the data and assess on another, to use cross-validation, or to report the selection procedure and treat the result as exploratory.
Example 7.41 (Adjusted R-squared in use). A model with and has . A fifth predictor raises to . Compute both adjusted values and decide.
Solution.
- Four parameters: .
- Five parameters: .
- The adjusted value rises very slightly, from to , so the predictor just about pays for itself.
- Sanity check via the criterion: by the proposition Adjusted R-squared penalises useless predictors, the increase means the new predictor's exceeds — which is a very weak bar, far below the needed for significance. So "improves adjusted " is not the same as "is worth keeping".
- The gain of in adjusted is not a reason to prefer the larger model. With this little to separate them, the simpler model is the better default.
7.9Extrapolation, and the limits of what a fit can say
Definition 7.42 (Extrapolation). Extrapolation is prediction at a predictor value outside the range spanned by the observed data.
The intervals of the theorem Confidence and prediction intervals already price part of the cost: the term grows quadratically as leaves the data, so both bands flare outward. But that widening assumes the model still holds, and the real danger is that it does not. A relationship linear over the observed range may bend, saturate, or reverse outside it, and the data contain no information whatever about which.
Pitfall. The most consequential limitation is that a regression coefficient is not a causal effect. Least squares fits a conditional mean; nothing in the derivation licenses a claim about what would happen under intervention.
Three mechanisms break the inference, and none is detectable from the fit:
- Confounding. A variable associated with both and creates an association between them with no causal path. Ice cream sales predict drownings; temperature drives both.
- Reverse causation. The regression is symmetric in a sense the causal claim is not — fitting on gives a slope , and both fits describe the same data equally well.
- Selection. If inclusion in the sample depends on both variables, an association appears among those selected even when none exists in the population.
Adding controls addresses only confounders you measured and thought of, and adding the wrong control — one on the causal pathway, or one caused by both variables — can create bias rather than remove it. Design is what licenses causal claims, which is the subject of the next chapter.
Example 7.43 (The same data, two regressions). For the running example, fit on as well as on , and compare.
Solution.
- on : slope , giving .
- on : slope , and intercept , giving .
- If the first line could simply be inverted, it would give — a slope of , not . The two fits are different lines.
- The product of the two slopes is , which is exact and general: . The lines coincide only when .
- Sanity check on what this means: each regression minimises errors in a different direction — vertical for on , horizontal for on — so asking "which variable causes which" cannot be settled by fitting, since both fits are equally valid descriptions. The asymmetry of a causal claim has no counterpart in the algebra.
- **Reading as a measure of association in general.** It measures *linear* association; over a symmetric range gives .
- Quoting a correlation or a fit without plotting the data. Anscombe's quartet shares every summary statistic across four data sets that demand four different models.
- **Treating as a verdict on the model.** It says how much variation the fit accounts for, not whether the form is right, and it rises with every predictor added.
- Confusing a confidence interval for the mean response with a prediction interval. The second carries an extra that never shrinks with more data.
- Reading a multiple-regression coefficient without saying what is held fixed. It is a partial effect, and can differ in sign from the simple-regression slope.
- Concluding a predictor is useless because its coefficient is insignificant under multicollinearity. Inflated standard errors say the data cannot apportion credit, not that the predictor is irrelevant.
- **Reporting -values from a model selected using the same data.** The surviving coefficients are biased away from zero precisely because they survived.
- Extrapolating. The bands widen quadratically outside the data, and even that assumes the model still holds, which the data cannot confirm.
- Calling a coefficient an effect. Confounding, reverse causation and selection all produce coefficients, and none is visible in the residual plots.