Contents / Probability / Foundations of Probability
Chapter 1
Foundations of Probability
Sample spaces and events, the Kolmogorov axioms and the rules derived from them, counting, conditional probability, independence, total probability and Bayes' theorem.
Introduction
Every other chapter in this subject rests on this one. A random variable is a function on a probability space, a distribution is an assignment of probabilities to its values, and an expectation is an average taken against those probabilities. The same is true one subject over: in the companion Statistics course an estimator is a random variable, a confidence interval is a statement about a probability, and a -value is a probability. None of those sentences means anything until "probability" has been given a precise definition, so that is what this chapter does.
The definition is surprisingly economical. Andrey Kolmogorov's 1933 formulation asks for only three things: probabilities are non-negative, the certain event has probability , and the probability of a union of non-overlapping events is the sum of their probabilities. That is the entire foundation. Everything else in this chapter — the complement rule, the addition rule, inclusion–exclusion, the law of total probability, Bayes' theorem — is derived, and this chapter derives them rather than asserting them. Where a rule needs a hypothesis, the proof is where you see the hypothesis being used, which is the only reliable way to remember that it is there.
Two ideas do more work than the rest. The first is conditional probability, which is not a formula to memorise but a definition: it says what it means to recompute a probability after learning something. The second is independence, which says what it means to learn nothing. Almost every modelling assumption you will make in later chapters — random sampling, independent errors, independent trials — is an independence assumption in disguise, and almost every catastrophic failure of a statistical model is an independence assumption that was false.
The chapter closes on Bayes' theorem and the base-rate effect, worked with real numbers on a medical test. The answer there is one that most readers, including most doctors in the studies that asked them, get badly wrong. It is worth arriving at it by proof rather than by intuition, which is the general argument for this chapter.
1.1Sample spaces, events, and the axioms
A probability model has three ingredients: a set of things that could happen, a collection of statements about what happened, and a rule assigning numbers to those statements. We take them in that order.
Definition 1.1 (Experiment, outcome, sample space). An experiment is any procedure whose result is not determined in advance. Each result it can produce is an outcome, and the set of all possible outcomes is the sample space, written (or ).
The outcomes in must be mutually exclusive — exactly one of them occurs — and exhaustive: no other result is possible.
The two conditions are what make a usable bookkeeping device, and they are a modelling choice rather than a fact about the world. For a single roll of a die, . For two rolls, the outcomes are ordered pairs and , which has elements. For "the number of emails arriving today", is countably infinite. For "the time until a component fails", is uncountable, and that case is genuinely harder — the last part of this section says why.
Definition 1.2 (Event). An event is a subset . The event occurs when the outcome that actually happened is an element of .
An event containing exactly one outcome is an elementary or simple event. itself is the certain event and the impossible event.
Because events are sets, the logical operations you want to perform on statements are already available as set operations. This dictionary is worth learning once rather than re-deriving each time.
Notation (Events as sets). For events and in a sample space :
- is " or " (inclusive: at least one occurs);
- , often written , is " and " (both occur);
- , also written or , is "not " (the complement of in );
- is " but not ";
- means " implies ": whenever occurs, so does ;
- means and are mutually exclusive (disjoint): they cannot both occur.
De Morgan's laws transfer between them: and . In words, "not (A or B)" is "neither", and "not (A and B)" is "at least one fails".
Example 1.3 (Events from two dice). A red die and a blue die are rolled, so with . Let be "the sum is ", be "the red die shows ", and be "both dice show the same face". Write each as a set, and describe , and .
Solution. , so .
, so .
, so .
: the sum is and the red die is , which pins the blue die to . One outcome.
: a double has an even sum, so it can never total . The events "sum is " and "doubles" are mutually exclusive.
contains the six outcomes of and the six of , but lies in both, so . That correction is the addition rule in miniature, and it is proved below.□
Now the numbers. Kolmogorov's axioms say what a probability is by listing what it must satisfy, and nothing else. Note what is absent: no mention of frequencies, of symmetry, of degrees of belief. Those are interpretations, and the mathematics is the same under all of them.
Definition 1.4 (The Kolmogorov axioms). A probability measure on a sample space is a function assigning a real number to each event , such that
- (Non-negativity) for every event ;
- (Normalisation) ;
- (Countable additivity) if are pairwise mutually exclusive — meaning whenever — then
The triple , where is the collection of events, is a probability space.
Three remarks on what the axioms do and do not say. First, the upper bound is not an axiom: it is a theorem, proved below from the three axioms. Stating it as a fourth axiom is a common but redundant presentation. Second, axiom 3 requires the events to be pairwise disjoint; applied to overlapping events it is false, and the addition rule exists precisely to repair it. Third, additivity is stated for countably many events rather than finitely many, which is a genuine strengthening and is what makes limits behave.
Note (Why not every subset is an event). For a finite or countable you may take to be all subsets of , and nothing goes wrong. For you cannot: there is no way to assign a length to every subset of the line while keeping countable additivity and translation invariance, a fact established by Vitali's construction of a non-measurable set. The fix is to restrict to a -algebra — a collection of subsets containing and closed under complements and countable unions — and to define only there. Every set you can describe in this course is in it. We will not mention again, but it is the reason "event" is a technical term rather than a synonym for "subset".
The first consequences are immediate, and they show the machine working.
Proposition 1.5 (The impossible event and finite additivity). For any probability measure :
- ;
- if are pairwise mutually exclusive, then .
Proof. (1) Take for every . These are pairwise disjoint, since , and their union is . Countable additivity gives
Write by non-negativity. If the right-hand side diverges while the left is finite, a contradiction; so .
(2) Given finitely many disjoint , extend the list by setting for . The extended list is still pairwise disjoint and has the same union, so countable additivity applies and the tail contributes by part (1).∎
So finite additivity is a consequence, not a separate assumption — and it needed , which in turn needed countable additivity. The axioms are more tightly interlocked than they look.
Proposition 1.6 (Complement rule). For any event , .
Proof. and are mutually exclusive, since no outcome is both in and outside it, and their union is . By finite additivity and normalisation,
Subtract .∎
The complement rule is the single most useful computational trick in the chapter, because "at least one" is almost always harder to count than "none". Whenever a problem says at least one, compute the probability of none and subtract.
Proposition 1.7 (Monotonicity, and the upper bound). If then , and moreover . Consequently for every event .
Proof. Split along : since ,
and the two pieces are disjoint, because one lies inside and the other inside . Finite additivity gives , which is the stated difference formula since . Non-negativity makes , hence .
For the bound, every event satisfies , so ; and is axiom 1.∎
Remark. Monotonicity is the formal version of "a more demanding requirement is less likely". If every outcome that makes true also makes true, cannot be the more probable of the two. A probability estimate that violates this — the celebrated "Linda is a bank teller and a feminist is more likely than Linda is a bank teller" experiment of Tversky and Kahneman — is not merely implausible, it is inconsistent with the axioms.
Now the rule for events that do overlap. Additivity fails for them, and the proof shows exactly by how much.
Theorem 1.8 (Addition rule for two events). For any events and ,
Proof. Decompose into three pairwise disjoint pieces:
which are " only", "both", and " only". They are disjoint by construction, and every outcome in lies in exactly one of them. Finite additivity gives
Applying the same split to alone, with disjoint pieces, so ; symmetrically . Substituting,
The subtraction is there because the outcomes in were counted once in and again in ; removing one copy leaves exactly one. When and are mutually exclusive, and the correction vanishes, recovering additivity.
Corollary 1.9 (Boole's inequality). For any events — disjoint or not —
Proof. For two events, the addition rule gives , since . Induction extends this to any finite union: if it holds for events, then .
For a countable union, disjointify: set and . The are pairwise disjoint with the same union as the , and , so monotonicity gives . Countable additivity then yields .∎
Boole's inequality is crude — it throws away every overlap — but it needs no assumptions whatsoever, which is what makes it usable when nothing is known about the dependence between the events. It reappears as the first step of the Borel-Cantelli lemmas, and in the companion Statistics course it is the whole of the Bonferroni correction for multiple testing.
Theorem 1.10 (Inclusion–exclusion for three events). For any events , , ,
Proof. Apply the two-event addition rule to and :
Expand the first term by the addition rule again: .
For the last term, distributivity of intersection over union gives , and a third use of the addition rule gives
because . Substituting both expansions and collecting signs gives the statement.∎
Intuition. Count everyone once. Adding counts anybody in two of the sets twice and anybody in all three of them three times. Subtracting the three pairwise overlaps fixes the doubles, but it also strips the triple-overlap people three times, leaving them counted times — so the final puts them back. The alternating signs are not a mnemonic; they are the bookkeeping of over- and under-correction.
Example 1.11 (A survey with three overlapping groups). In a survey of students, take statistics, take economics, and take computing. Of the overlaps, take statistics and economics, take statistics and computing, take economics and computing, and take all three. What fraction take at least one of the three, and what fraction take none?
Solution. Write , , for the three subjects. Inclusion–exclusion gives
So take at least one. By the complement rule, , or .
Sanity check by Boole's inequality: the sum exceeds , so the bound is vacuous here but not violated, and the true value must lie at or below it. A second check: the answer must be at least by monotonicity, since . Both hold.□
Finally, the special case that most introductory problems live in. It is a model, not a law, and the axioms are what justify it.
Definition 1.12 (The equally likely model). A sample space with finitely many outcomes is equiprobable when every elementary event has the same probability.
Proposition 1.13 (Counting formula for equiprobable spaces). If is equiprobable with outcomes, then for every event ,
Proof. Let be the common probability of each elementary event . The elementary events are pairwise disjoint with union , so finite additivity and normalisation give , hence . Now any event is the disjoint union of its elementary events, so finite additivity gives .∎
This is why counting matters, and it is the whole content of the next section: once the model is equiprobable, computing a probability means computing two cardinalities.
Pitfall. The formula is a theorem about equiprobable spaces, not a definition of probability. Two routine ways to misuse it: applying it to a sample space whose outcomes are not equally likely, and choosing a sample space that quietly destroys the symmetry. Rolling two dice and taking for the sum gives eleven outcomes, but and are not equal, so . Use the ordered pairs, where the symmetry of the physical setup really does make the outcomes equiprobable, and count within that.
1.2Counting outcomes
Counting is not a separate subject here; it is how probabilities in equiprobable spaces get computed. The results below are the minimum needed for probability, and each is a consequence of one principle.
Theorem 1.14 (Multiplication principle). If a procedure consists of stages, where stage can be completed in ways and, for each way of completing stages through , stage can be completed in ways, then the whole procedure can be completed in
ways.
Proof. Induct on . For the claim is the hypothesis. Suppose it holds for stages, so there are ways to complete them. Each of those partial outcomes extends to exactly full outcomes, and two full outcomes built on different partial outcomes are different. So the full outcomes are partitioned into groups of size , giving in total.∎
The phrase "for each way of completing the earlier stages" matters: the number of options at stage must not depend on which earlier choices were made, though which options they are may. Dealing cards is the standard case — after any first card there are always possibilities for the second, even though the identity of those changes.
Example 1.15 (Licence plates). A plate is three letters from a -letter alphabet followed by four digits. How many plates are there if repetition is allowed, and how many if all seven characters must be distinct within their groups (no repeated letter, no repeated digit)?
Solution. With repetition: .
Without repetition: the letters contribute and the digits , giving .
Sanity check: the second count must be smaller, since every distinct-character plate is also a plate — and . The ratio is about , so fewer than half of all plates have all-distinct characters.□
Definition 1.16 (Permutation). A permutation of objects from a set of distinct objects is an ordered selection of of them, without repetition. The number of such permutations is written .
Notation. The symbol is unfortunately overloaded: is a probability and a count. They are told apart by their arguments — a probability takes an event, a permutation count takes two integers. Some books write or to avoid the clash.
Proposition 1.17 (Permutation count). For ,
In particular : there are orderings of distinct objects.
Proof. Build the ordered selection one position at a time. Position can be filled in ways; whichever object was used, position can be filled in ways, and in general position in ways. The count at each stage does not depend on the earlier choices, so the multiplication principle applies and gives the product , which has factors. Multiplying and dividing by turns the product into .∎
Definition 1.18 (Combination). A combination of objects from distinct objects is an unordered selection of of them, without repetition — that is, a -element subset. The number of them is the binomial coefficient , read " choose ".
Theorem 1.19 (Combination count). For ,
Proof. Count the ordered selections of objects in two ways. Directly, there are of them.
Alternatively, build an ordered selection in two stages: first choose which objects to use, in ways by definition, then choose an order for them, in ways by the permutation count with . Every ordered selection arises exactly once this way, so by the multiplication principle there are of them.
Equating the two counts, , and dividing by gives the formula.∎
That argument — count one set two ways, equate — is called double counting, and it proves the next two facts more cleanly than algebra does.
Proposition 1.20 (Symmetry and Pascal's rule). For ,
Proof. Symmetry: choosing the objects to include is the same act as choosing the objects to leave out, so the two collections of subsets are in one-to-one correspondence and have the same size.
Pascal's rule. Fix one particular object . Every -subset either contains or does not. Those containing are determined by the remaining elements drawn from the other objects, giving ; those not containing are -subsets of the other objects, giving . The two cases are exclusive and exhaustive, so the counts add.∎
Method 1.21 (Choosing the right count). Before computing anything, answer two questions about the selection.
- Does order matter? If rearranging the chosen items gives a genuinely different outcome — offices, rankings, sequences, positions — order matters. If not — committees, hands of cards, subsets — it does not.
- Is repetition allowed? Drawing with replacement, or filling slots from an unlimited supply, allows repetition; dealing from a deck does not.
Then: ordered with repetition gives ; ordered without repetition gives ; unordered without repetition gives . (Unordered with repetition is , which probability needs only rarely.)
- Use the same convention on top and bottom. In , both cardinalities must count the same kind of object. Counting as ordered hands and as unordered hands is the most common arithmetic disaster in this chapter.
Example 1.22 (Ordered versus unordered from one pool). A club of people must (a) elect a president, a vice-president and a treasurer, all distinct offices, and (b) separately pick a -person social committee with no offices. Count each.
Solution. (a) Order matters, because being president is not the same job as being treasurer. This is a permutation:
(b) Order does not matter. This is a combination:
Sanity check: the only difference between the two questions is whether the orderings of a chosen trio count as distinct, and indeed , exactly as the proof of the combination count predicts.□
Example 1.23 (Exactly two aces). Five cards are dealt from a well-shuffled standard deck. Find the probability that the hand contains exactly two aces.
Solution. Take to be the set of all -card hands, unordered. Because the deck is well shuffled, all hands are equally likely, so the counting formula applies with
A hand with exactly two aces is built in two stages: choose which of the aces appear, then choose the other cards from the non-aces. By the multiplication principle,
Therefore
Sanity check: both counts are unordered, as the recipe demands. Note that "exactly two" required choosing the remaining three cards from the non-aces; drawing them from all remaining cards would have counted hands with three or four aces as well, and would also have double counted.□
Example 1.24 (The birthday problem). In a room of people with birthdays independently uniform over days, what is the probability that at least two share a birthday?
Solution. "At least one shared birthday" is awkward to count directly — the shares can occur in many patterns — so use the complement rule and count the opposite: all birthdays distinct.
The sample space of birthday lists has equally likely elements, by the multiplication principle with repetition allowed. Lists with all birthdays distinct are ordered selections without repetition, so there are of them. Hence
and by the complement rule
Sanity check on the surprise: the relevant count is not people but pairs, and people form pairs, each matching with probability . Boole's inequality caps the answer at , and the true value sits comfortably below it. The intuition that fails here compares with ; the right comparison is with .□
Pitfall. Counting the same outcome twice is far more common than missing one. If you compute "the number of -card hands with at least one ace" as — pick an ace, then any four other cards — you have counted a two-ace hand twice, once for each ace that could have been "the" chosen one. The safe routes are the complement, , or a sum over disjoint cases by the exact number of aces.
1.3Conditional probability
Probabilities are not properties of events alone; they are properties of events relative to a state of information. Learning something changes them. The card you are about to draw is a heart with probability , but if you are told it is red, the probability becomes — nothing about the card changed, only what you know.
Conditioning formalises that. Learning that occurred means every outcome outside is now impossible, so becomes the new sample space. Probabilities must be recomputed inside it, and they must be rescaled so that the new sample space still has total probability . Dividing by is exactly that rescaling.
Definition 1.25 (Conditional probability). Let and be events with . The conditional probability of given is
It is undefined when .
Two features of the definition repay attention. The numerator is , not : only the part of that survives inside counts. And the requirement is not a technicality to be waved past — you cannot condition on something that never happens, because the rescaling would divide by zero.
It is worth checking that this construction produces something that is itself a probability. If it did not, every theorem proved so far would have to be re-derived for conditional probabilities, and none of them would be safe to use.
Theorem 1.26 (Conditioning yields a probability measure). Fix an event with and define for every event . Then satisfies the Kolmogorov axioms on .
Proof. Non-negativity. is a quotient of a non-negative number by a strictly positive one, hence .
Normalisation. .
Countable additivity. Let be pairwise disjoint. Then the sets are also pairwise disjoint, since , and by distributivity. So
using countable additivity of in the middle step.∎
Corollary 1.27 (Rules transfer to conditional probabilities). For any event with : , and
Proof. Both are the complement rule and the addition rule applied to the probability measure , which is legitimate by Theorem Conditioning yields a probability measure: those rules were proved for an arbitrary probability measure, and is one.∎
Pitfall. The transfer works on the event in front of the bar, never behind it. is false in general, and it is a favourite trap: complementing the condition throws you into a different sample space entirely, where the numerator has no relation to the one you started with. For the medical test later in this chapter, while , and those two happen to sum to only by coincidence of the numbers chosen; with sensitivity and false-positive rate they would sum to .
Example 1.28 (Conditioning inside a die roll). Two fair dice are rolled. Given that the first die shows , what is the probability that the total is at least ? Compare with the unconditional probability.
Solution. Let be "total at least " and be "first die is ".
Unconditionally, , so .
Now condition. has outcomes, so . The intersection has outcomes, so . By the definition,
The same answer comes from the shrunken sample space directly: inside there are equally likely outcomes and of them qualify, giving . Knowing the first die is high doubled the probability, which is what it means for these events to be dependent.□
Rearranging the definition turns it into a rule for building joint probabilities out of sequential ones, which is how almost every multi-stage problem is actually computed.
Proposition 1.29 (Multiplication rule). If then , and symmetrically, if then .
Proof. Multiply both sides of the definition by , which is legitimate because . The second form is the same computation with the roles exchanged.∎
Theorem 1.30 (Chain rule). If , then
Proof. Induct on . The case is the multiplication rule. Assume the identity for events and write , so that by hypothesis and the multiplication rule applies to and :
Every conditioning event appearing in the inductive hypothesis contains , hence has positive probability by monotonicity, so the hypothesis expands into the first factors. Substituting gives the result.∎
Example 1.31 (Three hearts in a row). Three cards are dealt from a standard deck without replacement. Find the probability that all three are hearts.
Solution. Let be "the -th card is a heart". The draws are dependent: removing a heart changes what remains.
. Given the first card was a heart, hearts remain among cards, so . Given the first two were hearts, . By the chain rule,
Sanity check two ways. With replacement the draws would be independent and give ; removing hearts makes later hearts less likely, and indeed . Alternatively count unordered hands: , the same number, as it must be — ordering conventions cancel when used consistently on top and bottom.□
Intuition. Conditional probability is narrowing the world. asks: among cloudy days only, what fraction had rain? Sunny days are not counted as evidence against rain — they are not counted at all. That is also why and are different questions: one restricts to cloudy days and looks for rain, the other restricts to rainy days and looks for cloud. The second is nearly ; the first is not.
Pitfall. Reversing the bar is the most consequential error in applied statistics, and it has a name in the courtroom: the prosecutor's fallacy. The probability that a randomly chosen innocent person matches the forensic evidence, , may be one in a million. The probability that a person is innocent given a match, , can still be large, because the match probability must be weighed against how many people were searched. Converting between the two is exactly what Bayes' theorem does, and it cannot be done without the prior.
1.4Independence
Dependence was the interesting case above: conditioning changed the number. Independence is the boundary case where it does not — where learning leaves the probability of exactly where it was.
Definition 1.32 (Independence of two events). Events and are independent when
Otherwise they are dependent.
The product form is the right definition even though the conditional form is the better intuition, for two reasons: it is symmetric in and , so "independent" needs no ordering, and it remains meaningful when , where conditioning is undefined. The two agree whenever both make sense.
Proposition 1.33 (Conditional characterisation). If , then and are independent if and only if .
Proof. Suppose and are independent. Then
Conversely, if , multiply through by and use the multiplication rule: .∎
Proposition 1.34 (Independence passes to complements). If and are independent, so are and , and , and and .
Proof. Split along : , a disjoint union, so
using independence in the second step and the complement rule in the last. So and are independent. The case of and follows by symmetry of the definition, and applying the first result to the independent pair gives and .∎
This is the formal licence for a habit you will use constantly: if components fail independently, then they also survive independently, so .
Now the confusion the chapter must pre-empt. "Mutually exclusive" and "independent" are opposite ideas that students routinely fuse, and they are not merely different — for events that can actually happen, they are incompatible.
Theorem 1.35 (Mutually exclusive and independent are incompatible). Let and be mutually exclusive events with and . Then and are not independent.
Proof. Since , we have . Independence would require , and because both factors are strictly positive. So , a contradiction.∎
Intuition. Mutual exclusivity is the strongest possible dependence, not a form of independence. If and cannot both occur, then learning that happened tells you something decisive about : it did not happen. Formally . Independence means learning tells you nothing; exclusivity means it tells you everything.
For more than two events, the product rule must hold for every subcollection, not merely for pairs — and the difference is real.
Definition 1.36 (Mutual independence). Events are mutually independent when for every subset of indices with ,
They are pairwise independent when this holds for all with .
Mutual independence implies pairwise independence by definition. The converse fails, and the counterexample is small enough to check by hand.
Example 1.37 (Pairwise but not mutually independent). Toss two fair coins. Let = "the first coin is heads", = "the second coin is heads", and = "the two coins match". Show that are pairwise independent but not mutually independent.
Solution. The sample space is equiprobable, each outcome having probability . Then , and , so .
Each pairwise intersection is the single outcome :
so all three pairs are independent.
The triple intersection is also , so , whereas . Since , the three events are not mutually independent.
The failure is intelligible: any two of these events determine the third. Knowing the first coin is heads and that the coins match forces the second coin to be heads, so . Independence survives being checked two at a time and dies at three.□
Definition 1.38 (Conditional independence). Events and are conditionally independent given , where , when
Remark. Conditional independence neither implies nor is implied by ordinary independence. Two tests of the same patient are conditionally independent given disease status — the errors are made by the instrument, not by the patient — yet they are strongly dependent unconditionally, because a first positive result raises the chance of disease and therefore the chance of a second positive. That is precisely the structure that makes a repeated test informative, and it is used at the end of this chapter.
Example 1.39 (A redundant system). A system has three components that fail independently, each working with probability . Find the probability the system works if the components are wired in series (all must work) and if they are wired in parallel (at least one must work).
Solution. Let be "component works", with and the mutually independent.
Series. The system works exactly when occurs, so by mutual independence
Parallel. "At least one works" is awkward directly, so take the complement: the system fails only if all three fail. By Proposition Independence passes to complements, the events are mutually independent with , so
Sanity check: redundancy should help and duplication should hurt, and indeed . Note how much both answers rest on independence — if all three components share a power supply, the failures are positively dependent and is an overstatement, possibly a large one.□
Pitfall. Independence is an assumption to be justified, never a default. It is justified by the physical setup — separate coin tosses, separately sampled individuals, a randomised assignment — and not by the events merely feeling unrelated. Multiplying probabilities of dependent events is the error behind the Sally Clark case, in which two infant deaths in one family were treated as independent, producing a "one in million" figure that a shared genetic or environmental cause makes meaningless.
1.5The law of total probability
Some probabilities are hard to compute directly but easy to compute within each of several scenarios. A factory's overall defect rate is awkward; the defect rate of each machine is on a datasheet. The technique is to break the sample space into cases, work inside each, and recombine — and the recombination is a weighted average, with the weights being how likely each case is.
Definition 1.40 (Partition). Events form a partition of the sample space when they are pairwise mutually exclusive, for ; collectively exhaustive, ; and each has .
All three conditions earn their place. Exclusivity is what lets the pieces be added; exhaustiveness is what stops probability leaking out of the cases; positivity is what makes each conditional probability defined. The simplest partition is for any with , and it is the one used most often.
Theorem 1.41 (Law of total probability). If is a partition of , then for every event ,
Proof. Because the cover , intersecting with gives
using distributivity. The pieces are pairwise disjoint, since for . Countable additivity therefore gives , which is the first equality.
For the second, each has by the definition of a partition, so the multiplication rule applies term by term: .∎
Intuition. To find the overall chance of , split the world into non-overlapping cases that cover everything, find the chance of inside each case, then blend those numbers by how likely each case is. It is exactly a course grade computed as a weighted average of homework, midterm and final — and as there, the weights must sum to , which is what "partition" guarantees.
Corollary 1.42 (Weighted-average bounds). With a finite partition, lies between and .
Proof. Write and . Since every weight satisfies and by finite additivity,
and the same argument with gives the lower bound.∎
That corollary is the free sanity check on every total-probability computation: an answer outside the range of the conditional rates is arithmetic error, not a subtle effect.
Example 1.43 (Two machines, one defect rate). A factory runs two machines. Machine makes of the parts with a defect rate; machine makes the other with a defect rate. What fraction of all parts are defective?
Solution. Every part comes from exactly one machine, so partitions the sample space, with , and as required. The conditional defect rates are and .
By the law of total probability,
So of parts are defective.
Sanity check via Corollary Weighted-average bounds: the answer must lie in , and it does, sitting nearer because machine carries the larger weight.□
Example 1.44 (Drawing from a randomly chosen urn). Urn I holds red and white balls; urn II holds red and white. A fair coin selects the urn, and one ball is drawn from it. Find the probability the ball is red.
Solution. Let be the events "urn I chosen", "urn II chosen". They partition the sample space with , and , .
Sanity check: with equal weights the answer is the plain average of and , namely , and it lies between them. Note that pooling the urns into one bag of balls would give too — but only because the two urns hold equally many balls and the coin is fair. Change the urn sizes and the pooled calculation breaks while the law of total probability does not.□
Pitfall. The cases must partition the sample space, not merely be a list of interesting possibilities. Two failures recur. If the cases overlap, the shared outcomes are counted twice and the total can exceed . If they miss something — "the customer is a student, or a pensioner" leaves out everyone else — the weights sum to less than and the answer is silently too small. Before summing, check that the weights add to exactly ; it costs one line and catches both errors.
1.6Bayes' theorem and base rates
Conditional probabilities come in pairs pointing opposite ways, and the one you can measure is usually not the one you want. A laboratory can measure : take people known to have the disease and see how often the test fires. A patient holding a positive result wants . Bayes' theorem converts between the two, and the conversion requires one extra ingredient — the prevalence — which is exactly the ingredient intuition leaves out.
Theorem 1.45 (Bayes' theorem). If and , then
Proof. The multiplication rule computes in two ways, both legitimate because both conditioning events have positive probability:
Equate the right-hand sides and divide by .∎
Corollary 1.46 (Bayes' theorem with a partition). If is a partition of and , then for each ,
Proof. Apply Bayes' theorem to and , then replace the denominator by its expansion from the law of total probability, which applies because the form a partition.∎
So Bayes' theorem is two lines from the definition of conditional probability, and its denominator is the law of total probability. There is nothing else in it.
Notation (Prior, likelihood, posterior). In , read as a hypothesis and as data:
- is the prior — what you believed before seeing the data;
- is the likelihood — how probable the data would be if the hypothesis held;
- is the evidence or marginal likelihood, computed by total probability;
- is the posterior — the updated belief.
The whole formula says: posterior is proportional to likelihood times prior.
The next figure lays out the computation we are about to do. A probability tree makes the two routes to a positive result visible, and the base-rate effect is the observation that the lower branch carries more probability than the upper one.
Example 1.48 (A rare disease and an accurate test). A disease affects in people. A screening test has sensitivity and specificity , so its false-positive rate is . A randomly chosen person tests positive. What is the probability they have the disease?
Solution. Write for "has the disease". The prior is , so by the complement rule. The events partition the sample space, both having positive probability, so the law of total probability gives the evidence:
that is . Now Bayes' theorem:
About . Roughly nine out of ten positive results from this screen are false alarms, despite a test that is right of the time in both directions.
Sanity check by counting people. Among screened, about have the disease and of them test positive; about do not, and of them — people — test positive anyway. Of the positives, are genuine, and , matching. The two pieces also sum to the denominator: .□
The number people guess is usually close to , and the gap between and is the base rate doing its work. The reason is visible in the tree: the healthy branch starts with a thousand times more probability than the diseased one, so even a small error rate applied to it produces more positives than a near-perfect detection rate applied to the tiny diseased branch. Ignoring and reading the posterior off the sensitivity is the base-rate fallacy.
The update is enormous and the conclusion is still "probably not diseased". Both statements are true at once, and holding them together is the skill. In practice the resolution is a second, different test: the positive screen has moved the prior from to , and a confirmatory test now starts from that much more favourable base rate.
Proposition 1.50 (Odds form of Bayes' theorem). Define the odds of an event as . Then for any data with ,
Proof. Apply Bayes' theorem to and to separately:
Divide the first by the second. The common denominator cancels, leaving the stated product — which is why the odds form needs no total-probability computation at all.∎
Example 1.51 (The screening result in odds form). Redo the disease calculation using odds, and then find the posterior after a second independent positive test.
Solution. Prior odds: . Likelihood ratio for a positive: .
Posterior odds after one positive: . Converting back with ,
agreeing with the direct computation.
Now a second test, assumed conditionally independent of the first given disease status. Independence given and given makes the likelihood ratio of the pair the product of the individual ratios, , so
The direct check: and , and .
Sanity check on the structure: each positive multiplies the odds by , so evidence accumulates multiplicatively on the odds scale — which is precisely why the odds form is the one used in practice. Two positives take the patient from to .□
Pitfall. The second-test calculation rests on conditional independence given disease status, and that assumption fails when the two tests share a mechanism. Running the same assay twice on the same sample, on a patient whose blood chemistry is what confuses the assay, gives two positives that carry barely more information than one. The multiplicative update then overstates the evidence badly. A confirmatory test must be a different test for the arithmetic above to mean anything.
The posterior depends on the prevalence, and it is worth seeing how strongly. The figure below holds the test fixed — sensitivity and specificity both , as in the example — and sweeps the prevalence.
Example 1.53 (Spam filtering). A filter flags of spam messages and wrongly flags of legitimate ones. If of incoming mail is spam, what is the probability that a flagged message really is spam?
Solution. Let be "spam", be "flagged". Then , , .
Total probability: .
Bayes: .
Compare with the medical test, where a more accurate instrument gave a posterior of . The difference is entirely the base rate: spam is common, so the prior odds start high, and the likelihood ratio pushes them to , giving . A rare event needs a far stronger likelihood ratio to reach the same posterior.□
Summary. Axioms. A probability measure satisfies , , and countable additivity over pairwise disjoint events. Everything below is derived from those three, so every one of them holds in any probability model.
Derived rules, no hypotheses needed. ; finite additivity; ; monotonicity, , hence ; the addition rule ; Boole's inequality ; and inclusion–exclusion for three events with its correction.
Requires equally likely outcomes. holds only on a finite equiprobable sample space. Counting then uses the multiplication principle, for ordered selections without repetition, with repetition, and for unordered selections — with the same ordering convention used for and .
Requires . , which is itself a probability measure in , so the complement and addition rules apply in front of the bar but never behind it. The multiplication rule and the chain rule follow.
Requires independence. is the definition; when it is equivalent to . Independence passes to complements. Mutually exclusive events of positive probability are never independent. Mutual independence of events requires the product rule on every subcollection, not just on pairs.
Requires a partition. With disjoint, exhaustive and each of positive probability, the law of total probability gives , and must lie between the smallest and largest .
Requires . Bayes' theorem, , with the denominator supplied by total probability. Equivalently, posterior odds equal the likelihood ratio times the prior odds. The posterior depends on the prior: a -accurate test on a disease of prevalence gives a posterior of only , and a second conditionally independent positive raises it to .
- Forgetting to subtract the overlap in . Additivity is an axiom only for mutually exclusive events; for any others the correction is compulsory.
- Confusing mutually exclusive with independent. For events of positive probability they are incompatible: exclusivity forces while independence forces .
- Assuming independence rather than justifying it. Independence comes from the design — separate tosses, random sampling, randomised assignment — never from the events feeling unrelated.
- Checking independence only in pairs. Three events can be pairwise independent and still fail mutual independence, as two coins and the event "they match" show.
- Swapping and . These answer different questions, and converting between them requires the prior; doing it without one is the prosecutor's fallacy.
- Applying the complement rule behind the bar. is false; the correct identity is .
- Reading a posterior off the sensitivity. With a rare disease the false positives from the large healthy group outnumber the true positives, which is why accuracy can yield a posterior.
- Using on a sample space whose outcomes are not equally likely, such as the eleven possible sums of two dice.
- Mixing ordered and unordered counts in the same fraction. Both and must use one convention.
- Double counting in "at least one" problems. Use the complement, or sum over disjoint cases by exact number.
- Choosing permutations when order does not matter, or combinations when it does. Ask whether rearranging the selection produces a genuinely different outcome.
- Conditioning on an event of probability zero. is undefined when ; the definition divides by it.
- Using cases that do not partition the sample space in the law of total probability. Overlapping cases double count; incomplete cases lose probability. Check that the weights sum to .
- Multiplying likelihood ratios for repeated tests that are not conditionally independent. Two runs of the same assay on the same sample are not two pieces of evidence.
- Treating a large relative update as a large absolute probability. Multiplying a probability by is dramatic and still leaves below .