Contents / Linear Algebra / Vectors and Euclidean Space
Chapter 2
Vectors and Euclidean Space
Vector arithmetic, the dot and cross products, projections, and the geometry of lines and planes.
Introduction
Vectors are the atomic units of linear algebra. Everything that follows — matrices, linear transformations, eigenvalues, the singular value decomposition — is built on top of them. Before any of that makes sense you must be fluent with what a vector is, how vectors combine, and what the dot product measures.
There are two views of a vector, and a first course keeps switching between them on purpose. Geometrically, a vector in or is an arrow with a length and a direction. Algebraically, it is an ordered list of real numbers. The geometric view supplies the pictures and the intuition; the algebraic view supplies the computations and, crucially, keeps working in or , where there is no picture to draw. A great deal of linear algebra consists of proving that some geometric statement you believe about arrows in the plane is a consequence of algebra alone, and therefore true in every dimension.
The chapter runs in the order that dependency forces. First vectors themselves and their length. Then the two operations — addition and scalar multiplication — with the eight axioms they satisfy, because those axioms are the definition of a vector space that the later chapters generalise. Then the dot product, which is what turns from a bare list-space into a space with lengths and angles; the Cauchy–Schwarz inequality is the hinge on which that whole section turns. Then projection, which is the single most reused construction in the subject. Then the cross product and the scalar triple product in , which measure area and volume. Finally lines and planes, where all of the above is cashed out as concrete geometry.
2.1Vectors in
Definition 2.1 (Vector in ). Let be a positive integer. A vector in is an ordered -tuple of real numbers,
and denotes the set of all of them. The numbers are the components of . Two vectors are equal exactly when they have the same number of components and every corresponding pair of components is equal.
Equality is component-wise and nothing more: if and only if and . A vector in is never equal to a vector in , even one that looks like it with a zero tacked on; they are elements of different sets and no operation in this chapter mixes them. The most common first error in the subject is adding or dotting two vectors of different lengths, and the definition is what forbids it.
The geometric picture in low dimensions is an arrow. In , draw as an arrow whose tail sits at the origin and whose head sits at the point ; in , the same with three coordinates. Crucially, what identifies a vector is its displacement, not its starting point. An arrow from to has the same components — three across, four up — as the arrow from the origin to , so it is the same vector. Arrows that are translates of one another are one vector drawn in different places. This is exactly why vectors model velocities, forces and displacements: those quantities have size and direction but no location.
That said, it is often convenient to pin an arrow to the origin, and then a vector and a point become interchangeable. The position vector of a point is . Whether means "the point " or "the instruction: move three right and four up" is decided by context, and fluency means holding both at once.
Notation. Bold lower-case letters denote vectors and plain letters denote scalars (real numbers). Components are written with subscripts: . The arrow denotes the vector from the point to the point . Angle brackets and parentheses are both in common use for vectors; this chapter uses angle brackets for vectors and parentheses for points, but no mathematical distinction rests on the choice.
Definition 2.2 (Zero vector and negative). The zero vector in is . The negative of is .
The zero vector is a genuine vector, not a missing one, and it is the only vector with no direction: every other vector points somewhere. It will turn out to be the additive identity, and a surprising amount of linear algebra is the study of which vectors a given process sends to .
Two points determine a vector, and the recipe is head minus tail.
Proposition 2.3 (Displacement between points). If and , then
Proof. Travelling from to changes the -th coordinate from to , a change of . The displacement whose components are those changes, applied at , lands exactly at , and by definition is that displacement.∎
Subtract in that order and not the other: , and a sign error here propagates into every projection and distance computed afterwards.
Now the length. In the plane, the arrow is the hypotenuse of a right triangle with legs and , so Pythagoras gives its length as . In a second application of Pythagoras gives . Beyond there is no triangle to appeal to, so the same formula is promoted from a theorem to a definition.
Definition 2.4 (Norm). The norm (length, magnitude) of is
A vector with is a unit vector.
Two facts are immediate from the formula and worth naming, because everything later leans on them. First, always, since it is a square root of a sum of squares; and forces every , hence . So the only vector of length zero is the zero vector. Second, squaring removes the root:
and in practice you should work with for as long as possible and take the square root only at the end. Almost every proof in this chapter is an identity about squared norms.
This is the right moment to say what geometry in means for large , because students reasonably suspect a swindle. There is no picture of . What there is, is a definition of length that agrees with the visual one when , and a body of theorems — Cauchy–Schwarz, the triangle inequality, the Pythagorean theorem, the projection formula — proved from that definition by algebra that never mentions except as an index of summation. So when we later speak of the angle between two vectors in , we are not claiming to see it. We are saying: the number lies in (that is Cauchy–Schwarz, and it needs proof), so it is the cosine of exactly one angle in , and that angle obeys every law that angles obey in the plane. Any two vectors in span at most a -dimensional plane inside it, and inside that plane the picture is the ordinary one. That is the whole content of "high-dimensional geometry", and it is why the algebraic definitions are the ones to trust.
Intuition. A vector is a set of walking directions: "go blocks east and blocks north." The directions do not say where you start — only what steps to take — which is why the same vector can be drawn anywhere on the map. The norm is the straight-line distance from where you started to where you ended up: blocks by Pythagoras, even though you walked blocks of pavement to get there.
In the same sentence still makes sense if you imagine dials instead of a map: a vector is a list of adjustments, and its norm is the total size of the adjustment, measured so that one big change and a spread of small ones can be compared on the same scale.
Example 2.5 (Displacement and distance). Let and . Find and the distance from to .
Solution. Head minus tail, component by component:
Its norm is
and the norm of a displacement vector is exactly the straight-line distance between its endpoints, so and are apart.
Sanity check: the third coordinates of and agree, so the two points are at the same height and the third component of the displacement had to come out — it did. The answer is positive, as a distance must be.□
Example 2.6 (A norm in ). Find for , and find the two scalars for which .
Solution. Squares first, and note that the minus sign disappears when squared:
Scaling multiplies every component by , so it multiplies by and the norm by — this is proved as a theorem in the next section. Hence gives , so or .
Sanity check: there must be two answers, because and have equal length; a single answer would have signalled a dropped absolute value.□
Example 2.7 (Equidistant points). Find all values of for which the point is at distance from .
Solution. The displacement is , so
Setting this equal to gives , so and .
Sanity check: squaring the distance rather than solving with a square root in the way is what made this a one-line quadratic. Both answers are real because : the sphere of radius about does reach the line , and it must cut it twice.□
Pitfall. is not . The norm squares, sums and then takes a root, in that order. For the norm is , not . The "sum of absolute values" is a perfectly good measurement of size — it is the distance a taxi drives on a grid — but it is not the one that matches Pythagoras, and it is not the one the dot product induces.
2.2Vector addition, scalar multiplication and span
Definition 2.8 (Vector addition and scalar multiplication). For and a scalar , both operations act component-wise:
Subtraction is defined by .
Geometrically, addition is the triangle rule: slide the tail of to the head of ; the sum is the arrow from the tail of to the head of . Equivalently it is the parallelogram rule: draw and from a common tail, complete the parallelogram, and is the diagonal from that shared corner. The two descriptions give the same arrow because the opposite sides of a parallelogram are translates of each other, and translates are the same vector. The other diagonal is , drawn from the head of to the head of — which is just the displacement formula again, since when and are the heads.
Scalar multiplication stretches. If the arrow gets longer; if it shrinks; if it reverses and then scales by ; if it collapses to . In every case the arrow stays on the same line through the origin, which is what "same or opposite direction" means precisely.
Theorem 2.9 (Homogeneity of the norm). For every and every scalar ,
Proof. Compute directly:
Taking non-negative square roots of both sides gives , using and not .∎
The absolute value is the whole content of the theorem. Lengths cannot be negative, so doubling-and-reversing a vector of length produces a vector of length , not .
Now the axioms. Everything algebraic that you will ever do with vectors in this course follows from eight properties, and they are worth verifying once, concretely, in — because the later chapter on vector spaces takes exactly this list and declares it the definition of the subject. The point of that abstraction is that polynomials, matrices, sequences and functions satisfy the same eight rules, so every theorem proved from them applies to all of those at once. Here we check that really is the model case.
Theorem 2.10 (The eight vector space axioms for ). For all and all scalars :
- (addition is commutative);
- (addition is associative);
- (additive identity);
- (additive inverse);
- (distributive over vector addition);
- (distributive over scalar addition);
- (compatibility of scalar multiplication);
- (unit scalar acts trivially).
Proof. Every operation is defined component-wise, so each axiom reduces to the corresponding law for real numbers applied in each of the slots. Writing the -th component of each side:
- — commutativity of addition in .
- — associativity in .
- — is the additive identity in .
- — additive inverses in ; the resulting vector has every component , hence equals .
- — distributivity in .
- — distributivity in again, with the sum in the scalar slot.
- — associativity of multiplication in .
- — is the multiplicative identity in .
Since two vectors are equal precisely when all components agree, each line establishes the corresponding vector identity.∎
The proof is short, but do not mistake shortness for emptiness. Two things are being asserted. The first is that the axioms hold — that is the routine part. The second is that they hold for the reason given: because the operations are defined slot by slot and is a field. Change either fact and the list can fail. If you defined as your "addition", axiom 4 breaks immediately: has no inverse. If you took your scalars from rather than you would keep all eight axioms but lose the ability to normalise a vector, which is why the scalars must form a field.
Three familiar consequences are not on the list, and it is a good exercise in axiomatics to see that they do not need to be: , , and .
Proposition 2.11 (Consequences of the axioms). In , for every vector and scalar : , , and .
Proof. By axiom 6, . Adding to both sides and using axioms 2, 4 and 3 gives . The same argument with axiom 5 and gives . Finally, by axioms 8, 6 and the first result,
so is an additive inverse of ; inverses are unique (if and are both inverses of then ), so .∎
Notice that this proof never looked at a component. That is deliberate: it is the argument that will still work in the abstract chapter, where there are no components to look at.
Any nonzero vector can be rescaled to length .
Definition 2.12 (Normalisation). For , the unit vector in the direction of is
It really is a unit vector: by homogeneity, , where the scalar is positive so its absolute value is itself. The hypothesis is not decoration: the zero vector has norm , the formula divides by it, and — more to the point — has no direction to normalise.
Definition 2.13 (Standard basis vectors). In , let be the vector with a in slot and everywhere else. In these are also written
Every vector decomposes along them: , and in general . That expression is the prototype of the most important construction in the subject.
Definition 2.14 (Linear combination and span). A linear combination of is any vector
Their span, written , is the set of all such combinations.
Span is the answer to "where can I get to?" and its geometry in is worth memorising, because it is the picture behind solution sets of linear systems in the next chapter.
The span of the empty collection, or of , is : a single point, the origin. The span of one nonzero vector is , the line through the origin in the direction of . The span of two nonzero, non-parallel vectors is a plane through the origin: the parallelogram rule lets you reach every point of it, and no more, since every combination lies in the flat sheet they determine. If the two vectors are parallel, the second adds nothing and the span collapses back to a line. The span of three vectors in is all of when they do not lie in a common plane, and is a plane or a line or the origin when they do. Notice that every span contains — take all — so spans are flat objects through the origin, never shifted ones. A line that misses the origin is not a span, and that distinction is the whole difference between a subspace and an affine set.
Definition 2.15 (Linear independence). A list is linearly independent when the only scalars satisfying
are . Otherwise the list is linearly dependent.
Proposition 2.16 (Dependence means redundancy). A list with is linearly dependent if and only if at least one is a linear combination of the others. In that case removing does not change the span.
Proof. () Suppose with some . Solve for that term:
which is division by a nonzero scalar and so is legitimate. () If , move everything to one side to get , a combination whose -th coefficient is , hence nontrivial.
For the last claim, any combination using can have that term rewritten via the displayed formula, producing a combination of the remaining vectors only; so the two spans contain each other.∎
Two special cases are used constantly. Any list containing is dependent, since together with zero coefficients on everything else is a nontrivial combination giving . And two vectors are dependent exactly when one is a scalar multiple of the other — that is, when they are parallel.
Intuition. Adding vectors is following two sets of directions back to back: " east, north" then " east, south" leaves you at " east, north." Scalar multiplication is changing gear: multiply by and you cover twice the ground the same way; multiply by and you retrace your steps.
Independence is a team with no passengers. If three people each bring a genuinely new move, no one of them can be imitated by combining the other two. A dependent list has a passenger — one vector you could rebuild from the rest — so dropping it costs you nothing, which is exactly what the proposition above says about the span.
Example 2.17 (Arithmetic and normalisation). Let and . Compute , and find the unit vector in the direction of .
Solution. Scale first, then add component-wise:
For the unit vector, , so and
Sanity check: , and the signs of match those of because the scalar is positive.□
Example 2.18 (Membership in a span). Is in ?
Solution. Ask whether scalars exist with . Componentwise this is the system
The first equation forces and the third forces ; those are not choices, so the second equation is a consistency check rather than another equation to solve. It reads , which is true.
So and the answer is yes. Sanity check: recompute the combination, . Had the middle equation failed, the honest conclusion would have been that no such exist and lies off the plane spanned by the two vectors.□
Example 2.19 (Testing a list for dependence). Are , and independent?
Solution. By Proposition (Dependence means redundancy) it is enough to ask whether one is a combination of the others, and is the natural candidate. Solve .
The first component gives , since contributes nothing there. The third gives , so . Check the second component with , : , and the second component of is indeed .
So , equivalently — a nontrivial combination equal to zero. The list is dependent, and , a plane through the origin rather than all of .□
Example 2.20 (A parameter that creates dependence). For which is the list , dependent?
Solution. Two vectors are dependent exactly when one is a scalar multiple of the other. Comparing first components, the only possible multiplier is ; the second components agree with it (); so dependence requires .
For the list is : dependent, spanning a line. For every other the two vectors are independent and span a plane. Sanity check: at the combination is nontrivial, as required.□
Pitfall. "Linearly dependent" does not mean "all parallel". The list is dependent even though no two of its members are parallel — the dependence involves all three at once. Whenever , a list of vectors in is automatically dependent, a fact proved in the chapter on dimension.
2.3The dot product, length and angle
Addition and scalar multiplication cannot express a single geometric fact: nothing in the eight axioms mentions length or angle. One extra operation supplies both.
Definition 2.21 (Dot product). The dot product of is the scalar
The output is a number, not a vector. That is worth saying out loud because the cross product later in this chapter takes two vectors and returns a vector, and confusing the two is the single most common slip in the subject. An expression such as is meaningless: its left factor is a scalar and scalars have no dot product.
Theorem 2.22 (Algebraic properties of the dot product). For all and all scalars :
- (symmetry);
- (distributivity);
- (homogeneity);
- , with equality if and only if (positive definiteness).
Proof. All four are sums of scalar identities.
- , because multiplication in is commutative.
- , by distributivity in and the fact that a finite sum splits.
- , pulling the constant out of the sum.
- , which is by the definition of the norm. A sum of squares of real numbers is non-negative and vanishes only when every term vanishes, that is, only when every .
Properties 1–3 say the dot product is bilinear and symmetric: you may expand brackets in dot products exactly as in ordinary algebra, provided you never try to dot a scalar. Property 4 is the bridge to geometry, since it lets every norm be rewritten as a dot product and back:
Combining bilinearity with property 4 gives the expansion that the rest of this section runs on. For any ,
and replacing by ,
These are the vector analogues of , and the cross-term is where all the geometry hides.
Now the central inequality of the subject.
Theorem 2.23 (Cauchy–Schwarz inequality). For all ,
with equality if and only if one of the vectors is a scalar multiple of the other.
Proof. If both sides are and , so the statement holds; assume .
For every real , positive definiteness gives . Expand using bilinearity:
Because the leading coefficient is strictly positive, so is a genuine upward parabola. A parabola that is never negative has at most one real root, so its discriminant is at most zero:
that is, . Taking non-negative square roots gives the inequality.
Equality holds exactly when the discriminant is zero, which happens exactly when has a (double) real root . But forces , again by positive definiteness, so is a multiple of . Conversely if then both sides equal .∎
This proof is worth studying rather than memorising, for two reasons. It never mentions angles, so it cannot be accused of assuming the geometry it is about to justify; and it uses nothing about beyond the four properties of the dot product, so the identical argument proves Cauchy–Schwarz for any inner product space — for integrals of products of functions, for covariances of random variables. In statistics this inequality is why a correlation coefficient never exceeds in absolute value.
Corollary 2.24 (Triangle inequality). For all ,
Proof. Start from the expansion of and bound the cross term using :
Both and are non-negative, and is increasing, so taking square roots preserves the inequality.∎
Geometrically this says the direct route is never longer than a detour: the third side of a triangle is at most the sum of the other two. Equality requires , which by the equality case of Cauchy–Schwarz means the vectors are parallel and pointing the same way — the "triangle" is flat.
With homogeneity and the triangle inequality in hand, the norm satisfies the three properties that define a norm on any vector space.
Theorem 2.25 (The norm axioms). The function on satisfies, for all and all scalars :
- , and if and only if (positivity);
- (absolute homogeneity);
- (triangle inequality).
Proof. Property 1 is property 4 of the dot product after a square root; property 2 is the homogeneity theorem of the previous section; property 3 is the corollary just proved.∎
Any function on a vector space obeying those three rules is called a norm, and the one above — the Euclidean norm — is only the most familiar. The point of isolating the axioms is that every consequence of them, in particular every statement about distances, holds for all of them at once.
Definition 2.26 (Distance in ). The distance between and in is
Theorem 2.27 (The metric axioms). For all :
- , with equality if and only if ;
- (symmetry);
- (triangle inequality).
Proof. 1 follows from norm positivity applied to , which is exactly when . For 2, , so absolute homogeneity with gives equal norms. For 3, write and apply the norm triangle inequality to that sum.∎
The third axiom is the one with content, and it is the reason "shortest path" means anything in : inserting a waypoint can never shorten a journey.
Cauchy–Schwarz now lets us define angles, which is the payoff.
Definition 2.28 (Angle between vectors). For nonzero , the angle between them is the unique with
This definition is legitimate precisely because Cauchy–Schwarz guarantees the quotient lies in , and is a bijection from onto . Without the inequality the right-hand side might not be a cosine of anything, and the definition would be empty. In and it agrees with the angle you would measure with a protractor:
Theorem 2.29 (Geometric form of the dot product). If is the angle between nonzero and , then
In and , is the ordinary geometric angle at the common tail of the two arrows.
Proof. The displayed identity is the definition of rearranged. For the second claim, place and with a common tail and let be the geometric angle between them. The three arrows , and form a triangle with sides and with opposite the last side, so the law of cosines gives
The algebraic expansion computed earlier gives . Comparing the two and cancelling, . Since and is injective there, .∎
Read off the sign rule at once. Because , the sign of is the sign of : positive means (the vectors broadly agree), negative means (they broadly disagree), and zero means exactly.
Definition 2.30 (Orthogonality). Vectors and in are orthogonal when . A set of vectors is orthogonal when every pair in it is.
The definition is stated with the dot product rather than with the angle so that it also covers : the zero vector is orthogonal to everything, even though it makes no angle with anything. That convention is not a dodge — it is what makes statements like "the orthogonal complement is a subspace" true without exceptions.
Theorem 2.31 (Pythagorean theorem in ). Vectors are orthogonal if and only if
Proof. The expansion holds always. Subtracting from both sides, the stated equation is equivalent to , i.e. to .∎
Note that this is an "if and only if", which the school version is not: in , the Pythagorean identity characterises perpendicularity. The same expansion yields another identity worth knowing.
Proposition 2.32 (Parallelogram law). For all ,
Proof. Add the two expansions: the cross terms and cancel, leaving .∎
In words: the sum of the squares of a parallelogram's two diagonals equals the sum of the squares of its four sides. Remarkably, this identity characterises norms that come from a dot product; a norm failing it cannot be written as for any inner product.
Intuition. The dot product measures agreement of direction, scaled by both lengths. Point two torches the same way and the beams reinforce: the dot product is as large as the lengths allow. Turn one through a right angle and the beams share nothing: zero. Turn it all the way round and they cancel: maximally negative.
Cauchy–Schwarz is the quantitative version of "agreement cannot exceed total size". With and , the dot product must lie in , and it hits only in the one case where the vectors point exactly the same way.
Example 2.33 (Angle between two vectors). Find the angle between and .
Solution. Dot product first:
Norms: and . Then
using , and .
Sanity check: the dot product is positive, so the angle had to come out acute — it did. And , as Cauchy–Schwarz demands; a cosine above would prove an arithmetic slip.□
Example 2.34 (Forcing orthogonality). Find all for which and are orthogonal.
Solution. Orthogonality is one scalar equation:
so . Sanity check: with , and .
Note that the condition was linear in , so there is exactly one solution. Orthogonality to a fixed nonzero vector always cuts a hyperplane out of , and here the unknown moved along a line, which meets that hyperplane once.□
Example 2.35 (Working with norms without components). Suppose , and the angle between them is . Compute and .
Solution. By the geometric form,
Then the expansion of a squared difference gives
so .
Sanity check: the triangle inequality predicts , and the reverse triangle inequality predicts it is at least . Since , the answer is in range. No components were ever needed: the identities do all the work.□
Example 2.36 (A dependence hidden in the numbers). Show that and are orthogonal, and find the area of the rectangle they span.
Solution. Dot them: , so they are orthogonal.
Because the sides meet at a right angle, the parallelogram they span is a rectangle with side lengths
so its area is .
Sanity check: the Pythagorean theorem predicts ; directly, and . Consistent.□
Pitfall. does not mean one of the vectors is . Unlike real numbers, vectors have genuine "zero divisors" for this operation: with neither factor zero. Cancellation fails for the same reason — from you may conclude only that is orthogonal to , not that .
2.4Orthogonal projection
Here is the question. Given a nonzero direction , how much of a vector points along ? The answer should be a vector on the line , and the remainder should contain no trace of the direction — meaning it should be orthogonal to . Those two requirements determine completely, and the derivation is three lines.
Write for an unknown scalar ; this encodes the first requirement. The second requirement says
Expand by bilinearity: . Since we have and may divide:
The formula was not guessed; it was forced. And the derivation simultaneously proves that such a exists and that there is only one.
Definition 2.38 (Vector and scalar projection). Let . The vector projection of onto and the scalar projection (the signed length of the shadow) are
Keep the two straight. The vector projection is a vector parallel to ; the scalar projection is a number, and it is signed — negative when leans away from , which is exactly when the angle exceeds . They are related by , and consequently .
Note also that depends only on the line through , not on itself: replacing by multiplies the numerator by and the denominator by , and the extra factor restores the balance. The scalar projection does depend on the sign of but not on its length. And is indispensable: the zero vector determines no line to project onto, and the formula would divide by zero.
Theorem 2.39 (Orthogonal decomposition). Let . Every can be written in exactly one way as
with a scalar multiple of and orthogonal to . Explicitly,
Proof. Existence. Define and by the displayed formulas. Their sum is , and is a multiple of by construction. Orthogonality is the computation that produced the formula in the first place:
Uniqueness. Suppose with . Then . Dot both sides with : the right side gives , and the left gives . As , we get and hence .∎
Because and are orthogonal, the Pythagorean theorem applies to the decomposition:
which is a useful check on any projection computation — the two pieces must account for the original length, and neither piece can be longer than itself.
The projection is also the closest point of the line to , which is what makes it the right notion of approximation.
Theorem 2.40 (Best approximation on a line). Let and . Then for every scalar ,
with equality only when . Consequently the distance from the tip of to the line is .
Proof. Write and split . The first bracket is , orthogonal to and hence to , so Pythagoras gives
The extra term is zero exactly when .∎
That "drop a perpendicular and you land on the nearest point" is one line of algebra, and it generalises verbatim from a line to any subspace. It is the engine of least-squares regression, of Fourier series (where the "directions" are sines and cosines), and of the Gram–Schmidt process, which builds an orthogonal set by repeatedly subtracting off projections.
Intuition. Shine a light straight down on a tilted rod resting against a wall. The shadow on the floor is the projection of the rod onto the floor's direction: it records how much of the rod's length runs horizontally. Stand the rod upright and the shadow shrinks to a point — zero projection, the rod is orthogonal to the floor. Lay it flat and the shadow is the whole rod.
The perpendicular piece is the part of the rod the shadow loses, and its length is exactly the height of the rod's tip above the floor — that is the "distance to the line" claim, seen from the side.
Example 2.41 (Projection and decomposition). Project onto , and write as a sum of a piece parallel to and a piece orthogonal to it.
Solution. Two dot products:
Hence
Check orthogonality: .
Check lengths: and . Their sum is , as the Pythagorean identity requires.□
Example 2.42 (Distance from a point to a line through the origin). Find the distance from the point to the line through the origin in the direction .
Solution. Let be the position vector of . By the best-approximation theorem the distance is .
Sanity check: the line is , and the standard plane-geometry formula for the distance from to is . Agreement.□
Example 2.43 (A negative scalar projection). Find and for and , and interpret the sign.
Solution. The two ingredients are
Hence
The negative sign says the angle between and exceeds : the shadow falls on the ray opposite to , and its length is .
Sanity check: , which equals as it must. And , since a projection can never be longer than the vector it came from.□
Pitfall. and are different vectors — they do not even point the same way. The first is parallel to , the second parallel to . Read the subscript as "onto", and when a problem says "the projection of onto ", the vector you divide by is .
2.5The cross product, area and volume
Everything so far worked in every . This section does not: it is special to , where — and essentially only where — there is a product of two vectors that returns a vector.
Definition 2.44 (Cross product). For , the cross product is
The determinant in the first expression is a mnemonic — its first row holds vectors, not scalars — but expanding it along the top row produces exactly the three components on the right, and that is how everyone computes it in practice. Note the middle component: the pattern is , with the indices before . Writing there is the most frequent computational error with this formula, and it silently flips the sign of one component.
Theorem 2.45 (The cross product is orthogonal to both factors). For all ,
Proof. Compute the first directly:
Expanding, the six terms are , which cancel in pairs. The second identity is the same computation with the roles of and exchanged, or follows from it together with anticommutativity below.∎
This is the property that makes the cross product useful: it manufactures a normal direction to a plane out of two vectors lying in it, which is exactly what the last section of this chapter needs.
Theorem 2.46 (Algebraic properties of the cross product). For all and scalars :
- (anticommutativity);
- ;
- (distributivity);
- ;
- .
Proof. Each component of has the form . Swapping and turns this into , the negative of the original, which is property 1. Setting in property 1 gives , so and property 2 follows. For 3, the first component of the left side is , which expands to , the sum of the first components on the right; the other two components are identical in form. Property 4 follows because each component is linear in the entries of each factor, and 5 because every component has a factor from .∎
Two properties are conspicuously missing. The cross product is not commutative — property 1 says reversing the order reverses the vector — and it is not associative: with we get , while . So the expression is meaningless without brackets.
The length of is where the geometry is, and it comes from an identity that is pure algebra.
Lemma 2.47 (Lagrange's identity). For all ,
Proof. Expand . Each square contributes two "square" terms and one cross term, giving
where the first sum runs over the six ordered pairs with . Meanwhile
Adding the two lines, the cross terms cancel and the squares combine into over all nine pairs, which factors as .∎
Theorem 2.48 (Magnitude of the cross product). If is the angle between nonzero , then
Moreover if and only if and are parallel.
Proof. By Lagrange's identity and the geometric form of the dot product,
Take square roots; since we have , so no absolute value is needed. The cross product vanishes iff (or a factor is ), i.e. iff , which is exactly parallelism.∎
The direction of among the two unit normals is settled by the right-hand rule: point the fingers of your right hand along , curl them towards , and your thumb points along . Consistently, , , — cyclic in , with a sign change if you break the cycle.
Corollary 2.49 (Area of a parallelogram and a triangle). The parallelogram with adjacent sides has area , and the triangle with those two sides has area .
Proof. Take as the base, of length . The height is the distance from the tip of to the line through , which is . Base times height is . The triangle is half the parallelogram.∎
Multiplying by a third vector recovers volume.
Definition 2.50 (Scalar triple product). For , the scalar triple product is the number .
Theorem 2.51 (Triple product as a determinant).
Consequently the triple product is unchanged by a cyclic shift, , and changes sign when any two of the three vectors are swapped.
Proof. Dotting with gives
after rewriting the middle component as . That expression is precisely the cofactor expansion of the displayed determinant along its first row.
A cyclic shift of is two row swaps of the determinant, and each swap changes its sign, so two swaps restore it. A single swap changes the sign once.∎
Theorem 2.52 (Volume of a parallelepiped). The parallelepiped with adjacent edges has volume
Proof. Take the parallelogram spanned by and as the base; by the area corollary its area is . The solid's height is the distance from the tip of to the plane of the base, measured along the normal direction ; that is the absolute scalar projection
Volume is base area times height, and the factor cancels.∎
Corollary 2.53 (Coplanarity test). Three vectors lie in a common plane through the origin — equivalently, they are linearly dependent — if and only if
Proof. The parallelepiped they span degenerates to a flat figure exactly when its volume is , and by the volume theorem that happens exactly when the triple product vanishes. Flatness means all three edges lie in one plane through the common corner. If they do, one of them is a combination of the other two (or two of them are already parallel), so the list is dependent; conversely a dependent list has one vector in the span of the others, which is a plane or less.∎
Intuition. The cross product answers "which way is up?" given two arrows lying on a table, and how much table they enclose. Parallel arrows enclose no area, so their cross product is the zero vector — that is the same "no new direction" idea as linear dependence, wearing different clothes.
The triple product then asks how far a third arrow lifts off that table. If it stays on the table, the box you build has no height and no volume — and the three arrows are dependent. That is why one determinant answers a geometric question (volume) and an algebraic one (dependence) at the same time.
Example 2.54 (Cross product, area, and a normal direction). For and , compute , the area of the parallelogram they span, and the area of the triangle with those sides.
Solution. Component by component, watching the middle sign:
Then
so the parallelogram has area and the triangle has area .
Sanity check: dot the result with each factor. and . Both vanish, as the orthogonality theorem requires.□
Example 2.55 (Volume of a parallelepiped). Find the volume of the parallelepiped with edges , , .
Solution. Expand the determinant along the first row:
The volume is .
Sanity check: the value is nonzero, so the three edges are not coplanar and a genuine solid exists — consistent with the fact that no one of the three vectors is visibly a combination of the others. Had the sign come out negative, the volume would still be ; the sign only records the handedness of the edge ordering.□
Example 2.56 (Testing three vectors for coplanarity). Are , , coplanar?
Solution. Compute the triple product as a determinant along the first row:
The volume is zero, so the three vectors are coplanar and the list is linearly dependent.
Sanity check: find the dependence explicitly. Indeed , so is a nontrivial combination — dependence confirmed without any determinant.□
Example 2.57 (Area of a triangle from three points). Find the area of the triangle with vertices , , .
Solution. Convert points to edge vectors at :
Cross them:
Its norm is , so the triangle's area is .
Sanity check: and , so the computed vector really is normal to the triangle's plane. This normal is used again in the next section to write down the plane's equation.□
Pitfall. The cross product exists only in (and, in a different guise, ). Writing is meaningless. If a plane problem calls for a perpendicular in , use the rotation instead, which is orthogonal to because .
2.6Lines and planes in
Everything in this chapter now pays out as concrete geometry. A line is determined by a point and a direction; a plane by a point and a normal. Both descriptions turn straight into equations.
Definition 2.58 (Vector and parametric equations of a line). The line through the point with position vector and direction is the set of points with position vector
Writing and , the parametric equations are
The vector form says exactly what a line is: start at and travel any multiple of . The parameter is a coordinate along the line; gives and negative goes the other way. Neither nor is unique — any point of the line will do for , and any nonzero multiple of for the direction — so two very different-looking parametrisations can describe the same line. To decide whether they do, check that the directions are parallel and that one line's base point satisfies the other's equations.
Eliminating gives a description with no parameter at all.
Proposition 2.59 (Symmetric equations of a line). If are all nonzero, the line above consists of exactly the points with
Proof. From the parametric equations, , and whenever the denominators are nonzero. A point lies on the line precisely when some single produces all three coordinates, that is, precisely when the three expressions agree.∎
If one component of is zero — say — that coordinate is constant along the line and the symmetric form is written as , . Never write a zero denominator.
Now planes. A plane is pinned down not by directions inside it but by the one direction perpendicular to it.
Definition 2.60 (Normal vector). A normal vector to a plane is a nonzero vector orthogonal to every vector lying in that plane.
Theorem 2.61 (Point–normal and general equations of a plane). Let and let . The plane through with normal is the set of points with position vectors satisfying
equivalently
with . Conversely, every equation with describes a plane with normal .
Proof. A point lies in the plane exactly when the displacement lies in the plane's directions, which by the definition of a normal happens exactly when . Expanding that dot product componentwise gives the second form, and moving the constants to the right gives the third.
For the converse, given with, say, , the point satisfies it, and subtracting from returns the point–normal form with .∎
The coefficients of are the normal vector: that single observation answers most plane questions instantly. Two planes are parallel when their normals are parallel; they are perpendicular when their normals are orthogonal; and a line is parallel to a plane when its direction is orthogonal to the plane's normal.
Three non-collinear points determine a plane, and the cross product is how you find its normal.
Method 2.62 (Plane through three points). Given non-collinear , , :
- Form two edge vectors and .
- Set ; it is nonzero because the points are not collinear.
- Write and expand to .
- Check by substituting all three points.
Definition 2.63 (Angle between two planes). The angle between two planes is the angle with
where are normals to the two planes.
The absolute value is there because a plane has two opposite normals, so without it the answer would depend on an arbitrary sign; taking the modulus reports the acute angle, which is the one meant by "the angle between two planes". The same convention gives the angle between two lines from their direction vectors.
The two distance formulas are both projections in disguise.
Theorem 2.64 (Distance from a point to a plane). The distance from the point to the plane is
Proof. Let and let be any point of the plane, so . The distance from to the plane is measured along the normal direction, so it is the absolute scalar projection of onto :
which is the stated formula. That this really is the distance — i.e. that the nearest point of the plane is the foot of the perpendicular — is the best-approximation theorem applied to the decomposition along : moving within the plane changes only , and by Pythagoras any such change can only lengthen the total. Note the answer does not depend on which was chosen, since only survives.∎
Theorem 2.65 (Distance from a point to a line in ). The distance from a point to the line through with direction is
Proof. Put and let be the angle between and . The perpendicular distance from to the line is the leg opposite in the right triangle with hypotenuse , namely . By the magnitude theorem, , so dividing by gives exactly .
Equivalently, by best approximation; the cross-product form is the same number packaged to avoid computing the projection.∎
Intuition. A plane's equation is a single sentence: "your position, dotted with the normal, comes out to . " Every point of the plane gives the same reading; points on one side read more, points on the other read less, and the distance formula is just that reading's error divided by the length of the measuring stick — dividing by converts the reading into honest units of length.
Example 2.66 (Line through two points, in three forms). Find vector, parametric and symmetric equations for the line through and .
Solution. A direction is the displacement: , which may be simplified to since only the direction matters.
Vector form, based at :
Parametric: , , . Symmetric, solving each for :
Sanity check: gives , and substituting into the symmetric form gives .□
Example 2.67 (Plane through three points). Find an equation of the plane through , and .
Solution. The normal was computed in the previous section: with and ,
Point–normal form at :
Sanity check: gives ; gives ; gives . All three lie on it.□
Example 2.68 (Distance from a point to a plane). Find the distance from to the plane .
Solution. Here with . Evaluate the left-hand side at :
So
Sanity check: step one unit from along the unit normal , reaching . Substituting: . The point landed on the plane, so the distance is indeed .□
Example 2.69 (Where a line meets a plane). Find the point where the line crosses the plane .
Solution. Substitute the parametric coordinates , , into the plane equation:
Set , so and the intersection point is .
Sanity check: . A unique solution was expected because the direction is not orthogonal to the normal — their dot product is , not — so the line is not parallel to the plane and must meet it exactly once.□
Example 2.70 (Angle between two planes). Find the acute angle between the planes and .
Solution. Read the normals off the coefficients: and . Then
so .
Sanity check: the planes are not parallel (the normals are not multiples) and not perpendicular (the dot product is not ), so an angle strictly between and was expected.□
Example 2.71 (Distance from a point to a line in space). Find the distance from to the line through with direction .
Solution. First the displacement: . Cross it with the direction:
Its norm is , and , so
Sanity check: must not exceed , the distance from to the particular point on the line; indeed is smaller. The two agree only when is already perpendicular to , and here , so strict inequality was correct.□
Pitfall. The formula is only valid with the plane written as and the same coefficients used in . If the plane is given as , you may either use with , or divide the whole equation through by first — but never mix a reduced normal with an unreduced constant. The two consistent choices give the same distance; a mixed one does not.
Summary (The dot product, the cross product, and the geometry they give). In the dot product is symmetric, bilinear and positive definite, with . From those four properties alone follow Cauchy–Schwarz, with equality exactly when one vector is a multiple of the other; the triangle inequality; the Pythagorean characterisation ; and, because Cauchy–Schwarz puts the quotient in , the angle . For every splits uniquely as with , and is the distance from to the line — the projection is the nearest point on it.
Only in : is orthogonal to both factors, anticommutative but neither commutative nor associative, and Lagrange's identity gives , the area of the parallelogram on , vanishing exactly when they are parallel. The triple product is the determinant of the three rows, its absolute value is the volume of the parallelepiped, and it vanishes exactly when the three vectors are coplanar, that is, dependent. Cashed out: a line is with , a plane is , i.e. with read off the coefficients, and the two distance formulas and are projections in disguise.
- Confusing the dot and cross products. is a scalar and exists in every ; is a vector and exists only in . Expressions like are not wrong, they are meaningless.
- **Writing . ** The correct rule is ; norms are never negative, so with and the answer is .
- **Reading as "one of them is zero".** It means orthogonal. For the same reason, you cannot cancel a common factor out of a dot product equation.
- The middle component of the cross product. It is , not . Check every cross product by dotting the answer with both factors; both must give .
- Projecting the wrong way round. is parallel to , and the denominator is . Swapping the roles gives a different vector pointing in a different direction.
- **Normalising , or projecting onto it.** Both divide by zero. Check before either.
- Losing the absolute value in a distance. Both the point–plane and the point–line formulas take a modulus; a negative distance always signals a dropped .
- Assuming a span is a plane. The span of two vectors is a plane only when they are independent; if one is a multiple of the other, the span is a line. Always check for parallelism first.
- Mixing dimensions. and are both undefined. Vectors must live in the same before any operation combines them.
- Treating a line that misses the origin as a span. Every span contains . The line is a span only when is itself a multiple of .