Theorem 8.31 (Primary decomposition). Let be an complex matrix with distinct eigenvalues . Then
and each summand is -invariant with .
Contents / Linear Algebra / Generalized Eigenvectors and Jordan Form
Chapter 8
Minimal polynomials, the Cayley-Hamilton theorem, generalized eigenvectors, Jordan normal form, and the matrix exponential.
The eigenvalue chapter ended with a promise and a warning. The promise: if an matrix has independent eigenvectors, then in the basis they form it becomes diagonal, and every question about it — powers, exponentials, long-run behaviour — collapses into independent scalar questions. The warning: not every matrix has independent eigenvectors. The matrix
has the single eigenvalue repeated twice and only a one-dimensional space of eigenvectors. No change of basis will make it diagonal. It is not exotic, it is not a rounding artefact, and it will not go away.
This chapter says what to do instead. The answer is the Jordan normal form: every complex square matrix is similar to a matrix that is diagonal except for a scattering of 's directly above the diagonal, and that form is unique up to the order of its blocks. Getting there requires four genuinely new ideas — invariant subspaces, the minimal polynomial, generalized eigenvectors, and the structure of nilpotent operators — and each of them is worth having on its own. The Cayley–Hamilton theorem, one of the most quoted results in the subject, falls out along the way.
Throughout, we work over the complex numbers. That is not laziness: the whole theory depends on the characteristic polynomial factoring completely into linear factors, which over is guaranteed by the fundamental theorem of algebra and over is not. The last section says what survives if you insist on staying real.
Recall the two multiplicities attached to an eigenvalue. They are different numbers, and the whole chapter lives in the gap between them.
Definition 8.1 (Algebraic and geometric multiplicity). Let be with characteristic polynomial , and let be an eigenvalue of .
The algebraic multiplicity is the multiplicity of as a root of : the largest with .
The geometric multiplicity is the dimension of the eigenspace,
Theorem 8.2 (The multiplicities are ordered). For every eigenvalue of ,
Proof. The left inequality holds because an eigenvalue has at least one eigenvector. For the right, put and choose a basis of the eigenspace . Extend it to a basis of and let be the matrix with these columns. Since for , the first columns of are , so
for some blocks and . Similar matrices have the same characteristic polynomial, and the determinant of a block upper triangular matrix is the product of the determinants of the diagonal blocks, so
Thus divides , which says exactly that .∎
The proof is worth rereading, because the technique — take what you know, put it in the corner of a basis, and read the consequence off a block triangular matrix — is used four more times in this chapter.
Definition 8.3 (Defective). An eigenvalue is defective if . A matrix is defective if it has a defective eigenvalue, and diagonalizable otherwise.
Why "otherwise"? Because the eigenvectors belonging to distinct eigenvalues are always independent, so the eigenvectors of span a subspace of dimension , while the algebraic multiplicities always sum to (over , where factors completely). A full set of independent eigenvectors therefore exists precisely when for every . Every failure of diagonalizability is a missing eigenvector, and the deficit is measured by .
Intuition. Think of each eigenvalue as a room with chairs — that is how much of the -dimensional space it is entitled to. The eigenvectors are the people who actually show up: of them. A diagonalizable matrix fills every chair. A defective matrix leaves chairs empty, and the rest of this chapter is about who is allowed to sit in them: vectors that are not eigenvectors but are only one or two applications of away from being one.
Example 8.4 (A defective matrix). Show that is not diagonalizable.
Solution.
so is the only eigenvalue, with .
which has rank . Hence , spanned by .
That last step generalizes, and it is the fastest defectiveness test there is.
Proposition 8.5 (A repeated eigenvalue with no other). If has only one eigenvalue and , then is not diagonalizable.
Proof. If were diagonalizable, would be a diagonal matrix whose diagonal entries are the eigenvalues of , hence all equal to ; so and .∎
Example 8.6 (A matrix with a defective eigenvalue). Find both multiplicities of every eigenvalue of
Solution.
So and .
Row reduce: subtracting the third row from the second gives , and subtracting twice the third from the first gives . The reduced form has pivots in columns and , so and .
Example 8.7 (Same characteristic polynomial, different behaviour). The matrices
all have characteristic polynomial . Distinguish them.
Solution.
Pitfall. The characteristic polynomial is not a complete invariant. Two matrices with the same characteristic polynomial need not be similar, as , and above show. Conversely, similar matrices always have the same characteristic polynomial, the same eigenvalues with the same algebraic and geometric multiplicities, the same trace, determinant and rank.
Diagonalizing means breaking into one-dimensional pieces that preserves. When that is impossible, the next best thing is to break it into the smallest pieces does preserve.
Definition 8.8 (Invariant subspace). A subspace is invariant under (or -invariant) if for every . In that case restricts to a linear operator .
The trivial examples are and . The one-dimensional invariant subspaces are exactly the lines spanned by eigenvectors: is invariant precisely when is a multiple of . Between those extremes lie all the interesting ones.
Proposition 8.9 (Kernels and images of polynomials are invariant). For any polynomial , both and are -invariant.
Proof. The key fact is that commutes with , since is a sum of powers of . If then , so . If then .∎
Taking gives the subspaces this chapter is built from: for The case is the eigenspace; the larger are new.
Theorem 8.10 (Invariance means block triangular). Let be an -invariant subspace of dimension with . Choose a basis of , extend it to a basis of , and let . Then
where is the matrix of in the basis . Consequently
Proof. The -th column of is the coordinate vector of in the basis . For , invariance gives , so those coordinates involve only : the last entries of the column vanish. That is the claimed shape, and the top left block is by definition the matrix of the restriction. The determinant statement is the block triangular determinant formula applied to .∎
If we can split into invariant subspaces and take a basis adapted to the splitting, the off-diagonal block vanishes too and becomes block diagonal. That is the shape of the Jordan form, and the strategy of the chapter is now visible: find a decomposition of into invariant subspaces that are as small as possible.
Intuition. An invariant subspace is a room the dynamics cannot leave. If and lies in , then the whole future trajectory lies in , so understanding on is a self-contained problem. Splitting into invariant pieces is exactly splitting one -dimensional dynamical system into several smaller independent ones.
Example 8.11 (Invariant subspaces of a shift). Let , so , , . Verify that is invariant, and find the matrix of .
Solution.
Example 8.12 (An operator with no proper real invariant subspace). Show that the real rotation matrix has no -dimensional invariant subspace in , but does in .
Solution.
Theorem 8.13 (Existence of an eigenvector). Every complex matrix with has an eigenvalue and an eigenvector.
Proof. The characteristic polynomial is a polynomial of degree with complex coefficients, so by the fundamental theorem of algebra it has a root . Then , so is singular and has a nonzero null vector , which satisfies .∎
Corollary 8.14 (Triangularizability). Every complex square matrix is similar to an upper triangular matrix.
Proof. Induct on . For there is nothing to prove. For , take an eigenvector and let , a -dimensional invariant subspace. By the block triangular theorem there is with
with of size . By induction is upper triangular for some invertible ; conjugating by turns the display into an upper triangular matrix with in the corner and below it.∎
Triangular form already tells us a great deal — the eigenvalues sit on the diagonal — but it is far from unique, and it hides the geometric multiplicities. Jordan form is the triangular form with all remaining freedom squeezed out.
Theorem 8.15 (Cayley–Hamilton). Let be an matrix with characteristic polynomial . Then
the zero matrix. Explicitly, if , then .
Note the substitution convention. The constant term becomes , not the scalar , because the equation lives among matrices. With that understood, the statement says that satisfies its own characteristic equation.
The case can simply be checked.
Proposition 8.16 (Cayley–Hamilton for ). For with , we have .
Proof. Compute the square directly:
Subtract . The off-diagonal entries cancel exactly, and the diagonal entries become
So , which rearranges to .∎
Before the general proof, a warning about the argument almost everyone invents first.
Pitfall. The following is not a proof: ", so substituting gives . " Three things are wrong. Substituting a matrix for inside produces -style nonsense: is a matrix whose entries are scalars times , and replacing each by the matrix turns each entry into a matrix, so is no longer a matrix of numbers and its determinant is undefined. Second, even if the expression made sense, of it would be a scalar, not the zero matrix that must be. Third, the argument never uses anything about and would "prove" that for or any other polynomial obtained the same way. The identity is true, but it needs a real argument.
The real argument uses the adjugate, from the determinants chapter: for any square matrix ,
where the entry of is the cofactor of , a polynomial in the entries of .
Proof. Apply the adjugate identity to , whose entries are polynomials in of degree at most :
Each entry of is a cofactor of an submatrix, hence a polynomial in of degree at most . Collecting powers of , there are constant matrices with
Write . Expanding the left of and matching coefficients of each power of — legitimate because two matrices of polynomials are equal exactly when all their coefficient matrices agree — gives
Now multiply the equation indexed by on the left by and add everything up. The left-hand sides telescope:
because each cancels against the coming from the term indexed by . The right-hand sides add to . Hence .∎
Remark (A second proof, in one paragraph). Here is a completely different argument, worth knowing because it explains why the theorem is not an accident. If is diagonalizable, say with , then and because each is a root of . So the theorem is obvious for diagonalizable matrices. Now, the diagonalizable matrices are dense in the space of all complex matrices — any can be perturbed by an arbitrarily small amount to make its eigenvalues distinct, and distinct eigenvalues force diagonalizability — and the map is continuous, since it is built from sums and products of entries. A continuous map that vanishes on a dense set vanishes everywhere. Making "perturb to make the eigenvalues distinct" precise is a short exercise with the characteristic polynomial's discriminant; we take it as a sketch.
Cayley–Hamilton is not primarily a theoretical curiosity: it is a computational tool, because it expresses in terms of lower powers and therefore collapses every polynomial in to degree at most .
Example 8.17 (Reducing a high power). For , express as .
Solution.
Example 8.18 (The inverse from Cayley–Hamilton). Use Cayley–Hamilton to compute for .
Solution.
The same trick works in general: since with , the matrix is invertible exactly when , and then
Example 8.19 (A reduction). Let , whose characteristic polynomial is . Express in the form .
Solution.
and the check at now reads . At : . Both correct.□
Pitfall. The example above is left with its error in place on purpose: the arithmetic of a characteristic polynomial is the single most common place to go wrong, and checking the final identity at each eigenvalue costs ten seconds and catches it every time. Any polynomial identity must hold with replaced by each eigenvalue, because it holds on the corresponding eigenvector.
Cayley–Hamilton says satisfies some monic polynomial of degree . Often it satisfies one of much lower degree — satisfies — and the smallest such polynomial turns out to record exactly the information the characteristic polynomial throws away.
Definition 8.20 (Minimal polynomial). The minimal polynomial of a square matrix is the monic polynomial of least degree with .
Proposition 8.21 (Existence and uniqueness). Every square matrix has exactly one minimal polynomial, and for .
Proof. Existence: by Cayley–Hamilton the set of monic annihilating polynomials is non-empty, so it contains one of least degree. (Even without Cayley–Hamilton, the matrices are linearly dependent in the -dimensional space of matrices, which produces an annihilating polynomial.)
Uniqueness: if and are both monic of the same least degree with , then annihilates and has degree because the leading terms cancel. If were nonzero we could scale it to be monic, contradicting minimality of . So .
Finally, a monic polynomial of degree is the constant , and , so .∎
Theorem 8.22 (The minimal polynomial divides everything it should). If is any polynomial with , then . In particular divides the characteristic polynomial .
Proof. Divide with remainder: with or . Substituting ,
If , scaling it to be monic gives a monic annihilating polynomial of degree less than , contradicting minimality. So and . The last claim is Cayley–Hamilton plus this.∎
Theorem 8.23 (Same roots as the characteristic polynomial). A scalar is a root of if and only if is an eigenvalue of .
Proof. () Let with . For any polynomial one has , by applying repeatedly. Taking gives , and since , .
() Suppose , so with . By minimality , so there is a vector with . Then
so is an eigenvector with eigenvalue .∎
So and have the same roots; they differ only in the multiplicities. Since , and over the distinct eigenvalues, we get
Every exponent is at least , by the theorem just proved. The whole of Jordan theory is the statement that is the size of the largest Jordan block for .
Theorem 8.24 (Diagonalizability via the minimal polynomial). A complex square matrix is diagonalizable if and only if is a product of distinct linear factors, that is
with distinct.
Proof. () Suppose with diagonal and distinct diagonal entries . Put . Then is diagonal with entries , so and . Hence , and since must have all eigenvalues as roots, , a product of distinct linear factors.
() Suppose with distinct . Define the Lagrange polynomials
Each has degree and . The polynomial has degree at most and vanishes at the distinct points , so it is identically zero:
Fix and set . Then . Moreover
so each lies in the eigenspace . Every vector is therefore a sum of eigenvectors, so eigenvectors span and is diagonalizable.∎
Intuition. The characteristic polynomial counts how many chairs each eigenvalue owns; the minimal polynomial reports how badly each eigenvalue misbehaves. An exponent of means "this eigenvalue is perfectly well behaved: killing once is enough". An exponent of means "somewhere there is a vector that survives twice and dies only on the third application" — a chain of length .
Example 8.25 (Minimal polynomial of a diagonalizable matrix). Find for , given that .
Solution.
Example 8.26 (Minimal polynomial from the rank data). Let be the matrix with . Given , and , find .
Solution.
Example 8.27 (Deducing the minimal polynomial from a factorisation). A matrix satisfies and , . Find and decide whether is diagonalizable.
Solution.
Pitfall. and have the same roots, never a different set. A statement like " and " is impossible: is an eigenvalue and must be a root of . The only legal minimal polynomials there are and .
The missing eigenvectors of a defective matrix are not missing at all; they are one level deeper.
Definition 8.28 (Generalized eigenvector and generalized eigenspace). Let be an eigenvalue of . A nonzero vector is a generalized eigenvector of for if
The generalized eigenspace is
The smallest that works for a particular is the rank or height of ; height means is an ordinary eigenvector.
The union looks worrying — a union of subspaces is usually not a subspace — but the kernels here are nested, and nested unions are fine.
Lemma 8.29 (Kernels grow and then stop). Write . Then
and if for some , then for all . Consequently the chain stabilises at or before step , and
Proof. If then , giving the inclusions.
Suppose and take . Then , so , so , i.e. . Thus , and induction propagates the equality forward.
Until the chain stabilises, each inclusion is strict, so each step raises the dimension by at least . Starting from , the dimension would exceed after more than strict steps; so stabilisation happens at some index . Hence is the largest kernel, and .∎
In particular is a subspace, and since it is the kernel of a polynomial in it is -invariant. The crucial fact is that it is exactly the right size.
Theorem 8.30 (Dimension of the generalized eigenspace). For every eigenvalue of ,
the algebraic multiplicity.
Proof. Write and , of dimension , say. is -invariant, so by the block triangular theorem there is a basis in which
with the matrix of .
First, is nilpotent: it is the matrix of , and kills every vector of by definition. A nilpotent matrix has as its only eigenvalue, so has as its only eigenvalue and .
Second, is not an eigenvalue of . Suppose it were. Working in the basis with spanning , an eigenvector of for lifts to a vector , , with , i.e. . Then , so , so by the stabilisation lemma. But is a nonzero combination of , which meets only in . Contradiction.
So does not divide , and the multiplicity of in is exactly . That is, .∎
Intuition. Concretely, for the eigenvalue owns two chairs but has only one eigenvector, . The vector is not an eigenvector: . But one more application kills it: . So is a generalized eigenvector of height , it takes the empty chair, and has dimension . A generalized eigenvector is a vector that pushes one step down a ladder, instead of annihilating outright.
Theorem 8.31 (Primary decomposition). Let be an complex matrix with distinct eigenvalues . Then
and each summand is -invariant with .
Proof. Write with , and set
The polynomials have no common root — while every with vanishes at — so their greatest common divisor is . By Bézout's identity for polynomials there are polynomials with
Put . Then , so every decomposes as . Each piece lands where it should:
by Cayley–Hamilton, so . Hence .
The sum is direct by dimension count: by the previous theorem and the fact that splits, and a spanning sum of subspaces whose dimensions add to must be direct.
Invariance of each was noted after the stabilisation lemma.∎
This is the first half of the Jordan theorem. Choosing a basis adapted to the decomposition puts into block diagonal form
because is nilpotent. Everything now reduces to one question: what does a nilpotent matrix look like?
Example 8.32 (Computing a generalized eigenspace). For with , find and a generalized eigenvector of height .
Solution.
which has rank , so , as the theorem predicts. Thus . 3. A height- vector is one in but not . Take : it satisfies , and , while . 4. Sanity check: , so is a genuine eigenvector for — exactly the one found earlier. The chain is .□
Example 8.33 (The primary decomposition in action). With the same , exhibit the decomposition and check the dimensions.
Solution.
Pitfall. A generalized eigenvector of height is not an eigenvector, and for it. Writing for every element of is the commonest error in this chapter. What is true is , which is precisely why the Jordan form has 's above the diagonal rather than being diagonal.
Definition 8.34 (Nilpotent matrix and index). A square matrix is nilpotent if for some . The least such is the index of nilpotency of .
Proposition 8.35 (Basic facts about nilpotent matrices). Let be and nilpotent of index . Then:
Proof. (1) If with then , so . Over the characteristic polynomial splits, and all its roots are , so .
(2) annihilates and no smaller power does, by definition of the index; and gives .
(3) By the diagonalizability criterion, is diagonalizable iff has distinct linear factors; the only such power of is , i.e. .
(4) Telescoping, .∎
Fact (1) with Cayley–Hamilton gives at once: the index never exceeds the size.
Definition 8.36 (Jordan chain). Let be nilpotent and let have height , meaning and . The Jordan chain generated by is the list
of length . Its first element is an eigenvector (it is killed by ); its last element is the top of the chain.
Lemma 8.37 (A Jordan chain is independent). With , , as above, the vectors are linearly independent, and the subspace they span is -invariant of dimension .
Proof. Suppose with not all zero, and let be the first nonzero coefficient. Apply to the relation. Every term with index acquires with , hence vanishes, and the terms with have . What remains is
forcing , a contradiction. Invariance is clear: maps each listed vector to the next (and the last to ).∎
In the basis — note the order, eigenvector first — the restriction sends each basis vector to the previous one and the first to . Its matrix is therefore the matrix with 's on the superdiagonal and 's elsewhere. That is a nilpotent Jordan block. The structure theorem says every nilpotent operator is built entirely of such chains.
Theorem 8.38 (Structure of a nilpotent operator). Let be a nilpotent operator on a finite-dimensional space . Then has a basis consisting of finitely many Jordan chains for ; equivalently, is a direct sum of -invariant subspaces on each of which acts as a single nilpotent Jordan block.
Proof. Induct on . If there is nothing to prove, and if any basis is a collection of chains of length .
Otherwise let , which is -invariant and, since has nontrivial kernel, satisfies . By induction has a basis made of chains; say the chains have tops and lengths , so the basis of is .
Each lies in , so choose with . The chain generated by has length and extends the -th chain of by one step at the top.
The bottoms () are independent vectors of . Extend them to a basis of by adding vectors , where ; each is a chain of length .
Claim. The union
is a basis of . Count first: it has elements, using the rank–nullity theorem. So it suffices to prove independence.
Suppose a combination of the elements of is zero. Apply . Every dies, every (the chain bottoms coming from ) dies, and the rest map onto the basis of described above — indeed for . Independence of that basis of forces all the coefficients of with to vanish. What survives of the original relation is a combination of the chain bottoms and the , all of which lie in and were chosen to be independent there. So those coefficients vanish too.∎
Intuition. Picture the basis as a stack of columns, one column per chain, drawn with the eigenvector at the bottom and the top of the chain at the top. The operator moves every basis vector down one square and pushes the bottom row off the edge. The number of columns is — one eigenvector per column — and the height of the tallest column is the index of nilpotency. The whole structure is determined by the column heights, which is why the classification reduces to a list of block sizes.
Example 8.39 (Chains for the shift). Find the chain decomposition of .
Solution.
Example 8.40 (A nilpotent matrix in disguise). Show that is nilpotent, find its index, and produce a Jordan chain.
Solution.
Example 8.41 (Two chains). Let . Show , and find its chain structure.
Solution.
Definition 8.42 (Jordan block). For a scalar and an integer , the Jordan block is the matrix with on the diagonal, on the superdiagonal, and elsewhere:
A Jordan matrix is a block diagonal matrix whose diagonal blocks are Jordan blocks.
Read a block as , where is the nilpotent shift of the previous section. A single block has one eigenvalue with and : as defective as a matrix can be.
Theorem 8.43 (Jordan normal form). Every complex matrix is similar to a Jordan matrix :
where the are eigenvalues of , repetitions allowed. The Jordan matrix is unique up to the order in which the blocks are listed; it is called the Jordan normal form of .
Proof. Existence. By the primary decomposition theorem, over the distinct eigenvalues, each summand -invariant. On the operator is nilpotent, so by the structure theorem for nilpotent operators has a basis made of Jordan chains for . In such a basis, ordered eigenvector-first within each chain, is a direct sum of nilpotent shifts, so is a direct sum of blocks . Concatenating these bases over gives a basis of ; the matrix with those basis vectors as columns satisfies .
Uniqueness. This is the content of the rank formula below: the number of blocks of each size for each eigenvalue is determined by the ranks of the powers , which are invariants of the similarity class. So two Jordan matrices similar to the same have the same multiset of blocks.∎
The rank formula is the computational heart of the chapter. It starts with a one-line observation about a single block.
Lemma 8.44 (Rank of a power of a block). For any and ,
while for ,
Proof. is the shift, and has 's on the -th superdiagonal and zeros elsewhere; there are of them when , occupying distinct rows and columns, so the rank is ; and once .
For , is upper triangular with the nonzero constant on the diagonal, hence invertible; so are its powers, and their rank is .∎
Theorem 8.45 (The rank formula for block counts). Fix an eigenvalue of and set
Then the number of Jordan blocks for of size at least is
and the number of blocks of size exactly is
Proof. Similar matrices have equal ranks of corresponding powers, so we may compute with the Jordan form itself. Say the blocks for have sizes , and let the remaining blocks (for other eigenvalues) have total size . Rank is additive over a block diagonal matrix, so by the lemma
Subtract consecutive values. For each , equals if and otherwise. Hence
Finally .∎
Three consequences are worth naming, because in easy cases they settle the answer without any computation of powers.
Corollary 8.46 (Reading off the structure). For an eigenvalue of :
Proof. (1) is the displayed computation. (2) holds because the blocks for together form the matrix of restricted to , of dimension . For (3), of the minimal polynomials of the blocks, and since ; the least common multiple takes the largest exponent.∎
Summary (The classification data). For each eigenvalue , the Jordan structure is determined by, and determines, any one of these:
Two matrices are similar exactly when they have the same eigenvalues with the same block sizes. The characteristic polynomial gives of block sizes, the minimal polynomial gives the largest block size, and the geometric multiplicity gives the number of blocks — three numbers that determine the structure completely when , but not in general.
Intuition. The number says: "how many columns of the chain diagram survive to height . " Each application of shaves one square off every surviving column, and the drop in rank counts the columns that were still standing. The second difference is then the count of columns of height exactly — a discrete second derivative of the rank sequence.
Example 8.47 (Blocks from a rank sequence). A matrix has a single eigenvalue , and , , . Determine the Jordan form.
Solution.
Example 8.48 (A larger rank sequence). A matrix has the single eigenvalue with for . Find the block sizes.
Solution.
Example 8.49 (Two eigenvalues at once). A matrix has eigenvalues and with . It is given that , , , and . Find the Jordan form.
Solution.
Example 8.50 (Enumerating the possibilities). How many similarity classes of complex matrices have characteristic polynomial , and which have minimal polynomial ?
Solution.
Pitfall. Characteristic and minimal polynomials together do not always determine the Jordan form. For , and is consistent with block sizes and with and with and with . You need the geometric multiplicity, or the full rank sequence, to separate them.
Knowing is often enough. When you also need — to compute , , or to solve a differential system — you must produce the chains explicitly.
Method 8.51 (Finding a Jordan basis). Given an matrix :
Step 6 deserves a comment on ordering, because it is where signs and transposes go wrong. If a chain is , then
because and . Column of is the coordinate vector of , namely — which is precisely the -th column of , with the above the diagonal. Reversing the order inside a chain produces 's below the diagonal instead: also a legitimate normal form, but not the standard one.
Example 8.52 (A with one chain of length ). Find a Jordan basis and the Jordan form of
Solution.
and (rank must drop to since and there is one block of size ).
Example 8.53 (A with two eigenvalues). Find a Jordan basis and the Jordan form of .
Solution.
Example 8.54 (A with blocks of sizes and ). Find the Jordan form and a Jordan basis for
Solution.
Pitfall. Build the chains from the longest down, and always start a chain at its top — a vector of maximal height — then apply repeatedly. Starting from an eigenvector and trying to "solve upward" for with often fails, because a given eigenvector need not be in the image of ; only the ones at the bottom of a long chain are. Choosing the short chains first can leave you unable to complete a long one.
The payoff. Once , every power is , and is computed block by block.
Theorem 8.55 (Powers of a Jordan block). Write with the nilpotent shift, . Then for ,
so the entry in row , column is for , and all other entries vanish. For example
Proof. and commute, since commutes with everything, so the binomial theorem applies:
Terms with vanish because , which truncates the sum at . The entry description is the observation that has 's exactly in positions .∎
Notice the shape of the answer: the entries are times polynomials in of degree up to . That single fact explains the qualitative behaviour of every linear recursion: a Jordan block of size at eigenvalue contributes terms to the growth, so a repeated eigenvalue on the unit circle produces polynomial growth rather than a bounded orbit.
Example 8.56 (Powers of a defective directly). Find a closed formula for , .
Solution.
That shortcut is general and worth stating on its own, because it avoids computing entirely whenever there is only one eigenvalue.
Proposition 8.57 (Single-eigenvalue shortcut). If has the single eigenvalue and , then for every ,
Proof. and the two summands commute; apply the binomial theorem and drop the terms with .∎
Example 8.58 (A power). For , whose Jordan form is , compute the entry of .
Solution.
Example 8.59 (Growth rate of a recursion). A sequence satisfies where is with Jordan form . Describe the growth of .
Solution.
Remark (Functions of a matrix in general). The same computation defines for any function that is differentiable enough at the eigenvalues. Replacing — which is for — by the general Taylor coefficient gives
and . A function of a matrix depends only on the values of and its first derivatives at each eigenvalue, where is the largest block there. The next section is the case .
Definition 8.60 (Matrix exponential). For a square matrix ,
a series that converges entry-by-entry for every .
Proposition 8.61 (Properties of the exponential). For square matrices of the same size:
Proof. (1) , so the partial sums satisfy the identity and it passes to the limit.
(2) Commutativity lets the usual binomial rearrangement of the Cauchy product go through: . Taking gives .
(3) Differentiate the series term by term (justified by uniform convergence on bounded -intervals): .
(4) Powers of a block diagonal matrix are block diagonal with the blocks raised to the power.∎
Property (2) has a famous caveat: can fail spectacularly when and do not commute. Fortunately, in a Jordan block and do commute, which is exactly what makes the following formula work.
Theorem 8.62 (Exponential of a Jordan block). For the block ,
whose entry is for . For example
Proof. and commute, so , the last step because . The series for terminates at since .∎
Theorem 8.63 (Solution of a linear system). The initial value problem
has the unique solution .
Proof. It is a solution: by property (3), and . For uniqueness, let be any solution and set . Then
using that commutes with . So is constant, equal to , and .∎
Intuition. A diagonalizable gives solutions that are pure combinations of exponentials : each mode decays or grows at its own rate and nothing interacts. A Jordan block of size produces a solution containing . That extra factor of is the signature of resonance — it is the same that appears in the method of undetermined coefficients when the forcing term matches a root of the characteristic equation, and it is the reason a critically damped oscillator has a solution .
Example 8.64 (Exponential of a defective ). Compute for , and solve with .
Solution.
Example 8.65 (Exponential at a specific time). With the same , compute the entry of at , and the trace of .
Solution.
Example 8.66 (A system). Solve for with .
Solution.
Pitfall. is not the matrix of entrywise exponentials , and is not — the correct identity is , which follows by triangularizing. Nor is valid in general: take and , which do not commute, and the two sides differ.
The last item is the real form. Over , a matrix may have no real eigenvalues at all, and the Jordan form we have built is genuinely complex. There is a real substitute.
Theorem 8.67 (Real Jordan form). Let be a real square matrix. Then is similar, over , to a block diagonal matrix whose blocks are either ordinary Jordan blocks for the real eigenvalues , or real Jordan blocks for each conjugate pair with :
where the second display is the size- version, with identity matrices where a complex block has 's.
Proof. Complex eigenvectors come in conjugate pairs: if with and ( real), then . Separating real and imaginary parts of gives
so is a real -invariant plane, and in the basis the restriction has matrix ; using the basis instead yields the displayed above. Running the same substitution through a complex Jordan chain for turns each complex chain vector into a real pair and each into an ; the rest of the argument is the complex theorem applied to , whose real points are exactly the span of the real and imaginary parts.∎
Example 8.68 (Real form of a rotation-scaling). Find the real Jordan form of and describe the solutions of .
Solution.
Summary (The chapter in one page). For a complex matrix :
and on each the operator is nilpotent, hence a direct sum of Jordan chains. In a basis of chains, becomes , unique up to block order, with
The characteristic polynomial is , the minimal polynomial is , and is diagonalizable exactly when every block has size , equivalently when has no repeated root. Powers and exponentials follow block by block: