Contents / Calculus / Partial Derivatives
Chapter 11
Partial Derivatives
Functions of several variables, partial derivatives, the chain rule, gradients, and optimization with Lagrange multipliers.
Introduction
Single-variable calculus studies functions : one input, one output, and a graph you can draw on a page. Almost nothing in the world is that simple. The temperature in a room depends on where you stand, the pressure of a gas depends on its volume and its temperature, the cost of a product depends on the price of every ingredient. These are functions of several variables, and this chapter builds their differential calculus.
The plan is the same as in one variable — limits, then derivatives, then tangent lines, then optimisation — but every idea acquires more room to move. A function of two variables can be approached from infinitely many directions, so a limit is a stronger claim than before. Its graph is a surface, so a tangent line becomes a tangent plane. Its rate of change depends on which way you walk, and one vector, the gradient, packages every such rate at once.
The chapter ends with the two tools the rest of applied mathematics leans on constantly: the second derivatives test for locating maxima, minima and saddle points, and the method of Lagrange multipliers for optimising under a constraint. Multiple integrals, vector fields and the big integral theorems live in the chapters that follow; everything here is differentiation.
11.1Functions of several variables
Definition 11.1 (Function of two variables). A function of two variables is a rule that assigns to each ordered pair in a set a unique real number . The set is the domain of and the set of values is its range. The graph of is the set of points in with and .
We write and call and the independent variables, the dependent variable. When a formula is given without a domain, the domain is understood to be every for which the formula makes sense — no square roots of negatives, no logarithms of non-positive numbers, no division by zero. Finding it is a two-dimensional exercise: each restriction cuts out a region of the plane, and the domain is what survives all of them.
The graph of is a surface sitting over the domain. A linear function has a plane for its graph; has the upper hemisphere of radius ; has a bowl called a paraboloid, and has a saddle, curving up in one direction and down in the other. Surfaces are hard to draw, so the working tool is a flat picture instead.
Definition 11.2 (Level curves and contour maps). The level curves of are the curves in the plane with equations , one for each constant in the range of . A collection of level curves for equally spaced values of is a contour map of .
A level curve is the set of inputs on which takes the value ; it is the shadow in the -plane of the horizontal slice of the graph at height . Where the level curves crowd together the surface is steep, because a small step in the plane crosses many heights; where they are far apart the surface is nearly flat. Every idea in this chapter — partial derivatives, the gradient, constrained optimisation — has a picture on the contour map, and it is worth learning to read one.
Intuition. A walker's map of a mountain draws a line through every point at m, another through every point at m, and so on. No third dimension is printed, yet you can read the mountain off the map: contours packed tight mean a cliff, contours far apart mean a meadow, a closed ring means a summit or a hollow. A contour map of is exactly that map, with playing the role of altitude.
Everything extends to three or more inputs. A function assigns a number to each point of a region in space — the temperature at each point of a room, say. Its graph would live in four dimensions, so we never draw it; instead we draw its level surfaces , which are surfaces in space. For the level surfaces are concentric spheres, and knowing them is knowing the function.
Example 11.3 (Finding a domain). Find the domain of and of .
Solution. For , the square root requires and the denominator requires . The domain is the closed half-plane on and above the line , with the vertical line removed.
For , the logarithm requires , that is,
The domain is the interior of the ellipse with semi-axes and , boundary excluded. As a check, satisfies the inequality and is a perfectly good number, while lies on the ellipse and gives , which is undefined.□
Example 11.4 (Level curves of a paraboloid and of a saddle). Sketch the level curves of and of .
Solution. The level curves of are : for a circle of radius centred at the origin, for the single point , and for nothing at all. Since the radii grow ever more slowly, the circles for equally spaced bunch together as they move outward — the bowl steepens away from its base.
The level curves of are . For this is the pair of lines . For it is a hyperbola opening up and down, for a hyperbola opening left and right. The two families are separated by the crossed lines, which is the fingerprint of a saddle: walk along the -axis and rises, walk along the -axis and it falls.□
Example 11.5 (Reading a contour map). The level curves of for are drawn. Describe them, and say in which direction from the origin increases fastest.
Solution. Setting gives , an ellipse for :
For the semi-axes are (along ) and (along ); for they are and ; for they are and . Every ellipse is twice as tall as it is wide.
Contours are closer together along the -axis: moving from to takes only unit in the -direction but units in the -direction. So climbs fastest along the -axis, which agrees with the formula — the coefficient on makes count double.□
Example 11.6 (Level surfaces of a function of three variables). Describe the level surfaces of and find the domain of .
Solution. The level surface is , a paraboloid with vertex at . The level surfaces are all translates of a single bowl, sliding up as decreases.
For we need , so the domain is the set of points strictly above the paraboloid — the inside of the bowl, its surface excluded. The point is in the domain and ; the point sits on the paraboloid and is not.□
11.2Limits and continuity
Definition 11.7 (Limit of a function of two variables). Let be defined on a set containing points arbitrarily close to . We write
if for every there is a such that whenever and .
Compare this with the one-variable definition from Limits & Continuity. The expression is the distance from to , so the condition says: whenever the input is within of but not equal to it, the output is within of . Nothing is said about how the input approaches; the disc of radius contains every direction, every spiral, every curve into , and must be near on all of them at once.
That is a far stronger demand than in one variable, where a point can only be approached from the left or from the right. On the line, the limit exists when two one-sided limits agree. In the plane there are infinitely many "sides", and a single disagreement kills the limit.
Theorem 11.8 (Two-path test). If as along a path and along a path , with , then does not exist.
Proof. Suppose the limit existed and equalled . Points on close to are inside the -disc, so is within of on them, forcing . The same argument on forces . Then , contrary to assumption.∎
The test is the standard way to show a limit fails. Try the two axes first, then the line , then the family of lines ; if the answer depends on , you are done. If every line gives the same value, try a parabola — the third example shows why.
Showing that a limit exists is harder, because you must control every path at once. Two tools do it. The first is a squeeze: bound by something that visibly goes to zero. The second, which usually amounts to the same thing but organises it, is polar coordinates centred at : put , . Then means exactly , with free to do anything, and if the resulting expression tends to uniformly in — bounded by a function of alone that goes to — the limit is .
Intuition. Picture a summit and hikers arriving from every direction — up the north ridge, the south face, a spiral trail. The altitude at the summit is well defined only if every hiker reads the same number on the final step. Two hikers reporting different altitudes at the same point is the two-path test. A single altimeter reading that is within of for every hiker within of the summit is the definition of the limit.
Example 11.9 (Different limits along the axes). Show that does not exist.
Solution. Along the -axis, where and ,
so as along this path. Along the -axis, where ,
so . Two paths, two different limits: by the two-path test the limit does not exist. As a sanity check, along the function is identically , a third value again.□
Example 11.10 (Agreeing on the axes, disagreeing on a diagonal). Show that does not exist.
Solution. On the -axis and on the -axis , so both axes give the limit . That does not settle anything. Along the line ,
so the limit along this path is , and the limit does not exist. Along a general line the value is , a different constant for every slope: the function takes every value between and in every neighbourhood of the origin.□
Example 11.11 (Every line agrees, a parabola does not). Show that does not exist.
Solution. Along any line through the origin,
and along the -axis as well. Every straight line gives . But along the parabola ,
so the limit along this curve is . Two paths disagree, and the limit does not exist. Lines are not enough; the definition demands every path.□
Example 11.12 (Proving a limit exists with polar coordinates). Show that .
Solution. Substituting , turns the denominator into :
Since for every ,
and as regardless of . The limit is . In – terms, works. The same estimate can be reached without polar coordinates: gives , and .□
Example 11.13 (A logarithm tamed by polar coordinates). Evaluate .
Solution. In polar coordinates , so the expression is , which depends on alone. As this is a form; rewrite it as
by L'Hôpital's Rule. There is no to worry about, so the two-variable limit is .□
Pitfall. Substituting polar coordinates and letting is only valid when the result is controlled for every at once. If the expression after substitution still depends on in a way that does not vanish — for it becomes , which does not go to at all — then the limit does not exist, and the polar form has just handed you the paths that prove it.
Definition 11.14 (Continuity). A function of two variables is continuous at if
It is continuous on if it is continuous at every point of .
Continuity means what it meant before: the limit exists, the value exists, and they agree. The practical content is a list of functions you never have to check. Polynomials in and are continuous everywhere, because limits of sums and products are sums and products of limits, and , come straight from the definition. Rational functions are continuous wherever the denominator is nonzero. And if is continuous at and is a continuous function of one variable at , the composition is continuous at : near the value is near , and then is near . So , , and are continuous on all of without any further work.
Example 11.15 (Where is a function continuous?). Determine where is continuous, and whether
is continuous at the origin.
Solution. The quotient is a rational function, continuous wherever , and is continuous everywhere, so the composition is continuous on the set . On the line it is undefined, hence not continuous there. It cannot even be repaired: approaching from gives , from gives .
For : away from the origin it is a rational function with nonzero denominator, so continuous. At the origin the limit was shown to be in the polar-coordinates example, and by definition, so is continuous at the origin too. The function is continuous on all of .□
11.3Partial derivatives
Definition 11.16 (Partial derivatives). The partial derivatives of with respect to and to at are
whenever these limits exist.
In the second input is frozen at and only the first moves; the limit is the ordinary derivative of the one-variable function at . That observation is also the rule for computing: to find , regard as a constant and differentiate with respect to by every rule you already know; to find , regard as a constant and differentiate with respect to .
Notation. For the following all mean the same thing:
and likewise for . The curly is read "partial" and is never written : is not a quotient of two quantities and .
Geometrically, freezing slices the surface with a vertical plane; the slice is a curve , and is the slope of its tangent line at the point . Freezing gives a second curve whose tangent slope is . Two slices, two slopes — and, as the next section shows, those two tangent lines span the tangent plane.
As rates of change, is how fast changes per unit step in the -direction from , and how fast it changes per unit step in the -direction. If is the temperature of a plate, is the temperature gradient felt walking east and the one felt walking north.
Intuition. Stand on a hillside at a point where the ground is a plane tilted so that walking east climbs m per m and walking north climbs m per m. Those two numbers, and , are the partial derivatives of the altitude function at your feet. Neither is "the slope of the hill" — the hill has a different slope in every direction — but the section on the gradient shows that these two determine all the others.
Example 11.17 (Partial derivatives at a point). If , find and .
Solution. Holding constant and differentiating with respect to ,
Holding constant and differentiating with respect to (the term is a constant and disappears),
Check directly: with , , and .□
Example 11.18 (Partial derivatives with the chain rule inside). Find and for .
Solution. With fixed, is a constant multiple of , so
With fixed, has derivative , so
Example 11.19 (Implicit partial derivatives). If is defined implicitly as a function of and by , find and .
Solution. Differentiate both sides with respect to , treating as a constant and as a function of :
Differentiate with respect to , now treating as a constant; the product needs the product rule:
The chain rule section turns this procedure into a formula.□
Functions of three or more variables are handled the same way: freezes and and differentiates in . For ,
Since a partial derivative is again a function of and , it can be differentiated again. Four second-order partial derivatives arise:
and similarly and . Read the subscripts left to right: means differentiate in first, then in . In the notation the order is reversed, , because the operator nearest acts first. The two mixed partials and look as if they could differ, and the next example shows what happens in practice.
Example 11.20 (All four second partials). Find the second partial derivatives of .
Solution. From the first example, and . Differentiating each in both variables:
The mixed partials agree. That is not a coincidence.□
Theorem 11.21 (Clairaut's Theorem). Suppose is defined on a disc containing the point . If the functions and are both continuous on , then
The proof applies the Mean Value Theorem twice to the double difference , once in each order, and belongs to real analysis; the idea is that both mixed partials are limits of that same symmetric quantity divided by . What matters in practice is the hypothesis: continuity of the mixed partials. For every function built from polynomials, exponentials, logarithms and trigonometric functions on its natural domain it holds automatically, and you may compute whichever mixed partial is easier.
Remark. The continuity hypothesis is not decoration. The function
has and , so while . Its mixed partials exist everywhere but are discontinuous at the origin, and there they disagree.
Partial derivatives of order three and higher are defined the same way, and by Clairaut's Theorem applied repeatedly, whenever these are continuous. Only the number of differentiations in each variable matters.
Partial derivatives appear in the laws of physics chiefly through partial differential equations, equations relating an unknown function to its partial derivatives. Three of them describe most of classical physics. Laplace's equation
governs steady-state temperature, electrostatic potential and the flow of ideal fluids; its solutions are called harmonic functions. The wave equation
describes a vibrating string of displacement , with the wave speed. The heat equation
describes how temperature along a rod evens out over time. Solving them is a later course; verifying that a given function solves one is a partial-derivative computation.
Example 11.22 (Verifying a harmonic function). Show that satisfies Laplace's equation.
Solution. Differentiate twice in each variable:
Then , so is harmonic. The same computation shows is harmonic too.□
Example 11.23 (Verifying a solution of the wave equation). Show that satisfies the wave equation .
Solution. Treating as constant, and . Treating as constant, the chain rule contributes a factor each time:
So . Physically, is a sine wave travelling to the right at speed : at time the shape is the graph of shifted units right.□
Example 11.24 (Verifying a solution of the heat equation). Show that satisfies the heat equation .
Solution. , while and . Hence
The picture: the initial temperature profile keeps its shape and decays exponentially in amplitude, faster for larger — heat spreads from the warm crests into the cool troughs.□
Pitfall. The existence of both partial derivatives at a point does not make continuous there, let alone nice. The function with has and , because it vanishes identically on both axes, yet it has no limit at the origin. Partial derivatives only look along two lines; the next section introduces the right notion of "differentiable", which looks in every direction.
11.4Tangent planes and linear approximations
In one variable, the tangent line at is the line through with slope , and zooming in on a differentiable graph makes it indistinguishable from its tangent line. The same picture holds for a surface: zoom in on a smooth surface and it flattens into a plane, the tangent plane.
Let be the surface and a point on it, with . The tangent plane at should contain the tangent lines to the two slice curves (in the plane ) and (in the plane ). Any non-vertical plane through has an equation
for some constants and . Intersect it with the plane : the second term vanishes and we get , a line of slope . For this to be the tangent line to we need . Intersecting with in the same way forces . Two conditions, two unknowns, and the plane is determined.
Theorem 11.25 (Equation of the tangent plane). If has continuous partial derivatives, the tangent plane to the surface at the point is
Example 11.26 (A tangent plane to a paraboloid). Find the tangent plane to the elliptic paraboloid at the point .
Solution. With , and , so and . The tangent plane is
Check: the plane passes through , since . Near that point the surface and the plane are close: at the surface gives and the plane gives .□
The last computation is the whole point. The right-hand side of the tangent plane equation is a linear function of and , easy to evaluate, and it approximates near .
Definition 11.27 (Linearization and linear approximation). The linearization of at is the linear function
and the approximation for near is the linear (or tangent plane) approximation of at .
Example 11.28 (Using a linear approximation). Find the linearization of at and use it to approximate .
Solution. . By the product rule , so ; and , so . The linearization is
Hence . The true value is , so the approximation is off by about — reasonable for a step of size from the base point.□
There is a subtlety hiding here. The tangent plane formula uses only and , and the pitfall at the end of the previous section produced a function with that is not even continuous at the origin. Its "tangent plane" is a terrible approximation: the function takes the value arbitrarily close to . So the mere existence of the partial derivatives does not guarantee that the tangent plane approximates the surface. We need a definition of differentiability that asks for exactly that guarantee.
Write for the change in . In one variable, differentiability at is equivalent to with as : the change is the tangent-line prediction plus an error that is small compared with . The two-variable version reads the same way.
Definition 11.29 (Differentiability). The function is differentiable at if can be written as
where as .
In words: is differentiable when the tangent plane approximation is good, with an error that is small compared with the distance moved. The definition is awkward to verify directly, but there is a clean sufficient condition that covers every function met in practice.
Theorem 11.30 (Continuous partials imply differentiability). If the partial derivatives and exist near and are continuous at , then is differentiable at .
Proof. Split the change into two one-variable steps, first in then in :
In the first bracket only varies, so the Mean Value Theorem gives a between and with the bracket equal to . In the second bracket only varies, and there is a between and with the bracket equal to . Now set
Then exactly, and as the points and tend to , so continuity of and at makes .∎
Corollary 11.31 (Differentiability implies continuity). If is differentiable at , then is continuous at .
Proof. Every term on the right of the differentiability equation contains a factor or , so as ; that is, .∎
Polynomials, rational functions on their domains, and any combination of them with exponentials, logarithms and trigonometric functions have continuous partials, hence are differentiable wherever they are defined. The example has partials at the origin, but they are not continuous there, and the theorem does not apply — correctly, since the function is not differentiable there.
Intuition. Differentiability is the promise that the surface is locally flat. Rest a rigid sheet of card on a smooth hill at one point: it touches along the tangent plane and, in a small enough patch, the hill barely leaves the card. Now try the same on the surface near the origin, where it looks like a twisted ribbon rising to height in some directions and dropping to in others: no card fits. Two slopes exist, but no plane.
Once is differentiable, the approximation is worth its own notation. For a differentiable function , the differentials and are independent variables, and the total differential is
Taking and , the differential is the change in height of the tangent plane while is the change in height of the surface, and the definition of differentiability says exactly that .
Example 11.32 (Comparing and ). For , compare and when changes from to and changes from to .
Solution. The differential is . At with and ,
The actual change is
The differential is easier to compute and is off by only .□
The typical use is error estimation: measured quantities carry errors, a computed quantity inherits an error from them, and says how large.
Example 11.33 (Estimating the error in a computed volume). The base radius and height of a right circular cone are measured as cm and cm, each with a possible error of cm. Estimate the maximum error in the calculated volume, and the maximum relative error.
Solution. The volume is , so
The errors satisfy and ; the worst case is both at with the same sign:
That looks large, but , so the relative error is
about . Note how the exponent on doubles the relative contribution of the radius error.□
For three variables everything reads the same: is linearised by , and .
Example 11.34 (Approximating a value with three variables). Use differentials to approximate without a calculator.
Solution. Take at the base point , where . Each partial has the form , so at the base point
With , , ,
so the value is approximately . Squaring the true numbers gives , whose square root is ; the estimate is correct to three decimals.□
11.5The chain rule
In one variable the chain rule says that if and then : rates multiply along a chain. With several variables a quantity can depend on through more than one intermediate variable, and the rule acquires one term per route.
Theorem 11.35 (Chain Rule, case 1). Suppose is a differentiable function of and , where and are differentiable functions of . Then is a differentiable function of and
Proof. A change produces changes and , which produce a change in . Since is differentiable,
with as . Divide by and let . Then and by definition of the derivative. Because and are differentiable they are continuous, so and hence ; the last two terms therefore contribute , leaving the formula.∎
The formula is the one-variable chain rule applied once per route from to and summed. Think of a tree: at the top with branches down to and , each of which has a branch down to . Every root-to-leaf path is a product of derivatives along its edges, and is the sum over all paths. That description works for every case of the chain rule and is worth drawing whenever the bookkeeping gets heavy.
Example 11.36 (A rate along a curve). If , where and , find when .
Solution. By case 1,
At we have and , so
Check by substituting first: , and near the dominant term is , whose derivative is .□
Example 11.37 (Related rates for a gas). The pressure (kPa), volume (L) and temperature (K) of a mole of ideal gas satisfy . Find the rate at which the pressure is changing when K and increasing at K/s while L and increasing at L/s.
Solution. Write as a function of the two variables and , each a function of :
Substituting the given values,
The pressure is falling: the expansion of the volume outweighs the warming.□
When the intermediate variables themselves depend on two variables, the same tree has two leaves at the bottom, and we get a partial derivative for each.
Theorem 11.38 (Chain Rule, case 2). Suppose is differentiable, where and have first partial derivatives. Then
Proof. To compute we hold fixed, and then and are functions of the single variable . Case 1 applies verbatim, with the ordinary derivatives , replaced by the partials , because is being held constant. The formula for is the same argument with the roles of and exchanged.∎
Example 11.39 (Chain rule with two independent variables). If , where and , find and .
Solution. The four partial derivatives of the intermediate variables are
Case 2 gives
Theorem 11.40 (Chain Rule, general version). Suppose is a differentiable function of the variables , and each is a differentiable function of the variables . Then is a function of and, for each ,
The tree diagram now has at the top, the intermediate variables in a row beneath it, and the independent variables in a row beneath each of those. To find , follow every path from down to a leaf labelled , multiply the derivatives written on the edges of that path, and add up over paths. There are such paths, one through each intermediate variable, which is why the formula has terms.
Example 11.41 (Three intermediate and three independent variables). If , where , and , find when , , .
Solution. At the given point , , . The three paths from to give
Substituting , , , , , :
Example 11.42 (Rewriting an expression in polar coordinates). If has continuous second partials and , , show that
Solution. By case 2, with , , , ,
Square and combine:
The cross terms cancel and the Pythagorean identity does the rest. This identity is what lets physicists write the gradient's magnitude in polar coordinates.□
The chain rule also settles implicit differentiation, replacing the case-by-case procedure of the earlier chapters by a formula. Suppose an equation defines implicitly as a differentiable function of . Then for all in the domain, and differentiating both sides with respect to by case 1 (the two intermediate variables being itself and ),
Since , solving gives the first formula below; the same argument with and , differentiating in with held fixed so that , gives the second.
Theorem 11.43 (Implicit differentiation formulas). If defines implicitly as a differentiable function of and , then
If defines implicitly as a differentiable function of and and , then
When does an equation actually define such a function? The Implicit Function Theorem, proved in real analysis, gives the answer: if has continuous partials on a disc containing , with and , then near the equation defines as a differentiable function of , and its derivative is given by the formula. The three-variable version is identical with . The condition is exactly the condition that the formula makes sense — and geometrically, that the level curve is not vertical at the point.
Example 11.44 (Implicit differentiation, two variables). Find if .
Solution. Let . Then and , so
This is the same answer the direct method from the derivatives chapter gives, with the algebra done once and for all. The formula fails where , which is where the curve has a vertical tangent.□
Example 11.45 (Implicit differentiation, three variables). Find and if .
Solution. Let . Then
The symmetry of the original equation in and is reflected in the answers, which is a quick sanity check.□
Intuition. A bug crawls across a heated plate whose temperature is . The bug's position at time is . How fast does the temperature it feels change? Two effects add up: it is moving east at speed through a temperature gradient , contributing , and it is moving north at speed through a gradient , contributing . The chain rule is the statement that these two effects simply add. If the plate's temperature also drifted with time, a third term would join them — one term per way the input can change.
Pitfall. In case 2, and are not interchangeable symbols. Both are partial derivatives of " ", but with respect to different sets of variables: holds fixed, while holds fixed, and generally moves when does. Keep the tree in front of you and never cancel a against a as if they were fractions.
11.6Directional derivatives and the gradient vector
The partial derivatives and give the rate of change of in the directions of the two coordinate axes. There is nothing special about those directions; a hiker can walk north-east as easily as north. The directional derivative measures the rate of change in an arbitrary direction, given by a unit vector .
Definition 11.46 (Directional derivative). The directional derivative of at in the direction of the unit vector is
if this limit exists.
The point is a step of length from in the direction , so the quotient is the change in per unit distance travelled in that direction. Taking recovers , and recovers : the partial derivatives are the two special cases. The definition uses a unit vector so that really measures distance; with a longer vector the same limit would be scaled by its length.
Computing this limit from scratch every time would be painful. The chain rule turns it into a dot product.
Theorem 11.47 (Computing directional derivatives). If is differentiable at , then has a directional derivative there in every direction, and for a unit vector ,
Proof. Define the one-variable function . By the definition of the derivative,
On the other hand with and , so by the chain rule (case 1),
Setting gives , , and the two expressions for agree.∎
The right-hand side is a dot product of with the vector of partial derivatives, which deserves a name.
Definition 11.48 (Gradient). The gradient of is the vector function
With it, the directional derivative reads
Example 11.49 (Directional derivative in a given direction). Find the directional derivative of at in the direction making angle with the positive -axis.
Solution. The unit vector is . The partials are and , so at , . Then
The value is positive: heading at from (1,2), increases. It should, since the direction has a large component along and dominates.□
Example 11.50 (Direction given by a non-unit vector). Find the directional derivative of at in the direction of .
Solution. First the gradient: , so . The vector is not a unit vector; , so the direction is . Then
Forgetting to normalise would give , wrong by a factor of .□
For a function of three variables the definition and theorem read the same with a third component: and for unit in .
Example 11.51 (A directional derivative in space). If , find the directional derivative of at in the direction of .
Solution. , and at this is . The unit vector is , so
The formula answers a natural question: among all directions from a point, which one makes increase fastest, and how fast?
Theorem 11.52 (Maximising the directional derivative). Suppose is differentiable and . The maximum value of over all unit vectors is , attained when points in the direction of . The minimum value is , attained in the direction of .
Proof. By the geometric form of the dot product, , where is the angle between and . The factor ranges over , is exactly when , i.e. is parallel to , and is exactly when .∎
So the gradient vector points uphill, and its length is the steepness of the steepest ascent. Perpendicular to it, and does not change at all — that is the direction along the level curve, a fact made precise below.
Intuition. On the hillside from the partial-derivatives section, with eastward slope and northward slope , the gradient is . It points somewhat south of east; that is the direction a ball would not roll — it is straight uphill — and the ball rolls the opposite way. The steepest slope is , a bit more than the eastward slope alone, because a little southward tilt adds to it. Walk perpendicular to the gradient and you stay on your contour line, neither climbing nor descending.
Example 11.53 (Steepest ascent). For , find the rate of change of at in the direction from to . In what direction does increase fastest at , and what is that maximum rate?
Solution. , so . The vector from to is , of length , so the unit direction is and
The function increases fastest in the direction of , and the maximum rate is . The rate toward , being , is indeed less than .□
Example 11.54 (Temperature in space). The temperature at is . In which direction does the temperature increase fastest at , and what is the maximum rate of increase?
Solution. Write with . By the chain rule,
At , , so and
The temperature increases fastest in the direction of — toward the origin, where the denominator is smallest, as it should — and the maximum rate is degrees per unit distance.□
The claim that the gradient is perpendicular to the level curve is now a two-line consequence of the chain rule, and it extends to level surfaces in space.
Theorem 11.55 (The gradient is normal to level sets). Let be differentiable and let be the level surface . If lies on and , then is perpendicular to the tangent vector of every smooth curve on through . The same holds in the plane: is perpendicular to the level curve through .
Proof. Let be any smooth curve lying on , with . Because the curve lies on , for every . Differentiate this identity with respect to by the chain rule:
At this says : the gradient is perpendicular to the tangent vector at . The planar statement is the same argument with one fewer coordinate.∎
Since every tangent vector at is perpendicular to , they all lie in one plane, the tangent plane to at , whose normal vector is . Writing the plane through with that normal gives
and the normal line to at — the line through in the direction of — has symmetric equations
If the surface is given as a graph , take ; then , , , and the tangent plane equation becomes , exactly the formula derived from slices in the previous section. The two approaches agree, and the level-surface one is more general because it handles surfaces that are not graphs.
Example 11.56 (Tangent plane and normal line to an ellipsoid). Find the tangent plane and the normal line to the ellipsoid at the point .
Solution. The ellipsoid is the level surface of . Its gradient is
Tangent plane:
Normal line:
Check that the point lies on the plane: .□
Example 11.57 (Gradient perpendicular to a level curve). Verify directly that is perpendicular to the level curve of through .
Solution. The level curve through is , a circle of radius . The gradient is , so , which points radially outward from the origin — visibly perpendicular to the circle. For a check by computation, the circle can be parametrised as ; at the point the tangent vector is , and
Pitfall. The gradient of is a vector in the plane, not on the surface ; it lives in the domain and points along the direction of steepest ascent as read off the contour map. When you want a normal vector to the surface itself, use the three-dimensional gradient of . Confusing the two produces a "normal" that is missing its -component.
11.7Maximum and minimum values
Definition 11.58 (Local and absolute extrema). A function of two variables has a local maximum at if for all in some disc centred at ; the number is a local maximum value. If the inequality holds for all in the domain of , then has an absolute maximum at . Local and absolute minima are defined with .
Theorem 11.59 (Fermat's Theorem for two variables). If has a local maximum or minimum at and the first-order partial derivatives of exist there, then
Proof. Let . If has a local maximum at then has a local maximum at , so by the one-variable Fermat's Theorem ; but . Applying the same argument to gives .∎
Geometrically, at a local extremum the tangent plane is horizontal. As in one variable the theorem runs in one direction only: the partials vanishing is necessary for an extremum, not sufficient.
Definition 11.60 (Critical point). A point is a critical point of if and , or if one of these partial derivatives does not exist.
Example 11.61 (A critical point that is a minimum). Find the extreme values of .
Solution. and vanish only at : the single critical point. Completing the squares,
and since squares are non-negative, for every . So is a local minimum and in fact the absolute minimum; the graph is a paraboloid with vertex . There is no maximum, since far from the origin.□
Example 11.62 (A critical point that is neither). Find the extreme values of .
Solution. and , so the only critical point is , where . Along the -axis for , so is not a minimum; along the -axis , so it is not a maximum. The origin is a saddle point: a minimum along one line and a maximum along another. The graph is the hyperbolic paraboloid, shaped like a horse's saddle, and has no extreme values.□
Deciding which of these three things happens at a critical point is the job of the second derivatives test. It is the two-variable analogue of the second derivative test, and the quantity that decides is not alone but a combination of all three second partials.
Theorem 11.63 (Second Derivatives Test). Suppose the second partial derivatives of are continuous on a disc centred at , and suppose and . Let
- If and , then is a local minimum.
- If and , then is a local maximum.
- If , then is a saddle point of .
- If , the test gives no information.
Proof. Here is the idea, with the details of the error term left to real analysis. Write , , . Since the first partials vanish, the change in for a small step is governed by the second-order Taylor term, a quadratic form:
with an error small compared with because the second partials are continuous. So the sign of for small steps is the sign of , and the question is whether is always positive, always negative, or takes both signs. Suppose and complete the square in :
If , both coefficients and have the sign of , so has the sign of for every : a minimum when , a maximum when . (When we automatically have , since .) If , choose : then has the sign of ; now choose : then has the opposite sign. Two directions with opposite signs of is a saddle. If and then and , which also changes sign. When the quadratic form vanishes along some direction and the ignored higher-order terms decide, which is why the test is silent.∎
For a unit vector the quantity is the second directional derivative of in that direction, so the argument says: at a minimum every second directional derivative is positive, at a maximum every one is negative, and at a saddle some are positive and some negative. The determinant is what makes the sign of decidable from three numbers instead of infinitely many directions. Note the symmetry: when , and have the same sign, so either may be used in the test.
Example 11.64 (Classifying several critical points). Find and classify the critical points of .
Solution. Setting the partials to zero,
so and . Then gives or (the other roots of are not real). The critical points are , and .
The second partials are , , , so
At : , a saddle point. At : and , a local minimum with value . At : and , another local minimum, also with value . Since as , these two local minima are the absolute minimum.□
Example 11.65 (When ). Show that , and all have at the origin, yet behave differently there.
Solution. For each, the first partials vanish at the origin, so it is a critical point. The second partials , and all vanish at , so in each case, and the test says nothing.
Direct inspection settles it: , so the origin is a minimum of ; likewise it is a maximum of ; and while , so it is a saddle of . Three functions with identical second-order data and three different answers — the case genuinely cannot be decided by second derivatives.□
Example 11.66 (Shortest distance to a plane). Find the shortest distance from the point to the plane .
Solution. A point on the plane has , so its squared distance to is
Minimising the square avoids a square root and gives the same minimiser. The partials are
Solving and : subtracting, , so and . At this point , , , so with : a local minimum, and by the geometry the absolute one. The squared distance is
so the distance is . The point-to-plane formula agrees.□
Example 11.67 (Largest box with a given surface area). A rectangular box without a lid is to be made from m of cardboard. Find the maximum volume of such a box.
Solution. Let the box have base by and height . Then and the area of the base plus four sides is . Solving the constraint for ,
By the quotient rule, after simplification,
For both vanish only when and . Subtracting, , so , and then gives and . The single critical point is with m. That it is a maximum is clear from the physical situation — as the box becomes very flat or very thin — and the second derivatives test would confirm it. The box is .□
Intuition. Water settles at the bottom of a bowl. Which way is "down" at a critical point? Nowhere — that is what the vanishing partials say. Whether you are at the bottom of a bowl, the top of a dome or the middle of a mountain pass is a question about curvature, and answers it. A pass has the fingerprint : the ground curves up along the ridge and down along the road. A bowl and a dome both have , curving the same way in every direction, and says which.
Local extrema are found at critical points; absolute extrema need one more ingredient. In one variable a continuous function on a closed interval attains its extremes either at a critical point or at an endpoint. The two-variable analogue of a closed interval is a closed, bounded set: closed meaning it contains all its boundary points, bounded meaning it fits inside some disc. Discs, rectangles and triangles with their edges included are closed and bounded; the open disc is not closed, and the half-plane is not bounded.
Theorem 11.68 (Extreme Value Theorem for two variables). If is continuous on a closed, bounded set in , then attains an absolute maximum value and an absolute minimum value at some points and in .
The proof is real analysis (compactness), but the consequence is a procedure. An absolute extremum occurring in the interior of is also a local extremum, hence at a critical point by Fermat's Theorem; otherwise it occurs on the boundary. So:
Method 11.69 (Absolute extrema on a closed bounded set).
- Find the values of at the critical points of in the interior of .
- Find the extreme values of on the boundary of . Parametrise each boundary piece as a curve, so that restricted to it becomes a function of one variable on a closed interval, and apply the one-variable closed-interval method.
- The largest of the values from steps 1 and 2 is the absolute maximum; the smallest is the absolute minimum.
Example 11.70 (Absolute extrema on a rectangle). Find the absolute maximum and minimum values of on the rectangle .
Solution. Step 1. and vanish together only at , which is inside , with .
Step 2. The boundary has four edges. On the bottom edge , : , increasing, with minimum at and maximum at . On the right edge , : , decreasing from at to at . On the top edge , : , with minimum at and maximum at . On the left edge , : , from at to at .
Step 3. Collecting the candidates : the absolute maximum is at and the absolute minimum is , attained at both and .□
Example 11.71 (Absolute extrema on a triangle). Find the absolute extrema of on the closed triangular region with vertices , and .
Solution. Step 1. and give the single critical point . It lies inside , since , and .
Step 2. The boundary has three edges. On the edge , : , with at , so the candidates are , , . On the edge , : , with at , giving , , . On the hypotenuse , parametrise by with :
with at , so the candidates are , , .
Step 3. The candidate values are . The absolute maximum is at the vertex ; the absolute minimum is at the interior point . Sanity check: , , , so and the interior critical point is indeed a local minimum.□
Example 11.72 (Absolute extrema on a disc). Find the absolute extrema of on the closed disc .
Solution. Step 1. and give the critical point , inside the disc, with .
Step 2. Parametrise the boundary circle by , , :
using . Since ranges over , ranges over : minimum at , the point , and maximum at , the point .
Step 3. Comparing : the absolute maximum is at and the absolute minimum is at . The next section redoes the boundary step with Lagrange multipliers, which avoids the parametrisation.□
Pitfall. The second derivatives test classifies interior critical points only. A point on the boundary of a closed region can be an absolute extremum with no partial derivative vanishing at all — in the rectangle example the maximum sits at a corner where and . Never apply the test to boundary candidates, and never skip the boundary because the interior critical point "looks like" the answer.
11.8Lagrange multipliers
The box example of the previous section was a constrained problem: maximise subject to . We solved it by eliminating , which was possible only because the constraint could be solved for one variable and led to messy quotients even so. Lagrange's method handles constraints without solving them.
Consider maximising subject to . Draw the constraint curve and, on the same axes, the level curves for various . Maximising on the constraint means finding the largest for which the level curve still meets the curve . Picture increasing: the level curves sweep across the constraint curve, and at the last moment of contact the two curves touch rather than cross — they have a common tangent line. Their normal vectors are therefore parallel, and the normal vectors are the gradients: for some scalar .
The same argument works in space, and there it can be made precise with the tools of the gradient section. Suppose has an extreme value at on the level surface given by . Let be any smooth curve on passing through at , and let . Since is extreme at among points of , the one-variable function has an extremum at , so ; by the chain rule,
So is perpendicular to the tangent vector of every curve on through — it is normal to at . But we know that is normal to at too. Two vectors normal to the same surface at the same point are parallel, so if there is a number with .
Method 11.73 (Method of Lagrange multipliers). To find the maximum and minimum values of subject to the constraint , assuming these exist and on the constraint surface:
- Find all values of and such that
- Evaluate at all the points found in step 1. The largest value is the maximum of on the constraint; the smallest is the minimum.
The number is called a Lagrange multiplier. For two variables drop throughout.
Step 1 is a system of four equations, the three components of the vector equation and the constraint, in the four unknowns . It is rarely linear; the usual technique is to eliminate by dividing or cross-multiplying equations, watching for the cases where a divisor might be zero. The multiplier itself is often not needed, although it has a meaning: it is the rate at which the extreme value changes as the constraint level is relaxed.
Intuition. A fence runs across a hillside and you may walk only along the fence. Where is the highest point you can reach? Not necessarily where the fence crosses the summit — it may not — but at the spot where the fence runs along a contour line. Anywhere the fence cuts across contours you could gain height by walking on; only where fence and contour are tangent is there nowhere higher to go. Tangent curves share a normal direction, and the normals are (to the fence) and (to the contour). That is the equation .
Example 11.74 (A first Lagrange problem). Find the minimum value of subject to .
Solution. Take . Then and , so the equations are
The first two give , and the constraint then gives . The only candidate is , with . Since is the squared distance from the origin and grows without bound along the line, this candidate must be the minimum; there is no maximum. Geometrically, is the square of the distance from the origin to the line, as it should be.□
Example 11.75 (The box again, by Lagrange multipliers). A rectangular box without a lid is to be made from m of cardboard. Find the maximum volume by the method of Lagrange multipliers.
Solution. Maximise subject to . The equations and are
Multiply the first equation by , the second by and the third by ; each left side becomes :
Now (otherwise , forcing two of the dimensions to be , which contradicts the constraint). So , giving since , and , giving , so . Then , and the constraint reads , so , . The maximum volume is m, as before, with no quotient rule in sight.□
Example 11.76 (Extreme values on a circle). Find the extreme values of on the circle .
Solution. With , the equations are
The first says , so or . If , the constraint gives : the points . If , the second equation gives , so , and the constraint gives : the points . Evaluating,
The maximum on the circle is at and the minimum is at . On the contour map, the level curves of are ellipses that are wider than tall; the ellipse touches the unit circle at from inside and the ellipse touches it at from outside.□
Example 11.77 (Combining an interior search with Lagrange multipliers). Find the extreme values of on the closed disc .
Solution. Interior: and vanish only at , inside the disc, with .
Boundary: by the previous example, the extreme values of on the circle are and .
Comparing : the absolute maximum on the disc is at and the absolute minimum is at . The value is only the minimum on the circle — a candidate that lost. This is the pattern for any closed bounded region with a smooth boundary: critical points inside, Lagrange on the edge, compare.□
Example 11.78 (Closest and farthest points on a sphere). Find the points on the sphere that are closest to and farthest from the point .
Solution. The distance squared from to is , to be optimised subject to . The equations are
Solving each of the first three for the coordinate (, since would force ):
Substituting in the constraint, , so and . The two candidate points are
The first is a positive multiple of and so lies on the same side of the origin: it is the closest point, at distance . The second is the farthest, at distance . Both are on the line through the origin and , which is what geometry predicts.□
A constraint curve in space is usually given as the intersection of two surfaces, and . At an extremum of on the curve, the same chain-rule argument shows is perpendicular to the tangent vector of the curve. But and are both perpendicular to that tangent vector too, since the curve lies on both surfaces, so all three gradients lie in the plane normal to the curve at . If and are not parallel, they span that plane, and is a combination of them:
Together with and this is five equations in the five unknowns .
Example 11.79 (Two constraints). Find the maximum value of on the curve of intersection of the plane and the cylinder .
Solution. Take and . Since , and , the vector equation reads
The third gives at once; then and , so and (note ). The cylinder constraint gives
So , , and from the plane . The corresponding values are
The intersection curve is an ellipse, closed and bounded, so the Extreme Value Theorem guarantees both extremes exist: the maximum is and the minimum is .□
Pitfall. Lagrange's equations produce candidates, not classifications, and the recipe only works when an extreme value is known to exist — on a closed bounded constraint set by the Extreme Value Theorem, or by a geometric argument otherwise. On an unbounded constraint such as a line or a hyperbola, a lone candidate may be a minimum, a maximum, or neither, and the values must be compared with the behaviour of far out. Also check any point of the constraint where separately, since the method's derivation assumed there.
Summary. A limit demands the same value along every approach, so two paths with different limits prove it does not exist. If and are continuous on a disc containing , then (Clairaut). If and exist near and are continuous there, is differentiable at , hence continuous, and
is both the tangent plane and the linear approximation. The chain rule chains one factor per intermediate variable, , and implicitly and where the denominators are nonzero. For differentiable and a unit vector , ; this is largest, equal to , in the direction of , and is normal to the level set . Extrema with existing partials occur only at critical points, and where the second partials are continuous on a disc about a critical point, decides: with a local minimum, with a local maximum, a saddle, no information. A continuous on a closed bounded set attains an absolute maximum and minimum, found by comparing interior critical values with the boundary; under the constraint with , the candidates solve together with .
- Checking a limit along the two axes, finding they agree, and concluding the limit exists. Two paths agreeing proves nothing; two paths disagreeing proves non-existence. To prove existence, bound by something in alone.
- Trusting straight lines in the two-path test. The function tends to along every line through the origin and to along the parabola .
- Forgetting to hold the other variable constant, or forgetting the chain rule inside a partial derivative: of is , not .
- Reading in the wrong order. Subscripts read left to right ( first), the operator form reads right to left; both mean the same thing. By Clairaut's Theorem it rarely matters, but it does at points where the mixed partials are discontinuous.
- Believing that existence of and makes differentiable, or even continuous. The right sufficient condition is *continuity* of the partials, and is the standard counterexample.
- Using a non-unit vector in . Normalise first, or the answer is scaled by the vector's length.
- Treating as a vector in three-dimensional space lying on the surface. It is a vector in the domain plane. The normal vector to the surface is .
- In the chain rule, dropping a term because a variable "does not appear": if with and , then has two terms even when only one of them ends up nonzero.
- Applying the second derivatives test with and reporting an answer. The test is silent; inspect the function directly, as with versus .
- Concluding "local minimum" from alone. The sign of is only decisive once ; with the point is a saddle regardless.
- Finding absolute extrema on a closed region from interior critical points only. The boundary must be parametrised and searched separately, endpoints of each edge included, and boundary candidates are never fed into the second derivatives test.
- Dividing a Lagrange equation by (or by ) without treating the case (or ) separately. Solutions lost this way are often exactly the extremes.
- Reporting the points found by Lagrange multipliers as "the maximum and the minimum" without evaluating at each and comparing, and without confirming an extreme value exists at all on an unbounded constraint.