Rates of change of a multivariable function along each coordinate direction, combined into the gradient vector.
IntuitionTwo slopes at once on a curved landscape
Imagine standing on a hillside whose elevation is given by a two-variable function z=f(x,y), where x points east and y points north. Unlike a single-variable curve, there is no single slope at your feet: stepping due east changes your height at one rate, stepping due north changes it at another, and stepping at an angle blends the two. Freezing y and letting only x vary slices the hillside with an east-west vertical plane and reduces the problem to an ordinary one-variable derivative.
Interactive 3D surface plot illustrating coordinate cross-sections and tangent slopes.
Interactive 3D surface z=f(x,y): slicing along fixed y reveals the curve whose slope is ∂x∂f, while slicing along fixed x gives ∂y∂f.
UndergraduateFormal definitions: partial derivatives, tangent plane, and the gradient
Definition: Partial derivative and gradient vector
Let f:U⊆Rn→R be defined on an open set U. The partial derivative of f with respect to xi at a∈U, denoted ∂xi∂f(a) or fxi(a), is the limit keeping all other coordinates xj (j=i) fixed. Assembling all n partial derivatives into a single vector produces the gradient ∇f(a)=(∂x1∂f(a),…,∂xn∂f(a)).
Mere existence of partial derivatives at (x0,y0) is weaker than total differentiability. A function f(x,y) is totally differentiable at (x0,y0) if the increment Δf=f(x0+h,y0+k)−f(x0,y0) can be written as fx(x0,y0)h+fy(x0,y0)k+o(h2+k2) as (h,k)→(0,0). When f is totally differentiable, the graph z=f(x,y) admits a well-defined tangent plane at (x0,y0), and the directional derivative along any unit vector u=(u1,u2) with ∥u∥=1 is given by the dot product Duf=∇f⋅u.
If f is differentiable at a and ∇f(a)=0, then for any unit vector u (with ∥u∥=1), the directional derivative satisfies Duf(a)=∇f(a)⋅u=∥∇f(a)∥cosθ, where θ is the angle between ∇f(a) and u. Consequently, Duf(a) attains its maximum ∥∇f(a)∥ when u=∥∇f(a)∥∇f(a), and its minimum −∥∇f(a)∥ in the opposite direction.
Why is it true?
The dot product projects the gradient onto your chosen walking direction: you gain the full magnitude of the gradient only when you align perfectly with it, and zero change when you walk perpendicular to it along a contour line.
Proof
Define the single-variable function g(t)=f(a+tu). By definition of the directional derivative, Duf(a)=g′(0). Since f is totally differentiable at a, we have f(a+tu)−f(a)=∇f(a)⋅(tu)+o(∣t∣) as t→0. Dividing by t and taking the limit t→0 yields Duf(a)=∇f(a)⋅u.
By the geometric formula for the Euclidean inner product and the unit-length condition ∥u∥=1, we obtain ∇f(a)⋅u=∥∇f(a)∥∥u∥cosθ=∥∇f(a)∥cosθ. Since −1≤cosθ≤1, the maximum value ∥∇f(a)∥ is achieved uniquely at θ=0, where u=∥∇f(a)∥∇f(a), and the minimum −∥∇f(a)∥ is achieved at θ=π.
If f(x,y) has continuous second-order partial derivatives on a neighborhood of (x0,y0), then ∂x∂y∂2f(x0,y0)=∂y∂x∂2f(x0,y0). Moreover, if ∇f(x0,y0)=0 and D=fxx(x0,y0)fyy(x0,y0)−fxy(x0,y0)2, then (x0,y0) is a strict local minimum if D>0 and fxx>0, a strict local maximum if D>0 and fxx<0, and a saddle point if D<0.
Why is it true?
Symmetry of mixed partials guarantees the Hessian matrix is symmetric, and completing the square on its quadratic form reveals whether the surface bowls upward, domes downward, or twists like a saddle around a flat tangent plane.
Proof
For small nonzero h,k, consider the second difference Δ(h,k)=f(x0+h,y0+k)−f(x0+h,y0)−f(x0,y0+k)+f(x0,y0). Applying the one-variable Mean Value Theorem first to g(x)=f(x,y0+k)−f(x,y0) on [x0,x0+h] and then to fx along y yields Δ(h,k)=hkfyx(c1,d1), while applying it in the reverse order yields Δ(h,k)=hkfxy(c2,d2) for intermediate points converging to (x0,y0). Equating the two and letting (h,k)→(0,0) proves ∂x∂y∂2f(x0,y0)=∂y∂x∂2f(x0,y0) by continuity.
At a critical point where ∇f(x0,y0)=0, Taylor's formula gives f(x0+h,y0+k)−f(x0,y0)=21Q(h,k)+o(h2+k2), where Q(h,k)=fxxh2+2fxyhk+fyyk2. When fxx=0, completing the square writes Q(h,k)=fxx1[(fxxh+fxyk)2+Dk2] with D=fxxfyy−fxy2. If D>0, the bracketed sum is strictly positive for all (h,k)=(0,0), so the sign of Q(h,k) matches the sign of fxx (giving a strict minimum for fxx>0 and maximum for fxx<0). If D<0, Q(h,k) takes both positive and negative values along different lines through the origin, producing a saddle point.
UndergraduateReal-World Applications and Worked Examples
Partial derivatives and gradients drive modern science and engineering: in physics, Fourier's heat conduction law states that heat flux is proportional to the negative temperature gradient −∇T; in machine learning, training a neural network minimizes a loss function L(θ) by repeatedly stepping along −∇L(θ) (gradient descent); and in economics and mechanics, the Hessian test distinguishes stable equilibria (local minima of potential energy) from unstable saddle points.
Example: Directional derivative and steepest ascent of a temperature field
A metal plate has temperature T(x,y)=x2+3xy+y3 (in degrees Celsius) at coordinates (x,y) in meters. Find the gradient ∇T(1,2), the directional derivative at (1,2) in the direction of the vector v=(3,4), and the maximum rate of temperature increase at (1,2).
Solution
First, compute the partial derivatives by differentiating with respect to one variable while holding the other constant: Tx(x,y)=2x+3y and Ty(x,y)=3x+3y2. Evaluating at (1,2) gives Tx(1,2)=2(1)+3(2)=8 and Ty(1,2)=3(1)+3(2)2=15, so the gradient vector is ∇T(1,2)=(8,15).
Next, normalize v=(3,4) to obtain the unit direction vector u=∥v∥v=(53,54) since ∥v∥=32+42=5. The directional derivative is DuT(1,2)=∇T(1,2)⋅u=8⋅53+15⋅54=584=16.8. The maximum rate of increase at (1,2) occurs in the direction of ∇T(1,2) and equals ∥∇T(1,2)∥=82+152=289=17 degrees Celsius per meter.
Example: Classifying critical points with the Hessian determinant
Find and classify all critical points of the function f(x,y)=x3+y3−3xy using the second-derivative Hessian test.
Solution
Set both first partial derivatives to zero: fx(x,y)=3x2−3y=0 and fy(x,y)=3y2−3x=0. From the first equation we have y=x2; substituting into the second gives 3(x2)2−3x=3x(x3−1)=0, which has real solutions x=0 and x=1. Thus the two critical points are (0,0) and (1,1).
Now compute the second partials: fxx=6x, fyy=6y, and fxy=−3, giving the Hessian discriminant D(x,y)=fxxfyy−fxy2=36xy−9. At (0,0), we have D(0,0)=−9<0, so (0,0) is a saddle point. At (1,1), we have D(1,1)=36−9=27>0 and fxx(1,1)=6>0, so (1,1) is a strict local minimum with minimum value f(1,1)=−1.
Let f(x,y)=x3y2−4x+5y. What is the gradient vector ∇f(1,2)?
If ∇f(x0,y0)=(6,−8), what is the maximum possible directional derivative Duf(x0,y0) over all unit vectors u?
In machine learning, gradient descent updates parameters θ to minimize a differentiable loss function L(θ). Which update step moves θ in the direction of steepest local decrease of L?
Suppose (x0,y0) is a critical point of a smooth function f(x,y) where fxx(x0,y0)=−4, fyy(x0,y0)=−3, and fxy(x0,y0)=2. How is (x0,y0) classified?
References
Jerrold E. Marsden, Anthony J. Tromba (2012). Vector Calculus