MathLabs

Analysis

Partial derivatives and the gradient

Rates of change of a multivariable function along each coordinate direction, combined into the gradient vector.

IntuitionTwo slopes at once on a curved landscape

Imagine standing on a hillside whose elevation is given by a two-variable function z=f(x,y)z = f(x,y), where xx points east and yy points north. Unlike a single-variable curve, there is no single slope at your feet: stepping due east changes your height at one rate, stepping due north changes it at another, and stepping at an angle blends the two. Freezing yy and letting only xx vary slices the hillside with an east-west vertical plane and reduces the problem to an ordinary one-variable derivative.

Interactive 3D surface plot illustrating coordinate cross-sections and tangent slopes.
Interactive 3D surface z=f(x,y)z = f(x,y): slicing along fixed yy reveals the curve whose slope is ∂f∂x\frac{\partial f}{\partial x}, while slicing along fixed xx gives ∂f∂y\frac{\partial f}{\partial y}.

UndergraduateFormal definitions: partial derivatives, tangent plane, and the gradient

Definition: Partial derivative and gradient vector

Let f:U⊆Rn→Rf : U \subseteq \mathbb{R}^n \to \mathbb{R} be defined on an open set UU. The partial derivative of ff with respect to xix_i at a∈Ua \in U, denoted ∂f∂xi(a)\frac{\partial f}{\partial x_i}(a) or fxi(a)f_{x_i}(a), is the limit keeping all other coordinates xjx_j (j≠ij \ne i) fixed. Assembling all nn partial derivatives into a single vector produces the gradient ∇f(a)=(∂f∂x1(a),…,∂f∂xn(a))\nabla f(a) = \left(\frac{\partial f}{\partial x_1}(a), \dots, \frac{\partial f}{\partial x_n}(a)\right).

∂f∂xi(a)=lim⁡h→0f(a1,…,ai+h,…,an)−f(a1,…,an)h\frac{\partial f}{\partial x_i}(a) = \lim_{h \to 0} \frac{f(a_1, \dots, a_i + h, \dots, a_n) - f(a_1, \dots, a_n)}{h}

Mere existence of partial derivatives at (x0,y0)(x_0,y_0) is weaker than total differentiability. A function f(x,y)f(x,y) is totally differentiable at (x0,y0)(x_0,y_0) if the increment Δf=f(x0+h,y0+k)−f(x0,y0)\Delta f = f(x_0+h, y_0+k) - f(x_0,y_0) can be written as fx(x0,y0)h+fy(x0,y0)k+o(h2+k2)f_x(x_0,y_0)h + f_y(x_0,y_0)k + o\left(\sqrt{h^2+k^2}\right) as (h,k)→(0,0)(h,k) \to (0,0). When ff is totally differentiable, the graph z=f(x,y)z = f(x,y) admits a well-defined tangent plane at (x0,y0)(x_0,y_0), and the directional derivative along any unit vector u=(u1,u2)\mathbf{u} = (u_1, u_2) with ∥u∥=1\|\mathbf{u}\| = 1 is given by the dot product Duf=∇f⋅uD_{\mathbf{u}}f = \nabla f \cdot \mathbf{u}.

z=f(x0,y0)+∂f∂x(x0,y0)(x−x0)+∂f∂y(x0,y0)(y−y0),Duf(x0,y0)=∇f(x0,y0)⋅uz = f(x_0,y_0) + \frac{\partial f}{\partial x}(x_0,y_0)(x - x_0) + \frac{\partial f}{\partial y}(x_0,y_0)(y - y_0), \qquad D_{\mathbf{u}}f(x_0,y_0) = \nabla f(x_0,y_0) \cdot \mathbf{u}
Summary of multivariable differential operators and second-order objects
ObjectFormulaGeometric / physical role
Partial derivative∂f∂xi\frac{\partial f}{\partial x_i}Slope of the cross-section along coordinate axis xix_i.
Gradient vector∇f=(fx1,…,fxn)\nabla f = (f_{x_1}, \dots, f_{x_n})Points orthogonal to level sets toward steepest ascent; norm ∥∇f∥\|\nabla f\| is maximum rate.
Directional derivativeDuf=∇f⋅uD_{\mathbf{u}}f = \nabla f \cdot \mathbf{u}Instantaneous rate of change when moving with unit velocity u\mathbf{u}.
Hessian matrix & discriminantD=fxxfyy−fxy2D = f_{xx}f_{yy} - f_{xy}^2Measures local curvature to classify critical points into extrema or saddle points.

UndergraduateCore theorems: Steepest ascent, Schwarz symmetry, and the Hessian test

If ff is differentiable at aa and ∇f(a)≠0\nabla f(a) \ne \mathbf{0}, then for any unit vector u\mathbf{u} (with ∥u∥=1\|\mathbf{u}\| = 1), the directional derivative satisfies Duf(a)=∇f(a)⋅u=∥∇f(a)∥cos⁡θD_{\mathbf{u}}f(a) = \nabla f(a) \cdot \mathbf{u} = \|\nabla f(a)\| \cos\theta, where θ\theta is the angle between ∇f(a)\nabla f(a) and u\mathbf{u}. Consequently, Duf(a)D_{\mathbf{u}}f(a) attains its maximum ∥∇f(a)∥\|\nabla f(a)\| when u=∇f(a)∥∇f(a)∥\mathbf{u} = \frac{\nabla f(a)}{\|\nabla f(a)\|}, and its minimum −∥∇f(a)∥-\|\nabla f(a)\| in the opposite direction.

Why is it true?

The dot product projects the gradient onto your chosen walking direction: you gain the full magnitude of the gradient only when you align perfectly with it, and zero change when you walk perpendicular to it along a contour line.

Proof

Define the single-variable function g(t)=f(a+tu)g(t) = f(a + t\mathbf{u}). By definition of the directional derivative, Duf(a)=g′(0)D_{\mathbf{u}}f(a) = g'(0). Since ff is totally differentiable at aa, we have f(a+tu)−f(a)=∇f(a)⋅(tu)+o(∣t∣)f(a + t\mathbf{u}) - f(a) = \nabla f(a) \cdot (t\mathbf{u}) + o(|t|) as t→0t \to 0. Dividing by tt and taking the limit t→0t \to 0 yields Duf(a)=∇f(a)⋅uD_{\mathbf{u}}f(a) = \nabla f(a) \cdot \mathbf{u}.

By the geometric formula for the Euclidean inner product and the unit-length condition ∥u∥=1\|\mathbf{u}\| = 1, we obtain ∇f(a)⋅u=∥∇f(a)∥ ∥u∥cos⁡θ=∥∇f(a)∥cos⁡θ\nabla f(a) \cdot \mathbf{u} = \|\nabla f(a)\|\,\|\mathbf{u}\|\cos\theta = \|\nabla f(a)\|\cos\theta. Since −1≤cos⁡θ≤1-1 \le \cos\theta \le 1, the maximum value ∥∇f(a)∥\|\nabla f(a)\| is achieved uniquely at θ=0\theta = 0, where u=∇f(a)∥∇f(a)∥\mathbf{u} = \frac{\nabla f(a)}{\|\nabla f(a)\|}, and the minimum −∥∇f(a)∥-\|\nabla f(a)\| is achieved at θ=π\theta = \pi.

If f(x,y)f(x,y) has continuous second-order partial derivatives on a neighborhood of (x0,y0)(x_0,y_0), then ∂2f∂x∂y(x0,y0)=∂2f∂y∂x(x0,y0)\frac{\partial^2 f}{\partial x\partial y}(x_0,y_0) = \frac{\partial^2 f}{\partial y\partial x}(x_0,y_0). Moreover, if ∇f(x0,y0)=0\nabla f(x_0,y_0) = \mathbf{0} and D=fxx(x0,y0)fyy(x0,y0)−fxy(x0,y0)2D = f_{xx}(x_0,y_0)f_{yy}(x_0,y_0) - f_{xy}(x_0,y_0)^2, then (x0,y0)(x_0,y_0) is a strict local minimum if D>0D > 0 and fxx>0f_{xx} > 0, a strict local maximum if D>0D > 0 and fxx<0f_{xx} < 0, and a saddle point if D<0D < 0.

Why is it true?

Symmetry of mixed partials guarantees the Hessian matrix is symmetric, and completing the square on its quadratic form reveals whether the surface bowls upward, domes downward, or twists like a saddle around a flat tangent plane.

Proof

For small nonzero h,kh, k, consider the second difference Δ(h,k)=f(x0+h,y0+k)−f(x0+h,y0)−f(x0,y0+k)+f(x0,y0)\Delta(h,k) = f(x_0+h,y_0+k) - f(x_0+h,y_0) - f(x_0,y_0+k) + f(x_0,y_0). Applying the one-variable Mean Value Theorem first to g(x)=f(x,y0+k)−f(x,y0)g(x) = f(x,y_0+k) - f(x,y_0) on [x0,x0+h][x_0, x_0+h] and then to fxf_x along yy yields Δ(h,k)=hk fyx(c1,d1)\Delta(h,k) = hk\,f_{yx}(c_1, d_1), while applying it in the reverse order yields Δ(h,k)=hk fxy(c2,d2)\Delta(h,k) = hk\,f_{xy}(c_2, d_2) for intermediate points converging to (x0,y0)(x_0,y_0). Equating the two and letting (h,k)→(0,0)(h,k) \to (0,0) proves ∂2f∂x∂y(x0,y0)=∂2f∂y∂x(x0,y0)\frac{\partial^2 f}{\partial x\partial y}(x_0,y_0) = \frac{\partial^2 f}{\partial y\partial x}(x_0,y_0) by continuity.

At a critical point where ∇f(x0,y0)=0\nabla f(x_0,y_0) = \mathbf{0}, Taylor's formula gives f(x0+h,y0+k)−f(x0,y0)=12Q(h,k)+o(h2+k2)f(x_0+h,y_0+k) - f(x_0,y_0) = \frac{1}{2}Q(h,k) + o(h^2+k^2), where Q(h,k)=fxxh2+2fxyhk+fyyk2Q(h,k) = f_{xx}h^2 + 2f_{xy}hk + f_{yy}k^2. When fxx≠0f_{xx} \ne 0, completing the square writes Q(h,k)=1fxx[(fxxh+fxyk)2+Dk2]Q(h,k) = \frac{1}{f_{xx}}\left[(f_{xx}h + f_{xy}k)^2 + Dk^2\right] with D=fxxfyy−fxy2D = f_{xx}f_{yy} - f_{xy}^2. If D>0D > 0, the bracketed sum is strictly positive for all (h,k)≠(0,0)(h,k) \ne (0,0), so the sign of Q(h,k)Q(h,k) matches the sign of fxxf_{xx} (giving a strict minimum for fxx>0f_{xx} > 0 and maximum for fxx<0f_{xx} < 0). If D<0D < 0, Q(h,k)Q(h,k) takes both positive and negative values along different lines through the origin, producing a saddle point.

UndergraduateReal-World Applications and Worked Examples

Partial derivatives and gradients drive modern science and engineering: in physics, Fourier's heat conduction law states that heat flux is proportional to the negative temperature gradient −∇T-\nabla T; in machine learning, training a neural network minimizes a loss function L(θ)L(\theta) by repeatedly stepping along −∇L(θ)-\nabla L(\theta) (gradient descent); and in economics and mechanics, the Hessian test distinguishes stable equilibria (local minima of potential energy) from unstable saddle points.

Example: Directional derivative and steepest ascent of a temperature field

A metal plate has temperature T(x,y)=x2+3xy+y3T(x,y) = x^2 + 3xy + y^3 (in degrees Celsius) at coordinates (x,y)(x,y) in meters. Find the gradient ∇T(1,2)\nabla T(1,2), the directional derivative at (1,2)(1,2) in the direction of the vector v=(3,4)\mathbf{v} = (3,4), and the maximum rate of temperature increase at (1,2)(1,2).

Solution

First, compute the partial derivatives by differentiating with respect to one variable while holding the other constant: Tx(x,y)=2x+3yT_x(x,y) = 2x + 3y and Ty(x,y)=3x+3y2T_y(x,y) = 3x + 3y^2. Evaluating at (1,2)(1,2) gives Tx(1,2)=2(1)+3(2)=8T_x(1,2) = 2(1) + 3(2) = 8 and Ty(1,2)=3(1)+3(2)2=15T_y(1,2) = 3(1) + 3(2)^2 = 15, so the gradient vector is ∇T(1,2)=(8,15)\nabla T(1,2) = (8, 15).

Next, normalize v=(3,4)\mathbf{v} = (3,4) to obtain the unit direction vector u=v∥v∥=(35,45)\mathbf{u} = \frac{\mathbf{v}}{\|\mathbf{v}\|} = \left(\frac{3}{5}, \frac{4}{5}\right) since ∥v∥=32+42=5\|\mathbf{v}\| = \sqrt{3^2+4^2} = 5. The directional derivative is DuT(1,2)=∇T(1,2)⋅u=8⋅35+15⋅45=845=16.8D_{\mathbf{u}}T(1,2) = \nabla T(1,2) \cdot \mathbf{u} = 8 \cdot \frac{3}{5} + 15 \cdot \frac{4}{5} = \frac{84}{5} = 16.8. The maximum rate of increase at (1,2)(1,2) occurs in the direction of ∇T(1,2)\nabla T(1,2) and equals ∥∇T(1,2)∥=82+152=289=17\|\nabla T(1,2)\| = \sqrt{8^2 + 15^2} = \sqrt{289} = 17 degrees Celsius per meter.

Example: Classifying critical points with the Hessian determinant

Find and classify all critical points of the function f(x,y)=x3+y3−3xyf(x,y) = x^3 + y^3 - 3xy using the second-derivative Hessian test.

Solution

Set both first partial derivatives to zero: fx(x,y)=3x2−3y=0f_x(x,y) = 3x^2 - 3y = 0 and fy(x,y)=3y2−3x=0f_y(x,y) = 3y^2 - 3x = 0. From the first equation we have y=x2y = x^2; substituting into the second gives 3(x2)2−3x=3x(x3−1)=03(x^2)^2 - 3x = 3x(x^3 - 1) = 0, which has real solutions x=0x = 0 and x=1x = 1. Thus the two critical points are (0,0)(0,0) and (1,1)(1,1).

Now compute the second partials: fxx=6xf_{xx} = 6x, fyy=6yf_{yy} = 6y, and fxy=−3f_{xy} = -3, giving the Hessian discriminant D(x,y)=fxxfyy−fxy2=36xy−9D(x,y) = f_{xx}f_{yy} - f_{xy}^2 = 36xy - 9. At (0,0)(0,0), we have D(0,0)=−9<0D(0,0) = -9 < 0, so (0,0)(0,0) is a saddle point. At (1,1)(1,1), we have D(1,1)=36−9=27>0D(1,1) = 36 - 9 = 27 > 0 and fxx(1,1)=6>0f_{xx}(1,1) = 6 > 0, so (1,1)(1,1) is a strict local minimum with minimum value f(1,1)=−1f(1,1) = -1.

Let f(x,y)=x3y2−4x+5yf(x,y) = x^3 y^2 - 4x + 5y. What is the gradient vector ∇f(1,2)\nabla f(1, 2)?

If ∇f(x0,y0)=(6,−8)\nabla f(x_0, y_0) = (6, -8), what is the maximum possible directional derivative Duf(x0,y0)D_{\mathbf{u}}f(x_0, y_0) over all unit vectors u\mathbf{u}?

In machine learning, gradient descent updates parameters θ\theta to minimize a differentiable loss function L(θ)L(\theta). Which update step moves θ\theta in the direction of steepest local decrease of LL?

Suppose (x0,y0)(x_0, y_0) is a critical point of a smooth function f(x,y)f(x,y) where fxx(x0,y0)=−4f_{xx}(x_0,y_0) = -4, fyy(x0,y0)=−3f_{yy}(x_0,y_0) = -3, and fxy(x0,y0)=2f_{xy}(x_0,y_0) = 2. How is (x0,y0)(x_0, y_0) classified?

References

  1. Jerrold E. Marsden, Anthony J. Tromba (2012). Vector Calculus
  2. James Stewart (2015). Calculus: Early Transcendentals
  3. Augustin-Louis Cauchy (1847). Méthode générale pour la résolution des systèmes d'équations simultanées