MathLabs

Algebra

Bilinear and quadratic forms

Functions linear in each of two vector arguments, or quadratic in one, used to define lengths and angles.

IntuitionMeasuring shape with two vectors at once

The familiar dot product u⋅v\mathbf{u}\cdot\mathbf{v} takes two vectors and returns a number, linearly in each one; a bilinear form is this idea generalized, B(u,v)=u⊤AvB(\mathbf{u},\mathbf{v}) = \mathbf{u}^\top A \mathbf{v}, allowing the matrix AA to twist and weight the vectors before combining them. Feeding the same vector into both slots gives a quadratic form Q(x)=x⊤AxQ(\mathbf{x}) = \mathbf{x}^\top A \mathbf{x}, whose level sets Q(x)=cQ(\mathbf{x})=c trace out ellipses, hyperbolas, or degenerate lines depending on the shape of AA — the conic sections from school geometry, now in any dimension.

Rotatable 3D surface showing a quadratic form as a bowl, saddle, or ridge.
The surface z=Q(x,y)z=Q(x,y) of a two-variable quadratic form: a bowl means positive definite, a saddle means indefinite, and a ridge means a degenerate direction exists.

SchoolFrom the dot product to a general bilinear form

Definition: Bilinear and quadratic form

A bilinear form on Rn\mathbb{R}^n is a function B(u,v)B(\mathbf{u},\mathbf{v}) linear in each argument separately; in coordinates it is always B(u,v)=u⊤AvB(\mathbf{u},\mathbf{v}) = \mathbf{u}^\top A \mathbf{v} for an n×nn \times n matrix AA. When A⊤=AA^\top = A the form is symmetric, and its associated quadratic form is Q(x)=x⊤AxQ(\mathbf{x}) = \mathbf{x}^\top A \mathbf{x}. Conversely, any quadratic form determines its symmetric bilinear form uniquely via the polarization identity B(u,v)=12(Q(u+v)−Q(u)−Q(v))B(\mathbf{u},\mathbf{v}) = \tfrac{1}{2}(Q(\mathbf{u}+\mathbf{v}) - Q(\mathbf{u}) - Q(\mathbf{v})), so we always take AA to be symmetric.

B(u,v)=u⊤Av,Q(x)=x⊤AxB(\mathbf{u},\mathbf{v}) = \mathbf{u}^\top A \mathbf{v}, \qquad Q(\mathbf{x}) = \mathbf{x}^\top A \mathbf{x}

Changing basis by x=Py\mathbf{x} = P\mathbf{y} (with PP invertible) replaces the matrix AA not by a similarity P−1APP^{-1}AP, but by a congruence A′=P⊤APA' = P^\top A P. In a basis where A′A' is diagonal, the quadratic form becomes a pure sum of squares d1y12+⋯+dnyn2d_1 y_1^2 + \cdots + d_n y_n^2, and the counts of positive, negative, and zero coefficients form the signature (n+,n−,n0)(n_+, n_-, n_0).

A′=P⊤AP,Q(Py)=d1y12+⋯+dnyn2A' = P^\top A P, \qquad Q(P\mathbf{y}) = d_1 y_1^2 + \cdots + d_n y_n^2
Classification of real quadratic forms by their signature
TypeSignature (n+,n−,n0)(n_+, n_-, n_0) and geometry
Positive definite (Q(x)>0Q(\mathbf{x})>0 for x≠0\mathbf{x}\neq\mathbf{0})(n,0,0)(n,0,0); level sets are ellipsoids, graph is an upward bowl
Negative definite (Q(x)<0Q(\mathbf{x})<0 for x≠0\mathbf{x}\neq\mathbf{0})(0,n,0)(0,n,0); level sets are ellipsoids, graph is a downward dome
Indefinite (takes both signs)n+>0n_+>0 and n−>0n_->0; level sets are hyperboloids, graph is a saddle
Semidefinite / degeneraten0>0n_0>0; flat directions where QQ vanishes, level sets become cylinders

UndergraduateDiagonalization and Sylvester's law of inertia

Every real quadratic form Q(x)=x⊤AxQ(\mathbf{x}) = \mathbf{x}^\top A \mathbf{x} can be reduced by an invertible linear change of variables x=Py\mathbf{x}=P\mathbf{y} to a purely diagonal form d1y12+⋯+dnyn2d_1 y_1^2 + \cdots + d_n y_n^2.

Why is it true?

You do not even need eigenvalues to diagonalize a quadratic form: ordinary high-school completing the square, applied one variable at a time, already absorbs every cross term xixjx_i x_j into a square.

Proof

We proceed by induction on the number of variables nn; the case n=1n=1 is already a single term a11x12a_{11}x_1^2.

For n>1n > 1, if the entire form is zero there is nothing to do. Otherwise, suppose first that some diagonal coefficient is nonzero; after relabeling variables we may assume a11≠0a_{11} \neq 0. Grouping every term that involves x1x_1 gives a11x12+2x1(a12x2+⋯+a1nxn)a_{11}x_1^2 + 2x_1(a_{12}x_2 + \cdots + a_{1n}x_n) plus a quadratic form in x2,…,xnx_2,\dots,x_n alone.

Setting y1=x1+a12a11x2+⋯+a1na11xny_1 = x_1 + \frac{a_{12}}{a_{11}}x_2 + \cdots + \frac{a_{1n}}{a_{11}}x_n and yk=xky_k = x_k for k≥2k \ge 2 is an invertible linear change of variables, and a11y12a_{11}y_1^2 matches all terms involving x1x_1 up to a remainder that depends only on y2,…,yny_2,\dots,y_n; hence Q=a11y12+Q′(y2,…,yn)Q = a_{11}y_1^2 + Q'(y_2,\dots,y_n).

If instead every diagonal entry vanishes (aii=0a_{ii}=0 for all ii), pick a nonzero cross term 2a12x1x22a_{12}x_1 x_2 with a12≠0a_{12} \neq 0. The invertible substitution x1=u1+u2x_1 = u_1 + u_2, x2=u1−u2x_2 = u_1 - u_2 turns 2a12x1x22a_{12}x_1 x_2 into 2a12(u12−u22)2a_{12}(u_1^2 - u_2^2), creating a nonzero diagonal coefficient and reducing us to the previous case. Applying the induction hypothesis to Q′(y2,…,yn)Q'(y_2,\dots,y_n) completes the diagonalization.

No matter which invertible change of basis is used to diagonalize a real quadratic form QQ on Rn\mathbb{R}^n, the number n+n_+ of positive coefficients, the number n−n_- of negative coefficients, and the number n0n_0 of zero coefficients are always the same; the triple (n+,n−,n0)(n_+, n_-, n_0) is an intrinsic invariant of QQ.

Why is it true?

Switching to a new coordinate system can stretch the axes and change the individual magnitudes of the diagonal numbers did_i, but it can never turn an upward-curving direction into a downward-curving one without passing through a flat direction — so the counts of upward, downward, and flat axes are locked in forever.

Proof

Suppose QQ is diagonalized in two bases {e1,…,en}\{\mathbf{e}_1,\dots,\mathbf{e}_n\} and {f1,…,fn}\{\mathbf{f}_1,\dots,\mathbf{f}_n\}, with positive, negative, and zero counts (n+,n−,n0)(n_+, n_-, n_0) in the first basis and (p+,p−,p0)(p_+, p_-, p_0) in the second. Order each basis so the positive coefficients come first, then the negative ones, then the zeros.

Suppose toward a contradiction that n+>p+n_+ > p_+. Let V+=span(e1,…,en+)V_+ = \mathrm{span}(\mathbf{e}_1,\dots,\mathbf{e}_{n_+}), a subspace of dimension n+n_+ on which Q(x)>0Q(\mathbf{x}) > 0 for every nonzero x∈V+\mathbf{x} \in V_+ (since only the positive squares in the e\mathbf{e}-expansion are active on V+V_+).

Similarly, let W−=span(fp++1,…,fn)W_- = \mathrm{span}(\mathbf{f}_{p_+ + 1},\dots,\mathbf{f}_n), a subspace of dimension n−p+n - p_+ on which Q(x)≤0Q(\mathbf{x}) \le 0 for every x∈W−\mathbf{x} \in W_- (since only the negative and zero squares in the f\mathbf{f}-expansion are active on W−W_-).

Now count dimensions inside Rn\mathbb{R}^n: since n+>p+n_+ > p_+, we have dim⁡V++dim⁡W−=n++(n−p+)>n\dim V_+ + \dim W_- = n_+ + (n - p_+) > n. By the dimension formula for subspaces, two subspaces whose dimensions add up to more than nn cannot have trivial intersection; hence there exists a nonzero vector v∈V+∩W−\mathbf{v} \in V_+ \cap W_-.

Because v∈V+\mathbf{v} \in V_+ and v≠0\mathbf{v} \neq \mathbf{0}, we must have Q(v)>0Q(\mathbf{v}) > 0; because v∈W−\mathbf{v} \in W_-, we must simultaneously have Q(v)≤0Q(\mathbf{v}) \le 0, an outright contradiction. Therefore n+≤p+n_+ \le p_+, and by symmetry of the two bases p+≤n+p_+ \le n_+, so n+=p+n_+ = p_+. Applying the exact same argument to −Q-Q gives n−=p−n_- = p_-, and finally n0=n−n+−n−=p0n_0 = n - n_+ - n_- = p_0.

UndergraduateReal-World Applications and Worked Examples

Quadratic forms and their signatures decide whether a critical point of a multivariable function is a minimum, maximum, or saddle, and they encode the causal structure of spacetime in special relativity.

Example: Second-derivative test via the Hessian quadratic form

Classify the critical point (0,0)(0,0) of f(x,y)=x2+4xy+y2f(x,y) = x^2 + 4xy + y^2 by finding the signature of its quadratic form.

Solution

Near (0,0)(0,0) the function is already a pure quadratic form Q(x,y)=x2+4xy+y2Q(x,y) = x^2 + 4xy + y^2 with symmetric matrix A=(1221)A = \begin{pmatrix} 1 & 2 \\ 2 & 1 \end{pmatrix} (note the off-diagonal entry is half the coefficient of xyxy).

Complete the square in xx: x2+4xy+y2=(x+2y)2−4y2+y2=(x+2y)2−3y2x^2 + 4xy + y^2 = (x + 2y)^2 - 4y^2 + y^2 = (x + 2y)^2 - 3y^2. In the new coordinates u=x+2yu = x + 2y, v=yv = y, this is u2−3v2u^2 - 3v^2.

There is one positive square (+u2+u^2) and one negative square (−3v2-3v^2), so the signature is (1,1,0)(1,1,0) — the form is indefinite. Along v=0v=0 the function curves upward like +u2+u^2, while along u=0u=0 it curves downward like −3v2-3v^2, so (0,0)(0,0) is a saddle point, neither a local minimum nor a local maximum.

Example: The Minkowski spacetime interval and causal signature

In special relativity (with c=1c=1), the spacetime interval between an event (t,x,y,z)(t,x,y,z) and the origin is the quadratic form s2=t2−x2−y2−z2s^2 = t^2 - x^2 - y^2 - z^2. Find its signature, and explain via Sylvester's law why every inertial observer agrees on whether two events can be causally connected.

Solution

The form s2=t2−x2−y2−z2s^2 = t^2 - x^2 - y^2 - z^2 is already diagonal with matrix η=diag(1,−1,−1,−1)\eta = \mathrm{diag}(1,-1,-1,-1): one +1+1 and three −1-1 entries, so its signature is (1,3,0)(1,3,0) — indefinite, not positive definite like Euclidean distance.

Events with s2>0s^2 > 0 (timelike separation) lie inside the light cone and can be connected by a signal slower than light; events with s2<0s^2 < 0 (spacelike separation) lie outside and cannot influence each other; events with s2=0s^2 = 0 (lightlike) are connected only by a light ray.

Switching from one inertial observer to another is a linear change of coordinates (a Lorentz transformation) that preserves the form s2s^2, and by Sylvester's law of inertia the signature (1,3,0)(1,3,0) — one time direction and three space directions — cannot be altered by any invertible coordinate change. Consequently the sign of s2s^2 is invariant, so every observer agrees on which pairs of events are timelike, spacelike, or lightlike.

What is the symmetric matrix AA associated with the quadratic form Q(x,y)=3x2−6xy+5y2Q(x,y) = 3x^2 - 6xy + 5y^2?

What is the signature (n+,n−,n0)(n_+, n_-, n_0) of Q(x,y)=x2−2xy+y2Q(x,y) = x^2 - 2xy + y^2?

What does Sylvester's law of inertia state when a real quadratic form is diagonalized in two different bases?

If the Hessian matrix of a smooth function f(x,y)f(x,y) at a critical point has quadratic form with signature (2,0,0)(2,0,0), what kind of critical point is it?

References

  1. Roger A. Horn, Charles R. Johnson (2012). Matrix Analysis (2nd ed.)
  2. Gilbert Strang (2016). Introduction to Linear Algebra