Cauchy's four fundamental functional equations — additive f(x+y)=f(x)+f(y), exponential f(x+y)=f(x)f(y), logarithmic f(xy)=f(x)+f(y), and multiplicative f(xy)=f(x)f(y) — together with the substitution and injectivity/surjectivity toolkit that solves competition functional equations. We prove that every additive f:Q→Q has the form f(q)=cq, extend this to R under mild regularity, and show Jensen's equation f(2x+y)=2f(x)+f(y) reduces to the additive case. Applications include Shannon entropy's axiomatic derivation and the memoryless property of the exponential distribution.
IntuitionA Function Defined Only by a Rule
Most functions you meet are given by an explicit formula: f(x)=x2, f(x)=sinx. A functional equation instead describes a function only through a relationship it must satisfy for all inputs — for instance, "f turns sums into products of outputs" (f(x+y)=f(x)f(y)) — and your job is to deduce every possible f consistent with that single rule. This is a strange kind of detective work: instead of computing an answer, you plug in clever specific values (like x=y=0, or y=−x, or y=1/x) to squeeze out constraints until the function's entire shape is pinned down.
Linear function plot through the origin illustrating the additive Cauchy equation solution
The solution f(x)=cx (here c=2) of Cauchy's additive equation f(x+y)=f(x)+f(y): drag to see how changing the slope still keeps f(x+y)=f(x)+f(y) true for every linear function through the origin.
Cauchy studied four equations for f:R→R (or suitable subdomains): additivef(x+y)=f(x)+f(y), exponentialf(x+y)=f(x)f(y) (turns sums into products), logarithmicf(xy)=f(x)+f(y) (turns products into sums, domain restricted to positive reals), and multiplicativef(xy)=f(x)f(y). Each pair is linked by exp/log: if g solves the additive equation, f=exp∘g solves the exponential one; if f solves the multiplicative equation, g=log∘f solves the additive one (on positive reals).
f(x+y)=f(x)+f(y)⟺f=exp∘gf(x+y)=f(x)f(y)
The additive equation is the master case: every solution of the other three reduces to it via exp/log substitution (when the function is positive/nonzero as required), which is why Theorem 1 below, proved directly for the additive equation, secretly solves all four.
f(2x+y)=2f(x)+f(y)
The four Cauchy equations
Equation
Domain
General solution (regular case)
Additive
R→R
f(x)=cx
Exponential
R→R>0
f(x)=ax
Logarithmic
R>0→R
f(x)=clogx
Multiplicative
R>0→R
f(x)=xc
UndergraduateSolving Cauchy's Equation and Jensen's Equation
If f:Q→Q satisfies f(x+y)=f(x)+f(y) for all x,y∈Q, then f(q)=cq for all q∈Q, where c=f(1). If additionally f:R→R is continuous (or monotonic, or bounded on some interval), the same conclusion f(x)=cx holds for all x∈R.
Why is it true?
This theorem is the foundation of the entire functional-equations toolkit: it shows that a purely algebraic relation, with no continuity assumed on Q, already pins down f completely — and that on R, without some regularity assumption, wildly pathological non-linear solutions exist (built via a Hamel basis using the Axiom of Choice), so the regularity hypothesis is not a technicality but essential.
Proof
**Step 1: Determine f(0).** Set x=y=0 in f(x+y)=f(x)+f(y): f(0)=f(0)+f(0), so f(0)=0.
Step 2: Extend to positive integers. For a positive integer n, set x=(n−1) (or induct): f(n⋅1)=f((n−1)⋅1+1)=f((n−1)⋅1)+f(1). By induction on n, f(n)=nf(1) for every positive integer n (base case n=1 trivial, inductive step just applied).
Step 3: Extend to negative integers. Set y=−x in f(x+y)=f(x)+f(y): f(0)=f(x)+f(−x), and since f(0)=0, f(−x)=−f(x). Combined with Step 2, f(n)=nf(1) for every integer n (positive, negative, or zero), writing c=f(1).
Step 4: Extend to rationals. Let q=p/r with p∈Z, r∈Z>0. Since r⋅q=p (as integers under repeated addition, i.e. q+q+⋯+q (r times) =p), applying the integer case of Step 2/3 to the function evaluated r times gives f(rq)=rf(q) (by the same induction argument used for Step 2, now with x=q). But rq=p, so f(p)=rf(q), i.e. cp=rf(q), i.e. f(q)=c⋅rp=cq.
Step 5: Conclude the rational case. This shows f(q)=cq for every q∈Q, with c=f(1) — the entire function on Q is determined by its value at a single point.
**Step 6: Extend to R under continuity.** Suppose now f:R→R is additive and continuous at even one point (continuity everywhere then follows from additivity: f(x+h)−f(x)=f(h)→0 as h→0 if continuous at 0). For any real x, take a sequence of rationals qn→x. By Steps 1–5, f(qn)=cqn. Continuity gives f(x)=limnf(qn)=limncqn=cx.
**Step 7: Extend to R under monotonicity or local boundedness (sketch).** If f is monotonic, then for rationals q1<x<q2 squeezing any real x, monotonicity forces cq1≤f(x)≤cq2 (if c>0; reverse if c<0), and letting q1,q2→x pins f(x)=cx by the same squeeze. If instead f is bounded on some interval I, one shows f is bounded near 0 (using additivity to shift the interval), then f(x/n)→0 as n→∞ for fixed x forces continuity at 0, reducing to Step 6. In all three regularity cases (continuous, monotonic, bounded on an interval), the conclusion is the same: f(x)=cx for all x∈R.
If f:R→R satisfies Jensen's equation f(2x+y)=2f(x)+f(y) for all x,y∈R, then g(x):=f(x)−f(0) is additive (satisfies f(x+y)=f(x)+f(y)), so under continuity (or monotonicity, or boundedness on an interval), f(x)=cx+f(0) for some constant c.
Why is it true?
This shows Jensen's equation — which looks like a statement purely about midpoints and averages — is secretly Cauchy's additive equation wearing a disguise, so all the machinery of Theorem 1 (including the pathological non-regular solutions and the regularity conditions that rule them out) transfers over automatically.
Proof
**Step 1: Define g and check g(0)=0.** Let g(x)=f(x)−f(0). Then g(0)=f(0)−f(0)=0.
**Step 2: Rewrite Jensen's equation in terms of g.** Substituting f=g+f(0) into f(2x+y)=2f(x)+f(y) gives g(2x+y)+f(0)=2g(x)+f(0)+g(y)+f(0)=2g(x)+g(y)+f(0). The f(0) terms cancel, leaving g(2x+y)=2g(x)+g(y) — g satisfies the exact same Jensen equation.
Step 3: Derive the halving identity. Set y=0 in the equation for g: g(2x)=2g(x)+g(0)=2g(x) (using g(0)=0 from Step 1). So g(x/2)=g(x)/2 for every x, equivalently g(2u)=2g(u) for every u (substitute u=x/2).
**Step 4: Convert Jensen's equation for g into additivity.** For arbitrary x,y∈R, apply g's Jensen equation with the pair (x,y): g(2x+y)=2g(x)+g(y). By Step 3 with u=x+y, the left side equals g(x+y)/2 (since 2x+y is the halving of x+y). So 2g(x+y)=2g(x)+g(y), and multiplying both sides by 2: g(x+y)=g(x)+g(y).
Step 5: Conclude. This is exactly the additive Cauchy equation f(x+y)=f(x)+f(y) applied to g. By Theorem 1, if g (equivalently f, since they differ by the constant f(0)) is continuous, monotonic, or bounded on some interval, then g(x)=cx for some constant c=g(1)=f(1)−f(0). Substituting back, f(x)=g(x)+f(0)=cx+f(0), the general regular solution of Jensen's equation — an affine (not necessarily linear) function.
AdvancedReal-World Applications and Worked Examples
Functional equations are not just puzzles: Shannon's axiomatic derivation of entropy assumes that the uncertainty of two independent events with probabilities p and q satisfies H(pq)=H(p)+H(q), exactly Cauchy's logarithmic equation, forcing H(p)=−klogp for some constant k>0 — this single functional equation, plus continuity, is the reason entropy must be logarithmic. In probability theory, the memoryless property of waiting times (a bus is just as likely to arrive in the next 5 minutes whether you've waited 0 or 20 minutes already) is exactly Cauchy's exponential equation applied to the survival function, forcing the exponential distribution to be the unique continuous memoryless distribution.
Example: Deriving Shannon Entropy's Logarithmic Form
Assume the uncertainty function H:(0,1]→R≥0 of a single event with probability p satisfies H(pq)=H(p)+H(q) for independent events with probabilities p,q, and H is continuous. Show H(p)=−klogp for some constant k≥0.
Solution
Step 1: Substitute p=e−u, q=e−v for u,v≥0 and define G(u):=H(e−u). Then H(pq)=H(p)+H(q) becomes H(e−ue−v)=H(e−u)+H(e−v), i.e. H(e−(u+v))=G(u)+G(v), i.e. G(u+v)=G(u)+G(v) — Cauchy's additive equation for G on [0,∞).
Step 2: Since H is continuous and p↦e−u is continuous, G is continuous. By Theorem 1's continuity case, G(u)=ku for some constant k=G(1)=H(e−1).
Step 3: Since H≥0 (uncertainty is nonnegative) and probabilities p≤1 correspond to u=−logp≥0, we need G(u)=ku≥0 for all u≥0, forcing k≥0.
Step 4: Undo the substitution: H(p)=H(e−u)=G(u)=ku=k(−logp)=−klogp, as required.
Example: The Memoryless Property Forces the Exponential Distribution
Let S(t)=P(X>t) be the survival function of a continuous random variable X≥0, and suppose X is memoryless: P(X>s+t∣X>t)=P(X>s) for all s,t≥0. Show S(t)=e−λt for some λ>0, i.e. X is exponentially distributed.
Solution
Step 1: Rewrite the conditional probability using the definition of conditional probability: P(X>s+t∣X>t)=P(X>t)P(X>s+t,X>t)=P(X>t)P(X>s+t)=S(t)S(s+t) (using X>s+t⇒X>t for s≥0, so the joint event is just X>s+t).
Step 2: The memoryless assumption P(X>s+t∣X>t)=P(X>s) becomes S(t)S(s+t)=S(s), i.e. S(s+t)=S(s)S(t) for all s,t≥0 — exactly Cauchy's multiplicative-turned-exponential equation for S.
Step 3: S is monotonic (non-increasing, since it's a survival function: P(X>t) decreases as t increases) and continuous (since X is a continuous random variable), and 0≤S(t)≤1 so S is bounded. By the regularity extension in Theorem 1 (applied to g(t):=logS(t), which satisfies g(s+t)=g(s)+g(t) by taking logs of Step 2, and is monotonic/continuous since log is monotonic and S is), g(t)=−λt for some constant λ (writing −λ=g(1)=logS(1)).
Step 4: Undo the logarithm: S(t)=eg(t)=e−λt. Since S is non-increasing and S(0)=1, we need λ≥0; if λ=0 then S≡1, not a valid (non-degenerate) probability distribution, so λ>0. This is exactly the survival function of the exponential distribution with rate λ, proving the memoryless property forces X to be exponentially distributed.
For f:Q→Q satisfying f(x+y)=f(x)+f(y) for all rationals, what is f(3/2) in terms of c=f(1)?
Why can't Theorem 1's proof for Q be extended to R without any extra hypothesis (continuity, monotonicity, or boundedness)?
In the proof that Jensen's equation reduces to Cauchy's, what substitution g(x) is used?
The memoryless property of a continuous waiting-time distribution translates into which Cauchy equation for the survival function S(t)=P(X>t)?