MathLabs

Probability and statistics

Random variables and distributions

Numerical outcomes of random experiments, described by the probability of each possible value.

IntuitionIntuition: Turning Chance into Numbers

Flip a coin, roll a die, measure how tall the next person walking in is — each of these random experiments has an outcome you can't predict in advance. A random variable is simply a rule that assigns a number to every possible outcome, so instead of talking about vague outcomes we can compute with numbers: add them, average them, plot them.

Binomial-distribution bars with an overlaid Gaussian approximation curve.
The binomial distribution Bin(n,p)\mathrm{Bin}(n, p) (violet bars) alongside its Gaussian normal approximation (amber curve): adjust nn and pp to see how the mean npnp and spread np(1−p)\sqrt{np(1-p)} evolve.

SchoolDiscrete vs. Continuous Random Variables

Definition: Discrete random variable

A random variable XX is discrete if it takes a finite or countable list of values, each with its own probability. Its probability mass function (PMF) is P(X=x)P(X=x), and since XX must take some value, the probabilities of every possible value add up to 1.

∑xP(X=x)=1\sum_x P(X=x) = 1

The most famous discrete distribution counts successes in repeated independent trials. If a trial succeeds with probability pp and is repeated nn times, the number of successes XX follows a binomial distribution:

P(X=k)=(nk)pk(1−p)n−kP(X=k) = \binom{n}{k}p^k(1-p)^{n-k}

Here nn is the fixed number of trials, pp is the success probability of each trial, and (nk)\binom{n}{k} counts how many ways kk successes can be arranged among the nn trials, since any of those arrangements gives exactly kk successes and n−kn-k failures.

Definition: Continuous random variable

A continuous random variable can take any value in an interval of real numbers, so P(X=x)=0P(X=x)=0 for every single point xx. Instead of a PMF, it has a probability density function (PDF) f(x)f(x), and probabilities come from the area under ff over an interval.

f(x)=1σ2πe−(x−μ)2/2σ2f(x) = \frac{1}{\sigma\sqrt{2\pi}}e^{-(x-\mu)^2/2\sigma^2}

This is the normal (Gaussian) density: μ\mu is the mean, marking the center of the bell, and σ\sigma is the standard deviation, controlling how wide or narrow the bell is. Many measurement errors and natural quantities are well approximated by this curve.

Discrete vs. continuous random variables
AspectDiscreteContinuous
Describing functionProbability mass function P(X=x)P(X=x)Probability density function f(x)f(x)
Probability of one exact valueCan be positiveAlways 00
How probability is computedSum: ∑xP(X=x)\sum_x P(X=x)Integral: ∫abf(x) dx\int_a^b f(x)\,dx
Typical exampleBinomial (number of successes)Normal (measurement values)

UndergraduateExpectation and Variance

The expectation E[X]E[X] is the long-run average value of XX if the experiment were repeated many times, weighting each possible value by how likely it is:

E[X]=∑xxP(X=x)E[X] = \sum_x xP(X=x)

Each possible value xx is multiplied by its probability P(X=x)P(X=x) and the results are added up; expectation is linear, so E[aX+b]=aE[X]+bE[aX+b] = aE[X]+b for constants a,ba,b. Variance measures how spread out XX is around E[X]E[X], and is easiest to compute with the shortcut formula:

Var(X)=E[X2]−E[X]2\mathrm{Var}(X) = E[X^2] - E[X]^2

This says variance equals the average of X2X^2 minus the square of the average of XX; the standard deviation σ=Var(X)\sigma = \sqrt{\mathrm{Var}(X)} then puts the spread back in the same units as XX itself.

If XX follows a binomial distribution with nn trials and success probability pp, then E[X]=npE[X] = np.

Why is it true?

A binomial count is just a sum of many small yes/no trials, and expectation of a sum is always the sum of the expectations — no matter how the trials interact — so we only need the average contribution of a single trial.

Proof

Write X=X1+X2+⋯+XnX = X_1 + X_2 + \cdots + X_n, where XiX_i is 11 if the ii-th trial succeeds and 00 otherwise. Each XiX_i is a Bernoulli random variable with P(Xi=1)=pP(X_i=1)=p and P(Xi=0)=1−pP(X_i=0)=1-p, so its expectation is E[Xi]=1⋅p+0⋅(1−p)=pE[X_i] = 1\cdot p + 0\cdot(1-p) = p.

Expectation is linear for any random variables, whether or not they are independent: E[X1+X2+⋯+Xn]=E[X1]+E[X2]+⋯+E[Xn]E[X_1+X_2+\cdots+X_n] = E[X_1]+E[X_2]+\cdots+E[X_n]. This linearity follows directly from the definition of expectation as a weighted sum, since sums and weighted sums can always be reordered.

Applying linearity to X=X1+⋯+XnX = X_1+\cdots+X_n gives E[X]=E[X1]+⋯+E[Xn]=p+p+⋯+p⏟n times=npE[X] = E[X_1]+\cdots+E[X_n] = \underbrace{p+p+\cdots+p}_{n \text{ times}} = np, which is exactly the claimed formula.

For any random variable XX with finite mean, Var(X)=E[X2]−E[X]2\mathrm{Var}(X) = E[X^2] - E[X]^2.

Why is it true?

The definition of variance involves squaring a difference, which is awkward to compute directly; expanding that square and using linearity of expectation turns it into two easy quantities, E[X2]E[X^2] and E[X]E[X], that we already know how to compute.

Proof

By definition, Var(X)=E[(X−μ)2]\mathrm{Var}(X) = E[(X-\mu)^2] where μ=E[X]\mu = E[X]. Expand the square inside the expectation: (X−μ)2=X2−2μX+μ2(X-\mu)^2 = X^2 - 2\mu X + \mu^2.

Apply linearity of expectation term by term: E[(X−μ)2]=E[X2]−2μE[X]+μ2E[(X-\mu)^2] = E[X^2] - 2\mu E[X] + \mu^2. Since μ\mu is a constant (it does not depend on the outcome), it can be pulled out of the expectation in the middle and last terms.

Now substitute E[X]=μE[X] = \mu back in: E[X2]−2μ⋅μ+μ2=E[X2]−2μ2+μ2=E[X2]−μ2E[X^2] - 2\mu\cdot\mu + \mu^2 = E[X^2] - 2\mu^2 + \mu^2 = E[X^2] - \mu^2. Replacing μ\mu with E[X]E[X] gives exactly Var(X)=E[X2]−E[X]2\mathrm{Var}(X) = E[X^2] - E[X]^2, as claimed.

UndergraduateReal-World Applications and Worked Examples

Binomial variables model pass/fail counts everywhere: defective items on a production line, clicks on an ad, mutations in a strand of DNA. Normal variables model quantities built from many small independent effects: measurement error, human heights, portfolio returns, and background noise in electronics. Both let engineers, scientists and analysts turn a vague notion of "randomness" into a number they can compute with.

Example: Defective bulbs on a production line

A factory ships light bulbs in batches. Each bulb is independently defective with probability p=0.1p = 0.1. In a batch of n=8n = 8 bulbs, what is the probability that exactly k=2k = 2 are defective?

Solution

The number of defective bulbs XX is binomial with n=8n=8 and p=0.1p=0.1, since each bulb is an independent trial that is either defective or not.

Plug into the PMF: P(X=2)=(82)(0.1)2(0.9)6P(X=2) = \binom{8}{2}(0.1)^2(0.9)^6. The binomial coefficient is (82)=28\binom{8}{2} = 28, counting the ways to choose which 2 of the 8 bulbs are the defective ones.

Computing the powers, (0.1)2=0.01(0.1)^2 = 0.01 and (0.9)6≈0.531441(0.9)^6 \approx 0.531441, so P(X=2)=28×0.01×0.531441≈0.1488P(X=2) = 28 \times 0.01 \times 0.531441 \approx 0.1488. About a 15% chance that exactly two bulbs in the batch fail.

Example: Heights and the normal distribution

Adult male heights in a population are approximately normal with mean μ=170\mu = 170 (cm) and standard deviation σ=6\sigma = 6 (cm). Approximately what fraction of men have height between 164164 and 176176 cm?

Solution

Convert the two boundary heights to zz-scores using Z=X−μσZ = \frac{X-\mu}{\sigma}: for X=164X=164, z1=164−1706=−1z_1 = \frac{164-170}{6} = -1, and for X=176X=176, z2=176−1706=1z_2 = \frac{176-170}{6} = 1.

The interval 164<X<176164 < X < 176 is exactly the range within one standard deviation of the mean, i.e. −1<Z<1-1 < Z < 1 in standardized units.

For the standard normal curve, the empirical rule (from integrating the normal PDF) gives P(−1<Z<1)≈0.6827P(-1 < Z < 1) \approx 0.6827, so about 68% of men fall in that height range.

Let X∼Binomial(n=5,p=0.4)X \sim \mathrm{Binomial}(n=5, p=0.4). What is P(X=2)P(X=2)?

Which statement correctly distinguishes a discrete PMF from a continuous PDF?

A discrete random variable satisfies X∈{0,1,2}X \in \{0,1,2\} with P(X=0)=0.2P(X=0)=0.2, P(X=1)=0.5P(X=1)=0.5, P(X=2)=0.3P(X=2)=0.3. What is E[X]E[X]?

Standardized test scores are approximately normal with μ=500\mu = 500 and σ=100\sigma = 100. Using the empirical rule, about what fraction of test-takers score above 600600?

References

  1. Charles M. Grinstead, J. Laurie Snell (1997). Introduction to Probability
  2. Sheldon Ross (2019). A First Course in Probability