Probability and statistics
Random variables and distributions
Numerical outcomes of random experiments, described by the probability of each possible value.
IntuitionIntuition: Turning Chance into Numbers
Flip a coin, roll a die, measure how tall the next person walking in is — each of these random experiments has an outcome you can't predict in advance. A random variable is simply a rule that assigns a number to every possible outcome, so instead of talking about vague outcomes we can compute with numbers: add them, average them, plot them.
SchoolDiscrete vs. Continuous Random Variables
Definition: Discrete random variable
A random variable is discrete if it takes a finite or countable list of values, each with its own probability. Its probability mass function (PMF) is , and since must take some value, the probabilities of every possible value add up to 1.
The most famous discrete distribution counts successes in repeated independent trials. If a trial succeeds with probability and is repeated times, the number of successes follows a binomial distribution:
Here is the fixed number of trials, is the success probability of each trial, and counts how many ways successes can be arranged among the trials, since any of those arrangements gives exactly successes and failures.
Definition: Continuous random variable
A continuous random variable can take any value in an interval of real numbers, so for every single point . Instead of a PMF, it has a probability density function (PDF) , and probabilities come from the area under over an interval.
This is the normal (Gaussian) density: is the mean, marking the center of the bell, and is the standard deviation, controlling how wide or narrow the bell is. Many measurement errors and natural quantities are well approximated by this curve.
| Aspect | Discrete | Continuous |
|---|---|---|
| Describing function | Probability mass function | Probability density function |
| Probability of one exact value | Can be positive | Always |
| How probability is computed | Sum: | Integral: |
| Typical example | Binomial (number of successes) | Normal (measurement values) |
UndergraduateExpectation and Variance
The expectation is the long-run average value of if the experiment were repeated many times, weighting each possible value by how likely it is:
Each possible value is multiplied by its probability and the results are added up; expectation is linear, so for constants . Variance measures how spread out is around , and is easiest to compute with the shortcut formula:
This says variance equals the average of minus the square of the average of ; the standard deviation then puts the spread back in the same units as itself.
If follows a binomial distribution with trials and success probability , then .
Why is it true?
A binomial count is just a sum of many small yes/no trials, and expectation of a sum is always the sum of the expectations — no matter how the trials interact — so we only need the average contribution of a single trial.
Proof
Write , where is if the -th trial succeeds and otherwise. Each is a Bernoulli random variable with and , so its expectation is .
Expectation is linear for any random variables, whether or not they are independent: . This linearity follows directly from the definition of expectation as a weighted sum, since sums and weighted sums can always be reordered.
Applying linearity to gives , which is exactly the claimed formula.
For any random variable with finite mean, .
Why is it true?
The definition of variance involves squaring a difference, which is awkward to compute directly; expanding that square and using linearity of expectation turns it into two easy quantities, and , that we already know how to compute.
Proof
By definition, where . Expand the square inside the expectation: .
Apply linearity of expectation term by term: . Since is a constant (it does not depend on the outcome), it can be pulled out of the expectation in the middle and last terms.
Now substitute back in: . Replacing with gives exactly , as claimed.
UndergraduateReal-World Applications and Worked Examples
Binomial variables model pass/fail counts everywhere: defective items on a production line, clicks on an ad, mutations in a strand of DNA. Normal variables model quantities built from many small independent effects: measurement error, human heights, portfolio returns, and background noise in electronics. Both let engineers, scientists and analysts turn a vague notion of "randomness" into a number they can compute with.
Example: Defective bulbs on a production line
A factory ships light bulbs in batches. Each bulb is independently defective with probability . In a batch of bulbs, what is the probability that exactly are defective?
Solution
The number of defective bulbs is binomial with and , since each bulb is an independent trial that is either defective or not.
Plug into the PMF: . The binomial coefficient is , counting the ways to choose which 2 of the 8 bulbs are the defective ones.
Computing the powers, and , so . About a 15% chance that exactly two bulbs in the batch fail.
Example: Heights and the normal distribution
Adult male heights in a population are approximately normal with mean (cm) and standard deviation (cm). Approximately what fraction of men have height between and cm?
Solution
Convert the two boundary heights to -scores using : for , , and for , .
The interval is exactly the range within one standard deviation of the mean, i.e. in standardized units.
For the standard normal curve, the empirical rule (from integrating the normal PDF) gives , so about 68% of men fall in that height range.
Let . What is ?
Which statement correctly distinguishes a discrete PMF from a continuous PDF?
A discrete random variable satisfies with , , . What is ?
Standardized test scores are approximately normal with and . Using the empirical rule, about what fraction of test-takers score above ?
References
- Charles M. Grinstead, J. Laurie Snell (1997). Introduction to Probability
- Sheldon Ross (2019). A First Course in Probability