Mistake Master
Student view — seeing the site as a student does
Home Unit 2 · Probability, Random Variables, and Probability Distributions 2.1·2.2·2.3·2.4·2.5·2.6·2.7·2.8·2.9·2.10·2.11·2.12 Lesson
Skill Check 0 / 10 complete

Sketch it, shade it, then compute

For a continuous variable, probability is area, and the normal model supplies a curve whose areas are known once a value has been converted to standard deviations from the mean. The two failures are at either end of that sentence: applying the curve to a distribution that is not normal, and getting the area on the wrong side of the value because no sketch was drawn.

§1

For a continuous variable, probability is area under the curve.

A continuous random variable is described by a density curve: the curve never goes below the axis, the total area beneath it is exactly 1, and the probability that the variable falls in an interval is the area above that interval.

A consequence catches people out: for a continuous variable, $P(X = 5)$ is 0, because a single point spans no width. Only intervals carry probability, which is why $P(X < 5)$ and $P(X \le 5)$ are the same number here, unlike in the discrete case where the boundary value carried its own bar.

The normal distribution is one such curve: symmetric, single-peaked, and specified entirely by its mean $\mu$ and standard deviation $\sigma$. The mean locates the peak and the standard deviation sets the width, with the inflection points sitting one standard deviation on either side of the center.

§2

A z-score converts any normal variable to the standard one.

Standardizing subtracts the mean and divides by the standard deviation:

$$z = \frac{x - \mu}{\sigma}.$$

The result is the number of standard deviations $x$ sits from the mean, with the sign carrying the direction. Standardizing turns any normal distribution into the standard normal, mean 0 and standard deviation 1, whose areas are tabulated and built into every calculator.

Scores are approximately normal with $\mu = 500$ and $\sigma = 100$. A score of 650 gives $z = \frac{650 - 500}{100} = 1.5$, so about 6.7% of scores exceed it. A score of 400 gives $z = -1$, and about 15.9% of scores fall below it. The sign is not decoration: dropping it turns "one standard deviation below" into "one above" and moves the answer to the other tail.

Note what a z-score is not. It is not a probability: $z = 2$ does not mean 2%, or 0.02, or 98%. It is a position, and the probability is the area beyond that position.

§3

The empirical rule applies to normal distributions and to nothing else.

For a normal distribution, approximately:

  1. 68% of values lie within 1 standard deviation of the mean,
  2. 95% within 2 standard deviations,
  3. 99.7% within 3.

With $\mu = 500$ and $\sigma = 100$, about 68% of scores fall between 400 and 600, and about 95% between 300 and 700.

The rule is a property of the normal curve, not of data in general. Household income is strongly right-skewed, so applying it to a mean of 72,000 with a standard deviation of 55,000 predicts that 95% of households fall between $-38{,}000$ and 182,000, and a negative income is the tell: the interval extends past where the variable can even go. Before using the rule, establish that the distribution is approximately normal, from a graph of the data or from a stated model.

§4

Sketch first, and read the direction off the picture.

The mechanical habit that prevents most errors here is drawing the curve, marking the mean, marking the value, and shading the region the question asks for before computing anything. Then the arithmetic has something to check against.

Calculators and tables report a cumulative area, the probability below a value. A question about "more than" needs $1 - $ that area, and a question about "between" needs the difference of two of them. A shaded sketch makes it obvious which is required, while a bare $z = 1.5$ leaves the direction to memory.

Working backward runs the same steps in reverse. To find the score at the 90th percentile: locate the $z$ with 0.90 of the area below it, which is $z \approx 1.28$, then invert the standardizing formula, $x = \mu + z\sigma = 500 + 1.28(100) = 628$. The sketch is what makes clear that the answer must land above the mean, and it catches a sign error at once.

§5

Skill Check.

Ten scenarios. Pick the chips that match your answer, then check. A scenario marks complete the first time every part is right. Progress saves on this device.

0 of 10 scenarios complete