Mistake Master
Student view — seeing the site as a student does
Home Unit 2 · Probability, Random Variables, and Probability Distributions 2.1·2.2·2.3·2.4·2.5·2.6·2.7·2.8·2.9·2.10·2.11·2.12 Lesson
Skill Check 0 / 10 complete

A difference and a ratio are different claims

Once the counts are in, the summary is a choice: which relative frequencies to report, and how to compare them. Divide by the grand total and the numbers describe the whole sample; divide within a group and they describe that group. Then a difference of two conditional proportions and a ratio of the same two proportions can point at the same association and sound nothing alike.

§1

Three relative frequency tables come out of one table of counts.

A table of counts becomes a table of proportions in three different ways, and the choice of divisor is the whole content of the choice:

  1. Joint relative frequencies: every cell over the grand total. All four interior values sum to 1, because they partition the whole sample.
  2. Row conditional relative frequencies: every cell over its own row total. Each row sums to 1, and rows are compared with each other.
  3. Column conditional relative frequencies: every cell over its own column total. Each column sums to 1.

Take 500 employees classified by whether they completed an optional training and whether they later passed a certification. Trained: 150 passed, 50 did not, 200 in all. Untrained: 150 passed, 150 did not, 300 in all. Column totals: 300 passed, 200 did not.

The joint proportion $\frac{150}{500} = 0.30$ says 30% of all employees are trained passers. The row conditional $\frac{150}{200} = 0.75$ says 75% of trained employees passed. The column conditional $\frac{150}{300} = 0.50$ says half of the passers were trained. Three numbers from one cell, and the sum-to-1 check identifies which family a number belongs to: if the numbers reported for a group add to 1, they are that group's conditional distribution.

§2

Compare groups by subtracting conditional proportions.

The standard summary of association between two categorical variables is the difference in conditional proportions. Among the trained, the pass rate is 0.75; among the untrained, 0.50; the difference is $0.75 - 0.50 = 0.25$, or 25 percentage points.

Two vocabulary rules follow the arithmetic. A change from 50% to 75% is a rise of 25 percentage points, not "25 percent": as a percent change it is a 50% increase, since 25 is half of 50. And a difference of 0 means the conditional distributions match, which is what no association looks like in the sample.

Any conditional proportion is only as meaningful as the group it came out of. A 100% pass rate among 3 trained employees and a 75% rate among 200 are not comparable claims, and the group sizes belong in the report alongside the rates.

§3

A ratio and a difference can tell the same story at very different volumes.

The other common summary is the ratio of two conditional proportions, often called relative risk in a medical context: $\frac{0.75}{0.50} = 1.5$, so trained employees passed at 1.5 times the rate of untrained ones.

The two summaries can diverge dramatically when the proportions are small. A side effect that strikes 2 in 1,000 on placebo and 6 in 1,000 on a drug has a ratio of 3, which sounds alarming, and a difference of 4 per 1,000, which is 0.4 percentage points. Neither number is wrong. The ratio answers "how many times as likely", the difference answers "how many more people out of a hundred", and a report that gives only the ratio has chosen the louder one.

Read either summary back in context with both groups named: "the pass rate among trained employees was 25 percentage points higher than among untrained employees (75% versus 50%)". A percentage with no stated denominator group is not yet a claim.

§4

What a table of summaries cannot do.

Three limits ride along with every one of these summaries.

It cannot establish cause. Employees chose whether to take the training, so the trained group may differ in motivation, prior experience, and much else. An association in a table is a fact about these individuals; whether the training caused the difference depends on whether treatments were randomly assigned.

It cannot generalize on its own. A difference of 25 percentage points describes this sample. Whether it extends to a population depends on how the sample was selected, and whether a difference this large could arise from sampling variability alone is a question for inference, not for the table.

It cannot survive the wrong denominator. Every claim above is a statement about a specific group, and swapping the divisor silently changes the subject. Before comparing two percentages, confirm they were computed over the two groups being compared, and not one over a group and the other over everyone.

§5

Skill Check.

Ten scenarios. Pick the chips that match your answer, then check. A scenario marks complete the first time every part is right. Progress saves on this device.

0 of 10 scenarios complete