Mistake Master
Student view — seeing the site as a student does
Home Unit 2 · Probability, Random Variables, and Probability Distributions 2.1·2.2·2.3·2.4·2.5·2.6·2.7·2.8·2.9·2.10·2.11·2.12 Lesson
Skill Check 0 / 10 complete

Every proportion has a denominator

A two-way table is a grid of counts, and almost every mistake made with one is a mistake about which total sat underneath the division. The count in a cell can be turned into three different proportions, each answering a different question, and the graphs built from the table inherit the same choice: the display that shows association is the one that puts every group on the same vertical scale.

§1

A two-way table cross-classifies every individual once.

A two-way table (a contingency table) records two categorical variables measured on the same individuals. Rows carry the categories of one variable, columns the categories of the other, and each interior cell holds the number of individuals in that row and that column. The margins hold the row totals and column totals, and the grand total in the corner counts everyone.

Two hundred students are classified by grade and by whether they hold a part-time job:

  1. Juniors: 42 with a job, 98 without, 140 juniors in all.
  2. Seniors: 30 with a job, 30 without, 60 seniors in all.
  3. Column totals: 72 students with a job, 128 without, 200 students.

Every student is counted exactly once, in exactly one cell. That is the structural property the whole topic rests on: the four interior cells add to 200, each row adds to its row total, each column to its column total. A table whose interior does not add to the grand total has lost or duplicated individuals, and nothing computed from it will mean anything.

§2

One cell, three proportions, three different denominators.

The count 30 in the senior-with-a-job cell supports three separate proportions, and which one is correct depends entirely on the question:

  1. Joint: $P(\text{senior and job}) = \frac{30}{200} = 0.15$. Of all students, 15% are seniors who hold a job. Denominator: the grand total.
  2. Conditional: $P(\text{job} \mid \text{senior}) = \frac{30}{60} = 0.50$. Of the seniors, half hold a job. Denominator: the row total for the group being conditioned on.
  3. Marginal: $P(\text{job}) = \frac{72}{200} = 0.36$. Of all students, 36% hold a job, grade ignored. Denominator: the grand total, numerator from the margin.

The wording carries the denominator. "What percent of seniors have a job" conditions on seniors, so the senior total goes underneath. "What percent of students are seniors with a job" asks about everyone, so 200 goes underneath. And "what percent of the working students are seniors" conditions on the other variable: $\frac{30}{72} \approx 0.417$, a fourth number entirely. Before dividing, find the population the sentence names and put its total on the bottom.

A conditional distribution is the full set of conditional proportions within one group: among seniors, 50% have a job and 50% do not, and those two add to 1. Every conditional distribution sums to 1 within its own group, which is the arithmetic check that the right denominator was used.

§3

The graph that answers the question keeps every group on one scale.

Three displays are built from the same table, and they are not interchangeable:

  1. Side-by-side bar graph: one bar per category pair, drawn from raw counts or from relative frequencies. Drawn from counts, it compares group sizes as much as anything else.
  2. Segmented (stacked) bar graph: one bar per group, each bar scaled to 100% and divided into the categories of the other variable. Every bar has the same height, so the segments are the conditional distributions, side by side.
  3. Mosaic plot: a segmented bar graph whose bar widths are proportional to group size, so it shows the conditional distributions and the marginal distribution at once.

When the groups differ in size, raw-count displays mislead by design. The 140 juniors will out-bar the 60 seniors in almost every category simply because there are more of them. Converting to relative frequencies within each group removes group size from the picture and leaves the comparison the question actually asked for. This is why the segmented bar graph is the workhorse display here: equal bar heights force the comparison onto proportions.

§4

Association is a difference between conditional distributions.

Two categorical variables show an association when the conditional distribution of one variable changes across the categories of the other. The check is mechanical: compute the conditional distribution in each group, then compare.

Among juniors, $\frac{42}{140} = 0.30$ hold a job. Among seniors, $\frac{30}{60} = 0.50$. Those conditional distributions differ, so grade and employment are associated in this sample: knowing a student's grade changes what to expect about the job.

The comparison that does not work is the one most often reached for. Juniors supply 42 of the working students and seniors only 30, so the raw counts say juniors work more, while the rates say the opposite. The counts were never a fair comparison: they measure how many juniors there are as much as how often juniors work. If the conditional distributions match across every group, there is no association in the sample, however lopsided the raw cells look.

One more limit: an association found in a table is a statement about these individuals, and by itself says nothing about cause. That question is settled by how the data were collected, not by the size of the difference.

§5

Skill Check.

Ten scenarios. Pick the chips that match your answer, then check. A scenario marks complete the first time every part is right. Progress saves on this device.

0 of 10 scenarios complete