Mistake Master
Home Unit 1 · Exploring One-Variable Data and Collecting Data 1.1·1.2·1.3·1.4·1.5·1.6·1.7·1.8·1.9·1.10·1.11·1.12·1.13 Lesson
Skill Check 0 / 10 complete

Picturing a quantitative variable

Measured numbers get displays built on a number line: dotplots, stemplots, histograms. The histogram is the workhorse and the most misread graph in the course, because its bars look exactly like a bar chart's while obeying different rules. A bar here is a count over an interval, not a value, and once that lands, the rest of the topic is vocabulary.

§1

Dotplots, stemplots, and histograms all keep the number line.

A quantitative variable's values are measured amounts, so every display for it preserves the natural order, smallest to largest, along a number line. Three standard displays trade detail for scale:

  1. Dotplot. One dot per observation, placed over its value, with equal values stacked. Every individual data value is recoverable from the picture. Best for small data sets.
  2. Stemplot. Each value splits into a stem (leading digit or digits) and a leaf (usually the single next digit): with the key $4\,|\,1 = 41$, the row $4\,|\,1\ 3\ 8$ holds 41, 43, 48. Stems and leaves are ordered, so values are recoverable here too.
  3. Histogram. The axis is cut into equal-width intervals called bins, and each bar's height is the frequency (or relative frequency) of observations landing in its bin. Individual values are gone; what remains is the distribution's silhouette. Scales to any $n$.

All three answer the same first question: where do the values pile up, and how far do they straggle?

§2

A histogram bar is a count over an interval, not a data value.

Here is the misread that owns this topic. A bar standing over the interval from 60 to 80 with height 8 says one thing: eight observations landed somewhere in $[60, 80)$. It does not say the value 80 occurred, or that 8 is a data value, or that any observation equals 70 just because 70 sits under the bar.

Consequences worth rehearsing:

  1. The tallest bar marks the most heavily populated interval, not the largest value in the data. The largest value lives in the rightmost nonempty bin, wherever the tall bars sit.
  2. The number of observations is the sum of the heights, not the number of bars.
  3. Whether a specific value like 85 is in the data is unanswerable from a bar spanning 80 to 90: the bar reports its bin's count and nothing finer.

An empty bin, by contrast, is real information: a stretch of the number line where no observations fell. On a bar chart, spacing between bars is decoration; on a histogram, a gap is data.

§3

Bin width is a choice, and the choice changes the picture.

Nothing in the data dictates the bins. The same 40 measurements drawn with 2-unit bins can look jagged and twin-peaked, and with 8-unit bins look like one smooth mound. Neither picture is a lie; each is a summary at a different resolution.

Working habits that keep bin width from fooling you:

  1. Before reading shape, check the bin width and the axis scale. "Tall" and "wide" mean nothing until the scale is known.
  2. Trust features that survive reasonable changes of bin width. A second peak that appears at width 2 and vanishes at width 4 is weak evidence; a gap that persists at several widths is a feature of the data.
  3. Very narrow bins chase noise; very wide bins bury structure. A handful of bins to a couple dozen usually brackets the useful range.

A histogram can also be drawn with relative frequencies on the vertical axis. The silhouette is identical; only the labels change from counts to proportions.

§4

Skew is named by the tail, not the pile.

The shape vocabulary that Topic 1.6 builds on starts here, with one convention that students reliably reverse. A distribution is skewed right when its longer, thinner tail stretches toward larger values, and skewed left when the tail stretches toward smaller values. The name follows the tail. It does not follow the pile.

Incomes are the classic case: most households sit at the low end and a thin tail reaches toward the very rich. The pile is on the left, and the distribution is skewed right, because right is where the tail points. Say it from the tail every time and the reversal never happens.

When the left half roughly mirrors the right half, the shape is approximately symmetric. And when a distribution shows two prominent peaks, the word is bimodal: skew vocabulary does not apply, and forcing it loses the most interesting fact in the picture, that the data may contain two different kinds of thing.

§5

Skill Check.

Ten scenarios. Pick the chips that match your answer, then check. A scenario marks complete the first time every part is right. Progress saves on this device.

0 of 10 scenarios complete