Mistake Master
Student view — seeing the site as a student does
Home Unit 4 · Inference for Quantitative Data: Means 4.1·4.2·4.3·4.4·4.5·4.6·4.7·4.8·4.9·4.10 Lesson
Skill Check 0 / 10 complete

Not detected is not the same as not there

For a difference of means, one value decides the verdict: zero, the claim that the two population means are equal. An interval sitting entirely on one side of it establishes a direction. An interval straddling it establishes nothing about direction, and in particular does not establish that the two means are the same. The sign of the endpoints is unreadable until the order of subtraction is stated.

§1

Zero is the no-difference value.

An interval for $\mu_1 - \mu_2$ is judged against 0, because $\mu_1 - \mu_2 = 0$ says the two population means are equal:

  1. Entirely above 0: evidence that $\mu_1 > \mu_2$.
  2. Entirely below 0: evidence that $\mu_1 < \mu_2$.
  3. Containing 0: no convincing evidence of a difference, since positive differences, negative differences, and no difference are all plausible.

The teaching-methods interval $(2.40, 12.00)$ for A minus B lies entirely above 0, so there is convincing evidence that method A's mean score is higher, by somewhere between about 2.4 and 12.0 points.

§2

Containing zero is a failure to detect.

An interval of $(-1.2, 4.8)$ contains 0. What it establishes is that the data are consistent with no difference, and equally consistent with $\mu_1$ exceeding $\mu_2$ by up to about 4.8 and with $\mu_2$ exceeding $\mu_1$ by up to about 1.2.

Two conclusions are unavailable:

  1. "The two methods are equally effective" or "the means are equal". Zero is one plausible value among many, and a real 4-point gap sits inside this interval.
  2. "Method 1 is better, since most of the interval is positive." Values below 0 are inside, so a negative difference remains plausible.

The correct sentence names the failure: these data do not provide convincing evidence of a difference between the two population means. And with small samples that sentence carries little weight, because the interval would have had to be far narrower to rule anything out. Reporting the sample sizes alongside is what lets a reader judge that.

§3

Width tells you precision, not importance.

The margin of error $t^{*}\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}$ moves with three things: the confidence level, the two groups' variability, and the two sample sizes. It moves with nothing else, and in particular it does not account for nonsampling problems: two intact classes that differ in preparation, an instrument that reads high in one group, students who dropped out of one method.

Two readings to keep apart:

  1. A narrow interval excluding 0 establishes a difference precisely. Whether that difference matters is a separate judgment: $(0.2, 0.6)$ points on a 100-point exam is a real difference and a trivial one.
  2. A wide interval excluding 0 establishes a difference imprecisely: $(1, 40)$ points says A is ahead and refuses to say by how much.

So the interval answers two questions at once, and both belong in the report: is there a difference, and how large are the plausible values.

§4

The justification, in four parts.

  1. State: "We are 95% confident the interval from 2.40 to 12.00 points captures the true difference in mean score, method A minus method B, between the two populations of students."
  2. Locate 0: "The interval lies entirely above 0."
  3. Direction: "So there is convincing evidence that method A's mean score is higher."
  4. Magnitude: "by roughly 2.4 to 12.0 points."

Keep the hedge, and keep causation out unless the design supports it. Students randomly assigned to the two methods license "method A produces higher scores"; two intact classes license only "scores are higher among students taught by method A", with class composition and teacher as live alternative explanations.

§5

Skill Check.

Ten scenarios. Pick the chips that match your answer, then check. A scenario marks complete the first time every part is right. Progress saves on this device.

0 of 10 scenarios complete