Mistake Master
Student view — seeing the site as a student does
Home Unit 3 · Inference for Categorical Data: Proportions 3.1·3.2·3.3·3.4·3.5·3.6·3.7·3.8·3.9·3.10·3.11·3.12·3.13·3.14·3.15 Lesson
Skill Check 0 / 10 complete

Not different is not the same as the same

With the setup done, the arithmetic is one division and a tail area. What remains is the sentence, and it inherits both asymmetries from the one-sample case with a new way to go wrong: a large p-value here does not mean the two groups perform equally, and a small one says which direction the gap runs only if the order of subtraction is stated.

§1

The statistic is the gap divided by the pooled standard error.

Under $H_0: p_1 = p_2$, the difference $\hat{p}_1 - \hat{p}_2$ has mean 0, so

$$z = \frac{\hat{p}_1 - \hat{p}_2}{\sqrt{\hat{p}_c(1-\hat{p}_c)\left(\frac{1}{n_1} + \frac{1}{n_2}\right)}}.$$

The numerator has no $(p_1 - p_2)$ subtracted from it because that quantity is 0 under the null.

For the checkout study: $\hat{p}_1 - \hat{p}_2 = 0.46 - 0.36 = 0.10$, $\hat{p}_c \approx 0.4145$, and $SE \approx 0.0422$, so

$$z = \frac{0.10}{0.0422} \approx 2.37.$$

Against $H_a: p_1 > p_2$ the p-value is the right-tail area, about $0.0089$; against $H_a: p_1 \ne p_2$ it is about $0.0178$. At $\alpha = 0.05$ both reject.

§2

The conclusion carries the direction and the population.

Written out: "Since the p-value of 0.0089 is less than $\alpha = 0.05$, we reject $H_0$. There is convincing evidence that the completion rate among all shoppers using the new system is higher than among all shoppers using the old system."

Four things have to be in it: the comparison to $\alpha$, the decision, the direction, and the two populations. Dropping the direction leaves "there is a difference", which is true and less than the test established. Dropping the populations leaves a claim about the 550 shoppers actually observed, whose rates are known exactly.

Whether the verb may be causal depends on the design. Randomly assigned shoppers license "the new system produces a higher completion rate"; two observed groups license only "the completion rate is higher among shoppers using the new system", with the difference possibly explained by who chose which system.

§3

Failing to reject is not a finding of equality.

Suppose instead the p-value came out at 0.28. The correct report: there is not convincing evidence of a difference in completion rates between the two systems.

Three sentences that are unavailable, all versions of accepting the null:

  1. "The two systems perform equally well."
  2. "There is no difference between the two proportions."
  3. "The data show the new system is no better."

A large p-value says the observed gap is the kind two identical populations would produce fairly often. It does not rule out a real gap, and with modest sample sizes it rules out very little: a difference of 5 or 8 percentage points can easily go undetected. This is the power question from Topic 3.8, and it is why a non-significant two-sample result should be reported alongside its sample sizes.

§4

The test and the interval usually agree.

A two-sided test at level $\alpha$ and a confidence interval at level $1 - \alpha$ answer the same question two ways. For the checkout data, the two-sided p-value of about 0.0178 is below 0.05, and the 95% interval $(0.018, 0.182)$ excludes 0. Both say the same thing: 0 is not a plausible difference.

The correspondence is close but not exact for proportions, because the two procedures use different standard errors: the test pools and the interval does not. In borderline cases, a p-value just under 0.05 can accompany an interval whose endpoint sits just on the other side of 0. That is not an error in either one; it reflects the different assumptions each is built on.

The two also answer different questions, which is why both are worth reporting. The test says whether a difference is detectable. The interval says how large the plausible differences are, and 2 to 18 percentage points is information no p-value carries.

§5

Skill Check.

Ten scenarios. Pick the chips that match your answer, then check. A scenario marks complete the first time every part is right. Progress saves on this device.

0 of 10 scenarios complete