Mistake Master
Student view — seeing the site as a student does

Carrying Out a Test for a Population Proportion

▶︎  Watch it animatedinteractive step-through · ~3 min · optional

A test runs in four steps: state the parameter, hypotheses, and $\alpha$; plan by naming the procedure and checking conditions with the study's numbers; compute the test statistic and p-value; then compare, decide, and interpret in context. For $H_0: p = 0.40$ against $H_a: p < 0.40$ with $n = 250$ and 88 renewals, $z \approx -1.55$ and the p-value is about 0.061, so at $\alpha = 0.05$ we fail to reject and report that there is not convincing evidence the renewal rate is below 40%.

Two asymmetries define the topic. A large p-value never becomes support for the null: the data are consistent with 0.40 and equally with 0.38 and 0.36, so accepting $H_0$, proving no difference, or stating that the rate is 40% all convert an absence of evidence into a finding. And statistical significance measures detectability, not size: $\hat{p} = 0.508$ with $n = 40{,}000$ gives a two-sided p-value near 0.0014 for an eight-tenths-of-a-point effect, while $\hat{p} \approx 0.633$ with $n = 30$ gives about 0.072 for a 13-point one. The threshold is a line, not a cliff, so 0.049 and 0.051 describe nearly the same evidence.

compare, decide, interpret p-value < alpha REJECT H0 convincing evidence for Ha p-value > alpha FAIL TO REJECT H0 not convincing evidence for Ha three sentences that do not exist we accept H0 the data prove there is no difference the renewal rate IS 0.40 0.38 and 0.36 fit the data too; failing to rule a value out is not evidence that it is right
Only two decisions exist, and the right-hand one is not a verdict for the null. Many values besides 0.40 would also have produced ordinary-looking data, which is exactly why the null cannot be accepted.
significance and size are different questions sample effect p-value verdict at 0.05 n = 40,000 p-hat = 0.508 0.8 pts 0.0014 significant n = 30 p-hat = 0.633 13.3 pts 0.072 not significant significant answers CAN WE DETECT IT, never HOW BIG IS IT follow a significant result with an interval for the size and a null result with a question about the sample size
The columns disagree by design. Sample size drives detectability, so a difference of under a percentage point clears any threshold with enough data while a 13-point gap misses it with 30 observations.

The work

3 ways in · any order
Lesson
Carrying Out a Test for a Population Proportion

Runs the one-proportion z-test end to end, then settles the two asymmetries: a large p-value leaves the null standing without supporting it, and statistical significance measures detectability rather than size, with a conclusion that names the population and keeps its hedge.

Skill check · 10 scenarios
Diagnostic
10-item topic check

Ten items on carrying out a test: nulls accepted or proven, conclusions about the sampled group, significance read as importance, and thresholds treated as cliffs. Take it cold to find your habit, or after the lesson to check it is gone.

Not started · 10 items · ~15 min
Targeted Practice
Drill a single misconception

Pick one of the failure modes you missed and drill it on its own. The round is adaptive: two correct in a row clears it for now and moves you to the next. Two in a row is a checkpoint, not proof: if the error resurfaces later, the misconception comes back.

Take the diagnostic to identify your misconceptions