Mistake Master
Two Verdicts Exist, and Neither of Them Is "the Null Is True"
You'll learnto run a one-proportion z-test end to end — state, plan, do, conclude — and then to say what the verdict does and does not mean: why failing to reject never becomes accepting, why the 0.05 line is a convention rather than a cliff, and why significant measures detectable, not large.
The arithmetic of a test is four numbers long; the sentence at the end is where the marks are. A company claims a 40% renewal rate, 88 of 250 sampled customers renewed, and the p-value comes out at 0.0607 against a threshold of 0.05 — so the test fails to reject. What may be said now? Not that the rate is 0.40: the same data sit just as comfortably with 0.38, and more comfortably still with 0.36, and this animation runs those tests to prove it. Not that one customer matters much: with 87 renewals instead of 88 the same test rejects, so the two verdicts sit one customer apart on nearly identical evidence. And when a test does reject, that is a statement about detectability, not size — a 40,000-person sample pushes an effect of under one percentage point to a p-value of 0.001, while thirty observations cannot certify a thirteen-point gap. The lane on screen holds all of it: one dot, one line, and everything the dot's position does and does not license.