Mistake Master
Student view — seeing the site as a student does
Home Unit 3 · Inference for Categorical Data: Proportions 3.1·3.2·3.3·3.4·3.5·3.6·3.7·3.8·3.9·3.10·3.11·3.12·3.13·3.14·3.15 Lesson
Skill Check 0 / 10 complete

Two ways to be wrong, and they trade

A test decides from a sample, so it can be wrong in two different ways: raising an alarm when nothing is happening, or missing something that is. The two mistakes have different probabilities, different costs, and opposite responses to the significance level, so lowering one raises the other. Only more data buys both at once.

§1

Two truths, two decisions, two ways to be wrong.

Cross what is actually true with what the test decides and four cells appear, two correct and two errors:

  1. $H_0$ true, fail to reject: correct.
  2. $H_0$ true, reject: Type I error, a false alarm.
  3. $H_0$ false, fail to reject: Type II error, a missed detection.
  4. $H_0$ false, reject: correct, and this is the test's power.

Both errors are defined relative to the null, so naming them requires the null first. A drug company tests $H_0$: the new drug is no better than the standard, against $H_a$: it is better. A Type I error approves a drug that does not work; a Type II error shelves a drug that does. Describing an error as "being wrong" without saying which way, or attaching the consequence to the wrong cell, is the single most common failure here.

Only one error is possible in any given situation, and which one depends on a truth nobody knows. If the null is true, a Type II error cannot occur; if it is false, a Type I error cannot.

§2

Alpha and beta are conditional probabilities, on different conditions.

Each error has a probability, and each is conditional on a different state of the world:

$$\alpha = P(\text{reject } H_0 \mid H_0 \text{ true}), \qquad \beta = P(\text{fail to reject } H_0 \mid H_0 \text{ false}).$$

The significance level is therefore not just a decision threshold: it is the Type I error rate. Testing at $\alpha = 0.05$ means that if the null were true, about 5% of samples would lead to rejecting it.

$\beta$ is harder to pin down, because "the null is false" is not one situation. A test has one $\beta$ for a true proportion of 0.55, a smaller one for 0.65, and a smaller one still for 0.75: the further the truth is from the null value, the easier it is to detect. That is why $\beta$ is always quoted against a specific alternative value.

The power of a test is $1 - \beta$, the probability of correctly rejecting a false null. Like $\beta$, it is quoted against a stated alternative: "this test has power 0.80 to detect $p = 0.65$" means that if the true proportion were 0.65, the test would reject the null about 80% of the time.

§3

Lowering one error rate raises the other.

Picture the null distribution and the true distribution overlapping, with the rejection region cut off at a critical value. The area of the null curve past the cutoff is $\alpha$; the area of the true curve on the near side is $\beta$.

Sliding the cutoff outward shrinks $\alpha$ and grows $\beta$. Sliding it inward does the reverse. There is no position that shrinks both, which is why "use a smaller $\alpha$ so the test makes fewer mistakes" is wrong: it makes fewer false alarms and more missed detections.

Choosing $\alpha$ is therefore a judgment about which error costs more:

  1. Approving an unsafe drug is far worse than delaying a good one, so a small $\alpha$, perhaps 0.01, is appropriate.
  2. Missing an early warning of contamination is worse than investigating a false alarm, so a larger $\alpha$, perhaps 0.10, is appropriate.

What escapes the trade entirely is sample size. A larger $n$ narrows both distributions, so the curves overlap less and both error rates can fall together. Increasing $n$ is the only way to reduce $\beta$ without raising $\alpha$.

§4

Four things raise power, and one of them is not free.

Power rises when:

  1. The sample size increases. Both sampling distributions narrow, so they separate.
  2. The true effect is larger. A proportion far from the null value is easier to detect than one just beside it. This is a property of the world, not a choice.
  3. The variability is smaller. Less spread means less overlap.
  4. $\alpha$ increases. A wider rejection region catches more true effects, at the cost of more false alarms.

The last one is why power alone is never the goal: setting $\alpha = 0.50$ would give enormous power and reject half of all true nulls. The usable levers are sample size and study design.

Power is also the missing piece behind the previous topic's null results. A test with power 0.30 to detect an effect worth caring about will miss it most of the time, so failing to reject tells you very little. Reporting a non-significant result alongside its sample size is what lets a reader judge whether the test could have seen anything at all.

§5

Skill Check.

Ten scenarios. Pick the chips that match your answer, then check. A scenario marks complete the first time every part is right. Progress saves on this device.

0 of 10 scenarios complete