Mistake Master
Student view — seeing the site as a student does
For teachers › Field notes › p-value definition mangled

The p-value is not the probability the null is true: it assumes the null and asks about the data

A p-value is computed by taking the null hypothesis as given. Students reverse the conditional and report it as a probability about the hypothesis, which is the one thing it cannot be.

Field note AP Statistics · Unit 3 Published October 8, 2026

The p-value answers: if the null were true, how surprising would this sample be? Students answer a different question: given this sample, how likely is the null? The null is the condition, never the event, and the reversal is the single most common inference error on the exam.

01The mistake

A test returns $p = 0.03$. Ask what that number means. A large share of any class will say there is a 3% chance the null hypothesis is true, or a 97% chance the alternative is. Both are statements about the hypothesis, and the p-value is not one.

A second version is subtler and gets partial credit on a lot of homework: “the probability that the results happened by chance.” It sounds like a correct paraphrase and omits the conditional, which is the entire content of the definition. Without “assuming the null is true,” the sentence is about the data in general rather than about the data under a specific model.

A third version drops the word “or more extreme.” Students describe the p-value as the probability of getting exactly this result, which is a different and usually tiny number, and which makes the comparison to $\alpha$ meaningless.

On an exam this costs points even when the decision is right. A student can compare 0.03 to 0.05, reject correctly, and lose the interpretation point for a sentence that reverses the conditional. The decision and the interpretation are scored separately, so a reliable procedure does not protect them.

02Why it makes sense to the student

The reversed version is the question students actually care about. Nobody wants to know how surprising the data is; they want to know whether the claim is true. The p-value is the closest available number to the question they are asking, so they read it as an answer to that question.

Conditional probability is hard in both directions, and this is the same reversal students make everywhere else. $P(\text{data} \mid \text{hypothesis})$ and $P(\text{hypothesis} \mid \text{data})$ get swapped for the same reason $P(\text{positive test} \mid \text{disease})$ and $P(\text{disease} \mid \text{positive test})$ do. It is a general weakness showing up in a statistics context.

Classroom shorthand teaches the reversal. “Low p-value means reject the null” is a true and useful rule, and it is one short step from “low p-value means the null is probably false,” which is the misconception. The shorthand is doing damage precisely because it works.

And the structure of hypothesis testing is genuinely strange. Assuming something in order to argue against it is a proof by contradiction, which students meet in geometry and rarely again. The null is a temporary assumption rather than a belief, and nothing else in a statistics course asks them to hold a premise they expect to discard.

03The correction

Make the conditional the first words of every interpretation, out loud and in writing: “If the null hypothesis were true, the probability of getting a sample result at least this extreme is 0.03.” Require the “if” clause and require “at least this extreme.” The sentence is long and the length is the point; a short version is short because it dropped a condition.

Name the two probabilities side by side and ask which one the computation produced. Students who see $P(\text{data} \mid H_0)$ and $P(H_0 \mid \text{data})$ written as symbols usually recognize that the test statistic was built from the null's model, so the null had to be assumed before any number could be computed. That is the argument that sticks: the null cannot be the event when it was the input.

Then say what would be needed for the reversed question, since students deserve to know it is not an unanswerable one. Getting $P(H_0 \mid \text{data})$ requires a prior probability for the hypothesis, which this course does not supply and the data cannot provide. The reversal is not merely wrong notation, it asks for information that is not in the problem.

Drill the simulation picture. Build the null distribution by simulation, mark the observed statistic, and count the proportion of simulated samples at least that far out. Students who have produced a p-value by counting dots in a tail have a hard time believing it is a probability about a hypothesis, because they watched it be a proportion of samples.

A clean diagnostic: give a p-value and four interpretations, one correct, one reversing the conditional, one dropping “or more extreme,” and one describing it as the probability of the data by chance alone. Which wrong one a student picks tells you which part of the sentence to work on.

04A sample question

Diagnostic-style item

A researcher tests $H_0: p = 0.5$ against $H_a: p > 0.5$ and obtains $p\text{-value} = 0.04$. Which statement correctly interprets this p-value?

  • AThere is a 0.04 probability that the null hypothesis is true.
  • BAssuming $p = 0.5$, there is a 0.04 probability of obtaining a sample proportion at least as large as the one observed.
  • CThere is a 0.04 probability that the observed result occurred by chance alone.
  • DAssuming $p = 0.5$, there is a 0.04 probability of obtaining exactly the sample proportion that was observed.

05What each wrong answer reveals

  • A The conditional reversed. The dominant wrong answer and the one that costs the most points, because it is a statement about the hypothesis rather than about the data. Ask what had to be assumed before 0.04 could be computed. The answer is $p = 0.5$, which makes the null the condition and not the event.
  • B Correct. The null is assumed, and the probability is about sample results at least as extreme as the one observed. Both the conditional and the tail are present.
  • C The conditional dropped. This is the version that reads most like a correct paraphrase, which is why it is worth isolating. “By chance alone” gestures at the null model without specifying it, so the sentence does not say what chance process is being assumed. Ask the student which value of $p$ generated the 0.04; they cannot answer from their own sentence.
  • D The tail dropped. This student has the conditional right, which is most of the work, and lost “or more extreme.” The probability of exactly one sample proportion is a different and typically far smaller number, and with a continuous model it is zero. Showing that the two numbers are not close is enough; this is a repair rather than a rebuild.

A, C and D are three different sentences failing in three different places, and lumping them together as “p-value confusion” hides which word a student is missing. A has the direction wrong, C has the condition missing, and D has both right and the region wrong. Only A needs the conditional lesson.

06Try it in Mistake Master

Where this lives in the platform

Topic 3.6 (p-Values) is where the definition is established, and items there deliberately offer the reversed conditional alongside the correct sentence so that the two cannot be told apart by keyword matching. U3-ST8 pairs with U3-ST9 (accepting the null) and re-enters across Topics 3.7 and 3.13, where every conclusion has to be written in context and linked to the p-value. A student holding this code writes the conclusion as a claim about the hypothesis's probability, which is unscoreable no matter how the arithmetic went.