Mistake Master
Student view — seeing the site as a student does
For teachers › Field notes › Correlation read as causation

Correlation and causation: a strong r says how tightly the points follow a line, and nothing about why

Every student can recite that correlation does not imply causation. Asked to write a conclusion from an observational study with a correlation of 0.92, a large share of them will write a causal one anyway.

Field note AP Statistics · Unit 5 Published October 8, 2026

The slogan is memorized and the habit is not. A correlation coefficient measures linear association in the data you have; whether one variable causes the other depends on how the data was collected, and only a randomized experiment settles it. Students recite the rule and then violate it in the next sentence.

01The mistake

An observational study of 200 towns finds $r = 0.92$ between the number of fast-food outlets and the rate of heart disease. Ask for a conclusion. A large share of students write that fast food causes heart disease, or that reducing outlets would reduce disease. The same students, asked the vocabulary question directly ten minutes earlier, said correctly that correlation does not imply causation.

The gap between the recited rule and the written conclusion is the whole problem. This is not a misconception students will defend; it is a habit that reasserts itself the moment they are asked to say something useful about data. A conclusion that stops at “these variables are associated” feels to them like a refusal to answer.

Causal verbs smuggle it in. “Leads to,” “results in,” “improves,” “drives,” and “affects” are all causal, and a sentence built on any of them makes a causal claim regardless of what the student intended. “Associated with” and “tends to be higher when” are the available alternatives.

The strength of the correlation drives the error, which is the diagnostic fact worth knowing. Students who write a careful associational conclusion at $r = 0.4$ will write a causal one at $r = 0.92$, because a tight scatterplot feels like proof. Strength is about how closely the points follow the line and carries no information about mechanism at all.

02Why it makes sense to the student

Finding causes is what students think the subject is for. A conclusion that two things move together and stops there sounds like an unfinished answer, so they finish it. The impulse is not statistical ignorance; it is an ordinary expectation that analysis should produce an explanation.

The slogan is taught as a phrase to be recalled rather than as a constraint on writing. A student who can produce “correlation does not imply causation” on demand has met the apparent requirement, and nothing in that exchange trains them to check their own sentences for causal verbs.

Real headlines model the error constantly. Reporting on observational studies is causal almost by default, so students arrive having read hundreds of examples of exactly the inference they are being asked not to make, written by professionals in authoritative outlets.

And the confounding variable is invisible by construction. In the fast-food example, median income plausibly drives both variables, but it is not in the data set and nothing on the scatterplot hints at it. Students are being asked to reason about a variable they cannot see, which is a genuinely harder task than reading the one they can.

03The correction

Make the design the first thing read and the last thing checked. Observational study: association only. Randomized experiment: a causal conclusion is available. Asking “was treatment assigned at random?” before writing anything turns the rule into a step rather than a slogan.

Require a named confounding variable, not the phrase “there could be a confounder.” In the fast-food case, have students propose median income and explain how it would produce the observed correlation with no causal link between the two variables plotted. A student who can build the alternative explanation stops needing to be told the causal one is unjustified.

Ban the causal verbs in written conclusions and supply the replacements. This is a concrete, checkable writing rule, and it catches the error where it actually occurs. Students can self-edit for a list of five words far more reliably than they can self-edit for an inferential stance.

Separate strength from structure explicitly, because this is the piece students have not been told. The value of $r$ answers how tightly the points cluster around a line; the design answers whether the relationship is causal. Those are independent questions, and $r = 0.99$ from an observational study supports a causal claim exactly as much as $r = 0.3$ does, which is to say not at all.

Then give the one real exception so the rule does not feel absolute for the wrong reason: a randomized experiment with a strong association does support a causal conclusion, and that is the whole reason random assignment exists. Students who know when causation is available apply the restriction more carefully than students who think it is never allowed.

04A sample question

Diagnostic-style item

An observational study of 200 towns finds a correlation of $r = 0.92$ between the number of fast-food restaurants per capita and the rate of heart disease. Which conclusion is justified?

  • AFast-food restaurants cause heart disease, since the correlation is very strong.
  • BTowns with more fast-food restaurants per capita tend to have higher rates of heart disease, but this study cannot establish a causal link.
  • CNo conclusion is possible, since correlation never provides any information about the variables.
  • DHeart disease causes towns to have more fast-food restaurants, since the direction of causation cannot be determined.

05What each wrong answer reveals

  • A Strength read as evidence of mechanism. The dominant wrong answer, and the justification names the reasoning: $r = 0.92$ felt like proof. Ask for one variable that would plausibly raise both numbers at once. Median income does it, and a student who can state that alternative has refuted their own conclusion more convincingly than any rule would.
  • B Correct. The association is real and strong, and the observational design leaves causation unestablished. The verb is “tend to have,” which claims exactly what the data supports.
  • C The rule overapplied. This student has learned the restriction and turned it into a prohibition on saying anything. The correlation is real information: these variables are strongly associated, which is worth reporting and is often the finding that motivates an experiment. Worth correcting as firmly as A, because a student who thinks observational data is worthless has lost the use of most of the data in the world.
  • D Causation asserted in the reverse direction. The justification contradicts the claim, which is the interesting part: the student has correctly noticed that direction is undetermined and then committed to a direction anyway. They are closer than A on the reasoning and identical to A on the conclusion. Point at their own second clause.

A and D both assert causation and arrive there differently, so D needs only to follow its own justification to the end. C is the opposite failure and needs the opposite instruction: the association is a finding, and the restriction is on the mechanism, not on the data.

06Try it in Mistake Master

Where this lives in the platform

Topic 5.2 (Correlation) is where $r$ is introduced and where the strength-versus-structure separation has to hold, and items there pair high correlations with observational designs so that a causal conclusion is available to choose and wrong. U5-ST2 re-enters in Topic 5.5, where slope interpretation invites the same causal reading of a coefficient, and it pairs with U1-ST13 (random selection versus random assignment), which is the design-side half of the same idea. A student holding this code writes causal conclusions from every regression in the course.