Mistake Master
Student view — seeing the site as a student does
For teachers › Field notes › Paired and two-sample designs confused

Paired or two-sample: the design decides, and two columns of numbers look identical either way

Before and after measurements on the same subjects are one sample of differences. Students see two columns, count two groups, and reach for the two-sample procedure.

Field note AP Statistics · Unit 4 Published October 8, 2026

Paired data comes from one sample measured twice or from matched subjects. The analysis is a one-sample t procedure on the differences. Two independent samples get the two-sample procedure. The data table looks the same in both cases, so the choice has to come from reading how the data was collected.

01The mistake

Give students a table of 12 runners' times before and after a training program, in two columns. A large share will run a two-sample t test comparing the mean before-time to the mean after-time. The two columns describe the same 12 people, so the samples are not independent and the procedure's conditions fail at the first one.

The error runs the other way too, and that version is less often taught. Given two genuinely independent groups listed in two columns of equal length, some students subtract row by row and run a paired test. Equal column lengths are not pairing; pairing is a feature of how subjects were assigned or matched.

Both versions usually produce a plausible answer, which is what makes them expensive. A two-sample test on paired data runs to completion and returns a p-value. It is typically far too large, because the between-subject variation that pairing was designed to remove is still in the standard error, and a real effect gets buried in it.

The tell is a student who starts by computing two means. On paired data the first computation should be a column of differences, and a student who never writes that column has already chosen the wrong procedure regardless of what they do next.

02Why it makes sense to the student

Two columns of numbers look like two samples. The table is the first thing a student sees and it carries no information about how the data was collected, so the visual layout decides the procedure before the design description is read.

The phrase “two-sample” sounds like a description of the data rather than of the design. A student with two lists of numbers reasonably concludes they have two samples. The technical requirement is independence between the samples, which is a fact about the subjects and not about the columns.

Pairing is taught as a design topic in one unit and as an inference procedure in another, often weeks apart. The students who learn that matched pairs reduce variability are not necessarily the ones who connect that to which t procedure to run, because the connection lives in the gap between the two lessons.

And the paired procedure looks like a different thing than it is. It is a one-sample t test, but it is introduced in the chapter about comparing two groups, so students file it under two-group methods and then choose among them by how many groups they see. The name “paired t test” reinforces that filing.

03The correction

Make the first question about the subjects, not the numbers: did each value in one column come from the same individual, or a matched partner, as the value beside it? If yes, the design is paired and the differences are the data. Ask it before looking at the table, every time, until it is automatic.

Then make the paired analysis visibly a one-sample procedure. Have students write the column of differences and physically cross out the original two columns. What remains is one list of numbers, one mean, one standard deviation, and a one-sample t test of $\mu_d = 0$. Students who have crossed out the originals stop looking for a two-sample formula.

Show the cost with the same data analyzed both ways. On a realistic before-and-after data set the paired test finds a clear effect and the two-sample test finds nothing, because the two-sample standard error includes the variation between individuals that the pairing removed. The contrast in p-values is more persuasive than any statement about independence.

Give the hypotheses in their correct form and insist on the subscript. Paired: $H_0: \mu_d = 0$, about a mean difference. Two-sample: $H_0: \mu_1 - \mu_2 = 0$, about a difference of means. They are different parameters, and writing the right one is itself a scored step on the exam.

Build a sorting exercise rather than a procedure exercise. Hand out eight design descriptions with no data at all and have students label each paired or independent. Removing the numbers removes the cue they have been wrongly relying on, which is the whole point of the drill.

04A sample question

Diagnostic-style item

A coach records the 5-kilometer time of each of 15 runners before and after a six-week training program and wants to know whether mean time decreased. Which procedure is appropriate?

  • AA two-sample t test comparing the mean before-time to the mean after-time, since there are two sets of times.
  • BA one-sample t test on the 15 differences, since each pair of times comes from the same runner.
  • CA two-proportion z test, since the question asks whether times decreased.
  • DA one-sample t test on the 30 recorded times, since all of them come from the same group of runners.

05What each wrong answer reveals

  • A Two columns read as two independent samples. The dominant wrong answer, and the justification names the cue the student used: the shape of the table. Ask whether runner 7's before-time and after-time are independent measurements. They are not, and once a student says so the independence condition fails in front of them.
  • B Correct. Each runner supplies one difference, so the data is one sample of 15 differences and the procedure is a one-sample t test of $\mu_d = 0$ against a decrease.
  • C Procedure chosen from the verb. “Whether times decreased” reads as a yes-or-no question, which cues a proportion test. The response variable here is a time in minutes, which is quantitative, so no proportion exists to test. This is a variable-type error rather than a design error, and it is worth separating from A because the lesson is different: check what is being measured before checking how.
  • D Pairing collapsed instead of used. This student recognized that one group of runners is involved, which is the correct observation, and then pooled all 30 times into a single list. That throws away the pairing and tests nothing meaningful — the mean of all 30 times answers no question anyone asked. The repair is short: the data is the 15 differences, not the 30 measurements.

A and D are opposite failures on the same fact. A treats the pairing as absent and D notices it but discards it, so A needs the design question and D needs the differences column. C is unrelated and belongs to a variable-type lesson.

06Try it in Mistake Master

Where this lives in the platform

Topic 4.9 (Setting Up a Test for the Difference Between Two Population Means) is where the choice gets made, and items there include designs whose data tables are indistinguishable so that the decision has to come from the collection description. U4-ST7 re-enters in Topics 4.2 and 4.4, where the hypotheses are written and the wrong parameter is visible in the symbols. A student holding this code can execute both procedures correctly and still choose the wrong one on every item, which is why the diagnostic asks for the design before the arithmetic.