Home / AP® Exam / AP® Statistics / AP Statistics 4.7 Constructing a Confidence Interval for the Difference Between Two Population Means- Exam Style Questions – FRQs

AP Statistics 4.7 Constructing a Confidence Interval for the Difference Between Two Population Means- Exam Style Questions - FRQs - New Syllabus

Question

Systolic blood pressure is the amount of pressure that blood exerts on blood vessels while the heart is beating. The mean systolic blood pressure for people in the United States is reported to be 122 millimeters of mercury (mmHg) with a standard deviation of 15 mmHg.
The wellness department of a large corporation is investigating whether the mean systolic blood pressure of its employees is greater than the reported national mean. A random sample of 100 employees will be selected, the systolic blood pressure of each employee in the sample will be measured, and the sample mean will be calculated.
Let \(\mu\) represent the mean systolic blood pressure of all employees at the corporation. Consider the following hypotheses.
\(H_0 : \mu = 122\)
\(H_a : \mu > 122\)
(a) Describe a Type II error in the context of the hypothesis test.
(b) Assume that \(\sigma\), the standard deviation of the systolic blood pressure of all employees at the corporation, is \(15\) mmHg. If \(\mu = 122\), the sampling distribution of \(\bar{x}\) for samples of size 100 is approximately normal with a mean of 122 mmHg and a standard deviation of 1.5 mmHg. What values of the sample mean \(\bar{x}\) would represent sufficient evidence to reject the null hypothesis at the significance level of \(\alpha = 0.05\)?
The actual mean systolic blood pressure of all employees at the corporation is 125 mmHg, not the hypothesized value of 122 mmHg, and the standard deviation is 15 mmHg.
(c) Using the actual mean of 125 mmHg and the results from part (b), determine the probability that the null hypothesis will be rejected.
(d) What statistical term is used for the probability found in part (c)?
(e) Suppose the size of the sample of employees to be selected is greater than 100. Would the probability of rejecting the null hypothesis be greater than, less than, or equal to the probability calculated in part (c)? Explain your reasoning.

Most-appropriate topic codes (AP Statistics):

• Topic \(3.11\) — Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Proportions (Part \( \mathrm{a} \))
• Topic \(4.5\) — Carrying Out a Test for a Population Mean (Part \( \mathrm{b} \))
• Topic \(4.7\) — Constructing a Confidence Interval for the Difference Between Two Population Means (Parts \( \mathrm{c} \), \( \mathrm{d} \), \( \mathrm{e} \))
▶️ Answer/Explanation

(a)
A Type II error occurs when the alternative hypothesis is actually true, but we fail to reject the null hypothesis.
In this context, a Type II error would occur if the true mean systolic blood pressure of all employees is greater than 122 mmHg, but the hypothesis test does not detect this — that is, the null hypothesis (\(H_0 : \mu = 122\)) is not rejected.
In plain terms: the employees really do have higher-than-national-average blood pressure, but the test fails to conclude so.

(b)
Since this is a one-sided (right-tailed) test with known \(\sigma\), we reject \(H_0\) when the test statistic exceeds the critical value \(z^* = 1.645\) at \(\alpha = 0.05\).
The test statistic is \(z = \dfrac{\bar{x} – \mu_0}{\sigma / \sqrt{n}} = \dfrac{\bar{x} – 122}{15/\sqrt{100}} = \dfrac{\bar{x} – 122}{1.5}\).
Setting \(\dfrac{\bar{x} – 122}{1.5} > 1.645\) and solving:
\(\bar{x} – 122 > 1.645 \times 1.5 = 2.4675\)
\(\boxed{\bar{x} > 124.4675 \text{ mmHg}}\)
Any sample mean greater than approximately 124.47 mmHg provides sufficient evidence to reject \(H_0\) at the \(\alpha = 0.05\) level.

(c)
Now the true population mean is \(\mu = 125\) mmHg (with \(\sigma = 15\), \(n = 100\)), so the sampling distribution of \(\bar{x}\) is approximately normal with mean 125 and standard deviation \(\dfrac{15}{\sqrt{100}} = 1.5\).
We want the probability that \(\bar{x}\) falls in the rejection region found in part (b):
\(P(\bar{x} > 124.4675) = P\!\left(z > \dfrac{124.4675 – 125}{1.5}\right) = P(z > -0.355)\
Folk \boxed{P(z > -0.355) \approx 0.64}\)
There is approximately a 64% probability that the null hypothesis will be rejected when the true mean is 125 mmHg.

(d)
The probability found in part (c) — the probability of correctly rejecting a false null hypothesis — is called the power of the test.
\(\boxed{\text{Power of the test} \approx 0.64}\)

(e)
The probability of rejecting \(H_0\) would be greater than 0.64 if the sample size is larger than 100.
A larger sample size \(n\) reduces the standard error of \(\bar{x}\): \(\sigma_{\bar{x}} = \dfrac{\sigma}{\sqrt{n}}\), so the sampling distribution becomes narrower.
This means the critical value \(\bar{x}^* = \mu_0 + z^* \cdot \dfrac{\sigma}{\sqrt{n}}\) would be smaller (closer to 122), lowering the threshold needed to reject \(H_0\).
With a lower rejection threshold, it becomes more likely that \(\bar{x}\) exceeds that threshold when the true mean is 125 mmHg — so the power of the test increases.
\(\boxed{\text{Probability of rejecting } H_0 \text{ would be greater than 0.64.}}\)

Question

One of the two fire stations in a certain town responds to calls in the northern half of the town, and the other fire station responds to calls in the southern half of the town. One of the town council members believes that the two fire stations have different mean response times. Response time is measured by the difference between the time an emergency call comes into the fire station and the time the first fire truck arrives at the scene of the fire.
Data were collected to investigate whether the council member’s belief is correct. A random sample of 50 calls selected from the northern fire station had a mean response time of 4.3 minutes with a standard deviation of 3.7 minutes. A random sample of 50 calls selected from the southern fire station had a mean response time of 5.3 minutes with a standard deviation of 3.2 minutes.
(a) Construct and interpret a 95 percent confidence interval for the difference in mean response times between the two fire stations.
(b) Does the confidence interval in part (a) support the council member’s belief that the two fire stations have different mean response times? Explain.

Most-appropriate topic codes (AP Statistics):

• Topic \(4.7\) — Constructing a Confidence Interval for the Difference Between Two Population Means (Part \(\mathrm{a}\))
• Topic \(4.8\) — Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Means (Part \(\mathrm{b}\))

▶️ Answer/Explanation

(a)

Step 1 — Identify the procedure and check conditions
Let \(\mu_N\) = true mean response time (in minutes) for calls to the northern fire station, and \(\mu_S\) = true mean response time for calls to the southern fire station.
We will construct a two-sample \(t\)-interval for \(\mu_N – \mu_S\):
\( (\bar{x}_N – \bar{x}_S) \pm t^* \sqrt{\dfrac{s_N^2}{n_N} + \dfrac{s_S^2}{n_S}} \)

Conditions:
1. Independent random samples: The problem states that both samples were randomly selected, and the northern and southern calls are independent of each other.
2. Large samples: Both sample sizes are \(n_N = n_S = 50 > 30\), so by the Central Limit Theorem the sampling distribution of \(\bar{x}_N – \bar{x}_S\) is approximately normal, even without knowing the shape of the population distributions.

Step 2 — Mechanics
The known values are:
\( \bar{x}_N = 4.3,\quad s_N = 3.7,\quad n_N = 50 \)
\( \bar{x}_S = 5.3,\quad s_S = 3.2,\quad n_S = 50 \)
Point estimate of the difference:
\( \bar{x}_N – \bar{x}_S = 4.3 – 5.3 = -1.0 \text{ minutes} \)
Standard error of the difference:
\( SE = \sqrt{\dfrac{s_N^2}{n_N} + \dfrac{s_S^2}{n_S}} = \sqrt{\dfrac{3.7^2}{50} + \dfrac{3.2^2}{50}} = \sqrt{\dfrac{13.69}{50} + \dfrac{10.24}{50}} = \sqrt{0.2738 + 0.2048} = \sqrt{0.4786} \approx 0.6918 \)
Using conservative degrees of freedom \(df = 49\) (smaller of \(n_N – 1\) and \(n_S – 1\)), the critical value from the \(t\)-table is:
\( t^* \approx 2.010 \quad (df = 49,\ 95\%\ \text{confidence}) \)
The 95% confidence interval is:
\( -1.0 \pm 2.010 \times 0.6918 = -1.0 \pm 1.390 \)
\( \boxed{(-2.39,\ 0.39) \text{ minutes}} \)

Step 3 — Interpretation
Based on these samples, we are 95% confident that the true difference in mean response times (northern minus southern) is between \(-2.39\) minutes and \(0.39\) minutes.
In other words, the northern fire station’s mean response time could be anywhere from about 2.39 minutes faster to about 0.39 minutes slower than the southern fire station’s mean response time.

(b)
The confidence interval \((-2.39,\ 0.39)\) does not support the council member’s belief that the two fire stations have different mean response times.
The value \(0\) is contained within the interval, which means a true difference of \(\mu_N – \mu_S = 0\) (i.e., equal mean response times) is a plausible value based on the data. Because zero is a plausible value for the difference, we cannot conclude that a difference in mean response times actually exists.
It would be statistically incorrect to say the council member is definitely wrong — we simply don’t have sufficient evidence from this data to confirm a difference exists.

Question

The department of agriculture at a university was interested in determining whether a preservative was effective in reducing discoloration in frozen strawberries. A sample of 50 ripe strawberries was prepared for freezing. Then the sample was randomly divided into two groups of 25 strawberries each. Each strawberry was placed into a small plastic bag.
The 25 bags in the control group were sealed. The preservative was added to the 25 bags containing strawberries in the treatment group, and then those bags were sealed. All bags were stored at \(0^\circ\text{C}\) for a period of 6 months. At the end of this time, after the strawberries were thawed, a technician rated each strawberry’s discoloration from 1 to 10, with a low score indicating little discoloration.
The dotplots below show the distributions of discoloration rating for the control and treatment groups.
(a) The standard deviation of ratings for the control group is 2.141. Explain how this value summarizes variability in the control group.
(b) Based on the dotplots, comment on the effectiveness of the preservative in lowering the amount of discoloration in strawberries. (No calculations are necessary.)
(c) Researchers at the university decided to calculate a 95 percent confidence interval for the difference in mean discoloration rating between strawberries that were not treated with preservative and those that were treated with preservative. The confidence interval they obtained was \((0.16,\ 2.72)\). Assume that the conditions necessary for the \(t\)-confidence interval are met.
Based on the confidence interval, comment on whether there would be a difference in the population mean discoloration ratings for the treated and untreated strawberries.

Most-appropriate topic codes (AP Statistics):

• Topic 1.7 — Summary Statistics for One Quantitative Variable (Part \(\mathrm{a}\))
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Part \(\mathrm{b}\))
• Topic 4.7 — Constructing a Confidence Interval for the Difference Between Two Population Means (Part \(\mathrm{c}\))
• Topic 4.8 — Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Means (Part \(\mathrm{c}\))
▶️ Answer/Explanation

(a)

The standard deviation \(s = 2.141\) tells us roughly how far a typical discoloration rating in the control group falls from the group’s mean rating.
In other words, on average, a single strawberry’s discoloration rating in the control group deviates from the mean discoloration rating by about \(2.141\) points.
A larger standard deviation would indicate that the individual ratings are more spread out from the mean, while a smaller value would indicate the ratings are clustered more tightly around the mean.
\(\boxed{s = 2.141 \text{ is the typical distance of each discoloration rating from the mean in the control group.}}\)

(b)

The preservative does appear to have been effective in lowering the amount of discoloration in strawberries.
Looking at the dotplots, the treatment group’s distribution is clearly centered at a lower value than the control group’s distribution — the treatment group’s center is around 5, while the control group’s center is around 7.
Additionally, nearly all the summary statistics (minimum, \(Q_1\), median, and \(Q_3\)) are lower for the treatment group than for the control group, which consistently supports the conclusion that the preservative reduced discoloration.
\(\boxed{\text{The preservative appears effective: treatment ratings are consistently lower than control ratings.}}\)

(c)

We are given a 95% confidence interval for \(\mu_{\text{control}} – \mu_{\text{treatment}}\) as \((0.16,\ 2.72)\).
Since the entire interval lies above zero — meaning zero is not contained within \((0.16,\ 2.72)\) — we have convincing evidence that the two population means are not equal.
We are 95% confident that the population mean discoloration rating for untreated strawberries is between \(0.16\) and \(2.72\) units higher than the population mean rating for treated strawberries, indicating the preservative does make a real difference.
\(\boxed{0 \notin (0.16,\ 2.72) \Rightarrow \text{there is a significant difference in population mean discoloration ratings.}}\)

Question

Patients with heart-attack symptoms arrive at an emergency room either by ambulance or self-transportation provided by themselves, family, or friends. When a patient arrives at the emergency room, the time of arrival is recorded. The time when the patient’s diagnostic treatment begins is also recorded.
An administrator of a large hospital wanted to determine whether the mean wait time (time between arrival and diagnostic treatment) for patients with heart-attack symptoms differs according to the mode of transportation. A random sample of 150 patients with heart-attack symptoms who had reported to the emergency room was selected. For each patient, the mode of transportation and wait time were recorded. Summary statistics for each mode of transportation are shown in the table below.

(a) Use a 99 percent confidence interval to estimate the difference between the mean wait times for ambulance-transported patients and self-transported patients at this emergency room.
(b) Based only on this confidence interval, do you think the difference in the mean wait times is statistically significant? Justify your answer.

Most-appropriate topic codes (AP Statistics):

• Topic 4.7 — Constructing a Confidence Interval for the Difference Between Two Population Means (Part \(\mathrm{a}\))
• Topic 4.8 — Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Means (Part \(\mathrm{b}\))
▶️ Answer/Explanation

(a)

We use a two-sample \(t\)-interval for \(\mu_A – \mu_S\), the difference in mean wait times (Ambulance \(-\) Self).
Conditions:
The 150 patients were randomly selected, so it is reasonable to treat the ambulance and self-transport groups as independent random samples. Both sample sizes are large (\(n_A = 77 > 30\) and \(n_S = 73 > 30\)), so by the Central Limit Theorem the sampling distributions of the sample means are approximately normal.

Mechanics:
Using the conservative degrees of freedom \(df = \min(77-1,\, 73-1) = 72\) and \(t^* = 2.6459\) at the 99% level:
\((\bar{x}_A – \bar{x}_S) \pm t^* \sqrt{\frac{s_A^2}{n_A} + \frac{s_S^2}{n_S}}\)
\((6.04 – 8.30) \pm 2.6459\sqrt{\frac{4.30^2}{77} + \frac{5.16^2}{73}}\)
\(-2.26 \pm 2.6459\sqrt{\frac{18.49}{77} + \frac{26.63}{73}}\)
\(-2.26 \pm 2.6459\sqrt{0.2401 + 0.3648}\)
\(-2.26 \pm 2.6459 \times 0.7778\)
\(-2.26 \pm 2.0577\)
\(\boxed{(-4.318,\ -0.202)}\)
Interpretation: Based on this sample, we are 99% confident that the true difference in population mean wait times (Ambulance \(-\) Self) is between \(-4.318\) minutes and \(-0.202\) minutes. That is, ambulance-transported patients wait, on average, somewhere between about 0.2 and 4.3 minutes less than self-transported patients.

(b)

Yes, the difference in mean wait times is statistically significant at the \(\alpha = 0.01\) level.
Since the value \(0\) is not contained in the 99% confidence interval \((-4.318,\ -0.202)\), we can reject \(H_0: \mu_A – \mu_S = 0\) in favor of \(H_a: \mu_A – \mu_S \neq 0\) at the \(\alpha = 0.01\) significance level.
There is sufficient evidence to conclude that the mean wait time for ambulance-transported patients is significantly different from (specifically, shorter than) the mean wait time for self-transported patients.
\(\boxed{0 \notin (-4.318,\ -0.202) \implies \text{Reject } H_0 \text{ at } \alpha = 0.01}\)

Question

Lead, found in some paints, is a neurotoxin that can be especially harmful to the developing brain and nervous system of children. Children frequently put their hands in their mouth after touching painted surfaces, and this is the most common type of exposure to lead.
A study was conducted to investigate whether there were differences in children’s exposure to lead between suburban day-care centers and urban day-care centers in one large city. For this study, researchers used a random sample of 20 children in suburban day-care centers. Ten of these 20 children were randomly selected to play outside; the remaining 10 children played inside. All children had their hands wiped clean before beginning their assigned one-hour play period either outside or inside. After the play period ended, the amount of lead in micrograms (mcg) on each child’s dominant hand was recorded.
The mean amount of lead on the dominant hand for the children playing inside was \(3.75\) mcg, and the mean amount of lead for the children playing outside was \(5.65\) mcg. A \(95\) percent confidence interval for the difference in the mean amount of lead after one hour inside versus one hour outside was calculated to be \((-2.46, -1.34)\).
A random sample of 18 children in urban day-care centers in the same large city was selected. For this sample, the same process was used, including randomly assigning children to play inside or outside. The data for the amount (in mcg) of lead on each child’s dominant hand are shown in the table below.
(a) Use a 95 percent confidence interval to estimate the difference in the mean amount of lead on a child’s dominant hand after an hour of play inside versus an hour of play outside at urban day-care centers in this city. Be sure to interpret your interval.
(b) On the figure below,
• Using the vertical axis for the mean amount of lead, plot the mean for the amounts of lead on the dominant hand of children who played inside at the suburban day-care center and then plot the mean for the amounts of lead on the dominant hand of children who played inside at the urban day-care center.
• Connect these two points with a line segment.
• Plot the two means (suburban and urban) for the children who played outside at the two types of day-care centers.
• Connect these two points with a second line segment.
(c) From the study, what conclusions can be drawn about the impact of setting (inside, outside), environment (suburban, urban), and the relationship between the two on the amount of lead on the dominant hand of children after play in this city? Justify your answer.

Most-appropriate topic codes (AP Statistics):

• Topic 4.7 — Constructing a Confidence Interval for the Difference Between Two Population Means (Part a)
• Topic 4.8 — Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Means (Part a, Part c)
• Topic 5.1 — Graphical Representations Between Two Quantitative Variables (Part b)
▶️ Answer/Explanation

(a)
First check conditions for a two-sample \(t\)-interval: the two groups of urban children were assigned at random and independently to play inside or outside, and dotplots of each group’s data show no strong skew or outliers, so it’s reasonable to treat the underlying populations as approximately normal.

Summary statistics for the urban sample:
\( n_{\text{in}}=9,\quad \bar{x}_{\text{in}}=4.56,\quad s_{\text{in}}=0.846 \)
\( n_{\text{out}}=9,\quad \bar{x}_{\text{out}}=17.56,\quad s_{\text{out}}=4.61 \)
The two-sample \(t\)-confidence interval formula is
\( (\bar{x}_{\text{in}}-\bar{x}_{\text{out}})\pm t^*\sqrt{\dfrac{s_{\text{in}}^2}{n_{\text{in}}}+\dfrac{s_{\text{out}}^2}{n_{\text{out}}}} \)
Using the conservative degrees of freedom \(df=\min(n_{\text{in}}-1,\,n_{\text{out}}-1)=8\), so \(t^*=2.306\):
\( (4.56-17.56)\pm 2.306\sqrt{\dfrac{(0.846)^2}{9}+\dfrac{(4.61)^2}{9}} \)
\( -13.00\pm 2.306(1.564) \)
\( -13.00\pm 3.61 \)
\( \boxed{(-16.60,\ -9.40)\text{ mcg}} \)
Interpretation: we are 95% confident that, for the population of urban day-care children, the mean amount of lead on the dominant hand after an hour of play inside is between 9.40 and 16.60 mcg lower than after an hour of play outside. Since this interval doesn’t contain zero, the difference is meaningful — urban children who play outside pick up noticeably more lead on their hands.

(b)

Plot the points \((\text{Suburban},3.75)\) and \((\text{Urban},4.56)\), connect them with a line labeled “inside.” Then plot \((\text{Suburban},5.65)\) and \((\text{Urban},17.56)\), connect them with a line labeled “outside.” The “inside” line should be nearly flat and low on the graph, while the “outside” line should rise sharply from suburban to urban.
\( \boxed{\text{Inside line: nearly flat, low values; Outside line: steep increase from suburban to urban}} \)

(c)
Setting (inside vs. outside): In both suburban and urban environments, children who played outside ended up with more lead on their hands than children who played inside. This is supported by the fact that all four endpoints of the two confidence intervals (inside minus outside) are negative, and the graph shows the “outside” line sitting above the “inside” line everywhere.
Environment (suburban vs. urban): For both inside and outside play, urban children had more lead on their hands on average than suburban children. The graph shows both lines sloping upward from suburban to urban.
Relationship between the two: The effect of going inside versus outside depends heavily on the environment. In the suburban setting, the inside and outside means are fairly close together (3.75 vs. 5.65), but in the urban setting the gap is much larger (4.56 vs. 17.56). In other words, playing outside makes a much bigger difference in lead exposure in the urban environment than in the suburban environment.
\( \boxed{\text{Outside > Inside in both settings; Urban > Suburban in both settings; the inside/outside gap is much larger for urban than suburban}} \)

Question

The principal at Crest Middle School, which enrolls only sixth-grade students and seventh-grade students, is interested in determining how much time students at that school spend on homework each night. The table below shows the mean and standard deviation of the amount of time spent on homework each night (in minutes) for a random sample of 20 sixth-grade students and a separate random sample of 20 seventh-grade students at this school.
Based on dotplots of these data, it is not unreasonable to assume that the distribution of times for each grade were approximately normally distributed.
(a) Estimate the difference in mean times spent on homework for all sixth- and seventh-grade students in this school using an interval. Be sure to interpret your interval.
(b) An assistant principal reasoned that a much narrower confidence interval could be obtained if the students were paired based on their responses; for example, pairing the sixth-grade student and the seventh-grade student with the highest number of minutes spent on homework, the sixth-grade student and seventh-grade student with the next highest number of minutes spent on homework, and so on. Is the assistant principal correct in thinking that matching students in this way and then computing a matched-pairs confidence interval for the mean difference in time spent on homework is a better procedure than the one used in part (a)? Explain why or why not.

Most-appropriate topic codes (AP Statistics):

• Topic 4.7 — Constructing a Confidence Interval for the Difference Between Two Population Means (Part a)
• Topic 4.8 — Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Means (Part a)
• Topic 4.7 — Constructing a Confidence Interval for the Difference Between Two Population Means (Part b)
▶️ Answer/Explanation

(a)

We use a two-sample \(t\)-interval for the difference in means \((\mu_1 – \mu_2)\), where:
\(\mu_1\) = mean homework time for all sixth-graders at Crest Middle School
\(\mu_2\) = mean homework time for all seventh-graders at Crest Middle School
Assumptions checked: The two samples are independent random samples. The dotplots indicate that it is not unreasonable to assume approximate normality for both groups. So we may proceed.
The general form of the two-sample \(t\)-interval is:
\((\bar{X}_1 – \bar{X}_2) \pm t^* \sqrt{\dfrac{s_1^2}{n_1} + \dfrac{s_2^2}{n_2}}\)
Substituting the values (using sixth-grade as group 1 and seventh-grade as group 2):
\((27.3 – 47.0) \pm t^* \sqrt{\dfrac{10.8^2}{20} + \dfrac{12.4^2}{20}}\)
\(-19.7 \pm t^* \sqrt{\dfrac{116.64}{20} + \dfrac{153.76}{20}}\)
\(-19.7 \pm t^* \sqrt{5.832 + 7.688} = -19.7 \pm t^* \sqrt{13.52} = -19.7 \pm t^*(3.68)\)
Using \(t^* = 2.026\) based on approximately 37.297 degrees of freedom at a 95% confidence level:
\(-19.7 \pm 2.026 \times 3.68 = -19.7 \pm 7.45\)
\(\boxed{(-27.15,\ -12.25)}\)
Interpretation: Based on these samples, we can be 95% confident that the true difference in mean homework times (sixth-graders minus seventh-graders) for all students at Crest Middle School is between \(-27.15\) minutes and \(-12.25\) minutes. That is, seventh-graders spend, on average, between about 12.25 and 27.15 more minutes per night on homework than sixth-graders.

(b)

No, the assistant principal’s suggestion is not a valid procedure, and it would not produce a better confidence interval. Matching students based on their homework time responses — pairing the highest sixth-grader with the highest seventh-grader, and so on — is inappropriate for two key reasons.
First, valid matched-pairs designs require that pairs be formed before data are collected, based on some variable related to the response (such as prior GPA, study habits, or some pre-existing characteristic), not on the response variable itself. Pairing after the fact on the observed response creates an artificial association between the two independent samples.
Second, because the two samples were drawn independently with no natural connection between any sixth-grader and any seventh-grader, forcing them into pairs based on their ranked responses would artificially inflate the correlation between the paired differences. This would produce a confidence interval that is misleadingly narrow and fails to achieve the stated confidence level, giving a false impression of precision.

Scroll to Top