AP Statistics 1.10 The Investigative Question Revisited and Data Collection- Exam Style Questions - FRQs - New Syllabus
Question
Most-appropriate topic codes (AP Statistics):
• Topic \(1.13\) — Experimental Design (Part \( \mathrm{b} \))
▶️ Answer/Explanation
(a)
This is an observational study.
The researchers are simply gathering data by asking car owners to estimate their mileage, without actively imposing any treatments or randomly assigning participants to drive specific car models.
(b)
First, number the \(70\) days from \(1\) to \(70\).
Write the numbers \(1\) through \(70\) on identical slips of paper, place them into a hat, and mix them thoroughly.
Draw \(35\) slips of paper one by one without replacement.
The \(35\) days corresponding to the drawn numbers will be assigned the treatment of driving with the autopilot feature, and the remaining \(35\) days will be assigned to drive without the autopilot feature.
(c)
In order to generalize his findings to all Model D cars in his club, James cannot solely rely on an experiment conducted using only his own vehicle.
He would need to select a random sample of Model D cars (and their respective drivers) from the club’s membership to participate in his study.
Question
• Treatments
• Response variable
Most-appropriate topic codes (AP Statistics):
• Topic \(1.13\) — Experimental Design (Parts \( \mathrm{b} \), \( \mathrm{c} \))
▶️ Answer/Explanation
(a)
Experimental units: The 60 driveways (or the 60 new homes).
Treatments: The two types of concrete (concrete containing fibers and concrete without fibers).
Response variable: The severity of cracks in each driveway, measured on a scale of 0 to 10.
(b)
Number the 60 driveways from 1 to 60. Write the numbers 1 through 60 on identical slips of paper, place them in a hat, and mix thoroughly. Without looking, select 30 slips of paper one by one without replacement. The 30 driveways corresponding to the selected numbers will receive the concrete containing fibers, and the remaining 30 driveways will receive the concrete without fibers.
(c)
The primary benefit of random assignment is that it creates two treatment groups that are roughly equivalent at the beginning of the experiment. This helps to balance out the effects of potentially confounding variables, such as soil quality, driveway slope, or vehicle weight, across both groups. Because these other variables are distributed somewhat equally, the developer can confidently conclude that the statistically significant reduction in crack severity was actually caused by the addition of fibers to the concrete.
Question
• Experimental units:
• Response variable:
Most-appropriate topic codes (AP Statistics):
• Topic \(1.13\) — Experimental Design (Parts \( \mathrm{b} \), \( \mathrm{c} \))
▶️ Answer/Explanation
(a)
• Treatments: The new drug and the placebo.
• Experimental units: The 72 individual people (the 36 pairs of identical twins) participating in the experiment.
• Response variable: The improvement in acne severity, measured on a scale from 0 to 100.
(b)
A matched-pairs design controls for the variation in baseline acne severity and genetic/environmental factors among the subjects. Because identical twins share these traits, the difference in their acne improvement can be more directly attributed to the respective treatments rather than individual skin differences. This reduces the variability in the response and increases the statistical power to detect a true difference in effectiveness between the new drug and the placebo.
(c)
For each pair of identical twins, flip a fair coin. If the coin lands on heads, assign Twin A to receive the new drug and Twin B to receive the placebo. If the coin lands on tails, assign Twin A to receive the placebo and Twin B to receive the new drug. Repeat this random assignment process independently for all 36 pairs of twins.
Question



For mild allergy sufferers, Clinic B was more successful in treating allergies.
For severe allergy sufferers, Clinic B was more successful in treating allergies.
• Clinic B:
• Clinic B:
Most-appropriate topic codes (AP Statistics):
• Topic \(2.1\) — Tabular and Graphical Representations for the Distributions of Two Categorical Variables (Parts \( \mathrm{a} \), \( \mathrm{c} \))
• Topic \(2.2\) — Summary Statistics for Two Categorical Variables (Parts \( \mathrm{a} \), \( \mathrm{c} \), \( \mathrm{d} \))
▶️ Answer/Explanation
(a)(i)

(a)(ii)
Clinic A is more successful. The relative frequency of successful treatments for Clinic A ($0.633$ or $63.3\%$) is greater than the relative frequency of successful treatments for Clinic B ($0.515$ or $51.5\%$).
(b)
No. The researchers selected random samples of patient records, which means this is an observational study, not a randomized experiment. Because patients were not randomly assigned to Clinic A or Clinic B, we cannot establish a cause-and-effect relationship due to the potential presence of confounding variables.
(c)(i)
• Clinic A: Mild allergies are treated more successfully. The success rate for mild is $78 / (78 + 26) = 75\%$, while the success rate for severe is $11 / (11 + 24) = 31.4\%$.
• Clinic B: Mild allergies are treated more successfully. The success rate for mild is $32 / (32 + 1) = 97\%$, while the success rate for severe is $10 / (10 + 25) = 28.6\%$.
(c)(ii)
• Clinic A: Mild allergies are more likely to be treated. Clinic A treated 104 mild cases ($78+26$) compared to only 35 severe cases ($11+24$).
• Clinic B: Severe allergies are more likely to be treated. Clinic B treated 35 severe cases ($10+25$) compared to only 33 mild cases ($32+1$).
(d)
The conclusion is different because of Simpson’s Paradox. Clinic A’s overall success rate is higher because it treats a much larger proportion of mild allergy cases, which naturally have a higher success rate regardless of the clinic. Conversely, Clinic B treats a higher proportion of severe cases, which brings its overall average down, even though it performs better than Clinic A within each specific severity group.
Question
Most-appropriate topic codes (AP Statistics):
• Topic \(1.13\) — Experimental Design (Part \( \mathrm{c} \))
▶️ Answer/Explanation
(a)
Treatments: The four different concentrations of the fungus spray (\(0\,\text{ml/L}\), \(1.25\,\text{ml/L}\), \(2.5\,\text{ml/L}\), and \(3.75\,\text{ml/L}\)).
Experimental units: The \(20\) individual containers, each containing an equal number of insects.
Response variable: The number of insects that are still alive in each container one week after being sprayed.
(b)
Yes, the experiment definitely has a control group.
The containers sprayed with the \(0\,\text{ml/L}\) concentration form the control group because this specific mixture contains absolutely no fungus, serving as a baseline for comparison.
(c)
First, label each of the \(20\) containers with a unique integer from \(1\) to \(20\).
Next, use a random number generator to select \(15\) unique integers from \(1\) to \(20\) without replacement.
Assign the first five containers selected to receive the \(0\,\text{ml/L}\) treatment.
Assign the next five containers selected to receive the \(1.25\,\text{ml/L}\) treatment, and the next five to receive the \(2.5\,\text{ml/L}\) treatment.
Finally, the remaining five containers that were not selected will automatically receive the \(3.75\,\text{ml/L}\) treatment.
Question
Most-appropriate topic codes (AP Statistics):
• Topic \(1.11\) — Random Smapling (Part \( \mathrm{b} \))
• Topic \(1.12\) — Potential Problems with Sampling (Part \( \mathrm{c} \))
▶️ Answer/Explanation
(a)
Explanatory variable: The person’s degree of cigarette smoking — specifically, whether or not the individual smoked at least two packs of cigarettes per day (smoker of at least two packs per day versus non-smoker at that level).
Response variable: Whether or not the person developed Alzheimer’s disease during the course of the 23-year study.
(b)
This is an observational study, not an experiment. In an experiment, researchers would have to actively assign the treatment — in this case, the level of cigarette smoking — to the participants. Instead, the researchers simply tracked the existing medical histories of 21,123 men and women over 23 years. The smoking status of each person was passively observed and recorded, not controlled or manipulated by the researchers. Since no treatment was imposed, this is an observational study.
(c)
A confounding variable is one that is related to the explanatory variable and also independently influences the response variable, making it difficult to determine whether the explanatory variable alone is responsible for the observed association.
Exercise status could be a confounding variable here for two reasons working together:
First, people who exercise regularly tend to be more health-conscious overall, and as a result are less likely to smoke heavily. So exercise status is related to the explanatory variable — smoking status. Heavy smokers are, on average, less likely to exercise regularly than non-smokers.
Second, regular exercise may independently reduce the risk of developing Alzheimer’s disease. So exercise status is also related to the response variable — development of Alzheimer’s disease.
Because of both of these relationships, the observed association between heavy smoking and higher rates of Alzheimer’s could be at least partly explained by the fact that heavy smokers tend to exercise less — and it is the lack of exercise, not the smoking itself, that contributes to the increased risk. This makes it impossible to determine from this study alone whether the association between smoking and Alzheimer’s reflects a true causal relationship or is merely the result of exercise status being linked to both.
Question


Most-appropriate topic codes (AP Statistics):
• Topic \(1.7\) — Summary Statistics for a Quantitative Variable (Center, Spread, Shape) (Parts \(\mathrm{a}\), \(\mathrm{b}\), \(\mathrm{c}\))
• Topic \(1.9\) — Comparing Distributions of a Quantitative Variable (Parts \(\mathrm{a}\), \(\mathrm{b}\))
• Topic \(1.10\) — The Effect of Adding a Constant or Multiplying by a Constant on Summary Statistics (Part \(\mathrm{b}\))
• Topic \(3.1\) — Mean and Standard Deviation of a Linear Transformation (Part \(\mathrm{c}\))
▶️ Answer/Explanation
(a)
When comparing two distributions from boxplots, we examine center, spread, shape, and unusual features.
The cereals with one-cup serving sizes have a higher median sugar content per serving than the cereals with three-quarter-cup serving sizes. The one-cup distribution also has greater variability, as indicated by its larger range and larger interquartile range (IQR). In terms of shape, the one-cup distribution appears somewhat left-skewed because the median is closer to the upper quartile than to the lower quartile, while the three-quarter-cup distribution is more nearly symmetric. Neither distribution appears to contain extreme outliers.
(b)
Multiplying each sugar value in the three-quarter-cup group by \(\dfrac{4}{3}\) converts the measurements to sugar content per cup, allowing a fair comparison using equal serving sizes.
The adjusted boxplot shows that cereals with recommended serving sizes of three-quarter cup tend to contain more sugar per cup than cereals with recommended serving sizes of one cup. The median for the adjusted three-quarter-cup distribution is now noticeably higher than the median for the one-cup distribution. In addition, all measures of spread (range and IQR) for the adjusted distribution have increased by a factor of \(\dfrac{4}{3}\), reflecting the effect of multiplying every observation by a constant.
(c)
We would expect the mean sugar content per cup to be greater for cereals that list a serving size of three-quarter cup.
After adjustment, the three-quarter-cup distribution has a higher center than the one-cup distribution, as seen from its higher median. Because the mean generally follows the center of the distribution, the higher overall location of the adjusted three-quarter-cup distribution suggests a larger mean sugar content per cup.
\(\boxed{\bar{x}_{\frac{3}{4}\text{-cup}} > \bar{x}_{\text{1-cup}}}\)
Question
Most-appropriate topic codes (AP Statistics):
• Topic 1.10 — The Investigative Question Revisited and Data Collection (Part \(\mathrm{a}\))
• Topic 1.11 — Random Sampling (Part \(\mathrm{b}\))
▶️ Answer/Explanation
(a)
A control group gives the researchers a baseline comparison group — without any dietary supplement — so they can measure whether glucosamine or chondroitin actually makes a difference beyond what would happen due to the normal aging process alone.
Without a control group, we could not tell whether any improvements in joint and hip health were caused by the supplements or simply by other factors such as the passage of time, veterinary care, or natural variation between dogs.
In this study specifically, the control group allows us to isolate the true effect of glucosamine and chondroitin on reducing canine osteoarthritis by comparing both treatment groups against untreated dogs under the same conditions.
\(\boxed{\text{Control group provides a baseline to determine whether the supplements are truly effective.}}\)
(b)
First, assign each of the 300 dogs a unique number from \(001\) to \(300\).
Then, use a random number generator (calculator, statistical software, or a random number table) to randomly select 100 numbers from \(001\) to \(300\), ignoring any repeats — the dogs corresponding to these 100 numbers are assigned to the glucosamine group.
From the remaining 200 dogs, randomly select another 100 numbers using the same process — these dogs are assigned to the chondroitin group.
The final 100 remaining dogs are assigned to the control group and receive no dietary supplement.
\(\boxed{\text{Randomly assign dogs numbered } 001\text{–}300 \text{ into three equal groups of 100 using a random number generator.}}\)
(c)
The blocking variable should be the one that has a stronger association with the response variable — joint and hip health — so that dogs within each block are as similar (homogeneous) as possible.
Breed of dog is associated with the size of the dog, and size is known to be related to joint and hip health — larger breeds tend to have more joint problems than smaller breeds, so breed is likely to create more variability in the response.
Clinic, on the other hand, is less likely to be strongly associated with joint and hip health because most large veterinary practices see a wide variety of dog breeds and sizes, so dogs across clinics would not be meaningfully more similar to each other than dogs across breeds.
Therefore, we should block on breed of dog, since it is more strongly related to joint and hip health and will reduce variability more effectively within each block.
\(\boxed{\text{Block on breed of dog, as it has a stronger relationship to joint and hip health than clinic.}}\)
Question

Most-appropriate topic codes (AP Statistics):
• Topic 1.11 — Random Sampling (Part \(\mathrm{b}\))
• Topic 1.12 — Potential Problems with Sampling (Part \(\mathrm{b}\))
• Topic 1.10 — The Investigative Question Revisited and Data Collection (Part \(\mathrm{c}\))
▶️ Answer/Explanation
(a)
Reading the stemplot, the rural distribution is centered higher and is more spread out than the urban distribution.
For the rural students:
Mean \(\approx 40.45\) cal/kg
Median \(\approx 41\) cal/kg
Range \(= 19\)
SD \(\approx 6.04\)
IQR \(\approx 10\)
For the urban students:
Mean \(\approx 32.6\) cal/kg
Median \(\approx 32\) cal/kg
Range \(= 16\)
SD \(\approx 4.67\)
IQR \(\approx 7\)
So both the typical value and the spread are larger for the rural group. In terms of shape, the rural data look fairly symmetric and spread evenly between about 32 and 51 cal/kg, while the urban data appear skewed toward the larger values.
\( \boxed{\text{Rural: higher center and more spread; Urban: lower center, less spread, right-skewed}} \)
(b)
No. Each sample came from just one rural school and one urban school, so these two specific schools may not represent the much larger and more diverse population of all rural and urban ninth graders across the country. Because the schools themselves were not randomly chosen from all such schools, the results can’t be safely extended beyond these two schools.
\( \boxed{\text{No — only one school of each type was sampled, so results cannot be generalized nationally}} \)
(c)
Plan II is the better choice.
Both plans already adjust for body size by dividing calories by body weight, so that part is the same. The real issue is that a single day’s eating can be unusually high or low depending on what happened that day — a birthday party, a sick day, a weekend versus a school day, and so on. By recording food over a full 7-day period and averaging, Plan II smooths out this day-to-day variability and gives a more stable, precise picture of each student’s typical caloric intake.
\( \boxed{\text{Plan II — averaging over 7 days reduces day-to-day variability and gives a more precise estimate}} \)
Question

Most-appropriate topic codes (AP Statistics):
• Topic 1.10 — The Investigative Question Revisited and Data Collection (Parts a, b)
▶️ Answer/Explanation
(a)
Pairing consecutive volunteers by age (without regard to gender) gives the six blocks:

Since these researchers believe that the condition of hair changes with age but not gender, the volunteers are sorted from youngest to oldest. The volunteers in the sorted list are paired to form six blocks of size two. More specifically, the youngest two volunteers are placed in the first block. The next two volunteers in the sorted list are placed in the second block. This pairing continues until all six blocks of two are formed, with the oldest two volunteers in the sixth block.
(b)
Since both age and gender are believed to affect hair condition, volunteers must be blocked so that each block contains two people of the same gender and similar age.
First, separate by gender, then sort each group by age and pair consecutively
Males sorted by age: Volunteer 1 (age 21), Volunteer 11 (age 23), Volunteer 9 (age 44), Volunteer 3 (age 47), Volunteer 7 (age 58), Volunteer 6 (age 61)
Females sorted by age: Volunteer 2 (age 20), Volunteer 10 (age 24), Volunteer 8 (age 44), Volunteer 12 (age 46), Volunteer 4 (age 60), Volunteer 5 (age 62)
The six blocks are:

Since these researchers believe that the condition of hair changes with both age and gender, the women are sorted from youngest to oldest and then the men are sorted from youngest to oldest. The women (men) in the sorted list are paired to form the blocks of size two. More specifically, the youngest two women (men) are placed in a block. The next two youngest women (men) are placed in another block. Finally, the oldest two women (men) are placed in another block.
(c)
No, this is not an appropriate method for assigning treatments. Assigning the new formula to three entire blocks and the current formula to three other blocks does not allow a fair within-block comparison, because within each block both members would receive the same formula.
The correct approach in a matched-pairs block design is to randomly assign one person within each block to the new formula and the other to the current formula. This can be done as follows:
For each of the six blocks, flip a fair coin (or use a random number table/generator). If the result is heads (or the random digit is 0–4), assign the lower-numbered volunteer to the new formula and the other to the current formula. If tails (or digit 5–9), reverse the assignment. This ensures that within every block, one person receives the new formula and one receives the current formula, making a valid matched comparison possible.
\(\boxed{\text{Randomly assign one person per block to new formula; other receives current formula}}\)
