Home / AP® Exam / AP® Statistics / AP Statistics 1.13 Experimental Design- Exam Style Questions – FRQs

AP Statistics 1.13 Experimental Design- Exam Style Questions - FRQs - New Syllabus

Question

A car maker produces four different models of cars: A, B, C, and D. A group of researchers is investigating which model of car has the longest distance traveled per gallon of gas (mileage). Higher mileage is considered better than lower mileage. The researchers will conduct a study in which they contact several owners of each model of car and ask them to estimate their mileage.
(a) Is this an observational study or an experiment? Justify your answer in context.
Model D has an autopilot feature, in which the car controls its own motion with human supervision. James owns a Model D car and will investigate whether using the autopilot feature results in higher mileage than not using the autopilot. James will drive his car on \(70\) different days to and from work, using the same route at the same time each day. James will record the mileage each day.
(b) James will use a completely randomized design to conduct his investigation. Describe an appropriate method James could use to randomly assign the two treatments, driving using the autopilot feature and driving without using the autopilot feature, to \(35\) days each.
(c) After the investigation was completed, James verified that the conditions for inference were met and conducted a hypothesis test. He discovered the mean mileage when using the autopilot feature was significantly higher than the mean mileage when not using the autopilot feature. James is a member of a Model D club with thousands of members who all drive Model D cars. He will give a presentation at a Model D club members’ meeting later this year and would like to state that the results of his hypothesis test apply to all Model D cars in his club. Another member of the club who is a statistician tells James his findings do not apply to all Model D cars in the club. What change would James need to make to his original study to be able to generalize to all Model D cars in the club?

Most-appropriate topic codes (AP Statistics):

• Topic \(1.10\) — The Investigative Question Revisited and Data Collection (Parts \( \mathrm{a} \), \( \mathrm{c} \))
• Topic \(1.13\) — Experimental Design (Part \( \mathrm{b} \))
▶️ Answer/Explanation

(a)
This is an observational study.
The researchers are simply gathering data by asking car owners to estimate their mileage, without actively imposing any treatments or randomly assigning participants to drive specific car models.

(b)
First, number the \(70\) days from \(1\) to \(70\).
Write the numbers \(1\) through \(70\) on identical slips of paper, place them into a hat, and mix them thoroughly.
Draw \(35\) slips of paper one by one without replacement.
The \(35\) days corresponding to the drawn numbers will be assigned the treatment of driving with the autopilot feature, and the remaining \(35\) days will be assigned to drive without the autopilot feature.

(c)
In order to generalize his findings to all Model D cars in his club, James cannot solely rely on an experiment conducted using only his own vehicle.
He would need to select a random sample of Model D cars (and their respective drivers) from the club’s membership to participate in his study.

Question

A developer wants to know whether adding fibers to concrete used in paving driveways will reduce the severity of cracking, because any driveway with severe cracks will have to be repaired by the developer. The developer conducts a completely randomized experiment with 60 new homes that need driveways. Thirty of the driveways will be randomly assigned to receive concrete that contains fibers, and the other 30 driveways will receive concrete that does not contain fibers. After one year, the developer will record the severity of cracks in each driveway on a scale of 0 to 10, with 0 representing not cracked at all and 10 representing severely cracked.
 
(a) Based on the information provided about the developer’s experiment, identify each of the following.
• Experimental units
• Treatments
• Response variable
(b) Describe an appropriate method the developer could use to randomly assign concrete that contains fibers and concrete that does not contain fibers to the 60 driveways.
Suppose the developer finds that there is a statistically significant reduction in the mean severity of cracks in driveways using the concrete that contains fibers compared to the driveways using concrete that does not contain fibers.
(c) In terms of the developer’s conclusion, what is the benefit of randomly assigning the driveways to either the concrete that contains fibers or the concrete that does not contain fibers?

Most-appropriate topic codes (AP Statistics):

• Topic \(1.10\) — The Investigative Question Revisited and Data Collection (Part \( \mathrm{a} \))
• Topic \(1.13\) — Experimental Design (Parts \( \mathrm{b} \), \( \mathrm{c} \))
▶️ Answer/Explanation

(a)
Experimental units: The 60 driveways (or the 60 new homes).
Treatments: The two types of concrete (concrete containing fibers and concrete without fibers).
Response variable: The severity of cracks in each driveway, measured on a scale of 0 to 10.

(b)
Number the 60 driveways from 1 to 60. Write the numbers 1 through 60 on identical slips of paper, place them in a hat, and mix thoroughly. Without looking, select 30 slips of paper one by one without replacement. The 30 driveways corresponding to the selected numbers will receive the concrete containing fibers, and the remaining 30 driveways will receive the concrete without fibers.

(c)
The primary benefit of random assignment is that it creates two treatment groups that are roughly equivalent at the beginning of the experiment. This helps to balance out the effects of potentially confounding variables, such as soil quality, driveway slope, or vehicle weight, across both groups. Because these other variables are distributed somewhat equally, the developer can confidently conclude that the statistically significant reduction in crack severity was actually caused by the addition of fibers to the concrete.

Question

A dermatologist will conduct an experiment to investigate the effectiveness of a new drug to treat acne. The dermatologist has recruited 36 pairs of identical twins. Each person in the experiment has acne and each person in the experiment will receive either the new drug or a placebo. After each person in the experiment uses either the new drug or the placebo for 2 weeks, the dermatologist will evaluate the improvement in acne severity for each person on a scale from 0 (no improvement) to 100 (complete cure).
(a) Identify the treatments, experimental units, and response variable of the experiment.
• Treatments:
• Experimental units:
• Response variable:
Each twin in the experiment has a severity of acne similar to that of the other twin. However, the severity of acne differs from one twin pair to another.
(b) For the dermatologist’s experiment, describe a statistical advantage of using a matched-pairs design where twins are paired rather than using a completely randomized design.
(c) For the dermatologist’s experiment, describe how the treatments can be randomly assigned to people using a matched-pairs design in which twins are paired.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.10\) — The Investigative Question Revisited and Data Collection (Part \( \mathrm{a} \))
• Topic \(1.13\) — Experimental Design (Parts \( \mathrm{b} \), \( \mathrm{c} \))
▶️ Answer/Explanation

(a)
Treatments: The new drug and the placebo.
Experimental units: The 72 individual people (the 36 pairs of identical twins) participating in the experiment.
Response variable: The improvement in acne severity, measured on a scale from 0 to 100.

(b)
A matched-pairs design controls for the variation in baseline acne severity and genetic/environmental factors among the subjects. Because identical twins share these traits, the difference in their acne improvement can be more directly attributed to the respective treatments rather than individual skin differences. This reduces the variability in the response and increases the statistical power to detect a true difference in effectiveness between the new drug and the placebo.

(c)
For each pair of identical twins, flip a fair coin. If the coin lands on heads, assign Twin A to receive the new drug and Twin B to receive the placebo. If the coin lands on tails, assign Twin A to receive the placebo and Twin B to receive the new drug. Repeat this random assignment process independently for all 36 pairs of twins.

Question

Researchers are investigating the effectiveness of using a fungus to control the spread of an insect that destroys trees.
The researchers will create four different concentrations of fungus mixtures: \(0\) milliliters per liter (\(\text{ml/L}\)), \(1.25\,\text{ml/L}\), \(2.5\,\text{ml/L}\), and \(3.75\,\text{ml/L}\).
An equal number of the insects will be placed into \(20\) individual containers.
The group of insects in each container will be sprayed with one of the four mixtures, and the researchers will record the number of insects that are still alive in each container one week after spraying.
(a) Identify the treatments, experimental units, and response variable of the experiment.
Treatments:
Experimental units:
Response variable:
(b) Does the experiment have a control group? Explain your answer.
(c) Describe how the treatments can be randomly assigned to the experimental units so that each treatment has the same number of units.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.10\) — The Investigative Question Revisited and Data Collection (Parts \( \mathrm{a} \), \( \mathrm{b} \))
• Topic \(1.13\) — Experimental Design (Part \( \mathrm{c} \))
▶️ Answer/Explanation

(a)
Treatments: The four different concentrations of the fungus spray (\(0\,\text{ml/L}\), \(1.25\,\text{ml/L}\), \(2.5\,\text{ml/L}\), and \(3.75\,\text{ml/L}\)).
Experimental units: The \(20\) individual containers, each containing an equal number of insects.
Response variable: The number of insects that are still alive in each container one week after being sprayed.

(b)
Yes, the experiment definitely has a control group.
The containers sprayed with the \(0\,\text{ml/L}\) concentration form the control group because this specific mixture contains absolutely no fungus, serving as a baseline for comparison.

(c)
First, label each of the \(20\) containers with a unique integer from \(1\) to \(20\).
Next, use a random number generator to select \(15\) unique integers from \(1\) to \(20\) without replacement.
Assign the first five containers selected to receive the \(0\,\text{ml/L}\) treatment.
Assign the next five containers selected to receive the \(1.25\,\text{ml/L}\) treatment, and the next five to receive the \(2.5\,\text{ml/L}\) treatment.
Finally, the remaining five containers that were not selected will automatically receive the \(3.75\,\text{ml/L}\) treatment.

Question

The anterior cruciate ligament (ACL) is one of the ligaments that help stabilize the knee. Surgery is often recommended if the ACL is completely torn, and recovery time from the surgery can be lengthy. A medical center developed a new surgical procedure designed to reduce the average recovery time from the surgery. To test the effectiveness of the new procedure, a study was conducted in which 210 patients needing surgery to repair a torn ACL were randomly assigned to receive either the standard procedure or the new procedure.
(a) Based on the design of the study, would a statistically significant result allow the medical center to conclude that the new procedure causes a reduction in recovery time compared to the standard procedure, for patients similar to those in the study? Explain your answer.
(b) Summary statistics on the recovery times from the surgery are shown in the table.
Do the data provide convincing statistical evidence that those who receive the new procedure will have less recovery time from the surgery, on average, than those who receive the standard procedure, for patients similar to those in the study?

Most-appropriate topic codes (AP Statistics):

• Topic \(1.13\) — Experimental Design (Part \( \mathrm{a} \))
• Topic \(4.5\) — Carrying Out a Test for a Difference of Two Population Means (Part \( \mathrm{b} \))
▶️ Answer/Explanation

(a)
Yes, because the study used random assignment to assign participants to the two treatment groups.
Detailed Solution:
Random assignment creates comparable groups by balancing out potential confounding variables between the standard and new procedure groups.
Because of this design, a statistically significant difference in mean recovery times can indeed be attributed to the causal effect of the surgery type.
The inference applies to the population of patients similar to those who participated in the study.

(b)
Yes, the data provide convincing evidence that the new procedure results in lower average recovery time.
Detailed Solution:
We perform a two-sample $t$-test for a difference in means ($\mu_{standard} – \mu_{new} > 0$).
Calculating the standard error: $SE = \sqrt{\frac{34^2}{110} + \frac{29^2}{100}} = \sqrt{10.509 + 8.41} \approx 4.35$.
Calculating the $t$-statistic: $t = \frac{(217 – 186) – 0}{4.35} \approx 7.13$.
With a $t$-score of approximately $7.13$, the $p$-value is extremely small (virtually $0$), which is less than any standard significance level like $\alpha = 0.05$.
Therefore, we reject the null hypothesis and conclude there is overwhelming statistical evidence that the new procedure reduces recovery time.

Question

Consider an experiment in which two men and two women will be randomly assigned to either a treatment group or a control group in such a way that each group has two people. The people are identified as Man 1, Man 2, Woman 1, and Woman 2. The six possible arrangements are shown below.
Two possible methods of assignment are being considered: the sequential coin flip method, as described in part (a), and the chip method, as described in part (b). For each method, the order of the assignment will be Man 1, Man 2, Woman 1, Woman 2.
(a) For the sequential coin flip method, a fair coin is flipped until one group has two people. An outcome of tails assigns the person to the treatment group, and an outcome of heads assigns the person to the control group. As soon as one group has two people, the remaining people are automatically assigned to the other group.

(i) Complete the table below by calculating the probability of each arrangement occurring if the sequential coin flip method is used.

(ii) For the sequential coin flip method, what is the probability that Man 1 and Man 2 are assigned to the same group?
(b) For the chip method, two chips are marked “treatment” and two chips are marked “control.” Each person selects one chip at random without replacement.

(i) Complete the table below by calculating the probability of each arrangement occurring if the chip method is used.

(ii) For the chip method, what is the probability that Man 1 and Man 2 are assigned to the same group?
(c) Sixteen participants consisting of 10 students and 6 teachers at an elementary school will be used for an experiment to determine lunch preference for the school population of students and teachers. As the participants enter the school cafeteria for lunch, they will be randomly assigned to receive one of two lunches so that 8 will receive a salad, and 8 will receive a grilled cheese sandwich. The students will enter the cafeteria first, and the teachers will enter next. Which method, the sequential coin flip method or the chip method, should be used to assign the treatments? Justify your choice.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.13\) — Experimental Design (Part \( \mathrm{c} \))
• Topic \(2.6\) — Conditional Probability (Parts \( \mathrm{a} \), \( \mathrm{b} \))
• Topic \(2.7\) — Independent Events and Mutually Exclusive Events (Parts \( \mathrm{a} \), \( \mathrm{b} \))
▶️ Answer/Explanation

(a)(i)
Let T (tail) represent being assigned to the treatment group and H (head) represent being assigned to the control group. The process stops as soon as one group fills up. We trace each possible sequence of flips:

(a)(ii)
Man 1 and Man 2 are assigned to the same group only in Arrangements A (both in treatment) and D (both in control). So the probability is:

\(P(A) + P(D) = \dfrac{1}{4} + \dfrac{1}{4} = \boxed{\dfrac{1}{2}}\)

(b)(i)
Let T represent being assigned to the treatment group and C represent being assigned to the control group. Since chips are drawn without replacement from a pool of 2 T chips and 2 C chips, the probabilities change at each draw. Working through each arrangement:

(b)(ii)
Man 1 and Man 2 are in the same group only in Arrangements A and D. Therefore:

\(P(A) + P(D) = \dfrac{1}{6} + \dfrac{1}{6} = \boxed{\dfrac{1}{3}}\)

(c)
The chip method should be used. Here is the reasoning:

From parts (a)(i) and (b)(i), the chip method gives every arrangement an equal probability of \(\dfrac{1}{6}\), while the coin flip method assigns unequal probabilities — arrangements A and D each have probability \(\dfrac{1}{4}\), while B, C, E, and F each have probability \(\dfrac{1}{8}\).

From parts (a)(ii) and (b)(ii), the probability that both men end up in the same group is \(\dfrac{1}{2}\) under the coin method but only \(\dfrac{1}{3}\) under the chip method. Since students enter first and teachers enter next, the coin flip method is more likely to place all students together in one group — if teachers and students have different food preferences, this imbalance would make it impossible to tell whether any observed difference in lunch preference is due to the treatment (type of lunch) or the role of the participant (teacher vs. student).

The chip method, by giving all arrangements an equal chance, is therefore more appropriate for this experiment.

Question

Product advertisers studied the effects of television ads on children’s choices for two new snacks. The advertisers used two 30-second television ads in an experiment. One ad was for a new sugary snack called Choco-Zuties, and the other ad was for a new healthy snack called Apple-Zuties.
For the experiment, 75 children were randomly assigned to one of three groups, A, B, or C. Each child individually watched a 30-minute television program that was interrupted for 5 minutes of advertising. The advertising was the same for each group with the following exceptions.
• The advertising for group A included the Choco-Zuties ad but not the Apple-Zuties ad.
• The advertising for group B included the Apple-Zuties ad but not the Choco-Zuties ad.
• The advertising for group C included neither the Choco-Zuties ad nor the Apple-Zuties ad.
After the program, the children were offered a choice between the two snacks. The table below summarizes their choices.

(a) Do the data provide convincing statistical evidence that there is an association between type of ad and children’s choice of snack among all children similar to those who participated in the experiment?
(b) Write a few sentences describing the effect of each ad on children’s choice of snack.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.13\) — Experimental Design (Part \( \mathrm{b} \))
• Topic \(3.14\) — Setting Up a Chi-Square Test for Homogeneity or Independence (Part \( \mathrm{a} \))
• Topic \(3.15\) — Carrying Out a Chi-Square Test for Homogeneity or Independence (Part \( \mathrm{a} \))
▶️ Answer/Explanation

(a)

Step 1: State hypotheses.
\(H_0\): There is no association between the type of ad viewed and children’s choice of snack (the proportion choosing each snack is the same regardless of which ad is viewed).
\(H_a\): There is an association between the type of ad viewed and children’s choice of snack (the proportions differ based on which ad is viewed).

Step 2: Identify the procedure and check conditions.
The appropriate procedure is a chi-square test of homogeneity (also acceptable: chi-square test of independence).
Conditions:
1. Random: The 75 children were randomly assigned to the three groups — this condition is satisfied.
2. Large Counts: All expected cell counts must be at least 5. The expected counts are computed as:
\( E = \frac{(\text{row total}) \times (\text{column total})}{\text{table total}} \)
Expected counts table (observed counts with expected counts in parentheses):

The smallest expected count is \(6.33 \geq 5\), so the large counts condition is satisfied. Both conditions are met.

Step 3: Calculate the test statistic and \(p\)-value.
The chi-square test statistic is:

\( \chi^2 = \sum \frac{(O – E)^2}{E} \)
\( \chi^2 \approx \frac{(21-18.67)^2}{18.67} + \frac{(4-6.33)^2}{6.33} + \frac{(13-18.67)^2}{18.67} + \frac{(12-6.33)^2}{6.33} + \frac{(22-18.67)^2}{18.67} + \frac{(3-6.33)^2}{6.33} \)
\( \chi^2 \approx 0.292 + 0.860 + 1.720 + 5.070 + 0.595 + 1.754 \approx 10.291 \)
Degrees of freedom: \(df = (r-1)(c-1) = (3-1)(2-1) = 2\)
The \(p\)-value: \(P\!\left(\chi^2_{\,df=2} \geq 10.291\right) \approx 0.006\)

Step 4: State a conclusion in context.
Because the \(p\)-value \(\approx 0.006\) is much smaller than \(\alpha = 0.05\), we reject \(H_0\). The data provide convincing statistical evidence that there is an association between the type of ad viewed and children’s choice of snack, among all children similar to those who participated in the experiment.

(b)

When neither ad was shown (Group C), \(\dfrac{22}{25} = 88\%\) of the children chose Choco-Zuties, and only 12% chose Apple-Zuties — this reflects the baseline preference without any advertising.
When children saw the Choco-Zuties ad (Group A), 84% still chose Choco-Zuties, which is very close to the 88% in the no-ad group. So the Choco-Zuties ad had very little effect on children’s snack choice — they were already inclined toward the sugary option anyway.
When children saw the Apple-Zuties ad (Group B), only \(\dfrac{13}{25} = 52\%\) chose Choco-Zuties, and 48% chose Apple-Zuties. This is a large shift compared to the 12% who chose Apple-Zuties in the no-ad group, meaning the Apple-Zuties ad had a substantial effect in increasing children’s likelihood of choosing the healthy snack.

Question

Psychologists interested in the relationship between meditation and health conducted a study with a random sample of \(28\) men who live in a large retirement community. Of the men in the sample, \(11\) reported that they participate in daily meditation and \(17\) reported that they do not participate in daily meditation.
The researchers wanted to perform a hypothesis test of
\(H_0 : p_m – p_c = 0\)
\(H_a : p_m – p_c < 0,\)
where \(p_m\) is the proportion of men with high blood pressure among all the men in the retirement community who participate in daily meditation and \(p_c\) is the proportion of men with high blood pressure among all the men in the retirement community who do not participate in daily meditation.
(a) If the study were to provide significant evidence against \(H_0\) in favor of \(H_a\), would it be reasonable for the psychologists to conclude that daily meditation causes a reduction in blood pressure for men in the retirement community? Explain why or why not.
The psychologists found that of the \(11\) men in the study who participate in daily meditation, \(0\) had high blood pressure. Of the \(17\) men who do not participate in daily meditation, \(8\) had high blood pressure.
(b) Let \(\hat{p}_m\) represent the proportion of men with high blood pressure among those in a random sample of \(11\) who meditate daily, and let \(\hat{p}_c\) represent the proportion of men with high blood pressure among those in a random sample of \(17\) who do not meditate daily. Why is it not reasonable to use a normal approximation for the sampling distribution of \(\hat{p}_m – \hat{p}_c\)?
Although a normal approximation cannot be used, it is possible to simulate the distribution of \(\hat{p}_m – \hat{p}_c\). Under the assumption that the null hypothesis is true, \(10{,}000\) values of \(\hat{p}_m – \hat{p}_c\) were simulated. The histogram below shows the results of the simulation.
(c) Based on the results of the simulation, what can be concluded about the relationship between blood pressure and meditation among men in the retirement community?

Most-appropriate topic codes (AP Statistics):

• Topic \(1.13\) — Experimental Design (Part \( \mathrm{a} \))
• Topic \(3.12\) — Setting Up a Test for the Difference Between Two Population Proportions (Part \( \mathrm{b} \))
• Topic \(3.13\) — Carrying Out a Test for the Difference Between Two Population Proportions (Part \( \mathrm{c} \))
• Topic \(2.3\) — Estimating Probabilities Using Simulation (Part \( \mathrm{c} \))
▶️ Answer/Explanation

(a)

No, it would not be reasonable to conclude that daily meditation causes a reduction in blood pressure. This study is an observational study — the men themselves chose whether or not to meditate; no treatment was randomly assigned. Because there was no randomization of treatment, cause-and-effect conclusions cannot be drawn from the results. Men who choose to meditate may differ from men who don’t in other important ways that are also related to blood pressure, such as being more health-conscious, exercising more, or having lower stress in general. These potential confounding variables make it impossible to isolate meditation as the cause.

(b)

For a normal approximation to be valid for the sampling distribution of \(\hat{p}_m – \hat{p}_c\), we need the number of successes and failures in each group to each be at least \(10\). First, compute the combined sample proportion of successes:
$\hat{p} = \frac{0 + 8}{11 + 17} = \frac{8}{28} \approx 0.286$
Then check each group:
For the meditation group \((n_m = 11)\):
$n_m \hat{p} = 11 \times \frac{8}{28} \approx 3.14 < 10 \quad \text{(condition fails)}$
For the non-meditation group \((n_c = 17)\):
$n_c \hat{p} = 17 \times \frac{8}{28} \approx 4.86 < 10 \quad \text{(condition fails)}$
Since the expected number of successes in both groups is less than \(10\), the normal approximation condition is not met, and it is not reasonable to use a normal approximation for the sampling distribution of \(\hat{p}_m – \hat{p}_c\).

(c)

First, compute the observed value of the sample statistic from the data:
$\hat{p}_m – \hat{p}_c = \frac{0}{11} – \frac{8}{17} \approx -0.47$
From the simulation histogram, only \(76\) out of \(10{,}000\) simulated values were \(-0.47\) or less (the most extreme negative outcome), giving an approximate \(p\)-value of:
$p\text{-value} \approx \frac{76}{10{,}000} = 0.0076$

Since this \(p\)-value of \(0.0076\) is very small (less than any common significance level such as \(\alpha = 0.05\)), we reject \(H_0\). There is convincing statistical evidence that men in this retirement community who meditate daily have a lower rate of high blood pressure than men who do not meditate. However, because this is an observational study, we can only conclude that meditation is associated with lower blood pressure — we cannot conclude that meditation causes a reduction in blood pressure.

Question

People with acrophobia (fear of heights) sometimes enroll in therapy sessions to help them overcome this fear. Typically, seven or eight therapy sessions are needed before improvement is noticed. A study was conducted to determine whether the drug D-cycloserine, used in combination with fewer therapy sessions, would help people with acrophobia overcome this fear. Each of 27 people who participated in the study received a pill before each of two therapy sessions. Seventeen of the 27 people were randomly assigned to receive a D-cycloserine pill, and the remaining 10 people received a placebo. After the two therapy sessions, none of the 27 people received additional pills or therapy. Three months after the administration of the pills and the two therapy sessions, each of the 27 people was evaluated to see if he or she had improved.
 
(a) Was this study an experiment or an observational study? Provide an explanation to support your answer.
(b) When the data were analyzed, the D-cycloserine group showed statistically significantly more improvement than the placebo group did. Based on this result, would the researchers be justified in concluding that the D-cycloserine pill and two therapy sessions are as beneficial as eight therapy sessions without the pill? Justify your answer.
(c) A newspaper article that summarized the results of this study did not explain how it was determined which people received D-cycloserine and which received the placebo. Suppose the researchers allowed the therapists to choose which people received D-cycloserine and which received the placebo, and no randomization was used. Explain why such a method of assignment might lead to an incorrect conclusion.

Most-appropriate topic codes (AP Statistics):

• Topic 1.13 — Experimental Design (Part a)
• Topic 1.13 — Experimental Design (Part b)
• Topic 1.13 — Experimental Design (Part c)
▶️ Answer/Explanation

(a)
This study was an experiment.
An experiment is defined by the active imposition of treatments on subjects by the researchers. In this case, the researchers randomly assigned the 27 participants to receive either the active treatment (the D-cycloserine pill) or the control treatment (the placebo).
If it were an observational study, the researchers would have merely observed pre-existing behaviors or treatment choices without intervening or assigning treatments.

(b)
No, the researchers would not be justified in drawing that conclusion.
The experiment only compared two treatment groups: two therapy sessions with D-cycloserine and two therapy sessions with a placebo. There was no treatment group in the study that received eight therapy sessions without the pill.
Because the study did not include a group receiving eight therapy sessions, it is impossible to make a direct, statistically valid comparison between the D-cycloserine + two sessions group and the standard eight sessions approach based solely on this study’s data.

(c)
Allowing therapists to choose the treatment assignments instead of using random assignment introduces selection bias and confounding variables.
If therapists are allowed to choose, they might assign the active D-cycloserine pill to patients they believe need it most (e.g., patients with more severe acrophobia) or to those they feel are most likely to respond positively. Conversely, they might assign the placebo to less severe cases.
This means the two groups would no longer be comparable at the start of the study. If a difference in improvement is observed at the end of the study, it would be impossible to determine whether the difference was caused by the D-cycloserine pill or by the initial, systematic differences in the severity of the patients’ acrophobia (the confounding variable).

Question

Grass buffer strips are grassy areas that are planted between bodies of water and agricultural fields. These strips are designed to filter out sediment, organic material, nutrients, and chemicals carried in runoff water. The figure below shows a cross-sectional view of a grass buffer strip that has been planted along the side of a stream.
A study in Nebraska investigated the use of buffer strips of several widths between 5 feet and 15 feet. The study results indicated a linear relationship between the width of the grass strip (\(x\)), in feet, and the amount of nitrogen removed from the runoff water (\(y\)), in parts per hundred. The following model was estimated.
\(\hat{y} = 33.8 + 3.6x\)
(a) Interpret the slope of the regression line in the context of this question.
(b) Would you be willing to use this model to predict the amount of nitrogen removed for grass buffer strips with widths between 0 feet and 30 feet? Explain why or why not.
A scientist in California wants to know if there is a similar relationship in her area. To investigate this, she will place a grass buffer strip between a field and a nearby stream at each of eight different locations and measure the amount of nitrogen that the grass buffer strip removes, in parts per hundred, from runoff water at each location. Each of the eight locations can accommodate a buffer strip between 6 feet and 13 feet in width. The scientist wants to investigate which combination of widths will provide the best estimate of the slope of the regression line.
Suppose the scientist decides to use buffer strips of width 6 feet at each of four locations and buffer strips of width 13 feet at each of the other four locations. Assume the model, \(\hat{y} = 33.8 + 3.6x\), estimated from the Nebraska study is the true regression line in California and the observations at the different locations are normally distributed with standard deviation of 5 parts per hundred.
(c) Describe the sampling distribution of the sample mean of the observations on the amount of nitrogen removed by the four buffer strips with widths of 6 feet.
(d) Using your result from part (c), show how to construct an interval that has probability 0.95 of containing the sample mean of the observations from four buffer strips with widths of 6 feet.
For the study plan being implemented by the scientist in California, the graph on the left below displays intervals that each have probability 0.95 of containing the sample mean of the four observations for buffer strips of width 6 feet and for buffer strips of width 13 feet. A second possible study plan would use buffer strips of width 8 feet at four of the eight locations and buffer strips of width 10 feet at the other four locations. Intervals that each have probability 0.95 of containing the mean of the four observations for buffer strips of width 8 feet and for buffer strips of width 10 feet, respectively, are shown in the graph on the right below.
If data are collected for the first study plan, a sample mean will be computed for the four observations from buffer strips of width 6 feet and a second sample mean will be computed for the four observations from buffer strips of width 13 feet. The estimated regression line for those eight observations will pass through the two sample means. If data are collected for the second study plan, a similar method will be used.
(e) Use the plots above to determine which study plan, the first or the second, would provide a better estimator of the slope of the regression line. Explain your reasoning.
(f) The previous parts of this question used the assumption of a straight-line relationship between the width of the buffer strip and the amount of nitrogen that is removed, in parts per hundred. Although this assumption was motivated by prior experience, it may not be correct. Describe another way of choosing the widths of the buffer strips at eight locations that would enable the researchers to check the assumption of a straight-line relationship.

Most-appropriate topic codes (AP Statistics):

• Topic 5.3 — Linear Regression Models (Parts a, b)
• Topic 4.1 — Sampling Distributions for Sample Means (Parts c, d)
• Topic 5.5 — Least-Squares Regression (Part e)
• Topic 1.13 — Experimental Design (Part f)
▶️ Answer/Explanation

(a)

The slope of the regression line is \(3.6\).
This means that for each additional foot added to the width of the grass buffer strip, the amount of nitrogen removed from the runoff water increases by approximately \(3.6\) parts per hundred, on average.

(b)

No — this model should not be used for widths between 0 and 30 feet.
The Nebraska study only investigated buffer strips with widths between 5 feet and 15 feet, so the linear relationship was established only within that range.
Predicting for widths as small as 0 feet or as large as 30 feet would be extrapolation far beyond the data, making such predictions unreliable and potentially meaningless.

(c)

When the buffer strip width is \(x = 6\) feet, the true mean nitrogen removed is predicted by the model as:
\(\mu = 33.8 + 3.6(6) = 33.8 + 21.6 = 55.4 \text{ parts per hundred}\)
Since individual observations are normally distributed with standard deviation \(\sigma = 5\), the sampling distribution of the sample mean \(\bar{x}\) of four observations is normal with:
\(\mu_{\bar{x}} = 55.4 \text{ parts per hundred}\)
\(\sigma_{\bar{x}} = \frac{\sigma}{\sqrt{n}} = \frac{5}{\sqrt{4}} = 2.5 \text{ parts per hundred}\)
So the sampling distribution is \(N(55.4,\ 2.5)\).

(d)

Since the sampling distribution of \(\bar{x}\) is normal, use the critical value \(z^* = 1.96\) for a probability of 0.95.
The interval is constructed as:
\(\mu_{\bar{x}} \pm z^* \cdot \sigma_{\bar{x}} = 55.4 \pm 1.96 \times 2.5 = 55.4 \pm 4.9\)
\(\Rightarrow \left(55.4 – 4.9,\ \ 55.4 + 4.9\right) = (50.5,\ \ 60.3)\)
There is probability 0.95 that the sample mean of four 6-foot buffer strip observations falls between \(50.5\) and \(60.3\) parts per hundred.

(e)

The first study plan (widths 6 feet and 13 feet) provides the better estimator of the slope.
The estimated regression line must pass through the two sample means, so any variation in those sample means produces variation in the estimated slope.
In Study Plan 1, the two \(x\)-values (\(6\) ft and \(13\) ft) are spread far apart; even with vertical spread in the 0.95-probability intervals, the range of possible connecting slopes is relatively narrow — as seen in the left graph.
In Study Plan 2, the two \(x\)-values (\(8\) ft and \(10\) ft) are close together; the same vertical spread in the intervals produces a much wider range of possible slopes — as seen in the right graph.
Therefore, the sampling variability of the estimated slope \(\hat{b}\) is smaller under Study Plan 1, making it the better estimator of the true slope.

(f)

To check the linearity assumption, the researcher should use buffer strips of more than two different widths spread across the entire range of interest (6 to 13 feet).
For example, she could use eight different widths — one at each location — such as 6, 7, 8, 9, 10, 11, 12, and 13 feet.
With data at many distinct widths, a scatterplot of nitrogen removed versus strip width would reveal whether the relationship follows a straight line or shows curvature, directly allowing the researchers to assess the straight-line assumption.

Question

Agricultural experts are trying to develop a bird deterrent to reduce costly damage to crops in the United States. An experiment is to be conducted using garlic oil to study its effectiveness as a nontoxic, environmentally safe bird repellant. The experiment will use European starlings, a bird species that causes considerable damage annually to the corn crop in the United States. Food granules made from corn are to be infused with garlic oil in each of five concentrations of garlic—0 percent, 2 percent, 10 percent, 25 percent, and 50 percent. The researchers will determine the adverse reaction of the birds to the repellant by measuring the number of food granules consumed during a two-hour period following overnight food deprivation. There are forty birds available for the experiment, and the researchers will use eight birds for each concentration of garlic. Each bird will be kept in a separate cage and provided with the same number of food granules.
(a) For the experiment, identify
i. the treatments
ii. the experimental units
iii. the response that will be measured
(b) After performing the experiment, the researchers recorded the data shown in the table below.
i. Construct a graph of the data that could be used to investigate the appropriateness of a linear regression model for analyzing the results of the experiment.

ii. Based on your graph, do you think a linear regression model is appropriate? Explain.

Most-appropriate topic codes (AP Statistics):

• Topic 1.13 — Experimental Design (Part a)
• Topic 5.1 — Graphical Representations Between Two Quantitative Variables (Part b.i)
• Topic 5.3 — Linear Regression Models (Part b.ii)
▶️ Answer/Explanation

(a)

i. The treatments: The five different concentrations of garlic oil infused into the food granules, which are 0%, 2%, 10%, 25%, and 50%.
ii. The experimental units: The 40 individual European starlings (or the individual cages housing each bird).
iii. The response that will be measured: The number of food granules consumed by an individual bird during the two-hour period.

(b)
i. Graph construction:
To investigate linearity, we construct a scatterplot with Garlic Oil Concentration (%) on the horizontal $x$-axis and Mean Number of Food Granules Consumed on the vertical $y$-axis.

The points to plot are $(0, 58)$, $(2, 48)$, $(10, 29)$, $(25, 24)$, and $(50, 20)$.
ii. Linearity Assessment:
No, a linear regression model is not appropriate.
The scatterplot reveals a clear, distinct curved pattern rather than a straight line trend.
As the garlic oil concentration increases, the mean number of food granules consumed drops sharply at first and then begins to level off, indicating a non-linear relationship.

Question

Before beginning a unit on frog anatomy, a seventh-grade biology teacher gives each of the 24 students in the class a pretest to assess their knowledge of frog anatomy. The teacher wants to compare the effectiveness of an instructional program in which students physically dissect frogs with the effectiveness of a different program in which students use computer software that only simulates the dissection of a frog. After completing one of the two programs, students will be given a posttest to assess their knowledge of frog anatomy. The teacher will then analyze the changes in the test scores (score on posttest minus score on pretest).
(a) Describe a method for assigning the 24 students to two groups of equal size that allows for a statistically valid comparison of the two instructional programs.
(b) Suppose the teacher decided to allow the students in the class to select which instructional program on frog anatomy (physical dissection or computer simulation) they prefer to take, and 11 students choose actual dissection and 13 students choose computer simulation. How might that self-selection process jeopardize a statistically valid comparison of the changes in the test scores (score on posttest minus score on pretest) for the two instructional programs? Provide a specific example to support your answer.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.13\) — Experimental Design (Part \(\mathrm{a}\))
• Topic \(1.13\) — Experimental Design (Part \(\mathrm{b}\))
• Topic \(1.12\) — Potential Problems with Sampling (Part \(\mathrm{b}\))

▶️ Answer/Explanation

(a) — Completely Randomized Design
Assign each of the 24 students a unique two-digit number from \(01\) to \(24\).
Use a random number table or a random number generator to produce a sequence of two-digit numbers from \(01\) to \(24\), ignoring repeats and any numbers outside that range.
The first 12 distinct numbers that appear correspond to the students assigned to the physical dissection program; the remaining 12 students are assigned to the computer simulation program.
This ensures both groups are of equal size (\(n = 12\) each) and that assignment is governed entirely by chance, removing any systematic differences between the groups before the study begins.

Alternative: Randomized Block Design
Rank all 24 students from lowest to highest pretest score.
Form 12 blocks of 2 students each: Block 1 contains the two students with the lowest pretest scores, Block 2 the next two, and so on, with Block 12 containing the two students with the highest pretest scores.
Within each block, randomly assign one student to the physical dissection program and the other to the computer simulation program — for example, by flipping a fair coin or using a random number generator to decide which student in each pair gets which treatment.
This design controls for prior knowledge of frog anatomy (as measured by the pretest), making the comparison between treatments more precise.

(b)
When students self-select into groups, the two groups may differ systematically in ways that affect posttest performance — completely apart from which instructional method they received.
For example, suppose students who already know a great deal about frog anatomy tend to be the ones who choose the physical dissection program, because they are enthusiastic about frogs and eager to work with them hands-on. Since these students enter with higher prior knowledge, there is less room for them to improve between the pretest and the posttest, so their score changes (posttest \(-\) pretest) will tend to be smaller.
Meanwhile, the students who choose the computer simulation tend to know less about frog anatomy to begin with, so they have more room to improve, and their score changes will tend to be larger.
If the computer simulation group shows a larger average improvement, we cannot tell whether the simulation program itself is more effective or whether the difference simply reflects the fact that lower-knowledge students had more room to grow. The self-selection has introduced a confounding variable (prior knowledge of frog anatomy) that makes it impossible to isolate the effect of the instructional method alone.

Question

A French study was conducted in the 1990s to compare the effectiveness of using an instrument called a cardiopump with the effectiveness of using traditional cardiopulmonary resuscitation (CPR) in saving lives of heart attack victims. Heart attack patients in participating cities were treated with either a cardiopump or CPR, depending on whether the individual’s heart attack occurred on an even-numbered or an odd-numbered day of the month. Before the start of the study, a coin was tossed to determine which treatment, a cardiopump or CPR, was given on the even-numbered days. The other treatment was given on the odd-numbered days. In total, 754 patients were treated with a cardiopump, and 37 survived at least one year; while 746 patients were treated with CPR, and 15 survived at least one year.
(a) The conditions for inference are satisfied in the study. State the conditions and indicate how they are satisfied.
(b) Perform a statistical test to determine whether the survival rate for patients treated with a cardiopump is significantly higher than the survival rate for patients treated with CPR.

Most-appropriate topic codes (AP Statistics):

• Topic \(3.12\) — Setting Up a Test for the Difference Between Two Population Proportions (Part \(\mathrm{b}\): hypotheses and test selection)
• Topic \(3.13\) — Carrying Out a Test for the Difference Between Two Population Proportions (Part \(\mathrm{b}\): test statistic, \(p\)-value, and conclusion)
• Topic \(3.9\) — Sampling Distributions for the Difference Between Sample Proportions (Part \(\mathrm{a}\): large counts condition for inference)
• Topic \(1.13\) — Experimental Design (Part \(\mathrm{a}\): random assignment of treatments)
▶️ Answer/Explanation

Let \(p_A\) = true proportion of patients who survive at least one year if treated with the cardiopump.
Let \(p_B\) = true proportion of patients who survive at least one year if treated with CPR.

(a)
The two conditions required for a two-sample \(z\)-test comparing proportions in an experiment are:
Condition 1 — Random assignment of treatments: Before the study began, a coin was tossed to determine which treatment was assigned to even-numbered days and which to odd-numbered days. This coin toss serves as a reasonable approximation to randomly assigning the two treatments to the available subjects, so this condition is satisfied.
Condition 2 — Sufficiently large sample sizes: All four counts (successes and failures for each group) must be at least 5. Checking:
\(n_A \hat{p}_A = 37 \geq 5, \quad n_A(1-\hat{p}_A) = 754 – 37 = 717 \geq 5\)
\(n_B \hat{p}_B = 15 \geq 5, \quad n_B(1-\hat{p}_B) = 746 – 15 = 731 \geq 5\)
All four values are well above 5, so the large sample condition is satisfied.

(b)

Step 1 — Hypotheses:
\(H_0: p_A = p_B \quad \text{(or } p_A – p_B = 0\text{)}\)
\(H_a: p_A > p_B \quad \text{(or } p_A – p_B > 0\text{)}\)

Step 2 — Test: Two-sample \(z\)-test for proportions (one-sided).

Step 3 — Compute the test statistic:
First, compute the pooled sample proportion:
\(\hat{p} = \frac{n_A \hat{p}_A + n_B \hat{p}_B}{n_A + n_B} = \frac{37 + 15}{754 + 746} = \frac{52}{1500} \approx 0.0347\)
Now compute the \(z\)-statistic:
\(z = \frac{\hat{p}_A – \hat{p}_B}{\sqrt{\hat{p}(1-\hat{p})\left(\dfrac{1}{n_A} + \dfrac{1}{n_B}\right)}}\)
\(z = \frac{\dfrac{37}{754} – \dfrac{15}{746}}{\sqrt{(0.0347)(1 – 0.0347)\left(\dfrac{1}{754} + \dfrac{1}{746}\right)}} \approx 3.066\)
The corresponding one-sided \(p\)-value is:
\(p\text{-value} = P(Z > 3.066) \approx 0.0011\)

Step 4 — Conclusion:
Since the \(p\)-value of \(0.0011\) is much less than any reasonable significance level (such as \(\alpha = 0.05\) or \(\alpha = 0.01\)), we reject \(H_0\).
There is strong statistical evidence that the proportion of patients who survive at least one year is higher when treated with the cardiopump than when treated with CPR — that is, the cardiopump has a significantly higher survival rate than CPR.

Question

A manufacturer of toxic pesticide granules plans to use a dye to color the pesticide so that birds will avoid eating it. A series of experiments will be designed to find colors or patterns that three bird species (blackbirds, starlings, and geese) will avoid eating. Representative samples of birds will be captured to use in the experiments, and the response variable will be the amount of time a hungry bird will avoid eating food of a particular color or pattern.
(a) Previous research has shown that male birds do not avoid solid colors. However, it is possible that males might avoid colors displayed in a pattern, such as stripes. In an effort to prevent males from eating the pesticide, the following two treatments are applied to the pesticide granules.
Treatment 1: A red background with narrow blue stripes
Treatment 2: A blue background with narrow red stripes
To increase the power of detecting a difference in the two treatments in the analysis of the experiment, the researcher decided to block on the three species of birds (blackbirds, starlings, and geese). Assuming there are 100 birds of each of the three species, explain how you would assign birds to treatments in such a block design.
(b) Other than blocking, what could the researcher do to increase the power of detecting a difference in the two treatments in the analysis of the experiment? Explain how your approach would increase the power.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.13\) — Experimental Design (Part \(\mathrm{a}\): randomized block design and random assignment)
• Topic \(1.13\) — Experimental Design (Part \(\mathrm{a}\): blocking on bird species to reduce variability)
• Topic \(4.8\) — Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Means (Part \(\mathrm{b}\): power and ability to detect treatment differences)
• Topic \(4.10\) — Carrying Out a Test for the Difference Between Two Population Means (Part \(\mathrm{b}\): factors affecting power, including sample size and significance level)
▶️ Answer/Explanation

(a)

Form three blocks based on species: Block 1 = blackbirds, Block 2 = starlings, Block 3 = geese. Since there are 100 birds of each species, each block contains 100 birds.
Within each block, randomly assign the 100 birds to the two treatments as follows:
Label each bird in the block with a unique number from 00 to 99. Use a random number table, calculator, or statistical software to generate a list of 50 distinct two-digit numbers between 00 and 99. The birds whose labels match these 50 numbers are assigned Treatment 1 (red background with narrow blue stripes). The remaining 50 birds in the block are assigned Treatment 2 (blue background with narrow red stripes).
Repeat this exact randomization procedure independently within Block 2 (starlings) and Block 3 (geese).
This results in 50 birds per treatment within each species block, and the random assignment ensures that any differences observed between treatments are not due to systematic differences among the birds.

(b)

One effective way to increase the power of the test (other than blocking) is to increase the sample size.
Increasing the number of birds in the study reduces the standard error of the sampling distribution of the difference in sample means.
A smaller standard error means the test statistic will be larger for any given true difference between treatments, making it more likely that the test will detect a real difference if one exists — that is, the power of the test increases.
Another valid approach is to increase the significance level \(\alpha\) (for example, from \(\alpha = 0.01\) to \(\alpha = 0.05\)). Raising \(\alpha\) makes it easier to reject a false null hypothesis, which lowers the probability of a Type II error (\(\beta\)), and since power \(= 1 – \beta\), the power increases.

Question

Two treatments, A and B, showed promise for treating a potentially fatal disease. A randomized experiment was conducted to determine whether there is a significant difference in the survival rate between patients who receive treatment A and those who receive treatment B. Of 154 patients who received treatment A, 38 survived for at least 15 years, whereas 16 of the 164 patients who received treatment B survived at least 15 years.
(a) Treatment A can be administered only as a pill, and treatment B can be administered only as an injection. Can this randomized experiment be performed as a double-blind experiment? Why or why not?
(b) The conditions for inference have been met. Construct and interpret a 95 percent confidence interval for the difference between the proportion of the population who would survive at least 15 years if given treatment A and the proportion of the population who would survive at least 15 years if given treatment B.
In many of these types of studies, physicians are interested in the ratio of survival probabilities, \(\dfrac{p_A}{p_B}\), where \(p_A\) represents the true 15-year survival rate for all patients who receive treatment A and \(p_B\) represents the true 15-year survival rate for all patients who receive treatment B. This ratio is usually referred to as the relative risk of the two treatments.
For example, a relative risk of 1 indicates the survival rates for patients receiving the two treatments are equal, whereas a relative risk of 1.5 indicates that the survival rate for patients receiving treatment A is 50 percent higher than the survival rate for patients receiving treatment B. An estimator of the relative risk is the ratio of estimated probabilities, \(\dfrac{\hat{p}_A}{\hat{p}_B}\).
(c) Using the data from the randomized experiment described above, compute the estimate of the relative risk.
The sampling distribution of \(\dfrac{\hat{p}_A}{\hat{p}_B}\) is skewed. However, when both sample sizes \(n_A\) and \(n_B\) are relatively large, the distribution of \(\ln\!\left(\dfrac{\hat{p}_A}{\hat{p}_B}\right)\) — the natural logarithm of relative risk — is approximately normal with a mean of \(\ln\!\left(\dfrac{p_A}{p_B}\right)\) and a standard deviation of \(\sqrt{\dfrac{1-p_A}{n_A p_A}+\dfrac{1-p_B}{n_B p_B}}\), where \(p_A\) and \(p_B\) can be estimated by using \(\hat{p}_A\) and \(\hat{p}_B\).
When a 95 percent confidence interval for \(\ln\!\left(\dfrac{p_A}{p_B}\right)\) is known, an approximate 95 percent confidence interval for \(\dfrac{p_A}{p_B}\) — the relative risk of the two treatments — can be constructed by applying the inverse of the natural logarithm to the endpoints of the confidence interval for \(\ln\!\left(\dfrac{p_A}{p_B}\right)\).
(d) The conditions for inference are met for the data in the experiment above, and a 95 percent confidence interval for \(\ln\!\left(\dfrac{p_A}{p_B}\right)\) is \((0.3868,\ 1.4690)\). Construct and interpret a 95 percent confidence interval for the relative risk, \(\dfrac{p_A}{p_B}\), of the two treatments.
(e) What is an advantage of using the interval in part (d) over using the interval in part (b)?

Most-appropriate topic codes (AP Statistics):

• Topic \(1.13\) — Experimental Design (Part \(\mathrm{a}\): double-blind experiments and random assignment)
• Topic \(3.10\) — Constructing a Confidence Interval for the Difference Between Two Population Proportions (Part \(\mathrm{b}\))
• Topic \(3.11\) — Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Proportions (Parts \(\mathrm{b}\), \(\mathrm{e}\))
• Topic \(3.1\) — Estimators (Part \(\mathrm{c}\): estimating relative risk using \(\hat{p}_A/\hat{p}_B\))
• Topic \(3.11\) — Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Proportions (Part \(\mathrm{d}\): interpreting the confidence interval for relative risk)
▶️ Answer/Explanation

(a)

Yes, this experiment can be performed as a double-blind experiment by introducing placebos for each treatment group.
Patients assigned to treatment A (pill) would also receive a placebo injection. Patients assigned to treatment B (injection) would also receive a placebo pill. This way, every patient receives both a pill and an injection, but one of the two is a placebo.
Since neither the patients nor the physicians administering the treatments know which is the real treatment and which is the placebo, neither group is aware of the treatment assignment — making the experiment double-blind.

(b)

First, compute the sample proportions:
\(\hat{p}_A = \frac{38}{154} \approx 0.2468, \qquad \hat{p}_B = \frac{16}{164} \approx 0.0976\)
The 95% confidence interval for \(p_A – p_B\) is:
\(\hat{p}_A – \hat{p}_B) \pm z^* \sqrt{\frac{\hat{p}_A(1-\hat{p}_A)}{n_A} + \frac{\hat{p}_B(1-\hat{p}_B)}{n_B}}\)
\(0.2468 – 0.0976) \pm 1.96\sqrt{\frac{(0.2468)(0.7532)}{154} + \frac{(0.0976)(0.9024)}{164}}\)
\(0.1492 \pm 1.96(0.0418)\)
\(0.1492 \pm 0.0818\)
\(\boxed{(0.0674,\ 0.2310)}\)
We are 95% confident that the true difference in 15-year survival rates \((p_A – p_B)\) is between \(0.0674\) and \(0.2310\). Because the entire interval lies above zero, this provides evidence that treatment A has a higher 15-year survival rate than treatment B.

(c)

The estimated relative risk is:
\(\frac{\hat{p}_A}{\hat{p}_B} = \frac{38/154}{16/164} = \frac{0.2468}{0.0976} \approx 2.53\)
\(\boxed{\text{Estimated relative risk} \approx 2.53}\)
This means patients receiving treatment A are estimated to be about 2.53 times as likely to survive at least 15 years as patients receiving treatment B.

(d)

A 95% confidence interval for \(\ln\!\left(\dfrac{p_A}{p_B}\right)\) is given as \((0.3868,\ 1.4690)\).
To convert this to a confidence interval for the relative risk \(\dfrac{p_A}{p_B}\), apply the exponential (inverse of the natural log) to each endpoint:
\(e^{0.3868} \approx 1.47 \qquad \text{and} \qquad e^{1.4690} \approx 4.34\)
\(\boxed{\left(1.47,\ 4.34\right)}\)
We are 95% confident that patients receiving treatment A are between 1.47 and 4.34 times as likely to survive at least 15 years compared to patients receiving treatment B.

(e)

When the survival proportions are small (as here, approximately 0.25 and 0.10), the confidence interval for the relative risk is more informative and practically meaningful than the confidence interval for the difference in proportions.
Knowing that a patient’s chance of survival is between 1.47 and 4.34 times greater with treatment A is more vivid and easier to interpret clinically than knowing the absolute difference in proportions is somewhere between 0.07 and 0.23 — a range that may sound small even though it represents a substantial relative advantage.

Question

A researcher wants to conduct a study to test whether listening to soothing music for 20 minutes helps to reduce diastolic blood pressure in patients with high blood pressure, compared to simply sitting quietly in a noise-free environment for 20 minutes. One hundred patients with high blood pressure at a large medical clinic are available to participate in this study.
(a) Propose a design for this study to compare these two treatments.
(b) The null hypothesis for this study is that there is no difference in the mean reduction of diastolic blood pressure for the two treatments and the alternative hypothesis is that the mean reduction in diastolic blood pressure is greater for the music treatment. If the null hypothesis is rejected, the clinic will offer this music therapy as a free service to their patients with high blood pressure. Describe Type I and Type II errors and the consequences of each in the context of this study, and discuss which one you think is more serious.

Most-appropriate topic codes (AP Statistics):

• Topic 1.13 — Experimental Design (Part \(\mathrm{a}\))
• Topic 4.4 — Setting Up a Test for a Population Mean or Population Mean Difference (Part \(\mathrm{b}\))
• Topic 3.8 — Potential Errors When Performing Tests (Part \(\mathrm{b}\))
▶️ Answer/Explanation

(a)
Proposed Design (Completely Randomized Design):
Assign each of the 100 patients a unique number from 00 to 99. Using a random number table (or random number generator), select 50 unique numbers. The patients corresponding to the selected numbers will form Group 1 (Music Treatment); the remaining 50 patients will form Group 2 (Control — Noise-Free Environment).
Measure the diastolic blood pressure of every patient before the treatment begins. Then:
Group 1 patients sit quietly in a room where soothing music is played for 20 minutes.
Group 2 patients sit quietly in a noise-free environment for 20 minutes.
At the end of the 20-minute period, measure the diastolic blood pressure of every patient again. Compute the reduction in diastolic blood pressure (before \(-\) after) for each patient. Run a two-sample \(t\)-test to compare the mean reduction in the two groups.

Alternatively (Paired Design):
Each patient receives both treatments on two separate occasions, with a suitable washout period in between. For each patient, flip a coin to randomly assign which treatment is administered first. Compute the difference (music reduction \(-\) noise-free reduction) for each patient, and use a paired \(t\)-test to determine whether the mean difference is significantly greater than zero.

(b)
Let \(\mu_M\) = mean reduction in diastolic blood pressure under the music treatment, and \(\mu_C\) = mean reduction under the control (noise-free) treatment. The hypotheses are:
\(H_0: \mu_M = \mu_C \qquad H_a: \mu_M > \mu_C\)
Type I Error:
Rejecting \(H_0\) when it is actually true — that is, concluding that soothing music does reduce diastolic blood pressure more than sitting quietly, when in reality it does not.
Consequence: The clinic will offer music therapy as a free service to its patients even though the therapy is ineffective. This wastes clinic resources and money, providing no real medical benefit to patients.
Type II Error:
Failing to reject \(H_0\) when it is actually false — that is, concluding that soothing music does not reduce diastolic blood pressure more than sitting quietly, when in reality it does.
Consequence: The clinic will not offer music therapy, even though the therapy would genuinely help patients reduce their blood pressure. Patients are denied access to an effective, low-cost treatment.

Which error is more serious?
A reasonable case can be made for either error. One well-supported argument is that the Type II error is more serious: denying patients an effective treatment that could meaningfully improve their health — and potentially save lives — is a greater harm than wasting some clinic resources on an ineffective service. In the case of a Type I error, patients who receive music therapy are unlikely to be harmed by listening to music. But in the case of a Type II error, patients who could benefit from a real treatment are denied it altogether.

Question

As dogs age, diminished joint and hip health may lead to joint pain and thus reduce a dog’s activity level. Such a reduction in activity can lead to other health concerns such as weight gain and lethargy due to lack of exercise. A study is to be conducted to see which of two dietary supplements, glucosamine or chondroitin, is more effective in promoting joint and hip health and reducing the onset of canine osteoarthritis. Researchers will randomly select a total of 300 dogs from ten different large veterinary practices around the country. All of the dogs are more than 6 years old, and their owners have given consent to participate in the study. Changes in joint and hip health will be evaluated after 6 months of treatment.
(a) What would be an advantage to adding a control group in the design of this study?
(b) Assuming a control group is added to the other two groups in the study, explain how you would assign the 300 dogs to these three groups for a completely randomized design.
(c) Rather than using a completely randomized design, one group of researchers proposes blocking on clinics, and another group of researchers proposes blocking on breed of dog. How would you decide which one of these two variables to use as a blocking variable?

Most-appropriate topic codes (AP Statistics):

• Topic 1.13 — Experimental Design (Parts \(\mathrm{a}\), \(\mathrm{b}\), \(\mathrm{c}\))
• Topic 1.10 — The Investigative Question Revisited and Data Collection (Part \(\mathrm{a}\))
• Topic 1.11 — Random Sampling (Part \(\mathrm{b}\))
▶️ Answer/Explanation

(a)

A control group gives the researchers a baseline comparison group — without any dietary supplement — so they can measure whether glucosamine or chondroitin actually makes a difference beyond what would happen due to the normal aging process alone.
Without a control group, we could not tell whether any improvements in joint and hip health were caused by the supplements or simply by other factors such as the passage of time, veterinary care, or natural variation between dogs.
In this study specifically, the control group allows us to isolate the true effect of glucosamine and chondroitin on reducing canine osteoarthritis by comparing both treatment groups against untreated dogs under the same conditions.
\(\boxed{\text{Control group provides a baseline to determine whether the supplements are truly effective.}}\)

(b)

First, assign each of the 300 dogs a unique number from \(001\) to \(300\).
Then, use a random number generator (calculator, statistical software, or a random number table) to randomly select 100 numbers from \(001\) to \(300\), ignoring any repeats — the dogs corresponding to these 100 numbers are assigned to the glucosamine group.
From the remaining 200 dogs, randomly select another 100 numbers using the same process — these dogs are assigned to the chondroitin group.
The final 100 remaining dogs are assigned to the control group and receive no dietary supplement.
\(\boxed{\text{Randomly assign dogs numbered } 001\text{–}300 \text{ into three equal groups of 100 using a random number generator.}}\)

(c)

The blocking variable should be the one that has a stronger association with the response variable — joint and hip health — so that dogs within each block are as similar (homogeneous) as possible.
Breed of dog is associated with the size of the dog, and size is known to be related to joint and hip health — larger breeds tend to have more joint problems than smaller breeds, so breed is likely to create more variability in the response.
Clinic, on the other hand, is less likely to be strongly associated with joint and hip health because most large veterinary practices see a wide variety of dog breeds and sizes, so dogs across clinics would not be meaningfully more similar to each other than dogs across breeds.
Therefore, we should block on breed of dog, since it is more strongly related to joint and hip health and will reduce variability more effectively within each block.
\(\boxed{\text{Block on breed of dog, as it has a stronger relationship to joint and hip health than clinic.}}\)

Question

Researchers want to determine whether drivers are significantly more distracted while driving when using a cell phone than when talking to a passenger in the car. In a study involving 48 people, 24 people were randomly assigned to drive in a driving simulator while using a cell phone. The remaining 24 were assigned to drive in the driving simulator while talking to a passenger in the simulator. Part of the driving simulation for both groups involved asking drivers to exit the freeway at a particular exit. In the study, 7 of the 24 cell phone users missed the exit, while 2 of the 24 talking to a passenger missed the exit.
(a) Would this study be classified as an experiment or an observational study? Provide an explanation to support your answer.
(b) State the null and alternative hypotheses of interest to the researchers.
(c) One test of significance that you might consider using to answer the researchers’ question is a two-sample \(z\)-test. State the conditions required for this test to be appropriate. Then comment on whether each condition is met.
(d) Using an advanced statistical method for small samples to test the hypotheses in part (b), the researchers report a \(p\)-value of 0.0683. Interpret, in everyday language, what this \(p\)-value measures in the context of this study and state what conclusion should be made based on this \(p\)-value.

Most-appropriate topic codes (AP Statistics):

• Topic 1.13 — Experimental Design (Part \(\mathrm{a}\))
• Topic 3.12 — Setting Up a Test for the Difference Between Two Population Proportions (Part \(\mathrm{b}\))
• Topic 3.13 — Carrying Out a Test for the Difference Between Two Population Proportions (Part \(\mathrm{c}\))
• Topic 3.6 — p-Values (Part \(\mathrm{d}\))
• Topic 3.13 — Carrying Out a Test for the Difference Between Two Population Proportions (Part \(\mathrm{d}\))
▶️ Answer/Explanation

(a)

This study is classified as an experiment, not an observational study.
The key reason is that the researchers actively imposed treatments on the subjects — they randomly assigned drivers to one of two conditions: driving while using a cell phone, or driving while talking to a passenger.
In an observational study, researchers simply observe subjects without intervening; here, the environment was deliberately controlled and manipulated, which is the defining feature of an experiment.
\(\boxed{\text{Experiment — treatments (cell phone vs. passenger) were actively imposed on randomly assigned subjects.}}\)

(b)

Let \(p_{\text{cell}}\) = the population proportion of drivers who miss the exit while talking on a cell phone.
Let \(p_{\text{pass}}\) = the population proportion of drivers who miss the exit while talking to a passenger.
\(H_0: p_{\text{cell}} = p_{\text{pass}}\) (no difference in the proportion of drivers who miss the exit between the two groups)
\(H_a: p_{\text{cell}} > p_{\text{pass}}\) (a greater proportion of cell phone users miss the exit compared to passenger talkers)
\(\boxed{H_0: p_{\text{cell}} = p_{\text{pass}}, \quad H_a: p_{\text{cell}} > p_{\text{pass}}}\)

(c)

The two conditions required for a two-sample \(z\)-test for proportions are:
Condition 1 — Independent random samples or random assignment: The problem states that the 48 participants were randomly assigned to the two groups, so this condition is met.
Condition 2 — Large sample sizes (each of \(n_1\hat{p}_1\), \(n_1(1-\hat{p}_1)\), \(n_2\hat{p}_2\), \(n_2(1-\hat{p}_2)\) must be \(\geq 10\)):
For the cell phone group: \(n_1\hat{p}_1 = 24\cdot\dfrac{7}{24} = 7 < 10\)
For the passenger group: \(n_2\hat{p}_2 = 24\cdot\dfrac{2}{24} = 2 < 10\)
Both groups fail the large sample condition, so this condition is not met. A two-sample \(z\)-test is therefore not appropriate for this situation.
\(\boxed{\text{Condition 1 met (random assignment); Condition 2 NOT met } (n\hat{p} < 10 \text{ for both groups).}}\)

(d)

In everyday language, the \(p\)-value of \(0.0683\) means: assuming that cell phone use and talking to a passenger are equally distracting (i.e., \(H_0\) is true), there is about a \(6.83\%\) chance of observing a difference in missed-exit proportions as large as or larger than what was seen in this study just by random chance alone.
Since the \(p\)-value of \(0.0683\) is greater than \(\alpha = 0.05\), we fail to reject \(H_0\). We do not have statistically significant evidence at the 5% level to conclude that drivers using a cell phone are more distracted (as measured by missing the freeway exit) than drivers talking to a passenger.
\(\boxed{p\text{-value} = 0.0683 > 0.05 \Rightarrow \text{Fail to reject } H_0; \text{ insufficient evidence that cell phone use causes more distraction.}}\)

Question

A biologist is interested in studying the effect of growth-enhancing nutrients and different salinity (salt) levels in water on the growth of shrimps. The biologist has ordered a large shipment of young tiger shrimps from a supply house for use in the study. The experiment is to be conducted in a laboratory where 10 tiger shrimps are placed randomly into each of 12 similar tanks in a controlled environment. The biologist is planning to use 3 different growth-enhancing nutrients (A, B, and C) and two different salinity levels (low and high).
(a) List the treatments that the biologist plans to use in this experiment.
(b) Using the treatments listed in part (a), describe a completely randomized design that will allow the biologist to compare the shrimps’ growth after 3 weeks.
(c) Give one statistical advantage to having only tiger shrimps in the experiment. Explain why this is an advantage.
(d) Give one statistical disadvantage to having only tiger shrimps in the experiment. Explain why this is a disadvantage.

Most-appropriate topic codes (AP Statistics):

• Topic 1.13 — Experimental Design (Parts \(\mathrm{a}\), \(\mathrm{b}\))
• Topic 1.13 — Experimental Design (Control of Variability in Experimental Units) (Part \(\mathrm{c}\))
• Topic 1.13 — Experimental Design (Scope of Inference and Generalizability) (Part \(\mathrm{d}\))
▶️ Answer/Explanation

(a)

The biologist has 3 nutrients and 2 salinity levels, giving \(3 \times 2 = 6\) treatment combinations in total. The six treatments are:

(b)

Since 10 tiger shrimps have already been randomly placed into each of 12 similar tanks in a controlled environment, we must randomly assign the treatment combinations to the tanks. Each treatment combination will be randomly assigned to 2 of the 12 tanks. One way to do this is to generate a random number for each tank. The treatment combinations are then assigned by sorting the random numbers from smallest to largest.

After three weeks, record the growth (weight gain = after \(-\) before) for the shrimps in each tank, and compare the average growth across tanks for each of the six treatment combinations.

(c)

Statistical advantage: Reduced variability among experimental units.
Tiger shrimps are more similar to one another in their natural growth characteristics than a mix of different shrimp species would be. By using only tiger shrimps, the biologist eliminates the shrimp species as a source of variation among the experimental units (the tanks). With less background variability in the response (growth), any real differences in growth due to the nutrient or salinity treatments are easier to detect — that is, the experiment has greater power to identify treatment effects that are actually present.

(d)

Statistical disadvantage: Limited scope of inference (reduced generalizability).
Because the experiment uses only tiger shrimps, the conclusions can only be applied to tiger shrimps. Other species of shrimp may respond differently to the same nutrient and salinity treatments, so the biologist cannot generalize the results to shrimps in general. The scope of inference is restricted solely to tiger shrimps, which limits the practical usefulness of the study if the biologist’s broader goal is to understand shrimp growth across species.

Question

When a tractor pulls a plow through an agricultural field, the energy needed to pull that plow is called the draft. The draft is affected by environmental conditions such as soil type, terrain, and moisture.
A study was conducted to determine whether a newly developed hitch would be able to reduce draft compared to the standard hitch. (A hitch is used to connect the plow to the tractor.) Two large plots of land were used in this study. It was randomly determined which plot was to be plowed using the standard hitch. As the tractor plowed that plot, a measurement device on the tractor automatically recorded the draft at 25 randomly selected points in the plot.
After the plot was plowed, the hitch was changed from the standard one to the new one, a process that takes a substantial amount of time. Then the second plot was plowed using the new hitch. Twenty-five measurements of draft were also recorded at randomly selected points in this plot.
(a) What was the response variable in this study?
Identify the treatments.
What were the experimental units?
(b) Given that the goal of the study is to determine whether a newly developed hitch reduces draft compared to the standard hitch, was randomization used properly in this study? Justify your answer.
(c) Given that the goal of the study is to determine whether a newly developed hitch reduces draft compared to the standard hitch, was replication used properly in this study? Justify your answer.
(d) Plot of land is a confounding variable in this experiment. Explain why.

Most-appropriate topic codes (AP Statistics):

• Topic 1.13 — Experimental Design (Part \(\mathrm{a}\): response variable, treatments, and experimental units)
• Topic 1.13 — Experimental Design (Part \(\mathrm{b}\): random assignment of treatments)
• Topic 1.13 — Experimental Design (Part \(\mathrm{c}\): replication)
• Topic 1.13 — Experimental Design (Part \(\mathrm{d}\): confounding variables)
▶️ Answer/Explanation

(a)

Response variable: The response variable is the amount of draft — that is, the energy required to pull the plow through the agricultural field.
Treatments: The two treatments are the standard hitch and the newly developed hitch.
Experimental units: The experimental units are the two large plots of land.
\(\boxed{\text{Response variable: draft} \quad \text{Treatments: standard hitch, new hitch} \quad \text{Experimental units: two plots of land}}\)

(b)
Yes, randomization was used properly in this study. It was randomly determined which of the two plots of land would be plowed using the standard hitch, meaning the two treatments (hitches) were randomly assigned to the two experimental units (plots). This random assignment helps reduce the potential for bias in comparing the two hitches, since neither hitch was systematically favored by being assigned to a particular plot.
\(\boxed{\text{Yes — treatments were randomly assigned to experimental units}}\)

(c)
No, replication was not used properly in this study. Each treatment (hitch type) was applied to only one experimental unit (plot of land). Replication requires that each treatment be applied to more than one experimental unit, so that general patterns can be observed and random variation can be distinguished from true treatment effects. Since each hitch was tested on only a single plot, there is no replication, and it is not possible to draw reliable conclusions about the draft-reducing capability of the new hitch from this study alone.
\(\boxed{\text{No — each hitch was applied to only one plot; proper replication requires multiple plots per treatment}}\)

(d)
Plot of land is a confounding variable because each hitch was used exclusively on one plot of land. The two plots may differ in environmental characteristics such as soil type, terrain, and moisture level, all of which can affect the draft. As a result, if a difference in draft is observed between the two hitches, it is impossible to determine whether that difference is due to the type of hitch or due to the inherent differences between the two plots. The effect of the hitch is therefore confounded with — that is, mixed up with — the effect of the plot of land.
\(\boxed{\text{Hitch effect cannot be separated from plot effect} \implies \text{plot of land is a confounding variable}}\)

Question

Sunshine Farms wants to know whether there is a difference in consumer preference for two new juice products — Citrus Fresh and Tropical Taste. In an initial blind taste test, 8 randomly selected consumers were given unmarked samples of the two juices. The product that each consumer tasted first was randomly decided by the flip of a coin. After tasting the two juices, each consumer was asked to choose which juice he or she preferred, and the results were recorded.
(a) Let \(p\) represent the population proportion of consumers who prefer Citrus Fresh. In terms of \(p\), state the hypotheses that Sunshine Farms is interested in testing.
(b) One might consider using a one-proportion \(z\)-test to test the hypotheses in part (a). Explain why this would not be a reasonable procedure for this sample.
(c) Let \(X\) represent the number of consumers in the sample who prefer Citrus Fresh. Assuming there is no difference in consumer preference, find the probability for each possible value of \(X\). Record the \(x\)-values and the corresponding probabilities in the table below.
(d) When testing the hypotheses in part (a), Sunshine Farms will conclude that there is a consumer preference if too many or too few individuals prefer Citrus Fresh. Based on your probabilities in part (c), is it possible for the significance level (probability of rejecting the null hypothesis when it is true) for this test to be exactly 0.05? Justify your answer.
(e) The preference data for the 8 randomly selected consumers are given in the table below.
Based on these preferences and your previous work, test the hypotheses in part (a).
(f) Sunshine Farms plans to add one of these two new juices — Citrus Fresh or Tropical Taste — to its production schedule. A follow-up study will be conducted to decide which of the two juices to produce. Make one recommendation for the follow-up study that would make it better than the initial study. Provide a statistical justification for your recommendation in the context of the problem.

Most-appropriate topic codes (AP Statistics):

• Topic 3.5 — Setting Up a Test for a Population Proportion (Part \(\mathrm{a}\))
• Topic 3.2 — Sampling Distributions for Sample Proportions (Part \(\mathrm{b}\))
• Topic 2.10 — The Binomial Distribution (Part \(\mathrm{c}\))
• Topic 3.6 — p-Values (Part \(\mathrm{d}\))
• Topic 3.7 — Carrying Out a Test for a Population Proportion (Part \(\mathrm{e}\))
• Topic 1.13 — Experimental Design (Part \(\mathrm{f}\))
▶️ Answer/Explanation

(a)
Let \(p\) be the population proportion of consumers who prefer Citrus Fresh. The hypotheses are:
\(H_0: p = 0.5\)
\(H_a: p \neq 0.5\)
A two-sided alternative is appropriate because Sunshine Farms wants to detect any difference in preference, not just preference for one particular juice.

(b)
The conditions for a one-proportion \(z\)-test require that both \(np\) and \(n(1-p)\) be at least 5 (or 10). Here:
\(np = 8 \times 0.5 = 4 < 5\)
\(n(1-p) = 8 \times 0.5 = 4 < 5\)
Since both values are less than 5, the large-sample normal approximation is not valid, and using a one-proportion \(z\)-test would not be appropriate for a sample of only \(n = 8\).

(c)
Under \(H_0\), \(X \sim \text{Binomial}(n = 8,\ p = 0.5)\). The probabilities are computed using:
\(P(X = x) = \binom{8}{x}(0.5)^x(0.5)^{8-x} = \binom{8}{x}(0.5)^8\)

(d)
No, it is not possible for the significance level to be exactly 0.05. Because \(X\) is a discrete random variable, the tail probabilities can only take specific values — there is no rejection region that gives a type I error probability of exactly 0.05.
The most extreme rejection region \((X = 0 \text{ or } X = 8)\) gives:
\(\alpha = 2 \times 0.00391 = 0.00782 < 0.05\)
The next possible rejection region \((X \leq 1 \text{ or } X \geq 7)\) gives:
\(\alpha = 2 \times (0.00391 + 0.03125) = 2 \times 0.03516 = 0.07031 > 0.05\)
Since no rejection region produces a type I error probability of exactly 0.05, a significance level of exactly 0.05 is not achievable with this test.
\(\boxed{\alpha = 0.05 \text{ is not achievable — the achievable levels jump from } 0.00782 \text{ to } 0.07031}\)

(e)
From the data, 2 out of 8 consumers preferred Citrus Fresh, so \(X = 2\).
Since this is a two-sided test, the \(p\)-value is the probability of observing a result at least as extreme as \(X = 2\) in either tail:
\(p\text{-value} = P(X \leq 2) + P(X \geq 6)\)
\(= 2 \times [P(X=0) + P(X=1) + P(X=2)]\)
\(= 2 \times (0.00391 + 0.03125 + 0.10937)\)
\(= 2 \times 0.14453 = 0.28906\)
Since the \(p\)-value of \(0.289\) is much larger than any reasonable significance level (e.g., \(\alpha = 0.05\) or \(\alpha = 0.07031\)), we fail to reject \(H_0\). There is not statistically significant evidence of a consumer preference between Citrus Fresh and Tropical Taste.
\(\boxed{p\text{-value} \approx 0.289 \implies \text{Fail to reject } H_0; \text{ no significant consumer preference detected}}\)

(f)
The most important recommendation is to increase the number of consumers in the study. With only \(n = 8\) consumers, the test has very low power — even a large true difference in preference (like 75% vs. 25%) may not produce a statistically significant result. Increasing the sample size would reduce the standard error of the estimated proportion \(\hat{p}\), making it easier to detect a real difference, and would allow the use of the large-sample one-proportion \(z\)-test since \(np \geq 5\) and \(n(1-p) \geq 5\) would be satisfied. For example, with \(n = 80\) and \(X = 20\) (same sample proportion of 0.25), the \(z\)-statistic would be approximately:
\(z = \frac{0.25 – 0.5}{\sqrt{\frac{0.5(0.5)}{80}}} \approx -4.47\)
which gives a \(p\)-value near zero, allowing a clear conclusion to be reached.
\(\boxed{\text{Recommendation: Increase sample size to increase power and enable use of the } z\text{-test}}\)

Question

In search of a mosquito repellent that is safer than the ones that are currently on the market, scientists have developed a new compound that is rated as less toxic than the current compound, thus making a repellent that contains this new compound safer for human use. Scientists also believe that a repellent containing the new compound will be more effective than the ones that contain the current compound. To test the effectiveness of the new compound versus that of the current compound, scientists have randomly selected 100 people from a state.
Up to 100 bins, with an equal number of mosquitoes in each bin, are available for use in the study. After a compound is applied to a participant’s forearm, the participant will insert his or her forearm into a bin for 1 minute, and the number of mosquito bites on the arm at the end of that time will be determined.
(a) Suppose this study is to be conducted using a completely randomized design. Describe a randomization process and identify an inference procedure for the study.
(b) Suppose this study is to be conducted using a matched-pairs design. Describe a randomization process and identify an inference procedure for the study.
(c) Which of the designs, the one in part (a) or the one in part (b), is better for testing the effectiveness of the new compound versus that of the current compound? Justify your answer.

Most-appropriate topic codes (AP Statistics):

• Topic 1.13 — Experimental Design (Parts a, b, c)
• Topic 4.4 — Setting Up a Test for a Population Mean or Population Mean Difference (Part b)
• Topic 4.9 — Setting Up a Test for the Difference Between Two Population Means (Part a)
▶️ Answer/Explanation

(a) Completely Randomized Design

Randomization process:
Assign each of the 100 participants a unique random number using a random number generator. Sort the participants from smallest to largest by their assigned random number. The first 50 people on the sorted list are assigned to the new compound group, and the remaining 50 are assigned to the current compound group. To reduce bias, the compounds should be placed in identical, unmarked tubes so that neither participants nor researchers know which compound is being applied (double-blind). Each participant in each group then applies their assigned compound to one forearm and inserts it into a randomly assigned bin for 1 minute, after which the number of mosquito bites is counted.

Inference procedure:
Use a two-sample \(t\)-test (or construct a two-sample confidence interval for the difference in means) to compare the mean number of mosquito bites between the new-compound group and the current-compound group: \[ H_0: \mu_{\text{new}} = \mu_{\text{current}} \quad \text{vs.} \quad H_a: \mu_{\text{new}} < \mu_{\text{current}} \] where \(\mu_{\text{new}}\) and \(\mu_{\text{current}}\) are the mean number of bites under each compound.

(b) Matched-Pairs Design

Randomization process:
Each of the 100 participants serves as their own pair — both compounds are applied, one to each arm. For each participant, flip a coin (or use a random number generator) to decide which arm receives the new compound; the other arm receives the current compound. Each participant then inserts both arms simultaneously into a randomly assigned bin for 1 minute, and the number of bites on each arm is recorded. The compounds should again be placed in identical, unmarked tubes to maintain blinding.

Inference procedure:
Compute the difference in bites for each participant as \(d_i = (\text{bites, new compound}) – (\text{bites, current compound})\). Use a one-sample \(t\)-test on the differences (or a paired confidence interval for the mean difference): \[ H_0: \mu_d = 0 \quad \text{vs.} \quad H_a: \mu_d < 0 \] where \(\mu_d\) is the true mean difference in bites between the new and current compounds.

(c) Which design is better?

The matched-pairs design in part (b) is the better choice.

The key reason is that people naturally vary in how attractive they are to mosquitoes — some individuals get bitten far more than others regardless of which compound is used. In a completely randomized design, this person-to-person variability shows up as noise in the data, making it harder to detect a real difference between the two compounds. The matched-pairs design controls for this source of variability by having each person test both compounds, so individual differences in susceptibility cancel out when computing the within-person difference. This leads to a more precise and more powerful comparison of the two compounds.

Question

A researcher believes that treating seeds with certain additives before planting can enhance the growth of plants. An experiment to investigate this is conducted in a greenhouse. From a large number of Roma tomato seeds, 24 seeds are randomly chosen and 2 are assigned to each of 12 containers. One of the 2 seeds is randomly selected and treated with the additive. The other seed serves as a control. Both seeds are then planted in the same container. The growth, in centimeters, of each of the 24 plants is measured after 30 days. These data were used to generate the partial computer output shown below. Graphical displays indicate that the assumption of normality is not unreasonable.
(a) Construct a confidence interval for the mean difference in growth, in centimeters, of the plants from the untreated and treated seeds. Be sure to interpret this interval.
(b) Based only on the confidence interval in part (a), is there sufficient evidence to conclude that there is a significant mean difference in growth of the plants from untreated seeds and the plants from treated seeds? Justify your conclusion.

Most-appropriate topic codes (AP Statistics):

• Topic 4.2 — Constructing a Confidence Interval for a Population Mean or Population Mean Difference (Part a)
• Topic 4.3 — Justifying a Claim Based on a Confidence Interval for a Population Mean or Population Mean Difference (Part b)
• Topic 1.13 — Experimental Design (Matched-pairs design context)
▶️ Answer/Explanation

(a)

Step 1 — Identify the appropriate procedure:
Since the data consist of paired observations (one treated and one untreated seed per container), we use a one-sample \(t\)-confidence interval for the mean of the differences:
\( \bar{d} \pm t^* \cdot \frac{s_d}{\sqrt{n}} \)
Step 2 — Check conditions:
The 24 seeds were randomly chosen and randomly assigned within each container, so the differences are independent. The problem states that graphical displays indicate normality is not unreasonable, so the condition for using a \(t\)-procedure is satisfied.
Step 3 — Compute the interval:
From the computer output: \(\bar{d} = -2.015\), \(s_d = 1.163\), \(n = 12\).
Degrees of freedom: \(df = n – 1 = 11\).
For a 95% confidence interval, the critical value is \(t^* = 2.201\) (from the \(t\)-table with \(df = 11\)).
\( \bar{d} \pm t^* \cdot \frac{s_d}{\sqrt{n}} = -2.015 \pm 2.201 \times \frac{1.163}{\sqrt{12}} \)
\( = -2.015 \pm 2.201 \times 0.336 \)
\( = -2.015 \pm 0.739 \)
\( \boxed{(-2.754,\ -1.276)} \)

Step 4 — Interpret the interval:
We are 95% confident that the true mean difference in growth (untreated minus treated) is between \(-2.754\) cm and \(-1.276\) cm. In other words, on average, the untreated plants grew between about 1.28 cm and 2.75 cm less than the treated plants.

(b)

Step 1 — State the hypotheses (for reference):
\( H_0: \mu_d = 0 \quad \text{vs.} \quad H_a: \mu_d \neq 0 \)
where \(\mu_d\) is the true mean difference in growth between untreated and treated seeds.

Step 2 — Draw the conclusion:
Yes, there is sufficient evidence of a significant mean difference in growth. The 95% confidence interval \((-2.754,\ -1.276)\) does not contain zero. Since zero — the value that would indicate no difference — falls entirely outside the interval, we can reject \(H_0\) at the \(\alpha = 0.05\) significance level. The data provide convincing statistical evidence that the additive treatment produces greater growth than the control, with treated plants growing meaningfully taller on average.

Question

Researchers who are studying a new shampoo formula plan to compare the condition of hair for people who use the new formula with the condition of hair for people who use the current formula. Twelve volunteers are available to participate in this study. Information on these volunteers (numbered 1 through 12) is shown in the table below.
(a) These researchers want to conduct an experiment involving the two formulas (new and current) of shampoo. They believe that the condition of hair changes with age but not gender. Because researchers want the size of the blocks in an experiment to be equal to the number of treatments, they will use blocks of size 2 in their experiment. Identify the volunteers (by number) that would be included in each of the six blocks and give the criteria you used to form the blocks.
(b) Other researchers believe that hair condition differs with both age and gender. These researchers will also use blocks of size 2 in their experiment. Identify the volunteers (by number) that would be included in each of the six blocks and give the criteria you used to form the blocks.
(c) The researchers in part (b) decide to select three of the six blocks to receive the new formula and to give the other three blocks the current formula. Is this an appropriate way to assign treatments? If so, describe a method for selecting the three blocks to receive the new formula. If not, describe an appropriate method for assigning treatments.

Most-appropriate topic codes (AP Statistics):

• Topic 1.13 — Experimental Design (Parts a, b, c)
• Topic 1.10 — The Investigative Question Revisited and Data Collection (Parts a, b)
▶️ Answer/Explanation

(a)
Pairing consecutive volunteers by age (without regard to gender) gives the six blocks:

Since these researchers believe that the condition of hair changes with age but not gender, the volunteers are sorted from youngest to oldest. The volunteers in the sorted list are paired to form six blocks of size two. More specifically, the youngest two volunteers are placed in the first block. The next two volunteers in the sorted list are placed in the second block. This pairing continues until all six blocks of two are formed, with the oldest two volunteers in the sixth block.

(b)
Since both age and gender are believed to affect hair condition, volunteers must be blocked so that each block contains two people of the same gender and similar age.
First, separate by gender, then sort each group by age and pair consecutively
Males sorted by age: Volunteer 1 (age 21), Volunteer 11 (age 23), Volunteer 9 (age 44), Volunteer 3 (age 47), Volunteer 7 (age 58), Volunteer 6 (age 61)
Females sorted by age: Volunteer 2 (age 20), Volunteer 10 (age 24), Volunteer 8 (age 44), Volunteer 12 (age 46), Volunteer 4 (age 60), Volunteer 5 (age 62)
The six blocks are:

Since these researchers believe that the condition of hair changes with both age and gender, the women are sorted from youngest to oldest and then the men are sorted from youngest to oldest. The women (men) in the sorted list are paired to form the blocks of size two. More specifically, the youngest two women (men) are placed in a block. The next two youngest women (men) are placed in another block. Finally, the oldest two women (men) are placed in another block.

(c)
No, this is not an appropriate method for assigning treatments. Assigning the new formula to three entire blocks and the current formula to three other blocks does not allow a fair within-block comparison, because within each block both members would receive the same formula.
The correct approach in a matched-pairs block design is to randomly assign one person within each block to the new formula and the other to the current formula. This can be done as follows:
For each of the six blocks, flip a fair coin (or use a random number table/generator). If the result is heads (or the random digit is 0–4), assign the lower-numbered volunteer to the new formula and the other to the current formula. If tails (or digit 5–9), reverse the assignment. This ensures that within every block, one person receives the new formula and one receives the current formula, making a valid matched comparison possible.
\(\boxed{\text{Randomly assign one person per block to new formula; other receives current formula}}\)

Question

Because of concerns about employee stress, a large company is conducting a study to compare two programs (tai chi or yoga) that may help employees reduce their stress levels. Tai chi is a \(1{,}200\)-year-old practice, originating in China, that consists of slow, fluid movements. Yoga is a practice, originating in India, that consists of breathing exercises and movements designed to stretch and relax muscles. The company has assembled a group of volunteer employees to participate in the study during the first half of their lunch hour each day for a \(10\)-week period. Each volunteer will be assigned at random to one of the two programs. Volunteers will have their stress levels measured just before beginning the program and \(10\) weeks later at the completion of it.
(a) A group of volunteers who work together ask to be assigned to the same program so that they can participate in that program together. Give an example of a problem that might arise if this is permitted. Explain to this volunteer group why random assignment to the two programs will address this problem.
(b) Someone proposes that a control group be included in the design as well. The stress level would be measured for each volunteer assigned to the control group at the start of the study and again \(10\) weeks later. What additional information, if any, would this provide about the effectiveness of the two programs?
(c) Is it reasonable to generalize the findings of this study to all employees of this company? Explain.

Most-appropriate topic codes (AP Statistics):

• Topic 1.13 — Experimental Design (Parts a, b)
• Topic 1.11 — Random Sampling (Part c)
• Topic 1.12 — Potential Problems with Sampling (Part c)
▶️ Answer/Explanation

(a)
If volunteers who work together are all placed in the same program, there’s a risk that something specific to their workplace situation gets mixed up with the effect of the program itself. For example, suppose this group’s department recently had a deadline pushed back, which on its own would lower everyone’s stress level in that department regardless of which program they’re doing. If this entire group ends up in, say, the tai chi group, then the drop in stress they experience could mistakenly be credited to tai chi, when really it was caused by the lighter workload.

Random assignment fixes this issue. By randomly assigning volunteers to the two programs instead of letting groups choose, we spread people from this department across both the tai chi and yoga groups. This way, any unusual circumstance affecting that department’s stress levels — like the deadline change — gets “evened out” between the two treatment groups rather than being concentrated in just one. Randomization helps make sure the two groups are comparable at the start, so that any difference we see at the end can be attributed to the program rather than to some other confounding factor.

(b)
Yes, a control group would add useful information. Without one, the company could only compare tai chi to yoga directly — they could say which of the two programs led to a bigger drop in stress, but they couldn’t say whether either program actually caused a reduction in stress at all.

Here’s the issue: stress levels might naturally go down over a \(10\)-week period for reasons that have nothing to do with either program — for instance, if the overall work environment becomes less hectic during that time, everyone’s stress might drop a little just from that. A control group, which doesn’t participate in either program but still has its stress measured at the start and end of the \(10\) weeks, gives a baseline for what “no treatment” looks like under those same conditions.

By comparing each treatment group’s change in stress to the control group’s change, the company can tell how much of the reduction is actually attributable to tai chi or yoga specifically, rather than just background changes that would have happened anyway.

(c)
No, it is not reasonable to generalize these findings to all employees of the company. The participants in this study were volunteers, not a random sample of employees. People who choose to volunteer for a stress-reduction study might already be different from the typical employee — for example, they might be more motivated to manage their stress, more open to trying tai chi or yoga, or have more flexible schedules that let them give up part of their lunch hour.

Because the group wasn’t randomly selected from the entire employee population, there’s no guarantee that what works (or doesn’t work) for these volunteers would apply the same way to employees who didn’t volunteer. So while the random assignment within the study supports drawing cause-and-effect conclusions about tai chi versus yoga for people like these volunteers, it doesn’t justify extending those conclusions to the company as a whole.

Question

A study was conducted to determine if taking vitamin C reduces the occurrence of the flu. The study was conducted using 808 student volunteers who did not take a flu shot. The subjects were randomly assigned to one of two groups: a treatment group who received 1,000 milligrams of vitamin C daily or a control group who received a placebo flavored to taste like the vitamin C treatment. All participants were monitored to ensure that they adhered to their assigned treatment on a daily basis throughout the period of the study. At the end of the flu season, each subject’s medical record was reviewed by a physician to determine whether he or she had contracted the flu during the period of the study. The physician did not know which treatment each subject received. The results of the study are shown in the table below.
(a) Is this study an experiment or an observational study? Explain your answer.
(b) Based on this study, a health expert claims that there is evidence to suggest that vitamin C reduces the occurrence of the flu in the population of students who would volunteer for such a study. State the name of a test and the null and alternative hypotheses that the health expert could have used to support this claim. Do not carry out the test.

Most-appropriate topic codes (AP Statistics):

• Topic 1.13 — Experimental Design (Part a)
• Topic 3.12 — Setting Up a Test for the Difference Between Two Population Proportions (Part b)
• Topic 3.13 — Carrying Out a Test for the Difference Between Two Population Proportions (Part b)
• Topic 3.9 — Sampling Distributions for the Difference Between Sample Proportions (Part b)
▶️ Answer/Explanation

(a)

This study is an experiment, not an observational study.
In an experiment, the researchers deliberately impose a treatment on the subjects — here, the researchers assigned subjects to either the vitamin C group or the placebo group, rather than simply observing what subjects naturally chose to do.
Crucially, subjects were randomly assigned to the two treatment groups, which is the hallmark of a well-designed experiment and allows for causal conclusions to be drawn.
The use of a placebo and blind evaluation by a physician further strengthen this as a controlled experiment.

(b)

The health expert could use a two-proportion \(z\)-test to support this claim.
Let \(p_T\) be the true proportion of students (in the population of volunteers) who contract the flu when taking vitamin C, and let \(p_C\) be the true proportion who contract the flu when taking a placebo.
The hypotheses are:
\( H_0: p_T – p_C = 0 \quad \text{(vitamin C has no effect on flu occurrence)} \)
\( H_a: p_T – p_C < 0 \quad \text{(vitamin C reduces flu occurrence)} \)
Equivalently, this can be written as \(H_0: p_T = p_C\) versus \(H_a: p_T < p_C\), where a one-sided alternative is used because the claim is specifically that vitamin C reduces the occurrence of flu.

Question

There have been many studies recently concerning coffee drinking and cholesterol level. While it is known that several coffee-bean components can elevate blood cholesterol level, it is thought that a new type of paper coffee filter may reduce the presence of some of these components in coffee.
The effect of the new filter on cholesterol level will be studied over a 10-week period using 300 nonsmokers who each drink 4 cups of caffeinated coffee per day. Each of these 300 participants will be assigned to one of two groups: the experimental group, who will only drink coffee that has been made with the new filter, or the control group, who will only drink coffee that has been made with the standard filter. Each participant’s cholesterol level will be measured at the beginning and at the end of the study.
(a) Describe an appropriate method for assigning the subjects to the two groups so that each group will have an equal number of subjects.
(b) In this study, the researchers chose to include a group who only drank coffee that was made with the standard filter. Why is it important to include a control group in this study even though cholesterol levels will be measured at the beginning and at the end of the study?
(c) Which test would you conduct to determine whether the change in cholesterol level would be greater if people used the new filter rather than using the standard filter?
(d) Why would the researchers choose to use only nonsmokers in the study?

Most-appropriate topic codes (AP Statistics):

• Topic 1.13 — Experimental Design (Parts a, b, d)
• Topic 4.10 — Carrying Out a Test for the Difference Between Two Population Means (Part c)
• Topic 4.9 — Setting Up a Test for the Difference Between Two Population Means (Part c)
▶️ Answer/Explanation

(a)

Assign each of the 300 subjects a unique number from 001 to 300.
Use a random number table or a random number generator to randomly select 150 of these numbers — the subjects corresponding to those numbers will be placed in the experimental group (new filter), and the remaining 150 subjects will be assigned to the control group (standard filter).
This ensures that each subject has an equal chance of being assigned to either group, and that both groups end up with exactly 150 subjects.

(b)

Without a control group, if cholesterol levels change over the 10-week period, there is no way to know whether the change was caused by the new filter or by some other factor that changed during that time — for example, seasonal changes in diet or physical activity.
Simply measuring cholesterol at the beginning and end for one group only tells us that a change occurred, but not why it occurred.
By including a control group using the standard filter, researchers can compare the mean change in cholesterol between the two groups, allowing them to attribute any difference specifically to the new filter rather than to an outside confounding variable.
In short, the control group is what makes it possible to isolate the effect of the new filter from other changes that might naturally occur over the 10-week period.

(c)

The appropriate test is the two-sample \(t\)-test for means (or equivalently, the two-sample \(t\)-test for mean differences).
This test compares the mean change in cholesterol level for the new-filter group against the mean change in cholesterol level for the standard-filter group, using two independent samples of size 150 each.

(d)

Smoking is known to be related to cholesterol level, so if the study included both smokers and nonsmokers, smoking status could act as an extraneous source of variability — making it harder to detect the true effect of the coffee filter.
By restricting the study to nonsmokers only, the researchers create more homogeneous groups, reducing variability within each group and allowing for a more precise and sensitive comparison of the filter’s effect on cholesterol.
This is similar in spirit to blocking — controlling for a known nuisance variable (smoking) so that it does not obscure the treatment effect, though it does mean the results can only be generalized to nonsmokers.

Scroll to Top