Home / AP® Exam / AP® Statistics / AP Statistics 3.2 Sampling Distributions for Sample Proportions- Exam Style Questions – FRQs

AP Statistics 3.2 Sampling Distributions for Sample Proportions- Exam Style Questions - FRQs - New Syllabus

Question

A polling agency showed the following two statements to a random sample of 1,048 adults in the United States.
Environment statement: Protection of the environment should be given priority over economic growth.
Economy statement: Economic growth should be given priority over protection of the environment.
The order in which the statements were shown was randomly selected for each person in the sample. After reading the statements, each person was asked to choose the statement that was most consistent with his or her opinion. The results are shown in the table.

(a) Assume the conditions for inference have been met. Construct and interpret a 95 percent confidence interval for the proportion of all adults in the United States who would have chosen the economy statement.
(b) One of the conditions for inference that was met is that the number who chose the economy statement and the number who did not choose the economy statement are both greater than 10. Explain why it is necessary to satisfy that condition.
(c) A suggestion was made to use a two-sample \(z\)-interval for a difference between proportions to investigate whether the difference in proportions between adults in the United States who would have chosen the environment statement and adults in the United States who would have chosen the economy statement is statistically significant. Is the two-sample \(z\)-interval for a difference between proportions an appropriate procedure to investigate the difference? Justify your answer.

Most-appropriate topic codes (AP Statistics):

• Topic \(3.2\) — Sampling Distributions for Sample Proportions (Part \( \mathrm{b} \))
• Topic \(3.3\) — Constructing a Confidence Interval for a Population Proportion (Parts \( \mathrm{a} \), \( \mathrm{b} \))
• Topic \(3.4\) — Justifying a Claim Based on a Confidence Interval for a Population Proportion (Part \( \mathrm{a} \))
• Topic \(3.10\) — Constructing a Confidence Interval for the Difference Between Two Population Proportions (Part \( \mathrm{c} \))
▶️ Answer/Explanation

(a)

The appropriate procedure is a one-sample \(z\)-interval for a population proportion.
Let \(p\) = the true proportion of all adults in the United States who would have chosen the economy statement.
From the sample: \(\hat{p} = 0.37\), \(n = 1{,}048\), and the critical value for 95% confidence is \(z^* = 1.96\).
The confidence interval formula is:
\( \hat{p} \pm z^* \sqrt{\dfrac{\hat{p}(1-\hat{p})}{n}} \)
Substituting the values:
\( 0.37 \pm 1.96\sqrt{\dfrac{(0.37)(0.63)}{1{,}048}} \)
\( 0.37 \pm 1.96\sqrt{0.000222} \)
\( 0.37 \pm 1.96(0.0149) \)
\( 0.37 \pm 0.03 = (0.34,\ 0.40) \)
Interpretation: We are 95 percent confident that the interval from 0.34 to 0.40 captures the true proportion of all adults in the United States who would have chosen the economy statement.

(b)

This condition is necessary because the confidence interval formula for a population proportion relies on approximating the binomial distribution with a normal distribution. This approximation causes the sampling distribution of \(\hat{p}\) to be approximately normal, which is what allows us to use the \(z\)-critical value and construct a valid interval.
However, the normal approximation to the binomial works well only when both \(n\hat{p}\) and \(n(1-\hat{p})\) are at least 10. If either of these quantities is too small, the sampling distribution of \(\hat{p}\) will be noticeably skewed rather than approximately normal, and the confidence interval will not be reliable.

(c)

No, the two-sample \(z\)-interval for a difference between proportions is not an appropriate procedure here.
A key requirement for the two-sample \(z\)-interval is that the two proportions come from two independent samples. In this study, however, both proportions — the proportion choosing the environment statement and the proportion choosing the economy statement — come from the same single sample of 1,048 adults. Because each person was forced to choose between the two statements (or express no preference), the two proportions are not independent: knowing one proportion directly constrains the other. Since the independence condition is violated, the two-sample \(z\)-interval is not appropriate for this situation.

Question

A husband and wife, Mike and Lori, share a digital music player that has a feature that randomly selects which song to play. A total of 2,384 songs were loaded onto the player, some by Mike and the rest by Lori. Suppose that when the player was in the random-selection mode, 13 of the first 50 songs selected were songs loaded by Lori.
(a) Construct and interpret a 90 percent confidence interval for the proportion of songs on the player that were loaded by Lori.
(b) Mike and Lori are unsure about whether the player samples the songs with replacement or without replacement when the player is in random-selection mode. Explain why this distinction is not important for the construction of the interval in part (a).

Most-appropriate topic codes (AP Statistics):

• Topic \(3.3\) — Constructing a Confidence Interval for a Population Proportion (Part \(\mathrm{a}\))
• Topic \(3.4\) — Justifying a Claim Based on a Confidence Interval for a Population Proportion (Part \(\mathrm{a}\))
• Topic \(3.2\) — Sampling Distributions for Sample Proportions (Part \(\mathrm{b}\))
▶️ Answer/Explanation

(a)
Procedure: One-sample \(z\)-interval for a population proportion.
Conditions:
Random: The problem states the 50 songs were randomly selected. ✓
Large Counts: \(n\hat{p} = 50 \times 0.26 = 13 \geq 10\) and \(n(1-\hat{p}) = 50 \times 0.74 = 37 \geq 10\). ✓
Both conditions are satisfied, so we may proceed.
Calculation:
The sample proportion is:
\(\hat{p} = \frac{13}{50} = 0.26\)
For a 90% confidence interval, the critical value is \(z^* = 1.645\). The margin of error is:
\(ME = z^* \sqrt{\frac{\hat{p}(1-\hat{p})}{n}} = 1.645 \times \sqrt{\frac{0.26 \times 0.74}{50}}\)
\(ME = 1.645 \times \sqrt{\frac{0.1924}{50}} = 1.645 \times \sqrt{0.003848} = 1.645 \times 0.06203 \approx 0.102\)
The 90% confidence interval is:
\(\hat{p} \pm ME = 0.26 \pm 0.102 = \boxed{(0.158,\ 0.362)}\)
Interpretation: We are 90% confident that the true proportion of all songs on the digital music player that were loaded by Lori is between 0.158 and 0.362.
In plain terms — if we repeated this process many times, about 90% of the confidence intervals we’d construct would capture the actual fraction of songs that Lori loaded. Based on this one interval, our best guess is somewhere between roughly 16% and 36%.

(b)
The distinction between sampling with replacement and without replacement only matters meaningfully when the sample size is large relative to the population size. Here we check the ratio:
\(\frac{\text{Population size}}{\text{Sample size}} = \frac{2384}{50} = 47.7\)
Since the population is about 47.7 times larger than the sample, the sample of 50 represents only:
\(\frac{50}{2384} \approx 2.1\%\text{ of the population}\)
This is well below the standard 10% threshold. When the sample is this small relative to the population, removing a song from the pool (sampling without replacement) barely changes the probability of selecting any remaining song. For example:
Probability of selecting a specific song with replacement: \(\dfrac{1}{2384} \approx 4.195 \times 10^{-4}\)
Probability of selecting a specific song without replacement (after one song removed): \(\dfrac{1}{2383} \approx 4.196 \times 10^{-4}\)
The difference is negligible — on the order of \(10^{-7}\). Because the population is so large relative to the sample, the two sampling methods produce virtually identical results, and the confidence interval from part (a) is valid regardless of which method the player uses.

Question

Sunshine Farms wants to know whether there is a difference in consumer preference for two new juice products — Citrus Fresh and Tropical Taste. In an initial blind taste test, 8 randomly selected consumers were given unmarked samples of the two juices. The product that each consumer tasted first was randomly decided by the flip of a coin. After tasting the two juices, each consumer was asked to choose which juice he or she preferred, and the results were recorded.
(a) Let \(p\) represent the population proportion of consumers who prefer Citrus Fresh. In terms of \(p\), state the hypotheses that Sunshine Farms is interested in testing.
(b) One might consider using a one-proportion \(z\)-test to test the hypotheses in part (a). Explain why this would not be a reasonable procedure for this sample.
(c) Let \(X\) represent the number of consumers in the sample who prefer Citrus Fresh. Assuming there is no difference in consumer preference, find the probability for each possible value of \(X\). Record the \(x\)-values and the corresponding probabilities in the table below.
(d) When testing the hypotheses in part (a), Sunshine Farms will conclude that there is a consumer preference if too many or too few individuals prefer Citrus Fresh. Based on your probabilities in part (c), is it possible for the significance level (probability of rejecting the null hypothesis when it is true) for this test to be exactly 0.05? Justify your answer.
(e) The preference data for the 8 randomly selected consumers are given in the table below.
Based on these preferences and your previous work, test the hypotheses in part (a).
(f) Sunshine Farms plans to add one of these two new juices — Citrus Fresh or Tropical Taste — to its production schedule. A follow-up study will be conducted to decide which of the two juices to produce. Make one recommendation for the follow-up study that would make it better than the initial study. Provide a statistical justification for your recommendation in the context of the problem.

Most-appropriate topic codes (AP Statistics):

• Topic 3.5 — Setting Up a Test for a Population Proportion (Part \(\mathrm{a}\))
• Topic 3.2 — Sampling Distributions for Sample Proportions (Part \(\mathrm{b}\))
• Topic 2.10 — The Binomial Distribution (Part \(\mathrm{c}\))
• Topic 3.6 — p-Values (Part \(\mathrm{d}\))
• Topic 3.7 — Carrying Out a Test for a Population Proportion (Part \(\mathrm{e}\))
• Topic 1.13 — Experimental Design (Part \(\mathrm{f}\))
▶️ Answer/Explanation

(a)
Let \(p\) be the population proportion of consumers who prefer Citrus Fresh. The hypotheses are:
\(H_0: p = 0.5\)
\(H_a: p \neq 0.5\)
A two-sided alternative is appropriate because Sunshine Farms wants to detect any difference in preference, not just preference for one particular juice.

(b)
The conditions for a one-proportion \(z\)-test require that both \(np\) and \(n(1-p)\) be at least 5 (or 10). Here:
\(np = 8 \times 0.5 = 4 < 5\)
\(n(1-p) = 8 \times 0.5 = 4 < 5\)
Since both values are less than 5, the large-sample normal approximation is not valid, and using a one-proportion \(z\)-test would not be appropriate for a sample of only \(n = 8\).

(c)
Under \(H_0\), \(X \sim \text{Binomial}(n = 8,\ p = 0.5)\). The probabilities are computed using:
\(P(X = x) = \binom{8}{x}(0.5)^x(0.5)^{8-x} = \binom{8}{x}(0.5)^8\)

(d)
No, it is not possible for the significance level to be exactly 0.05. Because \(X\) is a discrete random variable, the tail probabilities can only take specific values — there is no rejection region that gives a type I error probability of exactly 0.05.
The most extreme rejection region \((X = 0 \text{ or } X = 8)\) gives:
\(\alpha = 2 \times 0.00391 = 0.00782 < 0.05\)
The next possible rejection region \((X \leq 1 \text{ or } X \geq 7)\) gives:
\(\alpha = 2 \times (0.00391 + 0.03125) = 2 \times 0.03516 = 0.07031 > 0.05\)
Since no rejection region produces a type I error probability of exactly 0.05, a significance level of exactly 0.05 is not achievable with this test.
\(\boxed{\alpha = 0.05 \text{ is not achievable — the achievable levels jump from } 0.00782 \text{ to } 0.07031}\)

(e)
From the data, 2 out of 8 consumers preferred Citrus Fresh, so \(X = 2\).
Since this is a two-sided test, the \(p\)-value is the probability of observing a result at least as extreme as \(X = 2\) in either tail:
\(p\text{-value} = P(X \leq 2) + P(X \geq 6)\)
\(= 2 \times [P(X=0) + P(X=1) + P(X=2)]\)
\(= 2 \times (0.00391 + 0.03125 + 0.10937)\)
\(= 2 \times 0.14453 = 0.28906\)
Since the \(p\)-value of \(0.289\) is much larger than any reasonable significance level (e.g., \(\alpha = 0.05\) or \(\alpha = 0.07031\)), we fail to reject \(H_0\). There is not statistically significant evidence of a consumer preference between Citrus Fresh and Tropical Taste.
\(\boxed{p\text{-value} \approx 0.289 \implies \text{Fail to reject } H_0; \text{ no significant consumer preference detected}}\)

(f)
The most important recommendation is to increase the number of consumers in the study. With only \(n = 8\) consumers, the test has very low power — even a large true difference in preference (like 75% vs. 25%) may not produce a statistically significant result. Increasing the sample size would reduce the standard error of the estimated proportion \(\hat{p}\), making it easier to detect a real difference, and would allow the use of the large-sample one-proportion \(z\)-test since \(np \geq 5\) and \(n(1-p) \geq 5\) would be satisfied. For example, with \(n = 80\) and \(X = 20\) (same sample proportion of 0.25), the \(z\)-statistic would be approximately:
\(z = \frac{0.25 – 0.5}{\sqrt{\frac{0.5(0.5)}{80}}} \approx -4.47\)
which gives a \(p\)-value near zero, allowing a clear conclusion to be reached.
\(\boxed{\text{Recommendation: Increase sample size to increase power and enable use of the } z\text{-test}}\)

Question

Researchers at a large health maintenance organization (HMO) are planning a study of a certain mild illness. They will select a random sample of patients who are ages 35 to 54 and see if they contract the illness in the next year. The researchers are interested in estimating the proportions of men and of women who are likely to develop the illness in each of 4 age-groups: 35–39, 40–44, 45–49, and 50–54.
The researchers plan to include 2,000 patients in the study. Suppose the researchers draw a random sample from all of the patients at this HMO who are ages 35 to 54 and find the following numbers within each gender and age-group.
(a) Suppose that at the end of the study, 10 percent of the females in the 40–44 age-group contracted the illness. Calculate a 95 percent confidence interval to estimate the population proportion of females in this age-group that contracted the illness.
Interpret this confidence interval in the context of this situation.
Interpret the confidence level of 95 percent.
(b) Suppose that at the end of the study, 10 percent of the males in the 40–44 age-group contracted the illness. The corresponding 95 percent confidence interval to estimate the population proportion of males in this age-group that contracted the illness is \((0.061,\ 0.139)\).
Note that this interval and the interval in part (a) are of different lengths even though the two sample proportions were identical. What would be an alternative way to allocate a sample of 2,000 subjects so that the 95 percent confidence interval widths for all male age-groups and for all female age-groups (i.e., for all 8 groups) would be the same when the sample proportions are the same? Justify your answer.
(c) Based on previous studies, researchers believe that the percentages of those who contract the illness will be similar for males and females, and therefore plan to ignore gender when selecting a sample for this study. Previous studies also indicate that the percentages of adults who will contract this illness in the 35–39, 40–44, 45–49, and 50–54 age-groups are anticipated to be 5%, 8%, 20%, and 35%, respectively. How should the sample of 2,000 subjects be allocated with respect to age-groups so that the widths of the 95 percent confidence intervals for the four groups will be approximately the same? Justify your answer.

Most-appropriate topic codes (AP Statistics):

• Topic 3.3 — Constructing a Confidence Interval for a Population Proportion (Part a)
• Topic 3.4 — Justifying a Claim Based on a Confidence Interval for a Population Proportion (Part a)
• Topic 3.1 — Estimators (Part b)
• Topic 3.2 — Sampling Distributions for Sample Proportions (Parts b, c)
▶️ Answer/Explanation

(a)

Procedure and Conditions:
We use a one-sample \(z\)-confidence interval for a proportion: \(\hat{p} \pm z^* \sqrt{\dfrac{\hat{p}(1-\hat{p})}{n}}\).
From the table, the sample of females in the 40–44 age-group has \(n = 370\). With 10% contracting the illness, \(\hat{p} = 0.10\).
Check conditions: The sample is a random sample from the HMO population. The expected number of successes is \(n\hat{p} = 370(0.10) = 37 > 5\) and the expected number of failures is \(n(1-\hat{p}) = 370(0.90) = 333 > 5\), so the sample size is large enough to proceed.

Mechanics:
\( \hat{p} \pm z^* \sqrt{\frac{\hat{p}(1-\hat{p})}{n}} = 0.10 \pm 1.96\sqrt{\frac{(0.10)(0.90)}{370}} \)
\( = 0.10 \pm 1.96\sqrt{\frac{0.09}{370}} = 0.10 \pm 1.96(0.01559) = 0.10 \pm 0.0306 \)
\( \boxed{(0.06943,\ 0.13057)} \)

Interpretation of the confidence interval:
We are 95% confident that the true proportion of the HMO’s 40–44 year old female patients who would contract this illness is between approximately 0.069 and 0.131 (i.e., between about 6.9% and 13.1%).

Interpretation of the 95% confidence level:
If we were to take many random samples of size 370 from this population and compute a 95% confidence interval from each sample, approximately 95% of those intervals would capture the true population proportion of 40–44 year old female HMO patients who contract the illness. In other words, the method we used to construct this interval will fail to contain the true proportion only about 5% of the time in repeated sampling.

(b)

The width of a confidence interval for a proportion is determined by the margin of error \(z^*\sqrt{\dfrac{\hat{p}(1-\hat{p})}{n}}\), which depends on both \(\hat{p}\) and \(n\).
When the sample proportions are equal (both 0.10), the only factor causing different interval widths is the different sample sizes — females had \(n = 370\) while males had \(n = 230\) in the 40–44 group, producing a wider interval for males.
To make the confidence interval widths equal for all 8 groups when sample proportions are equal, the sample sizes for all 8 groups must be equal.
With 2,000 total subjects and 8 groups, each group should receive a sample of size:
\( n = \frac{2000}{8} = \boxed{250 \text{ subjects per group}} \)
Equal sample sizes guarantee equal standard errors (and therefore equal interval widths) whenever the sample proportions are the same.

(c)

When sample proportions differ across groups, equal widths require not equal sample sizes but rather sample sizes proportional to \(\hat{p}(1-\hat{p})\) for each group — this keeps the standard error \(\sqrt{\dfrac{\hat{p}(1-\hat{p})}{n}}\) the same across all groups.
Compute \(p(1-p)\) for each age-group using the anticipated proportions:

The total of all \(p(1-p)\) values is:
\( 0.0475 + 0.0736 + 0.1600 + 0.2275 = 0.5086 \)
Allocate the 2,000 subjects proportionally:
\( n_{35\text{-}39} = \frac{0.0475}{0.5086} \times 2000 = \boxed{186.79} \)
\( n_{40\text{-}44} = \frac{0.0736}{0.5086} \times 2000 = \boxed{289.42} \)
\( n_{45\text{-}49} = \frac{0.1600}{0.5086} \times 2000 = \boxed{629.18} \)
\( n_{50\text{-}54} = \frac{0.2275}{0.5086} \times 2000 = \boxed{894.61} \)
These four sample sizes sum to 2,000 and ensure that \(\sqrt{\dfrac{p(1-p)}{n}}\) is approximately equal across all four age-groups, producing confidence intervals of equal width.

Scroll to Top