Home / AP® Exam / AP® Statistics / AP Statistics 3.11 Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Proportions- Exam Style Questions – FRQs

AP Statistics 3.11 Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Proportions- Exam Style Questions - FRQs - New Syllabus

Question

The manager of a large company that sells pet supplies online wants to increase sales by encouraging repeat purchases. The manager believes that if past customers are offered \(\$10\) off their next purchase, more than \(40 \text{ percent}\) of them will place an order. To investigate the belief, \(90\) customers who placed an order in the past year are selected at random. Each of the selected customers is sent an e-mail with a coupon for \(\$10\) off the next purchase if the order is placed within \(30\) days. Of those who receive the coupon, \(38\) place an order.
(a) Is there convincing statistical evidence, at the significance level of \(\alpha=0.05\), that the manager’s belief is correct? Complete the appropriate inference procedure to support your answer.
(b) Based on your conclusion from part (a), which of the two errors, Type I or Type II, could have been made? Interpret the consequence of the error in context.

Most-appropriate topic codes (AP Statistics):

• Topic \(3.5\) — Setting Up a Test for a Population Proportion (Part \( \mathrm{a} \))
• Topic \(3.6\) — \(p\)-Values (Part \( \mathrm{a} \))
• Topic \(3.7\) — Carrying Out a Test for a Population Proportion (Part \( \mathrm{a} \))
• Topic \(3.11\) — Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Proportions (Part \( \mathrm{b} \))
▶️ Answer/Explanation

(a)
Let \(p\) be the true proportion of customers who place an order. We test \(H_0: p = 0.40\) against \(H_a: p > 0.40\).
Conditions are met: it’s a random sample, the \(10\%\) rule is satisfied (assume \(\ge 900\) customers), and expected counts \((36, 54)\) are both \(\ge 10\).
The sample proportion is \(\hat{p} = \frac{38}{90} \approx 0.422\), giving a test statistic \(z = \frac{0.422 – 0.40}{\sqrt{0.4(0.6)/90}} \approx 0.430\) and a \(p\)-value of \(0.333\).
Since \(0.333 > 0.05\), we fail to reject \(H_0\); there is not convincing evidence the manager’s belief is correct.

(b)
Because we failed to reject the null hypothesis, a Type II error could have been made.
In context, this means the manager incorrectly thinks the coupon won’t bring in more than \(40\%\) of customers, deciding not to use it and ultimately missing out on a promotion that would have increased sales.

Question

Systolic blood pressure is the amount of pressure that blood exerts on blood vessels while the heart is beating. The mean systolic blood pressure for people in the United States is reported to be 122 millimeters of mercury (mmHg) with a standard deviation of 15 mmHg.
The wellness department of a large corporation is investigating whether the mean systolic blood pressure of its employees is greater than the reported national mean. A random sample of 100 employees will be selected, the systolic blood pressure of each employee in the sample will be measured, and the sample mean will be calculated.
Let \(\mu\) represent the mean systolic blood pressure of all employees at the corporation. Consider the following hypotheses.
\(H_0 : \mu = 122\)
\(H_a : \mu > 122\)
(a) Describe a Type II error in the context of the hypothesis test.
(b) Assume that \(\sigma\), the standard deviation of the systolic blood pressure of all employees at the corporation, is \(15\) mmHg. If \(\mu = 122\), the sampling distribution of \(\bar{x}\) for samples of size 100 is approximately normal with a mean of 122 mmHg and a standard deviation of 1.5 mmHg. What values of the sample mean \(\bar{x}\) would represent sufficient evidence to reject the null hypothesis at the significance level of \(\alpha = 0.05\)?
The actual mean systolic blood pressure of all employees at the corporation is 125 mmHg, not the hypothesized value of 122 mmHg, and the standard deviation is 15 mmHg.
(c) Using the actual mean of 125 mmHg and the results from part (b), determine the probability that the null hypothesis will be rejected.
(d) What statistical term is used for the probability found in part (c)?
(e) Suppose the size of the sample of employees to be selected is greater than 100. Would the probability of rejecting the null hypothesis be greater than, less than, or equal to the probability calculated in part (c)? Explain your reasoning.

Most-appropriate topic codes (AP Statistics):

• Topic \(3.11\) — Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Proportions (Part \( \mathrm{a} \))
• Topic \(4.5\) — Carrying Out a Test for a Population Mean (Part \( \mathrm{b} \))
• Topic \(4.7\) — Constructing a Confidence Interval for the Difference Between Two Population Means (Parts \( \mathrm{c} \), \( \mathrm{d} \), \( \mathrm{e} \))
▶️ Answer/Explanation

(a)
A Type II error occurs when the alternative hypothesis is actually true, but we fail to reject the null hypothesis.
In this context, a Type II error would occur if the true mean systolic blood pressure of all employees is greater than 122 mmHg, but the hypothesis test does not detect this — that is, the null hypothesis (\(H_0 : \mu = 122\)) is not rejected.
In plain terms: the employees really do have higher-than-national-average blood pressure, but the test fails to conclude so.

(b)
Since this is a one-sided (right-tailed) test with known \(\sigma\), we reject \(H_0\) when the test statistic exceeds the critical value \(z^* = 1.645\) at \(\alpha = 0.05\).
The test statistic is \(z = \dfrac{\bar{x} – \mu_0}{\sigma / \sqrt{n}} = \dfrac{\bar{x} – 122}{15/\sqrt{100}} = \dfrac{\bar{x} – 122}{1.5}\).
Setting \(\dfrac{\bar{x} – 122}{1.5} > 1.645\) and solving:
\(\bar{x} – 122 > 1.645 \times 1.5 = 2.4675\)
\(\boxed{\bar{x} > 124.4675 \text{ mmHg}}\)
Any sample mean greater than approximately 124.47 mmHg provides sufficient evidence to reject \(H_0\) at the \(\alpha = 0.05\) level.

(c)
Now the true population mean is \(\mu = 125\) mmHg (with \(\sigma = 15\), \(n = 100\)), so the sampling distribution of \(\bar{x}\) is approximately normal with mean 125 and standard deviation \(\dfrac{15}{\sqrt{100}} = 1.5\).
We want the probability that \(\bar{x}\) falls in the rejection region found in part (b):
\(P(\bar{x} > 124.4675) = P\!\left(z > \dfrac{124.4675 – 125}{1.5}\right) = P(z > -0.355)\ Folk\
\(\boxed{P(z > -0.355) \approx 0.64}\)
There is approximately a 64% probability that the null hypothesis will be rejected when the true mean is 125 mmHg.

(d)
The probability found in part (c) — the probability of correctly rejecting a false null hypothesis — is called the power of the test.
\(\boxed{\text{Power of the test} \approx 0.64}\)

(e)
The probability of rejecting \(H_0\) would be greater than 0.64 if the sample size is larger than 100.
A larger sample size \(n\) reduces the standard error of \(\bar{x}\): \(\sigma_{\bar{x}} = \dfrac{\sigma}{\sqrt{n}}\), so the sampling distribution becomes narrower.
This means the critical value \(\bar{x}^* = \mu_0 + z^* \cdot \dfrac{\sigma}{\sqrt{n}}\) would be smaller (closer to 122), lowering the threshold needed to reject \(H_0\).
With a lower rejection threshold, it becomes more likely that \(\bar{x}\) exceeds that threshold when the true mean is 125 mmHg — so the power of the test increases.
\(\boxed{\text{Probability of rejecting } H_0 \text{ would be greater than 0.64.}}\)

Question

Two treatments, A and B, showed promise for treating a potentially fatal disease. A randomized experiment was conducted to determine whether there is a significant difference in the survival rate between patients who receive treatment A and those who receive treatment B. Of 154 patients who received treatment A, 38 survived for at least 15 years, whereas 16 of the 164 patients who received treatment B survived at least 15 years.
(a) Treatment A can be administered only as a pill, and treatment B can be administered only as an injection. Can this randomized experiment be performed as a double-blind experiment? Why or why not?
(b) The conditions for inference have been met. Construct and interpret a 95 percent confidence interval for the difference between the proportion of the population who would survive at least 15 years if given treatment A and the proportion of the population who would survive at least 15 years if given treatment B.
In many of these types of studies, physicians are interested in the ratio of survival probabilities, \(\dfrac{p_A}{p_B}\), where \(p_A\) represents the true 15-year survival rate for all patients who receive treatment A and \(p_B\) represents the true 15-year survival rate for all patients who receive treatment B. This ratio is usually referred to as the relative risk of the two treatments.
For example, a relative risk of 1 indicates the survival rates for patients receiving the two treatments are equal, whereas a relative risk of 1.5 indicates that the survival rate for patients receiving treatment A is 50 percent higher than the survival rate for patients receiving treatment B. An estimator of the relative risk is the ratio of estimated probabilities, \(\dfrac{\hat{p}_A}{\hat{p}_B}\).
(c) Using the data from the randomized experiment described above, compute the estimate of the relative risk.
The sampling distribution of \(\dfrac{\hat{p}_A}{\hat{p}_B}\) is skewed. However, when both sample sizes \(n_A\) and \(n_B\) are relatively large, the distribution of \(\ln\!\left(\dfrac{\hat{p}_A}{\hat{p}_B}\right)\) — the natural logarithm of relative risk — is approximately normal with a mean of \(\ln\!\left(\dfrac{p_A}{p_B}\right)\) and a standard deviation of \(\sqrt{\dfrac{1-p_A}{n_A p_A}+\dfrac{1-p_B}{n_B p_B}}\), where \(p_A\) and \(p_B\) can be estimated by using \(\hat{p}_A\) and \(\hat{p}_B\).
When a 95 percent confidence interval for \(\ln\!\left(\dfrac{p_A}{p_B}\right)\) is known, an approximate 95 percent confidence interval for \(\dfrac{p_A}{p_B}\) — the relative risk of the two treatments — can be constructed by applying the inverse of the natural logarithm to the endpoints of the confidence interval for \(\ln\!\left(\dfrac{p_A}{p_B}\right)\).
(d) The conditions for inference are met for the data in the experiment above, and a 95 percent confidence interval for \(\ln\!\left(\dfrac{p_A}{p_B}\right)\) is \((0.3868,\ 1.4690)\). Construct and interpret a 95 percent confidence interval for the relative risk, \(\dfrac{p_A}{p_B}\), of the two treatments.
(e) What is an advantage of using the interval in part (d) over using the interval in part (b)?

Most-appropriate topic codes (AP Statistics):

• Topic \(1.13\) — Experimental Design (Part \(\mathrm{a}\): double-blind experiments and random assignment)
• Topic \(3.10\) — Constructing a Confidence Interval for the Difference Between Two Population Proportions (Part \(\mathrm{b}\))
• Topic \(3.11\) — Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Proportions (Parts \(\mathrm{b}\), \(\mathrm{e}\))
• Topic \(3.1\) — Estimators (Part \(\mathrm{c}\): estimating relative risk using \(\hat{p}_A/\hat{p}_B\))
• Topic \(3.11\) — Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Proportions (Part \(\mathrm{d}\): interpreting the confidence interval for relative risk)
▶️ Answer/Explanation

(a)

Yes, this experiment can be performed as a double-blind experiment by introducing placebos for each treatment group.
Patients assigned to treatment A (pill) would also receive a placebo injection. Patients assigned to treatment B (injection) would also receive a placebo pill. This way, every patient receives both a pill and an injection, but one of the two is a placebo.
Since neither the patients nor the physicians administering the treatments know which is the real treatment and which is the placebo, neither group is aware of the treatment assignment — making the experiment double-blind.

(b)

First, compute the sample proportions:
\(\hat{p}_A = \frac{38}{154} \approx 0.2468, \qquad \hat{p}_B = \frac{16}{164} \approx 0.0976\)
The 95% confidence interval for \(p_A – p_B\) is:
\(\hat{p}_A – \hat{p}_B) \pm z^* \sqrt{\frac{\hat{p}_A(1-\hat{p}_A)}{n_A} + \frac{\hat{p}_B(1-\hat{p}_B)}{n_B}}\)
\(0.2468 – 0.0976) \pm 1.96\sqrt{\frac{(0.2468)(0.7532)}{154} + \frac{(0.0976)(0.9024)}{164}}\)
\(0.1492 \pm 1.96(0.0418)\)
\(0.1492 \pm 0.0818\)
\(\boxed{(0.0674,\ 0.2310)}\)
We are 95% confident that the true difference in 15-year survival rates \((p_A – p_B)\) is between \(0.0674\) and \(0.2310\). Because the entire interval lies above zero, this provides evidence that treatment A has a higher 15-year survival rate than treatment B.

(c)

The estimated relative risk is:
\(\frac{\hat{p}_A}{\hat{p}_B} = \frac{38/154}{16/164} = \frac{0.2468}{0.0976} \approx 2.53\)
\(\boxed{\text{Estimated relative risk} \approx 2.53}\)
This means patients receiving treatment A are estimated to be about 2.53 times as likely to survive at least 15 years as patients receiving treatment B.

(d)

A 95% confidence interval for \(\ln\!\left(\dfrac{p_A}{p_B}\right)\) is given as \((0.3868,\ 1.4690)\).
To convert this to a confidence interval for the relative risk \(\dfrac{p_A}{p_B}\), apply the exponential (inverse of the natural log) to each endpoint:
\(e^{0.3868} \approx 1.47 \qquad \text{and} \qquad e^{1.4690} \approx 4.34\)
\(\boxed{\left(1.47,\ 4.34\right)}\)
We are 95% confident that patients receiving treatment A are between 1.47 and 4.34 times as likely to survive at least 15 years compared to patients receiving treatment B.

(e)

When the survival proportions are small (as here, approximately 0.25 and 0.10), the confidence interval for the relative risk is more informative and practically meaningful than the confidence interval for the difference in proportions.
Knowing that a patient’s chance of survival is between 1.47 and 4.34 times greater with treatment A is more vivid and easier to interpret clinically than knowing the absolute difference in proportions is somewhere between 0.07 and 0.23 — a range that may sound small even though it represents a substantial relative advantage.

Question

A large company has two shifts — a day shift and a night shift. Parts produced by the two shifts must meet the same specifications. The manager of the company believes that there is a difference in the proportions of parts produced within specifications by the two shifts. To investigate this belief, random samples of parts that were produced on each of these shifts were selected. For the day shift, 188 of its 200 selected parts met specifications. For the night shift, 180 of its 200 selected parts met specifications.
(a) Use a 96 percent confidence interval to estimate the difference in the proportions of parts produced within specifications by the two shifts.
(b) Based only on this confidence interval, do you think that the difference in the proportions of parts produced within specifications by the two shifts is significantly different from 0? Justify your answer.

Most-appropriate topic codes (AP Statistics):

• Topic 3.10 — Constructing a Confidence Interval for the Difference Between Two Population Proportions (Part \(\mathrm{a}\))
• Topic 3.11 — Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Proportions (Part \(\mathrm{b}\))
▶️ Answer/Explanation

(a)
Step 1: Identify the procedure and check conditions.
We will use a two-sample \(z\) confidence interval for \(p_D – p_N\), the difference in the proportions of parts meeting specifications for the day shift and night shift.
Let \(\hat{p}_D\) = proportion of day shift parts meeting specifications, and \(\hat{p}_N\) = proportion of night shift parts meeting specifications.
\(\hat{p}_D = \frac{188}{200} = 0.94 \qquad \hat{p}_N = \frac{180}{200} = 0.90\)
Conditions:
1. Independent random samples: The problem states that random samples of parts were selected from each shift. We assume that day shift and night shift production are independent — that is, each part is produced by one shift only and machine quality does not vary over time.
2. Large sample sizes (normal approximation valid):
\(n_D \hat{p}_D = 200(0.94) = 188 > 10\)
\(n_D(1-\hat{p}_D) = 200(0.06) = 12 > 10\)
\(n_N \hat{p}_N = 200(0.90) = 180 > 10\)
\(n_N(1-\hat{p}_N) = 200(0.10) = 20 > 10\)
All conditions are satisfied. It is reasonable to use the large-sample normal procedure.

Step 2: Compute the confidence interval.
The formula for the 96% confidence interval for \(p_D – p_N\) is:
\((\hat{p}_D – \hat{p}_N) \pm z^* \sqrt{\frac{\hat{p}_D(1-\hat{p}_D)}{n_D} + \frac{\hat{p}_N(1-\hat{p}_N)}{n_N}}\)
For a 96% confidence level, the critical value is \(z^* = 2.0537\).
\((0.94 – 0.90) \pm 2.0537\sqrt{\frac{(0.94)(0.06)}{200} + \frac{(0.90)(0.10)}{200}}\)
\(= 0.04 \pm 2.0537\sqrt{0.000282 + 0.00045}\)
\(= 0.04 \pm 2.0537\sqrt{0.000732}\)
\(= 0.04 \pm 2.0537(0.02705)\)
\(= 0.04 \pm 0.0556\)
\(\boxed{(-0.0156,\ 0.0956)}\)

Step 3: Interpret the interval.
Based on these samples, we are 96 percent confident that the true difference in the proportions of parts meeting specifications for the day shift and night shift is between \(-0.0156\) and \(0.0956\).

(b)
Since \(0\) is contained within the 96 percent confidence interval \((-0.0156,\ 0.0956)\), zero is a plausible value for the difference \(p_D – p_N\).
This means that at the \(\alpha = 0.04\) significance level, we do not have sufficient evidence to support the manager’s belief that there is a statistically significant difference between the proportions of parts meeting specifications for the two shifts.
\(\boxed{\text{No statistically significant difference detected at } \alpha = 0.04}\)

Scroll to Top