Home / AP® Exam / AP® Statistics / AP Statistics 3.8 Potential Errors When Performing Tests- Exam Style Questions – FRQs

AP Statistics 3.8 Potential Errors When Performing Tests- Exam Style Questions - FRQs - New Syllabus

Question

A city council wants to estimate the proportion of adult residents who are able to pass a standard physical fitness test. A random sample of 48 adults was selected, and each person was given the fitness test. Of the 48 adults sampled, 20 were able to pass the test.

The city council will provide funding for a new fitness center if there is convincing evidence that less than 35 percent of all adult residents are able to pass the physical fitness test. Let \(p\) represent the proportion of all adult residents in the city who are able to pass the physical fitness test.
(a) In the context of this study, describe a Type II error and its consequence.
(b) The city council decided to test the hypotheses \(H_0: p = 0.35\) versus \(H_a: p < 0.35\) using their sample data. The data resulted in a \(p\)-value of 0.97. What conclusion should the city council draw?
(c) After seeing the results of the survey, a council member remarked that the sample was flawed because it was recruited by asking for volunteers from a local running club, rather than choosing a completely random sample of city residents. Describe how this flaw would affect the inference about the proportion of all adult residents who can pass the test.

Most-appropriate topic codes (AP Statistics):

• Topic 3.8 — Potential Errors When Performing Tests (Part a)
• Topic 3.5 — Setting Up a Test for a Population Proportion (Part b)
• Topic 3.6 — p-Values (Part b)
• Topic 3.7 — Carrying Out a Test for a Population Proportion (Part b)
• Topic 1.12 — Potential Problems with Sampling (Part c)
▶️ Answer/Explanation

(a)
In the context of the study, a Type II error means failing to reject the null hypothesis that 35 percent of adult residents in the city are able to pass the fitness test when, in reality, the true proportion who can pass is actually less than 35 percent.
The real-world consequence of this error is that the city council would decide not to build or fund the new fitness center, even though the community actually needs it because the general health level is lower than desired.

(b)
Because the \(p\)-value of \(0.97\) is much larger than standard significance levels like \(\alpha = 0.05\), the council should fail to reject the null hypothesis.
There is not enough convincing evidence to conclude that the true proportion of adult residents in the city who can pass the test is less than 35 percent.
In fact, the observed sample proportion is:
\(\hat{p} = \dfrac{20}{48} \approx 0.417\)
Since \(0.417\) is actually higher than the baseline value of \(0.35\), it shifts in the opposite direction of what the alternative hypothesis was trying to establish.

(c)
Recruiting volunteers from a local running club creates a non-random, heavily biased sample because runners are typically in much better physical condition than the general public.
Consequently, the sample proportion of success \(\hat{p} = 0.417\) is almost certainly an overestimation of the true proportion \(p\) for all city residents.
This makes any resulting inference invalid, as it hides the true lack of physical fitness among the broader city population and could mistakenly convince the council that a new fitness center is unnecessary.

Question

A parent advisory board for a certain university was concerned about the effect of part-time jobs on the academic achievement of students attending the university. To obtain some information, the advisory board surveyed a simple random sample of \(200\) of the more than \(20,000\) students attending the university. Each student reported the average number of hours spent working part-time each week and his or her perception of the effect of part-time work on academic achievement. The data in the table below summarize the students’ responses by average number of hours worked per week (less than \(11\), \(11\) to \(20\), more than \(20\)) and perception of the effect of part-time work on academic achievement (positive, no effect, negative).
A chi-square test was used to determine if there is an association between the effect of part-time work on academic achievement and the average number of hours per week that students work. Computer output that resulted from performing this test is shown below.
(a) State the null and alternative hypotheses for this test.
(b) Discuss whether the conditions for a chi-square inference procedure are met for these data.
(c) Given the results from the chi-square test, what should the advisory board conclude?
(d) Based on your conclusion in part (c), which type of error (Type I or Type II) might the advisory board have made? Describe this error in the context of the question.

Most-appropriate topic codes (AP Statistics):

• Topic 3.14 — Setting Up a Chi-Square Test for Homogeneity or Independence (Part a)
• Topic 3.15 — Carrying Out a Chi-Square Test for Homogeneity or Independence (Part b)
• Topic 3.15 — Carrying Out a Chi-Square Test for Homogeneity or Independence (Part c)
• Topic 3.8 — Potential Errors When Performing Tests (Part d)
▶️ Answer/Explanation

(a)
\(H_0\): There is no association between the perceived effect of part-time work on academic achievement and the average time spent on part-time jobs.
\(H_a\): There is an association between the perceived effect of part-time work on academic achievement and the average time spent on part-time jobs.

(b)
The conditions for a chi-square inference procedure are met for these data:
1. Randomization: The problem states that the advisory board surveyed a “simple random sample” of students, so the random condition is met.
2. Expected Counts: All expected cell counts must be at least \(5\). Looking at the provided computer output, the expected counts (printed below the observed counts) are all greater than \(5\). The smallest expected count is \(6.825\), so this condition is met.

(c)
Because the \(p\)-value of \(0.007\) is less than standard significance levels (like \(\alpha = 0.05\)), we reject the null hypothesis.
The advisory board should conclude that there is convincing statistical evidence of an association between a student’s perceived effect of part-time work on academic achievement and the average number of hours per week that the student works.

(d)
Because the null hypothesis was rejected in part (c), the advisory board might have made a Type I error.
In the context of the question, a Type I error would be concluding that there is an association between the perceived effect of part-time work on academic achievement and the average time spent on part-time jobs when, in reality, there is no such association between these two variables.

Question

For many years, the medically accepted practice of giving aid to a person experiencing a heart attack was to have the person who placed the emergency call administer chest compression (CC) plus standard mouth-to-mouth resuscitation (MMR) to the heart attack patient until the emergency response team arrived. However, some researchers believed that CC alone would be a more effective approach.
In the 1990s a study was conducted in Seattle in which 518 cases were randomly assigned to treatments: 278 to CC plus standard MMR and 240 to CC alone. A total of 64 patients survived the heart attack: 29 in the group receiving CC plus standard MMR, and 35 in the group receiving CC alone. A test of significance was conducted on the following hypotheses.
\(H_0\): The survival rates for the two treatments are equal.
\(H_a\): The treatment that uses CC alone produces a higher survival rate.
This test resulted in a \(p\)-value of 0.0761.
(a) Interpret what this \(p\)-value measures in the context of this study.
(b) Based on this \(p\)-value and study design, what conclusion should be drawn in the context of this study? Use a significance level of \(\alpha = 0.05\).
(c) Based on your conclusion in part (b), which type of error, Type I or Type II, could have been made? What is one potential consequence of this error?

Most-appropriate topic codes (AP Statistics):

• Topic \(3.6\) — p-Values (Part \(\mathrm{a}\))
• Topic \(3.7\) — Carrying Out a Test for a Population Proportion (Part \(\mathrm{b}\))
• Topic \(3.8\) — Potential Errors When Performing Tests (Part \(\mathrm{c}\))

▶️ Answer/Explanation

(a)
The \(p\)-value of \(0.0761\) is the probability of observing a difference between the two sample survival proportions
\( \hat{p}_{\text{CC}} – \hat{p}_{\text{CC+MMR}} \)
as large as or larger than the one actually observed in this study (\(\frac{35}{240} – \frac{29}{278}\)), assuming that the survival rates for the two treatments are truly equal in the population.
In plain terms: if CC alone and CC + MMR actually produce the same survival rate, there is still about a \(7.61\%\) chance of seeing the CC-alone group outperform the CC + MMR group by as much as (or more than) it did in this study just by random chance alone.

(b)
Compare the \(p\)-value to the significance level:
\( p\text{-value} = 0.0761 > \alpha = 0.05 \)
Because the \(p\)-value exceeds \(\alpha\), we fail to reject \(H_0\).
There is not sufficient evidence at the \(\alpha = 0.05\) significance level to conclude that the treatment using CC alone produces a higher survival rate than CC plus standard MMR for heart attack patients.
Note that since the 518 cases were randomly assigned to the two treatments, this is a properly designed experiment — so a significant result would have allowed a causal conclusion. However, since we fail to reject \(H_0\), we simply cannot make that claim here.

(c)
Because we failed to reject \(H_0\) in part (b), the only error that could have occurred is a Type II error — failing to reject a null hypothesis that is actually false.
In this context, a Type II error would mean that CC alone truly does produce a higher survival rate than CC + MMR, but the study did not provide enough evidence to detect it.
A potential consequence of this error is that the medical community would continue recommending CC plus MMR as the standard treatment for heart attack patients, even though CC alone would actually save more lives. Heart attack patients would receive a less effective treatment, leading to preventable deaths that would not have occurred had CC alone been adopted as the standard practice.

Question

A researcher wants to conduct a study to test whether listening to soothing music for 20 minutes helps to reduce diastolic blood pressure in patients with high blood pressure, compared to simply sitting quietly in a noise-free environment for 20 minutes. One hundred patients with high blood pressure at a large medical clinic are available to participate in this study.
(a) Propose a design for this study to compare these two treatments.
(b) The null hypothesis for this study is that there is no difference in the mean reduction of diastolic blood pressure for the two treatments and the alternative hypothesis is that the mean reduction in diastolic blood pressure is greater for the music treatment. If the null hypothesis is rejected, the clinic will offer this music therapy as a free service to their patients with high blood pressure. Describe Type I and Type II errors and the consequences of each in the context of this study, and discuss which one you think is more serious.

Most-appropriate topic codes (AP Statistics):

• Topic 1.13 — Experimental Design (Part \(\mathrm{a}\))
• Topic 4.4 — Setting Up a Test for a Population Mean or Population Mean Difference (Part \(\mathrm{b}\))
• Topic 3.8 — Potential Errors When Performing Tests (Part \(\mathrm{b}\))
▶️ Answer/Explanation

(a)
Proposed Design (Completely Randomized Design):
Assign each of the 100 patients a unique number from 00 to 99. Using a random number table (or random number generator), select 50 unique numbers. The patients corresponding to the selected numbers will form Group 1 (Music Treatment); the remaining 50 patients will form Group 2 (Control — Noise-Free Environment).
Measure the diastolic blood pressure of every patient before the treatment begins. Then:
Group 1 patients sit quietly in a room where soothing music is played for 20 minutes.
Group 2 patients sit quietly in a noise-free environment for 20 minutes.
At the end of the 20-minute period, measure the diastolic blood pressure of every patient again. Compute the reduction in diastolic blood pressure (before \(-\) after) for each patient. Run a two-sample \(t\)-test to compare the mean reduction in the two groups.

Alternatively (Paired Design):
Each patient receives both treatments on two separate occasions, with a suitable washout period in between. For each patient, flip a coin to randomly assign which treatment is administered first. Compute the difference (music reduction \(-\) noise-free reduction) for each patient, and use a paired \(t\)-test to determine whether the mean difference is significantly greater than zero.

(b)
Let \(\mu_M\) = mean reduction in diastolic blood pressure under the music treatment, and \(\mu_C\) = mean reduction under the control (noise-free) treatment. The hypotheses are:
\(H_0: \mu_M = \mu_C \qquad H_a: \mu_M > \mu_C\)
Type I Error:
Rejecting \(H_0\) when it is actually true — that is, concluding that soothing music does reduce diastolic blood pressure more than sitting quietly, when in reality it does not.
Consequence: The clinic will offer music therapy as a free service to its patients even though the therapy is ineffective. This wastes clinic resources and money, providing no real medical benefit to patients.
Type II Error:
Failing to reject \(H_0\) when it is actually false — that is, concluding that soothing music does not reduce diastolic blood pressure more than sitting quietly, when in reality it does.
Consequence: The clinic will not offer music therapy, even though the therapy would genuinely help patients reduce their blood pressure. Patients are denied access to an effective, low-cost treatment.

Which error is more serious?
A reasonable case can be made for either error. One well-supported argument is that the Type II error is more serious: denying patients an effective treatment that could meaningfully improve their health — and potentially save lives — is a greater harm than wasting some clinic resources on an ineffective service. In the case of a Type I error, patients who receive music therapy are unlikely to be harmed by listening to music. But in the case of a Type II error, patients who could benefit from a real treatment are denied it altogether.

Question

When a law firm represents a group of people in a class action lawsuit and wins that lawsuit, the firm receives a percentage of the group’s monetary settlement. That settlement amount is based on the total number of people in the group—the larger the group and the larger the settlement, the more money the firm will receive.
A law firm is trying to decide whether to represent car owners in a class action lawsuit against the manufacturer of a certain make and model for a particular defect. If \(5\) percent or less of the cars of this make and model have the defect, the firm will not recover its expenses. Therefore, the firm will handle the lawsuit only if it is convinced that more than \(5\) percent of cars of this make and model have the defect. The firm plans to take a random sample of \(1{,}000\) people who bought this car and ask them if they experienced this defect in their cars.
(a) Define the parameter of interest and state the null and alternative hypotheses that the law firm should test.
(b) In the context of this situation, describe Type I and Type II errors and describe the consequences of each of these for the law firm.

Most-appropriate topic codes (AP Statistics):

• Topic 3.5 — Setting Up a Test for a Population Proportion (Part a)
• Topic 3.8 — Potential Errors When Performing Tests (Part b)
▶️ Answer/Explanation

(a)
The first step is to identify what unknown quantity the firm actually cares about. The firm isn’t interested in just the \(1{,}000\) cars in the sample — it wants to know the truth about the entire population of cars of this make and model. So let:
\( p = \) the proportion of all cars of this make and model that have the defect
Next, think about what the firm needs to be “convinced” of before it acts. The firm’s default assumption (the status quo it must be talked out of) is that the defect rate is low enough that the case isn’t worth taking — that is, \(5\%\) or less. The firm will only move forward if the evidence strongly suggests the rate is actually higher than \(5\%\). That gives:
\( H_0: p=0.05 \)
\( H_a: p>0.05 \)

(b)
To describe the errors, it helps to first restate what each hypothesis means in plain English:
\( H_0 \): \(5\%\) or less of the cars have the defect (not worth taking the case)
\( H_a \): more than \(5\%\) of the cars have the defect (worth taking the case)
A Type I error happens when \(H_0\) is true but the firm rejects it anyway:
The firm concludes that more than \(5\%\) of cars have the defect, when in reality \(5\%\) or fewer actually do.
Consequence: The firm decides to take the case based on this mistaken belief, spends time and money pursuing it, but since the true defect rate is \(5\%\) or less, the firm does not recover its expenses — resulting in a financial loss for the firm.
A Type II error happens when \(H_a\) is true but the firm fails to reject \(H_0\):
The firm is not convinced that more than \(5\%\) of cars have the defect, when in reality more than \(5\%\) actually do.
Consequence: The firm passes on the case, believing it wouldn’t be profitable, but since the true defect rate actually exceeds \(5\%\), the firm misses out on a lawsuit that could have earned them money.

Scroll to Top