AP Statistics 2.2 Summary Statistics for Two Categorical Variables- Exam Style Questions - FRQs - New Syllabus
Question

ii. State the appropriate null and alternative hypotheses for the hypothesis test you identified in (c-i). Do not perform the hypothesis test.
Most-appropriate topic codes (AP Statistics):
• Topic \(3.14\) — Setting Up a Chi-Square Test for Homogeneity or Independence (Part \( \mathrm{c} \))
• Topic \(3.15\) — Carrying Out a Chi-Square Test for Homogeneity or Independence (Part \( \mathrm{d} \))
▶️ Answer/Explanation
(a)
To find this probability, we sum the number of collectors who have a majority of regular cards AND have been collecting for \(11\) or more months (which covers the \(11-15\), \(16-20\), and \(21+\) columns).
Number of collectors \(= 71 + 76 + 112 = 259\).
\(P(\ge 11\text{ months and majority regular}) = \dfrac{259}{500} = 0.518\).
(b)
This is a conditional probability. We restrict our focus entirely to the column representing collectors with fewer than \(6\) months of collecting, which gives us a new total of \(91\) collectors.
Out of those \(91\) collectors, \(80\) have a majority of regular baseball cards.
\(P(\text{majority regular} \mid < 6\text{ months}) = \dfrac{80}{91} \approx 0.879\).
(c)
i. Because Michelle took a single random sample and is comparing two categorical variables from that single sample, she should use a chi-square test for independence.
ii. Null Hypothesis (\(H_0\)): There is no association between the number of months spent collecting baseball cards and majority card status for all baseball card collectors at the convention.
Alternative Hypothesis (\(H_a\)): There is an association between the number of months spent collecting baseball cards and majority card status for all baseball card collectors at the convention.
(d)
Because the \(p\)-value of \(0.0075\) is smaller than any reasonable significance level (such as \(\alpha = 0.05\)), Michelle should reject the null hypothesis.
The data provide convincing statistical evidence that there is a relationship between the number of months spent collecting baseball cards and which type of card is the majority in the collection for all baseball card collectors at the convention.
Question


ii. Which of the two high schools sold a greater number of large bottles? Justify your answer.
Most-appropriate topic codes (AP Statistics):
• Topic \(2.2\) — Summary Statistics for Two Categorical Variables (Part \( \mathrm{b} \))
▶️ Answer/Explanation
(a)

For the Elementary School bar graph, partition the segments at \(0.5\) for small bottles, \(0.8\) (\(0.5 + 0.3\)) for medium bottles, and \(1.0\) for large bottles.
For the Middle School bar graph, since the proportions are equal, partition the bar into three equal areas at approximately \(0.33\) and \(0.67\).
(b)
No, the elementary school administrator’s conclusion is incorrect.
Let \(x\) represent the total number of bottles sold by the elementary school.
This means the elementary school sold \(0.5x\) small bottles.
The middle school sold three times as many total bottles, which is \(3x\).
Since the proportion is equal across the three sizes, the middle school sold \(\frac{1}{3}(3x) = x\) small bottles.
Because \(x > 0.5x\), the middle school actually sold more small bottles than the elementary school.
(c)(i)
High School A sold a greater proportion of large bottles.
Looking at the y-axis of the mosaic plot, High School A’s proportion for large bottles is \(0.7\), which is greater than High School B’s proportion of \(0.6\).
(c)(ii)
High School B sold a greater number of large bottles.
In a mosaic plot, the total number of items is represented by the area of the segments.
Even though High School A had a larger proportion, the overall area of the rectangle representing large bottles for High School B is visibly larger than the area for High School A.
Question



For mild allergy sufferers, Clinic B was more successful in treating allergies.
For severe allergy sufferers, Clinic B was more successful in treating allergies.
• Clinic B:
• Clinic B:
Most-appropriate topic codes (AP Statistics):
• Topic \(2.1\) — Tabular and Graphical Representations for the Distributions of Two Categorical Variables (Parts \( \mathrm{a} \), \( \mathrm{c} \))
• Topic \(2.2\) — Summary Statistics for Two Categorical Variables (Parts \( \mathrm{a} \), \( \mathrm{c} \), \( \mathrm{d} \))
▶️ Answer/Explanation
(a)(i)

(a)(ii)
Clinic A is more successful. The relative frequency of successful treatments for Clinic A ($0.633$ or $63.3\%$) is greater than the relative frequency of successful treatments for Clinic B ($0.515$ or $51.5\%$).
(b)
No. The researchers selected random samples of patient records, which means this is an observational study, not a randomized experiment. Because patients were not randomly assigned to Clinic A or Clinic B, we cannot establish a cause-and-effect relationship due to the potential presence of confounding variables.
(c)(i)
• Clinic A: Mild allergies are treated more successfully. The success rate for mild is $78 / (78 + 26) = 75\%$, while the success rate for severe is $11 / (11 + 24) = 31.4\%$.
• Clinic B: Mild allergies are treated more successfully. The success rate for mild is $32 / (32 + 1) = 97\%$, while the success rate for severe is $10 / (10 + 25) = 28.6\%$.
(c)(ii)
• Clinic A: Mild allergies are more likely to be treated. Clinic A treated 104 mild cases ($78+26$) compared to only 35 severe cases ($11+24$).
• Clinic B: Severe allergies are more likely to be treated. Clinic B treated 35 severe cases ($10+25$) compared to only 33 mild cases ($32+1$).
(d)
The conclusion is different because of Simpson’s Paradox. Clinic A’s overall success rate is higher because it treats a much larger proportion of mild allergy cases, which naturally have a higher success rate regardless of the clinic. Conversely, Clinic B treats a higher proportion of severe cases, which brings its overall average down, even though it performs better than Clinic A within each specific severity group.
Question


ii. Which city had the smallest proportion of teens who consumed a soft drink in the previous week? Determine the value of the proportion.
ii. Identify the hypotheses of the test.
Most-appropriate topic codes (AP Statistics):
• Topic \(2.2\) — Summary Statistics for Two Categorical Variables (Parts \( \mathrm{a} \), \( \mathrm{b} \))
• Topic \(3.14\) — Setting Up a Chi-Square Test for Homogeneity or Independence (Part \( \mathrm{c} \))
▶️ Answer/Explanation
(a)
No, the researcher’s claim is incorrect because comparing mere counts is meaningless when the sample sizes across the cities are unequal. Instead, we must compare proportions: the proportion for Baltimore is \( \frac{727}{904} \approx 0.804 \), which is actually greater than the proportions for Detroit (\( \frac{1232}{1663} \approx 0.741 \)) and San Diego (\( \frac{1482}{2280} = 0.65 \)).
(b)(i)
To construct the segmented bar chart, you would calculate the relative frequencies for “Yes” and “No” for each city and stack them so each bar reaches \(1.0\) on the vertical axis. For example, Baltimore’s “Yes” segment would extend up to \( 0.804 \), with the “No” segment filling the rest up to \( 1.0 \).

(b)(ii)
San Diego had the smallest proportion of teens who consumed a soft drink, with a value of \( \frac{1482}{2280} = 0.65 \).
(c)(i)
A chi-square test for homogeneity is the appropriate inference procedure because we are investigating whether the distribution of a single categorical variable (consuming a soft drink) is the same across multiple independent populations (the three cities).
(c)(ii)
The null hypothesis \(H_0\) is that there is no difference in the true proportion of all teens who consumed a soft drink in the past week across the three cities. The alternative hypothesis \(H_a\) is that there is at least one difference in the proportion of all teens who consumed a soft drink in the past week across the three cities.
Question


Most-appropriate topic codes (AP Statistics):
• Topic \(2.2\) — Summary Statistics for Two Categorical Variables (Part \(\mathrm{b}\))
• Topic \(3.14\) — Setting Up a Chi-Square Test for Homogeneity or Independence (Part \(\mathrm{c}\))
▶️ Answer/Explanation
(a)
First, convert the raw counts to percentages within each gender so that males and females are directly comparable:
For Males \((n = 48)\):
• Never had a part-time job: \(\dfrac{21}{48} \approx 43.8\%\)
• Part-time job during summer only: \(\dfrac{15}{48} \approx 31.2\%\)
• Part-time job, not only during summer: \(\dfrac{12}{48} = 25.0\%\)
For Females \((n = 52)\):
• Never had a part-time job: \(\dfrac{31}{52} \approx 59.6\%\)
• Part-time job during summer only: \(\dfrac{13}{52} = 25.0\%\)
• Part-time job, not only during summer: \(\dfrac{8}{52} \approx 15.4\%\)
A side-by-side bar graph (or segmented bar graph) should be drawn with the three job-experience categories on the horizontal axis, percentage on the vertical axis, and separate bars (or segments) for Male and Female within each category. All axes must be labeled.

(b)
Looking at the bar graph, females were considerably more likely than males to have never had a part-time job — about \(59.6\%\) of females compared with \(43.8\%\) of males fall into that category.
On the other hand, males were more likely than females to have had a part-time job during the summer only (\(31.2\%\) vs. \(25.0\%\)), and more likely to have had a part-time job that was not only during the summer (\(25.0\%\) vs. \(15.4\%\)).
If there were no association between gender and job experience, we’d expect the bars for males and females to be roughly the same height in each category — but they clearly aren’t, so the sample data suggest there is an association between gender and job experience.
(c)
The appropriate test of significance is the chi-square test of association (or independence).
The chi-square test statistic is:
\( \chi^2 = \sum \frac{(\text{observed} – \text{expected})^2}{\text{expected}} \)
The null and alternative hypotheses are:
\(H_0\): There is no association between gender and job experience (for the population of high school seniors in the district).
\(H_a\): There is an association between gender and job experience (for the population of high school seniors in the district).
Equivalently:
\(H_0\): Gender and job experience are independent.
\(H_a\): Gender and job experience are not independent.
Since we have two categorical variables each with multiple categories, the chi-square test of association/independence is the correct choice — a two-sample \(z\)-test for proportions only compares two groups on a binary outcome, and wouldn’t capture all three job-experience categories at once.
Question

Most-appropriate topic codes (AP Statistics):
• Topic 2.6 — Conditional Probability (Part b)
• Topic 2.7 — Independent Events and Unions of Events (Part c)
• Topic 2.2 — Summary Statistics for Two Categorical Variables (Part a)
▶️ Answer/Explanation
(a)
We want the probability that a randomly selected person from the sample falls in the 31–45 age category.
The total number of people in the sample is 207, and the number in the 31–45 age group is 89.
\( P(\text{age } 31\text{–}45) = \frac{89}{207} \approx 0.42995 \)
\(\boxed{P(\text{age } 31\text{–}45) \approx 0.4300}\)
(b)
We now want the conditional probability that a person is in the 31–45 age category, given that their income is over \(\$50{,}000\).
From the table, the total number of people with income over \(\$50{,}000\) is 96, and among those, 35 are in the 31–45 age group.
\( P(\text{age } 31\text{–}45 \mid \text{income over } \$50{,}000) = \frac{35}{96} \approx 0.36458 \)
\(\boxed{P(\text{age } 31\text{–}45 \mid \text{income over } \$50{,}000) \approx 0.3646}\)
(c)
For two variables to be independent, knowing the value of one variable should not change the probability of the other — in other words, the marginal probability and the conditional probability must be equal.
From part (a), \(P(\text{age } 31\text{–}45) \approx 0.4300\), and from part (b), \(P(\text{age } 31\text{–}45 \mid \text{income over } \$50{,}000) \approx 0.3646\).
Since these two probabilities are not equal (\(0.4300 \neq 0.3646\)), knowing a person’s income category does change the probability of being in the 31–45 age group.
Therefore, annual income and age category are not independent for those in this sample.
