Home / AP® Exam / AP® Statistics / AP Statistics 2.1 Tabular and Graphical Representations for the Distributions of Two Categorical Variables- Exam Style Questions – FRQs

AP Statistics 2.1 Tabular and Graphical Representations for the Distributions of Two Categorical Variables- Exam Style Questions - FRQs - New Syllabus

Question

A local elementary school decided to sell bottles printed with the school district’s logo as a fund-raiser. The students in the elementary school were asked to sell bottles in three different sizes (small, medium, and large). The relative frequencies of the number of bottles sold for each size by the elementary school were \(0.5\) for small bottles, \(0.3\) for medium bottles, and \(0.2\) for large bottles.
A local middle school also decided to sell bottles as a fund-raiser, using the same three sizes (small, medium, and large). The middle school students sold three times the number of bottles that the elementary school students sold. For the middle school students, the proportion of bottles sold was equal for all three sizes.
(a) Complete the segmented bar graphs representing the relative frequencies of the number of bottles sold for each size by students at each school.
(b) An administrator at the elementary school concluded that the elementary school students sold more small bottles than the middle school students did. Is the elementary school administrator’s conclusion correct? Explain your response.
Two high schools are also selling the bottles and are competing to see which one sold more large bottles.
(c) A mosaic plot for the distribution of the number of bottles sold by each of the high schools is shown here.
i. Which of the two high schools sold a greater proportion of large bottles? Justify your answer.
ii. Which of the two high schools sold a greater number of large bottles? Justify your answer.

Most-appropriate topic codes (AP Statistics):

• Topic \(2.1\) — Tabular and Graphical Representations for the Distributions of Two Categorical Variables (Parts \( \mathrm{a} \), \( \mathrm{c} \))
• Topic \(2.2\) — Summary Statistics for Two Categorical Variables (Part \( \mathrm{b} \))
▶️ Answer/Explanation

(a)


For the Elementary School bar graph, partition the segments at \(0.5\) for small bottles, \(0.8\) (\(0.5 + 0.3\)) for medium bottles, and \(1.0\) for large bottles.
For the Middle School bar graph, since the proportions are equal, partition the bar into three equal areas at approximately \(0.33\) and \(0.67\).

(b)
No, the elementary school administrator’s conclusion is incorrect.
Let \(x\) represent the total number of bottles sold by the elementary school.
This means the elementary school sold \(0.5x\) small bottles.
The middle school sold three times as many total bottles, which is \(3x\).
Since the proportion is equal across the three sizes, the middle school sold \(\frac{1}{3}(3x) = x\) small bottles.
Because \(x > 0.5x\), the middle school actually sold more small bottles than the elementary school.

(c)(i)
High School A sold a greater proportion of large bottles.
Looking at the y-axis of the mosaic plot, High School A’s proportion for large bottles is \(0.7\), which is greater than High School B’s proportion of \(0.6\).

(c)(ii)
High School B sold a greater number of large bottles.
In a mosaic plot, the total number of items is represented by the area of the segments.
Even though High School A had a larger proportion, the overall area of the rectangle representing large bottles for High School B is visibly larger than the area for High School A.

Question

To compare success rates for treating allergies at two clinics that specialize in treating allergy sufferers, researchers selected random samples of patient records from the two clinics. The following table summarizes the data.

(a) (i) Complete the following table by recording the relative frequencies of successful and unsuccessful treatments at each clinic.

(ii) Based on the relative frequency table in part (a-i), which clinic is more successful in treating allergy sufferers? Justify your answer.
(b) Based on the design of the study, would a statistically significant result allow the researchers to conclude that receiving treatments at the clinic you selected in part (a-ii) causes a higher percentage of successful treatments than at the other clinic? Explain your answer.
A physician who worked at both clinics believed that it was important to separate the patients in the study by severity of the patient’s allergy (severe or mild). The physician constructed the following mosaic plot. The values in the mosaic plot represent the number of patients who were either successfully treated or unsuccessfully treated in each allergy severity group within each clinic. For example, the value 78 represents the number of patients successfully treated in the mild group within Clinic A.
Based on the mosaic plot, the physician concluded the following:
For mild allergy sufferers, Clinic B was more successful in treating allergies.
For severe allergy sufferers, Clinic B was more successful in treating allergies.
(c) (i) For each clinic, which allergy severity is treated more successfully? Justify your answer.
• Clinic A:
• Clinic B:
(ii) For each clinic, which allergy severity is more likely to be treated? Justify your answer.
• Clinic A:
• Clinic B:
(d) Using your answers from part (c), give a reasonable explanation of why the more successful clinic identified in part (a-ii) is the same as or different from the physician’s conclusion that Clinic B is more successful in treating both severe and mild allergies.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.10\) — The Investigative Question Revisited and Data Collection (Part \( \mathrm{b} \))
• Topic \(2.1\) — Tabular and Graphical Representations for the Distributions of Two Categorical Variables (Parts \( \mathrm{a} \), \( \mathrm{c} \))
• Topic \(2.2\) — Summary Statistics for Two Categorical Variables (Parts \( \mathrm{a} \), \( \mathrm{c} \), \( \mathrm{d} \))
▶️ Answer/Explanation

(a)(i)

(a)(ii)
Clinic A is more successful. The relative frequency of successful treatments for Clinic A ($0.633$ or $63.3\%$) is greater than the relative frequency of successful treatments for Clinic B ($0.515$ or $51.5\%$).

(b)
No. The researchers selected random samples of patient records, which means this is an observational study, not a randomized experiment. Because patients were not randomly assigned to Clinic A or Clinic B, we cannot establish a cause-and-effect relationship due to the potential presence of confounding variables.

(c)(i)
Clinic A: Mild allergies are treated more successfully. The success rate for mild is $78 / (78 + 26) = 75\%$, while the success rate for severe is $11 / (11 + 24) = 31.4\%$.
Clinic B: Mild allergies are treated more successfully. The success rate for mild is $32 / (32 + 1) = 97\%$, while the success rate for severe is $10 / (10 + 25) = 28.6\%$.

(c)(ii)
Clinic A: Mild allergies are more likely to be treated. Clinic A treated 104 mild cases ($78+26$) compared to only 35 severe cases ($11+24$).
Clinic B: Severe allergies are more likely to be treated. Clinic B treated 35 severe cases ($10+25$) compared to only 33 mild cases ($32+1$).

(d)
The conclusion is different because of Simpson’s Paradox. Clinic A’s overall success rate is higher because it treats a much larger proportion of mild allergy cases, which naturally have a higher success rate regardless of the clinic. Conversely, Clinic B treats a higher proportion of severe cases, which brings its overall average down, even though it performs better than Clinic A within each specific severity group.

Question

A research center conducted a national survey about teenage behavior. Teens were asked whether they had consumed a soft drink in the past week. The following table shows the counts for three independent random samples from major cities.
(a) Suppose one teen is randomly selected from each city’s sample. A researcher claims that the likelihood of selecting a teen from Baltimore who consumed a soft drink in the past week is less than the likelihood of selecting a teen from either one of the other cities who consumed a soft drink in the past week because Baltimore has the least number of teens who consumed a soft drink. Is the researcher’s claim correct? Explain your answer.
(b) Consider the values in the table.
i. Construct a segmented bar chart of relative frequencies based on the information in the table.

ii. Which city had the smallest proportion of teens who consumed a soft drink in the previous week? Determine the value of the proportion.
(c) Consider the inference procedure that is appropriate for investigating whether there is a difference among the three cities in the proportion of all teens who consumed a soft drink in the past week.
i. Identify the appropriate inference procedure.
ii. Identify the hypotheses of the test.

Most-appropriate topic codes (AP Statistics):

• Topic \(2.1\) — Tabular and Graphical Representations for the Distributions of Two Categorical Variables (Part \( \mathrm{b} \))
• Topic \(2.2\) — Summary Statistics for Two Categorical Variables (Parts \( \mathrm{a} \), \( \mathrm{b} \))
• Topic \(3.14\) — Setting Up a Chi-Square Test for Homogeneity or Independence (Part \( \mathrm{c} \))
▶️ Answer/Explanation

(a)
No, the researcher’s claim is incorrect because comparing mere counts is meaningless when the sample sizes across the cities are unequal. Instead, we must compare proportions: the proportion for Baltimore is \( \frac{727}{904} \approx 0.804 \), which is actually greater than the proportions for Detroit (\( \frac{1232}{1663} \approx 0.741 \)) and San Diego (\( \frac{1482}{2280} = 0.65 \)).

(b)(i)
To construct the segmented bar chart, you would calculate the relative frequencies for “Yes” and “No” for each city and stack them so each bar reaches \(1.0\) on the vertical axis. For example, Baltimore’s “Yes” segment would extend up to \( 0.804 \), with the “No” segment filling the rest up to \( 1.0 \).

(b)(ii)
San Diego had the smallest proportion of teens who consumed a soft drink, with a value of \( \frac{1482}{2280} = 0.65 \).

(c)(i)
A chi-square test for homogeneity is the appropriate inference procedure because we are investigating whether the distribution of a single categorical variable (consuming a soft drink) is the same across multiple independent populations (the three cities).

(c)(ii)
The null hypothesis \(H_0\) is that there is no difference in the true proportion of all teens who consumed a soft drink in the past week across the three cities. The alternative hypothesis \(H_a\) is that there is at least one difference in the proportion of all teens who consumed a soft drink in the past week across the three cities.

Question

An administrator at a large university is interested in determining whether the residential status of a student is associated with level of participation in extracurricular activities. Residential status is categorized as on campus for students living in university housing and off campus otherwise. A simple random sample of 100 students in the university was taken, and each student was asked the following two questions.
  • Are you an on campus student or an off campus student?
  • In how many extracurricular activities do you participate?
The responses of the 100 students are summarized in the frequency table shown.

(a) Calculate the proportion of on campus students in the sample who participate in at least one extracurricular activity and the proportion of off campus students in the sample who participate in at least one extracurricular activity.
On campus proportion:
Off campus proportion:
The responses of the 100 students are summarized in the segmented bar graph shown.
(b) Write a few sentences summarizing what the graph reveals about the association between residential status and level of participation in extracurricular activities among the 100 students in the sample.
(c) After verifying that the conditions for inference were satisfied, the administrator performed a chi-square test of the following hypotheses.

\(H_0\): There is no association between residential status and level of participation in extracurricular activities among the students at the university.

\(H_a\): There is an association between residential status and level of participation in extracurricular activities among the students at the university.

The test resulted in a \(p\)-value of \(0.23\). Based on the \(p\)-value, what conclusion should the administrator make?

Most-appropriate topic codes (AP Statistics):

• Topic \(2.1\) — Tabular and Graphical Representations for the Distributions of Two Categorical Variables (Parts \( \mathrm{a} \), \( \mathrm{b} \))
• Topic \(3.14\) — Setting Up a Chi-Square Test for Homogeneity or Independence (Part \( \mathrm{c} \))
• Topic \(3.15\) — Carrying Out a Chi-Square Test for Homogeneity or Independence (Part \( \mathrm{c} \))
▶️ Answer/Explanation

(a)

For on campus students, “at least one activity” means one activity or two or more activities, so we add those counts and divide by the total number of on campus students:
\(\hat{p}_{\text{on}} = \dfrac{17 + 7}{33} = \dfrac{24}{33} \approx 0.727\)
For off campus students, we do the same:
\(\hat{p}_{\text{off}} = \dfrac{25 + 12}{67} = \dfrac{37}{67} \approx 0.552\)
\(\boxed{\hat{p}_{\text{on}} \approx 0.727, \quad \hat{p}_{\text{off}} \approx 0.552}\)

(b)
Looking at the segmented bar graph, on campus residents appear more likely to participate in extracurricular activities than off campus residents. Specifically, on campus students have a higher proportion participating in one activity (about \(51.5\%\) vs. \(37.3\%\)) and a lower proportion participating in no activities (about \(27.3\%\) vs. \(44.8\%\)). The proportions participating in two or more activities are fairly similar between the two groups (on campus: \(\approx 21.2\%\), off campus: \(\approx 17.9\%\)).

(c)
The \(p\)-value of \(0.23\) is greater than conventional significance levels such as \(\alpha = 0.05\) or \(\alpha = 0.10\). Because the \(p\)-value is large, we fail to reject the null hypothesis \(H_0\).
The sample data do not provide sufficient evidence to conclude that there is an association between residential status and level of participation in extracurricular activities among all students at the university.

Question

The table below shows the political party registration by gender of all $500$ registered voters in Franklin Township.
(a) Given that a randomly selected registered voter is a male, what is the probability that he is registered for Party Y?
(b) Among the registered voters of Franklin Township, are the events “is a male” and “is registered for Party Y” independent? Justify your answer based on probabilities calculated from the table above.
(c) One way to display the data in the table is to use a segmented bar graph. The following segmented bar graph, constructed from the data in the party registration-Franklin Township table, shows party-registration distributions for males and females in Franklin Township.
In Lawrence Township, the proportions of all registered voters for Parties W, X, and Y are the same as for Franklin Township, and party registration is independent of gender. Complete the graph below to show the distributions of party registration by gender in Lawrence Township.

Most-appropriate topic codes (AP Statistics):

• Topic 2.6 — Conditional Probability (Part a)
• Topic 2.7 — Independent Events and Unions of Events (Part b)
• Topic 2.1 — Tabular and Graphical Representations for the Distributions of Two Categorical Variables (Part c)
▶️ Answer/Explanation

(a)
We need to find the conditional probability that a voter is registered for Party Y, given they are male. Looking only at the “Male” row, there are $200$ total males, and $48$ of them are registered for Party Y.
$P(\text{Party Y} \mid \text{Male}) = \frac{48}{200} = 0.24$

(b)
No, the events “is a male” and “is registered for Party Y” are not independent.
Two events $A$ and $B$ are independent if $P(A \mid B) = P(A)$. Let’s compare the conditional probability from part (a) to the overall marginal probability of being registered for Party Y.
$P(\text{Party Y}) = \frac{168}{500} = 0.336$
Since $P(\text{Party Y} \mid \text{Male}) = 0.24$ and $P(\text{Party Y}) = 0.336$, the probabilities are not equal ($0.24 \neq 0.336$). Knowing that a randomly selected voter is male changes the probability that they are registered for Party Y, so the events are dependent.

(c)
Because party registration is independent of gender in Lawrence Township, the distribution of party registration for both males and females must be identical to the overall marginal distribution of the town.
We first calculate the overall proportions for each party (which are the same as Franklin Township’s overall proportions):
Party W: $\frac{88}{500} = 0.176$
Party X: $\frac{244}{500} = 0.488$
Party Y: $\frac{168}{500} = 0.336$
To complete the segmented bar graph, you would draw identical bars for both the Male and Female categories with the following dividing lines:
• The segment for Party W starts at $0.0$ and ends at $0.176$.
• The segment for Party X starts at $0.176$ and ends at $0.176 + 0.488 = 0.664$.
• The segment for Party Y starts at $0.664$ and extends to $1.0$.

Question

An advertising agency in a large city is conducting a survey of adults to investigate whether there is an association between highest level of educational achievement and primary source for news. The company takes a random sample of 2,500 adults in the city. The results are shown in the table below.
(a) If an adult is to be selected at random from this sample, what is the probability that the selected adult is a college graduate or obtains news primarily from the internet?
(b) If an adult who is a college graduate is to be selected at random from this sample, what is the probability that the selected adult obtains news primarily from the internet?
(c) When selecting an adult at random from the sample of 2,500 adults, are the events “is a college graduate” and “obtains news primarily from the internet” independent? Justify your answer.

(d) The company wants to conduct a statistical test to investigate whether there is an association between educational achievement and primary source for news for adults in the city. What is the name of the statistical test that should be used?

What are the appropriate degrees of freedom for this test?

Most-appropriate topic codes (AP Statistics):

• Topic \(2.1\) — Tabular and Graphical Representations for the Distributions of Two Categorical Variables (Parts \(\mathrm{a}\), \(\mathrm{b}\), \(\mathrm{c}\))
• Topic \(2.6\) — Conditional Probability (Part \(\mathrm{b}\))
• Topic \(2.7\) — Independent Events and Unions of Events (Parts \(\mathrm{a}\), \(\mathrm{c}\))
• Topic \(3.14\) — Setting Up a Chi-Square Test for Homogeneity or Independence (Part \(\mathrm{d}\))
▶️ Answer/Explanation

(a)
Let \(C\) = event that the adult is a college graduate, and \(I\) = event that the adult obtains news primarily from the internet.
Using the Addition Rule:
\(P(C \cup I) = P(C) + P(I) – P(C \cap I)\)
Reading the values directly from the table:
\(P(C) = \frac{693}{2500}, \qquad P(I) = \frac{687}{2500}, \qquad P(C \cap I) = \frac{245}{2500}\)
\(P(C \cup I) = \frac{693}{2500} + \frac{687}{2500} – \frac{245}{2500} = \frac{693 + 687 – 245}{2500} = \frac{1135}{2500}\)
\(\boxed{P(C \cup I) = \frac{1135}{2500} = 0.454}\)
Don’t forget to subtract the overlap — college graduates who use the internet get counted in both the college graduate total and the internet total, so we subtract them once to avoid double-counting.

(b)
We want the conditional probability that an adult obtains news from the internet, given that the adult is a college graduate. From the table, among the 693 college graduates, 245 primarily use the internet:
\(P(I \mid C) = \frac{P(C \cap I)}{P(C)} = \frac{\dfrac{245}{2500}}{\dfrac{693}{2500}} = \frac{245}{693}\)
\(\boxed{P(I \mid C) = \frac{245}{693} \approx 0.354}\)
This is a conditional probability — we’ve already restricted our pool to only the 693 college graduates, so 693 becomes the new denominator. The 2,500 total cancels out entirely.

(c)
Two events are independent if and only if \(P(A \cap B) = P(A) \cdot P(B)\), which is equivalent to checking whether \(P(I \mid C) = P(I)\).
From the table:
\(P(I) = \frac{687}{2500} = 0.275\)
\(P(I \mid C) = \frac{245}{693} \approx 0.354\)
Since \(P(I \mid C) \approx 0.354 \neq 0.275 = P(I)\), the two events are not independent.
We can also verify using the multiplication rule directly:
\(P(C) \cdot P(I) = \frac{693}{2500} \times \frac{687}{2500} = \frac{476{,}091}{6{,}250{,}000} \approx 0.0762\)
\(P(C \cap I) = \frac{245}{2500} = 0.098\)
Since \(0.098 \neq 0.0762\), the events are confirmed to be not independent. In real terms, college graduates are noticeably more likely to get their news from the internet than the general adult population — that difference in rates is exactly what “not independent” means here.

(d)
The appropriate test is the Chi-Square Test of Association (or Independence).
This test is used when we want to determine whether there is an association between two categorical variables — here, educational achievement (3 categories) and primary news source (5 categories).
The degrees of freedom are calculated as:
\(\text{df} = (\text{number of rows} – 1) \times (\text{number of columns} – 1)\)
\(\text{df} = (5 – 1) \times (3 – 1) = 4 \times 2 = \boxed{8}\)
There are 5 rows (news source categories) and 3 columns (education levels), not counting the totals row and column. The degrees of freedom formula captures how many cells in the table are “free to vary” once the row and column totals are fixed.

Question

A simple random sample of 100 high school seniors was selected from a large school district. The gender of each student was recorded, and each student was asked the following questions.
1. Have you ever had a part-time job?
2. If you answered yes to the previous question, was your part-time job in the summer only?
The responses are summarized in the table below.

(a) On the grid below, construct a graphical display that represents the association between gender and job experience for the students in the sample.
(b) Write a few sentences summarizing what the display in part (a) reveals about the association between gender and job experience for the students in the sample.
(c) Which test of significance should be used to test if there is an association between gender and job experience for the population of high school seniors in the district?
State the null and alternative hypotheses for the test, but do not perform the test.

Most-appropriate topic codes (AP Statistics):

• Topic \(2.1\) — Tabular and Graphical Representations for the Distributions of Two Categorical Variables (Parts \(\mathrm{a}\), \(\mathrm{b}\))
• Topic \(2.2\) — Summary Statistics for Two Categorical Variables (Part \(\mathrm{b}\))
• Topic \(3.14\) — Setting Up a Chi-Square Test for Homogeneity or Independence (Part \(\mathrm{c}\))

▶️ Answer/Explanation

(a)
First, convert the raw counts to percentages within each gender so that males and females are directly comparable:
For Males \((n = 48)\):
• Never had a part-time job: \(\dfrac{21}{48} \approx 43.8\%\)
• Part-time job during summer only: \(\dfrac{15}{48} \approx 31.2\%\)
• Part-time job, not only during summer: \(\dfrac{12}{48} = 25.0\%\)
For Females \((n = 52)\):
• Never had a part-time job: \(\dfrac{31}{52} \approx 59.6\%\)
• Part-time job during summer only: \(\dfrac{13}{52} = 25.0\%\)
• Part-time job, not only during summer: \(\dfrac{8}{52} \approx 15.4\%\)

A side-by-side bar graph (or segmented bar graph) should be drawn with the three job-experience categories on the horizontal axis, percentage on the vertical axis, and separate bars (or segments) for Male and Female within each category. All axes must be labeled.

(b)
Looking at the bar graph, females were considerably more likely than males to have never had a part-time job — about \(59.6\%\) of females compared with \(43.8\%\) of males fall into that category.
On the other hand, males were more likely than females to have had a part-time job during the summer only (\(31.2\%\) vs. \(25.0\%\)), and more likely to have had a part-time job that was not only during the summer (\(25.0\%\) vs. \(15.4\%\)).
If there were no association between gender and job experience, we’d expect the bars for males and females to be roughly the same height in each category — but they clearly aren’t, so the sample data suggest there is an association between gender and job experience.

(c)
The appropriate test of significance is the chi-square test of association (or independence).
The chi-square test statistic is:
\( \chi^2 = \sum \frac{(\text{observed} – \text{expected})^2}{\text{expected}} \)
The null and alternative hypotheses are:
\(H_0\): There is no association between gender and job experience (for the population of high school seniors in the district).
\(H_a\): There is an association between gender and job experience (for the population of high school seniors in the district).
Equivalently:
\(H_0\): Gender and job experience are independent.
\(H_a\): Gender and job experience are not independent.
Since we have two categorical variables each with multiple categories, the chi-square test of association/independence is the correct choice — a two-sample \(z\)-test for proportions only compares two groups on a binary outcome, and wouldn’t capture all three job-experience categories at once.

Question

A simple random sample of adults living in a suburb of a large city was selected. The age and annual income of each adult in the sample were recorded. The resulting data are summarized in the table below.

(a) What is the probability that a person chosen at random from those in this sample will be in the 31–45 age category?
(b) What is the probability that a person chosen at random from those in this sample whose incomes are over \(\$50{,}000\) will be in the 31–45 age category? Show your work.
(c) Based on your answers to parts (a) and (b), is annual income independent of age category for those in this sample? Explain.

Most-appropriate topic codes (AP Statistics):

• Topic 2.1 — Tabular and Graphical Representations for the Distributions of Two Categorical Variables (All parts)
• Topic 2.6 — Conditional Probability (Part b)
• Topic 2.7 — Independent Events and Unions of Events (Part c)
• Topic 2.2 — Summary Statistics for Two Categorical Variables (Part a)
▶️ Answer/Explanation

(a)

We want the probability that a randomly selected person from the sample falls in the 31–45 age category.
The total number of people in the sample is 207, and the number in the 31–45 age group is 89.
\( P(\text{age } 31\text{–}45) = \frac{89}{207} \approx 0.42995 \)
\(\boxed{P(\text{age } 31\text{–}45) \approx 0.4300}\)

(b)

We now want the conditional probability that a person is in the 31–45 age category, given that their income is over \(\$50{,}000\).
From the table, the total number of people with income over \(\$50{,}000\) is 96, and among those, 35 are in the 31–45 age group.
\( P(\text{age } 31\text{–}45 \mid \text{income over } \$50{,}000) = \frac{35}{96} \approx 0.36458 \)
\(\boxed{P(\text{age } 31\text{–}45 \mid \text{income over } \$50{,}000) \approx 0.3646}\)

(c)

For two variables to be independent, knowing the value of one variable should not change the probability of the other — in other words, the marginal probability and the conditional probability must be equal.
From part (a), \(P(\text{age } 31\text{–}45) \approx 0.4300\), and from part (b), \(P(\text{age } 31\text{–}45 \mid \text{income over } \$50{,}000) \approx 0.3646\).
Since these two probabilities are not equal (\(0.4300 \neq 0.3646\)), knowing a person’s income category does change the probability of being in the 31–45 age group.
Therefore, annual income and age category are not independent for those in this sample.

Scroll to Top