AP Statistics 3.14 Setting Up a Chi-Square Test for Homogeneity or Independence- Exam Style Questions - FRQs - New Syllabus
Question

ii. State the appropriate null and alternative hypotheses for the hypothesis test you identified in (c-i). Do not perform the hypothesis test.
Most-appropriate topic codes (AP Statistics):
• Topic \(3.14\) — Setting Up a Chi-Square Test for Homogeneity or Independence (Part \( \mathrm{c} \))
• Topic \(3.15\) — Carrying Out a Chi-Square Test for Homogeneity or Independence (Part \( \mathrm{d} \))
▶️ Answer/Explanation
(a)
To find this probability, we sum the number of collectors who have a majority of regular cards AND have been collecting for \(11\) or more months (which covers the \(11-15\), \(16-20\), and \(21+\) columns).
Number of collectors \(= 71 + 76 + 112 = 259\).
\(P(\ge 11\text{ months and majority regular}) = \dfrac{259}{500} = 0.518\).
(b)
This is a conditional probability. We restrict our focus entirely to the column representing collectors with fewer than \(6\) months of collecting, which gives us a new total of \(91\) collectors.
Out of those \(91\) collectors, \(80\) have a majority of regular baseball cards.
\(P(\text{majority regular} \mid < 6\text{ months}) = \dfrac{80}{91} \approx 0.879\).
(c)
i. Because Michelle took a single random sample and is comparing two categorical variables from that single sample, she should use a chi-square test for independence.
ii. Null Hypothesis (\(H_0\)): There is no association between the number of months spent collecting baseball cards and majority card status for all baseball card collectors at the convention.
Alternative Hypothesis (\(H_a\)): There is an association between the number of months spent collecting baseball cards and majority card status for all baseball card collectors at the convention.
(d)
Because the \(p\)-value of \(0.0075\) is smaller than any reasonable significance level (such as \(\alpha = 0.05\)), Michelle should reject the null hypothesis.
The data provide convincing statistical evidence that there is a relationship between the number of months spent collecting baseball cards and which type of card is the majority in the collection for all baseball card collectors at the convention.
Question


ii. Which city had the smallest proportion of teens who consumed a soft drink in the previous week? Determine the value of the proportion.
ii. Identify the hypotheses of the test.
Most-appropriate topic codes (AP Statistics):
• Topic \(2.2\) — Summary Statistics for Two Categorical Variables (Parts \( \mathrm{a}\), \( \mathrm{b} \))
• Topic \(3.14\) — Setting Up a Chi-Square Test for Homogeneity or Independence (Part \( \mathrm{c} \))
▶️ Answer/Explanation
(a)
No, the researcher’s claim is incorrect because comparing mere counts is meaningless when the sample sizes across the cities are unequal. Instead, we must compare proportions: the proportion for Baltimore is \( \frac{727}{904} \approx 0.804 \), which is actually greater than the proportions for Detroit (\( \frac{1232}{1663} \approx 0.741 \)) and San Diego (\( \frac{1482}{2280} = 0.65 \)).
(b)(i)
To construct the segmented bar chart, you would calculate the relative frequencies for “Yes” and “No” for each city and stack them so each bar reaches \(1.0\) on the vertical axis. For example, Baltimore’s “Yes” segment would extend up to \( 0.804 \), with the “No” segment filling the rest up to \( 1.0 \).

(b)(ii)
San Diego had the smallest proportion of teens who consumed a soft drink, with a value of \( \frac{1482}{2280} = 0.65 \).
(c)(i)
A chi-square test for homogeneity is the appropriate inference procedure because we are investigating whether the distribution of a single categorical variable (consuming a soft drink) is the same across multiple independent populations (the three cities).
(c)(ii)
The null hypothesis \(H_0\) is that there is no difference in the true proportion of all teens who consumed a soft drink in the past week across the three cities. The alternative hypothesis \(H_a\) is that there is at least one difference in the proportion of all teens who consumed a soft drink in the past week across the three cities.
Question


Most-appropriate topic codes (AP Statistics):
• Topic \(3.15\) — Carrying Out a Chi-Square Test for Homogeneity or Independence (Entire Question)
▶️ Answer/Explanation
Step 1: State the Hypotheses
\(H_0\): Age group at diagnosis and gender are independent (not associated) in the population of people currently being treated for schizophrenia.
\(H_a\): Age group at diagnosis and gender are not independent (i.e., there is an association) in the population of people currently being treated for schizophrenia.
Step 2: Identify the Test and Check Conditions
The appropriate test is a chi-square test of independence.
The formula for the test statistic is:
\( \chi^2 = \sum \dfrac{(O – E)^2}{E} \Blocks\
Condition 1 — Random Sample: The problem states the sample was randomly selected.
Condition 2 — Large Expected Counts: All expected cell counts must be at least 5. The expected count for each cell is computed as:
\( E = \dfrac{(\text{row total}) \times (\text{column total})}{\text{grand total}} \)
The full table of expected counts (shown below observed counts) is:

All eight expected counts are at least 5, so the condition is satisfied.
Step 3: Calculate the Test Statistic and \(p\)-value
The degrees of freedom are:
\( df = (\text{rows} – 1)(\text{columns} – 1) = (2-1)(4-1) = 3 \)
The chi-square statistic is:
\( \chi^2 = \dfrac{(46-56.91)^2}{56.91} + \dfrac{(40-36.22)^2}{36.22} + \dfrac{(21-17.25)^2}{17.25} + \dfrac{(12-8.62)^2}{8.62} \)
\( \quad\quad + \dfrac{(53-42.09)^2}{42.09} + \dfrac{(23-26.78)^2}{26.78} + \dfrac{(9-12.75)^2}{12.75} + \dfrac{(3-6.38)^2}{6.38} \)
\( \chi^2 = 2.093 + 0.395 + 0.817 + 1.322 + 2.830 + 0.534 + 1.105 + 1.788 \)
\( \boxed{\chi^2 = 10.884} \)
The \(p\)-value is:
\( p\text{-value} = P(\chi^2 \geq 10.884) = 0.012 \quad \text{with } df = 3 \)
Step 4: State the Conclusion in Context
Since the \(p\)-value of \(0.012\) is less than the significance level \(\alpha = 0.05\), we reject \(H_0\).
The data provide convincing statistical evidence that there is an association between age group at diagnosis and gender in the population of people currently being treated for schizophrenia. In plain terms, men and women tend to be diagnosed at different ages — men are more concentrated in the 20–29 age group relative to what we’d expect if there were no relationship, while women are more spread across older age groups.
Question

Most-appropriate topic codes (AP Statistics):
• Topic \(3.14\) — Setting Up a Chi-Square Test for Homogeneity or Independence (Part \( \mathrm{a} \))
• Topic \(3.15\) — Carrying Out a Chi-Square Test for Homogeneity or Independence (Part \( \mathrm{a} \))
▶️ Answer/Explanation
(a)
Step 1: State hypotheses.
\(H_0\): There is no association between the type of ad viewed and children’s choice of snack (the proportion choosing each snack is the same regardless of which ad is viewed).
\(H_a\): There is an association between the type of ad viewed and children’s choice of snack (the proportions differ based on which ad is viewed).
Step 2: Identify the procedure and check conditions.
The appropriate procedure is a chi-square test of homogeneity (also acceptable: chi-square test of independence).
Conditions:
1. Random: The 75 children were randomly assigned to the three groups — this condition is satisfied.
2. Large Counts: All expected cell counts must be at least 5. The expected counts are computed as:
\( E = \frac{(\text{row total}) \times (\text{column total})}{\text{table total}} \)
Expected counts table (observed counts with expected counts in parentheses):

The smallest expected count is \(6.33 \geq 5\), so the large counts condition is satisfied. Both conditions are met.
Step 3: Calculate the test statistic and \(p\)-value.
The chi-square test statistic is:
\( \chi^2 = \sum \frac{(O – E)^2}{E} \)
\( \chi^2 \approx \frac{(21-18.67)^2}{18.67} + \frac{(4-6.33)^2}{6.33} + \frac{(13-18.67)^2}{18.67} + \frac{(12-6.33)^2}{6.33} + \frac{(22-18.67)^2}{18.67} + \frac{(3-6.33)^2}{6.33} \)
\( \chi^2 \approx 0.292 + 0.860 + 1.720 + 5.070 + 0.595 + 1.754 \approx 10.291 \)
Degrees of freedom: \(df = (r-1)(c-1) = (3-1)(2-1) = 2\)
The \(p\)-value: \(P\!\left(\chi^2_{\,df=2} \geq 10.291\right) \approx 0.006\)
Step 4: State a conclusion in context.
Because the \(p\)-value \(\approx 0.006\) is much smaller than \(\alpha = 0.05\), we reject \(H_0\). The data provide convincing statistical evidence that there is an association between the type of ad viewed and children’s choice of snack, among all children similar to those who participated in the experiment.
(b)
When neither ad was shown (Group C), \(\dfrac{22}{25} = 88\%\) of the children chose Choco-Zuties, and only 12% chose Apple-Zuties — this reflects the baseline preference without any advertising.
When children saw the Choco-Zuties ad (Group A), 84% still chose Choco-Zuties, which is very close to the 88% in the no-ad group. So the Choco-Zuties ad had very little effect on children’s snack choice — they were already inclined toward the sugary option anyway.
When children saw the Apple-Zuties ad (Group B), only \(\dfrac{13}{25} = 52\%\) chose Choco-Zuties, and 48% chose Apple-Zuties. This is a large shift compared to the 12% who chose Apple-Zuties in the no-ad group, meaning the Apple-Zuties ad had a substantial effect in increasing children’s likelihood of choosing the healthy snack.
Question
- Are you an on campus student or an off campus student?
- In how many extracurricular activities do you participate?


\(H_0\): There is no association between residential status and level of participation in extracurricular activities among the students at the university.
\(H_a\): There is an association between residential status and level of participation in extracurricular activities among the students at the university.
Most-appropriate topic codes (AP Statistics):
• Topic \(3.14\) — Setting Up a Chi-Square Test for Homogeneity or Independence (Part \( \mathrm{c} \))
• Topic \(3.15\) — Carrying Out a Chi-Square Test for Homogeneity or Independence (Part \( \mathrm{c} \))
▶️ Answer/Explanation
(a)
For on campus students, “at least one activity” means one activity or two or more activities, so we add those counts and divide by the total number of on campus students:
\(\hat{p}_{\text{on}} = \dfrac{17 + 7}{33} = \dfrac{24}{33} \approx 0.727\)
For off campus students, we do the same:
\(\hat{p}_{\text{off}} = \dfrac{25 + 12}{67} = \dfrac{37}{67} \approx 0.552\)
\(\boxed{\hat{p}_{\text{on}} \approx 0.727, \quad \hat{p}_{\text{off}} \approx 0.552}\)
(b)
Looking at the segmented bar graph, on campus residents appear more likely to participate in extracurricular activities than off campus residents. Specifically, on campus students have a higher proportion participating in one activity (about \(51.5\%\) vs. \(37.3\%\)) and a lower proportion participating in no activities (about \(27.3\%\) vs. \(44.8\%\)). The proportions participating in two or more activities are fairly similar between the two groups (on campus: \(\approx 21.2\%\), off campus: \(\approx 17.9\%\)).
(c)
The \(p\)-value of \(0.23\) is greater than conventional significance levels such as \(\alpha = 0.05\) or \(\alpha = 0.10\). Because the \(p\)-value is large, we fail to reject the null hypothesis \(H_0\).
The sample data do not provide sufficient evidence to conclude that there is an association between residential status and level of participation in extracurricular activities among all students at the university.
Question

Most-appropriate topic codes (AP Statistics):
• Topic \(3.15\) — Carrying Out a Chi-Square Test for Homogeneity or Independence (Step 3, Step 4)
▶️ Answer/Explanation
Step 1: State the hypotheses.
\(H_0\): There is no association between age group and whether or not a person consumes five or more servings of fruits and vegetables per day for adults in the United States (i.e., the two variables are independent).
\(H_a\): There is an association between age group and whether or not a person consumes five or more servings of fruits and vegetables per day for adults in the United States (i.e., the two variables are not independent).
Step 2: Identify the procedure and check conditions.
The appropriate test is a chi-square test of independence, with test statistic:
$\chi^2 = \sum \frac{(O – E)^2}{E}$
where the expected count for each cell is computed as:
$E = \frac{(\text{row total}) \times (\text{column total})}{\text{table total}}$
Condition 1 — Random sample: The problem states that the sample was randomly selected.
Condition 2 — Expected counts: All six expected counts are at least \(5\), as shown below:

The smallest expected count is \(240.2 \geq 5\).
Step 3: Compute the test statistic and \(p\)-value.
Degrees of freedom: \(df = (3-1)(2-1) = 2\)
$\chi^2 = \frac{(231-240.2)^2}{240.2} + \frac{(741-731.8)^2}{731.8} + \frac{(669-719.4)^2}{719.4} + \frac{(2242-2191.6)^2}{2191.6} + \frac{(1291-1231.4)^2}{1231.4} + \frac{(3692-3751.6)^2}{3751.6}$
$\chi^2 = 0.353 + 0.116 + 3.528 + 1.158 + 2.883 + 0.946$
$\boxed{\chi^2 = 8.983}$
$\boxed{p\text{-value} = P(\chi^2 \geq 8.983) \approx 0.011}$
Step 4: State the conclusion in context.
Since the \(p\)-value of \(0.011\) is less than \(\alpha = 0.05\), we reject \(H_0\). The data provide convincing statistical evidence that there is an association between age group and whether or not a person consumes five or more servings of fruits and vegetables per day for adults in the United States. In particular, adults aged 55 or older were more likely to consume five or more servings per day, while middle-aged adults (35–54 years) were less likely to do so relative to what would be expected under independence.
Question


Most-appropriate topic codes (AP Statistics):
• Topic 3.15 — Carrying Out a Chi-Square Test for Homogeneity or Independence (Part b)
• Topic 3.15 — Carrying Out a Chi-Square Test for Homogeneity or Independence (Part c)
• Topic 3.8 — Potential Errors When Performing Tests (Part d)
▶️ Answer/Explanation
(a)
\(H_0\): There is no association between the perceived effect of part-time work on academic achievement and the average time spent on part-time jobs.
\(H_a\): There is an association between the perceived effect of part-time work on academic achievement and the average time spent on part-time jobs.
(b)
The conditions for a chi-square inference procedure are met for these data:
1. Randomization: The problem states that the advisory board surveyed a “simple random sample” of students, so the random condition is met.
2. Expected Counts: All expected cell counts must be at least \(5\). Looking at the provided computer output, the expected counts (printed below the observed counts) are all greater than \(5\). The smallest expected count is \(6.825\), so this condition is met.
(c)
Because the \(p\)-value of \(0.007\) is less than standard significance levels (like \(\alpha = 0.05\)), we reject the null hypothesis.
The advisory board should conclude that there is convincing statistical evidence of an association between a student’s perceived effect of part-time work on academic achievement and the average number of hours per week that the student works.
(d)
Because the null hypothesis was rejected in part (c), the advisory board might have made a Type I error.
In the context of the question, a Type I error would be concluding that there is an association between the perceived effect of part-time work on academic achievement and the average time spent on part-time jobs when, in reality, there is no such association between these two variables.
Question

(d) The company wants to conduct a statistical test to investigate whether there is an association between educational achievement and primary source for news for adults in the city. What is the name of the statistical test that should be used?
Most-appropriate topic codes (AP Statistics):
• Topic \(2.6\) — Conditional Probability (Part \(\mathrm{b}\))
• Topic \(2.7\) — Independent Events and Unions of Events (Parts \(\mathrm{a}\), \(\mathrm{c}\))
• Topic \(3.14\) — Setting Up a Chi-Square Test for Homogeneity or Independence (Part \(\mathrm{d}\))
▶️ Answer/Explanation
(a)
Let \(C\) = event that the adult is a college graduate, and \(I\) = event that the adult obtains news primarily from the internet.
Using the Addition Rule:
\(P(C \cup I) = P(C) + P(I) – P(C \cap I)\)
Reading the values directly from the table:
\(P(C) = \frac{693}{2500}, \qquad P(I) = \frac{687}{2500}, \qquad P(C \cap I) = \frac{245}{2500}\)
\(P(C \cup I) = \frac{693}{2500} + \frac{687}{2500} – \frac{245}{2500} = \frac{693 + 687 – 245}{2500} = \frac{1135}{2500}\)
\(\boxed{P(C \cup I) = \frac{1135}{2500} = 0.454}\)
Don’t forget to subtract the overlap — college graduates who use the internet get counted in both the college graduate total and the internet total, so we subtract them once to avoid double-counting.
(b)
We want the conditional probability that an adult obtains news from the internet, given that the adult is a college graduate. From the table, among the 693 college graduates, 245 primarily use the internet:
\(P(I \mid C) = \frac{P(C \cap I)}{P(C)} = \frac{\dfrac{245}{2500}}{\dfrac{693}{2500}} = \frac{245}{693}\)
\(\boxed{P(I \mid C) = \frac{245}{693} \approx 0.354}\)
This is a conditional probability — we’ve already restricted our pool to only the 693 college graduates, so 693 becomes the new denominator. The 2,500 total cancels out entirely.
(c)
Two events are independent if and only if \(P(A \cap B) = P(A) \cdot P(B)\), which is equivalent to checking whether \(P(I \mid C) = P(I)\).
From the table:
\(P(I) = \frac{687}{2500} = 0.275\)
\(P(I \mid C) = \frac{245}{693} \approx 0.354\)
Since \(P(I \mid C) \approx 0.354 \neq 0.275 = P(I)\), the two events are not independent.
We can also verify using the multiplication rule directly:
\(P(C) \cdot P(I) = \frac{693}{2500} \times \frac{687}{2500} = \frac{476{,}091}{6{,}250{,}000} \approx 0.0762\)
\(P(C \cap I) = \frac{245}{2500} = 0.098\)
Since \(0.098 \neq 0.0762\), the events are confirmed to be not independent. In real terms, college graduates are noticeably more likely to get their news from the internet than the general adult population — that difference in rates is exactly what “not independent” means here.
(d)
The appropriate test is the Chi-Square Test of Association (or Independence).
This test is used when we want to determine whether there is an association between two categorical variables — here, educational achievement (3 categories) and primary news source (5 categories).
The degrees of freedom are calculated as:
\(\text{df} = (\text{number of rows} – 1) \times (\text{number of columns} – 1)\)
\(\text{df} = (5 – 1) \times (3 – 1) = 4 \times 2 = \boxed{8}\)
There are 5 rows (news source categories) and 3 columns (education levels), not counting the totals row and column. The degrees of freedom formula captures how many cells in the table are “free to vary” once the row and column totals are fixed.
Question


Most-appropriate topic codes (AP Statistics):
• Topic \(2.2\) — Summary Statistics for Two Categorical Variables (Part \(\mathrm{b}\))
• Topic \(3.14\) — Setting Up a Chi-Square Test for Homogeneity or Independence (Part \(\mathrm{c}\))
▶️ Answer/Explanation
(a)
First, convert the raw counts to percentages within each gender so that males and females are directly comparable:
For Males \((n = 48)\):
• Never had a part-time job: \(\dfrac{21}{48} \approx 43.8\%\)
• Part-time job during summer only: \(\dfrac{15}{48} \approx 31.2\%\)
• Part-time job, not only during summer: \(\dfrac{12}{48} = 25.0\%\)
For Females \((n = 52)\):
• Never had a part-time job: \(\dfrac{31}{52} \approx 59.6\%\)
• Part-time job during summer only: \(\dfrac{13}{52} = 25.0\%\)
• Part-time job, not only during summer: \(\dfrac{8}{52} \approx 15.4\%\)
A side-by-side bar graph (or segmented bar graph) should be drawn with the three job-experience categories on the horizontal axis, percentage on the vertical axis, and separate bars (or segments) for Male and Female within each category. All axes must be labeled.

(b)
Looking at the bar graph, females were considerably more likely than males to have never had a part-time job — about \(59.6\%\) of females compared with \(43.8\%\) of males fall into that category.
On the other hand, males were more likely than females to have had a part-time job during the summer only (\(31.2\%\) vs. \(25.0\%\)), and more likely to have had a part-time job that was not only during the summer (\(25.0\%\) vs. \(15.4\%\)).
If there were no association between gender and job experience, we’d expect the bars for males and females to be roughly the same height in each category — but they clearly aren’t, so the sample data suggest there is an association between gender and job experience.
(c)
The appropriate test of significance is the chi-square test of association (or independence).
The chi-square test statistic is:
\( \chi^2 = \sum \frac{(\text{observed} – \text{expected})^2}{\text{expected}} \)
The null and alternative hypotheses are:
\(H_0\): There is no association between gender and job experience (for the population of high school seniors in the district).
\(H_a\): There is an association between gender and job experience (for the population of high school seniors in the district).
Equivalently:
\(H_0\): Gender and job experience are independent.
\(H_a\): Gender and job experience are not independent.
Since we have two categorical variables each with multiple categories, the chi-square test of association/independence is the correct choice — a two-sample \(z\)-test for proportions only compares two groups on a binary outcome, and wouldn’t capture all three job-experience categories at once.
Question


Most-appropriate topic codes (AP Statistics):
• Topic 3.15 — Carrying Out a Chi-Square Test for Homogeneity or Independence (Parts \(\mathrm{a}\), \(\mathrm{b}\))
▶️ Answer/Explanation
(a)
Step 1 — Hypotheses
\(H_0\): The number of moose in each habitat type is proportional to the amount of acreage of that habitat type (moose have no preference for any particular habitat).
\(H_a\): The number of moose in at least one habitat type is not proportional to the acreage of that type (moose show a preference for at least one habitat type).
Step 2 — Test and Conditions
We use a chi-square goodness-of-fit test, with test statistic
\(\chi^2 = \displaystyle\sum \frac{(\text{Observed} – \text{Expected})^2}{\text{Expected}}\)
The conditions for inference are stated to be met.
Step 3 — Expected Counts and Test Statistic
The expected count for each habitat type is found by multiplying the total number of moose (117) by the proportion of total acreage:
\(E_1 = 0.340 \times 117 = 39.780\)
\(E_2 = 0.101 \times 117 = 11.817\)
\(E_3 = 0.104 \times 117 = 12.168\)
\(E_4 = 0.455 \times 117 = 53.235\)
The chi-square test statistic is:
\(\chi^2 = \dfrac{(25 – 39.780)^2}{39.780} + \dfrac{(22 – 11.817)^2}{11.817} + \dfrac{(30 – 12.168)^2}{12.168} + \dfrac{(40 – 53.235)^2}{53.235}\)
\(\chi^2 = \dfrac{(-14.780)^2}{39.780} + \dfrac{(10.183)^2}{11.817} + \dfrac{(17.832)^2}{12.168} + \dfrac{(-13.235)^2}{53.235}\)
\(\chi^2 = 5.491 + 8.775 + 26.133 + 3.290 = \boxed{43.689}\)
Degrees of freedom: \(\,df = 4 – 1 = 3\)
\(p\text{-value} = P(\chi^2_3 \geq 43.689) \approx 1.757 \times 10^{-9} \ll 0.05\)
Step 4 — Conclusion
Since the \(p\)-value is essentially zero and far below any conventional significance level (\(\alpha = 0.05\)), we reject \(H_0\). There is very strong evidence that the number of moose in each habitat type is not proportional to the acreage of that habitat, meaning moose show a statistically significant preference for certain habitat types.
(b)
The moose seem to prefer habitat types 2 and 3. Relative to the proportion of total acreage, a higher proportion of moose were observed in each of these habitat types than expected. In habitat types 1 and 4, the observed proportion of moose was less than the expected proportion of moose, indicating that these two habitat types are less desirable.
Question

Most-appropriate topic codes (AP Statistics):
• Topic 3.15 — Carrying Out a Chi-Square Test for Homogeneity or Independence (Part a)
• Topic 3.1 — Estimators (Part b)
• Topic 1.11 — Random Sampling (Part b)
▶️ Answer/Explanation
(a)
The appropriate procedure is a chi-square test for independence.
Hypotheses:
\(H_0\): Gender and satisfaction with hospital services are independent (no association).
\(H_a\): Gender and satisfaction with hospital services are not independent (there is an association)
Conditions:
— The sample is a random sample of 1,000 adult county residents.
— All expected cell counts must be at least 5.
Compute each:
\(E = \frac{(\text{row total})(\text{column total})}{\text{grand total}}\)
\(E(\text{Satisfied, Male}) = \dfrac{800 \times 464}{1000} = 371.2\)
\(E(\text{Satisfied, Female}) = \dfrac{800 \times 536}{1000} = 428.8\)
\(E(\text{Not Satisfied, Male}) = \dfrac{200 \times 464}{1000} = 92.8\)
\(E(\text{Not Satisfied, Female}) = \dfrac{200 \times 536}{1000} = 107.2\)
All expected counts are well above 5.
Test Statistic:
\(\chi^2 = \sum \frac{(O – E)^2}{E}\)
\(\chi^2 = \frac{(384 – 371.2)^2}{371.2} + \frac{(416 – 428.8)^2}{428.8} + \frac{(80 – 92.8)^2}{92.8} + \frac{(120 – 107.2)^2}{107.2}\)
\(\chi^2 = \frac{(12.8)^2}{371.2} + \frac{(-12.8)^2}{428.8} + \frac{(-12.8)^2}{92.8} + \frac{(12.8)^2}{107.2}\)
\(\chi^2 = 0.4413 + 0.3821 + 1.7655 + 1.5284 = 4.117\)
Degrees of freedom:
\(df = (r-1)(c-1) = (2-1)(2-1) = 1\)
P-value:
Using the \(\chi^2\) distribution with \(df = 1\):
\(p\text{-value} \approx 0.0424\)
Conclusion:
Since \(p\text{-value} = 0.0424 < \alpha = 0.05\), we reject \(H_0\). There is sufficient statistical evidence at the \(0.05\) significance level to conclude that there is an association between gender and satisfaction with hospital services for adult residents of this county.
\(\boxed{\chi^2 = 4.117,\quad df = 1,\quad p\text{-value} \approx 0.0424 \Rightarrow \text{Reject } H_0}\)
(b)
Yes, \(\dfrac{800}{1{,}000} = 0.80\) is a reasonable estimate for the proportion of all adult county residents who are satisfied with hospital services. The data were collected from a random sample of 1,000 adult county residents, which means the sample is likely representative of the population of all adult county residents. Because random sampling was used, the sample proportion \(\hat{p} = 0.80\) is an unbiased estimate of the true population proportion. Additionally, with a sample size of \(n = 1{,}000\), the estimate is based on a sufficiently large and randomly selected group, giving us reasonable confidence in its accuracy.
\(\boxed{\hat{p} = \frac{800}{1{,}000} = 0.80 \text{ is a reasonable estimate (random sample, large } n\text{)}}\)
Question

Most-appropriate topic codes (AP Statistics):
• Topic 3.15 — Carrying Out a Chi-Square Test for Homogeneity or Independence (Whole Question)
▶️ Answer/Explanation
This problem is asking whether two categorical variables — gender and opinion category — are related to each other, which calls for a Chi-Square test for independence.
State the hypotheses:
\( H_0: \) Response and gender are independent (there is no association between response and gender)
\( H_a: \) Response and gender are not independent (there is an association between response and gender)
Check the conditions:
The \(200\) students were randomly sampled, so the random condition is satisfied. To check the large sample size condition, we need the expected count for each cell, found using:
\( \text{Expected count}=\dfrac{(\text{row total})(\text{column total})}{\text{grand total}} \)
Computing each row and column total, then applying this formula to all \(10\) cells gives:

All of these expected counts are well above \(5\), so the sample size is large enough for the Chi-Square test to be valid.
Compute the test statistic:
The Chi-Square statistic compares each observed count to its expected count:
\( \chi^2=\sum\dfrac{(\text{Observed}-\text{Expected})^2}{\text{Expected}} \)
Adding up this quantity across all \(10\) cells:
\( \chi^2\approx0.907+0.500+0.500+0.278+2.722+0.742+0.409+0.409+0.227+2.227 \)
\( \chi^2\approx8.921 \)
The degrees of freedom for a table with \(2\) rows and \(5\) columns is:
\( df=(2-1)(5-1)=4 \)
Using \(\chi^2\approx8.921\) with \(df=4\), the corresponding P-value is approximately:
\( P\text{-value}\approx0.063 \)
State the conclusion:
Since the P-value \((\approx0.063)\) is greater than a standard significance level of \(\alpha=0.05\), we fail to reject \(H_0\).
\( \boxed{\text{There is not sufficient evidence to conclude that response is dependent on gender.}} \)
That said, since the P-value of \(0.063\) is only slightly above \(0.05\), results this extreme would occur by chance alone only about \(6\) times out of \(100\) if gender and response really were independent — so while we don’t reach the standard threshold for significance, this does provide some marginal evidence of an association between gender and response.
Question

- The contestant spins the wheel.
- If the result is a skunk, no money is won and the contestant’s turn is finished.
- If the result is a number, the corresponding amount in dollars is won. The contestant can then stop with those winnings or can choose to spin again, and his or her turn continues.
- If the contestant spins again and the result is a skunk, all of the money earned on that turn is lost and the turn ends.
- The contestant may continue adding to his or her winnings until he or she chooses to stop or until a spin results in a skunk.

Most-appropriate topic codes (AP Statistics):
• Topic 2.9 — Parameters of Random Variables (Part b)
• Topic 3.14 — Setting Up a Chi-Square Test for Homogeneity or Independence (Part c)
• Topic 3.15 — Carrying Out a Chi-Square Test for Homogeneity or Independence (Part c)
▶️ Answer/Explanation
(a)
There are four equally likely outcomes on the wheel: Skunk, \(\$100\), \(\$200\), and \(\$500\). So the probability of landing on a number (i.e., not a skunk) on any single spin is \(\dfrac{3}{4}\).
Since spins are independent, the probability of getting a number on all three of the first three spins is:
\( P(\text{number on all 3 spins}) = \left(\frac{3}{4}\right)^3 = \frac{27}{64} \approx 0.4219 \)
\(\boxed{P \approx 0.4219}\)
(b)
The contestant currently has \(\$800\) and chooses to spin a fourth time. The four equally likely outcomes on the fourth spin lead to the following total winnings:
The expected value of total winnings is:
\( E(\text{total winnings}) = 0\left(\frac{1}{4}\right) + 900\left(\frac{1}{4}\right) + 1000\left(\frac{1}{4}\right) + 1300\left(\frac{1}{4}\right) \)
\( = \frac{0 + 900 + 1000 + 1300}{4} = \frac{3200}{4} = \$800 \)
Alternatively, the expected gain from the fourth spin alone is:
\( E(\text{4th spin gain}) = (-800)\left(\frac{1}{4}\right) + 100\left(\frac{1}{4}\right) + 200\left(\frac{1}{4}\right) + 500\left(\frac{1}{4}\right) = \frac{-800+100+200+500}{4} = 0 \)
So the expected total winnings \(= \$800 + \$0 = \boxed{\$800}\).
Interestingly, the expected value of spinning again is exactly equal to the amount already won — so on average the fourth spin neither helps nor hurts.
(c)
Hypotheses:
\( H_0: p_1 = p_2 = p_3 = p_4 = \frac{1}{4} \quad \text{(all four outcomes are equally likely)} \)
\( H_a: \text{at least one } p_i \neq \frac{1}{4} \quad \text{(the four outcomes are not equally likely)} \)
Test: Chi-square goodness-of-fit test.
Conditions: The spins are independent (stated in the problem), and the expected count for each outcome is \(100 \times \frac{1}{4} = 25 > 5\), so the sample size is large enough to proceed.
Expected counts: 25 for each of the four outcomes.
Test statistic:
\( \chi^2 = \sum \frac{(\text{Observed} – \text{Expected})^2}{\text{Expected}} \)
\( = \frac{(33-25)^2}{25} + \frac{(21-25)^2}{25} + \frac{(20-25)^2}{25} + \frac{(26-25)^2}{25} \)
\( = \frac{64}{25} + \frac{16}{25} + \frac{25}{25} + \frac{1}{25} = \frac{106}{25} = 4.24 \)
Degrees of freedom: \(df = 4 – 1 = 3\)
P-value: \(p\text{-value} \approx 0.237\) (from chi-square table with \(df = 3\), the test statistic of 4.24 falls well below the critical value of 7.81 at \(\alpha = 0.05\)).
Conclusion: Since the \(p\text{-value} \approx 0.237 > 0.05\), we fail to reject \(H_0\). There is not convincing statistical evidence that the four outcomes on the wheel are not equally likely — the data are consistent with a fair wheel.
