Home / AP® Exam / AP® Statistics / AP Statistics 1.9 Comparisons of the Distributions for One Quantitative Variable- Exam Style Questions – FRQs

AP Statistics 1.9 Comparisons of the Distributions for One Quantitative Variable- Exam Style Questions - FRQs - New Syllabus

Question

The manager of an automotive company is interested in comparing the gas mileages for cars manufactured in Country A and cars manufactured in Country B. The manager selected a random sample of 100 cars manufactured in Country A and a random sample of 100 cars manufactured in Country B. The gas mileages for each sample, in miles per gallon (mpg), are summarized in the boxplots.
A. Compare the distributions of gas mileage for the sample of cars manufactured in Country A and the sample of cars manufactured in Country B.
B. For the distribution of gas mileage for the sample of cars manufactured in Country A, would you expect the mean to be greater than 18 mpg, less than 18 mpg, or equal to 18 mpg? Justify your answer.
C. The manager will create a new boxplot with the combined data from the sample of cars manufactured in Country A and the sample of cars manufactured in Country B.
i. What is the range of the combined data set? Justify your answer.
ii. What is a possible value of the median of the combined data set? Justify your answer by referencing the boxplots shown.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.7\) — Summary Statistics for One Quantitative Variable (Part \( \mathrm{C} \))
• Topic \(1.8\) — Graphical Representations of Summary Statistics for One Quantitative Variable (Parts \( \mathrm{B} \), \( \mathrm{C} \))
• Topic \(1.9\) — Comparisons of the Distributions for One Quantitative Variable (Part \( \mathrm{A} \))
▶️ Answer/Explanation

A.
• Center: The median gas mileage for Country B (\(32\,\text{mpg}\)) is higher than that of Country A (\(18\,\text{mpg}\)).
• Spread: The overall range for Country A (\(38 – 14 = 24\,\text{mpg}\)) is slightly wider than Country B (\(40 – 18 = 22\,\text{mpg}\)), but the Interquartile Range (IQR) for Country B (\(36 – 24 = 12\,\text{mpg}\)) is larger than Country A (\(22 – 14 = 8\,\text{mpg}\)).
• Outliers: Country A has a single extreme high value plotted as an outlier at \(38\,\text{mpg}\), while Country B displays no outliers.

B.
• The mean is expected to be greater than \(18\,\text{mpg}\).
• Because the distribution for Country A is skewed to the right and contains a high outlier at \(38\,\text{mpg}\), the mean will be pulled upward toward the long right tail while the median remains resistant.

C. i.
• The minimum value of the combined set is \(14\,\text{mpg}\) (from Country A) and the maximum value is \(40\,\text{mpg}\) (from Country B).
• \(\text{Combined Range} = \text{Maximum} – \text{Minimum} = 40 – 14 = 26\,\text{mpg}\).

C. ii.
• A possible value for the combined median is any value satisfying \(18\,\text{mpg} \le \text{Median} \le 32\,\text{mpg}\), such as \(24\,\text{mpg}\).
• Since both samples contain exactly 100 cars, combining them creates a total group of 200 cars where the new median rests between the 100th and 101st ordered values, forcing it to fall cleanly between the separate sample medians of \(18\,\text{mpg}\) and \(32\,\text{mpg}\).

Question

As part of a study on the chemistry of Alaskan streams, researchers took water samples from many streams with temperatures colder than 8°C and from many streams with temperatures warmer than 8°C. For each sample, the researchers measured the dissolved oxygen concentration, in milligrams per liter (mg/l).
(a) The researchers constructed the histogram shown for the dissolved oxygen concentration in streams from the sample with water temperatures colder than 8°C. Based on the histogram, describe the distribution of dissolved oxygen concentration in streams with water temperatures colder than 8°C.
(b) The researchers computed the summary statistics shown in the table for the dissolved oxygen concentration in streams from the sample with water temperatures warmer than 8°C. Use the summary statistics to construct a box plot for the dissolved oxygen concentration in streams with water temperatures warmer than 8°C. Do not indicate outliers.
(c) The researchers believe that streams with higher dissolved oxygen concentration are generally healthier for wildlife. Which streams are generally healthier for wildlife, those with water temperature colder than 8°C or those with water temperature warmer than 8°C? Using characteristics of the distribution of dissolved oxygen concentration for temperatures colder than 8°C and characteristics of the distribution of dissolved oxygen concentration for temperatures warmer than 8°C, justify your answer.
 

Most-appropriate topic codes (AP Statistics):

• Topic \(1.5\) — Graphical Representations for One Quantitative Variable (Parts \( \mathrm{a} \), \( \mathrm{b} \))
• Topic \(1.6\) — Descriptions for One Quantitative Variable Distributions (Part \( \mathrm{a} \))
• Topic \(1.9\) — Comparisons of the Distributions for One Quantitative Variable (Part \( \mathrm{c} \))
▶️ Answer/Explanation

(a)
Shape: The distribution is heavily skewed to the left (negatively skewed).
Center: The center of the distribution is around a median of between 11 and 12 mg/l (which is also the modal interval).
Spread: The dissolved oxygen concentrations vary from a minimum between 2 and 3 mg/l to a maximum between 13 and 14 mg/l.
Unusual Features: There appear to be potential low outliers between 2 and 6 mg/l.

(b)


To construct the box plot, follow these steps plotted against an appropriate scale (e.g., 2 to 14):
Draw a vertical line segment at the median: $M = 5.43$
Draw vertical line segments at the quartiles: $Q_1 = 4.39$ and $Q_3 = 6.12$, and connect them to form the central box.
Draw horizontal whiskers extending from the box to the minimum value at $2.10$ and the maximum value at $13.45$.

(c)
The streams with temperatures colder than 8°C are generally healthier for wildlife.
Justification: We must compare the centers and spreads. The median dissolved oxygen concentration for the colder streams (between 11 and 12 mg/l) is significantly higher than the median for the warmer streams ($5.43$ mg/l). Furthermore, the vast majority of colder streams have oxygen levels above 10 mg/l, while the third quartile ($Q_3$) for warmer streams is only $6.12$ mg/l. This means over 75% of warmer streams have lower oxygen concentrations than almost all of the colder streams.

Question

Studies have shown that foods rich in compounds known as flavonoids help lower blood pressure. Researchers conducted a study to investigate whether there was a greater reduction in blood pressure for people who consumed dark chocolate, which contains flavonoids, than people who consumed white chocolate, which does not contain flavonoids. Twenty-five healthy adults agreed to participate in the study and add 3.5 ounces of chocolate to their daily diets. Of the 25 participants, 13 were randomly assigned to the dark chocolate group and the rest were assigned to the white chocolate group. All participants had their blood pressure recorded, in millimeters of mercury (mmHg), before adding chocolate to their daily diets and again 30 days after adding chocolate to their daily diets.
The reduction in blood pressure (before minus after) for each of the participants in the two groups is shown in the dotplots below.
(a) Determine and compare the medians of the reduction in blood pressure for the two groups.
The researchers found the mean reduction in blood pressure for those who consumed dark chocolate is $\overline{x}_{dark}=6.08$ mmHg and the mean reduction in blood pressure for those who consumed white chocolate is $\overline{x}_{white}=0.42$ mmHg.
(b) One researcher indicated that because the difference in sample means of 5.66 mmHg is greater than 0 there is convincing statistical evidence to conclude that the population mean blood pressure reduction for those who consume dark chocolate is greater than for those who consume white chocolate. Why might the researcher’s conclusion, based only on the difference in sample means of 5.66 mmHg, not necessarily be true?
A simulation was conducted to investigate whether there is a greater reduction of blood pressure for those who consume dark chocolate than for those who consume white chocolate. The simulation was conducted under the assumption that no difference exists. The results of 120 trials of the simulation are shown in the following dotplot.
(c) Use the results of the simulation to determine whether the results from the 25 participants in the study provide convincing statistical evidence, at a 5 percent level of significance, that adding dark chocolate to a daily diet will result in a greater reduction in blood pressure, on average, than adding white chocolate to a daily diet. Justify your answer.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.9\) — Comparing the Distributions of One Quantitative Variable (Part \( \mathrm{a} \))
• Topic \(1.1\) — Introducing Statistics: Do the Data We Collected Match What We Expected? (Parts \( \mathrm{b} \), \( \mathrm{c} \))
▶️ Answer/Explanation

(a)
The median reduction in blood pressure for the dark chocolate group is 7 mmHg, and the median reduction in blood pressure for the white chocolate group is 0 mmHg. Therefore, the median reduction for the dark chocolate group is greater than the median reduction for the white chocolate group.

(b)
The researcher’s conclusion might not be true because the difference in sample means could simply be due to sampling variability (chance variation) arising from the random assignment of participants to the two groups. A difference of 5.66 mmHg can sometimes occur by chance even if the treatments are equally effective, so a statistical simulation or hypothesis test is necessary to determine if the result is statistically significant.

(c)
Yes, the results provide convincing statistical evidence. The observed difference in sample means is $6.08 – 0.42 = 5.66$ mmHg. According to the simulation dotplot, a simulated difference of 5.66 or greater occurred in only 3 of the 120 trials. The estimated p-value is $\frac{3}{120} = 0.025$. Because the p-value ($0.025$) is less than the significance level ($\alpha = 0.05$), we reject the null hypothesis and conclude there is convincing evidence that dark chocolate results in a greater reduction in blood pressure on average than white chocolate.

Question

Attendance at games for a certain baseball team is being investigated by the team owner. The following boxplots summarize the attendance, measured as average number of attendees per game, for \(47\) years of the team’s existence. The boxplots include the \(30\) years of games played in the old stadium and the \(17\) years played in the new stadium.
(a) Compare the distributions of average attendance between the old and new stadiums.
The following scatterplot shows average attendance versus year.
(b) Compare the trends in average attendance over time between the old and new stadium.
(c) Consider the following scatterplots.
i. Graph I shows the average attendance versus number of games won for each year. Describe the relationship between the variables.
ii. Graph II shows the same information as Graph I, but also indicates the old and new stadiums. Does Graph II suggest that the rate at which attendance changes as number of games won increases is different in the new stadium compared to the old stadium? Explain your reasoning.
(d) Consider the three variables: number of games won, year, and stadium. Based on the graphs, explain how one of those variables could be a confounding variable in the relationship between average attendance and the other variables.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.9\) — Comparing the Distributions of One Quantitative Variable (Part \( \mathrm{a} \))
• Topic \(5.1\) — Graphical Representations Between Two Quantitative Variables (Parts \( \mathrm{b} \), \( \mathrm{c} \))
• Topic \(5.3\) — Linear Regression Models (Parts \( \mathrm{c} \), \( \mathrm{d} \))
▶️ Answer/Explanation

(a)
The median average attendance is noticeably higher in the new stadium (around \(25,000\)) compared to the old stadium (around \(16,000\)).
While the interquartile ranges are fairly similar, indicating comparable variability in the middle \(50\%\) of attendance figures, the overall range is slightly wider for the new stadium.
Neither stadium’s distribution shows any clear outliers.

(b)
Average attendance in the old stadium remained relatively flat over time with no clear upward or downward trend.
In contrast, the new stadium exhibits a strong, positive, upward trend in average attendance over the years it was actively used.

(c)(i)
There is a strong, positive, linear relationship between the number of games won and the average attendance for each year.

(c)(ii)
No, the rate of change appears to be roughly the same for both stadiums.
If you drew separate lines of best fit through the points representing the old stadium and the new stadium, the slopes of both lines would be nearly identical, meaning attendance increases at a similar rate per win regardless of the venue.

(d)
The number of games won could serve as a confounding variable when assessing the relationship between the stadium type and average attendance.
Since the team won significantly more games playing in the new stadium, and winning is strongly associated with higher attendance, it’s impossible to tell if the attendance spike was caused by the appeal of the new stadium or simply because the team was performing better on the field.

Question

The chemicals in clay used to make pottery can differ depending on the geographical region where the clay originated. Sometimes, archaeologists use a chemical analysis of clay to help identify where a piece of pottery originated. Such an analysis measures the amount of a chemical in the clay as a percent of the total weight of the piece of pottery. The boxplots below summarize analyses done for three chemicals—X, Y, and Z—on pieces of pottery that originated at one of three sites: I, II, or III.
(a) For chemical Z, describe how the percents found in the pieces of pottery are similar and how they differ among the three sites.
(b) Consider a piece of pottery known to have originated at one of the three sites, but the actual site is not known.
(i) Suppose an analysis of the clay reveals that the sum of the percents of the three chemicals X, Y, and Z is \(20.5\%\). Based on the boxplots, which site—I, II, or III—is the most likely site where the piece of pottery originated? Justify your choice.
(ii) Suppose only one chemical could be analyzed in the piece of pottery. Which chemical—X, Y, or Z—would be the most useful in identifying the site where the piece of pottery originated? Justify your choice.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.9\) — Comparing the Distributions of One Quantitative Variable (Part \( \mathrm{a} \))
• Topic \(1.8\) — Graphical Representations of Summary Statistics for One Quantitative Variable (Parts \( \mathrm{a} \), \( \mathrm{b} \))
▶️ Answer/Explanation

(a)
Similarity: The median percent of chemical Z is approximately the same across all three sites, at roughly \(7\%\) of total weight. This means the center of the distribution of chemical Z does not help distinguish one site from another.
Difference: The spread (range) of the percent of chemical Z differs considerably among the three sites.
Site II has the smallest range, approximately \(2\%\) (from about \(6\%\) to \(8\%\)).
Site I has a moderate range, approximately \(6\%\) (from about \(4\%\) to \(10\%\)).
Site III has the largest range, approximately \(8\%\) (from about \(3\%\) to \(11\%\)).
So while the typical chemical Z content is similar at all three sites, Site II pottery is far more consistent in its Z content, while Site III pottery shows the greatest variability.

(b)(i)
The piece of pottery most likely originated at Site III.
To decide, we estimate the minimum and maximum possible sums of the percents of X, Y, and Z at each site by reading the minimum and maximum values from the boxplots:

For Site I, the possible sum ranges from \(21\%\) to \(33\%\) — since \(20.5\%\) falls below the minimum, a sum of \(20.5\%\) is unlikely for Site I.
For Site II, the possible sum ranges from \(12.9\%\) to \(19\%\) — since \(20.5\%\) falls above the maximum, a sum of \(20.5\%\) is also unlikely for Site II.
For Site III, the possible sum ranges from \(14\%\) to \(26.5\%\) — since \(20.5\%\) falls comfortably within this interval, Site III is the most plausible origin.
\(\boxed{\text{Most likely site: Site III}}\)

(b)(ii)
Chemical Y would be the most useful for identifying the site of origin.
Looking at the boxplots for chemical Y across the three sites:
Site I: chemical Y ranges from approximately \(11\%\) to \(15\%\)
Site II: chemical Y ranges from approximately \(1.9\%\) to \(4\%\)
Site III: chemical Y ranges from approximately \(6\%\) to \(8\%\)
These three ranges do not overlap at all — every possible value of chemical Y belongs to exactly one site’s range. So if we measure chemical Y in an unknown piece, we can pinpoint its origin with certainty.
By contrast, the distributions of chemicals X and Z show substantial overlap across the three sites, meaning a measured value of X or Z could plausibly belong to multiple sites, making those chemicals far less useful for classification.
\(\boxed{\text{Most useful chemical: Y}}\)

Question

A newspaper in Germany reported that the more semesters needed to complete an academic program at the university, the greater the starting salary in the first year of a job. The report was based on a study that used a random sample of 24 people who had recently completed an academic program. Information was collected on the number of semesters each person in the sample needed to complete the program and the starting salary, in thousands of euros, for the first year of a job. The data are shown in the scatterplot below.
(a) Does the scatterplot support the newspaper report about number of semesters and starting salary? Justify your answer.
The table below shows computer output from a linear regression analysis on the data.
(b) Identify the slope of the least-squares regression line, and interpret the slope in context.
An independent researcher received the data from the newspaper and conducted a new analysis by separating the data into three groups based on the major of each person. A revised scatterplot identifying the major of each person is shown below.
(c) Based on the people in the sample, describe the association between starting salary and number of semesters for the business majors.
(d) Based on the people in the sample, compare the median starting salaries for the three majors.
(e) Based on the analysis conducted by the independent researcher, how could the newspaper report be modified to give a better description of the relationship between the number of semesters and the starting salary for the people in the sample?

Most-appropriate topic codes (AP Statistics):

• Topic \(5.1\) — Graphical Representations Between Two Quantitative Variables (Parts \( \mathrm{a} \), \( \mathrm{c} \))
• Topic \(5.2\) — Correlation (Part \( \mathrm{e} \))
• Topic \(5.3\) — Linear Regression Models (Part \( \mathrm{b} \))
• Topic \(1.9\) — Comparisons of the Distributions for One Quantitative Variable (Part \( \mathrm{d} \))
▶️ Answer/Explanation

(a)

Yes, the scatterplot supports the newspaper report. The scatterplot shows a positive association between the number of semesters needed to complete an academic program and starting salary — as the number of semesters increases, starting salary tends to increase as well. This is consistent with the newspaper’s claim that more semesters are associated with a greater starting salary.

(b)

The slope of the least-squares regression line is \(b_1 = 1.1594\).
The least-squares regression equation is:
\( \hat{y} = 34.018 + 1.1594x \)
Interpretation: For each additional semester needed to complete an academic program, the predicted starting salary in the first year of a job increases by approximately €1,159.40 (i.e., 1.1594 thousand euros).

(c)

For the business majors alone, there is a strong, negative, linear association between the number of semesters and starting salary. Business majors who need more semesters to complete their academic program tend to have lower starting salaries — which is the opposite direction from the overall trend seen in the combined scatterplot.

(d)

Business majors have the lowest median starting salary, at approximately €38,000. Physics majors have the next highest median starting salary, at approximately €48,000. Chemistry majors have the highest median starting salary, at approximately €55,000.
So in order from lowest to highest median starting salary: Business \(\approx\) €38,000 < Physics \(\approx\) €48,000 < Chemistry \(\approx\) €55,000.

(e)

The newspaper report should be modified to account for the major of each person. The overall positive association in the original report is largely explained by the fact that different majors — chemistry, physics, and business — tend to require different numbers of semesters and also have very different starting salary levels. Chemistry majors take more semesters on average and also earn the highest salaries; business majors take fewer semesters and earn the lowest salaries. This creates an apparent positive association when all three majors are pooled together.
However, within each individual major, students who take a greater number of semesters to complete their program tend to have lower starting salaries, not higher. The newspaper report should therefore be revised to state that, while majors requiring more semesters overall tend to have higher starting salaries (chemistry highest, physics next, business lowest), within any given major, taking more semesters to complete the program is associated with a lower starting salary.

Question

Two large corporations, A and B, hire many new college graduates as accountants at entry-level positions. In \(2009\) the starting salary for an entry-level accountant position was \(\$36,000\) a year at both corporations. At each corporation, data were collected from \(30\) employees who were hired in \(2009\) as entry-level accountants and were still employed at the corporation five years later. The yearly salaries of the \(60\) employees in \(2014\) are summarized in the boxplots below.
(a) Write a few sentences comparing the distributions of the yearly salaries at the two corporations.
(b) Suppose both corporations offered you a job for \(\$36,000\) a year as an entry-level accountant.
i. Based on the boxplots, give one reason why you might choose to accept the job at corporation A.
ii. Based on the boxplots, give one reason why you might choose to accept the job at corporation B.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.9\) — Comparisons of the Distributions for One Quantitative Variable (Parts \( \mathrm{a} \), \( \mathrm{b} \))
▶️ Answer/Explanation

(a)
The median salary is approximately the same for both corporations. The range and interquartile range of the salaries are greater for Corporation A than for Corporation B. The two highest salaries at Corporation A are outliers while Corporation B has no outliers.

(b)(i)
Five years after starting, at least \(3\) out of \(30\) (\(10\%\)) of the salaries at Corporation A are greater than the maximum salary at Corporation B. If I accept the offer from Corporation A, I might be able to make a higher salary at Corporation A than at Corporation B.

(b)(ii)
Five years after starting, the minimum salary at Corporation B is greater than at Corporation A. In fact, at Corporation A it looks like some people are still making the starting salary of \(\$36,000\) and never received a raise in the five years since they were hired. So if I work at Corporation A, I might never receive a raise in salary.

Question

Tropical storms in the Pacific Ocean with sustained winds that exceed \(74\) miles per hour are called typhoons. Graph A below displays the number of recorded typhoons in two regions of the Pacific Ocean — the Eastern Pacific and the Western Pacific — for the years from \(1997\) to \(2010\).
(a) Compare the distributions of yearly frequencies of typhoons for the two regions of the Pacific Ocean for the years from \(1997\) to \(2010\).
(b) For each region, describe how the yearly frequencies changed over the time period from \(1997\) to \(2010\).
A moving average for data collected at regular time increments is the average of data values for two or more consecutive increments. The \(4\)-year moving averages for the typhoon data are provided in the table below. For example, the Eastern Pacific \(4\)-year moving average for \(2000\) is the average of \(22\), \(16\), \(15\), and \(21\), which is equal to \(18.50\).
(c) Show how to calculate the \(4\)-year moving average for the year \(2010\) in the Western Pacific. Write your value in the appropriate place in the table.
(d) Graph B below shows both yearly frequencies (connected by dashed lines) and the respective \(4\)-year moving averages (connected by solid lines). Use your answer in part (c) to complete the graph.
(e) Consider graph B.
i. What information is more apparent from the plots of the \(4\)-year moving averages than from the plots of the yearly frequencies of typhoons?
ii. What information is less apparent from the plots of the \(4\)-year moving averages than from the plots of the yearly frequencies of typhoons?

Most-appropriate topic codes (AP Statistics):

• Topic \(1.9\) — Comparisons of the Distributions for One Quantitative Variable (Part \( \mathrm{a} \))
• Topic \(1.5\) — Graphical Representations for One Quantitative Variable (Part \( \mathrm{b} \))
• Topic \(1.7\) — Summary Statistics for One Quantitative Variable (Part \( \mathrm{c} \))
• Topic \(1.5\) — Graphical Representations for One Quantitative Variable (Part \( \mathrm{d} \))
• Topic \(1.9\) — Comparisons of the Distributions for One Quantitative Variable (Part \( \mathrm{e} \))
▶️ Answer/Explanation

(a)

The Western Pacific consistently had a higher yearly frequency of typhoons than the Eastern Pacific throughout the entire period from \(1997\) to \(2010\). The center (median/mean) of the Western Pacific distribution is noticeably higher — roughly in the low-to-mid \(30\)s — compared to the Eastern Pacific, which clusters mostly in the high teens to low \(20\)s. The Western Pacific also showed greater year-to-year variability, with values ranging from \(18\) to \(39\), while the Eastern Pacific ranged from \(15\) to \(25\). Both distributions appear roughly similar in shape — neither strongly skewed — but the Western Pacific distribution is shifted considerably upward compared to the Eastern Pacific.

(b)

For the Eastern Pacific, typhoon frequencies were relatively stable throughout the period. After starting at \(22\) in \(1997\), the counts dipped slightly in the late \(1990\)s and early \(2000\)s, hovered in the high teens, then showed a slight increase around \(2006\) before settling back down to \(18\) in \(2010\). There is no strong overall upward or downward trend — the Eastern Pacific frequencies stayed roughly flat over the \(14\)-year period.

For the Western Pacific, typhoon frequencies showed a noticeable overall downward trend over the period. Starting at \(33\) in \(1997\), counts rose to a peak of \(39\) in \(2002\), then declined steadily, dropping sharply to \(18\) in \(2010\) — the same value as the Eastern Pacific in that year. The general trend in the Western Pacific is a decrease in typhoon frequency from the early \(2000\)s onward.

(c)

The \(4\)-year moving average for \(2010\) in the Western Pacific is the average of the four most recent yearly values: \(2007\), \(2008\), \(2009\), and \(2010\).
$\text{4-year moving average}_{2010} = \frac{28 + 27 + 28 + 18}{4} = \frac{101}{4} = \boxed{25.25}$
The value is written in the table as follows.

(d)

On Graph B, plot the point \((2010,\ 25.25)\) for the Western Pacific \(4\)-year moving average and connect it with a solid line to the previous moving average value of \(29.25\) at \(2009\). This completes the Western Pacific moving average line, which shows a continued decline ending at \(25.25\) in \(2010\).

(e)(i)

The \(4\)-year moving averages make the overall long-term trends in typhoon frequency more apparent. In particular, the moving average plot for the Western Pacific clearly shows the gradual downward trend in typhoon activity from the early \(2000\)s to \(2010\), which is harder to see in the raw yearly data due to year-to-year fluctuations. The moving averages smooth out the short-term noise and reveal the underlying direction of change over time.

(e)(ii)

The \(4\)-year moving averages make the year-to-year variability and individual fluctuations less apparent. For example, the sharp single-year spike or drop in a particular year (such as the Western Pacific dropping to \(26\) in \(2005\) or the Eastern Pacific jumping to \(25\) in \(2006\)) is dampened or obscured in the moving average plot. The raw yearly frequency plots show these individual extreme values much more clearly.

Question

Independent random samples of 500 households were taken from a large metropolitan area in the United States for the years 1950 and 2000. Histograms of household size (number of people in a household) for the years are shown below.
(a) Compare the distributions of household size in the metropolitan area for the years 1950 and 2000.
(b) A researcher wants to use these data to construct a confidence interval to estimate the change in mean household size in the metropolitan area from the year 1950 to the year 2000. State the conditions for using a two-sample \(t\)-procedure, and explain whether the conditions for inference are met.

Most-appropriate topic codes (AP Statistics):

• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Part a)
• Topic 4.2 — Constructing a Confidence Interval for a Population Mean or Population Mean Difference (Part b)
▶️ Answer/Explanation

(a)
When comparing distributions, we address shape, center, spread, and any unusual features:
Shape: Both distributions of household sizes are strongly skewed to the right (positively skewed).
Center: The household sizes in 1950 typical center at a higher value than in 2000. The median household size in 1950 is approximately 4 people, whereas the median household size in 2000 has shifted downward to approximately 3 people.
Spread: Household sizes were more variable in 1950 than in 2000. The total range for 1950 runs from 1 to 14 (a range of 13), while the range for 2000 is slightly narrower, running from 1 to 12 (a range of 11).
Outliers/Features: Both time periods show a cluster of typical values between 1 and 6 people, but 1950 has a longer tail extending to very large households compared to 2000.

(b)
The conditions required for a two-sample \(t\)-procedure, along with verification, are outlined below:
1. Independent Random Sampling Condition: The data must come from two independent random samples.
Status: Met. The problem states that “independent random samples of 500 households were taken.”
2. Normality Condition: The populations should be normally distributed, or the sample sizes must be large enough to apply the Central Limit Theorem.
Status: Met. Even though both histograms show distinct right-skewness, the sample sizes are \(n_{1950} = 500\) and \(n_{2000} = 500\). Since both sample sizes are well above the threshold of 30 (\(500 \ge 30\)), the Central Limit Theorem ensures that the sampling distribution of the difference in sample means is approximately normal.
3. Independence (10% Rule) Condition: The sample sizes should not exceed 10% of their respective population sizes if sampling without replacement.
Status: Met. It is highly reasonable to assume that 500 households is less than 10% of all available households in a “large metropolitan area” in the United States for both 1950 and 2000.

Question

A professional sports team evaluates potential players for a certain position based on two main characteristics, speed and strength.
(a) Speed is measured by the time required to run a distance of 40 yards, with smaller times indicating more desirable (faster) speeds. From previous speed data for all players in this position, the times to run 40 yards have a mean of 4.60 seconds and a standard deviation of 0.15 seconds, with a minimum time of 4.40 seconds, as shown in the table below.
Based on the relationship between the mean, standard deviation, and minimum time, is it reasonable to believe that the distribution of 40-yard running times is approximately normal? Explain.
(b) Strength is measured by the amount of weight lifted, with more weight indicating more desirable (greater) strength. From previous strength data for all players in this position, the amount of weight lifted has a mean of 310 pounds and a standard deviation of 25 pounds, as shown in the table below.
Calculate and interpret the z-score for a player in this position who can lift a weight of 370 pounds.
(c) The characteristics of speed and strength are considered to be of equal importance to the team in selecting a player for the position. Based on the information about the means and standard deviations of the speed and strength data for all players and the measurements listed in the table below for Players A and B, which player should the team select if the team can only select one of the two players? Justify your answer.

Most-appropriate topic codes (AP Statistics):

• Topic 2.11 — The Normal Distribution (Part a)
• Topic 1.7 — Summary Statistics for One Quantitative Variable (Part b)
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Part c)
▶️ Answer/Explanation

(a)
No, it is not reasonable to believe that the distribution of running times is approximately normal.
In a normal distribution, data extends several standard deviations below the mean. For this dataset, the minimum running time is \(4.40\) seconds, which yields a standardized distance from the mean of:
\(z = \dfrac{4.40 – 4.60}{0.15} = -1.33\)
Since a normal distribution expects approximately \(9.2\%\) of its observations to fall below \(1.33\) standard deviations beneath the mean, having a hard cutoff at \(1.33\) standard deviations indicates that the left tail is severely truncated. Thus, the distribution is likely skewed to the right.

(b)
To find the standardized score for a weight of \(370\) pounds, we use the z-score formula:
\(z = \dfrac{x – \mu}{\sigma}\)
\(z = \dfrac{370 – 310}{25} = \dfrac{60}{25} = 2.40\)
Interpretation: This player’s weightlifting performance is \(2.40\) standard deviations above the average weight lifted by all players in this position.

(c)
The team should select Player A.
To perform a fair comparison since both speed and strength carry equal importance, we compute the z-scores for both players on each metric:
Player A:
\(z_{\text{speed}} = \dfrac{4.42 – 4.60}{0.15} = -1.20\)
\(z_{\text{strength}} = \dfrac{370 – 310}{25} = 2.40\)
Since a lower running time indicates a more desirable speed, a negative z-score is a positive attribute. The combined standardized advantage for Player A is \(2.40 – (-1.20) = 3.60\) units of desirability (or we can think of a speed index where faster is positive, meaning a net sum of \(1.20 + 2.40 = 3.60\)).
Player B:
\(z_{\text{speed}} = \dfrac{4.57 – 4.60}{0.15} = -0.20\)
\(z_{\text{strength}} = \dfrac{375 – 310}{25} = 2.60\)
The combined standardized index advantage for Player B is \(0.20 + 2.60 = 2.80\).
Comparing the two candidates, Player A is dramatically faster than Player B (\(1.20\) standard deviations below the mean versus only \(0.20\) standard deviations below), while Player B is only slightly stronger than Player A (\(2.60\) standard deviations above the mean versus \(2.40\)). Therefore, Player A represents a significantly better overall draft value when both metrics are weighted equally.

Question

Records are kept by each state in the United States on the number of pupils enrolled in public schools and the number of teachers employed by public schools for each school year. From these records, the ratio of the number of pupils to the number of teachers (P-T ratio) can be calculated for each state. The histograms below show the P-T ratio for every state during the 2001-2002 school year. The histogram on the left displays the ratios for the 24 states that are west of the Mississippi River, and the histogram on the right displays the ratios for the 26 states that are east of the Mississippi River.
(a) Describe how you would use the histograms to estimate the median P-T ratio for each group (west and east) of states. Then use this procedure to estimate the median of the west group and the median of the east group.
(b) Write a few sentences comparing the distributions of P-T ratios for states in the two groups (west and east) during the 2001-2002 school year.
(c) Using your answers in parts (a) and (b), explain how you think the mean P-T ratio during the 2001-2002 school year will compare for the two groups (west and east).

Most-appropriate topic codes (AP Statistics):

• Topic 1.5 — Graphical Representations for One Quantitative Variable (Part a)
• Topic 1.6 — Descriptions for One Quantitative Variable Distributions (Part b)
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Part c)
▶️ Answer/Explanation

(a)
To find the median from a histogram, identify the interval that contains the middle observation by adding the frequencies of each bin from left to right until the cumulative total reaches half of the sample size.
For a group with \(n\) sorted observations, the median position is calculated using the formula \(\frac{n+1}{2}\).
For the West group, there are \(n = 24\) observations, meaning the median is the average of the 12th and 13th values.
By counting cumulative frequencies from the left: the 12-13 bin has 1, the 13-14 bin has 4 (total 5), the 14-15 bin has 6 (total 11), and the 15-16 bin has 3 (total 14).
Since the 12th and 13th observations fall inside the 15-16 interval, the estimated median P-T ratio for the West group is between 15 and 16.
For the East group, there are \(n = 26\) observations, meaning the median is the average of the 13th and 14th values.
Counting cumulative frequencies from the left: the 12-13 bin has 2, the 13-14 bin has 4 (total 6), the 14-15 bin has 4 (total 10), and the 15-16 bin has 11 (total 21).
Since the 13th and 14th observations also fall inside the 15-16 interval, the estimated median P-T ratio for the East group is also between 15 and 16.
\(\boxed{\text{West Median: } [15, 16], \text{ East Median: } [15, 16]}\)
From the histogram, cumulative frequencies for the two groups are shown in the table below.

Thus, the median P-T ratio for both groups is at least 15 students per teacher and at most 16 students per teacher.

(b)
Shape: The distribution of P-T ratios for the West group is unimodal and skewed to the right, whereas the distribution for the East group is unimodal and approximately symmetric.
Center: The centers are nearly identical, with both the West and East groups having a median P-T ratio located in the interval between 15 and 16 students per teacher.
Spread: There is visibly more variability in the P-T ratios for the West group than for the East group. The maximum possible range for the West group is \(22 – 12 = 10\), which is larger than the maximum possible range for the East group, which is \(19 – 12 = 7\).

(c)
The mean P-T ratio for the West group will likely be greater than the mean P-T ratio for the East group.
As established in part (a), both distributions have approximately the same median value between 15 and 16.
Because the distribution for the West group is strongly skewed to the right, its mean will be pulled upward to a value greater than its median.
In contrast, because the distribution for the East group is roughly symmetric, its mean will remain close to its median value.
Therefore, the mean of the West group will be higher than the mean of the East group.
\(\boxed{\text{Mean}_{\text{West}} > \text{Mean}_{\text{East}}}\)

Question

Hurricane damage amounts, in millions of dollars per acre, were estimated from insurance records for major hurricanes for the past three decades. A stratified random sample of five locations (based on categories of distance from the coast) was selected from each of three coastal regions in the southeastern United States. The three regions were Gulf Coast (Alabama, Louisiana, Mississippi), Florida, and Lower Atlantic (Georgia, South Carolina, North Carolina). Damage amounts in millions of dollars per acre, adjusted for inflation, are shown in the table below.
(a) Sketch a graphical display that compares the hurricane damage amounts per acre for the three different coastal regions (Gulf Coast, Florida, and Lower Atlantic) and that also shows how the damage amounts vary with distance from the coast.
(b) Describe differences and similarities in the hurricane damage amounts among the three regions.
Because the distributions of hurricane damage amounts are often skewed, statisticians frequently use rank values to analyze such data.
(c) In the table below, the hurricane damage amounts have been replaced by the ranks 1, 2, or 3. For each of the distance categories, the highest damage amount is assigned a rank of 1 and the lowest damage amount is assigned a rank of 3. Determine the missing ranks for the 10-to-20-miles distance category and calculate the average rank for each of the three regions. Place the values in the table below.
(d) Consider testing the following hypotheses.
\(H_0\): There is no difference in the distributions of hurricane damage amounts among the three regions.
\(H_a\): There is a difference in the distributions of hurricane damage amounts among the three regions.
If there is no difference in the distribution of hurricane damage amounts among the three regions (Gulf Coast, Florida, and Lower Atlantic), the expected value of the average rank for each of the three regions is 2. Therefore, the following test statistic can be used to evaluate the hypotheses above:
\[Q = 5\left[\left(\bar{R}_G – 2\right)^2 + \left(\bar{R}_F – 2\right)^2 + \left(\bar{R}_A – 2\right)^2\right]\]
where \(\bar{R}_G\) is the average rank over the five distance categories for the Gulf Coast (and \(\bar{R}_F\) and \(\bar{R}_A\) are similarly defined for the Florida and Lower Atlantic coastal regions).
Calculate the value of the test statistic \(Q\) using the average ranks you obtained in part (c).
(e) One thousand simulated values of this test statistic, \(Q\), were calculated, assuming no difference in the distributions of hurricane damage amounts among the three coastal regions. The results are shown in the table below. These data are also shown in the frequency plot where the heights of the lines represent the frequency of occurrence of simulated values of \(Q\).

Use these simulated values and the test statistic you calculated in part (d) to determine if the observed data provide evidence of a significant difference in the distributions of hurricane damage amounts among the three coastal regions. Explain.

Most-appropriate topic codes (AP Statistics):

• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Parts a, b)
• Topic 2.8 — Introduction to Random Variables and Probability Distributions (Part c)
• Topic 3.15 — Carrying Out a Chi-Square Test for Homogeneity or Independence (Parts d, e)

▶️ Answer/Explanation

(a) Graphical Display

A well-constructed graphical display for this data is a grouped bar chart, with the five distance-from-coast categories on the horizontal axis and damage amounts (in millions of dollars per acre) on the vertical axis. Three bars are drawn side by side within each distance group — one for Gulf Coast, one for Florida, and one for Lower Atlantic — with a clearly labeled key.

(b) Differences and Similarities

Similarity: In all three regions, hurricane damage amounts decrease consistently as distance from the coast increases. This pattern holds without exception across all five distance categories for every region.
Difference: For almost every distance category, Florida has the highest damage amounts, while the Lower Atlantic region generally has the lowest. The Gulf Coast falls in between, though at the 5-to-10-mile distance, the Gulf Coast actually has the highest damage of the three regions.

(c) Missing Ranks and Average Ranks
For the 10-to-20-miles distance category, compare the three damage amounts:
Florida: \(3.0\) million (highest) \(\Rightarrow\) rank \(= 1\)
Gulf Coast: \(1.7\) million (middle) \(\Rightarrow\) rank \(= 2\)
Lower Atlantic: \(0.3\) million (lowest) \(\Rightarrow\) rank \(= 3\)
The completed rank table is:

The average ranks are computed as follows:

\(\bar{R}_G = \frac{2+2+3+1+2}{5} = \frac{10}{5} = 2.0\)

\(\bar{R}_F = \frac{1+1+1+2+1}{5} = \frac{6}{5} = 1.2\)

\(\bar{R}_A = \frac{3+3+2+3+3}{5} = \frac{14}{5} = 2.8\)

(d) Calculating the Test Statistic \(Q\)
Substitute the average ranks from part (c) into the formula:
\(Q = 5\left[\left(\bar{R}_G – 2\right)^2 + \left(\bar{R}_F – 2\right)^2 + \left(\bar{R}_A – 2\right)^2\right]\)
\(Q = 5\left[(2.0 – 2)^2 + (1.2 – 2)^2 + (2.8 – 2)^2\right]\)
\(Q = 5\left[0 + (-0.8)^2 + (0.8)^2\right]\)
\(Q = 5\left[0 + 0.64 + 0.64\right]\)
\(\boxed{Q = 5 \times 1.28 = 6.4}\)

(e) Simulation-Based Conclusion

From the frequency table, simulated \(Q\) values of \(6.4\) or greater occurred in:
\(16 + 15 + 6 + 2 = 39 \text{ out of } 1{,}000 \text{ simulations}\)
This gives an approximate \(p\)-value of:
\(p\text{-value} \approx \frac{39}{1{,}000} = 0.039\)
Since the \(p\)-value of \(0.039\) is less than \(\alpha = 0.05\), we reject \(H_0\). The sample data provide reasonably strong evidence that there is a difference in the distributions of hurricane damage amounts among the three coastal regions (Gulf Coast, Florida, and Lower Atlantic).

Question

As a part of the United States Department of Agriculture’s Super Dump cleanup efforts in the early 1990s, various sites in the country were targeted for cleanup. Three of the targeted sites—River X, River Y, and River Z—had become contaminated with pesticides because they were located near abandoned pesticide dump sites. Measurements of the concentration of aldrin (a commonly used pesticide) were taken at twenty randomly selected locations in each river near the dump sites. The boxplots shown below display the five-number summaries for the concentrations, in parts per million (ppm) of aldrin, for the twenty locations that were sampled in each of the three rivers.
(a) Compare the distributions of the concentration of aldrin among the three rivers.
(b) The twenty concentrations of aldrin for River X are given below.
\(3.4\quad 4.0\quad 5.6\quad 3.7\quad 8.0\quad 5.5\quad 5.3\quad 4.2\quad 4.3\quad 7.3\)
\(8.6\quad 5.1\quad 8.7\quad 4.6\quad 7.5\quad 5.3\quad 8.2\quad 4.7\quad 4.8\quad 4.6\)
Construct a stemplot that displays the concentrations of aldrin for River X.
(c) Describe a characteristic of the distribution of aldrin concentrations in River X that can be seen in the stemplot but cannot be seen in the boxplot.
 

Most-appropriate topic codes (AP Statistics):

• Topic \(1.9\) — Comparisons of the Distributions for One Quantitative Variable (Part \( \mathrm{a} \))
• Topic \(1.5\) — Graphical Representations for One Quantitative Variable (Part \( \mathrm{b} \))
• Topic \(1.6\) — Descriptions for One Quantitative Variable Distributions (Part \( \mathrm{c} \))
▶️ Answer/Explanation

(a)
Comparing the medians reveals that the concentration of aldrin tends to be highest for River X and lowest for River Z.
About \(50\%\) of the concentrations of aldrin for Rivers X and Y are higher than all of the concentrations for River Z.
River X also displays the most variability in aldrin concentrations, as seen by the largest range and largest IQR, and River Z has the least variability, as judged by both IQR and range.
The shapes of the three distributions differ, in that the distribution appears to be skewed to the right for River X, roughly symmetric for River Y and slightly skewed to the left for River Z.

(b)
Aldrin concentrations (in ppm) for River X
Leaf unit \(= 0.1\) (for example, \(3 \mid 4\) represents \(3.4\,\text{ppm}\))
\(3 \mid 4\ 7\)
\(4 \mid 0\ 2\ 3\ 6\ 6\ 7\ 8\)
\(5 \mid 1\ 3\ 3\ 5\ 6\)
\(6 \mid \)
\(7 \mid 3\ 5\)
\(8 \mid 0\ 2\ 6\ 7\)

(c)
The stemplot shows a clear gap in the distribution of aldrin concentrations for River X.
This gap occurs between the values of \(5.6\) and \(7.3\,\text{ppm}\) of aldrin.
This gap is not apparent in the boxplot.

Question

To determine the amount of sugar in a typical serving of breakfast cereal, a student randomly selected 60 boxes of different types of cereal from the shelves of a large grocery store.
The student noticed that the side panels of some of the cereal boxes showed sugar content based on one-cup servings, while others showed sugar content based on three-quarter-cup servings. Many of the cereal boxes with side panels that showed three-quarter-cup servings were ones that appealed to young children, and the student wondered whether there might be some difference in the sugar content of the cereals that showed different-size servings on their side panels. To investigate the question, the data were separated into two groups. One group consisted of 29 cereals that showed one-cup serving sizes; the other group consisted of 31 cereals that showed three-quarter-cup serving sizes. The boxplots shown below display sugar content (in grams) per serving of the cereals for each of the two serving sizes.
(a) Write a few sentences to compare the distributions of sugar content per serving for the two serving sizes of cereals.
After analyzing the boxplots on the preceding page, the student decided that instead of a comparison of sugar content per recommended serving, it might be more appropriate to compare sugar content for equal-size servings. To compare the amount of sugar in serving sizes of one cup each, the amount of sugar in each of the cereals showing three-quarter-cup servings on their side panels was multiplied by \(\dfrac{4}{3}\). The bottom boxplot shown below displays sugar content (in grams) per cup for those cereals that showed a serving size of three-quarter-cup on their side panels.
(b) What new information about sugar content do the boxplots above provide?
(c) Based on the boxplots shown above on this page, how would you expect the mean amounts of sugar per cup to compare for the different recommended serving sizes? Explain.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.1\) — Analyzing Categorical Data / Representing Data Graphically (Parts \(\mathrm{a}\), \(\mathrm{b}\))
• Topic \(1.7\) — Summary Statistics for a Quantitative Variable (Center, Spread, Shape) (Parts \(\mathrm{a}\), \(\mathrm{b}\), \(\mathrm{c}\))
• Topic \(1.9\) — Comparing Distributions of a Quantitative Variable (Parts \(\mathrm{a}\), \(\mathrm{b}\))
• Topic \(1.10\) — The Effect of Adding a Constant or Multiplying by a Constant on Summary Statistics (Part \(\mathrm{b}\))
• Topic \(3.1\) — Mean and Standard Deviation of a Linear Transformation (Part \(\mathrm{c}\))
▶️ Answer/Explanation

(a)

When comparing two distributions from boxplots, we examine center, spread, shape, and unusual features.
The cereals with one-cup serving sizes have a higher median sugar content per serving than the cereals with three-quarter-cup serving sizes. The one-cup distribution also has greater variability, as indicated by its larger range and larger interquartile range (IQR). In terms of shape, the one-cup distribution appears somewhat left-skewed because the median is closer to the upper quartile than to the lower quartile, while the three-quarter-cup distribution is more nearly symmetric. Neither distribution appears to contain extreme outliers.

(b)

Multiplying each sugar value in the three-quarter-cup group by \(\dfrac{4}{3}\) converts the measurements to sugar content per cup, allowing a fair comparison using equal serving sizes.
The adjusted boxplot shows that cereals with recommended serving sizes of three-quarter cup tend to contain more sugar per cup than cereals with recommended serving sizes of one cup. The median for the adjusted three-quarter-cup distribution is now noticeably higher than the median for the one-cup distribution. In addition, all measures of spread (range and IQR) for the adjusted distribution have increased by a factor of \(\dfrac{4}{3}\), reflecting the effect of multiplying every observation by a constant.

(c)

We would expect the mean sugar content per cup to be greater for cereals that list a serving size of three-quarter cup.
After adjustment, the three-quarter-cup distribution has a higher center than the one-cup distribution, as seen from its higher median. Because the mean generally follows the center of the distribution, the higher overall location of the adjusted three-quarter-cup distribution suggests a larger mean sugar content per cup.
\(\boxed{\bar{x}_{\frac{3}{4}\text{-cup}} > \bar{x}_{\text{1-cup}}}\)

Question

A certain state’s education commissioner released a new report card for all the public schools in that state. This report card provides a new tool for comparing schools across the state. One of the key measures that can be computed from the report card is the student-to-teacher ratio, which is the number of students enrolled in a given school divided by the number of teachers at that school.
The data below give the student-to-teacher ratio at the 10 schools with the highest proportion of students meeting the state reading standards in the third grade and at the 10 schools with the lowest proportion of students meeting the state reading standards in the third grade.
(a) Display a dotplot for each group to compare the distribution of student-to-teacher ratios in the top 10 schools with the distribution in the bottom 10 schools. Comment on the similarities and differences between the two distributions.
(b) Any statistical test that is used to determine whether the mean student-to-teacher ratio is the same for the top 10 schools as it is for the bottom 10 schools would be inappropriate. Explain why in a few sentences.

Most-appropriate topic codes (AP Statistics):

• Topic 1.5 — Graphical Representations for One Quantitative Variable (Part \(\mathrm{a}\))
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Part \(\mathrm{a}\))
• Topic 1.11 — Random Sampling (Part \(\mathrm{b}\))
• Topic 4.4 — Setting Up a Test for a Population Mean or Population Mean Difference (Part \(\mathrm{b}\))
▶️ Answer/Explanation

(a)
First, let’s organize the data for both groups:
Highest Proportion group: \(7, 9, 12, 16, 16, 17, 17, 18, 21, 22\)
Lowest Proportion group: \(12, 12, 14, 14, 16, 16, 18, 19, 20, 20\)
The dotplots, displayed on a common scale from \(4\) to \(24\), are shown below:

Similarities: The two distributions are centered at approximately the same place. The median for the Highest Proportion group is \(\dfrac{16+17}{2} = 16.5\) and the median for the Lowest Proportion group is \(\dfrac{16+16}{2} = 16\), so both centers are very close to \(16\).
Differences: The distribution for the Highest Proportion group is much more spread out (variable) than the distribution for the Lowest Proportion group. The range for the Highest Proportion group is \(22 – 7 = 15\), while the range for the Lowest Proportion group is only \(20 – 12 = 8\). In other words, the top schools show much greater variability in their student-to-teacher ratios compared to the bottom schools.

(b)
The two groups of schools are not random samples drawn from two larger populations of interest.
The group of 10 schools with the highest proportion of students meeting the standards is itself the entire population of such schools — it is not a random sample from some larger population of high-performing schools.
Similarly, the group of 10 schools with the lowest proportion is itself the complete population of the lowest-performing schools in the state — not a random sample from a larger population.
Since statistical inference is designed to generalize conclusions from a sample to a broader population, and these two groups are not random samples but rather complete populations defined by their extreme values, applying any inferential procedure (such as a two-sample \(t\)-test) to these data would be inappropriate. There is no larger population to generalize to.

Question

The nerves that supply sensation to the front portion of a person’s foot run between the long bones of the foot. Tight-fitting shoes can squeeze these nerves between the bones, causing pain when the nerves swell. This condition is called Morton’s neuroma. Because most people have a dominant foot, muscular development is not the same in both feet. People who have Morton’s neuroma may have the condition in only one foot or they may have it in both feet.
Investigators selected a random sample of 12 adult female patients with Morton’s neuroma to study this disease further. The data below are measurements of nerve swelling as recorded by a physician. A value of 1.0 is considered “normal,” and 2.0 is considered extreme swelling. The population distribution of the swelling measurements is approximately normal for adult females who have Morton’s neuroma.
(a) A scatterplot of the ordered pairs (swelling in left foot, swelling in right foot), is shown below.
The scatterplot suggests there are two distinct groups of patients. Patients within each group share a common trait. Use the scatterplot above and the table to determine the common trait and explain how this trait differs for the two groups.
(b) A scatterplot of the ordered pairs (swelling in dominant foot, swelling in nondominant foot), is shown below.
What conclusion can be drawn from this scatterplot that is not apparent from the scatterplot in part (a)?
(c) Can you conclude that there is a difference between the mean swelling in the dominant foot and the mean swelling in the nondominant foot for adult females who have Morton’s neuroma in at least one foot? Give a statistical justification to support your answer.
(For easy reference, the table of data from above also appears at the bottom of this question.)
(d) The nerve swelling measurement is used to indicate whether a foot has Morton’s neuroma. Use the 24 measurements of nerve swelling to suggest a criterion for diagnosing Morton’s neuroma. Justify your suggestion graphically.
(For easy reference, the table of data from above also appears below.)

Most-appropriate topic codes (AP Statistics):

• Topic 5.1 — Graphical Representations Between Two Quantitative Variables (Parts \(\mathrm{a}\), \(\mathrm{b}\))
• Topic 4.4 — Setting Up a Test for a Population Mean or Population Mean Difference (Part \(\mathrm{c}\))
• Topic 4.5 — Carrying Out a Test for a Population Mean or Population Mean Difference (Part \(\mathrm{c}\))
• Topic 1.5 — Graphical Representations for One Quantitative Variable (Part \(\mathrm{d}\))
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Part \(\mathrm{d}\))
▶️ Answer/Explanation

(a)

The trait that distinguishes the two groups in the scatterplot is the dominant foot (left or right). All the points in the upper-left cluster represent patients whose dominant foot is the right foot, while all the points in the lower-right cluster represent patients whose dominant foot is the left foot. The dominant foot type is the common trait, and it differs between the two groups.

(b)

Two conclusions become clear from this scatterplot that were not visible before:
First, there is a positive linear relationship between swelling in the dominant foot and swelling in the nondominant foot — as swelling in the dominant foot increases, swelling in the nondominant foot tends to increase as well.
Second, and importantly, every single point lies below the line \(y = x\), which means swelling in the dominant foot is consistently greater than swelling in the nondominant foot for all patients in the sample. This pattern across both groups combined is something you simply could not see in the left-foot vs. right-foot scatterplot from part (a).

(c)

We perform a matched-pairs \(t\)-test on the differences \(d_i = \text{(dominant swelling)} – \text{(nondominant swelling)}\).
The 12 differences are:
\(0.30,\ 0.30,\ 0.45,\ 0.15,\ 0.30,\ 0.35,\ 0.25,\ 0.35,\ 0.20,\ 0.25,\ 0.40,\ 0.15\)
State hypotheses (where \(\mu_d\) is the mean difference, dominant minus nondominant):
\(H_0: \mu_d = 0\)
\(H_a: \mu_d \neq 0\)
Check conditions:
1. We are told a random sample was selected from the population of adult females with Morton’s neuroma.
2. A dotplot of the differences shows a roughly symmetric, unimodal distribution with no outliers — it is reasonable to treat the population of differences as approximately normal.
Compute the test statistic:
\(\bar{x}_d = 0.2875, \quad s_d = 0.0932, \quad n = 12, \quad df = 11\)
\(t = \dfrac{\bar{x}_d – 0}{\dfrac{s_d}{\sqrt{n}}} = \dfrac{0.2875 – 0}{\dfrac{0.0932}{\sqrt{12}}} = 10.68\)
\(p\text{-value} \approx 0.0000004 \approx 0\)
Since the \(p\)-value is essentially \(0\), which is far less than any reasonable significance level \(\alpha\), we reject \(H_0\). There is very convincing statistical evidence that the mean swelling in the dominant foot is different from (and specifically greater than) the mean swelling in the nondominant foot for adult females who have Morton’s neuroma in at least one foot.

(d)

To suggest a diagnostic criterion, we separate all 24 swelling measurements into two groups: the 17 foot measurements from feet that have Morton’s neuroma and the 7 foot measurements from feet that do not have Morton’s neuroma. A stacked dotplot of the two groups is shown below:

The dotplot makes it visually clear that all 7 feet without Morton’s neuroma have swelling measurements of \(1.40\) or below, while the feet with Morton’s neuroma have swelling values of \(1.40\) and above (with the measurements extending up to \(1.85\)). Based on this graphical display, a reasonable diagnostic criterion is:
\(\boxed{\text{Swelling measurement} \geq 1.4 \Rightarrow \text{diagnose Morton’s neuroma}}\)
A cutoff of approximately \(1.4\) or higher serves as a sensible threshold for diagnosing Morton’s neuroma, since it cleanly separates the feet with and without the condition in this dataset.

Question

The department of agriculture at a university was interested in determining whether a preservative was effective in reducing discoloration in frozen strawberries. A sample of 50 ripe strawberries was prepared for freezing. Then the sample was randomly divided into two groups of 25 strawberries each. Each strawberry was placed into a small plastic bag.
The 25 bags in the control group were sealed. The preservative was added to the 25 bags containing strawberries in the treatment group, and then those bags were sealed. All bags were stored at \(0^\circ\text{C}\) for a period of 6 months. At the end of this time, after the strawberries were thawed, a technician rated each strawberry’s discoloration from 1 to 10, with a low score indicating little discoloration.
The dotplots below show the distributions of discoloration rating for the control and treatment groups.
(a) The standard deviation of ratings for the control group is 2.141. Explain how this value summarizes variability in the control group.
(b) Based on the dotplots, comment on the effectiveness of the preservative in lowering the amount of discoloration in strawberries. (No calculations are necessary.)
(c) Researchers at the university decided to calculate a 95 percent confidence interval for the difference in mean discoloration rating between strawberries that were not treated with preservative and those that were treated with preservative. The confidence interval they obtained was \((0.16,\ 2.72)\). Assume that the conditions necessary for the \(t\)-confidence interval are met.
Based on the confidence interval, comment on whether there would be a difference in the population mean discoloration ratings for the treated and untreated strawberries.

Most-appropriate topic codes (AP Statistics):

• Topic 1.7 — Summary Statistics for One Quantitative Variable (Part \(\mathrm{a}\))
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Part \(\mathrm{b}\))
• Topic 4.7 — Constructing a Confidence Interval for the Difference Between Two Population Means (Part \(\mathrm{c}\))
• Topic 4.8 — Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Means (Part \(\mathrm{c}\))
▶️ Answer/Explanation

(a)

The standard deviation \(s = 2.141\) tells us roughly how far a typical discoloration rating in the control group falls from the group’s mean rating.
In other words, on average, a single strawberry’s discoloration rating in the control group deviates from the mean discoloration rating by about \(2.141\) points.
A larger standard deviation would indicate that the individual ratings are more spread out from the mean, while a smaller value would indicate the ratings are clustered more tightly around the mean.
\(\boxed{s = 2.141 \text{ is the typical distance of each discoloration rating from the mean in the control group.}}\)

(b)

The preservative does appear to have been effective in lowering the amount of discoloration in strawberries.
Looking at the dotplots, the treatment group’s distribution is clearly centered at a lower value than the control group’s distribution — the treatment group’s center is around 5, while the control group’s center is around 7.
Additionally, nearly all the summary statistics (minimum, \(Q_1\), median, and \(Q_3\)) are lower for the treatment group than for the control group, which consistently supports the conclusion that the preservative reduced discoloration.
\(\boxed{\text{The preservative appears effective: treatment ratings are consistently lower than control ratings.}}\)

(c)

We are given a 95% confidence interval for \(\mu_{\text{control}} – \mu_{\text{treatment}}\) as \((0.16,\ 2.72)\).
Since the entire interval lies above zero — meaning zero is not contained within \((0.16,\ 2.72)\) — we have convincing evidence that the two population means are not equal.
We are 95% confident that the population mean discoloration rating for untreated strawberries is between \(0.16\) and \(2.72\) units higher than the population mean rating for treated strawberries, indicating the preservative does make a real difference.
\(\boxed{0 \notin (0.16,\ 2.72) \Rightarrow \text{there is a significant difference in population mean discoloration ratings.}}\)

Question

The Better Business Council of a large city has concluded that students in the city’s schools are not learning enough about economics to function in the modern world. These findings were based on test results from a random sample of \(20\) twelfth-grade students who completed a \(46\)-question multiple-choice test on basic economic concepts. The data set below shows the number of questions that each of the \(20\) students in the sample answered correctly.
12 16 18 17 18 33 41 44 38 35
19 36 19 13 43 8 16 14 10 9
(a) Display these data in a stemplot.
(b) Use your stemplot from part (a) to describe the main features of this score distribution.
(c) Why would it be misleading to report only a measure of center for this score distribution?

Most-appropriate topic codes (AP Statistics):

• Topic 1.5 — Graphical Representations for One Quantitative Variable (Part \(\mathrm{a}\))
• Topic 1.6 — Descriptions for One Quantitative Variable Distributions (Part \(\mathrm{b}\))
• Topic 1.7 — Summary Statistics for One Quantitative Variable (Part \(\mathrm{c}\))
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Parts \(\mathrm{b}\), \(\mathrm{c}\))
▶️ Answer/Explanation

(a)
To build the stemplot, use the tens digit as the stem and the units digit as the leaf. The data range from \(8\) to \(44\), so the stems are \(0, 1, 2, 3, 4\). Place each data value’s units digit on the appropriate row.
\( \begin{array}{r|l} 0 & 8 \quad 9 \\ 1 & 2 \quad 6 \quad 8 \quad 7 \quad 8 \quad 9 \quad 6 \quad 9 \quad 3 \quad 4 \quad 0 \\ 2 & \\ 3 & 3 \quad 8 \quad 6 \quad 5 \\ 4 & 1 \quad 4 \quad 3 \\ \end{array} \)
Legend: \(1 \mid 2\) represents \(12\) questions answered correctly.
Ordered version (leaves sorted in ascending order):
\( \begin{array}{r|l} 0 & 8 \quad 9 \\ 1 & 0 \quad 2 \quad 3 \quad 4 \quad 6 \quad 6 \quad 7 \quad 8 \quad 8 \quad 9 \quad 9 \\ 2 & \\ 3 & 3 \quad 5 \quad 6 \quad 8 \\ 4 & 1 \quad 3 \quad 4 \\ \end{array} \)
Legend: \(1 \mid 2\) represents \(12\) questions answered correctly.

(b)
The most striking feature of this distribution is that the scores split into two distinct groups — one cluster concentrated in the teens (roughly \(8\)–\(19\)) and a second, smaller cluster in the high \(30\)s and low \(40\)s. There is a complete gap in the \(20\)s, with no student scoring between \(19\) and \(33\). The distribution is bimodal, suggesting that students tended to perform either quite poorly or quite well on the test, with no middle ground.

(c)
A measure of center such as the mean or median would fall somewhere around \(\bar{x} \approx 22.95\), which lies right in the gap of the \(20\)s — a region where no student actually scored. Reporting only the center gives the false impression that a typical student scored around \(23\), when in reality no student did. Because the data are bimodal with two separated groups, the center alone completely misses the two-cluster structure of the distribution and would mislead the reader into thinking the scores are roughly symmetric or unimodal.

Question

Two parents have each built a toy catapult for use in a game at an elementary school fair. To play the game, students will attempt to launch Ping-Pong balls from the catapults so that the balls land within a 5-centimeter band. A target line will be drawn through the middle of the band, as shown in the figure below. All points on the target line are equidistant from the launching location.
If a ball lands within the shaded band, the student will win a prize.
The parents have constructed the two catapults according to slightly different plans. They want to test these catapults before building additional ones. Under identical conditions, the parents launch 40 Ping-Pong balls from each catapult and measure the distance that the ball travels before landing. Distances to the nearest centimeter are graphed in the dotplots below.

(a) Comment on any similarities and any differences in the two distributions of distances traveled by balls launched from catapult A and catapult B.
(b) If the parents want to maximize the probability of having the Ping-Pong balls land within the band, which one of the two catapults, A or B, would be better to use than the other? Justify your choice.
(c) Using the catapult that you chose in part (b), how many centimeters from the target line should this catapult be placed? Explain why you chose this distance.

Most-appropriate topic codes (AP Statistics):

• Topic 1.6 — Descriptions for One Quantitative Variable Distributions (Part \(\mathrm{a}\))
• Topic 1.7 — Summary Statistics for One Quantitative Variable (Parts \(\mathrm{a}\), \(\mathrm{c}\))
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Parts \(\mathrm{a}\), \(\mathrm{b}\))
▶️ Answer/Explanation

(a)

Both distributions of distances are roughly symmetric and somewhat mound-shaped (bell-shaped). Looking at the centers, the median of Catapult A is approximately \(136\,\text{cm}\), which is slightly lower than the median of Catapult B at approximately \(138\,\text{cm}\). In terms of spread, Catapult A shows considerably more variability than Catapult B — the range of Catapult A is about \(30\,\text{cm}\), while the range of Catapult B is approximately \(11\,\text{cm}\). Additionally, there appear to be potential outliers in Catapult A’s distribution (e.g., a ball traveling approximately \(155\,\text{cm}\)), whereas Catapult B has no such extreme values.

(b)

Catapult B would be the better choice.

Since the target band is only \(5\,\text{cm}\) wide, the key factor is how tightly clustered the distances are around the center. Catapult B has a much smaller spread — most balls land between approximately \(133\,\text{cm}\) and \(143\,\text{cm}\) — meaning when placed correctly, a higher proportion of balls will fall within the narrow band. Catapult A’s larger variability means balls are scattered over a much wider range, making it far less likely that they will land consistently within the band.

(c)

Catapult B should be placed approximately \(\boxed{138\,\text{cm}}\) from the target line.

Since Catapult B’s distribution is roughly symmetric and mound-shaped, the median (approximately \(138\,\text{cm}\)) is a reliable measure of center and represents the most typical distance a ball will travel. Placing the catapult so that the target line is \(138\,\text{cm}\) away aligns the center of the distribution with the target, maximizing the chance that any given ball lands within the \(5\,\text{cm}\) band. Based on the sample data, approximately \(\frac{30}{40} = 0.75\) of the 40 balls launched from Catapult B landed within \(2.5\,\text{cm}\) on either side of \(138\,\text{cm}\), which further confirms this placement.

Question

The goal of a nutritional study was to compare the caloric intake of adolescents living in rural areas of the United States with the caloric intake of adolescents living in urban areas of the United States. A random sample of ninth-grade students from one high school in a rural area was selected. Another random sample of ninth graders from one high school in an urban area was also selected. Each student in each sample kept records of all the food he or she consumed in one day.
The back-to-back stemplot below displays the number of calories of food consumed per kilogram of body weight for each student on that day.
(a) Write a few sentences comparing the distribution of the daily caloric intake of ninth-grade students in the rural high school with the distribution of the daily caloric intake of ninth-grade students in the urban high school.
(b) Is it reasonable to generalize the findings of this study to all rural and urban ninth-grade students in the United States? Explain.
(c) Researchers who want to conduct a similar study are debating which of the following two plans to use.
Plan I: Have each student in the study record all the food he or she consumed in one day. Then researchers would compute the number of calories of food consumed per kilogram of body weight for each student for that day.
Plan II: Have each student in the study record all the food he or she consumed over the same 7-day period. Then researchers would compute the average daily number of calories of food consumed per kilogram of body weight for each student during that 7-day period.
Assuming that the students keep accurate records, which plan, I or II, would better meet the goal of the study? Justify your answer.

Most-appropriate topic codes (AP Statistics):

• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Part \(\mathrm{a}\))
• Topic 1.11 — Random Sampling (Part \(\mathrm{b}\))
• Topic 1.12 — Potential Problems with Sampling (Part \(\mathrm{b}\))
• Topic 1.10 — The Investigative Question Revisited and Data Collection (Part \(\mathrm{c}\))
▶️ Answer/Explanation

(a)
Reading the stemplot, the rural distribution is centered higher and is more spread out than the urban distribution.

For the rural students:

Mean \(\approx 40.45\) cal/kg
Median \(\approx 41\) cal/kg
Range \(= 19\)
SD \(\approx 6.04\)
IQR \(\approx 10\)

For the urban students:

Mean \(\approx 32.6\) cal/kg
Median \(\approx 32\) cal/kg
Range \(= 16\)
SD \(\approx 4.67\)
IQR \(\approx 7\)

So both the typical value and the spread are larger for the rural group. In terms of shape, the rural data look fairly symmetric and spread evenly between about 32 and 51 cal/kg, while the urban data appear skewed toward the larger values.

\( \boxed{\text{Rural: higher center and more spread; Urban: lower center, less spread, right-skewed}} \)

(b)
No. Each sample came from just one rural school and one urban school, so these two specific schools may not represent the much larger and more diverse population of all rural and urban ninth graders across the country. Because the schools themselves were not randomly chosen from all such schools, the results can’t be safely extended beyond these two schools.

\( \boxed{\text{No — only one school of each type was sampled, so results cannot be generalized nationally}} \)

(c)
Plan II is the better choice.

Both plans already adjust for body size by dividing calories by body weight, so that part is the same. The real issue is that a single day’s eating can be unusually high or low depending on what happened that day — a birthday party, a sick day, a weekend versus a school day, and so on. By recording food over a full 7-day period and averaging, Plan II smooths out this day-to-day variability and gives a more stable, precise picture of each student’s typical caloric intake.

\( \boxed{\text{Plan II — averaging over 7 days reduces day-to-day variability and gives a more precise estimate}} \)

Question

Let the random variable \(X\) represent the number of telephone lines in use by the technical support center of a software manufacturer at noon each day. The probability distribution of \(X\) is shown in the table below.
(a) Calculate the expected value (the mean) of \(X\).
(b) Using past records, the staff at the technical support center randomly selected 20 days and found that an average of 1.25 telephone lines were in use at noon on those days. The staff proposes to select another random sample of 1,000 days and compute the average number of telephone lines that were in use at noon on those days. How do you expect the average from this new sample to compare to that of the first sample? Justify your response.
(c) The median of a random variable is defined as any value \(x\) such that \(P(X \le x) \ge 0.5\) and \(P(X \ge x) \ge 0.5\). For the probability distribution shown in the table above, determine the median of \(X\).
(d) In a sentence or two, comment on the relationship between the mean and the median relative to the shape of this distribution.

Most-appropriate topic codes (AP Statistics):

• Topic 2.9 — Parameters of Random Variables (Part \(\mathrm{a}\))
• Topic 2.12 — Sampling Distributions and the Central Limit Theorem (Part \(\mathrm{b}\))
• Topic 1.7 — Summary Statistics for One Quantitative Variable (Part \(\mathrm{c}\))
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Part \(\mathrm{d}\))
▶️ Answer/Explanation

(a)
The expected value of a discrete random variable is found by multiplying each value by its probability and adding the results:
\( E(X)=\sum x_i\,p(x_i) \)
Plugging in the values from the table:
\( E(X)=0(0.35)+1(0.20)+2(0.15)+3(0.15)+4(0.10)+5(0.05) \)
\( E(X)=0+0.20+0.30+0.45+0.40+0.25 \)
\( \boxed{E(X)=1.6} \)

(b)
Both sample averages are estimates of the same population mean, \(\mu_X=1.6\), so they should be centered around the same value. What changes is how precise that estimate is.
The standard deviation of a sample mean is
\( \sigma_{\bar{x}}=\dfrac{\sigma}{\sqrt{n}} \)
Since \(1{,}000>20\), a sample of 1,000 days has a much smaller standard deviation for its average, meaning the sample average will tend to land closer to 1.6 than the first sample’s average of 1.25 did.
\( \boxed{\text{The new average should be closer to } 1.6\text{ than }1.25\text{, since larger samples have less variability}} \)

(c)
To find the median, build up the cumulative probabilities:

At \(x=1\), both \(P(X\le 1)=0.55\ge 0.5\) and \(P(X\ge 1)=0.65\ge 0.5\) are satisfied. At \(x=2\), \(P(X\ge 2)=0.45<0.5\), so \(x=2\) does not work.
\( \boxed{\text{Median} = 1} \)

(d)
Looking at the table, most of the probability is piled up at the small values (0, 1, 2) with a long tail of smaller probabilities stretching out to 4 and 5. This makes the distribution right-skewed, which pulls the mean up above the median — here the mean of 1.6 is noticeably larger than the median of 1, which is exactly the pattern you’d expect for a distribution skewed toward larger values.
\( \boxed{\text{The distribution is right-skewed, so the mean (1.6) is greater than the median (1)}} \)

Question

The graph below displays the scores of 32 students on a recent exam. Scores on this exam ranged from 64 to 95 points.
\(6\phantom{|}\) * *
\(6\phantom{|}\) * *
\(7\phantom{|}\) * * *
\(7\phantom{|}\) * * * *
\(8\phantom{|}\) * * * *
\(8\phantom{|}\) * * * * * *
\(9\phantom{|}\) * * * * * * *
\(9\phantom{|}\) * * * *
(a) Describe the shape of this distribution.
(b) In order to motivate her students, the instructor of the class wants to report that, overall, the class’s performance on the exam was high. Which summary statistic, the mean or the median, should the instructor use to report that overall exam performance was high? Explain.
(c) The midrange is defined as \(\dfrac{\text{maximum} + \text{minimum}}{2}\). Compute this value using the data on the preceding page.
Is the midrange considered a measure of center or a measure of spread? Explain.

Most-appropriate topic codes (AP Statistics):

• Topic 1.6 — Descriptions for One Quantitative Variable Distributions (Part a)
• Topic 1.7 — Summary Statistics for One Quantitative Variable (Parts b, c)
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Part b)
▶️ Answer/Explanation

(a)
The distribution is skewed to the left (skewed toward the lower values). You can see this from the stemplot: the longer tail stretches down into the 60s and lower 70s, while most of the data cluster in the upper 80s and 90s. There are relatively few low scores pulling the tail downward.

(b)
The instructor should report the median.
Because the distribution is skewed toward the lower values, the mean gets pulled in that direction — it will be lower than the median. The median, being resistant to the few very low scores, will better represent the “typical” high performance of the class. So ironically, to make performance look as high as possible, the median is the better choice here.

(c)
Step 1 — Compute the midrange:
The minimum score is \(64\) and the maximum score is \(95\), so:
\( \text{midrange} = \frac{\text{maximum} + \text{minimum}}{2} = \frac{95 + 64}{2} = \frac{159}{2} = 79.5 \)
\(\boxed{\text{midrange} = 79.5}\)

Step 2 — Identify as measure of center:
The midrange is a measure of center.

Step 3 — Rationale:
The maximum value tells us about the upper extreme and the minimum tells us about the lower extreme. By averaging these two values, we find the point that sits exactly halfway between the two extremes — this is a center point, not a spread. Measures of spread (like range, standard deviation, or IQR) describe how far apart the data are; the midrange instead gives a single representative middle value, placing it firmly in the category of measures of center.

Question

A consumer advocate conducted a test of two popular gasoline additives, A and B. There are claims that the use of either of these additives will increase gasoline mileage in cars. A random sample of 30 cars was selected. Each car was filled with gasoline and the cars were run under the same driving conditions until the gas tanks were empty. The distance traveled was recorded for each car.
Additive A was randomly assigned to 15 of the cars and additive B was randomly assigned to the other 15 cars. The gas tank of each car was filled with gasoline and the assigned additive. The cars were again run under the same driving conditions until the tanks were empty. The distance traveled was recorded and the difference in the distance with the additive minus the distance without the additive for each car was calculated.
The following table summarizes the calculated differences. Note that negative values indicate less distance was traveled with the additive than without the additive.
(a) On the grid below, display parallel boxplots (showing outliers, if any) of the differences of the two additives.
(b) Two ways that the effectiveness of a gasoline additive can be evaluated are by looking at either
• the proportion of cars that have increased gas mileage when the additive is used in those cars
or
• the mean increase in gas mileage when the additive is used in those cars.
i. Which additive, A or B, would you recommend if the goal is to increase gas mileage in the highest proportion of cars? Explain your choice.
ii. Which additive, A or B, would you recommend if the goal is to have the highest mean increase in gas mileage? Explain your choice.

Most-appropriate topic codes (AP Statistics):

• Topic 1.8 — Graphical Representations of Summary Statistics for One Quantitative Variable (Part a)
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Parts a, b)
• Topic 1.7 — Summary Statistics for One Quantitative Variable (Part b)
▶️ Answer/Explanation

(a)
First, check for outliers using the \(1.5 \times \text{IQR}\) rule.

For Additive A:
\(\text{IQR}_A = Q_3 – Q_1 = 4 – 1 = 3\)
\(1.5 \times \text{IQR}_A = 1.5 \times 3 = 4.5\)
Lower fence: \(Q_1 – 4.5 = 1 – 4.5 = -3.5\)
Upper fence: \(Q_3 + 4.5 = 4 + 4.5 = 8.5\)
Values below \(-3.5\): \(-10, -8\) → both are outliers
Values above \(8.5\): \(9\) → outlier

For Additive B:
\(\text{IQR}_B = Q_3 – Q_1 = 25 – (-2) = 27\)
\(1.5 \times \text{IQR}_B = 1.5 \times 27 = 40.5\)
Lower fence: \(-2 – 40.5 = -42.5\)
Upper fence: \(25 + 40.5 = 65.5\)
No values fall outside these fences → no outliers

The parallel boxplots should be drawn as follows:

(b)(i)
Additive A is the better recommendation when the goal is to improve mileage in the highest proportion of cars. Since \(Q_1 = 1 > 0\) for Additive A, we know that at least 75% of the cars in the sample showed a positive difference (i.e., improved mileage). For Additive B, \(Q_1 = -2 < 0\), so we can only be certain that at least 50% of the cars showed improvement — and it could be less than 75%. Therefore, Additive A is more reliable for maximizing the proportion of cars that benefit.
\(\boxed{\text{Recommend Additive A for highest proportion of cars with improved mileage}}\)

(b)(ii)
Additive B is the better recommendation when the goal is to maximize the mean increase in gas mileage. Although the median for Additive B (Median \(= 1\)) is lower than that of Additive A (Median \(= 3\)), the distribution for Additive B is strongly right-skewed, with very large values above \(Q_3\) such as \(35, 37,\) and \(40\). This skewness pulls the mean well above the median. In contrast, Additive A has two low outliers (\(-10, -8\)) pulling the mean downward. So the mean for Additive B is expected to be substantially higher than the mean for Additive A.
\(\boxed{\text{Recommend Additive B for highest mean increase in gas mileage}}\)

Question

A researcher thinks that modern Thai dogs may be descendants of golden jackals. A random sample of 16 animals was collected from each of the two populations. The length (in millimeters) of the mandible (jawbone) was measured for each animal. The lower quartile, median, and upper quartile for each sample are shown in the table below, along with all values below the lower quartile and all values above the upper quartile.

(a) Display parallel boxplots of mandible lengths (showing outliers, if any) for the modern Thai dogs and the golden jackals on the grid below.
Based on the boxplots, write a few sentences comparing the distributions of mandible lengths for the two types of dogs.
(b) Is it reasonable to use the sample of mandible lengths of modern Thai dogs to construct an interval estimate of the mean mandible length for the population of modern Thai dogs? Justify your answer. (Note: You do not have to compute the interval.)
(c) Is it reasonable to use the sample data of mandible lengths of modern Thai dogs and the sample data of mandible lengths of golden jackals to perform a two-sample \(t\)-test for the difference in mean mandible lengths for the two types of dogs? Justify your answer. (Note: You do not have to conduct the test.)

Most-appropriate topic codes (AP Statistics):

• Topic 1.8 — Graphical Representations of Summary Statistics for One Quantitative Variable (Part a)
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Part a)
• Topic 4.2 — Constructing a Confidence Interval for a Population Mean or Population Mean Difference (Part b)
• Topic 4.4 — Setting Up a Test for a Population Mean or Population Mean Difference (Part c)
▶️ Answer/Explanation

(a)

First, we check for outliers in each sample using the \(1.5 \times \text{IQR}\) rule.
Modern Thai Dogs:
\(\text{IQR} = Q_3 – Q_1 = 128 – 121 = 7\)
Lower fence: \(121 – 1.5(7) = 121 – 10.5 = 110.5\)
Upper fence: \(128 + 1.5(7) = 128 + 10.5 = 138.5\)
All values (minimum = 114, maximum = 132) fall within these fences. No outliers.
Golden Jackals:
\(\text{IQR} = Q_3 – Q_1 = 112 – 107 = 5\)
Lower fence: \(107 – 1.5(5) = 107 – 7.5 = 99.5\)
Upper fence: \(112 + 1.5(5) = 112 + 7.5 = 119.5\)
Values 122, 124, and 125 exceed the upper fence of 119.5. Outliers: 122, 124, 125.
The parallel boxplots (with the scale from 100 to 140 mm) are shown below:

Comparison of distributions: The distributions of mandible lengths for modern Thai dogs and golden jackals are quite different. Modern Thai dogs have a much larger typical mandible length — a median of 125 mm — compared to golden jackals, whose median is only 108 mm. The distribution for modern Thai dogs appears approximately symmetric with no outliers, whereas the distribution for golden jackals is heavily skewed to the right, with three high outliers (122, 124, and 125 mm). The variability (spread) of the two distributions is roughly similar in terms of IQR, but the overall range for golden jackals is larger once the outliers are included.

(b)

Yes, it is reasonable to construct a \(t\)-confidence interval for the mean mandible length of modern Thai dogs. The boxplot for this sample is roughly symmetric with no outliers, which provides support for the assumption that the underlying population distribution is approximately normal. Since the data come from a random sample and the normality condition is reasonably satisfied even with a sample size of only 16, using a one-sample \(t\)-interval is appropriate here.

(c)

No, it would not be reasonable to perform a two-sample \(t\)-test using both groups. While the modern Thai dog sample looks approximately normal, the golden jackal sample is clearly not. The boxplot for golden jackals is strongly skewed to the right and contains three high outliers (122, 124, 125) in a sample of only 16 animals — a substantial proportion of the data. With such a small sample size, the \(t\)-test is not robust enough to overcome this serious departure from normality, so the normality condition required for the two-sample \(t\)-test is not reasonably met for the golden jackal population. Therefore, performing the two-sample \(t\)-test with this data would not be appropriate.

Question

Since Hill Valley High School eliminated the use of bells between classes, teachers have noticed that more students seem to be arriving to class a few minutes late. One teacher decided to collect data to determine whether the students’ and teachers’ watches are displaying the correct time. At exactly \(12:00\) noon, the teacher asked \(9\) randomly selected students and \(9\) randomly selected teachers to record the times on their watches to the nearest half minute. The ordered data showing minutes after \(12:00\) as positive values and minutes before \(12:00\) as negative values are shown in the table below.
(a) Construct parallel boxplots using these data.
(b) Based on the boxplots in part (a), which of the two groups, students or teachers, tends to have watch times that are closer to the true time? Explain your choice.
(c) The teacher wants to know whether individual student’s watches tend to be set correctly. She proposes to test \(H_0:\mu=0\) versus \(H_a:\mu\neq0\), where \(\mu\) represents the mean amount by which all student watches differ from the correct time. Is this an appropriate pair of hypotheses to test to answer the teacher’s question? Explain why or why not. Do not carry out the test.

Most-appropriate topic codes (AP Statistics):

• Topic 1.8 — Graphical Representations of Summary Statistics for One Quantitative Variable (Part a)
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Part b)
• Topic 4.4 — Setting Up a Test for a Population Mean or Population Mean Difference (Part c)
▶️ Answer/Explanation

(a)
To build a boxplot, we need the five-number summary for each group. Since the data are already sorted, we just split each set in half around the median.
For the students (\(n=9\)):
\( \text{Min}=-4.5 \)
\( Q_1=\dfrac{-3.0+(-0.5)}{2}=-1.75 \)
\( \text{Median}=0 \)
\( Q_3=\dfrac{0.5+1.5}{2}=1.0 \)
\( \text{Max}=5.0 \)
For the teachers (\(n=9\)):
\( \text{Min}=-2.0 \)
\( Q_1=\dfrac{-1.5+(-1.5)}{2}=-1.5 \)
\( \text{Median}=-1.0 \)
\( Q_3=\dfrac{0+0}{2}=0 \)
\( \text{Max}=0.5 \)
Neither group has any outliers, since no value falls beyond \(1.5\times\text{IQR}\) from the nearer quartile. Plotting both sets of five-number summaries on the same number line (a common scale is essential so the two groups can be compared directly) gives:

Each boxplot uses the same horizontal scale, with “T” marking the teachers’ plot and “S” marking the students’ plot, so the two distributions can be lined up and compared directly.

(b)
The teachers’ watch times tend to be closer to the true noon time. Looking at the two boxplots, the teachers’ values are all squeezed into a fairly narrow band (roughly from \(-2.0\) to \(0.5\)), while the students’ values are spread out over a much wider range (from \(-4.5\) to \(5.0\)). Even though the teachers’ watches tend to run a little slow on average (the box is shifted slightly below \(0\)), their times stay much closer together and closer to zero than the students’ times do. Since “closer to the true time” is really about how small the time errors tend to be, the group with less spread — the teachers — is the better answer here.

(c)
No, this is not an appropriate pair of hypotheses for answering the teacher’s question. The teacher wants to know whether individual students’ watches tend to be set correctly, but \(H_0:\mu=0\) versus \(H_a:\mu\neq0\) only tests something about the average error across all student watches. It’s entirely possible for this mean to come out close to \(0\) even if no individual watch is actually correct — for instance, if some students’ watches run fast by a few minutes and others run slow by a few minutes, those errors could cancel out in the average, making \(\mu\) look like \(0\) overall. So a test about the population mean \(\mu\) doesn’t tell us anything about how far off each individual student’s watch tends to be from the true time.

Scroll to Top