Home / AP® Exam / AP® Statistics / AP Statistics 1.6 Descriptions for One Quantitative Variable Distributions- Exam Style Questions – FRQs

AP Statistics 1.6 Descriptions for One Quantitative Variable Distributions- Exam Style Questions - FRQs - New Syllabus

Question

As part of a study on the chemistry of Alaskan streams, researchers took water samples from many streams with temperatures colder than 8°C and from many streams with temperatures warmer than 8°C. For each sample, the researchers measured the dissolved oxygen concentration, in milligrams per liter (mg/l).
(a) The researchers constructed the histogram shown for the dissolved oxygen concentration in streams from the sample with water temperatures colder than 8°C. Based on the histogram, describe the distribution of dissolved oxygen concentration in streams with water temperatures colder than 8°C.
(b) The researchers computed the summary statistics shown in the table for the dissolved oxygen concentration in streams from the sample with water temperatures warmer than 8°C. Use the summary statistics to construct a box plot for the dissolved oxygen concentration in streams with water temperatures warmer than 8°C. Do not indicate outliers.
(c) The researchers believe that streams with higher dissolved oxygen concentration are generally healthier for wildlife. Which streams are generally healthier for wildlife, those with water temperature colder than 8°C or those with water temperature warmer than 8°C? Using characteristics of the distribution of dissolved oxygen concentration for temperatures colder than 8°C and characteristics of the distribution of dissolved oxygen concentration for temperatures warmer than 8°C, justify your answer.
 

Most-appropriate topic codes (AP Statistics):

• Topic \(1.5\) — Graphical Representations for One Quantitative Variable (Parts \( \mathrm{a} \), \( \mathrm{b} \))
• Topic \(1.6\) — Descriptions for One Quantitative Variable Distributions (Part \( \mathrm{a} \))
• Topic \(1.9\) — Comparisons of the Distributions for One Quantitative Variable (Part \( \mathrm{c} \))
▶️ Answer/Explanation

(a)
Shape: The distribution is heavily skewed to the left (negatively skewed).
Center: The center of the distribution is around a median of between 11 and 12 mg/l (which is also the modal interval).
Spread: The dissolved oxygen concentrations vary from a minimum between 2 and 3 mg/l to a maximum between 13 and 14 mg/l.
Unusual Features: There appear to be potential low outliers between 2 and 6 mg/l.

(b)


To construct the box plot, follow these steps plotted against an appropriate scale (e.g., 2 to 14):
Draw a vertical line segment at the median: $M = 5.43$
Draw vertical line segments at the quartiles: $Q_1 = 4.39$ and $Q_3 = 6.12$, and connect them to form the central box.
Draw horizontal whiskers extending from the box to the minimum value at $2.10$ and the maximum value at $13.45$.

(c)
The streams with temperatures colder than 8°C are generally healthier for wildlife.
Justification: We must compare the centers and spreads. The median dissolved oxygen concentration for the colder streams (between 11 and 12 mg/l) is significantly higher than the median for the warmer streams ($5.43$ mg/l). Furthermore, the vast majority of colder streams have oxygen levels above 10 mg/l, while the third quartile ($Q_3$) for warmer streams is only $6.12$ mg/l. This means over 75% of warmer streams have lower oxygen concentrations than almost all of the colder streams.

Question

A jewelry company uses a machine to apply a coating of gold on a certain style of necklace. The amount of gold applied to a necklace is approximately normally distributed. When the machine is working properly, the amount of gold applied to a necklace has a mean of 300 milligrams (mg) and standard deviation of 5 mg.
 
(a) A necklace is randomly selected from the necklaces produced by the machine. Assuming that the machine is working properly, calculate the probability that the amount of gold applied to the necklace is between 296 mg and 304 mg.
The jewelry company wants to make sure the machine is working properly. Each day, Cleo, a statistician at the jewelry company, will take a random sample of the necklaces produced that day. Each selected necklace will be melted down and the amount of the gold applied to that necklace will be determined. Because a necklace must be destroyed to determine the amount of gold that was applied, Cleo will use random samples of size $n=2$ necklaces.
Cleo starts by considering the mean amount of gold being applied to the necklaces. After Cleo takes a random sample of $n=2$ necklaces, she computes the sample mean amount of gold applied to the two necklaces.
(b) Suppose the machine is working properly with a population mean amount of gold being applied of 300 mg and a population standard deviation of 5 mg.
(i) Calculate the probability that the sample mean amount of gold applied to a random sample of $n=2$ necklaces will be greater than 303 mg.
(ii) Suppose Cleo took a random sample of $n=2$ necklaces that resulted in a sample mean amount of gold applied of 303 mg. Would that result indicate that the population mean amount of gold being applied by the machine is different from 300 mg? Justify your answer without performing an inference procedure.
Now, Cleo will consider the variation in the amount of gold the machine applies to the necklaces. Because of the small sample size, $n=2$, Cleo will use the sample range of the data for the two randomly selected necklaces, rather than the sample standard deviation.
Cleo will investigate the behavior of the range for samples of size $n=2$. She will simulate the sampling distribution of the range of the amount of gold applied to two randomly sampled necklaces. Cleo generates 100,000 random samples of size $n=2$ independent values from a normal distribution with mean $\mu=300$ and standard deviation $\sigma=5$. The range is calculated for the two observations in each sample. The simulated sampling distribution of the range is shown in Graph I. This process is repeated using $\sigma=8$ as shown in Graph II, and again using $\sigma=12$ as shown in Graph III.
(c) Use the information in the graphs to complete the following.
(i) Describe the sampling distribution of the sample range for random samples of size $n=2$ from a normal distribution with standard deviation $\sigma=5$, as shown in Graph I.
(ii) Describe how the sampling distribution of the sample range for samples of size $n=2$ changes as the value of the population standard deviation increases.
Recall that Cleo needs to consider both the mean and standard deviation of the amount of gold applied to necklaces to determine whether the machine is working properly. Suppose that one month later, Cleo is again checking the machine to make sure it is working properly. Cleo takes a random sample of 2 necklaces and calculates the sample mean amount of gold applied as 303 mg and the sample range as 10 mg.
(d) Recall that the machine is working properly if the amount of gold applied to the necklaces has a mean of 300 mg and standard deviation of 5 mg.
(i) Consider Cleo’s range of 10 mg from the sample of size $n=2$. If the machine is working properly with a standard deviation of 5 mg, is a sample range of 10 mg unusual? Justify your answer.
(ii) Do Cleo’s sample mean of 303 mg and range of 10 mg indicate that the machine is not working properly? Explain your answer.
 

Most-appropriate topic codes (AP Statistics):

• Topic \(1.6\) — Describing the Distribution of One Quantitative Variable (Part \( \mathrm{c} \))
• Topic \(2.11\) — The Normal Distribution (Part \( \mathrm{a} \))
• Topic \(2.12\) — Central Limit Theorem and Sampling Distributions for Sample Means (Parts \( \mathrm{b} \), \( \mathrm{d} \))
▶️ Answer/Explanation

(a)
To find the probability, standardize the given values to z-scores using the formula $z = \frac{x – \mu}{\sigma}$.
$P(296 < X < 304) = P\left(\frac{296-300}{5} < Z < \frac{304-300}{5}\right)$
This simplifies to $P(-0.8 < Z < 0.8) \approx 0.5762$.

(b) (i)
For $n=2$, the standard error of the mean is $\sigma_{\bar{x}} = \frac{5}{\sqrt{2}} \approx 3.535$.
$P(\bar{X} > 303) = P\left(Z > \frac{303-300}{3.535}\right) = P(Z > 0.849)$
$P(\bar{X} > 303) \approx 0.198$.

(b) (ii)
No, this result would not indicate the machine is malfunctioning. Because a sample mean of $303$ mg or higher has a probability of approximately $0.198$ (about $20\%$) assuming the machine is working properly, this result is fairly common and not unusual.

(c) (i)
The sampling distribution of the sample range is heavily right-skewed. The center is around a sample range of $4$ to $5$ mg, and the values vary from $0$ mg up to approximately $25$ mg.

(c) (ii)
As the population standard deviation increases, the center of the sampling distribution shifts to the right (indicating a larger expected range), and the distribution becomes more spread out, showing greater variability in the possible sample ranges.

(d) (i)
No, a sample range of $10$ mg is not unusual. Looking at Graph I (where $\sigma=5$), the bars at and to the right of $10$ mg make up a substantial portion of the total area (well over $5\%$), meaning a range of $10$ mg or more occurs quite frequently by chance.

(d) (ii)
No, these results do not indicate a problem. Based on part (b), a sample mean of $303$ mg is not unusual (happens about $20\%$ of the time), and based on part (d)(i), a sample range of $10$ mg is also quite typical. Since neither metric is statistically surprising, there is no convincing evidence to doubt the machine is working properly.

Question

The length of stay in a hospital after receiving a particular treatment is of interest to the patient, the hospital, and insurance providers. Of particular interest are unusually short or long lengths of stay. A random sample of \(50\) patients who received the treatment was selected, and the length of stay, in number of days, was recorded for each patient. The results are summarized in the following table and are shown in the dotplot.
(a) Determine the five-number summary of the distribution of length of stay.
(b) Consider two rules for identifying outliers, method A and method B. Let method A represent the \(1.5 \times IQR\) rule, and let method B represent the \(2\) standard deviations rule.
i. Using method A, determine any data points that are potential outliers in the distribution of length of stay. Justify your answer.
ii. The mean length of stay for the sample is \(7.42\) days with a standard deviation of \(2.37\) days. Using method B, determine any data points that are potential outliers in the distribution of length of stay. Justify your answer.
(c) Explain why method A might identify more data points as potential outliers than method B for a distribution that is strongly skewed to the right.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.6\) — Distribution of One Quantitative Variable (Part \( \mathrm{c} \))
• Topic \(1.7\) — Summary Statistics for One Quantitative Variable (Parts \( \mathrm{a} \), \( \mathrm{b} \))
▶️ Answer/Explanation

(a)
To find the five-number summary, we arrange the data from our sample and extract the necessary percentiles. Based on our \(50\) observations, it naturally breaks down as follows:

\( \text{Minimum} = 5 \text{ days} \)
\( Q_1 = 6 \text{ days} \)
\( \text{Median} = 7 \text{ days} \)
\( Q_3 = 8 \text{ days} \)
\( \text{Maximum} = 21 \text{ days} \)

(b)(i)
For Method A, we rely on the \( 1.5 \times IQR \) rule to flag potential outliers. We first calculate our IQR and resulting boundaries:
\( IQR = Q_3 – Q_1 = 8 – 6 = 2 \)
Lower boundary: \( Q_1 – 1.5 \times IQR = 6 – 1.5(2) = 3 \)
Upper boundary: \( Q_3 + 1.5 \times IQR = 8 + 1.5(2) = 11 \)
Since we have no data points below \( 3 \), there are no lower outliers. However, any values strictly greater than \( 11 \) are upper outliers, which directs us perfectly to the patients who stayed for \( 12 \) days and \( 21 \) days.

(b)(ii)
Method B instead uses the \( 2 \) standard deviations rule, which is centered around our given mean and standard deviation. Let’s map out this interval:
Lower boundary: \( \text{Mean} – 2 \times SD = 7.42 – 2(2.37) = 2.68 \)
Upper boundary: \( \text{Mean} + 2 \times SD = 7.42 + 2(2.37) = 12.16 \)
Scanning our data, there are no lengths of stay below \( 2.68 \). The only length of stay strictly above \( 12.16 \) is the \( 21 \)-day point. So, the only outlier using Method B is the single patient who stayed \( 21 \) days.

(c)
In a distribution that exhibits strong right skewness, non-resistant measures like the mean and standard deviation get pulled heavily toward the extreme values in the long right tail.
This drastic pull inflates the upper outlier boundary for Method B much more than it affects Method A, which is anchored by the highly resistant median and IQR.
Because the threshold for Method B gets stretched out so much further into the right tail, it loses its sensitivity, making it much harder to detect potential outliers compared to Method A.

Question

The sizes, in square feet, of the \(20\) rooms in a student residence hall at a certain university are summarized in the following histogram.
(a) Based on the histogram, write a few sentences describing the distribution of room size in the residence hall.
(b) Summary statistics for the sizes are given in the following table.
Determine whether there are potential outliers in the data. Then use the following grid to sketch a boxplot of room size.
(c) What characteristic of the shape of the distribution of room size is apparent from the histogram but not from the boxplot?

Most-appropriate topic codes (AP Statistics):

• Topic \(1.6\) — Describing the Distribution of One Quantitative Variable (Parts \( \mathrm{a} \), \( \mathrm{c} \))
• Topic \(1.7\) — Summary Statistics for One Quantitative Variable (Part \( \mathrm{b} \))
• Topic \(1.8\) — Graphical Representations of Summary Statistics for One Quantitative Variable (Part \( \mathrm{b} \))
▶️ Answer/Explanation

(a)
The distribution of the sample of room sizes is bimodal and roughly symmetric.
Most room sizes fall into two clusters: \(100\) to \(200\) square feet and \(250\) to \(350\) square feet.
The center of the distribution is between \(200\) and \(300\) square feet.
The range of the distribution is between \(150\) and \(250\) square feet.
There are no apparent outliers.

(b)
The interquartile range is:
\(IQR = Q_3 – Q_1 = 292 – 174 = 118\) square feet.
To find potential outliers, we check the lower and upper fences.
Lower Fence = \(Q_1 – 1.5(IQR) = 174 – 1.5(118) = -3\) square feet.
Upper Fence = \(Q_3 + 1.5(IQR) = 292 + 1.5(118) = 469\) square feet.
There are no potential outliers because the minimum room size of \(134\) square feet does not fall below \(-3\), and the maximum room size of \(315\) square feet does not exceed \(469\).

(c)
The histogram clearly shows the bimodal nature (two peaks or clusters) of the distribution of room sizes.
This bimodal characteristic is not apparent from the boxplot.

Question

The following histograms summarize the teaching year for the teachers at two high schools, A and B.
Teaching year is recorded as an integer, with first-year teachers recorded as \(1\), second-year teachers recorded as \(2\), and so on. Both sets of data have a mean teaching year of \(8.2\), with data recorded from \(200\) teachers at High School A and \(221\) teachers at High School B. On the histograms, each interval represents possible integer values from the left endpoint up to but not including the right endpoint.
(a) The median teaching year for one high school is \(6\), and the median teaching year for the other high school is \(7\). Identify which high school has each median and justify your answer.
(b) An additional \(18\) teachers were not included with the data recorded from the \(200\) teachers at High School A. The mean teaching year of the \(18\) teachers is \(2.5\). What is the mean teaching year for all \(218\) teachers at High School A?
(c) The standard deviation of the teaching year for the \(221\) teachers at High School B is \(7.2\). If one teacher is selected at random from High School B, what is the probability that the teaching year for the selected teacher will be within \(1\) standard deviation of the mean of \(8.2\)? Justify your answer.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.6\) — Describing the Distribution of One Quantitative Variable (Parts \( \mathrm{a} \), \( \mathrm{c} \))
• Topic \(1.7\) — Summary Statistics for One Quantitative Variable (Parts \( \mathrm{a} \), \( \mathrm{b} \), \( \mathrm{c} \))
▶️ Answer/Explanation

(a)
For High School A (\(n = 200\)), the median is the average of the \(100^{\text{th}}\) and \(101^{\text{st}}\) ordered values.

From the histogram, the number of teachers with teaching year in \((1, 7)\) is:
\(46 + 48 = 94 < 100\)
The number with teaching year in \((1, 10)\) is:
\(46 + 48 + 45 = 139 > 100\)
Since there are fewer than \(100\) values below \(7\) but more than \(100\) values below \(10\), both the \(100^{\text{th}}\) and \(101^{\text{st}}\) values fall in the interval \((7, 10)\).
\(\boxed{\text{High School A has median teaching year } = 7}\)

For High School B (\(n = 221\)), the median is the \(111^{\text{th}}\) ordered value.
From the histogram, the number of teachers with teaching year in \((1, 7)\) is:
\(79 + 34 = 113 > 111\)
Since there are more than \(111\) values below \(7\), the \(111^{\text{th}}\) value falls in the interval \((4, 7)\).
\(\boxed{\text{High School B has median teaching year } = 6}\)

Think of the median as a balancing point — you need to find where the “middle person” sits in the ordered list.
For School A, with 94 teachers below year 7 and the median being teacher #100, that teacher has to be somewhere in years 7–9.
For School B, there are already 113 teachers in years 1–6, so the 111th teacher — the median — must be in that same lower group.
The dramatically right-skewed shape of School B also gives a useful intuition: right-skewed distributions pull the mean upward away from the median, and since both schools share the same mean of 8.2, School B (the more skewed one) must have the bigger mean-median gap and therefore the lower median.

(b)
The mean of a combined group is a weighted average — you cannot simply average the two means.
Total teaching years for the original \(200\) teachers: \(\bar{x}_1 \cdot n_1 = 8.2 \times 200 = 1{,}640\)
Total teaching years for the additional \(18\) teachers: \(\bar{x}_2 \cdot n_2 = 2.5 \times 18 = 45\)
Combined mean for all \(218\) teachers: \(\bar{x}_{\text{combined}} = \frac{1{,}640 + 45}{200 + 18} = \frac{1{,}685}{218}\)
\(\boxed{\bar{x}_{\text{combined}} \approx 7.73 \text{ years}}\)
The 18 new teachers have a very low mean of 2.5 years — they are almost all brand-new — so they drag the overall mean down from 8.2 toward 7.73.
Because there are so few of them (only 18 out of 218), the pull is modest but real.
Always weight by group size when combining means; plain-averaging the means (i.e., \((8.2+2.5)/2=5.35\)) would be completely wrong here.

(c)
First, find the interval within \(1\) standard deviation of the mean:
\(\bar{x} \pm s = 8.2 \pm 7.2 \implies (1.0,\ 15.4)\)
Since teaching years are recorded as integers, the values that fall within this interval are \(1, 2, 3, \ldots, 15\) (since \(15 \leq 15.4 < 16\)).
From the High School B histogram, we count the teachers in the five histogram bars that contain years \(1\) through \(15\):

\((1,4):\; 79 \text{ teachers}\)
\((4,7):\; 34 \text{ teachers}\)
\((7,10):\; 28 \text{ teachers}\)
\((10,13):\; 29 \text{ teachers}\)
\((13,16):\; 19 \text{ teachers} \quad \text{(years 13, 14, 15 all} \leq 15.4\text{)}\)
\(\text{Total teachers with year} \in [1,\, 15] = 79 + 34 + 28 + 29 + 19 = 189\)
\(P(\text{within 1 SD of mean}) = \frac{189}{221}\)
\(\boxed{P \approx 0.8552}\)
The critical point here is that you must count directly from the histogram — the Empirical Rule (\(\approx 68\%\)) does not apply because High School B is strongly right-skewed, not bell-shaped.
Applying the Empirical Rule would give the wrong answer. Also note that the entire bar for \([13, 16)\) is included, because all integers in that bar (13, 14, 15) satisfy \(\leq 15.4\); only 16 would fall outside the interval, and there are no teachers listed as exactly year 16 in the \([13,16)\) bar.

Question

Robin works as a server in a small restaurant, where she can earn a tip (extra money) from each customer she serves. The histogram below shows the distribution of her 60 tip amounts for one day of work.
(a) Write a few sentences to describe the distribution of tip amounts for the day shown.
(b) One of the tip amounts was \$8. If the \$8 tip had been \$18, what effect would the increase have had on the following statistics? Justify your answers.
The mean:
The median:

Most-appropriate topic codes (AP Statistics):

• Topic \(1.6\) — Describing the Distribution of One Quantitative Variable (Part \( \mathrm{a} \))
• Topic \(1.7\) — Summary Statistics for One Quantitative Variable (Part \( \mathrm{b} \))
▶️ Answer/Explanation

(a)

The distribution of Robin’s tip amounts is skewed to the right (positively skewed), with the bulk of tips concentrated in the lower dollar amounts.
The center (median) is between \$2.50 and \$5.00, since the 30th and 31st values — which determine the median for \(n = 60\) — both fall in the second bar.
The spread ranges from approximately \$0 to over \$20, with most tips (about 78%) falling between \$0 and \$5.
There is a gap between \$15 and \$20 (no tips recorded in that range), and the single tip above \$20 (in the \$20–\$22.50 interval) appears to be a potential outlier.

(b)

The mean:
Replacing the \$8 tip with a \$18 tip increases the sum of all 60 tip amounts by \$10.
Since the mean is the sum divided by \(n\), the new mean \(= \bar{x}_{\text{old}} + \dfrac{10}{60} = \bar{x}_{\text{old}} + \dfrac{1}{6}\).
\(\boxed{\text{The mean would increase by } \$\tfrac{1}{6} \approx \$0.17 \text{ (about 17 cents).}}\)
The median:
With \(n = 60\) tips, the median is the average of the 30th and 31st values when sorted in order.
The first bar (\$0–\$2.50) contains 25 tips, and the second bar (\$2.50–\$5.00) contains 22 tips — so the 30th and 31st tips are both in the \$2.50–\$5.00 interval, giving a median between \$2.50 and \$5.00.
Since both \$8 and \$18 are greater than the current median, replacing one with the other does not shift any value across the median position.
\(\boxed{\text{The median would not change.}}\)

Question

An environmental group conducted a study to determine whether crows in a certain region were ingesting food containing unhealthy levels of lead. A biologist classified lead levels greater than \(6.0\) parts per million (ppm) as unhealthy. The lead levels of a random sample of \(23\) crows in the region were measured and recorded. The data are shown in the stemplot below.
(a) What proportion of crows in the sample had lead levels that are classified by the biologist as unhealthy?
(b) The mean lead level of the \(23\) crows in the sample was \(4.90\) ppm and the standard deviation was \(1.12\) ppm. Construct and interpret a \(95\) percent confidence interval for the mean lead level of crows in the region.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.5\) — Graphical Representations for One Quantitative Variable (Part \( \mathrm{a} \))
• Topic \(1.6\) — Descriptions for One Quantitative Variable Distributions (Part \( \mathrm{b} \) — checking normality condition)
• Topic \(4.2\) — Constructing a Confidence Interval for a Population Mean or Population Mean Difference (Part \( \mathrm{b} \))
• Topic \(4.3\) — Justifying a Claim Based on a Confidence Interval for a Population Mean or Population Mean Difference (Part \( \mathrm{b} \))
▶️ Answer/Explanation

(a)

From the stemplot, the crows with lead levels greater than \(6.0\) ppm are those with values \(6.3,\ 6.4,\ 6.6,\) and \(6.8\) ppm — that gives us exactly \(4\) crows out of the \(23\) sampled.
$\text{Proportion} = \frac{4}{23} \approx 0.174$
\(\boxed{\dfrac{4}{23} \approx 0.174}\)

(b)

Step 1: Identify the procedure and check conditions.
The appropriate procedure is a one-sample \(t\)-interval for a population mean, using the formula:
$\bar{x} \pm t^* \cdot \frac{s}{\sqrt{n}}$
Condition 1 — Random sample: The problem states the \(23\) crows were randomly selected, so this condition is met.
Condition 2 — Normality: The sample size of \(23\) is not large enough on its own, so we check the stemplot. The data show no strong skewness and no outliers, so it is reasonable to assume the population distribution of lead levels is approximately normal.

Step 2: Compute the confidence interval.
Given: \(\bar{x} = 4.90\) ppm, \(s = 1.12\) ppm, \(n = 23\)
Degrees of freedom: \(df = n – 1 = 22\)
Critical value at \(95\%\) confidence with \(22\) df: \(t^* = 2.074\)
$4.90 \pm 2.074 \times \frac{1.12}{\sqrt{23}}$
$4.90 \pm 2.074 \times 0.2336$
$4.90 \pm 0.484$
$\boxed{(4.416,\ 5.384) \text{ ppm}}$

Step 3: Interpret the interval.
We are \(95\%\) confident that the true mean lead level among all crows in this region is between \(4.416\) ppm and \(5.384\) ppm.

Question

Records are kept by each state in the United States on the number of pupils enrolled in public schools and the number of teachers employed by public schools for each school year. From these records, the ratio of the number of pupils to the number of teachers (P-T ratio) can be calculated for each state. The histograms below show the P-T ratio for every state during the 2001-2002 school year. The histogram on the left displays the ratios for the 24 states that are west of the Mississippi River, and the histogram on the right displays the ratios for the 26 states that are east of the Mississippi River.
(a) Describe how you would use the histograms to estimate the median P-T ratio for each group (west and east) of states. Then use this procedure to estimate the median of the west group and the median of the east group.
(b) Write a few sentences comparing the distributions of P-T ratios for states in the two groups (west and east) during the 2001-2002 school year.
(c) Using your answers in parts (a) and (b), explain how you think the mean P-T ratio during the 2001-2002 school year will compare for the two groups (west and east).

Most-appropriate topic codes (AP Statistics):

• Topic 1.5 — Graphical Representations for One Quantitative Variable (Part a)
• Topic 1.6 — Descriptions for One Quantitative Variable Distributions (Part b)
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Part c)
▶️ Answer/Explanation

(a)
To find the median from a histogram, identify the interval that contains the middle observation by adding the frequencies of each bin from left to right until the cumulative total reaches half of the sample size.
For a group with \(n\) sorted observations, the median position is calculated using the formula \(\frac{n+1}{2}\).
For the West group, there are \(n = 24\) observations, meaning the median is the average of the 12th and 13th values.
By counting cumulative frequencies from the left: the 12-13 bin has 1, the 13-14 bin has 4 (total 5), the 14-15 bin has 6 (total 11), and the 15-16 bin has 3 (total 14).
Since the 12th and 13th observations fall inside the 15-16 interval, the estimated median P-T ratio for the West group is between 15 and 16.
For the East group, there are \(n = 26\) observations, meaning the median is the average of the 13th and 14th values.
Counting cumulative frequencies from the left: the 12-13 bin has 2, the 13-14 bin has 4 (total 6), the 14-15 bin has 4 (total 10), and the 15-16 bin has 11 (total 21).
Since the 13th and 14th observations also fall inside the 15-16 interval, the estimated median P-T ratio for the East group is also between 15 and 16.
\(\boxed{\text{West Median: } [15, 16], \text{ East Median: } [15, 16]}\)
From the histogram, cumulative frequencies for the two groups are shown in the table below.

Thus, the median P-T ratio for both groups is at least 15 students per teacher and at most 16 students per teacher.

(b)
Shape: The distribution of P-T ratios for the West group is unimodal and skewed to the right, whereas the distribution for the East group is unimodal and approximately symmetric.
Center: The centers are nearly identical, with both the West and East groups having a median P-T ratio located in the interval between 15 and 16 students per teacher.
Spread: There is visibly more variability in the P-T ratios for the West group than for the East group. The maximum possible range for the West group is \(22 – 12 = 10\), which is larger than the maximum possible range for the East group, which is \(19 – 12 = 7\).

(c)
The mean P-T ratio for the West group will likely be greater than the mean P-T ratio for the East group.
As established in part (a), both distributions have approximately the same median value between 15 and 16.
Because the distribution for the West group is strongly skewed to the right, its mean will be pulled upward to a value greater than its median.
In contrast, because the distribution for the East group is roughly symmetric, its mean will remain close to its median value.
Therefore, the mean of the West group will be higher than the mean of the East group.
\(\boxed{\text{Mean}_{\text{West}} > \text{Mean}_{\text{East}}}\)

Question

As a part of the United States Department of Agriculture’s Super Dump cleanup efforts in the early 1990s, various sites in the country were targeted for cleanup. Three of the targeted sites—River X, River Y, and River Z—had become contaminated with pesticides because they were located near abandoned pesticide dump sites. Measurements of the concentration of aldrin (a commonly used pesticide) were taken at twenty randomly selected locations in each river near the dump sites. The boxplots shown below display the five-number summaries for the concentrations, in parts per million (ppm) of aldrin, for the twenty locations that were sampled in each of the three rivers.
(a) Compare the distributions of the concentration of aldrin among the three rivers.
(b) The twenty concentrations of aldrin for River X are given below.
\(3.4\quad 4.0\quad 5.6\quad 3.7\quad 8.0\quad 5.5\quad 5.3\quad 4.2\quad 4.3\quad 7.3\)
\(8.6\quad 5.1\quad 8.7\quad 4.6\quad 7.5\quad 5.3\quad 8.2\quad 4.7\quad 4.8\quad 4.6\)
Construct a stemplot that displays the concentrations of aldrin for River X.
(c) Describe a characteristic of the distribution of aldrin concentrations in River X that can be seen in the stemplot but cannot be seen in the boxplot.
 

Most-appropriate topic codes (AP Statistics):

• Topic \(1.9\) — Comparisons of the Distributions for One Quantitative Variable (Part \( \mathrm{a} \))
• Topic \(1.5\) — Graphical Representations for One Quantitative Variable (Part \( \mathrm{b} \))
• Topic \(1.6\) — Descriptions for One Quantitative Variable Distributions (Part \( \mathrm{c} \))
▶️ Answer/Explanation

(a)
Comparing the medians reveals that the concentration of aldrin tends to be highest for River X and lowest for River Z.
About \(50\%\) of the concentrations of aldrin for Rivers X and Y are higher than all of the concentrations for River Z.
River X also displays the most variability in aldrin concentrations, as seen by the largest range and largest IQR, and River Z has the least variability, as judged by both IQR and range.
The shapes of the three distributions differ, in that the distribution appears to be skewed to the right for River X, roughly symmetric for River Y and slightly skewed to the left for River Z.

(b)
Aldrin concentrations (in ppm) for River X
Leaf unit \(= 0.1\) (for example, \(3 \mid 4\) represents \(3.4\,\text{ppm}\))
\(3 \mid 4\ 7\)
\(4 \mid 0\ 2\ 3\ 6\ 6\ 7\ 8\)
\(5 \mid 1\ 3\ 3\ 5\ 6\)
\(6 \mid \)
\(7 \mid 3\ 5\)
\(8 \mid 0\ 2\ 6\ 7\)

(c)
The stemplot shows a clear gap in the distribution of aldrin concentrations for River X.
This gap occurs between the values of \(5.6\) and \(7.3\,\text{ppm}\) of aldrin.
This gap is not apparent in the boxplot.

Question

A consumer organization was concerned that an automobile manufacturer was misleading customers by overstating the average fuel efficiency (measured in miles per gallon, or mpg) of a particular car model. The model was advertised to get 27 mpg. To investigate, researchers selected a random sample of 10 cars of that model. Each car was then randomly assigned a different driver. Each car was driven for 5,000 miles, and the total fuel consumption was used to compute mpg for that car.
(a) Define the parameter of interest and state the null and alternative hypotheses the consumer organization is interested in testing.
One condition for conducting a one-sample \(t\)-test in this situation is that the mpg measurements for the population of cars of this model should be normally distributed. However, the boxplot and histogram shown below indicate that the distribution of the 10 sample values is skewed to the right.

(b) One possible statistic that measures skewness is the ratio \(\dfrac{\text{sample mean}}{\text{sample median}}\). What values of that statistic (small, large, close to one) might indicate that the population distribution of mpg values is skewed to the right? Explain.
(c) Even though the mpg values in the sample were skewed to the right, it is still possible that the population distribution of mpg values is normally distributed and that the skewness was due to sampling variability. To investigate, 100 samples, each of size 10, were taken from a normal distribution with the same mean and standard deviation as the original sample. For each of those 100 samples, the statistic \(\dfrac{\text{sample mean}}{\text{sample median}}\) was calculated. A dotplot of the 100 simulated statistics is shown below.
In the original sample, the value of the statistic \(\dfrac{\text{sample mean}}{\text{sample median}}\) was 1.03. Based on the value of 1.03 and the dotplot above, is it plausible that the original sample of 10 cars came from a normal population, or do the simulated results suggest the original population is really skewed to the right? Explain.
(d) The table below shows summary statistics for mpg measurements for the original sample of 10 cars.

Choosing only from the summary statistics in the table, define a formula for a different statistic that measures skewness.
What values of that statistic might indicate that the distribution is skewed to the right? Explain.

Most-appropriate topic codes (AP Statistics):

• Topic \(4.4\) — Setting Up a Test for a Population Mean or Population Mean Difference (Part \(\mathrm{a}\))
• Topic \(1.6\) — Descriptions for One Quantitative Variable Distributions (Part \(\mathrm{b}\))
• Topic \(2.3\) — Estimating Probabilities Using Simulation (Part \(\mathrm{c}\))
• Topic \(1.8\) — Graphical Representations of Summary Statistics for One Quantitative Variable (Part \(\mathrm{d}\))

▶️ Answer/Explanation

(a)
Let \(\mu\) = the true population mean fuel efficiency (in miles per gallon) for all cars of this particular model.
The consumer organization suspects the manufacturer is overstating the mpg, so the alternative hypothesis is lower-tailed:
\( H_0:\ \mu = 27\ \text{mpg} \)
\( H_a:\ \mu < 27\ \text{mpg} \)
The parameter must be defined as a population mean — not just a sample mean — and must be stated in context (fuel efficiency of this car model) to receive full credit.

(b)
Large values (greater than 1) of the ratio \(\dfrac{\text{sample mean}}{\text{sample median}}\) would indicate that the population distribution is skewed to the right.
The reason is that in a right-skewed distribution, the few unusually large values in the upper tail pull the mean upward, but the median — being a positional measure — is resistant to those extreme values and stays lower.
Therefore, when the distribution is right-skewed, we expect:
\( \text{sample mean} > \text{sample median} \implies \dfrac{\text{sample mean}}{\text{sample median}} > 1 \)
The further the ratio exceeds 1, the stronger the evidence of right-skewness.

(c)
The observed value of the statistic from the original sample is \(1.03\).
Looking at the dotplot of the 100 simulated statistics (all drawn from a normal population), we count that 14 out of 100 simulated values are greater than or equal to \(1.03\).
This gives a simulated \(p\)-value of approximately:
\( \hat{p} = \dfrac{14}{100} = 0.14 \)
Since \(0.14\) is larger than any commonly used significance level (such as \(\alpha = 0.05\) or \(\alpha = 0.10\)), we do not have convincing evidence that the original population is skewed to the right.
It is therefore plausible that the original sample of 10 cars came from a normally distributed population, and the observed right-skewness in the sample was simply due to random sampling variability.

(d)
Using only the five-number summary values, one reasonable skewness statistic is:
\( S = \dfrac{Q_3 – \text{Median}}{\text{Median} – Q_1} \)
Using the given data:
\( S = \dfrac{28 – 25.5}{25.5 – 24} = \dfrac{2.5}{1.5} \approx 1.67 \)
Values greater than 1 indicate right-skewness.
The reasoning is: in a right-skewed distribution, the data in the upper half are more spread out than in the lower half, so the distance from the median up to \(Q_3\) (the upper half of the middle 50%) will be larger than the distance from \(Q_1\) down to the median (the lower half of the middle 50%). This makes the numerator larger than the denominator, giving a ratio greater than 1.
Other acceptable statistics using only the five-number summary include:
\( S = \dfrac{\text{Maximum} – \text{Median}}{\text{Median} – \text{Minimum}}, \qquad S = \dfrac{\text{Maximum} – Q_3}{Q_1 – \text{Minimum}}, \qquad S = \dfrac{\frac{Q_1 + Q_3}{2}}{\text{Median}} \)
For all of these, values greater than 1 indicate right-skewness.

Question

A local arcade is hosting a tournament in which contestants play an arcade game with possible scores ranging from 0 to 20. The arcade has set up multiple game tables so that all contestants can play the game at the same time; thus contestant scores are independent. Each contestant’s score will be recorded as he or she finishes, and the contestant with the highest score is the winner.
After practicing the game many times, Josephine, one of the contestants, has established the probability distribution of her scores, shown in the table below.
Crystal, another contestant, has also practiced many times. The probability distribution for her scores is shown in the table below.
(a) Calculate the expected score for each player.
(b) Suppose that Josephine scores 16 and Crystal scores 17. The difference (Josephine minus Crystal) of their scores is \(-1\). List all combinations of possible scores for Josephine and Crystal that will produce a difference (Josephine minus Crystal) of \(-1\), and calculate the probability for each combination.
(c) Find the probability that the difference (Josephine minus Crystal) in their scores is \(-1\).
(d) The table below lists all the possible differences in the scores between Josephine and Crystal and some associated probabilities.
Complete the table and calculate the probability that Crystal’s score will be higher than Josephine’s score.

Most-appropriate topic codes (AP Statistics):

• Topic \(2.8\) — Introduction to Random Variables and Probability Distributions (Parts \(\mathrm{a}\), \(\mathrm{b}\), \(\mathrm{c}\), \(\mathrm{d}\))
• Topic \(2.9\) — Parameters of Random Variables (Part \(\mathrm{a}\))
• Topic \(2.7\) — Independent Events and Unions of Events (Parts \(\mathrm{b}\), \(\mathrm{c}\), \(\mathrm{d}\))
▶️ Answer/Explanation

(a)

The expected value (mean) of a discrete random variable is \(\mu = \sum x_i \cdot P(x_i)\). We apply this to each player’s distribution.
For Josephine:
\(\mu_J = 16(0.10) + 17(0.30) + 18(0.40) + 19(0.20)\)
\(\mu_J = 1.6 + 5.1 + 7.2 + 3.8 = \boxed{17.7}\)
For Crystal:
\(\mu_C = 17(0.45) + 18(0.40) + 19(0.15)\)
\(\mu_C = 7.65 + 7.20 + 2.85 = \boxed{17.7}\)
Both players have the same expected score of \(17.7\).

(b)

We need all pairs \((J, C)\) such that \(J – C = -1\), i.e., Josephine’s score is exactly 1 less than Crystal’s. Since scores are independent, the probability of each pair is the product of the individual probabilities.
\(J = 16,\ C = 17\): \(\quad P = (0.10)(0.45) = 0.045\)
\(J = 17,\ C = 18\): \(\quad P = (0.30)(0.40) = 0.120\)
\(J = 18,\ C = 19\): \(\quad P = (0.40)(0.15) = 0.060\)
These are the only three combinations that produce a difference of \(-1\).

(c)

Since the three combinations in part (b) are mutually exclusive, we simply add their probabilities:
\(P(J – C = -1) = 0.045 + 0.120 + 0.060\)
\(\boxed{P(J – C = -1) = 0.225}\)

(d)

First, we find the missing probability for a difference of \(-2\). Since all probabilities in the distribution must sum to 1:
\(P(J – C = -2) = 1 – 0.015 – 0.225 – 0.325 – 0.260 – 0.090\)
\(P(J – C = -2) = 1 – 0.915 = \boxed{0.085}\)
The completed distribution table is:


Crystal’s score is higher than Josephine’s when the difference \(J – C < 0\), i.e., when the difference is \(-3\), \(-2\), or \(-1\).
\(P(\text{Crystal} > \text{Josephine}) = P(J – C < 0) = 0.015 + 0.085 + 0.225\)
\(\boxed{P(\text{Crystal} > \text{Josephine}) = 0.325}\)

Question

The Better Business Council of a large city has concluded that students in the city’s schools are not learning enough about economics to function in the modern world. These findings were based on test results from a random sample of \(20\) twelfth-grade students who completed a \(46\)-question multiple-choice test on basic economic concepts. The data set below shows the number of questions that each of the \(20\) students in the sample answered correctly.
12 16 18 17 18 33 41 44 38 35
19 36 19 13 43 8 16 14 10 9
(a) Display these data in a stemplot.
(b) Use your stemplot from part (a) to describe the main features of this score distribution.
(c) Why would it be misleading to report only a measure of center for this score distribution?

Most-appropriate topic codes (AP Statistics):

• Topic 1.5 — Graphical Representations for One Quantitative Variable (Part \(\mathrm{a}\))
• Topic 1.6 — Descriptions for One Quantitative Variable Distributions (Part \(\mathrm{b}\))
• Topic 1.7 — Summary Statistics for One Quantitative Variable (Part \(\mathrm{c}\))
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Parts \(\mathrm{b}\), \(\mathrm{c}\))
▶️ Answer/Explanation

(a)
To build the stemplot, use the tens digit as the stem and the units digit as the leaf. The data range from \(8\) to \(44\), so the stems are \(0, 1, 2, 3, 4\). Place each data value’s units digit on the appropriate row.
\( \begin{array}{r|l} 0 & 8 \quad 9 \\ 1 & 2 \quad 6 \quad 8 \quad 7 \quad 8 \quad 9 \quad 6 \quad 9 \quad 3 \quad 4 \quad 0 \\ 2 & \\ 3 & 3 \quad 8 \quad 6 \quad 5 \\ 4 & 1 \quad 4 \quad 3 \\ \end{array} \)
Legend: \(1 \mid 2\) represents \(12\) questions answered correctly.
Ordered version (leaves sorted in ascending order):
\( \begin{array}{r|l} 0 & 8 \quad 9 \\ 1 & 0 \quad 2 \quad 3 \quad 4 \quad 6 \quad 6 \quad 7 \quad 8 \quad 8 \quad 9 \quad 9 \\ 2 & \\ 3 & 3 \quad 5 \quad 6 \quad 8 \\ 4 & 1 \quad 3 \quad 4 \\ \end{array} \)
Legend: \(1 \mid 2\) represents \(12\) questions answered correctly.

(b)
The most striking feature of this distribution is that the scores split into two distinct groups — one cluster concentrated in the teens (roughly \(8\)–\(19\)) and a second, smaller cluster in the high \(30\)s and low \(40\)s. There is a complete gap in the \(20\)s, with no student scoring between \(19\) and \(33\). The distribution is bimodal, suggesting that students tended to perform either quite poorly or quite well on the test, with no middle ground.

(c)
A measure of center such as the mean or median would fall somewhere around \(\bar{x} \approx 22.95\), which lies right in the gap of the \(20\)s — a region where no student actually scored. Reporting only the center gives the false impression that a typical student scored around \(23\), when in reality no student did. Because the data are bimodal with two separated groups, the center alone completely misses the two-cluster structure of the distribution and would mislead the reader into thinking the scores are roughly symmetric or unimodal.

Question

Two parents have each built a toy catapult for use in a game at an elementary school fair. To play the game, students will attempt to launch Ping-Pong balls from the catapults so that the balls land within a 5-centimeter band. A target line will be drawn through the middle of the band, as shown in the figure below. All points on the target line are equidistant from the launching location.
If a ball lands within the shaded band, the student will win a prize.
The parents have constructed the two catapults according to slightly different plans. They want to test these catapults before building additional ones. Under identical conditions, the parents launch 40 Ping-Pong balls from each catapult and measure the distance that the ball travels before landing. Distances to the nearest centimeter are graphed in the dotplots below.

(a) Comment on any similarities and any differences in the two distributions of distances traveled by balls launched from catapult A and catapult B.
(b) If the parents want to maximize the probability of having the Ping-Pong balls land within the band, which one of the two catapults, A or B, would be better to use than the other? Justify your choice.
(c) Using the catapult that you chose in part (b), how many centimeters from the target line should this catapult be placed? Explain why you chose this distance.

Most-appropriate topic codes (AP Statistics):

• Topic 1.6 — Descriptions for One Quantitative Variable Distributions (Part \(\mathrm{a}\))
• Topic 1.7 — Summary Statistics for One Quantitative Variable (Parts \(\mathrm{a}\), \(\mathrm{c}\))
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Parts \(\mathrm{a}\), \(\mathrm{b}\))
▶️ Answer/Explanation

(a)

Both distributions of distances are roughly symmetric and somewhat mound-shaped (bell-shaped). Looking at the centers, the median of Catapult A is approximately \(136\,\text{cm}\), which is slightly lower than the median of Catapult B at approximately \(138\,\text{cm}\). In terms of spread, Catapult A shows considerably more variability than Catapult B — the range of Catapult A is about \(30\,\text{cm}\), while the range of Catapult B is approximately \(11\,\text{cm}\). Additionally, there appear to be potential outliers in Catapult A’s distribution (e.g., a ball traveling approximately \(155\,\text{cm}\)), whereas Catapult B has no such extreme values.

(b)

Catapult B would be the better choice.

Since the target band is only \(5\,\text{cm}\) wide, the key factor is how tightly clustered the distances are around the center. Catapult B has a much smaller spread — most balls land between approximately \(133\,\text{cm}\) and \(143\,\text{cm}\) — meaning when placed correctly, a higher proportion of balls will fall within the narrow band. Catapult A’s larger variability means balls are scattered over a much wider range, making it far less likely that they will land consistently within the band.

(c)

Catapult B should be placed approximately \(\boxed{138\,\text{cm}}\) from the target line.

Since Catapult B’s distribution is roughly symmetric and mound-shaped, the median (approximately \(138\,\text{cm}\)) is a reliable measure of center and represents the most typical distance a ball will travel. Placing the catapult so that the target line is \(138\,\text{cm}\) away aligns the center of the distribution with the target, maximizing the chance that any given ball lands within the \(5\,\text{cm}\) band. Based on the sample data, approximately \(\frac{30}{40} = 0.75\) of the 40 balls launched from Catapult B landed within \(2.5\,\text{cm}\) on either side of \(138\,\text{cm}\), which further confirms this placement.

Question

The graph below displays the scores of 32 students on a recent exam. Scores on this exam ranged from 64 to 95 points.
\(6\phantom{|}\) * *
\(6\phantom{|}\) * *
\(7\phantom{|}\) * * *
\(7\phantom{|}\) * * * *
\(8\phantom{|}\) * * * *
\(8\phantom{|}\) * * * * * *
\(9\phantom{|}\) * * * * * * *
\(9\phantom{|}\) * * * *
(a) Describe the shape of this distribution.
(b) In order to motivate her students, the instructor of the class wants to report that, overall, the class’s performance on the exam was high. Which summary statistic, the mean or the median, should the instructor use to report that overall exam performance was high? Explain.
(c) The midrange is defined as \(\dfrac{\text{maximum} + \text{minimum}}{2}\). Compute this value using the data on the preceding page.
Is the midrange considered a measure of center or a measure of spread? Explain.

Most-appropriate topic codes (AP Statistics):

• Topic 1.6 — Descriptions for One Quantitative Variable Distributions (Part a)
• Topic 1.7 — Summary Statistics for One Quantitative Variable (Parts b, c)
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Part b)
▶️ Answer/Explanation

(a)
The distribution is skewed to the left (skewed toward the lower values). You can see this from the stemplot: the longer tail stretches down into the 60s and lower 70s, while most of the data cluster in the upper 80s and 90s. There are relatively few low scores pulling the tail downward.

(b)
The instructor should report the median.
Because the distribution is skewed toward the lower values, the mean gets pulled in that direction — it will be lower than the median. The median, being resistant to the few very low scores, will better represent the “typical” high performance of the class. So ironically, to make performance look as high as possible, the median is the better choice here.

(c)
Step 1 — Compute the midrange:
The minimum score is \(64\) and the maximum score is \(95\), so:
\( \text{midrange} = \frac{\text{maximum} + \text{minimum}}{2} = \frac{95 + 64}{2} = \frac{159}{2} = 79.5 \)
\(\boxed{\text{midrange} = 79.5}\)

Step 2 — Identify as measure of center:
The midrange is a measure of center.

Step 3 — Rationale:
The maximum value tells us about the upper extreme and the minimum tells us about the lower extreme. By averaging these two values, we find the point that sits exactly halfway between the two extremes — this is a center point, not a spread. Measures of spread (like range, standard deviation, or IQR) describe how far apart the data are; the midrange instead gives a single representative middle value, placing it firmly in the category of measures of center.

Scroll to Top