Home / AP® Exam / AP® Statistics / AP Statistics 1.11 Random Sampling- Exam Style Questions – FRQs

AP Statistics 1.11 Random Sampling- Exam Style Questions - FRQs - New Syllabus

Question

Aphids are tiny insects that feed on plants such as cabbage plants. A farmer wants to reduce the number of aphids in a cabbage field. A river is located 100 meters south of the cabbage field. The farmer divides the field into 25 regions of equal size, as shown in the diagram. Each region has approximately the same number of cabbage plants.
The farmer would like to estimate the proportion of cabbage plants in the field that are affected by aphids and believes that the extent of aphid damage is greater for the regions in the cabbage field closer to the river. To obtain the estimate, the farmer is considering three sampling methods.
  • Sampling method I: Select region 3, which is closest to the farmer’s house and farthest from the river. Examine every cabbage plant in the region for aphid damage.
  • Sampling method II: Randomly select one row (A, B, C, D, or E). For every region in the selected row, examine every cabbage plant for aphid damage.
  • Sampling method III: Randomly select one region from each of rows A, B, C, D, and E. For each selected region, examine every cabbage plant for aphid damage.
A. Explain whether sampling method I is an appropriate sampling method for the farmer to use to estimate the proportion of cabbage plants in the field that are damaged by aphids.
B. Using sampling method II, the farmer randomly selected row E and examined every cabbage plant in row E. If the farmer’s belief is correct, determine whether the selection of row E is likely to provide an overestimate or an underestimate of the proportion of cabbage plants in the field that are damaged by aphids. Justify your answer.
C. Using the information provided in the diagram of the cabbage field, describe how to implement sampling method III, which requires a random selection of one region from each of rows A, B, C, D, and E.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.11\) — Random Sampling (Parts \( \mathrm{A} \), \( \mathrm{B} \), \( \mathrm{C} \))
• Topic \(1.12\) — Potential Problems with Sampling (Parts \( \mathrm{A} \), \( \mathrm{B} \))
▶️ Answer/Explanation

A.
• Sampling method I is not an appropriate sampling method because it uses a convenience sample that lacks randomization.
• Since region 3 is located farthest from the river where aphid damage is believed to be lowest, this region is not representative of the whole field and will likely lead to an underestimate of the true population proportion.

B.
• The selection of row E is likely to provide an overestimate of the true proportion of damaged cabbage plants.
• Row E is the row positioned closest to the river, meaning every single plant checked in this sample belongs to the high-risk zone where the farmer expects aphid damage to be at its peak concentration.

C.
• Label the 5 individual regions within row A with unique identifiers from 1 to 5.
• Use a random number generator or draw numbered slips from a hat to select one single region from row A.
• Repeat this exact independent drawing process for row B (regions 6 to 10), row C (regions 11 to 15), row D (regions 16 to 20), and row E (regions 21 to 25) to complete a stratified sample containing exactly 5 distinct regions.

Question

Researchers will conduct a year-long investigation of walking and cholesterol levels in adults. They will select a random sample of \(100\) adults from the target population to participate as subjects in the study.
(a) One aspect of the study is to record the number of miles each subject walks per day. The researchers are deciding whether to have subjects wear an activity tracker to record the data or to have subjects keep a daily journal of the miles they walk each day. Describe what bias could be introduced by keeping the daily journal instead of wearing the activity tracker.
During the course of the study, the subjects will have their cholesterol levels measured each month by a doctor. The researchers will perform a significance test at the end of the study to determine whether the average cholesterol level for subjects who walk fewer miles each day is greater than for those who walk more miles each day.
(b) Selecting a random sample creates a reasonable representative sample of the target population. Explain the benefit of using a representative sample from the population.
(c) Suppose the researchers conduct the test and find a statistically significant result. Would it be valid to claim that increased walking causes a decrease in average cholesterol levels for adults in the target population? Explain your reasoning.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.11\) — Random Sampling (Parts \( \mathrm{b} \), \( \mathrm{c} \))
• Topic \(1.12\) — Potential Problems with Sampling (Part \( \mathrm{a} \))
▶️ Answer/Explanation

(a)
Keeping a daily journal introduces response bias because subjects are self-reporting. They might forget to log short walks or underestimate their distances, which would cause our sample’s average daily miles to be systematically biased too low compared to what they actually walked.

(b)
A representative sample is essential because it allows us to generalize our findings. By selecting randomly, our \(100\) adults look like the larger target population, meaning any inference we make about the relationship between walking and cholesterol isn’t skewed by an unrepresentative group.

(c)
No, we can’t claim causation here because this is an observational study, not an experiment. We didn’t randomly assign subjects to walk specific distances. Because of this lack of assignment, there could be confounding variables at play—like a person’s diet or general health consciousness—that simultaneously affect both how much they walk and their cholesterol levels.

Question

Emma is moving to a large city and is investigating typical monthly rental prices of available one-bedroom apartments. She obtained a random sample of rental prices for \(50\) one-bedroom apartments taken from a Web site where people voluntarily list available apartments.
(a) Describe the population for which it is appropriate for Emma to generalize the results from her sample.
The distribution of the \(50\) rental prices of the available apartments is shown in the following histogram.
(b) Emma wants to estimate the typical rental price of a one-bedroom apartment in the city. Based on the distribution shown, what is a disadvantage of using the mean rather than the median as an estimate of the typical rental price?
(c) Instead of using the sample median as the point estimate for the population median, Emma wants to use an interval estimate. However, computing an interval estimate requires knowing the sampling distribution of the sample median for samples of size \(50\). Emma has one point, her sample median, in that sampling distribution.
Using information about rental prices that are available on the Web site, describe how someone could develop a theoretical sampling distribution of the sample median for samples of size \(50\).
Because Emma does not have the resources to develop the theoretical sampling distribution, she estimates the sampling distribution of the sample median using a process called bootstrapping. In the bootstrapping process, a computer program performs the following steps.
• Take a random sample, with replacement, of size \(50\) from the original sample.
• Calculate and record the median of the sample.
• Repeat the process to obtain a total of \(15,000\) medians.
Emma ran the bootstrap process, and the following frequency table is the bootstrap distribution showing her results of generating \(15,000\) medians.
The bootstrap distribution provides an approximation of the sampling distribution of the sample median. A confidence interval for the median can be constructed using a percentage of the values in the middle of the bootstrap distribution.
(d) Use the frequency table to find the following.
i. Value of the \(5\text{th}\) percentile:
ii. Value of the \(95\text{th}\) percentile:
(e) Find the percentage of bootstrap medians in the table that are equal to or between the values found in part (d).
(f) Use your values from parts (d) and (e) to construct and interpret a confidence interval for the median rental price.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.11\) — Random Sampling (Part \( \mathrm{a} \))
• Topic \(1.12\) — Potential Problems with Sampling (Part \( \mathrm{a} \))
• Topic \(1.7\) — Summary Statistics for One Quantitative Variable (Parts \( \mathrm{b} \), \( \mathrm{d} \), \( \mathrm{e} \))
• Topic \(2.12\) — Central Limit Theorem and Sampling Distributions for Sample Means (Parts \( \mathrm{c} \), \( \mathrm{f} \))
▶️ Answer/Explanation

(a)
Because a random sample was used, it is appropriate to generalize the results to the population of all rental prices for one-bedroom apartments in the city that are listed on this particular website at the time the sample was taken.

(b)
The histogram indicates that the distribution of rental prices is strongly skewed to the right.
Because of this right skewness, the sample mean will be pulled substantially higher by a few very large rental prices.
Consequently, the sample mean would overestimate the typical rental price, whereas the sample median provides a more accurate representation of the center.

(c)
To develop a theoretical sampling distribution, Emma would need to obtain every possible sample of size \(50\) from the population of apartments listed on the website.
She would then compute the median rental price for each of these possible samples.
The entire collection of all these computed sample medians forms the theoretical sampling distribution for the sample median.

(d)
i. The \(5\text{th}\) percentile corresponds to the \((0.05)(15,000) = 750\text{th}\) value in the ordered data.
By keeping a running cumulative total of the frequencies down the columns, the \(750\text{th}\) value first falls in the row for the median of \(\$2,500\).
ii. The \(95\text{th}\) percentile corresponds to the \((0.95)(15,000) = 14,250\text{th}\) value in the ordered data.
Continuing to add the frequencies, the \(14,250\text{th}\) value falls in the row for the median of \(\$2,950\).

(e)
We need to sum the frequencies of all bootstrap medians from \(\$2,500\) to \(\$2,950\) inclusive.
\(\text{Sum} = 1,899 + 2 + \dots + 10 + 700 = 14,404\)
\(\text{Percentage} = \left(\dfrac{14,404}{15,000}\right) \times 100\% \approx 96.03\%\)

(f)
Based on our previous parts, an approximate \(96\%\) confidence interval for the median is \((\$2,500, \$2,950)\).
Interpretation: We are approximately \(96\%\) confident that the true median rental price of all one-bedroom apartments listed on this website for this city is between \(\$2,500\) and \(\$2,950\).

Question

Alzheimer’s disease results in a loss of cognitive ability beyond what is expected with typical aging. A local newspaper published an article with the following headline.
Study Finds Strong Association Between Smoking and Alzheimer’s
The article reported that a study tracked the medical histories of 21,123 men and women for 23 years. The article stated that, for those who smoked at least two packs of cigarettes a day, the risk of developing Alzheimer’s disease was 2.57 times the risk for those who did not smoke.
(a) Identify the explanatory and response variables in the study.
Explanatory variable:
Response variable:
(b) Is the study described in the article an observational study or an experiment? Explain.
(c) Exercise status (regular weekly exercise versus no regular weekly exercise) was mentioned in the article as a possible confounding variable. Explain how exercise status could be a confounding variable in the study.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.10\) — The Investigative Question Revisited and Data Collection (Parts \( \mathrm{a} \), \( \mathrm{b} \))
• Topic \(1.11\) — Random Smapling (Part \( \mathrm{b} \))
• Topic \(1.12\) — Potential Problems with Sampling (Part \( \mathrm{c} \))
▶️ Answer/Explanation

(a)

Explanatory variable: The person’s degree of cigarette smoking — specifically, whether or not the individual smoked at least two packs of cigarettes per day (smoker of at least two packs per day versus non-smoker at that level).
Response variable: Whether or not the person developed Alzheimer’s disease during the course of the 23-year study.

(b)

This is an observational study, not an experiment. In an experiment, researchers would have to actively assign the treatment — in this case, the level of cigarette smoking — to the participants. Instead, the researchers simply tracked the existing medical histories of 21,123 men and women over 23 years. The smoking status of each person was passively observed and recorded, not controlled or manipulated by the researchers. Since no treatment was imposed, this is an observational study.

(c)

A confounding variable is one that is related to the explanatory variable and also independently influences the response variable, making it difficult to determine whether the explanatory variable alone is responsible for the observed association.
Exercise status could be a confounding variable here for two reasons working together:
First, people who exercise regularly tend to be more health-conscious overall, and as a result are less likely to smoke heavily. So exercise status is related to the explanatory variable — smoking status. Heavy smokers are, on average, less likely to exercise regularly than non-smokers.
Second, regular exercise may independently reduce the risk of developing Alzheimer’s disease. So exercise status is also related to the response variable — development of Alzheimer’s disease.
Because of both of these relationships, the observed association between heavy smoking and higher rates of Alzheimer’s could be at least partly explained by the fact that heavy smokers tend to exercise less — and it is the lack of exercise, not the smoking itself, that contributes to the increased risk. This makes it impossible to determine from this study alone whether the association between smoking and Alzheimer’s reflects a true causal relationship or is merely the result of exercise status being linked to both.

Question

Corn tortillas are made at a large facility that produces \(100,000\) tortillas per day on each of its two production lines. The distribution of the diameters of the tortillas produced on production line A is approximately normal with mean \(5.9\) inches, and the distribution of the diameters of the tortillas produced on production line B is approximately normal with mean \(6.1\) inches. The figure below shows the distributions of diameters for the two production lines.
The tortillas produced at the factory are advertised as having a diameter of \(6\) inches. For the purpose of quality control, a sample of \(200\) tortillas is selected and the diameters are measured. From the sample of \(200\) tortillas, the manager of the facility wants to estimate the mean diameter, in inches, of the \(200,000\) tortillas produced on a given day.
Two sampling methods have been proposed.
Method 1: Take a random sample of \(200\) tortillas from the \(200,000\) tortillas produced on a given day. Measure the diameter of each selected tortilla.
Method 2: Randomly select one of the two production lines on a given day. Take a random sample of \(200\) tortillas from the \(100,000\) tortillas produced by the selected production line. Measure the diameter of each selected tortilla.
(a) Will a sample obtained using Method 2 be representative of the population of all tortillas made that day, with respect to the diameters of the tortillas? Explain why or why not.
(b) The figure below is a histogram of \(200\) diameters obtained by using one of the two sampling methods described. Considering the shape of the histogram, explain which method, Method 1 or Method 2, was most likely used to obtain a such a sample.
(c) Which of the two sampling methods, Method 1 or Method 2, will result in less variability in the diameters of the \(200\) tortillas in the sample on a given day? Explain.
Each day, the distribution of the \(200,000\) tortillas made that day has mean diameter \(6\) inches with standard deviation \(0.11\) inch.
(d) For samples of size \(200\) taken from one day’s production, describe the sampling distribution of the sample mean diameter for samples that are obtained using Method 1.
(e) Suppose that one of the two sampling methods will be selected and used every day for one year (\(365\) days). The sample mean of the \(200\) diameters will be recorded each day. Which of the two methods will result in less variability in the distribution of the \(365\) sample means? Explain.
(f) A government inspector will visit the facility on June \(22\) to observe the sampling and to determine if the factory is in compliance with the advertised mean diameter of \(6\) inches. The manager knows that, with both sampling methods, the sample mean is an unbiased estimator of the population mean. However, the manager is unsure which method is more likely to produce a sample mean that is close to \(6\) inches on the day of sampling. Based on your previous answers, which of the two sampling methods, Method 1 or Method 2, is more likely to produce a sample mean close to \(6\) inches? Explain.

Most-appropriate topic codes (AP Statistics):

• Topic \(1.11\) — Random Samples and Simple Random Sampling (Parts \( \mathrm{a} \), \( \mathrm{b} \), \( \mathrm{c} \))
• Topic \(4.1\) — Sampling Distributions for Sample Means (Parts \( \mathrm{d} \), \( \mathrm{e} \), \( \mathrm{f} \))
▶️ Answer/Explanation

(a)
No, a sample obtained using Method 2 will not be representative of all tortillas made that day. The sample obtained using Method 2 will only represent the tortillas from one production line, not from the entire population. Because the distributions of diameters for the two production lines are different, sampling from only one line misses the true characteristics of the combined production.

(b)
Method 1 was most likely used to select this sample. The bimodal shape in the histogram of sample data indicates that tortillas were selected from both production lines (with peaks around \(5.9\) and \(6.1\)), which is what would happen using Method 1. Method 2 would be likely to produce a unimodal distribution centered at either \(5.9\) inches or \(6.1\) inches.

(c)
Method 2 would result in less variability in the sample of \(200\) tortillas on a given day because the sample comes from only one production line. Since the distributions of diameters are not the same for the two production lines, selecting tortillas from both lines (as in Method 1) combines their differences and results in more variable sample data.

(d)
The sampling distribution of the sample mean diameter for samples obtained using Method 1 would be approximately normal because the sample size is large (\(n = 200 \ge 30\)), satisfying the Central Limit Theorem.

The mean of the sampling distribution is:
\( \mu_{\bar{x}} = \mu = 6 \text{ inches} \)

The standard deviation (standard error) of the sampling distribution is:
\( \sigma_{\bar{x}} = \dfrac{\sigma}{\sqrt{n}} = \dfrac{0.11}{\sqrt{200}} \approx 0.0078 \text{ inch} \)

(e)
Method 1 would result in less variability in the distribution of the \(365\) sample means. The sample means from Method 1 will all be clustered very closely around the true population mean of \(6\) inches. Conversely, the sample means from Method 2 will be clustered around \(5.9\) inches on some days and around \(6.1\) inches on other days, creating a much wider overall spread for the \(365\) daily means.

(f)
Method 1 is more likely to produce a sample mean close to \(6\) inches. Even though both methods are unbiased estimators in the long run, Method 1 consistently samples from the entire population and its sample mean has very little variability (\(SE \approx 0.0078\)). On the day of the inspection, Method 2 will likely produce a sample mean clustered near either \(5.9\) inches or \(6.1\) inches, which is relatively far from the advertised \(6\) inches.

Question

An administrator at a large university wants to conduct a survey to estimate the proportion of students who are satisfied with the appearance of the university buildings and grounds. The administrator is considering three methods of obtaining a sample of \(500\) students from the \(70{,}000\) students at the university.
(a) Because of financial constraints, the first method the administrator is considering consists of taking a convenience sample to keep the expenses low. A very large number of students will attend the first football game of the season, and the first \(500\) students who enter the football stadium could be used as a sample. Why might such a sampling method be biased in producing an estimate of the proportion of students who are satisfied with the appearance of the buildings and grounds?
(b) Because of the large number of students at the university, the second method the administrator is considering consists of using a computer with a random number generator to select a simple random sample of \(500\) students from a list of \(70{,}000\) student names. Describe how to implement such a method.
(c) Because stratification can often provide a more precise estimate than a simple random sample, the third method the administrator is considering consists of selecting a stratified random sample of \(500\) students. The university has two campuses with male and female students at each campus. Under what circumstance(s) would stratification by campus provide a more precise estimate of the proportion of students who are satisfied with the appearance of the university buildings and grounds than stratification by gender?

Most-appropriate topic codes (AP Statistics):

• Topic \(1.12\) — Potential Problems with Sampling (Part \( \mathrm{a} \))
• Topic \(1.11\) — Random Sampling (Part \( \mathrm{b} \))
• Topic \(1.11\) — Random Sampling (Part \( \mathrm{c} \) — stratified random sampling)
▶️ Answer/Explanation

(a)

The first \(500\) students who enter the football stadium are not likely to be representative of all \(70{,}000\) students at the university. Students who attend football games tend to have stronger school pride and a more positive connection to the university overall, which likely makes them more satisfied with the appearance of the buildings and grounds than the general student population. This means the sample would systematically overestimate the proportion of students who are satisfied — producing a positively biased estimate.

(b)

Step 1: Obtain a complete list of all \(70{,}000\) students at the university and assign each student a unique identification number from \(1\) to \(70{,}000\).
Step 2: Use a computer random number generator to generate \(500\) distinct random integers between \(1\) and \(70{,}000\), ignoring any repeated numbers that appear.
Step 3: Select the students whose assigned ID numbers match the \(500\) randomly generated numbers — these students form the simple random sample.

(c)

Stratification by campus would provide a more precise estimate than stratification by gender when the variability in students’ opinions about the appearance of the university buildings and grounds is greater between the two campuses than between the two genders. In other words, if the two campuses differ noticeably from each other in appearance (for example, one campus is newer or more attractive than the other), then students on different campuses will tend to have systematically different satisfaction levels — making campus a more effective stratification variable than gender.

Question

Two students at a large high school, Peter and Rania, wanted to estimate \(\mu\), the mean number of soft drinks that a student at their school consumes in a week. A complete roster of the names and genders for the 2,000 students at their school was available. Peter selected a simple random sample of 100 students. Rania, knowing that 60 percent of the students at the school are female, selected a simple random sample of 60 females and an independent simple random sample of 40 males. Both asked all of the students in their samples how many soft drinks they typically consume in a week.
 
(a) Describe a method Peter could have used to select a simple random sample of 100 students from the school.
Peter and Rania conducted their studies as described. Peter used the sample mean \(\bar{X}\) as a point estimator for \(\mu\). Rania used \(\bar{X}_{strat} = 0.6\bar{X}_{female} + 0.4\bar{X}_{male}\) as a point estimator for \(\mu\), where \(\bar{X}_{female}\) is the mean of the sample of 60 females and \(\bar{X}_{male}\) is the mean of the sample of 40 males.
Summary statistics for Peter’s data are shown in the table below.

(b) Based on the summary statistics, calculate the estimated standard deviation of the sampling distribution (sometimes called the standard error) of Peter’s point estimator \(\bar{X}\).
Summary statistics for Rania’s data are shown in the table below.
(c) Based on the summary statistics, calculate the estimated standard deviation of the sampling distribution of Rania’s point estimator \(\bar{X}_{strat}\).
A dotplot of Peter’s sample data is given below.
Comparative dotplots of Rania’s sample data are given below.
(d) Using the dotplots above, explain why Rania’s point estimator has a smaller estimated standard deviation than the estimated standard deviation of Peter’s point estimator.

Most-appropriate topic codes (AP Statistics):

• Topic 1.11 — Random Sampling (Part a)
• Topic 4.1 — Sampling Distributions for Sample Means (Parts b, c)
• Topic 4.1 — Sampling Distributions for Sample Means (Part d)
• Topic 1.11 — Random Sampling (Part d)
▶️ Answer/Explanation

(a)
To select a proper simple random sample, Peter could number all the students on the roster from 1 to 2,000.
Then, he can use a random number generator to produce numbers between 1 and 2,000.
He would ignore any repeated numbers and continue generating until he has a list of 100 unique random numbers.
The 100 students corresponding to those selected numbers will make up his sample.

(b)
The estimated standard deviation of the sampling distribution of the sample mean (standard error) for Peter’s simple random sample is calculated using the formula \(\text{SE}(\bar{X}) = \frac{s}{\sqrt{n}}\).
Substituting the values from Peter’s sample: \(\text{SE}(\bar{X}) = \frac{4.13}{\sqrt{100}}\).
\(\text{SE}(\bar{X}) = \frac{4.13}{10} = 0.413\).
The estimated standard deviation of Peter’s point estimator is \(0.413\).

(c)
Rania’s point estimator is given by \(\bar{X}_{strat} = 0.6\bar{X}_{female} + 0.4\bar{X}_{male}\).
Because the two samples (female and male) are independent, the variance of the stratified estimator is the sum of the variances of each part, scaled by their squared weights:
\(\text{Var}(\bar{X}_{strat}) = (0.6)^2 \text{Var}(\bar{X}_{female}) + (0.4)^2 \text{Var}(\bar{X}_{male})\).
The variance for the female sample mean is \(\frac{s_{f}^2}{n_f} = \frac{(1.80)^2}{60} = \frac{3.24}{60} = 0.054\).
The variance for the male sample mean is \(\frac{s_{m}^2}{n_m} = \frac{(2.22)^2}{40} = \frac{4.9284}{40} = 0.12321\).
Substitute these into the combined variance equation:
\(\text{Var}(\bar{X}_{strat}) = (0.36)(0.054) + (0.16)(0.12321) = 0.01944 + 0.0197136 = 0.0391536\).
The estimated standard deviation is the square root of the variance:
\(\text{SE}(\bar{X}_{strat}) = \sqrt{0.0391536} \approx 0.198\).

(d)
Based on the dotplots, there is a clear difference in the soft drink consumption patterns between genders. The dotplot for males indicates a visibly higher center (around 7-8 drinks) compared to the dotplot for females (centered around 2-3 drinks).
Because of this distinct separation between the two groups, a standard simple random sample like Peter’s will experience a large amount of variation as different samples will randomly capture varying proportions of males and females, producing a wider overall spread (standard deviation of 4.13).
By contrast, Rania’s stratified method guarantees that the sample consists of exactly 60% females and 40% males. The variability within each homogeneous gender group (1.80 for females and 2.22 for males) is much smaller than the overall variability of the mixed population.
Because stratification eliminates the between-group variation from the standard error calculation, Rania’s point estimator results in a substantially smaller estimated standard deviation.

Question

An apartment building has nine floors and each floor has four apartments. The building owner wants to install new carpeting in eight apartments to see how well it wears before she decides whether to replace the carpet in the entire building.
The figure below shows the floors of apartments in the building with their apartment numbers. Only the nine apartments indicated with an asterisk (*) have children in the apartment.
(a) For convenience, the apartment building owner wants to use a cluster sampling method, in which the floors are clusters, to select the eight apartments. Describe a process for randomly selecting eight different apartments using this method.
(b) An alternative sampling procedure would be to select a stratified random sample of eight apartments, where the strata are apartments with children and apartments without children. A stratified random sample of size eight might include two apartments with children and six apartments without children. In the context of this situation, give one statistical advantage of selecting such a stratified random sample as opposed to a cluster sample of two floors.

Most-appropriate topic codes (AP Statistics):

• Topic 1.11 — Random Sampling (Part a)
• Topic 1.11 — Random Sampling (Part b)
▶️ Answer/Explanation

(a)
To select a cluster sample of eight apartments using the floors as clusters, perform the following procedure:
Assign each of the nine floors a distinct integer label from $1$ to $9$.
Use a random number generator or a table of random digits to pick a random integer from $1$ to $9$. Select all four apartments on the corresponding floor.
Pick a second random integer from $1$ to $9$. If it matches the first number, ignore it and choose another until you get a different integer. Select all four apartments on that second floor.
The final sample will consist of the eight apartments located on the two uniquely chosen floors.

(b)
The direct advantage of using a stratified random sample over a cluster sample is that it guarantees both types of apartments—those with children and those without—are represented in the study.
Because carpet wear is highly likely to differ depending on whether children live in the apartment, it is vital to collect data on both scenarios to make an accurate overall assessment.
With the cluster sample, there is a distinct chance that the two floors randomly picked contain absolutely no apartments with children (such as floors 1, 3, or 5), entirely omitting a critical source of variation and blinding the owner to how well the carpet handles heavy child-related traffic.

Question

An automobile company wants to learn about customer satisfaction among the owners of five specific car models. Large sales volumes have been recorded for three of the models, but the other two models were recently introduced so their sales volumes are smaller. The number of new cars sold in the last six months for each of the models is shown in the table below.
The company can obtain a list of all individuals who purchased new cars in the last six months for each of the five models shown in the table. The company wants to sample 2,000 of these owners.
(a) For simple random samples of 2,000 new car owners, what is the expected number of owners of model E and the standard deviation of the number of owners of model E?
(b) When selecting a simple random sample of 2,000 new car owners, how likely is it that fewer than 12 owners of model E would be included in the sample? Justify your answer.
(c) The company is concerned that a simple random sample of 2,000 owners would include fewer than 12 owners of model D or fewer than 12 owners of model E. Briefly describe a sampling method for randomly selecting 2,000 owners that will ensure at least 12 owners will be selected for each of the 5 car models.

Most-appropriate topic codes (AP Statistics):

• Topic 2.8 — Introduction to Random Variables and Probability Distributions (Part a)
• Topic 2.10 — The Binomial Distribution (Parts a, b)
• Topic 1.11 — Random Sampling (Part c)
▶️ Answer/Explanation

(a)

Because the total population ($297,354$) is overwhelmingly large compared to the sample size ($2,000$), we can treat this as a binomial distribution even though sampling is without replacement.
The probability of selecting a model E owner is $p = \dfrac{2,323}{297,354} \approx 0.007812$.
The sample size is $n = 2000$.
Expected number (Mean):
$\mu_E = n \times p = 2000 \times 0.007812 \approx 15.62$ owners
Standard Deviation:
$\sigma_E = \sqrt{n \times p \times (1-p)} = \sqrt{2000 \times 0.007812 \times (1 – 0.007812)} = \sqrt{15.49} \approx 3.93$ owners

(b)

For the reason given in part (a), the binomial distribution with $n = 2,000$ and $p \approx 0.0078$ can be used here. The probability that the sample would contain fewer than 12 owners of model E is calculated from the binomial distribution to be $\sum_{x=0}^{11} \binom{2,000}{x} (0.0078)^x (0.9922)^{2,000-x} \approx 0.147$. This probability is small enough that the result (fewer than 12 owners of model E in the sample) is not likely, but this probability is also not small enough to consider the result very unlikely.

This binomial probability can also be evaluated using a normal approximation. This is reasonable because $n \times p = (2,000) \times (0.0078) = 15.6$ is larger than 10 and $n(1 – p) = (2,000) \times (0.9922) = 1,984.4$ is much larger than 10. Using the mean and standard deviation from part (a) gives

$P(X \le 11) \approx P\left( Z < \dfrac{12.0 – 15.62}{3.94} \right) = P(Z < -0.92) = 0.179$.

(c)

To guarantee at least 12 owners from each model, the company should use a stratified random sampling method.
The researcher should use the five car models as the strata.
They can determine how many individuals they want to sample from each model (stratum) as long as every model’s assigned sample size is $12$ or greater, and all five sizes add up to exactly $2,000$.
Then, they simply perform five separate simple random samples—one within each specific car model’s list of owners—to achieve the decided quota for that model.

Question

In response to nutrition concerns raised last year about food served in school cafeterias, the Smallville School District entered into a one-year contract with the Healthy Alternative Meals (HAM) company. Under this contract, the company plans and prepares meals for \(2,500\) elementary, middle, and high school students, with a focus on good nutrition. The school administration would like to survey the students in the district to estimate the proportion of students who are satisfied with the food under this contract.
Two sampling plans for selecting the students to be surveyed are under consideration by the administration. One plan is to take a simple random sample of students in the district and then survey those students. The other plan is to take a stratified random sample of students in the district and then survey those students.
(a) Describe a simple random sampling procedure that the administrators could use to select \(200\) students from the \(2,500\) students in the district.
(b) If a stratified random sampling procedure is used, give one example of an effective variable on which to stratify in this survey. Explain your reasoning.
(c) Describe one statistical advantage of using a stratified random sample over a simple random sample in the context of this study.
 

Most-appropriate topic codes (AP Statistics):

• Topic \(1.11\) — Random Sampling (Parts \( \mathrm{a} \), \( \mathrm{b} \), \( \mathrm{c} \))
▶️ Answer/Explanation

(a)
The administrators could number an alphabetical list of students from \(1\) to \(2,500\). They could then use a random number generator to select \(200\) unique random integers from \(1\) to \(2,500\). The students corresponding to those \(200\) numbers would be asked to participate in the survey.

(b)
One effective variable is school level (elementary, middle, high school). Students’ perceptions of food satisfaction likely differ by age and development level, making each school level a group with more internally consistent opinions than the district as a whole.

(c)
A primary advantage is reducing sampling variability by ensuring that each school level is represented proportionally in the survey. This ensures that the estimate of satisfaction is more precise because it accounts for the potential differences in opinion across different student age groups.

Question

A local school board plans to conduct a survey of parents’ opinions about year-round schooling in elementary schools. The school board obtains a list of all families in the district with at least one child in an elementary school and sends the survey to a random sample of 500 of the families. The survey question is provided below.
A proposal has been submitted that would require students in elementary schools to attend school on a year-round basis. Do you support this proposal? (Yes or No)
The school board received responses from 98 of the families, with 76 of the responses indicating support for year-round schools. Based on this outcome, the local school board concludes that most of the families with at least one child in elementary school prefer year-round schooling.
(a) What is a possible consequence of nonresponse bias for interpreting the results of this survey?
(b) Someone advised the local school board to take an additional random sample of 500 families and to use the combined results to make their decision. Would this be a suitable solution to the issue raised in part (a)? Explain.
(c) Suggest a different follow-up step from the one suggested in part (b) that the local school board could take to address the issue raised in part (a).

Most-appropriate topic codes (AP Statistics):

• Topic \(1.11\) — Random Sampling (Parts \(\mathrm{a}\), \(\mathrm{b}\), \(\mathrm{c}\))
• Topic \(1.12\) — Potential Problems with Sampling (Parts \(\mathrm{a}\), \(\mathrm{b}\), \(\mathrm{c}\))
▶️ Answer/Explanation

(a)

Only 98 out of 500 families responded — that is a response rate of just \(\dfrac{98}{500} = 19.6\%\), meaning \(80.4\%\) did not reply at all.
For the survey results to be unbiased, we would need the 402 non-responding families to have similar opinions to those who did respond. But that is unlikely to be true — families who feel strongly in favour of year-round schooling may be much more motivated to respond, while those who oppose or feel indifferent may simply ignore the survey.
As a result, the observed proportion \(\dfrac{76}{98} \approx 77.6\%\) in favour likely overestimates the true proportion of all families who support the proposal. The school board’s conclusion that “most families prefer year-round schooling” may therefore be misleading.

(b)

No, taking an additional random sample of 500 families and combining the results would not be a suitable solution.
• The core problem is not the size of the sample — it is the low response rate in the original survey.
• The original 402 non-responses represent a biased portion of the data that will still be present in the combined sample, regardless of how the second sample turns out.
• If the second survey is conducted in the same way (mailed questionnaire with no follow-up), it is very likely to suffer from the same nonresponse bias, leaving the combined results just as unreliable.
Simply increasing the number of people surveyed does not fix bias — it only gives you a larger biased sample.

(c)

The school board should directly contact the 402 families who did not respond to the original survey, using a different mode of communication such as telephone calls or in-person visits, to obtain their opinions.
• This targets the exact group whose missing responses are causing the bias.
• Combining these newly obtained responses with the original 98 would give a much more complete and representative picture of opinion across all families in the district.
Alternatively, the board could design a new survey from scratch using in-person interviews or telephone calls to ensure a much higher response rate from the outset, reducing the opportunity for nonresponse bias to arise.

Question

A certain state’s education commissioner released a new report card for all the public schools in that state. This report card provides a new tool for comparing schools across the state. One of the key measures that can be computed from the report card is the student-to-teacher ratio, which is the number of students enrolled in a given school divided by the number of teachers at that school.
The data below give the student-to-teacher ratio at the 10 schools with the highest proportion of students meeting the state reading standards in the third grade and at the 10 schools with the lowest proportion of students meeting the state reading standards in the third grade.
(a) Display a dotplot for each group to compare the distribution of student-to-teacher ratios in the top 10 schools with the distribution in the bottom 10 schools. Comment on the similarities and differences between the two distributions.
(b) Any statistical test that is used to determine whether the mean student-to-teacher ratio is the same for the top 10 schools as it is for the bottom 10 schools would be inappropriate. Explain why in a few sentences.

Most-appropriate topic codes (AP Statistics):

• Topic 1.5 — Graphical Representations for One Quantitative Variable (Part \(\mathrm{a}\))
• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Part \(\mathrm{a}\))
• Topic 1.11 — Random Sampling (Part \(\mathrm{b}\))
• Topic 4.4 — Setting Up a Test for a Population Mean or Population Mean Difference (Part \(\mathrm{b}\))
▶️ Answer/Explanation

(a)
First, let’s organize the data for both groups:
Highest Proportion group: \(7, 9, 12, 16, 16, 17, 17, 18, 21, 22\)
Lowest Proportion group: \(12, 12, 14, 14, 16, 16, 18, 19, 20, 20\)
The dotplots, displayed on a common scale from \(4\) to \(24\), are shown below:

Similarities: The two distributions are centered at approximately the same place. The median for the Highest Proportion group is \(\dfrac{16+17}{2} = 16.5\) and the median for the Lowest Proportion group is \(\dfrac{16+16}{2} = 16\), so both centers are very close to \(16\).
Differences: The distribution for the Highest Proportion group is much more spread out (variable) than the distribution for the Lowest Proportion group. The range for the Highest Proportion group is \(22 – 7 = 15\), while the range for the Lowest Proportion group is only \(20 – 12 = 8\). In other words, the top schools show much greater variability in their student-to-teacher ratios compared to the bottom schools.

(b)
The two groups of schools are not random samples drawn from two larger populations of interest.
The group of 10 schools with the highest proportion of students meeting the standards is itself the entire population of such schools — it is not a random sample from some larger population of high-performing schools.
Similarly, the group of 10 schools with the lowest proportion is itself the complete population of the lowest-performing schools in the state — not a random sample from a larger population.
Since statistical inference is designed to generalize conclusions from a sample to a broader population, and these two groups are not random samples but rather complete populations defined by their extreme values, applying any inferential procedure (such as a two-sample \(t\)-test) to these data would be inappropriate. There is no larger population to generalize to.

Question

As dogs age, diminished joint and hip health may lead to joint pain and thus reduce a dog’s activity level. Such a reduction in activity can lead to other health concerns such as weight gain and lethargy due to lack of exercise. A study is to be conducted to see which of two dietary supplements, glucosamine or chondroitin, is more effective in promoting joint and hip health and reducing the onset of canine osteoarthritis. Researchers will randomly select a total of 300 dogs from ten different large veterinary practices around the country. All of the dogs are more than 6 years old, and their owners have given consent to participate in the study. Changes in joint and hip health will be evaluated after 6 months of treatment.
(a) What would be an advantage to adding a control group in the design of this study?
(b) Assuming a control group is added to the other two groups in the study, explain how you would assign the 300 dogs to these three groups for a completely randomized design.
(c) Rather than using a completely randomized design, one group of researchers proposes blocking on clinics, and another group of researchers proposes blocking on breed of dog. How would you decide which one of these two variables to use as a blocking variable?

Most-appropriate topic codes (AP Statistics):

• Topic 1.13 — Experimental Design (Parts \(\mathrm{a}\), \(\mathrm{b}\), \(\mathrm{c}\))
• Topic 1.10 — The Investigative Question Revisited and Data Collection (Part \(\mathrm{a}\))
• Topic 1.11 — Random Sampling (Part \(\mathrm{b}\))
▶️ Answer/Explanation

(a)

A control group gives the researchers a baseline comparison group — without any dietary supplement — so they can measure whether glucosamine or chondroitin actually makes a difference beyond what would happen due to the normal aging process alone.
Without a control group, we could not tell whether any improvements in joint and hip health were caused by the supplements or simply by other factors such as the passage of time, veterinary care, or natural variation between dogs.
In this study specifically, the control group allows us to isolate the true effect of glucosamine and chondroitin on reducing canine osteoarthritis by comparing both treatment groups against untreated dogs under the same conditions.
\(\boxed{\text{Control group provides a baseline to determine whether the supplements are truly effective.}}\)

(b)

First, assign each of the 300 dogs a unique number from \(001\) to \(300\).
Then, use a random number generator (calculator, statistical software, or a random number table) to randomly select 100 numbers from \(001\) to \(300\), ignoring any repeats — the dogs corresponding to these 100 numbers are assigned to the glucosamine group.
From the remaining 200 dogs, randomly select another 100 numbers using the same process — these dogs are assigned to the chondroitin group.
The final 100 remaining dogs are assigned to the control group and receive no dietary supplement.
\(\boxed{\text{Randomly assign dogs numbered } 001\text{–}300 \text{ into three equal groups of 100 using a random number generator.}}\)

(c)

The blocking variable should be the one that has a stronger association with the response variable — joint and hip health — so that dogs within each block are as similar (homogeneous) as possible.
Breed of dog is associated with the size of the dog, and size is known to be related to joint and hip health — larger breeds tend to have more joint problems than smaller breeds, so breed is likely to create more variability in the response.
Clinic, on the other hand, is less likely to be strongly associated with joint and hip health because most large veterinary practices see a wide variety of dog breeds and sizes, so dogs across clinics would not be meaningfully more similar to each other than dogs across breeds.
Therefore, we should block on breed of dog, since it is more strongly related to joint and hip health and will reduce variability more effectively within each block.
\(\boxed{\text{Block on breed of dog, as it has a stronger relationship to joint and hip health than clinic.}}\)

Question

The goal of a nutritional study was to compare the caloric intake of adolescents living in rural areas of the United States with the caloric intake of adolescents living in urban areas of the United States. A random sample of ninth-grade students from one high school in a rural area was selected. Another random sample of ninth graders from one high school in an urban area was also selected. Each student in each sample kept records of all the food he or she consumed in one day.
The back-to-back stemplot below displays the number of calories of food consumed per kilogram of body weight for each student on that day.
(a) Write a few sentences comparing the distribution of the daily caloric intake of ninth-grade students in the rural high school with the distribution of the daily caloric intake of ninth-grade students in the urban high school.
(b) Is it reasonable to generalize the findings of this study to all rural and urban ninth-grade students in the United States? Explain.
(c) Researchers who want to conduct a similar study are debating which of the following two plans to use.
Plan I: Have each student in the study record all the food he or she consumed in one day. Then researchers would compute the number of calories of food consumed per kilogram of body weight for each student for that day.
Plan II: Have each student in the study record all the food he or she consumed over the same 7-day period. Then researchers would compute the average daily number of calories of food consumed per kilogram of body weight for each student during that 7-day period.
Assuming that the students keep accurate records, which plan, I or II, would better meet the goal of the study? Justify your answer.

Most-appropriate topic codes (AP Statistics):

• Topic 1.9 — Comparisons of the Distributions for One Quantitative Variable (Part \(\mathrm{a}\))
• Topic 1.11 — Random Sampling (Part \(\mathrm{b}\))
• Topic 1.12 — Potential Problems with Sampling (Part \(\mathrm{b}\))
• Topic 1.10 — The Investigative Question Revisited and Data Collection (Part \(\mathrm{c}\))
▶️ Answer/Explanation

(a)
Reading the stemplot, the rural distribution is centered higher and is more spread out than the urban distribution.

For the rural students:

Mean \(\approx 40.45\) cal/kg
Median \(\approx 41\) cal/kg
Range \(= 19\)
SD \(\approx 6.04\)
IQR \(\approx 10\)

For the urban students:

Mean \(\approx 32.6\) cal/kg
Median \(\approx 32\) cal/kg
Range \(= 16\)
SD \(\approx 4.67\)
IQR \(\approx 7\)

So both the typical value and the spread are larger for the rural group. In terms of shape, the rural data look fairly symmetric and spread evenly between about 32 and 51 cal/kg, while the urban data appear skewed toward the larger values.

\( \boxed{\text{Rural: higher center and more spread; Urban: lower center, less spread, right-skewed}} \)

(b)
No. Each sample came from just one rural school and one urban school, so these two specific schools may not represent the much larger and more diverse population of all rural and urban ninth graders across the country. Because the schools themselves were not randomly chosen from all such schools, the results can’t be safely extended beyond these two schools.

\( \boxed{\text{No — only one school of each type was sampled, so results cannot be generalized nationally}} \)

(c)
Plan II is the better choice.

Both plans already adjust for body size by dividing calories by body weight, so that part is the same. The real issue is that a single day’s eating can be unusually high or low depending on what happened that day — a birthday party, a sick day, a weekend versus a school day, and so on. By recording food over a full 7-day period and averaging, Plan II smooths out this day-to-day variability and gives a more stable, precise picture of each student’s typical caloric intake.

\( \boxed{\text{Plan II — averaging over 7 days reduces day-to-day variability and gives a more precise estimate}} \)

Question

A survey will be conducted to examine the educational level of adult heads of households in the United States. Each respondent in the survey will be placed into one of the following two categories:
• Does not have a high school diploma
• Has a high school diploma
The survey will be conducted using a telephone interview. Random-digit dialing will be used to select the sample.
(a) For this survey, state one potential source of bias and describe how it might affect the estimate of the proportion of adult heads of households in the United States who do not have a high school diploma.
(b) A pilot survey indicated that about 22 percent of the population of adult heads of households do not have a high school diploma. Using this information, how many respondents should be obtained if the goal of the survey is to estimate the proportion of the population who do not have a high school diploma to within \(0.03\) with \(95\) percent confidence? Justify your answer.
(c) Since education is largely the responsibility of each state, the agency wants to be sure that estimates are available for each state as well as for the nation. Identify a sampling method that will achieve this additional goal and briefly describe a way to select the survey sample using this method.

Most-appropriate topic codes (AP Statistics):

• Topic 1.12 — Potential Problems with Sampling (Part \(\mathrm{a}\))
• Topic 3.3 — Constructing a Confidence Interval for a Population Proportion (Part \(\mathrm{b}\))
• Topic 1.11 — Random Sampling (Part \(\mathrm{c}\))
▶️ Answer/Explanation

(a)
One issue is that random-digit dialing only reaches people who have a telephone. People without a high school diploma tend to have lower-paying jobs, so they may be less likely to be able to afford phone service. As a result, this group could be underrepresented in the sample.
\( \boxed{\text{Households without phones are missed, and these are more likely to lack a diploma, so the estimate may be too low}} \)

(b)
The margin of error for a proportion is
\( ME=z^*\sqrt{\dfrac{p(1-p)}{n}} \)
We want \(ME=0.03\) with \(95\%\) confidence, so \(z^*=1.96\), and we use the pilot estimate \(p=0.22\):
\( 0.03=1.96\sqrt{\dfrac{0.22(0.78)}{n}} \)
Solving for \(n\), first isolate the square root:
\( \sqrt{\dfrac{0.22(0.78)}{n}}=\dfrac{0.03}{1.96} \)
Square both sides:
\( \dfrac{0.22(0.78)}{n}=\left(\dfrac{0.03}{1.96}\right)^2 \)
Solve for \(n\):
\( n=\dfrac{0.22(0.78)}{\left(\dfrac{0.03}{1.96}\right)^2} \)
\( n=\left(\dfrac{1.96}{0.03}\right)^2(0.22)(0.78) \)
\( n\approx 732.47 \)
Since \(n\) must be a whole number and we need at least this many respondents, we round up:
\( \boxed{n=733\text{ respondents}} \)

(c)
A good approach here is stratified random sampling, using each state as a stratum. Within every state, a random sample of adult heads of households would be selected and surveyed, with the sample size in each state chosen based on the precision needed for that state. Once all the state-level samples are collected, the results can be combined to produce an overall national estimate.
\( \boxed{\text{Stratified random sampling — treat each state as a stratum and take a random sample within each state}} \)

Question

A rural county hospital offers several health services. The hospital administrators conducted a poll to determine whether the residents’ satisfaction with the available services depends on their gender. A random sample of 1,000 adult county residents was selected. The gender of each respondent was recorded and each was asked whether he or she was satisfied with the services offered by the hospital. The resulting data are shown in the table below.
(a) Using a significance level of 0.05, conduct an appropriate test to determine if, for adult residents of this county, there is an association between gender and whether or not they were satisfied with services offered by the hospital.
(b) Is \(\dfrac{800}{1{,}000}\) a reasonable estimate for the proportion of all adult county residents who are satisfied with the services offered by this hospital? Explain why or why not.

Most-appropriate topic codes (AP Statistics):

• Topic 3.14 — Setting Up a Chi-Square Test for Homogeneity or Independence (Part a)
• Topic 3.15 — Carrying Out a Chi-Square Test for Homogeneity or Independence (Part a)
• Topic 3.1 — Estimators (Part b)
• Topic 1.11 — Random Sampling (Part b)
▶️ Answer/Explanation

(a)
The appropriate procedure is a chi-square test for independence.
Hypotheses:
\(H_0\): Gender and satisfaction with hospital services are independent (no association).
\(H_a\): Gender and satisfaction with hospital services are not independent (there is an association)
Conditions:
— The sample is a random sample of 1,000 adult county residents.
— All expected cell counts must be at least 5.
Compute each:
\(E = \frac{(\text{row total})(\text{column total})}{\text{grand total}}\)
\(E(\text{Satisfied, Male}) = \dfrac{800 \times 464}{1000} = 371.2\)
\(E(\text{Satisfied, Female}) = \dfrac{800 \times 536}{1000} = 428.8\)
\(E(\text{Not Satisfied, Male}) = \dfrac{200 \times 464}{1000} = 92.8\)
\(E(\text{Not Satisfied, Female}) = \dfrac{200 \times 536}{1000} = 107.2\)
All expected counts are well above 5.
Test Statistic:
\(\chi^2 = \sum \frac{(O – E)^2}{E}\)
\(\chi^2 = \frac{(384 – 371.2)^2}{371.2} + \frac{(416 – 428.8)^2}{428.8} + \frac{(80 – 92.8)^2}{92.8} + \frac{(120 – 107.2)^2}{107.2}\)
\(\chi^2 = \frac{(12.8)^2}{371.2} + \frac{(-12.8)^2}{428.8} + \frac{(-12.8)^2}{92.8} + \frac{(12.8)^2}{107.2}\)
\(\chi^2 = 0.4413 + 0.3821 + 1.7655 + 1.5284 = 4.117\)
Degrees of freedom:
\(df = (r-1)(c-1) = (2-1)(2-1) = 1\)
P-value:
Using the \(\chi^2\) distribution with \(df = 1\):
\(p\text{-value} \approx 0.0424\)
Conclusion:
Since \(p\text{-value} = 0.0424 < \alpha = 0.05\), we reject \(H_0\). There is sufficient statistical evidence at the \(0.05\) significance level to conclude that there is an association between gender and satisfaction with hospital services for adult residents of this county.
\(\boxed{\chi^2 = 4.117,\quad df = 1,\quad p\text{-value} \approx 0.0424 \Rightarrow \text{Reject } H_0}\)

(b)
Yes, \(\dfrac{800}{1{,}000} = 0.80\) is a reasonable estimate for the proportion of all adult county residents who are satisfied with hospital services. The data were collected from a random sample of 1,000 adult county residents, which means the sample is likely representative of the population of all adult county residents. Because random sampling was used, the sample proportion \(\hat{p} = 0.80\) is an unbiased estimate of the true population proportion. Additionally, with a sample size of \(n = 1{,}000\), the estimate is based on a sufficiently large and randomly selected group, giving us reasonable confidence in its accuracy.
\(\boxed{\hat{p} = \frac{800}{1{,}000} = 0.80 \text{ is a reasonable estimate (random sample, large } n\text{)}}\)

Question

At a certain university, students who live in the dormitories eat at a common dining hall. Recently, some students have been complaining about the quality of the food served there. The dining hall manager decided to do a survey to estimate the proportion of students living in the dormitories who think that the quality of the food should be improved. One evening, the manager asked the first 100 students entering the dining hall to answer the following question.

Many students believe that the food served in the dining hall needs improvement. Do you think that the quality of food served here needs improvement, even though that would increase the cost of the meal plan?

_____ Yes_____ No_____ No opinion
(a) In this setting, explain how bias may have been introduced based on the way this convenience sample was selected and suggest how the sample could have been selected differently to avoid that bias.
(b) In this setting, explain how bias may have been introduced based on the way the question was worded and suggest how it could have been worded differently to avoid that bias.

Most-appropriate topic codes (AP Statistics):

• Topic 1.11 — Random Sampling (Part a)
• Topic 1.12 — Potential Problems with Sampling (Parts a, b)
▶️ Answer/Explanation

(a)

Since the manager used a convenience sample — the first 100 students entering the cafeteria — bias may have been introduced because students who arrive at the dining hall early may have opinions about food quality that differ systematically from other dormitory residents who come later or not at all. For example, students who are very hungry or who have strong feelings about the food may be more likely to arrive early, making this group unrepresentative of all dormitory students.
To avoid this bias, the manager should have selected a random sample of 100 dormitory residents — for instance, using a simple random sample from a list of all dormitory residents, a stratified random sample by dormitory building, or a systematic random sample with a random starting point. Any of these approaches gives every dormitory resident a known, non-zero chance of being selected, which eliminates the selection bias introduced by the convenience sample.

(b)

The question as worded contains two sources of wording bias. First, the opening statement — “Many students believe that the food served in the dining hall needs improvement” — is leading because it tells respondents what other students think, which may pressure them to agree and respond “Yes” even if they do not truly feel that way.
Second, the phrase “even though that would increase the cost of the meal plan” introduces a second bias in the opposite direction, making students less likely to say “Yes” because they are reminded of a financial consequence. These two biases may push responses in opposite directions, making the results unreliable in either direction.
A better, more neutral wording would simply ask: “Do you think that the quality of food served in the dining hall needs improvement?” This removes the leading statement and the cost reminder, allowing students to respond based only on their true opinion about food quality.

Question

In order to monitor the populations of birds of a particular species on two islands, the following procedure was implemented.
Researchers captured an initial sample of 200 birds of the species on Island A; they attached leg bands to each of the birds, and then released the birds. Similarly, a sample of 250 birds of the same species on Island B was captured, banded, and released. Sufficient time was allowed for the birds to return to their normal routine and location.
Subsequent samples of birds of the species of interest were then taken from each island. The number of birds captured and the number of birds with leg bands were recorded. The results are summarized in the following table.

Assume that both the initial sample and the subsequent samples that were taken on each island can be regarded as random samples from the population of birds of this species.
(a) Do the data from the subsequent samples indicate that there is a difference in proportions of the banded birds on these two islands? Give statistical evidence to support your answer.
(b) Researchers can estimate the total number of birds of this species on an island by using information on the number of birds in the initial sample and the proportion of banded birds in the subsequent sample. Use this information to estimate the total number of birds of this species on Island A. Show your work.
(c) The analyses in parts (a) and (b) assume that the samples of birds captured in both the initial and subsequent samples can be regarded as random samples of the population of birds of this species that live on the respective islands. This is a common assumption made by wildlife researchers. Describe two concerns that should be addressed before making this assumption.

Most-appropriate topic codes (AP Statistics):

• Topic 3.12 — Setting Up a Test for the Difference Between Two Population Proportions (Part a)
• Topic 3.13 — Carrying Out a Test for the Difference Between Two Population Proportions (Part a)
• Topic 3.1 — Estimators (Part b)
• Topic 1.11 — Random Sampling (Part c)
• Topic 1.12 — Potential Problems with Sampling (Part c)
▶️ Answer/Explanation

(a)

Step 1: State hypotheses.
Let \(p_A\) = true proportion of banded birds on Island A, and \(p_B\) = true proportion of banded birds on Island B.
\(H_0: p_A – p_B = 0 \qquad H_a: p_A – p_B \neq 0\)
Step 2: Identify the test and check assumptions.
We use a two-sample \(z\)-test for a difference in proportions. The test statistic is:
\(z = \dfrac{\hat{p}_A – \hat{p}_B}{\sqrt{\hat{p}(1-\hat{p})\left(\dfrac{1}{n_1} + \dfrac{1}{n_2}\right)}}\)
The problem states the samples are random. Since the two islands are separate, the samples are independent. We check the large sample condition using the pooled estimate:
\(\hat{p} = \dfrac{n_A\hat{p}_A + n_B\hat{p}_B}{n_A + n_B} = \dfrac{12 + 35}{180 + 220} = \dfrac{47}{400} = 0.1175\)
Expected counts: \(n_A\hat{p} = 21.15,\quad n_A(1-\hat{p}) = 158.85,\quad n_B\hat{p} = 25.85,\quad n_B(1-\hat{p}) = 194.15\)
All expected counts are well above 5, so the large sample condition is satisfied.
Step 3: Compute the test statistic and p-value.
\(\hat{p}_A = \dfrac{12}{180} = 0.067 \qquad \hat{p}_B = \dfrac{35}{220} = 0.159\)
\(z = \dfrac{0.067 – 0.159}{\sqrt{\dfrac{(0.1175)(0.8825)}{180} + \dfrac{(0.1175)(0.8825)}{220}}} = \dfrac{-0.092}{\sqrt{0.00105}} = \dfrac{-0.092}{0.032} = -2.875\)
\(\text{p-value} = 2 \times P(Z < -2.875) \approx 0.00429\)
Step 4: State conclusion in context.
Since the p-value of \(0.00429\) is less than \(\alpha = 0.05\), we reject the null hypothesis. There is convincing statistical evidence that the proportions of banded birds on the two islands are different — Island B has a notably higher proportion of banded birds than Island A.

(b)

We use the capture-recapture logic: the proportion of banded birds in the subsequent sample estimates the proportion of banded birds in the whole population.
For Island A, the number of birds banded in the initial sample is \(n_I = 200\), and the proportion of banded birds observed in the subsequent sample is:
\(\hat{p}_S = \dfrac{12}{180} \approx 0.06667\)
Setting this equal to the fraction of banded birds in the population:
\(\hat{p}_S \approx \dfrac{n_I}{\text{population size}}\)
Solving for the estimated population size:
\(\text{Estimated population size} = \dfrac{n_I}{\hat{p}_S} = \dfrac{200}{12/180} = \dfrac{200 \times 180}{12} = \dfrac{36{,}000}{12} = \boxed{3{,}000 \text{ birds}}\)

(c)

Two concerns that should be addressed before assuming the captures can be treated as random samples are:
Concern 1 — Differential catchability: Some birds may be more likely to be captured than others — for example, slower, older, or less wary birds might be caught at a higher rate than the general population. If the same birds that were easy to capture in the initial sample are also more likely to appear in the subsequent sample, then banded birds would be overrepresented in the subsequent sample, leading us to underestimate the true population size.
Concern 2 — Behavioural change after banding: Birds that were captured and banded in the initial sample may become more trap-shy (avoiding capture in the future) or, conversely, may be more conspicuous to predators due to the bands, altering their survival or behaviour. If banded birds are less likely to be recaptured, we would overestimate the population size. In either case, if banding changes the birds’ behaviour or survival, the subsequent sample can no longer be treated as a true random sample of the population.

Question

Because of concerns about employee stress, a large company is conducting a study to compare two programs (tai chi or yoga) that may help employees reduce their stress levels. Tai chi is a \(1{,}200\)-year-old practice, originating in China, that consists of slow, fluid movements. Yoga is a practice, originating in India, that consists of breathing exercises and movements designed to stretch and relax muscles. The company has assembled a group of volunteer employees to participate in the study during the first half of their lunch hour each day for a \(10\)-week period. Each volunteer will be assigned at random to one of the two programs. Volunteers will have their stress levels measured just before beginning the program and \(10\) weeks later at the completion of it.
(a) A group of volunteers who work together ask to be assigned to the same program so that they can participate in that program together. Give an example of a problem that might arise if this is permitted. Explain to this volunteer group why random assignment to the two programs will address this problem.
(b) Someone proposes that a control group be included in the design as well. The stress level would be measured for each volunteer assigned to the control group at the start of the study and again \(10\) weeks later. What additional information, if any, would this provide about the effectiveness of the two programs?
(c) Is it reasonable to generalize the findings of this study to all employees of this company? Explain.

Most-appropriate topic codes (AP Statistics):

• Topic 1.13 — Experimental Design (Parts a, b)
• Topic 1.11 — Random Sampling (Part c)
• Topic 1.12 — Potential Problems with Sampling (Part c)
▶️ Answer/Explanation

(a)
If volunteers who work together are all placed in the same program, there’s a risk that something specific to their workplace situation gets mixed up with the effect of the program itself. For example, suppose this group’s department recently had a deadline pushed back, which on its own would lower everyone’s stress level in that department regardless of which program they’re doing. If this entire group ends up in, say, the tai chi group, then the drop in stress they experience could mistakenly be credited to tai chi, when really it was caused by the lighter workload.

Random assignment fixes this issue. By randomly assigning volunteers to the two programs instead of letting groups choose, we spread people from this department across both the tai chi and yoga groups. This way, any unusual circumstance affecting that department’s stress levels — like the deadline change — gets “evened out” between the two treatment groups rather than being concentrated in just one. Randomization helps make sure the two groups are comparable at the start, so that any difference we see at the end can be attributed to the program rather than to some other confounding factor.

(b)
Yes, a control group would add useful information. Without one, the company could only compare tai chi to yoga directly — they could say which of the two programs led to a bigger drop in stress, but they couldn’t say whether either program actually caused a reduction in stress at all.

Here’s the issue: stress levels might naturally go down over a \(10\)-week period for reasons that have nothing to do with either program — for instance, if the overall work environment becomes less hectic during that time, everyone’s stress might drop a little just from that. A control group, which doesn’t participate in either program but still has its stress measured at the start and end of the \(10\) weeks, gives a baseline for what “no treatment” looks like under those same conditions.

By comparing each treatment group’s change in stress to the control group’s change, the company can tell how much of the reduction is actually attributable to tai chi or yoga specifically, rather than just background changes that would have happened anyway.

(c)
No, it is not reasonable to generalize these findings to all employees of the company. The participants in this study were volunteers, not a random sample of employees. People who choose to volunteer for a stress-reduction study might already be different from the typical employee — for example, they might be more motivated to manage their stress, more open to trying tai chi or yoga, or have more flexible schedules that let them give up part of their lunch hour.

Because the group wasn’t randomly selected from the entire employee population, there’s no guarantee that what works (or doesn’t work) for these volunteers would apply the same way to employees who didn’t volunteer. So while the random assignment within the study supports drawing cause-and-effect conclusions about tai chi versus yoga for people like these volunteers, it doesn’t justify extending those conclusions to the company as a whole.

Scroll to Top