Home / AP® Exam / AP® Statistics / AP Statistics 5.5 Least-Squares Regression- Exam Style Questions – FRQs

AP Statistics 5.5 Least-Squares Regression- Exam Style Questions - FRQs - New Syllabus

Question

Wildlife biologists are interested in the health of tule elk, a species of deer found in California. An important measurement of tule elk health is their weight. The weight of a tule elk is difficult to measure in the wild. However, chest circumference, which is believed to be related to the weight of a tule elk, can easily be measured from a safe distance using a harmless laser. A study was done to investigate whether chest circumference, in centimeters (cm), could be used to accurately estimate the weight, in kilograms (kg), of male tule elk. For the study, wildlife biologists captured 30 male tule elk, measured their chest circumference and weight, and then released the elk. The data for the 30 male tule elk are shown in the scatterplot.
(a) Describe the relationship between chest circumference and weight of male tule elk in context.
Following is the equation of the least-squares regression line relating chest circumference and weight for male tule elk.
$\text{predicted weight} = -350.3 + 3.7455(\text{chest circumference})$
(b) The weight of one male tule elk with a chest circumference of $145.9$ cm is $204.3$ kg.
(i) Using the equation of the least-squares regression line, calculate the predicted weight for this male tule elk. Show your work.
(ii) Calculate the residual for this male tule elk. Show your work.
(c) Interpret the slope of the least-squares regression line in context.
(d) The sambar, another species of deer, is similar in size to the tule elk. The slope of the population regression line relating chest circumference and weight for all male sambars is $4.5$ kilograms per centimeter. A wildlife biologist wants to determine whether the slope of the population regression line for male tule elk is different than that for male sambars. Let $\beta$ represent the slope of the population regression line for male tule elk. The wildlife biologist conducted a test of the following hypotheses using the sample of 30 tule elk.
$H_0: \beta = 4.5$
$H_a: \beta \ne 4.5$
The test statistic was calculated to be $3.408$. Assume all conditions for inference were met.
(i) Determine the p-value of the test.
(ii) At a significance level of $\alpha = 0.05$, what conclusion should the wildlife biologist make regarding the slope of the population regression line for male tule elk? Justify your response.
 

Most-appropriate topic codes (AP Statistics):

• Topic \(5.1\) — Graphical Representations Between Two Quantitative Variables (Part \( \mathrm{a} \))
• Topic \(5.3\) — Linear Regression Models (Parts \( \mathrm{b} \), \( \mathrm{c} \))
• Topic \(5.4\) — Residuals (Part \( \mathrm{d} \))
• Topic \(5.5\) — Least-Squares Regression (Part \( \mathrm{d} \))
▶️ Answer/Explanation

(a)
The relationship between chest circumference and weight of male tule elk is strong, positive, and linear. As the chest circumference of male tule elk increases, their weight tends to increase.

When describing a scatterplot, always cover the Direction, Form, and Strength (DFS), and make sure to include context by explicitly naming the variables (chest circumference and weight).

(b) (i)
$\text{predicted weight} = -350.3 + 3.7455(145.9)$
$\text{predicted weight} = 196.168 \text{ kg}$

(b) (ii)
$\text{residual} = \text{actual weight} – \text{predicted weight}$
$\text{residual} = 204.3 – 196.168$
$\text{residual} = 8.132 \text{ kg}$

(c)
For each additional centimeter increase in chest circumference, the predicted weight of a male tule elk increases by $3.7455$ kilograms.

(d) (i)
The degrees of freedom for regression slope inference is $df = n – 2$. With a sample size of $n = 30$, $df = 30 – 2 = 28$.
Using the t-distribution table with $df = 28$ and a test statistic of $t = 3.408$, the one-tail area is exactly $0.001$.
Since the alternative hypothesis ($H_a: \beta \ne 4.5$) is two-sided, the p-value is $2 \times 0.001 = 0.002$.

(d) (ii)
Because the p-value of $0.002$ is less than the significance level of $\alpha = 0.05$, we reject the null hypothesis.
There is convincing statistical evidence to conclude that the slope of the population regression line relating chest circumference to weight for male tule elk is different than $4.5$ kg/cm.

Question

A student measured the heights and the arm spans, rounded to the nearest inch, of each person in a random sample of \(12\) seniors at a high school. A scatterplot of arm span versus height for the \(12\) seniors is shown.
(a) Based on the scatterplot, describe the relationship between arm span and height for the sample of \(12\) seniors.
Let \(x\) represent height, in inches, and let \(y\) represent arm span, in inches. Two scatterplots of the same data are shown below. Graph \(1\) shows the data with the least squares regression line \(\hat{y}=11.74+0.8247x\), and graph \(2\) shows the data with the line \(y=x\).
(b) The criteria described in the table below can be used to classify people into one of three body shape categories: square, tall rectangle, or short rectangle.
i. For which graph, \(1\) or \(2\), is the line helpful in classifying a student’s body shape as square, tall rectangle, or short rectangle? Explain.
ii. Complete the table of classifications for the \(12\) seniors.
(c) Using the best model for prediction, calculate the predicted arm span for a senior with height \(61\) inches.

Most-appropriate topic codes (AP Statistics):

• Topic \(5.1\) — Graphical Representations Between Two Quantitative Variables (Parts \( \mathrm{a} \), \( \mathrm{b} \))
• Topic \(5.5\) — Least-Squares Regression (Part \( \mathrm{c} \))
▶️ Answer/Explanation

(a)
There is a moderately strong, positive, linear relationship between height and arm span so that taller students tend to have longer arm spans.

(b)(i)
The line in Graph \(2\) is the one that is helpful. For each student, the graph illustrates whether arm span is equal to height (points on the line), arm span is greater than height (points above the line), or arm span is less than height (points below the line).

(b)(ii)
The frequencies of the \(12\) seniors classified by body shape are:
Square: \(3\)
Tall Rectangle: \(4\)
Short Rectangle: \(5\)

(c)
The predicted arm span is calculated using the given least squares regression line equation:
\( \hat{y} = 11.74 + 0.8247x \)
\( \hat{y} = 11.74 + 0.8247(61) \)
\( \hat{y} = 62.05 \text{ inches} \)

Question

Jamal is researching the characteristics of a car that might be useful in predicting the fuel consumption rate (FCR); that is, the number of gallons of gasoline that the car requires to travel 100 miles under conditions of typical city driving. The length of a car is one explanatory variable that can be used to predict FCR. Graph I is a scatterplot showing the lengths of 66 cars plotted with the corresponding FCR. One point on the graph is labeled A.
Jamal examined the scatterplot and determined that a linear model would be a reasonable way to express the relationship between FCR and length. A computer output from a linear regression is shown below.
Linear Fit
\(\widehat{\text{FCR}} = -1.595789 + 0.0372614 \times \text{Length}\)
Summary of Fit
RSquare 0.250401
Root Mean Square Error 0.902382
Observations 66
(a) The point on the graph labeled A represents one car of length 175 inches and an FCR of 5.88. Calculate and interpret the residual for the car relative to the least squares regression line.
Jamal knows that it is possible to predict a response variable using more than one explanatory variable. He wants to see if he can improve the original model of predicting FCR from length by including a second explanatory variable in addition to length. He is considering including engine size, in liters, or wheel base (the length between axles), in inches. Graph II is a scatterplot showing the engine size of the 66 cars plotted with the corresponding residuals from the regression of FCR on length. Graph III is a scatterplot showing the wheel base of the 66 cars plotted with the corresponding residuals from the regression of FCR on length.
(b) In graph II, the point labeled A corresponds to the same car whose point was labeled A in graph I. The measurements for the car represented by point A are given below.
(i) Circle the point on graph III that corresponds to the car represented by point A on graphs I and II.
(ii) There is a point on graph III labeled B. It is very close to the horizontal line at 0. What does that indicate about the FCR of the car represented by point B?
(c) Write a few sentences to compare the association between the variables in graph II with the association between the variables in graph III.
(d) Jamal wants to predict FCR using length and one of the other variables, engine size or wheel base. Based on your response to part (c), which variable, engine size or wheel base, should Jamal use in addition to length if he wants to improve the prediction? Explain why you chose that variable.

Most-appropriate topic codes (AP Statistics):

• Topic \(5.3\) — Linear Regression Models (Parts \( \mathrm{a} \), \( \mathrm{d} \))
• Topic \(5.4\) — Residuals (Parts \( \mathrm{a} \), \( \mathrm{b} \), \( \mathrm{c} \), \( \mathrm{d} \))
• Topic \(5.5\) — Least-Squares Regression (Part \( \mathrm{a} \))
▶️ Answer/Explanation

(a)
Plug the length of 175 inches into the least squares regression equation to get the predicted FCR:
\(\widehat{\text{FCR}} = -1.595789 + 0.0372614 \times 175\)
\(\widehat{\text{FCR}} \approx 4.92 \text{ gallons per 100 miles}\)
Now compute the residual using the formula \(\text{residual} = \text{observed} – \text{predicted}\):
\(\text{residual} = 5.88 – 4.92 = 0.96\)
\(\boxed{\text{residual} \approx 0.96 \text{ gallons per 100 miles}}\)
The residual of \(0.96\) means that this car’s actual FCR is \(0.96\) gallons per 100 miles higher than what the least squares regression line would predict for a car of length 175 inches — so the model underestimates the fuel consumption for this particular car.

(b)(i)


Point A has a wheel base of 93 inches and a residual of approximately \(0.96\) gallons per 100 miles (from part a). So on Graph III, the point to circle is the one located at approximately \((93,\ 0.96)\).
\(\boxed{\text{Circle the point at wheel base} = 93 \text{ in., residual} \approx 0.96}\)

(b)(ii)
A residual very close to 0 means the observed FCR and the predicted FCR (from the regression of FCR on length) are nearly equal for that car — in other words, the length-based regression model predicts that car’s fuel consumption almost perfectly, leaving very little unexplained.
\(\boxed{\text{The car’s actual FCR} \approx \text{its FCR predicted by the length-based regression line}}\)

(c)
Graph II shows a moderate, positive, linear association between engine size and the residuals from the regression of FCR on length — as engine size increases, the residuals tend to increase as well. Graph III, on the other hand, shows little to no discernible pattern between wheel base and those same residuals; the points are scattered without any clear direction or trend. Overall, the association in Graph II is noticeably stronger than in Graph III.

(d)
Jamal should add engine size to the model along with length.
Because Graph II shows a stronger association between engine size and the residuals from the length-only regression, adding engine size will explain more of the leftover variability that length alone cannot account for. Wheel base shows almost no relationship with those residuals (Graph III), so including it would do little to improve the model’s predictions.
\(\boxed{\text{Choose engine size — it has a stronger association with the residuals, reducing unexplained variability more effectively.}}\)

Question

Wind windmills generate electricity by transferring energy from the wind to a turbine. A study was conducted to examine the relationship between wind velocity and electricity production. For a particular windmill, data were collected on the wind velocity, in miles per hour (mph), and the electricity production, in amperes, for 25 randomly selected days. A computer output for the regression analysis of electricity production predicted from wind velocity is given below.
(a) Use the computer output to determine the equation of the least squares regression line attributable to predicting electricity production from wind velocity.
(b) On a day with a wind velocity of 25 mph, how much more electricity would the windmill be expected to generate than on a day with a wind velocity of 15 mph? Show how you arrived at your answer.
(c) What proportion of the variation in electricity production is explained by its linear relationship with wind velocity?
(d) Is there statistically significant evidence that electricity production is linearly related to wind velocity? Explain your answer.

Most-appropriate topic codes (AP Statistics):

• Topic 5.3 — Linear Regression Models (Parts a & b)
• Topic 5.5 — Least-Squares Regression (Part c)
• Topic 5.3 — Linear Regression Models (Part d)
▶️ Answer/Explanation

(a)
From the table, the $y$-intercept (constant coefficient) is $0.137$ and the slope coefficient (wind velocity coefficient) is $0.240$.
$\widehat{\text{Electricity Production}} = 0.137 + 0.240 \times (\text{Wind Velocity})$

(b)
The slope coefficient, $b_1 = 0.240$, indicates that each additional mph of wind speed increases expected electricity generation by $0.240$ amperes.
The difference in wind velocities is $25 – 15 = 10 \text{ mph}$.
$\text{Expected Increase} = 10 \times 0.240 = 2.40 \text{ amperes}$
Alternatively, calculating predicted values individually gives $\hat{y}_{25} = 0.137 + 0.240(25) = 6.137$ and $\hat{y}_{15} = 0.137 + 0.240(15) = 3.737$, yielding a difference of $6.137 – 3.737 = 2.40 \text{ amperes}$.

(c)
The proportion of variation in the response variable explained by the linear model is given by the coefficient of determination, $R^2$.
From the output, $R\text{-Sq} = 87.3\%$.
Therefore, $0.873$ (or $87.3\%$) of the variation in electricity production is explained by the linear relationship with wind speed.

(d)
Yes, there is statistically significant evidence of a linear relationship.
The row for the “Wind Velocity” explanatory variable shows a $t$-test statistic of $12.63$ and a corresponding $p$-value of $0.000$.
Because the $p$-value is essentially $0$, which is less than any common significance level (such as $\alpha = 0.05$ or $\alpha = 0.01$), we reject the null hypothesis that the true population slope is zero ($\beta_1 = 0$) and conclude that wind velocity is a useful linear predictor of electricity output.

Question

Grass buffer strips are grassy areas that are planted between bodies of water and agricultural fields. These strips are designed to filter out sediment, organic material, nutrients, and chemicals carried in runoff water. The figure below shows a cross-sectional view of a grass buffer strip that has been planted along the side of a stream.
A study in Nebraska investigated the use of buffer strips of several widths between 5 feet and 15 feet. The study results indicated a linear relationship between the width of the grass strip (\(x\)), in feet, and the amount of nitrogen removed from the runoff water (\(y\)), in parts per hundred. The following model was estimated.
\(\hat{y} = 33.8 + 3.6x\)
(a) Interpret the slope of the regression line in the context of this question.
(b) Would you be willing to use this model to predict the amount of nitrogen removed for grass buffer strips with widths between 0 feet and 30 feet? Explain why or why not.
A scientist in California wants to know if there is a similar relationship in her area. To investigate this, she will place a grass buffer strip between a field and a nearby stream at each of eight different locations and measure the amount of nitrogen that the grass buffer strip removes, in parts per hundred, from runoff water at each location. Each of the eight locations can accommodate a buffer strip between 6 feet and 13 feet in width. The scientist wants to investigate which combination of widths will provide the best estimate of the slope of the regression line.
Suppose the scientist decides to use buffer strips of width 6 feet at each of four locations and buffer strips of width 13 feet at each of the other four locations. Assume the model, \(\hat{y} = 33.8 + 3.6x\), estimated from the Nebraska study is the true regression line in California and the observations at the different locations are normally distributed with standard deviation of 5 parts per hundred.
(c) Describe the sampling distribution of the sample mean of the observations on the amount of nitrogen removed by the four buffer strips with widths of 6 feet.
(d) Using your result from part (c), show how to construct an interval that has probability 0.95 of containing the sample mean of the observations from four buffer strips with widths of 6 feet.
For the study plan being implemented by the scientist in California, the graph on the left below displays intervals that each have probability 0.95 of containing the sample mean of the four observations for buffer strips of width 6 feet and for buffer strips of width 13 feet. A second possible study plan would use buffer strips of width 8 feet at four of the eight locations and buffer strips of width 10 feet at the other four locations. Intervals that each have probability 0.95 of containing the mean of the four observations for buffer strips of width 8 feet and for buffer strips of width 10 feet, respectively, are shown in the graph on the right below.
If data are collected for the first study plan, a sample mean will be computed for the four observations from buffer strips of width 6 feet and a second sample mean will be computed for the four observations from buffer strips of width 13 feet. The estimated regression line for those eight observations will pass through the two sample means. If data are collected for the second study plan, a similar method will be used.
(e) Use the plots above to determine which study plan, the first or the second, would provide a better estimator of the slope of the regression line. Explain your reasoning.
(f) The previous parts of this question used the assumption of a straight-line relationship between the width of the buffer strip and the amount of nitrogen that is removed, in parts per hundred. Although this assumption was motivated by prior experience, it may not be correct. Describe another way of choosing the widths of the buffer strips at eight locations that would enable the researchers to check the assumption of a straight-line relationship.

Most-appropriate topic codes (AP Statistics):

• Topic 5.3 — Linear Regression Models (Parts a, b)
• Topic 4.1 — Sampling Distributions for Sample Means (Parts c, d)
• Topic 5.5 — Least-Squares Regression (Part e)
• Topic 1.13 — Experimental Design (Part f)
▶️ Answer/Explanation

(a)

The slope of the regression line is \(3.6\).
This means that for each additional foot added to the width of the grass buffer strip, the amount of nitrogen removed from the runoff water increases by approximately \(3.6\) parts per hundred, on average.

(b)

No — this model should not be used for widths between 0 and 30 feet.
The Nebraska study only investigated buffer strips with widths between 5 feet and 15 feet, so the linear relationship was established only within that range.
Predicting for widths as small as 0 feet or as large as 30 feet would be extrapolation far beyond the data, making such predictions unreliable and potentially meaningless.

(c)

When the buffer strip width is \(x = 6\) feet, the true mean nitrogen removed is predicted by the model as:
\(\mu = 33.8 + 3.6(6) = 33.8 + 21.6 = 55.4 \text{ parts per hundred}\)
Since individual observations are normally distributed with standard deviation \(\sigma = 5\), the sampling distribution of the sample mean \(\bar{x}\) of four observations is normal with:
\(\mu_{\bar{x}} = 55.4 \text{ parts per hundred}\)
\(\sigma_{\bar{x}} = \frac{\sigma}{\sqrt{n}} = \frac{5}{\sqrt{4}} = 2.5 \text{ parts per hundred}\)
So the sampling distribution is \(N(55.4,\ 2.5)\).

(d)

Since the sampling distribution of \(\bar{x}\) is normal, use the critical value \(z^* = 1.96\) for a probability of 0.95.
The interval is constructed as:
\(\mu_{\bar{x}} \pm z^* \cdot \sigma_{\bar{x}} = 55.4 \pm 1.96 \times 2.5 = 55.4 \pm 4.9\)
\(\Rightarrow \left(55.4 – 4.9,\ \ 55.4 + 4.9\right) = (50.5,\ \ 60.3)\)
There is probability 0.95 that the sample mean of four 6-foot buffer strip observations falls between \(50.5\) and \(60.3\) parts per hundred.

(e)

The first study plan (widths 6 feet and 13 feet) provides the better estimator of the slope.
The estimated regression line must pass through the two sample means, so any variation in those sample means produces variation in the estimated slope.
In Study Plan 1, the two \(x\)-values (\(6\) ft and \(13\) ft) are spread far apart; even with vertical spread in the 0.95-probability intervals, the range of possible connecting slopes is relatively narrow — as seen in the left graph.
In Study Plan 2, the two \(x\)-values (\(8\) ft and \(10\) ft) are close together; the same vertical spread in the intervals produces a much wider range of possible slopes — as seen in the right graph.
Therefore, the sampling variability of the estimated slope \(\hat{b}\) is smaller under Study Plan 1, making it the better estimator of the true slope.

(f)

To check the linearity assumption, the researcher should use buffer strips of more than two different widths spread across the entire range of interest (6 to 13 feet).
For example, she could use eight different widths — one at each location — such as 6, 7, 8, 9, 10, 11, 12, and 13 feet.
With data at many distinct widths, a scatterplot of nitrogen removed versus strip width would reveal whether the relationship follows a straight line or shows curvature, directly allowing the researchers to assess the straight-line assumption.

Question

A real estate agent is interested in developing a model to estimate the prices of houses in a particular part of a large city. She takes a random sample of 25 recent sales and, for each house, records the price (in thousands of dollars), the size of the house (in square feet), and whether or not the house has a swimming pool. This information, along with regression output for a linear model using size to predict price, is shown below.
(a) Interpret the slope of the least squares regression line in the context of the study.
(b) The second house in the table has a residual of 49. Interpret this residual value in the context of the study.
The real estate agent is interested in investigating the effect of having a swimming pool on the price of a house.
(c) Use the residuals from all 25 houses to estimate how much greater the price for a house with a swimming pool would be, on average, than the price for a house of the same size without a swimming pool.
To further investigate the effect of having a swimming pool on the price of a house, the real estate agent creates two regression models, one for houses with a swimming pool and one for houses without a swimming pool. Regression output for these two models is shown below.
(d) The conditions for inference have been checked and verified, and a 95 percent confidence interval for the true difference in the two slopes is \((-0.099,\ 0.110)\). Based on this interval, is there a significant difference in the two slopes? Explain your answer.
(e) Use the regression model for houses with a swimming pool and the regression model for houses without a swimming pool to estimate how much greater the price for a house with a swimming pool would be than the price for a house of the same size without a swimming pool. How does this estimate compare with your result from part (c)?

Most-appropriate topic codes (AP Statistics):

• Topic \(5.3\) — Linear Regression Models (Parts \(\mathrm{a}\), \(\mathrm{e}\))
• Topic \(5.4\) — Residuals (Parts \(\mathrm{b}\), \(\mathrm{c}\))
• Topic \(5.5\) — Least-Squares Regression (Parts \(\mathrm{a}\), \(\mathrm{b}\), \(\mathrm{c}\), \(\mathrm{e}\))
• Topic \(4.8\) — Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Means (Part \(\mathrm{d}\))
▶️ Answer/Explanation

(a)
The slope of the least squares regression line is \(0.165\) (in thousands of dollars per square foot).
In context: for each additional square foot of house size, the predicted price of the house increases by \(0.165\) thousand dollars, or \$165, on average.
The slope tells us the rate at which the model expects price to grow with size — not a guarantee for any individual house, but the average trend across houses in this part of the city.

(b)
The residual value of 49 for this house indicates that its actual price is 49 thousand dollars higher than the model would predict for a house of its size.

(c)
We estimate the pool premium by comparing the average residuals of the two groups. If a group’s residuals average positive, the model consistently underestimates their prices; if negative, it overestimates.
Houses with a swimming pool (8 houses, residuals: \(6, 49, -18, 42, 1, 50, -23, 42\)):
\(\bar{e}_{\text{pool}} = \frac{6 + 49 + (-18) + 42 + 1 + 50 + (-23) + 42}{8} = \frac{149}{8} = 18.625 \text{ thousand dollars}\)
Houses without a swimming pool (17 houses, residuals: \(13, 26, -45, 22, 10, -46, -57, 1, -2, -69, 23, 44, -19, 26, -58, -52, 33\)):
\(\bar{e}_{\text{no pool}} = \frac{13 + 26 + (-45) + 22 + 10 + (-46) + (-57) + 1 + (-2) + (-69) + 23 + 44 + (-19) + 26 + (-58) + (-52) + 33}{17} = \frac{-150}{17} \approx -8.824 \text{ thousand dollars}\)
The estimated price premium for a swimming pool is the difference between these two averages:
\(\bar{e}_{\text{pool}} – \bar{e}_{\text{no pool}} = 18.625 – (-8.824) = \boxed{27.4 \text{ thousand dollars}}\)
This tells us that, for two houses of the same size, the one with a swimming pool is estimated to cost about \$27,400 more. The logic: pool houses have residuals that average \$18,625 above the model’s predictions, while no-pool houses sit \$8,824 below — that gap reflects the pool’s unmodeled contribution to price.

(d)
The 95% confidence interval for the true difference in slopes is \((-0.099,\ 0.110)\).
Since this interval contains zero, we cannot conclude there is a statistically significant difference between the two slopes at the 5% significance level. Zero is a plausible value for the true difference, which means it is entirely possible that the two population regression lines have the same slope.
In practical terms: the rate at which price increases with size appears to be the same for pool homes and non-pool homes — a pool shifts the price up by a roughly constant amount, but doesn’t change how sensitive the price is to square footage.

(e)
Since the two slopes are not significantly different, we pick a house size within the data range — say, \(\text{size} = 2{,}250\) sq ft (near the center of the distribution) — and compare predicted prices from both models.
Predicted price with pool:
\(\widehat{\text{Price}}_{\text{pool}} = -11.602 + 0.166 \times 2250 = -11.602 + 373.500 = 361.898 \text{ thousand dollars}\)
Predicted price without pool:
\(\widehat{\text{Price}}_{\text{no pool}} = -27.382 + 0.160 \times 2250 = -27.382 + 360.000 = 332.618 \text{ thousand dollars}\)
Estimated price premium for a pool:
\(361.898 – 332.618 = \boxed{29.280 \text{ thousand dollars} \approx \$29{,}280}\)
Comparison with part (c): The estimate from part (e), approximately \$29,280, is quite similar to the \$27,400 estimate obtained in part (c) from the residual averages. Both methods point to a pool adding roughly \$27,000–\$29,000 to the price of a house, giving us confidence that this is a reasonable estimate of the pool’s effect regardless of which approach we use.

Note — Alternative approach (difference in intercepts): Because the slopes were found not to be significantly different, we can also subtract the two fitted equations directly:
\((-11.602 + 0.166 \cdot x) – (-27.382 + 0.160 \cdot x) = 15.780 + 0.006 \cdot x\)
This gives the price difference as a function of size. For \(x = 2250\): \(15.780 + 0.006 \times 2250 = 15.780 + 13.500 = 29.280\), consistent with the calculation above.

Question

Administrators in a large school district wanted to determine whether students who attended a new magnet school for one year achieved greater improvement in science test performance than students who did not attend the magnet school. Knowing that more parents would want to enroll their children in the magnet school than there was space available for those children, the district administrators decided to conduct a lottery of all families who expressed interest in participating. In their data analysis, the administrators would then compare the change in test scores of those children who were selected to attend the magnet school with the change in test scores of those who applied to attend the magnet school but who were not selected.
The tables below show the scores on the same science pretest and the same science posttest for 20 students. Of the 20 students, 8 were randomly selected from the magnet school and 12 were randomly selected from those who applied to attend the magnet school but who were not selected and then attended their original school.
(a) Perform a test to determine whether students who attend the magnet school demonstrate a significantly higher mean difference in test scores \((\text{Posttest} – \text{Pretest})\) than students who applied to attend the magnet school but who were not selected and then attended their original school.
Administrators were also interested in using pretest scores on this test as a predictor of posttest scores on the test. The following computer output contains the results from separate regression analyses on the magnet school scores and on the original school scores. The accompanying graph displays the data and separate regression lines for the magnet and original schools.

(b)

(i) State the equation of the regression line for the magnet school and interpret its slope in the context of the question.
(ii) State the equation of the regression line for the original school and interpret its slope in the context of the question.
(c) To determine whether there is a significant correlation between pretest score and posttest score, a test of the following hypotheses will be performed.
\(H_0\): There is no correlation between pretest score and posttest score (true slope \(= 0\))
versus
\(H_a\): There is a correlation between pretest score and posttest score (true slope \(\neq 0\))
(i) Using the regression output, state the \(p\)-value and conclusion for this test at the magnet school. Assume the conditions for inference have been met.
(ii) Using the regression output, state the \(p\)-value and conclusion for this test at the original school. Assume the conditions for inference have been met.
(d) What additional information do the regression analyses give you about student performance on the science test at the two schools beyond the comparison of mean differences in part (a)?

Most-appropriate topic codes (AP Statistics):

• Topic 4.4 — Setting Up a Test for a Population Mean or Population Mean Difference (Part \(\mathrm{a}\))
• Topic 4.5 — Carrying Out a Test for a Population Mean or Population Mean Difference (Part \(\mathrm{a}\))
• Topic 5.3 — Linear Regression Models (Part \(\mathrm{b}\))
• Topic 5.2 — Correlation (Part \(\mathrm{c}\))
• Topic 5.5 — Least-Squares Regression (Parts \(\mathrm{b}\), \(\mathrm{c}\), \(\mathrm{d}\))
▶️ Answer/Explanation

(a)

Step 1 — Hypotheses
Let \(\mu_{\text{DiffM}}\) = the mean difference (posttest \(-\) pretest) for all students at the magnet school, and \(\mu_{\text{DiffO}}\) = the mean difference for all students who applied but were not selected and attended their original school.
\(H_0: \mu_{\text{DiffM}} = \mu_{\text{DiffO}}\)
\(H_a: \mu_{\text{DiffM}} > \mu_{\text{DiffO}}\)

Step 2 — Test and Conditions

We use a two-sample \(t\)-test for the difference of two means:
\(t = \dfrac{\bar{x}_M – \bar{x}_O}{\sqrt{\dfrac{s_M^2}{n_M} + \dfrac{s_O^2}{n_O}}}\)

  1. We need to assume randomness of the sampling used. It was stated in the stem that the students from the two different schools were randomly selected.
  2. We need to check the assumption that the distributions of differences (posttest – pretest) for each of the two schools are normally distributed. Based on histograms and boxplots of these differences, there are no outliers or extreme skewness. Because these graphs reveal no obvious departures from normality, it appears reasonable to proceed with the t-test.

Step 3 — Test Statistic and \(p\)-value
\(t = \dfrac{11.750 – 3.000}{\sqrt{\dfrac{(9.407)^2}{8} + \dfrac{(3.977)^2}{12}}} = \dfrac{8.750}{\sqrt{11.062 + 1.318}} = \dfrac{8.750}{\sqrt{12.380}} = \dfrac{8.750}{3.518} \approx 2.487\)
\(df \approx 8.69\), \(\quad p\text{-value} \approx 0.0177\)

Step 4 — Conclusion
Since \(p = 0.0177 < \alpha = 0.05\), we reject \(H_0\). There is convincing evidence that students who attend the magnet school have a higher mean improvement in science test scores than students who attended their original school.

(b)(i)

The regression equation for the magnet school is:
\(\hat{y} = 73.27 + 0.1811x\)
where \(x\) is the pretest score and \(\hat{y}\) is the predicted posttest score. The slope of \(0.1811\) means that for each additional point scored on the pretest by a magnet school student, the posttest score is predicted to increase by \(0.1811\) points, on average. The slope is positive but very close to zero, suggesting that pretest performance has almost no predictive power for posttest performance at the magnet school.

(b)(ii)

The regression equation for the original school is:
\(\hat{y} = 9.24 + 0.9204x\)
where \(x\) is the pretest score and \(\hat{y}\) is the predicted posttest score. The slope of \(0.9204\) means that for each additional point scored on the pretest by an original school student, the posttest score is predicted to increase by approximately \(0.9204\) points, on average — a nearly one-for-one relationship.

(c)(i) — Magnet School
From the regression output, the test statistic is \(t = 0.40\) with \(p\text{-value} = 0.706\).
Since \(0.706 > 0.05\), we fail to reject \(H_0\). There is insufficient evidence to conclude that there is a significant correlation between pretest score and posttest score at the magnet school. Pretest score is not a useful linear predictor of posttest score for magnet school students.

(c)(ii) — Original School

From the regression output, the test statistic is \(t = 6.09\) with \(p\text{-value} = 0.000\).
Since \(0.000 < 0.05\), we reject \(H_0\). There is strong evidence of a significant correlation between pretest score and posttest score at the original school. Pretest score is a very strong linear predictor of posttest score for original school students.

(d)

The two-sample \(t\)-test in part (a) told us only that the magnet school group had a higher average improvement — but it didn’t explain who benefited or by how much depending on their initial ability. The regression analyses reveal something much more interesting:
• At the magnet school, the slope is nearly zero (\(0.1811\)), and \(R^2 = 2.5\%\) — this means students score high on the posttest regardless of how they did on the pretest. A student who entered the magnet school with a low pretest score of 64 scored 89 on the posttest (an improvement of 25 points), while a student with a higher pretest score of 86 actually dropped 2 points. The magnet school appears to level the playing field and disproportionately benefits students who start with lower ability.
• At the original school, the slope is close to 1 (\(0.9204\)) and \(R^2 = 78.8\%\) — students essentially maintained their relative ranking, with high pretest scorers also achieving high posttest scores. There is very little “boost” effect for any student regardless of their starting point.
In short, the regression analyses reveal that the magnet school benefits students with low pretest scores the most, while the original school produces predictable but modest gains proportional to where students started.

Question

A study was designed to explore subjects’ ability to judge the distance between two objects placed in a dimly lit room. The researcher suspected that the subjects would generally overestimate the distance between the objects in the room and that this overestimation would increase the farther apart the objects were.
The two objects were placed at random locations in the room before a subject estimated the distance (in feet) between those two objects. After each subject estimated the distance, the locations of the objects were rerandomized before the next subject viewed the room.
After data were collected for 40 subjects, two linear models were fit in an attempt to describe the relationship between the subjects’ perceived distances \((y)\) and the actual distance, in feet, between the two objects.
Model 1: \(\hat{y} = 0.238 + 1.080 \times (\text{actual distance})\)
The standard errors of the estimated coefficients for Model 1 are 0.260 and 0.118, respectively.
Model 2: \(\hat{y} = 1.102 \times (\text{actual distance})\)
The standard error of the estimated coefficient for Model 2 is 0.393.
(a) Provide an interpretation in context for the estimated slope in Model 1.
(b) Explain why the researcher might prefer Model 2 to Model 1 in this context.
(c) Using Model 2, test the researcher’s hypothesis that in dim light participants overestimate the distance, with the overestimate increasing as the actual distance increases. (Assume appropriate conditions for inference are met.)
The researchers also wanted to explore whether the performance on this task differed between subjects who wear contact lenses and subjects who do not wear contact lenses. A new variable was created to indicate whether or not a subject wears contact lenses. The data for this variable were coded numerically (\(1 = \text{contact wearer}\), \(0 = \text{noncontact wearer}\)), and this new variable, named “contact,” was included in the following model.
Model 3: \(\hat{y} = 1.05 \times (\text{actual distance}) + 0.12 \times (\text{contact}) \times (\text{actual distance})\)
The standard errors of the estimated coefficients for Model 3 are 0.357 and 0.032, respectively.
(d) Using Model 3, sketch the estimated regression model for contact wearers and the estimated regression model for noncontact wearers on the grid below.
(e) In the context of this study, provide an interpretation of the estimated coefficients for Model 3.

Most-appropriate topic codes (AP Statistics):

• Topic 5.3 — Linear Regression Models (Parts \(\mathrm{a}\), \(\mathrm{b}\))
• Topic 5.5 — Least-Squares Regression (Part \(\mathrm{b}\))
• Topic 5.3 — Linear Regression Models (Part \(\mathrm{c}\))
• Topic 5.2 — Correlation (Part \(\mathrm{c}\))
• Topic 5.3 — Linear Regression Models (Parts \(\mathrm{d}\), \(\mathrm{e}\))
▶️ Answer/Explanation

(a)

The estimated slope of \(1.080\) in Model 1 means that for each additional foot of actual distance between the two objects, the perceived (estimated) distance is expected to increase by about \(1.080\) feet on average.
In other words, as the objects are placed farther apart in reality, subjects in dim light tend to perceive the distance as growing slightly faster than the true distance — at a rate of roughly \(1.080\) feet of perceived distance per foot of actual distance.
\( \boxed{\text{For every 1 ft increase in actual distance, perceived distance increases by approximately } 1.080 \text{ ft on average.}}\)

(b)

Model 2 is preferable because it has a \(y\)-intercept of zero, which makes physical sense in this context — if two objects are placed in the same location (actual distance \(= 0\)), a subject would be expected to perceive a distance of zero feet, not \(0.238\) feet as Model 1 would predict.
Since the intercept in Model 1 is not meaningfully different from zero (it is small relative to its standard error of \(0.260\)), removing it and using the simpler through-the-origin Model 2 produces a more interpretable and contextually appropriate model.
\(\boxed{\text{Model 2 is preferred because a zero intercept is physically sensible when actual distance} = 0.}\)

(c)

Let \(\beta\) be the true slope of the linear relationship between perceived distance and actual distance in Model 2. The researcher’s hypothesis that subjects overestimate with the overestimation growing as distance grows is equivalent to \(\beta > 1\).
Step 1 — Hypotheses:
\(H_0: \beta = 1\) (perceived distance increases at the same rate as actual distance — no overestimation growth)
\(H_a: \beta > 1\) (perceived distance increases faster than actual distance — overestimation grows with distance)
Step 2 — Test Statistic:
We use a \(t\)-test for the slope:
\(t = \dfrac{b – \beta_0}{s_b} = \dfrac{1.102 – 1}{0.393} = \dfrac{0.102}{0.393} \approx 0.260\)
\(\text{degrees of freedom} = n – 1 = 40 – 1 = 39\)
\(p\text{-value} = P(t_{39} > 0.260) \approx 0.398\)
Step 3 — Conclusion:
Since the \(p\)-value of \(0.398\) is much greater than \(\alpha = 0.05\), we fail to reject \(H_0\). We do not have statistically significant evidence to conclude that subjects overestimate the distance with the overestimation increasing as the actual distance increases.
\(\boxed{t \approx 0.260,\quad p\text{-value} \approx 0.398 > 0.05 \Rightarrow \text{Fail to reject } H_0.}\)

(d)

Substituting the two values of the indicator variable into Model 3:
For contact wearers (\(\text{contact} = 1\)):
\(\hat{y} = 1.05(\text{actual distance}) + 0.12(1)(\text{actual distance}) = 1.17 \times (\text{actual distance})\)
For noncontact wearers (\(\text{contact} = 0\)):
\(\hat{y} = 1.05(\text{actual distance}) + 0.12(0)(\text{actual distance}) = 1.05 \times (\text{actual distance})\)
Both lines pass through the origin. The contact wearers’ line has a steeper slope (\(1.17\)) than the noncontact wearers’ line (\(1.05\)), as shown in the graph above.
\(\boxed{\text{Contact: } \hat{y} = 1.17x;\quad \text{Noncontact: } \hat{y} = 1.05x \text{ (both through origin)}}\)

(e)

The coefficient \(1.05\) estimates the average increase in perceived distance (in feet) for each one-foot increase in actual distance for noncontact wearers — that is, for every foot farther apart the objects actually are, a noncontact wearer perceives them as about \(1.05\) feet farther apart on average.
The coefficient \(0.12\) estimates the additional average increase in perceived distance (in feet) for each one-foot increase in actual distance specifically for contact wearers, above and beyond the \(1.05\) rate for noncontact wearers — so contact wearers perceive distance as growing at a rate of \(1.05 + 0.12 = 1.17\) feet per foot of actual distance on average.
Taken together, the model tells us that both groups overestimate distance in dim light, but contact wearers overestimate by a slightly greater amount per foot of actual distance than noncontact wearers do.
\(\boxed{1.05: \text{ per-foot slope for noncontact wearers};\quad 0.12: \text{ additional per-foot slope for contact wearers.}}\)

Question

Scientists interested in preserving natural habitats and minimizing the possible extinction of certain bird species conducted a study to determine if it is better for conservation groups to purchase a few large nature preserves or many small preserves in order to meet these goals.
The scientists studied 13 randomly selected islands of different sizes to determine the risk of extinction for bird species. Islands are thought to be a good imitation of what would happen in a nature preserve because of their isolation. If a species lived on only one island, it was considered to be at risk. Scientists have determined that whether or not one species becomes extinct is independent of whether or not another species becomes extinct.
In 1990 scientists counted the number of at-risk species on each of the selected islands. They returned to each of these islands in the year 2000 to see whether the species still existed on the islands. Species that were present in 1990 but absent in 2000 were considered extinct. Data collected by the scientists are given in the table below.
(a) One scientist involved in the study believes that large islands (those with areas greater than 25 square kilometers) are more effective than small islands (those with areas of no more than 25 square kilometers) for protecting at-risk species. The scientist noted that for this study, a total of 19 of the 208 species on the large islands became extinct, whereas a total of 66 of the 299 species on the small islands became extinct. Assume that the probability of extinction is the same for all at-risk species on large islands and the same for all at-risk species on small islands. Do these data support the scientist’s belief? Give appropriate statistical justification for your answer.
(b) Another scientist who worked on this study thinks that the proportion of species that become extinct is more directly related to the size of the islands than simply to whether the islands are grouped as large or small. This scientist investigated the relationship between the proportion of extinct birds and the area, in square kilometers, of islands. A least squares analysis was conducted on the proportion extinct and \(\ln(\text{area})\). The regression analysis output, the scatterplot, and the residual plot are shown below.
Estimate the slope of the least squares regression line using a 95 percent confidence interval. Interpret your answer in the context of this situation.
(c) In part (a), the scientist assumed that the probability of a species becoming extinct is the same for each of the large islands. Similarly, the scientist assumed that the probability is the same for each of the small islands. Based on your answer in part (b), do you think this is a reasonable assumption? Explain.
(d) A conservation group with a long-term goal of preserving species believes that all at-risk species will disappear whenever land inhabited by those species is developed. It has an opportunity to purchase land in an area about to be developed. The group has a choice of creating one large nature preserve with an area of 45 square kilometers and containing 70 at-risk species, or 5 small nature preserves, each with an area of 3 square kilometers and each containing 16 at-risk species unique to that preserve. Which choice would you recommend and why?

Most-appropriate topic codes (AP Statistics):

• Topic 3.12 — Setting Up a Test for the Difference Between Two Population Proportions (Part \(\mathrm{a}\))
• Topic 3.13 — Carrying Out a Test for the Difference Between Two Population Proportions (Part \(\mathrm{a}\))
• Topic 5.5 — Least-Squares Regression (Part \(\mathrm{b}\))
• Topic 5.4 — Residuals (Part \(\mathrm{c}\))
• Topic 5.5 — Least-Squares Regression (Part \(\mathrm{d}\)
▶️ Answer/Explanation

(a)
We want to test whether the proportion of species going extinct is smaller on large islands than on small islands. Let \(p_L\) be the true proportion of at-risk species that become extinct on large islands, and \(p_S\) be the true proportion on small islands.
The hypotheses are:
\(H_0: p_L – p_S = 0\)
\(H_a: p_L – p_S < 0\)
We use a two-sample \(z\)-test for the difference in proportions. The sample proportions are:
\(\hat{p}_L = \frac{19}{208} \approx 0.091, \qquad \hat{p}_S = \frac{66}{299} \approx 0.221\)
Check conditions — all expected counts must be at least 5:
\(n_L\hat{p}_L = 19,\quad n_L(1-\hat{p}_L) = 189,\quad n_S\hat{p}_S = 66,\quad n_S(1-\hat{p}_S) = 233\)
All values are well above 5, so we may proceed.
The pooled sample proportion is:
\(\hat{p} = \frac{19+66}{208+299} = \frac{85}{507} \approx 0.168\)
The test statistic is:
\(z = \frac{\hat{p}_L – \hat{p}_S}{\sqrt{\hat{p}(1-\hat{p})\left(\dfrac{1}{n_L}+\dfrac{1}{n_S}\right)}} = \frac{0.091 – 0.221}{\sqrt{(0.168)(0.832)\left(\dfrac{1}{208}+\dfrac{1}{299}\right)}} = \frac{-0.130}{0.034} \approx -3.82\)
The corresponding \(p\)-value \(\approx 0.00006\), which is essentially \(0\).
Since the \(p\)-value is far less than any reasonable significance level, we reject \(H_0\). There is very strong statistical evidence that the proportion of species going extinct is smaller for large islands than for small islands, supporting the scientist’s belief.
\(\boxed{z \approx -3.82, \quad p\text{-value} \approx 0.00006 \quad \Rightarrow \quad \text{Reject } H_0}\)

(b)
We construct a 95% confidence interval for the slope \(\beta\) of the regression of proportion extinct on \(\ln(\text{area})\).
From the regression output: \(\hat{b} = -0.05323\), \(SE_b = 0.00618\), and \(df = n – 2 = 13 – 2 = 11\).
The critical value from the \(t\)-table with \(df = 11\) at the 95% level is \(t^* = 2.201\).
The confidence interval is:
\(\hat{b} \pm t^* \cdot SE_b = -0.05323 \pm 2.201(0.00618)\)
\(-0.05323 \pm 0.01360\)
\(\boxed{(-0.0668,\ -0.0396)}\)
We are 95% confident that for every 1-unit increase in \(\ln(\text{area})\), the mean proportion of species going extinct decreases by somewhere between \(0.0396\) and \(0.0668\). In plain terms, larger islands are associated with a meaningfully lower extinction rate, and this relationship is statistically significant.

(c)
The assumption is not reasonable. The regression analysis in part (b) shows that the proportion of species going extinct decreases steadily as island area increases — it is a continuous relationship, not a step function that jumps only between “large” and “small” groups.
Within the large island group, areas ranged from 31 to 46 sq km, and within the small island group, areas ranged from 1 to 9 sq km — meaning extinction probabilities varied considerably within each group as well.
Because extinction probability depends on actual area and not just on a binary large/small classification, the assumption that all large islands share one common extinction probability and all small islands share another is not supported by the data.

(d)
We use the regression model \(\widehat{\text{prop extinct}} = 0.28996 – 0.05323\ln(\text{area})\) to estimate extinction proportions for each option.
Option 1 — One large preserve (area = 45 sq km, 70 species):
\(\widehat{\text{prop extinct}} = 0.28996 – 0.05323\ln(45) = 0.28996 – 0.05323(3.807) \approx 0.28996 – 0.20261 \approx 0.0873\)
Expected extinctions: \(70 \times 0.0873 \approx 6.1\) species
Expected survivors: \(70 – 6.1 \approx \mathbf{63.9 \approx 64}\) species

Option 2 — Five small preserves (each area = 3 sq km, 16 species each; 80 total):
\(\widehat{\text{prop extinct}} = 0.28996 – 0.05323\ln(3) = 0.28996 – 0.05323(1.099) \approx 0.28996 – 0.05850 \approx 0.2315\)
Expected extinctions per preserve: \(16 \times 0.2315 \approx 3.7\) species
Total expected extinctions: \(5 \times 3.7 \approx 18.5\) species
Expected survivors: \(80 – 18.5 \approx \mathbf{61.5 \approx 62}\) species

The one large preserve is expected to save approximately 64 species, compared to about 62 species across the five small preserves. We recommend creating one large nature preserve, as it leads to a greater expected number of surviving species, and larger areas have been shown to have substantially lower extinction rates per species.
\(\boxed{\text{Recommend: One large preserve (45 sq km)} \Rightarrow \approx 64 \text{ species saved vs. } \approx 62 \text{ for five small preserves}}\)

Question

A manufacturer of dish detergent believes the height of soapsuds in the dishpan depends on the amount of detergent used. A study of the suds’ heights for a new dish detergent was conducted. Seven pans of water were prepared. All pans were of the same size and type and contained the same amount of water. The temperature of the water was the same for each pan. An amount of dish detergent was assigned at random to each pan, and that amount of detergent was added to the pan. Then the water in the dishpan was agitated for a set amount of time, and the height of the resulting suds was measured.
A plot of the data and the computer output from fitting a least squares regression line to the data are shown below.

(a) Write the equation of the fitted regression line. Define any variables used in this equation.
(b) Note that \(s = 1.99821\) in the computer output. Interpret this value in the context of this study.
(c) Identify and interpret the standard error of the slope.

Most-appropriate topic codes (AP Statistics):

• Topic 5.3 — Linear Regression Models (Part \(\mathrm{a}\))
• Topic 5.4 — Residuals (Part \(\mathrm{b}\))
• Topic 5.5 — Least-Squares Regression (Part \(\mathrm{c}\))
▶️ Answer/Explanation

(a)

Reading the coefficients directly from the computer output, the fitted regression line is:
\(\hat{y} = -2.679 + 9.5x\)
where \(\hat{y}\) represents the predicted (estimated) mean height of the soapsuds (in millimeters), and \(x\) represents the amount of detergent added to the pan (in grams).
\(\boxed{\hat{y} = -2.679 + 9.5x}\)

(b)

The value \(s = 1.99821\,\text{mm}\) is the standard deviation of the residuals.
In the context of this study, it measures a typical amount of variation in the observed heights of soapsuds from the heights predicted by the regression line — that is, for a given amount of detergent, the actual suds height typically differs from the predicted suds height by about \(1.998\,\text{mm}\).
\(\boxed{s = 1.99821\,\text{mm}}\)

(c)

The standard error of the slope is identified from the computer output as the SE Coef for the Amount row:
\(SE_b = 0.7553\,\text{mm per gram}\)
This value estimates the standard deviation of the sampling distribution of the estimated slope — in other words, it tells us how much the estimated slope \(\hat{b}_1\) would be expected to vary from experiment to experiment if the same study were repeated many times under identical conditions.
A small \(SE_b = 0.7553\) relative to the slope of \(9.5\) indicates that the estimated slope is very stable and reliable across repeated samples.
\(\boxed{SE_b = 0.7553\,\text{mm/g}}\)

Question

John believes that as he increases his walking speed, his pulse rate will increase. He wants to model this relationship. John records his pulse rate, in beats per minute (bpm), while walking at each of seven different speeds, in miles per hour (mph). A scatterplot and regression output are shown below.
(a) Using the regression output, write the equation of the fitted regression line.
(b) Do your estimates of the slope and intercept parameters have meaningful interpretations in the context of this question? If so, provide interpretations in this context. If not, explain why not.
(c) John wants to provide a 98 percent confidence interval for the slope parameter in his final report. Compute the margin of error that John should use. Assume that conditions for inference are satisfied.

Most-appropriate topic codes (AP Statistics):

• Topic 5.3 — Linear Regression Models (Parts a, b)
• Topic 5.5 — Least-Squares Regression (Parts a, b)
• Topic 5.4 — Residuals (Regression output interpretation)
▶️ Answer/Explanation

(a)
Reading directly from the regression output, the fitted regression equation is: \[ \widehat{\text{Pulse}} = 63.457 + 16.2809 \times \text{Speed} \] where Pulse is measured in beats per minute (bpm) and Speed is measured in miles per hour (mph).

(b)
Both estimates have meaningful interpretations in this context.
Slope interpretation:
The slope \(b_1 = 16.2809\) bpm/mph means that for each additional 1 mile per hour increase in John’s walking speed, his predicted pulse rate increases by approximately 16.28 beats per minute on average.
Intercept interpretation:
The intercept \(b_0 = 63.457\) bpm means that when John’s walking speed is 0 mph — that is, when he is standing still — his predicted pulse rate is approximately 63.5 beats per minute. This is a reasonable estimate of John’s resting pulse rate, so the intercept does carry a meaningful real-world interpretation here (unlike many regression contexts where the intercept falls outside the range of observed data).

(c)
The margin of error for a confidence interval for the slope is:
\( \text{Margin of error} = t^* \times SE_{b_1} \)
From the regression output, the standard error of the slope is \(SE_{b_1} = 0.8192\).
For a 98% confidence interval, the degrees of freedom are \(df = n – 2 = 7 – 2 = 5\).
From the \(t\)-table with \(df = 5\) and a 98% confidence level (tail probability \(= 0.01\)):
\( t^* = 3.365 \)
Therefore, the margin of error is:
\( \text{Margin of error} = 3.365 \times 0.8192 \)
\( \boxed{\text{Margin of error} \approx 2.757 \text{ bpm/mph}} \)
This means John’s 98% confidence interval for the true slope would be \(16.2809 \pm 2.757\), or approximately \((13.524,\ 19.038)\) bpm per mph. The relatively narrow interval, combined with the very high \(R^2 = 98.7\%\), confirms that speed is an extremely strong linear predictor of John’s pulse rate.

Question

A simple random sample of 9 students was selected from a large university. Each of these students reported the number of hours he or she had allocated to studying and the number of hours allocated to work each week. A least squares linear regression was performed and part of the resulting computer output is shown below.
The scatterplot below displays the data that were collected from the 9 students.
(a) After point \(P\), labeled on the graph on the previous page, was removed from the data, a second linear regression was performed and the computer output is shown below.
Does point \(P\) exercise a large influence on the regression line? Explain.
(b) The researcher who conducted the study discovered that the number of hours spent studying reported by the student represented by \(P\) was recorded incorrectly. The corrected data point for this student is represented by the letter \(Q\) in the scatterplot below.
Study Work 5 10 15 0 10 20 30 Q•
Explain how the least squares regression line for the corrected data (in this part) would differ from the least squares regression line for the original data.

Most-appropriate topic codes (AP Statistics):

• Topic 5.3 — Linear Regression Models (Part a)
• Topic 5.4 — Residuals (Part a)
• Topic 5.5 — Least-Squares Regression (Parts a, b)
• Topic 5.1 — Graphical Representations Between Two Quantitative Variables (Part b)
▶️ Answer/Explanation

(a)

Yes, point \(P\) does exercise a large influence on the regression line.
When point \(P\) is included, the slope of the regression line is \(b_1 = 0.4919\), and it is statistically significant (\(p = 0.040\)).
When point \(P\) is removed, the slope drops dramatically to \(b_1 = 0.1500\), which is no longer statistically significant (\(p = 0.709\)).
The intercept also changes from \(8.107\) to \(11.123\), and \(R^2\) falls sharply from \(47.6\%\) to only \(2.5\%\).
Because removing a single point caused such substantial changes in the slope, intercept, statistical significance, and \(R^2\), point \(P\) is clearly an influential point — it is extreme in the \(x\)-direction (a work value of about 30, far beyond the rest of the data), which gives it high leverage and strong pull on the regression line.

(b)

In the original data, point \(P\) is at approximately \((\text{Work} = 30,\ \text{Study} = 25)\), which is high on both variables and pulls the regression line upward to the right, producing a positive slope.
The corrected data point \(Q\) is at approximately \((\text{Work} = 30,\ \text{Study} = 6)\), which is far below the trend of the other data points at that work value.
With \(Q\) replacing \(P\), the regression line for the corrected data will have a negative slope rather than a positive slope — the corrected point now pulls the line downward on the right side.
The intercept would also be considerably larger for the corrected data, since the line must start higher on the \(y\)-axis to accommodate the negative slope while passing near the rest of the data.

Scroll to Top