AP Statistics 4.6 Sampling Distributions for the Difference Between Two Sample Means Study Notes - New Syllabus
AP Statistics 4.6 Sampling Distribution for the Difference Between Two Sample Means Study Notes – New Syllabus
AP Statistics 4.6 Sampling Distribution for the Difference Between Two Sample Means Study Notes – As per latest AP Statistics Syllabus.
LEARNING OBJECTIVES
- 4.6.A Calculate the mean and standard deviation of a sampling distribution for the difference between two sample means.
- 4.6.B Justify the appropriateness of conditions for the sampling distribution of the difference between two sample means.
- 4.6.C Interpret the mean, standard deviation, and probabilities for a sampling distribution for the difference between sample means.
ESSENTIAL KNOWLEDGE:
- 4.6.A.1 For two independent populations with population means \( \mu_1 \) and \( \mu_2 \) and population standard deviations \( \sigma_1 \) and \( \sigma_2 \), when the sampled values are independent, the sampling distribution of the difference in sample means \( \bar{x}_1-\bar{x}_2 \) has a mean
\( \mu_{(\bar{x}_1-\bar{x}_2)}=\mu_1-\mu_2 \)
and a standard deviation
\( \sigma_{(\bar{x}_1-\bar{x}_2)}=\sqrt{\dfrac{\sigma_1^2}{n_1}+\dfrac{\sigma_2^2}{n_2}} \) - 4.6.B.1 Sampling without replacement requires that two conditions must be met:
- 4.6.B.1.i The randomization condition—the data should be collected using two independent random samples.
- 4.6.B.1.ii The 10% condition—the size of each sample should be less than or equal to 10% of the respective population size: \( n_1\le10\%N_1 \) and \( n_2\le10\%N_2 \), where \( N_1 \) is the size of population 1, \( N_2 \) is the size of population 2, and the sample sizes are represented as \( n_1 \) and \( n_2 \).
- 4.6.B.2 If the data come from an experiment, the data only need to meet the randomization condition. The treatments must be randomly assigned to experimental units to meet the randomization condition.
- 4.6.B.3 The sampling distribution for the difference between sample means, \( \bar{x}_1-\bar{x}_2 \), can be modeled with a normal distribution if the two population distributions can each be modeled by a normal distribution.
- 4.6.B.4 The sampling distribution for the difference between sample means, \( \bar{x}_1-\bar{x}_2 \), can be modeled approximately by a normal distribution if the two population distributions cannot be modeled by a normal distribution but \( n_1\ge30 \) and \( n_2\ge30 \).
- 4.6.C.1 The mean, standard deviation, and probabilities for a sampling distribution for the difference between sample means should be interpreted within the context of specific populations.
4.6.A.1 Mean and Standard Deviation of the Sampling Distribution of the Difference Between Two Sample Means
When comparing the means of two independent populations, statisticians often examine the difference between the two sample means.

If independent random samples are selected from two populations, the statistic
\( \bar{x}_1-\bar{x}_2 \)
has its own sampling distribution.
This sampling distribution describes the values that the difference between the two sample means would take from repeated random sampling.
Provided the sampled values are independent, the sampling distribution has a predictable mean and standard deviation.
Mean of the Sampling Distribution
The mean of the sampling distribution of the difference between two sample means equals the difference between the two population means.
Formula
\( \mu_{\left(\bar{x}_1-\bar{x}_2\right)}=\mu_1-\mu_2 \)
Where:
- \( \mu_1 \) = Population mean of Population 1
- \( \mu_2 \) = Population mean of Population 2
- \( \mu_{\left(\bar{x}_1-\bar{x}_2\right)} \) = Mean of the sampling distribution of the difference between sample means
This means that the average difference between the sample means equals the true difference between the population means.
Standard Deviation of the Sampling Distribution
The standard deviation of the sampling distribution measures the variability in the differences between sample means.
Formula
\( \sigma_{\left(\bar{x}_1-\bar{x}_2\right)}=\sqrt{\dfrac{\sigma_1^2}{n_1}+\dfrac{\sigma_2^2}{n_2}} \)
Where:
- \( \sigma_1 \) = Population standard deviation of Population 1
- \( \sigma_2 \) = Population standard deviation of Population 2
- \( n_1 \) = Sample size from Population 1
- \( n_2 \) = Sample size from Population 2
- \( \sigma_{\left(\bar{x}_1-\bar{x}_2\right)} \) = Standard deviation of the sampling distribution
Conditions
- The two samples must be independent.
- The observations within each sample must be independent.
- The formula applies to independent random samples from two populations.
| Statistic | Formula | Interpretation |
|---|---|---|
| Mean | \( \mu_{\left(\bar{x}_1-\bar{x}_2\right)}=\mu_1-\mu_2 \) | The center of the sampling distribution equals the true difference between the population means. |
| Standard Deviation | \( \sigma_{\left(\bar{x}_1-\bar{x}_2\right)}=\sqrt{\dfrac{\sigma_1^2}{n_1}+\dfrac{\sigma_2^2}{n_2}} \) | Measures the variability of the differences between sample means. |
Important AP Exam Notes
- The two samples must be independent; otherwise, these formulas cannot be used.
- The mean of the sampling distribution is simply the difference between the two population means.
- The standard deviation depends on both population standard deviations and both sample sizes.
- Increasing either sample size decreases the standard deviation of the sampling distribution.
- Always subtract the population means in the same order as the sample means (\(\bar{x}_1-\bar{x}_2\)).
Example
Two independent populations have the following characteristics:
- Population 1: \( \mu_1=80,\; \sigma_1=12,\; n_1=36 \)
- Population 2: \( \mu_2=74,\; \sigma_2=15,\; n_2=25 \)
Calculate the mean and standard deviation of the sampling distribution of
\( \bar{x}_1-\bar{x}_2 \)
▶️ Answer / Explanation
Step 1: Calculate the mean.
\( \mu_{\left(\bar{x}_1-\bar{x}_2\right)}=\mu_1-\mu_2 \)
\( =80-74=6 \)
Step 2: Calculate the standard deviation.
\( \sigma_{\left(\bar{x}_1-\bar{x}_2\right)}=\sqrt{\dfrac{12^2}{36}+\dfrac{15^2}{25}} \)
\( =\sqrt{\dfrac{144}{36}+\dfrac{225}{25}} \)
\( =\sqrt{4+9} \)
\( =\sqrt{13}\approx3.61 \)
4.6.B.1 Conditions for the Sampling Distribution of the Difference Between Two Sample Means
Before using the sampling distribution of the difference between two sample means, statisticians must verify that the required conditions are satisfied.

When sampling is conducted without replacement, two important conditions must be met to ensure that the observations are independent and that the sampling distribution is valid.
These conditions are:
- Randomization Condition
- 10% Condition
4.6.B.1.i Randomization Condition
The data should be collected using two independent random samples, one from each population.
Random sampling helps eliminate bias and allows the results to be generalized to the populations.
The two samples must also be independent, meaning that the selection of individuals in one sample does not affect the selection of individuals in the other sample.
How to Verify
- The problem states that a simple random sample (SRS) was selected from Population 1.
- The problem states that a simple random sample (SRS) was selected from Population 2.
- The two samples were selected independently of one another.
4.6.B.1.ii 10% Condition
When sampling is performed without replacement, each sample size must be no more than 10% of its respective population size.
This condition ensures that observations within each sample can be treated as approximately independent.
Formulas
\( N_1 \ge 10n_1 \)
\( N_2 \ge 10n_2 \)
or equivalently
\( n_1 \le 0.10N_1 \)
\( n_2 \le 0.10N_2 \)
Where:
- \(N_1\) = Population size of Population 1
- \(N_2\) = Population size of Population 2
- \(n_1\) = Sample size from Population 1
- \(n_2\) = Sample size from Population 2
| Condition | Requirement | Purpose |
|---|---|---|
| Randomization Condition | Two independent random samples | Reduces bias and supports inference to the populations. |
| 10% Condition | \(N_1\ge10n_1\) and \(N_2\ge10n_2\) | Ensures approximate independence within each sample. |
Important AP Exam Notes
- Both samples must be selected independently.
- The Randomization Condition requires two independent random samples, not just one.
- The 10% Condition must be checked separately for each population.
- The 10% Condition applies only when sampling is done without replacement.
- Failure to verify these conditions may make statistical inference invalid.
Example
A researcher wants to compare the average mathematics test scores of students from two different schools.
A random sample of 40 students is selected from School A, which has 1,000 students.
A separate random sample of 35 students is selected from School B, which has 800 students.
Determine whether the Randomization Condition and the 10% Condition are satisfied.
▶️ Answer / Explanation
Step 1: Randomization Condition
The problem states that separate random samples were selected from both schools.
The two samples were selected independently.
✔ The Randomization Condition is satisfied.
Step 2: 10% Condition
School A
\(10n_1=10(40)=400\)
\(N_1=1000\)
Since
\(1000\ge400\),
✔ The condition is satisfied for School A.
School B
\(10n_2=10(35)=350\)
\(N_2=800\)
Since
\(800\ge350\),
✔ The condition is satisfied for School B.
Conclusion
Both the Randomization Condition and the 10% Condition are satisfied, so these conditions for using the sampling distribution of the difference between two sample means have been met.
4.6.B.2 Conditions for Experiments
The conditions for the sampling distribution of the difference between two sample means depend on whether the data come from an observational study or a randomized experiment.
For an experiment, only the Randomization Condition must be verified.
The treatments must be randomly assigned to the experimental units.
Random assignment helps create comparable treatment groups and reduces the effect of confounding variables.
Unlike random sampling, the 10% Condition is not required for experiments because participants are assigned to treatments rather than sampled without replacement from a population.
Randomization Condition for Experiments
The experiment must use random assignment of treatments to experimental units.
This ensures that the treatment groups are similar before the treatment is applied and allows researchers to make valid comparisons between the groups.
How to Verify
- The problem states that subjects were randomly assigned to treatments.
- The treatments were assigned using a random process, such as a random number generator or random selection.
| Study Type | Required Conditions |
|---|---|
| Observational Study | Randomization Condition + 10% Condition |
| Randomized Experiment | Randomization Condition only (random assignment) |
Random Sampling vs. Random Assignment
| Random Sampling | Random Assignment |
|---|---|
| Selects individuals from a population. | Assigns individuals to treatment groups. |
| Supports inference to the population. | Supports cause-and-effect conclusions. |
| Used in observational studies. | Used in experiments. |
Important AP Exam Notes
- For an experiment, verify that treatments were randomly assigned.
- The 10% Condition is not required for randomized experiments.
- Random assignment helps control for confounding variables and allows researchers to establish cause-and-effect relationships.
- Do not confuse random sampling with random assignment; they serve different purposes.
- On the AP Exam, always identify whether the study is an observational study or a randomized experiment before checking conditions.
Example
A researcher wants to compare the effectiveness of two study methods.
One hundred students volunteer to participate in the experiment. Each student is randomly assigned to either Study Method A or Study Method B.
Determine whether the required condition for comparing the two population means has been satisfied.
▶️ Answer / Explanation
Step 1: Identify the type of study.
This is a randomized experiment because students are assigned to treatment groups.
Step 2: Verify the condition.
The problem states that students were randomly assigned to the two study methods.
✔ The Randomization Condition is satisfied.
Step 3: Check whether the 10% Condition is needed.
Because this is a randomized experiment, the 10% Condition does not apply.
Conclusion
The required condition for comparing the two population means in this experiment has been satisfied because the treatments were randomly assigned.
4.6.B.3 Normal Sampling Distribution When Both Populations Are Normally Distributed
To make probability calculations or perform inference for the difference between two population means, the sampling distribution of
\( \bar{x}_1-\bar{x}_2 \)
should be approximately normally distributed.
If both population distributions can be modeled by a normal distribution, then the sampling distribution of the difference between the two sample means is also normally distributed, regardless of the sample sizes.
This result is true as long as the two samples are independent.
Condition for a Normal Sampling Distribution
If
- Population 1 is approximately normally distributed, and
- Population 2 is approximately normally distributed,
then
\( \bar{x}_1-\bar{x}_2 \)
has a normal sampling distribution, regardless of the values of
\( n_1 \) and \( n_2 \).
| Population Distribution | Sampling Distribution of \( \bar{x}_1-\bar{x}_2 \) |
|---|---|
| Both populations are approximately normal. | Exactly (or approximately) normal for any sample sizes. |
Why This Works
When both populations are normally distributed, every sample mean follows a normal distribution.
The difference between two normally distributed sample means is also normally distributed.
Therefore, the sampling distribution of
\( \bar{x}_1-\bar{x}_2 \)
is normal.
Important AP Exam Notes
- Both populations must be approximately normally distributed.
- The samples must be independent.
- When both populations are normal, there is no minimum sample size requirement.
- This condition allows probability calculations and inference using the normal model for the sampling distribution.
- Always verify the population distributions before using this result.
Example
A researcher compares the mean resting heart rates of two different species of birds.
The resting heart rates for both populations are known to be approximately normally distributed.
Independent random samples of \( n_1=12 \) and \( n_2=15 \) are selected.
Can the sampling distribution of \( \bar{x}_1-\bar{x}_2 \) be modeled using a normal distribution?
▶️ Answer / Explanation
Step 1: Check the population distributions.
The problem states that both populations are approximately normally distributed.
✔ This condition is satisfied.
Step 2: Check independence.
The samples were selected independently.
✔ The independence condition is satisfied.
Step 3: Draw a conclusion.
Since both population distributions are approximately normal, the sampling distribution of
\( \bar{x}_1-\bar{x}_2 \)
can be modeled using a normal distribution, even though the sample sizes are less than 30.
4.6.B.4 Approximate Normal Sampling Distribution Using the Central Limit Theorem
Sometimes the two population distributions are not normally distributed.
Even if this is the case, the sampling distribution of the difference between two sample means can still be modeled using a normal distribution if the sample sizes are sufficiently large.
This result follows from the Central Limit Theorem (CLT).
If the samples are independent and both sample sizes are large, the distribution of
\( \bar{x}_1-\bar{x}_2 \)
is approximately normal, even when the original population distributions are not normal.
Condition for an Approximate Normal Sampling Distribution
If the population distributions are not normally distributed, then the sampling distribution of
\( \bar{x}_1-\bar{x}_2 \)
can be modeled approximately by a normal distribution provided that
\( n_1 \ge 30 \) and \( n_2 \ge 30 \)
where
- \( n_1 \) = Sample size from Population 1
- \( n_2 \) = Sample size from Population 2
| Population Distribution | Sample Sizes | Sampling Distribution of \( \bar{x}_1-\bar{x}_2 \) |
|---|---|---|
| Not approximately normal | \( n_1 \ge 30,\; n_2 \ge 30 \) | Approximately normal (Central Limit Theorem) |
| Not approximately normal | One or both sample sizes less than 30 | Cannot automatically assume a normal sampling distribution. |
Why This Works
The Central Limit Theorem states that as the sample size increases, the sampling distribution of the sample mean becomes approximately normal, regardless of the shape of the population distribution.

Since this applies to both sample means, the difference
\( \bar{x}_1-\bar{x}_2 \)
is also approximately normally distributed when
\( n_1 \ge 30 \) and \( n_2 \ge 30 \).
Important AP Exam Notes
- If the populations are not normally distributed, check that both sample sizes satisfy
\( n_1 \ge 30 \) and \( n_2 \ge 30 \).
- The two samples must be independent.
- The Central Limit Theorem applies to both sample means, allowing the difference between them to be approximately normal.
- If one sample size is less than 30 and the corresponding population is not approximately normal, the normal model may not be appropriate.
- Always verify the sample sizes before using the normal model for the sampling distribution.
Example
A researcher compares the average daily screen time of students from two schools.
The distributions of screen time in both populations are strongly right-skewed.
Independent random samples are selected with \( n_1=40 \) and \( n_2=35 \).
Can the sampling distribution of \( \bar{x}_1-\bar{x}_2 \) be modeled using a normal distribution?
▶️ Answer / Explanation
Step 1: Check the population distributions.
The populations are not normally distributed.
Step 2: Check the sample sizes.
\( n_1=40\ge30 \)
\( n_2=35\ge30 \)
Both sample sizes are at least 30.
Step 3: Apply the Central Limit Theorem.
Because both sample sizes are at least 30 and the samples are independent, the sampling distribution of
\( \bar{x}_1-\bar{x}_2 \)
can be modeled using an approximately normal distribution.
4.6.C.1 Interpreting the Mean, Standard Deviation, and Probabilities for the Sampling Distribution of the Difference Between Two Sample Means
Once the sampling distribution of the difference between two sample means,

\( \bar{x}_1-\bar{x}_2 \)
has been established, its mean, standard deviation, and probabilities should always be interpreted in the context of the two populations being compared.
These values describe the behavior of the sampling distribution, not the original populations.
Interpreting the Mean
The mean of the sampling distribution is
\( \mu_{\left(\bar{x}_1-\bar{x}_2\right)}=\mu_1-\mu_2 \)
This value represents the expected difference between the sample means obtained from repeated random sampling.
Interpretation Template
On average, the difference between the sample means (\(\bar{x}_1-\bar{x}_2\)) is expected to be equal to the difference between the two population means.
Interpreting the Standard Deviation
The standard deviation of the sampling distribution is
\( \sigma_{\left(\bar{x}_1-\bar{x}_2\right)}=\sqrt{\frac{\sigma_1^2}{n_1}+\frac{\sigma_2^2}{n_2}} \)
The standard deviation describes the typical amount of variability in the differences between sample means from repeated random samples.
Interpretation Template
The differences between the sample means typically vary by approximately the value of the standard deviation from the true difference between the population means.
Interpreting Probabilities
Probabilities calculated from the sampling distribution describe the likelihood of obtaining a difference between two sample means that satisfies a particular condition.
These probabilities refer to the results of repeated random sampling, not to individual observations.
Interpretation Template
The probability represents the chance that the difference between the sample means from two independent random samples satisfies the stated condition.
| Statistic | Interpretation |
|---|---|
| Mean | The expected difference between the sample means equals the difference between the two population means. |
| Standard Deviation | Measures the typical variability of the differences between sample means from repeated random samples. |
| Probability | Represents the likelihood of obtaining a particular difference between sample means in repeated random sampling. |
Important AP Exam Notes
- Interpret all results in terms of the difference between two population means.
- The sampling distribution describes the behavior of \( \bar{x}_1-\bar{x}_2 \), not individual observations.
- The mean represents the expected value of the difference between sample means.
- The standard deviation measures sampling variability.
- Probability statements refer to repeated random samples, not to the probability that a population parameter changes.
Example
A researcher compares the average mathematics test scores of students from two schools.
Independent random samples are selected from both schools.
The sampling distribution of \( \bar{x}_1-\bar{x}_2 \) has \( \mu_{\left(\bar{x}_1-\bar{x}_2\right)}=5 \) and \( \sigma_{\left(\bar{x}_1-\bar{x}_2\right)}=2 \).
In addition, \( P(\bar{x}_1-\bar{x}_2>8)=0.067 \).
Interpret the mean, standard deviation, and probability.
▶️ Answer / Explanation
Mean
On average, the difference between the sample mean mathematics test scores (School 1 − School 2) is expected to be 5 points, which equals the true difference between the population means.
Standard Deviation
The differences between the sample mean mathematics test scores typically vary by about 2 points from the true difference between the population means due to random sampling variability.
Probability
The probability of obtaining a difference in sample mean mathematics test scores greater than 8 points from two independent random samples is 0.067, or about 6.7%.
