Home / AP Statistics 1.9 Comparisons of the Distributions for One Quantitative Variable Study Notes

AP Statistics 1.9 Comparisons of the Distributions for One Quantitative Variable Study Notes - New Syllabus

AP Statistics 1.9 Comparing Quantitative Distributions and Z-Scores Study Notes – New Syllabus

AP Statistics 1.9 Comparing Quantitative Distributions and Z-Scores Study Notes – As per latest AP Statistics Syllabus.

LEARNING OBJECTIVES

  • 1.9.A Compare multiple quantitative one-variable graphical representations.
  • 1.9.B Compare multiple quantitative one-variable graphical representations of summary statistics.
  • 1.9.C Justify a claim using multiple quantitative one-variable graphical representations.
  • 1.9.D Calculate z-scores with population parameters.
  • 1.9.E Compare z-scores as measures of relative position for distributions.

ESSENTIAL KNOWLEDGE:

  • 1.9.A.1 Graphical representations of a quantitative variable can be used to compare important features between two or more distributions of the same quantitative variable. Histograms, back-to-back stem-and-leaf plots, and dotplots may be used to compare center, variability, shape, outliers, clusters, or gaps in two or more distributions. Boxplots may be used to compare center, variability, outliers, and skewness (or symmetry).
  • 1.9.B.1 A comparison of graphical representations for two or more distributions can include any of the numerical summaries (e.g., mean, standard deviation, etc.).
  • 1.9.C.1 Multiple quantitative one-variable graphical representations may reveal information that can be used to justify claims about the variable in context.
  • 1.9.D.1 A standardized score measures the number of standard deviations a data value falls above or below the mean.
  • 1.9.D.2 A z-score is calculated as

    \( z=\dfrac{x_i-\mu}{\sigma} \)

    where \(x_i\) is the data value, \(\mu\) is the population mean, and \(\sigma\) is the population standard deviation. A z-score measures how many standard deviations a data value is above (positive z-score) or below (negative z-score) the mean. When the population mean and standard deviation are unknown, the sample mean and standard deviation may be used to determine a z-score.
  • 1.9.E.1 z-scores may be used to compare relative positions of individual values within a distribution or between distributions.

AP Statistics – Concise Summary Notes – All Topics

1.9.A.1 Comparing Quantitative Distributions Using Graphical Representations

Graphical representations reveal important characteristics that help compare distributions.

Different types of graphs provide different information.

Histograms, dotplots, and stem-and-leaf plots are especially useful for examining:     

  • Center
  • Spread
  • Shape
  • Outliers
  • Clusters
  • Gaps

Boxplots are especially useful for comparing:

  • Center
  • Spread
  • Outliers
  • Skewness or Symmetry

Comparing Center

The center represents the typical value of a distribution.

When comparing graphs:

  • A distribution located farther to the right generally has a larger center.
  • A distribution located farther to the left generally has a smaller center.

Measures often used:

  • Mean
  • Median

Comparing Variability (Spread)

Spread describes how dispersed the data values are.

  • A distribution with a wider spread has greater variability.
  • A distribution with a narrower spread has less variability.

Measures often used:

  • Range
  • IQR
  • Standard Deviation

Comparing Shape

The shape of a distribution may be:

  • Symmetric
  • Skewed Right
  • Skewed Left
  • Uniform
  • Unimodal
  • Bimodal

Histograms and dotplots are particularly useful for identifying shape.

Comparing Outliers

Outliers appear as observations that are unusually far from the rest of the data.

Outliers may affect:

  • Mean
  • Range
  • Standard Deviation

When comparing distributions, mention any unusual observations that appear.

Comparing Clusters and Gaps

Some distributions contain:

  • Clusters (groups of observations)
  • Gaps (intervals with no observations)

These features may indicate different subgroups within the data.

Graph TypeFeatures Best Compared
HistogramCenter, Spread, Shape, Clusters, Gaps, Outliers
DotplotCenter, Spread, Shape, Clusters, Gaps, Outliers
Back-to-Back Stem-and-Leaf PlotCenter, Spread, Shape, Outliers
BoxplotCenter, Spread, Outliers, Skewness

AP Statistics Comparison Template

When comparing two quantitative distributions, use the following order:

Center → Spread → Shape → Outliers

This structure appears frequently on AP Free-Response Questions.

A strong comparison should describe:

  • Which distribution has the larger center
  • Which distribution has greater variability
  • How the shapes differ
  • Whether outliers are present

Example

Two boxplots summarize test scores from two classes.

Class A:

\( \mathrm{Median=85,\ IQR=10} \)

Class B:

\( \mathrm{Median=78,\ IQR=18} \)

No outliers are shown on either boxplot.

Compare the two distributions.

▶️ Answer / Explanation

Center:

Class A has the larger median. Since: \( \mathrm{85>78} \)

Class A generally achieved higher test scores.

Spread:

Class B has the larger IQR.

Since:

\( \mathrm{18>10} \)

Class B’s scores are more variable.

Shape:

The shape cannot be determined from the information provided.

A complete boxplot would be needed to evaluate skewness or symmetry.

Outliers:

No outliers are present in either distribution.

Overall, Class A performed better on average, while Class B showed greater variability in scores.

Example

Two histograms display the daily exercise times of two groups of adults.

Group A has a roughly symmetric distribution centered around:

\( \mathrm{45} \) minutes.

Group B has a right-skewed distribution centered around:

\( \mathrm{35} \) minutes with several large values above: \( \mathrm{90} \) minutes.

Compare the two distributions.

▶️ Answer / Explanation

Center:

Group A has the larger center because its typical exercise time is around: \( \mathrm{45} \) minutes compared with: \( \mathrm{35} \) minutes for Group B.

Spread:

Group B appears to have greater variability because its values extend much farther to the right.

Shape:

Group A is approximately symmetric.

Group B is skewed right.

Outliers:

The unusually large exercise times above:

\( \mathrm{90} \)

minutes may be potential outliers in Group B.

Therefore, Group A has a higher typical exercise time, while Group B is more variable and more strongly skewed.


1.9.B Compare Multiple Quantitative One-Variable Graphical Representations of Summary Statistics

Graphical representations and summary statistics work together to describe quantitative distributions.

While graphs provide a visual display of the data, numerical summaries provide precise measurements that help compare distributions.

When comparing two or more distributions, statisticians often use:

  • Mean
  • Median
  • Range
  • IQR
  • Standard Deviation
  • Five-Number Summary

These numerical summaries help support conclusions that may be suggested by graphical displays such as:

  • Histograms
  • Dotplots
  • Stem-and-Leaf Plots
  • Boxplots

A strong statistical comparison combines graphical evidence with numerical summaries.


1.9.B.1 Comparing Distributions Using Numerical Summaries

A comparison of two or more distributions may include any numerical summary statistic.

These statistics provide objective evidence about the distributions and help justify statistical conclusions.

Common numerical summaries include:

StatisticUsed to Compare
MeanTypical value (center)
MedianTypical value (center)
RangeOverall spread
IQRSpread of middle 50%
Standard DeviationVariability around the mean
Five-Number SummaryCenter, spread, and position

When comparing distributions, numerical summaries help answer questions such as:

  • Which distribution has the higher center?
  • Which distribution is more variable?
  • Which distribution is more consistent?
  • Are the distributions similar or different?

Using Mean and Standard Deviation

When distributions are approximately symmetric and do not contain strong outliers, comparisons often focus on:

  • Mean
  • Standard Deviation

A larger mean indicates a larger typical value.

A larger standard deviation indicates greater variability.

Using Median and IQR

When distributions are skewed or contain outliers, comparisons often focus on:

  • Median
  • IQR

These statistics are resistant and provide a better description of the distribution.

AP Statistics Comparison Strategy

When comparing multiple distributions:

  1. Compare the centers.
  2. Compare the spreads.
  3. Describe any differences in shape.
  4. Mention any outliers.
  5. Support conclusions using numerical summaries.

This follows the AP Statistics comparison framework:

Center → Spread → Shape → Outliers

FeaturePossible Numerical Summary
CenterMean or Median
SpreadRange, IQR, Standard Deviation
ShapeMean-Median Relationship, Graphs
OutliersIQR Rule, Standard Deviation Rule

Example

Two classes took the same AP Statistics quiz.

StatisticClass AClass B
Mean8882
Median8981
Standard Deviation512

Compare the performance of the two classes using the numerical summaries.

▶️ Answer / Explanation

Center:

Class A has a higher mean and median than Class B.

Therefore, students in Class A generally scored higher on the quiz.

Spread:

Class B has the larger standard deviation.

Since:

\( \mathrm{12>5} \)

Class B’s scores are more spread out and less consistent.

Conclusion:

Class A performed better overall and had more consistent scores.

Class B had lower scores on average and greater variability.

Example

Two boxplots summarize weekly exercise times.

Group A:

\( \mathrm{Median=240} \) minutes, \( \mathrm{IQR=60} \) minutes

Group B:

\( \mathrm{Median=180} \) minutes, \( \mathrm{IQR=90} \) minutes

Use the numerical summaries to compare the two groups.

▶️ Answer / Explanation

Group A has the larger median.

Therefore, Group A typically exercises more each week.

Group B has the larger IQR.

Therefore, Group B shows greater variability in exercise time.

The middle 50% of exercise times are more spread out in Group B than in Group A.

Overall, Group A exercises more consistently and for a greater amount of time.


1.9.C Justify a Claim Using Multiple Quantitative One-Variable Graphical Representations

Graphs are powerful tools because they allow statisticians to visualize and compare distributions.

When multiple graphical representations of the same quantitative variable are available, they can provide evidence to support statistical claims.

Rather than relying on opinions or assumptions, statisticians use information displayed in graphs to justify conclusions.

Common graphical representations used to support claims include:

  • Histograms
  • Dotplots
  • Stem-and-Leaf Plots
  • Back-to-Back Stem-and-Leaf Plots
  • Boxplots

These graphs reveal important characteristics of distributions that can be used as statistical evidence.


1.9.C.1 Using Multiple Graphical Representations to Justify Claims

Multiple quantitative one-variable graphical representations may reveal information that supports claims about a variable in context.

A strong statistical claim should be supported by specific graphical evidence.

When examining graphs, statisticians commonly look for:

  • Center
  • Spread
  • Shape
  • Outliers
  • Clusters
  • Gaps

These characteristics provide evidence that can support conclusions about a distribution.

Claims About Center

Graphs can reveal which distribution has the larger typical value.

Evidence may include:

  • Larger median on a boxplot
  • Distribution shifted farther right on a histogram
  • Larger mean or median from a graphical summary

Claims About Variability

Graphs can reveal which distribution is more spread out.

Evidence may include:

  • Larger IQR on a boxplot
  • Wider histogram
  • Larger overall range

Claims About Shape

Graphs can reveal:

  • Symmetry
  • Skewness
  • Unimodality
  • Bimodality
  • Uniformity

Shape often provides important evidence about how the data are distributed.

Claims About Outliers

Graphs may reveal unusually large or unusually small observations.

Outliers often appear:

  • Beyond whiskers in boxplots
  • Separated from the main cluster in dotplots
  • Far from most observations in histograms

Outliers should always be mentioned when they affect the interpretation of a distribution.

FeaturePossible Graphical Evidence
CenterMedian position, location of distribution
SpreadIQR, range, overall width
ShapeSymmetry or skewness
OutliersIsolated observations
ClustersConcentrated groups of observations
GapsIntervals containing no observations

AP Statistics Claim Structure

When justifying a claim using graphs:

  1. State the claim.
  2. Identify graphical evidence.
  3. Interpret the evidence in context.

A strong response always connects the graphical feature directly to the claim being made.

Avoid statements such as:

“The graph looks higher.”

Instead write:

“The boxplot for Group A has a larger median than Group B, indicating that Group A typically has larger values.”

Example

Two boxplots compare weekly study times for two groups of students.

Group A:

  • Median = \( \mathrm{12} \) hours
  • IQR = \( \mathrm{4} \) hours

Group B:

  • Median = \( \mathrm{8} \) hours
  • IQR = \( \mathrm{9} \) hours

Use the graphical information to justify a claim about the study habits of the two groups.

▶️ Answer / Explanation

The boxplots suggest that Group A typically studies more than Group B.

This claim is supported by the medians:

\( \mathrm{12>8} \)

hours.

The boxplots also show that Group B has greater variability because:

\( \mathrm{IQR=9} \)

is larger than:

\( \mathrm{IQR=4} \)

hours.

Therefore, Group A generally studies more, while Group B’s study times vary more from student to student.

Example

Two histograms display daily commute times for workers in two cities.

  • City A has a distribution centered near: \( \mathrm{25} \) minutes and appears approximately symmetric.
  • City B has a distribution centered near: \( \mathrm{40} \) minutes and is skewed right with several very long commute times.

Use the histograms to justify a claim comparing the commuting patterns of the two cities.

▶️ Answer / Explanation

The histograms suggest that workers in City B generally have longer commute times than workers in City A.

This claim is supported by the centers of the distributions.

City B is centered near:

\( \mathrm{40} \)

minutes, whereas City A is centered near:

\( \mathrm{25} \)

minutes.

The histogram for City B is also skewed right, indicating that some workers have unusually long commute times.

Therefore, City B tends to have longer and more variable commute times than City A.

1.9.D.1 Standardized Scores (z-Scores)

A standardized score, or z-score, measures the number of standard deviations a data value lies above or below the population mean.

A z-score converts an original data value into a standardized value.

The z-score formula is:

\( z=\frac{x-\mu}{\sigma} \)

where:

  • \( z \) = z-score
  • \( x \) = observed data value
  • \( \mu \) = population mean
  • \( \sigma \) = population standard deviation

Important AP Statistics Ideas

  • A positive z-score means the observation is above the mean.
  • A negative z-score means the observation is below the mean.
  • The farther the z-score is from 0, the more unusual the observation is.
  • Z-scores allow comparisons between different distributions.

For example:

  • \( z=2 \) means the value is 2 standard deviations above the mean.
  • \( z=-1.5 \) means the value is 1.5 standard deviations below the mean.
Position Relative to MeanSign of z-Score
Above MeanPositive
At MeanZero
Below MeanNegative

Example

A population of exam scores has:

  • Population Mean: \( \mu=70 \)
  • Population Standard Deviation: \( \sigma=10 \)

A student earns a score of:

\( x=85 \)

Calculate the student’s z-score.

▶️ Answer / Explanation

Step 1: Use the z-score formula.

\( z=\frac{x-\mu}{\sigma} \)

Step 2: Substitute the values.

\( z=\frac{85-70}{10} \)

\( z=\frac{15}{10} \)

\( z=1.5 \)

Therefore:

\( \mathrm{z=1.5} \)

Interpretation:

The student’s score is:

\( \mathrm{1.5} \)

standard deviations above the population mean.

Example

A population of adult heights has:

  • Population Mean: \( \mu=170 \) cm
  • Population Standard Deviation: \( \sigma=8 \) cm

A person has a height of:

\( x=158 \) cm

Calculate the z-score and interpret its meaning.

▶️ Answer / Explanation

Step 1: Apply the z-score formula.

\( z=\frac{x-\mu}{\sigma} \)

\( z=\frac{158-170}{8} \)

\( z=\frac{-12}{8} \)

\( z=-1.5 \)

Therefore:

\( \mathrm{z=-1.5} \)

Interpretation:

The person’s height is:

\( \mathrm{1.5} \)

standard deviations below the population mean.

The negative z-score indicates the observation lies below average.

1.9.D.2 Calculating and Interpreting z-Scores

A z-score is a standardized score that measures how many standard deviations a data value lies above or below the mean.

The z-score is calculated by subtracting the population mean from the data value and then dividing by the population standard deviation.

z-Score Formula

\( z=\frac{x_i-\mu}{\sigma} \)

where:

  • \( x_i \) = data value (observation)
  • \( \mu \) = population mean
  • \( \sigma \) = population standard deviation
  • \( z \) = standardized score

The z-score tells us how far a data value is from the mean in units of standard deviations.

Interpretation of z-Scores

  • A positive z-score indicates the value is above the mean.
  • A negative z-score indicates the value is below the mean.
  • A z-score of 0 indicates the value is exactly equal to the mean.
z-ScoreMeaning
\( z=0 \)At the mean
\( z=1 \)1 standard deviation above the mean
\( z=-1 \)1 standard deviation below the mean
\( z=2 \)2 standard deviations above the mean
\( z=-2 \)2 standard deviations below the mean

Using Sample Statistics

Sometimes the population mean and population standard deviation are unknown.

In that situation, the sample mean and sample standard deviation may be used to calculate an approximate z-score.

Sample z-Score Formula

\( z=\frac{x_i-\bar{x}}{s} \)

where:

  • \( x_i \) = data value
  • \( \bar{x} \) = sample mean
  • \( s \) = sample standard deviation

This formula standardizes observations relative to the sample rather than the entire population.

Important AP Statistics Idea

The sign of the z-score tells the direction from the mean:

  • Positive → Above the mean
  • Negative → Below the mean

The magnitude of the z-score tells how unusual the observation is:

  • Small absolute value → Close to the mean
  • Large absolute value → Far from the mean

Example

The weights of a population of dogs have:

  • Population Mean: \( \mu=40 \) pounds
  • Population Standard Deviation: \( \sigma=5 \) pounds

A dog weighs:

\( x=50 \)

pounds.

Calculate the z-score and interpret its meaning.

▶️ Answer / Explanation

Step 1: Use the z-score formula.

\( z=\frac{x-\mu}{\sigma} \)

Step 2: Substitute the values.

\( z=\frac{50-40}{5} \)

\( z=\frac{10}{5} \)

\( z=2 \)

Interpretation:

The dog’s weight is:

\( \mathrm{2} \)

standard deviations above the population mean.

Example

A sample of students has:

  • Sample Mean: \( \bar{x}=72 \)
  • Sample Standard Deviation: \( s=8 \)

One student scored: \( x=60 \) on the exam.

Calculate the student’s z-score using the sample statistics.

▶️ Answer / Explanation

Step 1: Use the sample z-score formula.

\( z=\frac{x-\bar{x}}{s} \)

Step 2: Substitute the values.

\( z=\frac{60-72}{8} \)

\( z=\frac{-12}{8} \)

\( z=-1.5 \)

Interpretation:

The student’s score is:

\( \mathrm{1.5} \)

standard deviations below the sample mean.

The negative sign indicates that the score falls below average.


1.9.E Compare z-Scores as Measures of Relative Position for Distributions

A raw data value by itself does not always indicate how unusual or impressive an observation is.

For example, a score of: \( \mathrm{85} \) might be outstanding in one class but average in another.

To make meaningful comparisons, statisticians use z-scores.

Because z-scores measure distance from the mean in standard deviation units, they provide a standardized way to compare observations.

This allows comparisons:

  • Within the same distribution
  • Between different distributions

1.9.E.1 Using z-Scores to Compare Relative Position

A z-score describes an observation’s relative position within a distribution.

The z-score formula is:

\( z=\frac{x-\mu}{\sigma} \)

or when population parameters are unknown:

\( z=\frac{x-\bar{x}}{s} \)

A z-score indicates how many standard deviations an observation lies above or below the mean.

Because all z-scores are measured on the same scale, they can be used to compare observations from completely different distributions.

Comparing Relative Position Within a Distribution

Within a single distribution:

  • A larger z-score indicates a higher relative position.
  • A smaller z-score indicates a lower relative position.

Example:

  • Student A: \( z=2.0 \)
  • Student B: \( z=1.2 \)

Student A performed better relative to the rest of the group because Student A’s score is farther above the mean.

Comparing Relative Position Between Different Distributions

Z-scores are especially useful when comparing observations from different distributions.

Even if the original scales are different, z-scores place observations on a common standardized scale. The observation with the larger z-score has the stronger relative performance.

z-ScoreRelative Position
\( z=0 \)At the mean
\( z=1 \)Above average
\( z=2 \)Far above average
\( z=-1 \)Below average
\( z=-2 \)Far below average

Important AP Statistics Idea

When comparing observations from different distributions:

The larger z-score corresponds to the stronger relative position.

The original data values should not be compared directly because the distributions may have different centers and spreads.

ObservationRaw Valuez-ScoreBetter Relative Position?
AHigherSmallerNo
BLowerLargerYes

Example

Two students take the same exam.

Student A has:

\( z=1.8 \)

Student B has:

\( z=0.9 \)

Which student performed better relative to the rest of the class?

▶️ Answer / Explanation

The larger z-score indicates the stronger relative performance.

Since:

\( \mathrm{1.8>0.9} \)

Student A’s score is farther above the class mean.

Therefore, Student A performed better relative to the rest of the class.

Example

Maria scored: \( \mathrm{88} \) on a mathematics exam.

The mathematics scores have:

  • \( \mu=80 \)
  • \( \sigma=4 \)

Her z-score is:

\( z=\frac{88-80}{4}=2 \)

James scored: \( \mathrm{92} \) on a science exam.

The science scores have:

  • \( \mu=90 \)
  • \( \sigma=2 \)

His z-score is: \( z=\frac{92-90}{2}=1 \)

Who performed better relative to their classmates?

▶️ Answer / Explanation

Compare the z-scores:

Maria:

\( z=2 \)

James:

\( z=1 \)

Since:

\( \mathrm{2>1} \)

Maria’s score is farther above the mean of her distribution.

Therefore, Maria performed better relative to her classmates, even though James had the larger raw score.

This example shows why z-scores are useful when comparing observations from different distributions.

Scroll to Top