Home / AP Statistics 1.7 Summary Statistics for One Quantitative Variable Study Notes

AP Statistics 1.7 Summary Statistics for One Quantitative Variable Study Notes - New Syllabus

AP Statistics 1.7 Measures of Center and Variability Study Notes – New Syllabus

AP Statistics 1.7 Measures of Center and Variability Study Notes – As per latest AP Statistics Syllabus.

LEARNING OBJECTIVES

  • 1.7.A Calculate measures of center and position for quantitative data.
  • 1.7.B Calculate measures of variability for quantitative data.
  • 1.7.C Calculate different units of measurement for summary statistics.
  • 1.7.D Calculate outliers for quantitative data.
  • 1.7.E Compare multiple quantitative one-variable summary statistics.
  • 1.7.F Justify the selection of a summary statistic for describing quantitative data.

ESSENTIAL KNOWLEDGE:

  • 1.7.A.1 Two commonly used measures of center in the distribution of a quantitative variable are the mean and median.
  • 1.7.A.2 The mean is the sum of all the values divided by the number of values and can be found with and without using technology. For a sample, the mean is denoted by :

    \( \bar{x}=\dfrac{1}{n}\sum_{i=1}^{n}x_i \)

    where \(x_i\) represents the ith data point in the sample and \(n\) represents the number of data values in the sample.
  • 1.7.A.3 The median is the middle value when the data set is ordered from smallest to largest and can be found with and without using technology. One common method for determining the median of a data set with an even number of values is to use the mean of the two middle values. A common method for determining the median of a data set with an odd number of values is to use the value in the middle of all the values.
  • 1.7.A.4 In an ordered data set, the smallest value is the minimum value, and the largest value is the maximum value.
  • 1.7.A.5 The first quartile, denoted by Q1, is the median value of the lower half of the ordered data set from the minimum value to the position of the median. Approximately 25% of the values in the data set are less than or equal to Q1. The third quartile, denoted by Q3, is the median value of the upper half of the ordered data set from the position of the median to the maximum value. Approximately 75% of the values in the data set are less than or equal to Q3. The second quartile, Q2, is also the median of the data set. Q1 and Q3 form the boundaries for the middle 50% of values in an ordered data set.
  • 1.7.A.6 The pth percentile is the value that has p% of the data less than or equal to it when the data set is ordered from smallest to largest. The first and third quartiles are the 25th and 75th percentiles, respectively.
  • 1.7.B.1 Three commonly used measures of variability (or spread) in the distribution of a quantitative variable are the range, interquartile range, and standard deviation.
  • 1.7.B.2 The range is the difference between the maximum data value and the minimum data value.
  • 1.7.B.3 The interquartile range (IQR) is the difference between the third and first quartiles:

    \( Q_3-Q_1 \)
  • 1.7.B.4 The standard deviation is a typical deviation of the data values from their mean and can be found with and without using technology. The sample standard deviation is denoted by s and calculated by

    \( s=\sqrt{\dfrac{1}{n-1}\sum(x_i-\bar{x})^2} \)

    where \(x_i\) is the data value, \( \bar{x} \) is the mean, and \(n\) is the number of data values in the sample. The square of the sample standard deviation, \(s^2\), is called the sample variance.
  • 1.7.C.1 Changing units of measurement affects the values of the calculated statistics.
  • 1.7.D.1 There are many methods for determining potential outliers. Two methods frequently used are as follows:
    • 1.7.D.1.i An outlier is a value located more than 1.5 × IQR above the third quartile or more than 1.5 × IQR below the first quartile.
    • 1.7.D.1.ii An outlier is a value located more than 2 standard deviations above, or below, the mean.
  • 1.7.E.1 Summary statistics can be used to compare features of two or more independent samples, including center, variability, shape, and outliers.
  • 1.7.F.1 The median and IQR are considered a resistant (or robust) measure of center and measure of variability, respectively, because outliers do not greatly (if at all) affect their values. Because outliers can affect their values greatly, the mean is considered a nonresistant (or non-robust) measure of center, and the range and standard deviation are considered nonresistant (or non-robust) measures of variability.
  • 1.7.F.2 Summary statistics of a quantitative variable may reveal information that can be used to justify claims about the variable in context.

AP Statistics – Concise Summary Notes – All Topics


1.7.A Calculate Measures of Center and Position for Quantitative Data

When analyzing quantitative data, statisticians often want to identify a value that represents the center of the distribution.

Measures of center describe the typical value in a data set.

Measures of position describe where a particular value is located relative to the rest of the data.

Some of the most important measures of center and position include:

  • Mean
  • Median
  • Minimum and Maximum
  • Quartiles
  • Percentiles

These measures help summarize large data sets and provide information about the distribution of the data.


1.7.A.1 Measures of Center: Mean and Median

A measure of center is a numerical value that represents the middle or typical value of a distribution.

The two most commonly used measures of center are:   

  • Mean
  • Median

Mean

The mean is commonly called the average.

It uses every value in the data set and is calculated by adding all values and dividing by the number of observations.

Median

The median is the middle value in an ordered data set.

  • Half of the observations lie below the median and half lie above it.
  • Unlike the mean, the median is not strongly affected by unusually large or unusually small values.
MeasureDescription
MeanArithmetic average of all values
MedianMiddle value in ordered data

The choice between mean and median often depends on the shape of the distribution.

  • For symmetric distributions, the mean and median are usually similar.
  • For skewed distributions, the median is often preferred because it is resistant to outliers.

Example

The following data represent the number of hours students studied for an exam:

\( \mathrm{2,\ 4,\ 5,\ 6,\ 8} \)

Identify the two measures of center that could be used to describe this data set.

▶️ Answer / Explanation

The two commonly used measures of center are:

  • Mean
  • Median

Both measures describe the typical study time for the students.

The mean uses all values in the calculation, while the median identifies the middle value of the ordered data set.


1.7.A.2 Mean

The mean is the arithmetic average of a set of quantitative data.

It is calculated by adding all observations and dividing by the total number of observations.

For a sample, the mean is represented by:

$\bar{x}=\frac{\sum x_i}{n}$

where:

  • \( \mathrm{x_i} \) represents an individual data value
  • \( \mathrm{n} \) represents the sample size
  • \( \mathrm{\bar{x}} \) represents the sample mean

Steps for calculating the mean:

  1. Add all data values.
  2. Count the number of observations.
  3. Divide the sum by the number of observations.

The mean uses every value in the data set.

Because of this, extremely large or small values can greatly influence the mean.

Example

Calculate the mean of the following data set:

\( \mathrm{4,\ 6,\ 8,\ 10,\ 12} \)

▶️ Answer / Explanation

Step 1: Find the sum of the values.

\( \mathrm{4+6+8+10+12=40} \)

Step 2: Count the number of observations.

\( \mathrm{n=5} \)

Step 3: Apply the mean formula.

\( \mathrm{\bar{x}=\dfrac{40}{5}=8} \)

Therefore, the mean is:

\( \mathrm{8} \)


1.7.A.3 Median

The median is the middle value in an ordered data set.

  • Before finding the median, the data must be arranged from smallest to largest.
  • The method used depends on whether the number of observations is odd or even.

Formula / Method for Finding the Median

Step 1: Arrange the data from smallest to largest.

Step 2: Determine whether the number of observations (\(n\)) is odd or even.

If \(n\) is odd:

$ \text{Median}=\text{Middle Value} $

$\text{Position of Median}=\frac{n+1}{2}$

This gives the location of the middle value in the ordered data set.

If \(n\) is even

$ \text{Median}=\frac{\text{Two Middle Values}}{2} $

Equivalently,

$ \text{Median}=\frac{x_{\frac{n}{2}}+x_{\frac{n}{2}+1}}{2} $

where the data are arranged in ascending order.

Number of ValuesHow to Find the Median
OddChoose the middle value
EvenAverage the two middle values

The median is resistant to outliers because it depends only on the position of values rather than their magnitudes.

Example

Find the median of the following data set:

\( \mathrm{3,\ 5,\ 7,\ 9,\ 11,\ 13,\ 15} \)

▶️ Answer / Explanation

Step 1: Verify that the data are ordered.

\( \mathrm{3,\ 5,\ 7,\ 9,\ 11,\ 13,\ 15} \)

Step 2: Count the observations.

\( \mathrm{n=7} \)

Since there is an odd number of values, the median is the middle value.

The fourth value is:

\( \mathrm{9} \)

Therefore, the median is:

\( \mathrm{9} \)


1.7.A.4 Minimum and Maximum Values

In an ordered data set, the smallest value is called the minimum value and the largest value is called the maximum value.

These values identify the endpoints of the distribution.

The minimum and maximum values help describe:

  • The spread of the data
  • The overall range of values
  • The boundaries of the distribution

Definitions

  • Minimum: Smallest value in the ordered data set
  • Maximum: Largest value in the ordered data set

Formula

\( \text{Range}=\text{Maximum}-\text{Minimum} \)

Although range is studied later as a measure of variability, it is calculated directly from the minimum and maximum values.

MeasureMeaning
MinimumSmallest value in the data set
MaximumLargest value in the data set

Example

The following test scores were recorded:

\( \mathrm{58,\ 64,\ 70,\ 72,\ 75,\ 81,\ 89} \)

Identify the minimum value and maximum value.

▶️ Answer / Explanation

The smallest value is:

\( \mathrm{58} \)

Therefore, the minimum value is:

\( \mathrm{58} \)

The largest value is:

\( \mathrm{89} \)

Therefore, the maximum value is:

\( \mathrm{89} \)


1.7.A.5 Quartiles (Q1, Q2, Q3)

Quartiles divide an ordered data set into four approximately equal parts.

The three quartiles are:

  • \( \mathrm{Q_1} \) (First Quartile)
  • \( \mathrm{Q_2} \) (Second Quartile or Median)
  • \( \mathrm{Q_3} \) (Third Quartile)

Quartiles help describe the position of data values within a distribution.

First Quartile (Q1)

\( \mathrm{Q_1} \) is the median of the lower half of the ordered data.

Approximately: \( \mathrm{25\%} \)

of the observations are less than or equal to \( \mathrm{Q_1} \).

Second Quartile (Q2)

\( \mathrm{Q_2} \) is the median of the entire data set.

Approximately: \( \mathrm{50\%} \)

of the observations are less than or equal to \( \mathrm{Q_2} \).

Third Quartile (Q3)

\( \mathrm{Q_3} \) is the median of the upper half of the ordered data.

Approximately: \( \mathrm{75\%} \)

of the observations are less than or equal to \( \mathrm{Q_3} \).

Method for Finding Quartiles

  1. Order the data from smallest to largest.
  2. Find the median (\(Q_2\)).
  3. Find the median of the lower half (\(Q_1\)).
  4. Find the median of the upper half (\(Q_3\)).

\( \mathrm{Q_1} \) and \( \mathrm{Q_3} \) form the boundaries of the middle 50% of the data.

QuartilePercent of Data Below It
\( \mathrm{Q_1} \)25%
\( \mathrm{Q_2} \)50%
\( \mathrm{Q_3} \)75%

Example

Find \( \mathrm{Q_1} \), \( \mathrm{Q_2} \), and \( \mathrm{Q_3} \) for the following data set:

\( \mathrm{2,\ 4,\ 6,\ 8,\ 10,\ 12,\ 14,\ 16,\ 18} \)

▶️ Answer / Explanation

Step 1: Find the median.

\( \mathrm{Q_2=10} \)

Lower half:

\( \mathrm{2,\ 4,\ 6,\ 8} \)

Upper half:

\( \mathrm{12,\ 14,\ 16,\ 18} \)

Step 2: Find \( \mathrm{Q_1} \).

\( \mathrm{Q_1=\dfrac{4+6}{2}=5} \)

Step 3: Find \( \mathrm{Q_3} \).

\( \mathrm{Q_3=\dfrac{14+16}{2}=15} \)

Therefore:

\( \mathrm{Q_1=5,\ Q_2=10,\ Q_3=15} \)


1.7.A.6 Percentiles

A percentile indicates the position of a value within an ordered data set.

The pth percentile is the value that has approximately: \( \mathrm{p\%} \) of the observations less than or equal to it.

Percentiles help determine how a particular observation compares with the rest of the data.

Definition

\( \text{pth Percentile} = \text{Value with } p\% \text{ of observations at or below it} \)

Common percentiles include:

PercentileEquivalent Quartile
25th Percentile\( \mathrm{Q_1} \)
50th Percentile\( \mathrm{Q_2} \) (Median)
75th Percentile\( \mathrm{Q_3} \)

Percentiles are commonly used in:

  • Standardized testing
  • Growth charts
  • Ranking systems
  • Educational assessments

Example

A student scored at the 90th percentile on a standardized exam.

Interpret this percentile in context.

▶️ Answer / Explanation

Being at the 90th percentile means the student’s score was greater than or equal to approximately:

\( \mathrm{90\%} \)

of all scores in the distribution.

Only about:

\( \mathrm{10\%} \)

of students scored higher.

Therefore, the student performed better than most students who took the exam.


1.7.B Calculate Measures of Variability for Quantitative Data

While measures of center describe the typical value of a distribution, measures of variability describe how spread out the data values are.

Two data sets may have the same center but very different amounts of variability.

Measures of variability help statisticians understand:

  • How much the data vary
  • How closely values cluster around the center
  • Whether the distribution is tightly packed or widely spread

The three most commonly used measures of variability are:

  • Range
  • Interquartile Range (IQR)
  • Standard Deviation

1.7.B.1 Measures of Variability

Measures of variability describe the spread of a quantitative distribution.

  • A distribution with little variability has values that are close together.
  • A distribution with large variability has values that are spread farther apart.

MeasureDescription
RangeDistance from minimum to maximum
Interquartile Range (IQR)Spread of the middle 50% of data
Standard DeviationTypical distance from the mean

Different measures of variability are useful in different situations.

  • Range uses only the smallest and largest values.
  • IQR focuses on the middle 50% of the data.
  • Standard deviation uses every observation.

Example

Two classes have the same mean test score of: \( \mathrm{80} \)

Class A scores range from: \( \mathrm{78} \) to: \( \mathrm{82} \)

Class B scores range from: \( \mathrm{50} \) to: \( \mathrm{100} \)

Which class has greater variability?

▶️ Answer / Explanation

Class B has greater variability.

Although both classes have the same mean, Class B’s scores are spread over a much wider interval.

This indicates that the data values vary more from one another.


1.7.B.2 Range

The range measures the total spread of a data set.

It is calculated using only the minimum value and maximum value.

Formula

\( \text{Range}=\text{Maximum}-\text{Minimum} \)

where:

  • Maximum = largest value in the data set
  • Minimum = smallest value in the data set

A larger range indicates greater variability.

A smaller range indicates less variability.

One limitation of range is that it depends only on two values and can be greatly affected by outliers.

MinimumMaximumRange
123220

Example

The following data represent the number of books read by students:

\( \mathrm{3,\ 5,\ 8,\ 10,\ 12,\ 15} \)

Calculate the range.

▶️ Answer / Explanation

Step 1: Identify the minimum and maximum values.

Minimum:

\( \mathrm{3} \)

Maximum:

\( \mathrm{15} \)

Step 2: Apply the formula.

\( \text{Range}=15-3 \)

\( \text{Range}=12 \)

Therefore, the range is:

\( \mathrm{12} \)


1.7.B.3 Interquartile Range (IQR)

The Interquartile Range (IQR) measures the spread of the middle 50% of a data set.

  • Because it uses quartiles rather than extreme values, the IQR is resistant to outliers.
  • The IQR is the distance between the third quartile and the first quartile.

Formula

\( \mathrm{IQR}=Q_3-Q_1 \)

where:

  • \( \mathrm{Q_1} \) = First Quartile (25th percentile)
  • \( \mathrm{Q_3} \) = Third Quartile (75th percentile)

The IQR contains the middle: \( \mathrm{50\%} \) of all observations.

A larger IQR indicates greater variability among the middle half of the data.

A smaller IQR indicates that the middle half of the data is more concentrated.

MeasureFormula
Interquartile Range\( \mathrm{IQR=Q_3-Q_1} \)

Example

A data set has: \( \mathrm{Q_1=18} \) and \( \mathrm{Q_3=34} \)

Calculate the interquartile range.

▶️ Answer / Explanation

Step 1: Apply the IQR formula.

\( \mathrm{IQR}=Q_3-Q_1 \)

\( \mathrm{IQR}=34-18 \)

\( \mathrm{IQR}=16 \)

Therefore, the interquartile range is:

\( \mathrm{16} \)

This means the middle 50% of the observations span:

\( \mathrm{16} \)

units.


1.7.B.4 Standard Deviation

The standard deviation measures the typical distance that data values lie from the mean of a distribution.

It describes how spread out the data are around the center.

  • A small standard deviation indicates that most values are close to the mean.
  • A large standard deviation indicates that values are spread farther from the mean.
  • Unlike the range, which uses only two observations, standard deviation uses every value in the data set.

For a sample, the standard deviation is represented by: \( \mathrm{s} \)

Formula for Sample Standard Deviation

\( s=\sqrt{\frac{1}{n-1}\sum_{i=1}^{n}(x_i-\bar{x})^2} \)

where:

  • \( \mathrm{s} \) = sample standard deviation
  • \( \mathrm{x_i} \) = individual data value
  • \( \mathrm{\bar{x}} \) = sample mean
  • \( \mathrm{n} \) = sample size
  • \( \mathrm{(x_i-\bar{x})} \) = deviation from the mean

The standard deviation is calculated by:

  1. Finding the mean.
  2. Subtracting the mean from each observation.
  3. Squaring each deviation.
  4. Adding the squared deviations.
  5. Dividing by \( \mathrm{n-1} \).
  6. Taking the square root.

Because this process is lengthy, standard deviation is often calculated using technology in AP Statistics.

Sample Variance

The square of the sample standard deviation is called the sample variance.

It is represented by: \( \mathrm{s^2} \)

Formula for Sample Variance

\( s^2=\frac{1}{n-1}\sum_{i=1}^{n}(x_i-\bar{x})^2 \)

Variance and standard deviation contain the same information about variability, but standard deviation is usually easier to interpret because it is measured in the same units as the original data.

MeasureMeaning
\( \mathrm{s} \)Sample Standard Deviation
\( \mathrm{s^2} \)Sample Variance

Interpreting Standard Deviation

  • Small standard deviation → data values are close to the mean.
  • Large standard deviation → data values are spread farther from the mean.

For example:

  • A class with scores clustered around \( \mathrm{80} \) has a small standard deviation.
  • A class with scores ranging widely around \( \mathrm{80} \) has a large standard deviation.

Example

The following data represent the number of hours students studied for a quiz: \( \mathrm{4,\ 6,\ 8,\ 10,\ 12} \) The mean is: \( \mathrm{\bar{x}=8} \)

Calculate the sample standard deviation.

▶️ Answer / Explanation

Step 1: Calculate each deviation from the mean.

\(x_i\)\(x_i-\bar{x}\)\((x_i-\bar{x})^2\)
4-416
6-24
800
1024
12416

Step 2: Sum the squared deviations.

\( 16+4+0+4+16=40 \)

Step 3: Calculate the sample variance.

\( s^2=\frac{40}{5-1} \)

\( s^2=\frac{40}{4}=10 \)

Step 4: Take the square root.

\( s=\sqrt{10} \)

\( s\approx3.16 \)

Therefore, the sample standard deviation is:

\( s\approx3.16 \)

Interpretation:

The study times typically vary about:

\( \mathrm{3.16} \)

hours from the mean study time of:

\( \mathrm{8} \)

hours.


1.7.C Calculate Different Units of Measurement for Summary Statistics

Quantitative data can be measured using different units.

For example, a person’s height may be recorded in:

  • Inches
  • Feet
  • Centimeters
  • Meters

When the units of measurement change, the numerical values of the data also change.

As a result, summary statistics calculated from the data will change as well.

Understanding how summary statistics are affected by unit conversions is important when comparing data reported in different measurement systems.


1.7.C.1 Effect of Changing Units on Summary Statistics

Changing units of measurement affects the numerical values of calculated statistics.

Suppose every value in a data set is multiplied by a constant.

Then:

  • The mean is multiplied by the same constant.
  • The median is multiplied by the same constant.
  • The minimum and maximum are multiplied by the same constant.
  • The range is multiplied by the same constant.
  • The IQR is multiplied by the same constant.
  • The standard deviation is multiplied by the same constant.

In general:

\( \text{New Statistic} = (\text{Conversion Factor})(\text{Original Statistic}) \)

For example:

Since:

\( 1\text{ foot}=12\text{ inches} \) a height of: \( \mathrm{60} \) inches becomes: \( \frac{60}{12}=5 \) feet.

If every observation is divided by: \( \mathrm{12} \) then the mean, median, range, IQR, and standard deviation are also divided by: \( \mathrm{12} \)

StatisticEffect of Multiplying All Data by \(c\)
MeanMultiplied by \(c\)
MedianMultiplied by \(c\)
MinimumMultiplied by \(c\)
MaximumMultiplied by \(c\)
RangeMultiplied by \(c\)
IQRMultiplied by \(c\)
Standard DeviationMultiplied by \(c\)

Important AP Statistics Idea

Changing measurement units does not change the shape of a distribution.

Only the numerical values of the statistics change.

For example:

  • A right-skewed distribution remains right-skewed.
  • A symmetric distribution remains symmetric.
  • Outliers remain outliers.

Example

The heights of a group of students are measured in inches. The mean height is: \( \mathrm{66} \) inches. The standard deviation is: \( \mathrm{4} \) inches.

The heights are converted from inches to feet using:

\( 1\text{ foot}=12\text{ inches} \)

Find the new mean and the new standard deviation.

▶️ Answer / Explanation

Since all measurements are divided by: \( \mathrm{12} \)

the mean and standard deviation must also be divided by: \( \mathrm{12} \)

New Mean:

\( \frac{66}{12}=5.5 \) feet

New Standard Deviation:

\( \frac{4}{12}=0.333 \) feet

Therefore:

  • Mean = \( \mathrm{5.5} \) feet
  • Standard Deviation ≈ \( \mathrm{0.333} \) feet

The shape of the distribution remains unchanged because only the measurement units were converted.


1.7.D Calculate Outliers for Quantitative Data

An outlier is an observation that is unusually far from the rest of the data values in a distribution.

Outliers are important because they can:

  • Strongly affect the mean.
  • Increase variability.
  • Influence statistical conclusions.
  • Indicate unusual events or possible recording errors.

There is no single definition of an outlier.

Statisticians use several methods to identify potential outliers.

Two common methods used in AP Statistics are:

  • The IQR Method
  • The Standard Deviation Method

1.7.D.1 Methods for Identifying Potential Outliers

A potential outlier is a value that lies unusually far from the center of a distribution according to a specified rule.

The two most common rules are:

  • The 1.5 × IQR Rule
  • The 2 Standard Deviations Rule

These rules help identify observations that deserve further investigation.

Being identified as a potential outlier does not necessarily mean that the value should be removed from the data set.

The context of the study should always be considered.


1.7.D.1.i Outliers Using the 1.5 × IQR Rule

One common method for identifying potential outliers uses the interquartile range (IQR). 

Recall:

\( \mathrm{IQR=Q_3-Q_1} \)

A value is considered a potential outlier if it falls:

  • More than \( \mathrm{1.5(IQR)} \) below \( \mathrm{Q_1} \)
  • More than \( \mathrm{1.5(IQR)} \) above \( \mathrm{Q_3} \)

Lower Fence Formula

\( \mathrm{Lower\ Fence=Q_1-1.5(IQR)} \)

Upper Fence Formula

\( \mathrm{Upper\ Fence=Q_3+1.5(IQR)} \)

Any observation outside these fences is classified as a potential outlier.

ConditionPotential Outlier?
Below Lower FenceYes
Above Upper FenceYes
Between FencesNo

Example

A data set has: \( \mathrm{Q_1=20} \) and \( \mathrm{Q_3=40} \)

Determine whether: \( \mathrm{75} \) is a potential outlier using the 1.5 × IQR rule.

▶️ Answer / Explanation

Step 1: Calculate the IQR.

\( \mathrm{IQR=40-20=20} \)

Step 2: Calculate the fences.

Lower Fence:

\( \mathrm{20-1.5(20)=-10} \)

Upper Fence:

\( \mathrm{40+1.5(20)=70} \)

Step 3: Compare the value to the fences.

\( \mathrm{75>70} \)

Therefore:

\( \mathrm{75} \)

lies above the upper fence and is a potential outlier.


1.7.D.1.ii Outliers Using the Standard Deviation Rule

Another common method identifies potential outliers using the mean and standard deviation.

A value may be considered a potential outlier if it is more than:

\( \mathrm{2} \)

standard deviations above or below the mean.

Lower Boundary Formula

\( \mathrm{\bar{x}-2s} \)

Upper Boundary Formula

\( \mathrm{\bar{x}+2s} \)

where:

  • \( \mathrm{\bar{x}} \) = sample mean
  • \( \mathrm{s} \) = sample standard deviation

Values outside these boundaries may be classified as potential outliers.

ConditionPotential Outlier?
Below \( \mathrm{\bar{x}-2s} \)Yes
Above \( \mathrm{\bar{x}+2s} \)Yes
Between the BoundariesNo

Example

A sample has: \( \mathrm{\bar{x}=50} \) and \( \mathrm{s=8} \)

Determine whether: \( \mathrm{70} \)

is a potential outlier using the standard deviation rule.

▶️ Answer / Explanation

Step 1: Calculate the boundaries.

Lower Boundary:

\( \mathrm{50-2(8)=34} \)

Upper Boundary:

\( \mathrm{50+2(8)=66} \)

Step 2: Compare the value to the boundaries.

\( \mathrm{70>66} \)

Therefore:

\( \mathrm{70} \)

is more than two standard deviations above the mean and is considered a potential outlier.


1.7.E Compare Multiple Quantitative One-Variable Summary Statistics

Statisticians often compare two or more quantitative data sets to identify similarities and differences.

Rather than examining every individual observation, summary statistics provide a concise way to compare distributions.

When comparing quantitative data sets, statisticians focus on:

  • Center
  • Variability
  • Shape
  • Outliers

Comparisons should always be made in context.

For example:

  • Which group has the higher typical value?
  • Which group is more consistent?
  • Which group has greater spread?
  • Does one group contain unusual observations?

Summary statistics provide evidence that helps answer these questions.


1.7.E.1 Comparing Quantitative Data Using Summary Statistics

Summary statistics can be used to compare features of two or more independent samples.

Four important features should be considered:

  • Center
  • Variability
  • Shape
  • Outliers

1. Comparing Center

Measures of center describe the typical value of a distribution.

Common measures:

  • Mean (\( \bar{x} \))
  • Median

A larger mean or median indicates a higher typical value.

Example:

  • Mean score of Class A = \( \mathrm{82} \)
  • Mean score of Class B = \( \mathrm{76} \)

Class A performed better on average.

2. Comparing Variability

Measures of variability describe how spread out the data are.

Common measures:

  • Range
  • IQR
  • Standard Deviation

A larger measure indicates greater variability.

Example:

  • IQR of Class A = \( \mathrm{8} \)
  • IQR of Class B = \( \mathrm{20} \)

Class B’s scores are more spread out.

3. Comparing Shape

Distributions may differ in shape.

Possible descriptions include:

  • Symmetric
  • Skewed Right
  • Skewed Left
  • Unimodal
  • Bimodal
  • Uniform

Shape comparisons often come from graphical displays such as histograms or dotplots.

4. Comparing Outliers

One distribution may contain outliers while another does not.

Outliers can affect:

  • Mean
  • Range
  • Standard deviation

When comparing distributions, always mention any unusual observations.

FeatureSummary Statistics Used
CenterMean, Median
VariabilityRange, IQR, Standard Deviation
ShapeHistogram, Dotplot, Stem-and-Leaf Plot
OutliersIQR Rule or Standard Deviation Rule

AP Statistics Comparison Template

When comparing two distributions, AP Statistics students should discuss:

  • Center
  • Spread
  • Shape
  • Outliers

This is often remembered as:

CSSO

  • Center
  • Spread
  • Shape
  • Outliers

A complete AP response should address all four characteristics whenever possible.

Example

Two classes took the same mathematics exam.

StatisticClass AClass B
Mean8478
Median8577
IQR1018
Standard Deviation612

Compare the performance of the two classes using the summary statistics.

▶️ Answer / Explanation

Center:

Class A has a higher mean (\( \mathrm{84} \)) and median (\( \mathrm{85} \)) than Class B.

Therefore, Class A performed better overall.

Variability:

Class B has a larger IQR (\( \mathrm{18} \)) and a larger standard deviation (\( \mathrm{12} \)).

Therefore, Class B’s scores are more spread out and less consistent.

Shape:

The shape cannot be determined from the summary statistics alone.

A graphical display would be needed.

Outliers:

No information about outliers is provided.

Based on the available summary statistics, Class A performed better and showed more consistency than Class B.


1.7.F Justify the Selection of a Summary Statistic for Describing Quantitative Data

Different summary statistics provide different information about a quantitative distribution.

The choice of which summary statistic to use depends on the shape of the distribution and the presence of outliers.

Some statistics are strongly affected by unusual observations, while others remain relatively unchanged.

When selecting summary statistics, statisticians consider:

  • The shape of the distribution
  • The presence of outliers
  • The purpose of the analysis

Choosing appropriate summary statistics helps provide an accurate description of the data.


1.7.F.1 Resistant and Nonresistant Statistics

  • A resistant (robust) statistic is not greatly affected by unusually large or unusually small observations.
  • A nonresistant (non-robust) statistic can change substantially when outliers are present.

Understanding which statistics are resistant is important because outliers can sometimes distort the description of a distribution.

Resistant Measures

The following statistics are considered resistant:

  • Median
  • Interquartile Range (IQR)

These statistics depend primarily on the position of observations rather than their actual magnitudes.

As a result, extreme values have little effect on them.

Nonresistant Measures

The following statistics are considered nonresistant:

  • Mean
  • Range
  • Standard Deviation

These statistics use the actual numerical values of the observations.

Therefore, extreme observations can significantly affect their values.

StatisticResistant?Reason
MedianYesBased on position of values
IQRYesUses quartiles rather than extreme values
MeanNoUses every value directly
RangeNoDepends entirely on minimum and maximum
Standard DeviationNoUses squared deviations from the mean

Choosing the Appropriate Summary Statistics

For distributions that are:

  • Symmetric
  • Without significant outliers

the preferred measures are:

  • Mean
  • Standard Deviation

For distributions that are:

  • Skewed
  • Contain outliers

the preferred measures are:

  • Median
  • IQR

This is one of the most commonly tested concepts in AP Statistics.

Distribution TypePreferred Measure of CenterPreferred Measure of Spread
Symmetric, No OutliersMeanStandard Deviation
Skewed or Contains OutliersMedianIQR

Example

A distribution of household incomes is strongly skewed right because a small number of households earn extremely large incomes.

Which measures of center and variability should be used to describe this distribution?

▶️ Answer / Explanation

Because the distribution is skewed right, the mean and standard deviation may be heavily influenced by the extremely large incomes.

Therefore, the preferred measures are:

  • Median (center)
  • IQR (spread)

These measures are resistant to outliers and provide a more accurate description of the typical household income.


1.7.F.2 Using Summary Statistics to Justify Claims in Context

Summary statistics provide numerical evidence that can be used to support claims about a quantitative variable.

Rather than making conclusions based on opinion, statisticians use measures such as:

  • Mean
  • Median
  • Range
  • IQR
  • Standard Deviation

to justify statements about a distribution.

Claims may involve:

  • Typical values
  • Variability
  • Consistency
  • Comparisons between groups

A strong statistical justification should:

  • Reference specific summary statistics.
  • Interpret those statistics in context.
  • Connect the evidence directly to the claim.

Weak Claim:

“Class A did better.”

Strong Claim:

“Class A performed better because its median exam score was higher than Class B’s median exam score.”

StatisticPossible Claim Supported
Mean or MedianTypical value is larger or smaller
IQRMiddle 50% is more or less spread out
Standard DeviationData are more or less consistent

Example

Two basketball teams recorded the following median points scored per game:

  • Team A: \( \mathrm{82} \)
  • Team B: \( \mathrm{75} \)

Use the summary statistics to justify a claim about the teams’ scoring performances.

▶️ Answer / Explanation

Team A has a higher median score than Team B.

The median represents the middle score in each distribution.

Since:

\( \mathrm{82>75} \)

Team A typically scores more points per game than Team B.

Therefore, the summary statistics support the claim that Team A generally has stronger scoring performance.

Scroll to Top