Chapter 20

Univariate Statistics and Data Distributions

Sampling Techniques and Bias

Sampling Techniques and Bias

In statistics, we often want to learn about a large group without asking every single person or measuring every single item. This large group is called the population.

A smaller group chosen from the population is called a sample. If the sample is chosen well, it can give us useful information about the whole population.

For example, a school principal may want to know how students feel about school lunches. Instead of asking all 1,200 students, the principal might ask 100 students. Those 100 students are the sample.

The way a sample is chosen matters a lot. A good sampling method helps make results fair and accurate. A poor sampling method can lead to bias, which means the results are pushed in a certain direction and do not fairly represent the population.

1. Population, Sample, and Bias

Before learning the sampling techniques, let’s make sure these basic ideas are clear:

  • Population: the entire group you want to study
  • Sample: a smaller part of the population that you actually collect data from
  • Bias: a problem in how the sample is chosen that makes it unrepresentative

A sample should represent the population as closely as possible. If some groups are overrepresented or left out, the data may not reflect the truth.

2. Why Sampling Is Used

Sampling is useful because studying a whole population can be difficult, expensive, or take too much time.

  • A company may want to test a few products instead of every product made.
  • A scientist may measure some plants in a field instead of every plant.
  • A school may survey some students instead of all students.

Even though samples are smaller, they can still be helpful if they are selected carefully.

3. Main Sampling Techniques

There are several common ways to choose a sample. In 9th Grade, the most important ones are random, systematic, stratified, and convenience sampling.

A. Random Sampling

In a random sample, every member of the population has an equal chance of being chosen.

This is one of the fairest methods because no person or item is intentionally favored.

Examples of random sampling include:

  • Putting every student’s name in a box and drawing 50 names
  • Using a random number generator to pick survey participants

Why it is useful: Random sampling helps reduce bias because everyone has the same chance to be selected.

B. Systematic Sampling

In a systematic sample, you start at a random point and then choose every nth person or item.

For example, if a list of students is in alphabetical order, you might start with the 3rd student and then choose every 10th student after that.

If the sample size pattern is every 10th student, the rule is:

$$\text{Select every } 10\text{th student}$$

Why it is useful: It is simple and organized, especially when working from a list.

Warning: Systematic sampling can become biased if the list has a hidden pattern. For example, if every 10th item is different in some important way, the sample may not be fair.

C. Stratified Sampling

In a stratified sample, the population is divided into groups, called strata, based on an important characteristic. Then a random sample is taken from each group.

For example, a school may divide students by grade level: 9th, 10th, 11th, and 12th grade. Then it may randomly choose students from each grade.

This is useful when you want all important groups represented in the sample.

Why it is useful: Stratified sampling can give a more accurate picture when the population has clear groups.

D. Convenience Sampling

In a convenience sample, the sample is made up of the people or items that are easiest to reach.

Examples include:

  • Surveying only students in your math class
  • Asking the first 20 people you see at lunch
  • Polling shoppers who happen to walk into one store

Why it is risky: Convenience sampling is quick and easy, but it often leads to bias because the sample may not represent the whole population.

4. Understanding Bias

Bias happens when a sampling method favors certain outcomes or leaves out part of the population.

A biased sample does not truly reflect the population, so the conclusions drawn from it may be misleading.

Here are some common ways bias can happen:

  • Only sampling one group: asking only athletes about school sports funding
  • Convenience bias: surveying only your friends because they are easy to ask
  • Undercoverage: leaving out an important part of the population
  • Poor timing or location: asking about bus use only among students who stay after school

To reduce bias, the sample should include different types of people or items from the whole population.

5. How to Recognize the Sampling Method

When reading a question, look for key clues:

  • Random: words like “chosen at random,” “draw names,” or “random number generator”
  • Systematic: phrases like “every 5th person” or “every 20th item”
  • Stratified: population is divided into groups, then sampled from each group
  • Convenience: easiest people to reach are chosen

If a question asks whether a sample is biased, ask yourself:

  1. Does the sample represent the whole population?
  2. Was one group favored or ignored?
  3. Was the method fair, or just easy?

6. Worked Examples

Example 1: Identifying a Sampling Method

A teacher wants to survey students about homework time. She writes every student’s name on a slip of paper, mixes them, and picks 30 names.

Question: What sampling method is this?

Solution:

  • Every student’s name is included.
  • The names are mixed.
  • Any student could be picked.

This is a random sample.

Why: Every student has an equal chance of being chosen.

Example 2: Systematic Sampling

A factory checks light bulbs for defects. Starting with the 4th bulb produced, the manager tests every 15th bulb.

Question: What sampling method is being used?

Solution:

  • The manager begins at a starting point.
  • Then selects every 15th bulb.

This is systematic sampling.

Why: The sample follows a fixed pattern after a starting point.

Example 3: Stratified Sampling

A school wants to know how much time students spend on homework each night. The students are divided by grade: 9, 10, 11, and 12. Then 25 students are randomly selected from each grade.

Question: What sampling method is this, and why might it be a good choice?

Solution:

  • The school divided the population into groups by grade.
  • Then it randomly selected students from each group.

This is stratified sampling.

Why it is a good choice: Each grade level is represented, so the sample is more likely to reflect the whole school.

Example 4: Identifying Bias

A student wants to know whether teenagers in town like the new park. She surveys 40 teenagers who are already at the park.

Question: Is this sample likely to be biased?

Solution:

  • The population is all teenagers in town.
  • The sample only includes teenagers who are already at the park.
  • Teenagers at the park may be more likely to like it than those who do not go there.

Yes, this sample is biased.

Why: It does not fairly represent all teenagers in town. This is also a convenience sample because the student surveyed the easiest group to find.

7. Comparing the Sampling Methods

  • Random sampling: usually fair and helps reduce bias
  • Systematic sampling: organized and easy, but patterns in the list can cause bias
  • Stratified sampling: good for making sure important groups are included
  • Convenience sampling: fast, but often biased

8. Tips for Test Questions

  • If you see “every 8th student,” think systematic.
  • If you see “randomly chosen from each grade,” think stratified.
  • If you see “names drawn from a hat,” think random.
  • If you see “students in one classroom” or “people at one location,” think convenience and possibly bias.

9. Final Summary

Sampling is the process of selecting part of a population to study. The four main sampling methods are random, systematic, stratified, and convenience sampling.

A good sample should represent the population fairly. Bias happens when the sample is not representative, often because some groups are more likely to be included than others.

When identifying a sampling technique, pay attention to how the sample was chosen. When checking for bias, ask whether the sample gives all parts of the population a fair chance to be represented.

Put what you read to the test

You've worked through Sampling Techniques and Bias. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Observational Studies vs. Experiments

Observational Studies vs. Experiments

In statistics, people often want to answer questions like: Does getting more sleep improve test scores? or Do students who play sports miss fewer school days? To answer questions like these, we collect and study data.

There are two common ways to do this: observational studies and experiments. They may seem similar because both involve data, but they are different in a very important way.

The main difference is this: in an observational study, researchers only observe and record what is already happening. In an experiment, researchers apply a treatment to some people or objects and compare the results.

Understanding this difference helps you decide what conclusions you can make from data.

1. What is an observational study?

An observational study is when researchers collect information without changing anything. They watch, measure, survey, or record what people already do.

For example, a researcher might ask 100 students how many hours they sleep each night and what their math grades are. The researcher does not tell anyone how much to sleep. The researcher only records the data.

Key idea: in an observational study, there is no treatment given by the researcher.

  • The researcher observes existing behavior.
  • The researcher does not assign groups.
  • The researcher does not control who gets what.

2. What is an experiment?

An experiment is when researchers deliberately change one thing and study what happens. The thing being changed is often called a treatment.

For example, suppose a teacher wants to know whether a new study app improves quiz scores. The teacher gives half the students access to the app and the other half no app, then compares quiz scores.

Here, the teacher is not just watching what students choose to do. The teacher is creating groups and giving a treatment to one group.

  • One group gets the treatment.
  • Another group may be a control group, which does not get the treatment.
  • The results are compared.

3. Control group and treatment group

In many experiments, there are at least two groups:

  • Treatment group: the group that receives the change, condition, or item being tested.
  • Control group: the group that does not receive the treatment, so researchers have something to compare against.

Suppose 40 students are in a study about a new tutoring method.

  • 20 students use the new tutoring method.
  • 20 students continue with the usual method.

The first 20 are the treatment group. The second 20 are the control group.

4. Why does the difference matter?

The type of study affects what you are allowed to conclude.

An observational study can show an association, meaning two variables appear connected. For example, students who sleep more may also have higher test scores.

But that does not always mean that one thing directly causes the other. Maybe students who sleep more also have better study habits. Maybe they have less stress. There may be other reasons.

An experiment is stronger for testing cause and effect, because the researcher controls the treatment and compares groups.

So a simple rule is:

  • Observational study: can suggest a relationship.
  • Experiment: can provide stronger evidence about cause and effect.

5. How to tell them apart

Ask yourself these questions:

  1. Did the researcher just observe what was already happening?
  2. Or did the researcher assign a treatment or create groups?

If the researcher only watched, measured, or surveyed, it is usually an observational study.

If the researcher gave some people a treatment, changed conditions, or set up control and treatment groups, it is an experiment.

6. Quick comparison

  • Observational Study
    • Researcher observes only
    • No treatment is assigned
    • No control/treatment groups created by researcher
    • Used to find patterns or relationships
  • Experiment
    • Researcher applies a treatment
    • Groups are compared
    • Often has a control group and a treatment group
    • Used to test cause and effect more directly

Worked Example 1: Basic identification

Question: A school counselor records how many hours students spend on homework each night and compares that with their grades. Is this an observational study or an experiment?

Step 1: Ask whether the counselor changed anything.

The counselor did not assign homework times. The counselor only recorded what students already do.

Answer: This is an observational study.

Why? The researcher observed existing behavior and collected data without giving a treatment.

Worked Example 2: Identifying an experiment

Question: A coach wants to know whether a new warm-up routine improves running time. The coach assigns 15 athletes to use the new warm-up and 15 athletes to use the usual warm-up, then compares times. Is this an observational study or an experiment?

Step 1: Did the coach assign a treatment?

Yes. One group used the new warm-up routine.

Step 2: Is there a comparison group?

Yes. Another group used the usual warm-up.

Answer: This is an experiment.

Why? The coach created a treatment group and a control group.

Worked Example 3: Looking at conclusions

Question: In a survey, students who eat breakfast have an average test score of 84, while students who skip breakfast have an average test score of 76. The difference is

$$84 - 76 = 8$$

Can we say that eating breakfast definitely causes higher test scores?

Step 1: Identify the type of study.

This is a survey, so the researcher is observing what students already do. That means it is an observational study.

Step 2: Decide what conclusion is reasonable.

We can say there is an association between eating breakfast and higher test scores in this data.

But we cannot say for sure that breakfast caused the higher scores, because other factors could be involved.

Answer: No, we cannot say definitely that breakfast caused the higher scores from this study alone.

Worked Example 4: Classifying a more detailed situation

Question: A science teacher wants to test whether listening to soft music during work time helps students finish assignments faster. The teacher randomly places students into two groups. One group works with soft music, and the other group works in silence. The teacher records the completion times.

Is this an observational study or an experiment? Identify the groups.

Step 1: Did the teacher change the conditions?

Yes. The teacher set up one group with music and one group without music.

Step 2: Identify the treatment.

The treatment is working with soft music.

Step 3: Identify the groups.

  • Treatment group: students working with soft music
  • Control group: students working in silence

Answer: This is an experiment.

7. Common mistakes to avoid

  • Mistake 1: Thinking every study with groups is an experiment.

    If groups already existed and the researcher only observed them, it may still be an observational study.

  • Mistake 2: Thinking surveys are experiments.

    Most surveys are observational because they collect information without assigning treatment.

  • Mistake 3: Assuming association means cause.

    Just because two things happen together does not prove that one caused the other.

8. A simple memory trick

  • Observational study = Observe
  • Experiment = Apply

If the researcher only observes, it is an observational study.

If the researcher applies a treatment, it is an experiment.

Brief Summary

Observational studies and experiments are both ways to collect data, but they are not the same. In an observational study, the researcher only watches and records what is already happening. In an experiment, the researcher gives a treatment and compares a treatment group to a control group.

Observational studies are useful for finding patterns and relationships. Experiments are better for testing cause and effect. When deciding which type of study you are looking at, always ask: Did the researcher just observe, or did they assign a treatment?

Put what you read to the test

You've worked through Observational Studies vs. Experiments. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Histograms and Distribution Shape

Histograms and Distribution Shape

When we collect data, we often want to know more than just the smallest and largest values. We want to see how the data is spread out and where most of the values are found. A histogram is a graph that helps us do this.

In this lesson, you will learn how to read and describe histograms and how to recognize common distribution shapes, including normal, skewed left, skewed right, bimodal, and uniform.

1. What is a histogram?

A histogram is a graph that shows how numerical data is distributed. It uses bars that touch because the data is grouped into intervals, called bins or class intervals.

  • The horizontal axis shows the data intervals.
  • The vertical axis shows the frequency, which means how many data values fall in each interval.

For example, if test scores are grouped into intervals like 50 to 59, 60 to 69, and 70 to 79, a histogram can show how many students scored in each range.

2. Histogram vs. bar graph

A histogram and a bar graph may look similar, but they are used for different kinds of data.

  • A bar graph is used for categories, like favorite sport or eye color.
  • A histogram is used for numerical data grouped into intervals.
  • In a histogram, the bars touch.
  • In a bar graph, the bars usually have spaces between them.

3. How to read a histogram

To read a histogram, look at the height of each bar. Taller bars mean more data values are in that interval. Shorter bars mean fewer data values are there.

When reading a histogram, ask yourself these questions:

  1. Where are most of the data values?
  2. Are the values spread out or packed together?
  3. Is there one clear peak, two peaks, or no peak?
  4. Does the graph look balanced, or does one side stretch farther?

These questions help you describe the shape of the distribution.

4. What is distribution shape?

The distribution shape is the overall pattern the histogram makes. Different shapes tell us different things about the data.

The main shapes you need to know are:

  • Normal
  • Skewed right
  • Skewed left
  • Bimodal
  • Uniform

5. Normal distribution

A normal distribution has one peak in the middle and is roughly balanced on both sides. It looks like a hill or mound.

  • Most values are near the center.
  • Fewer values are at the low and high ends.
  • The left and right sides are about the same shape.

This shape is sometimes called bell-shaped.

Examples of data that may look normal include heights of students in a large school or scores on a test when most students do about average.

6. Skewed right distribution

A distribution is skewed right when most data values are on the lower side, but a few larger values stretch out to the right.

  • The graph has a longer tail to the right.
  • Most bars are taller on the left or middle-left.
  • A few high values pull the shape toward the right.

An example might be the number of hours students spend on video games in a week. Many students may play a moderate amount, but a few play much more than everyone else.

7. Skewed left distribution

A distribution is skewed left when most data values are on the higher side, but a few smaller values stretch out to the left.

  • The graph has a longer tail to the left.
  • Most bars are taller on the right or middle-right.
  • A few low values pull the shape toward the left.

An example might be scores on an easy quiz. Most students may score high, but a few students score much lower.

8. Bimodal distribution

A bimodal distribution has two clear peaks. This often means the data may come from two different groups.

  • There are two intervals or regions with high frequencies.
  • There is often a dip between the peaks.

For example, if a class combines the heights of 9th grade boys and girls, the histogram might show two peaks because the two groups may have different typical heights.

9. Uniform distribution

A uniform distribution has bars that are about the same height across the intervals.

  • No interval stands out much more than the others.
  • The data is spread fairly evenly.

An example might be the results of rolling a fair number cube many times, where each outcome appears about equally often.

10. Important words to use when describing a histogram

When writing about a histogram, use clear statistical words. Good descriptions often mention:

  • Center: where the data seems to cluster
  • Spread: how far the data extends from low to high values
  • Peaks: where the tallest bars are
  • Skew: whether one side stretches farther
  • Gaps: intervals with no data or very little data

For example, you might say: The histogram is skewed right, with most values between 10 and 20 and a few larger values above 30.

11. Steps for identifying the shape of a histogram

  1. Look for the tallest bar or bars.
  2. Check whether the graph has one peak, two peaks, or many bars of similar height.
  3. See whether the graph is balanced or has a long tail on one side.
  4. Decide whether the shape is normal, skewed left, skewed right, bimodal, or uniform.

Worked Example 1: Reading frequencies from a histogram table

A teacher groups quiz scores into intervals and records the frequencies below.

$$ \begin{array}{c|c} \text{Score Interval} & \text{Frequency} \\ \hline 50\text{ to }59 & 2 \\ 60\text{ to }69 & 5 \\ 70\text{ to }79 & 9 \\ 80\text{ to }89 & 6 \\ 90\text{ to }99 & 3 \end{array} $$

Step 1: Find the interval with the greatest frequency.

The greatest frequency is 9, in the interval 70 to 79. That means most students scored in the 70s.

Step 2: Look at the pattern on both sides.

The frequencies rise from 2 to 5 to 9, then fall to 6 and 3. This makes a mound shape with one peak.

Conclusion: The distribution is approximately normal because it has one peak near the center and is fairly balanced.

Worked Example 2: Identifying skewed right

Suppose a histogram for the number of books read over summer has these frequencies:

$$ \begin{array}{c|c} \text{Books Read} & \text{Frequency} \\ \hline 0\text{ to }2 & 12 \\ 3\text{ to }5 & 9 \\ 6\text{ to }8 & 5 \\ 9\text{ to }11 & 2 \\ 12\text{ to }14 & 1 \end{array} $$

Step 1: Notice where most of the data is.

Most students are in the lower intervals, especially 0 to 2 and 3 to 5.

Step 2: Look for a tail.

The frequencies keep getting smaller as the values increase. The graph would stretch to the right.

Conclusion: The distribution is skewed right.

Worked Example 3: Identifying skewed left

A class takes a very easy vocabulary quiz. The grouped scores are:

$$ \begin{array}{c|c} \text{Score Interval} & \text{Frequency} \\ \hline 50\text{ to }59 & 1 \\ 60\text{ to }69 & 2 \\ 70\text{ to }79 & 4 \\ 80\text{ to }89 & 8 \\ 90\text{ to }99 & 10 \end{array} $$

Step 1: Find where most scores are.

Most scores are high, especially in the 80 to 89 and 90 to 99 intervals.

Step 2: Check the tail.

There are only a few low scores, so the graph would have a tail stretching to the left.

Conclusion: The distribution is skewed left.

Worked Example 4: Bimodal or uniform?

Consider these grouped data values:

$$ \begin{array}{c|c} \text{Interval} & \text{Frequency} \\ \hline 0\text{ to }4 & 7 \\ 5\text{ to }9 & 2 \\ 10\text{ to }14 & 1 \\ 15\text{ to }19 & 6 \\ 20\text{ to }24 & 7 \end{array} $$

Step 1: Look for peaks.

There are high frequencies in 0 to 4 and 20 to 24, with another high interval at 15 to 19. The middle intervals are much lower.

Step 2: Decide if the bars are all similar height.

No. The bars are not about the same height, so it is not uniform.

Step 3: Decide if there are two high regions.

Yes. There are two clear clusters of higher frequencies separated by lower frequencies.

Conclusion: The distribution is bimodal.

12. Common mistakes to avoid

  • Do not confuse skew direction with where most data is. In a skewed right graph, the tail points right, even though most data is on the left.
  • Do not call every graph with one peak normal. A normal distribution should be roughly balanced.
  • Do not ignore gaps. A gap can suggest separate groups or unusual data.
  • Do not use a bar graph for grouped numerical data. A histogram is the correct display.

13. Quick shape guide

  • Normal: one middle peak, balanced sides
  • Skewed right: tail to the right
  • Skewed left: tail to the left
  • Bimodal: two peaks
  • Uniform: bars about the same height

14. Why histogram shape matters

The shape of a distribution helps us understand what is happening in real situations. It can show whether most values are typical, whether there are extreme values, or whether the data may come from more than one group.

For example, if a histogram of test scores is skewed left, that may mean the test was easy. If it is skewed right, that may mean the test was hard for many students. If it is bimodal, that may mean there were two different levels of preparation among students.

Summary

A histogram shows the frequency of numerical data grouped into intervals. The bars touch because the intervals are connected.

You can describe a histogram by looking at its center, spread, peaks, and whether it has a tail on one side. The main shapes to recognize are normal, skewed right, skewed left, bimodal, and uniform.

If you practice looking for peaks and tails, you will get better at quickly identifying the shape of a distribution and explaining what the data means.

Put what you read to the test

You've worked through Histograms and Distribution Shape. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Measures of Central Tendency

Measures of Central Tendency are ways to describe the center of a data set using a single value. In 9th Grade Maths, the three main measures of central tendency are the mean, median, and mode.

These measures help us answer questions like: “What is a typical value?” or “What number best represents the whole set of data?”

For example, if a class takes a quiz, the teacher may want one number that summarizes how the class performed. Depending on the data, the mean, median, or mode may be the best choice.

Why this matters: In real life, statistics are used in sports, business, science, and everyday decision-making. Choosing the right measure of center helps us understand data more accurately.

The three main measures are:

  • Mean: the average
  • Median: the middle value
  • Mode: the most frequent value

Let’s study each one carefully.

1. Mean

The mean is what most people call the average. To find it, add all the values in the data set and divide by the number of values.

The formula is:

$$\text{Mean} = \frac{\text{sum of all data values}}{\text{number of data values}}$$

Example: Find the mean of the numbers \(4, 6, 8, 10, 12\).

First add the numbers:

$$4+6+8+10+12=40$$

There are \(5\) numbers, so divide by \(5\):

$$\text{Mean}=\frac{40}{5}=8$$

So, the mean is 8.

Important note: The mean uses every value in the data set. Because of this, very large or very small values can change it a lot.

2. Median

The median is the middle value when the data is arranged in order from least to greatest.

Steps to find the median:

  1. Put the numbers in order.
  2. Find the middle number.
  3. If there are two middle numbers, add them and divide by \(2\).

Example with an odd number of values: Find the median of \(3, 7, 9, 11, 15\).

The numbers are already in order. There are \(5\) values, so the middle is the third number.

The median is 9.

Example with an even number of values: Find the median of \(2, 4, 8, 10\).

There are \(4\) values, so there are two middle numbers: \(4\) and \(8\).

Find their average:

$$\frac{4+8}{2}=\frac{12}{2}=6$$

So, the median is 6.

3. Mode

The mode is the value that appears most often in a data set.

Example: Find the mode of \(5, 7, 7, 9, 10\).

The number \(7\) appears twice, while the others appear once.

So, the mode is 7.

A data set can have:

  • One mode if one value appears most often
  • More than one mode if multiple values tie for appearing most often
  • No mode if no value repeats

Example of two modes: In \(2, 3, 3, 5, 5, 8\), both \(3\) and \(5\) appear twice. The set has two modes: \(3\) and \(5\).

Comparing mean, median, and mode

Each measure tells us something about the center of the data, but they are not always the same.

Consider the data set \(2, 4, 4, 6, 20\).

  • Mean: $$\frac{2+4+4+6+20}{5}=\frac{36}{5}=7.2$$
  • Median: the middle value is \(4\)
  • Mode: the most frequent value is \(4\)

Notice that the mean is \(7.2\), which is higher than most of the numbers. This happens because \(20\) is much larger than the other values.

A value that is much larger or much smaller than the rest is called an outlier. Outliers can pull the mean away from the center, but the median is usually less affected.

When should you use each measure?

  • Use the mean when the data does not have extreme values and you want to use every number in the set.
  • Use the median when the data has outliers or is not evenly spread.
  • Use the mode when you want to know the most common value.

Worked Example 1: Finding all three measures

Find the mean, median, and mode of \(6, 8, 8, 10, 13\).

Step 1: Mean

Add the values:

$$6+8+8+10+13=45$$

There are \(5\) numbers:

$$\text{Mean}=\frac{45}{5}=9$$

Step 2: Median

The numbers are already in order: \(6, 8, 8, 10, 13\).

The middle value is the third number, so the median is 8.

Step 3: Mode

The number \(8\) appears twice. The others appear once.

So, the mode is 8.

Answer:

  • Mean = \(9\)
  • Median = \(8\)
  • Mode = \(8\)

Worked Example 2: Median with an even number of values

Find the mean, median, and mode of \(4, 6, 9, 11, 11, 14\).

Step 1: Mean

$$4+6+9+11+11+14=55$$

There are \(6\) numbers:

$$\text{Mean}=\frac{55}{6}\approx 9.17$$

Step 2: Median

There are \(6\) values, so use the two middle numbers: \(9\) and \(11\).

$$\text{Median}=\frac{9+11}{2}=10$$

Step 3: Mode

The number \(11\) appears twice, more than any other number.

So, the mode is 11.

Answer:

  • Mean = \(\frac{55}{6}\approx 9.17\)
  • Median = \(10\)
  • Mode = \(11\)

Worked Example 3: Choosing the best measure

A basketball player scores these points in 6 games:

\(12, 14, 15, 15, 16, 40\)

Find the mean, median, and mode. Then decide which measure best represents the player’s usual score.

Step 1: Mean

$$12+14+15+15+16+40=112$$

$$\text{Mean}=\frac{112}{6}\approx 18.67$$

Step 2: Median

There are \(6\) values, so use the two middle values: \(15\) and \(15\).

$$\text{Median}=\frac{15+15}{2}=15$$

Step 3: Mode

The number \(15\) appears twice, so the mode is 15.

Step 4: Best measure

The score \(40\) is much higher than the rest, so it is an outlier. It pulls the mean up to about \(18.67\), even though most scores are near \(15\).

The median is the best measure of center here because it is not heavily affected by the outlier.

Worked Example 4: No mode

Find the mean, median, and mode of \(1, 3, 5, 7, 9\).

Mean:

$$\frac{1+3+5+7+9}{5}=\frac{25}{5}=5$$

Median:

The middle value is \(5\), so the median is 5.

Mode:

No number repeats, so there is no mode.

Common mistakes to avoid

  • Not putting data in order before finding the median
  • Forgetting to divide by the number of values when finding the mean
  • Choosing the largest number as the mode instead of the most frequent number
  • Ignoring outliers when deciding whether the mean is a good choice

Quick check questions

  1. What is the mean of \(2, 4, 6, 8\)?
  2. What is the median of \(1, 3, 7, 9, 11\)?
  3. What is the mode of \(4, 4, 5, 6, 6, 6, 8\)?
  4. Which measure is usually best when there is an outlier?

Answers:

  1. $$\frac{2+4+6+8}{4}=\frac{20}{4}=5$$
  2. The median is \(7\).
  3. The mode is \(6\).
  4. The median is usually best.

Summary

The mean is the average, found by adding all the values and dividing by how many there are. The median is the middle value when the numbers are arranged in order. The mode is the value that appears most often.

To choose the best measure, think about the data. If there are no extreme values, the mean is often useful. If there is an outlier, the median is usually more reliable. If you want the most common value, use the mode.

Understanding these three measures helps you describe data clearly and decide which number best represents the center of a data set.

Put what you read to the test

You've worked through Measures of Central Tendency. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Measures of Dispersion (Range, IQR, Standard Deviation)

Measures of Dispersion: Range, IQR, and Standard Deviation

When we look at a set of data, we do not only want to know the center of the data, such as the mean or median. We also want to know how spread out the values are. This spread is called dispersion or variability.

For example, two classes might both have an average test score of 75, but in one class most students scored close to 75, while in the other class some scored very high and some very low. The averages are the same, but the spread is different.

In this lesson, you will learn three important measures of dispersion:

  • Range
  • Interquartile Range (IQR)
  • Standard Deviation

Each one tells us something useful about how data is spread out.

1. Range

The range is the difference between the largest value and the smallest value in a data set.

Formula:

$$\text{Range} = \text{maximum} - \text{minimum}$$

The range is the simplest measure of spread. It gives a quick idea of how wide the data set is.

Example 1: Finding the Range

The data set is:

$$4,\ 7,\ 9,\ 10,\ 13$$

The smallest value is 4 and the largest value is 13.

$$\text{Range} = 13 - 4 = 9$$

So, the range is 9.

Why range is useful:

  • It is fast and easy to calculate.
  • It gives an overall picture of spread.

Limitation of range:

  • It only uses the smallest and largest values.
  • It can be greatly affected by one unusual value, called an outlier.

Because of this, range is helpful, but sometimes we need a measure that looks at the middle part of the data more carefully.

2. Interquartile Range (IQR)

The interquartile range, or IQR, measures the spread of the middle half of the data.

To find the IQR, we use quartiles:

  • Q1: the lower quartile, or the middle of the lower half of the data
  • Q3: the upper quartile, or the middle of the upper half of the data

Formula:

$$\text{IQR} = Q_3 - Q_1$$

The IQR is useful because it ignores the very smallest and very largest values. This means it is less affected by outliers than the range.

Steps for finding the IQR

  1. Put the data in order from smallest to largest.
  2. Find the median.
  3. Find the median of the lower half. This is \(Q_1\).
  4. Find the median of the upper half. This is \(Q_3\).
  5. Subtract: \(Q_3 - Q_1\).

Example 2: Finding the IQR

The data set is:

$$2,\ 4,\ 5,\ 7,\ 8,\ 10,\ 12$$

The data is already in order.

Step 1: Find the median.

The middle number is 7, so the median is 7.

Step 2: Find the lower half and upper half.

Lower half: $$2,\ 4,\ 5$$

Upper half: $$8,\ 10,\ 12$$

Step 3: Find \(Q_1\) and \(Q_3\).

The middle of the lower half is 4, so \(Q_1 = 4\).

The middle of the upper half is 10, so \(Q_3 = 10\).

Step 4: Calculate the IQR.

$$\text{IQR} = Q_3 - Q_1 = 10 - 4 = 6$$

So, the IQR is 6.

Important note: When the data set has an odd number of values, the median itself is not included in either half when finding \(Q_1\) and \(Q_3\).

Example 3: IQR with an even number of values

The data set is:

$$3,\ 5,\ 6,\ 8,\ 9,\ 11,\ 14,\ 15$$

Step 1: Find the median.

There are 8 values, so the median is the average of the 4th and 5th values:

$$\text{Median} = \frac{8 + 9}{2} = 8.5$$

Step 2: Split the data into two halves.

Lower half: $$3,\ 5,\ 6,\ 8$$

Upper half: $$9,\ 11,\ 14,\ 15$$

Step 3: Find \(Q_1\) and \(Q_3\).

$$Q_1 = \frac{5 + 6}{2} = 5.5$$

$$Q_3 = \frac{11 + 14}{2} = 12.5$$

Step 4: Find the IQR.

$$\text{IQR} = 12.5 - 5.5 = 7$$

So, the IQR is 7.

Why IQR is useful:

  • It shows the spread of the middle 50% of the data.
  • It is not strongly affected by outliers.
  • It works well when data has a few unusual values.

3. Standard Deviation

The standard deviation tells us how far the data values usually are from the mean.

If the standard deviation is small, the data values are close to the mean.

If the standard deviation is large, the data values are spread farther away from the mean.

Standard deviation uses all the values in the data set, so it gives a fuller picture of spread than the range.

For 9th Grade, the main goal is to understand what standard deviation means and how to calculate it for simple data sets.

Steps for finding standard deviation

  1. Find the mean.
  2. Subtract the mean from each value to find each deviation.
  3. Square each deviation.
  4. Find the mean of those squared deviations. This is called the variance.
  5. Take the square root of the variance.

So the formula is:

$$\text{Standard Deviation} = \sqrt{\frac{\text{sum of squared deviations}}{\text{number of values}}}$$

In symbols, for a data set with values \(x_1, x_2, x_3, \dots, x_n\) and mean \(\bar{x}\):

$$\sigma = \sqrt{\frac{(x_1-\bar{x})^2+(x_2-\bar{x})^2+\cdots+(x_n-\bar{x})^2}{n}}$$

You do not need to worry about the symbol names too much. Focus on the process.

Example 4: Finding Standard Deviation

Find the standard deviation of:

$$2,\ 4,\ 6$$

Step 1: Find the mean.

$$\bar{x} = \frac{2+4+6}{3} = \frac{12}{3} = 4$$

Step 2: Find each deviation from the mean.

  • For 2: \(2-4=-2\)
  • For 4: \(4-4=0\)
  • For 6: \(6-4=2\)

Step 3: Square each deviation.

  • \((-2)^2 = 4\)
  • \(0^2 = 0\)
  • \(2^2 = 4\)

Step 4: Find the mean of the squared deviations.

$$\text{Variance} = \frac{4+0+4}{3} = \frac{8}{3}$$

Step 5: Take the square root.

$$\text{Standard Deviation} = \sqrt{\frac{8}{3}} \approx 1.63$$

So, the standard deviation is about 1.63.

What does this mean?

It means the data values are usually about 1.63 units away from the mean of 4.

Comparing small and large spread

Look at these two data sets:

Set A: $$9,\ 10,\ 11$$

Set B: $$2,\ 10,\ 18$$

Both sets have the same mean:

$$\frac{9+10+11}{3} = 10$$

$$\frac{2+10+18}{3} = 10$$

But Set A is tightly grouped around 10, while Set B is much more spread out. So Set B has a larger standard deviation.

This shows why standard deviation is useful. It helps us see how closely the values cluster around the mean.

When to use each measure

  • Range: use for a quick idea of overall spread.
  • IQR: use when you want the spread of the middle 50% or when there may be outliers.
  • Standard deviation: use when you want to know how far values usually are from the mean.

A quick comparison

  • Range uses only the smallest and largest values.
  • IQR uses the middle half of the data.
  • Standard deviation uses every value in the data set.

Common mistakes to avoid

  • Do not forget to put the data in order before finding the median or quartiles.
  • Do not include the median in the lower and upper halves when there is an odd number of data values.
  • For range, always subtract smallest from largest.
  • For standard deviation, subtract the mean from each value before squaring.
  • Do not confuse variance with standard deviation. Standard deviation is the square root of the variance.

Why measures of dispersion matter

In real life, averages alone do not tell the whole story. A coach may want to know whether players' times are consistent. A teacher may want to know whether test scores are close together or very spread out. A company may want to know whether daily sales stay steady or change a lot.

Measures of dispersion help us answer these questions by showing how much variation there is in the data.

Lesson Summary

Range is the difference between the maximum and minimum values:

$$\text{Range} = \text{maximum} - \text{minimum}$$

IQR is the difference between the upper quartile and lower quartile:

$$\text{IQR} = Q_3 - Q_1$$

Standard deviation tells how far values usually are from the mean.

These three measures all describe spread, but they do it in different ways. Range is quick, IQR focuses on the middle half, and standard deviation shows how much the whole data set varies around the mean.

When you study a data set, it is often helpful to look at both a measure of center and a measure of dispersion. Together, they give a much clearer picture of the data.

Put what you read to the test

You've worked through Measures of Dispersion (Range, IQR, Standard Deviation). Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Box Plots and the 1.5 IQR Outlier Rule

Box Plots and the 1.5 IQR Outlier Rule

When we collect a set of numbers, we often want to understand where the data is centered, how spread out it is, and whether any values are unusually far away from the rest.

A box plot is a graph that helps us do all of that quickly. It is built from the five-number summary, and it can also show outliers using the 1.5 IQR rule.

In this lesson, you will learn how to:

  • find the five-number summary,
  • draw and read a box plot,
  • calculate the interquartile range, or IQR,
  • use the 1.5 IQR outlier rule,
  • decide which values are part of the main data and which are outliers.

1. The five-number summary

A box plot is based on five important values in a data set:

  • Minimum: the smallest value
  • First quartile \,\(Q_1\): the middle of the lower half of the data
  • Median \,\(Q_2\): the middle value of the whole data set
  • Third quartile \,\(Q_3\): the middle of the upper half of the data
  • Maximum: the largest value

These five values give a picture of how the data is spread out.

Important first step: Before finding any of these values, always put the data in numerical order from least to greatest.

2. What quartiles mean

The word quartile comes from a word meaning “four parts.” Quartiles split ordered data into four sections.

  • The median splits the whole data set into a lower half and an upper half.
  • \(Q_1\) is the median of the lower half.
  • \(Q_3\) is the median of the upper half.

If the data set has an odd number of values, the middle value is the median, and it is not included in either half when finding \(Q_1\) and \(Q_3\).

3. What a box plot shows

A box plot uses the five-number summary to make a simple graph.

  • The box goes from \(Q_1\) to \(Q_3\).
  • A line inside the box shows the median.
  • The parts stretching outward are called whiskers.

If there are no outliers, the whiskers usually go to the minimum and maximum. If there are outliers, the whiskers go only to the smallest and largest values that are not outliers. The outliers are marked separately.

4. The interquartile range (IQR)

The interquartile range, or IQR, measures the spread of the middle 50% of the data.

It is found by subtracting \(Q_1\) from \(Q_3\):

$$IQR = Q_3 - Q_1$$

A small IQR means the middle half of the data is packed closely together. A larger IQR means it is more spread out.

5. The 1.5 IQR outlier rule

An outlier is a value that is unusually far from the rest of the data.

To check for outliers, use these formulas:

$$\text{Lower fence} = Q_1 - 1.5(IQR)$$

$$\text{Upper fence} = Q_3 + 1.5(IQR)$$

Then compare each data value to these fences:

  • Any value less than the lower fence is an outlier.
  • Any value greater than the upper fence is an outlier.
  • Values between the fences are not outliers.

The fences are not always actual data values. They are just boundary numbers used to test whether a point is unusually far away.

6. Steps for solving box plot and outlier problems

  1. Put the data in order.
  2. Find the median.
  3. Find \(Q_1\) and \(Q_3\).
  4. Compute the IQR using \(Q_3 - Q_1\).
  5. Find the lower and upper fences using the 1.5 IQR rule.
  6. Identify any outliers.
  7. Use the non-outlier values to decide where the whiskers end.

Worked Example 1: Finding the five-number summary

Find the five-number summary for this data set:

\(3, 5, 7, 8, 10, 12, 13, 15, 18\)

Step 1: The data is already in order.

There are 9 values, so the median is the 5th value.

$$\text{Median} = 10$$

Step 2: Find the lower half and upper half.

Since there are an odd number of values, do not include the median in either half.

Lower half: \(3, 5, 7, 8\)

Upper half: \(12, 13, 15, 18\)

Step 3: Find \(Q_1\) and \(Q_3\).

For the lower half, the middle two values are 5 and 7, so:

$$Q_1 = \frac{5+7}{2} = 6$$

For the upper half, the middle two values are 13 and 15, so:

$$Q_3 = \frac{13+15}{2} = 14$$

Step 4: Find minimum and maximum.

Minimum \(= 3\)

Maximum \(= 18\)

Five-number summary:

  • Minimum: 3
  • \(Q_1 = 6\)
  • Median: 10
  • \(Q_3 = 14\)
  • Maximum: 18

Worked Example 2: Using the 1.5 IQR rule

Use the data set below to decide whether there are any outliers:

\(2, 4, 5, 6, 7, 8, 9, 10, 25\)

Step 1: Find the median.

There are 9 values, so the median is the 5th value:

$$\text{Median} = 7$$

Step 2: Find \(Q_1\) and \(Q_3\).

Lower half: \(2, 4, 5, 6\)

Upper half: \(8, 9, 10, 25\)

$$Q_1 = \frac{4+5}{2} = 4.5$$

$$Q_3 = \frac{9+10}{2} = 9.5$$

Step 3: Find the IQR.

$$IQR = Q_3 - Q_1 = 9.5 - 4.5 = 5$$

Step 4: Find the fences.

$$\text{Lower fence} = 4.5 - 1.5(5) = 4.5 - 7.5 = -3$$

$$\text{Upper fence} = 9.5 + 1.5(5) = 9.5 + 7.5 = 17$$

Step 5: Check the data values.

Any value less than \(-3\) or greater than \(17\) is an outlier.

The value \(25\) is greater than \(17\), so it is an outlier.

None of the other values are outside the fences.

Conclusion: The data set has one outlier, which is \(25\).

On a box plot, the right whisker would stop at \(10\), and the value \(25\) would be plotted as a separate point.

Worked Example 3: Even number of data values

Find the box plot values and any outliers for:

\(11, 12, 14, 15, 16, 18, 19, 22\)

Step 1: The data is already in order.

There are 8 values, so the median is the average of the 4th and 5th values:

$$\text{Median} = \frac{15+16}{2} = 15.5$$

Step 2: Split into halves.

Lower half: \(11, 12, 14, 15\)

Upper half: \(16, 18, 19, 22\)

Step 3: Find \(Q_1\) and \(Q_3\).

$$Q_1 = \frac{12+14}{2} = 13$$

$$Q_3 = \frac{18+19}{2} = 18.5$$

Step 4: Find the IQR.

$$IQR = 18.5 - 13 = 5.5$$

Step 5: Find the fences.

$$\text{Lower fence} = 13 - 1.5(5.5) = 13 - 8.25 = 4.75$$

$$\text{Upper fence} = 18.5 + 1.5(5.5) = 18.5 + 8.25 = 26.75$$

All values are between \(4.75\) and \(26.75\), so there are no outliers.

Five-number summary:

  • Minimum: 11
  • \(Q_1 = 13\)
  • Median: 15.5
  • \(Q_3 = 18.5\)
  • Maximum: 22

Since there are no outliers, the whiskers would go from 11 to 22.

Worked Example 4: Reading what a box plot means

Suppose a data set has this five-number summary:

  • Minimum: 20
  • \(Q_1 = 28\)
  • Median: 35
  • \(Q_3 = 42\)
  • Maximum: 60

Let us test whether the maximum might be an outlier.

$$IQR = 42 - 28 = 14$$

$$\text{Lower fence} = 28 - 1.5(14) = 28 - 21 = 7$$

$$\text{Upper fence} = 42 + 1.5(14) = 42 + 21 = 63$$

The maximum is \(60\), and \(60\) is less than \(63\), so it is not an outlier.

This tells us the box plot would have:

  • a box from 28 to 42,
  • a median line at 35,
  • a left whisker to 20,
  • a right whisker to 60.

7. Common mistakes to avoid

  • Forgetting to order the data first. Quartiles and median must be found from ordered data.
  • Including the median in both halves when the data set has an odd number of values. Do not include it in either half.
  • Using the minimum and maximum to find IQR. IQR uses only \(Q_1\) and \(Q_3\).
  • Thinking the whiskers always go to the minimum and maximum. If there are outliers, whiskers stop at the most extreme non-outlier values.
  • Mixing up the formulas. Remember:

$$IQR = Q_3 - Q_1$$

$$\text{Lower fence} = Q_1 - 1.5(IQR)$$

$$\text{Upper fence} = Q_3 + 1.5(IQR)$$

8. How box plots help us compare data

Box plots are useful because they show several things at once:

  • the center of the data using the median,
  • the spread of the middle half using the box,
  • whether the data appears more spread out on one side than the other,
  • whether there are outliers.

If one box plot has a much larger box than another, its middle 50% has a larger spread. If one has outliers and the other does not, that tells you one set has more extreme values.

9. Quick checklist

When solving a question, ask yourself:

  • Did I put the data in order?
  • Did I find the median correctly?
  • Did I split the data into lower and upper halves correctly?
  • Did I compute \(Q_1\) and \(Q_3\) correctly?
  • Did I calculate \(IQR\) correctly?
  • Did I use the 1.5 IQR rule to test for outliers?

Summary

A box plot is built from the five-number summary: minimum, \(Q_1\), median, \(Q_3\), and maximum. The box runs from \(Q_1\) to \(Q_3\), and the median is shown inside the box.

The interquartile range is found with \(IQR = Q_3 - Q_1\). To find outliers, use the 1.5 IQR rule:

$$\text{Lower fence} = Q_1 - 1.5(IQR), \qquad \text{Upper fence} = Q_3 + 1.5(IQR)$$

Any value outside these fences is an outlier. Once you can find quartiles, IQR, and fences, you can read and create box plots with confidence.

Put what you read to the test

You've worked through Box Plots and the 1.5 IQR Outlier Rule. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Normal Distribution Basics and Empirical Rule

Normal Distribution Basics and the Empirical Rule

In statistics, we often collect data and then look for patterns. Some sets of data cluster around a middle value, with fewer values far away from the middle. One very important pattern like this is called the normal distribution.

A normal distribution is a bell-shaped curve. It is highest in the middle and falls off evenly on both sides. Many real-life measurements are often modeled this way, such as heights, test scores, and measurement errors.

In this lesson, you will learn:

  • what a normal distribution looks like,
  • what the mean and standard deviation tell us,
  • how the Empirical Rule works, and
  • how z-scores help compare values from different data sets.

1. What is a normal distribution?

A normal distribution is a smooth curve with these important features:

  • It is symmetric, which means the left and right sides match.
  • The center of the curve is the mean, or average.
  • Most data values are close to the mean.
  • Fewer data values appear as you move farther away from the mean.

If a data set is normally distributed, the mean, median, and mode are all at the center or very close to it.

2. Mean and standard deviation

The mean is the average of the data. It tells us the center of the distribution.

The standard deviation tells us how spread out the data is. A small standard deviation means the data values are packed closely around the mean. A large standard deviation means the data is spread farther out.

We often use these symbols:

  • Mean: \(\mu\)
  • Standard deviation: \(\sigma\)

For a normal distribution, the mean is at the very center, and we measure distances from the center using standard deviations.

3. The Empirical Rule

The Empirical Rule is also called the 68-95-99.7 Rule. It tells us how data is spread in a normal distribution.

For a normal distribution:

  • About 68% of the data lies within 1 standard deviation of the mean.
  • About 95% of the data lies within 2 standard deviations of the mean.
  • About 99.7% of the data lies within 3 standard deviations of the mean.

In symbols, this means:

About 68% is between \(\mu-\sigma\) and \(\mu+\sigma\).

About 95% is between \(\mu-2\sigma\) and \(\mu+2\sigma\).

About 99.7% is between \(\mu-3\sigma\) and \(\mu+3\sigma\).

This rule is useful because it helps us quickly estimate how common or unusual a value is.

4. Visualizing the bell curve

Imagine the mean in the center. Then mark off standard deviations on both sides:

$$\mu-3\sigma \quad \mu-2\sigma \quad \mu-\sigma \quad \mu \quad \mu+\sigma \quad \mu+2\sigma \quad \mu+3\sigma$$

The percentages are spread like this:

  • About 34% from the mean to \(+1\sigma\)
  • About 34% from the mean to \(-1\sigma\)
  • About 13.5% between \(1\sigma\) and \(2\sigma\) on each side
  • About 2.35% between \(2\sigma\) and \(3\sigma\) on each side
  • Only about 0.15% beyond \(3\sigma\) on each side

That means values very far from the mean are rare.

5. Why the Empirical Rule matters

The Empirical Rule helps answer questions like:

  • Is this score typical or unusual?
  • What range contains most of the data?
  • About how many data values should fall in a certain interval?

It is especially helpful when exact data values are not listed, but the mean and standard deviation are known.

Worked Example 1: Finding the 68%, 95%, and 99.7% intervals

A set of test scores is normally distributed with mean \(\mu=80\) and standard deviation \(\sigma=5\).

Find the intervals for:

  • about 68% of scores,
  • about 95% of scores,
  • about 99.7% of scores.

Step 1: Find 1 standard deviation from the mean

$$80-5=75 \qquad 80+5=85$$

So about 68% of scores are between 75 and 85.

Step 2: Find 2 standard deviations from the mean

$$80-2(5)=80-10=70$$

$$80+2(5)=80+10=90$$

So about 95% of scores are between 70 and 90.

Step 3: Find 3 standard deviations from the mean

$$80-3(5)=80-15=65$$

$$80+3(5)=80+15=95$$

So about 99.7% of scores are between 65 and 95.

Answer:

  • 68%: 75 to 85
  • 95%: 70 to 90
  • 99.7%: 65 to 95

Worked Example 2: Estimating how many values fall in a range

The heights of 200 plants are approximately normally distributed with mean \(50\) cm and standard deviation \(4\) cm.

About how many plants have heights between \(46\) cm and \(54\) cm?

Step 1: Compare the endpoints to the mean

The mean is \(50\), and one standard deviation is \(4\).

$$50-4=46 \qquad 50+4=54$$

So the range 46 to 54 is within 1 standard deviation of the mean.

Step 2: Use the Empirical Rule

About 68% of the data lies within 1 standard deviation of the mean.

Step 3: Find 68% of 200

$$0.68 \times 200 = 136$$

Answer: About 136 plants have heights between 46 cm and 54 cm.

Worked Example 3: Finding how unusual a value is with a z-score

A z-score tells how many standard deviations a value is above or below the mean.

The formula is:

$$z=\frac{x-\mu}{\sigma}$$

where:

  • \(x\) is the data value,
  • \(\mu\) is the mean,
  • \(\sigma\) is the standard deviation.

Suppose a quiz score is \(92\), with mean \(80\) and standard deviation \(6\). Find the z-score.

Step 1: Substitute into the formula

$$z=\frac{92-80}{6}=\frac{12}{6}=2$$

Answer: The z-score is 2.

This means the score of 92 is 2 standard deviations above the mean.

Since it is 2 standard deviations above the mean, it is higher than most scores. By the Empirical Rule, about 95% of scores are within 2 standard deviations of the mean, so a score at \(+2\sigma\) is fairly high and somewhat unusual, but not extremely rare.

6. Understanding z-scores

Z-scores help standardize values. This means we can compare values from different situations, even if the original units are different.

Here is how to interpret a z-score:

  • \(z=0\): the value is exactly at the mean
  • Positive z-score: the value is above the mean
  • Negative z-score: the value is below the mean
  • Larger absolute value of \(z\): the value is farther from the mean

Examples:

  • \(z=1\) means 1 standard deviation above the mean
  • \(z=-1.5\) means 1.5 standard deviations below the mean
  • \(z=3\) means very far above the mean

Worked Example 4: Comparing two different performances

Jordan scored 84 on a math test where the mean was 76 and the standard deviation was 4.

Taylor scored 90 on a science test where the mean was 82 and the standard deviation was 8.

Who performed better compared to their class?

Step 1: Find Jordan's z-score

$$z=\frac{84-76}{4}=\frac{8}{4}=2$$

Step 2: Find Taylor's z-score

$$z=\frac{90-82}{8}=\frac{8}{8}=1$$

Step 3: Compare the z-scores

Jordan's z-score is 2. Taylor's z-score is 1.

Answer: Jordan performed better compared to the class because Jordan's score is 2 standard deviations above the mean, while Taylor's is only 1 standard deviation above the mean.

7. Connecting z-scores to the Empirical Rule

The Empirical Rule can also be written using z-scores:

  • About 68% of values have z-scores between \(-1\) and \(1\).
  • About 95% of values have z-scores between \(-2\) and \(2\).
  • About 99.7% of values have z-scores between \(-3\) and \(3\).

So if a value has a z-score of \(2.5\), it is outside the 95% range and is fairly unusual.

If a value has a z-score of \(0.3\), it is very close to the mean and is quite typical.

8. Common mistakes to avoid

  • Mixing up mean and standard deviation: The mean is the center. The standard deviation measures spread.
  • Forgetting both sides of the mean: The normal distribution is symmetric, so the same ideas apply above and below the mean.
  • Using the Empirical Rule for any shape of data: This rule works for data that is approximately normal, or bell-shaped.
  • Reading z-scores backwards: A positive z-score is above the mean, and a negative z-score is below the mean.

9. Quick review steps

When solving problems about a normal distribution, use this process:

  1. Identify the mean \((\mu)\) and standard deviation \((\sigma)\).
  2. Decide how many standard deviations away from the mean you need.
  3. Use the Empirical Rule if the problem asks for an approximate percent or count.
  4. Use the z-score formula if the problem asks how far a value is from the mean in standard deviation units.

Brief Summary

A normal distribution is a bell-shaped, symmetric data pattern centered at the mean. The standard deviation tells how spread out the data is.

The Empirical Rule says that about 68% of data is within 1 standard deviation of the mean, 95% is within 2, and 99.7% is within 3. Z-scores show how many standard deviations a value is from the mean, making it easier to tell whether a value is typical, unusual, or to compare scores from different data sets.

Put what you read to the test

You've worked through Normal Distribution Basics and Empirical Rule. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.