Chapter 11

Data Analysis and Statistics

Statistical Questions and Data Collection

Statistical Questions and Data Collection

In math, we sometimes ask questions so we can learn about a group. To answer these questions, we collect data. Data is information, such as numbers, choices, or counts.

But not every question is a statistical question. A statistical question is a question that expects different answers from different people or objects in a group.

For example, if you ask, “How many letters are in the word cat?” there is only one answer: 3. That is not a statistical question.

If you ask, “How many letters are in each student’s first name in our class?” the answers will probably be different. Some students may have 3 letters, some 5, and some 7. That is a statistical question.

When we ask a statistical question, we usually need to collect data from a group. Then we can study the answers to look for patterns, compare results, and understand what is typical.

What makes a question statistical?

  • It is about a group, not just one person or one thing.
  • It expects variation, which means the answers can be different.
  • It can be answered by collecting data.

Examples of statistical questions:

  • How many books did students in our class read last month?
  • What is the favorite fruit of students in 5th grade?
  • How many minutes do students at our school spend reading each night?

Examples of non-statistical questions:

  • How old is Maria?
  • What is the color of our classroom door?
  • How many days are in a week?

These are not statistical questions because they ask for one answer, not a set of different answers.

Understanding variation

Variation means the data values are not all the same. In real life, people and objects are often different in some way.

For example, if you ask, “How tall are the students in our class?” you should expect different heights. That difference is variation.

Variation is important because it helps us know that a question is statistical. If all answers must be the same, then the question is not statistical.

Collecting data

After writing a statistical question, the next step is to collect data carefully. Good data collection helps us get useful answers.

There are different ways to collect data:

  • Survey: asking people questions and recording their answers.
  • Count: finding how many of each kind there are.
  • Measure: finding length, height, time, or weight.
  • Observe: watching and recording what happens.

Steps for collecting data

  1. Write a clear statistical question.
  2. Decide who or what you will collect data from.
  3. Choose a fair way to gather the data.
  4. Record the data carefully.
  5. Check that your data matches the question.

Who should you ask?

The group you want to learn about is called a population. For example, if you want to know about all 5th grade students in your school, then all 5th grade students are the population.

Sometimes it is hard to ask everyone in a population. Instead, we ask a smaller group called a sample.

A sample should represent the larger group fairly. That means it should not include only one type of student if the whole group is more mixed.

Random sampling

One fair way to choose a sample is random sampling. Random means every person in the group has an equal chance of being chosen.

For example, if there are 100 students in 5th grade, you could put all 100 names in a box and draw 20 names without looking. That is a random sample.

Random sampling helps make the data more fair because you are not just choosing your friends or the students sitting nearest to you.

Why fair data matters

If you only survey one small part of a group, your results may not match the whole group.

For example, suppose you ask, “What is the favorite school lunch of 5th graders?” If you only ask students in one classroom, your answers may not represent all 5th graders.

If you ask students from several classes using random sampling, your data will likely be more useful.

Writing good survey questions

A good survey question should be clear and easy to understand. It should match exactly what you want to learn.

Here are some tips:

  • Use simple words.
  • Ask only one thing at a time.
  • Do not make the question confusing.
  • Make sure the answers can vary.

For example, “What is your favorite recess game?” is a clear question. But “Do you like fun games at recess and lunch?” is harder to answer because it asks about more than one time.

Worked Example 1: Is it statistical?

Question: “How many siblings does each student in our class have?”

Step 1: Is it about a group? Yes, it is about students in our class.

Step 2: Will the answers vary? Yes. Some students may have 0 siblings, some 1, some 2, and some more.

Step 3: Can data be collected? Yes, we can ask each student and record the answers.

Answer: This is a statistical question.

Now look at this question: “How many siblings does Ava have?”

This asks about one person and expects one answer. It is not a statistical question.

Worked Example 2: Choosing a good sample

Question: “What is the favorite sport of 5th graders at our school?”

You want to learn about all 5th graders, so the population is all 5th grade students at the school.

Suppose there are 90 students in 5th grade. Asking all 90 students would be best, but if that is not possible, you can use a sample.

A fair sample might be 18 students chosen randomly.

One way to do this:

  1. Write each student’s name on a slip of paper.
  2. Put all 90 slips in a container.
  3. Mix them well.
  4. Draw 18 names without looking.

This is better than asking only your friends, because your friends may like the same sports you do.

Worked Example 3: Collecting and recording data

Question: “How many minutes do students in our class read at home each night?”

This is a statistical question because the answers will probably vary.

Suppose you survey 8 students and get these answers in minutes:

$$15,\ 20,\ 30,\ 20,\ 10,\ 25,\ 30,\ 15$$

This list is your data. Each number is one data value.

You can organize the data in a table:

  • 10 minutes: 1 student
  • 15 minutes: 2 students
  • 20 minutes: 2 students
  • 25 minutes: 1 student
  • 30 minutes: 2 students

Notice the data has variation. The reading time is not the same for everyone.

Worked Example 4: Fixing a weak question

Weak question: “Do you like school?”

This question may be too general. Different students might think about different parts of school, such as math, lunch, or recess.

A better question is: “What is your favorite subject in school?”

This new question is clearer. It is also statistical because different students may choose different subjects.

Another better question could be: “How many minutes of homework do 5th grade students have on a typical school night?”

This question is also statistical because the answers can vary.

Common mistakes to watch for

  • Asking about only one person: That usually makes it non-statistical.
  • Writing a question with only one answer: If there is no variation, it is not statistical.
  • Choosing an unfair sample: Asking only one small group can give biased results.
  • Recording data carelessly: Missing or mixed-up answers can lead to mistakes.

How to check your own work

When you write a question, ask yourself:

  • Am I asking about a group?
  • Will the answers probably be different?
  • Can I collect data to answer it?

When you collect data, ask yourself:

  • Did I choose people fairly?
  • Did I record each answer correctly?
  • Does my data match the question I asked?

Summary

A statistical question is a question about a group that expects different answers. Those differences are called variation.

To answer a statistical question, we collect data by surveying, counting, measuring, or observing. If we cannot ask everyone, we can use a sample.

A random sample is a fair way to choose people because everyone has the same chance of being selected. Good questions and fair data collection help us learn true information about a group.

Put what you read to the test

You've worked through Statistical Questions and Data Collection. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Frequency Tables and Dot Plots

Frequency Tables and Dot Plots

When we collect data, we often end up with a long list of numbers. A frequency table helps us organize the numbers. A dot plot helps us show the numbers on a number line so we can look for patterns.

In this lesson, you will learn how to:

  • organize data in a frequency table,
  • make a dot plot from a set of data,
  • read a dot plot, and
  • find clusters and gaps in the data.

1. What is data?

Data is information we collect. For example, data could be test scores, number of books read, or the number of minutes students read each night.

Here is a set of data showing how many books 12 students read in one month:

2, 4, 3, 2, 5, 4, 2, 3, 6, 4, 3, 2

This list is called raw data because it has not been organized yet.

2. What is a frequency table?

A frequency table shows each value in the data and tells how many times it appears. The word frequency means how often.

To make a frequency table:

  1. List the numbers in order from least to greatest.
  2. Count how many times each number appears.
  3. Write that count next to the number.

Worked Example 1: Make a frequency table

Data: 2, 4, 3, 2, 5, 4, 2, 3, 6, 4, 3, 2

Step 1: List the values in order: 2, 3, 4, 5, 6

Step 2: Count each value:

  • 2 appears 4 times
  • 3 appears 3 times
  • 4 appears 3 times
  • 5 appears 1 time
  • 6 appears 1 time

The frequency table is:

2: 4
3: 3
4: 3
5: 1
6: 1

You can also check your work by adding the frequencies:

$$4+3+3+1+1=12$$

There are 12 data values, so the table matches the data.

3. What is a dot plot?

A dot plot shows data on a number line. Each dot stands for one piece of data. If a number appears more than once, the dots are stacked above that number.

To make a dot plot:

  1. Draw a number line that includes all the data values.
  2. Put one dot above a number for each time it appears.
  3. Stack dots neatly if there is more than one.

Worked Example 2: Make a dot plot from the frequency table

Use the same data:

2: 4, 3: 3, 4: 3, 5: 1, 6: 1

This dot plot would have:

  • 4 dots above 2
  • 3 dots above 3
  • 3 dots above 4
  • 1 dot above 5
  • 1 dot above 6

It can look like this:

6 |
5 |
4 | ●
3 | ● ●
2 | ● ● ●
1 | ● ● ● ● ●
   2 3 4 5 6

The important idea is that each dot means one data value.

4. How do frequency tables and dot plots help?

Both tools help us see data clearly.

  • A frequency table makes counting easy.
  • A dot plot makes patterns easy to see.

When reading a dot plot, look for:

  • Most common values: numbers with the most dots
  • Clusters: a group of data points close together
  • Gaps: places on the number line with no data
  • Least and greatest values: the smallest and largest numbers in the data

5. What are clusters and gaps?

A cluster is a group of data values that are close together. This shows where much of the data is found.

A gap is a value or set of values with no data. On a dot plot, there are no dots above those numbers.

Worked Example 3: Find clusters and gaps

Here is a new set of data showing the number of minutes students practiced piano:

10, 12, 12, 13, 14, 14, 15, 20, 20, 21

First, make the frequency table:

10: 1
11: 0
12: 2
13: 1
14: 2
15: 1
16: 0
17: 0
18: 0
19: 0
20: 2
21: 1

Now think about the dot plot. There would be many dots from 10 to 15, no dots from 16 to 19, and then dots again at 20 and 21.

What do we notice?

  • There is a cluster from 10 to 15.
  • There is a gap from 16 to 19.
  • There is another small group at 20 and 21.

This tells us most students practiced between 10 and 15 minutes, and nobody practiced 16, 17, 18, or 19 minutes.

6. Reading information from a dot plot

You can answer questions by counting dots.

For example, if a dot plot shows 3 dots above 8, then 3 data values are equal to 8.

If the leftmost dot is above 2 and the rightmost dot is above 9, then:

  • the least value is 2,
  • the greatest value is 9.

Worked Example 4: Answer questions from a frequency table and dot plot

Data: 1, 2, 2, 3, 3, 3, 5, 5, 6

Frequency table:

1: 1
2: 2
3: 3
4: 0
5: 2
6: 1

Questions:

  • How many times does 3 appear? It appears 3 times.
  • What is the least value? The least value is 1.
  • What is the greatest value? The greatest value is 6.
  • Is there a gap? Yes. There is a gap at 4 because there are 0 data values of 4.
  • Where is the cluster? The data is clustered around 2, 3, and 5, especially near 2 and 3.

7. Tips for making good frequency tables and dot plots

  • Write the data values in order.
  • Count carefully so no value is missed.
  • Make sure each dot stands for only one data value.
  • Stack dots straight up and down.
  • Include values with 0 frequency if they help show a gap.
  • Check that the total number of dots equals the total number of data values.

8. Common mistakes to avoid

  • Skipping a number on the number line
  • Putting two data values in one dot
  • Miscounting how many times a number appears
  • Forgetting gaps because numbers with 0 frequency were not noticed

9. Let’s review the steps

When you are given raw data:

  1. Put the numbers in order.
  2. Make a frequency table by counting each number.
  3. Draw a number line.
  4. Use dots to show the frequency of each value.
  5. Look for clusters, gaps, and values that appear most often.

Summary

A frequency table tells how many times each number appears in a data set. A dot plot shows the same data on a number line using dots.

These tools help us organize data and see patterns. We can use them to find the least and greatest values, the most common values, clusters, and gaps.

If you can count each value carefully and place the correct number of dots above each number, you can read and make frequency tables and dot plots successfully.

Put what you read to the test

You've worked through Frequency Tables and Dot Plots. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Bar Graphs and Histograms

Bar Graphs and Histograms are both ways to show data so it is easy to understand. They help us look for patterns, compare amounts, and answer questions about information.

Even though bar graphs and histograms may look similar, they are used for different kinds of data. Learning when to use each one is very important.

In this lesson, you will learn:

  • what a bar graph is
  • what a histogram is
  • how they are alike
  • how they are different
  • how to read and make each type of graph

1. What is a Bar Graph?

A bar graph shows data in separate groups or categories. The bars tell how many are in each category.

For example, if students vote for their favorite fruit, the categories might be apples, bananas, grapes, and oranges. These are names of groups, not number ranges.

In a bar graph:

  • each bar stands for one category
  • the height or length of the bar shows the amount
  • the bars are separated by spaces

The spaces matter because the categories are separate. Apples and bananas are different groups, so their bars do not touch.

Parts of a Bar Graph

  • Title: tells what the graph is about
  • Categories: the groups being compared
  • Scale: the numbers that show how much each bar is worth
  • Bars: the rectangles that show the data

Example 1: Reading a Bar Graph

A class voted for their favorite recess game.

  • Soccer: 8
  • Tag: 5
  • Jump Rope: 3
  • Basketball: 6

If we make a bar graph, each game is a category. The height of each bar shows how many students chose it.

Questions:

  1. Which game is the most popular?
  2. How many more students chose Soccer than Jump Rope?

Solution:

  1. Soccer is the most popular because its bar is tallest at 8.
  2. Find the difference: $$8 - 3 = 5$$ So, 5 more students chose Soccer than Jump Rope.

2. What is a Histogram?

A histogram shows numerical data that has been grouped into intervals, often called bins. A bin is a range of numbers.

For example, instead of favorite fruits, we might record the heights of plants in centimeters. Since heights are numbers that can vary across a range, we group them into bins like 0–9, 10–19, 20–29, and so on.

In a histogram:

  • the data is numerical
  • the numbers are grouped into ranges
  • the bars usually touch

The bars touch because the number ranges are connected. For example, 10–19 comes right before 20–29, so the data moves along a number line without gaps.

Parts of a Histogram

  • Title: tells what the graph is about
  • Bins: number ranges such as 0–4, 5–9, or 10–14
  • Scale: shows frequency, or how many data values are in each bin
  • Bars: show how many values fall in each range

Important Word: Frequency

Frequency means how many times something happens or how many data values are in a group.

If 4 test scores fall in the range 80–89, then the frequency of the 80–89 bin is 4.

Example 2: Reading a Histogram

A teacher records the number of minutes students read in one week. The data is grouped like this:

  • 0–9 minutes: 2 students
  • 10–19 minutes: 5 students
  • 20–29 minutes: 7 students
  • 30–39 minutes: 4 students

Questions:

  1. Which bin has the greatest frequency?
  2. How many students read less than 20 minutes?

Solution:

  1. The 20–29 minutes bin has the greatest frequency because 7 students are in that range.
  2. Less than 20 minutes means the first two bins: 0–9 and 10–19. Add them: $$2 + 5 = 7$$ So, 7 students read less than 20 minutes.

3. Bar Graph vs. Histogram

These graphs may look alike, but they are not the same.

  • A bar graph is used for categories.
  • A histogram is used for number ranges.

Main Differences

  • Bar graph: bars have spaces between them.
  • Histogram: bars usually touch.
  • Bar graph: compares categories like colors, pets, or foods.
  • Histogram: shows how many data values fall into ranges like 0–5, 6–10, or 11–15.

When to Use Each One

  • Use a bar graph when the data answers, “Which group?”
  • Use a histogram when the data answers, “How many values fall in this number range?”

4. How to Make a Bar Graph

  1. Write a title.
  2. Label the categories on one axis.
  3. Label the numbers on the other axis using a scale.
  4. Draw one bar for each category.
  5. Make sure the bars are the correct height or length.
  6. Leave spaces between the bars.

Example 3: Making a Bar Graph

Here is data about pets owned by students:

  • Dogs: 6
  • Cats: 4
  • Fish: 3
  • Birds: 2

To make a bar graph:

  1. Use the title Pets Owned by Students.
  2. Put the pet categories on the bottom.
  3. Use a number scale on the side, such as 0 to 6.
  4. Draw bars up to the correct values: Dogs to 6, Cats to 4, Fish to 3, Birds to 2.

What can we learn?

  • Dogs are the most common pet.
  • Birds are the least common pet.
  • There are $$6 - 4 = 2$$ more dogs than cats.

5. How to Make a Histogram

  1. Write a title.
  2. Look at the numerical data.
  3. Choose bins, or number ranges.
  4. Count how many data values fall into each bin.
  5. Label the bins on the horizontal axis.
  6. Label the frequency on the vertical axis.
  7. Draw bars that touch.

Example 4: Making a Histogram

The numbers of pages read by 12 students are:

3, 7, 8, 12, 14, 15, 16, 21, 22, 24, 27, 29

Let’s group them into bins:

  • 0–9
  • 10–19
  • 20–29

Now count the values in each bin.

Step 1: Count 0–9

The numbers are 3, 7, 8. That is 3 values.

Step 2: Count 10–19

The numbers are 12, 14, 15, 16. That is 4 values.

Step 3: Count 20–29

The numbers are 21, 22, 24, 27, 29. That is 5 values.

So the frequencies are:

  • 0–9: 3
  • 10–19: 4
  • 20–29: 5

In the histogram, the bars for these bins would touch each other.

What can we learn?

  • The 20–29 bin has the most students.
  • The fewest students are in the 0–9 bin.
  • More students read 20 or more pages than less than 20 pages.

6. Tips for Reading Graphs Carefully

  • Always read the title.
  • Check what the labels mean.
  • Look at the scale on the side.
  • Notice whether the graph uses categories or number ranges.
  • Check whether the bars have spaces or touch.

7. Common Mistakes to Avoid

  • Using a bar graph for number ranges instead of a histogram
  • Forgetting spaces between bars in a bar graph
  • Leaving spaces between bars in a histogram
  • Reading the scale incorrectly
  • Counting values into the wrong bin

8. Quick Check

Decide whether each situation should use a bar graph or a histogram.

  1. Favorite ice cream flavors in a class
  2. Heights of plants grouped into ranges
  3. Kinds of pets students own
  4. Test scores grouped as 60–69, 70–79, 80–89, 90–99

Answers:

  1. Bar graph
  2. Histogram
  3. Bar graph
  4. Histogram

Summary

A bar graph is used to compare categories, and its bars have spaces between them. A histogram is used to show numerical data grouped into bins, and its bars usually touch.

Both graphs help us organize data and see patterns. If you remember to ask, “Are these categories or number ranges?” you can choose the correct graph.

Put what you read to the test

You've worked through Bar Graphs and Histograms. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Line Graphs and Trends

Line Graphs and Trends

A line graph is a graph that shows how something changes over time. It helps us see patterns in data quickly. We use points to show the data, and then we connect the points with line segments.

Line graphs are especially useful when data is collected in order, such as by day, week, month, or year. For example, a line graph can show temperature during a week, the number of books read each month, or the distance walked each day.

When we study a line graph, we look for trends. A trend is the general direction the data is moving.

  • An upward trend means the data is generally increasing.
  • A downward trend means the data is generally decreasing.
  • If the graph goes up and down without a clear direction, the data has mixed changes.
  • If the graph stays about the same, it shows little or no change.

Parts of a Line Graph

A line graph usually has these important parts:

  • Title: tells what the graph is about.
  • Horizontal axis (x-axis): usually shows time, such as days or months.
  • Vertical axis (y-axis): shows what is being measured, such as inches of rain or number of visitors.
  • Scale: the numbers on each axis. The scale must be counted evenly.
  • Points and line segments: show the data and connect it in order.

To read a line graph correctly, first read the title. Next, check what each axis means. Then look at the scale, because each mark may stand for 1, 2, 5, 10, or another amount.

How to Make a Line Graph

  1. Write a title that tells what the data shows.
  2. Draw a horizontal axis and a vertical axis.
  3. Label the horizontal axis with time.
  4. Label the vertical axis with the thing being measured.
  5. Choose a scale that fits the data.
  6. Plot each data point in the correct place.
  7. Connect the points in order with line segments.

Understanding Trends

A trend tells us what is happening overall. We do not only look at one point. We look at many points together.

If the points move from lower to higher values over time, the graph shows an upward trend. If the points move from higher to lower values over time, the graph shows a downward trend.

Sometimes a graph goes up for a while and then down. That means the trend changes. It is important to notice where the changes happen.

Rate of Change

Rate of change tells how much a value changes from one time to the next. To find it, subtract the earlier amount from the later amount.

We can write this as:

$$\text{change} = \text{later value} - \text{earlier value}$$

If the answer is positive, the data increased. If the answer is negative, the data decreased.

For example, if a plant was 8 inches tall on Monday and 11 inches tall on Tuesday, the change is:

$$11 - 8 = 3$$

The plant grew 3 inches.

If the temperature was 72 degrees on Wednesday and 68 degrees on Thursday, the change is:

$$68 - 72 = -4$$

The temperature went down by 4 degrees.

Worked Example 1: Reading a Simple Line Graph

A student records the number of pages read each day.

  • Monday: 10 pages
  • Tuesday: 15 pages
  • Wednesday: 20 pages
  • Thursday: 18 pages
  • Friday: 25 pages

What trend do you see?

From Monday to Wednesday, the number of pages increases from 10 to 20. From Wednesday to Thursday, it drops from 20 to 18. Then it rises again to 25 on Friday.

This graph shows a mostly upward trend, even though there is a small drop on Thursday.

Worked Example 2: Finding the Change Between Two Points

A line graph shows the number of cups of water a class drank during sports week.

  • Day 1: 12 cups
  • Day 2: 14 cups
  • Day 3: 14 cups
  • Day 4: 17 cups
  • Day 5: 13 cups

How much did the number of cups change from Day 1 to Day 4?

Use subtraction:

$$17 - 12 = 5$$

The class drank 5 more cups on Day 4 than on Day 1.

What happened from Day 4 to Day 5?

$$13 - 17 = -4$$

The number of cups decreased by 4.

Worked Example 3: Making a Line Graph

Suppose a gardener measures a sunflower every week.

  • Week 1: 4 cm
  • Week 2: 7 cm
  • Week 3: 9 cm
  • Week 4: 12 cm

Step 1: Title the graph: Sunflower Height Each Week.

Step 2: Put weeks on the horizontal axis.

Step 3: Put height in centimeters on the vertical axis.

Step 4: Plot the points: \((1,4)\), \((2,7)\), \((3,9)\), and \((4,12)\).

Step 5: Connect the points in order.

What trend does the graph show?

The graph goes upward the whole time. This shows an upward trend. The sunflower is growing each week.

What is the change from Week 3 to Week 4?

$$12 - 9 = 3$$

The sunflower grew 3 cm from Week 3 to Week 4.

Worked Example 4: Looking for Bigger and Smaller Changes

A line graph shows the number of minutes a child practiced piano.

  • Monday: 20 minutes
  • Tuesday: 25 minutes
  • Wednesday: 35 minutes
  • Thursday: 30 minutes
  • Friday: 40 minutes

What is the greatest increase?

Find the change each day:

  • Tuesday minus Monday: $$25 - 20 = 5$$
  • Wednesday minus Tuesday: $$35 - 25 = 10$$
  • Thursday minus Wednesday: $$30 - 35 = -5$$
  • Friday minus Thursday: $$40 - 30 = 10$$

The greatest increase is 10 minutes. It happens from Tuesday to Wednesday and again from Thursday to Friday.

What is the overall trend?

The data starts at 20 minutes and ends at 40 minutes. Even though there is a drop on Thursday, the graph shows a general upward trend.

Tips for Reading Line Graphs Carefully

  • Always read the title first.
  • Check the labels on both axes.
  • Look at the scale before reading values.
  • Follow the points from left to right because time moves in order.
  • Notice where the graph rises, falls, or stays flat.
  • Compare points to find how much the data changed.

Common Mistakes to Avoid

  • Reading the wrong scale on the vertical axis.
  • Forgetting to plot points in time order.
  • Connecting points that are out of order.
  • Looking at only one point instead of the whole graph when finding a trend.
  • Mixing up an increase and a decrease when subtracting.

Quick Check

Imagine a line graph shows these temperatures during a week:

  • Monday: 60
  • Tuesday: 62
  • Wednesday: 65
  • Thursday: 63
  • Friday: 68

Ask yourself:

  • Is the trend mostly upward or downward?
  • How much did the temperature change from Monday to Friday?

The trend is mostly upward.

The change from Monday to Friday is:

$$68 - 60 = 8$$

The temperature increased by 8 degrees.

Summary

A line graph shows how data changes over time. It has a title, two labeled axes, a scale, and connected points.

We use line graphs to find trends. An upward trend means the data is increasing, and a downward trend means it is decreasing.

We can also find how much the data changes between two points by subtracting the earlier value from the later value. When you read a line graph carefully, you can understand patterns and changes in data.

Put what you read to the test

You've worked through Line Graphs and Trends. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Measures of Center: Mean, Median, and Mode

Measures of Center: Mean, Median, and Mode

When we collect data, we often want to describe what value is the most typical or central. Measures of center help us do that.

The three most common measures of center are mean, median, and mode. Each one tells us something a little different about a set of numbers.

In this lesson, you will learn what mean, median, and mode are, how to find each one, and when each one is useful.

1. What is a data set?

A data set is just a group of numbers collected together. For example, these could be the number of books students read in a month:

\(2, 4, 4, 5, 7\)

We can use this data set to find the mean, median, and mode.

2. Mean

The mean is the average. Another way to think about the mean is the fair share.

To find the mean:

  • Add all the numbers in the data set.
  • Divide by how many numbers there are.

The formula is:

$$\text{Mean} = \frac{\text{sum of all values}}{\text{number of values}}$$

Example 1: Finding the mean

Find the mean of \(3, 5, 7, 9\).

Step 1: Add the numbers.

$$3 + 5 + 7 + 9 = 24$$

Step 2: Count how many numbers there are.

There are \(4\) numbers.

Step 3: Divide.

$$\frac{24}{4} = 6$$

Answer: The mean is 6.

This means if the total \(24\) were shared equally among 4 numbers, each would be \(6\).

3. Median

The median is the middle number in a data set when the numbers are put in order from least to greatest.

To find the median:

  • First, put the numbers in order.
  • Find the middle number.

If there is an odd number of values, there is one middle number.

If there is an even number of values, there are two middle numbers. Add those two numbers and divide by 2.

Example 2: Finding the median with an odd number of values

Find the median of \(8, 3, 5, 1, 6\).

Step 1: Put the numbers in order.

$$1, 3, 5, 6, 8$$

Step 2: Find the middle number.

The middle number is 5.

Answer: The median is 5.

Example 3: Finding the median with an even number of values

Find the median of \(2, 4, 6, 8\).

Step 1: The numbers are already in order.

$$2, 4, 6, 8$$

Step 2: Find the two middle numbers.

The two middle numbers are \(4\) and \(6\).

Step 3: Find their average.

$$\frac{4+6}{2} = \frac{10}{2} = 5$$

Answer: The median is 5.

4. Mode

The mode is the number that appears the most often.

To find the mode:

  • Look at the data set.
  • Count how many times each number appears.
  • The number that appears the most is the mode.

Example 4: Finding the mode

Find the mode of \(4, 2, 4, 6, 4, 7, 2\).

Count how often each number appears:

  • \(2\) appears 2 times
  • \(4\) appears 3 times
  • \(6\) appears 1 time
  • \(7\) appears 1 time

Answer: The mode is 4 because it appears the most often.

5. A data set can have different kinds of mode

Sometimes a data set has:

  • One mode: one number appears most often.
  • More than one mode: two or more numbers tie for appearing most often.
  • No mode: no number repeats.

Example of more than one mode:

\(1, 2, 2, 3, 3, 4\)

Both \(2\) and \(3\) appear 2 times, so both are modes.

Example of no mode:

\(5, 6, 7, 8\)

Each number appears only once, so there is no mode.

6. Finding all three measures together

Now let’s find the mean, median, and mode for one data set.

Data set: \(2, 3, 3, 5, 7\)

Mean:

Add the numbers:

$$2 + 3 + 3 + 5 + 7 = 20$$

There are \(5\) numbers.

$$\frac{20}{5} = 4$$

The mean is 4.

Median:

The numbers are already in order:

$$2, 3, 3, 5, 7$$

The middle number is 3.

The median is 3.

Mode:

The number \(3\) appears 2 times. The other numbers appear 1 time.

The mode is 3.

So for this data set:

  • Mean = \(4\)
  • Median = \(3\)
  • Mode = \(3\)

7. Helpful tips

  • Mean: add, then divide.
  • Median: put numbers in order first.
  • Mode: look for the number used most.

A very common mistake is forgetting to put the numbers in order before finding the median.

Another common mistake is mixing up mean and median. Remember:

  • Mean uses all the numbers in the calculation.
  • Median is just the middle number.
  • Mode is the number that shows up the most.

8. When are these useful?

Measures of center help us describe data quickly.

  • A teacher might find the mean test score of a class.
  • A coach might look at the median number of points scored.
  • A store might find the mode shoe size sold most often.

Each measure gives a different way to understand the center of the data.

Summary

The mean is the average, found by adding all the numbers and dividing by how many numbers there are.

The median is the middle number when the data is in order.

The mode is the number that appears most often.

When you see a data set, ask yourself:

  • Do I need the fair share? Find the mean.
  • Do I need the middle value? Find the median.
  • Do I need the most common value? Find the mode.

With practice, you will be able to find all three measures of center with confidence.

Put what you read to the test

You've worked through Measures of Center: Mean, Median, and Mode. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Measures of Variability: Range

Measures of Variability: Range

When we collect data, we often want to know more than just the numbers in the list. We want to understand how the data is spread out.

One simple way to measure how spread out a set of data is called the range.

The range tells us the difference between the greatest value and the least value in a data set.

We can write it like this:

$$\text{Range} = \text{greatest value} - \text{least value}$$

This means you find the biggest number, find the smallest number, and subtract.

Why is range useful?

Range helps us understand whether the data values are close together or far apart.

  • If the range is small, the data values are closer together.
  • If the range is large, the data values are more spread out.

For example, look at these two sets of scores:

Set A: 8, 9, 9, 10, 10

Set B: 2, 5, 9, 12, 15

Both sets have numbers, but Set B is spread out much more. The range helps us see that clearly.

Steps for finding the range

  1. Look at the data set.
  2. Find the least value.
  3. Find the greatest value.
  4. Subtract: greatest value minus least value.

You do not add all the numbers. You do not use every number in the subtraction. You only use the smallest and largest values.

Worked Example 1

Find the range of: 3, 5, 6, 8, 10

Step 1: Find the least value. It is 3.

Step 2: Find the greatest value. It is 10.

Step 3: Subtract.

$$10 - 3 = 7$$

Answer: The range is 7.

Worked Example 2

Find the range of: 12, 7, 15, 9, 7, 11

The numbers are not in order, and that is okay. You can still find the least and greatest values.

Least value: 7

Greatest value: 15

$$15 - 7 = 8$$

Answer: The range is 8.

It can help to put the data in order first:

7, 7, 9, 11, 12, 15

Now it is easier to see the smallest and largest numbers.

Worked Example 3

A class measured the lengths of worms in centimeters: 4, 6, 5, 9, 7, 4, 8

Find the range.

Least value: 4

Greatest value: 9

$$9 - 4 = 5$$

Answer: The range is 5 centimeters.

Since the data is about length, the range should also use the same unit: centimeters.

Worked Example 4

The temperatures over five days were 68, 72, 70, 75, 71 degrees.

Find the range.

Least value: 68

Greatest value: 75

$$75 - 68 = 7$$

Answer: The range is 7 degrees.

What range tells us about data

Let’s compare two data sets:

Set 1: 20, 21, 20, 22, 21

Set 2: 10, 15, 20, 25, 30

For Set 1:

$$22 - 20 = 2$$

For Set 2:

$$30 - 10 = 20$$

Set 1 has a range of 2, so the numbers are close together.

Set 2 has a range of 20, so the numbers are much more spread out.

This is why range is called a measure of variability. It measures how much the data varies, or changes, from the smallest value to the largest value.

Common mistakes to avoid

  • Do not subtract in the wrong order. Always do greatest minus least.
  • Do not use the middle numbers. Only the smallest and largest values are needed.
  • Do not confuse range with how many numbers there are. Range is about spread, not counting.
  • Do not forget the unit when the data has one, like inches, centimeters, or degrees.

Quick Practice

Try these on your own:

  1. 5, 9, 12, 6, 8
  2. 14, 14, 14, 14
  3. 3, 11, 7, 19, 5

Answers:

1. Least is 5, greatest is 12, so $$12 - 5 = 7$$

2. Least is 14, greatest is 14, so $$14 - 14 = 0$$

3. Least is 3, greatest is 19, so $$19 - 3 = 16$$

If all the numbers are the same, the range is 0 because there is no difference between the smallest and largest values.

Summary

The range is a way to describe how spread out a data set is.

To find the range:

$$\text{Range} = \text{greatest value} - \text{least value}$$

Remember: find the biggest number, find the smallest number, and subtract. A small range means the data is close together. A large range means the data is more spread out.

Put what you read to the test

You've worked through Measures of Variability: Range. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Outliers and Data Skew

Outliers and Data Skew are important ideas in data analysis. They help us understand when a data set has a value that is very different from the rest, and how that unusual value can change what the “average” looks like.

In this lesson, you will learn how to:

  • find an outlier,
  • tell how an outlier changes the mean,
  • see why the median is often less affected,
  • and notice when data is skewed, or pulled to one side.

Let’s start with a quick review.

The mean is the sum of all the numbers divided by how many numbers there are.

$$\text{mean} = \frac{\text{sum of data values}}{\text{number of data values}}$$

The median is the middle number when the data is put in order from least to greatest.

If there are 2 middle numbers, the median is the number halfway between them.

An outlier is a value that is much bigger or much smaller than the other values in the data set.

For example, in the set \(4, 5, 5, 6, 5, 40\), the number \(40\) is far away from the other numbers. That makes it an outlier.

Why do outliers matter?

Outliers can change the mean a lot because the mean uses every number in the set.

The median often changes very little, because it depends on the middle value, not how far away the biggest or smallest number is.

This is why outliers can make a data set look different than it really is if we only look at the mean.

What does skew mean?

When a data set is skewed, it is pulled more to one side than the other.

  • If there is a very large outlier, the data may be pulled to the right.
  • If there is a very small outlier, the data may be pulled to the left.

You can think of skew as a “stretch” in one direction because of unusual values.

How to spot an outlier

  1. Put the numbers in order from least to greatest.
  2. Look for a number that is far from the rest.
  3. Ask: Does this value seem much smaller or much larger than all the others?
  4. If yes, it may be an outlier.

You do not need a special formula in 5th grade. Use careful thinking and compare the values.

Worked Example 1: Finding an outlier

A class measured the number of books read in one month:

\(2, 3, 3, 4, 4, 5, 18\)

First, look at the numbers. Most of them are between \(2\) and \(5\).

The number \(18\) is much greater than the rest.

So, 18 is an outlier.

This data set may be skewed to the right because of the large value \(18\).

Worked Example 2: How an outlier affects the mean and median

Suppose these are the quiz scores of 5 students:

\(80, 82, 83, 85, 84\)

First, put them in order:

\(80, 82, 83, 84, 85\)

Find the mean:

$$80 + 82 + 83 + 84 + 85 = 414$$

$$\text{mean} = \frac{414}{5} = 82.8$$

Find the median. The middle number is \(83\).

So before any outlier:

  • Mean = \(82.8\)
  • Median = \(83\)

Now imagine one score was written wrong as \(30\) instead of \(80\).

The new data set is:

\(30, 82, 83, 84, 85\)

Find the new mean:

$$30 + 82 + 83 + 84 + 85 = 364$$

$$\text{mean} = \frac{364}{5} = 72.8$$

Find the new median. The middle number is still \(83\).

So after the outlier:

  • Mean = \(72.8\)
  • Median = \(83\)

What happened?

The mean dropped from \(82.8\) to \(72.8\). That is a big change.

The median stayed at \(83\). It did not change at all.

This shows that outliers affect the mean more than the median.

Worked Example 3: A small outlier

Here are the times, in minutes, that 6 students spent finishing a puzzle:

\(9, 10, 10, 11, 12, 40\)

Most times are between \(9\) and \(12\), but \(40\) is much larger.

So \(40\) is an outlier.

Find the mean:

$$9 + 10 + 10 + 11 + 12 + 40 = 92$$

$$\text{mean} = \frac{92}{6} \approx 15.3$$

To find the median, the data is already in order:

\(9, 10, 10, 11, 12, 40\)

There are 6 numbers, so find the two middle numbers: \(10\) and \(11\).

The median is halfway between them:

$$\frac{10 + 11}{2} = 10.5$$

So:

  • Mean \(\approx 15.3\)
  • Median = \(10.5\)

The mean is much higher than the median because the outlier \(40\) pulled it upward.

This is another sign that the data is skewed to the right.

Worked Example 4: Comparing two data sets

Data Set A:

\(6, 7, 7, 8, 8, 9\)

Data Set B:

\(6, 7, 7, 8, 8, 20\)

Let’s compare them.

Data Set A

Sum:

$$6 + 7 + 7 + 8 + 8 + 9 = 45$$

Mean:

$$\frac{45}{6} = 7.5$$

Median: the middle two numbers are \(7\) and \(8\).

$$\frac{7 + 8}{2} = 7.5$$

So for Data Set A:

  • Mean = \(7.5\)
  • Median = \(7.5\)

Data Set B

Sum:

$$6 + 7 + 7 + 8 + 8 + 20 = 56$$

Mean:

$$\frac{56}{6} \approx 9.3$$

Median: the middle two numbers are still \(7\) and \(8\).

$$\frac{7 + 8}{2} = 7.5$$

So for Data Set B:

  • Mean \(\approx 9.3\)
  • Median = \(7.5\)

What do we notice?

Data Set B has an outlier, \(20\), and it makes the mean larger.

The median stays the same.

Again, the outlier affects the mean more than the median.

Mean or median: which is better?

When there are no outliers, the mean and median may both do a good job describing the center of the data.

When there is an outlier, the median is often a better way to describe the center, because it is not pulled as much by one unusual number.

For example, if most students read about 3 or 4 books, but one student read 18 books, the median gives a better picture of what is normal for the class.

Clues that data may be skewed

  • One number is much bigger or smaller than the others.
  • The mean and median are far apart.
  • The data seems stretched more on one side.

If the large values stretch out farther, the data is skewed right.

If the small values stretch out farther, the data is skewed left.

Let’s practice thinking

Suppose the data set is \(4, 5, 5, 6, 6, 25\).

  • Is there an outlier? Yes, \(25\).
  • Will the mean or median be affected more? The mean.
  • Will the data be skewed left or right? Right, because of the large value.

Now suppose the data set is \(1, 8, 8, 9, 9, 10\).

  • Is there an outlier? Yes, \(1\).
  • Will the mean or median be affected more? The mean.
  • Will the data be skewed left or right? Left, because of the small value.

Important ideas to remember

  • An outlier is a value far away from the rest of the data.
  • Outliers can change the mean a lot.
  • Outliers usually change the median less.
  • A large outlier can make data skew right.
  • A small outlier can make data skew left.

Summary

When you study data, do not just find the mean and median. Also look for values that seem very different from the rest.

If there is an outlier, the mean may not show the center very well because it gets pulled toward that unusual value.

The median is often more reliable when outliers are present. Looking for outliers and skew helps you understand what the data is really saying.

Put what you read to the test

You've worked through Outliers and Data Skew. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Analyzing Line Plot Clusters

Analyzing Line Plot Clusters

A line plot is a graph that shows data using marks above a number line. Each mark stands for one thing counted.

When we look at a line plot, we do more than count. We can also notice where the data is grouped together. A group of data values close to each other is called a cluster.

In this lesson, you will learn how to find clusters, notice values that are far away from the others, and tell which measurement happens most often.

What to look for on a line plot

  • Cluster: a group of marks close together
  • Most common measurement: the value with the most marks above it
  • Outlier: a value far away from most of the data

Think of a line plot like a row of houses. If many marks live close together, that area is a cluster. If one mark lives all alone far away, that may be an outlier.

How to analyze a line plot

  1. Look across the number line.
  2. Find places where many marks are close together.
  3. See if one value has more marks than the others.
  4. Check if any mark is far from the rest.

Example 1: Finding a cluster

Here is a line plot of pencil lengths in inches.

2: X
3: XX
4: XXX
5: XX
6: X

Let’s look at the marks.

  • There is 1 mark at 2.
  • There are 2 marks at 3.
  • There are 3 marks at 4.
  • There are 2 marks at 5.
  • There is 1 mark at 6.

The data is grouped most closely around 3, 4, and 5. That means the cluster is around 3 to 5.

The value 4 has the most marks. So, 4 is the most common measurement.

Example 2: Finding the most common measurement

Here is a line plot of how many books students read.

1: X
2: XXX
3: XXXX
4: XX
5: X

Let’s answer two questions.

What is the cluster?

Most of the marks are around 2, 3, and 4. So the cluster is from 2 to 4.

What measurement is most common?

The number 3 has 4 marks. That is more than any other value. So, 3 books is the most common measurement.

Example 3: Finding an outlier

Here is a line plot of toy car lengths.

2: XX
3: XXX
4: XX
5:
6:
7: X

Most of the marks are at 2, 3, and 4. That is the cluster.

But there is 1 mark at 7. It is far away from the others. That makes 7 an outlier.

An outlier is important because it does not fit with the main group.

Example 4: Thinking about the shape of the data

Here is a line plot of plant heights in inches.

1:
2: X
3: XXX
4: XXXX
5: XXX
6: X

This line plot has many marks in the middle. The data is grouped around 3, 4, and 5, so that is the cluster.

The most common measurement is 4 because it has the most marks.

There is no value far away from the rest, so there is no outlier here.

Helpful questions to ask

  • Where do I see many marks close together?
  • Which number has the most marks?
  • Is there a mark far away from the group?
  • Is the data mostly in the middle, or more on one side?

Tips for success

  • Count carefully.
  • Look for the biggest group first.
  • Do not choose a cluster from just one mark.
  • If one value is alone and far away, it may be an outlier.

Let’s practice thinking

If a line plot has many marks at 5, 6, and 7, then the cluster is around 5 to 7.

If 6 has the most marks, then 6 is the most common measurement.

If there is one mark at 1 and all the other marks are near 5, 6, and 7, then 1 is probably an outlier.

Summary

A line plot helps us see where data is grouped. A cluster is where many data points are close together. The most common measurement is the value with the most marks. An outlier is a value far away from the rest. When you study a line plot, look for the group, the tallest stack of marks, and any lonely mark far away.

Put what you read to the test

You've worked through Analyzing Line Plot Clusters. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.