Chapter 9

Statistics and Bivariate Data

Statistical Questioning and Bias

Lesson: Statistical Questioning and Bias

In statistics, we often use data to learn about a group of people, objects, or events. But before we collect data, we need to ask a good question and choose a fair way to gather information.

This lesson will help you understand statistical questions, variability, sampling, and bias. These ideas are important because bad questions or unfair samples can lead to misleading results.

1. What is a statistical question?

A statistical question is a question that expects different answers from different people or situations. In other words, it allows for variability.

Variability means that the data values are not all the same. They can vary, or change, from one person or item to another.

For example, the question "How many minutes do 7th graders spend on homework each night?" is a statistical question. Different students will give different answers, so the data will vary.

The question "How many days are in a week?" is not a statistical question. There is only one correct answer: 7. There is no variability.

Ask yourself: Does this question expect many possible answers, or just one fixed answer?

  • Statistical question: "How tall are the students in our class?"
  • Not statistical: "How tall is the door?"
  • Statistical question: "How many books do students read in a month?"
  • Not statistical: "What is the title of our math book?"

2. Why does variability matter?

Statistics is about studying data that changes. If every answer were exactly the same, there would not be much to analyze.

When we ask a statistical question, we expect the answers to spread out. Some answers may be small, some large, and many may be somewhere in the middle.

For example, if you ask, "How many pets do students in our grade have?", some students may have 0 pets, some may have 1 or 2, and a few may have more. That spread in answers is variability.

3. Population and sample

In statistics, the population is the whole group you want to learn about.

A sample is a smaller group taken from the population.

For example, if you want to know how students in your school feel about school lunch:

  • The population is all students in the school.
  • The sample might be 50 students who are asked to answer a survey.

We usually study a sample because asking every person in a population can take too much time.

4. What makes a sample representative?

A sample is representative if it fairly reflects the population. That means the sample should include different types of people or items from the whole group, not just one small part of it.

For example, if a school has students from grades 6, 7, and 8, a representative sample about school lunch should include students from all three grades, not only 8th graders.

If the sample is not representative, the results may not match the population very well.

5. Sampling methods

Different ways of choosing a sample can affect how fair the results are.

Good sampling method: random sampling

In a random sample, each member of the population has a fair chance of being chosen.

Examples of random sampling:

  • Putting all student names in a container and drawing some names
  • Using a random number generator to choose survey participants

Random sampling helps reduce unfairness and gives a better chance of getting a representative sample.

Less fair sampling methods

  • Convenience sample: choosing people who are easiest to reach
  • Voluntary response sample: only people who choose to respond are counted

These methods can lead to bias because they may leave out important parts of the population.

6. What is bias?

Bias is anything in a study that makes the results unfairly favor one outcome over another.

Bias can happen in different ways, but in 7th Grade statistics, two important kinds are:

  • Sampling bias
  • Question wording bias

7. Sampling bias

Sampling bias happens when the sample does not fairly represent the population.

Example: A student wants to know whether students at school like reading. She surveys only students in the library.

This sample is biased because students in the library may like reading more than other students. The sample does not represent the whole school fairly.

Another example: A survey asks students whether they play sports, but it is given only to students at basketball practice. That sample is likely to overestimate how many students play sports.

8. Question wording bias

Question wording bias happens when the way a question is asked pushes people toward a certain answer.

For example, look at this question:

"Don't you agree that our school should have a longer lunch period?"

This question is biased because it suggests that agreeing is the expected answer.

A better version would be:

"Do you think the lunch period should be longer, shorter, or stay the same?"

This version is more neutral. It does not pressure people to answer in one direction.

9. How to write a good statistical question

A good statistical question should:

  • Ask about a group, not just one person or one object
  • Expect different answers
  • Be clear and neutral
  • Match the population you want to study

Here is a simple checklist:

  1. Does the question involve data from many people or items?
  2. Will the answers vary?
  3. Is the wording fair?
  4. Is the sample chosen in a fair way?

Worked Example 1: Is it a statistical question?

Question: "What time do students in our class go to bed on school nights?"

Step 1: Does it ask about a group? Yes, it asks about students in the class.

Step 2: Will the answers vary? Yes. Different students go to bed at different times.

Conclusion: This is a statistical question.

Now compare it to: "What time does our school start?"

There is one set answer, so it is not a statistical question.

Worked Example 2: Identify sampling bias

A student wants to know how many hours 7th graders spend on video games each week. He surveys 20 students in the gaming club.

Step 1: What is the population? All 7th graders.

Step 2: Who is in the sample? Only students in the gaming club.

Step 3: Is the sample representative? Probably not. Students in the gaming club may play more video games than other 7th graders.

Conclusion: The study has sampling bias.

A better method would be to randomly choose 20 students from all 7th graders.

Worked Example 3: Improve a biased question

Original question: "Why is the new school rule unfair?"

This question is biased because it assumes the rule is unfair.

A better question would be:

"What do you think about the new school rule?"

Or, if you want choices:

"Do you think the new school rule is fair, unfair, or are you unsure?"

These revised questions are more neutral.

Worked Example 4: Choose the better study

A school wants to know whether students want more after-school clubs.

Study A: Survey students who are already staying after school for clubs.

Study B: Randomly survey students from different grades during lunch.

Step 1: Which sample is more representative of the whole school?

Answer: Study B.

Step 2: Why?

Study A is biased because students already in clubs may be more likely to want more clubs. Study B includes a wider range of students and uses a more fair method.

10. Watch out for common mistakes

  • Thinking every question about data is statistical. It must expect variability.
  • Surveying only friends or nearby people and assuming the sample is fair.
  • Using questions that sound like they want a certain answer.
  • Forgetting to match the sample to the population.

11. Quick comparison table

  • Statistical question: expects different answers
  • Non-statistical question: expects one definite answer
  • Representative sample: fairly reflects the population
  • Biased sample: unfairly leaves out or overuses part of the population
  • Neutral wording: does not push toward one answer
  • Biased wording: suggests how someone should answer

12. How this connects to real data

Once data is collected, students often find measures like the mean, median, or range, or make graphs such as dot plots and box plots. But these summaries are only useful if the data came from a good statistical question and a fair sample.

If the question is biased or the sample is unfair, then even careful calculations may lead to poor conclusions.

Summary

A statistical question is a question that expects varied answers. To answer it fairly, you need a representative sample, which often comes from random sampling.

Bias can happen when the sample is unfair or when the question is worded in a way that pushes people toward a certain answer. Good statistical studies use clear, neutral questions and fair sampling methods.

When you look at a survey or design your own, always ask: Does the question allow for variability? Is the sample representative? Is the wording unbiased?

Put what you read to the test

You've worked through Statistical Questioning and Bias. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Measures of Center and Outlier Effects

Measures of Center and Outlier Effects

When we collect data, we often want to describe what is typical or what the “center” of the data is. In 7th Grade, the three main measures of center are the mean, median, and mode.

But there is an important idea to remember: sometimes one unusual value can change the answer a lot. This unusual value is called an outlier. In this lesson, you will learn how to find mean, median, and mode, and how to decide which one best describes a data set.

1. The three measures of center

  • Mean: the average. Add all the values, then divide by the number of values.
  • Median: the middle value when the data is put in order from least to greatest.
  • Mode: the value that appears most often.

Each measure gives information about the center, but they do not always give the same result.

2. How to find each measure

Mean

To find the mean, use this rule:

$$ \text{mean} = \frac{\text{sum of all data values}}{\text{number of data values}} $$

Median

To find the median:

  1. Put the data in order from least to greatest.
  2. Find the middle number.
  3. If there are two middle numbers, add them and divide by 2.

Mode

To find the mode, look for the number that appears the most.

  • A data set can have one mode.
  • It can have more than one mode if two or more numbers appear the same greatest number of times.
  • It can have no mode if no value repeats.

3. What is an outlier?

An outlier is a data value that is much larger or much smaller than the other values in the set.

For example, in the data set 8, 9, 10, 10, 11, 50, the number 50 is far away from the rest of the data. So 50 is an outlier.

Outliers matter because they can pull the mean up or down. The median is usually affected much less.

4. Which measure is best?

Different situations call for different measures of center.

  • Use the mean when the data is fairly even and there are no extreme outliers.
  • Use the median when there is an outlier or when the data is skewed to one side.
  • Use the mode when you want to know the most common value.

The word skewed means the data is not balanced. More values may be stretched out on one side than the other. In these cases, the median often gives a better picture of a typical value.

Worked Example 1: Finding mean, median, and mode

Find the mean, median, and mode of the data set:

4, 6, 7, 7, 8

Step 1: Mean

Add the values:

$$ 4+6+7+7+8=32 $$

There are 5 values, so divide by 5:

$$ \text{mean} = \frac{32}{5} = 6.4 $$

Step 2: Median

The data is already in order:

4, 6, 7, 7, 8

The middle value is 7, so the median is 7.

Step 3: Mode

The number 7 appears most often, so the mode is 7.

Answer:

  • Mean: 6.4
  • Median: 7
  • Mode: 7

Worked Example 2: Finding the median with an even number of values

Find the median of this data set:

3, 5, 6, 8, 10, 12

There are 6 values, so there is no single middle number. The two middle numbers are 6 and 8.

Add them and divide by 2:

$$ \frac{6+8}{2} = \frac{14}{2} = 7 $$

So the median is 7.

Worked Example 3: How an outlier affects the mean

First, look at this data set:

10, 11, 12, 12, 15

Mean:

$$ 10+11+12+12+15=60 $$ $$ \text{mean} = \frac{60}{5} = 12 $$

Median:

The middle value is 12, so the median is 12.

Now suppose one value changes and the data becomes:

10, 11, 12, 12, 50

Now 50 is an outlier.

New mean:

$$ 10+11+12+12+50=95 $$ $$ \text{mean} = \frac{95}{5} = 19 $$

New median:

The middle value is still 12, so the median is 12.

What changed?

  • The mean changed from 12 to 19.
  • The median stayed 12.

This shows that the mean is strongly affected by an outlier, but the median is not affected as much.

Worked Example 4: Choosing the best measure of center

A class recorded how many minutes students read at home last night:

20, 25, 25, 30, 30, 35, 120

Let’s find the measures of center.

Mean:

$$ 20+25+25+30+30+35+120=285 $$ $$ \text{mean} = \frac{285}{7} \approx 40.7 $$

Median:

The data is already in order. The middle value is 30, so the median is 30.

Mode:

25 appears twice and 30 appears twice. So the data set has two modes: 25 and 30.

Which measure best describes a typical reading time?

The value 120 is much larger than the others, so it is an outlier. The mean is about 40.7, but most students read much less than that.

So the median, 30, is the best measure of a typical reading time in this data set.

5. Comparing the measures

Here is a simple way to think about the three measures:

  • Mean: uses every value in the data set.
  • Median: depends on the middle position.
  • Mode: shows what happens most often.

Because the mean uses every value, one very large or very small value can change it a lot. That is why outliers have the biggest effect on the mean.

6. Quick tips for choosing the best measure

  • If the data has no outliers and looks balanced, the mean is often a good choice.
  • If the data has an outlier, use the median to describe what is typical.
  • If you want to know the most common answer, use the mode.

7. Common mistakes to avoid

  • Do not find the median before putting the numbers in order.
  • Do not forget to divide by the number of values when finding the mean.
  • Do not assume every data set has a mode.
  • Do not choose the mean automatically if there is a clear outlier.

8. Summary

The mean, median, and mode are all measures of center. The mean is the average, the median is the middle value, and the mode is the most common value.

An outlier is a value far from the rest of the data. Outliers usually affect the mean the most, while the median is usually more reliable when there is an outlier or a skewed data set. When deciding what is most typical, choose the measure that best matches the shape of the data.

Put what you read to the test

You've worked through Measures of Center and Outlier Effects. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Measures of Variability (Range and IQR)

Measures of Variability: Range and Interquartile Range (IQR)

When we study data, we do not only want to know the middle or typical value. We also want to know how spread out the data is.

The amount that data is spread out is called its variability. In this lesson, you will learn two important measures of variability:

  • Range
  • Interquartile Range (IQR)

These measures help us answer questions like:

  • Are the values close together or far apart?
  • Is one data set more consistent than another?
  • Are there values that seem unusually high or low?

Why does variability matter?

Imagine two basketball players who both score an average of 12 points per game. One player scores close to 12 every game. The other player sometimes scores 2 points and sometimes 22 points. Even though their averages are the same, their data is very different.

Measures of variability help us describe that difference.


1. Range

The range tells how far apart the smallest and largest values are in a data set.

To find the range, use this rule:

$$\text{Range} = \text{greatest value} - \text{least value}$$

Steps for finding the range:

  1. Find the smallest number.
  2. Find the largest number.
  3. Subtract.

What range tells us:

  • A small range means the data values are closer together.
  • A large range means the data values are more spread out.

Important note: The range uses only the smallest and largest values. That means one unusual value can change the range a lot.


2. Interquartile Range (IQR)

The Interquartile Range, or IQR, measures the spread of the middle half of the data.

Instead of looking at only the smallest and largest values, IQR looks at where most of the data is grouped.

To understand IQR, we need to know about quartiles.

Quartiles divide ordered data into 4 equal parts.

  • Q1 is the middle of the lower half of the data.
  • Q3 is the middle of the upper half of the data.

Then we calculate:

$$\text{IQR} = Q_3 - Q_1$$

What IQR tells us:

  • A small IQR means the middle half of the data is packed closely together.
  • A large IQR means the middle half of the data is more spread out.

IQR is often more useful than range because it is less affected by extreme values.


How to Find the IQR

Follow these steps carefully:

  1. Put the data in order from least to greatest.
  2. Find the median of the whole data set.
  3. Separate the data into a lower half and an upper half.
  4. Find the median of the lower half. This is Q1.
  5. Find the median of the upper half. This is Q3.
  6. Subtract: $$Q_3 - Q_1$$

Tip: When the data set has an odd number of values, do not include the middle value when finding the lower and upper halves.


Worked Example 1: Finding the Range

Find the range of this data set:

\(4, 7, 9, 10, 12\)

Step 1: Find the least value: \(4\)

Step 2: Find the greatest value: \(12\)

Step 3: Subtract:

$$12 - 4 = 8$$

Answer: The range is 8.

This means the data stretches across 8 units from the smallest to the largest value.


Worked Example 2: Finding the IQR with an Odd Number of Data Values

Find the IQR of this data set:

\(2, 4, 5, 8, 9, 12, 15\)

Step 1: Order the data

The data is already in order.

Step 2: Find the median

There are 7 numbers, so the middle value is the 4th number: \(8\)

Step 3: Split into two halves

Lower half: \(2, 4, 5\)

Upper half: \(9, 12, 15\)

Notice that the median, \(8\), is not included in either half.

Step 4: Find \(Q_1\)

The median of \(2, 4, 5\) is \(4\), so \(Q_1 = 4\)

Step 5: Find \(Q_3\)

The median of \(9, 12, 15\) is \(12\), so \(Q_3 = 12\)

Step 6: Find the IQR

$$\text{IQR} = Q_3 - Q_1 = 12 - 4 = 8$$

Answer: The IQR is 8.

This means the middle 50% of the data has a spread of 8 units.


Worked Example 3: Finding the IQR with an Even Number of Data Values

Find the IQR of this data set:

\(3, 5, 6, 8, 10, 11, 13, 14\)

Step 1: Order the data

The data is already in order.

Step 2: Find the median

There are 8 values, so the median is the average of the 4th and 5th values.

Those values are \(8\) and \(10\).

$$\text{Median} = \frac{8+10}{2} = 9$$

Step 3: Split into two halves

Lower half: \(3, 5, 6, 8\)

Upper half: \(10, 11, 13, 14\)

Step 4: Find \(Q_1\)

Find the median of \(3, 5, 6, 8\):

$$Q_1 = \frac{5+6}{2} = 5.5$$

Step 5: Find \(Q_3\)

Find the median of \(10, 11, 13, 14\):

$$Q_3 = \frac{11+13}{2} = 12$$

Step 6: Find the IQR

$$\text{IQR} = 12 - 5.5 = 6.5$$

Answer: The IQR is 6.5.


Worked Example 4: Comparing Range and IQR

Suppose two classes take the same quiz.

Class A scores: \(70, 72, 74, 75, 76, 78, 80\)

Class B scores: \(50, 72, 74, 75, 76, 78, 100\)

Both sets are already in order.

Find the range for each class.

Class A:

$$80 - 70 = 10$$

Class B:

$$100 - 50 = 50$$

So Class B has a much larger range.

Now find the IQR for each class.

For both data sets, the median is \(75\).

Lower half for both: \(70, 72, 74\) in Class A and \(50, 72, 74\) in Class B

Upper half for both: \(76, 78, 80\) in Class A and \(76, 78, 100\) in Class B

Class A:

\(Q_1 = 72\), \(Q_3 = 78\)

$$\text{IQR} = 78 - 72 = 6$$

Class B:

\(Q_1 = 72\), \(Q_3 = 78\)

$$\text{IQR} = 78 - 72 = 6$$

What does this show?

  • The range is very different because Class B has extreme values: 50 and 100.
  • The IQR is the same because the middle half of the scores is spread the same way in both classes.

This is why IQR can give a better picture of what most of the data is doing.


When to Use Range and IQR

  • Use range when you want a quick idea of the total spread from smallest to largest.
  • Use IQR when you want to describe the spread of the middle half of the data.
  • IQR is usually better when the data has very high or very low values that do not match the rest.

Common Mistakes to Avoid

  • Forgetting to order the data first. Always put the numbers in order before finding quartiles.
  • Using the wrong numbers for range. Only use the smallest and largest values.
  • Including the median in both halves when there is an odd number of values. Leave it out when finding \(Q_1\) and \(Q_3\).
  • Mixing up median and quartiles. The median is the middle of the whole set. \(Q_1\) and \(Q_3\) are the middles of the two halves.

Quick Check

Try these on your own:

  1. Find the range of \(6, 9, 12, 14, 20\).
  2. Find the IQR of \(1, 3, 4, 6, 8, 9, 12\).

Answers:

  • Range: \(20 - 6 = 14\)
  • Median is \(6\), lower half is \(1, 3, 4\), upper half is \(8, 9, 12\), so \(Q_1=3\), \(Q_3=9\), and IQR is \(9-3=6\)

Summary

Measures of variability tell us how spread out data is.

  • Range is the difference between the greatest and least values.
  • IQR is the difference between \(Q_3\) and \(Q_1\).
  • Range uses the ends of the data set.
  • IQR uses the middle 50% of the data.

If a data set has unusual high or low values, the IQR often gives a better idea of how consistent the data really is.

Put what you read to the test

You've worked through Measures of Variability (Range and IQR). Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Mean Absolute Deviation (MAD)

Mean Absolute Deviation (MAD) is a way to measure how spread out a data set is.

It tells us, on average, how far the data values are from the mean (the average).

If the MAD is small, the data values are close to the mean. If the MAD is large, the data values are more spread out.

This makes MAD a useful measure of statistical spread.

Why do we use MAD?

Sometimes two data sets can have the same mean, but one set is tightly grouped while the other is spread out. MAD helps us notice that difference.

For example, these two sets both have mean 10:

  • Set A: 9, 10, 10, 11
  • Set B: 2, 10, 10, 18

Both sets average to 10, but Set B is much more spread out. MAD helps measure that spread.

What does “absolute deviation” mean?

A deviation is the distance between a data value and the mean.

To find a deviation, subtract the mean from the data value:

\(\text{deviation} = \text{data value} - \text{mean}\)

But some answers may be negative and some positive. Since we want distance, we use absolute value, which makes every distance positive.

So an absolute deviation is:

$$\text{absolute deviation} = |\text{data value} - \text{mean}|$$

The mean absolute deviation is the average of all the absolute deviations.

$$\text{MAD} = \frac{\text{sum of absolute deviations}}{\text{number of data values}}$$

Steps for finding MAD

  1. Find the mean of the data set.
  2. Find the distance from each data value to the mean.
  3. Make each distance positive by using absolute value.
  4. Add the absolute deviations.
  5. Divide by the number of data values.

Let’s work through some examples.

Example 1: A simple data set

Find the MAD of: 4, 6, 8

Step 1: Find the mean

$$\text{mean} = \frac{4+6+8}{3} = \frac{18}{3} = 6$$

Step 2: Find each absolute deviation from the mean

  • For 4: \(|4-6|=2\)
  • For 6: \(|6-6|=0\)
  • For 8: \(|8-6|=2\)

Step 3: Find the average of those absolute deviations

$$\text{MAD} = \frac{2+0+2}{3} = \frac{4}{3} = 1\frac{1}{3}$$

Answer: The MAD is \(1\frac{1}{3}\).

This means the data values are, on average, \(1\frac{1}{3}\) units away from the mean.

Example 2: A data set with more values

Find the MAD of: 3, 5, 7, 9, 11

Step 1: Find the mean

$$\text{mean} = \frac{3+5+7+9+11}{5} = \frac{35}{5} = 7$$

Step 2: Find the absolute deviations

  • \(|3-7|=4\)
  • \(|5-7|=2\)
  • \(|7-7|=0\)
  • \(|9-7|=2\)
  • \(|11-7|=4\)

Step 3: Average the absolute deviations

$$\text{MAD} = \frac{4+2+0+2+4}{5} = \frac{12}{5} = 2.4$$

Answer: The MAD is \(2.4\).

Example 3: Comparing two data sets

Both of these data sets have the same mean. Which one has the greater spread?

  • Set A: 8, 9, 10, 11, 12
  • Set B: 2, 6, 10, 14, 18

Step 1: Find the mean of each set

For Set A:

$$\text{mean} = \frac{8+9+10+11+12}{5} = \frac{50}{5} = 10$$

For Set B:

$$\text{mean} = \frac{2+6+10+14+18}{5} = \frac{50}{5} = 10$$

Both means are 10.

Step 2: Find the MAD of Set A

Absolute deviations from 10:

  • \(|8-10|=2\)
  • \(|9-10|=1\)
  • \(|10-10|=0\)
  • \(|11-10|=1\)
  • \(|12-10|=2\)

$$\text{MAD of Set A} = \frac{2+1+0+1+2}{5} = \frac{6}{5} = 1.2$$

Step 3: Find the MAD of Set B

Absolute deviations from 10:

  • \(|2-10|=8\)
  • \(|6-10|=4\)
  • \(|10-10|=0\)
  • \(|14-10|=4\)
  • \(|18-10|=8\)

$$\text{MAD of Set B} = \frac{8+4+0+4+8}{5} = \frac{24}{5} = 4.8$$

Answer: Set B has the greater spread because its MAD is larger.

This shows why MAD is helpful. Even though both sets have the same mean, Set B is much more spread out.

Example 4: MAD in a real-life situation

A student records the number of books read by 5 friends in one month: 2, 4, 4, 5, 10

Find the mean and the MAD.

Step 1: Find the mean

$$\text{mean} = \frac{2+4+4+5+10}{5} = \frac{25}{5} = 5$$

Step 2: Find the absolute deviations from the mean

  • \(|2-5|=3\)
  • \(|4-5|=1\)
  • \(|4-5|=1\)
  • \(|5-5|=0\)
  • \(|10-5|=5\)

Step 3: Find the MAD

$$\text{MAD} = \frac{3+1+1+0+5}{5} = \frac{10}{5} = 2$$

Answer: The mean is 5 books, and the MAD is 2 books.

This means the number of books read is, on average, 2 books away from the mean.

Important ideas to remember

  • The mean is the average of the data values.
  • A deviation tells how far a value is from the mean.
  • An absolute deviation is always positive or zero.
  • The MAD is the average of those absolute deviations.
  • A smaller MAD means the data are closer to the mean.
  • A larger MAD means the data are more spread out.

Common mistakes to avoid

  • Forgetting to find the mean first. MAD is based on distance from the mean.
  • Not using absolute value. Negative and positive deviations should not cancel each other out.
  • Dividing by the wrong number. Divide by the total number of data values.
  • Mixing up mean and MAD. The mean describes the center. MAD describes the spread.

Quick check

Try this on your own: Find the MAD of 6, 6, 8, 10.

Solution

First find the mean:

$$\text{mean} = \frac{6+6+8+10}{4} = \frac{30}{4} = 7.5$$

Now find the absolute deviations:

  • \(|6-7.5|=1.5\)
  • \(|6-7.5|=1.5\)
  • \(|8-7.5|=0.5\)
  • \(|10-7.5|=2.5\)

Find the average:

$$\text{MAD} = \frac{1.5+1.5+0.5+2.5}{4} = \frac{6}{4} = 1.5$$

So the MAD is \(1.5\).

Summary

Mean Absolute Deviation measures the average distance of data values from the mean.

To find it, first calculate the mean, then find each absolute deviation, and finally average those distances.

MAD helps us describe how much a data set varies. It is a useful tool for comparing how spread out different data sets are.

Put what you read to the test

You've worked through Mean Absolute Deviation (MAD). Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Box Plots and Quartile Distributions

Box Plots and Quartile Distributions

When we collect data, we often want to understand where the values are centered and how spread out they are. One useful graph for this is a box plot, also called a box-and-whisker plot.

A box plot shows the shape of a data set using five important numbers. These numbers help us quickly compare different groups of data.

In this lesson, you will learn how to find the five-number summary, how to divide data into quartiles, and how to draw and read a box plot.

1. What is a box plot?

A box plot is a graph that shows how data is distributed from smallest to largest. It uses a number line and marks five special values.

These five values are called the five-number summary:

  • Minimum: the smallest value
  • First quartile or Q_1: the middle of the lower half of the data
  • Median or Q_2: the middle value of the whole data set
  • Third quartile or Q_3: the middle of the upper half of the data
  • Maximum: the largest value

The box plot has:

  • a box from \(Q_1\) to \(Q_3\)
  • a line inside the box at the median
  • whiskers that stretch from the box to the minimum and maximum

2. What are quartiles?

The word quartile comes from the word quarter. Quartiles divide the data into four equal parts.

After the data is put in order:

  • The median splits the data into two halves.
  • \(Q_1\) splits the lower half into two parts.
  • \(Q_3\) splits the upper half into two parts.

That means each section of the box plot represents about 25% of the data.

3. How to find the five-number summary

Follow these steps:

  1. Put the data in order from least to greatest.
  2. Find the median.
  3. Find the lower half and the upper half.
  4. Find \(Q_1\), the median of the lower half.
  5. Find \(Q_3\), the median of the upper half.
  6. Identify the minimum and maximum.

Important note: If there is an odd number of data values, the middle number is the median and is not included in either half when finding \(Q_1\) and \(Q_3\).

4. Worked Example 1: Finding the five-number summary

Suppose the data set is:

3, 5, 6, 8, 9, 12, 14

The data is already in order.

There are 7 numbers, so the median is the 4th number:

$$\text{Median} = 8$$

Now split the data into two halves, not including the median:

  • Lower half: 3, 5, 6
  • Upper half: 9, 12, 14

Find the middle of each half:

$$Q_1 = 5 \qquad Q_3 = 12$$

Now list the five-number summary:

  • Minimum: 3
  • \(Q_1 = 5\)
  • Median: 8
  • \(Q_3 = 12\)
  • Maximum: 14

5. Worked Example 2: Data set with an even number of values

Find the five-number summary for:

2, 4, 7, 10, 13, 15, 18, 20

There are 8 numbers, so the median is the average of the 4th and 5th values:

$$\text{Median} = \frac{10+13}{2} = 11.5$$

Split the data into two equal halves:

  • Lower half: 2, 4, 7, 10
  • Upper half: 13, 15, 18, 20

Now find the median of each half:

$$Q_1 = \frac{4+7}{2} = 5.5$$

$$Q_3 = \frac{15+18}{2} = 16.5$$

So the five-number summary is:

  • Minimum: 2
  • \(Q_1 = 5.5\)
  • Median: 11.5
  • \(Q_3 = 16.5\)
  • Maximum: 20

6. How to draw a box plot

Once you know the five-number summary, you can draw the box plot.

  1. Draw a number line that fits all the data values.
  2. Mark the minimum, \(Q_1\), median, \(Q_3\), and maximum.
  3. Draw a box from \(Q_1\) to \(Q_3\).
  4. Draw a vertical line inside the box at the median.
  5. Draw whiskers from the box out to the minimum and maximum.

For Example 1, the box plot would use:

$$3 \qquad 5 \qquad 8 \qquad 12 \qquad 14$$

This means:

  • left whisker ends at 3
  • box starts at 5
  • median line is at 8
  • box ends at 12
  • right whisker ends at 14

7. What does the box tell us?

The box, from \(Q_1\) to \(Q_3\), shows the middle 50% of the data.

The length of the box tells us how spread out the middle part of the data is. A longer box means the middle values are more spread out. A shorter box means they are closer together.

The distance from minimum to maximum shows the overall spread of the data.

8. Interquartile Range (IQR)

The interquartile range, or IQR, tells how wide the middle 50% of the data is.

The formula is:

$$\text{IQR} = Q_3 - Q_1$$

For Example 1:

$$\text{IQR} = 12 - 5 = 7$$

So the middle half of the data has a spread of 7.

9. Worked Example 3: Finding IQR and making a box plot

Use this data set:

1, 2, 4, 6, 7, 8, 10, 12, 13

There are 9 numbers, so the median is the 5th number:

$$\text{Median} = 7$$

Do not include the median in either half.

  • Lower half: 1, 2, 4, 6
  • Upper half: 8, 10, 12, 13

Find the quartiles:

$$Q_1 = \frac{2+4}{2} = 3$$

$$Q_3 = \frac{10+12}{2} = 11$$

Minimum is 1 and maximum is 13.

So the five-number summary is:

  • Minimum: 1
  • \(Q_1 = 3\)
  • Median: 7
  • \(Q_3 = 11\)
  • Maximum: 13

Now find the IQR:

$$\text{IQR} = 11 - 3 = 8$$

To draw the box plot, place points at 1, 3, 7, 11, and 13, draw the box from 3 to 11, the median line at 7, and whiskers to 1 and 13.

10. How to compare two box plots

Box plots are especially useful when comparing two or more data sets. You can compare:

  • center by looking at the median
  • spread by looking at the range and IQR
  • balance by seeing whether one side of the box plot is longer than the other

If one box plot has a higher median, that group usually has larger values in the center.

If one box or whisker is much longer, that group has more spread in that part of the data.

11. Worked Example 4: Comparing two groups

Suppose two classes took the same quiz.

Class A five-number summary:

  • Minimum: 60
  • \(Q_1 = 70\)
  • Median: 78
  • \(Q_3 = 85\)
  • Maximum: 95

Class B five-number summary:

  • Minimum: 55
  • \(Q_1 = 65\)
  • Median: 78
  • \(Q_3 = 90\)
  • Maximum: 100

Let us compare them.

Step 1: Compare medians

Both classes have a median of 78, so their middle score is the same.

Step 2: Compare IQRs

For Class A:

$$\text{IQR}_A = 85 - 70 = 15$$

For Class B:

$$\text{IQR}_B = 90 - 65 = 25$$

Class B has a larger IQR, so the middle 50% of Class B's scores are more spread out.

Step 3: Compare ranges

For Class A:

$$95 - 60 = 35$$

For Class B:

$$100 - 55 = 45$$

Class B also has a larger overall spread.

Conclusion: The two classes have the same center, but Class B's scores are more spread out.

12. Common mistakes to avoid

  • Forgetting to put the data in order first. Quartiles must be found from ordered data.
  • Using the wrong median. Be careful when there are an even number of values.
  • Including the median in both halves when there is an odd number of values. Usually, the median is left out of both halves.
  • Mixing up \(Q_1\) and \(Q_3\). \(Q_1\) is in the lower half, and \(Q_3\) is in the upper half.
  • Thinking the box shows all the data. The box only shows the middle 50%.

13. Quick check questions

Try these on your own:

  1. Find the median of: 4, 6, 7, 9, 10
  2. Find the five-number summary of: 2, 3, 5, 8, 9, 11, 14
  3. If \(Q_1 = 12\) and \(Q_3 = 20\), what is the IQR?
  4. If two box plots have the same median, what else should you check to compare them?

Answers:

  • 1. Median = 7
  • 2. Minimum 2, \(Q_1 = 3\), Median 8, \(Q_3 = 11\), Maximum 14
  • 3. \(20 - 12 = 8\)
  • 4. Check the IQR, range, and shape of the box plot to compare spread.

14. Summary

A box plot is a quick way to show how data is spread out. It is built from the five-number summary: minimum, \(Q_1\), median, \(Q_3\), and maximum.

Quartiles divide the data into four parts, and the box shows the middle 50% of the data. The interquartile range, \(Q_3 - Q_1\), measures the spread of that middle half.

When comparing box plots, look at the median for center and the box and whiskers for spread. This helps you make good statistical comparisons between groups.

Put what you read to the test

You've worked through Box Plots and Quartile Distributions. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Histograms and Data Bins

Histograms and Data Bins

When we collect a lot of numerical data, it can be hard to understand just by looking at a long list of numbers. A histogram helps us organize the data so we can see patterns more clearly.

A histogram is a graph that shows how many data values fall into different intervals, also called bins. Histograms are especially useful for continuous data, which means data that can be measured and can take many values in a range, such as height, time, distance, or temperature.

For example, if we measured the running times of students, the times might be grouped into bins like 10–14 seconds, 15–19 seconds, and 20–24 seconds. Then we count how many times fall in each bin.

What makes a histogram different from a bar graph?

  • In a bar graph, the bars are separated, and the categories are usually different groups, like favorite sports.
  • In a histogram, the bars touch because the data is numerical and grouped into intervals that follow one another.

Main idea: A histogram shows the frequency, or count, of values in each bin.

Parts of a histogram

  • Title: tells what the data is about.
  • Horizontal axis: shows the bins, or intervals.
  • Vertical axis: shows the frequency, or how many data values are in each bin.
  • Bars that touch: show that the intervals are connected.

What are data bins?

Bins are groups of equal-width intervals used to organize data. If the data values go from 1 to 30, we might use bins of width 5:

$$1\text{–}5,\ 6\text{–}10,\ 11\text{–}15,\ 16\text{–}20,\ 21\text{–}25,\ 26\text{–}30$$

Each bin covers the same amount. In this case, the bin width is:

$$5$$

Using equal-width bins is important because it makes the graph fair and easy to compare.

How to make a histogram

  1. Find the smallest and largest data values.
  2. Choose a bin width.
  3. Create equal-width bins that cover all the data.
  4. Count how many data values fall into each bin.
  5. Draw bars for each bin. The height of each bar is the frequency.

Worked Example 1: Sorting data into bins

Suppose these are the numbers of minutes students read one evening:

$$12,\ 18,\ 21,\ 24,\ 25,\ 27,\ 29,\ 31,\ 33,\ 34,\ 38,\ 40$$

Let’s use bins of width 10:

$$10\text{–}19,\ 20\text{–}29,\ 30\text{–}39,\ 40\text{–}49$$

Now count the data in each bin:

  • 10–19: 12, 18 → frequency 2
  • 20–29: 21, 24, 25, 27, 29 → frequency 5
  • 30–39: 31, 33, 34, 38 → frequency 4
  • 40–49: 40 → frequency 1

The histogram would have bars with heights:

$$2,\ 5,\ 4,\ 1$$

This tells us most students read between 20 and 39 minutes.

Worked Example 2: Making sense of the histogram

Suppose a histogram of quiz scores has these frequencies:

  • 50–59: 1 student
  • 60–69: 3 students
  • 70–79: 7 students
  • 80–89: 6 students
  • 90–99: 2 students

What can we learn from this?

  • Most students scored in the 70s and 80s.
  • Very few students scored below 60.
  • The distribution has one main peak around the middle, so the data is clustered there.

Histograms help us answer questions like:

  • Where are most of the data values?
  • Are the values spread out or close together?
  • Does the graph look balanced, or does one side stretch out more?

Describing the shape of a histogram

After making a histogram, we look at its shape. The shape gives clues about the data.

  • Symmetric: The left and right sides look about the same.
  • Skewed right: Most data is on the lower end, with a tail stretching to the right.
  • Skewed left: Most data is on the higher end, with a tail stretching to the left.

If a histogram is symmetric, the data is balanced around the center. If it is skewed, one side has a longer tail than the other.

Worked Example 3: Recognizing shape

A teacher records how many minutes students practiced music this week. The histogram shows:

  • 0–9: 8 students
  • 10–19: 6 students
  • 20–29: 3 students
  • 30–39: 1 student

Most students practiced for a small number of minutes, and fewer students practiced for large numbers of minutes. The bars start tall on the left and get shorter to the right.

This distribution is skewed right because the tail stretches to the right.

Worked Example 4: Choosing bins carefully

Suppose these temperatures were recorded over 15 days:

$$61,\ 62,\ 63,\ 65,\ 66,\ 67,\ 68,\ 70,\ 71,\ 72,\ 74,\ 75,\ 76,\ 78,\ 79$$

If we use bins of width 5, we could choose:

$$60\text{–}64,\ 65\text{–}69,\ 70\text{–}74,\ 75\text{–}79$$

Now count each bin:

  • 60–64: 61, 62, 63 → 3
  • 65–69: 65, 66, 67, 68 → 4
  • 70–74: 70, 71, 72, 74 → 4
  • 75–79: 75, 76, 78, 79 → 4

The frequencies are very close:

$$3,\ 4,\ 4,\ 4$$

This histogram would look fairly even, showing that the temperatures are spread across the range without one bin being much taller than the others.

Tips for using histograms well

  • Use equal-width bins.
  • Make sure all the data values fit into the bins.
  • Be careful when counting values on the edges of bins.
  • Label both axes clearly.
  • Remember that histogram bars touch.

Common mistakes to avoid

  • Using bins with different widths.
  • Leaving gaps between bars.
  • Forgetting to include all data values.
  • Choosing bins that are too small or too large, making the graph hard to read.

How histograms help us make statistical arguments

Histograms do more than organize data. They help us explain what the data shows. For example, if one histogram has most of its bars on the high end, we can argue that the data values are generally larger. If another histogram is spread out, we can say the data has more variety.

By looking at the bins, frequencies, and shape, we can describe the data clearly and support our thinking with evidence from the graph.

Summary

A histogram is a graph used for numerical data that has been grouped into equal-width intervals called bins. The height of each bar shows the frequency of the data in that interval.

To read or create a histogram, choose equal-width bins, count how many values fall in each bin, and draw touching bars. Then study the shape of the graph to describe where the data is clustered and whether it is symmetric, skewed right, or skewed left.

Put what you read to the test

You've worked through Histograms and Data Bins. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Scatter Plots and Bivariate Data

Scatter Plots and Bivariate Data

In statistics, we often collect information to answer questions. Sometimes we study one kind of data, like the heights of students in a class. Other times, we study two related kinds of data at the same time. This is called bivariate data.

Bivariate data means data with two variables for each item or person. For example, for each student, we might record:

  • hours of study
  • test score

Because each student has both values, the data comes in pairs. We can write the pair as an ordered pair, like \((3, 85)\), where 3 is the number of study hours and 85 is the test score.

One of the best ways to display bivariate data is with a scatter plot.

A scatter plot shows pairs of numbers as points on a coordinate plane. Each point represents one set of related data.

Scatter plots help us look for patterns, such as:

  • whether the data points rise or fall from left to right
  • whether the points are close together or spread out
  • whether groups of points form clusters
  • whether some points do not fit the pattern

1. Understanding the Two Variables

In bivariate data, the two variables usually have different jobs.

  • The independent variable goes on the horizontal axis, or x-axis.
  • The dependent variable goes on the vertical axis, or y-axis.

The independent variable is the one you choose or measure first. The dependent variable is the one that may change depending on the independent variable.

For example, if we compare hours practiced and free-throw baskets made:

  • hours practiced = independent variable
  • baskets made = dependent variable

This means we would plot:

$$x = \text{hours practiced}, \qquad y = \text{baskets made}$$

2. How to Make a Scatter Plot

To make a scatter plot, follow these steps:

  1. Collect paired data.
  2. Draw a coordinate plane.
  3. Label the x-axis with the independent variable.
  4. Label the y-axis with the dependent variable.
  5. Choose a scale that fits the data.
  6. Plot each ordered pair as a point.

Each point should be placed carefully. For example, the point \((4, 12)\) means move 4 units right on the x-axis and 12 units up on the y-axis.

3. Looking for Patterns in a Scatter Plot

After plotting the points, the next step is to study the pattern.

There are three common types of association, or relationship, in scatter plots:

  • Positive association: As x increases, y tends to increase.
  • Negative association: As x increases, y tends to decrease.
  • No association: There is no clear pattern between x and y.

If points go upward from left to right, the association is positive.

If points go downward from left to right, the association is negative.

If points seem scattered without a pattern, there may be no association.

4. Clusters and Outliers

Sometimes points gather in one or more groups. These groups are called clusters.

A cluster can show that many data values are similar. For example, if many students studied between 2 and 4 hours and scored between 75 and 90, those points may form a cluster.

Sometimes one point is far away from most of the others. This is called an outlier.

An outlier may happen because:

  • the value is unusual but correct
  • there was a measuring mistake
  • the data was recorded incorrectly

Outliers are important because they can affect how we describe the data.

5. Worked Example 1: Plotting Paired Data

A teacher records the number of hours 5 students studied and their quiz scores.

The data is:

  • Student A: \((1, 65)\)
  • Student B: \((2, 70)\)
  • Student C: \((3, 78)\)
  • Student D: \((4, 85)\)
  • Student E: \((5, 92)\)

Step 1: Identify the variables.

  • x-axis: hours studied
  • y-axis: quiz score

Step 2: Plot the points \((1,65)\), \((2,70)\), \((3,78)\), \((4,85)\), and \((5,92)\).

Step 3: Describe the pattern.

As study hours increase, quiz scores also increase. This is a positive association.

6. Worked Example 2: Finding a Negative Association

A class records the outside temperature and the number of cups of hot chocolate sold at the school stand.

The data is:

  • \((30, 40)\)
  • \((40, 34)\)
  • \((50, 25)\)
  • \((60, 18)\)
  • \((70, 10)\)

Here:

  • x = temperature
  • y = cups of hot chocolate sold

When the temperature goes up, the number of cups sold goes down.

So this scatter plot shows a negative association.

This makes sense in real life because people usually buy more hot chocolate when it is colder.

7. Worked Example 3: Looking for No Association and a Cluster

A student compares shoe size and number of books read last month for several classmates.

The data is:

  • \((4, 3)\)
  • \((5, 7)\)
  • \((6, 2)\)
  • \((7, 6)\)
  • \((8, 4)\)
  • \((9, 7)\)

If these points are plotted, they do not clearly rise or fall from left to right.

That means there is no clear association between shoe size and books read.

Also, several points are in the middle of the graph, around 4 to 7 books and shoe sizes 5 to 8. That group can be described as a cluster.

8. Worked Example 4: Noticing an Outlier

A coach records hours of practice and number of successful soccer shots.

The data is:

  • \((1, 6)\)
  • \((2, 9)\)
  • \((3, 11)\)
  • \((4, 14)\)
  • \((5, 15)\)
  • \((6, 3)\)

Most of the points show that more practice leads to more successful shots. That suggests a positive association.

But the point \((6, 3)\) is very different from the others. A player practiced for 6 hours but only made 3 successful shots.

This point is an outlier.

The outlier might mean:

  • the player was tired that day
  • the data was measured incorrectly
  • something unusual happened

9. How to Describe a Scatter Plot

When answering questions about scatter plots, it helps to use clear sentences. You can describe:

  • the variables
  • the type of association
  • any clusters
  • any outliers

Here is a good sentence frame:

The scatter plot compares ___ and ___. It shows a positive/negative/no association. There is a cluster around ___ and an outlier at ___.

10. Important Things to Remember

  • A scatter plot uses ordered pairs to show bivariate data.
  • The x-axis usually shows the independent variable.
  • The y-axis usually shows the dependent variable.
  • Positive association means both variables tend to increase together.
  • Negative association means one variable tends to decrease as the other increases.
  • No association means there is no clear pattern.
  • Clusters are groups of points close together.
  • Outliers are points far from the rest.

11. Common Mistakes to Avoid

  • Mixing up the axes: Always put the independent variable on the x-axis first.
  • Plotting the points in the wrong order: In \((x,y)\), the first number is horizontal and the second number is vertical.
  • Drawing conclusions too quickly: Look at all the points before deciding on the pattern.
  • Ignoring outliers: A point far away from the others may be important.

12. Quick Check

Suppose the data pairs are:

  • \((2, 10)\)
  • \((3, 12)\)
  • \((4, 14)\)
  • \((5, 16)\)

What pattern would you expect?

As x increases, y also increases. So the scatter plot would show a positive association.

If another point \((6, 2)\) were added, it would likely be an outlier because it does not fit the pattern.

Summary

Bivariate data gives two values for each item, and a scatter plot shows those pairs on a coordinate plane.

Scatter plots help us see whether there is a positive association, negative association, or no association.

They also help us notice clusters and outliers, which give us more information about the data.

When reading or making a scatter plot, always label the axes correctly, plot ordered pairs carefully, and look at the overall pattern of the points.

Put what you read to the test

You've worked through Scatter Plots and Bivariate Data. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Trend Lines and Line of Best Fit

Trend Lines and Line of Best Fit

When we collect two sets of related data, we can show them on a scatter plot. A scatter plot helps us see whether the data points follow a pattern.

Sometimes the points seem to move upward, downward, or stay scattered without a clear pattern. When the points form a general straight-line pattern, we can draw a trend line, also called a line of best fit.

A line of best fit is a straight line drawn through a scatter plot that shows the overall direction of the data. It does not have to pass through every point. Instead, it should be close to as many points as possible.

This line can help us:

  • see the overall pattern in the data,
  • describe whether the relationship is positive or negative,
  • make predictions between known values, called interpolation,
  • make predictions beyond the known values, called extrapolation.

1. Understanding scatter plots

In a scatter plot, each point represents a pair of values. For example, one value might be hours studied, and the other might be test score.

If the points rise from left to right, the data shows a positive association. This means that as one value increases, the other value tends to increase too.

If the points fall from left to right, the data shows a negative association. This means that as one value increases, the other value tends to decrease.

If the points do not follow any clear upward or downward pattern, then there is little or no association.

2. What makes a good line of best fit?

A good line of best fit should match the pattern of the scatter plot. When drawing one by hand, try to:

  • follow the middle of the group of points,
  • have about the same number of points above the line as below it,
  • be as close as possible to most of the points.

You should not try to connect dot to dot. A line of best fit is about the overall trend, not every single point.

3. How to sketch a trend line

  1. Look at the scatter plot and decide whether the points go up, go down, or show no clear pattern.
  2. Place a ruler so it follows the middle of the points.
  3. Adjust the ruler so the points are balanced above and below the line.
  4. Draw the line across the data.

If one point is far away from the others, it may be an outlier. An outlier can affect where the line is drawn, so be careful when looking at the overall pattern.

4. Using a line of best fit to make predictions

Once you have a trend line, you can use it to estimate values.

  • Interpolation means predicting a value inside the range of the data.
  • Extrapolation means predicting a value outside the range of the data.

Interpolation is usually more reliable because it stays within the data you already have. Extrapolation is less certain because it guesses beyond the known information.

Worked Example 1: Recognizing the trend

A scatter plot shows the number of hours a student reads each week and the number of books they finish each month. The points mostly rise from left to right.

Question: What type of association does the scatter plot show?

Solution: Since the points rise from left to right, the data shows a positive association.

This means that as reading time increases, the number of books finished also tends to increase.

Worked Example 2: Drawing a line of best fit

Suppose a scatter plot shows these points:

a student's practice time and free-throw shots made:

a few points are near \((1, 3)\), \((2, 4)\), \((3, 5)\), \((4, 6)\), and \((5, 7)\).

Question: What should the trend line look like?

Solution: The points follow an upward pattern. A good line of best fit would be a straight line rising from left to right through the middle of these points.

It might pass near points like \((1,3)\) and \((5,7)\), but it does not need to go through every point exactly.

This line shows that more practice time is connected to more shots made.

Worked Example 3: Interpolation

A line of best fit on a scatter plot of study hours and quiz scores goes through points close to \((2, 70)\) and \((6, 90)\).

Question: Use the line to estimate the quiz score for a student who studied 4 hours.

Solution: The value 4 hours is between 2 and 6, so this is interpolation.

From 2 hours to 6 hours, the score rises from 70 to 90. That is an increase of

$$90 - 70 = 20$$

over

$$6 - 2 = 4$$

hours.

So the score increases about

$$20 \div 4 = 5$$

points per hour.

From 2 hours to 4 hours is 2 more hours, so add

$$2 \times 5 = 10$$

points to 70:

$$70 + 10 = 80$$

Estimated score: \(80\)

Worked Example 4: Extrapolation

A scatter plot shows the relationship between the age of a tree and its height. A line of best fit goes near \((3, 12)\) and \((7, 20)\).

Question: Estimate the height of the tree at age 9 years.

Solution: The age 9 is beyond 7, so this is extrapolation.

From age 3 to age 7, the height increases from 12 to 20. That is

$$20 - 12 = 8$$

feet over

$$7 - 3 = 4$$

years.

So the height increases about

$$8 \div 4 = 2$$

feet per year.

From 7 years to 9 years is 2 more years, so add

$$2 \times 2 = 4$$

feet to 20:

$$20 + 4 = 24$$

Estimated height: \(24\) feet

Because this is extrapolation, the estimate may be less reliable than a prediction inside the data range.

5. Important ideas to remember

  • A trend line shows the general direction of data in a scatter plot.
  • A line of best fit should go through the middle of the points.
  • It does not need to pass through every point.
  • Positive association: points rise from left to right.
  • Negative association: points fall from left to right.
  • Interpolation is a prediction within the data range.
  • Extrapolation is a prediction outside the data range.

6. Common mistakes

  • Connecting the dots: A line of best fit is one straight line showing the pattern, not a path through every point.
  • Ignoring the middle: The line should balance the points above and below it.
  • Trusting extrapolation too much: Predictions outside the data range may not be accurate.
  • Forgetting the direction: Always check whether the trend is positive, negative, or none.

Quick Check

  1. If a scatter plot falls from left to right, what kind of association is it?
  2. If you predict a value between two known data values, is that interpolation or extrapolation?
  3. Does a line of best fit need to pass through every point?

Answers:

  1. Negative association
  2. Interpolation
  3. No, it should show the overall pattern

Summary

A trend line, or line of best fit, is a straight line that shows the overall pattern in a scatter plot. It helps us understand whether two sets of data have a positive or negative association and lets us estimate values.

When using a line of best fit, remember to draw it through the middle of the data, not through every point. Use interpolation for estimates within the data range and be more careful with extrapolation outside the data range.

Put what you read to the test

You've worked through Trend Lines and Line of Best Fit. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Two-Way Frequency Tables

Two-Way Frequency Tables help us organize and study data about two different categories at the same time.

For example, suppose a class survey asks students two questions: Do you play a sport? and Do you prefer math or science? A two-way table helps us sort all the answers so we can compare groups and look for patterns.

This lesson will show you how to:

  • read a two-way frequency table,
  • find row totals, column totals, and the grand total,
  • calculate relative frequencies, and
  • use the table to decide whether there may be an association between the two categories.

Important vocabulary

  • Categorical data: data that is sorted into groups, such as yes/no, boy/girl, or favorite subject.
  • Frequency: the number of times something happens.
  • Two-way frequency table: a table that shows counts for two categories.
  • Relative frequency: a fraction, decimal, or percent that compares a count to a total.
  • Row total: the total across a row.
  • Column total: the total down a column.
  • Grand total: the total number of all data values in the table.
  • Conditional relative frequency: a relative frequency found using only one row total or one column total.
  • Association: when one category seems connected to another category.

1. What does a two-way table look like?

A two-way table has one category listed across the top and another category listed down the side. Inside the table are the counts.

Here is an example. A teacher asks 30 students whether they like reading fiction or nonfiction, and whether they prefer reading on paper or on a screen.

$$ \begin{array}{|c|c|c|c|} \hline & \text{Paper} & \text{Screen} & \text{Total} \\ \hline \text{Fiction} & 12 & 6 & 18 \\ \hline \text{Nonfiction} & 5 & 7 & 12 \\ \hline \text{Total} & 17 & 13 & 30 \\ \hline \end{array} $$

This table compares two variables:

  • type of reading: fiction or nonfiction
  • reading format: paper or screen

Each inside number tells how many students are in both categories at once.

  • The number 12 means 12 students prefer fiction and paper.

  • The number 7 means 7 students prefer nonfiction and screen.

2. Totals in a two-way table

The totals help us understand the data better.

  • Row total: add across a row.
  • Column total: add down a column.
  • Grand total: the total number of all people or objects in the survey.

In the table above:

  • The fiction row total is \(12+6=18\).
  • The nonfiction row total is \(5+7=12\).
  • The paper column total is \(12+5=17\).
  • The screen column total is \(6+7=13\).
  • The grand total is \(30\).

You can also check your work by making sure the row totals and column totals both add to the grand total:

$$18+12=30 \quad \text{and} \quad 17+13=30$$

3. What is relative frequency?

A relative frequency compares a part to a whole. It can be written as a fraction, decimal, or percent.

The basic formula is:

$$\text{Relative Frequency}=\frac{\text{part}}{\text{whole}}$$

For example, in the reading table, 12 students chose fiction and paper out of 30 students total. So the relative frequency is

$$\frac{12}{30}=\frac{2}{5}=0.4=40\%$$

This means 40% of all students chose fiction and paper.

4. Joint relative frequency

When we use one inside cell and compare it to the grand total, we are finding a joint relative frequency.

Examples from the table:

  • Fiction and paper: \(\frac{12}{30}=40\%\)
  • Fiction and screen: \(\frac{6}{30}=20\%\)
  • Nonfiction and paper: \(\frac{5}{30}\approx 16.7\%\)
  • Nonfiction and screen: \(\frac{7}{30}\approx 23.3\%\)

These tell the percent of the whole group in each category pair.

5. Marginal relative frequency

When we use a row total or column total and compare it to the grand total, we are finding a marginal relative frequency.

Examples:

  • Students who prefer fiction: \(\frac{18}{30}=60\%\)
  • Students who prefer nonfiction: \(\frac{12}{30}=40\%\)
  • Students who prefer paper: \(\frac{17}{30}\approx 56.7\%\)
  • Students who prefer screen: \(\frac{13}{30}\approx 43.3\%\)

These describe the totals along the edges, or margins, of the table.

6. Conditional relative frequency

A conditional relative frequency compares a part to a row total or a column total instead of the grand total.

This helps answer questions such as:

  • Of the students who prefer fiction, how many prefer paper?
  • Of the students who prefer screen, how many prefer nonfiction?

The word of is a clue. It tells you which total to use.

Example: Of the students who prefer fiction, what fraction prefer paper?

Use the fiction row total, 18:

$$\frac{12}{18}=\frac{2}{3}\approx 66.7\%$$

So about 66.7% of the fiction group prefer paper.

Another example: Of the students who prefer screen, what fraction prefer nonfiction?

Use the screen column total, 13:

$$\frac{7}{13}\approx 53.8\%$$

So about 53.8% of the screen group prefer nonfiction.

7. How can a two-way table show association?

We look for an association by comparing conditional relative frequencies.

If the percentages are very similar between groups, there may be little or no association.

If the percentages are noticeably different, there may be an association between the two categories.

In the reading example:

  • Of the fiction group, \(\frac{12}{18}\approx 66.7\%\) prefer paper.
  • Of the nonfiction group, \(\frac{5}{12}\approx 41.7\%\) prefer paper.

Because 66.7% and 41.7% are not very close, reading type and reading format may be associated.

This does not prove one causes the other. It only shows there may be a pattern in the data.

Worked Example 1: Reading a table

A club surveys 40 students about whether they are in band and whether they like art.

$$ \begin{array}{|c|c|c|c|} \hline & \text{Like Art} & \text{Do Not Like Art} & \text{Total} \\ \hline \text{In Band} & 14 & 6 & 20 \\ \hline \text{Not in Band} & 8 & 12 & 20 \\ \hline \text{Total} & 22 & 18 & 40 \\ \hline \end{array} $$

Question A: How many students are in band and like art?

Look at the cell where In Band and Like Art meet.

Answer: 14 students.

Question B: How many students are not in band?

Look at the row total for Not in Band.

Answer: 20 students.

Question C: How many students were surveyed in all?

Look at the grand total.

Answer: 40 students.

Worked Example 2: Finding missing totals

A survey asks students whether they have a pet and whether they walk to school.

$$ \begin{array}{|c|c|c|c|} \hline & \text{Walk} & \text{Do Not Walk} & \text{Total} \\ \hline \text{Has Pet} & 9 & 11 & ? \\ \hline \text{No Pet} & 7 & 13 & ? \\ \hline \text{Total} & ? & ? & ? \\ \hline \end{array} $$

Step 1: Find each row total.

  • Has Pet: \(9+11=20\)
  • No Pet: \(7+13=20\)

Step 2: Find each column total.

  • Walk: \(9+7=16\)
  • Do Not Walk: \(11+13=24\)

Step 3: Find the grand total.

$$20+20=40$$

or

$$16+24=40$$

The completed table is:

$$ \begin{array}{|c|c|c|c|} \hline & \text{Walk} & \text{Do Not Walk} & \text{Total} \\ \hline \text{Has Pet} & 9 & 11 & 20 \\ \hline \text{No Pet} & 7 & 13 & 20 \\ \hline \text{Total} & 16 & 24 & 40 \\ \hline \end{array} $$

Worked Example 3: Relative frequencies

Use the pet and walking table above.

Question A: What is the joint relative frequency for students who have a pet and walk?

Use the cell 9 and divide by the grand total 40:

$$\frac{9}{40}=0.225=22.5\%$$

Answer: 22.5% of all students have a pet and walk.

Question B: What is the marginal relative frequency for students who do not walk?

Use the column total 24 and divide by 40:

$$\frac{24}{40}=0.6=60\%$$

Answer: 60% of all students do not walk.

Question C: Of the students who have a pet, what percent walk?

This is conditional relative frequency, so use the Has Pet row total 20:

$$\frac{9}{20}=0.45=45\%$$

Answer: 45% of students who have a pet walk.

Worked Example 4: Looking for association

A cafeteria surveys 50 students about whether they bring lunch from home and whether they buy milk.

$$ \begin{array}{|c|c|c|c|} \hline & \text{Buy Milk} & \text{Do Not Buy Milk} & \text{Total} \\ \hline \text{Bring Lunch} & 18 & 7 & 25 \\ \hline \text{Do Not Bring Lunch} & 10 & 15 & 25 \\ \hline \text{Total} & 28 & 22 & 50 \\ \hline \end{array} $$

Question: Is there an association between bringing lunch and buying milk?

Step 1: Compare conditional relative frequencies.

Of the students who bring lunch, the fraction who buy milk is

$$\frac{18}{25}=72\%$$

Of the students who do not bring lunch, the fraction who buy milk is

$$\frac{10}{25}=40\%$$

Step 2: Compare the percentages.

Since 72% and 40% are quite different, there appears to be an association.

Conclusion: Students who bring lunch are more likely to buy milk than students who do not bring lunch.

8. Tips for solving problems

  • Read the row and column labels carefully before choosing a number.
  • Decide what total to use: row total, column total, or grand total.
  • If the question says out of all, use the grand total.
  • If the question says of the students who..., use that row total or column total.
  • Turn fractions into decimals or percents when asked.
  • To look for association, compare conditional relative frequencies, not just the counts.

9. Common mistakes to avoid

  • Using the wrong total. For conditional relative frequency, do not divide by the grand total unless the question says all students.
  • Mixing up rows and columns. Check the labels each time.
  • Comparing counts instead of percents. Percentages are better for deciding whether there may be an association.
  • Forgetting to check totals. Row totals and column totals should both match the grand total.

Brief Summary

A two-way frequency table organizes data for two categories in one chart. The inside cells show counts for both categories together, and the margins show totals.

Relative frequencies compare counts to totals. Joint relative frequency uses the grand total, marginal relative frequency uses a row or column total compared to the grand total, and conditional relative frequency uses a row total or column total as the whole.

To decide whether there may be an association between two categories, compare conditional relative frequencies. If the percents are very different, the categories may be related in the data.

Put what you read to the test

You've worked through Two-Way Frequency Tables. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Recognizing Misleading Data Visualizations

Recognizing Misleading Data Visualizations

Sometimes graphs and charts help us understand information quickly. They can show which group has more, which has less, and how amounts compare.

But not every graph tells the story fairly. Some graphs are misleading. That means they make the data look bigger, smaller, or more different than it really is.

In this lesson, you will learn how to spot graphs that are unfair or tricky. You will also learn what a fair graph should look like.

What is a data visualization?

A data visualization is a picture that shows data. Some common kinds are:

  • bar graphs
  • picture graphs
  • line plots
  • tables turned into charts

These visuals should make data easier to read. A good graph should be clear, neat, and honest.

What makes a graph misleading?

A graph is misleading when it gives the wrong idea about the data. The numbers may be correct, but the picture can still trick your eyes.

Here are some common ways a graph can be misleading:

  • the scale does not count evenly
  • the graph starts at a number other than 0 when it should not
  • pictures are different sizes in an unfair way
  • bars are not the same width
  • important labels are missing

1. Watch the scale

The scale is the set of numbers shown on the side of a graph. A fair scale should usually go up by the same amount each time.

For example, this is an even scale:

0, 2, 4, 6, 8, 10

Each step goes up by 2. That is clear and fair.

This scale is uneven:

0, 2, 4, 10, 12

The jumps are not the same. First it goes up by 2, then by 2, then by 6, then by 2. That can make bars look confusing or unfair.

When you read a graph, ask yourself:

  • Do the numbers increase by the same amount each time?
  • Can I tell what each line or mark means?

2. Check where the graph starts

Many bar graphs should start at 0. If they do not, the bars can look much more different than they really are.

Look at these two amounts:

Class A read 8 books. Class B read 10 books.

The difference is:

$$10 - 8 = 2$$

That is only 2 books more. But if a graph starts at 7 instead of 0, the bar for 10 may look much taller than the bar for 8. This can trick you into thinking Class B read a lot more books.

So always check the bottom of the graph. Does it begin at 0? If not, be extra careful.

3. Look at picture sizes

In a picture graph, each picture should stand for the same amount. If one picture is made much bigger in both height and width, it can look like much more than it should.

For example, suppose 1 apple picture means 2 apples. If one group uses a giant apple picture and another uses a small apple picture, your eyes may think the giant one means much more, even if both pictures are supposed to count the same way.

A fair picture graph should:

  • use the same size pictures
  • tell what each picture means
  • show parts of pictures clearly if needed

4. Check the bar widths and spaces

In a bar graph, the bars should usually be the same width. The spaces between bars should also be even.

If one bar is much wider, it may look more important, even if its value is not bigger. That can mislead the reader.

5. Make sure labels are clear

A graph should tell you:

  • what the graph is about
  • what each bar or picture stands for
  • what the scale means

If labels are missing, it is hard to understand the data correctly.

Questions to ask when looking at a graph

When you see a graph, you can be a data detective! Ask:

  • Does the graph have a title?
  • Are the labels clear?
  • Does the scale count by equal amounts?
  • Does the graph start at 0?
  • Are the bars the same width?
  • Are the pictures the same size?
  • Does the graph make the data look fair?

Worked Example 1: Finding an uneven scale

A bar graph shows the number of pets owned by students. The side scale says:

0, 1, 2, 5, 6

Is this fair?

Step 1: Check how the numbers increase.

From 0 to 1 is 1.

From 1 to 2 is 1.

From 2 to 5 is 3.

From 5 to 6 is 1.

Step 2: Decide if the jumps are equal.

No, they are not equal.

Answer: This graph may be misleading because the scale is uneven.

Worked Example 2: Starting above 0

A graph compares how many stickers two students have.

  • Lena: 12 stickers
  • Omar: 14 stickers

The graph starts at 10 instead of 0.

Step 1: Find the real difference.

$$14 - 12 = 2$$

Step 2: Think about how the bars might look.

If the graph begins at 10, Lena's bar goes from 10 to 12, which is 2 units tall. Omar's bar goes from 10 to 14, which is 4 units tall.

Step 3: Compare what your eyes see to the real data.

The taller bar looks about twice as tall, but Omar only has 2 more stickers.

Answer: The graph is misleading because it makes a small difference look much bigger.

Worked Example 3: Picture graph trouble

A picture graph shows favorite snacks. The key says 1 picture = 3 votes.

Crackers has 2 pictures. Grapes has 3 pictures.

But the grape pictures are drawn much larger than the cracker pictures.

Step 1: Use the key, not just your eyes.

Crackers: \(2 \times 3 = 6\) votes

Grapes: \(3 \times 3 = 9\) votes

Step 2: Find the difference.

$$9 - 6 = 3$$

Step 3: Think about the picture sizes.

If the grape pictures are much bigger, grapes may look like they got far more than 9 votes.

Answer: The graph is misleading because the pictures are not the same size. The key says what matters, not the giant drawings.

Worked Example 4: Is this graph fair?

A bar graph shows how many cans were collected in a food drive.

  • Team Red: 4 cans
  • Team Blue: 8 cans
  • Team Green: 6 cans

The graph has:

  • a title
  • labels for each team
  • a scale of 0, 2, 4, 6, 8
  • bars that are the same width

Step 1: Check the scale.

The scale goes up by 2 each time. That is even.

Step 2: Check the starting number.

It starts at 0. That is good.

Step 3: Check the bars and labels.

The bars are the same width, and the graph is labeled.

Answer: This graph seems fair and not misleading.

Tips for making your own fair graph

When you make a graph, be honest with your data. You want others to understand the information correctly.

  • Choose an even scale.
  • Start at 0 when it makes sense.
  • Use equal-size pictures.
  • Make bars the same width.
  • Add a title and labels.
  • Check that the graph matches the data.

Why this matters

Graphs are used in school, books, news, and everyday life. If a graph is misleading, people may believe something that is not really true.

That is why it is important to look carefully, ask questions, and use the numbers to help you decide.

Summary

A misleading data visualization is a graph or chart that makes data look different from what it really is. You can spot one by checking the scale, the starting point, the picture sizes, the bar widths, and the labels.

Remember: do not trust only your eyes. Trust the numbers too.

Put what you read to the test

You've worked through Recognizing Misleading Data Visualizations. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.