Chapter 17

Descriptive Statistics and Data Analysis

Data Classification and Sampling

Lesson: Data Classification and Sampling

In statistics, we often work with data, which are pieces of information collected about people, objects, or events. Before we analyze data, we need to understand what type of data we have and how the data were collected. These two ideas are very important because they affect how we summarize data and how trustworthy our conclusions are.

In this lesson, you will learn how to classify variables as qualitative or quantitative, and as discrete or continuous. You will also learn the main sampling methods: random, stratified, systematic, and cluster sampling.

1. What is a variable?

A variable is any characteristic that can be recorded for an individual in a study. An individual could be a person, a school, a car, a package, or anything else being studied.

Examples of variables include:

  • Age of a student
  • Favorite sport
  • Number of siblings
  • Height of a plant
  • Type of phone used

2. Qualitative vs. Quantitative Variables

The first big classification is whether a variable is qualitative or quantitative.

Qualitative variables describe qualities or categories. They usually do not involve meaningful numerical calculations.

Examples of qualitative variables:

  • Eye color
  • Blood type
  • Political party preference
  • Favorite subject

Even if categories are given numbers, the variable can still be qualitative. For example, if students rate lunch as 1 = poor, 2 = fair, 3 = good, these numbers represent categories, not actual amounts.

Quantitative variables represent numerical values where arithmetic makes sense. These variables measure or count something.

Examples of quantitative variables:

  • Test score
  • Weight
  • Distance traveled
  • Number of pets

A useful question is: Does the number describe an amount? If yes, the variable is probably quantitative. If it labels a category, it is qualitative.

3. Discrete vs. Continuous Variables

Discrete and continuous are classifications of quantitative variables.

Discrete variables come from counting. They usually take separate values, often whole numbers.

Examples of discrete variables:

  • Number of students in a classroom
  • Number of text messages received today
  • Number of goals scored in a game

You cannot usually have values like 2.5 students or 7.3 goals, so these values are counted one by one.

Continuous variables come from measuring. They can take any value in an interval, including decimals.

Examples of continuous variables:

  • Height
  • Time
  • Temperature
  • Mass

For example, a runner's time might be 12.4 seconds, 12.43 seconds, or 12.431 seconds depending on the precision of the measuring tool.

Important idea:

  • Qualitative: categories
  • Quantitative: numbers that measure or count
  • Discrete: counted values
  • Continuous: measured values

4. Population and Sample

In statistics, the population is the entire group we want information about. A sample is a smaller group selected from the population.

For example:

  • If a school wants to know how much sleep its students get, the population is all students in the school.
  • If the school surveys 150 students, those 150 students are the sample.

We use samples because studying an entire population is often too expensive, takes too much time, or is not practical.

5. Why Sampling Method Matters

A sample should represent the population fairly. If the sample is chosen poorly, the results may be biased. A biased sample tends to favor certain outcomes and does not reflect the whole population well.

For example, if a school surveys only students in an honors math class about homework time, the sample may not represent all students in the school.

6. Main Sampling Methods

There are several common sampling methods. You should know how each one works and when it is useful.

A. Random Sampling

In a random sample, every individual in the population has an equal chance of being selected.

Example: Number all 2,000 students in a school from 1 to 2000, then use a random number generator to choose 100 students.

Why it is useful: It helps reduce bias because selection is based on chance, not personal choice.

B. Stratified Sampling

In a stratified sample, the population is divided into groups called strata that share a common trait. Then a random sample is taken from each stratum.

Example: A school divides students by grade level (9th, 10th, 11th, 12th) and randomly selects students from each grade.

Why it is useful: It ensures that important subgroups are represented in the sample.

C. Systematic Sampling

In a systematic sample, you choose individuals according to a fixed pattern after a random starting point.

Example: From an alphabetized list of students, start at the 7th student and then choose every 20th student after that.

If the starting point is random and the list has no hidden pattern that affects results, this method can work well.

D. Cluster Sampling

In a cluster sample, the population is divided into natural groups, called clusters. Then some clusters are randomly selected, and all individuals in those selected clusters are surveyed.

Example: A district randomly selects 3 homeroom classes and surveys every student in those classes.

Why it is useful: It is often easier and cheaper when the population is spread out.

7. Comparing Stratified and Cluster Sampling

Students often confuse stratified and cluster sampling, so it is important to compare them carefully.

  • Stratified sampling: divide into groups based on a characteristic, then randomly choose some individuals from every group.
  • Cluster sampling: divide into natural groups, then randomly choose entire groups and include everyone in those groups.

Example:

  • Stratified: choose some students from each grade level.
  • Cluster: choose a few entire classrooms and survey every student in them.

8. How to Identify the Sampling Method

Use these clues:

  • If everyone has an equal chance directly, think random.
  • If the population is split into categories and sampled from each category, think stratified.
  • If every nth person is selected, think systematic.
  • If whole groups are chosen, think cluster.

9. Worked Example 1: Classifying Variables

Classify each variable as qualitative or quantitative. If it is quantitative, say whether it is discrete or continuous.

  1. Number of books in a backpack
  2. Brand of shoes worn by a student
  3. Time spent studying for a test
  4. Number of languages spoken

Solution

  • Number of books in a backpack: quantitative, discrete. It is a count.
  • Brand of shoes worn by a student: qualitative. It names a category.
  • Time spent studying for a test: quantitative, continuous. Time is measured and can include decimals.
  • Number of languages spoken: quantitative, discrete. It is counted in whole numbers.

10. Worked Example 2: Identifying a Sampling Method

A principal wants to survey students about cafeteria food. She separates students into freshmen, sophomores, juniors, and seniors, then randomly chooses 25 students from each group.

What sampling method is this?

Solution

The students are first divided into groups by grade level, and then a random sample is taken from each group. This is stratified sampling.

Why not cluster sampling? Because the principal did not choose whole groups and survey everyone in them. She selected some individuals from each group.

11. Worked Example 3: Distinguishing Systematic and Random Sampling

A company has a list of 800 employees. It randomly chooses a starting point at employee 12, then selects every 15th employee after that.

What sampling method is this?

Solution

This is systematic sampling because the sample is chosen using a fixed interval: every 15th employee.

The random starting point helps make the sample fairer, but the key feature is the regular pattern.

12. Worked Example 4: Evaluating a Sampling Plan

A student wants to estimate how many hours teenagers spend on social media each day. She asks 30 students from the school library right after school.

Is this likely to be a good sample?

Solution

This is probably not a good sample because it may be biased. Students in the library right after school may not represent all teenagers. For example, they may spend their time differently from students who play sports, work after school, or go straight home.

A better method would be to take a random sample of students from the entire school, or use a stratified sample if the student wants to make sure different grade levels are represented.

13. Choosing the Best Sampling Method

Different situations call for different methods.

  • Random sampling is a strong general method when a full list of the population is available.
  • Stratified sampling is useful when you want all important subgroups represented.
  • Systematic sampling is simple and efficient when data are listed in order and no harmful pattern exists.
  • Cluster sampling is useful when the population is naturally grouped and it is easier to sample entire groups.

14. Common Mistakes to Avoid

  • Do not call a variable continuous unless it is quantitative first.
  • Do not assume that any variable with numbers is quantitative. Sometimes numbers are just labels.
  • Do not confuse stratified and cluster sampling.
  • Do not ignore bias. Even a large sample can give poor results if it is chosen badly.

15. Quick Check

Try these on your own:

  1. Classify: hair color
  2. Classify: number of missed school days
  3. Classify: body temperature
  4. A coach selects every 5th athlete from a roster after a random start. What method is used?
  5. A researcher randomly chooses 4 classrooms and surveys every student in those rooms. What method is used?

Answers

  • Hair color: qualitative
  • Number of missed school days: quantitative, discrete
  • Body temperature: quantitative, continuous
  • Every 5th athlete: systematic sampling
  • 4 classrooms, everyone surveyed: cluster sampling

Summary

Data classification helps us understand what kind of information we are working with. Qualitative variables describe categories, while quantitative variables describe numerical amounts. Quantitative variables can be discrete if they are counted, or continuous if they are measured.

Sampling is how we choose part of a population to study. Random, stratified, systematic, and cluster sampling each have different uses. Understanding these methods helps you decide whether a sample is fair and whether statistical conclusions are reliable.

Put what you read to the test

You've worked through Data Classification and Sampling. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Survey Design and Bias

Survey Design and Bias is an important part of statistics because the quality of data depends on how the data is collected. Even if you calculate averages, percentages, or graphs correctly, your conclusions can still be misleading if the survey itself was poorly designed.

In this lesson, you will learn what makes a good survey, what bias means, and how to recognize common types of bias such as selection bias, response bias, and measurement bias. You will also learn how to improve a survey so that the results better represent the population.

What is a survey?

A survey is a method of collecting information by asking questions to a group of people. The goal is usually to learn something about a larger group, called the population, by studying a smaller group, called the sample.

For example, a school might want to know how students travel to school. It may not be practical to ask every student, so it asks a sample of students and uses their answers to estimate the habits of the whole school.

What is bias?

Bias is a systematic error that causes survey results to consistently lean in a certain direction. Bias is different from random chance. Random variation may cause small differences, but bias creates unfair or unrepresentative results.

If a survey is biased, the data may not reflect the true opinions or behaviors of the population. That means conclusions drawn from the survey can be inaccurate.

Why survey design matters

A good survey should:

  • Clearly identify the population being studied
  • Use a sample that represents that population
  • Ask questions in a clear and neutral way
  • Measure responses consistently and accurately
  • Avoid methods that pressure or confuse participants

If any of these parts are weak, bias can enter the study.

Main types of bias

In this lesson, we focus on three major types:

  • Selection bias
  • Response bias
  • Measurement bias

1. Selection bias

Selection bias happens when the sample is chosen in a way that does not fairly represent the population. Some groups may be more likely to be included, while others are left out.

This means the survey results may reflect only part of the population instead of the whole group.

Common causes of selection bias

  • Using a sample of convenience, such as surveying only nearby people
  • Surveying volunteers who choose to respond
  • Surveying only one subgroup of the population
  • Excluding people who are harder to reach

Example of selection bias

A student wants to know whether teenagers in the city support building more skate parks. She surveys people at a skate park.

This sample is biased because people at a skate park are more likely than average to support skate parks. The sample does not represent all teenagers in the city.

How to reduce selection bias

  • Choose the sample randomly when possible
  • Make sure all members of the population have a fair chance to be selected
  • Include people from different groups, locations, or times

2. Response bias

Response bias happens when people do not answer survey questions truthfully or their answers are influenced by the wording of the question or the survey situation.

Even if the right people are surveyed, the answers may still be inaccurate.

Common causes of response bias

  • Leading questions that push people toward a certain answer
  • Loaded questions that contain assumptions or emotional language
  • Questions about private topics, where people may not answer honestly
  • People trying to give socially acceptable answers instead of truthful ones

Example of response bias

Suppose a survey asks, “Don’t you agree that our school should finally improve its outdated cafeteria?”

This is a leading question. Words like “finally” and “outdated” suggest that the cafeteria needs improvement. This may pressure students to answer “yes,” even if they feel neutral.

Better version

A more neutral question would be: “How satisfied are you with the school cafeteria?”

This wording allows students to respond honestly without being pushed toward an answer.

How to reduce response bias

  • Use neutral wording
  • Avoid emotionally charged words
  • Keep questions clear and simple
  • Allow privacy when answering sensitive questions
  • Make answer choices balanced

3. Measurement bias

Measurement bias happens when the method of measuring or recording data is flawed. The tool, process, or observer may introduce error in a consistent way.

This type of bias is common not only in surveys but also in observational studies and experiments.

Examples of measurement bias

  • A scale that always reads 2 pounds too high
  • A thermometer that is not calibrated correctly
  • A survey question that is confusing, causing people to misunderstand what is being asked
  • Different people recording data in inconsistent ways

How to reduce measurement bias

  • Use accurate and tested measuring tools
  • Write questions clearly
  • Train data collectors to follow the same process
  • Check instruments and procedures before collecting data

Good survey questions

Strong survey questions are important because poor wording can create bias. A good survey question should be:

  • Clear: easy to understand
  • Neutral: does not push toward a certain answer
  • Specific: focused on one idea
  • Answerable: something the person can reasonably know

Avoid these common question problems

  • Leading questions: “How much do you enjoy our excellent school events?”
  • Double-barreled questions: “Do you like the teachers and the lunch?” This asks two things at once.
  • Vague questions: “Do you study often?” The word “often” means different things to different people.
  • Confusing answer choices: choices that overlap or leave out reasonable options

Representative samples

A survey is most useful when the sample is representative of the population. That means the sample reflects the variety of people in the full group.

For example, if a school has students from grades 9 through 12, a survey about school lunches should not question only seniors. A better sample would include students from all grades.

Random sampling

One of the best ways to reduce selection bias is random sampling. In a random sample, each member of the population has an equal or fair chance of being chosen.

Random sampling does not guarantee perfect results, but it helps the sample better represent the population.

Voluntary response and convenience samples

Two common sampling methods often create bias:

  • Voluntary response sample: people choose whether to participate
  • Convenience sample: people are selected because they are easy to reach

These methods are easy to use, but they often produce biased results.

For instance, an online poll on a news website about homework time is a voluntary response sample. People with strong opinions may be more likely to answer, so the results may not represent all students.

Nonresponse and missing opinions

Sometimes bias occurs because certain people do not respond at all. If the nonresponders differ from the responders, the results may be misleading.

For example, if a school emails a survey about after-school activities, students who rarely check email may be left out. Their opinions will not appear in the data.

This problem is closely related to selection bias because part of the population is underrepresented.

Worked Example 1: Identifying selection bias

A principal wants to know whether students want a later school start time. She surveys 100 students who are arriving late to school.

Question: Is this survey likely to be biased?

Solution:

  1. The population is all students at the school.
  2. The sample includes only students arriving late.
  3. Students who arrive late may be more likely to want a later start time than other students.

Conclusion: Yes, this survey shows selection bias. The sample is not representative of the whole school.

Improvement: Randomly select students from the entire school, not just late arrivals.

Worked Example 2: Identifying response bias

A survey asks, “Do you support the sensible decision to require school uniforms?”

Question: What type of bias might this cause?

Solution:

  1. The phrase “sensible decision” suggests that supporting uniforms is the smart choice.
  2. This wording may influence people to agree.

Conclusion: This creates response bias because the question is leading.

Improvement: Ask, “Do you support, oppose, or feel neutral about requiring school uniforms?”

Worked Example 3: Identifying measurement bias

A PE class records students’ sprint times, but one stopwatch consistently starts late by 0.3 seconds.

Question: What kind of bias is this?

Solution:

  1. The measuring tool is producing inaccurate results in a consistent way.
  2. All times recorded with that stopwatch will be off.

Conclusion: This is measurement bias.

Improvement: Check and calibrate the stopwatch or use the same accurate timing method for all students.

Worked Example 4: Critiquing a full survey

A local gym wants to know how often teenagers exercise each week. It posts a survey on its fitness app asking, “How many times do you work out each week to stay healthy and fit?”

Question: What possible biases are present?

Solution:

  1. The survey is posted on a fitness app, so the people who see it are likely already interested in exercise.
  2. This creates selection bias because the sample may not represent all teenagers.
  3. The phrase “to stay healthy and fit” may make people want to report more exercise than they actually do.
  4. This creates response bias.

Conclusion: The survey may have both selection bias and response bias.

Improvement: Survey a random sample of teenagers from different places, and ask a neutral question such as, “How many times do you exercise in a typical week?”

How to critique a survey step by step

When you are asked to evaluate a survey, use these questions:

  1. Who is the population?
  2. How was the sample chosen?
  3. Does the sample represent the population?
  4. Are the questions clear and neutral?
  5. Could people feel pressured to answer in a certain way?
  6. Was the data measured or recorded accurately?

This process helps you identify where bias may appear.

Comparing the three main biases

  • Selection bias: the wrong people are chosen
  • Response bias: people’s answers are influenced or not truthful
  • Measurement bias: the method of measuring or recording is flawed

A helpful way to remember them is:

  • Selection = who is in the sample
  • Response = how people answer
  • Measurement = how data is collected or measured

Why this matters in real life

Survey results are used in schools, businesses, science, health studies, and government decisions. If a survey is biased, people may make poor decisions based on inaccurate data.

For example, a company may launch a product based on a biased survey, or a school may change a policy based on results that do not reflect all students. Understanding bias helps you become a smarter reader of data and a better creator of surveys.

Brief summary

A survey is used to learn about a population by questioning a sample. For a survey to be useful, the sample must be representative, the questions must be neutral, and the data must be measured accurately.

The three main biases in this lesson are selection bias, response bias, and measurement bias. By learning to recognize these problems, you can judge whether data and conclusions are trustworthy.

Put what you read to the test

You've worked through Survey Design and Bias. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Measures of Central Tendency

Measures of Central Tendency are numbers that describe the center or typical value of a dataset. In statistics, the three main measures of central tendency are the mean, median, and mode.

These measures help us summarize data quickly. For example, if a class takes a test, instead of listing every score, we can describe the overall performance using a single central value.

In this lesson, you will learn how to find the mean, median, and mode for raw data and for grouped frequency distributions. You will also learn how to decide which measure is the most appropriate, especially when data is skewed.

1. Why measures of central tendency matter

Real-world data often contains many values. Measures of central tendency help answer questions like:

  • What is the average score?
  • What value lies in the middle?
  • Which value occurs most often?

Each measure gives a different idea of what the “center” means. Sometimes all three are similar, but sometimes they are very different.

2. Mean

The mean is what most people call the average. To find the mean, add all the data values and divide by the number of values.

The formula for the mean of raw data is:

$$ \text{Mean} = \frac{\text{sum of all values}}{\text{number of values}} = \frac{\sum x}{n} $$

Here, \(\sum x\) means “add all the data values,” and \(n\) is the total number of data points.

3. Median

The median is the middle value when the data is arranged in order from smallest to largest.

To find the median:

  1. Arrange the data in order.
  2. If there is an odd number of values, the median is the middle one.
  3. If there is an even number of values, the median is the average of the two middle values.

The median is useful because it is not strongly affected by very large or very small values.

4. Mode

The mode is the value that appears most often.

A dataset may have:

  • One mode if one value occurs most frequently.
  • More than one mode if several values tie for highest frequency.
  • No mode if all values occur the same number of times.

The mode is especially useful when data is categorical or when we want to know the most common value.

5. Worked Example 1: Finding mean, median, and mode from raw data

Consider the dataset:

\(4, 7, 7, 9, 10\)

Mean:

$$ \text{Mean} = \frac{4+7+7+9+10}{5} = \frac{37}{5} = 7.4 $$

Median:

The data is already in order: \(4, 7, 7, 9, 10\). There are 5 values, so the middle value is the 3rd value.

Median \(= 7\)

Mode:

The value \(7\) appears twice, more than any other value.

Mode \(= 7\)

So for this dataset:

  • Mean \(= 7.4\)
  • Median \(= 7\)
  • Mode \(= 7\)

6. Worked Example 2: Raw data with an even number of values

Find the mean, median, and mode for:

\(3, 5, 8, 8, 12, 14\)

Mean:

$$ \text{Mean} = \frac{3+5+8+8+12+14}{6} = \frac{50}{6} \approx 8.33 $$

Median:

The data is already in order. There are 6 values, so the median is the average of the 3rd and 4th values.

$$ \text{Median} = \frac{8+8}{2} = 8 $$

Mode:

The value \(8\) appears twice, while all other values appear once.

Mode \(= 8\)

So the results are:

  • Mean \(\approx 8.33\)
  • Median \(= 8\)
  • Mode \(= 8\)

7. Mean from a frequency table

Sometimes data is given in a frequency table. Instead of listing every value, the table tells us how many times each value appears.

To find the mean from a frequency table, use:

$$ \text{Mean} = \frac{\sum fx}{\sum f} $$

Here:

  • \(x\) is the data value,
  • \(f\) is the frequency,
  • \(fx\) means value \(\times\) frequency.

Worked Example 3: Mean, median, and mode from a frequency table

The number of books read by students in a month is shown below:

Values: \(1, 2, 3, 4, 5\)

Frequencies: \(2, 3, 4, 2, 1\)

First, find the mean.

$$ \sum f = 2+3+4+2+1 = 12 $$ $$ \sum fx = (1)(2) + (2)(3) + (3)(4) + (4)(2) + (5)(1) $$ $$ \sum fx = 2 + 6 + 12 + 8 + 5 = 33 $$ $$ \text{Mean} = \frac{33}{12} = 2.75 $$

Median:

There are 12 values in total, so the median is the average of the 6th and 7th values.

Let us list positions:

  • Value 1 appears in positions 1–2
  • Value 2 appears in positions 3–5
  • Value 3 appears in positions 6–9
  • Value 4 appears in positions 10–11
  • Value 5 appears in position 12

The 6th and 7th values are both \(3\), so:

$$ \text{Median} = 3 $$

Mode:

The highest frequency is 4, which belongs to value \(3\).

Mode \(= 3\)

So:

  • Mean \(= 2.75\)
  • Median \(= 3\)
  • Mode \(= 3\)

8. Grouped frequency distributions

When datasets are large, values are often grouped into class intervals, such as \(0\text{–}9\), \(10\text{–}19\), \(20\text{–}29\), and so on. This is called a grouped frequency distribution.

For grouped data, we usually estimate the mean using the midpoint of each class.

The midpoint of a class interval is:

$$ \text{Midpoint} = \frac{\text{lower class limit} + \text{upper class limit}}{2} $$

Then we use:

$$ \text{Estimated mean} = \frac{\sum fm}{\sum f} $$

where \(m\) is the midpoint of each class.

Worked Example 4: Estimated mean from grouped data

The test scores of students are grouped as follows:

  • \(0\text{–}9\): frequency 2
  • \(10\text{–}19\): frequency 5
  • \(20\text{–}29\): frequency 8
  • \(30\text{–}39\): frequency 3

Step 1: Find the midpoint of each class.

  • \(0\text{–}9\): midpoint \(= \frac{0+9}{2} = 4.5\)
  • \(10\text{–}19\): midpoint \(= \frac{10+19}{2} = 14.5\)
  • \(20\text{–}29\): midpoint \(= \frac{20+29}{2} = 24.5\)
  • \(30\text{–}39\): midpoint \(= \frac{30+39}{2} = 34.5\)

Step 2: Multiply each midpoint by its frequency.

$$ (4.5)(2) = 9 $$ $$ (14.5)(5) = 72.5 $$ $$ (24.5)(8) = 196 $$ $$ (34.5)(3) = 103.5 $$

Step 3: Add the frequencies and the products.

$$ \sum f = 2+5+8+3 = 18 $$ $$ \sum fm = 9+72.5+196+103.5 = 381 $$

Step 4: Find the estimated mean.

$$ \text{Estimated mean} = \frac{381}{18} \approx 21.17 $$

So the estimated mean score is about \(21.17\).

9. Median and mode for grouped data

For grouped data, the exact median and exact mode cannot usually be found because the exact individual values are not known. However, we can still identify:

  • the median class, which contains the middle value,
  • the modal class, which has the greatest frequency.

Using the data from the previous example:

  • Frequencies are \(2, 5, 8, 3\), so total \(n=18\).
  • The middle positions are the 9th and 10th values.

Cumulative frequencies are:

  • \(0\text{–}9\): 2
  • \(10\text{–}19\): 7
  • \(20\text{–}29\): 15
  • \(30\text{–}39\): 18

The 9th and 10th values fall in the class \(20\text{–}29\), so the median class is \(20\text{–}29\).

The highest frequency is 8, which also belongs to \(20\text{–}29\), so the modal class is \(20\text{–}29\).

10. Choosing the best measure

It is important to choose the measure that best represents the data.

Use the mean when:

  • the data is fairly balanced,
  • there are no extreme values that strongly affect the average.

Use the median when:

  • the data is skewed,
  • there are outliers,
  • you want the middle value rather than the average.

Use the mode when:

  • you want the most common value,
  • the data is categorical,
  • repeated values are important.

11. What is skewed data?

A dataset is skewed when values are not spread evenly. One side of the data stretches farther than the other.

For example, consider incomes in a small group:

\(20, 22, 24, 25, 120\)

Mean:

$$ \text{Mean} = \frac{20+22+24+25+120}{5} = \frac{211}{5} = 42.2 $$

Median:

The middle value is \(24\).

Here, the mean is \(42.2\), but most values are in the 20s. The large value \(120\) pulls the mean upward. So the mean does not represent the typical income very well.

In this case, the median is a better measure of center because it is less affected by the outlier.

12. Comparing mean, median, and mode

Here is a quick comparison:

  • Mean: uses every value; sensitive to outliers.
  • Median: middle value; resistant to outliers.
  • Mode: most frequent value; useful for most common category or score.

13. Common mistakes to avoid

  • Do not find the median without ordering the data first.
  • Do not confuse mean with median. The mean uses all values, while the median uses the middle position.
  • Do not forget frequencies. In a frequency table, repeated values must be counted properly.
  • For grouped data, remember that the mean is an estimate. It is based on class midpoints.
  • Do not automatically choose the mean. If the data is skewed, the median may be better.

14. Step-by-step strategy

When solving problems about central tendency, use this approach:

  1. Look at how the data is presented: raw list, frequency table, or grouped table.
  2. If needed, arrange raw data in order.
  3. Find the mean using the correct formula.
  4. Find the median by locating the middle position.
  5. Find the mode by identifying the highest frequency.
  6. Check whether the data has outliers or is skewed.
  7. Choose the measure that best describes the center.

15. Brief summary

The mean, median, and mode are the main measures of central tendency. The mean is the average, the median is the middle value, and the mode is the most frequent value.

For raw data, you can compute each measure directly. For frequency tables, use frequencies carefully, and for grouped data, use class midpoints to estimate the mean and identify the median class and modal class.

Most importantly, choose the measure that matches the shape of the data. If the data is skewed or has outliers, the median is often the best measure of center.

Put what you read to the test

You've worked through Measures of Central Tendency. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Measures of Dispersion

Measures of Dispersion describe how spread out a set of data is. Two data sets can have the same average but look very different if one is tightly grouped and the other is widely spread out.

In descriptive statistics, measures of center such as the mean and median tell us where the data is located, while measures of dispersion tell us how much the values vary. In this lesson, you will learn four important measures of spread:

  • Range
  • Interquartile Range (IQR)
  • Variance
  • Standard Deviation

Understanding these measures helps you compare data sets and decide whether values are clustered together or spread far apart.

1. Range

The range is the simplest measure of dispersion. It tells you the distance between the largest and smallest values in a data set.

$$\text{Range} = \text{Maximum} - \text{Minimum}$$

The range is easy to calculate, but it uses only two values. This means it can be strongly affected by an unusually large or small value.

Example: Find the range of the data set \(3, 7, 8, 10, 14\).

The maximum is \(14\) and the minimum is \(3\).

$$\text{Range} = 14 - 3 = 11$$

So, the range is 11.

2. Interquartile Range (IQR)

The interquartile range measures the spread of the middle 50% of the data. It is less affected by extreme values than the range.

To find the IQR, first find:

  • \(Q_1\): the first quartile, or lower quartile
  • \(Q_3\): the third quartile, or upper quartile

Then use:

$$\text{IQR} = Q_3 - Q_1$$

How to find quartiles:

  1. Put the data in order from least to greatest.
  2. Find the median.
  3. Find the median of the lower half; this is \(Q_1\).
  4. Find the median of the upper half; this is \(Q_3\).

Worked Example 1: Find the IQR of \(2, 4, 5, 7, 8, 10, 12, 15, 18\).

Step 1: The data is already in order.

Step 2: Find the median. Since there are 9 values, the middle value is the 5th value.

$$\text{Median} = 8$$

Step 3: Find the lower half and upper half. Do not include the median when splitting the data.

Lower half: \(2, 4, 5, 7\)

Upper half: \(10, 12, 15, 18\)

Step 4: Find \(Q_1\) and \(Q_3\).

$$Q_1 = \frac{4+5}{2} = 4.5$$

$$Q_3 = \frac{12+15}{2} = 13.5$$

Step 5: Find the IQR.

$$\text{IQR} = 13.5 - 4.5 = 9$$

So, the interquartile range is 9.

Why IQR is useful: It focuses on the center of the data and ignores extreme values. That makes it helpful when the data contains outliers.

3. Variance

The variance measures how far the data values are, on average, from the mean. It uses every value in the data set.

To calculate variance, follow these steps:

  1. Find the mean.
  2. Subtract the mean from each data value to get each deviation.
  3. Square each deviation.
  4. Find the average of those squared deviations.

For a full data set, the population variance is written as:

$$\sigma^2 = \frac{\sum (x-\mu)^2}{n}$$

where:

  • \(x\) represents each data value
  • \(\mu\) is the mean
  • \(n\) is the number of data values

At this level, you will often calculate variance directly from a small data set by using the steps rather than memorizing the symbols.

Worked Example 2: Find the variance of the data set \(2, 4, 6, 8\).

Step 1: Find the mean.

$$\mu = \frac{2+4+6+8}{4} = \frac{20}{4} = 5$$

Step 2: Find each deviation from the mean.

  • \(2-5=-3\)
  • \(4-5=-1\)
  • \(6-5=1\)
  • \(8-5=3\)

Step 3: Square each deviation.

  • \((-3)^2=9\)
  • \((-1)^2=1\)
  • \(1^2=1\)
  • \(3^2=9\)

Step 4: Add the squared deviations.

$$9+1+1+9=20$$

Step 5: Divide by the number of values, \(n=4\).

$$\sigma^2 = \frac{20}{4} = 5$$

So, the variance is 5.

Important note: Variance is measured in squared units. For example, if the data is measured in meters, the variance is measured in square meters. Because of this, variance is useful mathematically, but it is not always easy to interpret directly.

4. Standard Deviation

The standard deviation is the square root of the variance. It tells us the typical distance of the data values from the mean, in the same units as the original data.

$$\sigma = \sqrt{\text{variance}}$$

Using the population variance formula, standard deviation is:

$$\sigma = \sqrt{\frac{\sum (x-\mu)^2}{n}}$$

Because standard deviation is in the same units as the data, it is often easier to understand than variance.

Worked Example 3: Find the standard deviation of the data set \(2, 4, 6, 8\).

From the previous example, the variance is \(5\).

So the standard deviation is:

$$\sigma = \sqrt{5} \approx 2.24$$

So, the standard deviation is approximately 2.24.

This means the data values are typically about \(2.24\) units away from the mean.

Comparing Small and Large Spread

Look at these two data sets:

Set A: \(9, 10, 10, 11, 10\)

Set B: \(2, 6, 10, 14, 18\)

Both sets have mean \(10\), but Set A is tightly grouped around 10, while Set B is much more spread out. Measures of dispersion help us describe that difference clearly.

Set A will have a small range, small IQR, and small standard deviation. Set B will have larger values for those measures.

Worked Example 4: Compare the spread of two data sets using range and standard deviation.

Data Set A: \(5, 6, 7, 8, 9\)

Data Set B: \(1, 4, 7, 10, 13\)

Step 1: Find the range of each set.

For Set A:

$$\text{Range} = 9 - 5 = 4$$

For Set B:

$$\text{Range} = 13 - 1 = 12$$

So Set B has a larger range.

Step 2: Find the standard deviation of each set.

First, both means are:

$$\mu = 7$$

For Set A, the deviations are:

\(-2, -1, 0, 1, 2\)

The squared deviations are:

\(4, 1, 0, 1, 4\)

Sum:

$$4+1+0+1+4=10$$

Variance:

$$\sigma^2 = \frac{10}{5} = 2$$

Standard deviation:

$$\sigma = \sqrt{2} \approx 1.41$$

For Set B, the deviations are:

\(-6, -3, 0, 3, 6\)

The squared deviations are:

\(36, 9, 0, 9, 36\)

Sum:

$$36+9+0+9+36=90$$

Variance:

$$\sigma^2 = \frac{90}{5} = 18$$

Standard deviation:

$$\sigma = \sqrt{18} \approx 4.24$$

Conclusion: Set B has a much larger standard deviation, so its values are much more spread out from the mean than Set A.

When to Use Each Measure

  • Range: Quick measure of overall spread, but affected a lot by extreme values.
  • IQR: Good for describing the spread of the middle half of the data, especially when outliers exist.
  • Variance: Uses all data values and is useful for calculations, but is in squared units.
  • Standard deviation: Uses all data values and is easier to interpret because it is in the original units.

Common Mistakes to Avoid

  • Forgetting to put the data in order before finding the median or quartiles.
  • Using the wrong maximum or minimum when finding range.
  • For variance and standard deviation, forgetting to square the deviations.
  • Taking the square root too early before finishing the variance calculation.
  • Confusing variance and standard deviation. Remember: standard deviation is the square root of variance.

Key Ideas to Remember

  • Measures of dispersion describe how spread out data is.
  • A small measure of dispersion means the data values are close together.
  • A large measure of dispersion means the data values are spread farther apart.
  • Range and IQR are based on positions in the data.
  • Variance and standard deviation are based on distance from the mean.

Brief Summary

The range is the difference between the largest and smallest values. The IQR is the difference between the upper and lower quartiles and shows the spread of the middle 50% of the data.

The variance is the average of the squared distances from the mean, and the standard deviation is the square root of the variance. Together, these measures help us understand and compare how much data sets vary.

Put what you read to the test

You've worked through Measures of Dispersion. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Data Visualization

Data Visualization is the part of statistics where we turn raw numbers into graphs or organized displays so patterns are easier to see. In 11th Grade maths, three important visual tools are histograms, stem-and-leaf plots, and box-and-whisker plots. These help us understand how data is distributed, whether values are clustered together or spread out, and whether the data is symmetric or skewed.

When we visualize data, we are not just drawing pictures. We are looking for meaning. A good graph can help answer questions like:

  • Where do most of the data values fall?
  • Is the data spread out or tightly grouped?
  • Are there any unusually high or low values?
  • Is the distribution roughly symmetric, or is it skewed to one side?

Before using any graph, it is helpful to sort the data from least to greatest. Ordered data makes it easier to spot patterns and calculate important values like the median and quartiles.

1. Histograms

A histogram is a graph that shows how many data values fall within certain intervals, called bins or classes. It looks similar to a bar graph, but the bars in a histogram touch because the data is numerical and continuous over intervals.

To make a histogram:

  1. Choose equal-width intervals for the data.
  2. Count how many values fall into each interval.
  3. Draw bars whose heights match the frequencies.

For example, if test scores are grouped into intervals like 50–59, 60–69, 70–79, and so on, the histogram shows how many students scored in each range.

A histogram is useful because it quickly shows the shape of the distribution. Some common shapes are:

  • Symmetric: the left and right sides are roughly balanced.
  • Skewed right: most data is on the lower side, with a tail stretching to the right.
  • Skewed left: most data is on the higher side, with a tail stretching to the left.
  • Uniform: bars are about the same height.
  • Clustered: several values are concentrated in one or more areas.

2. Stem-and-Leaf Plots

A stem-and-leaf plot organizes data by place value. Each number is split into a stem and a leaf. Usually, the stem is all but the last digit, and the leaf is the last digit.

For example, the number 47 can be split as:

Stem: 4, Leaf: 7

A stem-and-leaf plot keeps the original data values while also showing the overall shape of the distribution. That makes it very useful for small to medium-sized datasets.

To create a stem-and-leaf plot:

  1. Write the stems in a vertical column.
  2. Place each leaf next to its correct stem.
  3. Arrange the leaves in order from least to greatest.

For instance, if the data values are 32, 35, 37, 41, 44, and 49, the plot would be:

3 | 2 5 7
4 | 1 4 9

This means the data values are 32, 35, 37, 41, 44, and 49.

3. Box-and-Whisker Plots

A box-and-whisker plot, also called a box plot, summarizes data using five key values. These are called the five-number summary:

  • Minimum
  • First quartile, or \(Q_1\)
  • Median
  • Third quartile, or \(Q_3\)
  • Maximum

The median divides the data into two equal halves. The quartiles divide the data into four equal parts.

In a box plot:

  • The box goes from \(Q_1\) to \(Q_3\).
  • A line inside the box marks the median.
  • The whiskers extend from the box to the minimum and maximum values.

The length of the box shows the interquartile range, or IQR, which measures the spread of the middle 50% of the data.

The formula for interquartile range is:

$$IQR = Q_3 - Q_1$$

Box plots are especially helpful for comparing datasets and for identifying spread, center, and possible skewness.

Understanding Distribution Shape and Skewness

One important goal of data visualization is to understand the distribution of the data. The distribution describes how the values are spread.

Symmetric distribution: If the graph looks similar on both sides of the center, the data is symmetric. In this case, the mean and median are often close together.

Right-skewed distribution: If the graph has a longer tail on the right, the data is right-skewed. This usually means a few larger values pull the distribution to the right.

Left-skewed distribution: If the graph has a longer tail on the left, the data is left-skewed. This usually means a few smaller values pull the distribution to the left.

In a box plot, skewness can often be seen if one whisker is much longer than the other or if the median is not centered in the box.

Worked Example 1: Making a Stem-and-Leaf Plot

The following quiz scores were recorded:

62, 75, 71, 68, 74, 63, 79, 72, 65, 67

Step 1: Order the data.

62, 63, 65, 67, 68, 71, 72, 74, 75, 79

Step 2: Separate stems and leaves.

  • Stem 6: leaves 2, 3, 5, 7, 8
  • Stem 7: leaves 1, 2, 4, 5, 9

Step 3: Write the plot.

6 | 2 3 5 7 8
7 | 1 2 4 5 9

Interpretation: The scores are spread through the 60s and 70s, with no extreme low or high values. The data appears fairly balanced.

Worked Example 2: Creating a Histogram from Frequency Data

A teacher groups project scores into intervals:

  • 50–59: 2 students
  • 60–69: 5 students
  • 70–79: 8 students
  • 80–89: 6 students
  • 90–99: 3 students

To draw the histogram, place the score intervals on the horizontal axis and frequency on the vertical axis. Then draw touching bars with heights 2, 5, 8, 6, and 3.

Interpretation: The tallest bar is 70–79, so most students scored in that range. The bars rise and then fall, so the distribution is roughly mound-shaped. It is fairly close to symmetric, though not perfectly.

Worked Example 3: Finding the Five-Number Summary and Drawing a Box Plot

Consider the data:

4, 6, 7, 8, 10, 12, 13, 15, 18

Step 1: Find the median.

There are 9 values, so the middle value is the 5th value:

Median = \(10\)

Step 2: Find \(Q_1\).

The lower half is 4, 6, 7, 8. The middle of these is:

$$Q_1 = \frac{6+7}{2} = 6.5$$

Step 3: Find \(Q_3\).

The upper half is 12, 13, 15, 18. The middle of these is:

$$Q_3 = \frac{13+15}{2} = 14$$

Step 4: Write the five-number summary.

  • Minimum = 4
  • \(Q_1 = 6.5\)
  • Median = 10
  • \(Q_3 = 14\)
  • Maximum = 18

Step 5: Find the interquartile range.

$$IQR = Q_3 - Q_1 = 14 - 6.5 = 7.5$$

Interpretation: The middle 50% of the data lies between 6.5 and 14. The distances on each side of the median are similar, so the data is approximately symmetric.

Worked Example 4: Describing Skewness from a Box Plot

Suppose a box plot has the following five-number summary:

  • Minimum = 12
  • \(Q_1 = 18\)
  • Median = 20
  • \(Q_3 = 24\)
  • Maximum = 40

Notice that the right whisker goes from 24 to 40, which is much longer than the left whisker from 12 to 18.

Interpretation: This suggests the distribution is right-skewed. Most values are lower, but a few larger values stretch the distribution to the right.

How to Choose the Best Display

Each graph has strengths. The best choice depends on the dataset and what you want to learn from it.

  • Use a histogram when you want to see the overall shape of numerical data, especially for larger datasets.
  • Use a stem-and-leaf plot when you want to see the shape of the data and still keep the exact values.
  • Use a box-and-whisker plot when you want a quick summary of center and spread, or when comparing two or more datasets.

Common Mistakes to Avoid

  • Using unequal interval widths in a histogram without clear labeling.
  • Forgetting to order leaves in a stem-and-leaf plot.
  • Confusing quartiles with the minimum or maximum in a box plot.
  • Describing a graph only by its highest point and ignoring spread or skewness.
  • Assuming every graph with one peak is perfectly symmetric.

What to Look for When Interpreting Any Data Display

  • Center: Where is the middle of the data?
  • Spread: How far apart are the values?
  • Shape: Is it symmetric, skewed, or clustered?
  • Unusual values: Are there gaps or extreme values?

If you build the habit of checking these four features, you will become much better at reading and comparing data displays.

Summary

Data visualization helps us organize and understand numerical information. Histograms show how often values fall in intervals, stem-and-leaf plots show the exact values while revealing the distribution, and box-and-whisker plots summarize the data using the five-number summary. By studying these graphs, we can describe the center, spread, and shape of data, including whether it is symmetric, right-skewed, or left-skewed. These skills help us draw better statistical conclusions from real-world information.

Put what you read to the test

You've worked through Data Visualization. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Outlier Detection

Outlier Detection is an important part of descriptive statistics. When we collect data, most values usually cluster in a main group, but sometimes one or two values are much smaller or much larger than the rest. These unusual values are called outliers.

Outliers matter because they can change how we describe a dataset. They can affect the mean, the spread, and sometimes the conclusions we make. In this lesson, you will learn how to identify outliers using the 1.5 IQR rule and how to describe their effect on statistical summaries.

What is an outlier?

An outlier is a data value that is unusually far away from the rest of the data. It does not automatically mean the value is wrong. Sometimes an outlier is caused by a mistake in recording data, but sometimes it is a real and meaningful result.

For example, in a set of quiz scores mostly between 72 and 91, a score of 15 would likely be an outlier. In a list of daily temperatures around 20 to 28 degrees, a temperature of 22 is not an outlier because it fits the pattern of the data.

Why do we use the 1.5 IQR rule?

One common way to detect outliers is with the interquartile range, or IQR. The IQR measures the spread of the middle half of the data. Since it focuses on the middle 50%, it is less affected by extreme values than the full range.

The 1.5 IQR rule uses quartiles:

  • : the first quartile, or the median of the lower half of the data
  • : the third quartile, or the median of the upper half of the data
  • IQR: the interquartile range, found by subtracting  from 

The formula is

$$IQR = Q_3 - Q_1$$

Then we calculate the lower fence and upper fence:

$$\text{Lower Fence} = Q_1 - 1.5(IQR)$$ $$\text{Upper Fence} = Q_3 + 1.5(IQR)$$

Any value below the lower fence or above the upper fence is considered an outlier.

Steps for detecting outliers

  1. Put the data in order from least to greatest.
  2. Find the median.
  3. Find , the median of the lower half.
  4. Find , the median of the upper half.
  5. Compute the IQR using  - .
  6. Find the lower and upper fences.
  7. Check whether any data values fall outside the fences.

Important note about quartiles

When finding quartiles, first find the median of the entire dataset. Then split the data into a lower half and an upper half. If the number of data points is odd, do not include the overall median in either half when finding  and .

Worked Example 1: Finding an outlier in a small dataset

Suppose the data are:

3, 5, 6, 7, 8, 9, 10, 12, 25

Step 1: Find the median

There are 9 values, so the median is the 5th value:

Median = 8

Step 2: Find  and 

Lower half: 3, 5, 6, 7

Upper half: 9, 10, 12, 25

 is the median of the lower half:

$$Q_1 = \frac{5+6}{2} = 5.5$$

 is the median of the upper half:

$$Q_3 = \frac{10+12}{2} = 11$$

Step 3: Find the IQR

$$IQR = Q_3 - Q_1 = 11 - 5.5 = 5.5$$

Step 4: Find the fences

$$\text{Lower Fence} = 5.5 - 1.5(5.5) = 5.5 - 8.25 = -2.75$$ $$\text{Upper Fence} = 11 + 1.5(5.5) = 11 + 8.25 = 19.25$$

Step 5: Identify outliers

Any value less than .75 or greater than 19.25 is an outlier. The value 25 is greater than 19.25, so 25 is an outlier.

Worked Example 2: A dataset with no outliers

Consider the data:

11, 12, 13, 15, 16, 18, 19, 20

Step 1: Find the median

There are 8 values, so the median is the average of the 4th and 5th values:

$$\text{Median} = \frac{15+16}{2} = 15.5$$

Step 2: Find  and 

Lower half: 11, 12, 13, 15

Upper half: 16, 18, 19, 20

$$Q_1 = \frac{12+13}{2} = 12.5$$ $$Q_3 = \frac{18+19}{2} = 18.5$$

Step 3: Find the IQR

$$IQR = 18.5 - 12.5 = 6$$

Step 4: Find the fences

$$\text{Lower Fence} = 12.5 - 1.5(6) = 12.5 - 9 = 3.5$$ $$\text{Upper Fence} = 18.5 + 1.5(6) = 18.5 + 9 = 27.5$$

Step 5: Check for outliers

All values are between 3.5 and 27.5, so there are no outliers.

Worked Example 3: Comparing the effect of an outlier on the mean and median

Look at these two datasets:

Dataset A: 4, 5, 6, 6, 7

Dataset B: 4, 5, 6, 6, 30

Both datasets have 5 values.

Dataset A

$$\text{Mean} = \frac{4+5+6+6+7}{5} = \frac{28}{5} = 5.6$$

Median = 6

Dataset B

$$\text{Mean} = \frac{4+5+6+6+30}{5} = \frac{51}{5} = 10.2$$

Median = 6

Notice what happened:

  • The mean changed a lot, from 5.6 to 10.2.
  • The median stayed the same at 6.

This shows that outliers strongly affect the mean, but the median is more resistant. That is why the median and IQR are often better summaries when outliers are present.

Worked Example 4: Full outlier analysis

A student records the number of minutes classmates spend studying on a certain night:

20, 25, 30, 35, 35, 40, 45, 50, 120

Step 1: Find the median

There are 9 values, so the median is the 5th value:

Median = 35

Step 2: Find  and 

Lower half: 20, 25, 30, 35

Upper half: 40, 45, 50, 120

$$Q_1 = \frac{25+30}{2} = 27.5$$ $$Q_3 = \frac{45+50}{2} = 47.5$$

Step 3: Find the IQR

$$IQR = 47.5 - 27.5 = 20$$

Step 4: Find the fences

$$\text{Lower Fence} = 27.5 - 1.5(20) = 27.5 - 30 = -2.5$$ $$\text{Upper Fence} = 47.5 + 1.5(20) = 47.5 + 30 = 77.5$$

Step 5: Identify outliers

The value 120 is greater than 77.5, so 120 is an outlier.

Step 6: Describe the effect on the data

Let us compare the mean and median.

$$\text{Mean with outlier} = \frac{20+25+30+35+35+40+45+50+120}{9} = \frac{400}{9} \approx 44.4$$

The median is 35.

The mean is much higher than the median because the outlier 120 pulls the mean upward. If we wanted a typical study time, the median of 35 minutes may describe the class better than the mean of about 44.4 minutes.

How outliers appear on box plots

A box plot is a graph that shows the median, quartiles, and possible outliers. The box goes from  to , and the line inside the box shows the median. The “whiskers” extend to the smallest and largest non-outlier values. Outliers are often plotted as separate points beyond the whiskers.

This makes box plots very useful for quickly spotting unusual values.

How outliers affect statistical summaries

  • Mean: usually affected a lot by outliers
  • Median: usually affected very little
  • Range: affected a lot because it uses the smallest and largest values
  • IQR: affected much less because it uses the middle 50% of the data

Because of this, when a dataset has outliers, the median and IQR are often better choices than the mean and range for describing the center and spread.

Be careful: an outlier is not always a mistake

After finding an outlier, you should think about what it means. Ask questions such as:

  • Was the data entered incorrectly?
  • Was the value measured in the wrong units?
  • Is the value unusual but still real?

For example, if one student studied for 120 minutes, that may be unusual, but it could still be true. In that case, the outlier should not simply be erased. Instead, it should be reported and interpreted carefully.

Common mistakes to avoid

  • Not ordering the data before finding quartiles
  • Using the median as part of both halves when the number of data values is odd
  • Forgetting that the IQR is calculated by  - 
  • Using 1.5 incorrectly in the fence formulas
  • Assuming every extreme value is automatically an error

Quick checklist for solving outlier problems

  1. Sort the data.
  2. Find the median, , and .
  3. Compute the IQR.
  4. Find the lower and upper fences.
  5. Identify any values outside the fences.
  6. Describe how those outliers affect the mean, median, range, or IQR.

Summary

An outlier is a value that is much smaller or larger than the rest of the data. To detect outliers, we use the 1.5 IQR rule: first find  and , then compute

$$IQR = Q_3 - Q_1$$

Next, calculate the fences:

$$Q_1 - 1.5(IQR) \quad \text{and} \quad Q_3 + 1.5(IQR)$$

Any value outside those fences is an outlier. Outliers can strongly affect the mean and range, but they usually have less effect on the median and IQR. That is why median and IQR are often the better summaries when outliers are present.

Put what you read to the test

You've worked through Outlier Detection. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Bivariate Data and Scatter Plots

Lesson: Bivariate Data and Scatter Plots

In statistics, we often want to study how two variables are related. For example, we might ask whether study time is connected to test scores, whether height is related to arm span, or whether temperature affects ice cream sales.

When data values come in pairs, we call the data bivariate data. The word bi means two, so bivariate data involves two measured quantities for each individual, object, or situation.

A very useful way to display bivariate data is with a scatter plot. A scatter plot helps us see whether the two variables seem related, and if they are, what kind of relationship they have.

1. What is bivariate data?

Bivariate data consists of ordered pairs, usually written as \((x, y)\).

  • \(x\) represents one variable, usually placed on the horizontal axis.
  • \(y\) represents the second variable, usually placed on the vertical axis.

For example, if a teacher records the number of hours each student studied and that student's test score, the data might look like this:

  • Student 1: \((2, 68)\)
  • Student 2: \((4, 75)\)
  • Student 3: \((5, 82)\)

Each ordered pair describes one student. The first number is the study time, and the second number is the test score.

2. What is a scatter plot?

A scatter plot is a graph made by plotting each ordered pair from a bivariate data set as a point on a coordinate plane.

To make a scatter plot:

  1. Identify the two variables.
  2. Choose which variable goes on the \(x\)-axis and which goes on the \(y\)-axis.
  3. Label both axes clearly with units if needed.
  4. Choose a reasonable scale for each axis.
  5. Plot each ordered pair as a point.

Unlike a line graph, you do not connect the points in a scatter plot. Each point represents one pair of values.

3. Reading a scatter plot

After graphing the data, look for a pattern. A scatter plot can suggest whether the variables are related.

There are three main types of association you should recognize:

  • Positive association: As \(x\) increases, \(y\) tends to increase.
  • Negative association: As \(x\) increases, \(y\) tends to decrease.
  • No clear association: The points do not show an obvious upward or downward trend.

For example:

  • If students who study more usually score higher, the scatter plot shows a positive association.
  • If the outside temperature rises while the number of hot chocolate sales falls, the scatter plot shows a negative association.
  • If shoe size and math grade do not seem connected, the scatter plot may show no clear association.

4. Linear and non-linear relationships

One important goal of a scatter plot is to decide whether the relationship looks linear or non-linear.

A relationship is linear if the points follow a pattern close to a straight line. This does not mean every point lies exactly on one line, but the overall trend looks straight.

A relationship is non-linear if the pattern curves or changes direction, so a straight line would not describe it well.

Here is how to tell the difference:

  • Linear: points cluster around a straight upward or downward trend
  • Non-linear: points form a curve, arc, or other shape that is not straight

For example, if distance traveled increases at a steady rate over time, the data may look linear. But if a ball is thrown into the air, its height over time may rise and then fall, creating a non-linear pattern.

5. Strength of association

Scatter plots also show how strong a relationship is.

  • Strong association: points lie close to a line or curve
  • Weak association: points are more spread out

So a scatter plot can be described in more than one way. For example, you might say:

  • strong positive linear association
  • weak negative linear association
  • strong non-linear association
  • no clear association

6. Outliers

An outlier is a data point that lies far away from the overall pattern of the rest of the data.

Outliers are important because they can affect how we interpret the relationship. An outlier might happen because of:

  • a recording mistake
  • an unusual event
  • a value that is real but uncommon

When looking at a scatter plot, always notice whether one or two points are much farther away than the others.

7. Choosing the variables for each axis

Usually, the variable that might help explain or predict the other one is placed on the \(x\)-axis. The response or outcome variable is usually placed on the \(y\)-axis.

For instance:

  • hours studied \(\rightarrow\) \(x\)-axis
  • test score \(\rightarrow\) \(y\)-axis

This makes sense because study time may help explain changes in test score.

Worked Example 1: Making a scatter plot

A coach records the number of practice hours and the number of free throws made by 6 players in one minute.

  • \((1, 5)\)
  • \((2, 6)\)
  • \((3, 8)\)
  • \((4, 9)\)
  • \((5, 11)\)
  • \((6, 12)\)

Step 1: Identify the variables.

  • \(x =\) practice hours
  • \(y =\) free throws made

Step 2: Plot the points on a coordinate plane.

Step 3: Describe the pattern.

As practice hours increase, the number of free throws made also increases. The points would lie close to an upward-sloping straight line.

Conclusion: The scatter plot shows a strong positive linear association.

Worked Example 2: Negative association

A store manager compares outside temperature and the number of winter hats sold in one day.

  • \((30, 25)\)
  • \((35, 22)\)
  • \((40, 18)\)
  • \((45, 14)\)
  • \((50, 10)\)
  • \((55, 7)\)

Here, \(x\) is temperature and \(y\) is hats sold.

As temperature increases, hat sales decrease. The points follow a downward trend that is close to a straight line.

Conclusion: The scatter plot shows a strong negative linear association.

Worked Example 3: Non-linear association

A science class measures the height of a small rocket at different times after launch.

  • \((0, 0)\)
  • \((1, 12)\)
  • \((2, 20)\)
  • \((3, 24)\)
  • \((4, 20)\)
  • \((5, 12)\)
  • \((6, 0)\)

The rocket rises, reaches a highest point, and then falls back down. If these points are plotted, they do not form a straight line. Instead, they form a curved shape.

Conclusion: The scatter plot shows a strong non-linear association.

This example is important because not every relationship is linear. A scatter plot helps you see that quickly.

Worked Example 4: Spotting an outlier

A class records hours studied and quiz scores:

  • \((1, 62)\)
  • \((2, 68)\)
  • \((3, 74)\)
  • \((4, 79)\)
  • \((5, 85)\)
  • \((6, 50)\)

The first five points suggest that more study time is linked to higher quiz scores. But the point \((6, 50)\) is much lower than expected.

This point does not fit the overall upward pattern, so it is an outlier.

Conclusion: The data mostly shows a positive linear association, but there is one outlier.

8. What scatter plots can and cannot tell us

A scatter plot can help us:

  • see whether two variables seem related
  • tell whether the relationship is positive, negative, or unclear
  • decide whether the pattern is linear or non-linear
  • notice outliers

However, a scatter plot does not prove that one variable causes the other. For example, even if two variables are related, that does not always mean one directly causes the change in the other.

9. Common mistakes to avoid

  • Mixing up the variables: Be sure the correct variable is on each axis.
  • Using an uneven scale: Choose a clear, consistent scale.
  • Connecting the points: Scatter plot points should usually stay separate.
  • Forcing a linear interpretation: If the data curves, describe it as non-linear.
  • Ignoring outliers: A strange point may matter.

10. Quick checklist for describing a scatter plot

When you are asked to interpret a scatter plot, use this checklist:

  1. Is the association positive, negative, or none?
  2. Is the pattern linear or non-linear?
  3. Is the association strong or weak?
  4. Are there any outliers?

A good complete description might be: "The scatter plot shows a strong positive linear association with no obvious outliers."

Brief Summary

Bivariate data involves two variables measured together as ordered pairs. A scatter plot graphs these pairs so that we can study the relationship between the variables. By looking at the direction, shape, strength, and any outliers, we can describe whether the data shows a positive or negative association, whether it is linear or non-linear, and how closely the points follow a pattern.

Put what you read to the test

You've worked through Bivariate Data and Scatter Plots. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Correlation and Regression

Correlation and Regression help us study the relationship between two numerical variables. For example, we might want to know whether study time is related to test scores, or whether hours of exercise are related to resting heart rate.

In this lesson, you will learn how to:

  • interpret the Pearson correlation coefficient;
  • find and understand the least squares regression line;
  • use a regression model to make predictions by interpolation and extrapolation;
  • decide when a linear model makes sense and when to be careful.

1. Looking for a relationship between two variables

When we compare two numerical variables, we often place the data on a scatter plot. Each point represents one pair of values \, \((x,y)\).

For example, if \(x\) is hours studied and \(y\) is test score, then a point \((4,82)\) means a student studied 4 hours and scored 82.

A scatter plot helps us see whether the variables are related.

  • Positive association: as \(x\) increases, \(y\) tends to increase.
  • Negative association: as \(x\) increases, \(y\) tends to decrease.
  • No clear association: the points do not follow a clear trend.

If the points lie roughly around a straight line, then a linear model may be appropriate.

2. Pearson correlation coefficient

The Pearson correlation coefficient, written as \(r\), measures the strength and direction of a linear relationship between two variables.

The value of \(r\) is always between \(-1\) and \(1\):

$$-1 \le r \le 1$$
  • If \(r>0\), the relationship is positive.
  • If \(r<0\), the relationship is negative.
  • If \(r\) is close to \(1\) or \(-1\), the linear relationship is strong.
  • If \(r\) is close to \(0\), the linear relationship is weak.

Common interpretations are:

  • \(r \approx 1\): very strong positive linear correlation
  • \(r \approx -1\): very strong negative linear correlation
  • \(r \approx 0\): little or no linear correlation

Important: correlation does not prove causation. If two variables are correlated, that does not automatically mean one causes the other.

For instance, ice cream sales and swimming activity may both increase in summer. They are related, but one does not directly cause the other. A third factor, temperature, affects both.

3. The least squares regression line

When data show a roughly linear trend, we can model the relationship using a line called the least squares regression line.

This line is often written as:

$$\hat{y} = a + bx$$

where:

  • \(\hat{y}\) is the predicted value of \(y\),
  • \(a\) is the y-intercept,
  • \(b\) is the slope.

The slope tells us how much the predicted value of \(y\) changes when \(x\) increases by 1 unit.

The y-intercept is the predicted value of \(y\) when \(x=0\). Sometimes this is meaningful, and sometimes it is not, depending on the situation.

The term least squares means the line is chosen so that the squared vertical distances from the data points to the line are as small as possible.

4. Formulas for the regression line

To calculate the least squares regression line, we often use these formulas:

$$b = r\left(\frac{s_y}{s_x}\right)$$ $$a = \bar{y} - b\bar{x}$$

where:

  • \(r\) is the correlation coefficient,
  • \(s_x\) is the standard deviation of the \(x\)-values,
  • \(s_y\) is the standard deviation of the \(y\)-values,
  • \(\bar{x}\) is the mean of the \(x\)-values,
  • \(\bar{y}\) is the mean of the \(y\)-values.

In many school problems, a calculator or software gives the regression equation directly. Even if technology does the calculations, you still need to understand what the equation means.

5. Interpreting the regression line

Suppose the regression line is:

$$\hat{y} = 52 + 6x$$

If \(x\) is hours studied and \(y\) is test score, then:

  • the slope \(6\) means each additional hour studied is associated with an increase of about 6 points in the predicted test score;
  • the intercept \(52\) means the predicted score for 0 hours of study is 52.

Remember that this is a prediction, not a guarantee. Real data points will usually not all lie exactly on the line.

6. Residuals

A residual measures how far an actual data value is from the predicted value.

$$\text{residual} = y - \hat{y}$$
  • A positive residual means the actual value is above the regression line.
  • A negative residual means the actual value is below the regression line.
  • A residual of 0 means the point lies exactly on the line.

Residuals help us judge how well the regression line fits the data.

7. Interpolation and extrapolation

Once we have a regression line, we can use it to make predictions.

  • Interpolation means predicting for an \(x\)-value within the range of the observed data.
  • Extrapolation means predicting for an \(x\)-value outside the range of the observed data.

Interpolation is usually safer because it stays within the data we actually observed.

Extrapolation can be risky because the pattern may not continue beyond the available data.

Worked Example 1: Interpreting correlation

A set of data has correlation coefficient \(r=0.84\).

Question: What does this tell us about the relationship between the variables?

Solution:

  • The value is positive, so the relationship is positive.
  • The value is close to 1, so the linear relationship is strong.

Conclusion: There is a strong positive linear correlation between the two variables.

If one variable increases, the other tends to increase as well.

Worked Example 2: Finding the regression line from summary statistics

Suppose for a dataset we know:

  • \(\bar{x}=5\)
  • \(\bar{y}=18\)
  • \(s_x=2\)
  • \(s_y=6\)
  • \(r=0.75\)

Find the least squares regression line.

Step 1: Find the slope.

$$b = r\left(\frac{s_y}{s_x}\right) = 0.75\left(\frac{6}{2}\right)=0.75(3)=2.25$$

Step 2: Find the intercept.

$$a = \bar{y} - b\bar{x} = 18 - 2.25(5) = 18 - 11.25 = 6.75$$

Regression equation:

$$\hat{y} = 6.75 + 2.25x$$

Interpretation: For each increase of 1 unit in \(x\), the predicted value of \(y\) increases by 2.25 units.

Worked Example 3: Using the regression line for prediction

A teacher finds that the relationship between hours studied \((x)\) and test score \((y)\) is modeled by:

$$\hat{y} = 52 + 6x$$

(a) Predict the score of a student who studies 4 hours.

Substitute \(x=4\):

$$\hat{y} = 52 + 6(4) = 52 + 24 = 76$$

So the predicted score is 76.

(b) If a student studied 4 hours and actually scored 81, find the residual.

$$\text{residual} = y - \hat{y} = 81 - 76 = 5$$

The residual is 5.

This means the student scored 5 points above the predicted value.

Worked Example 4: Interpolation vs. extrapolation

A regression model for plant growth was built using data for sunlight hours from 2 to 8 hours per day:

$$\hat{y} = 10 + 1.8x$$

(a) Predict the plant height when \(x=5\) hours of sunlight.

$$\hat{y} = 10 + 1.8(5) = 10 + 9 = 19$$

The predicted height is 19.

Since 5 is within the data range 2 to 8, this is interpolation.

(b) Predict the plant height when \(x=12\) hours of sunlight.

$$\hat{y} = 10 + 1.8(12) = 10 + 21.6 = 31.6$$

The model predicts 31.6.

However, 12 is outside the observed data range, so this is extrapolation.

We should be cautious because the linear pattern may not continue that far.

8. How to judge whether a linear model is appropriate

Before using correlation and regression, check the scatter plot.

A linear model is reasonable when:

  • the points follow a roughly straight-line pattern;
  • there are no extreme outliers that strongly distort the pattern;
  • the association looks reasonably consistent.

A linear model may not be appropriate when:

  • the pattern is curved instead of straight;
  • the data are scattered with no clear trend;
  • one or two unusual points heavily affect the graph.

9. Important cautions

  • Correlation does not imply causation.
  • A strong correlation can still have outliers.
  • A weak correlation does not always mean no relationship; the relationship might be non-linear.
  • Extrapolation can be unreliable.
  • The regression line predicts, but does not give exact values.

10. Step-by-step strategy for problems

  1. Look at the variables and identify which is \(x\) and which is \(y\).
  2. Use the scatter plot, if given, to see whether the relationship is positive, negative, or weak.
  3. Interpret the value of \(r\).
  4. If needed, find the regression line using: $$b = r\left(\frac{s_y}{s_x}\right), \qquad a = \bar{y} - b\bar{x}$$
  5. Write the equation in the form: $$\hat{y} = a + bx$$
  6. Substitute the given \(x\)-value to make a prediction.
  7. If asked, find the residual using: $$y - \hat{y}$$
  8. Decide whether the prediction is interpolation or extrapolation.

11. Quick practice questions

1. If \(r=-0.91\), what does that tell you?

Answer: a strong negative linear correlation.

2. If the regression line is \(\hat{y}=15-2x\), what does the slope mean?

Answer: for each increase of 1 in \(x\), the predicted value of \(y\) decreases by 2.

3. If the actual value is 27 and the predicted value is 24, what is the residual?

Answer: \(27-24=3\).

4. Why is extrapolation risky?

Answer: because it predicts beyond the observed data, where the same pattern may not continue.

Summary

Correlation measures the direction and strength of a linear relationship between two numerical variables. The Pearson correlation coefficient \(r\) ranges from \(-1\) to \(1\).

The least squares regression line, written as \(\hat{y}=a+bx\), gives a linear model for predicting values of \(y\) from values of \(x\). The slope describes the rate of change, and the intercept gives the predicted value when \(x=0\).

You can use the regression line for interpolation and extrapolation, but predictions outside the data range should be treated carefully. Always remember that correlation does not prove causation.

Put what you read to the test

You've worked through Correlation and Regression. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.