Chapter 13

Statistics

Grouped vs. Ungrouped Data Characteristics

Grouped vs. Ungrouped Data Characteristics

In statistics, data can be shown in different ways depending on how much information we want to keep and how large the data set is. Two important forms are ungrouped data and grouped data.

Understanding the difference between these two forms is important because it affects how we read the data, organize it, and calculate values such as averages and frequencies.

This lesson will help you learn how to tell grouped data from ungrouped data, how class intervals work, and how to identify class marks and class boundaries.

1. What is ungrouped data?

Ungrouped data is data listed as individual values. Each observation is shown separately, so we can see the exact data points.

For example, suppose 8 students scored the following marks on a quiz:

12, 15, 17, 15, 19, 14, 18, 15

This is ungrouped data because every score is written on its own, not collected into intervals.

Characteristics of ungrouped data:

  • Each value is shown clearly.
  • It is best for small data sets.
  • Exact values are known.
  • It is easier to find the precise minimum, maximum, and repeated values.

2. What is grouped data?

Grouped data is data that has been organized into class intervals. Instead of listing every individual value, we place values into groups.

For example, test marks for a large class might be shown like this:

  • \(0\text{–}10\): 3 students
  • \(10\text{–}20\): 7 students
  • \(20\text{–}30\): 12 students
  • \(30\text{–}40\): 8 students

This is grouped data because the marks are collected into intervals rather than shown one by one.

Characteristics of grouped data:

  • Data is organized into classes or intervals.
  • It is useful for large data sets.
  • It makes patterns easier to see.
  • Exact individual values are not shown.
  • Some detail is lost when data is grouped.

3. Main difference between grouped and ungrouped data

The biggest difference is that ungrouped data shows exact individual values, while grouped data summarizes values into intervals.

  • Ungrouped data keeps all the original information.
  • Grouped data gives a simpler summary of the data.

So, if you need exact values, ungrouped data is better. If you need to study a large set of values and look for trends, grouped data is more useful.

4. Discrete data points and continuous class intervals

Ungrouped data often appears as raw discrete data points. This means the values are listed one at a time.

For example, number of books read by students in a month could be:

2, 4, 1, 3, 2, 5, 4

These are separate values, so this is discrete and ungrouped.

Grouped data is often shown using continuous class intervals. These intervals cover a range of values.

For example, heights of students might be grouped as:

  • 140–149 cm
  • 150–159 cm
  • 160–169 cm

These are intervals, so this is grouped data.

5. What is a class interval?

A class interval is a range of values used to group data. Each group is called a class.

In the interval \(20\text{–}29\), the lower class limit is \(20\) and the upper class limit is \(29\).

If data is grouped as:

  • \(0\text{–}9\)
  • \(10\text{–}19\)
  • \(20\text{–}29\)

then each class interval has width \(10\).

Class width can be found by looking at the size of each interval. For equal intervals, the width stays the same throughout the table.

6. What is a class mark?

The class mark is the midpoint of a class interval. It is found by averaging the lower and upper class limits.

The formula is:

$$\text{Class mark} = \frac{\text{lower class limit} + \text{upper class limit}}{2}$$

For the class interval \(20\text{–}29\):

$$\text{Class mark} = \frac{20 + 29}{2} = \frac{49}{2} = 24.5$$

The class mark is used as a representative value for all the data in that class.

7. What are class boundaries?

Class boundaries are the true edges of a class interval when data is measured continuously.

They are especially helpful when there is no gap between classes.

For example, consider the classes:

  • \(10\text{–}19\)
  • \(20\text{–}29\)

If the data is measured to the nearest whole number, then the boundaries are found by subtracting \(0.5\) from the lower limit and adding \(0.5\) to the upper limit.

So:

  • \(10\text{–}19\) becomes \(9.5\text{–}19.5\)
  • \(20\text{–}29\) becomes \(19.5\text{–}29.5\)

Notice that the upper boundary of one class touches the lower boundary of the next class. This makes the grouped data continuous.

8. Why class boundaries matter

Class boundaries are important when drawing graphs such as histograms and cumulative frequency curves. These graphs need continuous intervals with no gaps.

If we only use class limits, it may look like there is a break between classes. Boundaries solve this problem.

9. Worked Example 1: Identify grouped and ungrouped data

Decide whether each set of data is grouped or ungrouped.

  1. \(5, 8, 6, 10, 7, 8\)
    • \(0\text{–}4\): 2
    • \(5\text{–}9\): 6
    • \(10\text{–}14\): 3

Solution:

1. \(5, 8, 6, 10, 7, 8\) is ungrouped data because each value is listed separately.

2. The second set is grouped data because the values are organized into class intervals with frequencies.

Worked Example 2: Find class marks

Find the class mark for each interval:

  • \(10\text{–}19\)
  • \(20\text{–}29\)
  • \(30\text{–}39\)

Solution:

Use:

$$\text{Class mark} = \frac{\text{lower limit} + \text{upper limit}}{2}$$

For \(10\text{–}19\):

$$\frac{10+19}{2} = \frac{29}{2} = 14.5$$

For \(20\text{–}29\):

$$\frac{20+29}{2} = \frac{49}{2} = 24.5$$

For \(30\text{–}39\):

$$\frac{30+39}{2} = \frac{69}{2} = 34.5$$

Answer: The class marks are \(14.5\), \(24.5\), and \(34.5\).

Worked Example 3: Find class boundaries

The class intervals are:

  • \(40\text{–}49\)
  • \(50\text{–}59\)
  • \(60\text{–}69\)

Find the class boundaries if the data is measured to the nearest whole number.

Solution:

For whole-number data, subtract \(0.5\) from each lower limit and add \(0.5\) to each upper limit.

  • \(40\text{–}49\) becomes \(39.5\text{–}49.5\)
  • \(50\text{–}59\) becomes \(49.5\text{–}59.5\)
  • \(60\text{–}69\) becomes \(59.5\text{–}69.5\)

Answer: The class boundaries are:

\(39.5\text{–}49.5\), \(49.5\text{–}59.5\), and \(59.5\text{–}69.5\).

Worked Example 4: Describe the data characteristics

A teacher records the ages of students as:

14, 15, 14, 16, 15, 15, 14, 16, 17, 15

Later, the teacher summarizes them in this table:

  • \(14\text{–}14\): 3
  • \(15\text{–}15\): 4
  • \(16\text{–}16\): 2
  • \(17\text{–}17\): 1

Question: Which form is ungrouped and which form is grouped? What is lost when the data is grouped?

Solution:

The first list is ungrouped data because every age is written individually.

The table is grouped data because the ages are organized into classes with frequencies.

When the data is grouped, we lose the original order of the values. We can still see how many students are each age, but we no longer see the exact sequence in which the ages were recorded.

10. Key ideas to remember

  • Ungrouped data shows each value separately.
  • Grouped data summarizes values into class intervals.
  • A class interval is a range used to group data.
  • A class mark is the midpoint of a class interval.
  • Class boundaries show the true edges of each class for continuous data.
  • Grouped data is easier to use for large sets, but it does not show every exact value.

11. Brief summary

Ungrouped data lists raw data values one by one, while grouped data collects them into intervals. Grouped data is useful for large data sets because it makes patterns easier to see, but it loses some exact detail.

When working with grouped data, you should be able to identify the class interval, class mark, and class boundaries. These ideas are important for tables, histograms, and cumulative frequency curves.

Put what you read to the test

You've worked through Grouped vs. Ungrouped Data Characteristics. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Mean of Grouped Data: Direct Method

Mean of Grouped Data: Direct Method

In statistics, the mean tells us the average value of a set of data. When we have a long list of individual values, finding the mean is simple: add all the values and divide by the number of values.

But sometimes data is not given one by one. Instead, it is arranged into class intervals such as 0–10, 10–20, 20–30, and so on. This is called grouped data. In that case, we use a special method to estimate the mean. One important method is the Direct Method.

This lesson will help you understand what grouped data is, how to find class marks, how to use frequencies, and how to calculate the mean step by step using the direct method.

1. What is grouped data?

Grouped data is data that has been collected and organized into groups or intervals. Each interval shows a range of values, and the frequency tells us how many observations fall in that interval.

For example, a table of marks may look like this:

  • 0–10: 3 students
  • 10–20: 5 students
  • 20–30: 7 students

Here, the marks are grouped into intervals, and the number of students in each interval is the frequency.

2. Idea behind the direct method

Since we do not know every exact value inside each class interval, we take the midpoint of each class as the representative value of that class. This midpoint is called the class mark.

The class mark of a class interval is found by:

$$x_i = \frac{\text{lower limit} + \text{upper limit}}{2}$$

Then we multiply each class mark by its frequency. This gives us the total contribution of that class to the average.

Finally, we use the formula:

$$\bar{x} = \frac{\sum f_i x_i}{\sum f_i}$$

where:

  • \(\bar{x}\) = mean
  • \(f_i\) = frequency of a class
  • \(x_i\) = class mark of that class
  • \(\sum f_i x_i\) = sum of all products of frequency and class mark
  • \(\sum f_i\) = total frequency

3. Steps to find the mean of grouped data by direct method

  1. Write the class intervals and frequencies.
  2. Find the class mark \(x_i\) of each interval.
  3. Multiply frequency and class mark to get \(f_i x_i\).
  4. Find \(\sum f_i\) and \(\sum f_i x_i\).
  5. Use the formula: $$\bar{x} = \frac{\sum f_i x_i}{\sum f_i}$$

4. Important note about class mark

If the class interval is 10–20, then the class mark is:

$$\frac{10+20}{2} = 15$$

If the interval is 20–30, then the class mark is:

$$\frac{20+30}{2} = 25$$

The class mark always lies in the middle of the interval.

Worked Example 1: Basic table

Find the mean of the following grouped data.

Class intervalFrequency \((f_i)\)
0–102
10–203
20–305

Step 1: Find class marks

  • For 0–10, \(x_i = \frac{0+10}{2} = 5\)
  • For 10–20, \(x_i = \frac{10+20}{2} = 15\)
  • For 20–30, \(x_i = \frac{20+30}{2} = 25\)

Step 2: Make the table

Class interval\(f_i\)\(x_i\)\(f_i x_i\)
0–102510
10–2031545
20–30525125

Step 3: Add the columns

$$\sum f_i = 2+3+5 = 10$$ $$\sum f_i x_i = 10+45+125 = 180$$

Step 4: Find the mean

$$\bar{x} = \frac{\sum f_i x_i}{\sum f_i} = \frac{180}{10} = 18$$

Answer: The mean is 18.

Worked Example 2: Marks scored by students

The marks of 20 students are grouped as follows. Find the mean.

MarksFrequency
10–204
20–306
30–405
40–503
50–602

Step 1: Find class marks

  • 10–20: \(15\)
  • 20–30: \(25\)
  • 30–40: \(35\)
  • 40–50: \(45\)
  • 50–60: \(55\)

Step 2: Multiply frequency by class mark

Marks\(f_i\)\(x_i\)\(f_i x_i\)
10–2041560
20–30625150
30–40535175
40–50345135
50–60255110

Step 3: Find totals

$$\sum f_i = 4+6+5+3+2 = 20$$ $$\sum f_i x_i = 60+150+175+135+110 = 630$$

Step 4: Use the formula

$$\bar{x} = \frac{630}{20} = 31.5$$

Answer: The mean marks are 31.5.

Worked Example 3: Slightly bigger data set

Find the mean of the following distribution.

Class intervalFrequency
5–156
15–258
25–3510
35–457
45–554

Step 1: Find class marks

  • 5–15: \(10\)
  • 15–25: \(20\)
  • 25–35: \(30\)
  • 35–45: \(40\)
  • 45–55: \(50\)

Step 2: Prepare the table

Class interval\(f_i\)\(x_i\)\(f_i x_i\)
5–1561060
15–25820160
25–351030300
35–45740280
45–55450200

Step 3: Calculate totals

$$\sum f_i = 6+8+10+7+4 = 35$$ $$\sum f_i x_i = 60+160+300+280+200 = 1000$$

Step 4: Find the mean

$$\bar{x} = \frac{1000}{35} = 28.57$$

Answer: The mean is approximately 28.57.

5. Why do we use class marks?

In grouped data, we do not know the exact values of all observations. For example, in the class interval 20–30, the actual values could be 21, 24, 29, and so on. To make calculation possible, we assume that all values in that class are centered around the midpoint.

This gives us an estimated mean, and for grouped data this is the standard way to calculate it.

6. Common mistakes to avoid

  • Using class limits instead of class marks: Do not multiply frequency by 10 or 20 directly for the class 10–20. First find the class mark, which is 15.
  • Forgetting to multiply: Make sure you calculate \(f_i x_i\) for every row.
  • Adding incorrectly: Check both \(\sum f_i\) and \(\sum f_i x_i\) carefully.
  • Wrong formula: The direct method formula is: $$\bar{x} = \frac{\sum f_i x_i}{\sum f_i}$$

7. Quick practice question

Find the mean of the following grouped data:

Class intervalFrequency
0–53
5–105
10–154

Solution:

Class marks are:

  • 0–5: \(2.5\)
  • 5–10: \(7.5\)
  • 10–15: \(12.5\)
Class interval\(f_i\)\(x_i\)\(f_i x_i\)
0–532.57.5
5–1057.537.5
10–15412.550
$$\sum f_i = 3+5+4 = 12$$ $$\sum f_i x_i = 7.5+37.5+50 = 95$$ $$\bar{x} = \frac{95}{12} \approx 7.92$$

Answer: The mean is approximately 7.92.

8. Final summary

To find the mean of grouped data by the direct method, first find the class mark of each interval. Then multiply each class mark by its frequency, add all these products, and divide by the total frequency.

The formula is:

$$\bar{x} = \frac{\sum f_i x_i}{\sum f_i}$$

This method is simple and reliable when data is given in grouped form. If you carefully find class marks and do the multiplication correctly, you can solve these questions with confidence.

Put what you read to the test

You've worked through Mean of Grouped Data: Direct Method. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Mean of Grouped Data: Assumed Mean Method

Mean of Grouped Data: Assumed Mean Method

When data is given in groups, we cannot see every individual value. Instead, the data is organized into class intervals such as 0–10, 10–20, 20–30, and so on, with their frequencies.

In such cases, to find the mean, we use the class marks (midpoints of the intervals) as representative values. One useful way to calculate the mean is the Assumed Mean Method.

This method makes calculation easier, especially when the class marks are large numbers. Instead of multiplying every class mark directly by its frequency, we choose one convenient value as the assumed mean and work with deviations from it.

Why do we use the assumed mean method?

  • It reduces large calculations.
  • It is helpful when class marks are close to each other.
  • It makes finding the mean faster and neater.

Step 1: Understand grouped data

A grouped frequency distribution usually has:

  • Class interval — the range of values
  • Frequency — the number of observations in that class
  • Class mark — the midpoint of the class interval

The class mark of a class interval is:

$$x_i = \frac{\text{lower limit} + \text{upper limit}}{2}$$

For example, for the class interval 20–30, the class mark is:

$$x_i = \frac{20+30}{2} = 25$$

Step 2: Idea of the assumed mean

Suppose the class marks are 15, 25, 35, 45, 55. Instead of using these directly, we choose one value, usually the middle one or any convenient class mark, as the assumed mean. Let this be represented by \(a\).

Then we find how far each class mark is from \(a\):

$$d_i = x_i - a$$

Here:

  • \(x_i\) = class mark
  • \(a\) = assumed mean
  • \(d_i\) = deviation from the assumed mean

After that, we multiply each deviation by its frequency:

$$f_i d_i$$

Then the mean is found using:

$$\bar{x} = a + \frac{\sum f_i d_i}{\sum f_i}$$

This is the formula for the assumed mean method.

Important note: The assumed mean is not the final mean. It is only a helpful starting value chosen to simplify the work.

Steps to find the mean by assumed mean method

  1. Write the class intervals and frequencies.
  2. Find the class mark \(x_i\) of each class.
  3. Choose a suitable assumed mean \(a\).
  4. Find deviations \(d_i = x_i - a\).
  5. Calculate \(f_i d_i\).
  6. Find \(\sum f_i\) and \(\sum f_i d_i\).
  7. Use the formula: $$\bar{x} = a + \frac{\sum f_i d_i}{\sum f_i}$$

Worked Example 1

Find the mean of the following grouped data.

Class intervals: 0–10, 10–20, 20–30, 30–40, 40–50
Frequencies: 3, 5, 7, 4, 1

Step 1: Find class marks

The class marks are:

  • 0–10: \(5\)
  • 10–20: \(15\)
  • 20–30: \(25\)
  • 30–40: \(35\)
  • 40–50: \(45\)

Choose assumed mean \(a = 25\).

Table:

\[ \begin{array}{|c|c|c|c|c|} \hline \text{Class Interval} & f_i & x_i & d_i = x_i - 25 & f_i d_i \\ \hline 0\text{–}10 & 3 & 5 & -20 & -60 \\ 10\text{–}20 & 5 & 15 & -10 & -50 \\ 20\text{–}30 & 7 & 25 & 0 & 0 \\ 30\text{–}40 & 4 & 35 & 10 & 40 \\ 40\text{–}50 & 1 & 45 & 20 & 20 \\ \hline \end{array} \]

Now,

$$\sum f_i = 3+5+7+4+1 = 20$$

$$\sum f_i d_i = -60-50+0+40+20 = -50$$

Use the formula:

$$\bar{x} = a + \frac{\sum f_i d_i}{\sum f_i}$$

$$\bar{x} = 25 + \frac{-50}{20}$$

$$\bar{x} = 25 - 2.5 = 22.5$$

Mean = 22.5

Worked Example 2

The marks obtained by 30 students are given below. Find the mean by the assumed mean method.

Class intervals: 10–20, 20–30, 30–40, 40–50, 50–60
Frequencies: 4, 6, 10, 7, 3

Step 1: Find class marks

The class marks are 15, 25, 35, 45, 55.

Choose assumed mean \(a = 35\).

\[ \begin{array}{|c|c|c|c|c|} \hline \text{Class Interval} & f_i & x_i & d_i = x_i - 35 & f_i d_i \\ \hline 10\text{–}20 & 4 & 15 & -20 & -80 \\ 20\text{–}30 & 6 & 25 & -10 & -60 \\ 30\text{–}40 & 10 & 35 & 0 & 0 \\ 40\text{–}50 & 7 & 45 & 10 & 70 \\ 50\text{–}60 & 3 & 55 & 20 & 60 \\ \hline \end{array} \]

Now,

$$\sum f_i = 4+6+10+7+3 = 30$$

$$\sum f_i d_i = -80-60+0+70+60 = -10$$

So,

$$\bar{x} = 35 + \frac{-10}{30}$$

$$\bar{x} = 35 - \frac{1}{3}$$

$$\bar{x} = 34.67 \text{ (approximately)}$$

Mean \(\approx 34.67\)

Worked Example 3

Find the mean of the following distribution.

Class intervals: 40–50, 50–60, 60–70, 70–80, 80–90, 90–100
Frequencies: 5, 8, 12, 10, 3, 2

Step 1: Find class marks

The class marks are 45, 55, 65, 75, 85, 95.

Choose assumed mean \(a = 65\).

\[ \begin{array}{|c|c|c|c|c|} \hline \text{Class Interval} & f_i & x_i & d_i = x_i - 65 & f_i d_i \\ \hline 40\text{–}50 & 5 & 45 & -20 & -100 \\ 50\text{–}60 & 8 & 55 & -10 & -80 \\ 60\text{–}70 & 12 & 65 & 0 & 0 \\ 70\text{–}80 & 10 & 75 & 10 & 100 \\ 80\text{–}90 & 3 & 85 & 20 & 60 \\ 90\text{–}100 & 2 & 95 & 30 & 60 \\ \hline \end{array} \]

Now,

$$\sum f_i = 5+8+12+10+3+2 = 40$$

$$\sum f_i d_i = -100-80+0+100+60+60 = 40$$

Therefore,

$$\bar{x} = 65 + \frac{40}{40}$$

$$\bar{x} = 65 + 1 = 66$$

Mean = 66

How to choose the assumed mean

You can choose any class mark as the assumed mean, but it is best to choose:

  • a class mark near the center, or
  • a class mark that makes deviations easy to calculate

If you choose a central value, the positive and negative deviations are usually smaller, which makes the arithmetic simpler.

Common mistakes to avoid

  • Forgetting to find class marks: Do not use class intervals directly in the formula. Use their midpoints.
  • Using the wrong assumed mean: The assumed mean should be a class mark, not a class interval.
  • Sign errors: Negative deviations must stay negative.
  • Adding frequencies incorrectly: Check \(\sum f_i\) carefully.
  • Adding \(f_i d_i\) incorrectly: Positive and negative values must be handled carefully.

Quick comparison with direct method

In the direct method, we use:

$$\bar{x} = \frac{\sum f_i x_i}{\sum f_i}$$

In the assumed mean method, we use:

$$\bar{x} = a + \frac{\sum f_i d_i}{\sum f_i}$$

Both methods give the same answer. The assumed mean method is often easier when the numbers are bigger or when calculations look lengthy.

Practice Question 1

Find the mean by the assumed mean method:

Class intervals: 0–10, 10–20, 20–30, 30–40, 40–50
Frequencies: 2, 6, 8, 3, 1

Practice Question 2

Find the mean by the assumed mean method:

Class intervals: 5–15, 15–25, 25–35, 35–45, 45–55
Frequencies: 4, 5, 9, 6, 2

Final Summary

The assumed mean method is a simple way to find the mean of grouped data. We first find the class marks, choose a convenient assumed mean \(a\), calculate deviations \(d_i = x_i - a\), and then use the formula:

$$\bar{x} = a + \frac{\sum f_i d_i}{\sum f_i}$$

This method saves time and reduces calculation work. If you carefully find the class marks, keep track of positive and negative deviations, and add correctly, you can easily find the mean of grouped data.

Put what you read to the test

You've worked through Mean of Grouped Data: Assumed Mean Method. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Mean of Grouped Data: Step-Deviation Method

Mean of Grouped Data: Step-Deviation Method

When data is given in grouped form, the exact individual values are not known. Instead, the data is arranged into class intervals such as 0–10, 10–20, 20–30, and so on, along with their frequencies. In such cases, we find the mean using the class marks and frequencies.

The step-deviation method is a short and efficient way to calculate the mean of grouped data, especially when the class intervals are equal and the class marks are large numbers. It reduces long calculations by working with smaller values.

In this lesson, you will learn what the step-deviation method is, why it works, how to use it, and how to solve questions step by step.

1. Recall: Mean of grouped data

For grouped data, we first find the class mark of each class interval. The class mark is the midpoint of the interval.

If a class interval is 10–20, then its class mark is

$$x_i = \frac{10+20}{2} = 15$$

If the frequencies are given by \(f_i\), then the mean by the direct method is

$$\bar{x} = \frac{\sum f_i x_i}{\sum f_i}$$

This method works well, but if the values of \(x_i\) are large, the multiplication becomes lengthy. That is why we use the assumed mean method or the even shorter step-deviation method.

2. Idea behind the step-deviation method

Suppose the class marks increase regularly, like 15, 25, 35, 45, 55. The difference between consecutive class marks is the same. This common difference is usually the class width, denoted by \(h\).

Instead of taking deviations directly from an assumed mean \(A\), we divide each deviation by \(h\). This gives smaller numbers and makes the arithmetic easier.

We define

$$u_i = \frac{x_i - A}{h}$$

Then the mean is calculated using

$$\bar{x} = A + h\left(\frac{\sum f_i u_i}{\sum f_i}\right)$$

This is called the step-deviation formula.

3. Symbols used

  • Class interval: the group, such as 20–30
  • Frequency \(f_i\): the number of observations in that class
  • Class mark \(x_i\): midpoint of the class
  • Assumed mean \(A\): a convenient class mark chosen to simplify work
  • Class width \(h\): size of each class interval
  • \(u_i = \frac{x_i-A}{h}\): scaled deviation

4. When should we use the step-deviation method?

This method is most useful when:

  • class intervals are of equal width,
  • class marks are large numbers,
  • you want shorter and faster calculations.

If the class widths are not equal, this method is usually not convenient in the standard form taught at this level.

5. Steps to find mean by step-deviation method

  1. Write the class intervals and frequencies.
  2. Find the class mark \(x_i\) for each class.
  3. Choose a suitable assumed mean \(A\), usually the middle class mark.
  4. Find the class width \(h\).
  5. Calculate \(u_i = \frac{x_i-A}{h}\).
  6. Find \(f_i u_i\).
  7. Find \(\sum f_i\) and \(\sum f_i u_i\).
  8. Use the formula:
$$\bar{x} = A + h\left(\frac{\sum f_i u_i}{\sum f_i}\right)$$

6. Important note about choosing \(A\)

You may choose any class mark as the assumed mean. However, it is best to choose a class mark near the center of the data. This keeps the values of \(u_i\) small, making calculations simpler.

7. Worked Example 1

Find the mean of the following grouped data using the step-deviation method.

Class intervalFrequency
0–105
10–208
20–3012
30–406
40–504

Step 1: Find class marks

The class marks are:

$$5, 15, 25, 35, 45$$

Step 2: Choose assumed mean and class width

Take \(A = 25\), since it is the middle class mark.

Class width \(h = 10\).

Step 3: Find \(u_i = \frac{x_i-A}{h}\)

  • For \(x_i = 5\), \(u_i = \frac{5-25}{10} = -2\)
  • For \(x_i = 15\), \(u_i = \frac{15-25}{10} = -1\)
  • For \(x_i = 25\), \(u_i = 0\)
  • For \(x_i = 35\), \(u_i = 1\)
  • For \(x_i = 45\), \(u_i = 2\)

Now make the table:

Class interval\(f_i\)\(x_i\)\(u_i\)\(f_i u_i\)
0–1055-2-10
10–20815-1-8
20–30122500
30–4063516
40–5044528

Now,

$$\sum f_i = 5+8+12+6+4 = 35$$ $$\sum f_i u_i = -10-8+0+6+8 = -4$$

Using the formula:

$$\bar{x} = A + h\left(\frac{\sum f_i u_i}{\sum f_i}\right)$$ $$\bar{x} = 25 + 10\left(\frac{-4}{35}\right)$$ $$\bar{x} = 25 - \frac{40}{35}$$ $$\bar{x} = 25 - 1.14 \approx 23.86$$

Answer: The mean is approximately \(23.86\).

8. Worked Example 2

Find the mean of the following data.

Class intervalFrequency
10–203
20–305
30–409
40–507
50–606

Step 1: Class marks

$$x_i = 15, 25, 35, 45, 55$$

Step 2: Choose \(A = 35\), \(h = 10\)

Step 3: Find \(u_i\) and \(f_i u_i\)

Class interval\(f_i\)\(x_i\)\(u_i = \frac{x_i-35}{10}\)\(f_i u_i\)
10–20315-2-6
20–30525-1-5
30–4093500
40–5074517
50–60655212

Now,

$$\sum f_i = 3+5+9+7+6 = 30$$ $$\sum f_i u_i = -6-5+0+7+12 = 8$$

Apply the formula:

$$\bar{x} = 35 + 10\left(\frac{8}{30}\right)$$ $$\bar{x} = 35 + \frac{80}{30}$$ $$\bar{x} = 35 + 2.67 \approx 37.67$$

Answer: The mean is approximately \(37.67\).

9. Worked Example 3

The heights of plants in cm are grouped as follows. Find the mean height.

Height (cm)Number of plants
140–1504
150–1607
160–17010
170–1808
180–1905

Step 1: Find class marks

$$x_i = 145, 155, 165, 175, 185$$

Step 2: Choose assumed mean and class width

Take \(A = 165\) and \(h = 10\).

Step 3: Compute \(u_i\)

$$u_i = \frac{x_i-165}{10}$$

So the values of \(u_i\) are \(-2, -1, 0, 1, 2\).

Class interval\(f_i\)\(x_i\)\(u_i\)\(f_i u_i\)
140–1504145-2-8
150–1607155-1-7
160–1701016500
170–180817518
180–1905185210

Then,

$$\sum f_i = 4+7+10+8+5 = 34$$ $$\sum f_i u_i = -8-7+0+8+10 = 3$$

Now use the formula:

$$\bar{x} = 165 + 10\left(\frac{3}{34}\right)$$ $$\bar{x} = 165 + \frac{30}{34}$$ $$\bar{x} \approx 165.88$$

Answer: The mean height is approximately \(165.88\) cm.

10. Worked Example 4

Find the mean for the following grouped data.

Class intervalFrequency
50–606
60–7010
70–8014
80–908
90–1002

Step 1: Class marks

$$x_i = 55, 65, 75, 85, 95$$

Step 2: Choose \(A = 75\), \(h = 10\)

Step 3: Find \(u_i\) and \(f_i u_i\)

Class interval\(f_i\)\(x_i\)\(u_i\)\(f_i u_i\)
50–60655-2-12
60–701065-1-10
70–80147500
80–9088518
90–10029524

Now,

$$\sum f_i = 6+10+14+8+2 = 40$$ $$\sum f_i u_i = -12-10+0+8+4 = -10$$

Apply the formula:

$$\bar{x} = 75 + 10\left(\frac{-10}{40}\right)$$ $$\bar{x} = 75 - 2.5 = 72.5$$

Answer: The mean is \(72.5\).

11. Why the method works

In the assumed mean method, we take deviations \(d_i = x_i - A\). In the step-deviation method, we write these deviations in a smaller form by dividing by the common class width \(h\):

$$u_i = \frac{d_i}{h} = \frac{x_i-A}{h}$$

This makes numbers like \(-20, -10, 0, 10, 20\) become \(-2, -1, 0, 1, 2\). The values are easier to multiply by frequency, so the work becomes faster and neater.

12. Common mistakes to avoid

  • Using class limits instead of class marks: always use the midpoint of each class.
  • Forgetting to divide by \(h\): step-deviation needs \(u_i = \frac{x_i-A}{h}\), not just \(x_i-A\).
  • Taking the wrong class width: for intervals like 20–30, 30–40, the width is 10.
  • Ignoring signs: negative and positive values of \(u_i\) must be handled carefully.
  • Using the wrong formula: the correct formula is
$$\bar{x} = A + h\left(\frac{\sum f_i u_i}{\sum f_i}\right)$$

13. Quick comparison of methods

  • Direct method: uses \(\bar{x} = \frac{\sum f_i x_i}{\sum f_i}\)
  • Assumed mean method: uses deviations \(d_i = x_i-A\)
  • Step-deviation method: uses scaled deviations \(u_i = \frac{x_i-A}{h}\)

The step-deviation method is usually the fastest when class intervals are equal.

14. Final checklist for solving problems

  • Did you find all class marks correctly?
  • Did you choose a suitable assumed mean \(A\)?
  • Did you use the correct class width \(h\)?
  • Did you calculate \(u_i\) correctly?
  • Did you find \(f_i u_i\), \(\sum f_i\), and \(\sum f_i u_i\) correctly?
  • Did you substitute into the formula carefully?

15. Summary

The step-deviation method is a shortcut for finding the mean of grouped data when class intervals are equal. We first find class marks, choose an assumed mean \(A\), divide deviations by the class width \(h\), and use the formula

$$\bar{x} = A + h\left(\frac{\sum f_i u_i}{\sum f_i}\right)$$

This method makes calculations shorter and easier, especially when the numbers are large. With careful use of class marks, signs, and the formula, you can find the mean quickly and accurately.

Put what you read to the test

You've worked through Mean of Grouped Data: Step-Deviation Method. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Mode of Grouped Data

Mode of Grouped Data

When data is arranged into groups or class intervals, we cannot always see the exact value that appears most often. In this case, we estimate the mode using the class with the greatest frequency and a special formula.

This lesson will show you what the mode of grouped data means, how to find the modal class, and how to calculate the estimated modal value using interpolation.

1. What is the mode?

The mode is the value that occurs most often in a set of data.

For example, in the data set \(2, 3, 3, 5, 7\), the mode is \(3\) because it appears more times than any other value.

But when data is grouped into intervals such as \(0\text{–}10\), \(10\text{–}20\), \(20\text{–}30\), we do not know the exact repeated values. We only know how many values lie in each interval. So we estimate the mode.

2. Modal class

The modal class is the class interval with the highest frequency.

Suppose we have this grouped table:

  • \(0\text{–}10: 4\)
  • \(10\text{–}20: 7\)
  • \(20\text{–}30: 12\)
  • \(30\text{–}40: 9\)

The highest frequency is \(12\), so the modal class is \(20\text{–}30\).

However, the mode is not simply \(25\) or just the class interval. We use a formula to estimate the exact modal value inside that interval.

3. Formula for the mode of grouped data

The formula is:

$$ \text{Mode} = l + \left(\frac{f_1 - f_0}{2f_1 - f_0 - f_2}\right)h $$

where:

  • \(l\) = lower boundary of the modal class
  • \(h\) = class width
  • \(f_1\) = frequency of the modal class
  • \(f_0\) = frequency of the class before the modal class
  • \(f_2\) = frequency of the class after the modal class

This formula uses the frequencies on both sides of the modal class to estimate where the highest point lies inside the class interval.

4. Why this formula makes sense

If the class before the modal class has a much smaller frequency, then the peak is likely to be closer to the left side of the modal class.

If the class after the modal class has a much smaller frequency, then the peak is likely to be closer to the right side of the modal class.

So the formula helps us place the mode more accurately than simply picking the middle of the modal class.

5. Important note about class boundaries

In grouped continuous data, we usually use class boundaries, not just the written class limits.

For example, if the classes are written as:

  • \(10\text{–}19\)
  • \(20\text{–}29\)
  • \(30\text{–}39\)

then the class boundaries are:

  • \(9.5\text{–}19.5\)
  • \(19.5\text{–}29.5\)
  • \(29.5\text{–}39.5\)

This is important because the mode formula uses the lower boundary \(l\) and the class width \(h\).

6. Steps to find the mode of grouped data

  1. Find the class with the highest frequency. This is the modal class.
  2. Write down:
    • \(l\): lower boundary of the modal class
    • \(h\): class width
    • \(f_1\): frequency of the modal class
    • \(f_0\): frequency before the modal class
    • \(f_2\): frequency after the modal class
  3. Substitute into the formula:
$$ \text{Mode} = l + \left(\frac{f_1 - f_0}{2f_1 - f_0 - f_2}\right)h $$
  1. Simplify carefully.
  2. State the answer clearly as an estimated mode.

7. Worked Example 1

Find the mode of the grouped data:

Class intervalFrequency
0–105
10–208
20–3014
30–409
40–504

Step 1: Find the modal class

The highest frequency is \(14\), so the modal class is \(20\text{–}30\).

Step 2: Identify the values

  • \(l = 20\)
  • \(h = 10\)
  • \(f_1 = 14\)
  • \(f_0 = 8\)
  • \(f_2 = 9\)

Step 3: Use the formula

$$ \text{Mode} = 20 + \left(\frac{14 - 8}{2(14) - 8 - 9}\right)(10) $$ $$ = 20 + \left(\frac{6}{28 - 17}\right)(10) $$ $$ = 20 + \left(\frac{6}{11}\right)(10) $$ $$ = 20 + \frac{60}{11} $$ $$ = 20 + 5.45 $$ $$ \text{Mode} \approx 25.45 $$

Answer: The estimated mode is \(25.45\).

8. Worked Example 2

Find the mode of the grouped data:

Class intervalFrequency
10–206
20–3011
30–4015
40–5013
50–607

Step 1: Modal class

The highest frequency is \(15\), so the modal class is \(30\text{–}40\).

Step 2: Write the values

  • \(l = 30\)
  • \(h = 10\)
  • \(f_1 = 15\)
  • \(f_0 = 11\)
  • \(f_2 = 13\)

Step 3: Substitute

$$ \text{Mode} = 30 + \left(\frac{15 - 11}{2(15) - 11 - 13}\right)(10) $$ $$ = 30 + \left(\frac{4}{30 - 24}\right)(10) $$ $$ = 30 + \left(\frac{4}{6}\right)(10) $$ $$ = 30 + 6.67 $$ $$ \text{Mode} \approx 36.67 $$

Answer: The estimated mode is \(36.67\).

Notice that the mode is closer to the upper end of the modal class because the frequency after the modal class, \(13\), is still quite high.

9. Worked Example 3 with class boundaries

Find the mode of the following grouped data:

MarksFrequency
0–93
10–197
20–2912
30–3910
40–495

Step 1: Modal class

The highest frequency is \(12\), so the modal class is \(20\text{–}29\).

Step 2: Use class boundaries

The class \(20\text{–}29\) has boundaries \(19.5\) to \(29.5\).

  • \(l = 19.5\)
  • \(h = 10\)
  • \(f_1 = 12\)
  • \(f_0 = 7\)
  • \(f_2 = 10\)

Step 3: Apply the formula

$$ \text{Mode} = 19.5 + \left(\frac{12 - 7}{2(12) - 7 - 10}\right)(10) $$ $$ = 19.5 + \left(\frac{5}{24 - 17}\right)(10) $$ $$ = 19.5 + \left(\frac{5}{7}\right)(10) $$ $$ = 19.5 + 7.14 $$ $$ \text{Mode} \approx 26.64 $$

Answer: The estimated mode is \(26.64\).

10. Common mistakes to avoid

  • Choosing the wrong modal class: Always select the class with the highest frequency.
  • Using class limits instead of boundaries when needed: For intervals like \(20\text{–}29\), use \(19.5\) as the lower boundary.
  • Mixing up \(f_0, f_1, f_2\):
    • \(f_1\) is the modal class frequency
    • \(f_0\) is the frequency before it
    • \(f_2\) is the frequency after it
  • Using the wrong class width: Make sure \(h\) is the width of one class interval.
  • Forgetting that the answer is an estimate: Since the data is grouped, the mode is not exact.

11. Quick comparison: mode in ungrouped and grouped data

  • For ungrouped data, the mode is the value that appears most often.
  • For grouped data, the exact values are hidden inside intervals, so we estimate the mode using a formula.

12. When is the mode useful?

The mode is useful when you want to know the most common or most typical value in a distribution.

For grouped data, it helps estimate the value around which the data is most heavily concentrated.

13. Final recap

To find the mode of grouped data, first identify the modal class, which is the class with the highest frequency.

Then use the interpolation formula:

$$ \text{Mode} = l + \left(\frac{f_1 - f_0}{2f_1 - f_0 - f_2}\right)h $$

Remember:

  • \(l\) is the lower boundary of the modal class
  • \(h\) is the class width
  • \(f_1\) is the modal class frequency
  • \(f_0\) and \(f_2\) are the neighboring frequencies

This gives an estimated value for the mode inside the modal class.

14. Practice checklist

Whenever you solve a question on the mode of grouped data, ask yourself:

  • Did I find the class with the highest frequency?
  • Did I use the correct lower boundary?
  • Did I identify \(f_0\), \(f_1\), and \(f_2\) correctly?
  • Did I use the correct class width?
  • Did I clearly state the final answer as an estimate?

If you can answer yes to all of these, you are very likely to get the question correct.

Put what you read to the test

You've worked through Mode of Grouped Data. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Median of Grouped Data

Median of Grouped Data

When data is organized into class intervals, we call it grouped data. For example, instead of listing every test score one by one, we may group the scores as 0–10, 10–20, 20–30, and so on.

In grouped data, we cannot always see the exact middle value directly. So we use a method to estimate the median. This estimate is very useful when there is a large amount of data.

The median is the value that lies in the middle of the data when all values are arranged in order. It divides the data into two equal parts:

  • about half the data is below the median,
  • about half the data is above the median.

For grouped data, we find the median in two main steps:

  1. Find the median class using cumulative frequency.
  2. Use the grouped-data median formula to estimate the exact median value.

1. Important ideas

Frequency tells us how many values are in each class interval.

Cumulative frequency is the running total of frequencies. It helps us locate where the middle value lies.

If the total frequency is \(N\), then the median is located at the \(\frac{N}{2}\)-th value.

After finding the class interval that contains the \(\frac{N}{2}\)-th value, we call it the median class.

2. Formula for the median of grouped data

Once the median class is known, use:

$$ \text{Median} = l + \left(\frac{\frac{N}{2} - cf}{f}\right) \times h $$

Here:

  • \(l\) = lower boundary of the median class
  • \(N\) = total frequency
  • \(cf\) = cumulative frequency before the median class
  • \(f\) = frequency of the median class
  • \(h\) = class width

This formula works by assuming the data is spread evenly inside the median class. This process is called interpolation.

3. How to find the median step by step

  1. Write the class intervals and frequencies in a table.
  2. Find the total frequency \(N\).
  3. Compute \(\frac{N}{2}\).
  4. Make a cumulative frequency column.
  5. Find the class where the cumulative frequency first becomes greater than or equal to \(\frac{N}{2}\). That is the median class.
  6. Identify \(l\), \(cf\), \(f\), and \(h\).
  7. Substitute into the formula.

4. Worked Example 1

Find the median of the following grouped data.

Class intervalFrequency
0–105
10–207
20–3012
30–406

Step 1: Find total frequency

$$N = 5 + 7 + 12 + 6 = 30$$

Step 2: Find \(\frac{N}{2}\)

$$\frac{N}{2} = \frac{30}{2} = 15$$

Step 3: Find cumulative frequencies

Class intervalFrequencyCumulative frequency
0–1055
10–20712
20–301224
30–40630

The 15th value lies in the class 20–30, because the cumulative frequency goes from 12 to 24 there.

So the median class is 20–30.

Step 4: Identify the values

  • \(l = 20\)
  • \(cf = 12\)
  • \(f = 12\)
  • \(h = 10\)

Step 5: Use the formula

$$ \text{Median} = 20 + \left(\frac{15 - 12}{12}\right) \times 10 $$ $$ = 20 + \left(\frac{3}{12}\right) \times 10 $$ $$ = 20 + 2.5 = 22.5 $$

Answer: The median is \(22.5\).

5. Worked Example 2

The table shows the weights of students.

Weight (kg)Frequency
40–504
50–606
60–7010
70–808
80–902

Step 1: Total frequency

$$N = 4 + 6 + 10 + 8 + 2 = 30$$

Step 2: Middle position

$$\frac{N}{2} = 15$$

Step 3: Cumulative frequency table

Weight (kg)FrequencyCumulative frequency
40–5044
50–60610
60–701020
70–80828
80–90230

The 15th value lies in 60–70, so this is the median class.

Step 4: Identify values

  • \(l = 60\)
  • \(cf = 10\)
  • \(f = 10\)
  • \(h = 10\)

Step 5: Substitute

$$ \text{Median} = 60 + \left(\frac{15 - 10}{10}\right) \times 10 $$ $$ = 60 + 5 = 65 $$

Answer: The median weight is \(65\text{ kg}\).

6. Worked Example 3

Find the median for the following distribution.

Class intervalFrequency
5–153
15–259
25–3515
35–457
45–556

Step 1: Total frequency

$$N = 3 + 9 + 15 + 7 + 6 = 40$$

Step 2: Find \(\frac{N}{2}\)

$$\frac{N}{2} = 20$$

Step 3: Cumulative frequency

Class intervalFrequencyCumulative frequency
5–1533
15–25912
25–351527
35–45734
45–55640

The 20th value lies in the class 25–35.

So:

  • \(l = 25\)
  • \(cf = 12\)
  • \(f = 15\)
  • \(h = 10\)

Now use the formula:

$$ \text{Median} = 25 + \left(\frac{20 - 12}{15}\right) \times 10 $$ $$ = 25 + \left(\frac{8}{15}\right) \times 10 $$ $$ = 25 + 5.33 \approx 30.33 $$

Answer: The median is approximately \(30.33\).

7. Notes about class boundaries

Sometimes questions use continuous class intervals such as 0–10, 10–20, 20–30. In these cases, we often use the lower limit directly as \(l\).

In some textbooks, class intervals may be written in a way that suggests a gap, such as 0–9, 10–19, 20–29. Then class boundaries may be needed. For example:

  • 0–9 becomes \(-0.5\) to \(9.5\)
  • 10–19 becomes \(9.5\) to \(19.5\)

Then the lower boundary of the median class is used in the formula. Always check how the class intervals are written.

8. Common mistakes to avoid

  • Using \(\frac{N+1}{2}\) instead of \(\frac{N}{2}\) for grouped data. For grouped data, we usually use \(\frac{N}{2}\).
  • Choosing the wrong median class because the cumulative frequency table was incorrect.
  • Using the cumulative frequency of the median class instead of the cumulative frequency before the median class.
  • Using the wrong class width \(h\).
  • Forgetting to use class boundaries when needed.

9. Quick check method

After finding your answer, ask:

  • Is the answer inside the median class?
  • Does it make sense as a middle value?

For example, if the median class is 20–30, your median should usually lie between 20 and 30.

10. Brief summary

The median of grouped data is found by locating the middle position \(\frac{N}{2}\) and using cumulative frequency to identify the median class.

Then we estimate the exact median with:

$$ \text{Median} = l + \left(\frac{\frac{N}{2} - cf}{f}\right) \times h $$

This method helps us find the middle value even when the data is grouped into intervals instead of listed one by one.

Put what you read to the test

You've worked through Median of Grouped Data. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Empirical Relationship Between Mean, Median, and Mode

Empirical Relationship Between Mean, Median, and Mode

In statistics, the mean, median, and mode are called measures of central tendency. They describe the center or typical value of a set of data.

Sometimes, especially when working with grouped data or large data sets, it may be difficult to find all three measures exactly. In such cases, we use an empirical relationship to estimate one measure if the other two are known.

The empirical relationship is:

$$\text{Mode} = 3\times \text{Median} - 2\times \text{Mean}$$

This formula is an approximate relationship, not an exact rule for every data set. It works best for data that is moderately skewed, meaning the data is not perfectly symmetrical but also not extremely uneven.

Let us first recall the meaning of each measure:

  • Mean: The average of all observations.
  • Median: The middle value when the data is arranged in order.
  • Mode: The value that occurs most often.

When the data is not evenly spread, these three values may be different. The empirical relationship helps us connect them.

Standard Formula

$$\text{Mode} = 3\text{Median} - 2\text{Mean}$$

This same formula can be rearranged if a different quantity is missing.

To find the median:

$$\text{Mode} = 3\text{Median} - 2\text{Mean}$$

$$3\text{Median} = \text{Mode} + 2\text{Mean}$$

$$\text{Median} = \frac{\text{Mode} + 2\text{Mean}}{3}$$

To find the mean:

$$\text{Mode} = 3\text{Median} - 2\text{Mean}$$

$$2\text{Mean} = 3\text{Median} - \text{Mode}$$

$$\text{Mean} = \frac{3\text{Median} - \text{Mode}}{2}$$

Important Note

This relation is useful when:

  • one of the three measures is missing,
  • the data is grouped,
  • you need a quick estimate,
  • or the mode is difficult to determine directly.

But remember: since it is empirical, the answer is usually an estimate.

How to Use the Formula

  1. Write the formula clearly.
  2. Substitute the known values.
  3. Solve step by step.
  4. Check whether the answer is reasonable.

Worked Example 1: Finding the Mode

The mean of a data set is 18 and the median is 20. Find the mode using the empirical relationship.

Step 1: Write the formula

$$\text{Mode} = 3\text{Median} - 2\text{Mean}$$

Step 2: Substitute the values

$$\text{Mode} = 3(20) - 2(18)$$

Step 3: Simplify

$$\text{Mode} = 60 - 36 = 24$$

Answer: The mode is approximately 24.

Worked Example 2: Finding the Median

The mean of a distribution is 26 and the mode is 32. Find the median.

Step 1: Use the rearranged formula

$$\text{Median} = \frac{\text{Mode} + 2\text{Mean}}{3}$$

Step 2: Substitute the values

$$\text{Median} = \frac{32 + 2(26)}{3}$$

Step 3: Simplify

$$\text{Median} = \frac{32 + 52}{3} = \frac{84}{3} = 28$$

Answer: The median is 28.

Worked Example 3: Finding the Mean

The median is 35 and the mode is 41. Find the mean.

Step 1: Use the rearranged formula

$$\text{Mean} = \frac{3\text{Median} - \text{Mode}}{2}$$

Step 2: Substitute the values

$$\text{Mean} = \frac{3(35) - 41}{2}$$

Step 3: Simplify

$$\text{Mean} = \frac{105 - 41}{2} = \frac{64}{2} = 32$$

Answer: The mean is 32.

Worked Example 4: A Word Problem with Grouped Data

In a grouped distribution, the mean mark of students is 48 and the median mark is 50. Estimate the mode.

Since this is grouped data, the exact mode may be hard to calculate directly. So we use the empirical formula.

$$\text{Mode} = 3\text{Median} - 2\text{Mean}$$

$$\text{Mode} = 3(50) - 2(48)$$

$$\text{Mode} = 150 - 96 = 54$$

Answer: The estimated mode is 54 marks.

Understanding the Relationship

This formula shows how the three measures are connected in many practical situations. If the mean is smaller than the median, the mode may become larger. If the mean is larger, the mode may become smaller.

This does not mean the formula always gives a perfect answer. It is mainly used to make a reasonable estimate when exact calculation is difficult.

Common Mistakes to Avoid

  • Using the wrong formula: Write the formula carefully before substituting values.
  • Confusing mean and median: Read the question closely.
  • Arithmetic errors: Be careful when multiplying by 2 or 3.
  • Forgetting it is approximate: The result is an estimate, especially for grouped data.

Quick Check

If mean = 22 and median = 25, then:

$$\text{Mode} = 3(25) - 2(22) = 75 - 44 = 31$$

If median = 30 and mode = 36, then:

$$\text{Mean} = \frac{3(30) - 36}{2} = \frac{90 - 36}{2} = \frac{54}{2} = 27$$

When is this Formula Helpful?

  • In questions on grouped frequency distributions
  • When one central tendency measure is missing
  • In estimation-based statistics problems
  • When checking whether your answer is reasonable

Summary

The empirical relationship between mean, median, and mode is:

$$\text{Mode} = 3\text{Median} - 2\text{Mean}$$

This formula is used to estimate one measure of central tendency when the other two are known. It is especially helpful in grouped data, where finding the exact mode may be difficult.

By rearranging the formula, you can also find the mean or median:

$$\text{Median} = \frac{\text{Mode} + 2\text{Mean}}{3}$$

$$\text{Mean} = \frac{3\text{Median} - \text{Mode}}{2}$$

Always remember that this is an approximate relationship, so it gives an estimate rather than an exact result in many real situations.

Put what you read to the test

You've worked through Empirical Relationship Between Mean, Median, and Mode. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Constructing Cumulative Frequency Distributions

Constructing Cumulative Frequency Distributions

When we collect a lot of data, it can be hard to understand the whole set by looking at individual values. In statistics, we often group data into class intervals such as 0–10, 10–20, 20–30, and so on. A frequency distribution tells us how many data values fall in each class.

A cumulative frequency distribution goes one step further. It shows a running total of the frequencies. This helps us answer questions like:

  • How many values are less than 30?
  • How many values are at most 50?
  • How many values are greater than 20?

There are two common types of cumulative frequency distributions:

  • Less-than cumulative frequency: the running total up to the upper boundary of each class.
  • More-than cumulative frequency: the running total starting from the lower boundary of each class and going upward.

Learning how to build these tables is important because they are used to draw cumulative frequency curves (also called ogives) and to estimate medians, quartiles, and percentiles from grouped data.

1. Review: Frequency distribution

A grouped frequency table usually has:

  • Class interval: the range of values in each group
  • Frequency: the number of values in that group

For example:

Suppose test scores are grouped like this:

ScoreFrequency
0–103
10–205
20–307
30–404
40–501

The total number of data values is:

$$3+5+7+4+1=20$$

This total is often written as \(N=20\).

2. What is cumulative frequency?

Cumulative frequency means adding frequencies as you move through the classes.

For a less-than table, you add from the first class downward:

  • first cumulative frequency = first frequency
  • second cumulative frequency = first + second frequency
  • third cumulative frequency = first + second + third frequency

In symbols, if the frequencies are \(f_1, f_2, f_3, \dots\), then the cumulative frequencies are:

$$f_1,\quad f_1+f_2,\quad f_1+f_2+f_3,\quad \dots$$

3. Constructing a less-than cumulative frequency distribution

A less-than cumulative frequency table answers questions such as:

  • How many values are less than 10?
  • How many values are less than 20?
  • How many values are less than 30?

To make this table:

  1. Write the upper class limits or upper class boundaries.
  2. Add the frequencies from top to bottom.
  3. Record each running total beside the correct upper value.

Using the score table above:

ScoreFrequency
0–103
10–205
20–307
30–404
40–501

The less-than cumulative frequencies are:

  • Less than 10: \(3\)
  • Less than 20: \(3+5=8\)
  • Less than 30: \(3+5+7=15\)
  • Less than 40: \(3+5+7+4=19\)
  • Less than 50: \(3+5+7+4+1=20\)

So the cumulative frequency table is:

Less thanCumulative Frequency
103
208
3015
4019
5020

Important idea: The last cumulative frequency in a less-than table should equal the total frequency \(N\).

4. Constructing a more-than cumulative frequency distribution

A more-than cumulative frequency table answers questions such as:

  • How many values are more than 0?
  • How many values are more than 10?
  • How many values are more than 20?

To make this table:

  1. Start with the total frequency.
  2. Use the lower class limits or lower class boundaries.
  3. Subtract frequencies as you move down the classes.

Using the same data, total frequency \(N=20\).

  • More than 0: \(20\)
  • More than 10: \(20-3=17\)
  • More than 20: \(20-(3+5)=12\)
  • More than 30: \(20-(3+5+7)=5\)
  • More than 40: \(20-(3+5+7+4)=1\)

So the more-than cumulative frequency table is:

More thanCumulative Frequency
020
1017
2012
305
401

Important idea: In a more-than table, the first cumulative frequency is the total frequency.

5. How less-than and more-than tables are connected

Both tables come from the same original frequencies. They just organize the running totals in different ways.

  • Less-than cumulative frequency increases as you move down the table.
  • More-than cumulative frequency decreases as you move down the table.

If your less-than values go upward and your more-than values go downward, that is a good sign that your table is correct.

6. Worked Example 1: Basic less-than cumulative frequency

The heights of plants are grouped below:

Height (cm)Frequency
0–52
5–106
10–158
15–204

Step 1: Find the total frequency.

$$N=2+6+8+4=20$$

Step 2: Add frequencies from top to bottom.

  • Less than 5: \(2\)
  • Less than 10: \(2+6=8\)
  • Less than 15: \(2+6+8=16\)
  • Less than 20: \(2+6+8+4=20\)

Answer:

Less thanCumulative Frequency
52
108
1516
2020

7. Worked Example 2: Basic more-than cumulative frequency

Use the same plant height data:

Height (cm)Frequency
0–52
5–106
10–158
15–204

Step 1: Start with the total frequency.

$$N=20$$

Step 2: Subtract frequencies as you move downward.

  • More than 0: \(20\)
  • More than 5: \(20-2=18\)
  • More than 10: \(20-(2+6)=12\)
  • More than 15: \(20-(2+6+8)=4\)

Answer:

More thanCumulative Frequency
020
518
1012
154

8. Worked Example 3: Making both cumulative frequency tables

A teacher groups the time students spent reading in one week:

Reading Time (hours)Frequency
0–25
2–49
4–67
6–83
8–101

Step 1: Total frequency

$$N=5+9+7+3+1=25$$

Step 2: Less-than cumulative frequency

  • Less than 2: \(5\)
  • Less than 4: \(5+9=14\)
  • Less than 6: \(5+9+7=21\)
  • Less than 8: \(5+9+7+3=24\)
  • Less than 10: \(25\)

Less-than table:

Less thanCumulative Frequency
25
414
621
824
1025

Step 3: More-than cumulative frequency

  • More than 0: \(25\)
  • More than 2: \(25-5=20\)
  • More than 4: \(25-(5+9)=11\)
  • More than 6: \(25-(5+9+7)=4\)
  • More than 8: \(25-(5+9+7+3)=1\)

More-than table:

More thanCumulative Frequency
025
220
411
64
81

9. Worked Example 4: Using class boundaries carefully

Sometimes grouped data is written with whole-number class intervals, but when drawing graphs, we may use class boundaries. For constructing cumulative frequency, the main idea stays the same: use the upper boundary for less-than tables and the lower boundary for more-than tables.

Suppose the masses of bags are grouped like this:

Mass (kg)Frequency
10–144
15–196
20–2410
25–295

If we treat the classes as continuous, the boundaries are:

  • 9.5–14.5
  • 14.5–19.5
  • 19.5–24.5
  • 24.5–29.5

Step 1: Total frequency

$$N=4+6+10+5=25$$

Step 2: Less-than cumulative frequency using upper boundaries

  • Less than 14.5: \(4\)
  • Less than 19.5: \(4+6=10\)
  • Less than 24.5: \(4+6+10=20\)
  • Less than 29.5: \(25\)

Step 3: More-than cumulative frequency using lower boundaries

  • More than 9.5: \(25\)
  • More than 14.5: \(25-4=21\)
  • More than 19.5: \(25-(4+6)=15\)
  • More than 24.5: \(25-(4+6+10)=5\)

This is especially useful when drawing a cumulative frequency curve.

10. Common mistakes to avoid

  • Mixing up frequency and cumulative frequency. Frequency is just one class count. Cumulative frequency is a running total.
  • Using the wrong class value. For less-than tables, use the upper class value or upper boundary. For more-than tables, use the lower class value or lower boundary.
  • Forgetting the total frequency. Check that the final less-than cumulative frequency equals \(N\).
  • Adding incorrectly. Since cumulative frequency depends on repeated addition, one small error can affect all later values.
  • Not keeping the order of classes. Always write classes from smallest to largest.

11. Quick checking methods

After building your cumulative frequency table, ask yourself:

  • Does the less-than table go upward each time?
  • Does the more-than table go downward each time?
  • Does the last less-than value equal the total frequency?
  • Does the first more-than value equal the total frequency?

If all of these are true, your table is probably correct.

12. Why cumulative frequency distributions matter

Cumulative frequency distributions help us see how data builds up across intervals. This makes it easier to understand the shape of the distribution and to compare groups.

They are also used to:

  • draw cumulative frequency curves
  • estimate the median
  • find quartiles
  • find percentiles

So even though a cumulative frequency table looks simple, it is a very powerful tool in statistics.

Summary

A cumulative frequency distribution is a running total of frequencies from a grouped frequency table. In a less-than cumulative frequency table, add frequencies from top to bottom and use upper class values. In a more-than cumulative frequency table, start with the total frequency and subtract as you move down, using lower class values.

If you remember which class values to use and carefully add or subtract, you can construct cumulative frequency distributions correctly and use them to study large sets of grouped data.

Put what you read to the test

You've worked through Constructing Cumulative Frequency Distributions. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Graphing Cumulative Frequency Curves (Ogives)

Graphing Cumulative Frequency Curves (Ogives)

In statistics, large sets of data are often grouped into class intervals such as 0–10, 10–20, 20–30, and so on. A cumulative frequency curve, also called an ogive, helps us see how the data builds up across these intervals.

An ogive shows the running total of the frequencies. Instead of just showing how many values are in each class, it shows how many values are up to a certain point. This makes it useful for understanding the distribution of data and for estimating values like the median, quartiles, and percentiles.

In this lesson, you will learn what cumulative frequency means, how to calculate it from a grouped frequency table, how to choose the correct points to plot, and how to draw and read an ogive.

1. What is cumulative frequency?

Frequency tells us how many data values fall in each class interval.

Cumulative frequency is the total frequency up to and including a class. It is found by adding frequencies as you go down the table.

For example, suppose we have this grouped data:

Class IntervalFrequency
0–104
10–207
20–305
30–409

The cumulative frequencies are found like this:

  • First class: \(4\)
  • Second class: \(4+7=11\)
  • Third class: \(11+5=16\)
  • Fourth class: \(16+9=25\)

So the cumulative frequency column is:

  • \(4\)
  • \(11\)
  • \(16\)
  • \(25\)

This means:

  • \(4\) values are less than the end of the first class,
  • \(11\) values are less than the end of the second class,
  • \(16\) values are less than the end of the third class,
  • \(25\) values are less than the end of the fourth class.

2. What is an ogive?

An ogive is a graph of cumulative frequency against the upper class boundaries.

The graph usually rises from left to right because cumulative frequency keeps increasing or stays the same. It often has a smooth stretched S-shape, although the exact shape depends on the data.

3. Why do we use class boundaries?

When graphing an ogive, we do not usually plot against the class midpoints. We plot against the class boundaries, especially the upper class boundaries.

For whole-number grouped data, boundaries are often halfway between the class limits. For example:

  • Class interval \(0\text{–}9\) has boundaries \(-0.5\) and \(9.5\)
  • Class interval \(10\text{–}19\) has boundaries \(9.5\) and \(19.5\)
  • Class interval \(20\text{–}29\) has boundaries \(19.5\) and \(29.5\)

If the intervals are already written continuously, such as \(0\text{–}10\), \(10\text{–}20\), \(20\text{–}30\), then the boundaries are usually just \(0, 10, 20, 30\), and so on.

Important idea: the cumulative frequency for a class is plotted at the end of that class, because it represents the total number of data values up to that point.

4. Steps for drawing an ogive

  1. Write the grouped frequency table.
  2. Find the cumulative frequencies by adding the frequencies one by one.
  3. Find the upper class boundaries for each class.
  4. Plot points using: $$ (\text{upper class boundary},\ \text{cumulative frequency}) $$
  5. Include the starting point at the lower boundary of the first class with cumulative frequency \(0\).
  6. Join the points with a smooth increasing curve, not a jagged bar graph.

5. The starting point matters

An ogive should begin at the lowest class boundary with cumulative frequency \(0\). This shows that before the data begins, no values have been counted yet.

For example, if the first class is \(0\text{–}10\), the graph starts at:

$$ (0,0) $$

If the first class is \(10\text{–}20\), the graph starts at:

$$ (10,0) $$

Worked Example 1: Making a cumulative frequency table

The table shows the number of hours students studied in a week.

HoursFrequency
0–53
5–106
10–158
15–204

Step 1: Find cumulative frequency

  • First class: \(3\)
  • Second class: \(3+6=9\)
  • Third class: \(9+8=17\)
  • Fourth class: \(17+4=21\)

So the completed table is:

HoursFrequencyCumulative Frequency
0–533
5–1069
10–15817
15–20421

Step 2: List the points to plot

Use the upper class boundaries:

  • \(5\) with cumulative frequency \(3\)
  • \(10\) with cumulative frequency \(9\)
  • \(15\) with cumulative frequency \(17\)
  • \(20\) with cumulative frequency \(21\)

Also include the starting point:

$$ (0,0) $$

So the points are:

$$ (0,0),\ (5,3),\ (10,9),\ (15,17),\ (20,21) $$

Step 3: Draw the ogive

Plot these points on graph paper. Put hours on the horizontal axis and cumulative frequency on the vertical axis. Then connect the points with a smooth increasing curve.

Worked Example 2: Using class boundaries with whole-number classes

The table shows test scores.

ScoreFrequency
10–192
20–295
30–399
40–497
50–593

Because these classes use whole numbers, we use class boundaries:

  • \(9.5\text{–}19.5\)
  • \(19.5\text{–}29.5\)
  • \(29.5\text{–}39.5\)
  • \(39.5\text{–}49.5\)
  • \(49.5\text{–}59.5\)

Step 1: Find cumulative frequencies

  • \(2\)
  • \(2+5=7\)
  • \(7+9=16\)
  • \(16+7=23\)
  • \(23+3=26\)

Step 2: Plot against upper class boundaries

The upper class boundaries are:

$$ 19.5,\ 29.5,\ 39.5,\ 49.5,\ 59.5 $$

So the plotted points are:

$$ (19.5,2),\ (29.5,7),\ (39.5,16),\ (49.5,23),\ (59.5,26) $$

Include the starting point at the lower boundary of the first class:

$$ (9.5,0) $$

Final set of points:

$$ (9.5,0),\ (19.5,2),\ (29.5,7),\ (39.5,16),\ (49.5,23),\ (59.5,26) $$

After plotting these, join them with a smooth curve.

Worked Example 3: Reading information from an ogive table

A grouped table shows the masses of bags in kilograms.

Mass (kg)Frequency
0–104
10–206
20–3010
30–408
40–502

Step 1: Find cumulative frequencies

  • \(4\)
  • \(4+6=10\)
  • \(10+10=20\)
  • \(20+8=28\)
  • \(28+2=30\)

The total frequency is:

$$ 30 $$

Step 2: Plot points for the ogive

$$ (0,0),\ (10,4),\ (20,10),\ (30,20),\ (40,28),\ (50,30) $$

Step 3: Estimate the median from the ogive

The median is the middle value. For \(30\) data values, the median is at:

$$ \frac{30}{2}=15 $$

So on the ogive, you would find cumulative frequency \(15\) on the vertical axis, move across to the curve, and then move down to the horizontal axis.

From the table, cumulative frequency \(15\) lies between:

  • \(10\) at \(20\) kg
  • \(20\) at \(30\) kg

So the median is somewhere between \(20\) kg and \(30\) kg.

Since \(15\) is halfway between \(10\) and \(20\), the estimate is about halfway between \(20\) and \(30\):

$$ 25\text{ kg} $$

This shows one important use of an ogive: it helps estimate the median and other values from grouped data.

6. How to read an ogive

Once the curve is drawn, you can use it to answer questions.

  • To find how many data values are less than a certain value: start on the horizontal axis, move up to the curve, then move across to the vertical axis.
  • To estimate a data value for a given cumulative frequency: start on the vertical axis, move across to the curve, then move down to the horizontal axis.

This is why an ogive is so useful: it lets you quickly estimate totals and positions in the data.

7. Common mistakes to avoid

  • Using ordinary frequency instead of cumulative frequency. Always add as you go down the table.
  • Plotting against class midpoints. For an ogive, use upper class boundaries.
  • Forgetting the starting point. Begin at the lower boundary of the first class with cumulative frequency \(0\).
  • Drawing straight bars. An ogive is a smooth curve, not a histogram.
  • Using wrong class boundaries. Be careful when intervals are written with whole-number limits like \(10\text{–}19\).

8. Quick check method

After making your cumulative frequency table, do these checks:

  • The cumulative frequency numbers should never go down.
  • The last cumulative frequency should equal the total frequency.
  • The x-values should be class boundaries, usually the upper ones.
  • The first plotted y-value should be \(0\) at the lowest boundary.

Worked Example 4: Full problem from start to finish

The table shows the time, in minutes, students took to finish a quiz.

Time (min)Frequency
0–101
10–204
20–307
30–405
40–503

Step 1: Find cumulative frequencies

  • \(1\)
  • \(1+4=5\)
  • \(5+7=12\)
  • \(12+5=17\)
  • \(17+3=20\)

The completed table is:

Time (min)FrequencyCumulative Frequency
0–1011
10–2045
20–30712
30–40517
40–50320

Step 2: Write the plotting points

Use upper class boundaries:

$$ 10,\ 20,\ 30,\ 40,\ 50 $$

So the points are:

$$ (10,1),\ (20,5),\ (30,12),\ (40,17),\ (50,20) $$

Add the starting point:

$$ (0,0) $$

Step 3: Draw the graph

Plot the points and join them smoothly.

Step 4: Answer a question from the ogive

How many students took less than \(35\) minutes?

On the graph, start at \(35\) on the horizontal axis, move up to the curve, and then across to the cumulative frequency axis.

Since \(35\) is halfway between \(30\) and \(40\), and the cumulative frequency rises from \(12\) to \(17\), the estimate is about:

$$ 14.5 $$

So about 14 or 15 students took less than \(35\) minutes.

9. When an ogive is especially helpful

  • When data is grouped into intervals
  • When you want to see how the total builds up
  • When estimating the median or quartiles
  • When comparing how quickly frequencies accumulate

10. Summary

A cumulative frequency curve, or ogive, is a graph that shows how frequencies add up across class intervals. To draw one, first calculate cumulative frequencies, then plot them against the upper class boundaries, starting with cumulative frequency \(0\) at the first lower boundary.

The curve should rise from left to right and is usually drawn smoothly. Ogives help you understand the distribution of grouped data and estimate important values such as how many observations are below a given number or where the median is located.

Put what you read to the test

You've worked through Graphing Cumulative Frequency Curves (Ogives). Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.

Determining the Median Graphically

Determining the Median Graphically

In statistics, the median is the middle value of a data set when the values are arranged in order. It divides the data into two equal parts: half the values are below it and half are above it.

When data is grouped into class intervals, it is often not possible to find the exact middle value just by looking at the table. In this case, we can estimate the median graphically using a cumulative frequency curve, also called an ogive.

This lesson will show you how to find the median from an ogive step by step.

1. What is an ogive?

An ogive is a graph of cumulative frequency against the upper class boundaries of grouped data.

Cumulative frequency means the running total of frequencies. As you move through the classes, you keep adding the frequencies.

For example, if the frequencies are 3, 5, 7, and 4, then the cumulative frequencies are:

  • \(3\)

  • \(3+5=8\)

  • \(8+7=15\)

  • \(15+4=19\)

So the cumulative frequencies are \(3, 8, 15, 19\).

2. Why does the median come from \(N/2\)?

If there are \(N\) total data values, then the median is the value that has half the data below it. That means we look for the \(\frac{N}{2}\)-th value.

On an ogive, the vertical axis shows cumulative frequency. So to find the median, we:

  1. Find the total frequency \(N\).

  2. Calculate \(\frac{N}{2}\).

  3. Mark \(\frac{N}{2}\) on the cumulative frequency axis.

  4. Draw a horizontal line to the ogive.

  5. From the point where it meets the curve, draw a vertical line down to the horizontal axis.

  6. Read the value on the horizontal axis. That value is the median.

This works because the ogive tells us how many data values are less than or equal to each value. The median is the point where half the data has been counted.

3. Steps for determining the median graphically

Use these steps every time:

  1. Make or read the grouped frequency table.

  2. Find the cumulative frequencies.

  3. Plot the ogive using upper class boundaries on the x-axis and cumulative frequencies on the y-axis.

  4. Find the total frequency, \(N\).

  5. Calculate \(\frac{N}{2}\).

  6. Locate \(\frac{N}{2}\) on the y-axis.

  7. Draw a horizontal line to meet the curve.

  8. From that point, drop a perpendicular to the x-axis.

  9. Read the x-value. This is the estimated median.

4. Important note about estimation

When you find the median from a graph, the answer is usually an estimate. This is because:

  • the data is grouped, not exact,

  • the curve is drawn by hand or read from a graph,

  • the intersection may fall between marked values.

So it is normal for the graph to give an approximate answer.

Worked Example 1: Reading the median from an ogive

A grouped data set has total frequency \(N=40\). An ogive has already been drawn. Find the median graphically.

Step 1: Find \(\frac{N}{2}\)

$$\frac{N}{2}=\frac{40}{2}=20$$

So we need the 20th value.

Step 2: Locate 20 on the y-axis

Find \(20\) on the cumulative frequency axis.

Step 3: Draw across to the curve

Draw a horizontal line from \(20\) until it touches the ogive.

Step 4: Drop down to the x-axis

From the meeting point, draw a vertical line down to the x-axis.

Suppose the line meets the x-axis at \(27\).

Answer: The median is about \(27\).

Worked Example 2: From a table to the median graphically

The table shows the marks scored by students in a test.

MarksFrequency
0–104
10–206
20–3010
30–408
40–502

Step 1: Find cumulative frequencies

MarksFrequencyCumulative Frequency
0–1044
10–20610
20–301020
30–40828
40–50230

Step 2: Find the total frequency

The total frequency is:

$$N=30$$

Step 3: Calculate \(\frac{N}{2}\)

$$\frac{N}{2}=\frac{30}{2}=15$$

So we must find the 15th value on the ogive.

Step 4: Plot the ogive

Plot cumulative frequency against the upper class boundaries:

  • \((10,4)\)

  • \((20,10)\)

  • \((30,20)\)

  • \((40,28)\)

  • \((50,30)\)

Then join the points with a smooth increasing curve.

Step 5: Read the median from the graph

Locate \(15\) on the y-axis, draw a horizontal line to the curve, and then drop a vertical line to the x-axis.

Since \(15\) lies between cumulative frequencies \(10\) and \(20\), the median will lie between \(20\) and \(30\) marks.

Suppose the graph gives approximately \(25\).

Answer: The estimated median mark is \(25\).

Worked Example 3: A larger grouped data set

The table shows the times, in minutes, that students spent reading.

Time (min)Frequency
0–53
5–107
10–1512
15–209
20–255
25–304

Step 1: Find cumulative frequencies

Time (min)FrequencyCumulative Frequency
0–533
5–10710
10–151222
15–20931
20–25536
25–30440

Step 2: Find total frequency

$$N=40$$

Step 3: Calculate half of the total

$$\frac{N}{2}=\frac{40}{2}=20$$

Step 4: Use the graph

On the cumulative frequency axis, mark \(20\). Draw across to the curve, then down to the time axis.

The 20th value lies in the class interval \(10\)–\(15\), because the cumulative frequency rises from \(10\) to \(22\) in that class.

Suppose the graph shows the median as about \(14\) minutes.

Answer: The estimated median reading time is \(14\) minutes.

5. How to know if your answer makes sense

After finding the median graphically, check whether your answer is reasonable.

  • The median should lie within the range of the data.

  • It should be near the middle of the distribution, not at an extreme end.

  • It should match the class interval where the \(\frac{N}{2}\)-th value falls.

For example, if \(\frac{N}{2}=15\) and the cumulative frequencies go from \(10\) to \(20\) in the class \(20\)–\(30\), then the median must be somewhere between \(20\) and \(30\).

6. Common mistakes to avoid

  • Using the frequency instead of the cumulative frequency. The graph must use cumulative frequency on the y-axis.

  • Forgetting to divide by 2. Always find \(\frac{N}{2}\), not just \(N\).

  • Reading the wrong axis. Start at \(\frac{N}{2}\) on the y-axis, go to the curve, then down to the x-axis.

  • Using class intervals instead of boundaries incorrectly. On an ogive, points are usually plotted at the upper class boundaries.

  • Expecting an exact whole number every time. Graphical answers are often estimates, such as \(24.5\) or \(13.8\).

7. Quick method to remember

You can remember the process like this:

Half total → across → curve → down → median

  • Find half the total frequency

  • Go across from the y-axis

  • Meet the ogive curve

  • Go down to the x-axis

  • Read the median value

8. Practice question

The frequency table below shows the heights of plants.

Height (cm)Frequency
0–102
10–205
20–309
30–407
40–503

Try these steps:

  1. Find the cumulative frequencies.

  2. Find the total frequency.

  3. Calculate \(\frac{N}{2}\).

  4. Draw or imagine the ogive.

  5. Use \(\frac{N}{2}\) to estimate the median graphically.

Answer check:

The cumulative frequencies are \(2, 7, 16, 23, 26\), so \(N=26\) and:

$$\frac{N}{2}=13$$

The 13th value lies in the \(20\)–\(30\) class, so the median should be between \(20\) cm and \(30\) cm. A graph would give an estimate somewhere in that interval.

Summary

To determine the median graphically, use the ogive and the total frequency \(N\). Find \(\frac{N}{2}\) on the cumulative frequency axis, draw across to the curve, and then drop a perpendicular to the x-axis. The x-value you read is the estimated median.

This method is especially useful for grouped data, where the exact middle value is not directly shown in the table.

Put what you read to the test

You've worked through Determining the Median Graphically. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.