Grouped vs. Ungrouped Data Characteristics
Grouped vs. Ungrouped Data Characteristics
In statistics, data can be shown in different ways depending on how much information we want to keep and how large the data set is. Two important forms are ungrouped data and grouped data.
Understanding the difference between these two forms is important because it affects how we read the data, organize it, and calculate values such as averages and frequencies.
This lesson will help you learn how to tell grouped data from ungrouped data, how class intervals work, and how to identify class marks and class boundaries.
1. What is ungrouped data?
Ungrouped data is data listed as individual values. Each observation is shown separately, so we can see the exact data points.
For example, suppose 8 students scored the following marks on a quiz:
12, 15, 17, 15, 19, 14, 18, 15
This is ungrouped data because every score is written on its own, not collected into intervals.
Characteristics of ungrouped data:
- Each value is shown clearly.
- It is best for small data sets.
- Exact values are known.
- It is easier to find the precise minimum, maximum, and repeated values.
2. What is grouped data?
Grouped data is data that has been organized into class intervals. Instead of listing every individual value, we place values into groups.
For example, test marks for a large class might be shown like this:
- \(0\text{–}10\): 3 students
- \(10\text{–}20\): 7 students
- \(20\text{–}30\): 12 students
- \(30\text{–}40\): 8 students
This is grouped data because the marks are collected into intervals rather than shown one by one.
Characteristics of grouped data:
- Data is organized into classes or intervals.
- It is useful for large data sets.
- It makes patterns easier to see.
- Exact individual values are not shown.
- Some detail is lost when data is grouped.
3. Main difference between grouped and ungrouped data
The biggest difference is that ungrouped data shows exact individual values, while grouped data summarizes values into intervals.
- Ungrouped data keeps all the original information.
- Grouped data gives a simpler summary of the data.
So, if you need exact values, ungrouped data is better. If you need to study a large set of values and look for trends, grouped data is more useful.
4. Discrete data points and continuous class intervals
Ungrouped data often appears as raw discrete data points. This means the values are listed one at a time.
For example, number of books read by students in a month could be:
2, 4, 1, 3, 2, 5, 4
These are separate values, so this is discrete and ungrouped.
Grouped data is often shown using continuous class intervals. These intervals cover a range of values.
For example, heights of students might be grouped as:
- 140–149 cm
- 150–159 cm
- 160–169 cm
These are intervals, so this is grouped data.
5. What is a class interval?
A class interval is a range of values used to group data. Each group is called a class.
In the interval \(20\text{–}29\), the lower class limit is \(20\) and the upper class limit is \(29\).
If data is grouped as:
- \(0\text{–}9\)
- \(10\text{–}19\)
- \(20\text{–}29\)
then each class interval has width \(10\).
Class width can be found by looking at the size of each interval. For equal intervals, the width stays the same throughout the table.
6. What is a class mark?
The class mark is the midpoint of a class interval. It is found by averaging the lower and upper class limits.
The formula is:
$$\text{Class mark} = \frac{\text{lower class limit} + \text{upper class limit}}{2}$$
For the class interval \(20\text{–}29\):
$$\text{Class mark} = \frac{20 + 29}{2} = \frac{49}{2} = 24.5$$
The class mark is used as a representative value for all the data in that class.
7. What are class boundaries?
Class boundaries are the true edges of a class interval when data is measured continuously.
They are especially helpful when there is no gap between classes.
For example, consider the classes:
- \(10\text{–}19\)
- \(20\text{–}29\)
If the data is measured to the nearest whole number, then the boundaries are found by subtracting \(0.5\) from the lower limit and adding \(0.5\) to the upper limit.
So:
- \(10\text{–}19\) becomes \(9.5\text{–}19.5\)
- \(20\text{–}29\) becomes \(19.5\text{–}29.5\)
Notice that the upper boundary of one class touches the lower boundary of the next class. This makes the grouped data continuous.
8. Why class boundaries matter
Class boundaries are important when drawing graphs such as histograms and cumulative frequency curves. These graphs need continuous intervals with no gaps.
If we only use class limits, it may look like there is a break between classes. Boundaries solve this problem.
9. Worked Example 1: Identify grouped and ungrouped data
Decide whether each set of data is grouped or ungrouped.
- \(5, 8, 6, 10, 7, 8\)
-
- \(0\text{–}4\): 2
- \(5\text{–}9\): 6
- \(10\text{–}14\): 3
Solution:
1. \(5, 8, 6, 10, 7, 8\) is ungrouped data because each value is listed separately.
2. The second set is grouped data because the values are organized into class intervals with frequencies.
Worked Example 2: Find class marks
Find the class mark for each interval:
- \(10\text{–}19\)
- \(20\text{–}29\)
- \(30\text{–}39\)
Solution:
Use:
$$\text{Class mark} = \frac{\text{lower limit} + \text{upper limit}}{2}$$
For \(10\text{–}19\):
$$\frac{10+19}{2} = \frac{29}{2} = 14.5$$
For \(20\text{–}29\):
$$\frac{20+29}{2} = \frac{49}{2} = 24.5$$
For \(30\text{–}39\):
$$\frac{30+39}{2} = \frac{69}{2} = 34.5$$
Answer: The class marks are \(14.5\), \(24.5\), and \(34.5\).
Worked Example 3: Find class boundaries
The class intervals are:
- \(40\text{–}49\)
- \(50\text{–}59\)
- \(60\text{–}69\)
Find the class boundaries if the data is measured to the nearest whole number.
Solution:
For whole-number data, subtract \(0.5\) from each lower limit and add \(0.5\) to each upper limit.
- \(40\text{–}49\) becomes \(39.5\text{–}49.5\)
- \(50\text{–}59\) becomes \(49.5\text{–}59.5\)
- \(60\text{–}69\) becomes \(59.5\text{–}69.5\)
Answer: The class boundaries are:
\(39.5\text{–}49.5\), \(49.5\text{–}59.5\), and \(59.5\text{–}69.5\).
Worked Example 4: Describe the data characteristics
A teacher records the ages of students as:
14, 15, 14, 16, 15, 15, 14, 16, 17, 15
Later, the teacher summarizes them in this table:
- \(14\text{–}14\): 3
- \(15\text{–}15\): 4
- \(16\text{–}16\): 2
- \(17\text{–}17\): 1
Question: Which form is ungrouped and which form is grouped? What is lost when the data is grouped?
Solution:
The first list is ungrouped data because every age is written individually.
The table is grouped data because the ages are organized into classes with frequencies.
When the data is grouped, we lose the original order of the values. We can still see how many students are each age, but we no longer see the exact sequence in which the ages were recorded.
10. Key ideas to remember
- Ungrouped data shows each value separately.
- Grouped data summarizes values into class intervals.
- A class interval is a range used to group data.
- A class mark is the midpoint of a class interval.
- Class boundaries show the true edges of each class for continuous data.
- Grouped data is easier to use for large sets, but it does not show every exact value.
11. Brief summary
Ungrouped data lists raw data values one by one, while grouped data collects them into intervals. Grouped data is useful for large data sets because it makes patterns easier to see, but it loses some exact detail.
When working with grouped data, you should be able to identify the class interval, class mark, and class boundaries. These ideas are important for tables, histograms, and cumulative frequency curves.
Put what you read to the test
You've worked through Grouped vs. Ungrouped Data Characteristics. Try answering a few questions to see what stuck — and what might deserve a quick reread before you move on.