How Do You Draw A Cumulative Frequency Graph

I remember my first real encounter with a cumulative frequency graph. I was in, I think, Year 10, and my teacher, Mrs. Davison (she had a penchant for floral cardigans and terrifyingly precise chalk dust control), unveiled this… thing. It looked like a slide that had melted down a hill, all curvy and upwardly mobile. My initial thought was, "What in the statistical Sam Hill is that?" We'd been doing bar charts and pie charts, which made sense. They showed you how many or what proportion. But this? This seemed to be showing you… how much has happened so far.
It was like Mrs. Davison was trying to explain the concept of "total accumulated goodness" or "cumulative doom" depending on the data set. Honestly, my brain felt a bit like it was trying to digest a particularly dense broccoli floret. But then, slowly, like a shy snail emerging from its shell, the penny dropped. And that’s what we’re going to dive into today, because trust me, once you get it, cumulative frequency graphs are actually pretty darn cool. And, dare I say it, useful.
So, picture this: you've got a bunch of data. Let's say it's the heights of all the students in your school. You've measured them all, and now you've got a giant list of numbers. You could make a frequency table, showing how many students are between 150cm and 155cm, how many are between 155cm and 160cm, and so on. That's all well and good. It tells you where the most students fall.
But what if you want to know something a little different? What if you want to know, for example, how many students are shorter than 160cm? A regular frequency table won't tell you that directly. You'd have to go back, add up all the counts for the groups below 160cm. Tedious, right? This is where our curvy friend, the cumulative frequency graph, swoops in to save the day. It’s essentially a visual shortcut to answering questions like, "How many are at or below this point?"
The "So Far" Storyteller
The core idea behind a cumulative frequency graph is that it shows you the total number of observations that fall below a certain value. It's like looking at your bank account balance at the end of each day. On day one, you have $X$. On day two, you have $X + Y$. On day three, you have $X + Y + Z$. Each day's balance is the cumulative total of all transactions up to that point. Our graph does the same thing, but for data points.
Think of it as a running tally. You’re not just counting what happened in one specific interval; you’re adding up all the counts from the beginning of your data range up to the upper limit of your current interval. It’s a historical record of your data, building up as you go along.
Let's Get Our Hands Dirty (Figuratively, of Course)
Alright, enough preamble. How do we actually draw one of these beasts? We need a set of data, and we need to organize it. Let’s stick with our hypothetical school heights. Imagine we’ve collected the data and put it into a frequency table, grouped into intervals. Something like this:
Height (cm) | Frequency (Number of students) ----------- | ----------------------------- 150 - 155 | 25 155 - 160 | 40 160 - 165 | 60 165 - 170 | 35 170 - 175 | 15
This is your standard, everyday frequency table. Useful, but not quite our cumulative beauty yet. To get there, we need to add a crucial column: the cumulative frequency. This is where the magic happens.
The cumulative frequency for an interval is the sum of the frequencies of that interval and all the intervals that come before it. Let's build that out:

Height (cm) | Frequency | Cumulative Frequency ----------- | --------- | -------------------- 150 - 155 | 25 | 25 (Just the first group) 155 - 160 | 40 | 25 + 40 = 65 (All up to 160cm) 160 - 165 | 60 | 65 + 60 = 125 (All up to 165cm) 165 - 170 | 35 | 125 + 35 = 160 (All up to 170cm) 170 - 175 | 15 | 160 + 15 = 175 (All up to 175cm)
See what's happening? The cumulative frequency for the 155-160 group is the frequency of the 150-155 group plus the frequency of the 155-160 group. And it just keeps adding up. By the time we get to the end, our cumulative frequency (175) should equal the total number of observations (which it does – 25 + 40 + 60 + 35 + 15 = 175. Phew, no errors!). This is a good sanity check.
Now, here's a slight twist. When we plot a cumulative frequency graph, we don't plot the cumulative frequency against the midpoint of the interval, like you might with some other graphs. Instead, we plot it against the upper boundary of each interval. Why? Because the cumulative frequency represents all the data up to that upper boundary. So, for our example:
- 25 students are shorter than or equal to 155cm.
- 65 students are shorter than or equal to 160cm.
- 125 students are shorter than or equal to 165cm.
- 160 students are shorter than or equal to 170cm.
- 175 students are shorter than or equal to 175cm.
So, our plotting points will be (155, 25), (160, 65), (165, 125), (170, 160), and (175, 175).
There’s one more point that's implied and often included to give the graph a proper start. Since no one can be shorter than the lower boundary of the very first interval (which is 150cm in our example), we can add a point (150, 0). This represents that zero students have a height of 150cm or less based on our grouped data. It anchors the graph nicely.
Plotting the Curves
Okay, we’ve got our data points. Now, let’s talk axes. On the horizontal axis (the x-axis, the one that goes across), you’ll put your variable. In our case, that's 'Height (cm)'. You need to make sure your scale starts at or before your lowest boundary (150cm) and goes up to at least your highest boundary (175cm). Label it clearly!
On the vertical axis (the y-axis, the one that goes up), you’ll put your cumulative frequency. Again, make sure your scale starts at 0 and goes up to at least your total number of observations (175). Label this one clearly too! Something like 'Cumulative Frequency (Number of Students)'.

Once your axes are drawn and labelled, it’s time to plot those points we figured out: (150, 0), (155, 25), (160, 65), (165, 125), (170, 160), (175, 175).
Now for the drawing part. This is where it differs from a simple scatter plot. You connect these points with a smooth, upward-sloping curve. It’s called an ogive (pronounced OH-jive, which sounds fancy, doesn’t it?). It should start at the initial point (150, 0) and gradually rise, leveling off at the very last point (175, 175). It should never go down. If it goes down, you’ve made a mistake somewhere, and frankly, the statistical gods will frown upon you.
Why a smooth curve? Because we’re assuming that the data within each interval is spread out somewhat evenly. We’re not jumping from 25 to 65 in a single, brutal step at exactly 160cm. The curve suggests a more gradual increase in the number of students as the height increases.
Sometimes, especially if you have a lot of data points, the curve can look a bit jagged. But generally, you aim for a nice, flowing line. Think of it as tracing the outline of a gentle hill.
What Can You Do With This Curvy Thing?
So, you’ve drawn it. Yay! But what’s the point? This is where the real utility comes in. A cumulative frequency graph is fantastic for estimating things like:
The Median
The median is the middle value when your data is ordered. In a cumulative frequency graph, it's the value of the variable (height, in our case) when the cumulative frequency is half of the total number of observations.

Our total number of students is 175. Half of 175 is 87.5. So, we go to the y-axis, find 87.5, draw a horizontal line across to where it hits our curve, and then draw a vertical line straight down to the x-axis. The value on the x-axis where this line lands is our estimated median height.
It might not be an exact number (especially if you didn't have data points right on 87.5), but it gives you a very good approximation of the central tendency. It tells you that roughly half the students are shorter than this height, and half are taller.
Quartiles and Percentiles
This is where it gets really powerful. Quartiles divide your data into four equal parts. The first quartile (Q1) is the value below which 25% of the data lies. The third quartile (Q3) is the value below which 75% of the data lies.
To find Q1, you look for 25% of your total frequency (0.25 * 175 = 43.75) on the y-axis, go across to the curve, and down to the x-axis. To find Q3, you look for 75% of your total frequency (0.75 * 175 = 131.25) on the y-axis, go across to the curve, and down to the x-axis.
You can also find any percentile. Want to know the height below which 90% of students fall? Look for 90% of 175 on the y-axis, and follow the same process.
These values give you a much richer picture of your data distribution than just the average (mean) or median alone. They tell you about the spread and skewness.

The Interquartile Range (IQR)
Once you’ve found Q1 and Q3, you can easily calculate the IQR (IQR = Q3 - Q1). This is a measure of the spread of the middle 50% of your data. It's less affected by extreme values (outliers) than the range (maximum - minimum), making it a robust measure of dispersion.
Estimating the Number of Observations within a Range
Remember our initial problem? How many students are shorter than 160cm? With the graph, it's a doddle! Find 160cm on the x-axis, go straight up to the curve, and then straight across to the y-axis. You’ll see it’s 65 students. Easy peasy.
What about students between 160cm and 170cm? Well, we know there are 125 students shorter than or equal to 165cm and 160 students shorter than or equal to 170cm. So, the number of students between 165cm and 170cm is 160 - 125 = 35. (Wait, that's not right! That's the frequency in that band. We want the cumulative. Okay, let's rephrase for clarity). To find the number of students between 160cm and 170cm, we find the cumulative frequency at 170cm (160) and subtract the cumulative frequency at 160cm (65). So, 160 - 65 = 95 students. See? You just use the cumulative frequencies of the upper and lower bounds of your desired range.
Common Pitfalls and How to Avoid Them
So, it’s not rocket science, but there are a few places where people tend to stumble:
- Confusing upper boundaries with midpoints: Remember, for cumulative frequency, it's always the upper boundary of the interval you plot against.
- Incorrectly calculating cumulative frequency: Double-check your additions. A small error here will mess up your whole graph. The final cumulative frequency must equal the total number of observations.
- Drawing a jagged line instead of a smooth curve: Unless your data is incredibly discrete and step-like, a smooth curve is generally preferred.
- Forgetting to label axes or provide a title: A graph without labels is like a sentence without words. Utterly useless!
- Not starting the graph at (0, lower boundary): This often makes the initial part of the distribution look misleading.
Don't be discouraged if your first attempt isn't perfect. The more you practice, the more intuitive it becomes. It’s like learning to ride a bike; a few wobbles are expected.
Cumulative frequency graphs, or ogives, are more than just pretty curves. They are powerful tools for understanding the distribution of your data, estimating key statistical measures, and answering questions about "how much of it all" has happened up to a certain point. So, next time you're faced with a pile of data, give the cumulative frequency graph a go. You might just find it’s a lot less daunting and a lot more illuminating than you first thought. And hey, you’ll have a fancy new statistical skill to brag about at your next (or perhaps a more statistically inclined) gathering. Happy graphing!
