Statistics - Calculation of mean, median, and mode for grouped frequency distributions using direct, assumed mean, and step-deviation methods
In statistical analysis, raw data collected from field observations is often voluminous, unorganized, and difficult to interpret directly. To extract meaningful insights, data is aggregated into grouped frequency distributions. However, analyzing a large table of class intervals and frequencies still requires a single numerical value that represents the central point or typical behavior of the entire dataset. This representative value is known as a measure of central tendency.
The three primary measures of central tendency studied in Class 10 Mathematics are:
- Mean: The mathematical average of all observations.
- Median: The middle-most value when data is arranged in order of magnitude.
- Mode: The value that occurs most frequently in the distribution.
Mastering the calculation of these three measures for grouped data—and understanding when to apply specific computational short-cuts such as the Assumed Mean and Step-Deviation methods—is essential for CBSE Class 10 Mathematics board examinations and forms the foundation for advanced data analysis, economics, and modern data science.
1. Fundamental Concepts & Terminology
Before calculating central tendencies, we must understand how grouped frequency tables are structured.
Class Intervals & Continuity
Data is organized into continuous ranges called class intervals.
- Exclusive Form (Continuous Data): . Here, the upper limit of one class equals the lower limit of the next class. An observation with value is included in , not .
- Inclusive Form (Discontinuous Data): . If data is given in this form, it must be converted to continuous form before calculating median and mode by subtracting from lower limits and adding to upper limits: .
Key Definitions
-
Class Limits: The minimum and maximum boundary values of a class interval.
-
Class Mark (): The midpoint of a class interval, which acts as the single representative value for all data points falling within that interval.
-
Class Size (): The difference between the upper limit and lower limit of a continuous class interval.
-
Frequency (): The number of observations falling within the class interval. The sum of all frequencies is denoted as or .
2. Calculation of Mean () of Grouped Data
The mean represents the arithmetic average of the dataset. For grouped data, we assume that the frequency of each class interval is centered at its class mark ().
Depending on the numerical magnitude of and , we use one of three computational methods:
┌─────────────────────────────────────────┐ │ Methods for Mean of Grouped Data │ └────────────────────┬────────────────────┘ │ ┌──────────────────────────┼──────────────────────────┐ │ │ │ ┌────────┴────────┐ ┌────────┴────────┐ ┌────────┴────────┐ │ Direct Method │ │ Assumed Mean │ │ Step-Deviation │ │ │ │ Method │ │ Method │ └─────────────────┘ └─────────────────┘ └─────────────────┘
Method 1: Direct Method
The Direct Method is applied when the values of class marks () and frequencies () are numerically small, allowing direct multiplication without risk of calculation errors.
Formula
Algorithm
- Construct a table with columns: Class Interval, Frequency (), Class Mark (), and Product ().
- Compute for each class interval.
- Multiply each by its corresponding to find .
- Sum all frequencies to get and sum all products to get .
- Divide by .
Method 2: Assumed Mean Method
When values are large, direct multiplication becomes tedious. The Assumed Mean Method reduces numerical size by shifting the origin. We select an arbitrary central class mark as the Assumed Mean () and compute deviations from it.
Formula
Where:
- (usually the middle class mark )
- (deviation of each class mark from )
Algorithm
- Calculate class marks () for all intervals.
- Choose one central value from as the assumed mean .
- Compute deviation for each class. (Deviations above will be negative; deviations below will be positive; deviation for itself is ).
- Compute the product for each row and sum them to obtain .
- Apply the formula: .
Method 3: Step-Deviation Method
When class sizes () are equal and the deviations share a common factor , we can further simplify calculations by scaling down the deviations.
Formula
Where:
Algorithm
- Compute class marks (), select assumed mean , and identify class size .
- Calculate step-deviations for each class interval. (For equal class intervals, will always follow a sequence like ).
- Compute for each row and find the sum .
- Multiply by and add .
3. Calculation of Mode of Grouped Data
The Mode is the value among observations that has the maximum frequency. In a grouped frequency distribution, we cannot determine the exact mode simply by looking at frequencies. Instead, we first locate the Modal Class.
Modal Class
The class interval corresponding to the maximum frequency () is called the modal class. The mode lies somewhere within this interval.
Formula
Where:
Important Note: Ensure class intervals are continuous before finding .
4. Calculation of Median of Grouped Data
The Median is the measure of central tendency that gives the value of the middle-most observation in the data.
Cumulative Frequency ()
To find the median, we construct a Cumulative Frequency Table of the 'less than' type by successively adding individual frequencies.
Median Class
- Compute total frequency .
- Find .
- Locate the class interval whose cumulative frequency () is just greater than or equal to . This interval is the Median Class.
Formula
Where:
5. Empirical Relationship Among Central Tendencies
For a moderately skewed distribution (a standard real-world data distribution), an empirical relationship exists between Mean, Median, and Mode:
Or rearranged:
This relationship allows us to estimate any one measure if the other two are known.
6. Summary Comparison Table
| Feature | Mean () | Median | Mode |
|---|---|---|---|
| Core Definition | Arithmetic average of all data values | Middle-most observation of data | Most frequently occurring value |
| Sensitivity to Outliers | Highly sensitive to extreme values | Unaffected by extreme values | Unaffected by extreme values |
| Prerequisites | Class marks () | Cumulative frequency () | Identification of max frequency () |
| Key Variable Needed | , , or | and preceding | |
| Best Used For | Symmetric data without extreme outliers | Skewed distributions (e.g., income, house prices) | Categorical data or finding most popular item |
7. Real-World Applications
1. Income Analysis & Economics (Why Median Matters)
Consider a small software startup with 10 employees where 9 junior developers earn ₹5,000/month and 1 CEO earns ₹5,00,000/month.
- Mean Salary: . This paints a misleading picture that employees earn well.
- Median Salary: ₹5,000. This provides a realistic representation of a typical worker's wage. Economists use the median rather than the mean to measure household income and national wealth distribution.
2. Retail, Manufacturing & Inventory Management (Why Mode Matters)
A shoe store manager wants to order new stock. If shoe sizes sold are measured:
- Sizes sold: .
- Mean size = . Ordering size shoes is impossible.
- Mode size = 8. The store stocks maximum inventory of size 8 because it is the most requested item.
3. School Grading Systems (Why Mean Matters)
A teacher evaluates a class of 40 students on a standardized test. The mean score reveals overall class mastery, helping administrators determine if the teaching methodology was effective across the entire group.
8. Step-by-Step Solved Examples
Example 1: Comprehensive Calculation of Mean using All Three Methods
Problem: The following table gives the distribution of marks obtained by 30 students in a mathematics test. Find the mean marks using:
- Direct Method
- Assumed Mean Method
- Step-Deviation Method
| Class Interval (Marks) | Number of Students () |
|---|---|
| 10 – 25 | 2 |
| 25 – 40 | 3 |
| 40 – 55 | 7 |
| 55 – 70 | 6 |
| 70 – 85 | 6 |
| 85 – 100 | 6 |
Solution:
Step 1: Construct the Master Working Table
Let us choose Assumed Mean (the midpoint of the class marks) and .
| Class Interval | Frequency () | Class Mark () | |||||
|---|---|---|---|---|---|---|---|
| 10 – 25 | 2 | 17.5 | 35.0 | -45.0 | -90 | -3 | -6 |
| 25 – 40 | 3 | 32.5 | 97.5 | -30.0 | -90 | -2 | -6 |
| 40 – 55 | 7 | 47.5 | 332.5 | -15.0 | -105 | -1 | -7 |
| 55 – 70 | 6 | 62.5 () | 375.0 | 0.0 | 0 | 0 | 0 |
| 70 – 85 | 6 | 77.5 | 465.0 | 15.0 | 90 | 1 | 6 |
| 85 – 100 | 6 | 92.5 | 555.0 | 30.0 | 180 | 2 | 12 |
| Total |
1. By Direct Method:
2. By Assumed Mean Method:
3. By Step-Deviation Method:
Conclusion: All three methods yield the exact same mean score of .
Example 2: Finding the Mode of Grouped Data
Problem: The following table shows the age distribution of patients admitted to a hospital during a year. Find the mode of the data.
| Age (in years) | 5 – 15 | 15 – 25 | 25 – 35 | 35 – 45 | 45 – 55 | 55 – 65 |
|---|---|---|---|---|---|---|
| Number of patients | 6 | 11 | 21 | 23 | 14 | 5 |
Solution:
-
Identify the Modal Class: Look at the frequencies: . The maximum frequency is , which corresponds to the class interval . Therefore, Modal Class = .
-
Extract Parameters:
- Lower limit of modal class () =
- Class size () =
- Frequency of modal class () =
- Frequency of class preceding modal class () =
- Frequency of class succeeding modal class () =
-
Apply the Formula:
Final Answer: The mode of the data is .
Example 3: Finding Missing Frequencies using Median
Problem: The median of the following data distribution is . Find the values of missing frequencies and , given that the total frequency is .
| Class Interval | 0 – 10 | 10 – 20 | 20 – 30 | 30 – 40 | 40 – 50 | 50 – 60 | Total |
|---|---|---|---|---|---|---|---|
| Frequency | 5 | 20 | 15 | 5 | 60 |
Solution:
Step 1: Construct the Cumulative Frequency Table
| Class Interval | Frequency () | Cumulative Frequency () |
|---|---|---|
| 0 – 10 | 5 | 5 |
| 10 – 20 | ||
| 20 – 30 | 20 | |
| 30 – 40 | 15 | |
| 40 – 50 | ||
| 50 – 60 | 5 | |
| Total |
Step 2: Formulate Equation (1) from Total Frequency
Step 3: Identify Median Class Given . Since lies within the range , the Median Class is .
Step 4: Extract Parameters
- Lower limit () =
- Class size () =
- Frequency of median class () =
- Cumulative frequency of class preceding median class () =
- Total observations () =
Step 5: Apply Median Formula
Step 6: Substitute into Equation 1
Final Answer: The missing frequencies are and .
9. Common Student Mistakes to Avoid
1. Wrong Cumulative Frequency () in Median Formula
- Mistake: Using the cumulative frequency of the median class itself instead of the class preceding it.
- Correction: Always select from the row above the median class.
2. Ignoring Discontinuous Class Intervals
- Mistake: Calculating lower limits () directly from discontinuous boundaries like without converting them to .
- Correction: If data is given in inclusive form, convert it to continuous form first. For , is , not .
3. Choosing Class Mark () as Modal Class or Median Class Value
- Mistake: Identifying value with the highest frequency and writing that as the Mode.
- Correction: Use the grouped mode formula. Mode is not simply a class mark.
4. Algebraic Sign Errors in Step-Deviation
- Mistake: Reversing the signs of (e.g., making values above assumed mean positive and below negative).
- Correction: Values of smaller than give negative and . Values greater than give positive values.
10. Practice Questions for Self-Assessment
Practice Question 1
Find the mean of the following distribution using the Step-Deviation Method:
| Class | 100 – 150 | 150 – 200 | 200 – 250 | 250 – 300 | 300 – 350 |
|---|---|---|---|---|---|
| Frequency | 4 | 5 | 12 | 2 | 2 |
Solution:
- Class marks (): .
- Choose Assumed Mean , class size .
- Calculate : .
- column:
- Sum of frequencies .
- Sum of products .
- Apply formula:
Answer: Mean = .
Practice Question 2
Calculate the Mode for the following frequency distribution:
| Marks | 0 – 10 | 10 – 20 | 20 – 30 | 30 – 40 | 40 – 50 |
|---|---|---|---|---|---|
| Number of Students | 5 | 8 | 15 | 10 | 2 |
Solution:
- Maximum frequency corresponds to interval .
- Modal class = .
- .
- Apply formula:
Answer: Mode = .
Practice Question 3
Compute the Median for the following distribution:
| Class Interval | 50 – 60 | 60 – 70 | 70 – 80 | 80 – 90 | 90 – 100 |
|---|---|---|---|---|---|
| Frequency | 3 | 7 | 11 | 6 | 3 |
Solution:
- Compute Cumulative Frequencies ():
- :
- :
- :
- :
- :
- Total .
- just greater than is , corresponding to Median Class .
- .
- Apply formula:
Answer: Median = .
11. Exam Revision & Frequently Asked Questions (FAQs)
Q1: What is the empirical relationship between Mean, Median, and Mode, and when is it used?
Answer: The empirical relationship is: It is used to estimate one measure of central tendency when two measures are already calculated, or in 1-mark board exam questions where two values are provided directly.
Q2: Can we apply the Step-Deviation method if class sizes () are unequal?
Answer: Strictly speaking, no. The standard Step-Deviation method assumes a constant class size across all intervals so that results in integers. If class sizes are unequal, use either the Direct Method or the Assumed Mean Method.
Q3: Why do all three methods for calculating the mean give the exact same answer?
Answer: The Assumed Mean and Step-Deviation methods are mathematically equivalent transformations of the Direct Method formula. They adjust the origin and scale to simplify arithmetic operations, but algebraic expansion yields the identical formula .
Q4: How do you choose the Assumed Mean ()? Can it be any number?
Answer: Mathematically, can be any real number. However, to maximize calculation efficiency, choose the exact midpoint/class mark () of the middle class interval. If there are an even number of classes, pick either of the two central class marks.