Published 2026-10-01
Chapter: Statistics

Statistics - Calculation of mean, median, and mode for grouped frequency distributions using direct, assumed mean, and step-deviation methods

In statistical analysis, raw data collected from field observations is often voluminous, unorganized, and difficult to interpret directly. To extract meaningful insights, data is aggregated into grouped frequency distributions. However, analyzing a large table of class intervals and frequencies still requires a single numerical value that represents the central point or typical behavior of the entire dataset. This representative value is known as a measure of central tendency.

The three primary measures of central tendency studied in Class 10 Mathematics are:

  1. Mean: The mathematical average of all observations.
  2. Median: The middle-most value when data is arranged in order of magnitude.
  3. Mode: The value that occurs most frequently in the distribution.

Mastering the calculation of these three measures for grouped data—and understanding when to apply specific computational short-cuts such as the Assumed Mean and Step-Deviation methods—is essential for CBSE Class 10 Mathematics board examinations and forms the foundation for advanced data analysis, economics, and modern data science.


1. Fundamental Concepts & Terminology

Before calculating central tendencies, we must understand how grouped frequency tables are structured.

Class Intervals & Continuity

Data is organized into continuous ranges called class intervals.

  • Exclusive Form (Continuous Data): 10−20,20−30,30−4010-20, 20-30, 30-40. Here, the upper limit of one class equals the lower limit of the next class. An observation with value 2020 is included in 20−3020-30, not 10−2010-20.
  • Inclusive Form (Discontinuous Data): 1−10,11−20,21−301-10, 11-20, 21-30. If data is given in this form, it must be converted to continuous form before calculating median and mode by subtracting 0.50.5 from lower limits and adding 0.50.5 to upper limits: 0.5−10.5,10.5−20.5,20.5−30.50.5-10.5, 10.5-20.5, 20.5-30.5.

Key Definitions

  1. Class Limits: The minimum and maximum boundary values of a class interval. Lower Class Limit =l,Upper Class Limit =u\text{Lower Class Limit } = l, \quad \text{Upper Class Limit } = u

  2. Class Mark (xix_i): The midpoint of a class interval, which acts as the single representative value for all data points falling within that interval. xi=Upper Limit+Lower Limit2=u+l2x_i = \frac{\text{Upper Limit} + \text{Lower Limit}}{2} = \frac{u + l}{2}

  3. Class Size (hh): The difference between the upper limit and lower limit of a continuous class interval. h=Upper Limit−Lower Limit=u−lh = \text{Upper Limit} - \text{Lower Limit} = u - l

  4. Frequency (fif_i): The number of observations falling within the ithi^{\text{th}} class interval. The sum of all frequencies is denoted as NN or ∑fi\sum f_i.


2. Calculation of Mean (xˉ\bar{x}) of Grouped Data

The mean represents the arithmetic average of the dataset. For grouped data, we assume that the frequency of each class interval is centered at its class mark (xix_i).

Depending on the numerical magnitude of xix_i and fif_i, we use one of three computational methods:

               ┌─────────────────────────────────────────┐
               │    Methods for Mean of Grouped Data     │
               └────────────────────┬────────────────────┘
                                    │
         ┌──────────────────────────┼──────────────────────────┐
         │                          │                          │
┌────────┴────────┐        ┌────────┴────────┐        ┌────────┴────────┐
│  Direct Method  │        │  Assumed Mean   │        │ Step-Deviation  │
│                 │        │     Method      │        │     Method      │
└─────────────────┘        └─────────────────┘        └─────────────────┘

Method 1: Direct Method

The Direct Method is applied when the values of class marks (xix_i) and frequencies (fif_i) are numerically small, allowing direct multiplication without risk of calculation errors.

Formula

xˉ=∑i=1nfixi∑i=1nfi\bar{x} = \frac{\sum_{i=1}^{n} f_i x_i}{\sum_{i=1}^{n} f_i}

Algorithm

  1. Construct a table with columns: Class Interval, Frequency (fif_i), Class Mark (xix_i), and Product (fixif_i x_i).
  2. Compute xi=Lower Limit+Upper Limit2x_i = \frac{\text{Lower Limit} + \text{Upper Limit}}{2} for each class interval.
  3. Multiply each fif_i by its corresponding xix_i to find fixif_i x_i.
  4. Sum all frequencies to get ∑fi\sum f_i and sum all products to get ∑fixi\sum f_i x_i.
  5. Divide ∑fixi\sum f_i x_i by ∑fi\sum f_i.

Method 2: Assumed Mean Method

When xix_i values are large, direct multiplication becomes tedious. The Assumed Mean Method reduces numerical size by shifting the origin. We select an arbitrary central class mark as the Assumed Mean (aa) and compute deviations from it.

Formula

xˉ=a+∑i=1nfidi∑i=1nfi\bar{x} = a + \frac{\sum_{i=1}^{n} f_i d_i}{\sum_{i=1}^{n} f_i}

Where:

  • a=Assumed Meana = \text{Assumed Mean} (usually the middle class mark xix_i)
  • di=xi−ad_i = x_i - a (deviation of each class mark from aa)

Algorithm

  1. Calculate class marks (xix_i) for all intervals.
  2. Choose one central value from xix_i as the assumed mean aa.
  3. Compute deviation di=xi−ad_i = x_i - a for each class. (Deviations above aa will be negative; deviations below aa will be positive; deviation for aa itself is 00).
  4. Compute the product fidif_i d_i for each row and sum them to obtain ∑fidi\sum f_i d_i.
  5. Apply the formula: xˉ=a+∑fidi∑fi\bar{x} = a + \frac{\sum f_i d_i}{\sum f_i}.

Method 3: Step-Deviation Method

When class sizes (hh) are equal and the deviations did_i share a common factor hh, we can further simplify calculations by scaling down the deviations.

Formula

xˉ=a+(∑i=1nfiui∑i=1nfi)×h\bar{x} = a + \left( \frac{\sum_{i=1}^{n} f_i u_i}{\sum_{i=1}^{n} f_i} \right) \times h

Where:

  • a=Assumed Meana = \text{Assumed Mean}
  • h=Class Size=u−lh = \text{Class Size} = u - l
  • ui=dih=xi−ahu_i = \frac{d_i}{h} = \frac{x_i - a}{h}

Algorithm

  1. Compute class marks (xix_i), select assumed mean aa, and identify class size hh.
  2. Calculate step-deviations ui=xi−ahu_i = \frac{x_i - a}{h} for each class interval. (For equal class intervals, uiu_i will always follow a sequence like …,−2,−1,0,1,2,…\dots, -2, -1, 0, 1, 2, \dots).
  3. Compute fiuif_i u_i for each row and find the sum ∑fiui\sum f_i u_i.
  4. Multiply ∑fiui∑fi\frac{\sum f_i u_i}{\sum f_i} by hh and add aa.

3. Calculation of Mode of Grouped Data

The Mode is the value among observations that has the maximum frequency. In a grouped frequency distribution, we cannot determine the exact mode simply by looking at frequencies. Instead, we first locate the Modal Class.

Modal Class

The class interval corresponding to the maximum frequency (f1f_1) is called the modal class. The mode lies somewhere within this interval.

Formula

Mode=l+(f1−f02f1−f0−f2)×h\text{Mode} = l + \left( \frac{f_1 - f_0}{2f_1 - f_0 - f_2} \right) \times h

Where:

  • l=Lower limit of the modal classl = \text{Lower limit of the modal class}
  • h=Size of the class interval (assuming equal class sizes)h = \text{Size of the class interval (assuming equal class sizes)}
  • f1=Frequency of the modal classf_1 = \text{Frequency of the modal class}
  • f0=Frequency of the class preceding the modal classf_0 = \text{Frequency of the class preceding the modal class}
  • f2=Frequency of the class succeeding the modal classf_2 = \text{Frequency of the class succeeding the modal class}

Important Note: Ensure class intervals are continuous before finding ll.


4. Calculation of Median of Grouped Data

The Median is the measure of central tendency that gives the value of the middle-most observation in the data.

Cumulative Frequency (cfcf)

To find the median, we construct a Cumulative Frequency Table of the 'less than' type by successively adding individual frequencies.

Median Class

  1. Compute total frequency N=∑fiN = \sum f_i.
  2. Find N2\frac{N}{2}.
  3. Locate the class interval whose cumulative frequency (cfcf) is just greater than or equal to N2\frac{N}{2}. This interval is the Median Class.

Formula

Median=l+(N2−cff)×h\text{Median} = l + \left( \frac{\frac{N}{2} - cf}{f} \right) \times h

Where:

  • l=Lower limit of the median classl = \text{Lower limit of the median class}
  • N=Total number of observations (∑fi)N = \text{Total number of observations } (\sum f_i)
  • cf=Cumulative frequency of the class PRECEDING the median classcf = \text{Cumulative frequency of the class PRECEDING the median class}
  • f=Frequency of the median classf = \text{Frequency of the median class}
  • h=Class size of the median classh = \text{Class size of the median class}

5. Empirical Relationship Among Central Tendencies

For a moderately skewed distribution (a standard real-world data distribution), an empirical relationship exists between Mean, Median, and Mode:

Mode=3Median−2Mean\text{Mode} = 3\text{Median} - 2\text{Mean}

Or rearranged: 3Median=Mode+2Mean3\text{Median} = \text{Mode} + 2\text{Mean}

This relationship allows us to estimate any one measure if the other two are known.


6. Summary Comparison Table

FeatureMean (xˉ\bar{x})MedianMode
Core DefinitionArithmetic average of all data valuesMiddle-most observation of dataMost frequently occurring value
Sensitivity to OutliersHighly sensitive to extreme valuesUnaffected by extreme valuesUnaffected by extreme values
PrerequisitesClass marks (xix_i)Cumulative frequency (cfcf)Identification of max frequency (f1f_1)
Key Variable Neededxix_i, did_i, or uiu_iN2\frac{N}{2} and preceding cfcff1,f0,f2f_1, f_0, f_2
Best Used ForSymmetric data without extreme outliersSkewed distributions (e.g., income, house prices)Categorical data or finding most popular item

7. Real-World Applications

1. Income Analysis & Economics (Why Median Matters)

Consider a small software startup with 10 employees where 9 junior developers earn ₹5,000/month and 1 CEO earns ₹5,00,000/month.

  • Mean Salary: (9×5000)+50000010=₹54,500\frac{(9 \times 5000) + 500000}{10} = \text{₹}54,500. This paints a misleading picture that employees earn well.
  • Median Salary: ₹5,000. This provides a realistic representation of a typical worker's wage. Economists use the median rather than the mean to measure household income and national wealth distribution.

2. Retail, Manufacturing & Inventory Management (Why Mode Matters)

A shoe store manager wants to order new stock. If shoe sizes sold are measured:

  • Sizes sold: 6,7,7,8,8,8,8,8,9,106, 7, 7, 8, 8, 8, 8, 8, 9, 10.
  • Mean size = 7.97.9. Ordering size 7.97.9 shoes is impossible.
  • Mode size = 8. The store stocks maximum inventory of size 8 because it is the most requested item.

3. School Grading Systems (Why Mean Matters)

A teacher evaluates a class of 40 students on a standardized test. The mean score reveals overall class mastery, helping administrators determine if the teaching methodology was effective across the entire group.


8. Step-by-Step Solved Examples

Example 1: Comprehensive Calculation of Mean using All Three Methods

Problem: The following table gives the distribution of marks obtained by 30 students in a mathematics test. Find the mean marks using:

  1. Direct Method
  2. Assumed Mean Method
  3. Step-Deviation Method
Class Interval (Marks)Number of Students (fif_i)
10 – 252
25 – 403
40 – 557
55 – 706
70 – 856
85 – 1006

Solution:

Step 1: Construct the Master Working Table

Let us choose Assumed Mean a=62.5a = 62.5 (the midpoint of the class marks) and h=15h = 15.

Class IntervalFrequency (fif_i)Class Mark (xix_i)fixif_i x_idi=xi−62.5d_i = x_i - 62.5fidif_i d_iui=di15u_i = \frac{d_i}{15}fiuif_i u_i
10 – 25217.535.0-45.0-90-3-6
25 – 40332.597.5-30.0-90-2-6
40 – 55747.5332.5-15.0-105-1-7
55 – 70662.5 (aa)375.00.0000
70 – 85677.5465.015.09016
85 – 100692.5555.030.0180212
Total∑fi=30\sum f_i = 30∑fixi=1860\sum f_i x_i = 1860∑fidi=−15\sum f_i d_i = -15∑fiui=−1\sum f_i u_i = -1

1. By Direct Method:

xˉ=∑fixi∑fi=186030=62\bar{x} = \frac{\sum f_i x_i}{\sum f_i} = \frac{1860}{30} = 62

2. By Assumed Mean Method:

xˉ=a+∑fidi∑fi=62.5+(−1530)=62.5−0.5=62\bar{x} = a + \frac{\sum f_i d_i}{\sum f_i} = 62.5 + \left( \frac{-15}{30} \right) = 62.5 - 0.5 = 62

3. By Step-Deviation Method:

xˉ=a+(∑fiui∑fi)×h=62.5+(−130)×15=62.5−1530=62.5−0.5=62\bar{x} = a + \left( \frac{\sum f_i u_i}{\sum f_i} \right) \times h = 62.5 + \left( \frac{-1}{30} \right) \times 15 = 62.5 - \frac{15}{30} = 62.5 - 0.5 = 62

Conclusion: All three methods yield the exact same mean score of 6262.


Example 2: Finding the Mode of Grouped Data

Problem: The following table shows the age distribution of patients admitted to a hospital during a year. Find the mode of the data.

Age (in years)5 – 1515 – 2525 – 3535 – 4545 – 5555 – 65
Number of patients6112123145

Solution:

  1. Identify the Modal Class: Look at the frequencies: 6,11,21,23,14,56, 11, 21, 23, 14, 5. The maximum frequency is 2323, which corresponds to the class interval 35−4535 - 45. Therefore, Modal Class = 35−4535 - 45.

  2. Extract Parameters:

    • Lower limit of modal class (ll) = 3535
    • Class size (hh) = 45−35=1045 - 35 = 10
    • Frequency of modal class (f1f_1) = 2323
    • Frequency of class preceding modal class (f0f_0) = 2121
    • Frequency of class succeeding modal class (f2f_2) = 1414
  3. Apply the Formula: Mode=l+(f1−f02f1−f0−f2)×h\text{Mode} = l + \left( \frac{f_1 - f_0}{2f_1 - f_0 - f_2} \right) \times h

    Mode=35+(23−212(23)−21−14)×10\text{Mode} = 35 + \left( \frac{23 - 21}{2(23) - 21 - 14} \right) \times 10

    Mode=35+(246−35)×10\text{Mode} = 35 + \left( \frac{2}{46 - 35} \right) \times 10

    Mode=35+(211)×10=35+2011=35+1.818=36.82 years\text{Mode} = 35 + \left( \frac{2}{11} \right) \times 10 = 35 + \frac{20}{11} = 35 + 1.818 = 36.82 \text{ years}

Final Answer: The mode of the data is 36.82 years36.82\text{ years}.


Example 3: Finding Missing Frequencies using Median

Problem: The median of the following data distribution is 28.528.5. Find the values of missing frequencies xx and yy, given that the total frequency is 6060.

Class Interval0 – 1010 – 2020 – 3030 – 4040 – 5050 – 60Total
Frequency5xx2015yy560

Solution:

Step 1: Construct the Cumulative Frequency Table

Class IntervalFrequency (fif_i)Cumulative Frequency (cfcf)
0 – 1055
10 – 20xx5+x5 + x
20 – 302025+x25 + x
30 – 401540+x40 + x
40 – 50yy40+x+y40 + x + y
50 – 60545+x+y45 + x + y
TotalN=60N = 60

Step 2: Formulate Equation (1) from Total Frequency ∑fi=45+x+y=60\sum f_i = 45 + x + y = 60 x+y=60−45  ⟹  x+y=15— (Equation 1)x + y = 60 - 45 \implies x + y = 15 \quad \text{--- (Equation 1)}

Step 3: Identify Median Class Given Median=28.5\text{Median} = 28.5. Since 28.528.5 lies within the range 20−3020 - 30, the Median Class is 20−3020 - 30.

Step 4: Extract Parameters

  • Lower limit (ll) = 2020
  • Class size (hh) = 1010
  • Frequency of median class (ff) = 2020
  • Cumulative frequency of class preceding median class (cfcf) = 5+x5 + x
  • Total observations (NN) = 60  ⟹  N2=3060 \implies \frac{N}{2} = 30

Step 5: Apply Median Formula Median=l+(N2−cff)×h\text{Median} = l + \left( \frac{\frac{N}{2} - cf}{f} \right) \times h

28.5=20+(30−(5+x)20)×1028.5 = 20 + \left( \frac{30 - (5 + x)}{20} \right) \times 10

28.5−20=(25−x20)×1028.5 - 20 = \left( \frac{25 - x}{20} \right) \times 10

8.5=25−x28.5 = \frac{25 - x}{2}

8.5×2=25−x  ⟹  17=25−x8.5 \times 2 = 25 - x \implies 17 = 25 - x

x=25−17=8x = 25 - 17 = 8

Step 6: Substitute x=8x = 8 into Equation 1 8+y=15  ⟹  y=78 + y = 15 \implies y = 7

Final Answer: The missing frequencies are x=8x = 8 and y=7y = 7.


9. Common Student Mistakes to Avoid

1. Wrong Cumulative Frequency (cfcf) in Median Formula

  • Mistake: Using the cumulative frequency of the median class itself instead of the class preceding it.
  • Correction: Always select cfcf from the row above the median class.

2. Ignoring Discontinuous Class Intervals

  • Mistake: Calculating lower limits (ll) directly from discontinuous boundaries like 1−10,11−201-10, 11-20 without converting them to 0.5−10.5,10.5−20.50.5-10.5, 10.5-20.5.
  • Correction: If data is given in inclusive form, convert it to continuous form first. For 11−2011-20, ll is 10.510.5, not 1111.

3. Choosing Class Mark (xix_i) as Modal Class or Median Class Value

  • Mistake: Identifying xix_i value with the highest frequency and writing that as the Mode.
  • Correction: Use the grouped mode formula. Mode is not simply a class mark.

4. Algebraic Sign Errors in Step-Deviation uiu_i

  • Mistake: Reversing the signs of uiu_i (e.g., making values above assumed mean positive and below negative).
  • Correction: Values of xix_i smaller than aa give negative did_i and uiu_i. Values greater than aa give positive values.

10. Practice Questions for Self-Assessment

Practice Question 1

Find the mean of the following distribution using the Step-Deviation Method:

Class100 – 150150 – 200200 – 250250 – 300300 – 350
Frequency451222

Solution:

  • Class marks (xix_i): 125,175,225,275,325125, 175, 225, 275, 325.
  • Choose Assumed Mean a=225a = 225, class size h=50h = 50.
  • Calculate ui=xi−22550u_i = \frac{x_i - 225}{50}: −2,−1,0,1,2-2, -1, 0, 1, 2.
  • fiuif_i u_i column:
    • 4×(−2)=−84 \times (-2) = -8
    • 5×(−1)=−55 \times (-1) = -5
    • 12×0=012 \times 0 = 0
    • 2×1=22 \times 1 = 2
    • 2×2=42 \times 2 = 4
  • Sum of frequencies ∑fi=4+5+12+2+2=25\sum f_i = 4 + 5 + 12 + 2 + 2 = 25.
  • Sum of products ∑fiui=−13+6=−7\sum f_i u_i = -13 + 6 = -7.
  • Apply formula: xˉ=225+(−725)×50=225+(−7×2)=225−14=211\bar{x} = 225 + \left( \frac{-7}{25} \right) \times 50 = 225 + (-7 \times 2) = 225 - 14 = 211

Answer: Mean = 211211.


Practice Question 2

Calculate the Mode for the following frequency distribution:

Marks0 – 1010 – 2020 – 3030 – 4040 – 50
Number of Students5815102

Solution:

  • Maximum frequency f1=15f_1 = 15 corresponds to interval 20−3020 - 30.
  • Modal class = 20−3020 - 30.
  • l=20,f1=15,f0=8,f2=10,h=10l = 20, f_1 = 15, f_0 = 8, f_2 = 10, h = 10.
  • Apply formula: Mode=20+(15−82(15)−8−10)×10\text{Mode} = 20 + \left( \frac{15 - 8}{2(15) - 8 - 10} \right) \times 10 Mode=20+(730−18)×10=20+7012=20+5.83=25.83\text{Mode} = 20 + \left( \frac{7}{30 - 18} \right) \times 10 = 20 + \frac{70}{12} = 20 + 5.83 = 25.83

Answer: Mode = 25.8325.83.


Practice Question 3

Compute the Median for the following distribution:

Class Interval50 – 6060 – 7070 – 8080 – 9090 – 100
Frequency371163

Solution:

  • Compute Cumulative Frequencies (cfcf):
    • 50−6050 - 60: f=3,cf=3f = 3, cf = 3
    • 60−7060 - 70: f=7,cf=10f = 7, cf = 10
    • 70−8070 - 80: f=11,cf=21f = 11, cf = 21
    • 80−9080 - 90: f=6,cf=27f = 6, cf = 27
    • 90−10090 - 100: f=3,cf=30f = 3, cf = 30
  • Total N=30  ⟹  N2=15N = 30 \implies \frac{N}{2} = 15.
  • cfcf just greater than 1515 is 2121, corresponding to Median Class 70−8070 - 80.
  • l=70,f=11,cf=10 (preceding),h=10l = 70, f = 11, cf = 10 \text{ (preceding)}, h = 10.
  • Apply formula: Median=70+(15−1011)×10=70+5011=70+4.55=74.55\text{Median} = 70 + \left( \frac{15 - 10}{11} \right) \times 10 = 70 + \frac{50}{11} = 70 + 4.55 = 74.55

Answer: Median = 74.5574.55.


11. Exam Revision & Frequently Asked Questions (FAQs)

Q1: What is the empirical relationship between Mean, Median, and Mode, and when is it used?

Answer: The empirical relationship is: Mode=3Median−2Mean\text{Mode} = 3\text{Median} - 2\text{Mean} It is used to estimate one measure of central tendency when two measures are already calculated, or in 1-mark board exam questions where two values are provided directly.


Q2: Can we apply the Step-Deviation method if class sizes (hh) are unequal?

Answer: Strictly speaking, no. The standard Step-Deviation method assumes a constant class size hh across all intervals so that ui=xi−ahu_i = \frac{x_i - a}{h} results in integers. If class sizes are unequal, use either the Direct Method or the Assumed Mean Method.


Q3: Why do all three methods for calculating the mean give the exact same answer?

Answer: The Assumed Mean and Step-Deviation methods are mathematically equivalent transformations of the Direct Method formula. They adjust the origin and scale to simplify arithmetic operations, but algebraic expansion yields the identical formula ∑fixi∑fi\frac{\sum f_i x_i}{\sum f_i}.


Q4: How do you choose the Assumed Mean (aa)? Can it be any number?

Answer: Mathematically, aa can be any real number. However, to maximize calculation efficiency, choose the exact midpoint/class mark (xix_i) of the middle class interval. If there are an even number of classes, pick either of the two central class marks.

NCERT Study Guide Directory

Textbook solutions, chapter notes & practice worksheets by grade

Interlinked syllabus