By the end of this chapter you'll be able to…

  • 1Construct a frequency distribution table from raw data and identify class marks
  • 2Calculate mean using: (i) direct method, (ii) assumed mean method, (iii) step deviation method
  • 3Calculate median for ungrouped and grouped data using the median formula
  • 4Find mode for ungrouped data; identify the modal class for grouped data
  • 5Draw a histogram and frequency polygon from a grouped frequency distribution
  • 6Choose the appropriate average (mean/median/mode) for a given context
💡
Why this chapter matters
Statistics is the most formula-dependent chapter in AP Class 9 Mathematics, but the formulas are few and the method is systematic. The mean (direct method and assumed mean method) is the highest-probability 4-5 mark question — a frequency table is given and the student must compute. The median for grouped data formula is a standard 3-4 mark question. Drawing a histogram and frequency polygon is a 3-4 mark graphical question. The step deviation method is the fastest calculation technique for equal-class-width data and is worth mastering.

Before you start — revise these

A 5-minute refresher here will save you 30 minutes of confusion below.

Statistics — Class 9 Mathematics

"Statistics turns NUMBERS into KNOWLEDGE. It is the science of making SENSE of data."

1. Data — Types and Collection

Primary Data: Collected DIRECTLY by the investigator for a SPECIFIC purpose. Survey, experiment. Secondary Data: ALREADY collected by someone else. Census data, reports, websites. 'Primary data is more RELIABLE but more EXPENSIVE and TIME-CONSUMING. Secondary data is cheaper but may not perfectly fit your purpose.'

2. Presentation of Data

Frequency Distribution Table

For large datasets: group into CLASS INTERVALS (e.g., 0-10, 10-20, 20-30). Tally marks count observations. Class size = Upper limit − Lower limit.

Graphical Representation

Bar Graph: For DISCRETE data. Bars have GAPS between them. Equal width bars. Height = frequency.

Histogram: For CONTINUOUS grouped data. Bars TOUCH each other (no gaps). Equal class intervals → height = frequency. Unequal class intervals → AREA = frequency (adjust height: height = frequency / class width).

Frequency Polygon: Join the MIDPOINTS of the tops of histogram bars with straight lines. Also plot at ZERO frequency at the beginning and end. 'A frequency polygon shows the SHAPE of the distribution more clearly than a histogram.'


3. Measures of Central Tendency

Mean (Arithmetic Average)

Ungrouped Data: X̄ = (x₁ + x₂ + ... + xₙ) / n = Σx / n.

Grouped Data — Direct Method: X̄ = Σfᵢxᵢ / Σfᵢ. Where xᵢ = class mark (midpoint) = (lower limit + upper limit) / 2.

Grouped Data — Assumed Mean Method: X̄ = A + (Σfᵢdᵢ / Σfᵢ). Where A = assumed mean (any class mark, usually middle one). dᵢ = xᵢ − A.

Grouped Data — Step Deviation Method: X̄ = A + h(Σfᵢuᵢ / Σfᵢ). Where uᵢ = (xᵢ − A) / h. h = class size. 'This is the SHORTEST method — use it for large datasets with equal class sizes.'

Median

Ungrouped Data: Arrange in ascending order. If n is ODD → median = (n+1)/2 th term. If n is EVEN → median = average of (n/2)th and (n/2+1)th terms.

Grouped Data: Median = L + [(N/2 − CF) / f] × h. Where L = lower limit of median class. N = Σfᵢ. CF = cumulative frequency just BEFORE the median class. f = frequency of median class. h = class size.

Mode

Ungrouped Data: The value that occurs MOST FREQUENTLY.

Grouped Data: Mode = L + [(f₁ − f₀) / (2f₁ − f₀ − f₂)] × h. Where f₁ = frequency of modal class (highest). f₀ = frequency before. f₂ = frequency after.


4. Worked Examples

Example 1 — Mean (Direct Method)

Class0-1010-2020-3030-4040-50
fᵢ5812105

xᵢ (midpoints): 5, 15, 25, 35, 45. Σfᵢxᵢ = 5×5 + 8×15 + 12×25 + 10×35 + 5×45 = 25+120+300+350+225 = 1020. Σfᵢ = 40. X̄ = 1020/40 = 25.5.

Example 2 — Median (Grouped)

Class0-1010-2020-3030-4040-50
fᵢ461073

N = 30. N/2 = 15. Cumulative frequencies: 4, 10, 20, ... Median class = 20-30 (15 falls in this class). L = 20. CF = 10. f = 10. h = 10. Median = 20 + [(15−10)/10]×10 = 20 + 5 = 25.


5. Common Mistakes to Avoid

  1. 'Bar graphs and histograms are the same' — Histogram bars TOUCH (continuous data). Bar graph bars have GAPS (discrete).
  2. Using xᵢ (class mark) directly as midpoints without computing — xᵢ = (lower + upper)/2.
  3. Confusing cumulative frequency (CF) in median formula — CF is BEFORE the median class, not the cumulative frequency OF the median class.

6. AP Exam Focus

TopicMarks
Mean — direct and assumed mean4-5
Median for grouped data3-4
Histogram / Frequency polygon3-4

Worked Example — Mean by Step Deviation

| Class | 0-10 | 10-20 | 20-30 | 30-40 | 40-50 | | fᵢ | 5 | 12 | 20 | 10 | 3 | Σf = 50

Let A = 25 (midpoint of middle class). h = 10. xᵢ (midpoints): 5, 15, 25, 35, 45. uᵢ = (xᵢ−A)/h: −2, −1, 0, 1, 2. Σfu = 5(−2)+12(−1)+20(0)+10(1)+3(2) = −10−12+0+10+6 = −6. X̄ = A + h(Σfu/Σf) = 25 + 10(−6/50) = 25 − 1.2 = 23.8.

Key Memory Aid — When to Use Which Average

Mean: Use when data is SYMMETRICAL. Affected by EXTREME values (outliers). Median: Use when data is SKEWED (income, house prices). Not affected by extremes. Mode: Use for CATEGORICAL data (most popular item, best-selling size).

Data Interpretation Skills for AP Exam

  • 'READ the question carefully — does it ask for mean, median, or mode?'
  • 'For grouped data: FIRST identify the correct class (median class, modal class). THEN apply the formula.'
  • 'In histograms: area = frequency. For unequal class widths, ADJUST the height.'
  • 'ALWAYS state the UNITS in your final answer.'

More Worked Examples

Example — Finding the Modal Class: Given frequencies: 0-10(4), 10-20(8), 20-30(15), 30-40(12), 40-50(6). Modal class = 20-30 (highest frequency = 15).

Example — Median from Ogive: Plot cumulative frequencies (4, 12, 27, 39, 45) vs upper limits (10, 20, 30, 40, 50). Draw smooth curve. Mark N/2 = 22.5 on y-axis. Draw horizontal line to curve. Drop vertical to x-axis → Median ≈ 25.

Example — Choosing the Right Average: A company reports 'average salary = ₹50,000.' But 2 executives earn ₹5,00,000 each and 8 workers earn ₹12,500 each. Mean = (2×5L + 8×12,500)/10 = (10L+1L)/10 = ₹1,10,000 (misleading — inflated by executives). Median = ₹12,500 (the TRUE typical salary). 'The MEDIAN tells the real story. Always ask: which average is being reported — and WHY?'

Common Graph Errors in AP Exams

  • Bar graph: bars must have EQUAL WIDTH. Gaps between bars. Start y-axis from ZERO.
  • Histogram: NO gaps (unless frequency = 0). Area ∝ frequency. For unequal class widths, height = frequency/class width.
  • Frequency polygon: Start and end at ZERO frequency (touch the x-axis at both ends).

CBSE/AP Exam Tip: When drawing a histogram, ALWAYS label both axes clearly. X-axis: 'Class Intervals (with units).' Y-axis: 'Frequency (number of students/items).' If asked to draw a frequency polygon WITHOUT a histogram, first construct the histogram, then join midpoints. For the assumed mean method, choose A as a middle class mark to minimise calculations with large numbers.

Final Reminder: Statistics is the MOST PRACTICAL math topic — you'll use it in science experiments, business, and everyday decision-making. Master the formulas now, and you'll use them for life.

Key formulas & results

Everything you need to memorise, in one card. Screenshot this for revision.

Central Tendency Formulas
CLASS MARK (midpoint): xᵢ = (lower limit + upper limit) / 2. MEAN — DIRECT METHOD: X̄ = Σfᵢxᵢ / Σfᵢ. MEAN — ASSUMED MEAN METHOD: X̄ = A + (Σfᵢdᵢ / Σfᵢ), where dᵢ = xᵢ − A (A = assumed mean, any class mark). MEAN — STEP DEVIATION METHOD: X̄ = A + h × (Σfᵢuᵢ / Σfᵢ), where uᵢ = (xᵢ − A)/h, h = class width. (Use this when class widths are equal — fewest calculations). MEDIAN (UNGROUPED): Arrange in ascending order. Odd n → middle term = (n+1)/2 th. Even n → average of (n/2)th and (n/2+1)th. MEDIAN (GROUPED): Median = L + [(N/2 − CF) / f] × h. L = lower class boundary of median class. N = Σfᵢ (total). CF = cumulative frequency just BEFORE median class. f = frequency of median class. h = class width. FIND MEDIAN CLASS: First compute N/2. Find first CF that is ≥ N/2. That class is the median class. MODE (UNGROUPED): Value with highest frequency. MODE (GROUPED): Modal class = class with highest frequency. Mode = L + [(f₁ − f₀) / (2f₁ − f₀ − f₂)] × h. f₁ = freq of modal class. f₀ = freq before. f₂ = freq after. GRAPHS: HISTOGRAM: bars TOUCH (continuous data), height = frequency (equal widths), or height = frequency density = freq/class width (unequal widths). FREQUENCY POLYGON: join midpoints of histogram tops with straight lines; extend to zero at both ends (midpoint of class before first and after last). BAR GRAPH: bars with GAPS (discrete data).
AP EXAM KEY TRAPS: (1) CF in the MEDIAN formula = cumulative frequency BEFORE (not of) the median class. (2) HISTOGRAM bars TOUCH. Bar graph has GAPS. (3) For step deviation: uᵢ values are integers (−2, −1, 0, 1, 2 for 5 classes) — very fast to work with. (4) When n is EVEN in ungrouped median: take AVERAGE of the two middle terms. (5) In frequency polygon: always start and end at x-axis (frequency = 0 outside the range). WHEN TO USE EACH AVERAGE: Mean = symmetrical data. Median = skewed data or when outliers exist (salaries, house prices). Mode = categorical data (most popular, most common).
⚠️

Common mistakes & fixes

These are the exact errors that cost students marks in board exams. Read them once, save yourself the trouble.

WATCH OUT
Using the cumulative frequency of the median class (instead of the class before it) in the median formula
In the grouped median formula: Median = L + [(N/2 − CF) / f] × h, the CF (cumulative frequency) is the cumulative frequency of all classes BEFORE the median class — it is the cumulative frequency just BELOW the median class, not the cumulative frequency including the median class itself. STEP-BY-STEP: (1) Compute N/2. (2) Make a cumulative frequency column. (3) Find the class where the cumulative frequency FIRST reaches or exceeds N/2 — this is the median class. (4) CF = cumulative frequency of the class BEFORE the median class (the row just above it). (5) L = lower class boundary of the median class. Example: if N=30, N/2=15, and cumulative frequencies are 4, 10, 20 (at classes 0-10, 10-20, 20-30), the median class is 20-30 (first class where CF ≥ 15). CF = 10 (cumulative freq of 10-20, the class BEFORE). Not 20.

Practice problems

Work through this chapter's problems as a readiness check — reveal each solution, mark yourself honestly, and get your gap report at the end.

Readiness check

Are you exam-ready for Statistics?

1 problems from this chapter. Try each one, reveal the worked solution, mark yourself honestly — get your gap report at the end.

1 questions~2 min

5-minute revision

The whole chapter, distilled. Read this the night before the exam.

  • CLASS MARK (midpoint): xᵢ = (lower limit + upper limit)/2. Example: for class 10-20, midpoint = 15.
  • MEAN — DIRECT METHOD: X̄ = Σfᵢxᵢ / Σfᵢ. Multiply each midpoint by its frequency, sum, divide by total frequency. Simple but tedious for large numbers.
  • MEAN — ASSUMED MEAN METHOD: X̄ = A + (Σfᵢdᵢ / Σfᵢ), where dᵢ = xᵢ − A, A = assumed mean (usually a midpoint near the middle). Reduces calculation size.
  • MEAN — STEP DEVIATION METHOD: X̄ = A + h × (Σfᵢuᵢ / Σfᵢ), where uᵢ = (xᵢ − A)/h, h = class width. FASTEST METHOD when all class widths are equal — uᵢ values are small integers like −2, −1, 0, 1, 2.
  • MEDIAN (UNGROUPED): Arrange data in ascending order. ODD n: median = ((n+1)/2)th value. EVEN n: median = average of (n/2)th and (n/2+1)th values.
  • MEDIAN (GROUPED): Median = L + [(N/2 − CF)/f] × h. L = lower boundary of median class. N = total frequency. CF = cumulative frequency BEFORE median class (NOT of). f = frequency of median class. h = class width.
  • FINDING MEDIAN CLASS: (1) Compute N/2. (2) Make CF column. (3) Find the FIRST class where CF reaches or exceeds N/2 — that is the median class.
  • MODE (UNGROUPED): The value occurring most frequently. MODE (GROUPED): Modal class = class with highest frequency. Mode = L + [(f₁ − f₀)/(2f₁ − f₀ − f₂)] × h. f₁ = freq of modal class, f₀ = freq before, f₂ = freq after.
  • HISTOGRAM: bars TOUCH each other (continuous data). Equal widths → height = frequency. Unequal widths → height = frequency DENSITY = frequency/width.
  • FREQUENCY POLYGON: join the MIDPOINTS of histogram tops with straight lines. Extend to ZERO at both ends (midpoint of class before first and after last). Useful for comparing two or more distributions on the same graph.
  • BAR GRAPH vs HISTOGRAM: bar graph has GAPS between bars (discrete data — different categories). Histogram has bars TOUCHING (continuous numerical data — class intervals).
  • CHOICE OF AVERAGE: Mean = symmetric data without extreme outliers. Median = skewed data with outliers (e.g., incomes, house prices). Mode = categorical data or 'most popular' (e.g., shoe sizes, favourite colours).

Andhra Pradesh (BIEAP) marks blueprint

Where the marks come from in this chapter — so you can plan your prep.

Where this shows up in the real world

This chapter isn't just an exam topic — it lives in the world around you.

AP government welfare schemes and income data

When the AP government targets welfare benefits (BPL cards, scholarships), it uses MEDIAN household income (not mean) because mean is distorted by ultra-high earners. Class 9 statistics — choosing the right average — directly affects policy. The Below Poverty Line (BPL) threshold is set using median household income data from National Sample Surveys (NSS). Understanding why median is preferred is crucial for understanding social welfare in AP.

School performance analysis

Teachers in AP schools analyse exam results using mean (overall class performance), median (typical student), and mode (most common score). A wide gap between mean and median (e.g., mean = 65, median = 50) indicates the average is being pulled up by a few high scorers — the typical student is struggling. AP's Samagra Shiksha programme uses such statistical analysis to identify schools needing remediation.

Cricket statistics and sports analytics

Cricket commentary uses statistics from this chapter constantly. A batsman's batting AVERAGE (= total runs / innings) is the MEAN. Cricket commentators also discuss MEDIAN runs (more 'typical' performance) and MODAL scores. Hawk-Eye and sports analytics platforms compute frequency distributions of bowling speeds, batting positions, and scoring patterns. The class 9 statistics is the foundation of modern sports analytics.

Exam strategy

Battle-tested tips from teachers and toppers for this chapter.

  1. Mean (4-5 marks): make a complete TABLE with columns Class, Class mark x, Frequency f, fx (and additional columns for assumed mean d, or step deviation u). Show all entries. Compute Σfx and Σf. Apply the formula and get the mean. Each step earns marks.
  2. Median (3-4 marks): identify median class explicitly. Write 'N = ___, N/2 = ___. CF = ___ (of the class before median class). L = ___. f = ___. h = ___.' Then substitute into formula. Show each substitution.
  3. Mode: identify modal class first (highest frequency). Then write 'f₁ = ___, f₀ = ___, f₂ = ___, L = ___, h = ___.' Substitute carefully. The (2f₁ − f₀ − f₂) in the denominator is the most error-prone step.
  4. Histogram (3-4 marks): use graph paper. Label both axes (x: class intervals, y: frequency). Draw bars touching each other. Equal width bars if class widths are equal. Title the histogram.
  5. Frequency polygon: can be drawn on a histogram (joining midpoints) or independently. Always extend to ZERO at both ends. Plot midpoints on x-axis, frequencies on y-axis.

Going beyond the textbook

For olympiad aspirants and curious learners — topics that build on this chapter.

  • Research VARIANCE and STANDARD DEVIATION — measures of how spread out the data is. Variance = Σf(x−mean)²/N. Standard deviation σ = √variance. Two data sets can have the same mean but very different spreads. Research the 68-95-99.7 rule for normal distributions.
  • Investigate Probability Distributions — extending statistics from descriptive (mean, median, mode of given data) to probabilistic (expected value, variance of random variables). The Normal Distribution (bell curve) is the most important — central to hypothesis testing, opinion polls, and quality control.
  • Explore Big Data analytics — modern data science (used by Flipkart, Amazon, Netflix in their AP operations) computes mean engagement times, median session durations, and mode content categories for billions of users. The basic Class 9 measures of central tendency are the foundational building blocks of machine learning and AI.
  • Research the LORENZ CURVE and GINI COEFFICIENT — graphical and numerical measures of income inequality. The Lorenz Curve plots cumulative share of income vs cumulative share of population. The Gini Coefficient (0 = perfect equality, 1 = total inequality) summarises inequality. India's Gini ≈ 0.35 (moderate). Research how these statistics are used by economists and policy makers in AP.

Where else this chapter is tested

CBSE board isn't the only one — other exams test this chapter too.

AP Board SSC (Class 10) — StatisticsVery High — Class 9 statistics directly extends to Class 10 mean by all 3 methods, median, mode, and ogive curves
JEE Main (Statistics)Medium — measures of central tendency and dispersion are JEE topics
NTSE (Mathematics)High — statistics problems are standard NTSE topics
AP EAPCET and Commerce streamVery High — statistics is a major chapter in EAPCET Mathematics and Class 11-12 Commerce statistics

Questions students ask

The real ones — pulled from the Q&A community and tutor sessions.

Use STEP DEVIATION when: (1) All class widths are EQUAL (same h). (2) Numbers are LARGE (midpoints in hundreds or thousands). (3) You want the FASTEST calculation. The step deviation method REDUCES the data to small integers (typically −3 to +3), making mental arithmetic possible. EXAMPLE: midpoints 100, 110, 120, 130, 140. Direct method: multiply each by frequencies and sum — large numbers. Step deviation: take A = 120, h = 10. uᵢ values become −2, −1, 0, 1, 2. Much faster arithmetic. RULE OF THUMB: if class widths are equal and midpoints have more than 2 digits, step deviation saves significant time. If midpoints are already small (1-2 digits), direct method is faster (no need to convert).

FOUR STEPS: (1) Sum all frequencies to get N. (2) Compute N/2. (3) Build a CUMULATIVE FREQUENCY column — running total of frequencies down the table. (4) Find the FIRST class where CF reaches or exceeds N/2 — THIS is the median class. EXAMPLE: Frequencies 5, 8, 15, 10, 2 → CFs 5, 13, 28, 38, 40. N = 40, N/2 = 20. Scan CF: 5 (no, <20), 13 (no, <20), 28 (yes, ≥20). The class with CF = 28 is the median class. CF FOR FORMULA = 13 (the CF of the class BEFORE the median class, not of the median class itself). This 'CF before' is the most common error — students use 28 (CF of median class) and get wrong answers.

A FREQUENCY POLYGON connects the midpoints of the tops of histogram bars. To make the polygon a CLOSED FIGURE (and to suggest that the frequency 'returns to zero' outside the data range), we extend it to imaginary classes BEFORE the first and AFTER the last — both with frequency 0. The midpoints of these imaginary classes are plotted on the x-axis (y = 0). EXAMPLE: if classes are 0-10, 10-20, 20-30 with frequencies 5, 8, 3, the midpoints are 5, 15, 25. Add imaginary midpoints −5 (before first) and 35 (after last) with frequency 0 each. Connect: (−5, 0), (5, 5), (15, 8), (25, 3), (35, 0). This gives a closed polygon resting on the x-axis. The extension makes it possible to compute the AREA UNDER the polygon, which equals the area of the histogram.

MEAN: best for SYMMETRIC distributions without outliers. Uses every data value. Example: heights of students, exam marks (most class distributions). Sensitive to outliers — one extreme value can shift the mean significantly. MEDIAN: best for SKEWED data or data with outliers. Uses only the middle value(s). Example: income data (a few billionaires inflate the mean but not the median). Real estate prices (a few mansions skew the mean). Median is more 'representative' of the typical value when extremes exist. MODE: best for CATEGORICAL data or 'most popular' questions. Example: most common shoe size sold (you can't have a 'mean' shoe size of 7.3). Most popular colour. Most frequent answer in a survey. Some data sets are MULTIMODAL (have more than one mode). RULE OF THUMB: 'What does typical mean here?' Mean = average. Median = middle. Mode = most common.

When CLASS WIDTHS ARE EQUAL, the height of each bar = frequency directly. When CLASS WIDTHS ARE UNEQUAL, you cannot use frequency directly — wider classes would appear too tall and mislead the visual comparison. Use FREQUENCY DENSITY instead: FD = frequency / class width. The HEIGHT of each bar = frequency density. The AREA of each bar then equals the frequency (area = FD × width = frequency). EXAMPLE: classes 0-5 (width 5, freq 10) and 5-25 (width 20, freq 40). Direct heights 10 vs 40 would mislead (suggesting the wider class is 4× more frequent). Frequency densities: 10/5 = 2, 40/20 = 2 — equal! The two classes actually represent the same density of data, just over different ranges. The histogram bars would be equal height (both 2) but different widths.
Verified by the tuition.in editorial team
Last reviewed on 28 May 2026. Written and reviewed by subject-matter experts — read about our process.
Editorial process →
Header Logo