Statistical Representation, Central Tendency & Dispersion
Weightage: Roughly 20 marks of the 40-mark Statistics section. The questions are computational and repeat in a narrow set of shapes, which makes this the most reliably scored block in Paper 3 after Logical Reasoning.
Why two measures are needed
Two production lines each average 100 units a day. The first produces between 98 and 102; the second between 40 and 160. The averages are identical and the situations are entirely different — the second is unmanageable and the first is not.
That is the whole justification for this chapter. A measure of central tendency says where the distribution sits. A measure of dispersion says how widely it spreads. Neither alone describes data adequately, and reporting an average without a measure of spread conceals exactly the information a decision-maker needs.
Statistical representation
Primary data is collected by the investigator directly for the purpose in hand. Secondary data already exists, having been collected by someone else for another purpose. Primary data is more reliable and more expensive; secondary data is cheaper and must be assessed for whether its definitions and coverage suit the present use.
Classification groups data by a characteristic — chronological by time, geographical by place, qualitative by attribute, quantitative by magnitude.
A frequency distribution groups values into classes with a count against each. Key terms:
- Class limits are the stated endpoints of a class. Class boundaries are the true limits after adjusting for the gap between classes: for classes 10-19 and 20-29, the boundaries are 9.5-19.5 and 19.5-29.5.
- Class mark or mid-point is the average of the two limits.
- Class width is the difference between the upper and lower boundaries.
- Inclusive classes include both limits, as in 10-19; exclusive classes exclude the upper limit, as in 10-20 and 20-30, where 20 falls in the second class.
Diagrams and graphs. A bar diagram compares magnitudes across categories; a pie chart shows composition as parts of a whole, each sector's angle being the category's proportion times 360 degrees. A histogram represents a continuous frequency distribution with adjacent rectangles whose areas are proportional to frequencies. A frequency polygon joins the mid-points of the tops of a histogram's bars. An ogive or cumulative frequency curve plots cumulative frequencies, and the median can be read off it as the value corresponding to half the total frequency.
Measures of central tendency
Arithmetic mean
For ungrouped data:
For grouped data, using class marks and frequencies :
Weighted mean, where observations carry different importance :
Combined mean of two groups:
Note that this is a weighted mean with the group sizes as weights, and that averaging the two means directly is correct only when the groups are equal in size — a standard trap.
Two properties of the arithmetic mean are examined:
- The sum of deviations from the mean is zero: . This is what makes the mean the balance point of the data.
- The sum of squared deviations is minimum when taken from the mean, which is why variance is defined about the mean.
The mean uses every observation, which is its strength and its weakness: it is the most informative average but is badly affected by extreme values.
Median
The median is the middle value when the data is arranged in order. For observations:
- If is odd, the median is the th value.
- If is even, it is the average of the th and th values.
For grouped data:
where is the lower boundary of the median class, the total frequency, the cumulative frequency before the median class, the frequency of the median class and its width.
The median is not affected by extreme values, which makes it the appropriate average for income and wealth data, where a few very large values would distort the mean. It can also be computed for open-ended distributions, where the mean cannot.
Mode
The mode is the most frequently occurring value. For grouped data:
where is the frequency of the modal class and and the frequencies of the preceding and following classes.
A distribution may have no mode, one mode, or several. The mode is the only average usable for qualitative data — the most common colour or size has a mode but no mean.
The empirical relationship
For a moderately skewed distribution:
This is an approximation, not an identity, and it holds only for moderately asymmetrical distributions. It is examined regularly because it allows any one of the three to be found from the other two.
In a perfectly symmetrical distribution, mean, median and mode coincide. In a positively skewed distribution the mean exceeds the median, which exceeds the mode; in a negatively skewed distribution the order reverses.
Geometric and harmonic mean
Geometric mean of values is the th root of their product:
It is the correct average for rates of growth and ratios, because growth compounds multiplicatively. Averaging annual growth rates arithmetically overstates the true average growth; the geometric mean does not.
Harmonic mean is the reciprocal of the arithmetic mean of the reciprocals:
It is the correct average for rates expressed per unit, such as speed over equal distances or price per unit for equal amounts spent.
For any set of positive values that are not all equal:
and for two observations, .
Partition values
Quartiles divide ordered data into four parts, deciles into ten and percentiles into a hundred. The second quartile, fifth decile and fiftieth percentile all equal the median. For grouped data the formula mirrors that of the median with the appropriate fraction of replacing .
Measures of dispersion
Absolute and relative measures
An absolute measure is expressed in the units of the data. A relative measure is a pure number, obtained by dividing an absolute measure by an appropriate average, and only relative measures permit comparison between data sets in different units or of very different magnitudes.
Range
Simple, but it depends on only two values and ignores everything between them, so it is severely affected by an outlier.
Quartile deviation
Also called the semi-interquartile range. It covers the middle half of the data and so is unaffected by extreme values, but it ignores the outer halves entirely.
Mean deviation
where is the mean, median or mode. The absolute values are essential: without them the deviations from the mean sum to zero and the measure would always be zero. Mean deviation is least when taken from the median.
Standard deviation and variance
The computational form is usually faster:
The standard deviation squares the deviations rather than taking absolute values, which handles the sign problem while remaining algebraically tractable — and the square root at the end restores the original units, which is why standard deviation rather than variance is quoted alongside a mean.
Properties of the standard deviation, all examined:
- It is independent of a change of origin: adding a constant to every observation leaves it unchanged, because every value and the mean shift equally so the deviations do not change.
- It is not independent of a change of scale: multiplying every observation by multiplies the standard deviation by .
- It is the least of all root-mean-square deviations, being minimum when taken from the mean.
- For consecutive natural numbers, .
Combined standard deviation of two groups:
where and . The terms matter: combining two groups introduces additional variability from the difference between their means, so the combined standard deviation is not simply an average of the two.
Coefficient of variation
This is the relative measure that permits comparison. A lower coefficient of variation indicates greater consistency, uniformity or stability; a higher one indicates greater variability.
The reason a relative measure is needed: a standard deviation of 5 is large for data averaging 20 and negligible for data averaging 5,000. Dividing by the mean removes the effect of scale, and any question asking which of two series is more consistent is asking for the coefficient of variation.
How this chapter is examined
Expect a combined mean or combined standard deviation computation; the empirical relationship used to find one average from the other two; identification of which average is appropriate for given data; a coefficient of variation comparison asking which series is more consistent; and a property question on the effect of adding to or multiplying every observation.
The recurring errors are averaging two group means without weighting by group size, forgetting the square root when moving from variance to standard deviation, omitting the terms in the combined standard deviation, and answering a consistency question with the standard deviation rather than the coefficient of variation. Before selecting any formula, establish whether the data is grouped or ungrouped.
