Probability, Distributions, Correlation, Regression & Index Numbers
Weightage: Roughly 20 marks of the 40-mark Statistics section. The formulas are numerous but each question type is narrow and repeats, which makes this a drilling chapter rather than an understanding-heavy one — with two exceptions where understanding decides the answer.
Probability
Definitions
A random experiment has more than one possible outcome and the outcome cannot be predicted in advance. The sample space is the set of all possible outcomes, and an event is a subset of the sample space.
The classical definition, applicable where outcomes are equally likely:
The relative frequency definition takes as the limiting proportion of times occurs in a long run of trials, and applies where outcomes are not equally likely.
Under the axiomatic approach, , , and the probability of the union of mutually exclusive events is the sum of their probabilities.
, and computing the complement is often far quicker than computing the event directly — particularly for events described as "at least one".
Addition and multiplication
reducing to when and are mutually exclusive, since .
reducing to when and are independent.
Mutually exclusive and independent are different
This is the first of the two points where understanding rather than memory decides the answer, and it is examined regularly.
Mutually exclusive means the events cannot occur together: . Drawing a card that is both a king and a queen is impossible.
Independent means the occurrence of one does not affect the probability of the other: .
Two events with non-zero probabilities that are mutually exclusive cannot be independent. If they are mutually exclusive, then knowing has occurred tells you certainly has not, which is the strongest possible dependence. The two concepts are not merely different; for non-trivial events they are incompatible.
Conditional probability and Bayes' theorem
Bayes' theorem reverses a conditional probability: it converts the probability of the evidence given a cause into the probability of the cause given the evidence.
Expected value
The expected value is the long-run average outcome, and it need not be a possible value of the variable. It is the basis of the guessing computation for this paper: with four options and a penalty of 0.25, .
Theoretical distributions
Binomial distribution
Applies to independent trials, each with two outcomes and a constant probability of success.
Since , the variance is always less than the mean in a binomial distribution — a property used to identify it.
Poisson distribution
Applies to rare events over a continuous interval of time or space, where is large and small.
The equality of mean and variance is the identifying property of the Poisson distribution, and it is the standard examination question on it.
Normal distribution
A continuous distribution, symmetrical and bell-shaped, defined by its mean and standard deviation. Its properties:
- Mean, median and mode coincide.
- It is symmetrical about the mean, so the two halves each carry a probability of 0.5.
- The curve is asymptotic to the horizontal axis, never touching it.
- Total area under the curve is 1.
- Approximately 68% of observations lie within one standard deviation of the mean, 95% within two, and 99.7% within three.
- Quartile deviation, mean deviation and standard deviation stand in the approximate ratio .
The standard normal variable is , which converts any normal distribution to one with mean 0 and standard deviation 1, allowing a single table to serve all cases.
Correlation
Correlation measures the strength and direction of a linear relationship between two variables.
Karl Pearson's coefficient:
Its properties:
- lies between and .
- It is independent of change of origin and of scale, so adding to or multiplying the variables does not change it.
- It is a pure number with no units.
- is perfect positive correlation, perfect negative, and no linear correlation.
Spearman's rank correlation, for ranked or ordinal data:
where is the difference between the ranks of each pair.
The coefficient of determination gives the proportion of variation in one variable explained by the other. An of 0.8 gives , so 64% of the variation is explained and 36% is not.
Correlation is not causation. A high correlation may reflect a causal relationship in either direction, a common cause acting on both, or pure coincidence. This is the second point where understanding rather than memory decides the answer.
Regression
Where correlation measures the strength of a relationship, regression estimates one variable from the other.
The two regression lines:
The two lines are different and are used for different purposes: to estimate from , use the line of on . They intersect at , so the means always satisfy both equations — which is how the means are recovered when only the two lines are given.
Properties of the regression coefficients:
The correlation coefficient is the geometric mean of the two regression coefficients, and it takes the sign of the coefficients.
Both regression coefficients must have the same sign, which is the sign of . It follows that if one is positive and the other negative, the data has been misreported. Also, since , the product cannot exceed 1, so both coefficients cannot exceed 1 in magnitude.
Regression coefficients are independent of change of origin but not of change of scale — unlike the correlation coefficient, which is independent of both.
Index numbers
An index number measures the relative change in a variable or group of variables over time or between places. The period against which comparison is made is the base period.
Price index formulas
For quantities and prices , with subscript 0 for the base period and 1 for the current period:
Laspeyres uses base year quantities as weights:
Paasche uses current year quantities:
Fisher's ideal index is the geometric mean of the two:
Laspeyres tends to overstate price rises and Paasche to understate them, because consumers substitute away from goods whose prices rise, so base-year quantities overweight those goods and current-year quantities underweight them. Fisher's index, lying between the two, moderates both biases.
Tests of adequacy
- Unit test — the index should be independent of the units in which prices are quoted. All the weighted formulas satisfy it; the simple aggregative index does not.
- Time reversal test — . Reversing the periods should invert the index.
- Factor reversal test — . The price index multiplied by the quantity index should equal the value index.
- Circular test — .
Fisher's index is called ideal because it satisfies both the time reversal and the factor reversal tests, which Laspeyres and Paasche do not. It fails the circular test, which only the simple aggregative index and the fixed-weight aggregative index satisfy.
Related concepts
The cost of living index measures the change in the cost of maintaining a given standard of living, computed by the aggregate expenditure method or the family budget method. Base shifting converts a series to a new base by dividing every value by the new base year's value and multiplying by 100. Deflating converts money values to real values by dividing by the price index and multiplying by 100.
How this chapter is examined
Probability appears as a straightforward computation using the addition or multiplication theorem, a conditional probability, or a distinction question on mutually exclusive against independent. Distributions appear as a mean or variance computation, or as identification from the relationship between mean and variance.
Correlation and regression appear as a computation of from the regression coefficients, recovery of the means from the intersection of the two lines, a property question, or a statement about causation. Index numbers appear as a Laspeyres or Paasche computation, or as a question on which index satisfies which test.
The recurring errors are treating mutually exclusive events as independent, confusing the two regression lines, using current-year quantities in Laspeyres, and inferring causation from correlation. Where a formula must be selected, identify first what is being weighted by what — that single question resolves both the regression and the index number choices.
