Unit 1 Towards Inferential Statistics

You already know how to calculate the average of a sample.

In this unit, you will learn why that average is almost certainly wrong.

At first glance, this seems like a strange claim. After all, there is only one way to calculate a sample mean, and if the arithmetic is correct, the answer must be correct as well. Yet the number you obtain from a sample is almost never exactly equal to the quantity you ultimately wish to know — the mean of the population from which that sample was drawn.

Suppose you wish to estimate the average income of households in India. Since surveying every household is impossible, you select a sample of 2,000 households and calculate its mean income. The calculation is straightforward, but the result immediately raises a question. How close is this sample mean to the true population mean?

The difficulty is that the sample you have drawn is only one of many that could have been observed. Had a different 2,000 households been selected, the sample mean would almost certainly have been different. Draw another sample, and the answer changes again. Every statistic calculated from a sample therefore reflects two things at once: the characteristics of the underlying population, and the randomness introduced by the sampling process.

This simple observation marks the boundary between descriptive and inferential statistics.

Until now, our objective has been to describe the data that we observe. Measures of central tendency, such as the mean and the median, summarise where observations are concentrated. Measures of dispersion, such as the variance and standard deviation, describe how widely they are spread. Histograms, bar charts and density plots reveal the overall shape of a distribution and help us identify skewness, outliers and other important features. Together, these tools provide a concise description of a dataset and are the natural starting point of every empirical investigation.

Most empirical research, however, is concerned not with the sample itself but with the population from which it was drawn. Economists survey a few thousand households to understand millions. Polling agencies interview a sample of voters to estimate national voting intentions. Medical researchers evaluate new treatments on a limited number of patients before recommending them to the wider population. In every case, the sample serves as evidence about a population that cannot be observed directly.

The central problem of inferential statistics is therefore straightforward to state: how can we use the information contained in a sample to learn about the population from which it came?

Throughout this course we adopt the frequentist approach to answering this question. Rather than treating the observed sample as unique, we imagine drawing many samples from the same population and calculating the same statistic repeatedly. Some sample means will be larger than others, some smaller. Understanding how statistics vary across these hypothetical samples allows us to assess how much confidence we should place in the estimate computed from the one sample we actually observe.

Before we can understand why, we need to understand where the uncertainty comes from. The answer lies in randomness.