2.7 Why is precision expensive?

If larger samples are better, why does anyone stop collecting data?

One practical consequence deserves its own section, because it shapes how survey work is actually done.

Return to the square root. Doubling the sample size does not halve the standard error; it divides it by \(\sqrt{2} \approx 1.41\). To halve the standard error, the sample must grow fourfold.

The standard error of the sample mean against sample size. Almost all of the available precision is bought in the first few hundred observations.

Figure 2.4: The standard error of the sample mean against sample size. Almost all of the available precision is bought in the first few hundred observations.

Precision is therefore expensive, and grows more expensive the more of it you buy. Each additional household costs the same to interview as the last, but contributes less than the last. Going from 100 households to 400 halves the standard error; going from 400 to 1,600 halves it again, at four times the cost.

This is why national surveys run to a few thousand households rather than a few hundred thousand. Past a certain point the additional precision is not worth the additional money, and the money is better spent on the things no sample size can fix — better questions, better interviewers, better coverage of the households that are hardest to reach.

That last point is worth holding onto. The failures in Section 2.2 were all failures of design, and none of them would have been helped by a larger sample. Survey B with ten thousand urban households is still wrong. When a result is disappointing, the instinct is to ask for more data; often the honest answer is that more of the same data will not help.