1.8 Before any mathematics can help
What must be true of the data before any of this applies?
Everything in this unit assumes that the sample carries information about the population we care about. That assumption is not guaranteed by any formula, and when it fails, no amount of statistical sophistication will announce the failure. The arithmetic runs perfectly on data that cannot answer the question.
Four ways it fails, none of which any calculation will warn you about.
The sample is not from the population you mean. The Annual Survey of Industries covers registered factories. Most Indian manufacturing employment is unregistered. Conclusions about “Indian manufacturing” drawn from that source are conclusions about a formal sliver of it, however large the sample.
The sample is not random. The class survey in Section 1.2 has thirty-two heights. They are not a random sample of students at the university, still less of young Indians: they are the people who took the course and attended that day. A confidence interval computed from them is arithmetically correct and tells you nothing about anyone else.
The data are measured badly. Incomes reported from memory are systematically understated. Rainfall attributed to a district may have been recorded at a station a hundred kilometres away. National accounts are revised for years after first publication. A sampling distribution describes how an estimate varies across samples; it says nothing about whether the quantity was measured correctly in the first place. Careful statistics performed on poor measurements remain poor statistics.
The question is causal but the data are not. Suppose states with more industry have higher literacy. Does industry raise literacy? Does literacy attract industry? Does something else drive both? The correlation is identical in all three cases, and no statistical procedure separates them without an argument about how the data came to exist.
These are not edge cases. They are the ordinary condition of applied economics, and recognising them is a skill quite separate from computation.
Good inference begins with good data and thoughtful study design, not with clever mathematics.