Unit 8 Multiple Regression: Holding Other Things Equal

Unit 7 ended with two students. One attended thirty classes, the other twenty, and the first earned the higher grade. The difficulty was that the two students differed in many other ways as well. One may have been better prepared, more motivated, more interested in the subject, or simply worked harder outside class. The regression slope faithfully described the difference in their average grades, but it could not tell us how much of that difference was due to attendance itself.

A natural response is to measure some of those other differences and compare students who are alike in every observed respect except attendance.

That is the idea behind multiple regression.

Conceptually, very little changes. We still model the conditional mean, fit the model by least squares, and measure uncertainty with standard errors, confidence intervals and hypothesis tests. The machinery is familiar.

What changes is the interpretation.

In a simple regression, the slope describes how average outcomes change as one variable changes. In a multiple regression, each coefficient describes how the outcome changes as one variable changes while the others are held fixed. That seemingly small change is what makes multiple regression one of the most useful tools in empirical economics.

It is also where many misunderstandings begin.

What does it really mean to “hold other things equal”? How does regression compare observations that are never exactly identical? When does controlling for additional variables reduce bias, and when can it make matters worse? Most importantly, does adding more variables bring us any closer to the causal comparison we wanted in Unit 7?

Those are the questions this unit takes up.