9.11 Summary

  • Two decisions precede any estimate: the form of the relationship and which variables enter. Neither is settled by the data.
  • A level–level slope depends on the units of both variables. Logarithms remove that dependence and often improve the approximation as well.
  • The four forms are level–level, log–level, level–log and log–log, with slopes read as units, per cent, units and per cent respectively. All follow from \(\Delta \ln y \approx \Delta y / y\).
  • A log–log slope is an elasticity and is unit-free. Cars: a 1% heavier car is 0.83% less fuel-efficient.
  • The log–level form is standard for wages. A year of schooling is worth about 9.4% before controls and 7.6% after them.
  • The percentage reading is an approximation. Above about 0.2, use \(100(e^{\beta} - 1)\).
  • A fan-shaped residual plot has two possible meanings: non-constant variance, or a badly scaled variable. In the earnings data, logs cut the ratio of residual spreads from 3.6 to 1.01, so it was the second.
  • A squared term captures diminishing returns. The model stays linear in the parameters, but the marginal effect becomes \(\beta_1 + 2\beta_2 x\), so \(\beta_1\) alone is meaningless — it is the effect at \(x = 0\). The turning point is \(-\beta_1/(2\beta_2)\), and it must be located before the curve is believed: for Indian output and capital per worker it lies at 292.8, far beyond the largest value ever observed.
  • A log flattens without ever turning; a quadratic must turn. Choose according to whether theory predicts a genuine peak.
  • Logs are undefined at zero, and dropping the zeros silently changes the question being answered. \(\ln(x+1)\) keeps them, but the constant is arbitrary, the estimate depends on it, and the coefficient is no longer quite an elasticity. A zero is often a different kind of outcome rather than a small number.
  • An irrelevant variable causes no bias. Its cost is precision, and the cost depends on how strongly it is related to the variable of interest, not on whether it affects the outcome.
  • \(R^2\) can never fall when a variable is added, so it cannot decide whether one belongs. Ten pure noise variables raised it from 0.5421 to 0.5464.
  • Adjusted \(R^2\) charges for each parameter and fell from 0.5401 to 0.5375 on the same regression. It is a comparison device, not a measure of fit, and a weak one.
  • \(R^2\) says nothing about bias, is not comparable across different dependent variables, and is low in most cross-sectional work for reasons that have nothing to do with model quality.
  • Controls are not all alike. A confounder should be included; a mediator changes the question; a collider creates bias where none existed.
  • Nothing in the data distinguishes the three. The decision rests on an argument about what causes what, made before the regression is run.