Summary
- Two decisions precede any estimate: the form of the relationship and
which variables enter. Neither is settled by the data.
- A level–level slope depends on the units of both variables. Logarithms remove
that dependence and often improve the approximation as well.
- The four forms are level–level, log–level, level–log and log–log, with slopes
read as units, per cent, units and per cent respectively. All follow from
\(\Delta \ln y \approx \Delta y / y\).
- A log–log slope is an elasticity and is unit-free. Cars: a 1% heavier car
is 0.83% less fuel-efficient.
- The log–level form is standard for wages. A year of schooling is worth
about 9.4% before controls and 7.6% after them.
- The percentage reading is an approximation. Above about 0.2, use
\(100(e^{\beta} - 1)\).
- A fan-shaped residual plot has two possible meanings: non-constant variance,
or a badly scaled variable. In the earnings data, logs cut the ratio of
residual spreads from 3.6 to 1.01, so it was the second.
- A squared term captures diminishing returns. The model stays linear in the
parameters, but the marginal effect becomes \(\beta_1 + 2\beta_2 x\), so
\(\beta_1\) alone is meaningless — it is the effect at \(x = 0\). The turning
point is \(-\beta_1/(2\beta_2)\), and it must be located before the curve is
believed: for Indian output and capital per worker it lies at 292.8, far
beyond the largest value ever observed.
- A log flattens without ever turning; a quadratic must turn. Choose according
to whether theory predicts a genuine peak.
- Logs are undefined at zero, and dropping the zeros silently changes the
question being answered. \(\ln(x+1)\) keeps them, but the constant is arbitrary,
the estimate depends on it, and the coefficient is no longer quite an
elasticity. A zero is often a different kind of outcome rather than a small
number.
- An irrelevant variable causes no bias. Its cost is precision, and the cost
depends on how strongly it is related to the variable of interest, not on
whether it affects the outcome.
- \(R^2\) can never fall when a variable is added, so it cannot decide whether
one belongs. Ten pure noise variables raised it from 0.5421 to 0.5464.
- Adjusted \(R^2\) charges for each parameter and fell from 0.5401 to 0.5375 on
the same regression. It is a comparison device, not a measure of fit, and a
weak one.
- \(R^2\) says nothing about bias, is not comparable across different dependent
variables, and is low in most cross-sectional work for reasons that have
nothing to do with model quality.
- Controls are not all alike. A confounder should be included; a
mediator changes the question; a collider creates bias where none
existed.
- Nothing in the data distinguishes the three. The decision rests on an argument
about what causes what, made before the regression is run.