Your regression coefficient is not a causal effect

Quick thought experiment. You regress earnings on years of education and get a positive coefficient. More education → higher earnings. Done. Done?

Now add parental income as a control. The coefficient on education shrinks. Add a measure of cognitive ability. Shrinks again. Add motivation, grit, neighbourhood quality. It keeps moving.

So which coefficient is the “real” effect of education?

This isn’t a trick question, it’s the central problem. Each specification encodes a different set of assumptions about what causes what, and the regression itself doesn’t tell you which assumptions are correct. You chose them when you chose your controls.

Why this matters more than you think

The returns to education debate consumed labor economics for decades, and it’s instructive for us because researchers weren’t making amateur mistakes. They were running careful regressions with thoughtful control variables and still getting wildly different answers. Estimates of the return to an additional year of schooling ranged from near zero to over 15%, depending on specification.

The debate only got traction when researchers reframed the question. Instead of asking “what’s the coefficient on education, conditional on controls?” they asked: “what would this specific person have earned if they hadn’t gotten that additional year of schooling?”

That second question sounds similar. It isn’t. The first is a statistical question: it has an answer for any set of controls you pick. The second is a causal question: it asks about a state of the world that didn’t happen. No amount of OLS, no matter how many controls you add, answers it without assumptions that need to be justified on their own terms.

The breakthroughs came when researchers found natural experiments: situations where some external force, unrelated to the outcome of interest, effectively randomised who got more education and who didn’t. Draft lotteries, compulsory schooling laws, quarter of birth. These aren’t designed experiments: nobody planned them for research purposes. But they create variation in education that is as good as random, which is what you need to credibly estimate a causal effect.

What happens when you get this wrong

Say the estimated return to education is 12% per additional year, based on OLS. A government uses this to justify a national tuition subsidy. The logic: if every year of education returns 12%, subsidising more years is an obvious win.

But that 12% is inflated. People who pursue more education tend to be more motivated, more able, and come from families that invest in them in ways the regression can’t observe. The regression attributes all of the earnings difference to education, but some of it, maybe a lot of it, would have shown up regardless. The true causal return might be 5%.

The government builds a business case on 12%. It funds accordingly. It sets targets accordingly. And when the program delivers returns closer to 5%, it looks like a failure. Not because education doesn’t work, but because the estimate confused “people with more education earn more” with “more education causes people to earn more.”

The program gets cut. Not scaled back or retargeted but cut, because the gap between promised and actual results erodes trust. The political narrative becomes “we invested heavily and it didn’t deliver,” when the real story is “we invested based on a number that was never causal in the first place.”

A more honest estimate up front (“the causal return is real but modest, so let’s size the investment accordingly”) would have produced a sustainable program with realistic expectations. Instead, a biased coefficient drove a policy cycle of overpromise and collapse.

Machine learning doesn’t fix this

If you’re thinking that more sophisticated methods like random forests, gradient boosting, neural networks would solve the problem, they won’t. These tools are excellent at prediction: given someone’s characteristics, what are they likely to earn? But prediction and causation are different tasks. Prediction asks “what will happen?” Causation asks “what will happen if we intervene?” ML is built for the first question, not the second. A model can perfectly predict who will earn more after additional education without telling you whether the education caused the higher earnings. In fact, ML models will happily learn the same confounded relationships that bias OLS: they’ll just do it with more flexibility and less transparency. The problem isn’t that our models aren’t complex enough. It’s that no amount of pattern recognition in observational data, no matter how sophisticated, can substitute for a credible identification strategy.

The problem in our work

We don’t write academic papers, but we make the same inferential move constantly. “Controlling for X, program Y is associated with a Z percent improvement in outcome W.” That sentence structure implies causation while using language that technically hedges. And our readers (leadership, stakeholders, etc) don’t hear the hedge. They read “program Y causes a Z percent improvement” and make decisions accordingly.

The question isn’t whether we should be more careful with language. It’s whether our analytical approach actually supports the claim we’re implying.

Ask yourself: when you chose your control variables in your last analysis, could you explain why each one was necessary and why nothing important was missing? Not statistically but conceptually. What’s the story about what causes what, and does your model reflect it?

If that story isn’t explicit, your regression is hiding your assumptions, not eliminating them.

Further reading:

  • Angrist, J. & Krueger, A. (1991). Does Compulsory School Attendance Affect Schooling and Earnings? Quarterly Journal of Economics, 106(4), 979–1014. https://doi.org/10.2307/2937954.
  • Card, D. (1999). The Causal Effect of Education on Earnings. In Ashenfelter, O. & Card, D. (Eds.), Handbook of Labor Economics, Vol. 3, Ch. 30, 1801–1863. https://doi.org/10.1016/S1573-4463(99)03011-4.
  • Angrist, J. & Pischke, J. (2009). Mostly Harmless Econometrics: An Empiricist’s Companion. Princeton University Press.
  • Angrist, J. & Pischke, J. (2014). Mastering ‘Metrics: The Path from Cause to Effect. Princeton University Press. Start here, Ch. 6 covers the returns to education example.