How to Read a Program Evaluation

By Francis Secada · March 17, 2015

The slogan that a good evaluation is built to test

"Speed Kills" is one of those phrases that feels less like a claim and more like a law of nature. Everyone knows it. Raise the speed limit and more people die; lower it and fewer do. For decades that was the conventional wisdom of highway safety, and it drove real policy — states writing and enforcing lower limits on the shared belief that slower roads are safer roads. When an intuition is that loud and that morally satisfying, it stops sounding like a hypothesis at all.

But that is exactly the kind of belief a program evaluation exists to test. I want to walk through one that did — David J. Houston's study of the 65-MPH speed limit, published in Evaluation Review in June 1999 — because it is a nearly perfect worked example of how to read any evaluation. The point is not really about highways. It is about learning to see the machinery underneath a finding, so that the next time someone hands you "the data proves X," you know which questions to ask before you believe them.

Start with the question actually being asked

The first thing to notice is that the question is more precise than the slogan. "Speed Kills" implies that raising limits raises fatalities, full stop. But there are two very different ways that could be true, and they don't mean the same thing.

Raw fatality counts can rise simply because people drive more. If a faster, more convenient road induces more total miles traveled, you could see more deaths in absolute numbers while each individual mile is no more dangerous than before. That is a volume story, not a safety story. The safety question — the one policymakers actually care about — is about the rate: fatalities per unit of exposure, not the headline body count.

So the evaluation reframed the question as a causal one worth its salt: does raising the limit increase the fatality rate? Houston measured deaths per one billion vehicle miles traveled. That single choice of denominator is where a good evaluation either earns your trust or loses it. Match the outcome to the population it's supposed to describe, and you're measuring danger. Match it to the wrong one, and you're just measuring traffic.

The design: a natural experiment, disciplined

Here is the elegant part. You cannot run a randomized trial on speed limits — you can't assign half the country to drive fast and half to drive slow. But states adopted the higher limit at different times and on different roads, which hands the evaluator a natural experiment: staggered adoption across states, observed over time. Houston used a pooled time-series model, grouping the data into road categories and measuring across the interval — 750 observations in all.1

The categories are the whole game, and they're worth stating plainly. He didn't just look at "roads." He separated:

  • rural interstate highways (where the higher limit actually applied),

  • rural non-interstate roads,

  • all roads except rural interstates, and

  • all roads combined.

Why bother splitting it four ways? Because a policy that only changes the speed limit on rural interstates should, if it does anything, show its clearest effect on rural interstates. Lumping every road together would let a big effect on the affected roads get diluted — or hidden — inside a national average. And comparing the affected roads against the unaffected ones is precisely how you catch spillover: deaths that don't disappear so much as move somewhere else.

The outcome data came from FARS — the Fatality Analysis Reporting System, compiled by the National Highway Traffic Safety Administration.2 That provenance matters more than it looks. FARS is assembled by a government body strictly for research, which is a very different thing from numbers produced by a private firm or an advocacy group with a conclusion to sell. When you read an evaluation, ask where the outcome data came from and who had an interest in how it turned out. Rigorous, disinterested data is the floor the whole analysis stands on.

And crucially, the model controlled for the things that move fatalities but have nothing to do with speed limits: seatbelt laws (as a dummy variable), the minimum legal drinking age, police and safety spending per capita, the share of drivers aged 18–24 as a proxy for young drivers, per-capita alcohol consumption, population density, temperature, and income. This is the counterfactual made concrete. If you don't hold the background conditions still, you can credit the speed limit with an improvement it had nothing to do with, or blame it for one it didn't cause. Controlling for what else moves the outcome is how you isolate the signal from the noise.

Where the slogan breaks

Now the finding. On rural interstate highways — the roads that actually got the higher limit — fatalities did rise. So far, "Speed Kills" looks vindicated. But on the other road categories, fatalities fell significantly, enough that the net effect came out as a negative correlation between the 65-MPH limit and total traffic fatalities.

Read that carefully, because it is the crux. The intuition wasn't exactly wrong about the affected roads; it was wrong about the whole. What appears to be happening is a spillover: raising the interstate limit produces a marginal benefit on the rural non-interstate roads, likely because a faster interstate pulls traffic onto it, and interstates — with more continuous driving and less stopping — carry their own fatality dynamics. Deaths rise in one category while falling more in the others. The system's overall fatality rate drops even as the affected road gets more dangerous.

This is the distinction every reader of evaluations has to be able to draw: is a change a real effect, or is it displacement — a shifting-around of the same underlying activity? They look identical if you only stare at one category. They tell opposite policy stories. If lowering the interstate limit just pushes drivers back onto roads where the aggregate outcome is worse, then "slow the highways down" isn't the safety win it looks like. Houston's finding — consistent with Lave and Elias (1994), and backed by the usual diagnostics of R², F-tests, and Hausman tests — is that lowering highway speed limits is simply not the most effective way to reduce fatalities.3 States chasing safety would do better to aim their resources at more effective interventions.

The same discipline, a different program

I ran into the identical trap from the other side while building a logic model for New York's Medicaid long-term-care waiver programs — the ones that keep chronically ill and disabled people in their homes instead of in nursing facilities. The tempting comparison is a nursing home's tidy Medicaid-reimbursable rate, a daily average somewhere around $280–$400 depending on region, against home-care costs that vary wildly from patient to patient. Some home patients carry daily Medicaid rates higher than the nursing-home rate, and if you stop there the program looks like a loser.

But that's the wrong denominator again. Nursing facilities also layer clinical services billed on top of the standard rate, on a fee-for-service schedule with no real cost controls for the state. The waiver programs, by contrast, run on prior approval, so the state knows every expenditure ahead of time. The honest measure isn't the most expensive individual case; it's total annual Medicaid cost divided by the number of enrolled patients across the region — average cost per enrollee, with outliers examined rather than cherry-picked. Measured that way, the aggregate can come out well below facility care even with some pricey individual cases. Same lesson as the highways: pick the unit of measurement that matches the question, and a story can completely reverse.

A reading checklist

So when the next evaluation lands on your desk claiming to have proved something, run it through the same few questions the speed-limit study answers so cleanly:

  • Match the outcome to the affected population. Is this a rate or a raw count? Does the measure describe the people and the roads the policy actually touched, or a diluted national average?

  • Find the counterfactual. What else was moving the outcome, and has the analysis controlled for it before crediting the program?

  • Distinguish displacement from a real effect. Did the problem shrink, or did it just move somewhere the first chart wasn't looking?

  • Interrogate the data's provenance. Who compiled it, and did they have a stake in the result?

  • Let the data discipline the intuition — not the other way around. The louder and more moral the slogan, the more you should want to see it survive a design built specifically to break it.

"Speed Kills" is a good instinct and a bad analysis. The value of a real evaluation is that it takes an intuition we're all sure of, builds a design capable of proving it wrong, and then tells us what actually happened — even when the answer is more complicated, and less satisfying, than the bumper sticker.

Notes

  1. David J. Houston, study of the 65-MPH speed limit, "Evaluation Review," Vol. 23, No. 3 (June 1999) — a pooled time-series design of 750 observations measuring fatalities per one billion vehicle miles traveled.
  2. Fatality Analysis Reporting System (FARS), compiled by the National Highway Traffic Safety Administration.
  3. Lave and Elias (1994) — corroborating pooled time-series findings on highway speed limits and fatalities.