Governments and foundations test new approaches through pilot programs. Successful pilots frequently fail to become permanent policy, and the reasons are largely structural.

Pilots run under favorable conditions

A pilot typically has committed staff, close attention from leadership, dedicated funding and a site chosen partly because it was ready to participate.

Those advantages are not incidental to the result. They are part of why the program worked, and none of them is guaranteed at scale.

Evaluations that measure the intervention while treating the surrounding conditions as background can therefore overstate what the same design would achieve elsewhere.

Selection shapes the population served

Participants often enroll voluntarily, and people who enroll differ from those who do not in motivation, stability and awareness of the program.

A design that works well for people who sought it out may perform differently when applied to everyone eligible, including those hardest to reach.

Some pilots address this by randomizing or by enrolling automatically, though both approaches add cost and complexity that pilot budgets do not always cover.

Funding changes character at scale

Pilots are usually funded from grants or one-time appropriations. Permanent programs require recurring money in a budget that already has claimants.

Moving from a demonstration line item to an ongoing entitlement or appropriation is a political decision, not an administrative one, and it competes with everything else in the budget.

A program that costs little in one county may look very different when multiplied statewide, and the aggregate figure often changes the conversation entirely.

Administration is where designs break

At scale a program needs eligibility rules, appeals, fraud controls, staff training, data systems and coordination with existing agencies.

Those additions can alter the intervention itself, since verification requirements and paperwork change the experience for participants and may reduce take-up.

Programs that depended on informal discretion by a small trusted team are particularly vulnerable, because discretion is difficult to standardize across many offices.

Evidence is used differently than expected

Positive results do not automatically produce adoption, and negative results do not automatically end a program, because evidence enters a process that also weighs cost, politics and administrative capacity.

Supporters and critics frequently read the same evaluation differently, disputing whether the measured outcomes were the ones that matter and whether the comparison group was appropriate.

Understanding this is useful for reading claims in either direction, since a pilot result is one input into a decision rather than a determination of what policy should follow.