From framing a policy problem to weighing options, allocating scarce resources, and judging effects and equity. A rigorous, interactive guide to the disciplined analysis that turns evidence into decisions — built for graduate students, evaluators, and educational leaders who must act under real constraints.
Policy analysis asks which course of action best serves a public purpose before we act; program evaluation asks whether what we did actually worked, for whom, and at what cost after we act. Together they form a continuous loop of reasoned decision-making under uncertainty.
The stages heuristic (Lasswell, Jones, Anderson) frames policymaking as a cycle. Real politics is messier and iterative — but the stages give you a spine for locating where a decision sits and what analysis it needs.
Name the condition, establish that it is a public problem, and frame it. How a problem is defined largely determines which solutions become thinkable.
The problem competes for scarce attention among policymakers and the public. Focusing events, framing, and problem/politics/policy streams (Kingdon) open a window.
Alternatives are designed and analyzed — instruments, mandates, incentives, capacity-building — and their likely consequences projected and compared.
An option is authorized through legislative, regulatory, or administrative choice, acquiring legitimacy and the resources to proceed.
The policy is enacted by agencies, schools, and street-level actors whose discretion shapes what the policy actually becomes on the ground.
Effects, costs, and equity are assessed; findings feed back into maintenance, revision, or termination — and the cycle begins again.
The disciplined study of decisions builds in four movements — from the logic of inquiry itself, through quantitative and qualitative evidence, to the applied craft of program analysis and policy evaluation. Select a strand to see its purpose, core tools, and the questions it answers.
Establishes how we come to warranted knowledge for decisions: research questions, concepts and their measurement, the difference between association and cause, validity of inference, and the ethics of studying people. It is the grammar the later strands speak.
Central question: "How do we reason from evidence to a defensible claim we can act on?"
Turns numbers into evidence about scope, difference, and effect: descriptive statistics, sampling and inference, regression, and the design logic behind experiments and quasi-experiments used to estimate causal effects of programs and policies.
Central question: "How much, how many, and does the difference we see exceed chance and rival explanations?"
Explains how and why a policy plays out as it does. Through interviews, observation, and documents it recovers meaning, mechanism, and context — indispensable for understanding implementation, unintended consequences, and the experience of those a policy touches.
Central question: "What is actually happening here, how do people make sense of it, and through what mechanisms?"
Integrates the prior strands into applied practice: articulating a program's theory of change, framing evaluation questions, choosing a design fit for the decision, appraising costs and benefits, and judging merit, worth, and equity to inform a real choice.
Central question: "Did this program work — for whom, at what cost, and what should we do next?"
A logic model is the backbone of evaluation: it makes explicit how resources become activities, activities become results, and results — if the theory holds — become impact. Step through building one for a sample after-school peer tutoring initiative, then read any stage for its definition.
Click any stage above, or walk through the chain one link at a time.
Each stage is linked by an implicit "if we invest these inputs, then we can run these activities; if these activities produce these outputs, then these outcomes should follow." Naming the links exposes where a theory of change can break.
Outputs count what the program produced (sessions held, students served). Outcomes capture change in participants (skills, behavior, achievement). Confusing the two is the most common evaluation error.
Every model rests on assumptions (tutors attend, matches are stable) and external factors (home support, school climate). Surfacing them tells you what to monitor and what could confound your results.
No single framework fits every evaluation. Each answers a different question and enters the program's life at a different moment. Match the framework to your purpose and your stage.
| Framework | Core question | Structure | Best used for |
|---|---|---|---|
| Formative | How can we improve it while it runs? | Ongoing feedback during design/delivery | Refining a program in progress |
| Summative | Did it work; should it continue? | Judgment of outcomes after maturity | Accountability, continuation, scaling decisions |
| CIPP (Stufflebeam) | Is the program well-conceived and run? | Context · Input · Process · Product | Comprehensive, improvement-oriented evaluation |
| Kirkpatrick's four levels | Did training change behavior & results? | Reaction · Learning · Behavior · Results | Training & professional-development programs |
| Theory-based / outcomes | Do outcomes follow the change theory? | Logic model / theory of change tested | Explaining why a program does or doesn't work |
| Needs assessment | What is the gap we should address? | Gap between current & desired state | Front-end planning before a program exists |
| Utilization-focused (Patton) | Will intended users actually use it? | Designed around specific users & uses | Ensuring findings inform real decisions |
| Developmental (Patton) | How do we adapt as the program evolves? | Real-time feedback amid emergence & complexity | Innovations and systems change without a fixed model |
| Empowerment (Fetterman) | Can stakeholders evaluate themselves? | Participants own the evaluation & build capacity | Community programs; self-determination & ownership |
| Goal-free (Scriven) | What did it actually do — intended or not? | Evaluator brackets stated goals; looks at all effects | Surfacing side effects and unintended consequences |
| RE-AIM | Will it work at scale in the real world? | Reach · Effectiveness · Adoption · Implementation · Maintenance | Judging external validity and scalability of an intervention |
Assess needs, assets, and problems in the setting to set defensible goals and priorities.
Appraise competing strategies, work plans, and budgets for feasibility and fit.
Monitor how faithfully and how well the plan is being carried out in practice.
Measure and interpret outcomes — intended and unintended, short and long term.
Bardach's "eightfold path" ends by projecting outcomes and confronting trade-offs. These five criteria are the yardsticks analysts apply — rarely satisfied all at once, which is precisely why the trade-offs must be made explicit rather than hidden.
The extent to which the option actually achieves the intended goal or reduces the problem.
Benefits relative to costs — cost-benefit (dollar-valued) or cost-effectiveness (cost per unit of outcome).
How burdens and benefits are distributed across groups — fairness, not just aggregate gain.
Administrative, technical, and fiscal capacity to actually implement the option as designed.
Whether relevant stakeholders and authorizers will tolerate or support the option — the criterion that determines whether a technically sound policy ever survives contact with politics.
Policy is enacted through budgets. How resources are raised, deployed, and accounted for expresses a system's real priorities — and whether stated commitments to equity survive the ledger.
An aggregate "it works" can conceal that a policy worked for some and failed others — or widened the very gaps it meant to close. An equity lens is not an add-on; it changes what counts as success.
Inequities are produced and reproduced by policies, funding structures, and institutional routines — not only individual attitudes. Analysis must interrogate the system, not just the actor.
Break every result down by race, income, disability, language, and gender. An improved mean can mask a widening disparity beneath it.
Equality gives everyone the same; equity gives each what they need to reach a common standard. Adequacy and vertical equity operationalize the difference.
Policies ripple. Incentives get gamed, resources shift, and burdens fall unevenly. Evaluate for the effects you didn't intend, not only the ones you promised.
Eugene Bardach's widely taught sequence gives a repeatable spine for an analysis — iterative in practice, but disciplined in structure.
State it as a deficit or opportunity, quantified where possible, framed to admit real solutions — and beware defining the solution into the problem.
Gather data and prior research that bear on the problem's magnitude and on what has worked elsewhere. Spend effort where it changes the answer.
Lay out a realistic menu of policy options — including the status quo as a baseline — before narrowing.
Choose the evaluative standards — effectiveness, efficiency, equity, feasibility, acceptability — that will discipline the comparison.
For each alternative, forecast what would actually happen against each criterion — the hardest and most honest step.
Where no option dominates, make the trade-offs across criteria explicit and comparable.
Apply judgment to reach a recommendation — and test it by being able to defend it to a skeptic.
Communicate the analysis clearly to the decision-maker — the "New York taxi driver" test of coherence and candor.
A policy goal is inert until it is attached to an instrument — a mechanism through which the state acts on behavior. Christopher Hood's Tools of Government sorts them by the resource they draw on (the NATO scheme); Vedung's shorthand calls them carrots, sticks, and sermons. Most real policies bundle several. Naming the instrument disciplines the analysis of feasibility and side effects.
Money that rewards or penalizes: grants, per-pupil funding weights, performance bonuses, vouchers, tax credits. Powerful but costly, and prone to gaming when tied to a single indicator.
Legal command: mandates, standards, licensure, class-size caps, accountability sanctions. Fast to enact and low-cost to the treasury, but blunt and dependent on monitoring and compliance.
Persuasion and disclosure: report cards, public data, awareness campaigns, best-practice guidance, nudges. Cheap and light-touch, but weak where interests or capacity block a response.
The state acts through its own people and organizations: running schools, deploying coaches, building data systems, and training the workforce that any reform ultimately depends on.
Both ask whether an option is worth its cost, but they answer differently. Cost-effectiveness (CEA) compares cost per unit of a single outcome, so outcomes need not be dollar-valued. Cost-benefit (CBA) converts every effect to money, allowing a net-benefit verdict and comparison across unlike programs. Work a small example (Levin & McEwan's method) to see how each reads the same two programs.
| Program | Cost / student·yr | Effect (reading) | CEA: cost per 0.10 SD | CBA: benefit @ $3,000 per 0.10 SD | Net benefit · B/C |
|---|---|---|---|---|---|
| A — Class-size reduction | $2,400 | +0.15 SD | $2,400 ÷ 1.5 = $1,600 | 1.5 × $3,000 = $4,500 | $4,500 − $2,400 = +$2,100 · 1.9 |
| B — Peer tutoring | $600 | +0.20 SD | $600 ÷ 2.0 = $300 | 2.0 × $3,000 = $6,000 | $6,000 − $600 = +$5,400 · 10.0 |
CEA says Program B buys a 0.10 SD gain for $300 versus $1,600 for A — roughly five times more outcome per dollar, with no need to price learning. CBA goes further: both clear their cost (net benefits are positive), but B returns $10 per $1 spent against A's $1.90. When outcomes can be credibly monetized, CBA ranks unlike options; when they cannot, CEA still discriminates cleanly.
Discounting — benefits arrive in the future and must be valued in today's terms: present value = future value ÷ (1 + r)t. A $10,000 lifetime-earnings gain realized in 10 years, discounted at 3%, is worth only $10,000 ÷ 1.0310 ≈ $7,441 now. Opportunity cost — every dollar spent on A is a dollar not spent on its best alternative use: the $2,400 for one class-size slot could instead fund four tutoring slots. Cost is what you give up, not just what you write a check for.
Feasibility and political acceptability are not abstractions — they are people with power and interests. A power/interest grid (Eden & Ackermann; Mendelow) sorts stakeholders into four quadrants, each implying a different engagement strategy. The point is to anticipate resistance and coalition before, not after, a decision.
School board, superintendent, teachers' union, state education agency. They can make or break adoption and care intensely. Engage directly, negotiate, and secure their buy-in first.
Legislature, major funders, taxpayer groups. Able to intervene decisively if roused, but not currently focused. Keep them informed enough that the policy never becomes a surprise grievance.
Parents, students, front-line teachers, community organizations. Deeply affected but individually weak; consult them, honor their voice, and recognize they can become powerful once organized.
Vendors, peripheral agencies, the general public. Low effort for now, but watch for shifts — a focusing event can move a group into a higher-stakes quadrant overnight.
The terms are used loosely, but the distinction is real and useful. A logic model is the operational map — the linear inputs → activities → outputs → outcomes → impact chain (which you built above). A theory of change is the causal argument behind it — why each link should hold, what preconditions must be in place, and what assumptions the whole edifice rests on (Weiss, 1995).
A policy is only as good as its delivery. Implementation science studies the gap between a program on paper and the program students actually receive — and gives evaluators the constructs to measure it (Durlak & DuPre; Dane & Schneider). Without them, a null result cannot distinguish a bad idea from a good idea poorly delivered.
The degree to which the program is carried out as its model intends — the core components present and adhered to. Low fidelity confounds every outcome claim.
The quantity actually delivered and taken up — sessions attended, hours engaged. A treatment nobody received cannot show an effect; measure exposure, not just offer.
Which of the intended population were engaged, how well delivery was executed, and what local adaptations were made — the difference between a faithful variant and program drift.
Bardach's eighth step — "tell your story" — usually takes the form of a decision memo. Analysis that a busy decision-maker cannot follow does not get used. The memo is a genre with a reliable spine: lead with the recommendation, show the reasoning, and be candid about what you don't know.