Evaluation Frameworks
Kirkpatrick, old and new. Phillips, Brinkerhoff, Thalheimer, Kaufman, Anderson, Stufflebeam, and a 1970 model from Warr, Bird and Rackham. The models disagree about emphasis and interpretation far more than they disagree about data. Every one of them, at some point, needs someone to ask participants structured questions and record the answers. That part is the one that usually doesn't happen.
Reaction, learning, behavior, results. The shared language of training evaluation for seventy years, which is both its strength and the reason people attack it. Most stakeholders already think in its terms whether they know the name or not.
Reach for it when: you need a framework everyone in the room already accepts. Which is most of the time.
Read the full guideThe official revision. Plan backwards from Level 4, measure the "required drivers" that make behavior change survive, watch leading indicators instead of waiting a year, and judge success as Return on Expectations rather than ROI. The practical effect is a longer, better list of questions at every level.
Reach for it when: you want Kirkpatrick's familiarity with a sharper measurement plan, or stakeholders ask about ROE.
Read the full guideKirkpatrick's four levels plus a fifth: return on investment. The distinctive machinery is isolation, separating the program's effect from everything else that changed, usually via participant estimates discounted for confidence. Produces a percentage a CFO will engage with.
Reach for it when: finance holds the budget and wants the business case in their own units. Reserve the full treatment for your most expensive programs.
Read the full guideIgnore the average, study the extremes. A short screening survey finds the most and least successful participants; interviews with both produce documented success stories and a diagnosis of what blocked everyone else. Unusually persuasive with executives, because stories with names attached beat bar charts.
Reach for it when: you need to convince leadership a program deserves continued investment, or to find out why transfer keeps failing.
Read the full guideEight tiers of evidence strength, from attendance at the bottom to organizational effects at the top. Built as a critique of how the industry actually measures: completions and smile sheets. Less a process to run than a standard to hold your evidence against.
Reach for it when: your client's L&D team is methodologically sharp and the credibility of the evidence itself is under scrutiny.
Read the full guideKirkpatrick with the first level split in two (the materials versus the delivery) and a "mega" level above organizational results for societal impact. The Level 1 split is quietly useful for question design; the mega level matters mainly where the mission is the societal outcome.
Reach for it when: the client is public-sector or mission-driven, or you want cleaner diagnostics out of your reaction questions.
Read the full guideContext, Input, Reaction, Outcome. Half the model evaluates decisions made before delivery: was the need real, was the design right. Its other lasting idea is treating participant reactions as improvement suggestions rather than verdicts.
Reach for it when: honestly, rarely as a full process. Its instincts about pre-delivery evaluation survive inside newer models.
Read the full guideEvaluates the learning function rather than individual programs: check alignment with strategy, then choose the kind of evidence (ROI, expectations, benchmarks, capacity) your decision-makers actually value, and evaluate in that currency.
Reach for it when: the client is a UK HR team, or the real question is what to measure rather than how a single program performed.
Read the full guideContext, Input, Process, Product. A decision-oriented model from the education world: evaluate the needs a program addresses, the soundness of its design, the quality of its delivery, and its outcomes, to decide what to do with it. Stufflebeam's slogan was that evaluation's purpose is "not to prove, but to improve."
Reach for it when: the work is education-sector, public-sector, or accreditation-shaped, where the question is whether to adopt, continue, or restructure a program. In corporate L&D you'll meet it rarely, but its process-versus-product distinction quietly shows up in good evaluation design everywhere.
| Model | The question it asks | What it demands of you | Where surveys fit |
|---|---|---|---|
| Kirkpatrick | Did they like it, learn it, use it, and did it matter? | Low to moderate. Levels 1-3 are practical; level 4 needs business data. | Levels 1-3 are survey territory end to end. |
| New World Kirkpatrick | Are the drivers in place, and did we meet expectations? | Moderate. Expectations must be captured up front. | Adds relevance, confidence, commitment, and required-driver questions. |
| Phillips | Was it worth the money? | High. Isolation, cost capture, data conversion. | Levels 1-4 inputs, including the estimate-attribution-confidence chain. |
| Brinkerhoff SCM | Who succeeded, who didn't, and why? | Moderate. Interview time and verification discipline. | The Phase 1 screening survey, which must be named, not anonymous. |
| LTEM | How strong is our evidence, really? | Varies by tier. Tiers 4-6 need assessment instruments. | Tier 3 done well; delayed self-report as tier 7 evidence. |
| Kaufman | Did it work, from the materials out to society? | Low to moderate; the mega level is rarely operationalized. | Split input/process reaction questions; levels 2-4 as Kirkpatrick. |
| CIRO | Were the right choices made, and did they pay off? | Moderate. Needs access to pre-program decisions. | Relevance and format questions, plus open-text improvement asks. |
| Anderson | Is learning aligned with strategy, and what evidence counts here? | Moderate. An organization-level exercise. | Stage 2 runs on the same program-level instruments as the others. |
| CIPP | What should we decide about this program? | Moderate to high. Broad scope across a program's life. | Needs assessment (context), delivery feedback (process), outcomes (product). |
Pick the model your client's stakeholders already find credible. That sounds cynical and isn't. The frameworks overlap heavily in what they measure; where they differ is in how the findings are framed and for whom. HR audiences think in Kirkpatrick. Finance audiences want Phillips. An executive sponsor who is wavering responds to Brinkerhoff stories. A methodologically serious L&D team will respect you for knowing LTEM's tiers and being straight about which ones your evidence reaches.
A consultancy can serve all of these from largely the same instrument. The reaction questions, the retrospective before-and-after scales, the delayed application questions, the open-text result descriptions: collect those well and you can report them in whichever frame the audience needs.
The expensive mistake is the opposite one: spending a quarter debating frameworks while collecting nothing. A model is a way of interpreting data. None of them work without it.
ImpactCheck is the collection layer: branded micro-surveys for the moments every model cares about. Post-program reaction questions. Retrospective before-and-after scales for learning and behavior change. Delayed surveys for application and transfer. Named surveys when a method, like Brinkerhoff's screening, needs to know who answered. Anonymous surveys when candor matters more.
It is not an interview tool, an assessment platform, or an ROI calculator, and the models that need those things need them from somewhere else. What the consultant gets from us is clean, well-asked, well-timed feedback data that fits whichever framework the engagement calls for.
Branded micro-surveys with retrospective mode, delayed sends, and named or anonymous collection. Built for consultancies that report into someone else's framework.
Start your free trial30-day free trial. No credit card required.