The Kirkpatrick Model: 4 Levels, Examples, and Free Template
By Devlin Peck · Updated
The Kirkpatrick Model is a four-level framework for evaluating the effectiveness of training programs: Reaction, Learning, Behavior, and Results. Donald Kirkpatrick introduced it in 1959, and it remains the most widely used training evaluation model in corporate learning and instructional design.
This guide covers all four levels with worked examples and sample questions, a planning template, the updated New World Kirkpatrick Model (including the 2026 changes), the model's main criticisms, and how it compares to alternatives like Phillips ROI and LTEM.
What Is the Kirkpatrick Model?
The Kirkpatrick Model is a training evaluation framework that measures the impact of a learning program across four levels: Level 1 (Reaction), Level 2 (Learning), Level 3 (Behavior), and Level 4 (Results). Each level answers a different question about whether the training is working, and each level becomes more difficult (and more valuable) to measure than the one before it.
The model is most often applied to corporate training and eLearning programs, but it's flexible enough to evaluate any kind of learning intervention. Modern practitioners, including the Kirkpatrick family's own organization, treat the model as cyclical rather than linear, and they plan in reverse: define the Level 4 business results first, then work backward to design the training itself.
The Four Levels at a Glance
Here's how the four levels compare across the questions they answer, the methods used to measure them, and when to do that measurement.
| Level | Question Answered | Common Methods | When to Measure | Difficulty |
|---|---|---|---|---|
| 1. Reaction | Did learners find the training engaging, relevant, and useful? | Post-training surveys, pulse checks, focus groups | During and immediately after training | Low |
| 2. Learning | Did learners gain the intended knowledge, skills, attitudes, confidence, and commitment? | Quizzes, tests, demonstrations, simulations, pre/post assessments | During and at the end of training | Moderate |
| 3. Behavior | Are learners applying what they learned on the job? | Observation, performance data, 360-degree feedback, supervisor reviews | Starting within a few weeks of training; ongoing | High |
| 4. Results | Is the training producing measurable outcomes for the business? | KPIs, ROE, ROP, Contributive ROI, leading and lagging indicators | 3 to 12 months post-training, ongoing | Very High |
How Widely Used Is the Kirkpatrick Model?
Kirkpatrick is the dominant evaluation framework in corporate L&D by a wide margin, even though most organizations only apply the lower levels. The numbers tell both stories at once:
| Statistic | Figure | Source |
|---|---|---|
| Organizations that use Kirkpatrick as their learning evaluation framework, making it the most common by far | 71% | ATD Research, 2025 (n=222) |
| Organizations that evaluate the business results (Level 4) of their learning programs | 35% | ATD, Evaluating Learning: Getting to Measurements That Matter, 2016 (n=199) |
| Practitioners in the Kirkpatrick certified community worldwide | 14,000+ | Kirkpatrick Partners |
That gap between the 71% who use the model and the 35% who actually measure business results is the single most important thing to understand about training evaluation in practice. We'll come back to it in the criticisms section.
Who Created the Kirkpatrick Model?
Donald Kirkpatrick developed the model in his 1954 PhD dissertation at the University of Wisconsin and published it as a series of journal articles in 1959. The framework has been refined by his family and successors ever since.
| Year | Milestone |
|---|---|
| 1954 | Donald Kirkpatrick introduces the four levels in his PhD dissertation at the University of Wisconsin |
| 1959 | The framework is published as a series of articles in the journal Training and Development |
| 1994 | Kirkpatrick's book Evaluating Training Programs: The Four Levels makes the model the de facto industry standard |
| 2010s | Jim Kirkpatrick and Wendy Kayser Kirkpatrick launch the New World Kirkpatrick Model, emphasizing Level 3, business partnership, and planning from Level 4 backward |
| 2016 | Jim and Wendy publish Kirkpatrick's Four Levels of Training Evaluation with ATD Press, the definitive New World reference |
| 2026 | Vanessa Milara Alzate expands the model to include the performance environment and extends it beyond L&D toward enterprise performance |
The sections below walk through each level in order from 1 to 4, which makes them easier to understand. When you actually plan an evaluation, you'll work in the opposite direction, and we'll cover that workflow after the levels.
What Is Level 1: Reaction?
Level 1 (Reaction) measures how participants respond to the training: how engaging, relevant, and useful they found it. This is the most commonly collected evaluation data, usually gathered through a short post-training survey (sometimes called a "smile sheet").
One important update from the New World model: relevance predicts behavior change far better than satisfaction or engagement. A learner can enjoy a workshop and still not apply any of it. But if they tell you the content was directly relevant to their job, they're far more likely to actually use it. Kirkpatrick Partners' current guidance is explicit on this point: relevancy is the Level 1 measure that best indicates whether application and transfer will occur. Build your Level 1 surveys around relevance, not just enjoyment.
Level 1 evaluation also shouldn't wait until the end of a program. Formative pulse checks during the training, like quick polls, check-in questions, or facilitator observations, let you adjust in the moment instead of finding out a week later that learners were lost in module two.
How to Measure Level 1
- Post-training surveys (delivered via email, LMS, or in-session form)
- In-the-moment polls and pulse checks
- Short interviews or focus groups with a sample of participants
- Facilitator observations during live sessions
Sample Level 1 Questions
Use these as a starting point for your own questionnaire. Most are rated on a 1 to 5 Likert scale; the last two are open-ended.
- The training was directly relevant to my job.
- I will be able to apply what I learned in my work.
- The content was clear and well-organized.
- The instructor (or eLearning experience) kept me engaged.
- The pace of the training was appropriate.
- The practice activities helped me build confidence.
- I would recommend this training to a colleague.
- I feel confident I can use what I learned after this session.
- What part of the training was most useful to you?
- What would you change to make it more useful?
Level 1 Example: Screen Sharing Training
A technical support call center rolls out new screen-sharing software and runs a one-hour webinar teaching agents when to use it, how to initiate a session, and how to handle legal disclaimers. At the end, agents complete a short online survey rating relevance, clarity, and confidence, plus two open-ended questions on what worked and what didn't. The training team uses the results to flag the disclaimer section as confusing and revise it before the next cohort.
What Is Level 2: Learning?
Level 2 (Learning) measures whether learners actually acquired the knowledge, skills, and attitudes the training was designed to teach. This is the cornerstone of most instructional design work. It's where quizzes, demonstrations, and final assessments live.
The New World Kirkpatrick Model expands Level 2 beyond the original "KSA" (knowledge, skills, attitudes) trio. It adds two new indicators: confidence (do learners believe they can do this back on the job?) and commitment (do they intend to actually do it?). Both are strong predictors of whether Level 3 behavior change will happen, and both can be measured with a few well-placed survey questions.
Pre-tests are also worth the effort when feasible. Measuring knowledge before and after training is the cleanest way to attribute gains to the training itself rather than to existing experience.
How to Measure Level 2
- Multiple-choice quizzes and written tests (best for knowledge and cognitive skills)
- Skill demonstrations, role plays, and simulations (best for procedural and interpersonal skills)
- Pre- and post-assessments to measure gain
- Confidence and commitment surveys (e.g., "How confident are you that you can do X on Monday?")
- Case-study analyses and scenario-based questions
Role plays and demonstrations are the gold standard for measuring skills, but they historically didn't scale: someone had to sit with each learner, watch the performance, and score it by hand. AI conversation simulations have changed this. Learners practice a voice or text conversation with an AI character, and every attempt produces a full transcript plus a score against the rubric you defined. Full disclosure: this is exactly what my own tool, devlin.ai, is built for. You describe the scenario in plain English, get a working simulation, embed it in Storyline, Rise, or your LMS, and the transcripts and scores report back through SCORM or xAPI.
A question I hear from teams running these simulations: do the scores count as Level 2 or Level 3 evidence? Level 2. The practice is still happening inside a learning experience. Level 3 is when you look at and evaluate transcripts of real conversations, once people are doing the work with actual clients and customers. And remember that a simulation on its own, where the learner practices a tough conversation and gets some brief feedback, is valuable but only half the loop. The other half is what happens back on the job, which is exactly where Level 3 picks up.
Sample Level 2 Questions & Assessment Formats
Knowledge check (multiple choice):
- Which of the following is the correct first step in initiating a screen-sharing session? (A/B/C/D)
- Before sharing your screen with a customer, you must: (select all that apply)
Skill demonstration prompts:
- "Walk me through initiating a screen-sharing session with this practice customer."
- "Show me how you would handle a customer who declines the screen-sharing request."
Confidence and commitment (1 to 5 scale):
- I feel confident I can initiate a screen-sharing session correctly on my next live call.
- I plan to use screen sharing on appropriate calls starting this week.
Level 2 Example: Screen Sharing Assessment
After the call center webinar, agents complete a 10-question multiple-choice quiz on the screen-sharing process and legal disclaimers. They must score 80% or higher to receive certification. They also complete a live role play with their supervisor, initiating a session, walking through the disclaimer, and screen sharing successfully, before they're authorized to use the tool on real customer calls.
A quick contrast: for a coffee-roastery cleaning workshop, written tests don't cut it. Physical procedural skills are best measured by direct observation, like watching each operator clean the machine end to end.
What Is Level 3: Behavior?
Level 3 (Behavior) measures whether people are applying what they learned on the job. Learning something in a classroom and applying it back at your desk are two different things, and Level 3 is where the model starts producing data you can actually act on.
Two practical notes. First, measurement should begin within a few weeks of training, not after 90 days. You'll still find guides that tell you to wait 3 to 6 months before measuring Level 3. This advice is outdated: Kirkpatrick Partners' current guidance is that Level 3 measurement cannot begin at 90 days, because with too much time between learning and measurement "you risk losing the opportunity to adjust" the program. A common cadence is to start observing 2 to 4 weeks post-training and continue through the following months.
Second, behavior change doesn't happen because of training alone. Don Kirkpatrick identified four conditions necessary for behavior change: desire to change, knowledge of what to do, the right climate (a supportive manager, removed obstacles), and rewards for doing it. This framing has held up well in the academic literature; Rouse's 2011 review in Perspectives in Health Information Management walks through these conditions and the transfer barriers that block them. The New World model formalizes the idea as required drivers: the reinforcement, accountability, and support systems that turn learning into behavior, including performance support tools like job aids that carry the learning into the workflow.
How to Measure Level 3
- Direct on-the-job observation by supervisors or peers
- Performance metrics already tracked by business systems (call data, sales activity, ticket resolution times)
- 360-degree feedback from managers, peers, and direct reports
- Self-reported behavior surveys at intervals (30, 60, 90 days)
- Action plans and follow-up coaching conversations
- xAPI (Experience API / Tin Can) data for tracking informal and on-the-job activity
Sample Level 3 Questions & Observation Methods
Self-report (sent 30 and 60 days post-training):
- How often have you used screen sharing on customer calls in the past 30 days?
- What's been the biggest obstacle to applying what you learned?
- What support from your manager or team would help you apply this more consistently?
Supervisor observation checklist (sample items):
- Agent identified an appropriate moment to offer screen sharing
- Agent read the disclaimer accurately and in full
- Agent troubleshot connection issues without disengaging the customer
Level 3 Example: On-the-Job Screen Sharing Behavior
The screen-sharing software is integrated with the call center's performance management platform, so every screen-share session is logged automatically. Three weeks after training, the team pulls a report: what percentage of eligible calls included a screen share? Agents below a threshold get a coaching conversation, not a reprimand. The goal is to find out what's blocking transfer (forgot the steps? worried about customer reaction? no manager reinforcement?) and fix it.
What Is Level 4: Results?
Level 4 (Results) measures whether the training is moving the needle on the outcomes the business cares about: sales, customer satisfaction, retention, safety incidents, output, error rates. This is where training proves its worth, and it's the level most organizations skip. It's also worth asking this question before you build the program at all, which is the territory covered in whether training is the right solution in the first place.
One concept that makes Level 4 more practical is the distinction between leading and lagging indicators. Lagging indicators (quarterly revenue, annual turnover) tell you what already happened. Leading indicators (number of qualified demos booked, first-call resolution rate) predict where the lagging numbers are headed and let you course-correct earlier.
The other big shift is away from chasing a precise ROI number. Calculating the dollar return on a training program with real certainty is extremely hard; there are too many confounding variables. The current Kirkpatrick Model instead presents three complementary measures that together form what Kirkpatrick Partners calls an integrated return:
- Return on Expectations (ROE): Did the training deliver what stakeholders said success would look like? Success criteria are defined collaboratively with stakeholders at the start of the project, not imposed afterward.
- Return on Performance (ROP): The evidence layer that connects expectations to reality. ROP examines whether people actually performed the critical behaviors on the job (Level 3) and whether those behaviors contributed to the targeted outcomes (Level 4), which supports continuous improvement along the way.
- Contributive ROI (cROI): When stakeholders need a financial number, cROI acknowledges that training contributes to business results alongside leadership, systems, culture, and timing, rather than claiming sole credit for them.
Phillips' ROI Methodology (covered in the alternatives section below) takes a different approach and tries to isolate training's financial impact directly. Both views are valid; choose based on what your stakeholders find credible.
How to Measure Level 4
- Tracking organization-level KPIs against pre-training baselines
- Control-group comparisons where feasible
- ROE check-ins with sponsors against pre-defined success criteria
- Customer feedback, NPS, and CSAT data
- Operational metrics (output, error rates, safety incidents, turnover)
Sample Level 4 Metrics & Questions
For the screen-sharing initiative:
- CSAT (customer satisfaction) on calls that included screen sharing vs. calls that did not
- First-call resolution rate
- Average handle time on issues where screen sharing was used
For stakeholder ROE conversations:
- Did we hit the customer-satisfaction targets you defined at the start of the project?
- What evidence would convince you this program is worth continuing?
- What unintended outcomes, positive or negative, have you observed?
Level 4 Example: Customer Satisfaction Impact
Six months after the screen-sharing rollout, the team compares CSAT scores on calls that used screen sharing against a matched sample that didn't. Calls with screen sharing show a measurable lift in CSAT and a reduction in average handle time. Combined with the Level 3 data showing strong adoption, the training team makes a credible Contributive ROI argument: the program is one of several factors driving the CSAT improvement, and the evidence supports continued investment.
What Is the New World Kirkpatrick Model?
The New World Kirkpatrick Model, developed by Jim Kirkpatrick and Wendy Kayser Kirkpatrick, is the current version of the framework. It keeps the original four levels but fixes the biggest weakness of the 1950s version: the assumption that good training automatically leads to behavior change and business results.
The key additions:
- Plan from Level 4 backward. Every project starts by defining business results, then identifies the critical behaviors that produce those results, then the learning required for those behaviors, then the experience that delivers the learning.
- Required drivers at Level 3. Reinforcement, encouragement, rewards, and accountability systems that have to exist in the work environment for behavior change to happen. Without them, even excellent training fails to transfer.
- Confidence and commitment at Level 2. Two learning indicators that predict transfer better than knowledge alone.
- ROE, ROP, and Contributive ROI at Level 4. Stakeholder-defined and contribution-based measures of success that replace the often-unreliable hunt for a precise ROI percentage.
The model took another significant step in 2026. Vanessa Milara Alzate expanded it to include the performance environment as an explicit component: the systems, culture, and external factors that determine whether learning transfers to behavior and then to results. On the current Kirkpatrick Partners model page, Level 3 now includes critical behaviors, performance support, and the performance environment as named components. Alzate also extended the framework's usefulness beyond L&D, applying it across the entire enterprise; Kirkpatrick Partners describes this direction in The New Kirkpatrick Model: From Training Evaluation to Enterprise Performance. The practical takeaway for instructional designers: the model's center of gravity keeps moving away from "did people like the course?" and toward "does the whole system support performance?"
If you're learning the model today, learn the New World version. The original 1959 framework is foundational, but the updated model reflects what practitioners have figured out over six decades of trying to make it work in real organizations.
How Do You Use the Kirkpatrick Model?
You use the Kirkpatrick Model by planning backward: start at Level 4 with the business result, work down to the behaviors and learning that produce it, and build your evaluation instruments before the training launches. Here's the workflow most experienced evaluators follow.
Define Level 4 outcomes with stakeholders
What measurable business result is this training supposed to support? Get a specific target where possible ("sell 800,000 units in year one," "reduce safety incidents by 20%," "improve CSAT by 5 points") and document it as your Return on Expectations. Demand hard numbers, not anecdotes: people having a harder time on calls isn't the same as conversion data showing new hires underperforming against a benchmark. And stay open to the possibility that the best solution isn't training at all; sometimes it's a tool that removes the complexity so people can focus on the work that matters. Run a proper needs assessment before you commit to training.
Identify Level 3 critical behaviors
Working with subject matter experts and frontline managers, list the specific on-the-job behaviors that drive the Level 4 result. Be ruthless and focus on the few behaviors that matter most. Action mapping is built for exactly this critical-behaviors step, and Cathy Moore's Map It approach is the definitive treatment of it. Pair this with a training needs assessment that ties the request to a business metric.
Plan the required drivers
For each critical behavior, identify the reinforcement, accountability, and support that needs to exist post-training. Who reinforces it? What's measured? What's rewarded? Without this step, training won't transfer.
Define Level 2 learning objectives
What knowledge, skills, attitudes, confidence, and commitment do learners need to perform the critical behaviors? Write objectives that point directly at Level 3 behaviors. This is where the analysis work that produces solid learning objectives pays off.
Design the Level 1 experience
Build the training to be relevant first, engaging second. Make sure learners can see the link between what they're doing in the session and what they'll do back on the job.
Build evaluation instruments before launch
Write your surveys, quizzes, observation checklists, and metric dashboards before the training rolls out. If you wait until afterward, you'll miss baseline data and lose credibility with stakeholders.
Implement with formative pulse checks
Don't wait for end-of-training feedback. Check in during the program, after each module, and at intervals post-training (30, 60, 90 days).
Measure, report, and iterate
Share results with sponsors against the ROE you defined in step 1. Where the program is working, scale it. Where it's not, find out whether the issue is design, delivery, or missing required drivers, and then fix the root cause.
Step 1 is where most evaluation efforts live or die, and how far you get depends on who's in the room. In my early client projects, I would try to have this conversation, and the clients would keep bringing the metrics back to the number of courses they needed. Now that I work on more strategic projects and interface directly with Directors and VPs, we can have the Level 4 conversations about which metrics we need to impact at an organizational level. If you're stuck talking course counts, the fix usually isn't a better argument. It's getting closer to the people who own the business metrics.
This workflow pairs well with broader ID frameworks like ADDIE and other instructional design models. Kirkpatrick handles the evaluation strategy; ADDIE handles the design and development workflow.
Ready to put the workflow to work? The builder below walks you through all four levels for your specific program and hands you an evaluation plan you can copy straight into your project doc.
Interactive tool
Build your evaluation plan
Pick your program type, tell me what the business needs, and choose the instruments that fit. You'll leave with a four-level evaluation plan you can copy into your project doc.
Fair warning before you start: most teams only ever measure Levels 1 and 2, because Levels 3 and 4 are harder. But the data that matters most lives at Levels 3 and 4, and behavior change depends more on manager reinforcement than on training quality. If you can't commit to the follow-through, trim the plan rather than pretend.
Prefer to read? The full measurement menu behind this tool
| Program type | Level | What to measure | Instrument | Timing | Effort |
|---|---|---|---|---|---|
| Compliance training | 1. Reaction | Perceived relevance to the learner's actual job | Post-training survey built around relevance, not enjoyment | Immediately after training | low |
| Compliance training | 1. Reaction | Confusion points in the content | In-session pulse checks after each policy section | During training | low |
| Compliance training | 1. Reaction | What learners would change | Two open-ended questions at the end of the survey | Immediately after training | low |
| Compliance training | 2. Learning | Knowledge of the policy and required steps | Scored quiz with a pass threshold, scenario questions over recall questions | End of training | medium |
| Compliance training | 2. Learning | Ability to apply the policy in context | Short case analyses graded against a rubric | End of training | medium |
| Compliance training | 2. Learning | Confidence and commitment to comply | Two survey items: can you do this Monday, and do you intend to | End of training | low |
| Compliance training | 3. Behavior | Compliant behavior in real work | Supervisor spot-check observations with a short checklist | Starting 2 to 4 weeks post-training, ongoing | high |
| Compliance training | 3. Behavior | Process adherence captured by systems | Existing system logs and exception reports | Monthly pulls from week 4 onward | medium |
| Compliance training | 3. Behavior | Barriers to compliance | 30 and 60 day self-report survey on obstacles | 30 and 60 days post-training | low |
| Compliance training | 3. Behavior | Manager reinforcement of the policy | Required-drivers audit: is anyone reinforcing, measuring, or rewarding this | 4 weeks post-training | medium |
| Compliance training | 4. Results | Violation and incident rate | Incident reports and audit findings vs pre-training baseline | Quarterly, 3 to 12 months post-training | medium |
| Compliance training | 4. Results | Stakeholder-defined success | Return on Expectations check-in with legal or compliance sponsor | Defined at kickoff, reviewed quarterly | low |
| Compliance training | 4. Results | Audit readiness | Internal audit scores against pre-training baseline | Next audit cycle | medium |
| New-hire onboarding | 1. Reaction | Relevance of week-one content to the actual role | End-of-week-one survey focused on relevance and clarity | End of first week | low |
| New-hire onboarding | 1. Reaction | Pacing and overload | Quick pulse check after each onboarding module | During onboarding | low |
| New-hire onboarding | 1. Reaction | What's missing from the program | Short interviews with a sample of recent hires | End of onboarding | medium |
| New-hire onboarding | 2. Learning | Role knowledge and core procedures | Knowledge check aligned to the role's first real tasks | End of each module | medium |
| New-hire onboarding | 2. Learning | Ability to perform core tasks | Supervised task walkthrough or role play before going live | End of onboarding | high |
| New-hire onboarding | 2. Learning | Confidence and commitment for the first solo week | Confidence and commitment survey items | End of onboarding | low |
| New-hire onboarding | 3. Behavior | Independent performance of core tasks | Manager observation checklist on the job | 2 to 4 weeks after onboarding ends | high |
| New-hire onboarding | 3. Behavior | Ramp progress on real work | Existing performance systems: tickets closed, calls handled, tasks shipped | Weekly from week 2 through month 3 | medium |
| New-hire onboarding | 3. Behavior | Obstacles new hires hit applying the training | 30/60/90 day self-report survey | 30, 60, and 90 days | low |
| New-hire onboarding | 3. Behavior | Manager and buddy reinforcement | Required-drivers check: is the manager coaching to the same standards onboarding taught | 30 days post-onboarding | medium |
| New-hire onboarding | 4. Results | Time to full productivity | Ramp metric vs pre-program baseline | Per cohort, 3 to 6 months | medium |
| New-hire onboarding | 4. Results | Early retention | 90-day and 1-year retention vs baseline cohorts | Quarterly | low |
| New-hire onboarding | 4. Results | Early quality and error rates | QA scores or error rates for new hires' first 90 days | Monthly, first two quarters | medium |
| Sales enablement | 1. Reaction | Relevance to live deals | Post-session survey asking reps if this maps to their current pipeline | Immediately after training | low |
| Sales enablement | 1. Reaction | Engagement during rollout | In-session polls and facilitator observation | During training | low |
| Sales enablement | 1. Reaction | Field credibility of the content | Focus group with a few senior reps | Within a week of training | medium |
| Sales enablement | 2. Learning | Pitch and objection-handling skill | Recorded role play scored against a rubric | End of training | high |
| Sales enablement | 2. Learning | Product and process knowledge | Short scenario-based quiz | End of training | low |
| Sales enablement | 2. Learning | Commitment to use the new approach | Commitment survey item plus a named next deal to try it on | End of training | low |
| Sales enablement | 3. Behavior | Use of the new approach in real deals | CRM activity data: talk tracks logged, plays run, demos booked | Weekly from week 2 onward | medium |
| Sales enablement | 3. Behavior | Quality of application on live calls | Call recording reviews or manager ride-alongs with a checklist | Starting 2 to 4 weeks post-training | high |
| Sales enablement | 3. Behavior | Manager reinforcement in pipeline reviews | Required-drivers check: are managers coaching to the new approach | 4 weeks post-training | medium |
| Sales enablement | 4. Results | Leading indicators of revenue | Qualified demos booked, proposals out, stage-conversion rates | Monthly from month 1 | medium |
| Sales enablement | 4. Results | Win rate and quota attainment | Lagging sales KPIs vs pre-training baseline | Quarterly, 3 to 12 months | low |
| Sales enablement | 4. Results | Stakeholder-defined success | Return on Expectations check-in with the sales VP | Defined at kickoff, reviewed quarterly | low |
| Technical skills | 1. Reaction | Relevance and pacing for mixed skill levels | Post-training survey with relevance and pace items | Immediately after training | low |
| Technical skills | 1. Reaction | Confusion points in hands-on sections | Pulse checks after each lab or practice block | During training | low |
| Technical skills | 1. Reaction | Perceived usefulness of the practice environment | Two open-ended survey questions on the labs | Immediately after training | low |
| Technical skills | 2. Learning | Hands-on skill with the tool or procedure | Skill demonstration: perform the task end to end, observed | End of training | high |
| Technical skills | 2. Learning | Knowledge gain over baseline | Pre- and post-assessment with matched items | Before training and at the end | medium |
| Technical skills | 2. Learning | Confidence to use the skill unsupervised | Confidence survey items after the demonstration | End of training | low |
| Technical skills | 3. Behavior | Actual tool usage on the job | System logs and usage analytics, xAPI if you have it | Weekly from week 2 onward | medium |
| Technical skills | 3. Behavior | Quality of applied work | QA reviews or peer observation with a checklist | Starting 2 to 4 weeks post-training | high |
| Technical skills | 3. Behavior | Obstacles to applying the skill | 30 and 60 day self-report on blockers | 30 and 60 days post-training | low |
| Technical skills | 4. Results | Operational metrics the skill should move | Resolution time, error rate, or output vs baseline | Monthly, 3 to 12 months | medium |
| Technical skills | 4. Results | Downstream customer impact | CSAT or NPS on work that used the new skill vs work that didn't | 3 to 6 months post-training | high |
| Technical skills | 4. Results | Stakeholder-defined success | Return on Expectations check-in against kickoff criteria | Defined at kickoff, reviewed at 6 months | low |
| Leadership development | 1. Reaction | Relevance to each leader's real situation | Post-session survey centered on relevance items | After each session | low |
| Leadership development | 1. Reaction | Session-by-session engagement | Cohort check-ins and facilitator observation | During the program | low |
| Leadership development | 1. Reaction | Perceived applicability of each model taught | One open-ended question per session: what will you use this week | After each session | low |
| Leadership development | 2. Learning | Judgment in realistic scenarios | Case-study analysis scored against a rubric | End of program | medium |
| Leadership development | 2. Learning | Skill in coaching conversations | Observed role play of a real conversation they need to have | During and end of program | high |
| Leadership development | 2. Learning | Confidence and commitment to lead differently | Confidence and commitment items plus a written action plan | End of program | low |
| Leadership development | 3. Behavior | Changed behavior as their team experiences it | 360-degree feedback vs pre-program baseline | Baseline before, repeat at 3 to 6 months | high |
| Leadership development | 3. Behavior | Follow-through on action plans | Structured coaching conversations reviewing the plan | 30, 60, and 90 days | medium |
| Leadership development | 3. Behavior | Direct-report experience over time | Short pulse survey to each leader's team | 60 and 120 days post-program | medium |
| Leadership development | 4. Results | Team retention and engagement | Retention and engagement scores on participants' teams vs comparable teams | 6 to 12 months | medium |
| Leadership development | 4. Results | Internal mobility of participants | Promotion and readiness data vs baseline | 12 months | low |
| Leadership development | 4. Results | Stakeholder-defined success | Return on Expectations review with the executive sponsor | Defined at kickoff, reviewed at 6 and 12 months | low |
| Something else | 1. Reaction | Perceived relevance to the job | Post-training survey built around relevance, not satisfaction | Immediately after training | low |
| Something else | 1. Reaction | Confusion and drop-off points | In-the-moment pulse checks during the program | During training | low |
| Something else | 1. Reaction | What worked and what to change | Short interviews or focus group with a sample of learners | Within a week of training | medium |
| Something else | 2. Learning | Knowledge and skill gain over baseline | Pre- and post-assessment with matched items | Before training and at the end | medium |
| Something else | 2. Learning | Ability to perform the target skill | Skill demonstration or scenario assessment, observed | End of training | high |
| Something else | 2. Learning | Confidence and commitment to apply it | Two survey items: can you do this on Monday, and will you | End of training | low |
| Something else | 3. Behavior | On-the-job application of the training | Direct observation by supervisors or peers with a short checklist | Starting 2 to 4 weeks post-training | high |
| Something else | 3. Behavior | Behavior captured by existing systems | Performance data the business already tracks | Monthly from week 4 | medium |
| Something else | 3. Behavior | Obstacles to applying it | 30/60/90 day self-report survey | 30, 60, and 90 days | low |
| Something else | 3. Behavior | Reinforcement in the work environment | Required-drivers audit: who reinforces, what's measured, what's rewarded | 4 weeks post-training | medium |
| Something else | 4. Results | The business KPI this training supports | KPI tracking vs pre-training baseline | 3 to 12 months, ongoing | medium |
| Something else | 4. Results | Early signals the KPI will move | Leading indicators reviewed monthly | Monthly from month 1 | medium |
| Something else | 4. Results | Stakeholder-defined success | Return on Expectations check-in against criteria set at kickoff | Defined at kickoff, reviewed quarterly | low |
| Something else | 4. Results | Effect vs a comparison group | Control-group comparison where feasible | 3 to 6 months post-training | high |
Fair warning before you start: most teams only ever measure Levels 1 and 2, because Levels 3 and 4 are harder. But the data that matters most lives at Levels 3 and 4, and behavior change depends more on manager reinforcement than on training quality. If you can't commit to the follow-through, trim the plan rather than pretend.
Kirkpatrick Model Template and Sample Questionnaire
Use this template to plan your evaluation with the Kirkpatrick Model: for each level, define what success looks like, how you'll measure it, and when. It's the same evaluation-planning structure we teach at Peck Academy, my licensed career school for people transitioning into instructional design.
Level 1: Reaction
| Success Criteria | Planned Method(s) | Timing |
|---|---|---|
| [Insert success criteria here] | [Insert proposed method(s) here] | [Insert proposed timing here] |
Level 2: Learning
| Success Criteria | Planned Method(s) | Timing |
|---|---|---|
| [Insert success criteria here] | [Insert proposed method(s) here] | [Insert proposed timing here] |
Level 3: Behavior
| Success Criteria | Planned Method(s) | Timing |
|---|---|---|
| [Insert success criteria here] | [Insert proposed method(s) here] | [Insert proposed timing here] |
Level 4: Results
| Success Criteria | Planned Method(s) | Timing |
|---|---|---|
| [Insert success criteria here] | [Insert proposed method(s) here] | [Insert proposed timing here] |
Sample Questionnaire by Level
The sample questions earlier in this guide can be assembled into a complete questionnaire. A practical approach:
- End-of-training survey (Level 1): 8 to 10 items covering relevance, clarity, engagement, confidence, and two open-ended questions.
- Knowledge assessment (Level 2): 10 to 15 multiple-choice items aligned to learning objectives, plus 2 or 3 confidence and commitment items on a 5-point scale.
- 30/60/90-day follow-up (Level 3): 5 to 8 self-report items on frequency of application, perceived obstacles, and support received.
- Stakeholder ROE check-in (Level 4): 3 to 5 questions asking sponsors to evaluate progress against the success criteria defined at kickoff.
What Are the Criticisms of the Kirkpatrick Model?
The main criticisms of the Kirkpatrick Model are that the causal links between levels are weaker than the model implies, that organizations rarely measure Levels 3 and 4, that the model doesn't isolate training from environmental factors, and that Level 4 attribution is difficult to defend. Knowing these limitations makes you better at defending your evaluation choices to a skeptical stakeholder.
- The causal-link assumption is weak. The model implies that positive Level 1 reactions lead to Level 2 learning, which leads to Level 3 behavior, which leads to Level 4 results. Decades of research show these relationships are far weaker than the model suggests. Will Thalheimer's critique, which led him to build the LTEM alternative covered below, is the sharpest version of this argument: strong Level 1 scores tell you almost nothing about whether Level 3 transfer will happen.
- Most organizations stop at Levels 1 and 2. Because Levels 3 and 4 are harder to measure, they get skipped. According to ATD's 2016 report Evaluating Learning: Getting to Measurements That Matter, only 35% of the 199 organizations surveyed evaluated the business results of their learning programs. The data that matters most is the data least often collected.
- The model doesn't isolate training from the environment. If Level 3 behavior doesn't change, was the training bad or did the work environment fail to support it? The original model doesn't answer this clearly. The New World model's required drivers and, as of 2026, the performance environment address this directly, and it's also why the problem often isn't a lack of training at all.
- Attribution at Level 4 is hard. Business results are influenced by economic conditions, product changes, leadership, market timing, and dozens of other variables. Claiming training caused a specific outcome is rarely defensible. ROE, ROP, and Contributive ROI are more credible framings.
- The model is descriptive, not prescriptive. It tells you what to measure but not how to design effective training. It needs to be paired with instructional design models and learning science.
None of this invalidates the framework. It's still the most useful starting point we have. But pretending the model is airtight does the profession no favors.
What Are the Alternatives to the Kirkpatrick Model?
The main alternatives are the Phillips ROI Methodology, LTEM, the CIRO Model, Anderson's Value of Learning Model, and Brinkerhoff's Success Case Method. Each was developed to address a Kirkpatrick limitation or take a different angle on the same problem.
| Model | What It Adds | Best For |
|---|---|---|
| Phillips ROI Methodology | Adds a fifth level (financial ROI) and provides a specific methodology for isolating training's impact and converting results to monetary value | Organizations that need a defensible dollar-figure ROI |
| LTEM (Thalheimer) | An eight-tier Learning-Transfer Evaluation Model built specifically to fix Kirkpatrick's weak causal links, with tiers that distinguish attendance and activity from decision-making competence and transfer | Teams that want evaluation grounded in learning research rather than reaction data |
| CIRO Model (Warr, Bird, Rackham) | Four stages: Context, Input, Reaction, Output. Adds front-end analysis (context and input) that Kirkpatrick assumes you've already done | Programs that need formal needs analysis built into the evaluation framework |
| Anderson's Value of Learning Model | Three-stage model emphasizing strategic alignment of learning with business priorities before measurement | Senior L&D leaders aligning a training portfolio with strategy |
| Brinkerhoff's Success Case Method | Identifies the most and least successful cases and studies them in depth, rather than averaging across all participants | Quickly understanding what makes training transfer (or fail to transfer) in real conditions |
In practice, most organizations end up using Kirkpatrick as the spine and borrowing from Phillips (for ROI), LTEM (for rigor about what learning data actually means), CIRO (for front-end context), or Brinkerhoff (for case-based insight) when the situation calls for it.
Frequently asked questions
What are the four levels of the Kirkpatrick Model?
Level 1 (Reaction) measures how learners respond to the training. Level 2 (Learning) measures what they know and can do. Level 3 (Behavior) measures whether they apply it on the job. Level 4 (Results) measures the impact on business outcomes.
What is the Kirkpatrick Model used for?
The Kirkpatrick Model is used to evaluate whether training programs work: whether learners found them relevant, learned from them, changed their on-the-job behavior, and produced business results. Practitioners also use it in reverse as a planning tool, defining the Level 4 business outcome first and designing the training backward from it.
Who developed the Kirkpatrick Model and when?
Donald Kirkpatrick developed the model as part of his 1954 PhD dissertation at the University of Wisconsin and published it through a series of articles in 1959. His son Jim Kirkpatrick and daughter-in-law Wendy Kayser Kirkpatrick updated it into the New World Kirkpatrick Model in the 2010s, and Vanessa Milara Alzate expanded it in 2026 to include the performance environment.
When should each level be measured?
Level 1: during and immediately after training. Level 2: at the end of training (and pre-training as a baseline where possible). Level 3: starting within 2 to 4 weeks of training, not after 90 days, and continuing over the following months. Level 4: 3 to 12 months post-training, with ongoing tracking.
What's the difference between the original and New World Kirkpatrick Model?
The New World model adds required drivers (the post-training reinforcement that enables behavior change), confidence and commitment as Level 2 indicators, an emphasis on planning from Level 4 backward, and ROE, ROP, and Contributive ROI as more practical alternatives to financial ROI. The 2026 expansion adds the performance environment as an explicit Level 3 component and extends the model beyond L&D toward enterprise performance.
What is the difference between ADDIE and the Kirkpatrick Model?
ADDIE is an instructional design process (Analysis, Design, Development, Implementation, Evaluation) that guides how you build training. The Kirkpatrick Model is an evaluation framework that measures whether the training worked. They pair naturally: Kirkpatrick supplies the evaluation strategy inside ADDIE's evaluation phase. See instructional design models and theories for how the frameworks fit together.
Is the Kirkpatrick Model still relevant?
Yes. ATD's 2025 research found it is the most common evaluation framework, used by 71% of organizations surveyed. The New World version addresses most of the modern critiques of the original, and the four-level vocabulary is universal among practitioners.
What are the main criticisms of the Kirkpatrick Model?
The main criticisms are that the causal links between levels are weaker than the model implies, that organizations rarely measure Levels 3 and 4 in practice (ATD's 2016 research found only 35% evaluate business results), that the original model doesn't account for environmental and transfer factors, and that Level 4 attribution is difficult to defend.
How does Kirkpatrick compare to the Phillips ROI Model?
Phillips adds a fifth level (financial ROI) and a specific methodology for isolating training's monetary impact. Kirkpatrick (especially the New World version) prefers Return on Expectations, Return on Performance, and Contributive ROI, which acknowledge that training contributes to outcomes alongside other factors rather than claiming sole financial credit.
Putting the Kirkpatrick Model into Practice
The Kirkpatrick Model isn't a checklist. It's a discipline. Used well, it forces you to answer two hard questions before you build a single slide: what does the business actually need? And what will people need to do differently for that to happen? Everything else follows from there.
If you only do one thing with this framework, plan from Level 4 backward. This single habit separates evaluation that informs decisions from evaluation that fills out a checkbox. And it points at where I think the whole field is headed:
We're moving from content delivery to practice delivery. And practice is where performance actually changes.Devlin Peck
Practice is also where Levels 2 and 3 finally become measurable instead of aspirational. If this way of working appeals to you, where the business result comes first and the course is just one possible means to it, you're already thinking like a performance consultant. Here's what operating as a performance consultant actually looks like.
Explore Training Evaluation & Performance
Every guide in this series, all free.
Start with analysis
- 4 Types of Analysis for Instructional Design
You've likely heard about the importance of analysis for instructional design, but what are you supposed to analyze? This article outlines four common types of analysis for instructional design and why they're useful.
- How to Conduct a Needs Assessment
Are you ready to design training for issues that training will actually solve? This article explains how to conduct a needs assessment to get to the root of human performance problems.
- What is a Training Needs Assessment?
The training needs assessment, also referred to as "needs analysis" or "front-end analysis," is a process used by organizations to identify their human performance needs. A performance consultant typically conducts a needs assessment for an organization when...
Tie training to performance
- What is Performance Consulting?
Learn more about performance consulting, a viable approach to improving the performance of employees and organizations.
- What is Performance Support and Why is it Effective?
This post explores performance support as a technique to improve human performance. It outlines why this approach is so effective, as well as some of the terminology used to discuss the different types of performance support.
- Is Lack of Training a Problem?
L&D teams run into many problems when they assume that lack of training is a problem. In this video, we explore the different causes for human performance issues.
- The Ethics of Designing Training: Should You Build That Course?
The real ethics problem in instructional design is building courses that no analysis supports. See which codes of ethics actually exist, when training is the right solution, how to push back on course orders, and the new duties AI creates.
Action mapping with Cathy Moore
- Action Mapping: Cathy Moore's Model Explained with Examples
Action mapping is Cathy Moore's goal-first training design model. Learn the five steps, see a worked example, and compare it with ADDIE.
- Book Review: Map It by Cathy Moore
This review explores Cathy Moore's flagship book: Map It, The Hands-On Guide to Strategic Training Design. I discuss the book's impact on the industry, its core tenets, and its strengths and weaknesses.