📍 Independent. Unsponsored. Reliable.

The Kirkpatrick Model: The 4 Levels of Training Evaluation Explained

The Kirkpatrick Model is a four-level framework for evaluating whether training actually works. It measures Reaction, Learning, Behavior, and Results, in that order, so L&D teams can trace a straight line from “people liked the …

Diagram of the Kirkpatrick Model's 4 levels: Reaction, Learning, Behavior, and Results

The Kirkpatrick Model is a four-level framework for evaluating whether training actually works. It measures Reaction, Learning, Behavior, and Results, in that order, so L&D teams can trace a straight line from “people liked the workshop” to “the business changed because of it.”

Developed by Dr. Donald Kirkpatrick in 1959, the model has since been updated twice: once as the “New World Kirkpatrick Model” in the 2010s, and again in 2026 under Kirkpatrick Partners’ current leadership. This guide covers all four levels with one running example, explains exactly what changed in each update, and is honest about where the model still falls short.

What Is the Kirkpatrick Model?

The Kirkpatrick Model is a training evaluation framework that measures a learning program’s effectiveness across four levels: how participants reacted to it, what they learned, whether they changed their behavior on the job, and whether that behavior produced measurable business results. L&D teams use it to justify training budgets, catch weak programs early, and improve course design.

Unlike a single survey score, the model treats evaluation as a chain of evidence. Reaction data helps explain learning data, learning data helps explain behavior change, and behavior change helps explain business results. Skip a level and you lose the ability to explain why results did, or did not, happen. That chain is also why the Kirkpatrick model of training evaluation is still taught in nearly every instructional design program despite being more than 65 years old.

Who Created the Kirkpatrick Model, and When?

Dr. Donald Kirkpatrick, a professor at the University of Wisconsin, introduced the four levels in a 1959 series of articles for the US Training and Development Journal, then expanded the ideas into his 1994 book, Evaluating Training Programs. The four-level structure he described then is functionally the same one taught today.

His son, Dr. Jim Kirkpatrick, and daughter-in-law, Wendy Kayser Kirkpatrick, took over stewardship of the model in the 2000s and modernized it into what they branded the New World Kirkpatrick Model. Today, Kirkpatrick Partners, the company that holds the trademark and certification programs built on the model, is led by Vanessa Milara Alzate, who pushed through a further update in 2026 that this guide covers in detail below.

What Are the 4 Levels of the Kirkpatrick Model?

The Kirkpatrick 4 levels are Reaction, Learning, Behavior, and Results, and each one answers a progressively harder question about a training program: did people engage with it, did they learn from it, did they change what they do, and did that change matter to the business. The table below maps each level to what it measures and how teams typically measure it.

Level What It Measures Typical Methods Example Metric
Level 1: Reaction Engagement, relevance, and satisfaction with the experience Post-session survey, pulse check, interviews Relevance score (percent who say they can apply it)
Level 2: Learning Knowledge, skill, attitude, confidence, and commitment gained Pre/post test, role play, teachback, simulation Skill assessment pass rate
Level 3: Behavior Whether critical behaviors are applied on the job Manager observation, call monitoring, 90-day follow-up survey Percent of learners performing the critical behavior
Level 4: Results Business outcomes the training was meant to influence Operational reporting, dashboards, leading-indicator tracking Change in the target KPI (for example, first-call resolution)

Level 1: Reaction, What Does It Measure?

Level 1 measures the degree to which learners find a training experience engaging, relevant, and worth their time, not just whether they enjoyed it. A low relevance score is an early warning sign that Level 3 behavior change is unlikely, no matter how high the satisfaction rating is.

Most organizations only measure this level, usually with an end-of-session “smile sheet.” That data is easy to collect but weak on its own. A learner can rate a workshop five stars and still never apply anything from it, which is exactly the gap Level 2 through Level 4 exist to close.

Level 2: Learning, What Does It Measure?

Level 2 measures whether participants actually acquired the intended knowledge, skills, attitude, confidence, and commitment needed to apply the training on the job. Confidence and commitment matter here because a learner can score well on a knowledge test and still lack the self-belief or intent to use it under pressure.

Typical measurement includes pre- and post-tests, but role plays, teachbacks, and scenario-based simulations are generally better predictors of on-the-job performance than multiple-choice quizzes. This is also the level where instructional design quality shows up most directly. If learning objectives and assessments do not align with the Level 3 behaviors you actually want, the test scores will look fine while behavior stays flat.

Level 3: Behavior, What Does It Measure?

Level 3 measures whether learners are performing the specific, critical behaviors the training targeted, and whether the surrounding environment supports and rewards that performance. This is where most evaluation programs quietly stop, because it requires observation weeks or months after the training ends, not a form filled out on the last day.

Measurement typically happens 30 to 90 days post-training through manager observation, call or ticket monitoring, checklists, or 360-degree feedback. Behavior change also depends heavily on “required drivers”, the coaching, job aids, manager check-ins, and incentives that reinforce the new behavior after the course ends. Training without those drivers in place is one of the most common reasons Level 3 results disappoint even when Level 2 test scores were strong.

Level 4: Results, What Does It Measure?

Level 4 measures whether the targeted organizational outcome, such as reduced cost, improved quality, higher retention, or increased sales, actually occurred as a result of the training and the performance support around it. This is the level executives care about and the one that is hardest to isolate from other business factors.

Because full Level 4 data can take a quarter or more to materialize, the New World update introduced “leading indicators”: short-term, observable signals such as early customer satisfaction shifts or reduced error rates that suggest the desired result is on track before the final numbers are in. Planning a program by starting at Level 4 and working backward is the single biggest practical shift the New World Kirkpatrick Model made, and it is covered fully in the next section.

How Does the Kirkpatrick Model Work in Practice? A Worked Example

The clearest way to see the Kirkpatrick model of training evaluation at work is to run one program through all four levels. Take a customer support team rolling out a de-escalation skills course: Level 1 asks if reps found the scenarios relevant to real calls, Level 2 tests whether they can apply the de-escalation script in a role play, Level 3 checks whether they actually use it on live calls a month later, and Level 4 tracks whether escalation and churn rates fall as a result.

At Level 1, reps complete a short survey right after the session rating relevance and engagement, not just whether the trainer was likeable. At Level 2, they are scored on a live role play against a rubric, which also captures a self-reported confidence score, a better predictor of follow-through than a written quiz alone.

At Level 3, thirty days later, quality analysts score a sample of real calls against the same rubric, while team leads run weekly coaching check-ins as the “required driver” that keeps the behavior alive. At Level 4, the team tracks a leading indicator (first-call resolution rate, visible within two weeks) alongside the lagging result the program was actually funded to move (escalation rate and customer churn, visible after a full quarter).

What Changed in the New World Kirkpatrick Model?

The New World Kirkpatrick Model, formalized by Dr. Jim Kirkpatrick and Wendy Kayser Kirkpatrick in their 2016 book, Kirkpatrick’s Four Levels of Training Evaluation, kept the original four levels but added specific new components inside each one and reversed the order in which programs are planned. It did not add a fifth level or change what the four levels are called.

The concrete additions were: Level 1 gained engagement and relevance alongside plain satisfaction, Level 2 gained confidence and commitment alongside knowledge, skills, and attitude, Level 3 was restructured around “critical behaviors” plus “required drivers” (the systems and coaching that reinforce behavior) and explicit recognition that most learning happens on the job rather than in formal training, and Level 4 gained “leading indicators” as short-term proxies for slower-moving business results.

The other major change was procedural rather than definitional: the New World model tells practitioners to design backward, starting with the Level 4 business result a stakeholder actually cares about, then working back through Level 3 behaviors, Level 2 learning objectives, and Level 1 relevance, rather than designing a course first and hoping the business impact follows. It also introduced Return on Expectations (ROE) as a way to define success in stakeholder terms up front, as an alternative to forcing every program into a hard financial ROI calculation.

Start Evaluation Before Day One

Write your Level 3 critical behavior and Level 4 target metric before you build a single slide. If you cannot name the one behavior change and the one business number a program should move, the course itself is not ready to be designed yet.

What’s New in the 2026 Kirkpatrick Model Update?

In 2026, Kirkpatrick Partners CEO Vanessa Milara Alzate extended the model again by formally adding the “Performance Environment” as a factor that sits underneath and around all four levels, not just Level 3. Kirkpatrick Partners’ own description of the model now states that the performance environment “acts as both a foundation and an influencer” that can support or block progress at every level, covering factors like leadership, culture, timing, and competing initiatives.

The 2026 update also refined how the model talks about ROI. Alongside Return on Expectations, it introduced Return on Performance (ROP), which checks whether critical behaviors were applied and actually contributed to results, and Contributive ROI (cROI), a more honest alternative to trying to prove training was the sole cause of a financial outcome. Together these reframe Level 4 around contribution and evidence rather than a single, hard-to-defend dollar figure.

The broader shift is scope: Alzate’s 2026 update on evaluation as enterprise performance, also laid out in her book Building a Culture of Evaluation, positions the framework as a tool for enterprise-wide performance intelligence, not only training evaluation. For a training-technology buyer, the practical takeaway is small: the four levels have not changed, but “Level 3” now formally includes the surrounding environment as something you should assess and report on, not treat as background noise.

How Do You Measure Level 3 Behavior and Level 4 Results?

You measure Level 3 through direct observation, manager scoring, or system data (call recordings, ticket logs, CRM activity) captured 30 to 90 days after training, and Level 4 through the operational or financial metric the program was funded to move, tracked over a full business cycle. Both levels are harder than Level 1 or 2 because the signal is slower, noisier, and easier for a busy manager to skip collecting.

Building rubrics, observation checklists, and a realistic data-collection cadence for these two levels is a deep enough topic to deserve its own resource. Our Kirkpatrick Level 3 and Level 4 practice guide walks through templates, sampling approaches, and how to get managers to actually complete observations instead of letting them lapse.

How Do You Calculate Training ROI With the Kirkpatrick Model?

The Kirkpatrick model tells you whether a result occurred at Level 4; it does not, by itself, convert that result into a dollar figure. To do that, you need a defined training ROI formula that turns the Level 4 metric into a financial return, using the cost of the program as the denominator.

In practice, that means isolating the portion of a result reasonably attributable to training (what cROI is meant to make explicit), pricing that improvement, and comparing it to fully loaded program cost. Our guide on how to calculate training ROI walks through that math step by step, including how to handle results that are only partly caused by the training itself.

How Do You Report Kirkpatrick Results to Executives and the CFO?

Executives generally want one page: the business metric that moved, the confidence you have that training contributed to it, and the cost to get there, not a breakdown of all four levels. Translating Level 1 through 4 data into that single narrative is a distinct skill from collecting the data in the first place.

Our guide to demonstrating training ROI to a CFO covers how to frame Level 3 and Level 4 evidence in financial language a finance leader will actually trust, including where to be upfront about attribution limits rather than overselling causation.

Lead With The Metric, Not The Model

Never open an executive readout with “here is our Kirkpatrick Level 4 data.” Open with the business number that moved, then use the levels underneath as supporting evidence only if someone asks how you know training caused it.

What Are the Criticisms of the Kirkpatrick Model?

The most common criticism is that Level 4 results are genuinely hard to isolate from other factors, market conditions, tooling changes, seasonality, that also affect business outcomes, which is exactly why the Phillips ROI Methodology exists as an informal fifth level focused specifically on isolating and monetizing training’s contribution.

A second, well-documented problem is that most organizations never get past Level 1. Industry estimates suggest roughly 80 percent of training events include a Level 1 reaction survey, while Level 3 and Level 4 evaluation are rare because they take longer, cost more, and require cooperation from managers outside L&D.

A sharper critique, laid out in a widely read Forbes critique of the model, argues the levels are effectively backward: learners tend to rate the training they enjoyed most highly, and that training is not reliably the training that changes behavior most. The New World model’s backward-planning fix (start at Level 4, design down to Level 1) addresses this in principle, but only if a team actually follows that order instead of designing the course first and bolting evaluation on afterward.

Kirkpatrick Model Examples by Training Type

The four levels apply the same way regardless of training type, but what counts as the “critical behavior” and the Level 4 metric changes a lot by use case. The table below gives a realistic Level 1, Level 3, and Level 4 example for four common training categories.

Training Type Level 1 Example Level 3 Example Level 4 Example
Compliance training Relevance rating on realistic scenario content Correct handling observed in audits or spot checks Reduction in policy violations or audit findings
Sales skills Confidence rating after negotiation role play Use of the sales methodology on recorded calls Win rate or average deal size change
Leadership development Engagement with case studies and peer coaching 360-degree feedback on coaching behaviors Team engagement scores or voluntary turnover
New hire onboarding Relevance of first-week content to actual role Manager checklist of core tasks performed unsupervised Time to full productivity and 90-day retention

Conclusion

If you take one thing from this guide, take the backward-planning habit. Before you build another course, write down the Level 4 business metric a stakeholder actually cares about and the one Level 3 behavior that would move it, then design the training and the Level 1 questions around those, not the other way around.

From there, pick the level your organization is weakest at measuring, most teams it is Level 3, and fix that one first rather than trying to overhaul evaluation across all four levels at once. A single, well-run behavior check 30 days after your next launch will tell you more than another year of five-star reaction scores.

If ROI reporting to leadership is the immediate blocker, start with the training ROI formula and CFO-reporting guides linked above; they are built to plug directly into the Level 3 and Level 4 data this model asks you to collect.

FAQ

Q1. What are the 4 levels of the Kirkpatrick Model?

The four levels are Reaction (did learners find it engaging and relevant), Learning (did they gain the intended knowledge, skills, and confidence), Behavior (are they applying it on the job), and Results (did it produce a measurable business outcome). Each level builds on the one before it, forming a chain of evidence.

Q2. What is the New World Kirkpatrick Model?

It is Jim and Wendy Kirkpatrick’s 2016 update to the original four levels. It kept the same four levels but added engagement and relevance to Level 1, confidence and commitment to Level 2, required drivers to Level 3, and leading indicators to Level 4, plus a rule to plan backward from Level 4.

Q3. What changed in the 2026 Kirkpatrick Model update?

Kirkpatrick Partners CEO Vanessa Milara Alzate added the “Performance Environment” as a factor influencing all four levels, introduced Return on Performance (ROP) and Contributive ROI (cROI) alongside Return on Expectations, and expanded the model’s use beyond training into broader enterprise performance evaluation.

Q4. How do you measure Kirkpatrick Level 3 behavior change?

Measure it 30 to 90 days after training through manager observation, call or ticket monitoring, checklists, or 360-degree feedback focused on a small number of specific, pre-defined critical behaviors. Pairing this with coaching or job aids, the “required drivers”, makes the behavior change more likely to stick.

Q5. What is the difference between the Kirkpatrick Model and the Phillips ROI Methodology?

Kirkpatrick’s four levels tell you whether learning and behavior change occurred and whether a business result followed. The Phillips ROI Methodology adds a fifth step that isolates training’s specific contribution to that result and converts it into a monetary ROI percentage.

Q6. Is the Kirkpatrick Model still relevant in 2026?

Yes. It remains the most widely taught training evaluation framework, and it is still being actively updated, most recently in 2026 with the addition of the performance environment concept, rather than being replaced by a newer model.

Q7. Who created the Kirkpatrick Model?

Dr. Donald Kirkpatrick introduced it in a 1959 series of articles and his 1994 book, Evaluating Training Programs. His son Jim Kirkpatrick and daughter-in-law Wendy Kayser Kirkpatrick later modernized it, and Kirkpatrick Partners, now led by Vanessa Milara Alzate, continues to update it today.

Rohan Mehta

Written by Rohan Mehta

Rohan ran operations for a mid-size commercial training company before turning to writing full-time, so his advice on scheduling, instructor logistics, and revenue-per-course tends to come from having actually lived the spreadsheet chaos he now writes about avoiding. He covers the business side of training delivery, the parts that don’t show up in a course catalog but determine whether a training company is profitable. He’s opinionated about TMS platforms and will tell you exactly why.

Table of contents