Blog
L&D & Training
September 21, 2026

How to Evaluate a Training Program

Learning and Development EvangelistΒ at Synthesia

Create AI videos with 240+ avatars in 160+ languages

Most L&D teams can tell you how many people participated in a training, and if people liked it, but what they can't tell you is whether it changed behavior. (Trust me, I know.)

I used to facilitate an in-person manager development program. On the last day, my team would hand out surveys (and cupcakes). The survey asked questions like, "What was most useful?" and "What would have made this program better?"

We accepted this one-time input as the way we measured training effectiveness, which in hindsight, wasn't great. Our team was stretched thin and had little bandwidth to do much more. But we were investing so much in this program, that we needed a better way to understand the program's impact.

We knew people liked the program (or maybe just the cupcakes). What we didn't know was whether they managed differently after the program. Those are questions we should've been answering all along. And we would have been if we used a model like Kirkpatrick's.

How to evaluate trainingΒ 

Everyone knows you're supposed to evaluate training. It's embedded in instructional design frameworks, like ADDIE and SAM, and people like me seem to talk about it ad nauseam. So why aren't we, as L&D practitioners, better at it?

I have a few theories, but the primary one is that we overcomplicate evaluation. Sometimes, I just need to know whether to keep a training video in my LMS' content library. I need enough information to make an informed decision, and my measurement approach should reflect that.

That's why I use the Kirkpatrick Model. Here's how it works.

Level 1: Reaction

The degree to which the target audience finds an experience or initiative favorable, engaging, relevant, and supportive of the work needed to achieve the targeted outcomes.

Level 1 is where most L&D teams start, and stop, their evaluation of training. This often takes the form of surveys, pulse checks, and live polls.

Research shows that there's little correlation between participant sentiment and learning transfer (you don't need to like a training for it to be effective). But that doesn't make this data useless.

I like Reaction data, especially when I'm piloting new training. I mean, why shouldn't we strive for learning to be enjoyable, or at the very least, engaging? I'm more likely to apply something I learned when I'm actively engaged, than when I'm a passive participant. If participants are checking out during a training, or find it irrelevant to their roles, that's input I can use to reconsider the design and delivery format of the training.

Reaction data may not revolutionize my approach to mandatory training, like compliance or safety training, but it certainly shapes how I think about optional development opportunities.

Level 2: Learning

The degree to which the target audience acquires the intended knowledge, skills, attitude, confidence, and commitment necessary to achieve the targeted outcomes.

Level 2 is where evaluation starts to get more challenging. Ideally, you're measuring the transfer of knowledge or skills using a pre- and post-assessment.

This could be a traditional assessment, like an eLearning certification program that has you take a quiz before you start and at the end of the program. This is the cleanest, and easiest way to conduct Level 2 evaluation, but it's not always that simple.

Where things can get more complicated is in the evaluation of skills, which require either someone to conduct the pre- and post-assessment, or the use of a more scalable model (like our Roleplay Sessions).

Measuring things like "attitude, confidence, and commitment" is much more subjective and, I think, nearly impossible to assess consistently. Someone may present their knowledge or skill confidently in a controlled learning environment and then struggle when asked to replicate it on the job. Account for the gap between confidence and performance in how you evaluate this level.

Level 3: Behavior

The degree to which the target audience performs the critical behaviors in its environment and is supported in and accountable for its performance.

Level 3 evaluation is about assessing what happens once someone goes back to work. This could look like a new hire's time-to-competency after onboarding, or how a manager's direct reports rate their psychological safety after a manager development program.

I'm not going to sugarcoat it: Level 3 evaluation is hard. That's because behavior can be difficult to observe (unless you have a team of behavioral analysts on staff assessing daily performance). That's why L&D teams often look at business metrics (Level 4) as proxy metrics for behavior change. For instance, if a new sales hire is able to move deals more quickly through the pipeline, you might infer the objection training has been effective because they're able to better handle objections in real conversations.

Since you can't always observe behavior directly, it helps to build measurement into the program from the start. That could look like offering a structured Week 1 onboarding program, followed by reinforcement in the flow of work, and an assessment at the end of the first month by the manager.

Level 4: Results

The degree to which targeted organizational outcomes occur as a result of an experience or initiative and performance support.

Level 4 is the gold standard of L&D evaluation. It's what we all wax on about, being able to show the relationship between training and performance in terms the business understands.

That's not easy. Kirkpatrick Partners actually call it a common misconception that Level 4 is the hardest level to measure. They recommend finding leading and lagging indicators that are already tracked somewhere in your organization. In sales, for example, a leading indicator might be the number of leads generated, and a lagging indicator would be the total revenue those leads eventually produced. Ideally, those metrics correlate with your sales enablement training.

I'd push back on how simple that makes it sound. It's nearly impossible to attribute leading and lagging indicators to a training, even the best ones, with real certainty. Two reps could go through the same training, have different levels of manager reinforcement, and still put up very different numbers. How do you isolate the impact of a manager who reinforces the training more than another?

That's the real challenge of Level 4: separating what the training did from everything else happening around it. A before-and-after comparison can show that a metric moved, but it can't tell you on its own whether training caused the move, or something else did. Even the best sales teams see revenue dips that have nothing to do with L&D.

Is the Kirkpatrick Model linear?

The Kirkpatrick Model is often visualized as a pyramid (regular or inverted), suggesting a hierarchical relationship, where you go from Reaction to Results in a linear fashion. Kirkpatrick Partners recommend treating the model as a cycle, where insights gathered at each level fuel an engine of improvement.

How to decide what level of measurement you need

I use the Kirkpatrick Model to align the importance of the decision I'm trying to make with the depth of evaluation. If I run a training pilot and the group is disengaged, then that Level 1 Reaction data is usually enough for me to know I need to work on the design and delivery of the training. I don't need to wait until I scale the training to know that.

While Kirkpatrick Partners encourage you to start with Level 4 as a design principle (i.e., know your business outcome before you build the training), it isn't a mandate that you have to evaluate every training at that level. L&D teams are often working in less than ideal scenarios. You may need to evaluate a training you didn't design.

Here are a few examples of how to align your decision-making with the depth of evaluation.

Question it answersDecision to make
Level 1: ReactionDid people respond well to the training?Whether the format or delivery needs adjusting
Level 2: LearningDid they gain the intended knowledge or skills?Whether the content itself needs revising
Level 3: BehaviorAre they applying it on the job?Whether to reinforce, add support, or investigate why not
Level 4: ResultsDid it move a business outcome?Whether to scale, continue, or stop the program

What about training ROI?

In the case of the manager development program I mentioned earlier, I needed to know whether to continue running the program. I needed Level 4 data, which I certainly wasn't going to get from my participant surveys.

The reality was the organizational data I had access to was ever-changing (systems and structures). Over my tenure, data hygiene practices were slowly implemented. I'm still not sure I would've been able to find a consistent lagging and leading metric to use cohort over cohort, but I should've tried. Without that signal, it was impossible to know whether to keep investing in the program.

But that's also where training ROI comes in, because most stakeholders care about the business outcomes within the context of the financial investment. They want to know whether the costs outweighed the benefits, and vice versa.

That's why the Phillips ROI training model was created. It adds a fifth level to the Kirkpatrick Model, so you can build a case for the training's costs and benefits in terms the business already tracks.

Amy Vidor

Amy Vidor, PhD, is the Learning and Development Evangelist at Synthesia, where she researches learning trends and helps organizations apply AI at scale. With 15 years of experience, she has advised companies, governments, and universities on skills.

Go to author's profile

Frequently asked questions

What is the Kirkpatrick Model of evaluation?

The Kirkpatrick Model evaluates training at four levels: Reaction, Learning, Behavior, and Results. It assesses how people felt about the training, whether they gained knowledge or skills, and if so, whether they're applying that knowledge or skills in work, and whether doing so is having a measurable effect.

Why did Kirkpatrick create this model?

Donald Kirkpatrick was invited to publish a series of articles called "Techniques for Evaluating Training Programs" in 1959 for the American Society for Training and Development (ASTD). Each of the articles addressed a different technique focused on one of four areas: reaction, learning, behavior, and results. The discussion was based on his research as a professor at the University of Wisconsin about applied learning results.

It wasn't until 1994, when Kirkpatrick published Evaluating Training Programs: The Four Levels, that the Model became more commonly known.

Is the Kirkpatrick Model still relevant?

Yes, in 2026, the firm responsible for maintaining the legacy of the Model, Kirkpatrick Partners, even updated it to address the role of external factors impacting learning transfer and business outcomes. They also continue to evolve the adaptability of the Model to encourage enterprise performance intelligence.

Is the Kirkpatrick Model still relevant?

Learner satisfaction, or Level 1: Reaction Data in the Kirkpatrick Model, only tells you how participants felt about the training. Workplace learning research has found little correlation between how much participants liked a training and learning transfer.

How often should you evaluate a training program?

How often you should evaluate a training program depends on the complexity and maturity of the training. For instance, if you're developing a one-off training video, you'll want to measure the training shortly after launch (e.g., one week) to monitor for any improvement opportunities.

For more complex trainings, like onboarding or manager development, evaluate by cohort: run a pilot, evaluate it, then apply the same cadence to every subsequent group that goes through the program.

What's the difference between training evaluation and effectiveness?

Training evaluation is the broader process of assessing a training program's outcomes. Training effectiveness is how well a training supports learning transfer.

Video template title
Video template
Create video from template