Measurement Guide

Measuring training transfer: what the research says and what to ask

Transfer is the whole point. Not whether people enjoyed the program or passed the quiz, but whether anything different happens in their actual work, weeks later, when nobody from L&D is watching. It's also the thing evaluation most often skips, because measuring it means going back to participants after everyone has moved on.

First, about that "only 10% of training transfers" statistic

You'll meet this number in vendor decks and conference talks. It traces back to a 1982 article by Donald Georgenson, where it appears as a rhetorical question, an off-the-cuff estimate, not a research finding. Nobody measured it. It has been repeated for forty years because it's alarming and round.

The actual research picture is less quotable and more useful. Blume, Ford, Baldwin and Huang's 2010 meta-analysis found transfer varies a lot, and varies predictably: it's higher when the work environment supports it, when trainees believe the skills are useful, and when there are chances to practice. Transfer isn't a fixed leakage rate. It's an outcome with levers.

Which is good news for anyone whose job is to measure it, because the levers tell you what to ask about.

Baldwin and Ford's transfer model

The framework most transfer research builds on is Baldwin and Ford's 1988 review in Personnel Psychology. They defined transfer as two things: generalization (using what you learned in real situations, not just classroom ones) and maintenance (still using it months later). And they organized the influences into three groups.

Trainee characteristics

Ability, motivation, and belief that the skills matter. Partly settled before anyone designs anything.

Training design

Realistic practice, varied examples, principles over procedures. The part L&D controls directly.

Work environment

Manager support and the opportunity to use the skill. In study after study, the biggest lever, and the one training itself can't pull.

The uncomfortable implication: a program can be well designed, well delivered, and well received, and still produce no transfer because participants returned to managers who never mentioned it again. If your measurement only looks at the program, you'll misdiagnose that failure every time.

How to measure it with surveys

Direct observation is the gold standard and almost nobody's budget covers it. The practical instrument is a delayed self-report survey, and four design choices decide whether it produces evidence or noise.

Wait six to twelve weeks

A survey on the last day measures intentions. Transfer needs time to happen or fail. Six weeks is early enough that memory is intact, late enough that "I've used it" means something. For maintenance, a second pass at three to six months tells you whether the change stuck.

Ask about specific behaviors, not the program

"Have you applied the training?" invites a polite yes. Name the behavior instead:

"I give my direct reports specific, actionable feedback within a week of observing an issue."

Use before-and-now ratings in one sitting

Rate each behavior twice, as it was before the program and as it is now. The retrospective format avoids the response-shift problem where training changes participants' internal standards and wrecks pre/post comparisons. The before-to-now gap, per behavior, is your transfer evidence.

Measure the environment too

Since the work environment is the biggest lever, ask about it in the same survey:

  • "Since the program, my manager has discussed it with me at least once."
  • "I've had genuine opportunities to use what I learned."
  • "If you haven't used parts of it, what got in the way?" (open text)

When transfer scores are low, these questions tell you why, and the answer is frequently a finding the client needs to hear about their own managers rather than a verdict on the program.

Be straight about what this instrument is: self-reported transfer, not observed transfer. In LTEM terms, tier 7 evidence rather than proof. It costs perhaps a hundredth of an observation study, and it beats the alternative most programs choose, which is measuring nothing after the last day.

When transfer fails, find out where

A delayed survey across a cohort gives you the shape of the problem. Segment it and you get the location. Split transfer scores by facilitator, by cohort, by session, and the pattern usually points somewhere specific: one group whose manager championed the program and one whose manager didn't, one delivery that landed and one that didn't.

For the deepest version of that diagnosis, Brinkerhoff's Success Case Method uses a short screening survey to find the strongest and weakest transfer cases, then interviews both. The screening survey is the same delayed instrument described here, with names attached.

Related reading

Measure what happens after the last day

Delayed retrospective surveys with behavior-specific questions, environment questions, and per-facilitator segmentation built in.

Start your free trial

30-day free trial. No credit card required.