Systems & AdoptionApril 21, 2026·5 min read

How to Evaluate AI Training Programs (A Buyer's Checklist)

Most AI training teaches tools, measures completion, and produces no change in how anyone works. Here are the questions to ask a vendor before you buy, and the ones that separate real skill-building from expensive awareness-raising.

By Patin Team · Examples are illustrative composites

Most AI training fails in a specific and predictable way: it teaches the interface of a tool, measures whether people finished the course, and produces no detectable change in how anyone works a month later. Everyone involved reports success. Nothing is different.

If you're buying AI training for a team, these are the questions that separate programmes that build skill from programmes that raise awareness at considerable expense.

1. What specific capability will someone have afterwards that they don't have now?

Ask for the answer as a testable claim, not a topic. "Understands prompt engineering" is a topic. "Can look at a request that produced a vague answer and identify which of four things was missing" is a capability — and it's checkable.

If a vendor can't state the outcome in a form you could verify, they haven't designed for one.

2. Does anyone practise, or do they only watch?

This is the single biggest predictor. Watching someone use a tool well produces the feeling of competence without the substance. Skill comes from attempting something, being wrong, finding out why, and attempting again.

Ask what proportion of the programme is the learner doing something versus the learner watching. If it's mostly video, you're buying awareness.

3. What happens when a learner gets it wrong?

The answer to this question tells you whether the programme was built by someone who understands learning.

Weak answer: they see the correct answer. Better answer: they see why their answer didn't work, in terms of the specific thing they did. A learner who submits a vague prompt should be told which category of information was missing — not shown a model answer and left to reverse-engineer the difference.

4. Is it tied to the work these people actually do?

Generic AI training uses generic examples, and generic examples are the ones people can't transfer. A finance manager practising on a marketing scenario has to do two jobs: learn the skill, and translate it. Most people don't complete the second job.

Ask whether content adapts to role. Then ask to see it adapt.

5. How is progress measured — completion, or capability?

Completion rates measure whether people clicked through. They're easy to collect and tell you nothing. Ask whether the programme can show you that someone's work improved: scores that move, attempts that get better, specific skills that went from failing to passing.

If the only number you'll get is a completion percentage, you'll have no way to tell a successful rollout from an expensive one.

6. Does it teach judgement, or just technique?

Technique is when to use a persona and how to structure output. Judgement is knowing which tasks shouldn't be handed to AI at all, spotting when an answer is confidently wrong, and deciding what a tool is allowed to do without asking.

Technique gets stale as tools change. Judgement doesn't. A programme that only teaches technique will need replacing within a year.

7. What's the time commitment, honestly?

A two-day workshop produces a burst of enthusiasm and very little retained skill. Skill-building needs distributed practice — short sessions, repeatedly, over weeks.

Be suspicious of anything that promises transformation in a single sitting, and equally suspicious of anything demanding hours a week from people who don't have them. The realistic shape is small and frequent.

Dan — the rollout that looked like a success

Dan is head of L&D at a 900-person insurance company. He ran an AI training programme to a 94% completion rate and reported it as a win.

Three months later, an internal survey asked people how their work had changed. The most common answer was that it hadn't. A handful of enthusiasts were using AI heavily — the same people who'd been using it before the training. Everyone else had watched the videos and gone back to their week.

What he'd bought was a tour of ChatGPT's interface. What his people needed was practice at deciding which of their tasks were worth handing over. He now opens any vendor conversation by asking what the learner does in the first ten minutes. If the answer is "watches", the conversation ends.

Meera — the pilot that told her something

Meera runs enablement for a professional services firm. Rather than buying for 400 people, she bought for 12 and set one condition: at the end, each person had to show a piece of real work they'd done differently.

Four could. Eight produced examples that were essentially the same work with AI-generated filler in it. When she looked at what separated them, it wasn't tool fluency — everyone could operate the tools. It was that the four had learned to describe their task precisely, and the eight were still writing one-line requests and accepting whatever came back.

She went back to the vendor with that finding. The next cohort spent its first session doing nothing but rewriting weak requests.

The one thing

The question that predicts the outcome better than any other: what does the learner actually do, and what happens when they get it wrong?

A programme with a good answer to that will build skill even if its tool coverage is thin. A programme without one will produce a completion rate and nothing else, however polished the videos are.

Reading about it only gets you so far

Patin turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.

Just want the writing? .