How to Measure Whether AI Training Actually Worked
Attendance and satisfaction scores tell you nothing. Four measures that do, how to capture a baseline before you start, and what a realistic return looks like.

Attendance was 94 per cent. Average satisfaction was 4.6 out of 5. Neither of those numbers tells you anything about whether the training changed how your business operates, and if they are the only numbers you have, you cannot tell whether to do it again.
This article covers four measures that do work, how to capture the baseline before you start, and what a realistic return looks like.
This article is part of our guide to AI training for UK businesses.
Capture the baseline first
This is the step almost everybody skips, and skipping it makes the rest impossible. Afterwards, everyone remembers the before state as worse than it was, and the resulting numbers are unconvincing to exactly the people you need to convince.
Two weeks before the training, spend an hour collecting four things.
Time on three named tasks. Pick three specific, frequent, real tasks that AI might help with. Writing a proposal. Summarising a client meeting. Producing the monthly report. Have the people who do them record actual minutes for a fortnight. Crude, and more convincing than any dashboard.
A five-question quiz on data handling. Anonymous. What happens to what you type into the free tier? Which of these may go into an AI tool? Who do you ask if unsure? The score will be low, which is the point.
A count of AI tools in use. Including unapproved ones. See shadow AI for how to ask that question without getting a useless answer.
Incidents and near-misses in the last six months. This will usually be zero, which is not the same as none having happened.
The four measures
1. Time on the named tasks
Repeat the fortnight of timing, four to six weeks after training. Compare.
What good looks like: 20 to 40 per cent reduction on tasks well suited to AI, and no change at all on tasks that are not. That second half matters. If every task shows an improvement, someone is reporting what they think you want.
Convert to money using fully loaded hourly cost, which for UK employees is roughly 1.3 to 1.5 times salary. If twelve people save four hours a month at £24 an hour, that is £13,800 a year against a training cost of perhaps £6,000.
2. The data handling quiz, repeated
Same five questions, same anonymity, six weeks later.
This is the measure that matters most for risk and the one nobody runs. If people cannot say what happens to what they type, the training did not do the single most valuable thing it was supposed to do, regardless of how the session felt.
Target: near-universal correct answers on what must never go into a tool, and on who to ask. Anything less and it needs revisiting, which is cheap at that stage.
3. Reported near-misses, going up
Counterintuitive and important. A rise in reported near-misses is usually good news. It means people now recognise a problem when they see one and feel able to say so.
A business reporting zero incidents before and after training has learned nothing. A business going from zero to four reported near-misses has acquired visibility it did not have, and each of those four is a problem caught early.
Judge this alongside actual incidents, which should be flat or falling.
4. Whether people can say where AI does not help
Ask an open question six weeks later: name a task you tried this on where it was not worth using.
People who can answer specifically have genuinely internalised the boundaries. People who cannot are either not using the tools or using them indiscriminately, and both need a follow-up.
This is the single best proxy for whether the training was honest. Sessions that oversell produce enthusiasm and no boundaries, and it shows up here first.
What not to measure
Satisfaction scores. They measure whether the session was enjoyable, which correlates poorly with whether anything changed.
Individual usage figures. Publishing them converts a capability question into a compliance one, and it punishes people whose work genuinely does not benefit. Aggregate is fine. Named leaderboards are counterproductive.
Prompts written, or any volume metric. Rewards activity over judgement. The best outcome for some tasks is that nobody uses AI on them.
Certificates. They record attendance. So does a register, more cheaply.
What a realistic return looks like
For a fifty-person business spending £7,000 to £19,000 in a first year, as set out in the training guide:
Measurable time savings of somewhere between £10,000 and £40,000 a year, concentrated in ten to twenty people rather than spread evenly. Most of the organisation will show little change, and that is normal.
Risk reduction that cannot be proven, because you never see the incident that did not happen. One avoided confidentiality breach exceeds the entire training cost, and you will never be able to point at it. Put it in the case as judgement rather than arithmetic, clearly labelled. See how to calculate the ROI of automating a business process for why separating the two matters.
Payback typically within four to nine months on the measurable portion alone.
Anything promising a 10x return in a quarter is selling something.
Key Takeaways
- Capture a baseline two weeks before training: task times, a data-handling quiz, tool count, incidents. Without it nothing afterwards is provable.
- Expect 20 to 40 per cent time reduction on suited tasks and no change on unsuited ones. Improvement everywhere means someone is telling you what you want to hear.
- Repeat the data-handling quiz. It is the most important measure and the one almost nobody runs.
- Reported near-misses going up is good news. It means people can now see problems and feel able to report them.
- Ask people where AI did not help. Specific answers are the best evidence that the training was honest.
Frequently Asked Questions
How soon can we measure?
Four to six weeks after training for time and knowledge measures. Any sooner and you are measuring novelty. Any later and other changes in the business start confounding the result.
What if the numbers show no improvement?
That is useful and it usually has a specific cause: the tools were not actually approved, there was no follow-up, or the tasks chosen were not ones AI helps with. Diagnose before repeating. Running the same session again rarely produces a different outcome.
Do we need this if we are only training for compliance reasons?
For Article 4 purposes you need a record of what was covered, who attended and when. The measures here go beyond that. They are worth doing anyway, because if the training did not change behaviour you have a record of compliance and an unchanged risk, which is the worst of both.
Want training with measurement built in from the start? Talk to Halo Technology Lab. Our support and training service includes the baseline capture, because without it neither of us can tell whether it worked.
Enjoyed this? Get the next one by email
Practical AI playbooks, build logs and tool teardowns. One email a week, free, unsubscribe in one click.
Related Articles
Training a Team That Thinks AI Is Coming for Their Job
Resistance is usually a reasonable response to a badly handled rollout. How to run AI training when people are worried, and why pretending nothing will change makes it worse.
Workshop, Programme or Course? Choosing the Right Shape of AI Training
A half-day workshop, an ongoing programme and a self-serve course solve different problems. What each is good at, what each costs, and how to tell which one your team needs.
AI Literacy for Non-Technical Leaders: Ten Things You Actually Need to Understand
You do not need to know how a model works. You do need to know why it makes things up, what it does with your data, and when to distrust it. The ten concepts that matter.