Thought leadership
Competence, Confidence, Compliance: The Three Measures That Prove Practice Worked

Ask most organisations how their people development is going and you'll get an activity report. Completion rates. Attendance. Course satisfaction scores. Hours logged.
None of that tells you whether a single manager can hold a redundancy conversation without making it worse. None of it tells you whether your new AE can handle the objection that kills half your deals. None of it tells you whether the person who clicked through your anti-bribery module last March would recognise a bribe if one were offered over dinner.
This is not a new complaint. Brandon Hall Group's Learning Measurement Study found fewer than one in six organisations effectively track the business impact of learning. The Kirkpatrick model has had four levels since 1959, and most organisations still stop at levels one and two — did they like it, and did they remember it — because levels three and four, behaviour and results, have historically been too hard and too expensive to observe at scale.
The result is that the function responsible for workforce capability has, for decades, had the worst data in the building. Finance can tell you margin by product line to two decimal places. Operations can tell you throughput by shift. L&D can tell you that 94% of people finished the course.
Here is the part most buyers haven't yet clocked. AI role-play doesn't just close the practice gap. It closes the data gap. Because the practice happens as a live, spoken conversation that is scored as it goes, it produces a behavioural data set that has never previously existed inside an organisation — not because nobody wanted it, but because capturing it used to require a human assessor sitting in the room with a clipboard.
That data set resolves into three measures: competence, confidence and compliance.
Here is where it comes from, why those three are the right three, and what each one is worth to the business.
Every conversation is a data set
Start with what a session actually is.

A live session. The learner is playing the manager. The character is real-time, voice-led, and reacting to what the learner actually says.
To the person doing it, this is a ten-minute conversation. They speak, a character responds, it gets uncomfortable, they recover or they don't. It feels like an experience, not an assessment. That's the point — people behave the way they actually behave when they aren't performing for a scorer.
Underneath, every second of it is instrumented. A single session generates:
- A full transcript of what was said, by whom, in what order, with timings.
- Per-skill assessment against the behavioural framework for that scenario — banded from Novice to Expert, not a single blunt score.
- Objective tracking — was the outcome the conversation existed to achieve actually reached, or not.
- Policy adherence checks — where a policy or procedure is linked to the scenario, each requirement is scored as met, missed, or violated, with the moment it happened.
- Character state — whether the AI character's trust and composure held or broke, and at which line.
- Self-rated confidence, before and after — the same question asked twice, so the delta is attributable to this session rather than to a mood.
- Attempt history — the same person, the same scenario, over time, so improvement is measurable rather than assumed.

The output of one session. Note the two things happening at once: the learner gets coaching they can act on, and the organisation gets evidence — a score, a delta, an objective marked not met, and the exact words that caused it.
Two things are worth sitting with here.
The first is that the data is a by-product. Nobody fills in a form. No manager observes. No facilitator writes up notes at 6pm. The measurement is the same event as the learning, which is why it can run at the scale of an entire workforce rather than a nominated cohort of twelve.
The second is the sheer difference in resolution. A compliance module produces one data point per person per year: completed, yes or no, plus a quiz score that mostly measures reading comprehension. One ten-minute role-play produces a scored transcript against ten skills, an objective outcome, a set of policy checks, and a confidence delta — and most people do it more than once.
That is the shift enterprise buyers keep missing. It isn't that the avatars are impressive. It's that a conversation, run at scale and scored consistently, is the richest behavioural instrument a workforce has ever been pointed at.
The three measures

Every session in the studio rolls into the same three measures — with the exclusions shown openly, so the numbers survive scrutiny.
1. Competence: can they actually do it?
Competence is performance in the moment the skill is tested. Not whether someone knows the framework. Whether they can use it while a customer is talking over them.
The standard proxy for competence is a quiz. It's a poor one. A multiple-choice question measures recall of a correct answer in a calm room with no time pressure and no one pushing back. That is not the skill. The skill is holding structure under emotional load, choosing words while someone is crying or shouting, and recovering when the first attempt lands badly.
The alternative proxy — self-assessment — is worse. The most cited evidence here comes from medicine: a systematic review in JAMA compared physicians' assessments of their own competence against externally observed measures. Of twenty comparisons, thirteen showed little, no, or an inverse relationship. The authors also noted that the worst accuracy showed up among two groups in particular: the least skilled, and the most confident.
Read that again, because it's the whole argument for measuring more than one thing. The people least able to judge their own capability are the ones most likely to tell you they're fine.
Practice removes the proxy. You aren't inferring competence from a test score or a self-report — you're scoring the behaviour itself, against a defined bar, with the transcript attached.
Speed to competence. Competence isn't only a level. It's a rate.
For any organisation hiring, restructuring, launching a product or entering a market, the question isn't just can they do it — it's how many weeks until they can. Published benchmarks put average ramp for an account executive at around five to six months, with enterprise reps on complex deals taking considerably longer, and that number has been getting worse, not better, as buying processes have grown more complicated. Every one of those months is full salary against partial output.
The same maths applies well outside sales: contact centre agents, clinicians, claims handlers, new managers taking their first team. Time-to-productivity is one of the most expensive numbers in the business and one of the least managed — mostly because nobody could see the underlying capability curve, only the lagging output that eventually resulted from it.
Practice compresses that curve, because repetitions are the mechanism. A new hire who has run the difficult conversation thirty times before their first real one is not the same hire as one who has read about it. And when practice is scored, the curve becomes visible: you can see the median number of sessions and days a cohort takes to reach the bar, and you can see which cohorts are stalling long before it shows up in the quarterly numbers.
Ties to: ramp time and cost of ramp, quota attainment, conversion rate, average handling time, first-contact resolution, quality scores, error and rework rates, early-tenure attrition.
2. Confidence: will they do it?
Competence tells you whether someone can. Confidence tells you whether they will — willingly, promptly, and without the three days of dread beforehand.
This is not a soft measure with a hard business consequence attached; it is a hard business measure that happens to feel soft. The meta-analytic evidence on self-efficacy — the belief that you can execute a specific task — is about as robust as organisational psychology gets. Stajkovic and Luthans pooled 114 studies and found a weighted average correlation of .38 between self-efficacy and work-related performance. Belief predicts performance, particularly on complex tasks.
Managers have always known this intuitively through the skill/will matrix, popularised by Max Landsberg in The Tao of Coaching: someone can have the capability and still not use it. High skill, low will is a real and common quadrant — and low will is very often not a motivation problem at all. It's a confidence problem wearing motivation's clothes.
The behavioural cost is avoidance, and avoidance is where the damage compounds:
- The performance conversation postponed for six months until it becomes a grievance.
- The safeguarding concern nobody escalates because they aren't sure how to raise it.
- The customer objection the rep talks around instead of addressing, and the deal that quietly dies.
- The vulnerable customer whose disclosure gets processed as a transaction because the agent doesn't feel equipped to go there.
And there's the inverse failure mode, which is the more dangerous one: high confidence, low competence. Someone who feels ready and isn't. That's the person who walks into a live disciplinary certain they'll handle it and creates a tribunal claim by lunchtime.
You only see either pattern if you measure both, separately, on the same person, at the same time. Confidence measured alone is a feelings survey. Competence measured alone misses why capable people don't act. Measured together they tell you what to do: coach the confident-but-not-capable, encourage the capable-but-not-confident, and stop treating those two people as the same performance problem.
Because Real Talk Studio asks the same 0–10 readiness question immediately before and immediately after every session, the movement is attributable, repeatable and poolable across thousands of sessions — rather than an annual engagement item nobody can act on.
Ties to: attrition and regretted attrition, engagement and manager effectiveness scores, escalation and speak-up rates, grievance volumes, time-to-first-difficult-conversation, discretionary effort, internal mobility and promotion readiness.
3. Compliance: did they do it the way it has to be done?
Compliance in our terms means adherence — to a policy, a regulatory expectation, or a defined procedure for how a conversation should be run. Consumer Duty language on vulnerable customers. Safeguarding escalation steps. The bribery policy. The disciplinary procedure. The clinical consent sequence. The mandated disclosures on a sales call.
Almost all of it currently lives in a PDF that has been emailed to everyone and read by almost no one, followed by a module with a knowledge check at the end.
The problem is not that the policy is wrong. It's that a completion record and a compliant workforce are different things that produce identical paperwork. Ponemon Institute research widely cited in the risk community puts the total cost of non-compliance at roughly 2.7 times the cost of maintaining compliance — and none of that cost is avoided by the certificate. It's avoided by what someone actually says when the moment arrives.
Link the policy to the scenario and every session scores against it line by line: what was got right, what was missed, and what was an outright violation — with the transcript moment attached to each. That's not a proxy for policy adherence. It is policy adherence, observed.
There is now a legal edge to this in the UK, too. Since October 2024 the Worker Protection Act has required employers to take reasonable steps to prevent sexual harassment, with employment tribunals able to uplift compensation by up to 25% where that duty hasn't been met, and sexual harassment awards are uncapped. The EHRC's statutory Code makes clear that a policy on its own isn't sufficient. Employment lawyers commenting on the standard have been blunt about the evidential hierarchy: a short online awareness module confirming someone has read the policy is materially weaker evidence than substantive training with documented assessment. And the bar is scheduled to rise again — the Employment Rights Act 2025 lifts the duty from "reasonable steps" to "all reasonable steps" and extends it to third-party harassment.
Put a person through a simulated harassment disclosure, a bribery approach, or a vulnerable-customer restriction conversation, and two things happen that a PDF cannot deliver. First, you find out how they would actually behave — including the ones who would say the wrong thing with total confidence. Second, you generate a dated, scored, individually attributable record that the behaviour was tested, not merely published.
That record is a genuine asset. It is the difference, in an investigation or a tribunal bundle, between "we had a policy" and "we tested this person against this policy on this date, here is how they performed, and here is what we did about the gap."
Ties to: regulatory findings and audit outcomes, incident and near-miss rates, complaint volumes, tribunal exposure and settlement costs, insurance and due diligence positions, remediation cost, brand and reputational risk.
Why simulation moves these measures when content doesn't
The training transfer problem has been documented since Baldwin and Ford's foundational 1988 review: knowledge delivered in a room or a fifteen-minute click-through rarely survives contact with the job.
Simulation is the exception with the strongest evidence base behind it, mostly built in healthcare because that's where the stakes forced the investment. Cook and colleagues' JAMA meta-analysis of technology-enhanced simulation — over 600 studies — found large pooled effects on knowledge and skills and meaningful effects on behaviour and on patient outcomes. McGaghie's comparative review found simulation with deliberate practice outperformed traditional clinical education. A 2023 update comparing simulation against traditional teaching found a substantial overall effect in its favour.
The mechanism is not mysterious. Practice puts the person in the moment where the skill is tested, at volume, with feedback, and without a real human bearing the cost of the attempt that goes wrong.
What AI adds is not the pedagogy — role-play with trained actors has always worked. It's the economics and the instrumentation. Actor-based assessment is excellent and rationed: a handful of people, once a year, if the budget survives, and the output is a facilitator's written impression. AI role-play makes the same thing available to everyone, on demand, at the difficulty required, in over 200 languages — and scores it identically every time.
Consistent scoring is what converts practice from an experience into data. It's also what makes the data defensible: the same bar, applied to everyone, without the observer variance that makes human assessment hard to stand behind.
From three measures to business measures
The three Cs are only worth anything if they connect to numbers the executive team already cares about. They do, and the mapping isn't a stretch:
The right way to use this is to baseline before you scale. Take one cohort, measure the three Cs at the start, run the practice, measure again, and hold that against the operational metric you actually care about. That's a defensible ROI case rather than a vendor's assertion — and it's the case L&D has rarely been able to make.
What this looks like in practice
None of this is a bolt-on reporting module. It's what the platform produces as a consequence of people practising.
At organisation level, every session in the studio — named or anonymous, including embeds and single-use links — rolls into the three measures, with the exclusions stated openly so the numbers survive a challenge from someone who is looking for a reason to dismiss them.

At cohort level, the same three measures are filtered to identified learners, so you can compare this quarter's intake against last, one region against another, or the group who did the programme against the group who haven't yet. Speed to competence appears here as a hard number — median sessions and median days to reach the bar — which is the metric that turns practice into a capacity planning input rather than a training expense.

At individual level, you can see who is improving, who has plateaued, and who is at risk — early enough to intervene. This is the view that changes the manager conversation from "have you done your training?" to "you've plateaued at 61 across four attempts on this scenario, let's look at where it's going wrong."
All of it is exportable via webhooks, embeddable in your LMS via SCORM and xAPI, filterable by cohort, and dated. Evidence, not anecdote.
The driving licence
Here's the analogy we keep coming back to.
Your current training data is the theory test. It confirms someone sat down, read the material and picked the right answers. Useful, necessary, and no basis whatsoever for handing over the keys.
Practice is the road test. It's the bit where someone demonstrates, under observation, that they can actually do it — and where a qualified assessor either signs the licence or doesn't.
Every organisation issues licences constantly. You put a new manager in front of a team. You put a new agent on the phone to a distressed customer. You put a rep in front of a seven-figure account. You do it on the strength of an interview, a CV and a completion record, and then you hope.
The three measures are how you stop hoping. Competence tells you they can. Confidence tells you they will. Compliance tells you they'll do it the way it has to be done — and leaves you with the evidence that you checked.
This isn't about avatars. The avatars are just how the conversation happens. The point is that the conversation, scored, is data you have never had.
Want to see what this looks like for your own team? Run a scenario at realtalkstudio.com or book a walkthrough of the analytics.
Sources
- Brandon Hall Group, Learning Measurement Study (2020) — organisational tracking of learning business impact.
- Kirkpatrick Partners — the four levels of training evaluation.
- Davis DA, Mazmanian PE, Fordis M, et al. "Accuracy of Physician Self-assessment Compared With Observed Measures of Competence: A Systematic Review." JAMA. 2006;296(9):1094–1102.
- Stajkovic AD, Luthans F. "Self-Efficacy and Work-Related Performance: A Meta-Analysis." Psychological Bulletin. 1998;124(2):240–261.
- Landsberg M. The Tao of Coaching. Profile Books.
- Baldwin TT, Ford JK. "Transfer of Training: A Review and Directions for Future Research." Personnel Psychology. 1988.
- Cook DA, Hatala R, Brydges R, et al. "Technology-Enhanced Simulation for Health Professions Education: A Systematic Review and Meta-Analysis." JAMA. 2011;306(9):978–988.
- McGaghie WC, Issenberg SB, Cohen ER, Barsuk JH, Wayne DB. "Does Simulation-Based Medical Education With Deliberate Practice Yield Better Results Than Traditional Clinical Education?" Academic Medicine. 2011;86:706–711.
- Ponemon Institute — cost of non-compliance versus cost of compliance.
- Worker Protection (Amendment of Equality Act 2010) Act 2023; EHRC statutory Code of Practice (October 2024); Employment Rights Act 2025.