AI Role-Play Simulations: What They Are and When They Work
By Devlin Peck · Updated
Part of the AI in Instructional Design guide
An AI role-play simulation is a practice conversation in which an AI plays a character, a skeptical buyer, like an upset patient or a defensive direct report, and the learner plays themselves. The learner talks or types their way through the scenario, the character responds in real time, and the learner gets a score + feedback on how they handled it.
When I started in instructional design back in 2017, I specialized in designing scenario-based eLearning and branching scenarios, often using a tool like Storyline 360, back before AI reshaped design workflows.
But in 2026, the tide is changing, which is part of why I get asked so often whether AI will replace instructional designers. I have been on a lot of calls with learning leaders, directors, and senior instructional designers. Nearly every one of them is trying to work AI-powered scenarios and role-play into their learning programs. Most of them have not figured out exactly how to do it yet. This briefing is the straightforward comparison I give those leaders: what AI role-play training does well, what the old approach still does well, and how to tell which conversation skills are worth simulating. It is part of my full guide to AI in instructional design.
How are AI role-play simulations different from branching scenarios?
A branching scenario is a script. An AI role-play simulator improvises within guardrails that you define. With branching, your team writes every line the character can say and every choice the learner can make, then builds the tree. With AI role-play, you define the character, the situation, and the evaluation criteria, and the AI handles the infinite ways a real conversation can go.
| Branching scenario | AI role-play simulation | |
|---|---|---|
| Authoring effort | Weeks: script every path, build every branch | Hours: describe character, scenario, and criteria |
| Learner input | Picks from 2 to 4 pre-written choices | Says or types whatever they would actually say |
| Realism | As good as the script, and learners can game it | Improvises and pushes back like a person |
| Replayability | Same tree every time | Different conversation every attempt |
| Measurement | Which branch they clicked | Scored performance against your criteria, plus transcript |
| Predictability | Total: every path is authored | High but not total: the AI improvises within bounds |
That last row is the real trade-off. A branching scenario can never surprise you, which also means it can never surprise the learner. If total scriptedness is a regulatory requirement for a given course, branch it. For everything else, the realism gap is enormous, and learners feel it immediately.
The authoring row deserves emphasis too. You can now build a first draft of a simulation in an afternoon, instead of spending months learning to develop something similar in an authoring tool. That collapses the biggest cost that used to make conversation practice a luxury item.
Here is why this category excites me, and it is not that we can recreate old formats faster. It is that we can build something we literally could not build before: a learner talking to a character who has a personality, emotional responses, and motives. Practicing a difficult conversation with a stakeholder who pushes back. Rehearsing feedback for a direct report who gets defensive. Not a branching scenario with three pre-written paths. A real conversation that adapts in real time.
One caution: AI role-play removes the production bottleneck, not the design discipline. You still need real choices, real consequences, and scenarios drawn from what top performers actually do. For where this format fits among the other simulation types, see my full guide to AI training simulations.
Which skills are AI role-play simulations good for?
AI roleplays are best for skills where the hard part is the conversation itself:
- Sales: discovery calls, objection handling, negotiation, renewal conversations.
- Customer service: de-escalation, delivering bad news, staying on policy under pressure.
- Leadership and management: feedback conversations, coaching sessions, performance reviews, conflict mediation.
- Healthcare and care roles: patient education, informed-consent conversations, difficult family conversations.
- HR and compliance: receiving a harassment report correctly, accommodation conversations, exit interviews.
These solutions are best when the learner already knows what they should do, but they haven't practiced doing it while a realistic person pushes back. That gap between knowing and doing is exactly what AI roleplay training closes, and employees respond to training that closes it. According to Axonify's State of Workplace Training study, 92% of employees say the right kind of formal workplace training positively impacts their job engagement. And per Lorman's employee training statistics roundup, 59% of employees believe training directly improves their job performance. "The right kind" is the operative phrase in both: practice against a realistic counterpart is what makes conversation training the right kind.
When should you not use an AI role-play simulation?
You shouldn't use AI role-play simulations when the skill is not conversational, or when improvisation adds risk instead of realism:
- Procedural and technical training. Software walkthroughs, equipment operation, and checklists want demonstration and hands-on practice, not dialogue. That being said, many orgs are rolling out AI simulations combined with software simulations so that reps can practice talking to a simulated customer while navigating the same software they'd use on the call.
- Compliance attestation. If the requirement is "every employee saw this exact content and acknowledged it," you need the exact content, not an improvised conversation.
- Pure knowledge checks. If a quiz answers the question, use a quiz. Simulations earn their cost on performance, not knowledge recall.
- Scenarios your organization cannot let the AI improvise. A few legal scripts must be delivered word for word. Script those, and simulate the conversation around them.
What does the learner actually experience?
In a modern implementation, the learner opens the simulation inside the eLearning course they were already taking, meets a character with a name, a voice, and a disposition, and has the conversation. Voice-capable characters laugh, hesitate, and get defensive, which is what makes the practice feel real enough to raise a learner's heart rate in a safe place. An AI coach can support them during the conversation and run a debrief afterward, scored against the criteria your team defined.
It is humbling in the way real practice is humbling. In some of the live workshops where I build and test voice simulations in front of a live audience, I feel the nerves and anxiety on the simulated call just as I would on a real call where I'm out of my depth. The pressure is much more real when the AI character is waiting for you respond, getting impatient if you don't hold an actual conversation with them. That is the point: a score you cannot charm your way to, with specific feedback at the end that you can act on for the next attempt.
Full disclosure: devlin.ai is my company. It is an AI text and voice conversation simulation builder: describe a scenario in plain English and you get a working simulation that embeds in Storyline, Rise, or any LMS, with AI evaluation, transcripts, and scores reporting back to the course. That description of the learner experience is how simulations built with it behave. Whatever tool you evaluate, hold it to that bar: in-course, in-character, coached, and scored.
If you want to see one of these built and played end to end, I did exactly that on a live stream, including voice mode with a skeptical manager character:
How is learner performance scored and measured?
The AI evaluates the full conversation transcript against criteria your team defines, and the score, pass/fail result, and transcript report back to your course and LMS. This is the single biggest upgrade over both branching scenarios and human role-play: you get performance data instead of click data, and it scales past what any human observer could review.
Here is how it works in my platform: you define the scoring criteria, each with a point value you determine. A master prompt takes in those criteria plus the full transcript of the conversation that just happened, then assigns full credit, partial credit, or zero credit per criterion. The learner sees which criteria they hit and which they missed, with evidence grounded in the actual transcript.
Two procurement questions are really important here, too. First, is the scoring consistent? We regression-test our evaluation engine against a library of graded practice sessions so evaluator accuracy stays reliable over time, and cited evidence gets verified against the actual transcript rather than invented. Ask any vendor how they keep their evaluator honest; if the answer is a shrug, then the scores are decoration. Second, does the data land where your systems need it? Look for SCORM and xAPI reporting into your LMS, plus a workspace-level analytics view of sessions, completions, and pass rates. Results can even update Storyline 360 variables, like score, pass/fail, and character mood, so the course content itself can change states based on how the conversation went.
I cannot give you a universal pass-rate benchmark for every simulation. The right target depends on the skill, the stakes, and whether the simulation sits before or after the instruction.
What keeps the AI character on track?
Well-designed guardrails constrain the character to its role, ground it's references in your source material, and make every exchange visible in transcripts. "Can a learner break it?" is a fair question, and it deserves a real mechanical answer, not reassurance.
In our case, the answer has four parts. The character gets a dedicated system prompt that is entirely about staying in character, and we made an early design decision to separate the character completely from the evaluator that grades the performance afterward, so neither job degrades the other. You can ground the character's knowledge in specific documents, like PDFs and PowerPoint files, so it speaks from your material instead of guessing. We hardened those prompts over months with beta users to minimize the chance of a character stepping out of its role.
Even with all of that, "high but not total predictability" is still the honest phrase when it comes to character steering. That is why the transcript is your audit trail: every exchange is recorded and reviewable, so an odd response is visible to your team rather than lost in a hallway conversation. For enterprise deployments, also confirm the boring-but-critical controls: SSO/SAML for access, and configurable data-retention windows for learner transcripts.
Put these on your vendor checklist regardless of which tool you evaluate: How is the character bounded? Can it be grounded in our documents? Are full transcripts available? What are the access and retention controls?
How much do AI role-play simulations cost?
Most platforms in this category price one of two ways: an annual per-seat license, or usage-based pricing where you pay for what learners actually consume. Per-seat licensing is predictable but charges you the same for the employee who practices weekly and the one who never logs in. Usage-based pricing tracks actual practice, which tends to fit pilot-first buyers better.
devlin.ai uses the second model: usage-based AI credits, so you effectively pay per conversation. The Team plan serves small learning teams and Enterprise serves larger organizations, and you get full visibility into how credits are consumed across the workspace, so cost becomes predictable as usage scales up.
One budgeting note, because I see instructional designers want to start with text-based simulations (which use fewer credits), whereas leaders often see the value in voice. Voice mode is what makes the whole experience: the hesitation, the tone, and the pressure of speaking out loud is where the behavior change comes from, so plan your pilot around voice conversations rather than treating voice as a premium add-on. The real question is not the per-conversation cost anyway; it is whether the conversations move a business metric. I break down how to run that math in the ROI of practice-based training.
How should an L&D leader pilot AI role-play?
Answer five questions about one specific skill you want to train. You get a verdict: build an AI role-play simulation, build a branching scenario instead, or use a different method entirely. The logic comes straight from the criteria in this article.
If a practitioner on your team is already excited about this category, give them one high-value conversation skill and let them build the pilot. Then judge the result on evidence: transcripts, scores, and whether the business metric the skill feeds actually moves.
Remember those calls I mentioned at the top: nearly every learning leader I talk to is attempting something like this, and most are stuck. The ones who get unstuck are not the ones with the biggest budgets. They are the ones who pick one skill instead of trying to solve it all at once, then they ship a pilot, view the scores, and read the transcripts. My briefings on evaluating AI simulation tools and rolling out AI simulations across your organization cover the two halves that come next: choosing the platform, and scaling past the pilot.