Rolling Out AI Simulations: What Your Team Actually Needs
By Devlin Peck · Updated
Part of the AI in Instructional Design guide
To roll out AI simulations, your team needs less than you'd think. Rolling out AI training simulations does not require new headcount, an AI team, or a transformation program. It requires one instructional designer who owns the practice layer, a leader who clears the security review and sets budget guardrails, and a quality bar that keeps AI role-play simulations worth practicing. The field agrees: Donald Taylor's L&D Global Sentiment Survey 2026, drawing on over 3,500 respondents from more than 100 countries, found that interest in AI has peaked and it is now being used selectively. Focused solutions beat trying to add 'AI everywhere,' and this briefing is here to help you focus on releasing something valuable.
I have been on a lot of calls with hiring managers, learning leaders, and senior IDs this year, and nearly all of them are trying to add AI-powered scenarios and role play to their learning experiences. Most have not figured out the operating model yet. This playbook is the operating model. It is part of my full guide to AI training simulations, and it assumes you have already run the pilot from my evaluation briefing.
Who does what in a rollout?
There are four roles, all of which you may already have.
| Role | Who | Responsibility |
|---|---|---|
| Practice owner | An instructional designer on your team, usually the pilot champion | Designs scenarios and scoring criteria, reads learner transcripts, iterates weekly |
| Sponsor | You | Picks the business metrics, clears security review, defends the budget |
| SMEs | The best performers in the target role | Supply the real conversations: the objections, the phrasing, what good actually sounds like |
| IT and security | Existing function | One-time vendor review, SSO, and data-flow sign-off |
The practice owner is the hire-nothing approach. Modern simulation tools let an instructional designer describe a character, scenario, and rubric in plain English, with no code and no AI specialists, so that the skills gap is instructional design judgment. Your team likely already has this. It is one piece of the broader shift in how AI is reshaping instructional design work, and my AI Upskilling Track is a free path for the designer stepping into this role.
I also suggest budgeting the role in hours, not headcount. With current tools, a first draft of a simulation takes an afternoon. When my team recently built one of our own flagship simulation projects, the whole build took about 10 hours, and most of it went into the Storyline design and development around the sim. The AI simulations were the easy part.
Notice the champion dynamic: in most organizations the rollout lead is the practitioner who brought the tool to you in the first place. Formalizing their ownership early is the single cheapest thing a leader can do to make the rollout succeed. It also fits where the field is going: the work is changing shape, not disappearing, and the teams handling the change well are the ones that named an owner.
What does a 30-60-90 day rollout look like?
Days 1 to 30 make the pilot production-ready, days 31 to 60 are where you run one real cohort, and days 61 to 90 are where you refine before you widen. For calibration: the teams I see typically go from kickoff to a rolled-out pilot in about four weeks, with pilot cohorts ranging from 15 learners to 15,000 depending on how much data decision-makers need before committing to a bigger rollout. Ninety days is a realistic horizon for the full sequence.
Days 1 to 30: productionize the pilot
Finish security review and SSO, formalize the practice owner's role, and rebuild the pilot scenario to production quality with SME input. Define the scoring rubric and pass standard in writing, and wire the simulation into the course or LMS path where the target audience already trains.Days 31 to 60: first real cohort
Launch to one full team with a manager who wants it. The practice owner reads transcripts weekly and tunes the scenario. You review the dashboard biweekly: completion, pass rates, attempts to mastery. Collect the metric baseline you established in the pilot.Days 61 to 90: expand by use case, not by org chart
Add the second and third scenario for the same audience before adding new audiences; depth beats breadth while your quality cadence matures. Stay inside the sweet spot: today's AI simulations are strongest at character-based conversation practice (an upset customer, a coaching conversation, a stakeholder who pushes back), not vocational or safety walkthroughs. Publish the first results readout against the business metric, and set the renewal decision date.
The day-90 readout is where most rollouts get vague, so the measurement section below covers exactly what goes in it.
What will IT and security ask?
They'll likely ask where conversation data goes, how long transcripts are kept, who can access them, how learners authenticate, and what attestations the vendor holds. In the vendor reviews I see, three questions come up almost every time: will AI models be trained on our data, where is the data stored, and does the vendor hold independent attestations. Put all five on your checklist:
- Model training on your data. Get the vendor's written answer on whether learner conversations are used to train AI models, and under what terms their model providers process conversation text.
- Data storage and flow. Where transcripts and scores live, which subprocessors touch them, and how data moves between the simulation, your LMS, and your reporting stack.
- Independent attestations. SOC 2 or equivalent, plus published security documentation your team can review without an NDA dance.
- Transcript retention and access. Who can read learner conversations and for how long. Note the tension: transcripts are the raw material of your quality system, so the right answer is a deliberate retention window, not zero retention.
- Authentication. SSO through your identity provider, with automatic provisioning so the rollout does not stall on account creation.
Full disclosure before the specifics: devlin.ai is my company. Here is how we handle these, as a reference point for what a clean vendor answer looks like. We have a SOC 2 Type 1 report (Type 2 expected in October 2026), and we publish a trust center with our security documentation. Admins on our Team and Enterprise workspaces set their own retention window for learner transcripts and evaluation feedback. Voice session audio is never stored, and voice transcripts are kept for only 30 days by the voice subprocessor . Enterprise workspaces sign in through SAML SSO with identity providers like Okta or Microsoft Entra ID, and teammates are added automatically on first sign-in. Whatever tool you evaluated, your vendor should be able to answer at this level of specificity in one email.
How do you measure success in the first 90 days?
Track three layers: usage, performance, and one business metric against the baseline you set in the pilot, and put all three in a day-90 readout.
| Metric layer | Example metrics | Where it comes from | What decision it informs |
|---|---|---|---|
| Usage | Completion rate, repeat attempts, practice minutes | Platform dashboard, LMS reporting | Whether practice is happening where you embedded it |
| Performance | Score trend against the rubric, attempts to mastery, pass rate against your own standard | AI evaluation, transcripts | Whether the scenario is producing skill growth |
| Business | The pilot's baseline metric: expert hours spent on live role plays, ramp time, QA scores | The system that owned your baseline | Whether to renew and expand |
It helps to know what the scores in that middle layer actually measure. In devlin.ai, the designer defines the scoring criteria and the points attached to each one; the evaluation then runs the full conversation transcript against those criteria and awards full, partial, or zero credit per criterion. We regression-test the evaluation engine against a library of graded practice sessions so scoring stays consistent. The practical implication for you as sponsor: the pass standard is a design decision your practice owner makes for your context, not an industry constant, so treat pass-rate movement as feedback on the scenario and the cohort rather than comparing it to some external benchmark.
When leaders ask which single number to anchor the renewal decision on, my answer is pass rate against your own standard, and ideally the metric the simulations were brought in to move. If the point was reducing the hours your experts spend running live role plays with trainees, the day-90 readout should show trainees performing as well or better while experts spend less time training them. That readout feeds directly into the ROI framework for practice-based training when the budget conversation starts.
How do you keep quality high?
Simulations are living content, and the maintenance model looks more like managing a channel than shipping a course:
- Weekly transcript sampling. The practice owner reads five to ten transcripts a week. Transcripts surface everything: scenarios that are too easy, characters that drift, criteria that misfire.
- A scenario registry with owners. Every live simulation has a named owner and a review date. Unowned scenarios rot.
- SME refresh on a cadence. Quarterly, ask top performers what has changed in the real conversations. Update scenarios the same week; with modern tools a saved change is live in about a minute, without republishing the course.
- Retire ruthlessly. A scenario nobody assigns or passes information through is dead weight. Archive it.
The fastest way to calibrate your own quality bar is to take a simulation yourself. Here is a simulation integrated into a Storyline course where you handle a conversation with an underperforming employee.
What usually goes wrong?
The failure modes are predictable, which means they are avoidable:
- Practice as a destination. If learners must leave their course or LMS to practice, most never do. Embed simulations where training already happens.
- No named owner. The rollout that belongs to "the team" belongs to nobody by day 60.
- Boiling the ocean. Ten mediocre scenarios across five departments lose to three excellent ones for a single team. Expand from strength.
- Ignoring the transcripts. Dashboards summarize; transcripts explain. Leaders who read a few transcripts a month make better calls than leaders who only watch pass rates.
- Skipping the manager. If the target team's manager treats practice as optional, it is. This is the strongest-evidenced lever in the research: Grossman and Salas's 2011 review of training transfer in the International Journal of Training and Development found supervisor support has some of the strongest evidence among work-environment factors for whether training shows up on the job. Recruit the manager before the cohort launches.
- Unbounded usage budgets. Simulation platforms price in usage-based AI credits, so your company effectively pays per conversation, and voice consumes credits faster than text: a single minute of voice practice can use more credits than an entire text conversation. Voice is also what makes the practice feel real, so budget for it rather than around it. The failure is skipping the guardrails, because a popular simulation without caps produces a surprise bill and a defensive CFO. In devlin.ai, admins set a daily credit cap on each simulation and coach, get alerted when a simulation hits 80% of its limit, and can see exactly how credits are consumed across the workspace. Whatever tool you run, insist on those controls before the first cohort.
This playbook reflects the rollouts I have watched succeed and stall across the more than 2,000 people building simulations with devlin.ai, and the pattern holds regardless of the tool you pick.
Before you greenlight the first cohort, run your plan through this checker. It scores your rollout against the preconditions above and tells you what to fix first.
Rollout Readiness Checker
Six preconditions, straight from this playbook. Answer honestly. You get a green light or a ranked list of what to fix before the first cohort.