Devlin Peck

Rolling Out AI Simulations: What Your Team Actually Needs

By Devlin Peck · Updated

Part of the AI in Instructional Design guide

To roll out AI simulations, your team needs less than you'd think. Rolling out AI training simulations does not require new headcount, an AI team, or a transformation program. It requires one instructional designer who owns the practice layer, a leader who clears the security review and sets budget guardrails, and a quality bar that keeps AI role-play simulations worth practicing. The field agrees: Donald Taylor's L&D Global Sentiment Survey 2026, drawing on over 3,500 respondents from more than 100 countries, found that interest in AI has peaked and it is now being used selectively. Focused solutions beat trying to add 'AI everywhere,' and this briefing is here to help you focus on releasing something valuable.

I have been on a lot of calls with hiring managers, learning leaders, and senior IDs this year, and nearly all of them are trying to add AI-powered scenarios and role play to their learning experiences. Most have not figured out the operating model yet. This playbook is the operating model. It is part of my full guide to AI training simulations, and it assumes you have already run the pilot from my evaluation briefing.

Who does what in a rollout?

There are four roles, all of which you may already have.

RoleWhoResponsibility
Practice ownerAn instructional designer on your team, usually the pilot championDesigns scenarios and scoring criteria, reads learner transcripts, iterates weekly
SponsorYouPicks the business metrics, clears security review, defends the budget
SMEsThe best performers in the target roleSupply the real conversations: the objections, the phrasing, what good actually sounds like
IT and securityExisting functionOne-time vendor review, SSO, and data-flow sign-off

The practice owner is the hire-nothing approach. Modern simulation tools let an instructional designer describe a character, scenario, and rubric in plain English, with no code and no AI specialists, so that the skills gap is instructional design judgment. Your team likely already has this. It is one piece of the broader shift in how AI is reshaping instructional design work, and my AI Upskilling Track is a free path for the designer stepping into this role.

I also suggest budgeting the role in hours, not headcount. With current tools, a first draft of a simulation takes an afternoon. When my team recently built one of our own flagship simulation projects, the whole build took about 10 hours, and most of it went into the Storyline design and development around the sim. The AI simulations were the easy part.

Notice the champion dynamic: in most organizations the rollout lead is the practitioner who brought the tool to you in the first place. Formalizing their ownership early is the single cheapest thing a leader can do to make the rollout succeed. It also fits where the field is going: the work is changing shape, not disappearing, and the teams handling the change well are the ones that named an owner.

What does a 30-60-90 day rollout look like?

Days 1 to 30 make the pilot production-ready, days 31 to 60 are where you run one real cohort, and days 61 to 90 are where you refine before you widen. For calibration: the teams I see typically go from kickoff to a rolled-out pilot in about four weeks, with pilot cohorts ranging from 15 learners to 15,000 depending on how much data decision-makers need before committing to a bigger rollout. Ninety days is a realistic horizon for the full sequence.

  1. Days 1 to 30: productionize the pilot

    Finish security review and SSO, formalize the practice owner's role, and rebuild the pilot scenario to production quality with SME input. Define the scoring rubric and pass standard in writing, and wire the simulation into the course or LMS path where the target audience already trains.
  2. Days 31 to 60: first real cohort

    Launch to one full team with a manager who wants it. The practice owner reads transcripts weekly and tunes the scenario. You review the dashboard biweekly: completion, pass rates, attempts to mastery. Collect the metric baseline you established in the pilot.
  3. Days 61 to 90: expand by use case, not by org chart

    Add the second and third scenario for the same audience before adding new audiences; depth beats breadth while your quality cadence matures. Stay inside the sweet spot: today's AI simulations are strongest at character-based conversation practice (an upset customer, a coaching conversation, a stakeholder who pushes back), not vocational or safety walkthroughs. Publish the first results readout against the business metric, and set the renewal decision date.

The day-90 readout is where most rollouts get vague, so the measurement section below covers exactly what goes in it.

What will IT and security ask?

They'll likely ask where conversation data goes, how long transcripts are kept, who can access them, how learners authenticate, and what attestations the vendor holds. In the vendor reviews I see, three questions come up almost every time: will AI models be trained on our data, where is the data stored, and does the vendor hold independent attestations. Put all five on your checklist:

Full disclosure before the specifics: devlin.ai is my company. Here is how we handle these, as a reference point for what a clean vendor answer looks like. We have a SOC 2 Type 1 report (Type 2 expected in October 2026), and we publish a trust center with our security documentation. Admins on our Team and Enterprise workspaces set their own retention window for learner transcripts and evaluation feedback. Voice session audio is never stored, and voice transcripts are kept for only 30 days by the voice subprocessor . Enterprise workspaces sign in through SAML SSO with identity providers like Okta or Microsoft Entra ID, and teammates are added automatically on first sign-in. Whatever tool you evaluated, your vendor should be able to answer at this level of specificity in one email.

How do you measure success in the first 90 days?

Track three layers: usage, performance, and one business metric against the baseline you set in the pilot, and put all three in a day-90 readout.

Metric layerExample metricsWhere it comes fromWhat decision it informs
UsageCompletion rate, repeat attempts, practice minutesPlatform dashboard, LMS reportingWhether practice is happening where you embedded it
PerformanceScore trend against the rubric, attempts to mastery, pass rate against your own standardAI evaluation, transcriptsWhether the scenario is producing skill growth
BusinessThe pilot's baseline metric: expert hours spent on live role plays, ramp time, QA scoresThe system that owned your baselineWhether to renew and expand

It helps to know what the scores in that middle layer actually measure. In devlin.ai, the designer defines the scoring criteria and the points attached to each one; the evaluation then runs the full conversation transcript against those criteria and awards full, partial, or zero credit per criterion. We regression-test the evaluation engine against a library of graded practice sessions so scoring stays consistent. The practical implication for you as sponsor: the pass standard is a design decision your practice owner makes for your context, not an industry constant, so treat pass-rate movement as feedback on the scenario and the cohort rather than comparing it to some external benchmark.

When leaders ask which single number to anchor the renewal decision on, my answer is pass rate against your own standard, and ideally the metric the simulations were brought in to move. If the point was reducing the hours your experts spend running live role plays with trainees, the day-90 readout should show trainees performing as well or better while experts spend less time training them. That readout feeds directly into the ROI framework for practice-based training when the budget conversation starts.

How do you keep quality high?

Simulations are living content, and the maintenance model looks more like managing a channel than shipping a course:

The fastest way to calibrate your own quality bar is to take a simulation yourself. Here is a simulation integrated into a Storyline course where you handle a conversation with an underperforming employee.

Handle With Care — a text-based sim where managers practice a conversation with an underperforming employee. Open it full-screen.

What usually goes wrong?

The failure modes are predictable, which means they are avoidable:

  1. Practice as a destination. If learners must leave their course or LMS to practice, most never do. Embed simulations where training already happens.
  2. No named owner. The rollout that belongs to "the team" belongs to nobody by day 60.
  3. Boiling the ocean. Ten mediocre scenarios across five departments lose to three excellent ones for a single team. Expand from strength.
  4. Ignoring the transcripts. Dashboards summarize; transcripts explain. Leaders who read a few transcripts a month make better calls than leaders who only watch pass rates.
  5. Skipping the manager. If the target team's manager treats practice as optional, it is. This is the strongest-evidenced lever in the research: Grossman and Salas's 2011 review of training transfer in the International Journal of Training and Development found supervisor support has some of the strongest evidence among work-environment factors for whether training shows up on the job. Recruit the manager before the cohort launches.
  6. Unbounded usage budgets. Simulation platforms price in usage-based AI credits, so your company effectively pays per conversation, and voice consumes credits faster than text: a single minute of voice practice can use more credits than an entire text conversation. Voice is also what makes the practice feel real, so budget for it rather than around it. The failure is skipping the guardrails, because a popular simulation without caps produces a surprise bill and a defensive CFO. In devlin.ai, admins set a daily credit cap on each simulation and coach, get alerted when a simulation hits 80% of its limit, and can see exactly how credits are consumed across the workspace. Whatever tool you run, insist on those controls before the first cohort.

This playbook reflects the rollouts I have watched succeed and stall across the more than 2,000 people building simulations with devlin.ai, and the pattern holds regardless of the tool you pick.

Before you greenlight the first cohort, run your plan through this checker. It scores your rollout against the preconditions above and tells you what to fix first.

Rollout Readiness Checker

Six preconditions, straight from this playbook. Answer honestly. You get a green light or a ranked list of what to fix before the first cohort.

  1. 1. Do you have a named practice owner with allocated hours?

    One instructional designer who designs scenarios and scoring criteria, reads transcripts, and iterates weekly. Hours, not headcount.

  2. 2. Is the target team's manager recruited and bought in?

    The first cohort launches to one full team with a manager who wants it.

  3. 3. Is the simulation embedded where training already happens?

    Wired into the course or LMS path the target audience already trains in.

  4. 4. Are the security review and budget guardrails done?

    Vendor review, SSO, and daily credit caps with usage alerts before the first cohort.

  5. 5. Are the business metric and baseline defined in writing?

    The pilot's baseline metric: expert hours on live role plays, ramp time, or QA scores.

  6. 6. Is a weekly transcript review cadence scheduled?

    The practice owner reads five to ten transcripts a week and tunes the scenario.

Answer all six questions to get your verdict.

Frequently asked questions

Do we need to hire anyone to roll out AI simulations?

Usually no. The critical role is a practice owner, an instructional designer who designs scenarios and reads transcripts, and modern tools require no code or AI expertise. Budget part of one existing ID's time rather than a new position.

How many simulations should we launch with?

One production-quality scenario for one team, then two or three more for the same audience before expanding to new audiences. Depth first: your quality cadence needs to mature before it can support breadth.

How much ongoing maintenance do AI simulations need?

Plan for a few hours a week of transcript reading and scenario tuning per active audience, plus a quarterly SME refresh. The work is light but must be owned; unmaintained scenarios lose realism and learner trust.

What does an AI simulation rollout cost?

Mostly reallocated instructional designer time plus usage-based vendor credits: simulation platforms effectively charge per conversation, so cost scales with practice volume rather than seat count. Platforms like devlin.ai offer Team and Enterprise plans for organizations, and daily credit caps per simulation keep usage forecastable.

Do AI simulations create data privacy risk?

They create reviewable conversation data, which is exactly why the security review focuses on retention windows, access controls, and SSO. Ask the vendor where transcripts are stored, who can read them, how long they are kept, whether learner conversations train AI models, and what independent attestations back the answers.