Devlin Peck

Build vs Buy: How to Evaluate AI Simulation Tools for Your Team

By Devlin Peck · Updated

Part of the AI in Instructional Design guide

Buy, unless conversational AI is your core business. That is my answer to the build-vs-buy question for AI training simulations, and I say it as someone who builds these systems for a living. This briefing gives you the reasoning, the evaluation criteria, a vendor comparison, and a two-week pilot plan, so you can evaluate tools like a buyer instead of a spectator.

It is part of my full guide to AI training simulations, written for the L&D leader who has to sign off, not the practitioner who found the tool. If you want the wider picture of how AI is changing instructional design, start there. This page assumes you already believe practice matters and are deciding how to deliver it.

The decision is live in more organizations than you might think. In my recent conversations with hiring managers, learning leaders, directors, and senior IDs, nearly everyone is trying to add AI-powered scenarios and role play to their learning programs, and most have not figured out exactly how yet. This gap is where bad build decisions and bad purchases both get made.

Should you build or buy AI training simulations?

Buy. The prototype is cheap, but owning a production training system is a permanent engineering commitment, and the gap between the two is where the real cost lives.

The build option looks deceptively close. Your team can get a prototype conversation running against a model API in a week, and this prototype is exactly what makes leaders overestimate the build path.

The six costs the prototype hides

What does building really cost?

The generic build-vs-buy research puts numbers on what the prototype hides, and none of them favor an L&D team building on the side.

The hidden costThe number or realitySource
Ongoing maintenance15 to 20% of the initial build cost, every yearAakash Gupta's product-leader guide to buy vs build
Compliance certification$100,000+ and 6 to 12 months; vendors amortize this across their whole customer baseAakash Gupta
Engineering talent$200k+ fully loaded per senior engineer, spent on non-core workAakash Gupta
Production hardeningThe last 20% (security, governance, observability, reliability) is 80% of the effortHatchWorks' 2026 build-vs-buy framework
Build to learn vs build to runPrototypes and experiments are one category; production systems with SLAs, audits, and long-term maintenance are another, and teams that do not label which one they are funding get burnedHatchWorks
The core tradeBuilding gets you exactly what you want, at greater expense and effort; buying gets you proven capability quickly, at the cost of some customization and controlThoughtworks' build-vs-buy e-book

The build-to-learn distinction matters because it lets you say yes to your engineer's demo without funding a product team. A prototype that teaches your organization what AI practice feels like is a legitimate spend. A production system your compliance team, your learners, and your renewal budget depend on is a different commitment entirely.

When does building genuinely make sense?

If your organization ships conversational AI products, none of the list above scares you, and building might fit. You already have the engineers, the evaluation discipline, and the safety review process, and the training simulation becomes one more surface for a capability you maintain anyway.

For everyone else, the strongest 2026 answer is the hybrid: buy the platform, build the glue. Your scenarios, your scoring criteria, your Storyline integration, your reporting views are yours to build and own; the conversation engine, evaluation infrastructure, and voice stack are the vendor's problem. HatchWorks frames this as buying the core and building what makes it yours, and it is where most serious teams land. A tool that embeds inside your existing courses is the hybrid path: you keep control of the learning experience without converting your L&D budget into an under-resourced software team.

Want to pressure-test where your organization lands? Answer a few questions about your team and get a build, buy, or hybrid verdict you can paste straight into a decision brief.

Answer six questions about your org. I will give you a Build, Buy, or Hybrid call using the same framework as the article, with a one-line reason for each answer.

1Is conversational AI part of your core product?
2What engineering capacity can you commit permanently?
3What compliance and security posture do you need?
4When do the first learners need to be live?
5Do you have unique data or differentiation a vendor tool cannot express?
6Must practice run live inside existing Storyline, Rise, or LMS courses?
Answer all six questions first.

What criteria matter when evaluating AI simulation tools?

Eleven criteria decide it: course integration, authoring speed, character realism, coaching and scoring, analytics, iteration speed, security, pricing at scale, vendor trajectory, exit path, and language coverage. The last two are the ones most buyers forget.

CriterionThe question to askWhy it matters
Course integrationDoes it run inside our existing Storyline, Rise, or LMS content?If practice lives in a separate destination, most learners never arrive
Authoring speedCan an ID build a working simulation this week, in plain English, without code?Adoption dies when every scenario needs a developer
Character realismDo characters improvise, push back, and stay in role? Is there voice with real emotion?Realism is what makes AI role-play simulations transfer to the job
Coaching and scoringIs there in-conversation support and a scored debrief against criteria we define?Practice without feedback is just chat
AnalyticsCan I see pass rates, scores, attempts, and unique learners on a dashboard?This is your ROI evidence
Iteration speedWhen we edit a scenario, how fast is the change live? Republishing required?Weekly-iteration tools beat quarterly-republish tools
SecurityIs there a SOC 2 report? Where does learner data go? Is it used for model training?Your security team will ask; ask first
Pricing at scalePer seat, per learner, or per conversation? What happens at renewal and at rollout volume?Pilot pricing and rollout pricing can differ sharply
Vendor trajectoryIs the product improving monthly? Who is behind it?You are buying a roadmap in a fast-moving category
Exit pathCan we export our workspace, scenarios, and learner data if we leave?Lock-in is the quiet cost of a fast-moving category
Language coverageWhich languages do the role plays actually run in?Vendors differ sharply; some cover four languages, others sixteen. Global teams should check before piloting

On character realism, do not take a demo reel's word for it. Here is a finished simulation: play the manager giving feedback to an underperforming employee, and judge the pushback yourself.

Handle With Care — a text-based sim where managers practice a conversation with an underperforming employee. Open it full-screen.

On pricing at scale, model real voice volume before you sign. On devlin.ai, voice runs roughly 130 to 150 credits per minute, while a full ten-turn text simulation usually lands around 70 credits, so one minute of voice can cost more than an entire text conversation. The rule of thumb I give buyers: multiply 150 credits by the minutes of voice practice per learner, then by your learner count, and you have a conservative year-one consumption estimate. Voice is what makes the practice work, so price for it directly rather than modeling around it.

On exit path: full disclosure, devlin.ai is my company and it competes in this category, which is exactly why I am comfortable publishing the hard questions; see my Synthesia Roleplay Sessions vs devlin.ai comparison for how we stack up against a specific competitor. We built full workspace export so no customer is locked in, and I think every vendor should have to answer this entire table.

How do devlin.ai and Synthesia Roleplay Sessions compare?

They are different bets. devlin.ai gives designers full control and embeds practice inside the courses you already have; Synthesia Roleplay Sessions gives you interactive video avatars inside the Synthesia platform your team may already use for training video.

Synthesia is the biggest name to enter this category, and their roleplay feature deserves a plain look. The details below come from their feature page as of August 2026; the feature is new and moving fast, so verify before you buy.

devlin.aiSynthesia Roleplay Sessions
Practice formatVoice and text conversations with AI charactersInteractive video avatar role plays with real-time coaching
Where practice livesEmbedded in Storyline, Rise, or any LMS, with SCORM and xAPI reporting and a Storyline 360 variable bridgeStandalone web app with Synthesia login, plus SCORM delivery into your LMS
Role-play languages16 (voice and text)English, German, Spanish, and French (Synthesia Studio video supports 160+)
ScoringDesigner-defined rubrics and criteria, backed by a regression-tested evaluation enginePer-scenario skills rubric with per-skill scores tracked across attempts
Pricing modelUsage-based AI credits, pay per conversation; Team and Enterprise plansTalk to sales
Exit pathFull workspace exportNot stated on the feature page

The fit question is simpler than the table. If your team already runs its training video on Synthesia and real-time video avatars are essential, Roleplay Sessions is the natural choice: practice lives in the platform you already operate. devlin.ai wins when you want maximum flexibility: use it as a standalone platform, or embed simulations inside your existing eLearning, including Storyline courses where the variable bridge can drive score, pass/fail, and even character mood.

And here is the candor I think you should demand from every vendor, starting with me: devlin.ai's simulations are optimized for character-based conversation practice, like an upset customer, a coaching conversation, or a stakeholder pushing back. They are not built for vocational or safety-procedure walkthroughs. Ask every vendor on your shortlist where their tool is weak. The ones who answer quickly are the ones worth piloting.

How should you run a pilot?

Two weeks of building, then a defined evaluation window. The goal is evidence, not a feeling.

  1. Pick one high-value conversation skill

    Choose a skill with a visible business metric: sales objection handling, support de-escalation, manager feedback. Avoid piloting on low-stakes content; a pilot that cannot show value was designed not to.
  2. Let your practitioner build it

    Your instructional designer describes the character, scenario, and scoring criteria. Clock the authoring time; it is one of your evaluation data points. For calibration: a recent project of mine took about 10 hours end to end, and the AI simulations were the easy part. Most of the time went into the surrounding Storyline design and development. A first draft of the simulation itself should take an afternoon, not a sprint.
  3. Define pass criteria before launch

    Agree on what a passing performance looks like and what pass rate you expect after two attempts. Write it down before the first learner enters.
  4. Run 20 to 50 learners through it

    Enough for a real read on completion, pass rates, and repeat attempts. Read a sample of transcripts yourself; they are the most revealing artifact in the pilot.
  5. Debrief against the criteria table

    Score the tool on authoring speed, realism, coaching quality, analytics, and iteration speed. Then decide on rollout, and read my briefing on rolling out AI simulations after the pilot before you do.

The fastest way to calibrate what authoring speed should mean: describe one of your own scenarios in plain English and see what comes back.

Describe a scenario and try the simulation devlin.ai builds from it. Open it full-screen at devlin.ai.

The rollout side, including who owns scenarios and what a 30-60-90 plan looks like, is covered in the rollout briefing linked above. The budget defense is covered in the ROI case for practice-based training.

What security questions should procurement ask?

Send the vendor these five questions verbatim; a serious vendor answers all five in one email.

  1. Can you share your current SOC 2 report and the status of any audit in progress?
  2. Which AI model providers process our learner conversations, and under what data-processing terms?
  3. Is any of our data used to train models, yours or a third party's?
  4. What learner-identifiable data do you store, for how long, and how is deletion handled?
  5. What is your incident-response commitment, and when did you last exercise it?

Evasion on any of them is your answer. As for what good looks like: mature vendors publish their posture without being asked. Synthesia's roleplay page lists SOC 2 Type II, ISO 27001, 27701, and 42001, GDPR compliance, and SSO and SCIM support. On our side, devlin.ai holds a SOC 2 Type 1 report, with the Type 2 report expected in October 2026, and supports SSO/SAML. Whoever you evaluate, get the current audit stage in writing, because "SOC 2" on a homepage can mean several different things.

Frequently asked questions

Is it cheaper to build AI simulations in-house?

Almost never once you count the full cost. The prototype is cheap; the evaluation engine, safety work, voice infrastructure, LMS integration, analytics, and permanent maintenance are not. Buying converts an open-ended engineering liability into a subscription you can cancel.

What does a good AI simulation tool cost?

Pricing in this category is typically per seat, per active learner, or usage-based per conversation, and it varies with voice usage and volume. The more useful number is fully loaded pilot cost, which should be under a month of one instructional designer's time plus a small subscription.

How long should an evaluation take?

About 30 days: two weeks to build and launch one scenario, two weeks to gather learner data and read transcripts. If a vendor cannot get you to a live pilot in two weeks, that is itself evaluation data.

Can't we just use ChatGPT for role-play practice?

A general assistant can improvise a conversation, but it gives you no designer-defined scoring, no LMS reporting, no learner analytics, and no enterprise data-processing terms. The evaluation, integration, and governance problems this article describes are exactly what you would be rebuilding on top of it, which puts you back on the build path with none of the benefits of owning it.