Build vs Buy: How to Evaluate AI Simulation Tools for Your Team
By Devlin Peck · Updated
Part of the AI in Instructional Design guide
Buy, unless conversational AI is your core business. That is my answer to the build-vs-buy question for AI training simulations, and I say it as someone who builds these systems for a living. This briefing gives you the reasoning, the evaluation criteria, a vendor comparison, and a two-week pilot plan, so you can evaluate tools like a buyer instead of a spectator.
It is part of my full guide to AI training simulations, written for the L&D leader who has to sign off, not the practitioner who found the tool. If you want the wider picture of how AI is changing instructional design, start there. This page assumes you already believe practice matters and are deciding how to deliver it.
The decision is live in more organizations than you might think. In my recent conversations with hiring managers, learning leaders, directors, and senior IDs, nearly everyone is trying to add AI-powered scenarios and role play to their learning programs, and most have not figured out exactly how yet. This gap is where bad build decisions and bad purchases both get made.
Should you build or buy AI training simulations?
Buy. The prototype is cheap, but owning a production training system is a permanent engineering commitment, and the gap between the two is where the real cost lives.
The build option looks deceptively close. Your team can get a prototype conversation running against a model API in a week, and this prototype is exactly what makes leaders overestimate the build path.
The six costs the prototype hides
- Character quality is prompt engineering, forever. Characters drift, break role, or go flat, and keeping them believable across model updates is ongoing skilled work.
- Evaluation is its own product. Scoring a conversation against rubrics, consistently and defensibly, is harder than running the conversation. Here is what this took on our side at devlin.ai: a master prompt that takes in the scoring criteria a designer defines, reads the full transcript of the conversation that just happened, and assigns full, partial, or zero credit per criterion. Then, to keep those scores trustworthy, we regression-test the evaluation engine, pin models so an upstream update cannot silently change scoring, ground every score in real turns from the transcript, test for bias, and support human overrides. Your two engineers would be rebuilding all of it.
- Safety and boundaries. You own every off-script thing your character says to an employee.
- Voice, latency, and cost. Real-time speech with emotional range is a specialist engineering problem, and model costs need managing at scale.
- Integration and analytics. Learners need it inside courses and the LMS; you need dashboards, transcripts, and completion data. That is a second product.
- Maintenance has no finish line. Models deprecate, APIs change, and the two engineers who built it get pulled onto other work.
What does building really cost?
The generic build-vs-buy research puts numbers on what the prototype hides, and none of them favor an L&D team building on the side.
| The hidden cost | The number or reality | Source |
|---|---|---|
| Ongoing maintenance | 15 to 20% of the initial build cost, every year | Aakash Gupta's product-leader guide to buy vs build |
| Compliance certification | $100,000+ and 6 to 12 months; vendors amortize this across their whole customer base | Aakash Gupta |
| Engineering talent | $200k+ fully loaded per senior engineer, spent on non-core work | Aakash Gupta |
| Production hardening | The last 20% (security, governance, observability, reliability) is 80% of the effort | HatchWorks' 2026 build-vs-buy framework |
| Build to learn vs build to run | Prototypes and experiments are one category; production systems with SLAs, audits, and long-term maintenance are another, and teams that do not label which one they are funding get burned | HatchWorks |
| The core trade | Building gets you exactly what you want, at greater expense and effort; buying gets you proven capability quickly, at the cost of some customization and control | Thoughtworks' build-vs-buy e-book |
The build-to-learn distinction matters because it lets you say yes to your engineer's demo without funding a product team. A prototype that teaches your organization what AI practice feels like is a legitimate spend. A production system your compliance team, your learners, and your renewal budget depend on is a different commitment entirely.
When does building genuinely make sense?
If your organization ships conversational AI products, none of the list above scares you, and building might fit. You already have the engineers, the evaluation discipline, and the safety review process, and the training simulation becomes one more surface for a capability you maintain anyway.
For everyone else, the strongest 2026 answer is the hybrid: buy the platform, build the glue. Your scenarios, your scoring criteria, your Storyline integration, your reporting views are yours to build and own; the conversation engine, evaluation infrastructure, and voice stack are the vendor's problem. HatchWorks frames this as buying the core and building what makes it yours, and it is where most serious teams land. A tool that embeds inside your existing courses is the hybrid path: you keep control of the learning experience without converting your L&D budget into an under-resourced software team.
Want to pressure-test where your organization lands? Answer a few questions about your team and get a build, buy, or hybrid verdict you can paste straight into a decision brief.
Answer six questions about your org. I will give you a Build, Buy, or Hybrid call using the same framework as the article, with a one-line reason for each answer.
What criteria matter when evaluating AI simulation tools?
Eleven criteria decide it: course integration, authoring speed, character realism, coaching and scoring, analytics, iteration speed, security, pricing at scale, vendor trajectory, exit path, and language coverage. The last two are the ones most buyers forget.
| Criterion | The question to ask | Why it matters |
|---|---|---|
| Course integration | Does it run inside our existing Storyline, Rise, or LMS content? | If practice lives in a separate destination, most learners never arrive |
| Authoring speed | Can an ID build a working simulation this week, in plain English, without code? | Adoption dies when every scenario needs a developer |
| Character realism | Do characters improvise, push back, and stay in role? Is there voice with real emotion? | Realism is what makes AI role-play simulations transfer to the job |
| Coaching and scoring | Is there in-conversation support and a scored debrief against criteria we define? | Practice without feedback is just chat |
| Analytics | Can I see pass rates, scores, attempts, and unique learners on a dashboard? | This is your ROI evidence |
| Iteration speed | When we edit a scenario, how fast is the change live? Republishing required? | Weekly-iteration tools beat quarterly-republish tools |
| Security | Is there a SOC 2 report? Where does learner data go? Is it used for model training? | Your security team will ask; ask first |
| Pricing at scale | Per seat, per learner, or per conversation? What happens at renewal and at rollout volume? | Pilot pricing and rollout pricing can differ sharply |
| Vendor trajectory | Is the product improving monthly? Who is behind it? | You are buying a roadmap in a fast-moving category |
| Exit path | Can we export our workspace, scenarios, and learner data if we leave? | Lock-in is the quiet cost of a fast-moving category |
| Language coverage | Which languages do the role plays actually run in? | Vendors differ sharply; some cover four languages, others sixteen. Global teams should check before piloting |
On character realism, do not take a demo reel's word for it. Here is a finished simulation: play the manager giving feedback to an underperforming employee, and judge the pushback yourself.
On pricing at scale, model real voice volume before you sign. On devlin.ai, voice runs roughly 130 to 150 credits per minute, while a full ten-turn text simulation usually lands around 70 credits, so one minute of voice can cost more than an entire text conversation. The rule of thumb I give buyers: multiply 150 credits by the minutes of voice practice per learner, then by your learner count, and you have a conservative year-one consumption estimate. Voice is what makes the practice work, so price for it directly rather than modeling around it.
On exit path: full disclosure, devlin.ai is my company and it competes in this category, which is exactly why I am comfortable publishing the hard questions; see my Synthesia Roleplay Sessions vs devlin.ai comparison for how we stack up against a specific competitor. We built full workspace export so no customer is locked in, and I think every vendor should have to answer this entire table.
How do devlin.ai and Synthesia Roleplay Sessions compare?
They are different bets. devlin.ai gives designers full control and embeds practice inside the courses you already have; Synthesia Roleplay Sessions gives you interactive video avatars inside the Synthesia platform your team may already use for training video.
Synthesia is the biggest name to enter this category, and their roleplay feature deserves a plain look. The details below come from their feature page as of August 2026; the feature is new and moving fast, so verify before you buy.
| devlin.ai | Synthesia Roleplay Sessions | |
|---|---|---|
| Practice format | Voice and text conversations with AI characters | Interactive video avatar role plays with real-time coaching |
| Where practice lives | Embedded in Storyline, Rise, or any LMS, with SCORM and xAPI reporting and a Storyline 360 variable bridge | Standalone web app with Synthesia login, plus SCORM delivery into your LMS |
| Role-play languages | 16 (voice and text) | English, German, Spanish, and French (Synthesia Studio video supports 160+) |
| Scoring | Designer-defined rubrics and criteria, backed by a regression-tested evaluation engine | Per-scenario skills rubric with per-skill scores tracked across attempts |
| Pricing model | Usage-based AI credits, pay per conversation; Team and Enterprise plans | Talk to sales |
| Exit path | Full workspace export | Not stated on the feature page |
The fit question is simpler than the table. If your team already runs its training video on Synthesia and real-time video avatars are essential, Roleplay Sessions is the natural choice: practice lives in the platform you already operate. devlin.ai wins when you want maximum flexibility: use it as a standalone platform, or embed simulations inside your existing eLearning, including Storyline courses where the variable bridge can drive score, pass/fail, and even character mood.
And here is the candor I think you should demand from every vendor, starting with me: devlin.ai's simulations are optimized for character-based conversation practice, like an upset customer, a coaching conversation, or a stakeholder pushing back. They are not built for vocational or safety-procedure walkthroughs. Ask every vendor on your shortlist where their tool is weak. The ones who answer quickly are the ones worth piloting.
How should you run a pilot?
Two weeks of building, then a defined evaluation window. The goal is evidence, not a feeling.
Pick one high-value conversation skill
Choose a skill with a visible business metric: sales objection handling, support de-escalation, manager feedback. Avoid piloting on low-stakes content; a pilot that cannot show value was designed not to.Let your practitioner build it
Your instructional designer describes the character, scenario, and scoring criteria. Clock the authoring time; it is one of your evaluation data points. For calibration: a recent project of mine took about 10 hours end to end, and the AI simulations were the easy part. Most of the time went into the surrounding Storyline design and development. A first draft of the simulation itself should take an afternoon, not a sprint.Define pass criteria before launch
Agree on what a passing performance looks like and what pass rate you expect after two attempts. Write it down before the first learner enters.Run 20 to 50 learners through it
Enough for a real read on completion, pass rates, and repeat attempts. Read a sample of transcripts yourself; they are the most revealing artifact in the pilot.Debrief against the criteria table
Score the tool on authoring speed, realism, coaching quality, analytics, and iteration speed. Then decide on rollout, and read my briefing on rolling out AI simulations after the pilot before you do.
The fastest way to calibrate what authoring speed should mean: describe one of your own scenarios in plain English and see what comes back.
The rollout side, including who owns scenarios and what a 30-60-90 plan looks like, is covered in the rollout briefing linked above. The budget defense is covered in the ROI case for practice-based training.
What security questions should procurement ask?
Send the vendor these five questions verbatim; a serious vendor answers all five in one email.
- Can you share your current SOC 2 report and the status of any audit in progress?
- Which AI model providers process our learner conversations, and under what data-processing terms?
- Is any of our data used to train models, yours or a third party's?
- What learner-identifiable data do you store, for how long, and how is deletion handled?
- What is your incident-response commitment, and when did you last exercise it?
Evasion on any of them is your answer. As for what good looks like: mature vendors publish their posture without being asked. Synthesia's roleplay page lists SOC 2 Type II, ISO 27001, 27701, and 42001, GDPR compliance, and SSO and SCIM support. On our side, devlin.ai holds a SOC 2 Type 1 report, with the Type 2 report expected in October 2026, and supports SSO/SAML. Whoever you evaluate, get the current audit stage in writing, because "SOC 2" on a homepage can mean several different things.