Build vs Buy: How to Evaluate AI Simulation Tools for Your Team
By Devlin Peck · Updated
Part of the AI in Instructional Design guide
Buy, unless conversational AI is your core business. That is my answer to the build-vs-buy question for AI training simulations, and I say that as someone who builds AI systems for a living. This briefing explains the reasoning, then gives you the criteria and the pilot plan for evaluating tools like a buyer instead of a spectator.
It is part of my full guide to AI training simulations, written for the L&D leader who has to sign off, not the practitioner who found the tool.
Should you build or buy?
The build option looks deceptively close. Your team can get a prototype conversation running against a model API in a week, and that prototype is exactly what makes leaders overestimate the build path. The distance between a prototype and a training system your company can rely on is where the real cost lives:
- Character quality is prompt engineering, forever. Characters drift, break role, or go flat, and keeping them believable across model updates is ongoing skilled work.
- Evaluation is its own product. Scoring a conversation against rubrics, consistently and defensibly, is harder than running the conversation.
- Safety and boundaries. You own every off-script thing your character says to an employee.
- Voice, latency, and cost. Real-time speech with emotional range is a specialist engineering problem, and model costs need managing at scale.
- Integration and analytics. Learners need it inside courses and the LMS; you need dashboards, transcripts, and completion data. That is a second product.
- Maintenance has no finish line. Models deprecate, APIs change, and the two engineers who built it get pulled onto other work.
If your organization ships conversational AI products, none of that scares you, and building might genuinely fit. For everyone else, the build path converts your L&D budget into an under-resourced software team, and it is worth asking a bigger question first: will AI replace instructional designers, or is your team's time better spent evaluating tools than building them.
What criteria matter when evaluating AI simulation tools?
| Criterion | The question to ask | Why it matters |
|---|---|---|
| Course integration | Does it run inside our existing Storyline, Rise, or LMS content? | If practice lives in a separate destination, most learners never arrive |
| Authoring speed | Can an ID build a working simulation this week, in plain English, without code? | Adoption dies when every scenario needs a developer |
| Character realism | Do characters improvise, push back, and stay in role? Is there voice with real emotion? | Realism is what makes practice transfer |
| Coaching and scoring | Is there in-conversation support and a scored debrief against criteria we define? | Practice without feedback is just chat |
| Analytics | Can I see pass rates, scores, attempts, and unique learners on a dashboard? | This is your ROI evidence |
| Iteration speed | When we edit a scenario, how fast is the change live? Republishing required? | Weekly-iteration tools beat quarterly-republish tools |
| Security | Is there a SOC 2 report? Where does learner data go? Is it used for model training? | Your security team will ask; ask first |
| Pricing model | Per seat, per learner, or per simulation? What happens at renewal and at scale? | Pilot pricing and rollout pricing can differ sharply |
| Vendor trajectory | Is the product improving monthly? Who is behind it? | You are buying a roadmap in a fast-moving category, and my overview of AI in design tracks where these tools are headed |
Two of these deserve leader-level attention because your practitioners cannot evaluate them from inside the tool: security and pricing at scale. Get the SOC 2 report and the data-flow answer in writing, and price the tool at year-two volume, not pilot volume.
Full disclosure: devlin.ai is my company and it competes in this category, which is exactly why I am comfortable publishing the hard questions. We hold a clean SOC 2 Type 1 report with Type 2 in progress, and I think every vendor should have to answer this table.
How should you run a pilot?
Two weeks of building, then a defined evaluation window. The goal is evidence, not a feeling.
Pick one high-value conversation skill
Choose a skill with a visible business metric: sales objection handling, support de-escalation, manager feedback. Avoid piloting on low-stakes content; a pilot that cannot show value was designed not to.Let your practitioner build it
Your instructional designer describes the character, scenario, and scoring criteria, the kind of work I cover in my broader guide to AI in instructional design. Clock the authoring time; it is one of your evaluation data points.Define pass criteria before launch
Agree on what a passing performance looks like and what pass rate you expect after two attempts. Write it down before the first learner enters.Run 20 to 50 learners through it
Enough for a real signal on completion, pass rates, and repeat attempts. Read a sample of transcripts yourself; they are the most honest artifact in the pilot.Debrief against the criteria table
Score the tool on authoring speed, realism, coaching quality, analytics, and iteration speed. Then decide on rollout, and read my rollout briefing before you do.
The rollout side, including who owns scenarios and what a 30-60-90 plan looks like, is covered in Rolling Out AI Simulations. The budget defense is covered in The ROI of Practice-Based Training.
What security questions should procurement ask?
Send the vendor these five, verbatim:
- Can you share your current SOC 2 report and the status of any audit in progress?
- Which AI model providers process our learner conversations, and under what data-processing terms?
- Is any of our data used to train models, yours or a third party's?
- What learner-identifiable data do you store, for how long, and how is deletion handled?
- What is your incident-response commitment, and when did you last exercise it?
A serious vendor answers all five in one email. Evasion on any of them is your answer.