Devlin Peck

Build vs Buy: How to Evaluate AI Simulation Tools for Your Team

By Devlin Peck · Updated

Part of the AI in Instructional Design guide

Buy, unless conversational AI is your core business. That is my answer to the build-vs-buy question for AI training simulations, and I say that as someone who builds AI systems for a living. This briefing explains the reasoning, then gives you the criteria and the pilot plan for evaluating tools like a buyer instead of a spectator.

It is part of my full guide to AI training simulations, written for the L&D leader who has to sign off, not the practitioner who found the tool.

Should you build or buy?

The build option looks deceptively close. Your team can get a prototype conversation running against a model API in a week, and that prototype is exactly what makes leaders overestimate the build path. The distance between a prototype and a training system your company can rely on is where the real cost lives:

If your organization ships conversational AI products, none of that scares you, and building might genuinely fit. For everyone else, the build path converts your L&D budget into an under-resourced software team, and it is worth asking a bigger question first: will AI replace instructional designers, or is your team's time better spent evaluating tools than building them.

What criteria matter when evaluating AI simulation tools?

CriterionThe question to askWhy it matters
Course integrationDoes it run inside our existing Storyline, Rise, or LMS content?If practice lives in a separate destination, most learners never arrive
Authoring speedCan an ID build a working simulation this week, in plain English, without code?Adoption dies when every scenario needs a developer
Character realismDo characters improvise, push back, and stay in role? Is there voice with real emotion?Realism is what makes practice transfer
Coaching and scoringIs there in-conversation support and a scored debrief against criteria we define?Practice without feedback is just chat
AnalyticsCan I see pass rates, scores, attempts, and unique learners on a dashboard?This is your ROI evidence
Iteration speedWhen we edit a scenario, how fast is the change live? Republishing required?Weekly-iteration tools beat quarterly-republish tools
SecurityIs there a SOC 2 report? Where does learner data go? Is it used for model training?Your security team will ask; ask first
Pricing modelPer seat, per learner, or per simulation? What happens at renewal and at scale?Pilot pricing and rollout pricing can differ sharply
Vendor trajectoryIs the product improving monthly? Who is behind it?You are buying a roadmap in a fast-moving category, and my overview of AI in design tracks where these tools are headed

Two of these deserve leader-level attention because your practitioners cannot evaluate them from inside the tool: security and pricing at scale. Get the SOC 2 report and the data-flow answer in writing, and price the tool at year-two volume, not pilot volume.

Full disclosure: devlin.ai is my company and it competes in this category, which is exactly why I am comfortable publishing the hard questions. We hold a clean SOC 2 Type 1 report with Type 2 in progress, and I think every vendor should have to answer this table.

How should you run a pilot?

Two weeks of building, then a defined evaluation window. The goal is evidence, not a feeling.

  1. Pick one high-value conversation skill

    Choose a skill with a visible business metric: sales objection handling, support de-escalation, manager feedback. Avoid piloting on low-stakes content; a pilot that cannot show value was designed not to.
  2. Let your practitioner build it

    Your instructional designer describes the character, scenario, and scoring criteria, the kind of work I cover in my broader guide to AI in instructional design. Clock the authoring time; it is one of your evaluation data points.
  3. Define pass criteria before launch

    Agree on what a passing performance looks like and what pass rate you expect after two attempts. Write it down before the first learner enters.
  4. Run 20 to 50 learners through it

    Enough for a real signal on completion, pass rates, and repeat attempts. Read a sample of transcripts yourself; they are the most honest artifact in the pilot.
  5. Debrief against the criteria table

    Score the tool on authoring speed, realism, coaching quality, analytics, and iteration speed. Then decide on rollout, and read my rollout briefing before you do.

The rollout side, including who owns scenarios and what a 30-60-90 plan looks like, is covered in Rolling Out AI Simulations. The budget defense is covered in The ROI of Practice-Based Training.

What security questions should procurement ask?

Send the vendor these five, verbatim:

  1. Can you share your current SOC 2 report and the status of any audit in progress?
  2. Which AI model providers process our learner conversations, and under what data-processing terms?
  3. Is any of our data used to train models, yours or a third party's?
  4. What learner-identifiable data do you store, for how long, and how is deletion handled?
  5. What is your incident-response commitment, and when did you last exercise it?

A serious vendor answers all five in one email. Evasion on any of them is your answer.

Frequently asked questions

Is it cheaper to build AI simulations in-house?

Almost never once you cost it honestly. The prototype is cheap; the evaluation engine, safety work, voice infrastructure, LMS integration, analytics, and permanent maintenance are not. Buying converts an open-ended engineering liability into a subscription you can cancel.

What does a good AI simulation tool cost?

Pricing in this category is typically per seat or per active learner, and it varies with voice usage and volume. The more useful number is fully loaded pilot cost, which should be under a month of one instructional designer's time plus a small subscription.

How long should an evaluation take?

About 30 days: two weeks to build and launch one scenario, two weeks to gather learner data and read transcripts. If a vendor cannot get you to a live pilot in two weeks, that is itself evaluation data.