Two clinical simulation platforms can both claim to be AI-enabled and still differ substantially in the value they deliver to your program. The differences often come down to questions that go unasked:
- Does the platform strengthen clinical reasoning across an entire clinical encounter, from history through patient management with multi-role interactions, or does it focus only on the patient conversation?
- How much control do faculty have over case content?
- How rigorously does the tool assess performance against the curriculum and current professional standards for clinical practice?
- In case development, is AI confined to scaling the work of clinical experts, or does the platform rely on AI to generate clinical content and judgments on its own?
If your program is evaluating AI-enabled clinical simulation platforms this year, it is worth seeing firsthand how a platform performs against this checklist and your own set of expectations. Use the following checklist to ask critical questions during your next vendor conversation, along with what to look for in each answer.
Checklist to evaluate AI-enabled clinical simulation platforms
Platform focus
1. Does the platform help build clinical reasoning across an end-to-end clinical encounter or does it stop at the patient conversation? Check for multi-role interactions where learners work with a patient, and also carry the case through physical exam, diagnostic workup, differential diagnosis, and management decisions.
3. Does the platform respond dynamically while keeping AI patients clinically accurate and behaviorally consistent? Ask for evidence that subject matter experts and learners test each case before launch.
4. Does the interface guide learners in real time, or does structured feedback only arrive after the encounter ends? Ask for evidence of real-time support during the case and structured feedback at the end of each encounter.
5. Does the platform give faculty and administrators a performance dashboard that tracks learners over time? Look for longitudinal data at the individual and cohort level that shows where reasoning is strong or breaking down, so faculty can make informed decisions about remediation, learning, and curriculum design.
Case exposure
6. Does the case library offer clinical realism? Look for scenarios that go beyond patient conversation, including multi-role scenarios that intentionally prepare learners for real-world experiences.
7. Does the case library cover a broad range of scenario types? Ask to see cases that range from clinical skills practice, like X-ray and ECG interpretation, to exercising clinical judgment, across specialties such as Neurology and OBGYN.
8. Can the simulation library be personalized to match your curriculum? Benchmark the ability to map existing cases accordingly and request new cases built around your curriculum, population focus, and program-specific requirements, rather than reliance on a fixed library.
9. Are cases mapped to the competency frameworks your program is accountable to? Check for mapping to the standards your accreditors use, e.g. NONPF, AAMC, AACOM, ACGME, AAPA, and others.
10. Who creates the clinical content, and what role does AI play? Verify that cases are authored and reviewed by credentialed clinicians, with AI limited to delivering and scaling their expert-designed content rather than leading development.
Assessment rigor
Assessment rigor in an AI-enabled simulation platform depends on two things: the quality of the standards learners are scored against, and how reliably the AI applies them. Faculty should ask who writes the scoring criteria, how accuracy was validated against expert graders, whether scores stay consistent, and how much visibility faculty have into every score.
Standards
11. Is performance scored against multiple rubrics tied to specific competencies, or does it collapse into a single number? Confirm separate scores for competencies such as history taking, differential development, communication, and management.
12. Who writes the rubrics, and can the AI change them? Insist on rubrics built with clinical experts from published, evidence-based frameworks, and AI that cannot add, drop, or reweight criteria on its own.
13. Is the same rubric framework used across all relevant cases, or does assessment vary widely from case to case? Look for one consistent framework, so learner performance is comparable from case to case.
Reliability
14. How was scoring accuracy measured, and against what benchmark? Ask to see a reference set of learner transcripts scored by clinical experts, with disagreements resolved by consensus.
15. Will the same learner transcript receive the same score every time? Verify evidence of repeated testing of the same transcripts before any new model or prompt is released.
16. How were proficiency thresholds set? Look for cut points based on real learner data at a comparable stage of training, reviewed with faculty.
Oversight
17. Can faculty see and audit every score? Check for access to full transcripts, rubrics, and the AI's written feedback for every learner.
Implementation and support quality
18. Will the vendor's team respond quickly and personally to your program's requests? Ask for direct access to the clinicians and technologists who build the platform, and the capacity to develop solutions specific to your program. Vendors that are part of very large corporations may route faculty requests through layers of support. Very small vendors may lack the resources to act on those requests.
How does DDx by Sketchy stand up to this checklist?
Platform focus
Every DDx conversation, including spoken interaction with realistic patient avatars, responds to what a student actually says instead of a fixed script. Learners at the University of Cincinnati found DDx "felt far more realistic" because it was "much easier to communicate with the patient" than their previous simulation tool.
Before publication, every case goes through QA testing, the final phase of case creation. Subject matter experts and students in relevant training programs complete each new case and check that patient personas stay within the case.
DDx's performance dashboard gives faculty and administrators longitudinal data at the individual and cohort level, broken down by competency, so they can see where reasoning is strong and where it is breaking down. The dashboard also provides remediation guidance, helping faculty identify which learners or cohorts need targeted support and where curriculum adjustments may be needed.
- Case exposure
DDx's cases are written by an interprofessional team of credentialed medical providers. Cases for NPs arrive already mapped to NONPF domains and AACN Essentials, and faculty control curriculum alignment by choosing which cases suit their courses, pedagogy, and student body. When a program needs something the library doesn't cover, the DDx team edits existing cases and builds custom cases to match the program's curriculum and requirements.
Assessment rigor
- Standards. DDx scores learners in two layers:
- Key finding detection checks whether a learner elicited or addressed the specific findings the case author defined, such as asking about a relevant symptom or ordering a key test.
- A rubric is tied to the instructional focus of the case, such as clinical reasoning or communication.
AI does the scoring and writes the feedback, using expert-written rubrics and key findings as strict guidelines. It can't add criteria, remove them, or change how they are weighted.
Every rubric is mapped to nationally recognized competency frameworks, including ACGME, AACOM, AAMC, AAPA, and NONPF, so assessment is consistent and transparent across learners and programs. Rubrics are based on published literature and widely used frameworks. Our communication rubrics, for example, draw on frameworks such as SHARE for shared decision-making.
Our Clinical Reasoning Assessment was developed with faculty experts from the US and Canada. It includes clearly defined performance anchors across key dimensions of diagnostic and management reasoning, based on published, validated tools such as ART and IDEA.
- Reliability. Before any scoring model or prompt reaches learners, we test it against a reference set of learner transcripts scored by DDx clinical experts. Each key finding is reviewed by multiple team members and finalized only when all reviewers agree, so the set represents full expert consensus. Because generative models can vary from one run to the next, we score each transcript in the reference set thousands of times with the same model and prompt. This confirms stability and consistency before any change is implemented.
Each DDx rubric defines what a learner is scored on, and learner data determines where the performance levels fall:
- Preliminary: 60–69%
- Progressing: 70–89%
- Proficient: 90% and above
We set these cut points using the score distribution per criterion from more than 10,000 completed cases, and reviewed them with faculty from NP, PA, and MD/DO programs. Our rubrics keep the standard fixed for every learner, while the cut points reflect how learners at a similar stage of training actually perform.
- Oversight. We monitor DDx cases continuously to make sure grading stays accurate and that each case keeps performing as designed after it goes live. Learners and faculty can rate a case and leave feedback after every attempt. Learners can also report an issue with a specific key finding, either during the conversation or any time afterward while reviewing their performance. Our team reviews case scores and feedback every week. When a case's scores fall outside the expected range, or someone reports an issue, the content team reviews the case and applies changes if needed. Faculty always have access to student conversation transcripts, full rubrics, and individual and cohort-level performance data.
Implementation and support quality
- DDx is built by a small, dedicated, agile team of clinicians and technologists who respond quickly to requests from program faculty. We're small enough to know you and care about your program, and large enough to develop the personalized solutions you need.
Set up a quick meeting with the DDx team and put us to the test.
FAQ
What should faculty evaluate when comparing AI-enabled clinical simulation platforms?
Faculty should compare four areas:
- Platform focus: clinical reasoning versus recall.
- Case exposure: library depth and faculty control over cases.
- Assessment rigor: expert-built, competency-mapped rubrics applied consistently by AI.
- Implementation and support: how quickly and personally the vendor's team responds to faculty requests.
What should faculty ask a simulation vendor before starting a pilot?
Programs should ask:
- Is the platform adaptive and built for clinical reasoning, or mainly recall-based?
- Does the case library cover more than patient conversation?
- Can cases be edited or custom-built to match their curriculum?
- Does assessment tie to multiple competency-based rubrics rather than a single score?
Because most platforms describe themselves as AI-enabled, these four areas reveal what a pilot will actually deliver.
How can faculty tell whether an AI simulation platform scores learners accurately? Faculty can tell whether an AI simulation platform scores learners accurately by asking how the vendor validated its scoring against expert graders. Strong answers describe a reference set of learner transcripts scored by clinical experts, a consensus process for resolving disagreements, repeated testing for run-to-run consistency, and faculty access to every transcript and rubric.
Who writes and reviews the clinical content behind DDx's cases? DDx's cases are written by an interprofessional team of credentialed medical providers. Their clinical judgment sets the guardrails that keep AI-generated conversation accurate, clinically contextual, and consistent with professional ethics. Every simulated encounter reflects the real nursing, NP, PA, or medical scope of practice.