Testing a roleplaying scenario sounds straightforward until you try to define what “passing” looks like.
Imagine two learners completing the same workplace conversation.
One says: “I understand why that would be frustrating. Let me see what I can do.”
The other says: “Yeah, I can see why you’re annoyed. Give me a second and I’ll look into it.”
Those are different answers. They could both be perfectly good answers, but it creates an interesting testing problem.
There may not be one correct response
Traditional software testing works exceptionally well when we know exactly what should happen. Give a calculator 2 + 2 and you know what answer you want.
A conversation does not work that way.
There might be dozens of appropriate ways to acknowledge someone’s frustration. There might be hundreds. If we test an AI roleplay by looking for one exact sentence, we have basically defeated the purpose of using conversational AI.
But the opposite approach creates another problem. If almost anything can count as acceptable, what exactly are we testing?
I have dealt with that problem from the testing side before.
Before SkillSaige, I was the project manager on a proof of concept for a more traditional prompt-based AI training tool. That put me pretty deep into both the prompt engineering and the QA process.
A lot of testing looked something like this.
Run the scenario. Find a response that does not behave the way you expected. Add instructions to the prompt. Run it again. Sometimes that fixes the problem. Often it creates a new one.
It becomes a game of whack-a-mole, and because an LLM can produce different outputs from the same starting point, even something you thought you had fixed still needs to be tested again from multiple directions.
Those experiences helped shape how we thought about SkillSaige from the very beginning. We knew we did not want to build a training system where QA meant continually expanding a prompt and hoping we had covered enough edge cases.
Our CEO Sayre wrote more about the broader testing problem on her blog, particularly why LLMs cannot always be tested the same way as conventional software.
For training, I think the answer starts by testing something other than the exact words.
Test the objective, not the script
Consider a scenario where the learner is practicing a difficult conversation with an employee.
The training objective might include behaviors like:
- Clearly explaining the issue
- Avoiding unnecessarily confrontational language
- Establishing a reasonable next step
There are countless sentences someone could use to accomplish those objectives.
So the system should not care whether the learner used the exact wording we imagined when writing the course.
It should care whether they accomplished the objective.
That gives us something much more useful to test. Can the system recognize multiple reasonable ways of accomplishing the same goal? Does an ineffective response actually create a different outcome? Does the scenario continue behaving according to its rules even when the conversation itself takes an unexpected path?
Those are much more meaningful questions than whether the AI said exactly what we expected.
Test the boundaries too
Good testing also means deliberately trying to break the scenario.
What happens if the learner ignores the problem? What happens if they are rude? What happens if they say something completely unrelated?
One of the biggest risks with AI roleplay is that the simulated character becomes too accommodating.
The learner says something that should make the conversation worse, but the AI tries to be helpful and moves things forward anyway. Now the learner may be practicing the wrong behavior and still receiving a successful outcome.
That is not just an AI problem. It is a training problem.
A simulation needs to preserve cause and effect.
The learner’s choices should matter.
Repeat the same test in different ways
There is another important wrinkle.
Because the conversation can vary, testing something once is rarely enough.
You want to try different wording. You want to approach the same objective from different directions. You want to repeat the scenario and make sure that conversational variation does not change the underlying rules.
The character does not need to say the same sentence every time. In fact, it probably shouldn’t.
But if the learner has not resolved the customer’s concern, the customer should not suddenly behave as though the problem has been solved. That distinction is incredibly important.
The words can vary. The logic should remain dependable.
A roleplay is successful when the experience holds together
That is ultimately what we are testing. Not whether the model can generate a convincing paragraph. Not whether every conversation follows the same script.
We are testing whether the simulation continues to behave like the situation it is supposed to represent.
Can learners try different strategies? Do those strategies produce sensible consequences? Do the learning objectives remain intact? Can the system handle the messy variety of ways real people communicate without losing track of what the exercise is actually teaching?
If so, you have something much more useful than a chatbot with a good prompt. You have a training simulation.