Does Practicing on an AI Actually Work?
What the studies actually say about AI role play simulations
“The voice went on and on, mild and deliberate, inflexibly gentle; the small girls listened intently, framing in their minds little pious sentences with which to surprise their parents, and the boy yawned against the whitewash. Everything has an end.” — Graham Greene, The Power and the Glory
I’ve been writing for a while about academic role-playing exercises and conversational simulations that are designed to push back. I call them friction bots. When I talk to other educators about their value, especially professors, I often get the question, “Ok, but what does the research say?”
Stands to reason. What good is a method without empirical data to go along with it? So I spent some time this week rounding up some of the published studies over the last few years that tracked the efficacy of LLM-based role-playing simulations as a learning tool.
Before I get to sharing them, I want to acknowledge my bias. I’ve been touting this approach in one form or another for 2-3 years and now am in the process of building out a platform that aims to make all of this not just possible for teachers, but easy. So, you should read my own translation of these experiments with a grain of salt. Always consider the author.
That said, I sell this idea because I believe in it. I use this AI-powered approach in my own classroom. My students clamor for more. I don’t use it every day or every week, but I think it has the potential to change education in a positive way. So, to the extent that this resonates with you, I hope you can see that my bias comes more from a place of experimentation than unadulterated hype.
Thanks for listening. Onward we go.
1. Medical education — the virtual standardized patient RCT
Perhaps the most intuitive application of the approach is in the medical field, where research indicates that simulations can surface “clinical thinking ability” in students. In JMIR Medical Education Journal, researchers created 3-D virtual simulated patients that present system-preset internal-medicine cases in mimicked clinical scenarios.
They engaged with 60 medical students interning at a hospital. Thirty were trained on the system and engaged with it, the other thirty received traditional academic case-based training. They measured effectiveness via basic knowledge assessments and virtual system scoring.
The VSP students showed a theoretical score improvement roughly six and a half points higher than the control group (17.07 mean vs. 10.67). The authors concluded that the virtual experience reinforced foundational internal-medicine knowledge, enhanced clinical thinking ability, and improved capacity for independent work. They called the system feasible, practical, and cost-effective.
My take: This is a field where simulated patient care has long been an accepted practice. Medical education experts rely on these activities to develop “bedside manner” and an ability to take complex knowledge and, for example, explain it to a patient who can’t understand jargon. The simulations often test the doctor’s ability to ask the right questions, which stems from one’s own background knowledge in the field. But more than that, they test patience, communication skills, and empathy — the soft skills side of a profession that requires an immense amount of concrete knowledge and skill.
2. Counseling — the CARE study
The LLM simulation approach also appears to show promise in the counseling field.
In “Can LLM-Simulated Practice and Feedback Upskill Human Counselors?” researchers at Stanford University and Georgia Tech created an AI patient. The LLM simulates a client in a counseling session presenting “real-world challenges.”
The researchers put 94 novice counselors into two groups:
Group 1: Practice with the AI patient alone without receiving feedback
Group 2: Practice with the AI patient plus AI-generated feedback that offered suggested alternative responses with rationales.
They then measured behavioral performance, self-efficacy, and qualitative reflections. Their conclusions were that the practice-plus-feedback group improved in client-centered microskills, while the practice-alone group showed no improvements. Furthermore, with respect to empathy, the practice-alone group declined over time and performed significantly worse than the feedback group. And, when researchers interviewed participants, the ones who received feedback adopted a client-centered listening approach. Practice-alone participants remained solution-oriented.
My take: This study goes further than most. Instead of measuring the efficacy of an LLM-based role-playing simulation against traditional forms of instruction, the study places both groups in LLM sims and asks – “What difference does AI-generated feedback on the simulated performance itself make?”
It is not entirely controversial to say “feedback is good.” But it is somewhat controversial to say “AI-generated feedback is good.” So the question I have coming out of this study is, “why did the AI-generated feedback appear to work in this case?”
I think this is a question that is ripe for further research. AI-generated feedback on essays, for example, face reasonable detraction. Might it be possible that AI-generated feedback on a different set of skills actually is effective?
3. History — the medieval peasant simulation
Stepping out of the medical and psychiatric fields, there is growing evidence that LLM’s make it possible to simulate perspective in a way that we have not had access to in the past.
In Teaching History, Benjamin Breen of UC Santa Cruz describes integrating an LLM-enabled simulation into a one-week module on the social history of medieval peasantry. In this one, the student plays the character. Breen developed a prompt that reliably generated a randomized persona of a medieval peasant, a “playable character” the student inhabits for a day in the life, set in a randomly chosen location between Iceland and the Levant, sometime between 900 and 1400 CE.
In the first session, Breen delivered a lecture, circulated primary sources (both textual and visual), and completed a live playthrough on the classroom projector before students touched it. Students then ran the simulation themselves, with follow-on assignments including “prompt revision,” in which they altered the prompt to retell events from a different perspective.
The findings are qualitative (classroom case studies, no control group). Students described the simulations as encouraging historical empathy, the realization that the people they study were once real human beings with normal lives. Breen concluded that the simulations worked best when integrated into a larger cycle of discussion, analysis, reflection, and research, and not as one-off engagement activities.
My take: The feedback I received from our beta partners over the last six months of experimentation follows a similar pattern. The inclusion of pre-and-post reflections, accompanied by discussion and analysis, made a significant difference in the perceived efficacy on the various tests.
It also creates a portfolio of student learning artifacts that create a narrative of their own experience. When consumed together, the educator sees an arc that tells the story of their student’s journey.
4. Teacher education — AI students in a virtual classroom
One of the criticisms that many educators levy against the teacher training programs that they experienced in undergraduate or graduate level education settings is that the traditional form of study doesn’t teach you how to manage a classroom. The pre-service teacher may learn quite a bit of meaningful pedagogy and write many papers on what is effective and what is not, but how do they get prepared to manage the outlandish and disruptive comment from the student in the back-row? How might they practice keeping students on-task when the weather is nice and it’s Friday at 1:30pm? How might they receive “reps” in responding to genuine misconceptions that you didn’t plan for?
LLM-based role-playing simulations show great promise in this area.
In Education Sciences, researchers introduced TeacherGen@i, a generative AI-enhanced virtual reality simulation for pre-service teachers. The characters are the students: virtual agents that respond verbally in real time, display dynamic facial expressions, and sit in a virtual classroom where the trainee’s own course slides appear on a display. The system ran on the participants’ personal laptops.
The study was an explanatory case study with a mixed-methods design, embedded in a university teacher-preparation course. Pre-service teachers taught lessons to the AI students, and the researchers collected usability surveys and reflective journals, analyzed through thematic coding and computational linguistic analysis.
The findings suggest the simulation facilitated meaningful development of teaching competencies, specifically instructional decision-making, classroom communication, and student engagement. The authors also documented design limitations: cognitive load, user interface problems, and insufficient instructional scaffolding. They frame the work as exploratory.
Related work in the same space embedded GPT-4 student agents in a 3D Roblox classroom, where pre-service teachers practiced problem-solving lessons, with pilot studies indicating high usability and learner engagement.
My take: We’ve run teacher education simulations as well. However, we’ve focused on conversations that happen outside of the classroom.
Just last week, faculty at Portland Community College engaged with a friction bot that presents one of two difficult colleagues. One is the type to agree to everything but never actually execute the plan. Another is the type who disagrees with everything. In each case, the task aligns with the conflict point. Can you get the commitment-averse educator to firmly agree to follow through on a plan? How? In the second, how do you navigate disagreement without damaging the relationship? Can you find a center ground that is both true to what you believe but also takes into account the unique perspective of your counterpart?
One PCC faculty member described it as “hitting close to home.” The realism of these experiences at a minimum provides a space for rehearsal of difficult conversations. It’s “a safe place to fall on your face.” Wouldn’t that be nice?
5. Negotiation — ACE
No one ever taught me how to negotiate. I remember failing miserably to secure a raise in one of my first year-end reviews. A few years later, I studied negotiation methods before entering a similar meeting. I didn’t have a place to practice, but I had a plan, and that’s more than I can say for most situations. I succeeded in getting some of what I wanted, but not all. I look back and wonder; what if I had a place to practice that negotiation? In a place devoid of fear of judgment or failure? Might I have performed better in the moment?
In that vein, researchers at Columbia University built ACE (Assistant for Coaching Negotiation), presented at EMNLP 2024. The character is a negotiation chatbot that serves as the learner’s bargaining counterpart, with a coaching layer built on top of it.
To build the system, the team collected a dataset of negotiation transcripts between MBA students, trained negotiators working through realistic single-issue bargaining scenarios. With expert consultation, they designed an annotation scheme for detecting negotiation mistakes. ACE uses that scheme to identify the user’s errors and deliver targeted feedback, alongside examples written by experts.
They then tested it with 374 participants, each completing two rounds of negotiation, across three conditions: ACE’s feedback, no feedback, and an alternative feedback method. ACE significantly improved negotiation performance compared with both other conditions. Participants in the ACE condition also reported higher perceived improvement in their second negotiation, a mean of 4.23 on a five-point scale. The authors built the system because negotiation skills are traditionally taught through resource-intensive seminars that most learners never get access to.
My take: For our part, we have run pre-hire job candidate screening simulations for multiple businesses in the same regard. In one, a candidate had to negotiate with a room full of conflicted stakeholders. In another, the candidate had to defend their investment thesis to a difficult investment committee member. Negotiation-adjacent on that last one, to be sure, but a practice space for flexing the same muscles.
What’s it all mean?
This the tip of the iceberg. Let’s recap some of what we know:
Role-playing has been proven time and time again to be good for learning.
Rehearsal is good for the brain, too.
LLM’s are really good at mimicking human beings.
Human-human “simulations” are great! We shouldn’t abandon them. But what LLM’s may introduce is a way to scale the practice.
The advancement of LLM’s in the past few years mean that simulations can now be executed in a far more dynamic, adaptable, and immersion-inducing manner.
What we need are the right pedagogical mechanisms around the LLM simulation for this to really be useful. Beyond that, the “craft” of producing these virtual rooms, simulated conversations, and role-playing dialogues needs to advance and be codified across a wide range of educational contexts.
One more note: These are just a few of the studies out there. I worked with Claude to produce a one-sheet of relevant studies in the field. You can find it here, nothing fancy.
Furthermore, there are many educators already running these simulations informally/as experiments. Substack is filled with them, though I didn’t have time to curate and place them all inside this article. If you know of any good ones out there (or if you have one yourself), please drop them in the comments.
What I’m Reading/A Note on the Epigraph
Graham Greene’s The Power and the Glory. My uncle is a big Greene fan, but this is my first foray into his work. What I’m most enjoying is Greene’s ability to shift perspectives quickly and often without losing me. I’m only about 4-5 chapters in. Here’s how it’s going:
We begin with a nameless dentist in a raggedy Mexican town who ends up interacting with the only remaining Catholic priest in a war-torn state where all the other Catholic churches and priests have been eradicated. By Chapter Four, the dentist is gone and replaced by the lieutenant chasing him, a fellow priest who gave into the rebel’s demands and broke his celibate vows, a surreptitiously devout Catholic family in the same town, and an English sea captain who finds out his daughter has been harboring the fugitive priest without his knowledge.
Greene has me spinning like a dreidel through each of their minds, and yet I am not mad about it. I’ve tried to write in the same way. For my money it is one of the greatest challenges a writer can face.
The quote in the epigraph was chosen less for its thematic meaning and more for its style. I wonder what Strunk and White would have said about it. Actually, who am I kidding? I don’t care. I think it’s beautiful, and that’s enough for me.
Until next time.



Your open question, why AI feedback worked here when feedback on essays so often disappoints, may have its answer buried in the study’s failure case.
The counselors who practiced alone did not merely fail to improve. They became less empathic. In the authors’ discussion, the reason is revealing: the AI patient responded no differently to empathic and non-empathic statements. The counselors read the room correctly and concluded that empathy was inert. They gradually stopped using it.
The simulation taught its own blind spot. What the simulated patient could not register, the learner learned to stop doing.
That changes the role of feedback. Practice alone was never norm-free; it trained participants on the simulator’s defaults, and those defaults worked against the skill the exercise was meant to build. Feedback succeeded because it overrode a lesson the raw interaction was already teaching through its silence.
The sharper validation question may therefore be this: what does the simulated interlocutor reinforce by responding, and what does it extinguish by remaining flat?
I've been working on building AI coaches. In my case, I teach techniques to access one's own resourcefulness or to handle conflict. When people experience a technique a few times, they learn how to both do it on themselves when they need it and to coach others using the technique. If you decide you want a closer look, reach out and I can show you.