Make Them Teach It
Notes on the AI teach-back, from a calculus classroom in Connecticut
“When it comes to assessing intelligence, do we honestly believe that someone who has a higher IQ score is somehow smarter? No, of course not.” — Ken Ono, Mathematician, Axiom Math and The University of Virginia
In April, a calculus teacher in Connecticut named Bill Fallon gave his twenty-one seniors an assignment that didn’t involve calculating anything. They had to explain what a derivative actually means to a classmate named “Devon,” who kept almost getting it and then slipping back.
Devon is a bot. He’s seventeen and in AP Calculus AB, with one specific hang-up. He knows the derivative gives you the slope. But a curve doesn’t have a single slope, since it’s steeper in some places and flatter in others. So the slope of what?
Bill and I built him over a handful of conversations, working out what the character should be stuck on and where the whole thing should sit in the unit. He ran it in his classroom, we surveyed his students, and afterward the two of us went through the transcripts together and talked about what worked and what didn’t. The rest of this is that story.
Before I go any further, let me acknowledge that I have a stake in this. I’m building a platform that makes these simulations easy for teachers to run, we launch next month, and I’d obviously prefer you finish this thinking they work. So please, read accordingly.
I picked a math class on purpose. When I describe using the approach of conversational simulations to combat AI cheating and to build durable skills, the response I usually get is that it sounds like an English department exercise. It’s dialogic, usually involves argumentation or exposition skills, and leans heavily on a student’s grasp of language. But it expands beyond that. Math is the hardest case to make, which is why I engaged in the experiment with Bill.
None of this is new, pedagogically speaking. In 2009 a Stanford group published work on what they named the protégé effect, built around a program called Betty’s Brain. Students who believed they were teaching a computer agent spent more time on the material and learned more than students studying for themselves, and the effect was largest among the lower-achieving kids.
In June 2023, Ethan and Lilach Mollick laid out seven ways to assign AI in a classroom, and one of the seven was AI-student, where the model plays the learner and the student does the teaching.
What I’ve tried to add to the process is targeted resistance, a record, and an evaluation mechanism. The character – in this case, “Devon” – pushes back on explanations they can’t use. The conversation leaves an artifact that a teacher can read and evaluate for understanding afterward.
What Devon does
Devon opens the same way every time:
Then he waits.
If you’d like to take “Devon” for a test drive, here he is: [Sim Link]. It’s supposed to be annoying, so notice the moment you want to quit. That moment is most of what I’ve been studying the last few years.
Your chats are anonymous, no personal information is captured, and it ain’t flashy, so try and leverage your imagination rather than focus on the tech stack.
How you make a “Devon”
When people hear AI role-play, they assume that all you have to do is type act like a confused student into a chat window and hit go.
The truth is; you can. But in my experience, it falls apart quickly, because the model capitulates. Your learner says something half-formed, the character says oh, that makes sense now, and it ends in a compliment. Most of the design is spent stopping the model from being nice. The other half is about teaching it how to react when the student makes their attempt at teaching.
Bill’s half of this was knowing what a calculus student actually gets stuck on. Mine was building the character with enough depth to make the conversation useful. I just finished an MFA in creative writing, where I spent a year running these simulations on myself and keeping a diary about what happened to my own mind. That paper is out for review now. But it’s also how AI Friction Labs was born.
When designing these characters, the first requirement is that the character has to want something – it can’t just be something. In this case, Devon wants to understand what the derivative describes. He has a test coming, and the student that enters the simulation is the gatekeeper to that understanding. Everything else follows from that tension, including the frustration when a student repeats the same explanation without adjusting, adding, or deleting. Devon’s voice – so wait, okay but, no hold on – acts as a set of immersive clues.
Then come the rules, which have to leave him somewhere to move. If a student’s answer is delivered as jargon with nothing underneath it, Devon says he doesn’t know what that means. He asks one question per turn and then stops. He never praises (unlike the Khanmigos of the world). Underneath all of it sits a ladder of five stages of understanding with a ceiling he isn’t allowed to pass. In other words, Devon can get close but can’t finish the thought himself, because the moment he can, the student is off the hook.
He is also built to slip back. A student explains the tangent line, Devon says he’s got it, and two turns later the confusion can resurface. That’s on purpose, and it’s meant to teach that explaining something to someone else is never a one-off. It’s also the part students hate the most, which I’ll get to.
Here’s the full build, free to take: [Link to System Prompt]. Run it as a CustomGPT or Gemini Gem and you’ll get something that half works. The character drifts after a while and there’s no record of what happened afterward. What we’re building are mechanisms to view that artifact, evaluate it easily, and engage in a dialogue with your student about what we as human beings can learn from these experiences.
Two students, same phrase
Twenty-one students ran the simulation twenty-nine times over one week. Every transcript was scored by an AI evaluation that Bill and I then read back against the conversations themselves. Scores averaged 82 out of 100 and ran from 23 to 96. Five students went back for a second attempt without being told to. All five improved, by an average of 33.6 points.
Two of the transcripts, three days apart, are worth reading side by side – a version of what I call “Comparative Transcript Analysis.” Both students used the same phrase, telling Devon the derivative gives you the slope of an “imaginative line” at a point.
Devon didn’t take it from either of them. I hear those words. But what even is a tangent line, really? And how does a curve, which is all bendy, have an “exact steepness” at just one point?
The first student defended the phrase. Then defended it again. Then, cornered, told Devon it was impossible to find the exact slope and that any line works as a tangent. Devon lost the thread, and the session ended with him worse off than he started. He scored 50, with 25 out of 30 for persistence, since he never stopped trying, and 10 out of 30 for explanation.
The second student dropped the “imaginative line” phrase. Devon then rejected “instantaneous rate of change,” then “limits.” So the student stopped using them and started talking about direction, about where the point was heading. It took him forty messages, but he got Devon to a point of understanding and scored a 93 on the AI evaluation.
Worth noting; each student used the same “imaginative line” phrase at turn three, but ended up finishing forty-three points apart. One was rewarded for persistence, patience, and creative problem-solving. That’s a crucial point about this method. It won’t just give you a record of what they know about the subject, but also how they behave when resistance in that field is presented.
In the case of the successful student, they didn’t stick with the same approach, they found a new one. They saw Devon’s confusion as information, whereas others saw the confusion as a reason to stop. The AI wasn’t working, they thought. But that type of failure creates a different kind of learning experience that, I believe, focuses on human development alongside the acquisition of academic content.
Bill put it this way:
“Things are going to frustrate you in life. People are going to frustrate you. This bot is frustrating you. How do you work around that and make it so that you are productive and helping the person as best you can? That is a whole other level of thought and learning.”
What I got wrong
One of the things we learned from this beta test was that one rep wasn’t enough. Students needed feedback and a chance to try again. Our mechanisms weren’t built for that process at the time, but we’ve learned from it and are working on it as we speak.
That concept reminds me of the popular “Ungrading” approach where students receive unlimited attempts to show mastery. It also reminds me of Jason Gulya’s “Try-Again Classroom” concept, a salve for modern times if there ever was one.
Another important note relates to the AI scoring mechanism. Evaluating a conversation isn’t objective. Our AI scorer can never be 100% right. But that’s where the educator comes in. A cross-referencing approach can lead to a faster, more viable route to “grading the chat” than hand-grading (which I have done before and is a doozy.) Bill’s take was that the evaluation scores were not only useful but largely accurate, for what it’s worth.
And, to repeat the point, it allows you to choose whether you are giving math feedback or human feedback. Do you want to correct the language they use to describe the tangent line? Or do you want to teach them about resilience, patience, persistence, and a growth mindset? Either are available, and the transcript is the avenue by which those teaching points are delivered.
If you’re interested in learning more, we are hosting a free virtual launch event on August 19th. You can register for it here - we hope to see you there: [AI Friction Labs Faculty Demo Event Registration]
In the meantime, let’s keep looking for creative ways to redesign assessments – not just so that we can end-around AI cheating (which is important), but also so that we can test students on the types of skills they will have to deploy in their day-to-day lives when they leave our classrooms.
A Note on the Epigraph
I went down a bit of a math rabbit hole over the last few weeks. I’ve got some more content coming on the wave of mathematical conjectures that have either been solved or disproven by AI recently.
But before that, I was caught by this video from EO Studio sharing Ken Ono’s thoughts on AI’s impact on intelligence, mathematics, and the economy.
To hear someone with so much experience in the field of mathematics and technology speak this calmly about AI and its impact on our world put my mind at ease, if only for a moment. I recommend giving it a watch for yourself, but also consider sharing with students. Here’s the full quote from the epigraph from the beginning of the video:
“We live at a time where the world places so much emphasis on benchmarks…There’s a lot of anxiety. When it comes to assessing intelligence, do we honestly believe that someone who has a higher IQ score is somehow smarter? No, of course not. It should people at ease. If you have to live up to the standards set by someone else, then you are not living for yourself. You’re not giving yourself credit. I consider that toxic.”
Until next time.
The free virtual launch event is Wednesday, August 19th at 1pm EST. We’ll demo the platform and you can spar with a friction bot yourself, which I’d recommend, since it’s a different thing on the receiving end. Register here: [AI Friction Labs Faculty Demo Event Registration]






