Last week, Sal Khan told Chalkbeat something that quietly landed in test prep circles: Khanmigo, the AI chatbot he spent years evangelizing as a revolution in learning, was “a non-event” for most students. “They just didn’t use it much.”
His analogy for why was sharper than anything in his 2023 TED Talk. Imagine, he said, that he walked into a classroom, sat in the back, and waited for students to seek him out. “Some will; most won’t.” That, he acknowledged, has been the Khanmigo experience.
Anyone who has worked with teenagers for more than five minutes could have told him this would happen. That’s not a cheap shot. It’s actually the most interesting part of the story.
The premise was always the problem
Khanmigo was built on the assumption that students would engage with a tool that didn’t give them the answer. It would Socratically guide them toward understanding instead, asking questions, nudging them forward. A thoughtful pedagogical design in theory. In practice, it required students to already possess the thing that struggling students most conspicuously lack: knowing what they don’t know.
Kristen DiCerbo, Khan Academy’s chief learning officer, explained to Chalkbeat that the tool can only respond to what students ask. Students, it turns out, are not great at asking. This is not a criticism of students—it’s a description of a cognitive reality anyone in education recognizes. When you don’t have the conceptual scaffolding to understand what you’re confused about, you can’t generate a useful question. An AI sitting patiently in the back of the room is no help if you don’t know you need to raise your hand.
That’s not a bug in Khanmigo’s design. It’s a structural limitation of reactive tools applied to students who need proactive intervention.
What the research actually shows
The evidence base for AI in education is, at this point, mixed, and anyone telling you otherwise is selling something.
The most useful study I’ve seen is a Wharton paper published in PNAS in 2025. Researchers gave nearly a thousand Turkish high school math students access to two AI tools: a standard ChatGPT interface and a version designed to give hints rather than answers. Students using unguarded ChatGPT solved 48% more problems during practice. When the AI was removed, they performed 17% worse than students who’d never had access at all. The scaffolding had been doing their cognitive work for them. The hint-based version eliminated the effect entirely.
Design matters enormously. Obvious in retrospect, but apparently not obvious enough in advance.
A Harvard study published in Scientific Reports the same year found that a purpose-built AI tutor outperformed in-class active learning for college physics students. It worked because it was built around the same pedagogical principles as the course. The students were also already motivated, enrolled, and engaged. They showed up. That last part is doing a lot of work in that finding.
The OECD’s 2026 Digital Education Outlook puts the central tension plainly: AI can support learning when guided by clear teaching principles, but without that structure, it boosts performance on a given task without producing real learning gains. Doing better on the homework is not the same as learning the material.

The student type problem
Every time I hear a claim about AI transforming tutoring, I think the same thing: it depends almost entirely on who the student is.
The research keeps returning to the same divide. Students with strong academic self-regulation who can manage their own attention, recognize when they’re confused, and ask coherent questions get real value from well-designed AI tools. Students without those skills tend to use ChatGPT to complete the work rather than understand it. A 2025 study found that AI dependence was tied to lower critical thinking with cognitive fatigue as the link. Students who leaned on AI the most ended up more exhausted and less capable.
The students most likely to use an AI tutor productively are the ones who need it least.
I don’t think that’s an argument against the technology. But it is an argument for being honest about what it can and can’t do and for whom.
What this means for families
The question I get most often from KFT families is some version of: can’t my kid just use one of these tools instead? Sometimes the answer is yes, genuinely. A motivated, self-directed student working on content they basically understand can get real value from a good AI tool for practice and feedback. I think that’s a legitimate use case, and I say so when I mean it.
But most students who end up in tutoring are there because they are not self-directing their way through confusion. They’re stuck. Often they’re also avoidant, which means the confusion has been sitting there long enough to harden into something that looks like attitude. An AI tool that waits for the right question can’t diagnose what went wrong two steps back in the conceptual chain. It won’t notice that a student’s algebra errors are actually reading comprehension errors. It can’t pick up on a one-word answer that means a kid has checked out entirely.
A good tutor is not primarily in the content delivery business. Content is everywhere. What changes in a session is that someone notices something and names it before the student can deflect. That part doesn’t happen in the back of the room.

The gap that hasn’t closed
Khan’s analogy is actually a near-perfect description of what AI tutoring currently is, even at its best: a resource that works for students who would have sought out help anyway. The students who walk to the back of the room. Those students, in my experience, were going to be okay. They have the metacognitive foundation to use almost any resource well, including the ones they’re already using without anyone prompting them.
The students who don’t walk to the back of the room are the ones who need someone to come find them.
That gap between a reactive tool and a proactive human is not a problem AI has solved. Sal Khan now seems to agree.
Further reading:
The education of Sal Khan and the limits of his chatbot — Chalkbeat
Generative AI without guardrails can harm learning — PNAS, 2025
AI tutoring outperforms in-class active learning: an RCT — Scientific Reports / Harvard, 2025
What the research shows about generative AI in tutoring — Brookings, 2026
OECD Digital Education Outlook 2026
This piece first appeared on our Substack. Read it there or subscribe to get new essays on testing, admissions, and learning in your inbox.