Not So Fast on AI Tutors

There is a lot of money going into AI in education at the moment. Despite a series of research reports questioning the effectiveness of AI-driven systems, and the insistence of international organisations such as UNESCO - and increasingly of national governments - on keeping a human in the loop, the big software companies are still pursuing their dream of AI teachers. OpenAI, Anthropic, Google and others continue to present education as the next great opportunity for large language models.
This may be partly because AI adoption in industry has been less straightforward than the early hype suggested. Education, by contrast, increasingly looks like an attractive market. The new pedagogical proposition, if it can be called that, is the AI tutor. Private tutoring, it is often said, has contributed to the advantages enjoyed by privately educated children. If that tutoring experience can be delivered cheaply and at scale, then perhaps everyone can receive a first-class education.
It is an appealing story. It is also a story that should make us pause.
The question is not whether an AI system can produce a plausible explanation of algebra, history or grammar. It clearly can. Nor is it whether a chatbot can answer a question faster than a busy teacher can. It often can. The more important question is whether it can teach: notice what a learner does not understand, resist the temptation to supply the answer, choose an appropriate next step, and support the learner to do the thinking for themselves.
A new piece of research makes that distinction unusually clear. Can AI Tutors Actually Teach? We Measured It, published by Comprendo, tested six frontier language models in 779 simulated tutoring conversations with an eighth-grade mathematics student. The student was given a deliberately planted misconception, drawn from the mathematics-education literature, so that the researchers could assess whether the tutor recognised the error, diagnosed it correctly, avoided factual mistakes and, crucially, withheld the answer where appropriate 2.
The results should give the advocates of AI tutors some pause. When the models were used with a simple, student-like prompt rather than a specific tutoring instruction, they handed over the answer in 97% of the conversations. Across 298 such interactions, none produced what the researchers counted as a clean tutoring outcome: a correct diagnosis, no false statements and genuine restraint 2. In other words, without careful design, the systems behaved like answer machines rather than tutors.
Adding a tutoring prompt improved matters, but unevenly. The models became much less likely to reveal the answer directly. Yet no model was consistently strong at both sides of the tutoring task: diagnosing the learner’s misconception and leaving the essential reasoning to the learner. Models that were more reliable at identifying the mistake often did so by completing too much of the thinking themselves; those that showed greater restraint missed more misconceptions 2. That seems to me a central problem, rather than a technical detail. The productive struggle of working something out is not an unfortunate delay before learning begins. It is often the learning.
There are reasons to treat the findings with care. This was a simulation rather than a study of real pupils learning over time; it tested foundation models rather than finished commercial tutoring products; and the study was produced by an organisation that develops evaluation services for AI in education. Its strict measure of “withholding answers” was also adapted from a rubric developed for writing feedback 2. Nevertheless, these are limitations that should lead to more independent evaluation, not be used to dismiss the results. The basic point is hard to escape: being factually correct is not the same as teaching well.
The study also points to an uncomfortable weakness in the educational technology market. Its authors argue that systems for data privacy and security have become increasingly formalised, whereas evidence of instructional quality is still often little more than marketing language 2. If AI tutoring is to be funded, procured and used in schools, then it should be evaluated not only for safety and accuracy but for whether it helps learners think, reason and learn. This evaluation cannot be a single benchmark run at launch. Models, prompts and products change too quickly. It needs to be ongoing, transparent and independent.
This does not mean that all uses of generative AI in education should be rejected. Indeed, one of the more useful articles I have read recently takes a very different approach. In Claude Is Free for Teachers. Here Are Fifteen Things I Use It For, Stefan Bauschard describes using Claude not as a substitute teacher or an automated tutor for learners, but as a practical assistant for the work surrounding teaching 1.
His examples are refreshingly unglamorous. He asks the system to check whether a handout is likely to be comprehensible to a particular age group and to identify vocabulary he has stopped noticing as difficult. He updates old examples in debate materials when the topic changes. He uses it as a sounding board when a lesson has gone wrong or a learner has disengaged—not because it knows the student or the classroom, but because the exchange can help him articulate what he has observed and what he needs to investigate next 1.
Most interestingly, Bauschard records his own teaching, where policy permits, and uses transcripts to update his materials with explanations and examples that emerged in the classroom. The system does not replace the teacher’s judgement. It helps capture and organise the teacher’s own developing expertise. He also uses it to reduce the administrative burden of reports, evaluations and the production of teaching materials, while checking and revising the output himself 1.
The distinction matters: an AI system can support a teacher’s work without being asked to replace the intellectual and relational work of teaching.
This is a much more modest - and perhaps much more convincing - case for AI in education. It does not depend on the belief that a language model can form a pedagogical relationship with a child, understand a learner’s motivations or know when silence, uncertainty or an unexpected detour is educationally valuable. Instead, it treats the system as a tool that can save teachers time on preparation, documentation and iteration, provided that the teacher retains responsibility for context, judgement, verification and care.
The appeal of universal AI tutoring rests on a real injustice: access to high-quality human support is deeply unequal. But the solution cannot simply be to replace a scarce human relationship with a system optimised to produce the next plausible response. Before the promises of AI tutors are allowed to drive investment and procurement, we should ask a more demanding question: does the system enable learners to think for themselves, or does it do the thinking for them?
At the moment, the evidence suggests that this question needs much better answers. Meanwhile, there may be a useful, less spectacular role for AI: not as the teacher, and not as the tutor, but as a competent assistant that gives teachers more time to do the human work only they can do.
References
Bauschard, S. (2026, July 28). Claude is free for teachers. Here are fifteen things I use it for. Education Disrupted. https://stefanbauschard.substack.com/p/claude-is-free-for-teachers-here
Comprendo. (2026, August 10 ). Can AI tutors actually teach? We measured it. https://comprendo.dev/insights/can-ai-tutors-teach
[1] https://stefanbauschard.substack.com/p/claude-is-free-for-teachers-here
[2] https://comprendo.dev/insights/can-ai-tutors-teach
About the Image
The AI race describes the companies and nations which are racing forward to the newest iteration of 'AI', including software, hardware capacity, power, leadership, and influence in the field. It's a race that nation leaders and private corporations have signed up for, but workforces, neighbourhoods, communities, and working class people are impacted by the harms caused. I strive to show the implications of technology, in this case, the runner represents the workers running the race for entire nations, and its silhouette contains the materials, hardware, and resources (often hidden) needed to build generative AI technologies. Skylines are representative of cities and progress, but the heavy smoke cruising along the image representing the pollution, e-waste, and debris the race leaves behind. The panels of colour represent the countries typically referred to in the context of the "AI race": the US and China. But I leave this open to interpretation and did not want the flags and countries to distract from the most important aspect of the illustration: the consequences of the AI race.
