AI Tutors Should Teach, Not Just Answer

12 min read

Summary

Recent field experiments and meta-analyses point to a practical design principle: make learners think, practise and explain instead of simply making answers easier to obtain.

AI Tutors Should Teach, Not Just Answer

The difference between help and learning

An answer can solve the problem in front of a learner. A tutor has a harder job: help the learner become more capable when the tutor is no longer there.

That distinction matters now that a plausible explanation is only a prompt away. Generative AI can rewrite a paragraph, solve an equation and produce a confident response in seconds. Those abilities can remove friction from learning. They can also remove the thinking that learning requires. A system that makes today's work easier may leave tomorrow's learner no better prepared.

Recent research does not justify either of the simplest stories about AI in education. It does not show that conversational AI inevitably damages learning. Nor does it show that a fluent chatbot is automatically an effective tutor. Instead, early experiments point toward a design question: what does the learner have to think, retrieve, explain and practise while the AI is present? 4 5 6

The most useful standard is therefore not how impressive an AI response appears. It is whether the interaction develops knowledge and strategies that the learner can use independently.

One experiment, two very different outcomes

A 2025 classroom-randomised field experiment in high-school mathematics makes the distinction unusually visible. Classrooms comprising nearly 1,000 students were assigned to three conditions: ordinary classroom resources, access to a general-purpose GPT interface, or access to a GPT-based tutor with safeguards intended to keep the interaction educational. While AI was available, practice grades rose by 48% in the general GPT condition and by 127% in the safeguarded tutor condition, relative to the control arm. But on an unaided examination immediately after practice in the same session, students who had used the general GPT interface scored 17% below the control group. The safeguarded tutor largely mitigated that drop: its estimated difference from control was only 0.004 points lower on the 0-to-1 scale, too small to distinguish reliably from no difference. 1

These figures should not be turned into a universal verdict. The study took place in one school, in one subject, with particular prompts, safeguards and assessments. Its unaided exam measured performance immediately after practice, not delayed retention. Still, its central contrast is powerful: performance with AI and performance without it are not the same outcome.

If a learner can submit the right answer while a tool is present, we know that the combined human-and-tool system succeeded. We do not yet know what changed in the learner. To answer that question, the assessment must eventually remove or reduce the support and ask for unaided retrieval, explanation or transfer.

Two practice groups are compared with control. With AI support, practice grades were 48% higher for GPT Base and 127% higher for safeguarded GPT Tutor. On the unaided exam immediately afterward, the GPT Base group scored 17% lower; the safeguarded group showed no clear difference from control, with an estimated difference of minus 0.004.
Assisted and unaided performance are different outcomes. The exam followed practice in the same session with AI removed. It measured immediate unaided performance, not delayed retention.

Scaffolding can change the result

Other trials show why the design of the interaction matters. In a randomised crossover study with 194 eligible students in an introductory Harvard physics course, a carefully designed AI tutor produced larger short-term pre-to-post learning gains than an active-learning class for two lessons. Median tutor use was 49 minutes, compared with a 60-minute classroom learning period, although only the tutor time was individually tracked. The tutor used structured explanations, questions, hints and pacing principles rather than acting as an unrestricted answer generator. 2

That is promising evidence for one research-based tutor in one course. It is not evidence that AI generally outperforms teachers, that a two-lesson gain will persist for months, or that every learner would prefer the same mode. The comparison also sets up a false choice if it is read as “AI or teachers.” Teachers establish goals, interpret context, build trust, notice disengagement and make ethical judgements that a controlled lesson comparison cannot capture.

A World Bank randomised field trial offers a different context. In a six-week, teacher-facilitated after-school programme in Edo State, Nigeria, first-year senior-secondary students used generative AI as part of English and digital-skills instruction. The programme produced a 0.31 standard-deviation gain on a combined outcome and a 0.23 standard-deviation gain in English. 3 The treatment was not an autonomous chatbot dropped into a classroom. Teachers supported the work, the programme had defined content, and the paper was published as a working paper. Those details are part of the result, not footnotes to be discarded.

Across newer systematic reviews, conventional pooled effects of generative-AI-supported learning are often positive, but the literature is young and heterogeneous. Studies vary in duration, quality, discipline, comparison group, assessment and whether the AI is still available when outcomes are measured. Short interventions and immediate post-tests are common. In a 2026 STEM review, a robust Bayesian analysis suggested that the conventional positive average could largely reflect publication bias, while the prediction interval ranged from harm to benefit. 4 5 6

The responsible conclusion is neither celebration nor panic. It is that tutoring outcomes are design-dependent, and that independent performance must be measured deliberately.

What an AI tutor should actually do

Long before large language models, research on intelligent tutoring systems found positive average effects across many subjects and comparison conditions. Those systems were usually narrower than today’s chatbots. They modelled a domain, tracked steps, selected problems and delivered feedback within a constrained environment. Meta-analyses suggest that such purpose-built systems can improve learning, while also showing that effects depend on the comparison and the quality of implementation. 7 8 9

The lesson is not that every new tutor should imitate older software. It is that instructional structure is a feature, not an obstacle between the learner and an answer.

A learning-preserving tutor can use a four-part loop:

  1. Diagnose. Ask for an attempt before intervening. A wrong answer is not merely an error to replace; it is evidence about a misconception, missing prerequisite or unproductive strategy.
  2. Cue. Offer the smallest hint likely to restart productive thinking. A question, partial example or reminder can be more useful than a complete solution.
  3. Practise. Require the learner to do something observable: retrieve a principle, explain a step, compare alternatives, draw a representation or solve a related problem.
  4. Fade. Reduce support as competence grows, then check whether performance survives without the tutor.
A four-step tutoring cycle: diagnose the learner’s attempt, cue with the smallest useful hint, require practice, then fade support before diagnosing again.
The learning-preserving loop. A tutor should progressively return the work to the learner.

This loop changes the meaning of “personalisation.” Personalisation is not a stream of agreeable prose addressed to an individual. It is the disciplined adjustment of task, hint, explanation and timing in response to evidence of learning. A tutor should know when to explain, when to ask, when to wait and when to step away.

Product choices that protect thinking

Several concrete choices follow from that principle.

Require an attempt before revealing a worked answer. The attempt can be short, but it should expose the learner’s current model. “I don’t know yet” is useful evidence too. The tutor can then lower the entry point without pretending that a completed solution represents understanding.

Separate hints from answers. Progressive disclosure allows a learner to request the next layer of support. The first response might restate the goal; the second might identify a relevant principle; the third might model one step. The complete worked solution remains available when instruction or accessibility requires it, but it is not the default first move.

Ask for explanation and transfer. A correct answer can be guessed, copied or pattern-matched. Asking “Why?”, requesting an alternative method, or changing a surface feature reveals more. Transfer questions are especially important because a tutor can otherwise overfit the learner to one conversation.

Make uncertainty visible. Language models can be fluent when wrong. Grounding, source links, constrained tools and answer checking can reduce error, but no interface should convert uncertainty into theatrical confidence. Learners need a way to inspect evidence, report a problem and understand when human review matters. Scholarly and international guidance repeatedly identifies accuracy, bias, privacy and overreliance as central risks. 11 12

Preserve productive struggle without romanticising frustration. Difficulty is useful only when it is connected to the target knowledge and remains recoverable. Repeated failure, inaccessible language or missing prerequisites are not desirable difficulties. The tutor should adjust the grain size of a task and provide a route back into successful effort.

Design for teachers and learners together. A teacher-facing view can surface common misconceptions, stalled learners and the evidence behind recommendations. It should not automate consequential judgements from opaque signals. The stronger model is a division of labour: the system handles repeatable practice and immediate low-stakes feedback; educators retain contextual, relational and curricular judgement.

A five-question product review: Thinking, Accuracy, Adaptation, Fading, Transfer.
A five-question product review. The outcome to optimise is independent capability, not conversation volume.

Measure the learner, not the conversation

AI products naturally produce attractive operational metrics: messages sent, minutes engaged, questions answered and satisfaction ratings. These can reveal usability. They do not establish learning.

A serious evaluation should ask at least four different questions:

  • Can learners perform immediately? This checks near-term acquisition.
  • Can they perform after a delay? This tests whether access to the knowledge lasts.
  • Can they solve a meaningfully different problem? This tests transfer rather than imitation.
  • Can they perform when the AI is absent or deliberately restricted? This separates human learning from combined human-tool performance.

Comparison conditions matter too. Beating no instruction, an unstructured chatbot, a textbook chapter and an expert teacher are different claims. Time-on-task, teacher support, prior knowledge and access conditions can all change the interpretation. Bloom’s famous “2 sigma” challenge made individual tutoring an aspirational benchmark, but it was a research agenda, not a licence to label any one-to-one interface equivalent to expert tutoring. 10

Evidence from products already used in classrooms can complement controlled studies by showing what implementation looks like in practice. Khan Academy's research hub, for example, brings together studies connected to its platform and offers useful context about classroom use. But the collection itself is not proof that AI tutoring works. The important questions remain the same: Who conducted each study? What was the comparison group? What did learners retain without the tool? 13

What the evidence does not show

The current literature does not show that AI tutors generally outperform teachers. It does not show that gains from one lesson, course or country will transfer unchanged to another. It does not establish long-term safety, equity or retention for rapidly changing models. It rarely captures the full cost of implementation, teacher preparation, data governance or unequal access.

Meta-analytic averages do not solve these problems. A pooled effect combines interventions that may share a technology label while differing in pedagogy. Publication bias, short follow-up and inconsistent outcome measures remain concerns. Even a well-run randomised trial estimates the effect of a particular programme under particular conditions; it does not certify an entire product category. 4 5 6

There is also a deeper boundary. Education includes identity, care, culture, judgement and participation in a community, not only the efficient transfer of answers. An AI tutor may support parts of that work, but a conversational imitation of attention is not the same as a human relationship.

The standard worth building toward

The most credible promise for AI tutoring is modest and demanding: use computation to make high-quality practice, feedback and explanation more available while protecting the learner’s agency and independent capability.

That promise changes a product review. Instead of asking whether the model can answer, ask whether the system requires thinking. Instead of optimising only for session completion, examine delayed and unaided performance. Instead of replacing uncertainty with confidence, expose sources and limits. Instead of removing educators from the loop, give them better evidence for the decisions only they can make.

An answer engine ends when the response appears. A tutor succeeds later, when the learner can continue alone.

That delayed independence is the outcome against which every impressive interaction should ultimately be judged.

References

  1. Bastani, H., et al. “Generative AI without guardrails can harm learning: Evidence from high school mathematics.” Proceedings of the National Academy of Sciences, 122(26), e2422633122. Source.
  2. Kestin, G., et al. “AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting.” Scientific Reports, 15, 17458. Source.
  3. De Simone, M. E., et al. “From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria.” World Bank Policy Research Working Paper 11125. Source.
  4. Deng, R., Jiang, M., Yu, X., Lu, Y., & Liu, S. “Does ChatGPT enhance student learning? A systematic review and meta-analysis of experimental studies.” Computers & Education, 227, 105224. Source.
  5. Han, X., Peng, H., & Liu, M. “The impact of GenAI on learning outcomes: A systematic review and meta-analysis of experimental studies.” Educational Research Review, 100714. Source.
  6. Boolzen, C., et al. “Evidence of impact and interpretational limits of generative AI in STEM education: a systematic review and meta-analysis on cognitive learning outcomes.” Artificial Intelligence Review. Source.
  7. Ma, W., Adesope, O. O., Nesbit, J. C., & Liu, Q. “Intelligent tutoring systems and learning outcomes: A meta-analysis.” Journal of Educational Psychology, 106(4), 901–918. Source.
  8. Kulik, J. A., & Fletcher, J. D. “Effectiveness of intelligent tutoring systems: A meta-analytic review.” Review of Educational Research, 86(1), 42–78. Source.
  9. VanLehn, K. “The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems.” Educational Psychologist, 46(4), 197–221. Source.
  10. Bloom, B. S. “The 2 Sigma Problem: The Search for Methods of Group Instruction as Effective as One-to-One Tutoring.” Educational Researcher, 13(6), 4–16. Source.
  11. Kasneci, E., et al. “ChatGPT for good? On opportunities and challenges of large language models for education.” Learning and Individual Differences, 103, 102274. Source.
  12. UNESCO. Guidance for Generative AI in Education and Research. Source.
  13. Khan Academy. Research: studies and evidence published by Khan Academy and research partners. Source.

Share this post

Spread the word

Comments

Ulearngo
Ulearngo provides study and exam preparation tools that help students learn effectively and prepare confidently for upcoming examinations.

Ulearngo is independent and is not affiliated with or endorsed by any examination board, government agency, university, or admissions body.

Products

Resources

Company

Legal

Copyright © 2026 Ulearngo. All rights reserved.