At some point this year a parent will ask you about Alpha School. They will have read that it condenses academic instruction into a two-hour daily block of AI tutoring, that the rest of the day goes on workshops with coaches who are not necessarily qualified teachers, and that fees run from $40,000 to $75,000. They will want to know whether it works, and whether you are behind.
Gerald LeTendre, who researches AI and teachers at Penn State, has written the clearest short account I've seen of what the evidence actually supports. It is worth reading before that conversation. It is also worth reading for a reason he doesn't spell out: the studies that look most impressive for AI tutoring are, on closer inspection, studies about how capable the student was before they opened the app.
The claim
The pitch is familiar by now. Schooling is one-size-fits-all, students learn at different speeds, and a model can give every child a tutor that adapts to them. Sal Khan has made a version of this argument for years. Bill Gates has suggested AI will take over much of the teacher's job within a decade. Some of the marketing is blunter, along the lines of "School is broken, and we're here to fix it."
The trouble with the framing is that it sets a responsive, motivating AI against a caricature of classroom teaching as a lecture nobody listens to. Set the comparison up that way and you have decided the result before you collect any data.
What the evidence supports
Human tutoring works, and the research on it is old and strong. A 2020 review of many tutoring programmes found consistent gains across subjects and age ranges. The versions that work best are regular, long-term, in small groups, in school, delivered by someone skilled, and tied to what the student is doing in class. Relationship matters to whether the student turns up and stays engaged.
Computer tutoring is not new either. Intelligent tutoring systems have been running since the 1970s, and they raise attainment when used as a supplement to teaching. Compared directly against human tutors, the pooled evidence shows no meaningful difference. Generative AI has improved the experience, mainly by letting students interact in ordinary language rather than through a fixed menu of exercises, and a 2025 study found positive effects across subjects and year groups.
Then there is the Harvard physics study, which gets cited more than any of the others. Students using a purpose-built AI tutor reported learning faster and feeling more motivated than students in a well-designed active-learning class. Read the conditions. The students were already highly motivated and had strong study skills. The tutor had been built by the professors who taught the course, so its explanations matched the curriculum exactly.
That is not a finding about AI tutors in general. It is a finding about what happens when a capable, self-regulating learner meets a tool that has been aligned to their course by the person assessing them.
Read the caveats as a curriculum
Here is where this becomes an AI literacy question rather than a procurement one.
Strip out the conditions that made those studies work and ask which of them a school can actually change. You cannot ship every student a tutor built by their own subject teacher, though you can get closer than most schools try. What you can change is what the student brings to the interaction.
A student gets value from a model as a tutor when they can do a specific set of things. Notice the gap in their own understanding precisely enough to ask about it. Recognise a confident explanation that is wrong. Ask for the working rather than the answer, then check the step they are able to check. Resist the version of the task that produces something submittable in ninety seconds. Know the point at which the model has stopped helping and a person is needed.
None of that is innate, and none of it is generic. It is the difference between a tutoring session that builds understanding and one that produces a neat set of notes about nothing.
The uncomfortable part
Now run it the other way. The students who most need tutoring are the least equipped to catch a plausible wrong explanation, because catching it requires the knowledge they came to the session to acquire. A Year 9 who half understands moles in chemistry cannot audit a fluent paragraph about moles. The tool sounds equally certain either way.
So the personalisation claim quietly inverts. The tool amplifies the learner's existing capacity to interrogate it. Give it to a motivated sixth former with good study habits and it is a strong supplement. Give it to a student who is behind, unsupervised, and it is a fast answer machine that leaves the gap exactly where it was while making the work look finished.
That is the equity problem in AI tutoring, and it is not solved by buying a better product.
What this means for the tools you already have
The findings that should shape practice are the quieter ones. When human tutors in low-income middle schools were given AI support, outcomes improved. When experienced teachers use AI for lesson planning, they revise the output critically to fit their curriculum, which is precisely the skill a less experienced teacher hasn't built yet. Both point the same way. The gain sits with the person who knows enough to push back.
Three things worth doing this term, none of which require a purchase.
Run a wrong-answer hunt in a topic the class has already mastered. Ask the model to explain it, then mark the explanation as a class. Students who have caught a model being wrong once about something they know are permanently different users of it.
Ask students to record the question they asked, not just what came back. The quality of the question is the assessable part, and it tells you more about their understanding than the output does.
Give teachers the same session. A department that can spot where a generated lesson drifts from the specification is the department getting value from these tools. That capacity is built by practice, not by a twilight on prompt engineering.
Where the literacy lives
Notice that none of this can be taught in a standalone AI lesson. Spotting a bad explanation of osmosis requires knowing osmosis. Judging whether a worked solution is sound requires the maths. Deciding whether a summary of a source has flattened the argument requires the history.
AI literacy of this kind is subject literacy, made explicit and assessed. Which is why the question about Alpha School has a better answer than yes or no. The two-hour tutoring block is not the interesting variable. What the student can do while they sit in it is.
Source: Gerald K. LeTendre, "Despite the growth of AI schools, AI tutors aren't better than humans," The Conversation / UPI, 13 July 2026.
AILitKit surfaces the AI literacy already inside the lessons your teachers are planning. Three guides are free on the starter tier.