In July 2026, the Gates Foundation published a refreshed education strategy built partly around artificial intelligence: tools to help teachers individualize instruction, and AI assistants for college advisers buried in paperwork. Reporting on it for The 74, Kevin Mahnken quoted Carly Robinson, a Stanford researcher who studies student support. Her formulation was careful and, we think, correct.
"A huge share of their time goes to red tape rather than actually working with kids," she wrote. "If AI can absorb that logistical load, you're not replacing the counselor. You're directing them to be able to do more of the important work with students."
This is the sensible position, and the numbers behind it are real. American school counselors carry an average caseload of 372 students against a professional recommendation of 250. Teachers who use AI weekly report saving around six hours a week. If software can take the forms and the scheduling and hand back the hours, that is a straightforward good.
But notice the phrase carrying the weight of the whole argument. The important work. It is never defined, and its vagueness is not harmless. Suppose you cannot say precisely what the important work is. Then the sentence justifies whatever gets automated next, one task at a time, with the irreplaceable part always assumed to be whatever happens to be left over. The useful question is not whether to use these tools. It is what humans must keep doing while they do.
That question has a better answer than most of the debate suggests, and it does not come from opinion. It comes from the research on why tutoring works at all.
A student gets a problem wrong. A good answer engine will tell him what the right answer is and walk him through the correct steps, patiently, at any hour, without irritation. On that narrow task it is excellent, and often better than a tired human at nine in the evening.
What it does not do is ask why he got it wrong. There is a difference between a student who misread the question, a student with a specific broken idea about negative numbers, and a student who understood but has not slept and is quietly falling apart. Those three need three different responses. Only one of them is a mathematics problem.
Teaching one student at a time, that distinction is most of the job. You notice the pause before an answer. You notice that a confident student has gone quiet for two weeks. You notice the wrong answer that is interesting, because it reveals exactly which piece of the foundation is missing. Then you decide, deliberately, not to give the answer, because this is the student who needs to struggle for ninety more seconds.
None of that is logistics. All of it is the work.
Figures from the studies and bodies named. Together they describe a technology that helps, and helps most when a person is steering it.
In 2025 a team led by researchers at Wharton and their colleagues published a study in the Proceedings of the National Academy of Sciences. Nearly a thousand high-school students in Turkey were split into three groups for a mathematics unit. One practiced with a standard AI assistant. One practiced with a version deliberately restricted to hints and guiding questions, built so it would not simply hand over answers. One had no AI at all.
During practice, the AI groups pulled far ahead. Then the tool was taken away for the exam, which is the only measurement that mattered.
Read that again, because it is the whole argument in one result. Both groups used the same underlying technology. What differed was that a person had decided in advance that the tool would withhold the answer and make the student work. That decision, and not the model, was the difference between learning and the appearance of learning.
Withholding the answer is not a natural act. It runs against every instinct a helpful system has, and against what the student in front of you is asking for. It requires someone who has decided what this particular child needs and is willing to be briefly unhelpful in order to give it. That is the first piece of the important work, and we can now name it precisely.
Once you take the Turkish result seriously, the division becomes clearer than the public debate usually allows. Some of teaching is purely logistical, and handing it to software is a gift. Some of it looks like information transfer but is judgment about a specific person, and automating it removes the thing that made it work.
The left column should be automated aggressively. The argument of this piece is not caution about technology. It is that the left column is the easy half, and moving it to software only matters if the hours it returns are spent on the right column. Hours recovered and then spent on larger class sizes have not improved anything. They have just been recovered.
Tutoring is the most heavily studied form of teaching there is, which makes it the best available window into what drives learning. The headline finding is strong. Across 96 randomized trials, tutoring produces an average gain of about 0.29 standard deviations, one of the more reliable effects in education research.
The revealing part is what happens next. When researchers examined 282 trials and asked what to expect from tutoring delivered at large scale, the effect fell to between a third and a half of the headline figure, declining steadily as programs got bigger. Something in tutoring does not survive being scaled, and identifying it tells you where the value lives.
The field's own answer is not ambiguous. The standard guidance defines effective tutoring as intensive, relationship-based, individualized instruction, and lists relationships first among its features. A named design principle is that a student should meet the same tutor throughout, because consistency builds the relationship that carries the learning. Group size matters for the same reason. One-to-one tuition delivers around five months of additional progress and small groups around four. Transcript studies find the one-to-one advantage comes partly from time spent on personalization and relationship, not pure content coverage.
The last of those figures deserves attention, because it comes from Robinson herself. In two randomized trials, elementary students were given access to a well-regarded AI reading tutor. Left to themselves, they used it for two to five minutes. Her conclusion was that the challenge is not building good tools but getting students to engage with them, and that availability is not use. A separate study of hers found that messages to students alone did nothing to increase take-up of free tutoring, while messages to parents and students together raised it by 46 percent.
This is worth stating plainly. The researcher quoted arguing that AI should absorb the logistical load is herself a scholar of relationships in education, and her own findings identify the constraint. It is not tool quality. It is whether a person is in the loop making the tool matter to a particular child.
We have an interest here. Harland sells human teaching time, so an argument for the irreplaceability of human teachers is an argument for our own product. That is exactly why the concessions below need to be specific rather than decorative.
Unguarded AI can reduce learning. In the PNAS trial, students who practiced with an unrestricted assistant scored about 17 percent worse than the no-AI control once it was removed. A human-designed constraint prevented that.
The gains come from pairing, not replacement. Every rigorous success we found had a person in the loop: tutors using an AI assistant, human tutors supervising AI-drafted replies, or a teacher facilitating in the room.
Something relational does not scale. Tutoring's measured effect falls by half or more as programs grow, and the field's own design guidance puts relationships and tutor consistency at the center.
AI tutoring produces real learning. A World Bank trial in Nigeria measured gains of roughly 0.23 to 0.31 standard deviations in six weeks. Older intelligent tutoring systems already rivaled human tutors on well-defined tasks. This is not marketing.
The equity case is strong and we do not have an answer to it. AI support reaches students who would otherwise receive none, and the largest gains in the Stanford trial went to students of the weakest tutors. Something imperfect beats nothing.
The direct link is not proven. Relationship and consistency are well evidenced as drivers of engagement and attendance. That they raise test scores directly is a reasonable chain of inference, not a settled finding, and the field's own researchers say so.
Human teaching also deserves less mystique than it usually gets. The widely repeated claim that tutoring produces two standard deviations of improvement has not held up. The more careful figure is closer to 0.79, and well-built software has matched human tutors on narrow, well-defined tasks. The argument here is not that people are better at everything. It is that a short list of things currently has no machine substitute, and that list is where the value of a teacher now sits.
Practical questions for reading any school or provider's claims about AI, including ours.
An academy that sells human teaching hours writing that human teaching cannot be automated is not an unbiased source, and you should read this piece knowing that. We have tried to earn it by conceding the strongest points against us plainly: AI tutoring works, it reaches students who have nothing, and human tutoring is less miraculous than the folklore claims.
We also use these tools. They draft practice sets, they handle administration, they take work off our teachers that no one should be doing by hand. We are not arguing that the technology is a threat to be resisted. We are arguing from having watched what it does well and what it cannot yet do at all.
What one-on-one teaching gives us is an unusually direct view of the boundary. When you sit with the same student every week, you find out quickly which parts of the work a machine could take and which parts collapse without a person. The answer has been consistent. It can carry the load. It cannot carry the judgment about this child, and it cannot be the person he does not want to disappoint.
Harland teaches one student at a time, with a subject-specialist teacher matched to the student and kept consistent. Our pedagogy is content-based learning: understanding is built through real subject content, not drilled in isolation. The structure exists to protect exactly the work this piece describes.
If you are weighing what a tutor should give your child that software already gives them for free, that is the right question to be asking. Tell us where your child is, and we will be honest about which parts we can help with.
Start the conversation