The Harland Review · The Modern Student

The Important Work

機器能分擔的,與不能取代的
For
Parents weighing what AI changes about teaching
Reading time
13 minutes
Last updated
August 2026
The Reframe

Everyone agrees AI should free teachers for the important work. Almost nobody says what that is.

In July 2026, the Gates Foundation published a refreshed education strategy built partly around artificial intelligence: tools to help teachers individualize instruction, and AI assistants for college advisers buried in paperwork. Reporting on it for The 74, Kevin Mahnken quoted Carly Robinson, a Stanford researcher who studies student support. Her formulation was careful and, we think, correct.

"A huge share of their time goes to red tape rather than actually working with kids," she wrote. "If AI can absorb that logistical load, you're not replacing the counselor. You're directing them to be able to do more of the important work with students."

This is the sensible position, and the numbers behind it are real. American school counselors carry an average caseload of 372 students against a professional recommendation of 250. Teachers who use AI weekly report saving around six hours a week. If software can take the forms and the scheduling and hand back the hours, that is a straightforward good.

But notice the phrase carrying the weight of the whole argument. The important work. It is never defined, and its vagueness is not harmless. Suppose you cannot say precisely what the important work is. Then the sentence justifies whatever gets automated next, one task at a time, with the irreplaceable part always assumed to be whatever happens to be left over. The useful question is not whether to use these tools. It is what humans must keep doing while they do.

That question has a better answer than most of the debate suggests, and it does not come from opinion. It comes from the research on why tutoring works at all.

What We See
The question a machine does not think to ask.

A student gets a problem wrong. A good answer engine will tell him what the right answer is and walk him through the correct steps, patiently, at any hour, without irritation. On that narrow task it is excellent, and often better than a tired human at nine in the evening.

What it does not do is ask why he got it wrong. There is a difference between a student who misread the question, a student with a specific broken idea about negative numbers, and a student who understood but has not slept and is quietly falling apart. Those three need three different responses. Only one of them is a mathematics problem.

Teaching one student at a time, that distinction is most of the job. You notice the pause before an answer. You notice that a confident student has gone quiet for two weeks. You notice the wrong answer that is interesting, because it reveals exactly which piece of the foundation is missing. Then you decide, deliberately, not to give the answer, because this is the student who needs to struggle for ninety more seconds.

None of that is logistics. All of it is the work.

At a Glance

What the research shows

Figures from the studies and bodies named. Together they describe a technology that helps, and helps most when a person is steering it.

Tutoring, overall effect
0.29 SD
Pooled effect across 96 randomized trials, per Nickow, Oreopoulos and Quan, 2024
What happens at scale
⅓ to ½
How much of that effect survives in large programs, across 282 trials, per Kraft, Schueler and Falken
Teacher hours recovered
5.9 hrs
Weekly time saved by teachers using AI regularly, per Gallup and the Walton Family Foundation, 2025
Counselor caseload
372 : 1
US average against a recommended 250 to 1, per the American School Counselor Association
AI helping a tutor
+9 pts
Topic mastery gain for students of the weakest tutors when tutors used an AI assistant, per Stanford's Tutor CoPilot trial
AI without guardrails
−17%
Exam performance against a no-AI control once the tool was removed, per Bastani and colleagues, PNAS 2025
Agreeableness
49%
How much more often leading models affirmed a user's actions than humans did, across 11 models, per a 2026 Stanford-led study
Left alone with a good tool
2–5 min
Time elementary students spent with an AI reading tutor when nobody was embedding it in their learning, per Robinson and colleagues
The Finding That Should Shape This Debate

The tool that helped most during practice hurt most on the test

In 2025 a team led by researchers at Wharton and their colleagues published a study in the Proceedings of the National Academy of Sciences. Nearly a thousand high-school students in Turkey were split into three groups for a mathematics unit. One practiced with a standard AI assistant. One practiced with a version deliberately restricted to hints and guiding questions, built so it would not simply hand over answers. One had no AI at all.

During practice, the AI groups pulled far ahead. Then the tool was taken away for the exam, which is the only measurement that mattered.

Generative AI and learning · Bastani et al., PNAS 2025 · ~1,000 students
During practice
Performance on practice problems, with the tool available
Standard AI assistant+48%
AI restricted to hints+127%
No AIbaseline
On the exam, tool removed
Performance against the students who never used AI
Standard AI assistant−17%
AI restricted to hintsharm largely removed
No AIcontrol
The students who learned least were the ones the machine helped most. What separated the two AI groups was a single design decision: whether the tool was allowed to give the answer.

Read that again, because it is the whole argument in one result. Both groups used the same underlying technology. What differed was that a person had decided in advance that the tool would withhold the answer and make the student work. That decision, and not the model, was the difference between learning and the appearance of learning.

Withholding the answer is not a natural act. It runs against every instinct a helpful system has, and against what the student in front of you is asking for. It requires someone who has decided what this particular child needs and is willing to be briefly unhelpful in order to give it. That is the first piece of the important work, and we can now name it precisely.

Drawing The Line

What can be absorbed, and what cannot

Once you take the Turkish result seriously, the division becomes clearer than the public debate usually allows. Some of teaching is purely logistical, and handing it to software is a gift. Some of it looks like information transfer but is judgment about a specific person, and automating it removes the thing that made it work.

Absorb This
Load a machine can carry
  • Scheduling, reminders, forms, and application tracking
  • Generating unlimited practice problems at a set difficulty
  • Answering a factual question at eleven at night
  • Drafting a first version of a lesson plan or a rubric
  • Marking work where the answer is unambiguously right or wrong
  • Summarizing which students missed which question types
Roughly six hours a week of a teacher's time, on the available evidence. Real, and worth taking.
Cannot Be Absorbed
The important work
  • Deciding to withhold the answer because this student needs to struggle
  • Diagnosing why a student is stuck, not just that he is
  • Noticing disengagement, avoidance, or distress nobody reported
  • Holding a standard when a student would rather be told he is fine
  • Knowing this child well enough for the judgment to be right
  • Being someone whose expectations a student does not want to disappoint
Not one of these is information transfer. All of them require knowing a particular person over time.

The left column should be automated aggressively. The argument of this piece is not caution about technology. It is that the left column is the easy half, and moving it to software only matters if the hours it returns are spent on the right column. Hours recovered and then spent on larger class sizes have not improved anything. They have just been recovered.

Why We Can Be Specific

The research names the same thing

Tutoring is the most heavily studied form of teaching there is, which makes it the best available window into what drives learning. The headline finding is strong. Across 96 randomized trials, tutoring produces an average gain of about 0.29 standard deviations, one of the more reliable effects in education research.

The revealing part is what happens next. When researchers examined 282 trials and asked what to expect from tutoring delivered at large scale, the effect fell to between a third and a half of the headline figure, declining steadily as programs got bigger. Something in tutoring does not survive being scaled, and identifying it tells you where the value lives.

The field's own answer is not ambiguous. The standard guidance defines effective tutoring as intensive, relationship-based, individualized instruction, and lists relationships first among its features. A named design principle is that a student should meet the same tutor throughout, because consistency builds the relationship that carries the learning. Group size matters for the same reason. One-to-one tuition delivers around five months of additional progress and small groups around four. Transcript studies find the one-to-one advantage comes partly from time spent on personalization and relationship, not pure content coverage.

0.29 SD
Average effect of tutoring across 96 randomized trials
Nickow, Oreopoulos & Quan, 2024
⅓–½
How much of that effect remains once programs are scaled up, across 282 trials
Kraft, Schueler & Falken
2–5 min
Time students spent with a good AI tutor when no adult was embedding it in their learning
Robinson and colleagues, 2025

The last of those figures deserves attention, because it comes from Robinson herself. In two randomized trials, elementary students were given access to a well-regarded AI reading tutor. Left to themselves, they used it for two to five minutes. Her conclusion was that the challenge is not building good tools but getting students to engage with them, and that availability is not use. A separate study of hers found that messages to students alone did nothing to increase take-up of free tutoring, while messages to parents and students together raised it by 46 percent.

This is worth stating plainly. The researcher quoted arguing that AI should absorb the logistical load is herself a scholar of relationships in education, and her own findings identify the constraint. It is not tool quality. It is whether a person is in the loop making the tool matter to a particular child.

What The Evidence Will and Will Not Bear

The honest limits of this argument

We have an interest here. Harland sells human teaching time, so an argument for the irreplaceability of human teachers is an argument for our own product. That is exactly why the concessions below need to be specific rather than decorative.

What the evidence supports

Unguarded AI can reduce learning. In the PNAS trial, students who practiced with an unrestricted assistant scored about 17 percent worse than the no-AI control once it was removed. A human-designed constraint prevented that.

The gains come from pairing, not replacement. Every rigorous success we found had a person in the loop: tutors using an AI assistant, human tutors supervising AI-drafted replies, or a teacher facilitating in the room.

Something relational does not scale. Tutoring's measured effect falls by half or more as programs grow, and the field's own design guidance puts relationships and tutor consistency at the center.

Where it should stop

AI tutoring produces real learning. A World Bank trial in Nigeria measured gains of roughly 0.23 to 0.31 standard deviations in six weeks. Older intelligent tutoring systems already rivaled human tutors on well-defined tasks. This is not marketing.

The equity case is strong and we do not have an answer to it. AI support reaches students who would otherwise receive none, and the largest gains in the Stanford trial went to students of the weakest tutors. Something imperfect beats nothing.

The direct link is not proven. Relationship and consistency are well evidenced as drivers of engagement and attendance. That they raise test scores directly is a reasonable chain of inference, not a settled finding, and the field's own researchers say so.

Human teaching also deserves less mystique than it usually gets. The widely repeated claim that tutoring produces two standard deviations of improvement has not held up. The more careful figure is closer to 0.79, and well-built software has matched human tutors on narrow, well-defined tasks. The argument here is not that people are better at everything. It is that a short list of things currently has no machine substitute, and that list is where the value of a teacher now sits.

Common Situations

What to ask, as a parent

Practical questions for reading any school or provider's claims about AI, including ours.

A school announces it is adopting AI tutoring.
The question is not whether the tool is good. It is what the recovered time is being spent on. If AI absorbs marking and admin and those hours go back to students, that is the case working as intended. If class sizes rise or contact hours fall because software is covering the difference, the logic has been used to justify the opposite of what it promised.
Your child is using an AI assistant for homework.
The distinction that matters is whether it gives answers or withholds them. A tool that explains a completed solution produces the fluent, confident practice performance that collapsed on the exam in the Turkish study. A tool used to get unstuck, then closed, is a different activity. Ask your child to solve the next one without it.
The tool tells your child the work is excellent.
Treat praise from a chatbot as close to meaningless. A 2026 study across eleven leading models found they affirmed users' actions 49 percent more often than humans did, and that people who used agreeable AI became more convinced they were right. A system optimized to be liked cannot reliably tell your child the work is not good enough yet.
A provider says AI lets them serve more students.
This may be true and may be good, particularly where the alternative is no support at all. But ask the follow-up: does your child still have a consistent person who knows him, or a rotating cast plus better software? The evidence on what disappears when tutoring scales suggests that answer matters more than the technology in either direction.
Why We Can Write This Honestly
We are not a neutral party, and we use these tools.

An academy that sells human teaching hours writing that human teaching cannot be automated is not an unbiased source, and you should read this piece knowing that. We have tried to earn it by conceding the strongest points against us plainly: AI tutoring works, it reaches students who have nothing, and human tutoring is less miraculous than the folklore claims.

We also use these tools. They draft practice sets, they handle administration, they take work off our teachers that no one should be doing by hand. We are not arguing that the technology is a threat to be resisted. We are arguing from having watched what it does well and what it cannot yet do at all.

What one-on-one teaching gives us is an unusually direct view of the boundary. When you sit with the same student every week, you find out quickly which parts of the work a machine could take and which parts collapse without a person. The answer has been consistent. It can carry the load. It cannot carry the judgment about this child, and it cannot be the person he does not want to disappoint.

How Harland Helps

A teacher who knows your child, doing the part that cannot be automated.

Harland teaches one student at a time, with a subject-specialist teacher matched to the student and kept consistent. Our pedagogy is content-based learning: understanding is built through real subject content, not drilled in isolation. The structure exists to protect exactly the work this piece describes.

01
The same teacher, long enough to know the student.
The research on tutoring names a consistent tutor as a design principle, not a nicety, because judgment about a student depends on knowing that student. Your child works with a primary teacher matched by subject, and if the fit is wrong we change the teacher rather than asking the child to adapt.
02
We decide when not to give the answer.
The clearest finding in this piece is that withholding the answer is what separated learning from its appearance. That decision is made deliberately, student by student, by someone who knows how much productive struggle this particular child can carry today. It is the least automatable thing we do.
03
Someone notices what the data does not report.
A dashboard reports that a student missed questions. A person notices when a settled student turns wary, when avoidance creeps in, or that the problem this week is not academic at all. Every lesson generates a written record parents can see, but the noticing happens in the room.

Who is teaching your child?

If you are weighing what a tutor should give your child that software already gives them for free, that is the right question to be asking. Tell us where your child is, and we will be honest about which parts we can help with.

Start the conversation
PUBLISHED 19 August 2026 · LAST UPDATED 19 August 2026 · All figures verified against the study or body named at the date of publication.
Sources

Where these figures come from

Reporting on the Gates Foundation's refreshed education strategy and its AI priorities, and the source of the Carly Robinson quotation. The 74 discloses that it receives financial support from the Gates Foundation.
Bastani, Bastani, Sungu, Ge, Kabakcı and Mariman, Proceedings of the National Academy of Sciences (2025)
"Generative AI without guardrails can harm learning." Around 994 Turkish high-school students; practice gains of 48 and 127 percent, and a fall of roughly 17 percent against the control on the exam once the unrestricted tool was removed.
Nickow, Oreopoulos and Quan, American Educational Research Journal (2024)
Systematic review and meta-analysis of 96 randomized tutoring experiments; pooled effect of about 0.29 standard deviations.
Kraft, Schueler and Falken (2024, 2026)
"What Impacts Should We Expect From Tutoring at Scale?" Expanded meta-analysis of 282 randomized trials finding effects of roughly one-third to one-half the headline figure once aligned to large-scale programs.
Robinson and colleagues, Stanford SCALE Initiative
Randomized trials on AI tutor take-up (two to five minutes of use without adult embedding) and on nudges to increase tutoring uptake (a 46 percent increase when messages reached parents and students together).
Wang, Ribeiro, Robinson, Loeb and Demszky, Tutor CoPilot (2024)
Randomized trial of an AI assistant used by 900 tutors with 1,800 students; four percentage points higher topic mastery overall and nine points for students of the lowest-rated tutors.
Robinson, Kraft, Loeb and Schueler, EdResearch for Action Brief 30 (June 2024)
Design principles for high-impact tutoring, defining it as relationship-based and naming a consistent tutor as a design principle. The authors note that individual program features have not been directly compared.
One-to-one tuition associated with about five months of additional progress and small-group tuition about four months.
Gallup and the Walton Family Foundation (June 2025)
Survey of 2,232 US public-school teachers; weekly AI users reported saving an average of 5.9 hours per week.
American School Counselor Association
National average counselor caseload of 372 students in 2024–25 against a recommended ratio of 250 to 1.
De Simone and colleagues, World Bank Policy Research Working Paper 11125 (May 2025)
Randomized trial of GPT-4 tutoring in Benin City, Nigeria; measured gains of about 0.23 standard deviations in English and 0.31 overall across six weeks, with a teacher facilitating sessions.
Cheng and colleagues, Science (2026)
Study across eleven leading models finding they affirmed users' actions 49 percent more often than humans did, and that interacting with agreeable systems increased users' confidence that they were right.