After one hour with a tutor, what can a student solve without the tutor beside them? A preprint posted September 23 puts that question to 2,383 adults working through GRE-style math and verbal problems. In this particular test, AI tutoring produced a measurable immediate gain over a no-tutoring comparison and, when results were pooled, a gain statistically equivalent to expert human tutoring. That is a result about one hour and an immediate exam—not proof that an AI can replace a teacher.

The team behind StudentBench, funded by Handshake AI, made the student rather than the chatbot the unit of interest. Each participant took a pre-test, spent an hour in a randomly assigned condition, then took a post-test with different questions covering the same concepts. Students received text-based AI tutoring, a live video call with a human expert, or an educational video unrelated to the GRE. The questions were written for the study by former exam creators, reducing the chance that a model had simply seen the post-test answers in training.

The number that describes learning

The analysis covered 2,469 math or verbal sessions from 2,383 people. Across both sections, AI tutoring improved the post-minus-pre score by an adjusted 6.15 percentage points more than the video comparison—roughly one-and-a-half to two additional correct answers on a 27-question test. The human group also did better than that comparison. The AI-versus-human adjusted difference was minus 0.58 points, with a 90% confidence interval from minus 2.18 to plus 1.03 points. The researchers had specified an equivalence margin of plus or minus 4.09 points; the whole interval fell inside it.

In plain English
Equivalence is not the same as a finding that two tutors get exactly identical results. Researchers choose a range of differences small enough to count as practically similar, then ask whether the uncertainty around the measured difference stays inside that range. Here it did for the combined GRE result. It did not establish equivalence for the verbal section on its own.

The study found equivalence for quantitative reasoning separately, but not verbal reasoning separately. It tested a pooled set of AI tutors, not the claim that any chatbot available to any student will teach just as well. The human arm was also much smaller than the AI arm: 140 human-tutored sessions versus 2,139 AI-tutored sessions, with 190 video-control sessions. Those unequal groups matter when interpreting how precise the comparison can be.

A tutor that asks, rather than answers

The AI setup did more than accept questions. It saw each student's pre-test mistakes, planned a lesson, generated practice problems and then tutored through conversation. The human tutor received those mistakes too. This gave both sides a way to target a weak concept rather than deliver the same lecture to everyone.

Imagine a student who can calculate an average but stumbles when a word problem asks which data belong in it. A tutor could ask the student to explain their choice, offer a new practice question and correct the mistaken step. A system that merely supplies the final average may help finish the current problem while leaving the misconception intact. StudentBench checked unaided post-test performance, which is a stronger measure of immediate learning than counting correct answers while the chatbot is open.

The paper also compares lesson plans, practice-problem design and conversational behaviors. Expert tutors rated generated materials, and the researchers examined tutoring transcripts. Those measures help explain differences among models but are not themselves measurements of learning. Faster AI replies were associated with more dialogue and more correct practice in the math sessions. Association does not demonstrate that speeding up replies alone would cause students to learn more.

Cheap to run is not free to provide

One tested model's mean inference bill—the priced API computation for lesson planning, practice generation and an hour of chat—was about $0.067 per session. The researchers calculate a 918-fold cost-per-point advantage for that model against a human tutor priced at $75 an hour. It is a deliberately narrow comparison: it does not include software development, access to devices and connectivity, quality assurance, support, safeguarding or a school’s integration costs. Nor does the paper test whether such services reach people currently excluded from one-to-one tutoring.

There are three important unanswered learning questions. First, the test was immediate: the authors did not measure retention months later. Second, there was no group that spent the hour solving GRE practice problems alone, so the extra contribution of the conversational tutor over self-practice remains unmeasured. Third, participants were paid adults who could read and write English and had access to the Handshake platform. Results cannot simply be transferred to children, classrooms, other languages or subjects.

These limits do not erase the finding. They make it useful: a comparatively large randomized comparison measured what students could do after a bounded tutoring session, not how persuasive a generated answer sounded. The next decisive test would keep the unaided assessment, add an active solo-practice group and return weeks later. That would tell educators whether the immediate gain survives—and whether the conversation, rather than simply the extra hour of practice, earned its place.

Keep exploring

Source-based analysis; no independent tutoring test. Handshake AI funded the cited study. AI-assisted. Sources checked.