IMPORTANT: yes
The Two-Sigma Problem Gets an AI Answer
In 1984, educational psychologist Benjamin Bloom posed what became known as the "two-sigma problem": one-on-one tutoring produces learning gains roughly two standard deviations above conventional classroom instruction, yet most families cannot afford it. Nearly four decades later, a new study suggests that large language models may finally make Bloom's vision broadly affordable, without sacrificing effectiveness.
Curtis Northcutt and colleagues at Handshake AI introduce StudentBench, an evaluation platform and research project that put 13 different AI tutoring systems head-to-head against expert human tutors on the GRE. Across 2,383 participants and 2,469 tutoring sessions, the researchers found that AI tutoring produced learning gains statistically equivalent to those of human experts (p = .015). In five of the seven GRE domains tested, the best-performing AI tutor actually surpassed the mean human tutor. The paper was submitted on September 23, 2026, and the full dataset, code, and platform are publicly available.
What StudentBench Actually Measures
The study design is unusually disciplined. Students took a 27-question pre-test modeled on actual GRE formats, created by former ETS exam writers and Kaplan tutors with at least five years of experience. No previous GRE questions were reused, because LLMs may have seen them in training data. After the pre-test, students were randomly assigned to one of three conditions: AI tutoring, human tutoring via live video call, or a no-tutoring control where they watched educational videos unrelated to the GRE. All students then took a post-test with different questions covering the same concepts and difficulty levels.
Crucially, the AI tutors operated with minimal software scaffolding. Each was defined by just two prompts: one for lesson planning and practice problem generation, and one for live interactive tutoring. The AI tutors decided what to teach, planned lessons, generated practice problems, and tutored students in real time, all without access to the post-test. This design choice matters because it isolates the LLM's teaching capability from the quality of the software wrapper around it.
The 13 AI tutors spanned a wide capability range, including frontier models like GPT-5.5 Pro and Gemini 3.1 Pro, lightweight models like Gemini 3.5 Flash, and open-weight models like Gemma 4 31B. New tutors entered the study as they became available during the two-month summer 2026 data collection window.
The Learning Gains: Where AI Matches and Beats Humans
After adjusting for pre-test scores, the AI-control difference in learning gain was 6.86 percentage points in Quantitative (95% CI [4.02, 9.69]) and 5.47 percentage points in Verbal (95% CI [2.46, 8.47]). Human tutoring also outperformed the control in both sections. The combined AI-versus-control advantage was 6.15 percentage points, roughly equivalent to 1.5 to 2 more correct answers out of 27.
The equivalence test used a standard two one-sided tests framework at α = 0.05, with equivalence bounds of ±0.25 pooled standard deviations. Pooled AI and human tutoring were statistically equivalent across the combined Quantitative and Verbal sessions (p = .015), and the result also held under the tighter ±0.20-SD bounds (p = .023). The combined AI-human difference was −0.58 percentage points (90% CI [−2.18, 1.03]), well within the equivalence margin.
But the pooled result masks important variation across domains. The seven GRE domains tested were data analysis, geometry, arithmetic, and algebra on the Quantitative side, and sentence equivalence, text completion, and reading comprehension on the Verbal side. In five of these seven domains, the highest mean learning gain came from an AI tutor rather than a human. Google AI tutors led in data analysis, geometry, and reading comprehension. GPT-5.5 Pro led in arithmetic. Kimi K2.6 led in algebra. Anthropic models led in sentence equivalence and text completion. Human tutoring retained the highest mean in only algebra and sentence equivalence.
One notable finding: among students in the top quartile of starting proficiency, Gemini tutors dominated. Four of the five highest mean learning gains for high-performers came from Gemini models, led by Gemini 3.5 Flash. This exploratory result hints that different model families may have differentiated strengths depending on the student's starting level.
How the AI Tutors Were Evaluated Beyond Learning Gains
Beyond the learning-gain study, the researchers conducted a second study with 51 expert human tutors who completed 2,028 pairwise reviews of AI-generated lesson plans and practice problems. Each comparison paired two lesson plans generated by different AI tutors from the same student pre-test. Reviewers, recruited from ETS or Kaplan with at least five years of tutoring experience, rated each plan on eight criteria: concept relevance, concept grouping, concept prioritization, time allocation, practice alignment, appropriate difficulty, example accuracy, and test-taking strategies.
The expert evaluations produced clear leaderboards. Anthropic models, particularly Opus variants, were strongly preferred for lesson planning and practice-problem creation. The leaderboards distinguished models convincingly, with non-overlapping confidence intervals for most model pairs. A striking finding: AI tutors rated highly on practice-problem design by experts also tended to receive fewer student disputes. Students could flag answers they disagreed with, and GPT-5.4 mini had the highest dispute rate at approximately 11% of answered Quantitative problems. The correlation between expert preference and fewer student disputes was ρ = 0.71 in Quantitative (p = .035) and ρ = 0.77 in Combined (p = .033).
A third evaluation examined conversational pedagogy. The researchers identified six fixed teaching behaviors drawn from established learning science literature: scaffolding cues, requests for explanation, interactive questions, early attempt requests, long solutions after student replies, and reasoning checks. Using deterministic text-analysis rules on 1,971 AI and 135 human tutoring transcripts, they found that AI tutors clustered by model family with no overlap between families, suggesting pedagogical characteristics may reflect company-wide training practices.
The Cost Equation: 918 Times Cheaper
Perhaps the most consequential finding concerns cost. StudentBench measured not just learning outcomes but the cost to increase a student's GRE learning gain by one percentage point, calculated as mean session cost divided by mean learning gain.
Gemma 4 31B achieved learning gains statistically equivalent to human tutoring (p = .044) at approximately 918 times lower cost per percentage point. Its mean inference cost for the entire session, including lesson planning, practice generation, and one hour of interactive tutoring, was just $0.067. In contrast, human tutoring used a $75-per-hour reference rate, translating to $4.81 per percentage point of learning gain. GPT-5.5 Pro, by comparison, cost $21.24 per session.
The cost spectrum across the 12 shared AI tutors ranged from $0.067 to $21.24 per session. Gemini 3.5 Flash delivered a similar learning gain to GPT-5.5 Pro while costing 20 times less. Six AI tutors passed individual equivalence tests against human tutoring at p < .05, and all had a mean session cost under $5.
Latency, Engagement, and the Chain of Causality
The study also uncovered a causal chain linking AI speed to student learning. For Quantitative sessions, faster AI replies were strongly associated with more student messages (Spearman ρ = −0.81, p = .0056 across all 12 AI tutors). More student messages predicted more correct practice problems, and more correct practice predicted larger learning gains. All three associations were significant at p < .002 in Quantitative, with the chain holding across the combined analysis.
AI reply times ranged from 1.9 seconds for GPT-5.4 mini to 31.0 seconds for GPT-5.5 Pro. The paper notes that the association between latency and engagement was less pronounced in Verbal, where the reply-time-to-messages correlation was not significant (p = .86). This suggests that the speed-engagement-learning chain may operate differently across subject matter, a nuance worth watching.
Limitations and Open Questions
The equivalence result was established for the combined Quantitative and Verbal sections but not for Verbal alone at the ±0.25-SD margin. Human tutoring retained the highest whole-section mean in Verbal, even though several AI tutors exceeded human means in individual Verbal domains. The study used minimal prompts by design, which the authors acknowledge may understate what better-engineered prompts could achieve. A separate pilot explored prompt design: minimal prompts produced a pooled mean gain of 14.27 percentage points compared with 8.85 for expanded prompts, though these comparisons were exploratory and drawn from different time periods.
The conversational pedagogy evaluation used text-analysis rules on typed chat transcripts; human sessions were spoken conversations over video calls, and the same rules may behave differently in spoken versus typed modalities. The student dispute rates record disagreement with AI-generated answers, not independently verified errors. And the study's scope is limited to GRE preparation; it does not test whether these findings generalize to other subjects, younger students, or longer learning horizons.
The authors also note that the AI-human equivalence observed here may represent a lower bound, as models continue to improve and inference costs fall. Between May and September 2026 alone, Gemini Flash's Terminal-Bench score rose from 76% to 89%, and Gemini 3.6 Flash's token prices dropped 50%.
What This Means in Practice
The StudentBench platform is freely available at studentbench.org, with the complete dataset on Hugging Face and the code on GitHub. Any student can use the platform, and the authors have made data, code, and all study materials openly available to support reproducibility and future research.
The practical implications are significant. A student can now receive tutoring that matches the learning gains of an expert human tutor at a fraction of a cent per percentage point of improvement. This does not mean human tutors are obsolete, but it does mean that the "two-sigma problem" Bloom identified may finally have a scalable answer. For a working developer or educator, the availability of the platform and data means the field can move beyond asking whether AI can tutor and start asking which AI tutors work best for which students, in which subjects, and under what conditions.
The framework the authors propose, recursive human self-improvement (RHSI), adds a broader theoretical lens: if AI tutors help humans learn more effectively, and humans in turn can build better AI tutors, a positive feedback loop emerges. StudentBench establishes the first necessary condition for that loop, using the most rigorous available evidence: real students, real exams, real human tutors, and a randomized controlled design.
The next question is not whether AI can teach, but how the education ecosystem adapts to make this capability accessible, trustworthy, and continuously improving.