"IMPORTANT: yes"

Large language models have advanced rapidly, achieving strong capabilities in mathematical reasoning and code generation. Solving a problem, however, is different from teaching someone else to solve it. Frontier models score above 56% against expert-written tutoring rubrics, while strong problem-solving ability does not directly translate into effective teaching. This gap matters for human learners: receiving effective task assistance does not necessarily translate into learning, and may even reduce cognitive engagement or subsequent independent performance. Developing models that help people learn to think and solve problems independently is therefore increasingly important, beyond improving individual dimensions of teaching skill.

Problem Formulation for Adaptive Teaching

One approach to AI-assisted learning is to build educational systems around existing LLMs, injecting human-designed pedagogical strategies through prompts or external scaffolding. Such systems can provide useful instructional structure, but they primarily shape the model's behavior at inference time. To develop teaching as an intrinsic capability of the model itself, a growing body of work instead trains LLMs specifically as teachers. Supervised fine-tuning and preference optimization came first, but they rely on the quality and coverage of instructional demonstrations or preference data. More recent work has turned to reinforcement learning so that teachers can learn from their own interactions with simulated students. Yet these approaches share a fundamental limitation: their training signals encode predefined notions of what good teaching looks like, rather than being grounded directly in diverse students' learning outcomes. In practice, students differ in their preferences and needs, and effective teaching cannot be fully specified without accounting for this heterogeneity.

To close this gap, we introduce Sherpa, a multi-turn reinforcement learning framework that trains a teacher to improve student task-solving across diverse student archetypes. Rather than predefining what constitutes good teaching, Sherpa separates pedagogical specification from teacher optimization: diverse instructional needs are instantiated on the student side, while the teacher is optimized based on student improvement. Each teaching episode follows a preparation–tutoring–testing cycle: given a problem, the teacher first prepares a solution, tutors the student over multiple turns of interaction, and is rewarded based on the student's improvement at test time. Student models exhibit similar preferences for instructional strategies even after prompting or across model families, making them insufficiently diverse as archetypes for our environment. We therefore design a gating mechanism that makes instructional strategies consequential: guidance that meets the student's preference passes the gate and is forwarded to the student backbone model to elicit substantive reasoning, whereas incompatible guidance is blocked and receives feedback requesting more appropriate instruction. The teacher must therefore infer each student's needs from interaction and adapt its teaching accordingly. In this way, Sherpa allows diverse teaching behaviors to emerge while keeping the optimization objective simple and outcome-based.

Student Archetypes With Distinct Learning Preferences

To design student archetypes with different latent variables that genuinely require different guidance strategies, two principles are needed. First, student improvement should depend on whether the teacher's instruction meets the demands represented by the student's latent variable. Second, student replies should reflect the preference represented by the latent variable. Together, these principles make the student's latent demands both consequential for learning and inferable from interaction.

In Sherpa, we instead use a weak model that cannot solve the problem unaided as the student backbone. This makes learning gains easier to attribute to tutoring, but the weaker model is also less reliable at expressing distinct behaviors through prompting alone. We therefore use gates to control which teacher instructions reach the student backbone. Specifically, we decompose adaptive teaching into two components: providing useful guidance and matching that guidance to the student's latent demands, and implement it with a guidance gate and an adaptive gate.

The guidance gate uses an LLM checker to determine whether each teacher instruction reveals the ground-truth final answer. Instructions that do so are blocked; intermediate reasoning and partial solutions remain allowed. This design permits substantive guidance while leaving the student to work out the final answer. The adaptive gate assigns students different preferences for teaching strategies drawn from widely studied pedagogical interventions: attempt diagnosis, causal justification, subgoal decomposition, step demonstration, independent verification, and contrastive comparison. Unlike the guidance demands shared by all student archetypes, these different demands for teaching strategies distinguish their latent variables. The teacher does not know the student's preferences before the interaction and therefore needs to adapt their teaching approach based on the feedback.

Formal gating mechanism: at turn t, a teacher instruction is accepted only if it passes both the guidance and adaptive gates. Let m_t ∈ {0,1} be the gate-acceptance indicator, with m_t = 1 when the instruction passes both gates and m_t = 0 when it is rejected. When m_t = 0, the student archetype directly returns a scripted reply stating why the instruction failed (e.g., "Don't directly show me the answer.", "Tell me the step-by-step plan."). When m_t = 1, the student backbone generates a reply based on the student context and the teacher instruction. The update rule for the student context appends accepted exchanges and keeps the context unchanged when the gate rejects.

Seven student archetypes are instantiated: six with distinct preferences enforced by their respective adaptive gates, and one without preference (None, Non.). The descriptions and names below distinguish them by their preferences. The training setup uses four of these archetypes (None, Non.; Attempt, Att.; Subgoal, Sub.; Contrast, Ctr.), making them in-distribution, while the remaining three (Causal, Cau.; Step, Stp.; and Verification, Ver.) are used for out-of-distribution evaluation.

  • Attempt diagnosis: identify and explain a specific issue in the student's latest attempt.
  • Causal justification: explain why a mathematical step or claim holds.
  • Subgoal decomposition: identify the current subgoal and relate it to the overall solution.
  • Step demonstration: demonstrate and explain one new step before returning control to the student.
  • Independent verification: check a claim through an alternative route.
  • Contrastive comparison: compare plausible alternatives and explain their decisive difference.
  • None, Non.: no specific preference; the student accepts any valid guidance.

Unlike prompted students, which are given the same six preferences explicitly in their prompts, these archetypes enforce preferences through the gating mechanism. This design ensures that the student's reply reflects the preference and that improvement depends on whether the teacher's instruction meets the demands represented by the student's latent variable.

Multi-Turn RL for Teacher Training

The teaching process follows a classroom cycle with three stages, which together constitute a teaching episode: preparation, in which the teacher prepares the material; tutoring, in which the teacher chats with the student; and testing, in which the student completes a test. Based on this design, we optimize the teacher via multi-turn reinforcement learning.

Preparation

The teacher privately solves the assigned problem and retains a solution draft as background knowledge, just like human teachers prepare a problem's solution and underlying reasoning before teaching it. For LLM teachers, particularly small models, keeping the solution draft available throughout tutoring helps reduce drift as the dialogue grows and improves teaching performance.

Tutoring

The teacher initiates the dialogue with a student archetype, with the problem and the solution draft at the beginning of the teacher context. At each turn, the teacher generates a complete response, consisting of private reasoning and a student-visible instruction. The private part allows the teacher to reason freely without being constrained by student preferences. After receiving the student reply, the teacher context is always updated. A turn budget is set for tutoring, while the teacher may also end the tutoring earlier, mirroring the conclusion of a lesson.

Testing

After tutoring, we test the student by providing the interaction history together with the original problem, and asking it to produce a complete solution. We estimate accuracy through repeated testing. To measure the student's unaided performance, we evaluate the same student model without providing the interaction history accumulated during tutoring.

Objective and Reward

The objective is to maximize the student's improvement on the problem. Let p(x,z,C_H^S) denote the probability that student archetype E_z answers x correctly with its context after H turns, and p(x,z,C_0^S) the probability for the same frozen student model with no interaction history. We maximize the difference between these two probabilities. The distribution over problems and student archetypes is denoted by D, and P_θ is the trajectory distribution induced by the teacher policy and student archetype. Because the student backbone is frozen, p(x,z,C_0^S) is independent of θ, so maximizing expected improvement is equivalent to maximizing expected post-tutoring success.

For a tutoring trajectory on problem x with student archetype E_z, the trajectory reward R is the difference between the empirical estimators of p(x,z,C_H^S) and p(x,z,C_0^S). These estimates are obtained by repeatedly testing the student with and without interaction history, respectively.

The algorithm design builds on critic-free RL algorithms. For each problem, we sample a group of trajectories sharing the student archetype and preparation context, each with a reward. Since rejected turns do not enter the student context, they make no contribution to student improvement. We therefore apply a mask to their turn-level returns without making any assumption about their quality. Let m_{i,t} ∈ {0,1} indicate whether turn t in trajectory i passes the guidance and adaptive gates. We define the masked return as U_{i,t} = m_{i,t} R_i. This design admits a local surrogate-objective interpretation: under the rollout policy, U_{i,t} serves as the return weight in a masked policy-gradient surrogate. Leave-one-out centering and turn-level normalization yield the advantage function, which is broadcast to all tokens in the teacher response and used to optimize the teacher using asynchronous PPO.

Gating Mechanisms for Adaptive Instruction

At turn t, a teacher instruction is accepted only if it passes both the guidance and adaptive gates. The guidance gate uses an LLM checker to determine whether each teacher instruction reveals the ground-truth final answer. Instructions that reveal the answer are blocked; intermediate reasoning and partial solutions remain allowed.

The adaptive gate assigns students different preferences for teaching strategies. At turn t, after the teacher generates its instruction, the gate checks whether the instruction matches the student's latent preference. If it passes, m_t = 1 and the student backbone generates a reply. If it fails, m_t = 0 and the student returns a scripted reason stating why the instruction failed (e.g., "Don't directly show me the answer.", "Tell me the step-by-step plan."). The scripted reason stays in the teacher context so that the policy can condition on it at the next turn.

This gating mechanism ensures that the teacher must infer each student's needs from interaction and adapt its teaching accordingly. The teacher cannot rely on a single explanation style; it must adjust its approach based on the feedback received.

Experimental Results Across Archetypes and Benchmarks

The authors evaluate Sherpa on mathematics problems from the MATH dataset, with filtering to remove less teachable cases. Problems beyond the teacher's capabilities and those the student can already solve unaided are removed. This filtering retains 759 training and 528 evaluation problems from the corresponding dataset splits. The teacher backbone is Qwen3-8B and the student backbone is Qwen3-1.7B.

Results on student archetypes show that SherpaDP, trained with diverse-preference students, achieves an overall test accuracy of 69.0%, outperforming the untrained backbone, PedagogicalRL, and SherpaNP on every student archetype. SherpaNP, trained with no-preference students, also improves on the untrained backbone, raising overall test accuracy from 48.6% to 52.0%, comparable to PedagogicalRL's 52.7%. This demonstrates the effectiveness of using student improvement to guide teacher training. Furthermore, SherpaDP performs substantially better than SherpaNP across student archetypes, demonstrating the importance of diverse student archetypes in training. The gains in accuracy differ across archetypes, indicating that improvements do not transfer equally across them.

On MathTutorBench, Sherpa raises the overall pedagogy score from 52.5% to 79.2%, outperforming PedagogicalRL on all four pedagogical metrics and achieving the highest average pedagogy score among the evaluated models. Training raises instruction-following win rates from 74.0% to 90.3% in the standard setting and from 70.6% to 87.2% in the hard setting. Scaffolding scores also increase from 33.1% to 72.0% and from 32.4% to 67.3%, respectively. These gains indicate that training improves both adherence to students' instructional requests and the ability to provide guidance that supports their progress. Meanwhile, mistake-location accuracy declines across all trained teachers, and most also regress on Socratic questioning, suggesting a side effect of specific training for teaching. Nonetheless, SherpaDP retains higher mistake-location accuracy than PedagogicalRL.

Human Pairwise Comparison and Alignment

To evaluate the effectiveness of our model in real human teaching, we construct dialogue contexts from labeled conversations. Human teachers substantially prefer our model: it receives 728 judgments versus 187 for the base model, corresponding to a 79.6% win rate excluding ties. The advantage holds across all four preference types (75.6-85.5%), including the two preferences unseen during training. Moreover, 87 of 96 teachers favor SherpaDP more often across their judgments, and majority voting prefers it on 270 of 320 response pairs versus 34 for the base model.

Participants' explanations most often attribute this preference to better guidance tailored to the student's requested approach, and sometimes to better identification and correction of mistakes. A minority prefer the base response when they find our explanations potentially overwhelming, suggesting that improved adaptivity can occasionally come at the cost of simplicity.

Evaluating the Training Design

The authors compare their student archetype design with prompted students, which are given the same six preferences explicitly in their prompts. Prompted students show neither a clear diagonal advantage in improvement nor consistently lower complaint rates along the diagonal. On the other hand, the archetype design shows both desired properties: improvement is concentrated along the diagonal, while complaint rates are lower for matched pairs. This indicates that the gating-based design more effectively elicits the student behaviors needed for the training environment compared to explicit prompts.

Different reward designs were also ablated. Adding gate penalties based on predefined criteria harms generalization, whereas training based on student improvement allows broader teaching capabilities to emerge. A small penalty leaves overall performance largely unchanged, whereas a larger penalty markedly reduces performance, including generalization to out-of-distribution archetypes. These results suggest that adding penalties based on predefined criteria harms generalization, whereas training based on student improvement allows broader teaching capabilities to emerge.

What This Means for AI Tutors

The Sherpa framework offers a practical pathway for building AI tutors that adapt to diverse learners. The key insight is that separating pedagogical specification from teacher optimization allows the model to learn adaptive teaching behaviors without needing predefined criteria for what good teaching looks like. The gating mechanism ensures that instructional strategies are consequential for student learning, while the multi-turn RL objective directly optimizes for student improvement.

For practitioners, the main components are the guidance gate (prevents answer disclosure), the adaptive gate (matches teaching strategy to student preference), and the multi-turn RL training loop (optimizes teacher based on student improvement). The framework uses seven student archetypes, but the design can be extended with additional preferences drawn from education literature. The mined dataset and code are publicly available, enabling reproducibility and further exploration.

One practical consideration is the choice of student backbone. The authors use a weak model that cannot solve the problem unaided, which makes learning gains easier to attribute to tutoring but may limit the student's problem-solving capacity. Alternative designs could use stronger student models while preserving the gating mechanism to control which teacher instructions reach the student.

Another consideration is the turn budget for tutoring. The authors set a turn budget while allowing the teacher to end the tutoring earlier. The optimal turn budget may vary by problem and student archetype, and adaptive turn selection could be an area for future work.

Finally, the authors acknowledge that specific training for teaching can have side effects, such as declines in mistake-location accuracy and Socratic questioning. However, the overall gains in pedagogy scores and human preference suggest that the benefits outweigh these costs, and the framework's outcome-based objective helps mitigate unintended consequences.

Read the paper on arXiv