People learn not only by repeating successful actions but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? This question drives the investigation of Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO's 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.
From Reward Signals to Retrospection Signals
Reinforcement learning from human feedback and related reward-based fine-tuning methods have become the standard for aligning language model agents with desired behaviors. These approaches typically require external reward signals, human demonstrations, or learned value functions to update the policy. ROFT departs from this paradigm by dispensing with any external reward signal entirely. Instead, the training signal originates from the model's own prior attempt at a task. After the model generates a completion for a given problem, it is prompted to produce a retrospective explanation—essentially, a post-hoc justification or analysis of what happened during the attempt. This explanation is then used as the target for next-token prediction fine-tuning. The key insight is that the act of explaining an experience, even when the experience was initially unsuccessful, can restructure the model's internal representations in a way that makes future successful completions more likely. By treating the retrospective explanation as the unit of learning, ROFT bypasses the need for reward modeling, policy gradient computation, or any form of environment interaction during training.
The procedural flow is deceptively simple: sample a model trajectory, prompt the model to explain its own behavior, and fine-tune on the token sequence of that explanation. Despite its simplicity, the method produces measurable improvements on complex reasoning benchmarks. The paper attributes this to the inductive bias built into the explanation format: when a model is forced to articulate its reasoning steps, it implicitly reviews which actions led toward the goal and which diverged, creating a self-supervised credit-assignment mechanism. This mechanism operates without any explicit reward signal, making ROFT substantially simpler to implement and debug than RL-based alternatives.
Experimental Setup and Baseline Comparisons2>
The experiments use Qwen3.5-4B as the base model and evaluate on two subsets of SWE-bench: Verified and Pro. SWE-bench is a benchmark for software engineering tasks, specifically focused on fixing issues in real GitHub repositories. The base model attempts each problem, and its success or failure is recorded. ROFT then fine-tunes the model on explanations generated from these attempts. The training loop operates online: after each fine-tuning step, the next batch of problems is sampled, and the process repeats. No verifier is used during training to label successes or failures; the only signal is the model's own generated explanation.
The comparison baseline is GRPO, a reinforcement policy optimization method that also fine-tunes the model on software-engineering tasks but uses reward-based updates. GRPO requires a verifier to provide scalar rewards (typically test-pass/fail signals) after each attempt, and these rewards shape the policy update. In the reported results, ROFT reaches 49.2% solve rate on SWE-bench Verified after 20 updates, while GRPO requires 40 updates to reach 48.0%. On SWE-bench Pro, the gap is similar: 26.8% for ROFT after 20 updates versus 25.3% for GRPO after 40 updates. Beyond the final performance, ROFT shows faster early progress, meaning that each update contributes more to the cumulative improvement than in the GRPO runs. This early progress is notable because it suggests that the retrospective explanation signal is more efficient per training step than the reward signal used by GRPO.
Learning From Unsuccessful Trajectories
A particularly striking finding is that ROFT can learn to solve individual tasks on which all 64 sampled base-model attempts initially failed. In these cases, the starting model has zero success rate on the problem, yet after ROFT fine-tuning on the model's own explanations, the fine-tuned model can solve the problem. This demonstrates that learning can begin without any initially successful trajectories, a property that few fine-tuning procedures possess. Standard supervised fine-tuning on failed attempts typically does not yield improvement because there is no correct target to copy; RL methods also struggle when the reward signal is uniformly zero. ROFT's success in this regime indicates that the explanation generation process itself provides enough structural information to guide the model toward better behavior, even when the initial attempts are entirely wrong.
The paper offers a behavioral analysis of how ROFT assigns credit to actions. The retrospective explanation, while not providing a dense reward signal, does indicate which parts of the reasoning path were productive and which were not. The fine-tuning objective, by predicting the next token of this explanation, indirectly shapes the model's policy toward actions that are more likely to be included in a coherent, successful explanation. In practice, this means that good actions are encouraged and incorrect ones are discouraged, even though no explicit reward signal ever mentions "correct" or "incorrect." The credit assignment is soft and distributed, emerging from the language modeling objective on the explanation text.
Prompting for Directness and Attempt Length
The authors also investigate how the content of the retrospective explanation affects subsequent behavior. They experiment with prompting the model to emphasize more direct solutions in its explanation. When the prompt encourages the model to focus on the most efficient path to the correct answer, the subsequent attempts after fine-tuning are shorter, even though no explicit length penalty is applied. This result suggests that the explanation prompt shapes the model's reasoning policy in a way that prunes unnecessary or redundant steps. The model learns to allocate its compute and reasoning depth more efficiently, producing shorter but more effective attempts. This effect is observed without any constraint on output length, making it a natural byproduct of the explanation-driven fine-tuning signal.
Broader Implications for Agent Training
ROFT's approach opens a new direction for agent training that is orthogonal to the current trend of scaling reinforcement learning. The method is minimally invasive: it requires only a prompt modification to generate explanations and a standard fine-tuning loop. No reward model, no value function, no importance sampling, and no environment interaction beyond the initial trajectory generation. This simplicity makes ROFT an attractive option for teams that want to improve agent performance without the engineering overhead of RL pipelines. Moreover, because the signal is self-generated, it can be applied in settings where human feedback is scarce or expensive, such as specialized code-fixing domains or low-resource languages.
The finding that learning to explain can also improve learning to do suggests a broader principle: the act of articulating one's reasoning reorganizes the model's knowledge in ways that are beneficial for future reasoning tasks. This principle, well-documented in human learning theory, appears to translate to language-model agents. The paper motivates further study of explanation-to-action transfer, including how different explanation formats (step-by-step, high-level summary, error analysis) might interact with various fine-tuning objectives.
Limitations and Open Questions
The paper acknowledges several limitations. The experiments are confined to the SWE-bench software engineering benchmark; it is unclear whether ROFT's benefits extend to other agent domains such as mathematical reasoning, long-horizon planning, or multi-turn dialogue. The base model size (4B parameters) is relatively small, and it is possible that the explanation signal is more impactful for smaller models that have more room for representation shifts. The method also relies on the model's ability to generate coherent explanations, which may degrade for very difficult or out-of-distribution tasks. Finally, the paper does not explore the computational cost of generating explanations at each training step, though the authors note that the explanation generation step is lightweight compared to the RL baselines that require multiple environment rollouts per update.
Practical Takeaways for Engineers
For engineers looking to improve agent performance without deploying a full RL infrastructure, ROFT offers a low-friction alternative. The implementation steps are: (1) after each model trajectory, append a prompt asking the model to explain its reasoning or why it succeeded/failed; (2) collect the explanation token sequence; (3) perform a standard next-token prediction fine-tuning step on that sequence; (4) repeat. The paper's results suggest that just 20 such updates can match or exceed 40 updates of a reward-based method. The approach is especially well-suited for code-editing agents, theorem-proving assistants, and any system where the model's internal reasoning can be captured in natural language.
Developers should also consider the phrasing of the explanation prompt. The experiments show that prompting the model to emphasize direct, concise solutions yields shorter subsequent attempts, which can reduce inference cost and improve user experience. Even without a length penalty, the model learns to prefer more efficient reasoning paths when the explanation signal encourages it. This insight can be combined with other fine-tuning tricks, such as data formatting or curriculum learning, to further boost performance.
Conclusion
Retrospection-Only Fine-Tuning demonstrates that self-generated explanations can serve as a powerful training signal for language-model agents, rivaling or surpassing reward-based fine-tuning while requiring far less infrastructure. The method's success in learning from entirely unsuccessful trajectories, its efficiency per update, and its ability to shape more direct reasoning patterns all point to a simple but profound principle: forcing a model to explain its own behavior reorganizes its knowledge in ways that make future behavior more effective. As agentic systems become more prevalent, methods like ROFT that leverage self-supervised signals from the model's own experience will likely play an important role in their development and alignment.
Read the paper on arXiv