Reinforcement learning with verifiable rewards has become the dominant recipe for improving the reasoning abilities of language models. RLVR samples solutions to problems, scores each solution with a programmatic verifier, and shifts probability toward those that pass. In principle, a model can sample novel ideas absent from its prior training data. In practice, however, augmenting RLVR with strong novelty incentives has seen limited success and can degrade model quality. Because verifiable rewards supervise only a narrow slice of the model's knowledge and behavior, such degradations are difficult to recover from. Instead, we decouple exploration from optimization in a framework we call Exploration-Distillation (ExpDis). We train one or more explorer policies with a novelty bonus in the reward, filter their trajectories for correctness and quality, and distill them into a separate student policy. The student policy is then trained without a novelty bonus. We repeat the above procedure for several rounds, alternating between exploration and optimization. This decoupling allows us to aggressively scale exploration without degrading the student policy. Across seven mathematical reasoning benchmarks and two model families, ExpDis outperforms DAPO at the same wall-clock budget. Moreover, we observe improved pass@k scaling, indicating that ExpDis produces models that generate more diverse correct solutions.
How RLVR promises discovery but caps exploration with correctness pressure
Reinforcement learning with verifiable rewards (RLVR) has become the dominant recipe for improving the reasoning abilities of language models. RLVR samples solutions to problems, scores each solution with a programmatic verifier, and shifts probability toward those that pass. By iteratively sampling new solutions and reinforcing successful ones, RLVR can in principle discover and then reinforce new solutions that were absent from the model's training data. In practice, however, it can only reinforce what the model samples, so discovery depends on exploration.
A natural way to boost exploration is to add an explicit novelty bonus to the reward, following a long line of intrinsic-motivation methods in RL. When the model explores new solution strategies, some of them may succeed but often most of them fail, leading to a temporarily degraded model as we now are left with a model that mostly generates failed solutions. Though our initial policy is a pretrained LLM with vast knowledge including virtually all of Wikipedia, the reward supervises only a tiny slice of its knowledge. This means an LLM could forget most of Wikipedia and still achieve high reward. The reward therefore cannot detect damage that exploration causes outside that slice, let alone repair such degradations. We argue that this failure is not a property of novelty as a signal, but of where the novelty bonus is applied.
The novelty bonus trap: why adding intrinsic motivation to the reward backfires
Adding a novelty bonus to the RLVR reward seems like a straightforward way to encourage the model to try unfamiliar reasoning paths. In practice, however, the method has seen limited success. When the model explores new solution strategies, some of them may succeed but often most of them fail, leading to a temporarily degraded model as the policy shifts probability toward those failed solutions. Because the verifiable reward supervises only a narrow slice of the model's knowledge and behavior, achieving high reward does not prevent general capability degradation. An LLM could forget most of its broad training distribution and still achieve high reward on the narrow math task, meaning the reward signal cannot detect or repair the degradation.
This is the core problem that the paper identifies: the reward function simply does not span the model's full capabilities. Adding a novelty bonus directly to the DAPO reward increases diversity but degrades performance on prior capabilities relative to the base model. The explorer can be driven to far higher diversity without harming the student only when exploration and optimization are decoupled, allowing the student to absorb the explorer's discoveries while remaining insulated from the novelty-driven pressure that corrupts the deployed model.
Exploration-Distillation: a separate explorer policy with a novelty bonus
Exploration-Distillation trains an explorer policy with a novelty bonus to generate diverse trajectories, filters them for quality and correctness, and distills the resulting trajectories into a student policy. The student policy never sees a novelty reward. It learns from the explorer only through filtered trajectories, and is then trained with standard RLVR using only a correctness reward. The two models communicate through filtered trajectories rather than parameters, so the student policy absorbs the explorer's discoveries without inheriting its degradations induced by exploration.
Decoupling makes exploration safe to scale aggressively, since the student policy never receives the novelty bonus and the filter keeps failed trajectories out of its training data. We can scale breadth by running several explorers in parallel and distilling their pooled trajectories into one student policy. We can also scale depth by repeating the pipeline for multiple rounds that each start from the previous student. We scale both of these axes independently and find that their composition enables sustained performance gains even when all experiments are wall-clock matched.
Parallel explorer policies and pooling diverse reasoning traces
Rather than train a single explorer policy, we can train K independent explorer policies and pool their trajectories before filtering and performing distillation. Compared to the single-explorer variant, training parallel explorers can increase the diversity of trajectories we curate for the student. This diversity increases coverage over solutions and boosts the probability of discovering successful ones.
We set the novelty bonus weight identically for all peer explorer policies, as we find little benefit from varying them. In our experiments, we divide compute among explorers, requiring no additional wall-clock time over a single explorer. The more compute a training run uses, the more parallel explorers we expect to be optimal.
Iterative Exploration-Distillation: repeating rounds of expansion and consolidation
We can scale the above method by performing many rounds of alternating exploration and distillation. After producing a student policy, we can use this student as a strong initialization for both the next explorer policy and subsequent student policy. Formally: to train student π_t, we first train one or more explorers μ from π_{t-1}, filter the trajectories, distill those trajectories onto π_{t-1} via supervised finetuning and then perform standard RLVR on the distilled checkpoint without an exploration bonus to produce π_t.
We partition the training prompts into R non-overlapping shards of equal size. Each round draws its training prompts from its assigned shard, pairing a stronger initialization with prompts it has not yet trained on. Because exploration is split across several rounds, we can set its strength per round through the novelty-bonus weight λ, which we anneal on a fixed schedule from 0.75 to 0.25 over four rounds. Early rounds, where correct trajectories are still undiscovered, get strong novelty pressure whereas later rounds get lower novelty pressure to consolidate what has been found.
Seven benchmarks and two model families: ExpDis beats DAPO at equal compute
We evaluate ExpDis on mathematical reasoning benchmarks with Qwen3-1.7B, Qwen3-4B, and Ministral-3-3B-Instruct-2512 as base models. At the same amount of compute, ExpDis outperforms DAPO in both pass@1 and pass@64. Furthermore, spreading this fixed budget over multiple parallel explorers and several rounds of the ExpDis procedure yields further gains. Adding a novelty bonus directly to DAPO, in contrast, increases diversity but degrades performance on general knowledge benchmarks. In ExpDis, the explorer can be pushed to far higher diversity without harming the student, which improves in both diversity and accuracy and avoids degradation.
A single round of ExpDis already outperforms DAPO trained for 4× as many steps. By spreading the same compute over several rounds of alternating exploration and optimization, we observe further gains. All ExpDis runs outperform this prolonged DAPO run in both average performance and pass@k up to 64 generations per problem.
pass@k scaling: more diverse correct solutions emerge
RLVR is known to collapse entropy and sharpen the sampling distribution as training proceeds. Counteracting this collapse within a single model requires a delicate balance, since the same model must both explore diverse reasoning strategies but must simultaneously be optimized for correctness and avoiding incorrect solutions. Adding a novelty bonus to DAPO illustrates this trade-off: it increases both lexical and answer diversity relative to standard DAPO, yet slightly reduces mathematical reasoning performance.
We find that we can push the explorer toward far greater diversity than a single model could tolerate while maintaining or improving on downstream accuracy. Explorer policy can be driven to more diversity without harming the student. Since we filter for correctness and quality before distilling, a moderate increase in explorer diversity brought on by tuning the novelty bonus weight does not negatively impact the student policy. ExpDis generations contain more semantically diverse mathematical reasoning strategies.
Protecting general capabilities: why the student stays strong without a novelty bonus
A verifiable reward, typically applied to math and code tasks, directly supervises only a narrow slice of a model's capabilities. In our experiments, DAPO with the RND novelty bonus degrades below the base model on prior capability benchmarks despite improving on math. Conversely, ExpDis gets the benefits of exploration without degradation. We report results for each general capability benchmark including MMLU-Pro, MMLU-Redux, GPQA-Diamond, ZebraLogic, and IFEval. This serves as a proxy for capabilities not directly supervised by our training dataset and reward to study the extent to which exploration pressures can degrade a model's knowledge over time.
While the drops we observe are modest at our scale, we speculate that larger-scale RLVR runs that apply exploration bonuses directly to the trained model may accumulate substantially larger degradations over time, whereas decoupling avoids this risk of corrupting the deployed model. The student both receives the benefit of exploration via distillation, which is off-policy, and also the benefit of on-policy training during its DAPO stage, inheriting none of the degradation from exploration pressure.
Policy evolution across multiple ExpDis rounds
We study how the explorer and student policies evolve across several iterations of ExpDis. The explorers repeatedly expand the space of reasoning strategies available to the student, and the student keeps a growing share of that expansion while leaving its costs behind. Within each round, the novelty bonus pushes the explorer entropy and semantic diversity well above the previous student, and subsequent distillation on the shortest correct trajectories contracts them again. We find that both the student's diversity and accuracy ratchet upward together across subsequent rounds.
Explorer policies reach roughly the same semantic diversity in every round, but the gap between them and the student they produce nearly closes by the fourth round. One explanation is that each round's explorers start from a stronger student and therefore solve more of the training problems, so more problems contribute a correct trajectory to the distillation set and the student inherits a broader range of the explorer policy's solutions. The student does not retain the lexical diversity of its explorers: it declines over rounds even as the student's accuracy improves. This suggests that lexical diversity is not a useful target for exploration.
What working developers can take from ExpDis for their own RL pipelines
For developers working with RLVR, the key insight is that exploration pressure and optimization pressure should not share the same model instance when the goal is to discover new reasoning strategies while preserving prior capabilities. By training a separate explorer policy with a novelty bonus, filtering its trajectories for correctness and quality, and distilling those trajectories into a student that trains without a novelty bonus, it becomes possible to scale exploration aggressively without corrupting the student policy. The framework is compatible with any exploration mechanism, since the student learns only from the explorer's filtered trajectories. Since the explorer is never deployed, it can also use mechanisms too aggressive to apply to the deployed model directly. Exploration methods that would destabilize standard RLVR may therefore become viable under ExpDis. Developers can allocate compute across parallel explorers and iterative rounds, observing that depth contributes more to pass@1 improvements while breadth widens the set of problems with correct trajectories. The annealing schedule for the novelty bonus weight from 0.75 to 0.25 over four rounds follows the common practice of decreasing exploration pressure as training progresses, allowing early rounds to discover novel strategies and later rounds to consolidate them.