Reinforcement learning with verifiable rewards improved reasoning capabilities of large language models yet predictions stayed sensitive to task-irrelevant prompt features. Researchers investigated this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Analysis revealed substantial variation in token-level sensitivity and showed that suppressing high-drift token candidates during decoding improves reasoning accuracy without updating model weights.
These findings highlight a limitation of Group Relative Policy Optimization GRPO which assigns the same outcome-derived advantage to every response token and may reinforce potential spurious dependence alongside useful reasoning. Motivated by this observation the authors introduced Semifactual Credit-Augmented Policy Optimization SCAPO a causally inspired variant of GRPO that incorporates semifactual stability into token-level credit assignment. SCAPO measures token probability drift for fixed responses under semifactual interventions and uses normalized stability scores to reduce advantages for relatively unstable tokens during early training while granting no additional credit for stability alone.
On Qwen3-4B-Base and Qwen3-1.7B-Base SCAPO improved AIME 2024-2026 accuracy over GRPO by 5.63 and 4.17 percentage points respectively. At both model scales SCAPO achieved the best results on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks among compared methods. These results suggest that semifactual stability provides an effective training signal for improving reasoning and generalization through finer-grained credit assignment in RLVR.
Semifactual interventions expose token-level sensitivity
The authors began by probing how LLM reasoning responds to small prompt perturbations that keep the core problem and correct answer unchanged. By applying semifactual interventions they observed that individual tokens vary widely in how much their probability shifts when the prompt changes. Some tokens are highly sensitive to superficial wording while others remain stable. This variation means a uniform credit signal cannot distinguish between tokens that contribute reliable reasoning steps and those that merely echo surface patterns.
Group Relative Policy Optimization's uniform credit assignment
Group Relative Policy Optimization GRPO is a widely used method for reinforcement learning from verifiable rewards. In GRPO the advantage given to each token depends only on the final reward of the entire response. Every token in a response receives the same advantage signal regardless of when it appears or how it influences the model's chain of thought. The paper argues this design may reinforce spurious token-level dependencies that correlate with reward by chance rather than by causal contribution to correct reasoning.
Semifactual Credit-Augmented Policy Optimization introduced
Semifactual Credit-Augmented Policy Optimization SCAPO builds on GRPO by adding a stability-based modulation to the token-level advantage. The key idea is to measure how much a token's probability drifts when the prompt is altered under semifactual constraints that preserve the problem and answer. Tokens that show large drift receive a reduced advantage while stable tokens keep their full credit. This adjustment is applied without changing the model's architecture or requiring additional forward passes during decoding.
Normalized stability scores modulate token advantages
SCAPO computes a stability score for each token position by comparing the model's output distribution under the original prompt and under a semifactual variant. These scores are normalized across the response so they sum to one. The final advantage for each token is the GRPO advantage multiplied by its normalized stability score. This product ensures that tokens essential for consistent reasoning retain strong gradients while noisy or fragile tokens receive weaker updates.
AIME and mathematics benchmark improvements
Experiments on the AIME 2024-2026 mathematics benchmark showed consistent gains for SCAPO over GRPO. On the Qwen3-4B-Base model SCAPO raised exact-match accuracy by 5.63 percentage points. On the smaller Qwen3-1.7B-Base model the gain was 4.17 percentage points. Both models used the same training protocol and reward function; the only difference was the stability-based advantage modulation. On additional mathematics benchmarks SCAPO also placed first or tied for first across most evaluated sets, outperforming GRPO and other RLVR variants.
Out-of-distribution generalization results
Beyond in-distribution mathematics problems the authors tested on out-of-distribution benchmarks that include harder problem distributions and domain-shifted geometry questions. SCAPO achieved the best results among all compared methods on every out-of-distribution evaluation set. The improvement was most pronounced on benchmarks that require multi-step reasoning under distribution shift, suggesting that the stability-based credit assignment helps the model learn more robust reasoning patterns rather than overfitting to specific prompt styles.
Broader impact on reasoning and generalization
The paper concludes that incorporating semifactual stability into token-level credit assignment is a simple yet effective modification to existing RLVR pipelines. Because the stability computation reuses the same forward passes needed for the reward signal it adds virtually no extra latency. The authors release their code and configuration to encourage follow-up work on stability-aware policy optimization for other reasoning domains.