Reasoning models have become remarkably capable at solving complex mathematical, scientific, and coding problems. But they pay for this capability with compute. A single reasoning trace can span tens of thousands of tokens, making inference slow and expensive. The dominant approaches to this problem fall into two categories. Some methods add early-stopping mechanisms at inference time, monitoring confidence or uncertainty signals to halt generation when the model appears certain. Others modify training directly, using reinforcement learning with length penalties or fine-tuning on curated short traces to explicitly encourage brevity. A team from the University of Maryland and Capital One has found a third path that avoids both patterns. Their method, called ConfSFT, trains models to predict their own confidence at intermediate reasoning steps using a self-supervised procedure. The training loss contains no term for reasoning length, no stopping signal, and no efficiency objective. Yet the resulting models generate up to 25 percent fewer tokens at matched accuracy across Gemma, Qwen, Nemotron, and GPT-OSS families on mathematical, scientific, and coding benchmarks.

The confidence signal inside reasoning trajectories

At any point during generation, a language model assigns probabilities to every possible next token. When a model has arrived at a correct intermediate answer, these probabilities concentrate on the right tokens. When reasoning remains uncertain, the distribution stays spread out. This internal signal can be extracted without gold labels or external judges. The authors define confidence at an intermediate step as the geometric mean of token probabilities assigned to the model's current trial answer. To obtain this trial answer, they append a fixed prompt such as "The final answer is" to the reasoning prefix and greedily decode. The resulting confidence score is a length-normalized likelihood computed entirely from the model's own distribution.

This signal carries real information. On Nemotron traces from AIME 2024, higher confidence at intermediate states correlates strongly with trial answer correctness. It also correlates with answer stability, meaning the intermediate guess increasingly matches the model's eventual final answer. Crucially, the expected accuracy gain from continuing to reason drops sharply as confidence rises. When confidence exceeds 0.95, further reasoning yields almost no average improvement, yet models routinely generate thousands of additional tokens. The authors quantify this with a utility metric: the difference between final answer correctness and intermediate answer correctness. The expected utility approaches zero at high confidence, confirming that models overthink problems they have already effectively solved.

How ConfSFT works

The method operates in iterative rounds. Each round begins by sampling reasoning rollouts from a small set of training problems using the current policy. No special instructions about efficiency or stopping are given. The rollouts are standard reasoning traces generated with the same decoding settings used at inference.

From each rollout, the method identifies decision points using textual markers. The default marker is "Wait", which many reasoning models emit spontaneously between logical sub-steps. At each decision point, the system constructs a confidence label by probing for a trial answer and computing its geometric mean token probability. This continuous score is quantized onto a grid of 50 percentage levels (2 percent increments) and formatted as a textual label like "72%".

A training example consists of the problem prompt, the reasoning prefix up to the decision point, a fixed priming phrase ("From 0% (very low) to 100% (very high), my confidence in the answer so far is"), and the quantized confidence label. Critically, only the confidence label positions are unmasked in the loss. The problem statement, reasoning prefix, and priming phrase are all masked. The model learns to predict the confidence label conditioned on the intermediate reasoning state using standard next-token cross-entropy. The objective contains no term for reasoning length, stopping, or efficiency.

Because both rollouts and confidence labels are policy-generated, they are refreshed at each training round as the policy changes. The training set contains only 600 problems from AIME 2000-2023, partitioned into eight groups allowing up to eight rounds. At inference, the fine-tuned model uses standard generation with no confidence elicitation, no early-exit mechanism, and no verifier. The model does not spontaneously emit confidence values or the priming prefix during its reasoning traces.

Results across model families and benchmarks

The authors evaluate ConfSFT on four model families: Gemma-4-E2B, Qwen3-4B, Nemotron-Nano-8B, and GPT-OSS-20B. Training uses 600 AIME problems with AIME 2024 held out for validation. Evaluation spans five benchmarks: AIME 2025 (mathematics), GSM8K (grade-school math), GPQA-Diamond (scientific reasoning), LiveCodeBench (coding), and HumanEval (code generation). Sixteen completions are sampled per problem.

ConfSFT reduces average generated tokens by 11.1% on Nemotron, 10.3% on Gemma, 19.2% on Qwen, and 11.1% on GPT-OSS, with all reductions statistically significant while accuracy changes remain indistinguishable from zero. The gains transfer across domains despite training exclusively on mathematics problems. On GPQA-Diamond, ConfSFT reduces tokens on all four model families while preserving accuracy. On LiveCodeBench and HumanEval, the same pattern holds.

Comparison with explicit efficiency methods reveals the core finding. A&Z, which adds response length directly into its RL objective, and On-Policy SFT, which fine-tunes on selected concise trajectories, both explicitly target shorter reasoning. On Qwen3-4B, On-Policy SFT reduces tokens by 17.2% at 69.3% average accuracy, while ConfSFT achieves 19.2% at 69.5% accuracy. On Nemotron, ConfSFT similarly matches or exceeds explicit efficiency baselines. The method achieves comparable gains without any supervision favoring brevity.

Inference-time early stopping tells a different story. DEER, which uses confidence to terminate reasoning, produces larger raw token reductions but with severe accuracy costs on coding benchmarks. On Nemotron, LiveCodeBench accuracy collapses from 44.2% to 17.4% and HumanEval from 89.7% to 47.1%. DEER also degrades GSM8K accuracy from 91.8% to 89.2% and GPQA-Diamond from 52.5% to 49.8%. Its effectiveness varies wildly across tasks. ConfSFT achieves efficiency without accuracy penalties and without external stopping criteria.

In absolute terms, GPT-OSS-20B drops from 5,802 to 5,235 average tokens on AIME 2025, saving 567 tokens per problem. Qwen3-4B drops from 13,268 to 11,365, saving roughly 1,900 tokens. At inference scale, these savings compound substantially with zero deployment infrastructure changes.

Training dynamics and reasoning composition

Tracking training dynamics shows a clean relationship. As confidence prediction error (mean absolute error against self-supervised targets) drops, average generated tokens fall while accuracy remains flat. The model learns to predict confidence and reasoning becomes more efficient simultaneously. The predicted confidence distribution also transforms from concentrating on a few discrete levels to spanning a broad range, indicating the model learns graded uncertainty judgments rather than defaulting to narrow responses.

A deeper analysis uses Schoenfeld's Episode Theory to decompose reasoning traces into eight cognitive categories: Read, Analyze, Plan, Implement, Explore, Verify, Monitor, and Answer. The authors compare ConfSFT with explicit efficiency methods including L1-Max, ThinkPrune, A&Z, and On-Policy SFT. They measure the change in token share for each episode category and summarize overall reallocation using total variation distance.

Explicit efficiency methods cause large shifts in reasoning composition. L1-Max and ThinkPrune substantially redistribute computation, with large changes in Implement, Explore, and Verify categories. These methods alter what the model thinks about, not just how much it thinks. ConfSFT induces relatively small changes in episode shares, preserving the base policy's high-level reasoning structure. The model still reads, analyzes, plans, implements, explores, verifies, monitors, and answers in roughly the same proportions, but with fewer tokens at each step. This suggests confidence supervision achieves uniform compression across all reasoning behaviors rather than selectively suppressing particular ones.

Ablations and design choices

The authors test whether gains are specific to confidence or arise from fine-tuning on intermediate states generally. Three alternative supervision signals are evaluated: position-based targets derived from token location, binary correctness targets requiring gold answers, and shuffled confidence labels that preserve label distribution but break state-label correspondence. Shuffled confidence largely eliminates or reverses efficiency gains, confirming the precise state-label mapping matters. Position and binary correctness yield weaker, less consistent improvements across models. The graded confidence signal with preserved correspondence is doing the specific work.

The decision-point marker also proves flexible. While "Wait" is the default, paragraph boundaries ("\n\n") work as a generic structural marker. On Gemma-4-E2B, paragraph-boundary decision points still reduce tokens by 8.4% while maintaining accuracy, demonstrating the approach does not depend on models emitting a particular token.

Limitations and open questions

The training data consists of 600 mathematical competition problems. While results transfer to scientific and coding benchmarks, the paper does not test domains with substantially different reasoning structures such as multi-document analysis or open-ended dialogue. The confidence definition uses geometric mean token probabilities, one specific choice among many possible formulations. The method requires base models to produce sufficiently structured reasoning traces with detectable intermediate states. Models generating very short or unstructured traces may not offer enough decision points. The iterative on-policy training adds overhead, though the 600-problem scale keeps costs modest compared to full RL pipelines. Experiments cover models from 2B to 20B parameters. Whether confidence supervision produces similar gains at larger scales or with models trained from scratch remains untested.

What this means for practitioners

Efficient reasoning does not require explicitly teaching models to be brief. A practitioner can apply ConfSFT with a few hundred problems, a standard next-token loss, and no inference infrastructure changes. The fine-tuned model uses the same decoding configuration as the base model. The 600-problem requirement is remarkably low compared to typical supervised fine-tuning or RL datasets. The self-supervised confidence labels require no human annotation.

Preserving reasoning composition is an underappreciated advantage. Systems that explicitly penalize length can inadvertently strip valuable reasoning steps, producing shorter but lower-quality explanations. ConfSFT reduces token counts while keeping the cognitive architecture intact, which matters for applications where reasoning must be auditable or comprehensible.

The broader conceptual contribution is that metacognitive signals like confidence may serve as powerful latent training objectives for shaping behavior. Efficiency emerges as a downstream consequence rather than a direct target. This opens a design space where other desirable properties such as caution, verification, or openness to revision might emerge from supervising related metacognitive signals.

ConfSFT does not replace all existing efficiency methods. In contexts requiring inference-time adaptability, early-stopping mechanisms may still be valuable. But the results establish that a substantial fraction of efficiency gains sought by more complex approaches can be achieved through a simpler principle: teach the model to know what it knows, and efficiency follows.

Read the paper on arXiv