The study of reasoning in base models has taken a surprising turn. Rather than relying on reinforcement learning to shape behavior, researchers have found that specific starting token cues can unlock reasoning capabilities already present in the model weights. This finding suggests that the association between cue tokens and reasoning patterns is baked into the training data itself, and can be triggered at inference time without any weight updates.
Cue-Driven Reasoning in Base Models
The central observation is that fixing particular starting tokens makes a base model's performance competitive with its reinforcement learning (RL)-trained counterparts on math and coding benchmarks. For the OLMo-3-7B model, appending ".\n\nOkay" as a cue raises MATH-500 pass@1 accuracy from 42% to 78%. For Qwen3-14B, the cue "Alright," lifts MATH-500 pass@1 from 72% to 87%. These numbers place the cued base models on par with, or exceeding, models that have undergone RL fine-tuning.
The effect is not merely additive. When the same cues are left in place during RL training, the trained models assign higher probability to those cue sequences. However, if the cues are fixed post-training, RL's performance advantage over the base model shrinks significantly, indicating that much of RL's gain is captured by learning to emit these natural starting patterns.
From Arbitrary Words to Reasoning Triggers
Perhaps the most striking result is that arbitrary words can be turned into effective reasoning cues through causal intervention on the training data. The authors demonstrate that the word "chicken", when associated with the right downstream context, becomes a trigger for structured math reasoning. Similarly, the prompt instruction "Think duck duck goose" elicits reasoning behavior nearly as effectively as the established "Think step by step".
The intervention works by modifying the association strength between a token and the reasoning trajectories that follow it in the training distribution. The mechanism is not semantic in the usual sense—"chicken" has no mathematical meaning—but rather statistical: the token becomes a marker that the model has learned to associate with particular reasoning chains present in the pre-training corpus.
Hidden State Correlations With Document Types
Beyond behavior, the authors probe the internal representations induced by different cues. Hidden state activations at the start of the response correlate with distinct document types from the training set. A cue tied to programming examples produces hidden states that resemble code documentation; a cue tied to proofs produces states resembling theorem-proof structures. This correlation suggests that the model's reasoning pathways are partially grounded in the kinds of text it saw during pre-training, and that cue tokens act as switches between these pathways.
Safety Implications
The study also examines language model safety through the lens of cueing. Different starting tokens elicit distinct refusal and compliance behaviors, and these behaviors map onto the types of training data that conditioned the model's responses. A cue associated with safety-conscious documentation produces more frequent refusals on risky queries, while a cue associated with permissive coding assistants produces more compliance. The effect persists even when the underlying model weights are unchanged, confirming that prompt framing can steer safety-critical behavior in predictable ways.
Practical Takeaways
For practitioners, the findings imply that RL from human feedback (RLHF) may not be the only—or even the most efficient—path to robust reasoning. Carefully chosen start-of-response tokens, derived from or validated on the target task distribution, can activate competent reasoning in off-the-shelf base models. This reduces the cost and data requirements associated with full RL pipelines.
For interpretability researchers, the work provides a concrete method for probing how training data associations manifest in model activations. By intervening on specific tokens and measuring the resulting changes in hidden states and output distributions, it is possible to trace which parts of the corpus the model has latched onto for particular behaviors.
Conclusion
The paper establishes that base models carry reasoning capabilities that can be triggered by specific token cues, and that these cues owe their effectiveness to associations formed during pre-training. The ability to turn arbitrary words into reasoning triggers, and to predict safety behaviors from cue patterns, opens new directions for both model deployment and mechanistic analysis. The findings suggest that the gap between base models and RL-trained models is narrower than previously assumed, provided the right input cues are used.
Read the paper on arXiv