Large language models implicitly infer attributes of their users and adjust their behavior accordingly, yet these internal user models remain difficult to inspect and causally manipulate. A team from the Technical University of Munich introduces Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external annotations. Unlike conventional probing, BSD isolates not only information present in activations, but a state whose causal role can be directly tested. Across multiple model families, BSD faithfully recovers user beliefs and enables substantially stronger interventions than matched hidden-state steering. Perhaps most crucially, the authors find that refusal depends not only on the request, but on the model's inferred user intent: changing this belief alters refusal while holding the request fixed. The authors further uncover a striking cross-model regularity: independently trained LLMs converge on a shared geometry for representing their users. Together, these results reveal implicit user models as readable and causally writable internal states with direct implications for AI safety, shaping how models condition safety decisions on whom they believe they are interacting with.

The opacity of internal user models

An LLM responds not only to what is being asked, but also to whom it believes is asking. Models infer attributes of their users from conversational context and condition their behavior on these inferences, affecting personalization, compliance, and even safety-relevant behavior. Yet the internal user model driving these effects remains largely opaque. The authors therefore ask a mechanistic question: what does an LLM believe about its user, and can that belief be isolated and causally modified?

Existing work provides important pieces of this picture. Linear probes show that attributes such as demographics can be decoded from hidden activations, while activation steering demonstrates that some of these representations can be causally manipulated. More expressive approaches such as LatentQA train auxiliary decoders to extract open-ended user representations from model activations. However, these approaches typically begin from controlled attributes, synthetic cues, or representations optimized primarily for readout. Existing approaches therefore largely treat reading and intervention as separate stages; BSD instead learns a single bottleneck explicitly through its ability to support both. This distinction matters because a model's belief about its user has no external ground truth: the relevant object is not necessarily who the user is, but who the model believes the user to be.

Belief Self-Distillation

The method extracts a low-dimensional user vector from a frozen model and trains it to reproduce the model's own beliefs when written back into a second frozen copy of the same model. The teacher observes the conversation and encodes its hidden state at a chosen layer and the last token position. A low-dimensional read-projection compresses this hidden state into a user vector. The student receives only the attribute probe and never observes the conversation; its only conversation-specific information is the user vector, which is written into the student through a learned write-projection. Only the read and write maps are trained. The training objective minimizes the Kullback–Leibler divergence between the teacher's answer distribution and the student's answer distribution after the user vector is injected. Because the student never observes the conversation, the information required to reproduce the teacher's belief must pass through the user vector. The objective therefore selects a compact representation according to the information the model itself uses to characterize the user. Together, the read and write maps form a read–write interface to the model's user state.

The attribute set instantiated by BSD spans 13 attributes spanning two groups. Socio-demographic attributes (e.g., continent, gender, education) follow the standard axes studied in prior work on implicit personalization. Motivated by recent findings on assistant personas, the authors introduce safety-relevant attributes for user models – such as truth-seeking intent, interaction style, and level of AI trust – to study how these traits modulate the model's safety behavior.

Reading and writing the user state

The authors evaluate the user vector at both ends of its read-write interface. The read evaluation tests fidelity: how much of the teacher's belief remains linearly decodable from the compressed state relative to the raw hidden states. The write evaluation tests causality: whether intervening on the user vector shifts the model's beliefs toward a chosen target direction.

To test the decodability of the extracted beliefs, the authors train linear probes on the distilled user vectors. Each probe predicts the teacher's inferred belief for an attribute, i.e., the argmax of the belief distribution. At rank 128, the user vector compresses the hidden state by a factor of 32× across all evaluated models. To test whether this compression preserves readout, the authors compare probes trained on the user vector against probes trained directly on the uncompressed hidden states: since the user vector is computed from the hidden states, the raw states contain at least as much decodable user information and serve as a natural upper bound. Additionally, probes fitted on random vectors serve as a baseline for the lower performance bound.

Regarding causal interventions, the authors steer the model using contrastive activation addition. For a given attribute pair, e.g., from the source class male to the target class female, the mean user vector for each class across the training data is computed and the steering vector is defined as the difference. During inference, the intervention is performed by adding a unit-normalized steering vector, injected through the write map, to the model's activations. The authors compare three interventions: contrastive activation addition computed on the distilled user vector, the uncompressed raw hidden state, and a random vector as the null control. Since the write map has orthonormal columns, this injection is norm-preserving, so the unit random vector is a magnitude-matched null baseline. For each attribute-value pair, 50 test items are sampled and the shift in answers is measured using the multiple-choice questions, reporting the flip rate: the fraction of predictions flipped across all multiple-choice layouts after steering. Aggregated over all 86 attribute pairs, this results in 4,300 steering evaluations per model. The scalar intervention strength is swept over a shared grid across all methods and the best flip rate per intervention method is reported.

Across all three models, the distilled user state preserves nearly all of the linearly decodable user information available in the raw hidden state, while providing a substantially stronger interface for causal intervention. For all three models, the gap in macro-F1 score between the raw hidden states and the user vector remains within 4%. Regarding causal interventions, steering on the user vector consistently outperforms steering on the raw hidden states. This difference is most pronounced in Llama-3.1-8B, where raw state steering achieves a 21% flip rate, whereas user vector steering successfully flips 78% of the predictions. Since all interventions share the same norm, this performance gap is driven purely by the learned steering direction. The low random baseline (13%) further confirms that unit-scale perturbations alone do not drive the flips. The advantage persists in OLMo-3-7B and Qwen3-8B, with user vector achieving flip rates 23% and 13% higher than the raw hidden states, respectively. Intervening on the user vector can qualitatively steer the model's responses in open-ended generation tasks without altering the prompt text, such as flipping a story's protagonist or a film recommendation.

The steering advantage is not explained merely by compressing beliefs into a low-dimensional space. Holding both the user vector and the target steering direction fixed, the learned write map is replaced with a random orthonormal projector. Across all three models, the learned map outperforms the random projector on the large majority of attribute-pair interventions. BSD therefore learns not only what user information to preserve, but also where to write it so that the resulting belief can causally influence the model.

Implicit versus explicit user representations

Previous work on user representations often studies implicit personalization by injecting user cues into the prompt, either through explicit attribute statements or stereotypical demographic associations. The authors compare BSD against such baselines, using controlled cues and recovering them from hidden representations. In contrast, BSD infers the model's beliefs about the user directly from natural conversational context, without injecting synthetic user attributes.

To compare the generalization between explicit cues and implicit beliefs, the authors cross-evaluate three sets of probes: those trained on distilled user vectors, those trained on the raw implicit beliefs hidden states, and those trained on the explicit-cue hidden states – across both the explicit cue test set and the inferred belief test set. The results reveal a sharp asymmetry in cross-distribution generalization. Probes trained on explicitly injected user cues perform almost perfectly when evaluated on similarly constructed explicit examples, but fail to transfer to user beliefs inferred from natural conversations. In contrast, probes trained on naturally inferred beliefs generalize substantially better across both settings. This suggests that representations induced by synthetic declarations or stereotypical cues need not coincide with the user representations that emerge naturally from conversation.

Inferred user intent modulates refusal

An analysis of the collinearity of the steering directions reveals that, across all evaluated models, the four safety attributes converge to a single shared axis with a mean cosine similarity of 0.91. The steering vectors for adversarial to benign, skeptical to trusting, low to high error tolerance, and hostile to conversational interaction style are highly aligned. This indicates that the model organizes representations for these attributes in a single direction. Furthermore, the authors hypothesize that this axis acts as a mechanism the model uses to classify users and adjust its safety-related behavior.

To test whether a user's position along this safety axis predicts model refusal, the authors extract user vectors generated by prompts from the forbidden question set and project them onto the User Intent direction. A linear classifier trained on those scalar projections can accurately predict whether the model will refuse the prompt: a repeated stratified cross-validation across 390 prompts achieves an AUC of 0.94 and an accuracy of 88% in predicting refusal.

Using a steering framework, the authors ask whether changing the model's belief about the user causally alters its safety behavior while the request itself remains fixed. Shifting the model's inferred user intent from adversarial toward benign substantially reduces refusal from 98% to 62%, even though the harmful request itself is unchanged. This effect cannot be explained by activation perturbation alone: magnitude-matched random steering has a substantially weaker effect, and steering an unrelated user belief, Gender, closely follows the random baseline. The safety effect is therefore specific to the model's inferred user intent rather than a generic consequence of editing its activations. Importantly, refusal decreases but does not disappear. The latent user model therefore modulates rather than solely determines the safety decision: the model conditions refusal jointly on what is being asked and on whom it believes it is interacting with.

Shared geometry across models

Transformer representation spaces are known to be anisotropic. In the distilled space, the representations are strongly concentrated around the centroid, with a mean cosine similarity of 0.89 to 0.92 across the three models, leaving only 20% that is user-specific. The centroid itself therefore represents the default user. Decoded through each model's probes, the three models default to a user who is male, European, middle-income, apolitical, emotionally neutral, and benign, decoded with high confidence only for benign intent and with lower confidence on the rest. This default is likely shaped by the corpus: since WildChat is raw ordinary conversations, it reflects the model's expected user over real-world interactions rather than a neutral prior.

Beyond the default, the authors ask whether the models agree on the variation around it, the user-specific residual that distinguishes one user from another. Across the full test set the three models are strongly aligned (mean pairwise Centered Kernel Alignment 0.75, range 0.73 to 0.79; permutation null 0.01). Fitting a single orthogonal Procrustes map between the two models' class centroids and applying it to transfer user vectors across models, the target model adopts the source's belief in 53–54% of cases on the disagreement set, well above the 30% random-rotation baseline and stable across all three directed pairs.

Limitations and future work

BSD requires a predefined attribute set and multiple-choice probes to define the distillation target, so the learned representation is shaped by the beliefs we choose to elicit and need not capture the model's full user state; beliefs may also depend on the elicitation procedure, despite multiple templates and option permutations. However, BSD needs no per-user labels: the attribute set determines what to ask about, while the model determines what it believes, and once trained, the user vector is extracted without probes. The authors study linear maps at a single layer in three similarly sized instruction-tuned models, leaving generality across architectures, scales, layers, and longer interactions open. Future work could pursue open-ended belief discovery, track how user states evolve over extended interactions, and test whether the compact state can serve as a persistent, editable memory across turns.

What this means for AI safety and practitioner implications

The most immediate practical implication is that a model's perception of its user is part of its safety mechanism. Inferred user intent both predicts refusal and causally changes it when edited, while unrelated user attributes do not reproduce the effect. BSD exposes a safety-relevant internal state that influences refusal independently of the request itself. For practitioners, BSD's read–write interface makes latent user models observable in a way that was previously difficult. The 58,000-conversation training corpus is moderately sized, and once the projector is trained, the user vector is extracted without additional probes. However, BSD requires a predefined attribute set and the learned representation is shaped by the beliefs we choose to elicit; the representation need not capture the model's full user state. Several attributes studied (e.g., gender, political orientation, income) are sensitive, and their coarse value sets follow prior work rather than endorse these categorizations. The extracted vectors should not be used to make decisions about individuals. While the white-box access and model-specific projector required for BSD's intervention represent a constrained setting, the result suggests that safety mechanisms should remain robust to perturbations of latent user representations.

The broader conceptual contribution is that implicit user models are not incidental features of individual systems, but structured, causally consequential representations that may generalize across models. The shared geometry across independently trained LLMs reveals common structure in how different models represent the people they interact with. This finding has implications for transfer learning and cross-model analysis of user behavior. The authors view the exposure of safety vulnerabilities as a motivating result: making latent states observable is necessary for audit and robustness, even though the same write access can also weaken safety behavior when the model's perceived user intent is manipulated.

Read the paper on arXiv

Read the paper on arXiv