The fundamental assumption that neural network inference can be made truly independent of training data has been called into question. Membership inference, the problem of determining whether a particular datum was part of a model's training set, has traditionally been addressed through statistical analysis of output probabilities. This paper argues that the execution footprint of a model—its interaction with hardware during inference—carries a data-dependent signature that survives even the most aggressive software-level constant-time protections. The authors demonstrate that transformer tokenization, performed once during training, permanently alters the memory access patterns the model exhibits at inference time. These alterations manifest in page table access patterns and translation lookaside buffers, creating a hardware-level side channel that reveals whether an input belongs to the training distribution. The practical consequence is a membership inference tool that requires no surrogate model, operates at low cost, and achieves substantially higher robustness than prior software-only approaches.
Hardware as a Data-Dependent Side Channel
Prior work on membership inference has focused on output distributions, loss values, or shadow model training. Each approach assumes that the model's computational graph executes identically regardless of its training history. This assumption is falsified by the present study, which performs a cycle-level examination of how large language models and vision transformers interact with modern microarchitecture components. The examination covers integrated accelerators, TLBs, and page table walks, across model scales that span several orders of magnitude.
The central finding is that the data distribution seen during training leaves a lasting imprint on inference-time memory access. The root cause traced by the authors is the tokenizer, which runs during training and maps vocabulary items to token IDs. This mapping determines, for each token position, which embedding vector must be fetched from the model's weight matrix. The embedding lookup is not a uniform operation; its memory address depends on the token ID, and the sequence of token IDs produced by the tokenizer depends on the training distribution the model internalized. During inference, when the model again encounters inputs drawn from versus outside its training distribution, the resulting sequence of vocabulary fetches differs. This difference propagates to the page table and TLB, where the pattern of page accesses and evictions diverges in a way that is both data-dependent and, crucially, input-dependent without involving any runtime branches or dynamic optimizations.
Root Cause: Tokenization-Induced Locality Shifts
The authors conduct a systematic root cause analysis, isolating the transformer's tokenization steps as the sole training-time operation that alters inference-time locality. During training, the tokenizer learns a mapping from tokens to IDs that reflects the statistical structure of the training corpus. This mapping is static and fixed after pre-training, but its consequence is ongoing: every inference begins with a tokenization step whose output sequence is a function of the training distribution.
The authors show that this sequence of token IDs changes the linear addresses accessed when fetching embedding vectors from the weight matrix. In a system with virtual memory, each such address maps to a page, and the pattern of page references determines TLB hit rates, page table walk frequency, and cache residency. The authors provide quantitative evidence that TLB miss rates and page table access counts differ significantly between in-distribution and out-of-distribution inputs, even for models explicitly trained with constant-time objectives and masked confidence outputs. The effect scales with model size, becoming more pronounced as the number of parameters and the width of the embedding matrices grow.
TranScope: A Microarchitecture Tool for Membership Detection
Building on the root cause analysis, the authors introduce TranScope, a microarchitecture-based tool for detecting membership information. TranScope requires no surrogate model, incurs minimal runtime overhead, and operates entirely from performance counters and TLB miss registers that are available on commodity hardware. The tool measures the pattern of page table accesses and TLB behavior during a single forward pass, then classifies the input as in-distribution or out-of-distribution based on learned thresholds that capture the characteristic access patterns induced by the model's tokenizer.
The performance results are striking. On the PETAL benchmark, the best previously reported AUC was 0.6. TranScope achieves an AUC of 0.9, a substantial improvement that places membership detection on hardware ground comparable to the best statistical methods, but with the practical advantages of requiring no shadow models, no gradient access, and no auxiliary training. The robustness of the signal across model sizes, tokenizer architectures, and dataset types suggests that this hardware footprint is a general property of transformer-based models, not a quirk of a particular architecture or pre-training corpus.
Implications for Privacy and Copyright Enforcement
The authors frame their findings as reopening hardware as both an opportunity and a new channel. As an opportunity, TranScope provides, for the first time, a hardware-based method for checking copyright violation. If a model is suspected of having been trained on proprietary data, auditors can present inputs that are known to be in versus out of the training distribution and measure the hardware's response. The reliability of this approach, without needing to retrain shadow models or query the model's parameters, makes it a practical tool for legal and compliance contexts.
As a new channel for membership inference, TranScope adds to the existing arsenal of statistical, gradient-based, and property inference methods. Its distinguishing feature is that it operates at the microarchitectural level, meaning it is largely orthogonal to software-level defenses. Defenses that mask output confidence or add noise to gradients do not necessarily affect page table access patterns or TLB behavior. This means that hardware-based detection can serve as a complement, and potentially a bypass, to software-only mitigation strategies.
Scope and Limitations of the Observation
The authors confirm that the effect holds across both language models and vision transformers, and across integrated as well as discrete accelerators. However, the signal is not universal in the sense that its magnitude varies with model architecture, tokenizer vocabulary size, and the specific microarchitecture features available for measurement. On some configurations, particularly very small models or those with aggressively flattened memory hierarchies, the AUC drops but remains above random guessing. The authors also note that the effect requires the tokenizer to be active during inference, which is the standard case but excludes tokenizer-free or embedding-lookup-shared configurations.
Another limitation is the physical access requirement. TranScope requires the ability to measure performance counters and TLB state during model execution, which implies either running the model on controlled hardware or obtaining performance data from a co-located execution environment. This places the method in a threat model that assumes some level of hardware access, though the authors argue that such access is plausible for many audit and compliance scenarios.
Technical Mechanism in Detail
The transformer's tokenization process maps each input token to a learned embedding vector. The location of this vector in the weight matrix is determined by the token ID, which is itself a function of the tokenizer's vocabulary and the input text. During inference, the model sequentially fetches embedding vectors for each token position. In a system with virtual memory, each fetch generates a memory request whose address is translated through the page table and TLB. The sequence of page references produced by these fetches is not random; it is patterned by the tokenizer's ID assignment, which reflects the statistical distribution of the training corpus.
When the input text resembles text from the training distribution, the tokenizer produces ID sequences that, on average, reference pages already warmed up by prior inferences on similar inputs. When the input is out-of-distribution, the ID sequence references different pages, causing different TLB entries to be used, evicted, or accessed. The resulting microarchitectural state—TLB hit/miss patterns, page table walk counts, cache line residency—differs systematically. TranScope captures this difference by counting TLB misses per page, tracking page table walk frequency, and computing a simple statistical feature vector that summarizes the access pattern. A threshold, learned on a small calibration set, then classifies the input.
Related Privacy Work and Orthogonal Defenses
Membership inference has been studied from many angles. Statistical methods examine the model's output distribution, particularly the confidence scores assigned to the correct class. Gradient-based methods query the model with crafted inputs and examine how loss changes with respect to training data proximity. Property inference attacks attempt to learn whether the model satisfies some global property, such as "was this data point included in training?" Each of these approaches has been met with defenses: output perturbation, gradient masking, and differential privacy training objectives, respectively.
TranScope is largely orthogonal to these defenses because it does not query the model's outputs or gradients. It measures hardware state that is computed regardless of the model's confidence masking or confidence calibration. A model that outputs uniform probabilities over classes will still exhibit different page table access patterns depending on whether its tokenizer was shaped by training data that included the input. This insensitivity to software-level defenses is both the tool's greatest practical value and its greatest policy concern.
Conclusion
The paper establishes that the data a model was trained on leaves a measurable, data-dependent signature in its inference-time hardware footprint. This signature persists even in models designed for constant-time execution and masked confidence, because its source is the tokenizer's mapping from tokens to embedding vectors—a mapping fixed during training but whose microarchitectural consequences play out on every inference. TranScope, the first microarchitecture-based membership detection tool, achieves an AUC of 0.9 on the PETAL benchmark, significantly surpassing the previous best of 0.6 and requiring no surrogate model or auxiliary training. The work redefines the privacy landscape for large models, reintroducing hardware as a domain both for new defensive opportunities, such as copyright verification, and for new inference channels, such as hardware-based membership testing. The authors' systematic root cause analysis, identifying tokenization-induced locality shifts as the mechanism, provides a clear target for future work on hardware-aware privacy guarantees and microarchitecture-level mitigations.
Read the paper on arXiv