Linear Vision Transformers (ViTs) aim to replace the softmax attention mechanism with a linear-complexity alternative, enabling more efficient token routing at scale. However, a persistent obstacle is that linear ViTs typically require training from scratch and underperform their softmax counterparts. The prevailing assumption has been that the attention mechanism itself is the primary carrier of pre-trained knowledge, so much prior work has attempted to transfer attention weights directly. This paper challenges that assumption, demonstrating that simply copying attention weights between softmax and linear operators is not merely ineffective but can actually hurt performance relative to random initialization. Instead, the authors show that the token routing behavior of attention must be recovered through distillation with a carefully designed loss, while the MLP sublayers—responsible for carrying learned representations—are operator-agnostic and transfer perfectly via direct copying. The central insight, succinctly framed as "copy what stays the same and distill what differs," provides a clear recipe for bridging the softmax-to-linear boundary.

Prior Work and the Initialization Gap

Foundation models in vision have largely been built on softmax attention, which scales quadratically with the number of tokens. Linear attention variants, such as Performer or Linformer, replace the softmax with an approximation that reduces complexity to linear time, making them attractive for long-sequence or high-resolution settings. Yet these architectures rarely inherit pre-trained weights from softmax ViTs; instead, they are trained from scratch, incurring significant compute and data costs. Prior work on attention transfer has shown that attention statistics can be matched between models, under the assumption that attention is the primary vector of transferable knowledge. The authors investigate whether this holds when moving from softmax to linear attention, and their findings reveal a more nuanced picture: the weight matrices themselves are operator-specific, and naively copying them does not preserve the original model's behavior.

The gap is particularly striking because the MLP components—feed-forward networks that process token representations—do not depend on the attention mechanism's mathematical form. Their weights are therefore portable across architectures. The paper's key observation is that while the MLPs can simply be copied, the attention module requires a distillation procedure that aligns the routing behavior, not just the parameter values. This division of labor—copy the stable parts, distill the variable parts—forms the backbone of the proposed initialization strategy.

Distillation of Token Routing Behavior

The distillation procedure central to this work is not a standard feature matching or KL-divergence loss on logits. Instead, the authors design a loss that encourages the linear attention's token routing to mimic that of the softmax teacher. The intuition is that softmax attention distributes weight across tokens based on soft probability scores, while linear attention uses a different kernelized similarity. By optimizing the linear ViT's attention outputs to match the softmax teacher's routing patterns—measured through a task-aware loss—the model learns to approximate the same effective token interactions, even though the underlying computations differ. This distillation step is where most of the accuracy recovery happens.

Critically, the distillation is not performed on the attention weights directly. Rather, the loss operates on the resulting token representations after attention aggregation. If the linear ViT's attended tokens produce representations close to those of the softmax teacher, the downstream MLPs—which are copied verbatim—can then operate on features that are already well-aligned. The paper provides empirical evidence that this two-step process: (1) copy MLP weights, (2) distill attention routing—yields linear ViTs that not only match but sometimes surpass softmax ViTs on the same downstream tasks.

Experimental Results Across Configurations

The authors evaluate their initialization method across a range of linear ViT variants, model sizes (from small to large), and datasets including ImageNet-1K and JFT-300M. In every setting, the combined approach of copied MLPs plus distilled attention closes the performance gap between linear and softmax ViTs. In many cases, the resulting linear ViT exceeds the softmax baseline after fine-tuning. The results are notably consistent: the method does not rely on accidental hyperparameter luck but reflects a structural property of how knowledge transfers across attention operators.

Metrics reported include top-1 image classification accuracy, with baselines ranging from randomly initialized linear ViTs to fully trained softmax ViTs. The paper shows that simply copying MLP weights from a softmax ViT already recovers a large fraction of the original model's performance—often within 2–3 points of the fully trained softmax baseline. Adding the distillation step then pushes the linear ViT over the top, frequently surpassing the softmax reference. The ablation studies confirm that each component contributes indispensably: removing the distillation step degrades performance sharply, while copying only the attention weights (without MLP transfer) yields marginal or negative gains.

Limitations and Trade-offs

The paper acknowledges that the distillation procedure adds a fine-tuning overhead compared to truly zero-initialization of a linear ViT. The lookahead depth supervision or iterative routing alignment, while effective, requires additional compute during the initialization phase. However, this cost is incurred only once, and the resulting model can then be fine-tuned or deployed at inference time with the same efficiency as any linear ViT. Another trade-off is that the distillation loss design is sensitive to the choice of teacher-student temperature and the specific token routing metric used; the authors note that different choices can modulate how much of the softmax behavior is recovered, though the broad pattern of results remains stable across reasonable settings.

Furthermore, the study primarily focuses on image classification benchmarks. While the authors assert that the same principles should extend to other vision tasks—such as object detection or semantic segmentation—the experimental evidence is currently confined to classification. Readers working on downstream tasks may need to validate that the copied-MLP + distilled-attention recipe generalizes beyond the imageNet regime.

Practical Impact for Developers

For a working developer the practical takeaway is straightforward: if you are adopting a linear attention ViT but have access to a pre-trained softmax model, do not attempt to copy the attention weights directly. Instead, copy the MLP layers as-is, and then run a distillation routine that aligns the linear attention's token routing to the softmax teacher's output. The paper provides code and ablations that make this process reproducible. The result is a linear ViT that starts from a much stronger checkpoint, reducing the amount of scratch training required to reach target accuracy.

This approach is especially valuable in resource-constrained environments where pre-trained softmax ViTs are available but linear variants are preferred for their linear scaling inference cost. By reusing the "same" (MLPs) and "distilling the difference" (attention routing), teams can transition to more efficient token routing without sacrificing the benefit of prior pre-training. The method also opens a path for hybrid models: one could imagine freezing the copied MLPs and only fine-tuning the distilled attention module for a target task, further reducing compute.

Conclusion

Copy the Same, Distill the Difference reframes the problem of transferring knowledge between attention operators by making a seemingly simple but underexplored observation: some components change when the operator changes, while others remain constant. By treating MLP weights as portable and attention behavior as something to be distilled, the authors produce linear ViTs that not only close the performance gap with softmax ViTs but occasionally surpass them. The findings are robust across model sizes, variant architectures, and datasets, and the resulting initialization strategy is implementable with modest engineering effort. For anyone building or adapting vision transformers, this work provides both a conceptual framework and a practical recipe for crossing the softmax-to-linear boundary.

Read the paper on arXiv