Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax version. How to initialize linear ViTs both efficiently and effectively still remains unclear. In this work, we explicitly ask: given that most foundation ViTs are built on the mainstream Softmax attention, can linear ViTs benefit from their pre-trained weights? Recent works on Attention Transfer show that attention is the effective transferable component between Softmax ViTs, suggesting attention alone suffices for such reuse. However, we find the opposite for Softmax-to-linear transfer. The attention weights are operator-specific: copying them barely helps, and is sometimes even worse than random initialization. Instead, the attention's token routing behavior can be recovered through distillation with a proper loss design, letting linear ViTs reduce the gap and even match Softmax ones. In contrast, the MLP weights, which carry the learned representation, are operator-agnostic: they can be transferred by simple direct copying, which already carries most of the benefit of the pre-trained weights. Thus, copying MLPs can serve as an effective foundation for Softmax-to-linear transfer: paired with the distilled attention, linear ViTs eventually close the remaining gap and even surpass Softmax ones. These findings hold consistently across various linear ViT variants, different model sizes, and diverse datasets. We hope this study deepens the understanding of reusing pre-trained weights across attention operators: copy what stays the same and distill what differs, to recover the benefit across the Softmax-to-linear boundary.

2. Document Verification & Archival Data
- Contributing Researchers: Huaiyuan Qin, Muli Yang, Gabriel James Goenawan, Shiqi Huang, Min Kass Chong, Wahyu Wiratama, Peng Hu, Chen Gong, Wu Liu, Xi Peng, Chun Jian Ho, Hongyuan Zhu
- Submission Date: September 28, 2026
- Full Preprint Document: Download Official PDF
- Permanent Archive Record: arXiv:2609.35745v1
3. Academic Citation Reference
Standard Reference (APA Format):
Huaiyuan Qin, et al. (2026). Copy the Same, Distill the Difference: Initializing Linear Vision Transformers. arXiv:2609.35745v1. https://arxiv.org/abs/2609.35745v1
Academic Field: Computer Vision | Document Identifier: arXiv:2609.35745v1
BibTeX Entry:
@article{arxiv_2609.35745v1,
author = {Huaiyuan Qin and Muli Yang and Gabriel James Goenawan and Shiqi Huang and Min Kass Chong and Wahyu Wiratama and Peng Hu and Chen Gong and Wu Liu and Xi Peng and Chun Jian Ho and Hongyuan Zhu},
title = {{Copy the Same, Distill the Difference: Initializing Linear Vision Transformers}},
journal = {arXiv preprint arXiv:2609.35745v1},
year = {2026},
url = {https://arxiv.org/abs/2609.35745v1}
}
Henüz yorum yapılmamış. İlk yorumu siz yapın!