Telescopic Language Models

One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point. We train a Telescopic Language Model (TLM) to be that continuum: a nested-capacity Transformer supervised by stochastic prefix supervision with a full anchor. At every step, one randomly truncated prefix of the capacity axis is trained against the full next-token target, alongside one full-capacity pass, so the trained artifact is a valid language model at every depth. Two forward-backward passes per step, no architectural change, nothing extra at inference. Fixed-exit suites such as Matryoshka Language Model Suites (MLMS) occupy one point in this design space, and the point has a cost: supervising only a few fixed exits leaves the nested model at chance level everywhere else (perplexity 10^2-10^5 in our baselines). On a 200M proxy suite (20B FineWeb-Edu tokens, identical data stream for all methods), a single TLM run is a valid language model at every one of its twenty layer prefixes, in perplexity and on perplexity-sensitive downstream tasks, reducing the area under the quality-budget curve by 43-44% relative to the fixed-exit suites while matching them at full capacity, at ~12% lower GPU cost per run. The prefix sampling density is a dial: concentrating it on a few depths recovers fixed-exit quality there at the price of the continuum, so the operating points become a training-time choice rather than an architectural one. These results indicate that the training objective, not the nesting itself, is what makes a model elastic.

2. Document Verification & Archival Data

  • Contributing Researchers: Zhilin Guo, Boqiao Zhang, Hakan Aktas, Kyle Fogarty, Nursena Koprucu Aslan, Wenzhao Li, Canberk Baykal, Albert Miao, Siyu Hong, Yixiao Liu, Adam Wu, Ashish Kumar Singh, Sakar Khattar, Chenliang Zhou, Weihao Xia, Cristina Nader Vasconcelos, Cengiz Oztireli
  • Submission Date: September 28, 2026
  • Full Preprint Document: Download Official PDF
  • Permanent Archive Record: arXiv:2609.35769v1

3. Academic Citation Reference

Standard Reference (APA Format):

Zhilin Guo, et al. (2026). Telescopic Language Models. arXiv:2609.35769v1. https://arxiv.org/abs/2609.35769v1

Academic Field: Computation and Language (NLP) | Document Identifier: arXiv:2609.35769v1


BibTeX Entry:

Kod
@article{arxiv_2609.35769v1,
  author    = {Zhilin Guo and Boqiao Zhang and Hakan Aktas and Kyle Fogarty and Nursena Koprucu Aslan and Wenzhao Li and Canberk Baykal and Albert Miao and Siyu Hong and Yixiao Liu and Adam Wu and Ashish Kumar Singh and Sakar Khattar and Chenliang Zhou and Weihao Xia and Cristina Nader Vasconcelos and Cengiz Oztireli},
  title     = {{Telescopic Language Models}},
  journal   = {arXiv preprint arXiv:2609.35769v1},
  year      = {2026},
  url       = {https://arxiv.org/abs/2609.35769v1}
}

Yorumlar (0)

Henüz yorum yapılmamış. İlk yorumu siz yapın!

Yorum Bırakın