2026-10-01
•
⏱ 2 dk okuma
•
en
Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace.
2026-10-01
•
⏱ 2 dk okuma
•
en
Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs.
2026-10-01
•
⏱ 3 dk okuma
•
en
Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax version.
2026-10-01
•
⏱ 2 dk okuma
•
en
Evaluating finance research agents requires rubrics that reflect expert standards and fix the values correct as of an information cutoff.
2026-10-01
•
⏱ 2 dk okuma
•
en
A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow.