Full-bandwidth transformer
Summary
A new transformer variant, the full-bandwidth transformer, feeds back top-layer hidden states through a gated linear unit to improve reasoning and efficiency without altering the core architecture. Trained up to 400B tokens, it matches standard transformers trained with 1.5x more data while producing shorter reasoning traces.
View Cached Full Text
Cached at: 08/14/26, 03:26 AM
Paper page - Full-bandwidth transformer
Source: https://huggingface.co/papers/2608.08888
Abstract
Full-bandwidth transformers use latent feedback of top-layer hidden states to improve reasoning and efficiency without altering the core architecture.
Autoregressive transformerscompute along two axes: horizontally across generated tokens, and vertically through model depth.Dense attentiongives each token broad horizontal access to the past, but the vertical feedback channel between decoding steps remains narrow: only the sampled token returns to the bottom of the stack, while the top-layer hidden state is discarded. We introduce thefull-bandwidth transformer, which widens this channel withlatent feedback: at each decoding step, the previous top-layer hidden state is fused with the sampled token embedding through agated linear unitand fed back as the next input.Latent feedbacklets non-verbalized computation re-enter the stack with a renewed depth budget, while preserving the standard transformer architecture,KV cache, and language-modeling objective. To trainfull-bandwidth transformers without losing parallelteacher forcing, we use ascheduled multi-pass objectivethat introduceslatent feedbacklate in pretraining and mixes a small fraction of deeper feedback passes for stability. We train 1B-parameterfull-bandwidth transformers up to 400B tokens and find thatlatent feedbackimproves validation loss, 5-shot language-model evaluation, math and coding generation, and instruction-tuned performance. With negligible per-token decoding overhead,full-bandwidth transformers match or approach standard transformers trained with roughly 1.5times more tokens, and manage to produce shorter reasoning traces at equal or better accuracy.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2608\.08888
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.08888 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.08888 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.08888 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Microsoft's Full-bandwidth Transformers (26 minute read)
The paper introduces full-bandwidth transformers, which use latent feedback to enhance autoregressive models by allowing non-verbalized computation to re-enter the stack, improving performance with negligible decoding overhead.
A new transformer passes its hidden state to the next token instead of recomputing it (25 minute read)
The paper introduces LIFT (Latent Information Feedback Transformer), which propagates deep-layer hidden states across generation steps instead of relying solely on decoded tokens, using teacher-forced pretraining with states derived from an off-the-shelf LM's next-token distribution. Experiments on 135M–1B models show LIFT outperforms standard Transformers on language modeling, reasoning, and procedural tasks under token-matched budgets, with code and models released.
Variable-Width Transformers
Proposes a nonuniform width allocation transformer (hourglass shape) that outperforms uniform baselines in language modeling, reducing FLOPs and KV cache size.
ResBM: a new transformer-based architecture for low-bandwidth pipeline-parallel training, achieving 128× activation compression [R]
ResBM introduces a transformer-based architecture with residual encoder-decoder bottlenecks for pipeline-parallel training, achieving 128× activation compression while maintaining convergence. The work advances decentralized, internet-grade distributed training by reducing inter-stage communication overhead.
Transformer co-author validates post-transformer cost efficiency breakthrough
A 150M-parameter non-transformer architecture achieves state-of-the-art cost-efficiency on ARC-AGI-1, validated by Transformer co-author Łukasz Kaiser, suggesting that recurrent latent reasoning can replace brute-force scaling.