progressive-head-schedule

Tag

Cards List
#progressive-head-schedule

Prism Transformer: Progressive Head Schedules for Hierarchical Attention Processing

arXiv cs.LG · 2026-06-29 Cached

The Prism Transformer replaces uniform multi-head attention with a progressive head schedule that increases head count across layers, enabling a local-to-global hierarchy without extra parameters or FLOPs. It consistently outperforms standard Transformers on language modeling and zero-shot benchmarks at 124M, 354M, and 757M scales.

0 favorites 0 likes
← Back to home

Submit Feedback