Tag
This paper introduces SiPE, a lightweight method that injects syntactic priors from dependency parses into transformer positional embeddings, improving syntactic generalization (up to 10.3% on SyntaxGym) and language understanding (up to 8.2% on GLUE) without increasing inference cost.
This paper introduces Bifocal Attention, which decouples positional encoding into geometric (standard RoPE) and spectral (learnable harmonic operators) modalities to address the 'Spectral Rigidity' of fixed RoPE, improving algorithmic generalization beyond the training window.