Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers
Summary
This paper introduces SiPE, a lightweight method that injects syntactic priors from dependency parses into transformer positional embeddings, improving syntactic generalization (up to 10.3% on SyntaxGym) and language understanding (up to 8.2% on GLUE) without increasing inference cost.
View Cached Full Text
Cached at: 08/12/26, 08:21 AM
Paper page - Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers
Source: https://huggingface.co/papers/2608.06111
Abstract
SiPE integrates a lightweight syntactic prior from dependency parses into positional embeddings across transformer architectures, improving syntactic generalization and language understanding without altering self-attention or increasing inference cost.
Positional embeddings(PE) inTransformersencode token distance and order but are largely agnostic tosyntactic structure. We introduce Syntax-informedPositional Embeddings(SiPE), which learns a lightweight syntactic prior fromdependency parsesduring pretraining and injects it across all three dominant PE families (absolute, relative, rotary), for both encoders and decoders, leavingself-attentionand the rest of the architecture untouched. We isolate where and how the prior should enter the model, and find it depends on the architecture: for autoregressive decoders that use relative PE, the prior is strongest when coupled multiplicatively with the relative-position term of the attention score, outperforming injection into the input embeddings, intoself-attention, or into the positional and attention terms jointly---while for encoders it is best added directly to the input embeddings, composing with each encoder’s native positional mechanism. We find that models pre-trained withSiPEimprove on theSyntaxGymbenchmark by up to 10.3% while simultaneously reducingperplexityby 9.0% over a base model with no syntactic supervision---a metric nearly every existing syntax-injection method instead degrades. Crucially, these gains extend beyond syntactic generalization:SiPEalso improves real-world language understanding, raising scores on theGLUEbenchmark by up to 8.2% over a model trained without it. Unlike existing syntactic language models that marginalize over many parses at inference or discard syntax at runtime,SiPEconditions on a single parse, establishing a new Pareto frontier between syntactic supervision and inference cost.
View arXiv pageView PDFProject pageGitHub0Add to collection
Get this paper in your agent:
hf papers read 2608\.06111
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.06111 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.06111 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.06111 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Position Encoding in Transformers: From Absolute and Relative Methods to Rotary Position Embeddings and Long-Context Scaling
A technical survey on position encoding methods in Transformers, covering absolute and relative methods, RoPE, and long-context scaling techniques like Position Interpolation, NTK-aware scaling, YaRN, and LongRoPE.
Mitigating Position Bias in Transformers via Layer-Specific Positional Embedding Scaling
Introduces LPES, a layer-specific positional embedding scaling method that mitigates the 'lost-in-the-middle' problem in LLMs by assigning distinct scaling factors per layer using a genetic algorithm with Bézier curves, achieving up to 11.2% accuracy gain without fine-tuning or latency increase.
Syntax vs. Semantics: How Transformers Learn Deep Dependencies
This paper introduces a mechanistic framework analyzing transformer learning dynamics, identifying gradient starvation as a barrier to deep semantic dependencies and validating chain-of-thought strategies for effective learning.
Enhancing Transformer-based Routing by Encoding Distance via Relative Positional Encoding
This paper explores Relative Positional Encoding (RPE) as an additive bias in Transformer architectures to solve the Team Orienteering Problem, demonstrating consistent improvements in collected rewards and optimality gaps over vanilla Transformer architectures.
RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
This paper provides a theoretical proof that Rotary Positional Embeddings (RoPE) in Transformer-based language models lose their locality bias and ability to distinguish token order in long contexts, with attention scores becoming no better than random. The authors show that increasing the RoPE base trades off position vs. token distinction and that multi-head, multi-layer architectures cannot compensate for this fundamental limitation.