Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers

Hugging Face Daily Papers Papers

Summary

This paper introduces SiPE, a lightweight method that injects syntactic priors from dependency parses into transformer positional embeddings, improving syntactic generalization (up to 10.3% on SyntaxGym) and language understanding (up to 8.2% on GLUE) without increasing inference cost.

Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to syntactic structure. We introduce Syntax-informed Positional Embeddings (SiPE), which learns a lightweight syntactic prior from dependency parses during pretraining and injects it across all three dominant PE families (absolute, relative, rotary), for both encoders and decoders, leaving self-attention and the rest of the architecture untouched. We isolate where and how the prior should enter the model, and find it depends on the architecture: for autoregressive decoders that use relative PE, the prior is strongest when coupled multiplicatively with the relative-position term of the attention score, outperforming injection into the input embeddings, into self-attention, or into the positional and attention terms jointly---while for encoders it is best added directly to the input embeddings, composing with each encoder's native positional mechanism. We find that models pre-trained with SiPE improve on the SyntaxGym benchmark by up to 10.3% while simultaneously reducing perplexity by 9.0% over a base model with no syntactic supervision---a metric nearly every existing syntax-injection method instead degrades. Crucially, these gains extend beyond syntactic generalization: SiPE also improves real-world language understanding, raising scores on the GLUE benchmark by up to 8.2% over a model trained without it. Unlike existing syntactic language models that marginalize over many parses at inference or discard syntax at runtime, SiPE conditions on a single parse, establishing a new Pareto frontier between syntactic supervision and inference cost.
Original Article
View Cached Full Text

Cached at: 08/12/26, 08:21 AM

Paper page - Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers

Source: https://huggingface.co/papers/2608.06111

Abstract

SiPE integrates a lightweight syntactic prior from dependency parses into positional embeddings across transformer architectures, improving syntactic generalization and language understanding without altering self-attention or increasing inference cost.

Positional embeddings(PE) inTransformersencode token distance and order but are largely agnostic tosyntactic structure. We introduce Syntax-informedPositional Embeddings(SiPE), which learns a lightweight syntactic prior fromdependency parsesduring pretraining and injects it across all three dominant PE families (absolute, relative, rotary), for both encoders and decoders, leavingself-attentionand the rest of the architecture untouched. We isolate where and how the prior should enter the model, and find it depends on the architecture: for autoregressive decoders that use relative PE, the prior is strongest when coupled multiplicatively with the relative-position term of the attention score, outperforming injection into the input embeddings, intoself-attention, or into the positional and attention terms jointly---while for encoders it is best added directly to the input embeddings, composing with each encoder’s native positional mechanism. We find that models pre-trained withSiPEimprove on theSyntaxGymbenchmark by up to 10.3% while simultaneously reducingperplexityby 9.0% over a base model with no syntactic supervision---a metric nearly every existing syntax-injection method instead degrades. Crucially, these gains extend beyond syntactic generalization:SiPEalso improves real-world language understanding, raising scores on theGLUEbenchmark by up to 8.2% over a model trained without it. Unlike existing syntactic language models that marginalize over many parses at inference or discard syntax at runtime,SiPEconditions on a single parse, establishing a new Pareto frontier between syntactic supervision and inference cost.

View arXiv pageView PDFProject pageGitHub0Add to collection

Get this paper in your agent:

hf papers read 2608\.06111

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2608.06111 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2608.06111 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2608.06111 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably

arXiv cs.CL

This paper provides a theoretical proof that Rotary Positional Embeddings (RoPE) in Transformer-based language models lose their locality bias and ability to distinguish token order in long contexts, with attention scores becoming no better than random. The authors show that increasing the RoPE base trades off position vs. token distinction and that multi-head, multi-layer architectures cannot compensate for this fundamental limitation.