Tag
This paper introduces SiPE, a lightweight method that injects syntactic priors from dependency parses into transformer positional embeddings, improving syntactic generalization (up to 10.3% on SyntaxGym) and language understanding (up to 8.2% on GLUE) without increasing inference cost.
The paper proposes GiLT (Graph-Infused Layers Transformer Language Model), which improves syntactic generalization by modulating attention weights using features from dependency graphs constructed incrementally during token prediction, outperforming baselines while maintaining competitive perplexity.