Directional Contextual Representations for Dependency Relations: Why Cross-Direction Pairing Fails
Summary
This paper investigates why cross-direction pairing in bidirectional LSTMs underperforms same-direction pairing for dependency relation-type classification, using frozen-trunk diagnostics to analyze representational redundancy and directional information decay.
View Cached Full Text
Cached at: 08/24/26, 04:24 AM
# Directional Contextual Representations for Dependency Relations: Why Cross-Direction Pairing Fails
Source: [https://arxiv.org/html/2608.20647](https://arxiv.org/html/2608.20647)
Sai Krishna Arthanari,JaeHyeong Chang,Chengzhe Sun,Siwei LyuAffiliation:Institute for Artificial Intelligence and Data Science \(IAD\), University at Buffalo Buffalo, NY, USA arthanarisaikrishna@gmail\.com, jchang46@buffalo\.edu, csun22@buffalo\.edu, siweilyu@buffalo\.edu
###### Abstract
Splitting a bidirectional LSTM’s contextual representation into a forward\-onlyFiF\_\{i\}\(strictly a function of tokens1\.\.i1\.\.i\) and a backward\-onlyBiB\_\{i\}\(strictly a function of tokensi\.\.ni\.\.n\) beats either alone and beats a fused self\-attention representation for dependency relation\-type classification\. But a specific, natural extension of this idea – pairing a token’s forward state against a*candidate*’s backward state \(“cross\-direction” pairing,FiF\_\{i\}vs\.BjB\_\{j\}\) – consistently*underperforms*same\-direction pairing, and the penalty*grows*, not shrinks, with token distance, both paired\-bootstrap significant\. We diagnose why using a frozen\-trunk methodology: architectural information leakage between directions is impossible by construction \(a single\-layer BiLSTM, verified by code inspection\); 93% of the same\-vs\-cross gap survives freezing the trunk and training only fresh heads, ruling out training\-co\-adaptation as the primary cause; linear regression shows partial representational redundancy betweenFiF\_\{i\}andBiB\_\{i\}\(R2=0\.324R^\{2\}\{=\}0\.324vs\.0\.0280\.028for a shuffled control\) and a linear probe shows partial anticipatory encoding of upcoming tokens inFiF\_\{i\}\(36\.5% vs\. 17\.2% majority baseline\) – real effects, but neither alone, nor combined, cleanly explains the full gap\. Extended frozen\-trunk diagnostics \(a positional probe and a distance\-decay probe\) show directional information is genuinely stored but not exactly positioned, and propagates only a few tokens before decaying to baseline – consistent with, and mechanistically underneath, the distance\-growth finding\. We validate the core claim three ways: a parameter\-matched Transformer backbone \(a fresh end\-to\-end BiLSTM still significantly outperforms it, Cohen’sddshrinking from 0\.216 to 0\.085 but staying significant\); a second UD English treebank of a substantially different genre \(GUM\), on which both the directional\-splitting and cross\-direction\-failure findings replicate closely; and a literature search that found no prior work performing this specific cross\-direction ablation, reported plainly as a novelty claim rather than assumed\.
###### Index Terms:
dependency parsing, self\-attention, bidirectional LSTM, directional representations, Universal Dependencies, empirical evaluation
## IIntroduction
Contextual encoders that process a sequence in both directions – bidirectional LSTMs, and, less explicitly, self\-attention – are ubiquitous in NLP, but the specific contribution of*directional*structure to relational tasks like dependency parsing is rarely isolated\. Most architectures fuse forward and backward \(or all\-to\-all\) information into a single representation before any relational comparison happens, so it is not obvious whether keeping directions separate through the comparison step itself matters, and if so, in which specific combinations\. This paper asks that question narrowly and empirically: does directional structure in a contextual representation carry recoverable relational information, and specifically, is comparing a token’s state from*one*direction against a candidate’s state from the*opposite*direction informative, or not? TheO\(n2\)O\(n^\{2\}\)cost of unrestricted self\-attention\[[1](https://arxiv.org/html/2608.20647#bib.bib1)\]is useful background for why representation structure is worth studying carefully, but reducing that cost is not this paper’s question – we make no adaptive\-computation or efficiency claim here\.
We use a single\-layer bidirectional LSTM trunk, whose two directions architecturally guarantee a forward\-only representationFiF\_\{i\}\(strictly a function of tokens1\.\.i1\.\.i\) and a backward\-only representationBiB\_\{i\}\(strictly a function of tokensi\.\.ni\.\.n\), with no attention or mixing layer between the recurrence and the split – a property we verify, not assume\. On Universal Dependencies English EWT, we test a sequence of specific, falsifiable claims and preview them here so the results read as confirmation or refutation, not an undifferentiated list of numbers:\(1\)does splitting intoFF/BBbeat either alone or a fused self\-attention baseline?\(2\)does*cross\-direction*pairing – comparingFiF\_\{i\}to a candidate’sBjB\_\{j\}– add signal beyond same\-direction pairing?\(3\)if it does not, why – architectural leakage, representational redundancy, a training artifact, or something else?\(4\)does the resulting picture hold under real scrutiny: a parameter\-matched alternative backbone, a second dataset, and a check of the relevant literature?
- •\(1\) Yes, unfusedF\+BF\{\+\}Bwins\.CombiningFFandBBbeats either alone and beats a fused self\-attention representation for relation\-type classification, though bootstrap validation shows this specific margin is statistically robust but a small standardized effect\.
- •\(2\) No, cross\-direction pairing loses, and loses more with distance\.FiF\_\{i\}\-vs\-BjB\_\{j\}pairing is consistently the weakest pairwise construction tested, and its penalty relative to same\-direction pairing grows monotonically with token distance – the opposite of a “masked by short\-range dominance” account – both bootstrap\-significant\.
- •\(3\) Partially diagnosed, not fully explained\.Architectural leakage is ruled out by construction; the gap is mostly not a training\-co\-adaptation artifact \(93% survives freezing the trunk\); representational redundancy and anticipatory encoding are both real but partial effects, and extended positional/distance probes show information is stored but imprecisely positioned and short\-range – together consistent with, but not a complete mechanistic account of, the failure\.
- •\(4\) Holds under scrutiny\.A parameter\-matched Transformer backbone narrows but does not close a fresh BiLSTM’s advantage; a second UD treebank of a different genre \(GUM\) replicates both the directional\-splitting and cross\-direction\-failure findings closely; and a literature search found no prior work on this specific cross\-direction ablation\.
## IIRelated Work
Self\-attention\.Vaswani et al\.\[[1](https://arxiv.org/html/2608.20647#bib.bib1)\]introduced scaled dot\-product self\-attention, where every token attends to every other via one fused, all\-to\-all computation – useful background for why keeping directions separate is a non\-trivial design choice, though this paper does not address attention’s computational cost\.Bidirectional representations\.Peters et al\.\[[2](https://arxiv.org/html/2608.20647#bib.bib2)\]\(ELMo\) build contextualized word representations from a bidirectional LSTM’s internal states, combining forward and backward directions for transfer; we instead keepFiF\_\{i\}/BiB\_\{i\}separate through the pairwise\-comparison step itself, to isolate which directional combinations carry relational signal – directly comparable prior art for the splitting question, though ELMo does not test cross\-direction pairwise comparison specifically\.Biaffine parsing\.Dozat and Manning\[[3](https://arxiv.org/html/2608.20647#bib.bib3)\]score candidate dependency arcs with a biaffine function of two tokens’ representations, trained with a per\-sentence softmax over candidate heads; our backbone\-comparison arc scorer \(Section[V](https://arxiv.org/html/2608.20647#S5)\) follows this same structured\-prediction recipe, on an explicit, hand\-specified feature rather than a learned biaffine transform, keeping the pairwise representation interpretable\.Universal Dependencies\.Nivre et al\.\[[4](https://arxiv.org/html/2608.20647#bib.bib4)\]describe UD; we use English EWT as the primary testbed and English GUM, a substantially different genre, for cross\-dataset validation\.
Cross\-direction pairing: a literature check\.We searched specifically for prior work comparing a token’s forward\-direction state against another token’s backward\-direction state \(or symmetric variants\) in a bidirectional recurrent representation, and for prior probing work establishing an asymmetry between what forward and backward directions predict about neighboring tokens\. We found no directly relevant prior work under this or related framings\. We report the cross\-direction pairwise ablation in this paper as, to our knowledge, novel, while noting plainly that this search was not exhaustive\.
## IIIMethodology
We deliberately use a single\-layer bidirectional LSTM, not a Transformer, as the trunk throughout the core ablation \(Sections[V](https://arxiv.org/html/2608.20647#S5)A–C\), stated explicitly rather than left implicit\. The requirement is thatFiF\_\{i\}/BiB\_\{i\}be provably, architecturally disjoint – a single\-layer BiLSTM gives this for free, verifiable by inspection\. A Transformer does not: self\-attention mixes every position and direction from layer one, so extracted “forward”/“backward” views would need independent verification against the very contamination this study isolates\. We use the BiLSTM as a clean\-room diagnostic instrument, not a smaller Transformer stand\-in; a fresh Transformer backbone is used only as an external validation point \(Section[V](https://arxiv.org/html/2608.20647#S5)D\), not as an alternative substrate for theFF/BB\-splitting ablation itself \(Section[VII](https://arxiv.org/html/2608.20647#S7)\)\.
### III\-AForward, backward, and positional representations
For a sentence ofnntokens we build a shared trunk: word embeddings \(dimension 100, vocabulary from training tokens with minimum frequency 2, plus<unk\>/<pad\>\) concatenated with UPOS embeddings \(dimension 32\), a 132\-dimensional per\-token input, fed through a single\-layer bidirectional LSTM \(hidden size 128 per direction,nn\.LSTM\(\.\.\., bidirectional=True\), nonum\_layersoverride\)\. The forward\-direction hidden stateFi∈ℝ128F\_\{i\}\\in\\mathbb\{R\}^\{128\}and backward\-direction hidden stateBi∈ℝ128B\_\{i\}\\in\\mathbb\{R\}^\{128\}are sliced directly from the raw bidirectional output with no attention or mixing layer in between:
Fi=LSTM→\(x1,…,xi\),Bi=LSTM←\(xi,…,xn\)\.F\_\{i\}=\\mathrm\{LSTM\}\_\{\\rightarrow\}\(x\_\{1\},\\ldots,x\_\{i\}\),\\qquad B\_\{i\}=\\mathrm\{LSTM\}\_\{\\leftarrow\}\(x\_\{i\},\\ldots,x\_\{n\}\)\.Because the LSTM is single\-layer andFiF\_\{i\}/BiB\_\{i\}are read directly off its two output halves,FiF\_\{i\}is by construction a strict function of tokens1\.\.i1\.\.ionly, andBiB\_\{i\}of tokensi\.\.ni\.\.nonly – no mechanism exists for information from the “wrong” side to leak into either representation, which we rely on directly in the diagnostic analysis \(Section[V](https://arxiv.org/html/2608.20647#S5)C\)\.Pi∈ℝ64P\_\{i\}\\in\\mathbb\{R\}^\{64\}is the classic sinusoidal positional encoding, a fixed function of absolute indexii:Pi\[2k\]=sin\(i/100002k/64\)P\_\{i\}\[2k\]=\\sin\(i/10000^\{2k/64\}\),Pi\[2k\+1\]=cos\(i/100002k/64\)P\_\{i\}\[2k\{\+\}1\]=\\cos\(i/10000^\{2k/64\}\)\. Figure[1](https://arxiv.org/html/2608.20647#S3.F1)summarizes the trunk\.
x1x\_\{1\}x2x\_\{2\}⋯\\cdotsxix\_\{i\}⋯\\cdotsxnx\_\{n\}embedder\(word\+UPOS\)1\-layer BiLSTMFiF\_\{i\}forward outputBiB\_\{i\}backward outputPiP\_\{i\}sinusoidal \(fixed\)Hi=concat\(Fi,Bi,Pi\)H\_\{i\}=\\mathrm\{concat\}\(F\_\{i\},B\_\{i\},P\_\{i\}\)
Fig\. 1:Trunk architecture\. A single\-layer BiLSTM over token embeddings exposesFiF\_\{i\}andBiB\_\{i\}directly, with no attention or mixing layer between the LSTM and the split\.PiP\_\{i\}is computed independently of the LSTM\.Hi=concat\(Fi,Bi,Pi\)H\_\{i\}=\\mathrm\{concat\}\(F\_\{i\},B\_\{i\},P\_\{i\}\)is the per\-token diagnostic representation used by the positional and distance\-decay probes \(Section[V](https://arxiv.org/html/2608.20647#S5)C\)\.
### III\-BPairwise representations tested
For a candidate pair\(i,j\)\(i,j\)\(an arbitrary ordered token pair, or a gold head\-dependent pair for Section[V](https://arxiv.org/html/2608.20647#S5)A\) we tested the following pairwise feature constructionsRijR\_\{ij\}, each fed to an identically\-shaped classifier head \(single hidden layer, ReLU, dropout 0\.2, linear output\) so comparisons isolate the representation, not classifier capacity:
Rijf\_only\\displaystyle R^\{\\mathrm\{f\\\_only\}\}\_\{ij\}=\[Fi,Fj,Fi⊙Fj,\|Fi−Fj\|\]\\displaystyle=\[F\_\{i\},F\_\{j\},F\_\{i\}\{\\odot\}F\_\{j\},\|F\_\{i\}\{\-\}F\_\{j\}\|\]Rijb\_only\\displaystyle R^\{\\mathrm\{b\\\_only\}\}\_\{ij\}=\[Bi,Bj,Bi⊙Bj,\|Bi−Bj\|\]\\displaystyle=\[B\_\{i\},B\_\{j\},B\_\{i\}\{\\odot\}B\_\{j\},\|B\_\{i\}\{\-\}B\_\{j\}\|\]Rijf\_plus\_b\\displaystyle R^\{\\mathrm\{f\\\_plus\\\_b\}\}\_\{ij\}=\[Fi,Bi,Fj,Bj,Fi⊙Bj,Bi⊙Fj,\\displaystyle=\[F\_\{i\},B\_\{i\},F\_\{j\},B\_\{j\},F\_\{i\}\{\\odot\}B\_\{j\},B\_\{i\}\{\\odot\}F\_\{j\},OPEN\|Fi−Fj\|,\|Bi−Bj\|\]\\displaystyle\\hskip 8\.50012pt\|F\_\{i\}\{\-\}F\_\{j\}\|,\|B\_\{i\}\{\-\}B\_\{j\}\|\]Rijfull\_ref\\displaystyle R^\{\\mathrm\{full\\\_ref\}\}\_\{ij\}=Rijf\_plus\_b⊕\(Pi−Pj\)\\displaystyle=R^\{\\mathrm\{f\\\_plus\\\_b\}\}\_\{ij\}\\oplus\(P\_\{i\}\-P\_\{j\}\)Rijsame\_f\\displaystyle R^\{\\mathrm\{same\\\_f\}\}\_\{ij\}=Rijf\_only\(binary\-task naming\)\\displaystyle=R^\{\\mathrm\{f\\\_only\}\}\_\{ij\}\\hskip 8\.50012pt\\text\{\(binary\-task naming\)\}Rijsame\_b\\displaystyle R^\{\\mathrm\{same\\\_b\}\}\_\{ij\}=Rijb\_only\(binary\-task naming\)\\displaystyle=R^\{\\mathrm\{b\\\_only\}\}\_\{ij\}\\hskip 8\.50012pt\\text\{\(binary\-task naming\)\}Rijcross\_fb\\displaystyle R^\{\\mathrm\{cross\\\_fb\}\}\_\{ij\}=\[Fi,Bj,Fi⊙Bj,\|Fi−Bj\|\]\\displaystyle=\[F\_\{i\},B\_\{j\},F\_\{i\}\{\\odot\}B\_\{j\},\|F\_\{i\}\{\-\}B\_\{j\}\|\]Rijcross\_bf\\displaystyle R^\{\\mathrm\{cross\\\_bf\}\}\_\{ij\}=\[Bi,Fj,Bi⊙Fj,\|Bi−Fj\|\]\\displaystyle=\[B\_\{i\},F\_\{j\},B\_\{i\}\{\\odot\}F\_\{j\},\|B\_\{i\}\{\-\}F\_\{j\}\|\]Rijcombined\_e\\displaystyle R^\{\\mathrm\{combined\\\_e\}\}\_\{ij\}=\[Fi,Bj,Bi,Fj,Pi−Pj\]\\displaystyle=\[F\_\{i\},B\_\{j\},B\_\{i\},F\_\{j\},P\_\{i\}\{\-\}P\_\{j\}\]Rijff\_plus\_fb\_concat\\displaystyle R^\{\\mathrm\{ff\\\_plus\\\_fb\\\_concat\}\}\_\{ij\}=Rijsame\_f⊕Rijcross\_fb\\displaystyle=R^\{\\mathrm\{same\\\_f\}\}\_\{ij\}\\oplus R^\{\\mathrm\{cross\\\_fb\}\}\_\{ij\}Rijresidual\_fb\_ff\\displaystyle R^\{\\mathrm\{residual\\\_fb\\\_ff\}\}\_\{ij\}=Rijcross\_fb−Rijsame\_f\\displaystyle=R^\{\\mathrm\{cross\\\_fb\}\}\_\{ij\}\-R^\{\\mathrm\{same\\\_f\}\}\_\{ij\}where⊙\\odotis elementwise product,⊕\\oplusconcatenation\.Rfull\_refR^\{full\\\_ref\}\(f\_plus\_b\_plus\_pf\\\_plus\\\_b\\\_plus\\\_p\) is the strongest variant in Section[V](https://arxiv.org/html/2608.20647#S5)A and is reused unchanged as the feature formula for the full\-pairwise arc scorer in the backbone comparison \(Section[V](https://arxiv.org/html/2608.20647#S5)D\)\. InRresidual\_fb\_ffR^\{residual\\\_fb\\\_ff\}the first 128 dims \(Fi−FiF\_\{i\}\-F\_\{i\}\) cancel to zero by construction; the remaining 384 carry the signal\. As a fused\-representation baseline we also trained a small \(2\-layer, 4\-head,dmodel=256d\_\{model\}\{=\}256\) Transformer encoder over the same embeddings \(sinusoidal position added elementwise, standard practice\), producing one fusedHiH\_\{i\}per token with no forward/backward split:Rijstandard\_attention=\[Hi,Hj,Hi⊙Hj,\|Hi−Hj\|\]R^\{standard\\\_attention\}\_\{ij\}=\[H\_\{i\},H\_\{j\},H\_\{i\}\{\\odot\}H\_\{j\},\|H\_\{i\}\{\-\}H\_\{j\}\|\]\.
### III\-CPer\-token diagnostic representation
To probe what a token’s own state encodes about position and neighboring tokens – independent of any specific comparison pair – we use
Hi=concat\(Fi,Bi,Pi\)∈ℝ320\.H\_\{i\}=\\mathrm\{concat\}\(F\_\{i\},B\_\{i\},P\_\{i\}\)\\in\\mathbb\{R\}^\{320\}\.HiH\_\{i\}is used here solely as the input to two frozen\-trunk diagnostic probes \(Section[V](https://arxiv.org/html/2608.20647#S5)C\): a positional probe \(doesHiH\_\{i\}, or itsFiF\_\{i\}/BiB\_\{i\}/Fi\+BiF\_\{i\}\{\+\}B\_\{i\}sub\-parts, encode absolute position, and how exactly\) and a distance\-decay probe \(how far does directional information about specific neighboring tokens propagate\)\. These probes characterize what the representations store; we make no claim here about usingHiH\_\{i\}for selective or adaptive computation\.
### III\-DBackbone comparison: BiLSTM vs\. Transformer, parameter\-matched
To test whether theFF/BB\-splitting advantage \(Section[V](https://arxiv.org/html/2608.20647#S5)A, a classification task\) extends to a genuine structured\-prediction task, we trained full\-pairwise arc scorers for unlabeled dependency head\-finding:score\(i,j\)=MLP\(Rijfull\_ref\)\\mathrm\{score\}\(i,j\)=\\mathrm\{MLP\}\(R^\{full\\\_ref\}\_\{ij\}\), trained end\-to\-end \(trunk and scorer optimized jointly\) with a per\-sentence softmax cross\-entropy over candidate headsjjin the same sentence, following the same structured\-prediction recipe as\[[3](https://arxiv.org/html/2608.20647#bib.bib3)\]\. We compare three backbones under this identical recipe: the BiLSTM trunk above; a Transformer backbone \(thestandard\_attentiontrunk from Section[V](https://arxiv.org/html/2608.20647#S5)A’s fused baseline\); and, since that Transformer has∼\\sim52% more parameters than the BiLSTM – a real confound for a backbone comparison – a*parameter\-matched*Transformer \(dmodel=144d\_\{model\}\{=\}144, 2 layers, 4 heads,dimff=288\\dim\_\{ff\}\{=\}288\) reduced until its total parameter count sits within 3% of the BiLSTM’s\.
## IVExperimental Setup
All experiments use Universal Dependencies English EWT \(en\_ewt\-ud\-\{train,dev,test\} \.conllu\) as the primary testbed, with English GUM \(a substantially different genre: academic, fiction, how\-to, news, interview, and travel\-guide text, vs\. EWT’s informal web/blog/email/review text\) for cross\-dataset validation \(Section[V](https://arxiv.org/html/2608.20647#S5)E\)\. Both are parsed with a minimal hand\-written CoNLL\-U parser skipping comment lines and multiword/empty\-node IDs \(\-/\.in the ID\), retaining per token its form, UPOS, integer head index, and DEPREL, reduced to its base label \(before any:\); the 15 most frequent training\-split base labels become classes, the rest map toother\(16 classes\)\. Arcs withHEAD==0\(virtual root\) are excluded throughout\. EWT splits: 12,544/2,001/2,077 sentences, 192,034/23,147/23,017 arcs\. GUM splits: 11,314/1,575/1,464 sentences, 188,909/26,933 train/test arcs\.
Training protocol, held fixed across the study for comparability: AdamW, learning rate10−310^\{\-3\}, weight decay10−510^\{\-5\}; batch size 32 sentences; 8 epochs; classifier head shapeLinear\(pair\_dim→256\)→ReLU→Dropout\(0\.2\)→Linear\(256→num\_classes\)\\mathrm\{Linear\}\(\\mathrm\{pair\\\_dim\}\\rightarrow 256\)\\rightarrow\\mathrm\{ReLU\}\\rightarrow\\mathrm\{Dropout\}\(0\.2\)\\rightarrow\\mathrm\{Linear\}\(256\\rightarrow\\mathrm\{num\\\_classes\}\); seeds 42–44 as the default \(3 seeds per configuration\), extended to seeds 42–46 \(5 seeds\) for the two most load\-bearing claims – the relation\-classification ablation \(Section[V](https://arxiv.org/html/2608.20647#S5)A\) and the backbone comparison \(Section[V](https://arxiv.org/html/2608.20647#S5)D\)\. Hardware: a single NVIDIA RTX 4090 \(24GB\)\. The binary edge\-existence task \(Sections[V](https://arxiv.org/html/2608.20647#S5)B–C\) uses a fixed 1:1 positive/negative dataset \(gold arcs vs\. randomly sampled non\-arc index pairs per sentence, negative\-sampling seed 42, reused unchanged across variants and training seeds so comparisons isolate model variance, not data variance\)\.
## VResults
Four questions, traced in order: \(A\) doesFF/BBsplitting beat fused? \(B–C\) does cross\-direction pairing add signal, and why does it fail? \(D–E\) does this hold under real scrutiny?
### A\. Does splittingFF/BBbeat either alone or fused?
Table[I](https://arxiv.org/html/2608.20647#S5.T1)reports dependency relation\-type classification \(16\-way, gold arcs\) test accuracy and macro\-F1 for each pairwise representation \(single seed 42\), plus a 5\-seed \(42–46\) replication of four variants confirming the single\-seed figures\.
TABLE I:Relation\-type classification, test set\. Single\-seed columns are the primary run; 5\-seed column \(mean±\\pmstd, seeds 42–46\) confirms it\.Combining unfusedFFandBBgives the largest jump over either direction alone, beating fused self\-attention by∼\\sim0\.7 points with∼\\sim35% fewer parameters\. Paired bootstrap \(seed 42, 5000 resamples,n=23,017n\{=\}23\{,\}017\) confirmsf\_plus\_b\_plus\_pbeatsf\_only\(\+0\.0080, 95% CI\(\+0\.0059,\+0\.0101\)\(\+0\.0059,\+0\.0101\), Cohen’sd=0\.049d\{=\}0\.049\),b\_only\(\+0\.0106, CI\(\+0\.0083,\+0\.0127\)\(\+0\.0083,\+0\.0127\),d=0\.064d\{=\}0\.064\), andstandard\_attention\(\+0\.0076, CI\(\+0\.0054,\+0\.0097\)\(\+0\.0054,\+0\.0097\),d=0\.045d\{=\}0\.045\) – all significant, all small standardized effects\. This establishes the starting point: directional splitting works\. The rest of the paper asks the sharper question of*which*directional combinations carry the signal\.
### B\. Does cross\-direction pairing add signal?
Table[II](https://arxiv.org/html/2608.20647#S5.T2)reports binary edge\-existence detection \(is there a dependency arc betweeniiandjjat all\), mean±\\pmstd over 3 seeds, on a 1:1 balanced gold\-arc\-vs\.\-random\-pair dataset\.
TABLE II:Binary edge existence, test set, mean±\\pmstd over 3 seeds\.Cross\-direction pairing is not an additional source of signal – it is the clear*loser*:cross\_fbtrailssame\_f/same\_bby∼\\sim3 points andcross\_bftrails by∼\\sim7–8 points, gaps 30–80×\\timeslarger than seed\-to\-seed noise \(std≤0\.0025\\mathrm\{std\}\\leq 0\.0025throughout\)\. Paired bootstrap \(seed 42, sentence\-independent example resampling, 5000 iterations\) confirmscross\_fbtrailssame\_fby−\-0\.0308 \(95% CI\(−0\.0336,−0\.0279\)\(\-0\.0336,\-0\.0279\), Cohen’sd=−0\.100d\{=\}\-0\.100\) – significant, small standardized effect\. The best variant overall iscombined\_e, a raw concatenation with no products or absolute differences, not a crafted interaction term\.
Does this weakness concentrate at short range, where a same\-direction shortcut might dominate the aggregate, masking a real long\-range cross\-direction advantage? Table[III](https://arxiv.org/html/2608.20647#S5.T3)tests this directly by token distance\|i−j\|\|i\-j\|\.
TABLE III:cross\_fb−\-same\_f accuracy/F1 gap by token distance \(test, mean over 3 seeds\)\.The cross\-direction penalty grows monotonically with distance – nearly 7×\\timeslarger in F1 at 21\+ than at 1–2 – the opposite of the “masked by short\-range dominance” account\. Figure[2](https://arxiv.org/html/2608.20647#S5.F2)plots both gaps\. Paired bootstrap on the difference of gaps \(gap at 21\+ minus gap at 1–2, stratified resampling within each bucket, 5000 iterations\) gives−\-0\.0371 \(95% CI\(−0\.0516,−0\.0227\)\(\-0\.0516,\-0\.0227\)\) – the growth trend itself is statistically robust, not just numerically monotonic\.
Fig\. 2:cross\_fb−\-same\_f accuracy and F1 gap by token\-distance bucket \(test set, mean over 3 seeds\)\. Both gaps are monotonically negative and grow with distance – the opposite of the “masked by short\-range dominance” prediction\.
### C\. Why does cross\-direction pairing fail?
We probe four candidate explanations on one frozen trunk \(combined\_e, seed 42, zero further training\)\.\(a\) Architectural leakageis ruled out by code inspection alone \(Section[V](https://arxiv.org/html/2608.20647#S5): the trunk is single\-layer with no mixing before theFF/BBsplit; see also Methodology\)\.\(b\) Representational redundancy:linear regressionFi→BiF\_\{i\}\\to B\_\{i\}achievesR2=0\.324R^\{2\}\{=\}0\.324out\-of\-sample vs\.R2=0\.028R^\{2\}\{=\}0\.028for a shuffled\-partner control \(11\.5×\\timeshigher, but far from theR2→1R^\{2\}\{\\to\}1full collapse would require\)\.\(c\) Training\-dynamics artifact:fresh classifier heads trained on the frozen trunk reproduce93%of the end\-to\-end same\_f\-vs\-cross\_fb gap \(2\.96 of 3\.18 points; Table[IV](https://arxiv.org/html/2608.20647#S5.T4)\), largely ruling out co\-adaptation as the primary cause\.\(d\) Anticipatory encoding:FiF\_\{i\}predicts the*next*token’s UPOS at 36\.5% accuracy \(majority baseline 17\.2%, same\-token ceiling 98\.1%\) – real but partial anticipatory signal, and not, on its own, an explanation for why cross\-direction pairing is actively*worse*rather than merely no\-better\.
TABLE IV:Frozen\-trunk vs\. end\-to\-end same\_f−\-cross\_fb accuracy gap\.Extended diagnostic 1: is position stored, and is it exact?Does the LSTM recurrence already implicitly encode absolute position, makingPiP\_\{i\}redundant, or doesPiP\_\{i\}add real recoverable precision? We probe position \(decile bucket, 10 classes; and normalizedi/\(n−1\)i/\(n\{\-\}1\), regression\) from four sources on the frozen trunk \(Table[V](https://arxiv.org/html/2608.20647#S5.T5)\), each with both an MLP head \(as used throughout this section\) and a pure linear head \(no hidden layer\), to test whether the encoded information is linearly accessible\.
TABLE V:Position\-probe bucket accuracy, MLP vs\. linear head, mean over 3 seeds \(majority baseline 0\.1367\)\. MAE columns are for the MLP head’s normalized\-position regression \(median baseline MAE 0\.2718\)\.Neither answer is clean\. The recurrence already encodes real position without any explicit signal \(F\+BF\{\+\}Breaches∼\\sim3×\\timesthe majority baseline with noPiP\_\{i\}at all under the MLP head\), soPiP\_\{i\}is not filling a total void – butPiP\_\{i\}is not redundant either:HiH\_\{i\}jumps a further∼\\sim19 points overF\+BF\{\+\}Balone \(MLP head\)\. The linear head always trails the MLP, sometimes narrowly \(direction\-adjacent sources\) and sometimes by a lot \(HiH\_\{i\}:−\-15\.4 points, the largest linear\-vs\-MLP gap we observed\) – the position information exists but is not fully linearly accessible, especially oncePiP\_\{i\}’s sinusoidal encoding is involved, since decoding it into a discrete bucket is an inherently nonlinear operation\.
Extended diagnostic 2: how far does directional information propagate?We predictUPOS\(i\+Δ\)\\mathrm\{UPOS\}\(i\{\+\}\\Delta\)fromFiF\_\{i\}alone \(symmetricallyUPOS\(i−Δ\)\\mathrm\{UPOS\}\(i\{\-\}\\Delta\)fromBiB\_\{i\}\) forΔ∈\{1,2,3,5,10,20\}\\Delta\\in\\\{1,2,3,5,10,20\\\}, against a majority baseline and a same\-token direct\-observation ceiling \(Fi\+Δ→UPOS\(i\+Δ\)F\_\{i\+\\Delta\}\\to\\mathrm\{UPOS\}\(i\{\+\}\\Delta\),≈\\approx0\.98 throughout, confirming the probe architecture is not the bottleneck; Table[VI](https://arxiv.org/html/2608.20647#S5.T6)\)\.
TABLE VI:Distance\-decay probe accuracy vs\. majority baseline, frozen trunk, seed 42\.Both directions decay sharply, hitting their majority baseline byΔ=5\\Delta\{=\}5\(forward\) orΔ=3\\Delta\{=\}3\(backward\) – directional information about a specific token’s category propagates only a handful of positions before vanishing into noise, a concrete mechanistic number underneath part B’s task\-level distance\-growth finding: cross\-direction pairing has less and less real signal to draw on as\|i−j\|\|i\-j\|grows, exactly where its penalty grows largest\.
### D\. Does this hold under a fresh, parameter\-matched backbone?
Table[VII](https://arxiv.org/html/2608.20647#S5.T7)compares three full\-pairwise arc scorers trained end\-to\-end \(trunk and scorer jointly\) for real unlabeled dependency head\-finding \(UAS\), 5 seeds each\.
TABLE VII:Full\-pairwise arc scorer backbone comparison, mean±\\pmstd over 5 seeds\.Parameter\-matching narrows the gap \(test UAS\+3\.04\+3\.04points unmatched→\\to\+2\.60\+2\.60points matched\) but does not close it\. Table[VIII](https://arxiv.org/html/2608.20647#S5.T8)reports paired bootstrap \(seed 42, sentence\-level resampling – whole sentences, not tokens – 5000 iterations\) for both comparisons\.
TABLE VIII:Paired bootstrap, BiLSTM−\-Transformer UAS \(seed 42, sentence\-level\)\.Both CIs exclude zero: at near\-identical parameter count, the BiLSTM still significantly outperforms the Transformer on this task, though the standardized effect shrinks from moderate \(d=0\.216d\{=\}0\.216\) to small \(d=0\.085d\{=\}0\.085\)\. This extends part A’s classification\-only finding to a real structured\-prediction task under a fair parameter budget\.
### E\. Does this hold on a different genre?
Table[IX](https://arxiv.org/html/2608.20647#S5.T9)replicates the two core claims \(parts A and B\) on UD English GUM \(academic, fiction, how\-to, news, interview, and travel\-guide text – substantially different from EWT’s informal web/blog/email/review text; 11,314/1,575/1,464 train/dev/test sentences, 188,909/26,933 train/test arcs, comparable scale to EWT, identical protocol and seeds\)\.
TABLE IX:EWT vs\. GUM, core claims \(test set\)\.Both claims replicate closely on a substantially different genre: bidirectional splitting still beats fused attention \(GUM’s full 5\-variant ordering matches EWT’s:f\_plus\_b\_plus\_p\>\>f\_plus\_b\>\>f\_only\>\>standard\_attention\>\>b\_only\), and cross\-direction pairing still loses by a comparable, if slightly larger, margin\. A third UD treebank in a different*language*was not attempted, named here as an open question rather than skipped silently \(Section[VII](https://arxiv.org/html/2608.20647#S7)\)\.
## VIDiscussion
What replicated\.Directional splitting is genuinely informative and generalizes across genre \(part A, part E\): unfusedF\+BF\{\+\}Bbeats either alone and beats fused self\-attention, and the same ordering holds on GUM\. Cross\-direction pairing’s failure is equally robust: consistently the weakest construction, its penalty growing \(not concentrating\) with distance, both bootstrap\-confirmed and replicating on GUM\. The backbone comparison \(part D\) shows this is not an artifact of the specific backbone either – a fresh, end\-to\-end, parameter\-matched Transformer still loses to the BiLSTM on a real structured\-prediction task\.
What remains only partially explained\.Part C rules out architectural leakage cleanly and training\-co\-adaptation largely \(93% of the gap survives trunk\-freezing\), but the positive mechanism is incomplete: representational redundancy \(R2=0\.324R^\{2\}\{=\}0\.324\) and anticipatory encoding \(36\.5% vs\. 17\.2%\) are both real, above\-control effects, yet neither in isolation – nor, by construction, obviously in combination – accounts for why cross\-direction pairing is actively*worse*than same\-direction pairing rather than merely no\-better\. The extended positional and distance\-decay probes sharpen the picture \(information is stored but imprecisely positioned, and propagates only 3–5 tokens\) without fully closing this gap\. We consider the mechanism partially, not fully, diagnosed, and say so plainly rather than overclaim a tidy causal story\.
On novelty\.We found no prior work performing this specific cross\-direction pairwise ablation, and report it as such – a claim we hold loosely, since our literature search was targeted, not exhaustive\.
## VIILimitations
Two datasets, both English \(EWT and GUM\); a third treebank in a different language was not attempted, an explicit open question rather than a silent gap\. The coreFF/BB\-splitting ablation \(parts A–C\) uses a single architecture family, a single\-layer bidirectional LSTM, chosen because it makesFiF\_\{i\}/BiB\_\{i\}separation architecturally exact; the Transformer backbone in part D is used only as an external validation point on a downstream task, not as an alternative substrate for testing directional splitting itself – whether a Transformer\-derived forward/backward split \(e\.g\. via causal masking or paired unidirectional models\) would show the same cross\-direction failure is untested\. The literature search underlying our novelty claim \(cross\-direction pairwise ablation\) was targeted at specific query framings, not exhaustive\. The mechanism behind cross\-direction pairing’s failure remains partially, not fully, diagnosed: redundancy and anticipatory encoding are real contributing factors but do not, individually or combined, fully account for the size of the gap\.
## VIIIConclusion
Splitting a bidirectional LSTM’s representation into forward\-only and backward\-only parts is genuinely useful for dependency relation discovery, beating both single\-direction and fused alternatives, and this generalizes across genre\. But a specific, natural extension – cross\-direction pairwise comparison – consistently fails, and fails more as token distance grows, both statistically robust findings\. We diagnosed this failure using a frozen\-trunk methodology: architectural leakage is impossible by construction, the gap is mostly not a training artifact, and partial representational redundancy and anticipatory encoding are real contributing factors that do not, individually, fully explain it – a mechanism we report as partially, not fully, understood\. The core claims hold under a parameter\-matched alternative backbone and a substantially different\-genre dataset, and, to our knowledge, this specific cross\-direction ablation has not been reported before\. We present this as a focused, honestly\-scoped empirical result rather than a general theory of directional representations\.
## References
- \[1\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, L\. Kaiser, and I\. Polosukhin, “Attention is all you need,” in*Advances in Neural Information Processing Systems 30 \(NeurIPS 2017\)*, Long Beach, CA, USA, Dec\. 2017, pp\. 5998–6008\.
- \[2\]M\. E\. Peters, M\. Neumann, M\. Iyyer, M\. Gardner, C\. Clark, K\. Lee, and L\. Zettlemoyer, “Deep contextualized word representations,” in*Proc\. 2018 Conf\. North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(NAACL\-HLT 2018\)*, Vol\. 1 \(Long Papers\), New Orleans, LA, USA, 2018, pp\. 2227–2237\.
- \[3\]T\. Dozat and C\. D\. Manning, “Deep biaffine attention for neural dependency parsing,” in*Proc\. 5th International Conference on Learning Representations \(ICLR 2017\)*, Toulon, France, Apr\. 2017\.
- \[4\]J\. Nivre, M\.\-C\. de Marneffe, F\. Ginter, J\. Hajič, C\. D\. Manning, S\. Pyysalo, S\. Schuster, F\. Tyers, and D\. Zeman, “Universal Dependencies v2: An evergrowing multilingual treebank collection,” in*Proc\. 12th Language Resources and Evaluation Conference \(LREC 2020\)*, Marseille, France, 2020, pp\. 4034–4043\.Similar Articles
Correct When Paired, Wrong When Split: Decoupling and Editing Modality-Specific Neurons in MLLMs
This paper identifies and addresses the 'editing decoupling failure' in Multimodal LLMs, where knowledge updates via multimodal inputs fail to generalize to unimodal queries. The authors propose DECODE, a method to disentangle and localize modality-specific neurons for more effective knowledge editing.
Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks
This paper identifies a blind spot in long-context LLM reasoning benchmarks: they fail to control task position within the context, allowing positional failures to go undetected. The authors propose Context Rot Evaluation (CRE) to systematically vary task position, filler content, and context length, revealing severe accuracy drops for some models when reasoning tasks are placed in the middle of long contexts.
Vernier: Probing Representational Misalignment Behind Lexical Gaps in Causal Reasoning
This paper investigates why instruction-tuned language models give different answers to causal reasoning questions when variable names are replaced with placeholders, finding that the issue stems from representational misalignment rather than information loss. The authors introduce Vernier, a method using paired-view weight updates and mechanism inspection to reveal that answer-relevant content is still present in the placeholder view but misaligned.
Compositional Reasoning in Language Models under Reinforcement Learning Post-Training
This paper proposes a dependency-graph framework to formalize compositional reasoning in language models and evaluates the impact of reinforcement learning post-training, finding an asymmetry where composed-skill training transfers more readily to decomposed tasks than vice versa.
Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment
This paper introduces the Hybrid Reward-Cyclic (HRC) model and Dynamic Self-Play Preference Optimization (DSPPO) to address the cyclic nature of human preferences in LLM alignment, achieving improved performance over Bradley-Terry and General Preference Model baselines.