Paragraph Boundaries Are Not White Space:Compression Depth as the Signature of Hierarchical Structure

Hugging Face Daily Papers Papers

Summary

This paper introduces hierarchical rotary positional encoding (hRoPE) to represent text hierarchy and uses compression depth as a signature of paragraph structure, showing it correlates with genuine hierarchical organization better than other metrics.

Standard positional encodings represent position as a one-dimensional reading-order coordinate, but reading order alone does not determine hierarchical textual structure. We use a hierarchical rotary positional encoding (hRoPE) that represents paragraph, sentence, and token indices as separate channels, hold the token sequence fixed, intervene on the paragraph coordinate p1, and measure cross-paragraph attention with a token-distance-exact estimator. Attention is compressed relative to a token-distance-matched baseline in every corpus, but compression alone is not diagnostic of true structure: an architecturally identical channel with density-matched random labels is compressed too, more shallowly. What distinguishes real structure is the depth of compression, which is greater and corpus-dependent while the control's is not. Comparing eight corpus-only quantities across three constructs (lexical persistence, paragraph length, embedding-based coherence), none fully reproduces the cross-corpus ordering of depth, though embedding-based coherence comes closest. Compression depth, not its location, is the reproducible signature of genuine paragraph structure in our setting.
Original Article
View Cached Full Text

Cached at: 09/28/26, 12:04 PM

Paper page - Paragraph Boundaries Are Not White Space:Compression Depth as the Signature of Hierarchical Structure

Source: https://huggingface.co/papers/2609.23551

Abstract

Standardpositionalencodingsrepresentpositionasaone-dimensionalreading-ordercoordinate,butreadingorderalonedoesnotdeterminehierarchicaltextualstructure.Weuseahierarchicalrotarypositionalencoding(hRoPE)thatrepresentsparagraph,sentence,andtokenindicesasseparatechannels,holdthetokensequencefixed,interveneontheparagraphcoordinatep1,andmeasurecross-paragraphattentionwithatoken-distance-exactestimator.Attentioniscompressedrelativetoatoken-distance-matchedbaselineineverycorpus,butcompressionaloneisnotdiagnosticoftruestructure:anarchitecturallyidenticalchannelwithdensity-matchedrandomlabelsiscompressedtoo,moreshallowly.Whatdistinguishesrealstructureisthedepthofcompression,whichisgreaterandcorpus-dependentwhilethecontrol’sisnot.Comparingeightcorpus-onlyquantitiesacrossthreeconstructs(lexicalpersistence,paragraphlength,embedding-basedcoherence),nonefullyreproducesthecross-corpusorderingofdepth,thoughembedding-basedcoherencecomesclosest.Compressiondepth,notitslocation,isthereproduciblesignatureofgenuineparagraphstructureinoursetting.

View arXiv pageView PDFGitHub0Add to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.23551 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.23551 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.23551 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Text Distance from Nested and Hierarchical Repetitions: A Compression-Based Perspective

arXiv cs.CL

This paper presents a new method for structural sequence analysis using the Ladderpath approach to extract nested and hierarchical repetitions, defining three distance measures that outperform gzip-based NCD and BERT in out-of-distribution and few-shot text classification tasks, offering a lightweight and interpretable alternative.