DanLing NestedTensor: 用于深度学习的可组合多变张量

arXiv cs.LG 论文

摘要

DanLing NestedTensor 提出了一种用于可组合多变张量的 PyTorch 张量抽象,减少了可变大小深度学习输入中的填充浪费,并在 A100 GPU 上实现了显著加速(例如 2.74×–3.39×)和内存效率提升。

arXiv:2609.30379v1 Announce Type: new Abstract: Variable-size inputs are common in deep learning, but dense batching allocates a shared envelope and spends computation on padding. The cost multiplies across varying axes: an explicit pair state allocates $BN_{\max}^2$ positions instead of $\sum_i N_i^2$. Packing removes that waste, but composing packed operations still requires the logical axes and sample boundaries a flat buffer no longer exposes. We present DanLing NestedTensor, a PyTorch tensor abstraction that makes multi-ragged structure a property of the tensor itself. Packed values carry tensor-backed partitions and logical dimension order, so broadcasting creates ragged axes, feature transformations retain them, and reductions consume them. The same representation carries through autograd and both eager and compiled execution. On an A100, the geometric-mean speedup over same-mode padding is 2.74$\times$ eager and 3.39$\times$ compiled across four BERT scales, and 1.97$\times$ eager across four FCN backbones. A four-block Pairformer-style workload runs 2.40-4.32$\times$ faster than a padded reference using native PyTorch kernels across square length regimes in eager execution, with peak allocation falling from 38.08 to 5.41 GiB on its high-variation batch. The tensor interface lets model code built from its supported operators compose efficient variable-size computation without managing offsets at any call site. Code will be released publicly upon publication.
查看原文
查看缓存全文

缓存时间: 2026/09/29 09:35

# DanLing NestedTensor: Composable Multi-Ragged Tensors for Deep Learning
Source: [https://arxiv.org/html/2609.30379](https://arxiv.org/html/2609.30379)
###### Abstract

Variable\-size inputs are common in deep learning, but dense batching allocates a shared envelope and spends computation on padding\. The cost multiplies across varying axes: an explicit pair state allocatesB​Nmax2BN\_\{\\max\}^\{2\}positions instead of∑iNi2\\sum\_\{i\}N\_\{i\}^\{2\}\. Packing removes that waste, but composing packed operations still requires the logical axes and sample boundaries a flat buffer no longer exposes\. We present DanLing NestedTensor, a PyTorch tensor abstraction that makes multi\-ragged structure a property of the tensor itself\. Packed values carry tensor\-backed partitions and logical dimension order, so broadcasting creates ragged axes, feature transformations retain them, and reductions consume them\. The same representation carries through autograd and both eager and compiled execution\. On an A100, the geometric\-mean speedup over same\-mode padding is 2\.74×\\timeseager and 3\.39×\\timescompiled across four BERT scales, and 1\.97×\\timeseager across four FCN backbones\. A four\-block Pairformer\-style workload runs 2\.40–4\.32×\\timesfaster than a padded reference using native PyTorch kernels across square length regimes in eager execution, with peak allocation falling from 38\.08 to 5\.41 GiB on its high\-variation batch\. The tensor interface lets model code built from its supported operators compose efficient variable\-size computation without managing offsets at any call site\. Code is available at[https://github\.com/ZhiyuanChen/DanLing](https://github.com/ZhiyuanChen/DanLing)\.

Figure 1:Speed and memory efficiency relative to same\-mode padded execution\.One panel per workload; one point per model scale and mode, against that cell’s own padded baseline at\(1,1\)\(1,1\)\. Shape and colour identify the method, fill the mode \(hollow eager, solid compiled\)\. Both axes use steady\-state execution \(Section[4\.1](https://arxiv.org/html/2609.30379#S4.SS1)\), so the coordinates share one measurement window; Appendix Table[3](https://arxiv.org/html/2609.30379#A2.T3)reports compilation counts and full\-trace cost\. Inline text gives DanLing’s geometric mean per mode\. All completed paired measurements are shown; other outcomes use the codes of Appendix[C](https://arxiv.org/html/2609.30379#A3)\.## 1Introduction

Data vary in size; dense tensor batches do not\. Sequence packing already handles the one\-dimensional case, and for language models it is close to sufficient\. Much of what people train is not one\-dimensional: a protein structure model refines anNi×NiN\_\{i\}\\times N\_\{i\}residue\-pair state through dozens of blocks\([Jumper et al\., 2021](https://arxiv.org/html/2609.30379#bib.bib15);[Abramson et al\., 2024](https://arxiv.org/html/2609.30379#bib.bib1)\), and vision models are trained at variable resolution and aspect ratio\([Dehghani et al\., 2023](https://arxiv.org/html/2609.30379#bib.bib8)\)\. Convolution must keep height and width separate to address neighbourhoods within each image\. Padding places these examples inside a common envelope, consuming memory and computation at empty positions\([Fegade et al\., 2022](https://arxiv.org/html/2609.30379#bib.bib13)\)\. The cost multiplies across varying axes: a pair representation over a length\-NiN\_\{i\}sample holdsNi2N\_\{i\}^\{2\}entries, so a padded batch allocatesB​Nmax2BN\_\{\\max\}^\{2\}of them rather than∑iNi2\\sum\_\{i\}N\_\{i\}^\{2\}; a sample at half the maximum length wastes three quarters of its block rather than half\. On our Pairformer\-style workload padded allocation barely moves: 34\.76–38\.09 GiB whether lengths are near\-uniform or badly skewed \(Table[2](https://arxiv.org/html/2609.30379#S4.T2)\)\.

Collecting that saving is harder than allocating a smaller buffer, because the pair state is not where the program ends\. A single–pair–single block broadcasts two single representations into a pair tensor, a residual feature transformation updates it, and a reduction over one pair axis returns singles\. Packing the valid pair cells removes the padded storage, but the flat buffer no longer shows which cells share a sample and a row\. The reduction needs that grouping, and its result must carry a structure the next operator can accept\. When this structure is maintained outside the tensor, model code must supply indexing and reconstruction logic at each structural transition\.

Existing approaches address different parts of this problem\. Bucketing groups examples of similar size, which reduces padding but ties batch composition to input geometry and can change the order and statistics of optimisation\([Morishita et al\., 2017](https://arxiv.org/html/2609.30379#bib.bib23);[Kocmi and Bojar, 2017](https://arxiv.org/html/2609.30379#bib.bib16)\)\. Packing keeps the chosen batch, and variable\-length attention kernels execute attention on its packed storage\([Dao et al\., 2022](https://arxiv.org/html/2609.30379#bib.bib6);[Dao, 2024](https://arxiv.org/html/2609.30379#bib.bib7)\)\. Neither specifies how logical axes and sample boundaries are maintained when a model constructs, transforms and reduces an explicit pair state, and PyTorch’storch\.jaggedadmits only one ragged dimension\([PyTorch contributors, 2026c](https://arxiv.org/html/2609.30379#bib.bib27)\)\.

DanLing NestedTensor treats this bookkeeping the way ordinary tensors treat shape and dtype: as state the tensor carries, not state the model rebuilds\. The packed payload travels with exact element sizes, hierarchical partitions and logical dimension order, and each handler updates that structure as it computes, so autograd and compilation consume the same representation\. Figure[1](https://arxiv.org/html/2609.30379#S0.F1)summarises the speed and memory this yields; composition is the sharper test\. A four\-block Pairformer\-style application, which creates, transforms and reduces two ragged pair axes in every block, compiles and trains under this representation with default compiler settings, while the padded and hand\-packed implementations we tested do not—although the simpler block compiles for all three\.

This paper makes three contributions\.

- •A representation that composes\.Multiple varying logical axes are carried together, including non\-leading axes and rectangular pair states, without conflating them with the flattened storage axis\. Each handler consults and updates those axes as it already consults shape—broadcasting creates one, feature transformations preserve it, reductions consume it—and returns a structure the next operator can use, so a complete single–pair–single program composes without offset bookkeeping at any call site\. Handlers of this kind cover 404 distinct operators across PyTorch’s two dispatch protocols \(Sections[3\.1](https://arxiv.org/html/2609.30379#S3.SS1)and[3\.2](https://arxiv.org/html/2609.30379#S3.SS2)\)\.
- •Differentiation and compilation across the representation boundary\.Bridges in both directions keep the gradient connected where a structured wrapper meets its packed payload, so a compiled block can return singles and pairs to an eager loss and still receive gradients through changes of rank and layout\. Compilation fixes the structural schema at each program point while sizes and partitions stay runtime tensor inputs, so supported length\-dynamic programs reuse one compiled graph rather than specialising against each length vector \(Section[3\.4](https://arxiv.org/html/2609.30379#S3.SS4), Table[3](https://arxiv.org/html/2609.30379#A2.T3)\)\.
- •Efficiency, and where it comes from\.Four BERT scales give a 3\.39×\\timescompiled geometric\-mean speedup over padding and FCN\-ResNet50 drops from 28\.38 to 12\.25 GiB; a four\-block Pairformer\-style workload runs 4\.32×\\timesfaster than a padded reference using native PyTorch kernels on its high\-variation batch\. Against explicit packing, a matched\-granularity control separates cost from strategy: at equal call granularity the tensor interface adds at most 1\.60% latency, while batch\-level segmented execution runs 1\.53–3\.94×\\timesfaster than per\-sample calls \(Sections[4\.2](https://arxiv.org/html/2609.30379#S4.SS2)and[4\.4](https://arxiv.org/html/2609.30379#S4.SS4)\)\.

## 2Related Work

#### Ragged representations\.

TensorFlow RaggedTensor, Awkward Array and FBGEMM jagged operators represent variable\-size data through numerical buffers and structural metadata\([TensorFlow contributors, 2026](https://arxiv.org/html/2609.30379#bib.bib31);[Pivarski et al\., 2020](https://arxiv.org/html/2609.30379#bib.bib24);[FBGEMM contributors, 2026](https://arxiv.org/html/2609.30379#bib.bib12)\)\. PyTorch’s nativetorch\.jaggedcombines a tensor interface with packed storage, differentiation and compilation for its supported operations; its documented layout admits one ragged dimension\([PyTorch contributors, 2026c](https://arxiv.org/html/2609.30379#bib.bib27)\)\.

#### Ragged and dynamic compilation\.

CoRa provides a programming interface and compiler for ragged computation, using techniques including selective padding and operation splitting\([Fegade et al\., 2022](https://arxiv.org/html/2609.30379#bib.bib13)\)\. PyTorch 2 integrates graph capture, differentiation and backend compilation, while Relax makes symbolic shapes explicit across computational graphs and tensor programs\([Ansel et al\., 2024](https://arxiv.org/html/2609.30379#bib.bib2);[Lai et al\., 2025](https://arxiv.org/html/2609.30379#bib.bib18)\)\. NestedTensor integrates its representation with PyTorch, exposing changing sizes and partitions as tensor data while preserving the structural schema of each program point\.

#### Packing and specialised kernels\.

Sequence packing and NaViT retain independent\-example computation while reducing padding in language and vision workloads\([Krell et al\., 2021](https://arxiv.org/html/2609.30379#bib.bib17);[Dehghani et al\., 2023](https://arxiv.org/html/2609.30379#bib.bib8)\)\. FlashAttention and FlashAttention\-2 optimise attention execution, and FlexAttention provides a programmable fused attention interface\([Dao et al\., 2022](https://arxiv.org/html/2609.30379#bib.bib6);[Dao, 2024](https://arxiv.org/html/2609.30379#bib.bib7);[Dong et al\., 2025](https://arxiv.org/html/2609.30379#bib.bib10)\)\. NestedTensor reuses these execution paths while keeping projections, normalisation and residual operations packed around them\. KeOps evaluates formula\-defined pairwise computations lazily\([Charlier et al\., 2021](https://arxiv.org/html/2609.30379#bib.bib5)\), whereas our pair states persist across multiple feature transformations before a structural reduction\.

## 3NestedTensor: Representation and Execution

A single–pair–single block illustrates how NestedTensor carries logical structure through successive operations\. For nonempty inputsXi∈ℝNi×CX\_\{i\}\\in\\mathbb\{R\}^\{N\_\{i\}\\times C\}, the block computes

Pi\[a,b,:\]\\displaystyle P\_\{i\}\[a,b,:\]=uθ\(Xi\[a,:\]\)\+vθ\(Xi\[b,:\]\),\\displaystyle=u\_\{\\theta\}\(X\_\{i\}\[a,:\]\)\+v\_\{\\theta\}\(X\_\{i\}\[b,:\]\),\(1\)P~i\\displaystyle\\widetilde\{P\}\_\{i\}=Pi\+hθ​\(LNCp⁡\(Pi\)\),\\displaystyle=P\_\{i\}\+h\_\{\\theta\}\\\!\\left\(\\operatorname\{LN\}\_\{C\_\{p\}\}\(P\_\{i\}\)\\right\),X~i\[a,:\]\\displaystyle\\widetilde\{X\}\_\{i\}\[a,:\]=Xi\[a,:\]\+gθ\(1Ni∑b=0Ni−1P~i\[a,b,:\]\)\.\\displaystyle=X\_\{i\}\[a,:\]\+g\_\{\\theta\}\\\!\\left\(\\frac\{1\}\{N\_\{i\}\}\\sum\_\{b=0\}^\{N\_\{i\}\-1\}\\widetilde\{P\}\_\{i\}\[a,b,:\]\\right\)\.Hereuθ,vθu\_\{\\theta\},v\_\{\\theta\}project to pair widthCpC\_\{p\},hθh\_\{\\theta\}is a SwiGLU\([Shazeer, 2020](https://arxiv.org/html/2609.30379#bib.bib30)\)feature update, andgθg\_\{\\theta\}maps the reduced features back to widthCC\. The block returns both updated singles and pairs, so the pair state is model state that later operations consume, not a score matrix that exists only inside a fused attention call\. Figure[2](https://arxiv.org/html/2609.30379#S3.F2)draws what the tensor holds at each stage of the block for a two\-sample batch, with the exact packed shapes and offsets, and Figure[3](https://arxiv.org/html/2609.30379#S3.F3)gives the measured implementation \(Appendix[A\.1](https://arxiv.org/html/2609.30379#A1.SS1)\)\.

### 3\.1Packed Representation of Variable\-Size Batches

#### Problem setting\.

A batch𝒳=\(X0,…,XB−1\)\\mathcal\{X\}=\(X\_\{0\},\\ldots,X\_\{B\-1\}\)contains individually dense tensors of common rankDD; raggedness is variation between samples, not irregularity inside one\. Letni,dn\_\{i,d\}be the extent of element dimensionddin sampleii\. An ordered tupleℛ=\(r1,…,rq\)\\mathcal\{R\}=\(r\_\{1\},\\ldots,r\_\{q\}\)declares the ragged axes\. The complementary static axes𝒮=\(s1,…,sD−q\)\\mathcal\{S\}=\(s\_\{1\},\\ldots,s\_\{D\-q\}\)have common extentscj=ni,sjc\_\{j\}=n\_\{i,s\_\{j\}\}across samples\.

#### Numerical storage\.

The permutationπ=\(ℛ,𝒮\)\\pi=\(\\mathcal\{R\},\\mathcal\{S\}\)places ragged axes before static axes\. We collapse each sample’s ragged prefix and concatenate the resulting rows\. Writingpi=∏j=1qni,rjp\_\{i\}=\\prod\_\{j=1\}^\{q\}n\_\{i,r\_\{j\}\}for a sample’s packed row count,T=∑ipiT=\\sum\_\{i\}p\_\{i\}andoi=∑k<ipko\_\{i\}=\\sum\_\{k<i\}p\_\{k\}, the value tensorV∈ℝT×c1×⋯×cD−qV\\in\\mathbb\{R\}^\{T\\times c\_\{1\}\\times\\cdots\\times c\_\{D\-q\}\}satisfies

V\[oi:oi\+1\]=reshape\(permuteπ\(Xi\),\(pi,c1,…,cD−q\)\)\.V\[o\_\{i\}:o\_\{i\+1\}\]=\\reshape\\\!\\left\(\\permute\_\{\\pi\}\(X\_\{i\}\),\(p\_\{i\},c\_\{1\},\\ldots,c\_\{D\-q\}\)\\right\)\.\(2\)Images\(C,Hi,Wi\)\(C,H\_\{i\},W\_\{i\}\)pack as\(∑iHi​Wi,C\)\(\\sum\_\{i\}H\_\{i\}W\_\{i\},C\), pairs\(Ni,Mi,C\)\(N\_\{i\},M\_\{i\},C\)as\(∑iNi​Mi,C\)\(\\sum\_\{i\}N\_\{i\}M\_\{i\},C\), and non\-leading layouts\(S,Ni,C\)\(S,N\_\{i\},C\)as\(∑iNi,S,C\)\(\\sum\_\{i\}N\_\{i\},S,C\)\.

Figure 2:Structure through a single–pair–single block\.For lengths\(2,3\)\(2,3\), broadcasting expands five single positions into thirteen pair cells; column reduction returns five singles\. First\-level row splits count rows per sample, second\-level splits delimit each pair row’s cells, and sample offsets delimit complete elements in packed\-cell space\. First\-level row splits and sample offsets coincide for singles but differ for pairs\.
#### Structural metadata\.

The wrapper retains the exact size matrixAi,d=ni,dA\_\{i,d\}=n\_\{i,d\}, sample offsetsoo, dimension orderπ\\pi, and row splits at each ragged level\. Layouts with explicitly declared ragged axes retain the partitions as tensors, so later operations and reconstructed wrappers can use them directly\. Although sizes can generate these partitions, retaining them avoids recovering per\-sample Python shapes at each reconstruction boundary, and exact sizes distinguish zero\-volume shapes such as\(0,3\)\(0,3\)and\(0,7\)\(0,7\)whose offsets coincide\. The public logical shape describes the padded envelope; numerical storage remainsVV\.

pair=self\.left\_projection\(left\)\.unsqueeze\(\-2\)\+self\.right\_projection\(right\)\.unsqueeze\(\-3\)

pair=pair\+self\.transition\(pair\)

single=left\+self\.reduction\_projection\(pair\.mean\(dim=\-2\)\)

pair=self\.left\_projection\(left\)\[rows\]\+self\.right\_projection\(right\)\[columns\]

Figure 3:Nothing in the model code refers to the packing\.The measured implementation of Equation[1](https://arxiv.org/html/2609.30379#S3.E1), with its three method bodies inlined:left\_projectionandright\_projectionareuθu\_\{\\theta\}andvθv\_\{\\theta\},transitionis the pre\-normalised updatehθ∘LNCph\_\{\\theta\}\\circ\\operatorname\{LN\}\_\{C\_\{p\}\}, andreduction\_projectionisgθg\_\{\\theta\}\. The tensor carries the partitions of Figure[2](https://arxiv.org/html/2609.30379#S3.F2)through all three lines without appearing in any of them\. The last line is the explicit\-packed baseline’s equivalent of the first, reaching the same values but requiring caller\-suppliedrowsandcolumnsgather indices\.

### 3\.2Structure\-Aware Operator Dispatch

A handler interprets its operands in logical coordinates, runs the numerical operation in packed coordinates, and constructs the structure of its result, so that result is ready for the next operator without model\-level reconstruction\. NestedTensor installs these handlers on high\-level PyTorch functions through\_\_torch\_function\_\_and on ATen operations through\_\_torch\_dispatch\_\_, covering 404 distinct operators across the two protocols \(Appendix[B](https://arxiv.org/html/2609.30379#A2)\)\. For ragged coordinates\(a1,…,aq\)\(a\_\{1\},\\ldots,a\_\{q\}\)in sampleii, the packed row is

ℓi​\(a1,…,aq\)=oi\+∑j=1qaj​∏k=j\+1qni,rk\.\\ell\_\{i\}\(a\_\{1\},\\ldots,a\_\{q\}\)=o\_\{i\}\+\\sum\_\{j=1\}^\{q\}a\_\{j\}\\prod\_\{k=j\+1\}^\{q\}n\_\{i,r\_\{k\}\}\.\(3\)Static axes map individually to axes ofVV, while the ragged coordinates jointly determine its leading index\.

#### Creating axes through broadcasting\.

The projected singles in Equation[1](https://arxiv.org/html/2609.30379#S3.E1)acquire complementary singleton dimensions, giving element shapes\(Ni,1,Cp\)\(N\_\{i\},1,C\_\{p\}\)and\(1,Ni,Cp\)\(1,N\_\{i\},C\_\{p\}\)\. Their broadcast creates\(Ni,Ni,Cp\)\(N\_\{i\},N\_\{i\},C\_\{p\}\)pair states\. The handler associates each pair cell with its sample\-local row and column, gathers the corresponding projected features, and constructs the square output sizes and partitions\. Packed row counts change fromNiN\_\{i\}toNi2N\_\{i\}^\{2\}, and in Figure[2](https://arxiv.org/html/2609.30379#S3.F2)the sample offsets from\[0,2,5\]\[0,2,5\]to\[0,4,13\]\[0,4,13\]\. The first pair partition retains the single\-row grouping, while the second records the columns within each row\.

#### Retaining structure through feature transformations\.

Activations and feature\-wise normalisation retain the pair axes and partitions\. A projection updates the static feature extent, while the residual addition requires matching logical pair structure\. These operations reuse the existing partitions instead of inferring new ones from their current extents\.

#### Consuming one logical axis\.

For a rectangular pair statePi∈ℝNi×Mi×CpP\_\{i\}\\in\\mathbb\{R\}^\{N\_\{i\}\\times M\_\{i\}\\times C\_\{p\}\}, a column sumYiY\_\{i\}groups cells with the same sample and row coordinate:

VY\[ηi\+a,:\]=∑b=0Mi−1VP\[oi\+aMi\+b,:\],ηi=∑k<iNk\.V\_\{Y\}\[\\eta\_\{i\}\+a,:\]=\\sum\_\{b=0\}^\{M\_\{i\}\-1\}V\_\{P\}\[o\_\{i\}\+aM\_\{i\}\+b,:\],\\qquad\\eta\_\{i\}=\\sum\_\{k<i\}N\_\{k\}\.\(4\)The output has shape\(Ni,Cp\)\(N\_\{i\},C\_\{p\}\)in each sample, with offsetsη\\etaand the remaining row axis\. The square\-pair mean in Equation[1](https://arxiv.org/html/2609.30379#S3.E1)additionally divides each group byNiN\_\{i\}\. Its projection can then be added toXiX\_\{i\}because the handler has restored the single representation’s logical structure as well as its values\.

#### Compatibility and semantics\.

Elementwise alignment checks logical extents and dimension order, including the grouping represented by partitions\. For example, shapes\(2,6,C\)\(2,6,C\)and\(3,4,C\)\(3,4,C\)both flatten to twelve feature rows but are incompatible for direct elementwise addition\. Dense operands are aligned to logical axes before permutation or gathering\. Each handler is specified so that interpreting its packed output logically agrees with running the reference operation on the logically interpreted inputs, which for sample\-local operations is the corresponding dense element computation; Appendix[A](https://arxiv.org/html/2609.30379#A1)states that contract and the semantics of batch and global operations\.

### 3\.3Efficient Execution on Packed Storage

#### Dense work across valid positions\.

Feature\-wise operations reuse dense kernels onVV\. Projections computeV′=V​W𝖳\+bV^\{\\prime\}=VW^\{\\mathsf\{T\}\}\+b, and normalisation acts on each row’s static feature tail\. This preserves a large numerical batch while eliminating padded positions: the feature transformations in Equation[1](https://arxiv.org/html/2609.30379#S3.E1)process∑iNi2\\sum\_\{i\}N\_\{i\}^\{2\}pair cells rather thanB​Nmax2BN\_\{\\max\}^\{2\}\.

#### Partition\-dependent computation\.

Ragged\-axis reductions derive group indices from retained coordinates and use indexed accumulation or scatter reduction, with means additionally using group counts\. Pair broadcasting uses coordinate maps to gather the projected row and column features before their numerical combination\. Large pair\-coordinate maps are constructed on the values device from the per\-sample lengths and sample offsets, which are batch\-sized, and their construction, gathers and backward accumulation contribute to runtime and temporary storage\. The partitions also provide the segment boundaries consumed by variable\-length kernels\.

#### Attention through packed blocks\.

Projections produce token\-major queries, keys, and values with static head and feature axes\. The variable\-length attention path passes these buffers and cumulative lengths to FlashAttention\([Dao et al\., 2022](https://arxiv.org/html/2609.30379#bib.bib6);[Dao, 2024](https://arxiv.org/html/2609.30379#bib.bib7)\), and the output adopts the query structure, allowing different query and key lengths in cross\-attention\. FlexAttention paths map packed indices to local positions and restrict interactions to matching sample identifiers\([Dong et al\., 2025](https://arxiv.org/html/2609.30379#bib.bib10)\)\. Projections, feature normalisations, and residual operations share the packed representation around attention, avoiding conversions at each block boundary\.

#### Convolution over variable\-size images\.

Convolution requires neighbourhoods in each image’s own coordinate system\. The spatial path gathers sample\-local input neighbourhoods, including the receptive\-field halo, into batches of equal\-size tiles; native convolution processes those tiles and valid outputs are scattered back into packed storage, with coordinates and boundary padding defined within each image so that no interaction crosses a sample boundary\. Backward accumulates overlapping input\-gradient contributions and sums parameter gradients across tiles\. This replaces a batch\-wide padded envelope with tile buffers, trading padded computation for halo duplication and gather/scatter work; Appendix[B](https://arxiv.org/html/2609.30379#A2)gives the halo geometry and the tiling tradeoff\.

#### Structural costs\.

Execution cost divides into numerical kernels, structural indexing and data movement; packed execution pays off when the saved padded work outweighs the indexing and movement it adds\. Retained partitions additionally let an operator issue one segmented call over the batch instead of one call per sample\.

### 3\.4Autograd and Dynamic Compilation

#### Gradients across representations\.

Gradient connectivity must be preserved where a structured wrapper meets its packed payload\. Autograd records its graph over the payload while the model’s forward returns the wrapper, and a compiled block can hand structured singles and pairs to an eager consumer, so that boundary is crossed in both directions within one step and has to be differentiable itself\. A wrapper\-to\-packed bridge passes the numerical payload to such a consumer and, in backward, rebuilds a gradient carrying the same partitions and dimension order as the forward value\. The inverse bridge attaches packed numerical results to their output wrappers and aligns incoming gradients with the packed dimension order\. Together they preserve the dependency through changes in rank, layout and representation; for Equation[1](https://arxiv.org/html/2609.30379#S3.E1)they connect the eager loss back to the compiled projections, pair update and reduction\.

#### Runtime sizes and structural schemas\.

Compilation fixes the structural schema at each program point: rank, declared ragged axes, logical dimension order, and static feature layout\. The primary length\-dynamic evaluation also fixes batch size; we report no compiled results for traces whose batch size changes\. Different program points may have different schemas: in Equation[1](https://arxiv.org/html/2609.30379#S3.E1)broadcasting creates a pair and reduction returns singles, so its three program points do not share one schema, though each is fixed across calls\. Per\-sample sizes and partitions remain runtime tensor inputs, so supported length\-dynamic programs reuse a compiled graph rather than specialising on each length vector\. Output reconstruction combines the fixed schema with these runtime tensors, and custom\-operation boundaries represent data\-dependent packed extents symbolically\. Explicit declarations preserve an axis’s meaning even when observed sizes happen to coincide, so a batch whose lengths are momentarily equal still compiles to the ragged schema\.

## 4Experiments

The evaluation asks when the abstraction improves speed and memory over padding, whether it carries a complete multi\-ragged computation, and where the benefit comes from\.

### 4\.1Setup

#### Workloads and methods\.

The models are randomly initialised systems workloads, measured for execution cost rather than task quality\. The dataset study spans BERT\([Devlin et al\., 2019](https://arxiv.org/html/2609.30379#bib.bib9)\)on IMDB\([Maas et al\., 2011](https://arxiv.org/html/2609.30379#bib.bib21)\), GPT\-2\([Radford et al\., 2019](https://arxiv.org/html/2609.30379#bib.bib28)\)on WikiText\-103\([Merity et al\., 2016](https://arxiv.org/html/2609.30379#bib.bib22)\), an encoder–decoder Transformer\([Vaswani et al\., 2017](https://arxiv.org/html/2609.30379#bib.bib34)\)on WMT14\([Bojar et al\., 2014](https://arxiv.org/html/2609.30379#bib.bib3)\), a variable\-resolution ViT\([Dosovitskiy et al\., 2021](https://arxiv.org/html/2609.30379#bib.bib11)\)on ImageNet\-1K\([Russakovsky et al\., 2015](https://arxiv.org/html/2609.30379#bib.bib29)\), DETR\-R50\([Carion et al\., 2020](https://arxiv.org/html/2609.30379#bib.bib4)\)on COCO\([Lin et al\., 2014](https://arxiv.org/html/2609.30379#bib.bib19)\)and FCN\-ResNet\([He et al\., 2016](https://arxiv.org/html/2609.30379#bib.bib14);[Long et al\., 2015](https://arxiv.org/html/2609.30379#bib.bib20)\)on ADE20K\([Zhou et al\., 2019](https://arxiv.org/html/2609.30379#bib.bib35)\); Appendix[C](https://arxiv.org/html/2609.30379#A3)specifies the model scales and configurations\. The main comparison is padded execution, DanLing, and nativetorch\.nestedusing itstorch\.jaggedlayout\. Three analysis controls isolate separate effects: comparing attention\-only with whole\-program packing assesses the benefit of keeping computation packed beyond attention; explicit packing is a direct\-programming reference that manages offsets in model code; and matched\-granularity execution separates the cost of the tensor interface from the benefit of batch\-level call granularity\. Each paired comparison uses the same examples in the same order, with the same initial weights, objective and logical batch partitions\. Eager and Inductor\-compiled execution are parallel experiments, each normalised to padding in the same mode\.

#### Measurement\.

The experiments use an A100\-SXM4\-80GB with PyTorch2\.13\.0\+cu132and BF16 parameters and floating\-point inputs\. TF32 and stochastic layers are disabled for the comparisons\. Host\-wall timing covers forward computation, loss, backward and runtime structural work; preparation, input transfer and optimiser updates are outside this compute boundary\. Comparisons against padding use memory efficiencyEm=Mpad/MmE\_\{m\}=M\_\{\\mathrm\{pad\}\}/M\_\{m\}and speedupSm=Tpad/TmS\_\{m\}=T\_\{\\mathrm\{pad\}\}/T\_\{m\}, both improving above one\. Results use one of two windows throughout:*full\-trace*pools every measured step, keeping any in\-trace compilation in the total, while*steady\-state*pools the steps on which*neither*method compiled, so both sides average identical batches\. Time and peak allocated memory are taken over the selected steps; the windows coincide in eager execution\. Figure[1](https://arxiv.org/html/2609.30379#S0.F1)reports steady\-state; Figure[4](https://arxiv.org/html/2609.30379#S4.F4)and the result tables report full\-trace; Appendix Table[3](https://arxiv.org/html/2609.30379#A2.T3)gives both where they diverge, and Appendix[B](https://arxiv.org/html/2609.30379#A2)the retained fraction per cell\. Configuration ranges describe variation across model scales rather than independent\-run uncertainty\.

#### Validation protocol\.

Values, logical structure and gradients are checked against matching references at the same precision\. The performance matrix retains all completed paired measurements; other outcomes are reported separately using the codes defined in Appendix[C](https://arxiv.org/html/2609.30379#A3)\.

### 4\.2Speed and Memory Across Workloads

Across the four BERT scales, DanLing gives geometric\-mean speedups of 2\.74×\\timesin eager and 3\.39×\\timesin compiled execution\. BERT\-Base uses 61\.95% less peak allocated memory than padding\. The four eager FCN backbones give a 1\.97×\\timesgeometric\-mean speedup, with a 28\.38\-to\-12\.25 GiB allocation reduction for ResNet50\. GPT\-2 on WikiText\-103 gives an eager geometric\-mean speedup of 1\.32×\\times\. Peak reserved memory favours packing less and, for GPT\-2 Small to Large, exceeds padding \(Appendix[C](https://arxiv.org/html/2609.30379#A3)\)\.

The vision results separate the two kinds of benefit: variable\-resolution ViT has a 0\.91×\\timeseager geometric\-mean throughput ratio while improving allocated\-memory efficiency by 1\.48×\\times\. DETR is close to padded throughput \(1\.05×\\times\) with 4\.27×\\timesmemory efficiency, against the matched padded reference that preserves DETR’s valid\-boundary semantics; Appendix[C\.7](https://arxiv.org/html/2609.30379#A3.SS7)’s conventional padded workflow changes those semantics and runs faster still\.

Eager WMT\-Base saves substantial memory but runs at 0\.67×\\timespadded throughput\. Its compiled speedup is 1\.60×\\timesin steady state and 58\.18×\\timesover the full trace, reflecting four additional padded recompilations that DanLing’s fixed compiled schema avoids\. Native jagged reaches steady\-state throughput close to DanLing’s in the completed compiled GPT\-2 and WMT14 cells, while their eager results differ more widely across workloads\.

### 4\.3Composing Multiple Ragged Axes

The single–pair–single program of Equation[1](https://arxiv.org/html/2609.30379#S3.E1)directly exercises the paper’s central design, with both of its outputs used by an eager, sample\-weighted squared loss\. The same program is expressed in square, rectangular and non\-leading layouts, withB=8B=8, single width 384 and pair width 128\. Padded, explicit\-packed and DanLing executions share the feature transformations and objective\.

Table 1:Single–pair–single program, eager and compiled\.Full\-trace mean ms/step / peak allocated GiB; the final column is speedup over padding\. Both modes run the same logical workloads\.NestedTensor runs 1\.79–2\.26×\\timesfaster than padding across these layouts and modes \(Table[1](https://arxiv.org/html/2609.30379#S4.T1)\)\. The compiled program reuses its graph across the tested changing partitions and extents\. NestedTensor and explicit packing both run at whole\-batch granularity here, the latter through gather indices and a segmented reduction; on this program of cheap operators NestedTensor’s latency is 25–34% higher, whereas the matched\-granularity Pairformer control of Section[4\.4](https://arxiv.org/html/2609.30379#S4.SS4)observes at most 1\.60% on fused kernels\.

#### A larger pair application\.

A four\-block Pairformer\-style workload extends the comparison to repeated triangle multiplication, starting\- and ending\-node triangle attention, single attention and feature transitions\([Abramson et al\., 2024](https://arxiv.org/html/2609.30379#bib.bib1)\), at single width 384 and pair width 128\. Padding runs the triangle updates as nativeeinsumcontractions and masked attention over the envelope\. For square pairs, explicit packing, which again manages offsets itself, invokes OOps \(a fused triangle\-kernel library\) dense operators per sample, while NestedTensor invokes the corresponding segmented operators over the batch; both run pair\-biased single attention per sample on unpacked slices\.

Table 2:Four\-block Pairformer\-style application, eager execution\.Full\-trace mean ms/step / peak allocated GiB; the final column is speedup over padding\. The square regimes span length variability; padding runs native PyTorch kernels\.The measured square regimes give 2\.40–4\.32×\\timesspeedups over padding, which combine packing, fused kernels and call granularity; explicit\-packed and NestedTensor allocations are similar to each other and both well below padded\. The non\-leading and rectangular variants run per sample in both packed implementations, which then differ by no more than the interface overhead of Section[4\.4](https://arxiv.org/html/2609.30379#S4.SS4)\.

#### Compiled execution\.

Under the default Inductor configuration, DanLing completes compiled execution in all four square regimes, while the tested padded and explicit\-packed implementations fail during compilation\. They fail at different points: padded execution fails an Inductor divisibility check on a symbolic pair extent, and explicit packing fails fake\-tensor propagation on its data\-dependent gather indices \(Appendix[B](https://arxiv.org/html/2609.30379#A2)\)\. Explicit packing’s failure is the kind the runtime\-tensor metadata of Section[3\.4](https://arxiv.org/html/2609.30379#S3.SS4)avoids; padding’s is a limitation of Inductor’s fusion\. The complete pair application therefore compiles, not only the single–pair–single block of Equation[1](https://arxiv.org/html/2609.30379#S3.E1)\. Disabling Inductor fusion entirely, though not epilogue fusion alone, lets the padded implementation compile as a full graph, but only 1\.13–1\.17×\\timesfaster than its eager execution; DanLing compiled is 2\.23–6\.08×\\timesfaster again, a gap combining padding removal with fusion and kernel choice\. Relative to its own eager execution, DanLing is 1\.06–2\.20×\\timesfaster in steady state, with two of the four regimes recompiling once during the measured trace\.

### 4\.4Performance Analysis

#### Whole\-program packing\.

Figure[4](https://arxiv.org/html/2609.30379#S4.F4)separates four implementations of the same BERT workload, from padding the whole model to keeping every surrounding operation packed\.

Figure 4:Packing attention versus packing the complete program\.The four controls use the same IMDB BERT\-Base workload\. Attention\-only packing includes unpadding and repadding; explicit packing and DanLing keep surrounding operations packed\. Eager and compiled panels use full\-trace mean step time \(Section[4\.1](https://arxiv.org/html/2609.30379#S4.SS1)\)\.For random batching, replacing padded attention helps, but the larger gain comes from also keeping the surrounding feature computation packed\. Attention\-only packing uses DanLing’s variable\-length backend with every other operation left padded, so it separates the backend from the rest: 1\.47×\\timespadded throughput eager and 1\.51×\\timescompiled, against DanLing’s 3\.94×\\timesand 4\.31×\\times\. In compiled execution, DanLing and explicit packing differ by at most 0\.19% in full\-trace throughput under either batching policy, so a tensor interface can approach direct packed performance without model\-level offset bookkeeping; eager execution retains a larger gap, as the multi\-ragged program also shows\.

#### Separating interface cost from execution strategy\.

Once padded cells are gone, how the remaining work is issued still matters\. On the Pairformer workload NestedTensor is 1\.50–3\.94×\\timesfaster than explicit packing, the opposite direction to the single–pair–single program of Section[4\.3](https://arxiv.org/html/2609.30379#S4.SS3), because that baseline also differs in how it calls the kernels\. A third implementation separates the two effects, keeping the NestedTensor interface but issuing one dense call per sample as the baseline does\. At matched granularity the interface adds at most 1\.60% to latency across the four square regimes, while moving to one segmented call per operator over the whole batch gives 1\.53–3\.94×\\times\(Appendix Table[7](https://arxiv.org/html/2609.30379#A3.T7)\); the speedups above are that gain net of the interface overhead, up to run\-to\-run variation\.

#### Batching policy\.

How much redundant work exists to remove depends on the batch: for sequences the valid fraction isρseq=∑iNi/B​Nmax\\rho\_\{\\mathrm\{seq\}\}=\\sum\_\{i\}N\_\{i\}/BN\_\{\\max\}, with the image and pair analogues in Appendix[C](https://arxiv.org/html/2609.30379#A3)\. Random and bucketed IMDB traces hold the same examples, partitioned differently, and bucketing raises token occupancy from 31\.0% to 97\.7%\. The eager DanLing/padded throughput ratio then changes from 3\.94×\\timesto 1\.01×\\times, and the compiled ratio from 4\.31×\\timesto 1\.17×\\times\. Packing and bucketing are therefore complementary: the representation preserves the chosen batch, while the gain available depends on that batch’s geometry\.

#### Model scale and shape variability\.

BERT speedups increase with model scale before levelling off at Base and Large, while length heterogeneity moves a different workload: across the square Pairformer regimes atB=8B=8the speedup over padding rises from 2\.40×\\timesunder near\-uniform lengths to 4\.32×\\timesunder high variability, because padding’s cost tracks the batch maximum whatever the distribution\. At that near\-uniform end the lengths span only 208–224 and pair occupancy is 94\.4% \(Appendix Table[4](https://arxiv.org/html/2609.30379#A3.T4)\), leaving little pair\-cell padding to remove; its speedup combines packing, kernel choice and call granularity, which this comparison does not separate\. The extreme regime additionally doubles the batch toB=16B=16and is reported separately \(Table[2](https://arxiv.org/html/2609.30379#S4.T2); Appendix Figure[5](https://arxiv.org/html/2609.30379#A3.F5)\)\.

## 5Discussion and Conclusion

NestedTensor keeps logical structure available after packing and transforms that structure together with the numerical values, through a common PyTorch tensor interface\. The complete four\-block pair application compiles and trains under this representation with default compiler settings where the padded and hand\-packed implementations we measured do not, although the simpler single–pair–single block compiles for all three\. The measurements then separate where the efficiency comes from: on BERT, packing the whole program rather than attention alone; on the pair workload, issuing batch\-level segmented calls rather than per\-sample ones\. How much there is to gain stays conditional—bucketing recovers most of it on IMDB, and memory savings can coexist with slower execution—but the composition it rests on does not\.

Structure that survives into execution is useful twice: it removes padded work, and it leaves the execution layer enough information to choose how the remaining work is issued\.

### AI use statement

Generative AI assistants were used to support software development, experiment tooling, analysis, and manuscript preparation\. They were not used for research ideation or for literature search and related\-work discovery: the research question, the design of the representation and its operator rules, and the choice and framing of the experiments are the author’s own, and every cited work was located and read by the author\. No reported quantity was produced by a language model: every number in the text, tables and figures is emitted by one script from the frozen benchmark records\. AI\-assisted code was reviewed by running it and comparing its output against a reference implementation; AI\-assisted text was checked against the stored artifacts, and its cross\-references, macros and citations verified, by automated scans over the sources\. The author reviewed all AI\-assisted work and takes responsibility for the implementation, analysis, claims, and final manuscript\.

### Reproducibility statement

The representation and its structural transformations are specified in Section[3](https://arxiv.org/html/2609.30379#S3), with reference semantics in Appendix[A](https://arxiv.org/html/2609.30379#A1)and execution paths in Appendix[B](https://arxiv.org/html/2609.30379#A2)\. Appendix[C](https://arxiv.org/html/2609.30379#A3)records the environment, model configurations, measurement boundaries and the outcome codes used when a cell has no completed measurement; Section[4\.1](https://arxiv.org/html/2609.30379#S4.SS1)defines the two timing windows, and Appendix[B](https://arxiv.org/html/2609.30379#A2)reports how many steps the steady\-state window retains\.

### Ethics statement

The experiments use existing public NLP and vision datasets under their respective access terms and licenses\. They measure execution efficiency without collecting new human\-subject data or deploying trained models\. The models are randomly initialised, so no model capable of downstream use is produced or released\.

## References

- Abramsonet al\.\(2024\)J\. Abramson, J\. Adler, J\. Dunger, R\. Evans, T\. Green, A\. Pritzel, O\. Ronneberger, L\. Willmore, A\. J\. Ballard, J\. Bambrick, S\. W\. Bodenstein, D\. A\. Evans, C\. Hung, M\. O’Neill, D\. Reiman, K\. Tunyasuvunakool, Z\. Wu, A\. Žemgulytė, E\. Arvaniti, C\. Beattie, O\. Bertolli, A\. Bridgland, A\. Cherepanov, M\. Congreve, A\. I\. Cowen\-Rivers, A\. Cowie, M\. Figurnov, F\. B\. Fuchs, H\. Gladman, R\. Jain, Y\. A\. Khan, C\. M\. R\. Low, K\. Perlin, A\. Potapenko, P\. Savy, S\. Singh, A\. Stecula, A\. Thillaisundaram, C\. Tong, S\. Yakneen, E\. D\. Zhong, M\. Zielinski, A\. Žídek, V\. Bapst, P\. Kohli, M\. Jaderberg, D\. Hassabis, and J\. M\. JumperAccurate structure prediction of biomolecular interactions with AlphaFold 3\.Nature630\(8016\),pp\. 493–500\(en\)\.Cited by:[§1](https://arxiv.org/html/2609.30379#S1.p1.1),[§4\.3](https://arxiv.org/html/2609.30379#S4.SS3.SSS0.Px1.p1.1)\.
- Anselet al\.\(2024\)J\. Ansel, E\. Yang, H\. He, N\. Gimelshein, A\. Jain, M\. Voznesensky, B\. Bao, P\. Bell, D\. Berard, E\. Burovski, G\. Chauhan, A\. Chourdia, W\. Constable, A\. Desmaison, Z\. DeVito, E\. Ellison, W\. Feng, J\. Gong, M\. Gschwind, B\. Hirsh, S\. Huang, K\. Kalambarkar, L\. Kirsch, M\. Lazos, M\. Lezcano, Y\. Liang, J\. Liang, Y\. Lu, C\. K\. Luk, B\. Maher, Y\. Pan, C\. Puhrsch, M\. Reso, M\. Saroufim, M\. Y\. Siraichi, H\. Suk, S\. Zhang, M\. Suo, P\. Tillet, X\. Zhao, E\. Wang, K\. Zhou, R\. Zou, X\. Wang, A\. Mathews, W\. Wen, G\. Chanan, P\. Wu, and S\. ChintalaPyTorch 2: faster machine learning through dynamic Python bytecode transformation and graph compilation\.InProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2ASPLOS ’24: 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2,New York, NY, USA,pp\. 929–947\.Cited by:[§2](https://arxiv.org/html/2609.30379#S2.SS0.SSS0.Px2.p1.1)\.
- Bojaret al\.\(2014\)O\. Bojar, C\. Buck, C\. Federmann, B\. Haddow, P\. Koehn, J\. Leveling, C\. Monz, P\. Pecina, M\. Post, H\. Saint\-Amand, R\. Soricut, L\. Specia, and A\. TamchynaFindings of the 2014 workshop on statistical machine translation\.InProceedings of the Ninth Workshop on Statistical Machine Translation,O\. Bojar, C\. Buck, C\. Federmann, B\. Haddow, P\. Koehn, C\. Monz, M\. Post, and L\. Specia \(Eds\.\),Baltimore, Maryland, USA,pp\. 12–58\.External Links:[Document](https://dx.doi.org/10.3115/v1/W14-3302),[Link](https://aclanthology.org/W14-3302/)Cited by:[§4\.1](https://arxiv.org/html/2609.30379#S4.SS1.SSS0.Px1.p1.1)\.
- Carionet al\.\(2020\)N\. Carion, F\. Massa, G\. Synnaeve, N\. Usunier, A\. Kirillov, and S\. ZagoruykoEnd\-to\-end object detection with transformers\.InComputer Vision – ECCV 2020,A\. Vedaldi, H\. Bischof, T\. Brox, and J\. Frahm \(Eds\.\),pp\. 213–229\.External Links:ISBN 978\-3\-030\-58452\-8Cited by:[§4\.1](https://arxiv.org/html/2609.30379#S4.SS1.SSS0.Px1.p1.1)\.
- Charlieret al\.\(2021\)B\. Charlier, J\. Feydy, J\. A\. Glaunès, F\. Collin, and G\. DurifKernel operations on the GPU, with autodiff, without memory overflows\.Journal of Machine Learning Research22\(74\),pp\. 1–6\.External Links:[Link](http://jmlr.org/papers/v22/20-275.html)Cited by:[§2](https://arxiv.org/html/2609.30379#S2.SS0.SSS0.Px3.p1.1)\.
- Daoet al\.\(2022\)T\. Dao, D\. Fu, S\. Ermon, A\. Rudra, and C\. RéFlashAttention: fast and memory\-efficient exact attention with IO\-awareness\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35,pp\. 16344–16359\.External Links:[Document](https://dx.doi.org/10.52202/068431-1189),[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/67d57c32e20fd0a7a302cb81d36e40d5-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.30379#S1.p3.1),[§2](https://arxiv.org/html/2609.30379#S2.SS0.SSS0.Px3.p1.1),[§3\.3](https://arxiv.org/html/2609.30379#S3.SS3.SSS0.Px3.p1.1)\.
- Dao \(2024\)T\. DaoFlashAttention\-2: faster attention with better parallelism and work partitioning\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=mZn2Xyh9Ec)Cited by:[§1](https://arxiv.org/html/2609.30379#S1.p3.1),[§2](https://arxiv.org/html/2609.30379#S2.SS0.SSS0.Px3.p1.1),[§3\.3](https://arxiv.org/html/2609.30379#S3.SS3.SSS0.Px3.p1.1)\.
- Dehghaniet al\.\(2023\)M\. Dehghani, B\. Mustafa, J\. Djolonga, J\. Heek, M\. Minderer, M\. Caron, A\. P\. Steiner, J\. Puigcerver, R\. Geirhos, I\. Alabdulmohsin, A\. Oliver, P\. Padlewski, A\. A\. Gritsenko, M\. Lucic, and N\. HoulsbyPatch n’ pack: NaViT, a vision transformer for any aspect ratio and resolution\.InThirty\-seventh Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=VpGFHmI7e5)Cited by:[§1](https://arxiv.org/html/2609.30379#S1.p1.1),[§2](https://arxiv.org/html/2609.30379#S2.SS0.SSS0.Px3.p1.1)\.
- Devlinet al\.\(2019\)J\. Devlin, M\. Chang, K\. Lee, and K\. ToutanovaBERT: pre\-training of deep bidirectional transformers for language understanding\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),Minneapolis, Minnesota,pp\. 4171–4186\.External Links:[Document](https://dx.doi.org/10.18653/v1/N19-1423),[Link](https://aclanthology.org/N19-1423)Cited by:[Appendix C](https://arxiv.org/html/2609.30379#A3.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2609.30379#S4.SS1.SSS0.Px1.p1.1)\.
- Donget al\.\(2025\)J\. Dong, B\. Feng, D\. Guessous, Y\. Liang, and H\. HeFlexAttention: a programming model for generating fused attention variants\.\.InProceedings of Machine Learning and Systems,M\. Zaharia, G\. Joshi, and Y\. Lin \(Eds\.\),Vol\.7,pp\.\.External Links:[Link](https://proceedings.mlsys.org/paper_files/paper/2025/file/61a9278dfef5f871b5e472389f8d6fa1-Paper-Conference.pdf)Cited by:[§2](https://arxiv.org/html/2609.30379#S2.SS0.SSS0.Px3.p1.1),[§3\.3](https://arxiv.org/html/2609.30379#S3.SS3.SSS0.Px3.p1.1)\.
- Dosovitskiyet al\.\(2021\)A\. Dosovitskiy, L\. Beyer, A\. Kolesnikov, D\. Weissenborn, X\. Zhai, T\. Unterthiner, M\. Dehghani, M\. Minderer, G\. Heigold, S\. Gelly, J\. Uszkoreit, and N\. HoulsbyAn image is worth 16x16 words: transformers for image recognition at scale\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=YicbFdNTTy)Cited by:[Appendix C](https://arxiv.org/html/2609.30379#A3.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2609.30379#S4.SS1.SSS0.Px1.p1.1)\.
- FBGEMM contributors \(2026\)FBGEMM contributorsJagged tensor operators\.Note:FBGEMM documentationAccessed 2026\-09\-16External Links:[Link](https://docs.pytorch.org/FBGEMM/fbgemm_gpu/overview/jagged-tensor-ops/JaggedTensorOps.html)Cited by:[§2](https://arxiv.org/html/2609.30379#S2.SS0.SSS0.Px1.p1.1)\.
- Fegadeet al\.\(2022\)P\. Fegade, T\. Chen, P\. Gibbons, and T\. MowryThe CoRa tensor compiler: compilation for ragged tensors with minimal padding\.InProceedings of Machine Learning and Systems,D\. Marculescu, Y\. Chi, and C\. Wu \(Eds\.\),Vol\.4,pp\. 721–747\.External Links:[Link](https://proceedings.mlsys.org/paper_files/paper/2022/file/afe8a4577080504b8bec07bbe4b2b9cc-Paper.pdf)Cited by:[§1](https://arxiv.org/html/2609.30379#S1.p1.1),[§2](https://arxiv.org/html/2609.30379#S2.SS0.SSS0.Px2.p1.1)\.
- Heet al\.\(2016\)K\. He, X\. Zhang, S\. Ren, and J\. SunDeep residual learning for image recognition\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),Cited by:[§4\.1](https://arxiv.org/html/2609.30379#S4.SS1.SSS0.Px1.p1.1)\.
- Jumperet al\.\(2021\)J\. Jumper, R\. Evans, A\. Pritzel, T\. Green, M\. Figurnov, O\. Ronneberger, K\. Tunyasuvunakool, R\. Bates, A\. Žídek, A\. Potapenko, A\. Bridgland, C\. Meyer, S\. A\. A\. Kohl, A\. J\. Ballard, A\. Cowie, B\. Romera\-Paredes, S\. Nikolov, R\. Jain, J\. Adler, T\. Back, S\. Petersen, D\. Reiman, E\. Clancy, M\. Zielinski, M\. Steinegger, M\. Pacholska, T\. Berghammer, S\. Bodenstein, D\. Silver, O\. Vinyals, A\. W\. Senior, K\. Kavukcuoglu, P\. Kohli, and D\. HassabisHighly accurate protein structure prediction with AlphaFold\.Nature596\(7873\),pp\. 583–589\(en\)\.Cited by:[§1](https://arxiv.org/html/2609.30379#S1.p1.1)\.
- Kocmi and Bojar \(2017\)T\. Kocmi and O\. BojarCurriculum learning and minibatch bucketing in neural machine translation\.InProceedings of the International Conference Recent Advances in Natural Language Processing, RANLP 2017,R\. Mitkov and G\. Angelova \(Eds\.\),Varna, Bulgaria,pp\. 379–386\.External Links:[Document](https://dx.doi.org/10.26615/978-954-452-049-6%5F050),[Link](https://aclanthology.org/R17-1050/)Cited by:[§1](https://arxiv.org/html/2609.30379#S1.p3.1)\.
- Krellet al\.\(2021\)M\. M\. Krell, M\. Kosec, S\. P\. Perez, and A\. FitzgibbonEfficient sequence packing without cross\-contamination: accelerating large language models without impacting performance\.Note:Revised version: arXiv:2107\.02027v2, 2022External Links:2107\.02027,[Link](https://arxiv.org/abs/2107.02027)Cited by:[§2](https://arxiv.org/html/2609.30379#S2.SS0.SSS0.Px3.p1.1)\.
- Laiet al\.\(2025\)R\. Lai, J\. Shao, S\. Feng, S\. Lyubomirsky, B\. Hou, W\. Lin, Z\. Ye, H\. Jin, Y\. Jin, J\. Liu, L\. Jin, Y\. Cai, Z\. Jiang, Y\. Wu, S\. Park, P\. Srivastava, J\. Roesch, T\. C\. Mowry, and T\. ChenRelax: composable abstractions for end\-to\-end dynamic machine learning\.InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2,ASPLOS ’25,New York, NY, USA,pp\. 998–1013\.External Links:[Document](https://dx.doi.org/10.1145/3676641.3716249),ISBN 9798400710797,[Link](https://doi.org/10.1145/3676641.3716249)Cited by:[§2](https://arxiv.org/html/2609.30379#S2.SS0.SSS0.Px2.p1.1)\.
- Linet al\.\(2014\)T\. Lin, M\. Maire, S\. Belongie, J\. Hays, P\. Perona, D\. Ramanan, P\. Dollár, and C\. L\. ZitnickMicrosoft COCO: common objects in context\.InComputer Vision – ECCV 2014,D\. Fleet, T\. Pajdla, B\. Schiele, and T\. Tuytelaars \(Eds\.\),Cham,pp\. 740–755\.External Links:ISBN 978\-3\-319\-10602\-1Cited by:[§4\.1](https://arxiv.org/html/2609.30379#S4.SS1.SSS0.Px1.p1.1)\.
- Longet al\.\(2015\)J\. Long, E\. Shelhamer, and T\. DarrellFully convolutional networks for semantic segmentation\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),Cited by:[§4\.1](https://arxiv.org/html/2609.30379#S4.SS1.SSS0.Px1.p1.1)\.
- Maaset al\.\(2011\)A\. L\. Maas, R\. E\. Daly, P\. T\. Pham, D\. Huang, A\. Y\. Ng, and C\. PottsLearning word vectors for sentiment analysis\.InProceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies,D\. Lin, Y\. Matsumoto, and R\. Mihalcea \(Eds\.\),Portland, Oregon, USA,pp\. 142–150\.External Links:[Link](https://aclanthology.org/P11-1015/)Cited by:[§4\.1](https://arxiv.org/html/2609.30379#S4.SS1.SSS0.Px1.p1.1)\.
- Merityet al\.\(2016\)S\. Merity, C\. Xiong, J\. Bradbury, and R\. SocherPointer sentinel mixture models\.External Links:1609\.07843,[Link](https://arxiv.org/abs/1609.07843)Cited by:[§4\.1](https://arxiv.org/html/2609.30379#S4.SS1.SSS0.Px1.p1.1)\.
- Morishitaet al\.\(2017\)M\. Morishita, Y\. Oda, G\. Neubig, K\. Yoshino, K\. Sudoh, and S\. NakamuraAn empirical study of mini\-batch creation strategies for neural machine translation\.InProceedings of the First Workshop on Neural Machine Translation,T\. Luong, A\. Birch, G\. Neubig, and A\. Finch \(Eds\.\),Vancouver,pp\. 61–68\.External Links:[Document](https://dx.doi.org/10.18653/v1/W17-3208),[Link](https://aclanthology.org/W17-3208/)Cited by:[§1](https://arxiv.org/html/2609.30379#S1.p3.1)\.
- Pivarskiet al\.\(2020\)J\. Pivarski, P\. Elmer, and D\. LangeAwkward arrays in Python, C\+\+, and Numba\.EPJ Web of Conferences245,pp\. 05023\.Cited by:[§2](https://arxiv.org/html/2609.30379#S2.SS0.SSS0.Px1.p1.1)\.
- PyTorch contributors \(2026a\)PyTorch contributorsCUDA semantics\.Note:PyTorch main documentationAccessed 2026\-09\-16External Links:[Link](https://docs.pytorch.org/docs/main/notes/cuda.html)Cited by:[Appendix C](https://arxiv.org/html/2609.30379#A3.SS0.SSS0.Px3.p1.1)\.
- PyTorch contributors \(2026b\)PyTorch contributorsPyTorch benchmark\.Note:PyTorch TutorialsAccessed 2026\-09\-16External Links:[Link](https://docs.pytorch.org/tutorials/recipes/recipes/benchmark.html)Cited by:[Appendix C](https://arxiv.org/html/2609.30379#A3.SS0.SSS0.Px3.p1.1)\.
- PyTorch contributors \(2026c\)PyTorch contributorsTorch\.nested\.Note:PyTorch main documentationAccessed 2026\-09\-16External Links:[Link](https://docs.pytorch.org/docs/main/nested.html)Cited by:[§1](https://arxiv.org/html/2609.30379#S1.p3.1),[§2](https://arxiv.org/html/2609.30379#S2.SS0.SSS0.Px1.p1.1)\.
- Radfordet al\.\(2019\)A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, and I\. SutskeverLanguage models are unsupervised multitask learners\.Cited by:[§4\.1](https://arxiv.org/html/2609.30379#S4.SS1.SSS0.Px1.p1.1)\.
- Russakovskyet al\.\(2015\)O\. Russakovsky, J\. Deng, H\. Su, J\. Krause, S\. Satheesh, S\. Ma, Z\. Huang, A\. Karpathy, A\. Khosla, M\. Bernstein, A\. C\. Berg, and L\. Fei\-FeiImageNet large scale visual recognition challenge\.International Journal of Computer Vision115\(3\),pp\. 211–252\(en\)\.Cited by:[§4\.1](https://arxiv.org/html/2609.30379#S4.SS1.SSS0.Px1.p1.1)\.
- Shazeer \(2020\)N\. ShazeerGLU variants improve transformer\.External Links:2002\.05202,[Link](https://arxiv.org/abs/2002.05202)Cited by:[§3](https://arxiv.org/html/2609.30379#S3.p1.3)\.
- TensorFlow contributors \(2026\)TensorFlow contributorsRagged tensors\.Note:TensorFlow Core documentationAccessed 2026\-09\-16External Links:[Link](https://www.tensorflow.org/guide/ragged_tensor)Cited by:[§2](https://arxiv.org/html/2609.30379#S2.SS0.SSS0.Px1.p1.1)\.
- Touvronet al\.\(2021\)H\. Touvron, M\. Cord, M\. Douze, F\. Massa, A\. Sablayrolles, and H\. JégouTraining data\-efficient image transformers & distillation through attention\.InProceedings of the 38th International Conference on Machine Learning,M\. Meila and T\. Zhang \(Eds\.\),Proceedings of Machine Learning Research, Vol\.139,pp\. 10347–10357\.External Links:[Link](https://proceedings.mlr.press/v139/touvron21a.html)Cited by:[Appendix C](https://arxiv.org/html/2609.30379#A3.SS0.SSS0.Px2.p1.1)\.
- Turcet al\.\(2019\)I\. Turc, M\. Chang, K\. Lee, and K\. ToutanovaWell\-read students learn better: on the importance of pre\-training compact models\.External Links:1908\.08962,[Link](https://arxiv.org/abs/1908.08962)Cited by:[Appendix C](https://arxiv.org/html/2609.30379#A3.SS0.SSS0.Px2.p1.1)\.
- Vaswaniet al\.\(2017\)A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. PolosukhinAttention is all you need\.InAdvances in Neural Information Processing Systems,I\. Guyon, U\. V\. Luxburg, S\. Bengio, H\. Wallach, R\. Fergus, S\. Vishwanathan, and R\. Garnett \(Eds\.\),Vol\.30,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf)Cited by:[§4\.1](https://arxiv.org/html/2609.30379#S4.SS1.SSS0.Px1.p1.1)\.
- Zhouet al\.\(2019\)B\. Zhou, H\. Zhao, X\. Puig, T\. Xiao, S\. Fidler, A\. Barriuso, and A\. TorralbaSemantic understanding of scenes through the ADE20K dataset\.International Journal of Computer Vision127\(3\),pp\. 302–321\(en\)\.Cited by:[§4\.1](https://arxiv.org/html/2609.30379#S4.SS1.SSS0.Px1.p1.1)\.

## Appendix ARepresentation, Semantics, and the Running Program

#### Structural invariants\.

Ragged axes in this work vary between individually dense elements of a common rank\. The representation retains static feature axes, an ordered tuple of ragged axes, and the batch\-position convention\. For the batch\-leading exposition, element axisddis public axisd\+1d\+1\. The implementation also supports its sequence\-first convention; the batch axis is not an arbitrary permutable element dimension\. Sample offsets satisfyo0=0o\_\{0\}=0,oi\+1−oi=pio\_\{i\+1\}\-o\_\{i\}=p\_\{i\}, andoB=To\_\{B\}=T\. Zero\-volume elements may have coincident consecutive offsets, while the exact size matrix distinguishes their remaining logical dimensions\.

Sizes and offsets play different roles\. For example, element shapes\(2,6,C\)\(2,6,C\)and\(3,4,C\)\(3,4,C\)both occupy twelve packed feature rows but cannot be aligned as the same logical pair tensor\. Hierarchical row splits identify the groups consumed by reductions\. For rectangular elements\(Ni,Mi,C\)\(N\_\{i\},M\_\{i\},C\), the innermost splits contain1\+∑iNi1\+\\sum\_\{i\}N\_\{i\}entries; they need not beO⁡\(B\)O\(B\)metadata\. Canonical sizes and offsets stay on the CPU, with device mirrors and execution indices used by GPU consumers\.

#### Reference semantics\.

Let𝒟\\mathcal\{D\}denote the logical interpretation of a structured tensor; it acts componentwise on argument tuples and leaves dense tensors unchanged\. A deterministic functional handlerf^\\widehat\{f\}is specified against its referencefreff\_\{\\mathrm\{ref\}\}by𝒟⁡\(f^​\(𝐙\)\)=fref​\(𝒟⁡\(𝐙\)\)\\mathcal\{D\}\(\\widehat\{f\}\(\\mathbf\{Z\}\)\)=f\_\{\\mathrm\{ref\}\}\(\\mathcal\{D\}\(\\mathbf\{Z\}\)\), the contract cited in Section[3\.2](https://arxiv.org/html/2609.30379#S3.SS2)\. Sample\-local operations are compared with independent dense element computations using the same weights, masks, coordinate conventions, and loss weighting\. Global reductions and operations combining samples require their corresponding batch reference\. A mean over all stored entries weights entries equally\. The supplied API’s single batch\-axis reduction instead performs a complete reduction within each element and stacks the results: for elements\[1,2\]\[1,2\]and\[3,4\]\[3,4\], itssum\(dim=0\)gives\[3,7\]\[3,7\]\. This convention is distinct from a dense column\-wise batch reduction\. The cross\-entropy handler uses a trailing class dimension, so a reference must align the objective accordingly\.

The validity mask uses true for valid positions, while attention interfaces can use different boolean conventions\. The adapter translates those conventions and combines padding validity with model masks\. Padded accessors are interoperability operations; the canonical numerical payload remains packed\. That contract is a specification; derivative and structural agreement are tested and reported separately\.

### A\.1Measured single–pair–single program

The measured mechanism uses the projections, residual pre\-normalised SwiGLU update, and column mean in Equation[1](https://arxiv.org/html/2609.30379#S3.E1)\. Pair construction is the sum of the two projected single features; there is no additional activation at that construction step\. Both returned outputs are consumed outside the compiled forward by

ℒ=∑iαi​‖X~i‖F2S​C​∑iNi\+∑iβi​‖P~i‖F2S​Cp​∑iNi​Mi,αi,βi\>0\.\\mathcal\{L\}=\\frac\{\\sum\_\{i\}\\alpha\_\{i\}\\\|\\widetilde\{X\}\_\{i\}\\\|\_\{F\}^\{2\}\}\{SC\\sum\_\{i\}N\_\{i\}\}\+\\frac\{\\sum\_\{i\}\\beta\_\{i\}\\\|\\widetilde\{P\}\_\{i\}\\\|\_\{F\}^\{2\}\}\{SC\_\{p\}\\sum\_\{i\}N\_\{i\}M\_\{i\}\},\\qquad\\alpha\_\{i\},\\beta\_\{i\}\>0\.\(5\)HereMi=NiM\_\{i\}=N\_\{i\}for square pairs,S=1S=1for leading layouts, andS=3S=3for the non\-leading variant\. The rectangular variant has distinct row and column inputs and divides its column mean byMiM\_\{i\}\. The sample weights areαi=1\+i/B\\alpha\_\{i\}=1\+i/Bandβi=1\.5−i/\(2​B\)\\beta\_\{i\}=1\.5\-i/\(2B\)fori=0,…,B−1i=0,\\ldots,B\-1\. The sample weights and losses are identical across representations; the squared sums accumulate in at least FP32\. This mechanism has no trainable readout after the returned pair state\. The separate Pairformer application does include a trainable pair readout in its eager loss\.

The mechanism usesB=8B=8,C=384C=384,Cp=128C\_\{p\}=128, and transition expansion four\. Eight validation patterns cover changed partitions with fixed totals, changed packed extents and maxima, and fresh equal structures\. Each measured method/layout/mode cell has three traces of two steps\. All eight patterns pass the recorded structure, output, gradient and one\-update checks for all 18 combinations\.

## Appendix BImplementation and Compilation Scope

#### Execution paths\.

Operator availability, a packed implementation, and fullgraph compilation are distinct properties\. The two registries hold 708 keys for 404 distinct operators, since one operator owns several keys, astorch\.add,Tensor\.add,Tensor\.\_\_add\_\_andaten::add\.Tensordo; operators are counted after dropping module paths and overload suffixes, and an in\-place operator counts separately from its out\-of\-place form because it mutates packed storage through its own handler\. Feature\-wise operations reuse dense packed kernels\. Ragged reductions and broadcasts use coordinate/group maps; selected attention paths use native variable\-length or FlexAttention execution\. Spatial paths gather tile neighbourhoods and invoke dense convolution, or use direct pooling kernels; selected pooling paths instead address packed pixels directly through per\-sample shapes and offsets\. Pointwise convolution reuses the dense packed path after aligning the channel axis with the static feature dimension\. For output tile dimensions\(th,tw\)\(t\_\{h\},t\_\{w\}\), stridess, dilationδ\\deltaand kernel sizekk, the gathered input tile has dimensionsuh=\(th−1\)​sh\+δh​\(kh−1\)\+1u\_\{h\}=\(t\_\{h\}\-1\)s\_\{h\}\+\\delta\_\{h\}\(k\_\{h\}\-1\)\+1anduw=\(tw−1\)​sw\+δw​\(kw−1\)\+1u\_\{w\}=\(t\_\{w\}\-1\)s\_\{w\}\+\\delta\_\{w\}\(k\_\{w\}\-1\)\+1\. Tile size therefore balances numerical batching against overlapping halo reads, boundary tiles and gather/scatter work; chunking bounds the number of simultaneous tile buffers, the dense backend may use channels\-last storage, and backward can regenerate input tiles from the saved packed input for weight gradients\. Some bindings, including segmentedcdistandcumprod, retain per\-sample native kernel invocations behind a custom\-operation boundary\. Eager fallback and unsupported compiled argument combinations remain part of the implementation’s declared scope; a workload counts as passing only on its own measured outcome, never on operator availability\.

#### Dynamic metadata and differentiation\.

The principal length\-dynamic contract fixes batch size and the structural schema at each program point\. The flattening protocol exposes one differentiable payload child alongside nondifferentiable structural metadata\. Tensor\-backed sizes and row splits are runtime inputs; standalone FakeTensor construction retains the Python metadata needed where symbolic data\-dependent extents are unavailable\. Eager reconstructions propagate weak dynamic annotations before entering a compiled region\. This prevents a ragged extent from being accidentally identified with an unrelated metadata dimension of the same observed size\. The optimised eager propagation recognises the state installed by the public marker and preserves a public\-API path for missing marker metadata, tensor subclasses, and FakeTensor execution\.

Wrapper\-to\-packed and packed\-to\-wrapper autograd bridges preserve the corresponding numerical dependency and packed dimension order, so changing partitions reuse the compiled graph for the evaluated program and gradients remain connected across the compiled/eager boundary\.

#### Current compiled application outcomes\.

The ordinary\-operation mechanism completes in all three layouts\. The dataset matrix records upstream native\-jagged BERT guard errors, failed padded vision compilation, and NestedTensor FCN fake\-propagation failures\. Under the default Inductor configuration the full Pairformer compiles for DanLing but not for either baseline \(Section[4\.3](https://arxiv.org/html/2609.30379#S4.SS3)\)\. The recorded failures are identical across the four square regimes: padded raisesInductorErrorwithCantSplit:128\*s83notdivisibleby\(\(s83\*\*2\)//s83\), and explicit packing raisesTorchRuntimeErrorduring fake\-tensor propagation\. The padded failure has no user frame: Inductor’s scheduler raises it while fusing nodes whose iteration extents ares2s^\{2\}and128​s128s, a constraint that holds for every positive integerssbut that it cannot establish\. Padded therefore compiles under static shapes, and under dynamic shapes once fusion is disabled \(max\_fusion\_size=1;epilogue\_fusion=Falsealone still fails\), running at 522–556 ms/step in one harness whose padded eager reproduces Table[2](https://arxiv.org/html/2609.30379#S4.T2)\. Explicit packing fails before any fusion decision and was not retested\. Of DanLing’s four compiled regimes, two recompile once inside the measured trace, unlike the sequence workloads, where the compiled schema is reused throughout\. For FCN, padded ResNet18/101 reach the recompilation limit, ResNet152 exhausts the recorded two\-hour budget without a completed outcome, and the standalone ResNet50 padded compiled cell has no result artifact\. The four NestedTensor FCN compiled cells fail during convolution fake propagation\.

#### Steady\-state execution versus in\-trace recompilation\.

A measured step whose shape was not already covered by a compiled graph triggers a new frontend/backend capture; eager execution has no such event, so the distinction only matters for compiled runs\. Figure[1](https://arxiv.org/html/2609.30379#S0.F1)’s compiled\-mode speedup pools only steps that did not themselves trigger a capture, isolating kernel and packing efficiency from compiler behaviour\. Every steady\-state ratio pools the steps on which neither method compiled, so numerator and denominator average identical batches\. That intersection retains all paired steps except in six compiled cells: 63 of 64 for WikiText Small, Medium and Large, 124 of 128 for WMT14\-Base, 125 of 128 for WMT14\-Big, and four of six for WikiText\-XL, whose short trace makes it the only cell where the restriction is material \(it moves the ratio by 4\.1%; elsewhere the shift is at most 0\.48%\)\. Table[3](https://arxiv.org/html/2609.30379#A2.T3)reports the complement for the compiled cells whose full\-trace and steady\-state ratios diverge most \(Section[4\.2](https://arxiv.org/html/2609.30379#S4.SS2)\): how many graphs each method compiled over its lifetime, including calibration, and how many of those compilations were shape\-driven recompilations occurring after calibration, alongside the resulting steady\-state and full\-trace per\-step times\. In these recorded traces the padded implementations recompile more often than DanLing and native jagged, which reuse one compiled schema in five of the six cells; WikiText\-XL is the exception, where its six\-step trace leaves every method recompiling once or twice\.

Table 3:Dynamic\-compilation evidence\.“Graphs” counts lifetime frontend/backend invocations, including calibration; “recompile” counts only those occurring after calibration, during the measured trace\. “Compile step” is the mean measured time of a recompiling step, execution and compilation combined; it is not an isolated compiler stopwatch\. Steady\-state time pools non\-recompiling steps; full\-trace time pools every measured step\.

## Appendix CBenchmark Protocol and Complete Dataset Matrix

#### Environment and source cohorts\.

The PyTorch build is2\.13\.0\+cu132, commitcf30153c4c131c8164ee7798e5022d810682e2cb, with CUDA 13\.2 and Triton 3\.7\.1 on A100\-SXM4\-80GB GPUs\. Parameters use BF16; TF32 is disabled, float32 matmul precision ishighest, and stochastic layers are disabled for these comparisons\. The validation optimiser is SGD at10−510^\{\-5\}, with BF16 parameters and no separate FP32 master\-weight copy\.

The reported tables and figures are not all measurements of a single checkout: the dataset and execution\-control timings, the mechanism snapshot, and the square Pairformer timings were each frozen at a different source revision\. Every exported table and figure is tied to frozen manifests recording the exact source revision and input\-file hashes, distributed with the artifact\. The Pairformer cells additionally depend on the OOps kernel library, whose public release postdates these measurements\.

Figure 5:Model\-scale and shape\-heterogeneity sensitivity\.\(a\)BERT model\-scale speedup \(padded / DanLing mean step time\) atB=64B=64, eager \(hollow\) and compiled \(solid\)\.\(b\)Eager mean step time for the same four\-block Pairformer\-style workload as Table[2](https://arxiv.org/html/2609.30379#S4.T2), for padded, explicit\-packed and DanLing across four length\-heterogeneity regimes at fixed model size, from nearly uniform lengths to one long chain among many short ones; the rightmost regime additionally doubles the batch toB=16B=16\. Every point compares methods within the same execution mode\.The image and pair occupancy fractions accompanyingρseq\\rho\_\{\\mathrm\{seq\}\}in Section[4\.4](https://arxiv.org/html/2609.30379#S4.SS4)areρimage=∑iHi​Wi/B​Hmax​Wmax\\rho\_\{\\mathrm\{image\}\}=\\sum\_\{i\}H\_\{i\}W\_\{i\}/BH\_\{\\max\}W\_\{\\max\}andρpair=∑iNi​Mi/B​Nmax​Mmax\\rho\_\{\\mathrm\{pair\}\}=\\sum\_\{i\}N\_\{i\}M\_\{i\}/BN\_\{\\max\}M\_\{\\max\}\. Table[4](https://arxiv.org/html/2609.30379#A3.T4)lists the lengths andρpair\\rho\_\{\\mathrm\{pair\}\}behind Tables[1](https://arxiv.org/html/2609.30379#S4.T1)and[2](https://arxiv.org/html/2609.30379#S4.T2)\.

Table 4:Length regimes of the pair workloads\.Lengths span every measured sample \(rows×\\timescolumns for rectangular pairs\);ρpair\\rho\_\{\\mathrm\{pair\}\}pools valid and padded pair cells over the measured batches\. A cell’s eager and compiled runs, and the matched\-granularity run of a square regime \(Table[7](https://arxiv.org/html/2609.30379#A3.T7)\), use the same length schedule\.
#### Architecture and batch sizes\.

IMDB uses BERT Tiny/Small configurations from[Turc et al\. \(2019\)](https://arxiv.org/html/2609.30379#bib.bib33)and Base/Large configurations from[Devlin et al\. \(2019\)](https://arxiv.org/html/2609.30379#bib.bib9), with\(L,H,C\)\(L,H,C\)equal to\(2,2,128\)\(2,2,128\),\(4,8,512\)\(4,8,512\),\(12,12,768\)\(12,12,768\), and\(24,16,1024\)\(24,16,1024\), all atB=64B=64\. WikiText uses GPT\-2 Small/Medium/Large/XL shapes with\(L,H,C\)\(L,H,C\)equal to\(12,12,768\)\(12,12,768\),\(24,16,1024\)\(24,16,1024\),\(36,20,1280\)\(36,20,1280\), and\(48,25,1600\)\(48,25,1600\)\. Its main document\-length cap is 4096, withB=8B=8except XL atB=2B=2\. The WMT Base/Big settings use six encoder and six decoder layers, widths 512/1024, and 8/16 heads atB=64B=64\. ViT Ti/S use the DeiT\-Ti/S scale configurations of[Touvron et al\. \(2021\)](https://arxiv.org/html/2609.30379#bib.bib32), while B/L follow[Dosovitskiy et al\. \(2021\)](https://arxiv.org/html/2609.30379#bib.bib11); all use patch size 16, widths 192/384/768/1024, depths 12/12/12/24, and 3/6/12/16 heads atB=32B=32\. DETR uses ResNet50, six encoder and six decoder layers, hidden width 256, eight heads, and 100 queries atB=32B=32\. FCN uses ResNet18/50/101/152 and 150 classes atB=32B=32\.

#### Measurement accounting\.

Timing follows PyTorch’s benchmarking and CUDA synchronisation guidance\([PyTorch contributors, 2026b](https://arxiv.org/html/2609.30379#bib.bib25);[PyTorch contributors, 2026a](https://arxiv.org/html/2609.30379#bib.bib26)\)\. Every cell fixes one execution mode\. Batch identities and sizes come from the frozen sibling manifest, not placeholder configuration lengths or a mutable external manifest\. Completed traces are checked for matching IDs, repetition IDs and step counts\. Latency sums all measured steps; throughput divides all measured examples by that elapsed time\. Memory takes the maximum allocated or reserved bytes across all measured traces\. No trace is removed for containing compilation\. Validation and calibration timings are outside the compute measurement, and their cache priming means that a timed capture is not a measurement of all compilation costs from a pristine process\. The primary tables record one repetition and descriptive step variation\. Inference about independent\-run uncertainty or maximum feasible batch size is outside these records\.

#### Execution outcomes\.

Where a cell has no completed measurement,ndmarks no run on record,blmarks a padded reference that itself failed validation \(blocking the comparison\),gmarks a dynamic\-shape guard error during compiled dispatch,cfmarks any other compilation failure,nrmarks a run that ended without a completed outcome, andnemarks a method not evaluated because the workload falls outside its structural support \(e\.g\. native jagged on COCO/ADE20K\); unfinished measurements keepnrrather than a diagnosed cause\. The complete grids below report each configuration as two rows: mean throughput in samples/s, then peak allocated / reserved memory in GiB\. Reserved memory additionally counts blocks the caching allocator holds for reuse, and favours packing less than allocated memory: for GPT\-2 Small to Large, DanLing and native jagged reserve more than padding\.

### C\.1IMDB / BERT

### C\.2WikiText\-103 / GPT\-2

### C\.3WMT14 / Transformer

### C\.4ImageNet / ViT

### C\.5COCO / DETR

### C\.6ADE20K / FCN\-ResNet

### C\.7Conventional spatial padding

Table 5:Matched and conventional spatial padding\.Cells are full\-trace mean ms/step / peak allocated GiB\. The conventional workflow changes boundary and interpolation semantics, so it is a different computation rather than a faster version of the same one\.The conventional FCN workflow applies interpolation over the padded envelope, whereas the matched reference applies it at the valid image extent\. Intermediate convolution and pooling boundaries also differ\. Frozen BatchNorm eliminates one source of coupling but does not identify every source of the observed differences\. DETR comparisons include predictions, matching, loss and gradients\. WMT has no corresponding spatial\-boundary comparator\.

### C\.8Execution and granularity controls

Table 6:Five\-way execution controls\.Cells are full\-trace samples/s\. Random and bucketed policies are distinct frozen batch partitions of the same measured example cohort\.Attention\-only packing includes Q/K/V unpadding and output repadding\. Explicit packing and DanLing keep surrounding feature computations packed\. The attention\-only, explicit\-packed and DanLing controls request native variable\-length attention; padded attention runs masked scaled dot\-product attention over the padded envelope, and native jagged follows its own SDPA dispatch\. Compiled IMDB reaches similar DanLing and explicit\-packed throughput in both policies\. The IMDB random\-policy cells repeat the BERT\-Base configuration of Figure[1](https://arxiv.org/html/2609.30379#S0.F1)in a separate run, so their ratios differ slightly from that figure’s\. WMT combines a larger eager implementation gap with timed compilation outliers in several compiled methods\. Method\-level kernel and metadata differences remain part of these comparisons\.

The matched\-granularity control supports two further comparisons on the Pairformer workload: DanLing with per\-sample calls against explicit packing with per\-sample calls isolates the interface, and DanLing with segmented calls against DanLing with per\-sample calls isolates the execution strategy\. Table[7](https://arxiv.org/html/2609.30379#A3.T7)reports the three step times for each square regime\.

Table 7:Matched\-granularity control on the four\-block Pairformer\-style workload, eager execution\.Full\-trace mean ms/step\. Explicit packing and DanLing per\-sample call each OOps triangle operator once per sample; DanLing segmented calls it once per batch\. Interface cost is the per\-sample DanLing slowdown relative to explicit packing; granularity gain is per\-sample over segmented DanLing time\. The three methods share one run per regime, separate from Table[2](https://arxiv.org/html/2609.30379#S4.T2), so both ratios come from that run and its times differ slightly from that table’s\. Ratios use unrounded times\.

相似文章

利用张量特征训练网络加速高维函数学习

arXiv cs.LG

本文提出一种加速高维函数深度神经网络训练的方法,通过引入上下文特征(包括来自分解预训练DNN的秩1特征和张量特征),并利用随机张量分解将存储成本降低数个数量级。

DeepLoop:循环Transformer的深度缩放

arXiv cs.LG

DeepLoop为循环Transformer引入了一种残差缩放方法,该方法根据参数访问进行调整,从而在物理块跨多轮重用时提高稳定性和性能。