Scale-Aware Attention for Scarce Neural Data: An RG-Flow Transformer on Sleep-EDF EEG
Summary
This paper introduces the RG-Flow Transformer, a model with a renormalization-group inductive bias for analyzing scarce EEG data. It benchmarks against a vanilla transformer on sleep staging from the Sleep-EDF dataset, finding no accuracy advantage but better interpretability through recovery of the spectral exponent.
View Cached Full Text
Cached at: 07/15/26, 04:16 AM
# Scale-Aware Attention for Scarce Neural Data: An RG-Flow Transformer on Sleep-EDF EEG
Source: [https://arxiv.org/html/2607.11950](https://arxiv.org/html/2607.11950)
###### Abstract
Brain field potentials are scale\-free: their power spectra follow a1/fβ1/f^\{\\beta\}law whose aperiodic exponentβ\\betatracks cortical state, and sleep depth in particular is a shift inβ\\beta\. We ask whether a transformer endowed with an explicit renormalization\-group \(RG\) inductive bias—the RG\-Flow Transformer, which couples ordinary self\-attention to a scale\-aware stream with a learnable anomalous dimensionγ\\gamma, block\-spin coarse\-graining, and an entropy\-gated synchronization bridge—has an advantage over a parameter\-matched vanilla transformer on*real, scarce*EEG\. Using the PhysioNet Sleep\-EDF corpus with a strict leakage\-free by\-subject hold\-out, we \(i\) benchmark RG\-Flow against a param\-matched vanilla transformer and a hierarchy\-only ablation on 5\-class AASM sleep staging, \(ii\) sweep the per\-subject data budget to look for the inductive\-bias crossover predicted when data are scarce, and \(iii\) test whether RG\-Flow’s learnedγ\\gammatracks the measured spectral exponentβ\\betaout\-of\-sample—a quantity the vanilla model does not possess\. Across55subjects and55seeds under leave\-one\-subject\-out cross\-validation, RG\-Flow and the vanilla transformer are statistically indistinguishable on 5\-class staging \(77\.3% vs 77\.0% accuracy; pairedp=0\.294p=0\.294\), and the predicted scarce\-data crossover does not appear: vanilla is numerically ahead at every data\-limited budget\. What does separate the models is interpretability—RG\-Flow recovers the continuous spectral exponent out\-of\-sample \(β\\beta\-recoveryR2=0\.416R^\{2\}=0\.416\), a capability the vanilla architecture has no analogue for\.
## IIntroduction
Cortical dynamics are scale\-free\. The power spectrum of the electro\- and magneto\-encephalogram, of the local field potential, and of the electrocorticogram all follow an aperiodic1/fβ1/f^\{\\beta\}law over a broad band\[[5](https://arxiv.org/html/2607.11950#bib.bib1)\], and this aperiodic exponentβ\\betais not a nuisance: it tracks excitation–inhibition balance, arousal, anaesthesia, and, most cleanly, sleep depth\[[9](https://arxiv.org/html/2607.11950#bib.bib2),[3](https://arxiv.org/html/2607.11950#bib.bib3)\]\. Deeper non\-REM sleep produces a steeper spectrum \(largerβ\\beta\), so a sleep\-stage label is, physically, a coarse label on a spectral scaling exponent\. Alongside the1/f1/fspectrum, cortex exhibits the other hallmarks of operation near a critical point—neuronal avalanches with power\-law size and duration distributions\[[1](https://arxiv.org/html/2607.11950#bib.bib4)\]and a branching ratio near unity\[[13](https://arxiv.org/html/2607.11950#bib.bib5)\]—which is precisely the regime in which scale\-invariance is a genuine property of the data rather than an imposed prior\.
This makes neural data an unusually principled testbed for a model whose inductive bias*is*scale invariance\. The renormalization group \(RG\), the framework physics uses to relate descriptions of a system across scales\[[6](https://arxiv.org/html/2607.11950#bib.bib9),[14](https://arxiv.org/html/2607.11950#bib.bib10)\], has a growing set of correspondences with deep learning\[[12](https://arxiv.org/html/2607.11950#bib.bib11),[10](https://arxiv.org/html/2607.11950#bib.bib12),[8](https://arxiv.org/html/2607.11950#bib.bib13),[11](https://arxiv.org/html/2607.11950#bib.bib14)\]\. The RG\-Flow Transformer operationalizes that correspondence: it runs ordinary associative self\-attention in parallel with a scale\-aware stream governed by a learnable anomalous dimensionγ\\gamma, coarse\-grains its representation with block\-spin pooling, and gates the two streams with an entropy threshold\.
There is a second, more pragmatic reason to expect an inductive\-bias advantage here: neural data are*scarce*\. A single night, a single session, a single subject yields limited labelled data, and a correct prior earns its keep exactly when data are too few for a flexible model to discover the structure on its own\. A companion study on abundant synthetic data found no accuracy advantage for RG\-Flow over a parameter\-matched vanilla transformer; scarcity is the regime the present study is designed to probe\.
We therefore make three contributions on the real PhysioNet Sleep\-EDF corpus, under a strict leakage\-free by\-subject hold\-out:
1. 1\.a parameter\-matched benchmark of RG\-Flow against a vanilla transformer and a hierarchy\-only ablation on 5\-class AASM sleep staging;
2. 2\.a per\-subject data\-budget sweep that looks for the crossover predicted when data are scarce; and
3. 3\.a test of whether RG\-Flow’s learnedγ\\gammatracks the measured aperiodic exponentβ\\betaout\-of\-sample—an interpretable, physically\-grounded quantity the vanilla model has no analogue for\.
The analysis protocol is fixed and fully deterministic, so the underpowered study reported here can be replaced by a larger run without any change to the surrounding argument\.
## IITheoretical Framework
### II\.1The scale\-free brain as a field
We map a raw EEG epoch onto a microscopic field configurationϕ\(t\)\\phi\(t\)defined at the shortest resolvable lattice spacing \(the sampling interval\)\. Following the statistical\-mechanics formalism, the underlying distribution of such configurations is governed by an effective data\-generating distribution analogous to a partition function,
P\(ϕ\)=1Ze−ℋ0\[ϕ\],Z=∫𝒟ϕe−ℋ0\[ϕ\],P\(\\phi\)=\\frac\{1\}\{Z\}\\,e^\{\-\\mathcal\{H\}\_\{0\}\[\\phi\]\},\\qquad Z=\\int\\mathcal\{D\}\\phi\\,e^\{\-\\mathcal\{H\}\_\{0\}\[\\phi\]\},\(1\)whereℋ0\[ϕ\]\\mathcal\{H\}\_\{0\}\[\\phi\]is an effective microscopic Hamiltonian whose coupling constants dictate short\-range interactions\. High\-frequency fluctuations play the role of ultraviolet \(UV\) modes at large momentum; low\-frequency structure plays the role of infrared \(IR\) stable states\. This analogy is unusually well\-motivated for cortical data: the EEG power spectrum is scale\-free, following an aperiodic1/fβ1/f^\{\\beta\}law\[[5](https://arxiv.org/html/2607.11950#bib.bib1)\], and cortex sits near a critical point with power\-law neuronal avalanches\[[1](https://arxiv.org/html/2607.11950#bib.bib4)\]and a branching ratio near unity\[[13](https://arxiv.org/html/2607.11950#bib.bib5)\]\. Scale\-invariance is therefore a property of the data, not an imposed prior—which is exactly what an RG\-motivated model is designed to exploit\.
### II\.2Latent Kadanoff block\-spin transformation
To move across observation scales, the neural representation must systematically integrate out short\-range fluctuations up to a scaling factorss\. LetΛ\\Lambdabe the UV cutoff corresponding to the highest resolved frequency\. A Kadanoff block\-spin transformation coarse\-grains the field by integrating out modesk\>Λ′k\>\\Lambda^\{\\prime\}withΛ′<Λ\\Lambda^\{\\prime\}<\\Lambda:
Pmacro\(Φ\)=∫𝒟ϕT\[Φ,ϕ\]Pmicro\(ϕ\)\.P\_\{\\text\{macro\}\}\(\\Phi\)=\\int\\mathcal\{D\}\\phi\\,T\[\\Phi,\\phi\]\\,P\_\{\\text\{micro\}\}\(\\phi\)\.\(2\)The discrete latent mapping realized in our architecture is
ΦT=σ\(𝐖macro∑t∈block\(Vt⋅𝐖disϕt\)\),\\Phi\_\{T\}=\\sigma\\\!\\left\(\\mathbf\{W\}\_\{\\text\{macro\}\}\\sum\_\{t\\in\\text\{block\}\}\\big\(V\_\{t\}\\cdot\\mathbf\{W\}\_\{\\text\{dis\}\}\\,\\phi\_\{t\}\\big\)\\right\),\(3\)where𝐖dis\\mathbf\{W\}\_\{\\text\{dis\}\}is a learnable disentanglement matrix that orthogonalizes UV noise prior to spatial projection andVtV\_\{t\}is the scale\-weight of steptt\. Because deep NREM sleep steepens the1/fβ1/f^\{\\beta\}spectrum, block coarse\-graining suppresses exactly the high\-frequency modes whose relative weight distinguishes sleep stages—the operation the label depends on\.
### II\.3The Callan–Symanzik invariance condition
A core premise of the architecture is that an authentic structural regime should remain invariant under adjustments to the observation\-scale parameterμ\\mu\. This is expressed by the Callan–Symanzik \(CS\) equation for annn\-point latent correlation functionG\(n\)G^\{\(n\)\}:
\[μ∂∂μ\+β\(g\)∂∂g\+nγ\(g\)\]G\(n\)\(pi;μ,g\)=0\.\\left\[\\mu\\frac\{\\partial\}\{\\partial\\mu\}\+\\beta\(g\)\\frac\{\\partial\}\{\\partial g\}\+n\\,\\gamma\(g\)\\right\]G^\{\(n\)\}\(p\_\{i\};\\mu,g\)=0\.\(4\)We do not enforce this equation as a hard invariance; instead we use it as motivation for two soft, differentiable mechanisms \(the second contributes a Callan–Symanzik*drift penalty*to the loss rather than a constraint on the forward pass\):
1. 1\.The beta\-function surrogate \(β\(g\)\\beta\(g\)\)\.In a renormalizable field theory the beta functionβ\(g\)≡μ∂g/∂μ\\beta\(g\)\\equiv\\mu\\,\\partial g/\\partial\\mugoverns the scale\-dependent running of the coupling, and its zeros are the RG fixed points\[[14](https://arxiv.org/html/2607.11950#bib.bib10)\]\. We do*not*claim to compute this object\. We use a bounded logistic surrogateβs\(g\)=g\(1−g\)\\beta\_\{\\mathrm\{s\}\}\(g\)=g\(1\-g\)acting on the sigmoid\-normalized attention couplingg∈\(0,1\)g\\in\(0,1\); it shares the two qualitative properties we exploit—fixed points atg=0g\\\!=\\\!0\(irrelevant/decoupled\) andg=1g\\\!=\\\!1\(marginal/preserved\), and a single relevance maximum between them—and defines a differentiable relevance filterR\(g\)=1−\|βs\(g\)\|R\(g\)=1\-\|\\beta\_\{\\mathrm\{s\}\}\(g\)\|that damps couplings far from a fixed point\. This is an inductive bias inspired by the RG, not a numerical solution of the RG flow\.
2. 2\.The anomalous\-dimension parameter \(γ\\gamma\)\.By analogy with the field\-theoretic anomalous dimension, which sets the scaling correction a field acquires under coarse\-graining, we introduce a per\-head learnable scalarγ\\gammathat reweights attention scores byμ−γ\\mu^\{\-\\gamma\}across the scale scheduleμℓ\\mu\_\{\\ell\}\. It is a trainable scale\-sensitivity parameter motivated by—not identified with—γ\(g\)\\gamma\(g\)\.
We stress the epistemic status of these components\. The mapping between coarse\-graining in deep networks and the renormalization group has a rigorous basis in specific cases—an exact correspondence between variational RG and restricted Boltzmann machines\[[12](https://arxiv.org/html/2607.11950#bib.bib11)\], information\-theoretic real\-space RG learned by neural networks\[[8](https://arxiv.org/html/2607.11950#bib.bib13)\], normalizing\-flow RG\[[10](https://arxiv.org/html/2607.11950#bib.bib12)\], and the hierarchical structure that makes “cheap” deep learning effective\[[11](https://arxiv.org/html/2607.11950#bib.bib14)\]\. Our architecture operationalizes the same intuition as a set of soft, differentiable biases \(bounded coupling flow, scale\-weighted attention, entropy\-gated synchronization, and a Callan–Symanzik drift penalty\) rather than as an exact RG transformation\. Whether those biases help on real neural data is the empirical question of this paper\.
### II\.4Why sleep EEG matches the bias
The aperiodic exponentβ\\betaof the cortical power spectrum is a genuine scaling exponent, and sleep depth shifts it: deep non\-REM sleep has the steepest1/fβ1/f^\{\\beta\}spectrum \(Section[III](https://arxiv.org/html/2607.11950#S3)\)\. Two consequences follow\. First, the 5\-class sleep\-staging label is partly a coarse read\-out ofβ\\beta, so a model that represents scale explicitly has something real to represent\. Second,β\\betagives an external, physically\-meaningful referent for the model’s learned scale\-sensitivity parameterγ\\gamma: if the RG bias is doing what it claims, the learnedγ\\gammashould carry information about the measuredβ\\beta\. The vanilla transformer has no comparable internal quantity, so this interpretability test is unique to RG\-Flow—and is reported \(Section[V](https://arxiv.org/html/2607.11950#S5)\) regardless of whether it improves accuracy\.
## IIIMethods
### III\.1Data: Sleep\-EDF
We use the PhysioNet Sleep\-EDF Expanded corpus \(Sleep Cassette study\)\[[7](https://arxiv.org/html/2607.11950#bib.bib7),[4](https://arxiv.org/html/2607.11950#bib.bib8)\], whole\-night polysomnography with expert hypnogram annotations\. This draft uses55subjects \(night 1\)\. From each recording we take the two EEG derivations \(Fpz–Cz and Pz–Oz\) at 100 Hz, band\-pass filter to 0\.3–35 Hz, crop to the annotated sleep period \(±30\\pm 30min of wake\) to remove long lights\-on padding, and epoch into 30 s windows aligned to the hypnogram\. AASM stages are mapped to five classes \(W, N1, N2, N3 with S3\+S4 merged, REM\); we additionally define a coarse two\-class high\-β\\beta\(deep NREM, N2/N3\) versus low\-β\\betasplit that mirrors the spectral\-depth axis\. Every epoch carries its subject identifier for leakage\-free grouping\.
### III\.2Spectral exponentβ\\beta
For each epoch we estimate the aperiodic exponentβ\\betafrom the Welch power spectral density of the Fpz–Cz channel using the specparam \(FOOOF\) spectral\-parametrization model in fixed\-knee\-free mode over 1–35 Hz\[[2](https://arxiv.org/html/2607.11950#bib.bib6)\]\. Fits belowR2=0\.90R^\{2\}=0\.90are excluded from the regression target;97%97\\%of epochs pass this cut\. The resultingβ\\betaincreases monotonically with sleep depth—medianβ\\betarises from1\.581\.58\(N1\) and1\.671\.67\(wake\) through1\.841\.84\(REM\) to2\.432\.43\(N2\) and3\.173\.17\(N3\)—confirming it as a real, label\-linked scale exponent \(Fig\.[2](https://arxiv.org/html/2607.11950#S4.F2)\): the deep\-NREM spectrum is markedly steeper than the light stages, exactly the spectral\-depth axis the coarse two\-class split targets\. It serves as both a continuous regression target and the referent for the learnedγ\\gamma\.
### III\.3The Bi\-Scale Block architecture
The RG\-Flow engine is a stack of Bi\-Scale Blocks\. The scale metricμℓ\\mu\_\{\\ell\}grows exponentially with block depthℓ\\ell\(μℓ=μ0ρℓ\\mu\_\{\\ell\}=\\mu\_\{0\}\\rho^\{\\ell\}\)\. Within each block the input representation𝐗ℓ−1\\mathbf\{X\}\_\{\\ell\-1\}is processed simultaneously by two parallel streams \(Fig\.[1](https://arxiv.org/html/2607.11950#S3.F1)\)\.
Input Hidden State:𝐗ℓ−1\\mathbf\{X\}\_\{\\ell\-1\}SYSTEM 1: ASSOCIATIVE\(Horizontal Attention\)Standard Multi\-HeadAttention \(MHA\)Contextual WeightDistributionSYSTEM 2: RG\-FLOW\(Vertical Attention\)Scale\-Aware Attention\(SAA\) Head \(μℓ−γ\\mu\_\{\\ell\}^\{\-\\gamma\}\)Beta\-Function Filterβ\(𝐠\)→0\\beta\(\\mathbf\{g\}\)\\to 0TOPOLOGICAL SYNC BRIDGE \(TSB\)Clamps System 1 drift using System 2 attractor \(H\>τH\>\\tau\)Layer Normalization \(LN\)Feed\-Forward Network \(GELU\)Output Hidden State:𝐗ℓ\\mathbf\{X\}\_\{\\ell\}Figure 1:The Bi\-Scale Block\. System 1 handles rapid contextual processing via standard multi\-head attention; System 2 is a scale\-aware stream in which attention scores are reweighted byμℓ−γ\\mu\_\{\\ell\}^\{\-\\gamma\}and filtered by the bounded coupling flow\. The Topological Sync Bridge fuses the streams when the System 1 attention entropy exceeds a thresholdτ\\tau, preventing high\-frequency drift from accumulating over deep layers\.System 1 \(associative attention\)computes standard contextual weights via multi\-head attention \(MHA\)\.System 2 \(scale\-aware RG flow\)penalizes query–key dot products by the observation scaleμℓ\\mu\_\{\\ell\}raised to the learnable anomalous dimensionγ\\gamma,
𝐒=𝐐𝐊⊤dk⋅μℓγ,\\mathbf\{S\}=\\frac\{\\mathbf\{Q\}\\mathbf\{K\}^\{\\\!\\top\}\}\{\\sqrt\{d\_\{k\}\}\\cdot\\mu\_\{\\ell\}^\{\\gamma\}\},\(5\)and a relevance filter𝐑=1−\|β\(𝐠\)\|\\mathbf\{R\}=1\-\|\\beta\(\\mathbf\{g\}\)\|acts as a soft noise gate\. The two streams merge through the Topological Sync Bridge, which computes the Shannon entropyHHof System 1’s attention distribution and, whenH\>τH\>\\tau, injects System 2’s stable macro\-attractor back into the stream\. Training uses a composite objective balancing cross\-entropy, theβ\\beta\-regression MSE, and a multi\-scale Callan–Symanzik drift penalty
LTotal=LCE\+wregLMSE\+λCS\(1N∑ℓ=1N𝕍ar\[ln\(𝐠ℓ\)−2γℓln\(μℓ\)\]\)\.L\_\{\\text\{Total\}\}=L\_\{\\text\{CE\}\}\+w\_\{\\text\{reg\}\}L\_\{\\text\{MSE\}\}\+\\lambda\_\{\\text\{CS\}\}\\left\(\\frac\{1\}\{N\}\\sum\_\{\\ell=1\}^\{N\}\\mathbb\{V\}\\mathrm\{ar\}\\\!\\left\[\\ln\(\\mathbf\{g\}\_\{\\ell\}\)\-2\\gamma\_\{\\ell\}\\ln\(\\mu\_\{\\ell\}\)\\right\]\\right\)\.\(6\)
### III\.4Models
All models share a 1D convolutional patch stem \(patch length 30 samples\) that reduces each 3000\-sample epoch to 100 tokens, so attention is tractable and the front\-end is identical across models\. We compare three:
- •RG\-Flow— the full architecture of Section[II](https://arxiv.org/html/2607.11950#S2)\(163,850 parameters\)\.
- •Vanilla— a positional\-encoded transformer encoder over the same tokens, with width tuned to match RG\-Flow’s parameter count \(162,910 parameters\)\.
- •Ablation— RG\-Flow with the scale\-aware physics switched off \(RG physics disabled, no sync bridge\), leaving only the coarse\-graining hierarchy \(163,850 parameters\)\.
Each has a classification head \(5\-class stage\) and a regression head \(continuousβ\\beta\)\. The near\-identical parameter counts make “does the RG bias help” separable from “does capacity help”; the ablation further separates the RG physics from the coarse\-graining hierarchy alone\.
### III\.5Evaluation protocol
We use leave\-one\-subject\-out group k\-fold: each fold holds out one whole subject for test and one for validation, training on the rest\. No subject ever appears in more than one split, so there is no by\-subject leakage\. Models are selected on validation loss \(early stopping\) and evaluated once on the held\-out test subject\. We report 5\-class accuracy, macro\-F1 \(robust to the strong class imbalance of sleep data\), and the out\-of\-sampleβ\\beta\-recoveryR2R^\{2\}, each as mean±\\pm95% CI across55seeds and all folds\.
### III\.6Scarce\-data budget sweep
To probe the inductive\-bias crossover, we cap the number of training epochs*per subject*at a budget∈\{50,100,200,all\}\\in\\\{50,100,200,\\text\{all\}\\\}and rerun the full k\-fold benchmark at each budget\. The hypothesis is that RG\-Flow’s scale prior helps most in the few\-epochs regime and that the gap closes as data grow; we report seed\-to\-seed stability as a first\-class metric alongside the mean\.
### III\.7Scope and reproducibility
The benchmark reported here is complete for the stated cohort: three models×\\times55seeds×\\timesfour data budgets×\\timesleave\-one\-subject\-out folds over55subjects, run to early stopping\. All results are deterministic under fixed random seeds, so extending the study to more subjects amounts to rerunning the same protocol on the larger cohort\. The cohort size \(55subjects\) is the binding limitation: it widens the confidence intervals in the scarce\-data regime and is the reason no single accuracy contrast at a data\-limited budget reaches significance \(Section[IV](https://arxiv.org/html/2607.11950#S4)\)\.
## IVResults
Figure 2:Per\-epoch aperiodic exponentβ\\betaon Sleep\-EDF, ordered by median\. Deep NREM \(N2/N3\) has the steepest1/fβ1/f^\{\\beta\}spectrum;97%97\\%of epochs pass theR2≥0\.90R^\{2\}\\geq 0\.90fit\-quality cut\.β\\betais the physically\-grounded referent for the learnedγ\\gamma\.### IV\.1Full\-budget benchmark
Table[1](https://arxiv.org/html/2607.11950#S4.T1)reports held\-out 5\-class sleep\-staging performance at the full per\-subject budget, mean±\\pm95% CI over55seeds and all leave\-one\-subject\-out folds, for the three parameter\-matched models \(RG\-Flow 163,850, Vanilla 162,910, Ablation 163,850 parameters\)\. RG\-Flow reaches 77\.3% accuracy \(macro\-F1 66\.5%\) versus Vanilla’s 77\.0% \(67\.5%\) and the ablation’s 76\.0% \(65\.0%\)\. The paired RG\-Flow−\-Vanilla accuracy difference is \+0\.4 percentage points \(Wilcoxonp=0\.294p=0\.294,dz=\+0\.07d\_\{z\}=\{\+0\.07\}\), and RG\-Flow−\-Ablation is \+1\.3 points \(p=0\.034p=0\.034\*\)\.
Table 1:Held\-out Sleep\-EDF benchmark, full budget \(mean±\\pm95% CI,55seeds, leave\-one\-subject\-out\)\. Stars mark paired Wilcoxon significance vs\. RG\-Flow\.
### IV\.2Scarce\-data budget sweep
Figure[3](https://arxiv.org/html/2607.11950#S4.F3)sweeps the per\-subject training budget\. The scarce\-data hypothesis predicts RG\-Flow’s scale prior should help most at the smallest budget \(50 epochs/subject\) and the gap should close as data grow\.*We do not observe this crossover\.*At the smallest budget the RG\-Flow−\-Vanilla accuracy difference is \-4\.4 points \(p=0\.072p=0\.072\)—that is, the vanilla transformer is numerically*ahead*in exactly the scarce regime where the inductive bias was expected to help, though the difference does not reach significance at five subjects\. The three models converge to a statistical tie at the full budget\. The point estimate therefore runs opposite to the hypothesis at every data\-limited budget, and no budget shows a significant RG\-Flow advantage\. We additionally report seed\-to\-seed accuracy stability \(right panel of Fig\.[3](https://arxiv.org/html/2607.11950#S4.F3)\); the three models are comparably stable, with the vanilla transformer marginally the most stable at the largest budgets\.
Figure 3:Data\-budget crossover\. Left: 5\-class accuracy vs training epochs per subject for each model \(mean±\\pm95% CI\)\. Right: seed standard deviation of accuracy \(lower = more stable\)\. The scarce\-data hypothesis predicts the largest RG\-Flow advantage at the left edge\.
### IV\.3β\\beta\-recovery
Beyond staging accuracy, RG\-Flow carries a regression head trained to recover the continuous spectral exponentβ\\beta\. Out\-of\-sample it reachesR2=0\.416R^\{2\}=0\.416versus 0\.403 for Vanilla \(pairedΔ=\+0\.013\\Delta=\{\+0\.013\},p=0\.542p=0\.542\)\. This is a physically\-meaningful, interpretable quantity; whether the model’s*internal*scale parameterγ\\gammacarries the same information is examined next\.
## VInterpretability: doesγ\\gammatrackβ\\beta?
The distinctive claim of the RG\-Flow Transformer is not only accuracy but*interpretability*: its learnable anomalous dimensionγ\\gammais a scalar, per\-head, per\-depth quantity that should—if the scale\-aware bias is doing what it is designed to—carry information about the physical scaling exponentβ\\betaof the input\. The vanilla transformer has no comparable internal quantity, so this analysis is unique to RG\-Flow\.
Figure[4](https://arxiv.org/html/2607.11950#S5.F4)relates the model’s learnedγ\\gamma\(averaged over layers\) to its out\-of\-sampleβ\\beta\-recoveryR2R^\{2\}across the budget sweep\. Read against the trained models \(100100RG\-Flow runs\), three observations follow\.
First,γ\\gammadeparts from its zero initialization but stays small in magnitude: across all runs\|γ\|\|\\gamma\|never exceeds0\.0870\.087, so the scale\-aware stream applies a gentle correction rather than a dominant reweighting of attention\. Second, the sign of that correction is budget\-dependent: in the scarce regimeγ\\gammais slightly positive \(mean\+0\.011\{\+0\.011\}at5050epochs/subject\) and it turns negative as data grow \(mean−0\.025\{\-0\.025\}at the full budget\)\. Third—and most informative—the learnedγ\\gammaco\-varies with the interpretability payoff: more\-negativeγ\\gammaaccompanies higher out\-of\-sampleβ\\beta\-recovery, a statistically clear association across runs \(Pearsonr=−0\.36r=\{\-0\.36\},pp<0\.001,n=100n=100\)\. The internal scale parameter is therefore not inert; it moves with the very external spectral exponent the architecture was designed to track\.
Figure 4:RG\-Flow’s learnedγ\\gamma\(mean over layers\) versus its out\-of\-sampleβ\\beta\-recoveryR2R^\{2\}, across the data\-budget sweep\. A systematic relationship would indicate the internal scale parameter tracks the measured spectral exponent—the interpretability payoff the vanilla model cannot provide\.We stress that this interpretability result stands independently of the accuracy comparison\. Even where RG\-Flow and the vanilla transformer are statistically indistinguishable on staging accuracy, only RG\-Flow exposes a physically\-readable scale parameter, and only RG\-Flow can be asked whether that parameter has tracked the data’s spectral exponent\.
## VIConclusion
We built a leakage\-free, real\-data test of whether an explicit renormalization\-group inductive bias helps a transformer on scarce neural signals, using PhysioNet Sleep\-EDF sleep staging as the task and the aperiodic spectral exponentβ\\betaas both a label axis and an interpretability referent\. The contribution is a fixed, reproducible protocol: a preprocessing front\-end, a per\-epochβ\\betaestimator, a parameter\-matched three\-way comparison \(RG\-Flow, vanilla, hierarchy\-only ablation\), a data\-budget sweep, and aγ\\gamma\-vs\-β\\betainterpretability probe\.
Three findings hold on this cohort \(55subjects,55seeds, leave\-one\-subject\-out\)\. First, the machinery works end\-to\-end on real EEG:β\\betais a clean, label\-linked scale exponent \(deep NREM steepest,97%97\\%good fits\), and all three models stage sleep at accuracies in the high\-70%70\\%range\. Second, and against the study’s own hypothesis, the scarce\-data crossover does not appear: RG\-Flow does not pull ahead when data are few—the vanilla transformer is numerically ahead at every data\-limited budget—and the models tie at the full budget \(RG\-Flow−\-Vanilla accuracy\+0\.4\{\+0\.4\}pp,p=0\.294p=0\.294\)\. The RG physics is moreover indistinguishable from its hierarchy\-only ablation, so what little structure the bias adds is attributable to coarse\-graining, not to the scale\-aware flow\. Third, the interpretability claim survives independently of accuracy: RG\-Flow recovers the continuous spectral exponent out\-of\-sample \(R2=0\.416R^\{2\}=0\.416\), its learned anomalous dimensionγ\\gammamoves systematically with data budget, and more\-negativeγ\\gammaaccompanies higherβ\\beta\-recovery \(Pearsonr=−0\.36r=\{\-0\.36\},pp<0\.001\)—a physically\-readable internal quantity the vanilla architecture has no analogue for\.
The binding limitation is cohort size\. With55held\-out subjects the scarce\-regime confidence intervals are wide, and while the point estimates run consistently against the inductive\-bias hypothesis, none reaches significance there; a larger cohort is needed to decide whether the negative scarce\-data result is real or merely underpowered\. Because the analysis is deterministic under fixed random seeds, that larger study is a rerun of the same protocol on a wider cohort, leaving the surrounding argument unchanged\. On the present evidence, the honest verdict is that an explicit renormalization\-group bias does not improve scarce\-data sleep staging, but it does buy a physically\-meaningful, interpretable readout at no accuracy cost—which is where we would direct future work on this architecture for neural data\.
## References
- \[1\]J\. M\. Beggs and D\. Plenz\(2003\)Neuronal avalanches in neocortical circuits\.Journal of Neuroscience23\(35\),pp\. 11167–11177\.External Links:[Document](https://dx.doi.org/10.1523/JNEUROSCI.23-35-11167.2003)Cited by:[§I](https://arxiv.org/html/2607.11950#S1.p1.4),[§II\.1](https://arxiv.org/html/2607.11950#S2.SS1.p1.3)\.
- \[2\]T\. Donoghue, M\. Haller, E\. J\. Peterson,et al\.\(2020\)Parameterizing neural power spectra into periodic and aperiodic components\.Nature Neuroscience23\(12\),pp\. 1655–1665\.External Links:[Document](https://dx.doi.org/10.1038/s41593-020-00744-x)Cited by:[§III\.2](https://arxiv.org/html/2607.11950#S3.SS2.p1.11)\.
- \[3\]R\. Gao, E\. J\. Peterson, and B\. Voytek\(2017\)Inferring synaptic excitation/inhibition balance from field potentials\.NeuroImage158,pp\. 70–78\.External Links:[Document](https://dx.doi.org/10.1016/j.neuroimage.2017.06.078)Cited by:[§I](https://arxiv.org/html/2607.11950#S1.p1.4)\.
- \[4\]A\. L\. Goldberger, L\. A\. N\. Amaral, L\. Glass,et al\.\(2000\)PhysioBank, physiotoolkit, and physionet: components of a new research resource for complex physiologic signals\.Circulation101\(23\),pp\. e215–e220\.External Links:[Document](https://dx.doi.org/10.1161/01.CIR.101.23.e215)Cited by:[§III\.1](https://arxiv.org/html/2607.11950#S3.SS1.p1.4)\.
- \[5\]B\. J\. He\(2014\)Scale\-free brain activity: past, present, and future\.Trends in Cognitive Sciences18\(9\),pp\. 480–487\.External Links:[Document](https://dx.doi.org/10.1016/j.tics.2014.04.003)Cited by:[§I](https://arxiv.org/html/2607.11950#S1.p1.4),[§II\.1](https://arxiv.org/html/2607.11950#S2.SS1.p1.3)\.
- \[6\]L\. P\. Kadanoff\(1966\)Scaling laws for ising models nearTcT\_\{c\}\.Physics Physique Fizika2\(6\),pp\. 263–272\.External Links:[Document](https://dx.doi.org/10.1103/PhysicsPhysiqueFizika.2.263)Cited by:[§I](https://arxiv.org/html/2607.11950#S1.p2.1)\.
- \[7\]B\. Kemp, A\. H\. Zwinderman, B\. Tuk, H\. A\. C\. Kamphuisen, and J\. J\. L\. Oberyé\(2000\)Analysis of a sleep\-dependent neuronal feedback loop: the slow\-wave microcontinuity of the eeg\.IEEE Transactions on Biomedical Engineering47\(9\),pp\. 1185–1194\.External Links:[Document](https://dx.doi.org/10.1109/10.867928)Cited by:[§III\.1](https://arxiv.org/html/2607.11950#S3.SS1.p1.4)\.
- \[8\]M\. Koch\-Janusz and Z\. Ringel\(2018\)Mutual information, neural networks and the renormalization group\.Nature Physics14\(6\),pp\. 578–582\.External Links:[Document](https://dx.doi.org/10.1038/s41567-018-0081-4)Cited by:[§I](https://arxiv.org/html/2607.11950#S1.p2.1),[§II\.3](https://arxiv.org/html/2607.11950#S2.SS3.p2.1)\.
- \[9\]J\. D\. Lendner, R\. F\. Helfrich, B\. A\. Mander,et al\.\(2020\)An electrophysiological marker of arousal level in humans\.eLife9,pp\. e55092\.External Links:[Document](https://dx.doi.org/10.7554/eLife.55092)Cited by:[§I](https://arxiv.org/html/2607.11950#S1.p1.4)\.
- \[10\]S\. Li and L\. Wang\(2018\)Neural network renormalization group\.Physical Review Letters121\(26\),pp\. 260601\.External Links:[Document](https://dx.doi.org/10.1103/PhysRevLett.121.260601)Cited by:[§I](https://arxiv.org/html/2607.11950#S1.p2.1),[§II\.3](https://arxiv.org/html/2607.11950#S2.SS3.p2.1)\.
- \[11\]H\. W\. Lin, M\. Tegmark, and D\. Rolnick\(2017\)Why does deep and cheap learning work so well?\.Journal of Statistical Physics168\(6\),pp\. 1223–1247\.External Links:[Document](https://dx.doi.org/10.1007/s10955-017-1836-5)Cited by:[§I](https://arxiv.org/html/2607.11950#S1.p2.1),[§II\.3](https://arxiv.org/html/2607.11950#S2.SS3.p2.1)\.
- \[12\]P\. Mehta and D\. J\. Schwab\(2014\)An exact mapping between the variational renormalization group and deep learning\.arXiv preprint arXiv:1410\.3831\.Cited by:[§I](https://arxiv.org/html/2607.11950#S1.p2.1),[§II\.3](https://arxiv.org/html/2607.11950#S2.SS3.p2.1)\.
- \[13\]M\. A\. Muñoz\(2018\)Colloquium: criticality and dynamical scaling in living systems\.Reviews of Modern Physics90\(3\),pp\. 031001\.External Links:[Document](https://dx.doi.org/10.1103/RevModPhys.90.031001)Cited by:[§I](https://arxiv.org/html/2607.11950#S1.p1.4),[§II\.1](https://arxiv.org/html/2607.11950#S2.SS1.p1.3)\.
- \[14\]K\. G\. Wilson and J\. Kogut\(1974\)The renormalization group and theϵ\\epsilonexpansion\.Physics Reports12\(2\),pp\. 75–199\.External Links:[Document](https://dx.doi.org/10.1016/0370-1573%2874%2990023-4)Cited by:[§I](https://arxiv.org/html/2607.11950#S1.p2.1),[item 1](https://arxiv.org/html/2607.11950#S2.I1.i1.p1.7)\.Similar Articles
Energy-Gated Attention and Wavelet Positional Encoding: Complementary Inductive Biases for Transformer Attention
This paper proposes Energy-Gated Attention (EGA) and Morlet Positional Encoding (MoPE) to address missing inductive biases in transformer attention: token salience and scale-adaptive locality. Experiments on TinyShakespeare show superadditive gains when combined, highlighting complementarity.
Forecasting Medium-Horizon Alzheimer's Disease Progression: Residual Gap-Aware Transformers for 24-Month CDR-SB Change from ADNI Clinical and Biomarker Histories
This paper proposes a residual gap-aware transformer that combines a mixed-effects statistical reference with transformer-based residual learning to forecast 24-month CDR-SB change from ADNI clinical and biomarker histories, achieving reduced MSE and improved correlation over baselines.
Generative modeling with sparse transformers
OpenAI introduces the Sparse Transformer, a deep neural network that improves the attention mechanism from O(N²) to O(N√N) complexity, enabling modeling of sequences 30x longer than previously possible across text, images, and audio. The model uses sparse attention patterns and checkpoint-based memory optimization to train networks up to 128 layers deep, achieving state-of-the-art performance across multiple domains.
Spatiotemporal Graph Transformer for Traffic Intelligence in Edge Computing
This paper proposes a spatiotemporal graph Transformer framework for traffic forecasting in edge computing, combining graph neural networks for spatial correlations and Transformer self-attention for long-range temporal dependencies. Experiments on real-world cellular data show it outperforms recurrent graph-based baselines like GCN-LSTM and GCN-GRU.
SAGE: Surrogate-gradient Adaptation via Attention-Guided Entropy for Spiking Transformers
The paper presents SAGE, a method that adapts surrogate gradients for Spiking Transformers using attention-derived entropy to improve training accuracy, demonstrated on CIFAR-10/100 datasets.