Why SWAVE May Not Be All You Need:A Concept-Evolution Retrospective on Complex-Valued Recurrent Language Models

arXiv cs.LG Papers

Summary

This paper presents a retrospective on the design evolution of SWave, a complex-valued recurrent language model, detailing which architectural components were retained, reframed, superseded, or proved non-load-bearing, along with formal characterizations of failure modes like cos-domination collapse.

arXiv:2606.18324v1 Announce Type: new Abstract: SWave is a complex-valued recurrent language model (169.26M parameters, D=384, L=16, T=2048) trained on FineWeb-Edu using 2xH100 NVL. It was designed around three founding premises: that representing language as complex waves rather than real-valued numbers enables richer information encoding; that a Cayley-parameterised unitary transition provides a mathematical guarantee against state decay or explosion; and that a hidden state which rotates rather than shrinks preserves signal integrity over arbitrarily long contexts. The core of SWave evolved substantially across three development phases. The Resonance Head was found to structurally admit imaginary-channel collapse as a global loss minimum (a failure mode we term cos-domination collapse) and was superseded by an untied head with independent real and imaginary embedding tables from the Phase-Associative Memory (PAM) architecture. This resolved the degenerate minimum and enabled stable 200,000-step training (best-step PPL 22.0 at step 89,861). ComplexNorm and the Wave Propagation Scan proved load-bearing throughout all three phases and were retained to the final architecture. ProtectGatedScan was reframed as a structural prior rather than a learned behaviour. The four multi-scale retention concepts showed no measurable improvement under controlled evaluation and were found non-load-bearing. The ComplexGatedUnit was superseded by a real-valued squared-ReLU channel mixer with fewer parameters. The auxiliary training objectives showed no benefit once structural constraints were resolved. The investigation yields a formal characterisation of cos-domination collapse, a parallel scan with a log-space backward pass for numerical stability, six transferable engineering principles for complex-valued recurrent training, and a plan-to-code traceability methodology for catching structural divergences that conventional test suites miss.
Original Article
View Cached Full Text

Cached at: 06/18/26, 05:41 AM

# Why SWave May Not Be All You Need: A Concept-Evolution Retrospective on Complex-Valued Recurrent Language Models
Source: [https://arxiv.org/html/2606.18324](https://arxiv.org/html/2606.18324)
Swathika N EdgeVerve Systems Limited swathika\.n@edgeverve\.comSahil Dilip Panse EdgeVerve Systems Limited SahilDilip\_Panse@edgeverve\.com

###### Abstract

SWaveis a complex\-valued recurrent language model \(169\.26M parameters,D=384D=384,L=16L=16,T=2048T=2048\) trained on FineWeb\-Edu using2×2\\timesH100 NVL\. It was designed around three founding premises: that representing language as complex waves rather than real\-valued numbers enables richer information encoding; that a Cayley\-parameterised unitary transition provides a mathematical guarantee against state decay or explosion; and that a hidden state which*rotates*rather than shrinks preserves signal integrity over arbitrarily long contexts\. The core ofSWaveevolved substantially across three development phases\. The Resonance Head was found to structurally admit imaginary\-channel collapse as a global loss minimum \(a failure mode we term*cos\-domination collapse*\) and was superseded by an untied head with independent real and imaginary embedding tables drawn from the Phase\-Associative Memory \(PAM\) architecture\(Vishwakarma et al\.,[2026](https://arxiv.org/html/2606.18324#bib.bib15)\)\. This resolved the degenerate minimum and enabled stable 200,000\-step training \(best\-step PPL 22\.0 at step 89,861\)\. ComplexNorm and the Wave Propagation Scan proved load\-bearing throughout all three phases and were retained to the final architecture\. ProtectGatedScan was reframed as a structural prior rather than a learned behaviour\. The four multi\-scale retention concepts, despite their design motivation, did not produce differential cross\-entropy improvement under controlled evaluation and were found to be non\-load\-bearing\. The ComplexGatedUnit was superseded by a real\-valued squared\-ReLU channel mixer, which achieved equivalent performance with fewer parameters\. The auxiliary training objectives did not demonstrate measurable benefit once the structural constraints they were designed to compensate were resolved at the source\.

The investigation yields a formal characterisation of cos\-domination collapse, a parallelisable scan with a log\-space backward pass for numerical stability, six transferable engineering principles for complex\-valued recurrent training, and a plan\-to\-code traceability methodology for catching structural divergences that conventional test suites miss\. The documented concept lifecycle \(what was retained, what was reframed, what was superseded, and what proved non\-load\-bearing\) provides a reference case for future complex\-valued model design\.

Keywords:complex\-valued recurrent neural networks; language modelling; loss landscape analysis; phase\-associative memory; architecture retrospective; design\-concept lifecycle\.

## 1Introduction

Transformer\-based language models face two well\-known scalability constraints\. First, the attention mechanism incursO​\(N2\)O\(N^\{2\}\)computation and anO​\(N\)O\(N\)\-memory KV cache that grows linearly with context length, making very long sequences economically prohibitive\. Second, the linear\-recurrence alternatives \(RWKV, Mamba/S4\) that address the quadratic cost do so through exponential decay: the state contracts at each step, so early\-context information fades over long sequences\.SWavewas designed to escape both constraints simultaneously, matching RWKV’sO​\(N\)O\(N\)training cost andO​\(1\)O\(1\)inference memory while eliminating the decay that causes SSMs to forget\.

Recurrent sequence models with complex\-valued hidden states offer a theoretically motivated route toO​\(1\)O\(1\)\-memory inference with norm\-preserving long\-range state retention\(Arjovsky et al\.,[2016](https://arxiv.org/html/2606.18324#bib.bib1); Wisdom et al\.,[2016](https://arxiv.org/html/2606.18324#bib.bib16)\)\. The associative\-memory interpretation draws on complex Hopfield networks\(Noest,[1992](https://arxiv.org/html/2606.18324#bib.bib7)\)and Holographic Reduced Representations\(Plate,[1995](https://arxiv.org/html/2606.18324#bib.bib11)\)\. The unit\-magnitude constraintht=ei​φ​ht−1\+xth\_\{t\}=e^\{i\\varphi\}h\_\{t\-1\}\+x\_\{t\}preserves‖h‖\\\|h\\\|exactly across arbitrarily long sequences, a property that real\-valued recurrences can only approximate through careful regularisation\. Building on this foundation,SWaveset out to explore what a full\-featured complex\-valued language model would look like: one that could match Transformer\-scale training while retainingO​\(1\)O\(1\)memory per sequence position\.

##### Founding premises\.

The design was motivated by three core ideas\.Wave\-based processing: tokens are embedded as complex numbers, so the hidden state carries both amplitude and phase, enabling richer encoding than real\-valued representations\.Cayley unitary memory: the state transition is Cayley\-parameterised to enforce\|αt\|=1\|\\alpha\_\{t\}\|=1, providing a formal guarantee that state energy neither decays nor explodes across long sequences\.Bounded\-decay context: because the state*rotates*with bounded decay rather than unconstrained exponential shrinkage at each step, early\-context signals are attenuated far less than in standard real\-valued recurrences, where decay is unbounded\. Sixteen design concepts were developed to realise this vision; the paper documents what each became\.

##### What we tried\.

Sixteen design concepts were developed across six groups: output head \(Resonance Head, Wave Embedding\), state dynamics \(Wave Propagation Scan, AmplitudeGate, Unitary Rotation, Cayley Transform, ProtectGatedScan\), channel mixing \(ComplexGatedUnit, WaveMixer\), normalisation \(ComplexNorm\), multi\-scale retention \(Wavelet State Hierarchy, Phase Bus, Echo Memory, Resonant Router\), and diagnostic infrastructure \(Orthogonal Initialisation, Wave Diagnostics, Wave Rewind\)\. Each concept is presented with its hypothesis and outcome in Section[2](https://arxiv.org/html/2606.18324#S2)\. Development proceeded in three phases \(Table[1](https://arxiv.org/html/2606.18324#S1.T1)\)\.

##### Development phases\.

Phase 1 \(Original Idea\) established the design concepts and the tied resonance head architecture\. Phase 2 \(PAM Baseline\) resolved a structural issue in the output head by adoptingPAMprimitives \(Phase 2 is a near\-direct architectural adoption ofPAMrather than an independent invention;SWave’s specific contribution at this phase is empirical validation at 169M parameters and a vector\-state rather than matrix\-state design\), ran a 200k\-step training run confirming stability, and refined the scan, normalisation, and gradient monitoring infrastructure\. Phase 3 \(Integration\) brought the Phase 2 core back into the concept\-rich Phase 1 architecture, evaluated each retention concept on a stable foundation, and progressively replaced Phase 2 primitives withMamba/RWKVload\-bearing equivalents alongside a falsifier\-driven development methodology\.

Table 1:Three development phases ofSWave\.![Refer to caption](https://arxiv.org/html/2606.18324v1/x1.png)Figure 1:Cross\-entropy loss during training across the three development phases\.\(a\) Phase 1 \(Original Idea\):The tied resonance head produces unstable training that never exceeds CE 7\.29 \(PPL 1471\) before diverging to CE 25\.3 by step 5,850, a signature of cos\-domination collapse\.\(b\) Phase 2 \(PAM Baseline\):The untied head resolves the degenerate minimum; training runs stably for 200,000 steps, reaching best CE 3\.09 \(PPL 22\.0\)\.\(c\) Phase 3 \(Integration\):The stable Phase 2 core integrated into the original architecture reaches best CE 2\.75 \(PPL 15\.6\), with noisier dynamics reflecting the more complex configuration\. Shaded curves are raw logs; solid curves are smoothed\.All results use a single model configuration:D=384D=384,L=16L=16,T=2048T=2048,V=100,277V=100\{,\}277\(cl100k\_base tokeniser\), 169\.26M parameters, trained on FineWeb\-Edu\(Penedo et al\.,[2024](https://arxiv.org/html/2606.18324#bib.bib9)\)on2×2\\timesNVIDIA H100 NVL\.

##### Verdict taxonomy\.

Each design concept is assigned one of five verdicts based on pre\-specified quantitative criteria \(Table[2](https://arxiv.org/html/2606.18324#S1.T2)\)\.

Table 2:Verdict taxonomy applied to all design concepts\.
##### Scope\.

This paper does not claim a new state\-of\-the\-art architecture\. The value lies in the documented analysis of why specific design decisions evolved as they did, and the transferable methodology for identifying structural divergences before they compound across multiple training runs\.

## 2Architecture

SWaveis a complex\-valued recurrent language model\. Each hidden state in Phase 2 is a complex vectorz∈ℂDz\\in\\mathbb\{C\}^\{D\}, stored as a pair of real tensors\(zr,zi\)∈ℝD×ℝD\(z^\{r\},z^\{i\}\)\\in\\mathbb\{R\}^\{D\}\\times\\mathbb\{R\}^\{D\}\. The Phase 2 forward pass at each layer applies a sequence mixer \(ProtectGatedScan\) followed by a channel mixer \(ComplexGatedUnit\), withComplexNormin a sandwich arrangement around each module\. Phase 3 integrated these Phase 2 components into the original Phase 1 architecture and subsequently replaced all four core primitives; Section[3\.2](https://arxiv.org/html/2606.18324#S3.SS2)documents those changes\.

Each design concept below is presented with its original hypothesis, followed by a unified account of how it evolved across the three phases, and a final verdict\.

### 2\.1Output Head

The output head maps the complex hidden state to a vocabulary distribution\.

#### 2\.1\.1Resonance Head \(Tied Phase Head\)

##### Hypothesis\.

Each vocabulary itemvvis assigned a learnable phaseθv∈ℝ\\theta\_\{v\}\\in\\mathbb\{R\}; the logit is:

ℓv​\(hr,hi;θv\)=cos⁡\(θv\)​hr\+sin⁡\(θv\)​hi=cos⁡\(θh−θv\),\\ell\_\{v\}\(h\_\{r\},h\_\{i\};\\,\\theta\_\{v\}\)=\\cos\(\\theta\_\{v\}\)\\,h\_\{r\}\+\\sin\(\\theta\_\{v\}\)\\,h\_\{i\}=\\cos\(\\theta\_\{h\}\-\\theta\_\{v\}\),\(1\)where the final equality holds whenh=ei​θhh=e^\{i\\theta\_\{h\}\}is unit\-norm\. Vocabulary retrieval was conceived as a phase\-alignment operation: tokens whose learned phase is close to the hidden state’s current phase score high, analogous to resonance in a tuned oscillator\.

##### Journey\.

The first training run \(Phase 1\) reproduced a predicted failure mode at step 2,000: the cos/sin term ratio at the output head reached245×245\\times, vocabulary phases were essentially frozen \(θdrift,mean=6\.7×10−4\\theta\_\{\\mathrm\{drift,mean\}\}=6\.7\\times 10^\{\-4\}rad\), and cross\-entropy showed no monotone descent over any 500\-step window\. Formal analysis established why this is unavoidable under the tied parameterisation: for any token distribution expressible byhrh\_\{r\}alone, the configuration\(hr,hi=0\)\(h\_\{r\},\\,h\_\{i\}=0\)attains the same cross\-entropy minimum as anyhi≠0h\_\{i\}\\neq 0configuration, because the gradient∂ℒCE/∂hi=∑vsin⁡\(θv\)​\(pv−yv\)\\partial\\mathcal\{L\}\_\{\\mathrm\{CE\}\}/\\partial h\_\{i\}=\\sum\_\{v\}\\sin\(\\theta\_\{v\}\)\(p\_\{v\}\-y\_\{v\}\)vanishes assin⁡\(θv\)→0\\sin\(\\theta\_\{v\}\)\\to 0for dominant vocabulary items\. The tied constraintcos2⁡θv\+sin2⁡θv=1\\cos^\{2\}\\theta\_\{v\}\+\\sin^\{2\}\\theta\_\{v\}=1makes thehi≡0h\_\{i\}\\equiv 0sub\-manifold a global, not local, minimum\. Five successive interventions targeting optimisation dynamics confirmed that the failure was architectural rather than optimisation\-dependent: gradient clipping, warmup adjustment, learning rate reduction, auxiliary losses, and initialisation changes each reproduced the245×245\\timessignature without altering the loss landscape geometry\.

The structural resolution came fromPAM\(Vishwakarma et al\.,[2026](https://arxiv.org/html/2606.18324#bib.bib15)\)in Phase 2: replacing the tied head with two independent real matricesEr,Ei∈ℝV×DE\_\{r\},E\_\{i\}\\in\\mathbb\{R\}^\{V\\times D\}, initialised with𝒩​\(0,0\.022\)\\mathcal\{N\}\(0,0\.02^\{2\}\), so that the logit becomesℓv=Er​\[v\]⊤​hr\+Ei​\[v\]⊤​hi\\ell\_\{v\}=E\_\{r\}\[v\]^\{\\top\}h^\{r\}\+E\_\{i\}\[v\]^\{\\top\}h^\{i\}\. This drops the head\-term initialisation ratio from775×775\\timesto≈1\.0×\{\\approx\}1\.0\\times, giving both channels equal logit variance from step 0\. The 200k\-step training run confirmed structural stability:ρ=RMS​\(hi\)/RMS​\(hr\)∈\[0\.79,1\.22\]\\rho=\\mathrm\{RMS\}\(h\_\{i\}\)/\\mathrm\{RMS\}\(h\_\{r\}\)\\in\[0\.79,1\.22\]throughout\. In Phase 3, aPhaseAttentionHeadvariant with three learned bridge projections\(Wρ,Wφr,Wφi\)\(W\_\{\\rho\},W\_\{\\varphi^\{r\}\},W\_\{\\varphi^\{i\}\}\)was developed and verified via plan\-to\-code audit; the plain untied head was confirmed as the production default\.

![Refer to caption](https://arxiv.org/html/2606.18324v1/x2.png)Figure 2:Cos\-domination collapse signature in Phase 1 training logs\.\(a\)The phase\-parameter gradient norm‖∇φ‖\\\|\\nabla\\varphi\\\|\(red\) exceeds the embedding gradient norm‖∇We‖\\\|\\nabla W\_\{e\}\\\|\(dark\) by orders of magnitude throughout the early steps, indicating the loss surface is almost entirely shaped by the phase parameters\.\(b\)The phase\-to\-embedding gradient ratio peaks at728×728\\timesat step 50 and remains chronically elevated, confirming that the tied head’s loss landscape structurally driveshi→0h\_\{i\}\\to 0regardless of auxiliary rescues\. The degenerate minimum is architectural, not optimisation\-level\.
##### Verdict\.

Superseded\. The tied resonance head is superseded by the untiedPAMhead\. The cos\-domination collapse analysis is the primary finding of this work, empirically observed in training and formally characterised\.

#### 2\.1\.2Wave Embedding \(Tied Unit\-Circle Embedding\)

##### Hypothesis\.

Each tokenvvis embedded as a point on the unit circle:hr=cos⁡θvh^\{r\}=\\cos\\theta\_\{v\},hi=sin⁡θvh^\{i\}=\\sin\\theta\_\{v\}, withcos2⁡θv\+sin2⁡θv=1\\cos^\{2\}\\theta\_\{v\}\+\\sin^\{2\}\\theta\_\{v\}=1enforced per \(token, channel\)\. The two components carry distinct semantic roles: amplitude encodes the importance or salience of a token, while phase encodes its semantic direction, with semantically related tokens expected to occupy nearby angles\. This wave\-like representation was expected to support constructive interference \(related concepts reinforcing each other in the hidden state\) and destructive interference \(noise and irrelevant signals cancelling\) through the phase relationships between complex\-valued representations\. Embeddings living on a phase manifold were also expected to align naturally with the resonance head’s phase\-matching retrieval\.

##### Journey\.

The unit\-circle constraint was conceived in Phase 1 as the input\-side complement to the resonance head\. The two concepts were tightly coupled: the constraint meant embeddings had magnitudes of exactly 1, which combined with the tied head’s cos/sin parameterisation to makehi≡0h\_\{i\}\\equiv 0a global rather than local loss minimum; the constraint removed the escape route that free magnitudes would have provided\. When the resonance head was replaced by the untied head in Phase 2, the unit\-circle constraint was simultaneously released: the independent tablesEr,Ei∈ℝV×DE\_\{r\},E\_\{i\}\\in\\mathbb\{R\}^\{V\\times D\}carry unconstrained magnitudes, and the structural pathology dissolved\. No change was made in Phase 3\.

##### Verdict\.

Withdrawn\. The amplitude/phase semantic framing \(amplitude encoding token salience, phase encoding semantic direction\) is a coherent design hypothesis that was not independently evaluated\. The unit\-circle magnitude constraint became structurally entangled with the tied resonance head: releasing it was part of the fix for that head, not a rejection of the embedding concept itself\. Whether free\-magnitude complex embeddings with explicit phase priors provide richer representations than standard real\-valued embeddings remains an open question\.

### 2\.2State Dynamics: The Recurrent Core

The heart ofSWaveis a first\-order linear recurrence over a complex hidden state\. Four design concepts shaped how that recurrence was gated, rotated, and implemented efficiently\.

#### 2\.2\.1Wave Propagation Scan

##### Hypothesis\.

A complex first\-order linear recurrenceht=At⊙ht−1\+xth\_\{t\}=A\_\{t\}\\odot h\_\{t\-1\}\+x\_\{t\}, withAtA\_\{t\}input\-dependent and parallelisable via the associative operator\(a1,b1\)⊕\(a2,b2\)=\(a1​a2,a2​b1\+b2\)\(a\_\{1\},b\_\{1\}\)\\oplus\(a\_\{2\},b\_\{2\}\)=\(a\_\{1\}a\_\{2\},\\;a\_\{2\}b\_\{1\}\+b\_\{2\}\)\.O​\(log⁡T\)O\(\\log T\)depth parallel prefix computation was expected to replace sequentialO​\(T\)O\(T\)depth, enabling efficient GPU training at full sequence length\.

##### Journey\.

The Phase 1 implementation used a sequential Python loop overT=2048T=2048timesteps, dispatching∼150,000\{\\sim\}150\{,\}000CUDA kernels per forward pass and achieving only4040–42%42\\%GPU utilisation due to dispatch overhead\. Phase 2 brought three successive improvements that together made the scan practical at scale: first, a chunked scan with chunk sizeC=64C=64reduced the loop fromTTto≈96\{\\approx\}96iterations; second, a closed\-form vectorisation of the intra\-chunk recurrence using cumulative products,

c​\[t\]\\displaystyle c\[t\]=∏s=1tr​\[s\]\(cumprod\),\\displaystyle=\\prod\_\{s=1\}^\{t\}r\[s\]\\quad\(\\texttt\{cumprod\}\),\(2\)h​\[t\]\\displaystyle h\[t\]=c​\[t\]⋅cumsum​\(x​\[t\]c​\[t\]\),\\displaystyle=c\[t\]\\cdot\\mathrm\{cumsum\}\\\!\\left\(\\frac\{x\[t\]\}\{c\[t\]\}\\right\),\(3\)reduced kernel launches per scan call from∼158\{\\sim\}158to66\(a26×26\\timesreduction\), recovering GPU utilisation from4040–42%42\\%to near\-peak; third, the scan body is formulated in log\-space so that all cumulative products stay in\(0,1\]\(0,1\], bounding the largest backward intermediate byexp⁡\(0\)=1\\exp\(0\)=1and avoiding fp32 overflow at the production sequence length and halflife range\. Log\-α\\alphais clamped atlog⁡\(1−10−5\)\\log\(1\-10^\{\-5\}\)to prevent decay saturation at long halflives, verified by a machine\-verifiable falsifier \(atan2gradient finite at origin\)\. Both requirements were identified during numerical analysis of the fp32 dynamic range atT=2048T=2048\. The associative scan operator carried forward into Phase 3 unchanged, generalised to the two\-stream setting ofTwoStreamScan\(Section[3\.2](https://arxiv.org/html/2606.18324#S3.SS2)\)\.

##### Verdict\.

Survived\. The associative scan operator is load\-bearing throughout all three phases; the parallelisation improvements are transferable to any first\-order linear recurrence\.

#### 2\.2\.2AmplitudeGate

##### Hypothesis\.

A standalone magnitude gate controlling write strength:gatet=σ​\(Wg⋅xt\)\\mathrm\{gate\}\_\{t\}=\\sigma\(W\_\{g\}\\cdot x\_\{t\}\),hgated=gatet⊙hth\_\{\\mathrm\{gated\}\}=\\mathrm\{gate\}\_\{t\}\\odot h\_\{t\}, with decay hardcoded at1\.01\.0\. Selective write suppression was expected to prevent uninformative inputs from polluting the hidden state\.

##### Journey\.

In Phase 1, the AmplitudeGate operated as a standalone post\-recurrence module\. Its core limitation became apparent: with decay fixed at1\.01\.0, the gate could suppress writes but could not decelerate state decay, so state preservation under uninformative inputs was incomplete\. Phase 2 addressed this with theProtectGatedScan, which unified the gating function into the recurrence itself via a dual\-purpose protect gatept=σ​\(Wp​\|zt\|\+bp\)p\_\{t\}=\\sigma\(W\_\{p\}\|z\_\{t\}\|\+b\_\{p\}\)withbp=−3\.0b\_\{p\}=\-3\.0:

γt\\displaystyle\\gamma\_\{t\}=e−Δ​tt​\(1−pt\)\+pt,\\displaystyle=e^\{\-\\Delta t\_\{t\}\}\(1\-p\_\{t\}\)\+p\_\{t\},\(4\)Vt′\\displaystyle V^\{\\prime\}\_\{t\}=Vt⋅\(1−pt\)\.\\displaystyle=V\_\{t\}\\cdot\(1\-p\_\{t\}\)\.\(5\)The sameptp\_\{t\}simultaneously suppresses the write \(Vt′V^\{\\prime\}\_\{t\}\) and decelerates decay \(γt\\gamma\_\{t\}\), so highptp\_\{t\}freezes the state rather than merely limiting the write, providing a strictly more complete form of state preservation than the original gate\. Thebp=−3\.0b\_\{p\}=\-3\.0bias initialisespt≈0\.047p\_\{t\}\\approx 0\.047, biasing toward preservation from step 0\. The ProtectGatedScan was itself superseded byTwoStreamScanin Phase 3, but the dual\-purpose gating insight carried forward into the new design\.

##### Verdict\.

Superseded\. AmplitudeGate’s conceptual contribution, explicit write suppression, was subsumed and generalised by the dual\-purpose protect gate\.

#### 2\.2\.3Unitary Rotation

##### Hypothesis\.

Step\-wise unitary rotation of the hidden state:

hnewr\\displaystyle h^\{r\}\_\{\\mathrm\{new\}\}=cos⁡φ​hr−sin⁡φ​hi\+xr,\\displaystyle=\\cos\\varphi\\,h^\{r\}\-\\sin\\varphi\\,h^\{i\}\+x^\{r\},\(6\)hnewi\\displaystyle h^\{i\}\_\{\\mathrm\{new\}\}=sin⁡φ​hr\+cos⁡φ​hi\+xi\.\\displaystyle=\\sin\\varphi\\,h^\{r\}\+\\cos\\varphi\\,h^\{i\}\+x^\{i\}\.\(7\)Exact unitary transitions were expected to prevent gradient vanishing/explosion\(Arjovsky et al\.,[2016](https://arxiv.org/html/2606.18324#bib.bib1)\)and encode positional information as a continuous rotation in the complex plane\.

##### Journey\.

The Phase 1 implementation applied the rotation as a standalone module at each recurrence step\. The additive\+x\+xterm entangled the rotation with the write operation, making the update difficult to interpret or analyse independently\. In Phase 2, the rotation was factored out of the recurrence and applied as a phase hook on the key tensorKtK\_\{t\}before the conjugate write; the formulacos⁡\(m​θd\)​zr−sin⁡\(m​θd\)​zi\\cos\(m\\theta\_\{d\}\)z^\{r\}\-\\sin\(m\\theta\_\{d\}\)z^\{i\}is mathematically identical but cleanly separated from the gating logic, and has a clear positional\-encoding interpretation analogous to RoPE\(Su et al\.,[2021](https://arxiv.org/html/2606.18324#bib.bib14)\)\. This factored form carried into Phase 3 as a structural prior in the log\-α\\alphamulti\-timescale spectrum\.

##### Verdict\.

Survived\(reframed\)\. The phase\-preserving constraint survives inComplexNormand in boundedγt\\gamma\_\{t\}; the original “zero decay” claim is reframed as “phase\-preserving with bounded decay,” which is what the implementation actually delivers\.

#### 2\.2\.4Cayley Transform for Unitary Matrices

##### Hypothesis\.

Parameterise unitary weight matrices via the Cayley map:U=\(I−A\)​\(I\+A\)−1U=\(I\-A\)\(I\+A\)^\{\-1\}, whereAAis skew\-symmetric\. The Cayley map covers the full unitary group without the rank deficiency of matrix exponential approximations, offering exact parameterisation of all unitary transformations\.

##### Journey\.

When evaluated in Phase 1, the Cayley map required a matrix inversion\(I\+A\)−1\(I\+A\)^\{\-1\}atD=384D=384, costingO​\(D3\)O\(D^\{3\}\)per forward pass \(∼57​M\{\\sim\}57MFLOPs atD=384D=384;∼68​B\{\\sim\}68BFLOPs atD=4096D=4096\) and exhibiting numerical sensitivity as\(I\+A\)\(I\+A\)approaches singular\. A diagonal approximation \(restrictingAAto a skew\-diagonal, i\.e\. scalar rotation angles per channel\) reduces this toO​\(D\)O\(D\)\(384384multiplications instead of57​M57M\), recovering the zero\-decay property at negligible cost; however, it also removes the inter\-channel mixing that full unitarity provides\. In Phase 2, the full\-matrix path was replaced by a spectral decompositionRm=V​diag​\(ei​ω​m\)​VHR^\{m\}=V\\,\\mathrm\{diag\}\(e^\{i\\omega m\}\)\\,V^\{H\}, computed viatorch\.linalg\.eighon the Hermitian generatorH=−i​KH=\-iK,K=\(A−AH\)/2K=\(A\-A^\{H\}\)/2, with exact unitarity guaranteed and fully differentiable through PyTorch\. However, across all phases the strict unitary constraint proved less useful in practice than the bounded decay with a learned protect gate, which achieves near\-unitary behaviour where appropriate without the computational overhead of a full matrix decomposition\.

##### Verdict\.

Not load\-bearing\. The spectral form captures the stability benefit of the Cayley formulation efficiently; the full\-matrix unitary constraint is not required to realise that benefit in the tested regime, pointing to the spectral approximation as the operative principle\.

### 2\.3Channel Mixing: The Per\-Position Nonlinearity

Between recurrent state updates, the model applies a per\-position nonlinear transformation\. Two design concepts shaped this\.

#### 2\.3\.1Channel Mixer \(SwiGLU FFN→\\toComplexGatedUnit\)

##### Hypothesis\.

The standard gated FFN fromShazeer \([2020](https://arxiv.org/html/2606.18324#bib.bib12)\):\(W1​x⊙SiLU​\(W2​x\)\)⋅W3\(W\_\{1\}x\\odot\\mathrm\{SiLU\}\(W\_\{2\}x\)\)\\cdot W\_\{3\}\. Using a well\-validated real\-valued nonlinearity was expected to allow borrowing published hyperparameter settings and provide a stable channel mixer from the outset\.

##### Journey\.

In Phase 1, applying SwiGLU to complexz=\(zr,zi\)z=\(z^\{r\},z^\{i\}\)by channel\-splitting \(treating real and imaginary parts as independent real vectors\) broke the complex multiplication structure and discarded the phase relationships the rest of the architecture was designed to preserve\. Phase 2 replaced SwiGLU with theComplexGatedUnit\(CGU\), a five\-step operation that works natively on complex inputs:

zup\\displaystyle z\_\{\\mathrm\{up\}\}=Wup​z,\\displaystyle=W\_\{\\mathrm\{up\}\}\\,z,\(8\)zact\\displaystyle z\_\{\\mathrm\{act\}\}=modReLU​\(zup\),\\displaystyle=\\mathrm\{modReLU\}\(z\_\{\\mathrm\{up\}\}\),\(9\)zϕ\\displaystyle z\_\{\\phi\}=Wϕ​z/\|Wϕ​z\|,\\displaystyle=W\_\{\\phi\}\\,z\\;/\\;\|W\_\{\\phi\}\\,z\|,\(10\)zrot\\displaystyle z\_\{\\mathrm\{rot\}\}=zact⊙zϕ,\\displaystyle=z\_\{\\mathrm\{act\}\}\\odot z\_\{\\phi\},\(11\)CGU​\(z\)\\displaystyle\\mathrm\{CGU\}\(z\)=Wdown​\(zrot⊙σ​\(\|Wg​z\|\)\),\\displaystyle=W\_\{\\mathrm\{down\}\}\\\!\\left\(z\_\{\\mathrm\{rot\}\}\\odot\\sigma\(\|W\_\{g\}z\|\)\\right\),\(12\)wheremodReLU​\(z\)=ReLU​\(\|z\|−b\)⋅z/\|z\|\\mathrm\{modReLU\}\(z\)=\\mathrm\{ReLU\}\(\|z\|\-b\)\\cdot z/\|z\|preserves phase while thresholding magnitude,zϕz\_\{\\phi\}is a unit\-phase gate rotating the activated state inℂ\\mathbb\{C\}, andσ​\(\|Wg​z\|\)\\sigma\(\|W\_\{g\}z\|\)is a real\-valued magnitude gate\. In Phase 3, as the hidden\-state carrier shifted from complex to real, CGU was in turn superseded byChannelMixBlock\(RWKV\-V5 channel mix\(Peng et al\.,[2023](https://arxiv.org/html/2606.18324#bib.bib8)\)\), which is structurally matched to the real carrier\.

##### Verdict\.

Superseded\. Each transition \(SwiGLU to CGU, CGU to ChannelMixBlock\) reflects a tightening alignment between the channel mixer and the carrier type\.

#### 2\.3\.2WaveMixer \(Token\-Shift Blend\)

##### Hypothesis\.

Blend current and previous tokens before the channel mixer:xk=μk⊙xcur\+\(1−μk\)⊙xprevx\_\{k\}=\\mu\_\{k\}\\odot x\_\{\\mathrm\{cur\}\}\+\(1\-\\mu\_\{k\}\)\\odot x\_\{\\mathrm\{prev\}\}, withμk,μv,μr\\mu\_\{k\},\\mu\_\{v\},\\mu\_\{r\}learnable per\-channel blend scalars and a gated output=σ​\(r\)×tanh⁡\(k\)×v=\\sigma\(r\)\\times\\tanh\(k\)\\times v, following the RWKV\-V4 token\-shift approach\(Peng et al\.,[2023](https://arxiv.org/html/2606.18324#bib.bib8)\)\. Learnable blend weights were expected to encode optimal per\-channel mixing as training progressed\.

##### Journey\.

Designed in Phase 1 as a dynamic blend mechanism, the WaveMixer was not active in the Phase 2 PAM\-aligned configuration\. When re\-introduced in Phase 3 with asymmetric initialisation\(μk,μv,μr\)=\(0\.3,0\.5,0\.7\)\(\\mu\_\{k\},\\mu\_\{v\},\\mu\_\{r\}\)=\(0\.3,0\.5,0\.7\)to break the symmetry of the three\-vector parametrisation, theμ\\muvectors remained bit\-identical to initialisation across 500 steps and all 16 blocks\. The mechanism did not learn; the asymmetric initialisation itself acted as a structural prior encoding a fixed blend rather than a learned one\. Phase 3 E7 promotion removed the token\-shift vectors from the default model graph\.

##### Verdict\.

Reframed→\\toSuperseded\.111The arrow notation indicates a concept that was initially reframed \(the asymmetric initialisation was understood as a structural prior\) and subsequently removed from the default path entirely, warranting the stronger Superseded verdict\.The dynamic blend hypothesis is not supported by the training evidence; the asymmetric initialisation functions as a static structural prior\. Removed from the default path\.

### 2\.4Normalisation: Protecting Phase Geometry

#### 2\.4\.1ComplexNorm

##### Hypothesis\.

Standard RMSNorm applied independently to real and imaginary parts would distort the phase relationship∠​z=atan2​\(zi,zr\)\\angle z=\\mathrm\{atan2\}\(z^\{i\},z^\{r\}\)by scaling the two components by different factors\.ComplexNormnormalises by the joint complex magnitude,

rms​\(z\)=meand​\(\|zd\|2\)\+ε,z~d=sd⋅zd/rms​\(z\),\\mathrm\{rms\}\(z\)=\\sqrt\{\\mathrm\{mean\}\_\{d\}\(\|z\_\{d\}\|^\{2\}\)\+\\varepsilon\},\\qquad\\tilde\{z\}\_\{d\}=s\_\{d\}\\cdot z\_\{d\}\\,/\\,\\mathrm\{rms\}\(z\),\(13\)with learnable per\-channel gains∈ℝDs\\in\\mathbb\{R\}^\{D\}\(init1\.01\.0\), preserving∠​z~d=∠​zd\\angle\\tilde\{z\}\_\{d\}=\\angle z\_\{d\}by construction\.

##### Journey\.

ComplexNorm was designed and implemented in Phase 1 as the normalisation layer for the complex carrier\. Its critical role became clearer in Phase 2, when engineering investigations revealed that unconstrained amplitude growth was driving phase\-gradient instability: at step 37,250, phase drift was negligible \(θdrift,mean=0\.00144\\theta\_\{\\mathrm\{drift,mean\}\}=0\.00144rad\) while the phase\-group gradient norm escalated to150150–290290, because phase gradients scale as\|h\|2\|h\|^\{2\}\. Promoting ComplexNorm to a sandwich arrangement, applied before each sublayer and after each residual add, bounded residual magnitudes at every layer\. Without this, geometric amplification of∼1\.5×\{\\sim\}1\.5\\times/layer compounds to1\.516≈656×1\.5^\{16\}\\approx 656\\timesacross a 16\-layer stack\. The sandwich\-norm pattern is established in Stable LM 3B and Gemma 2\(Gemma Team,[2024](https://arxiv.org/html/2606.18324#bib.bib2)\);SWave’s contribution is the quantitative diagnosis specific to the complex carrier, where phase\-gradient scaling as\|h\|2\|h\|^\{2\}makes amplitude control essential\. In Phase 3, when the hidden\-state carrier shifted from complex to real,ComplexNormwas replaced by standard real\-valuedRMSNormas the appropriate normalisation for the real carrier; the sandwich arrangement was retained throughout\.

##### Verdict\.

Survived\. Phase\-preserving normalisation is load\-bearing for the complex carrier; the sandwich arrangement is a transferable engineering principle retained across all three phases\.

### 2\.5Multi\-Scale Retention

Four design concepts were developed to giveSWavericher multi\-scale memory, enabling the model to simultaneously reason over short, medium, and long temporal contexts within a single recurrent pass\. All four were evaluated in the Phase 3 controlled ablation and found to be non\-load\-bearing\.

#### 2\.5\.1Wavelet State Hierarchy

##### Hypothesis\.

Maintain hidden states at three temporal strides\{1,4,16\}\\\{1,4,16\\\}and blend them with an input\-dependent softmax router:blend=softmax​\(W⋅\[\|hf\|,\|hm\|,\|hc\|\]\)\\mathrm\{blend\}=\\mathrm\{softmax\}\(W\\cdot\[\|h\_\{f\}\|,\|h\_\{m\}\|,\|h\_\{c\}\|\]\)\. Motivated by wavelet decomposition theory, different stride levels were expected to capture different temporal scales of the input signal, mirroring the hierarchical temporal structure of natural language\.

##### Journey\.

Designed in Phase 1 and retained as an opt\-in flag in Phase 2, the Wavelet State Hierarchy was evaluated in a controlled 500\-step ablation in Phase 3 alongside the other three retention concepts\. The CE descent curves showed no differential improvement across the evaluation window\. The investigation pointed to a missing inductive bias: multi\-stride states without frequency\-structured priors \(cf\. HiPPO initialisation\(Gu et al\.,[2020](https://arxiv.org/html/2606.18324#bib.bib3)\)in Mamba\) do not constitute a multi\-scale mechanism\. Different strides provide different receptive fields but no structural frequency bias, leaving the router without a substantive signal to route on\.

##### Verdict\.

Not load\-bearing\. Multi\-stride states require frequency\-structured initialisation \(e\.g\. HiPPO\-style priors\) to give each stride a distinct functional role; the investigation identified this as the missing inductive bias\.

#### 2\.5\.2Phase Bus \(Cross\-Layer EMA Communication\)

##### Hypothesis\.

A cross\-layer communication channel propagating a phase signal via exponential moving average:bus←ema⋅bus\+\(1−ema\)⋅scale⋅h\\mathrm\{bus\}\\leftarrow\\mathrm\{ema\}\\cdot\\mathrm\{bus\}\+\(1\-\\mathrm\{ema\}\)\\cdot\\mathrm\{scale\}\\cdot h,h←h\+read​\_​scale⋅bush\\leftarrow h\+\\mathrm\{read\\\_scale\}\\cdot\\mathrm\{bus\}, with learnable per\-block write and read scales\. Deeper layers injecting a summary phase signal into earlier layers was expected to enable global coherence across the stack atO​\(1\)O\(1\)cost\.

##### Journey\.

Designed in Phase 1 and retained as opt\-in in Phase 2, the Phase Bus was evaluated in Phase 3 and showed no differential improvement\. The residual stream already handles cross\-layer communication effectively at this scale\. Without a frequency structure analogous to RWKV’s geometric halflife schedules, the EMA decays at a single undifferentiated rate and adds no organisation that the residual stream cannot already provide\.

##### Verdict\.

Not load\-bearing\. The residual stream provides sufficient inter\-layer communication at this scale; a differentiated per\-layer frequency structure would be needed for the Phase Bus to add organisation beyond what the residual path already carries\.

#### 2\.5\.3Echo Memory \(Resonance Retrieval\)

##### Hypothesis\.

A content\-addressable retrieval path over a learned basis of complex keys:resk=Re​\(⟨h,bk⟩\)\\mathrm\{res\}\_\{k\}=\\mathrm\{Re\}\(\\langle h,b\_\{k\}\\rangle\),weights=softmax​\(res/d\)\\mathrm\{weights\}=\\mathrm\{softmax\}\(\\mathrm\{res\}/\\sqrt\{d\}\), with a gate initialised atσ​\(−3\.0\)≈0\.047\\sigma\(\-3\.0\)\\approx 0\.047\. Resonance scoring was expected to give the model associative retrieval without the quadratic cost of attention\.

##### Journey\.

Designed in Phase 1, Echo Memory requires a stable basis\{bk\}\\\{b\_\{k\}\\\}that tracks the hidden state throughout training\. In the early training regime, hidden\-state variance is high; causal cumulative mean retrieval with a sigmoid gate was used to satisfy this stability requirement\. In Phase 3, evaluation in the controlled ablation showed that the long\-time\-constant heads in the log\-α\\alphaspectrum already implement an implicit retrieval path, revealing that explicit basis parameters are one of several routes to the same functional goal\.

##### Verdict\.

Not load\-bearing\. The multi\-timescale scan spectrum covers the retrieval role implicitly; Echo Memory’s dedicated basis keys represent an alternative architectural route to the same function\.

#### 2\.5\.4Resonant Router

##### Hypothesis\.

A soft router over the wavelet state levels based on per\-level amplitudes:ampk=mean​\(hk2\)\\mathrm\{amp\}\_\{k\}=\\sqrt\{\\mathrm\{mean\}\(h\_\{k\}^\{2\}\)\},blend=softmax​\(Wr⋅\[amps\]\)\\mathrm\{blend\}=\\mathrm\{softmax\}\(W\_\{r\}\\cdot\[\\mathrm\{amps\}\]\)\. If different levels encoded different temporal scales, the router was expected to learn content\-dependent weighting across them\.

##### Journey\.

Implemented as specified in Phase 1 and retained as opt\-in in Phase 2, the Resonant Router was evaluated in Phase 3\. The router collapsed to single\-mode: the softmax output converged to one near\-1 component with the rest near\-0, consistent with mode collapse in mixture\-of\-experts routing without a load\-balancing loss\(Shazeer et al\.,[2017](https://arxiv.org/html/2606.18324#bib.bib13)\)\. The router can specialise to one level and stay there because there is no mechanism encouraging it to distribute across levels\.

##### Verdict\.

Not load\-bearing\. Soft routing requires a load\-balancing objective to distribute across levels; with that training signal added, the routing mechanism would have the incentive structure its design assumes\.

### 2\.6Architecture Utilities: Initialisation, Diagnostics, Inference

Three design concepts addressed the training and monitoring infrastructure surrounding the model rather than the forward computation itself\.

#### 2\.6\.1Orthogonal Initialisation

##### Hypothesis\.

Initialise all weight matrices as orthogonal:U,S,V⊤=SVD​\(𝒩​\(0,1/d\)\)U,S,V^\{\\top\}=\\mathrm\{SVD\}\(\\mathcal\{N\}\(0,1/d\)\),Winit=UW\_\{\\mathrm\{init\}\}=U\. Orthogonal matrices preserve input norm at initialisation, preventing the variance explosion or collapse that random Gaussian initialisation can produce in deep stacks\.

##### Journey\.

The original Phase 1 design attributed a∼31%\{\\sim\}31\\%PPL improvement to orthogonal initialisation\. This claim was not independently re\-verified at Phase 2 scale, but the practice was adopted universally:nn\.init\.orthogonal\_\(\)×1/2\\times 1/\\sqrt\{2\}was applied to allComplexLinearweight components, where the1/21/\\sqrt\{2\}scaling restores unit spectral norm for the combined complex operator \(two orthogonal real components combine to give spectral norm2\\sqrt\{2\}\)\. A machine\-verifiable falsifier verifiesW​W⊤≈IWW^\{\\top\}\\approx Iholds after initialisation\. The practice carried forward unchanged into Phase 3\.

##### Verdict\.

Survived\(claim partially verified\)\. Universally applied and verified by falsifier; the quantitative PPL improvement claim from Phase 1 was not independently re\-verified at Phase 2 scale\.

#### 2\.6\.2Wave Diagnostics

##### Hypothesis\.

A complex\-valued hidden state naturally exposes physics\-based health signals that real\-valued models cannot produce: energyE=mean​\(\|h\|2\)E=\\mathrm\{mean\}\(\|h\|^\{2\}\)measures whether the state is collapsing or exploding; phase coherenceC=‖mean​\(ei​θ\)‖C=\\\|\\mathrm\{mean\}\(e^\{i\\theta\}\)\\\|measures whether the phase distribution is ordered or chaotic; and winding numberW=∑Δ​θ/2​πW=\\sum\\Delta\\theta/2\\pitracks accumulated rotational drift\. This structural observability, unavailable in real\-valued architectures, was expected to give operators real\-time, interpretable visibility into model health without post\-hoc probing, withC<0\.2C<0\.2proposed as a threshold for detecting incoherent \(hallucination\-prone\) generation states\.

##### Journey\.

All three diagnostics were logged during Phase 1 training\. By Phase 2, the monitoring infrastructure evolved: the phase\-balance ratioρ=RMS​\(hi\)/RMS​\(hr\)\\rho=\\mathrm\{RMS\}\(h^\{i\}\)/\\mathrm\{RMS\}\(h^\{r\}\)replaced coherenceCCas the primary monitor, becauseρ\\rhohas a more interpretable scale \(ρ≈1\\rho\\approx 1means both channels are active;ρ→0\\rho\\to 0means imaginary collapse\) and is easier to threshold operationally\. The Phase 2 run confirmedρ∈\[0\.79,1\.22\]\\rho\\in\[0\.79,1\.22\]throughout 200,000 steps\. Energy and gradient norms were retained as secondary telemetry; winding number was deprioritised\. A concrete example of the diagnostic value: at step 37,250, theta\_drift\_mean was0\.001440\.00144rad \(negligible\) while the phase\-group gradient norm escalated to150150–290290\. Naïve diagnosis would target phase dynamics; the correct root cause was upstream amplitude drift, because phase gradients scale as\|h\|2\|h\|^\{2\}and the normalisation gain lived in the nodecay parameter group, growing unconstrained\. Fix: increase energy regularisation lambda20×20\\times, tighten the energy target from1\.351\.35to1\.101\.10, add per\-group gradient clipping \(phase group 1\.0, others 5\.0\)\. This chain \(small observable drift symptom, large gradient\-norm signal, upstream amplitude root cause\) is transferable to any complex\-valued stack where phase and magnitude parameters share an optimiser\. In Phase 3, the diagnostic infrastructure expanded further: per\-bucket gradient telemetry across 9 buckets and afirst\_nan\_attributionevent were added, replacing the need for post\-mortem re\-runs after gradient divergence\.

##### Verdict\.

Reframed\. The diagnostic philosophy is load\-bearing and evolved throughout; the specific signals \(CC, winding number\) were replaced by more operationally interpretable alternatives \(ρ\\rho, per\-bucket gnorm\)\. The hallucination\-detection capability claim is deferred pending deployment in an inference context\.

#### 2\.6\.3Wave Rewind \(Inference Correction Buffer\)

##### Hypothesis\.

An 8\-step inverse\-rotation buffer for inference correction:hprev=e−i​φ⊙hcurh\_\{\\mathrm\{prev\}\}=e^\{\-i\\varphi\}\\odot h\_\{\\mathrm\{cur\}\}, buffered for rollback when phase coherence falls below a threshold\. The model would be able to rewind its hidden state to a higher\-confidence configuration on demand\.

##### Journey\.

Designed in Phase 1 as a complement to the coherence diagnostic \(Section[2\.6\.2](https://arxiv.org/html/2606.18324#S2.SS6.SSS2)\), Wave Rewind was predicated on the possibility of phase decoherence events that would warrant rollback\. In Phase 2, the protect gate’s near\-closed initialisation \(pt≈0\.047p\_\{t\}\\approx 0\.047\) ensured the state was rarely overwritten aggressively; highptp\_\{t\}freezes rather than overwrites, so prior context is preserved structurally at every step\. The type of decoherence event the rewind was designed to recover from did not arise in the Phase 2 or Phase 3 training runs\. The mechanism was never implemented in production\.

##### Verdict\.

Not load\-bearing\. The protect gate proved sufficient to prevent the decoherence events Wave Rewind targeted; the mechanism remains available for regimes where decoherence events arise despite the gate\.

## 3Training

Both Phase 2 and Phase 3 were trained on FineWeb\-Edu\(Penedo et al\.,[2024](https://arxiv.org/html/2606.18324#bib.bib9)\)with the same base configuration:D=384D=384,L=16L=16,T=2048T=2048,V=100,277V=100\{,\}277, 169\.26M parameters, on2×2\\timesNVIDIA H100 NVL, using AdamW with cosine LR decay, gradient clipping at norm threshold5\.05\.0, and checkpointing every 2,500 steps\.

### 3\.1Phase 2: PAM Baseline

The Phase 2 training run established the stable empirical baseline for this paper\. The model was trained from random initialisation for 200,000 steps with peak LR1\.0×10−41\.0\\times 10^\{\-4\}, warmup 1,000 steps, decayed to5\.0×10−55\.0\\times 10^\{\-5\}by step 200,000\. Step time≈0\.57\{\\approx\}0\.57s/step; total wall time≈19\.8\{\\approx\}19\.8hours\. The two AdamW parameter groups were: decay \(standardℓ2\\ell\_\{2\}regularisation\) and nodecay \(bias,scale,b,E\_r,E\_i\)\.

##### Training dynamics\.

Almost all useful learning occurred in the first 10–25k steps\. Rolling\-mean CE plateaued at≈4\.6\{\\approx\}4\.6nats by step 100k; the remaining steps contributed a modest \(∼17%\{\\sim\}17\\%\) PPL improvement attributable primarily to LR decay\. Activation RMS grew from≈0\.72\{\\approx\}0\.72at step 500 to≈3\.97/3\.45\{\\approx\}3\.97/3\.45at step 200k \(5\.5×5\.5\\times\), with gradient norm rising from0\.80\.8–2\.52\.5early to3\.03\.0–10\+10\+in the tail, consistent with residual normalisation headroom being a limiting factor\. Best CE3\.093\.09nats \(PPL 22\.0\) at step 89,861; no NaNs or crashes\.

![Refer to caption](https://arxiv.org/html/2606.18324v1/x3.png)Figure 3:Phase 2 \(PAM Baseline\) cross\-entropy loss and cosine learning rate schedule over 200,000 steps\. The bulk of CE reduction occurs in the first 25,000 steps, after which the model enters a slow\-improvement phase tracking the LR decay curve\. Best CE 3\.09 \(PPL 22\.0\) is reached at step 89,861, well before the LR minimum\. The extended tail \(steps 90k–200k\) provides modest additional improvement, suggesting model capacity was largely utilised in the first quarter of training\.

### 3\.2Phase 3: Integration

Phase 3 was trained for 200,000 steps with peak LR3\.0×10−53\.0\\times 10^\{\-5\}, warmup 2,000 steps, cosine decay; batch size 3; gradient clipping at 5\.0\. The architecture isTwoStreamScan\(H=8H=8heads, halflives 5–5,000 tokens\),ChannelMixBlock\(squared\-ReLU\), real\-valuedRMSNormsandwich, plainnn\.Linear\(DD,VV\) head, real carrier throughout; token\-shift and stability loss disabled\. The model reaches best CE 2\.75 \(PPL 15\.6\) at step 161k\.

![Refer to caption](https://arxiv.org/html/2606.18324v1/x4.png)Figure 4:Phase 3 \(Integration\) training dynamics over 200,000 steps\.\(a\)Cross\-entropy loss and cosine LR schedule\. Best CE 2\.75 \(PPL 15\.6\) at step 161k confirms that the integrated architecture improves on the Phase 2 baseline \(PPL 22\.0\)\. The noisier trajectory reflects the more complex configuration relative to Phase 2\.\(b\)Gradient norm and embedding RMS over training\. Real embedding RMS \(red\) rises steadily while imaginary RMS \(purple, dashed\) remains flat throughout, consistent with real\-valued head operation\. Gradient norm spikes in the mid\-to\-late run reflect adaptation within the more complex multi\-head configuration\.![Refer to caption](https://arxiv.org/html/2606.18324v1/x5.png)Figure 5:\(a\)Best perplexity achieved across the three development phases\. Phase 1’s tied resonance head is structurally limited to PPL 1,471 before diverging; Phase 2 reaches PPL 22\.0 after resolving the collapse; Phase 3 reaches PPL 15\.6, confirming that the Phase 2 core generalises to the broader architecture\.\(b\)Smoothed CE convergence for Phase 2 and Phase 3 on a common step axis\. Phase 3’s noisier trajectory reflects the more complex configuration; both phases converge to comparable CE floors with Phase 3 reaching a modestly lower best value\.![Refer to caption](https://arxiv.org/html/2606.18324v1/x6.png)Figure 6:Gradient norm by component during Phase 3 \(Integration\) training\. Scan, channel mix, embeddings, and log\-decay parameters maintain comparable scales throughout, with no component dominating or vanishing\. This stands in contrast to Phase 1, where phase\-parameter gradients exceeded all others by three orders of magnitude \(Figure[2](https://arxiv.org/html/2606.18324#S2.F2)\), confirming that the integrated architecture produces a structurally stable gradient landscape\.
### 3\.3The Phase 3 Plan\-to\-Code Audit

The plan\-to\-code audit that emerged from Phase 3 is the most directly transferable methodological contribution of this project\.

##### The methodology\.

The Phase 3 design plan specified aPhaseAttentionHeadwith three learned bridge projections\(Wρ,Wφr,Wφi\)\(W\_\{\\rho\},W\_\{\\varphi^\{r\}\},W\_\{\\varphi^\{i\}\}\)and included a risk table with pre\-registered numerical predictions: a correctly implemented head should yield step\-0 CE≈ln⁡\(100,277\)≈11\.5\\approx\\ln\(100\{,\}277\)\\approx 11\.5nats\. This pre\-registered threshold acts as a falsifier: any value far from11\.511\.5nats at step 0 indicates a structural deviation regardless of whether unit tests pass\.

##### Audit outcome\.

The final run uses a fully verified implementation, confirmed by two machine\-verifiable falsifiers:atan2gradient finite at the origin, andW​W⊤≈IWW^\{\\top\}\\approx Iafter orthogonal initialisation\. Pre\-registered numerical thresholds written before implementation detect design\-intent deviations at step 0, before any training compute is spent, complementing unit tests, which verify execution but not intent\.

## 4Design Concept Outcomes

TableLABEL:tab:reckoningsummarises the verdict for each design concept\.

Table 3:Outcome summary for allSWavedesign concepts\.Design ConceptVerdictOutcome summaryWave EmbeddingWithdrawnAmplitude/phase semantic hypothesis not independently evaluated; unit\-circle constraint was released as part of the resonance head fix, not as a rejection of the embedding concept\.WaveMixerReframed→\\toSupersededBlend vectors do not learn from initialisation; asymmetric init acts as a structural prior; removed from default path\.AmplitudeGateSupersededWrite suppression subsumed by the dual\-purpose protect gate, which additionally controls state decay\.Unitary RotationSurvivedPhase\-preserving constraint retained via ComplexNorm and boundedγt\\gamma\_\{t\}; factored into a positional\-encoding role\.Cayley TransformNot load\-bearingO​\(D3\)O\(D^\{3\}\)inversion impractical at scale; spectral form adopted; full\-matrix constraint not required to realise the stability benefit\.Wave Propagation ScanSurvivedCore recurrence operator preserved across all phases; parallelisation improvements transferable to any linear recurrence\.Wavelet State HierarchyNot load\-bearingRequires frequency\-structured priors to realise multi\-scale behaviour; clarifies the prerequisite for stride\-based retention\.Phase BusNot load\-bearingResidual stream sufficient at this scale; differentiated per\-layer frequency structure would unlock the Phase Bus’s intended role\.Echo MemoryNot load\-bearingImplicit timescale coverage in the scan spectrum covers the retrieval role; explicit basis keys offer an alternative path\.Resonant RouterNot load\-bearingRequires a load\-balancing objective to distribute across modes; identifies the missing training signal for stride\-based routing\.Resonance Head \(tied\)SupersededTied parameterisation makeshi≡0h\_\{i\}\\equiv 0a global CE minimum; superseded by untiedEr,EiE\_\{r\},E\_\{i\}tables\.ComplexNormSurvivedPhase\-preserving normalisation load\-bearing for the complex carrier; sandwich arrangement a transferable design principle\.Wave DiagnosticsReframedMonitoring philosophy survives; specific signals evolved toρ\\rhoand per\-bucket gnorm for operational interpretability\.Orthogonal InitSurvivedUniversally applied;W​W⊤≈IWW^\{\\top\}\\approx Iverified by falsifier\.Wave RewindNot load\-bearingProtect gate proved sufficient for the tested regime; mechanism remains available for contexts where decoherence events arise\.SwiGLU FFNSupersededReal\-scalar formulation breaks complex coupling; superseded by ComplexGatedUnit, then by ChannelMixBlock\.Summary\.Survived: Wave Propagation Scan, ComplexNorm, Unitary Rotation, Orthogonal Initialisation\.Reframed: Wave Diagnostics, WaveMixer \(further superseded\)\.Superseded: AmplitudeGate, Resonance Head, SwiGLU FFN\.Not load\-bearing: Cayley Transform, Wavelet State Hierarchy, Phase Bus, Echo Memory, Resonant Router, Wave Rewind\.Withdrawn: Wave Embedding, Self\-correcting generation via Wave Rewind, Hallucination detection viaC<0\.2C<0\.2\.

##### Capability claims reckoning\.

Beyond the design concepts, six capability\-level promises were made at the start of the project\. Table[4](https://arxiv.org/html/2606.18324#S4.T4)documents their status\.

Table 4:Reckoning of six capability\-level promises against evidence\. Two are reframed based on measured results; one is structurally satisfied; three were not measured in any documented run\.

## 5Discussion and Open Questions

##### Does the complex carrier contribute to performance?

Phase 2 \(complex carrier, 200k steps\) reaches PPL 22\.0; Phase 3 \(real carrier, 200k steps, more capable primitives\) reaches PPL 15\.6\. The Phase 3 improvement is attributable to the architectural substitutions \(TwoStreamScan with 8 multi\-timescale heads, squared\-ReLU channel mixer, sandwich RMSNorm\) rather than to the carrier change alone, since multiple components changed simultaneously\. Isolating the carrier’s individual contribution would require an ablation that holds all other components fixed and varies only the carrier type\. That experiment was not run, so the independent contribution of the complex carrier to performance remains an open question\.

##### Do complex embeddings provide richer representations?

The Wave Embedding concept \(free\-magnitude complex embeddings with explicit phase priors encoding semantic direction\) was not independently evaluated because its entanglement with the tied resonance head meant releasing the unit\-circle constraint was part of fixing that head, not a test of the embedding concept itself\. Whether free\-magnitude complex embeddings with explicit phase priors provide richer representations than standard real\-valued embeddings remains an open question for future work\.

## 6Related Work

##### Complex\-valued and unitary recurrent networks\.

Unitary RNNs\(Arjovsky et al\.,[2016](https://arxiv.org/html/2606.18324#bib.bib1); Wisdom et al\.,[2016](https://arxiv.org/html/2606.18324#bib.bib16)\)enforce unit\-magnitude state transitions to prevent vanishing and exploding gradients, establishing the theoretical foundation for norm\-preserving recurrences\. Phase\-Associative Memory\(Vishwakarma et al\.,[2026](https://arxiv.org/html/2606.18324#bib.bib15)\)extends this to language modelling with a content\-addressable write mechanism based on Hermitian inner products, drawing on the tradition of complex Hopfield networks\(Noest,[1992](https://arxiv.org/html/2606.18324#bib.bib7)\)and Holographic Reduced Representations\(Plate,[1995](https://arxiv.org/html/2606.18324#bib.bib11)\)\.

##### Selective state spaces and linear recurrence models\.

The S4/S5 lineage\(Gu et al\.,[2022](https://arxiv.org/html/2606.18324#bib.bib4)\)andMamba\(Gu and Dao,[2024](https://arxiv.org/html/2606.18324#bib.bib5)\)demonstrate that structured state spaces with input\-dependent decay achieve competitive performance on long\-sequence tasks while retainingO​\(1\)O\(1\)inference memory\.RWKV\(Peng et al\.,[2023](https://arxiv.org/html/2606.18324#bib.bib8)\)shows that linear recurrences with per\-head geometric halflife schedules, receptance gating, and token\-shift mixing are effective language model backbones at scale\.

##### Architecture retrospectives and empirical methodology\.

Kaplan et al\. \([2020](https://arxiv.org/html/2606.18324#bib.bib6)\)establishes the empirical tradition of documenting scaling behaviour across a large design space\.Pineau et al\. \([2021](https://arxiv.org/html/2606.18324#bib.bib10)\)argues for documented, reproducible research processes in machine learning\.

## 7Conclusion

SWavedemonstrates that a complex\-valued recurrent language model can be trained stably at 169M parameters over 200,000 steps, reaching PPL 15\.6 with eight multi\-timescale heads, a two\-stream recurrence, and a real\-valued channel mixer\. Being purely recurrent, inference requires only the current hidden state:O​\(1\)O\(1\)memory per sequence position by construction, independent of context length\.

The central finding is a formal characterisation of cos\-domination collapse: tied phase\-matching output heads structurally admit imaginary\-channel collapse as a global loss minimum, independent of optimisation choices\. The actionable constraint for any complex\-valued language model is to verify that the output head parameterisation does not admit this degenerate minimum before any training is run\. The scan engineering \(chunked parallel scan, log\-space backward,log⁡α\\log\\alphasaturation clamp\) resolves the three principal numerical hazards of first\-order complex linear recurrences and is transferable to any SSM\-style model with learned per\-step decay\.

Two patterns carry beyond this work\. Verify the loss landscape geometry before tuning optimisation; in complex\-valued models, phase\-parameter gradient norms can exceed embedding gradient norms by728×728\\times, requiring separate per\-group learning rate scheduling\. And structural resemblance to an effective mechanism does not transfer its inductive biases: each retention concept that did not contribute mirrored an established mechanism without the specific prior that makes the original effective\. Taken together, these results offer a replicable reference point for future complex\-valued recurrent model design\.

## References

- Arjovsky et al\. \(2016\)Arjovsky, M\., Shah, A\., and Bengio, Y\. \(2016\)\.Unitary Evolution Recurrent Neural Networks\.In*Proceedings of ICML*\.
- Gemma Team \(2024\)Gemma Team \(2024\)\.Gemma 2: Improving Open Language Models at a Practical Size\.*arXiv:2408\.00118*\.
- Gu et al\. \(2020\)Gu, A\., Dao, T\., Ermon, S\., Rudra, A\., and Ré, C\. \(2020\)\.HiPPO: Recurrent Memory with Optimal Polynomial Projections\.In*Proceedings of NeurIPS*\.
- Gu et al\. \(2022\)Gu, A\., Goel, K\., and Ré, C\. \(2022\)\.Efficiently Modeling Long Sequences with Structured State Spaces\.In*Proceedings of ICLR*\.
- Gu and Dao \(2024\)Gu, A\. and Dao, T\. \(2024\)\.Mamba: Linear\-Time Sequence Modeling with Selective State Spaces\.*arXiv:2312\.00752*\.
- Kaplan et al\. \(2020\)Kaplan, J\., McCandlish, S\., Henighan, T\., et al\. \(2020\)\.Scaling Laws for Neural Language Models\.*arXiv:2001\.08361*\.
- Noest \(1992\)Noest, A\. J\. \(1992\)\.Associative memory as a complex Hopfield network\.*Neural Networks*, 5\(2\):365–376\.
- Peng et al\. \(2023\)Peng, B\., Alcaide, E\., Anthony, Q\., et al\. \(2023\)\.RWKV: Reinventing RNNs for the Transformer Era\.In*Proceedings of EMNLP*\.
- Penedo et al\. \(2024\)Penedo, G\., Kydlíček, H\., allal, L\. B\., et al\. \(2024\)\.The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale\.*arXiv:2406\.17557*\.
- Pineau et al\. \(2021\)Pineau, J\., Vincent\-Lamarre, P\., Sinha, K\., et al\. \(2021\)\.Improving Reproducibility in Machine Learning Research\.*Journal of Machine Learning Research*, 22\(164\):1–20\.
- Plate \(1995\)Plate, T\. A\. \(1995\)\.Holographic Reduced Representations\.*IEEE Transactions on Neural Networks*, 6\(3\):623–641\.
- Shazeer \(2020\)Shazeer, N\. \(2020\)\.GLU Variants Improve Transformer\.*arXiv:2002\.05202*\.
- Shazeer et al\. \(2017\)Shazeer, N\., Mirhoseini, A\., Maziarz, K\., et al\. \(2017\)\.Outrageously Large Neural Networks: The Sparsely\-Gated Mixture\-of\-Experts Layer\.In*Proceedings of ICLR*\.
- Su et al\. \(2021\)Su, J\., Lu, Y\., Pan, S\., et al\. \(2021\)\.RoFormer: Enhanced transformer with rotary position embedding\.*arXiv:2104\.09864*\.
- Vishwakarma et al\. \(2026\)Vishwakarma, S\. et al\. \(2026\)\.Phase\-Associative Memory for sequence modelling\.*arXiv:2604\.05030*\.
- Wisdom et al\. \(2016\)Wisdom, S\., Powers, T\., Hershey, J\., Le Roux, J\., and Atlas, L\. \(2016\)\.Full\-capacity unitary recurrent neural networks\.In*Proceedings of NeurIPS*\.

Similar Articles

Prediction Dynamics in Depth-Recurrent Language Models

arXiv cs.CL

This paper analyzes the dynamics of predictions in depth-recurrent language models, deriving a decomposition of score changes to explain how intermediate latent updates can preserve final answers despite fluctuations in scores.