Tag
Introduces the Recurrent Looped Transformer (RLT), a new AI architecture with infinite reasoning depth, discussed in the context of achieving safe superintelligence.
This paper presents a method to autonomously generate the full 'Bad Apple' video using a small recurrent dynamical system with 417k parameters, trained from a single initial state. It details the architecture, training techniques like rollout horizon curriculum and state perturbation noise, and shares code and models on GitHub.
Introduces a Lindblad-inspired multi-timescale reservoir architecture that separates rotation and dissipation for independent control of mixing, memory, and stability, achieving competitive results on benchmarks like NARMA-20 and Lorenz-63.
The paper introduces the context-ready transformer, a recurrent architecture that pre-contextualizes tokens before the transformer block, achieving significant inference speedups (e.g., 1.7x on A100) while matching or exceeding standard transformer performance with fewer layers.
This paper presents a retrospective on the design evolution of SWave, a complex-valued recurrent language model, detailing which architectural components were retained, reframed, superseded, or proved non-load-bearing, along with formal characterizations of failure modes like cos-domination collapse.