Tag
The author is seeking collaborators for scaling and independent evaluation of a new recurrent language model architecture, with a preprint and code available.
This paper introduces the Structured Recurrent Mixer (SRM), an architecture enabling algebraic conversion between parallel training and recurrent inference without specialized kernels. Experiments show SRMs achieve significantly higher throughput and concurrency compared to Transformers, with effective performance in reinforcement learning tasks.