Tag
ALPHABET is a compact linear-time sequence model that compresses temporal history into stable complex pole modes, achieving competitive performance with fewer parameters and faster inference.
WorldToken presents a time-first sequence modeling method for robotic imitation learning, using causal Transformers and diffusion action heads to fuse heterogeneous observations, with performance evaluations on RoboCasa and RMBench benchmarks.
This paper uses mechanistic interpretability to analyze how LLaMA 3.1 8B models numerical sequences, revealing that it internally computes and stores first differences to extrapolate patterns.
The paper explores using complex-valued states inspired by quantum theory in sequence models, showing it can reduce optimization steps in Mamba and attention-based models with better sample efficiency.
This paper introduces Modular TTT, a framework that represents test-time training inner learners as directed acyclic graphs, enabling systematic ablation and composition of components. The authors train 410M and 1.45B parameter models on 100B tokens, achieving performance comparable to GatedDeltaNet.
This paper introduces Complementary Matrix Gating (CMG) for QKAN-based fast-weight programmers, enabling coordinate-wise memory control for quantum dynamics forecasting. The method shows consistent improvements and low mean-squared errors on quantum simulation benchmarks.
HantaWatch is a federated learning framework for hantavirus genomic surveillance that enables collaborative training of sequence-based models without sharing raw data, integrating k-mer feature extraction and adaptive optimization to support risk screening and expert prioritization.
Flexformer proposes a flexible linear Transformer with fully learnable attention kernels using random Fourier features, achieving linear complexity while matching or exceeding softmax attention performance on language modeling and sequence classification tasks.
This paper compares xLSTM, Mamba-2, and Gated DeltaNet on complex sequence modeling tasks and finds xLSTM superior due to its enhanced state tracking and memory dynamics, validated on synthetic length-generalization tasks.
This paper presents a memory–stability–expressivity trilemma for trainable dissipative oscillator networks, showing that damping governs all three and limits trainability, with experimental validation on a 20-oscillator network confirming the theoretical bounds.
Proposes the Mamba-Assisted Closure (MAC) framework, a Mamba-based sequence model for non-Markovian closure in reduced-order modeling of high-dimensional dynamical systems, outperforming GRU-based and Markovian methods on Burgers' equation and Lorenz '96 systems.
This paper introduces generic triple-latent recurrent models that compress token pair interactions into a latent state, and a gated associative retrieval variant that improves exact recall. The hybrid model outperforms Transformers on byte-level WikiText-2 and a tokenized language benchmark, achieving up to 41.9% associative recall versus 25%.
This paper proposes Q-align DT, a framework that aligns return-to-go with Q-values to improve controllability and performance in offline reinforcement learning, achieving superior results on D4RL benchmarks.
Proposes Interdomain Attention, a new method that integrates state space models into attention via kernel methods, achieving efficient long-context modeling with a fixed-size state and outperforming SSMs and softmax attention in language modeling experiments up to 1.3B parameters.
The author presents SM1, a variant of Mamba1 with d_state=1, using two native PyTorch ops to replace the selective scan, reducing memory by 16x compared to d_state=16. The closed-form solution eliminates the state dimension, enabling efficient inference with constant memory per token.
Introduces Next-Latent Prediction (NextLat), a self-supervised objective that trains transformers to predict their next latent state, encouraging compact internal world models and improving generalization across sequence modeling tasks.