Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers
Summary
The paper investigates replacing matrix multiplication in Transformer layers with an associative algebra product to reduce arithmetic cost while retaining parameters, demonstrating feasibility with improved throughput but some performance trade-offs.
View Cached Full Text
Cached at: 09/29/26, 08:09 AM
Paper page - Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers
Source: https://huggingface.co/papers/2609.32814 We study whether the matrix multiplication used in Transformer projections can be replaced by a different, cheaper associative product while keeping the same weight bank and parameter count.
We construct an associative algebra product with a lower bilinear rank than standard matrix multiplication. For the q=2 case, the product has rank 6, compared with rank 7 for Strassen’s 2×2 algorithm. This allows us to replace linear projections without introducing additional nonlinearities, while reducing their arithmetic cost.
We evaluate the approach at three levels: GPU projection kernels across several Transformer shapes, including Qwen and DeepSeek; and two approximately 110M-parameter decoder-only LMs trained for 12.3B tokens, differing only in the MLP multiplication law. The algebraic model achieves a 6.2–7.8% increase in end-to-end generation throughput across four prompt domains, while obtaining lower scores on GSM8K, MBPP, and IFEval.
The results provide a small-scale feasibility and trainability study of changing the multiplication law of Transformer projections while retaining the full parameter bank.
Similar Articles
Transformer co-author validates post-transformer cost efficiency breakthrough
A 150M-parameter non-transformer architecture achieves state-of-the-art cost-efficiency on ARC-AGI-1, validated by Transformer co-author Łukasz Kaiser, suggesting that recurrent latent reasoning can replace brute-force scaling.
Training transformers where every layer W = V·Uᵀ from initialization reveals a corpus-determined optimal rank - looking for arXiv endorser (cs.LG) [D]
This paper proposes Native Factorized Weights for transformers, where every linear layer is trained as a product of two low-rank matrices from initialization. Experiments show a corpus-determined optimal rank that minimizes validation loss and a generalization band, outperforming dense baselines with fewer parameters.
@rohanpaul_ai: Interesting, this paper shows that Transformers may not need separate key and value projections to work well. This pape…
This paper investigates whether Transformers need separate key and value projections, finding that sharing them can reduce KV cache by 50% with only 3.1% higher perplexity, and further cuts when combined with GQA and MQA.
Parallel Manifold Steering: Efficient Adaptation of Large Associative Memories via Residual Energy Shaping
This paper proposes H-Res, a method to adapt large transformer models by shaping the energy landscape of associative memories without modifying weights or adding prompts, preserving memory capacity and outperforming LoRA.
Tensor Cache: Eviction-conditioned Associative Memory for Transformers
Tensor Cache introduces a two-level caching mechanism that compresses evicted key-value pairs from sliding-window attention into a fixed-size associative memory, improving long-context language modeling without unbounded memory growth.