Tag
The paper investigates replacing matrix multiplication in Transformer layers with an associative algebra product to reduce arithmetic cost while retaining parameters, demonstrating feasibility with improved throughput but some performance trade-offs.