Tag
An experimental language model architecture uses hypersurfaces for dynamic weight updating to reduce training parameters, achieving better performance than unrolled baselines while using only 16% of the parameters. The approach is tested on the FineWeb-Edu dataset and includes a GitHub implementation.