Tag
A native Rust and Vulkan training backend has been developed for 143 Transformer architectures, enabling CUDA-free model training and inference with support for various hardware vendors.
The paper reveals that the growth of Weibull weight-scale in transformer training is determined by data predictability, specifically through bigram conditional entropy, and establishes a predictive law for this growth.
An interactive, explorable explanation of various parallelization schemes for training transformers, adapted from academic content on scaling models.
Introduces SymExpLin (SEL), a weight reparameterization that combines symmetric-exponential and linear pathways to improve optimization in neural networks, reducing training steps by up to 1.49x on transformers.
A developer reverse-engineered Apple's private APIs to enable training neural networks directly on the Apple Neural Engine (ANE) in M4 Macs and iPhones, bypassing CoreML and GPU. The project demonstrates that ANE hardware is capable of training, though with limitations like low utilization and CPU fallbacks for some operations.
Introduces KV-Compression Aware Training (KV-CAT), a method that encourages transformers to learn compressible key-value caches during training, improving memory efficiency for long-context tasks without sacrificing performance.