transformer-training

Tag

Cards List
#transformer-training

I built a native Vulkan training backend for 143 modern Transformer architectures — no CUDA or PyTorch required

Reddit r/LocalLLaMA ↗ · 2026-09-16

A native Rust and Vulkan training backend has been developed for 143 Transformer architectures, enabling CUDA-free model training and inference with support for various hardware vendors.

0 favorites 0 likes
#transformer-training

Data Predictability Shapes Weibull Weight-Scale Growth in Transformer Training

arXiv cs.LG ↗ · 2026-08-26 Cached

The paper reveals that the growth of Weibull weight-scale in transformer training is determined by data predictability, specifically through bigram conditional entropy, and establishes a predictive law for this growth.

0 favorites 0 likes
#transformer-training

Parallelizing Transformer Training (39 minute read)

TLDR AI ↗ · 2026-08-21 Cached

An interactive, explorable explanation of various parallelization schemes for training transformers, adapted from academic content on scaling models.

0 favorites 0 likes
#transformer-training

Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization

arXiv cs.LG ↗ · 2026-07-14 Cached

Introduces SymExpLin (SEL), a weight reparameterization that combines symmetric-exponential and linear pathways to improve optimization in neural networks, reducing training steps by up to 1.49x on transformers.

0 favorites 0 likes
#transformer-training

@0x0SojalSec: Apple hid 15.8 TFLOPS of raw AI power in every M4 Mac & iPhone. They only let you use the Neural Engine for inference. …

X AI KOLs Timeline ↗ · 2026-06-15 Cached

A developer reverse-engineered Apple's private APIs to enable training neural networks directly on the Apple Neural Engine (ANE) in M4 Macs and iPhones, bypassing CoreML and GPU. The project demonstrates that ANE hardware is capable of training, though with limitations like low utilization and CPU fallbacks for some operations.

0 favorites 0 likes
#transformer-training

@jiqizhixin: What if your AI’s memory didn’t have to balloon with every extra sentence? University of Oxford, Technion, AITHYRA, and…

X AI KOLs Timeline ↗ · 2026-06-14 Cached

Introduces KV-Compression Aware Training (KV-CAT), a method that encourages transformers to learn compressible key-value caches during training, improving memory efficiency for long-context tasks without sacrificing performance.

0 favorites 0 likes
← Back to home

Submit Feedback