Tag
Profiles Dan Fu, a key contributor to high-performance kernels like FlashAttention, Hyena, Monarch Mixer, and ThunderKittens, now a distinguished researcher at Together AI whose work is used in ChatGPT, Claude, and Gemini.
Dan Fu announces a live chat with Olive Song at aiDotEngineer about MiniMax 3, covering its training decisions.
Together AI open-sources OSCAR, an attention-aware 2-bit KV cache quantization system that enables efficient long-context LLM serving by redistributing quantization error according to attention importance.