@HanGuo97: LLM training is built on fast MatMuls. But many surrounding ops still run as memory-bound kernels. CODA reparameterizes…

X AI KOLs Following Papers

Summary

CODA reparameterizes memory-bound operations in LLM training to fuse them into the matmul epilogue, achieving near state-of-the-art performance with LLM-generated kernels.

LLM training is built on fast MatMuls. But many surrounding ops still run as memory-bound kernels. CODA reparameterizes them to hide in the matmul’s shadow, fused into its epilogue before results leave the chip. Bonus: LLMs can write fast CODA kernels too (approaching SoLs). https://t.co/cOTeMUr4py
Original Article
View Cached Full Text

Cached at: 05/22/26, 05:49 AM

LLM training is built on fast MatMuls. But many surrounding ops still run as memory-bound kernels.

CODA reparameterizes them to hide in the matmul’s shadow, fused into its epilogue before results leave the chip.

Bonus: LLMs can write fast CODA kernels too (approaching SoLs). https://t.co/cOTeMUr4py

Similar Articles

I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens

Reddit r/LocalLLaMA

YOON1v released Apex-2, a from-scratch decoder-only Mixture-of-Experts LLM with 3.87B total / 1.45B active parameters, pretrained on 86.5B tokens and SFT-tuned, with full code, architecture writeup, and benchmark results (HumanEval 43.9) released on Hugging Face and GitHub.

I built a tool to edit my agents text. I need honest feedback

Reddit r/AI_Agents

The author built a tool combining an open-source CLI that hooks into commands like git commit and file writes with an MCP server that queues agent-generated text, so the user can review and edit drafts before approval. They're seeking feedback on whether this workflow matches how others handle AI-written commits, emails, and PR descriptions.

Format-Aware Fusion for Fast FP4 Pretraining

arXiv cs.LG

提出了 format-aware fusion 方法,将 FP4 量化中的 scale 计算、操作数打包与反向状态保存与 GEMM 融合,在 Llama-3 8B 预训练中达到最高 37.9K tokens/s/GPU,MXFP4 路线在略高于 bfloat16 训练损失终点的情况下实现 37.2K tokens/s/GPU(86.3% MFU),显著超过 Transformer Engine 的 27.6K。

@shao__meng: Latest CMU Fall Course 11-768: AI Agents — Lecture 12 Notes Released: Reinforcement Learning Systems, Taught by @gneubig. Course: https://cmu-agents.com, Slides: https:/…

X AI KOLs Timeline

Lecture 12 notes from CMU course 11-768 AI Agents are now public, covering Reinforcement Learning Systems: how to train LLM Agents with RL on multi-GPU clusters, coordination between training and inference engines, KV/Prefix Caching, Weight Sync, common pitfalls in Agentic RL (sequence extension, chat templates, TITO, sync vs async), and system design choices ranging from single programs to cloud-native microservices.