Tag
The article promotes a poster presentation at PyTorch Conference North America on enabling the open-source vime RL post-training framework on AMD Instinct GPUs using ROCm, and provides registration details for the conference.
The author built FlashMLA for consumer-grade Blackwell sm_120, achieving 2-3x performance gains over PyTorch SDPA in attention-heavy workloads like long-context training and sparse prefill.
The article discusses how open labs are shifting towards continued post-training on existing AI models for incremental improvements, which benefits the local community by ensuring better compatibility with inference engines like llama.cpp.
This tweet introduces 'Awesome LLM Books', a curated GitHub list of 22 high-quality books for LLM development, evaluated by strict criteria including relevance, content quality, and social proof. Each book entry includes author, publisher, rating, and links, helping developers quickly find suitable resources.
A new optimization technique for open-source RL training engines introduces prompt caching during training, achieving up to 7.5x speedup on long-prompt, short-response workloads by reducing redundant compute.
Anthropic has published the Claude Cookbook, a curated collection of 81 practical developer guides spanning AI agents, RAG, evaluations, multimodal apps, and production workflows. The resource offers actionable code examples and best practices for building and deploying applications with Claude.
Former Anthropic scientist Shunyu Yao revealed details on the R&D of Claude 3.7 in a podcast, along with Anthropic's strategic shift to heavily bet on coding capabilities, and compared the differences in decision-making structures between Anthropic and OpenAI.
The author reflects on a visit to China's AI labs, comparing cultural differences between Chinese and American labs in building LLMs. Chinese labs benefit from a culture of collective work and student involvement, while American labs face challenges from individual ego and career ambitions.
Marin is an open-source research program and software platform dedicated to the transparent development of foundation models, covering data curation to model training and evaluation.