[Full Workshop] Reinforcement Learning, Kernels, Reasoning, Quantization & Agents — Daniel Han
Summary
At the AI Engineer World Congress, Daniel Han delivered an in-depth talk on the practical experiences of reinforcement learning, model fine-tuning, quantization, and agents. He reviewed the evolution of open-source models from Llama to DeepSeek R1 and analyzed the five key stages of modern model training.
View Cached Full Text
Cached at: 06/25/26, 01:32 PM
Similar Articles
@danielhanchen: I’m running a 3 hour advanced workshop at AI Engineer World’s Fair! 2026 has greatly changed how one should learn lower…
Daniel Han is hosting a 3-hour advanced workshop at the AI Engineer World's Fair, sharing insights on the history of open-source large models, classification of training stages (pre-training, intermediate training, supervised fine-tuning, post-training, reinforcement fine-tuning), and the leap in reasoning models. He also introduced his team's open-source contributions to fine-tuning optimization.
@shao__meng: Latest CMU Fall Course 11-768: AI Agents — Lecture 12 Notes Released: Reinforcement Learning Systems, Taught by @gneubig. Course: https://cmu-agents.com, Slides: https:/…
Lecture 12 notes from CMU course 11-768 AI Agents are now public, covering Reinforcement Learning Systems: how to train LLM Agents with RL on multi-GPU clusters, coordination between training and inference engines, KV/Prefix Caching, Weight Sync, common pitfalls in Agentic RL (sequence extension, chat templates, TITO, sync vs async), and system design choices ranging from single programs to cloud-native microservices.
@shao__meng: https://x.com/shao__meng/status/2106715778078974378
NYU's Fall 2026 graduate seminar, “Language Models, Reinforcement Learning, Reasoning,” threads frontier research from Transformers to world models across a 13-week course — covering pretraining scaling laws, RLVR/GRPO post-training, alignment and safety, and real-world engineering practices such as Kimi K3. The instructor, Pavel Izmailov, is from Anthropic and previously contributed to OpenAI's o1 and Superalignment teams.
@Michaelzsguo: This is one of the best deep discussions I've seen recently about the fundamentals of reinforcement learning and its relationship to modern AI. Eric Jang and Dwarkesh turned a seemingly retro exercise—rebuilding AlphaGo with today's tools—into a very clear masterclass: why 'search +...'
A detailed discussion on reinforcement learning and its connection to modern AI, using the reconstruction of AlphaGo with modern tools as a clear example of search and self-play. Key takeaways include neural network amortization of search, credit assignment challenges in LLMs vs AlphaGo, and implications for automated research.
@dair_ai: https://x.com/dair_ai/status/2053495521243799717
DAIR AI's weekly roundup highlights top research papers including HeavySkill, which improves model performance via internalized parallel reasoning, and Sakana AI's Conductor, which uses RL to optimize agent orchestration. It also covers Meta FAIR's work on self-improving pretraining.