@danielhanchen: I’m running a 3 hour advanced workshop at AI Engineer World’s Fair! 2026 has greatly changed how one should learn lower…

X AI KOLs Following Events

Summary

Daniel Han is hosting a 3-hour advanced workshop at the AI Engineer World's Fair, sharing insights on the history of open-source large models, classification of training stages (pre-training, intermediate training, supervised fine-tuning, post-training, reinforcement fine-tuning), and the leap in reasoning models. He also introduced his team's open-source contributions to fine-tuning optimization.

I’m running a 3 hour advanced workshop at AI Engineer World’s Fair! 🚀 2026 has greatly changed how one should learn lower-level technicals like kernels, agentic RL, reward hacking, cont learning. What would you like to see? Last year @aiDotEngineer: https://t.co/3j8qcMD9u8
Original Article
View Cached Full Text

Cached at: 06/24/26, 08:00 AM

I’m running a 3 hour advanced workshop at AI Engineer World’s Fair! 🚀

2026 has greatly changed how one should learn lower-level technicals like kernels, agentic RL, reward hacking, cont learning.

What would you like to see?

Last year @aiDotEngineer: https://t.co/3j8qcMD9u8


TL;DR: Daniel Han shared insights on the history of open-source LLMs, classification of training stages, and recent reasoning model breakthroughs at AI Engineer World’s Fair, along with his team’s open-source contributions to fine-tuning optimization.

Self-Introduction and Project Background

Daniel Han first apologized for being late and introduced himself as a member of the AI Engineer community. His team is known for being active on X, having fixed gradient accumulation bugs, proposed “asynchronous offload gradient checkpointing,” and collaborated with Hugging Face, Google, Meta, Mistral, etc., to fix bugs in models like Gemma, Llama, and Mistral MoE. They contribute to the entire open-source ecosystem, such as contributions to llama.cpp and collaboration with Qwen and Mistral on model releases.

The team’s monthly downloads on Hugging Face have exceeded 10 million, and GitHub stars have reached 40,000. Their core work is making fine-tuning faster and more memory-efficient. The GitHub package includes free Colab and Kaggle notebooks (free Google GPU, Kaggle 30 hours/week free GPU) for inference, continued pre-training, supervised fine-tuning, etc. They also upload quantized models (e.g., 1.58-bit versions) to Hugging Face, which are small and retain most accuracy, runnable on local devices.

History of Open-Source Models and Key Milestones

The Catalytic Effect of Llama’s Open-Source Release

The earliest was Meta’s Llama research paper, initially with only research access. After weights were accidentally leaked, the entire open-source movement was catalyzed. Llama 1 was trained on only 1.4 trillion tokens, with loss decreasing over training time. Now models are much larger: Gemma 3 was trained on 14 trillion tokens, Llama 4 on 30 trillion.

Open-Source vs. Closed-Source: Catching Up and Divergence

Maxim’s graph shows that on MMLU 5-shot, the slope for open-source models is steeper, and eventually Llama 3.1 405B reached GPT-4 level; open-source has caught up to closed-source. However, after September 2024, there was an “open-source winter”: o1-preview demonstrated reasoning chains with a jump in capability, and the open-source community couldn’t replicate it for four months. It wasn’t until January 2025 that DeepSeek R1 was released, proving that open-source models can also be trained to match o1/o3 capabilities.

Classification and Naming of Training Stages

Daniel explained the current stages of model training:

  1. Pre-training: Predicts the next word using all public data (Wikipedia, web pages, etc.).
  2. Mid-training: Gives higher weight to high-quality data (e.g., Wikipedia) or performs long-context extension.
  3. SFT / Instruction Fine-tuning: Converts a base model into a chat model (e.g., ChatGPT, Claude 4 Opus, Gemini 2.5 Pro, etc.).
  4. Post-training: Preference fine-tuning, DPO, RLHF, etc.
  5. RLVR: Reinforcement Learning with Verifiable Rewards – a new paradigm.

Model naming is inconsistent: common ones include PT (pre-training), IT (instruction fine-tuning), Instruct, Chat, Base, etc. The open-source community should standardize naming.

Looking to the Future from History: Two Waves of Capability Leap

  • First Leap: SFT / RLHF leap – through good supervised fine-tuning and reinforcement learning, model performance significantly improved.
  • Second Leap: RL leap – using reinforcement learning methodology (e.g., R1) to further greatly boost performance.
  • What’s next? Daniel thinks reasoning might be the last step, because the DeepSeek R1 paper indicates models already have reasoning ability, they just need reinforcement. But each time closed-source models make a step function, it’s uncertain whether a plateau is coming.

Yann LeCun’s Cake Analogy

The community often cites Yann LeCun’s view from 2016: unsupervised learning (pre-training) is the cake, supervised fine-tuning is the icing, and reinforcement learning is the cherry. But RL data is scarce; large model labs optimize through pre-training + iterative fine-tuning.

Explanation of Random Initialization

Models start from random parameters (e.g., GPT-4’s 7 billion parameters are all random numbers) and gradually move weights to useful states through training.


Source: @danielhanchen: I’m running a 3 hour advanced workshop at AI Engineer World’s Fair! | YouTube (https://www.youtube.com/watch?v=OkEGJ5G3foU)

Similar Articles

@shao__meng: https://x.com/shao__meng/status/2106715778078974378

X AI KOLs Timeline

NYU's Fall 2026 graduate seminar, “Language Models, Reinforcement Learning, Reasoning,” threads frontier research from Transformers to world models across a 13-week course — covering pretraining scaling laws, RLVR/GRPO post-training, alignment and safety, and real-world engineering practices such as Kimi K3. The instructor, Pavel Izmailov, is from Anthropic and previously contributed to OpenAI's o1 and Superalignment teams.

@Jolyne_AI: I found a solid hands-on tutorial for large language models on GitHub: the "Hands-On Large Model Series." It takes you from zero to mastering the entire tech stack. Using a combination of videos, documents, and code, it links key capabilities like fine-tuning, deployment, RAG, and Agent into a reproducible learning path — each knowledge point can be directly practiced...

X AI KOLs Timeline

Recommend the "Hands-on Large Model Series" tutorial on GitHub, which systematically explains practical techniques such as fine-tuning, deployment, RAG, and Agent through videos, documents, and code, suitable for AI developers to improve their engineering skills.