@yuwen_lu_: I'm halfway through, damn why did no one ever tell me RL is this fun

X AI KOLs Timeline Tools

Summary

Sanbu 散步 released a modern RL tutorial Hands-On Modern RL, covering from CartPole+PPO basics to LLM post-training (RLHF, DPO, GRPO) and Agentic RL, code-first, English version coming soon.

I'm halfway through, damn why did no one ever tell me RL is this fun
Original Article
View Cached Full Text

Cached at: 05/31/26, 07:03 AM

Halfway through, holy crap, why has no one ever told me RL is this fun?

Sanbu 散步 (@sanbuphy): Spent some time writing a RL tutorial, Hands-On Modern RL. The path starts with CartPole + PPO, then moves to LLM post-training (RLHF, DPO, GRPO), and Agentic RL. Code-first, formulas used to explain phenomena. English version coming soon. Currently a draft. RLHF and Agentic RL sections are under local review. PRs, Issues & GPU sponsorship welcome:

Similar Articles

@duange6099: If you're also looking to get started with Claude Code, I highly recommend bookmarking this collection of resources. Feel free to share your experiences and learn together. 1. Most comprehensive Chinese tutorial (recommended to read first) Covers from installation to hands-on projects https://github.com/shareAI-lab/le…

X AI KOLs Timeline

A Twitter user shared a carefully curated collection of Claude Code learning resources, including Chinese tutorials, command tips, skill libraries, etc., for beginners to get started.