@neural_avb: Next video is on training tiny (<1B) models for preference tuning. Plus how to generate preference datasets with local …
Summary
Announces an upcoming video on training tiny models for preference tuning, covering reward models, RLHF, DPO, ORPO with Unsloth and TRL.
View Cached Full Text
Cached at: 05/26/26, 01:10 PM
Next video is on training tiny (<1B) models for preference tuning. Plus how to generate preference datasets with local models.
Covers reward models, RLHF, DPO, ORPO with Unsloth and TRL. Releasing sometime this week! https://t.co/iFuBj5oaIT
Similar Articles
@neural_avb: Locally generating GRPO-like rollouts with my SLM, and using this tiny RM as the rubric. Next I'll be RL training on fr…
Neural_avb releases a lightweight Answer-eq Reward Model for RL training on QA tasks, claiming 80% agreement with external judge LM and faster than F1/ROUGE/BertScore.
@neural_avb: Watch this 45 min video to learn how to create synthetic datasets and train tiny (100M params) local language models th…
A 45-minute video tutorial on creating synthetic datasets and training tiny (100M parameter) local language models for narrow tasks, with code and resources provided.
@neural_avb: Lurking the Reasoning Training docs rn. Time to write a verifiers env and Unsloth/TRL that shit! Video soon if it all g…
The user is working on implementing reasoning training with verifiers using Unsloth and TRL, reporting progress on locally generating GRPO-like rollouts with a small SLM and a tiny RM, and promises a video soon.
@neural_avb: This post-training article came out earlier this year and completely flew under my radar. Highly recommended for my GRP…
A recommendation of a post-training article on GRPO/RLVR that was overlooked earlier this year, aimed at those interested in reinforcement learning from verifiable rewards.
@dwarkesh_sp: What does the next training paradigm look like? 0:00:00 – The big research bet the labs are making 0:02:12 – Grindabili…
A discussion on the next training paradigm for AI, covering research bets, grindability, RLVR, and a vision for 2027.