Tag
A technical note on performing RL post-training across 14 Macs distributed in 4 countries, highlighting distributed compute for reinforcement learning.