@seclink: 罗老师真的牛,又有一个大创新,号召大厂都学习需学习...

X AI KOLs Following 模型

摘要

Fuli Luo 分享了关于 MiMo-V2.6 模型在强化学习方面的扩展研究,重点关注计算规模、环境和工具链的提升。

罗老师真的牛,又有一个大创新,号召大厂都学习需学习...
查看原文
查看缓存全文

缓存时间: 2026/09/17 04:08

罗老师真的牛,又有一个大创新,号召大厂都学习需学习…

Fuli Luo (@_LuoFuli): Nearly half a year of silence. We spent it studying one problem: how far RL can scale.

MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task

相似文章