Tag
The author expresses anticipation for future Grok AI models (4.7 and 4.8), praising Grok 4.6 for being steerable, fast, and avoiding over-engineering.
This paper proposes using Masked Diffusion Language Models (MDLMs) as text-based world models for agentic reinforcement learning, showing that their any-order denoising objective avoids prefix mode collapse and leads to stronger performance than autoregressive baselines.