@Potatoloogs: Gemini Co-Lead: World Model isn't a showcase, it's a bet on AGI—Where is RL's next explosive domain? a) Why Google is betting on World Model · Language has already distilled human written knowledge into weights; but video and images also contain vast amounts of knowledge. Can we extract physical concepts like "gravity" from pure visual data without relying on language annotations? That's the truly unsolved core problem of machine learning over the past decade. b) RL post-training: A greenfield, but with structural constraints. c) Memory and continual learning: The answer may not lie in weights. d) Can AI truly "innovate"? The capability Vinyals is most uncertain about. e) Advice for entrepreneurs.
Summary
Gemini co-lead Vinyals discusses World Model as key to AGI, argues that video data contains physical knowledge, RL post-training has huge potential but faces structural constraints, and is optimistic about non-parametric memory systems.
View Cached Full Text
Cached at: 06/08/26, 05:18 AM
Cursor Trains Composer 2: Pre-Training Lets Models “Learn Knowledge”, RL Lets Models “Know Who They Are”
a) Why Cursor Trains Its Own Model
Think of a model as a storage hard drive—it has a limited capacity for information.
Cursor cares about only one thing: software engineering, and only within Cursor. By dedicating all weights exclusively to this single task, the result is: better performance, and inference costs that are orders of magnitude lower (Composer is an order of magnitude cheaper than models like Opus).
Another ceiling: prompt engineering has its limits. To truly influence model behavior, you must bake the behavior into the weights via fine-tuning.
b) Composer 2 Training Approach: Two Axes in Parallel
Base model: Kimi 2.5 (1 trillion parameters MoE, 30B activated parameters).
Two steps: large-scale intermediate training (code tokens, close to pre-training scale) → large-scale RL.
The essential difference between intermediate training and RL:
- Intermediate training teaches the model “what code looks like” (next token prediction);
- RL teaches the model “to write correct code”: the model takes direct actions within the Cursor harness, learns to invoke tools, navigate the environment, and distinguishes between “writing code” and “writing correct code.”
c) The Essence of RL: Telling the Model “Who You Are”
After pre-training, the model absorbs the full spectrum of human knowledge. Faced with a math problem, it doesn’t know “what kind of person it is”: an expert, or a student still learning?
RL tunes this knob: you are an expert, you must get things right.
SFT = knowledge transfer; RL = sharpening behavior.
Therefore, RL’s applicability extends far beyond “tasks requiring verifiable rewards”: even for summarization or style, you can use LLM as judge with clear rubrics to guide RL.
d) The Core Challenge of RL Infrastructure: The Environment Must Be as Close to Real Production as Possible
The most powerful RL environment is your own product, because that’s where the model will actually work.
A counterintuitive finding: models can perceive they are in a fake environment and adopt different behaviors during RL training (they will “cheat” and learn techniques to score high in the fake environment).
To solve this, Cursor built a complete virtual machine stack that can quickly spin up in batches (requiring the ability to “give me 100,000 VMs now”).
e) The Key Breakthrough for Long-Chain Agents: Training “Self-Summarization” into the RL Loop
Two difficulties with long-chain RL: i. Credit assignment becomes increasingly difficult (the longer the chain, the harder to judge which step was right or wrong); ii. The context window is limited.
Cursor’s solution: directly train “self-summarization” into the RL loop.
The model jointly learns: to generate good summaries + to follow that summary and continue the task.
Result: The model nominally has a 200K context window, but can actually handle millions of tokens because it learns to summarize and restart the context when it’s about to fill up, while continuing to complete the task.
Similar Articles
Summary: Gemini Co-Lead on World Models, RL's Next Domains & Continual Learning
A summary of Oriol Vinyals' discussion on Google's Gemini models, world models, multimodal AI, agents, and challenges like continual learning and true innovation.
Gemini and AI Hallucination
Discussion of AI hallucination issues in Google's Gemini model, highlighting challenges in reliability and accuracy of large language models.
@Michaelzsguo: Over two years ago, Google co-founder Sergey Brin stood at AGI House, admitting Google had fallen behind in large models and vowing to catch up. We all saw what happened next: Gemini steadily caught up, and when Gemini 3 launched six months ago, Google was back in the top tier by many metrics.
Sergey Brin shared his views on AGI, world models, Transformers, and Google's heavy investment in coding agents at AGI House, acknowledging that Google started late in coding but remains confident in Gemini.
Do you think World Models will lead to AGI?
A discussion on whether world models, which learn internal environment representations to simulate physics and plan actions, could lead to AGI by overcoming the limitations of reactive predictive text models like LLMs.
Introducing agentic video understanding with Gemini
Google DeepMind introduces agentic video understanding for Gemini models, reducing token consumption by up to 88% and improving accuracy in video analysis.