Goat Simulator and Gemini

Reddit r/singularity News

Summary

Google's SIMA 2 agent, powered by Gemini, shows a major leap in 3D game task success (65% vs 31%), trained via Genie 3's infinite world generation. The article notes recent deployments like Project Genie and Gemini Robotics 2, suggesting ongoing agentic AI progress.

While casually pretending I wasn't doomscrolling, I came across a really interesting model from Google SIMA 2. It’s an agent that lives inside a 3D world. It sees the game just like we do raw screen pixels, zero access to the game’s backend code, and its output is just keyboard and mouse inputs. But unlike current models, it actually understands what it’s looking at: it plans steps, explains what it’s about to do, and fine-tunes itself on the fly when unexpected things happen. It has Gemini under the hood, which makes all the difference a 65% task success rate compared to SIMA 1’s 31%. But as usual, it’s not all sunshine and rainbows... Training relies on games, which are finite, so the model can easily end up just memorizing mechanics or specific actions which kind of defeats the purpose. Why do we need a model that only knows how to play Goat Simulator? That’s why Google built Genie 3 a world model generator where SIMA 2 can actually train right now. Infinitely generating random environments with different rules and events forces Gemini not to just memorize Goat Simulator, but to actually learn how to learn. And this isn't just some tech demo on a blog anymore: on January 29, 2026, Google opened Project Genie to AI Ultra subscribers, and in February, Waymo used Genie 3 to build its own World Model for simulating autonomous driving. The tech is actually being deployed seriously just not in games yet. Training in diverse virtual worlds means we could eventually plug SIMA 2 into almost anything as a teammate, or even as a dynamic event generator. Just imagine what games could look like if they were truly alive and filled with AI-generated dynamic events. Anyone who plays roleplaying games in places like JanitorAI or CharacterAI knows how massive that scale can get. We're already seeing a glimpse of this replayability potential in the Skyrim mod "Mantella Bring NPCs to Life with AI": you’ve probably seen tons of videos of people playing with it and how mind-blowing it is. Quests and storylines could be procedurally generated on the fly, unique every single time and that's just at the text/dialogue level. Imagine when this tech levels up on a much broader scale. (Maybe that’s why Nvidia is betting so heavily on neural networks and even if I don't buy that Huang is doing it out of the goodness of his heart for gamers, it's still undeniably cool). Unfortunately, we haven't heard much about SIMA 2 itself since November 13, 2025. Or have we? On July 30, 2026, DeepMind released Gemini Robotics 2. What do we know about it? It’s an AI with a physical body except instead of a 3D simulation, it operates in the real world. And what do we see? Literally the exact same stack as SIMA 2, just brought to reality: perception, reasoning, planning, and execution. Plus, version 2 controls full humanoids walking, crouching, fine-finger control not just a pair of robotic arms over a table (as shown on Apptronik’s Apollo 2). This strongly suggests Google’s research in this area is alive and well, and that we’ll eventually see SIMA 3 and Genie 4 down the line. The real question is: when? My bet? Probably around the release of Gemini 4 (or whatever the next frontier generation will be). Starting around version 3.5, Google finally realized they need to lean heavily into agentic AI. SIMA 2 is the definition of an agentic model: it perceives visual inputs, plans steps, acts, self-improves, and communicates with the user. And its huge performance jump came directly from having Gemini as its core brain. This means the next leap for these agents will roll in right alongside the next generation of the base model, not in isolation. Judging by the current pace, we’ll probably have to wait a bit: the flagship Pro last got an update back in February, 3.5 Pro feels delayed, and mostly Flash models are rolling out right now. We just need Google not to fumble the bag—but at least the roadmap and developmental logic are clearly there, and we just have to wait it out. Business as usual, really. Why did I write this? Honestly, I’m just tired of the constant doomposting and sketchy leaks from random clout-chasers on this subreddit, so I wanted to share something actually interesting to chew on. Good luck, everyone, and thanks for reading!
Original Article

Similar Articles

Gemini 2.5: Our most intelligent AI model

Google DeepMind Blog

Google announced Gemini 2.5, its most intelligent AI model, with Gemini 2.5 Pro Experimental leading LMArena benchmarks by significant margins and demonstrating enhanced reasoning and coding capabilities through improved thinking model architecture.

Gemini 3.5: frontier intelligence with action

Google DeepMind Blog

Google announces Gemini 3.5, a new family of AI models focused on agentic workflows and coding, starting with 3.5 Flash which delivers frontier performance at high speed.

Gemini 2.5: Our most intelligent models are getting even better

Google DeepMind Blog

Google announces Gemini 2.5 series updates, including improved 2.5 Pro and Flash models with new capabilities like Deep Think (enhanced reasoning mode), native audio output, and computer use abilities via Project Mariner. The models now lead on WebDev Arena and LMArena leaderboards.