World Labs' Fei-Fei Li on Creating Large World Models

Reddit r/singularity News

Summary

Fei-Fei Li explains that World Labs focuses on building large world models to unlock spatial intelligence, considering this the next frontier after language models, and argues its value from perspectives of evolutionary history, application scenarios, and technology classification, while expressing a pragmatic attitude towards AI safety and the necessity of educational reform.

No content available
Original Article
View Cached Full Text

Cached at: 06/05/26, 11:11 AM

TL;DR: Li Fei-Fei explains that World Labs, which she co-founded, is focused on building large world models to unlock spatial intelligence. She argues this is the next frontier after language models, supporting this view with evolutionary history, application scenarios, and a technical taxonomy (renderers, planners, simulators). She also expresses a pragmatic attitude toward AI safety and the necessity of educational reform. ## Large World Models and Spatial Intelligence Everyone is focused on large language models like ChatGPT and Claude, but you've raised a billion dollars to build something different. Large world models provide the rationale. What are you betting on that others aren't? Well, this is the startup I co-founded, and we're fully committed to spatial intelligence. The means to achieve spatial intelligence is to build large world models. So what's our reasoning? It's a story spanning 500 million years: animal intelligence began with seeing and moving in the physical world. Evolution started with us as animals, knowing what the world is, who we are, how to move, how to interact. A large part of human life—work, private life—involves perceiving, understanding, reasoning, and interacting with the world, including creativity, productivity, and the imagined worlds of virtual reality. Therefore, unlocking this capability in machines—the ability to generate unfamiliar 3D and 4D worlds, to reason within any world, and to teach agents or robots, or help humans interact with the world—is exactly what spatial intelligence means. That's our focus. ## What World Models Can Do That Language Models Cannot Eliminate text. Put out a fire with text, or fry an omelet? Hmm. I think, in terms of creativity, people design things. Whether it's designing interior spaces, machines, homes, or stories. Much of it goes beyond text. We also use agents, whether in virtual worlds for entertainment like games, or for more serious industrial applications like digital twin design, inspection, or optimization tasks. Or we build robots to help us do many things, from firefighting to medical scenarios to manufacturing—these are all downstream applications of unlocking spatial intelligence and building world models. ## What Will the "ChatGPT Moment" for World Models Look Like? That's a good question, Emily. Because chatting is a consumer activity. The ChatGPT moment usually describes a viral, public consumption moment that feels close to something I can do in the world model domain. The spatial intelligence we're trying to unlock. I'm still trying to figure out if there's a corresponding consumer moment, because the applications we discuss often first go to professionals—professional creators, designers, developers, researchers, and engineers—who use them for robotics, industrial design, etc. So there may not be a consumer moment, but there might be. I'd love a simpler way to design my home, like clicking once to change the curtain color. Sounds cool. ## Taxonomy and Advantages of World Models: Renderers, Planners, Simulators In the past six months, Jonathan started studying world models. Google launched Project Genie. Nvidia has its own world model, Cosmos. Nvidia is also one of your investors. What do you have that they don't? Among competitors, where is your biggest advantage? Well, first, we started this work in 2024. I still remember when we went out to talk about our model and spatial intelligence, just a year later people were still talking about large language models. So we really had a first-mover advantage in understanding that this would be the next frontier. I'm very excited about that. So what do they have that we don't? First, I think we have an incredible team. We have conviction. They certainly don't have the godmother. But the world is big, and I think just like large language models, many companies will do incredible work in world models. Just 24 hours ago, we got tired of the term "world model" being so confused and used in many different ways, so we actually published a blog explaining a functional taxonomy of world models, rather than mixing everything together. In my view, there are currently three ways of calling something a world model related to spatial intelligence. The first I call a **renderer**, a model that generates beautiful pixels on the screen, mostly like video generation models, with consumers primarily being human eyes. While the model promises beautiful pixels, it doesn't necessarily promise physical, dynamical, and geometric correctness. Because it only serves human visual consumption, not necessarily for computation or other tasks. Another kind of world model we call a **planner**, more applicable to machines and robots. It outputs the correct next action based on the input world state or action. This kind of world model is common in robotics applications, and you'll hear it in that context. The third, which I think is key among the three, is the **simulator**. It is consumed by both humans and machines, trying to respect the structure, physics, and dynamics of the world, truly simulating 3D, 4D information, and semantic information. A simulator can become a renderer or a planner, but this layer, in my view, is the critical path to unlocking spatial intelligence. That's what our lab is working on. ## Robotics and World Models: Bridging Hype and Reality This all ties to robotics. I'd like to hear your thoughts on the field, especially humanoid robots. Humanoid robotics funding reached $6 billion. But they still can't load my dishwasher as fast as you, or pick up my Amazon package. Can world models and World Labs bridge the gap between hype and reality? That's a pointed question, Emily. First, this is my job. I understand. First, robotics will be one of the most important revolutions in human industrialization. $6 billion is too small, right? Look at the investment in autonomous driving, look at the investment in language models—the required funds are far more than $6 billion. I'm not saying it's like that now. I think investment needs time, and I hope it's not hype but thoughtful investment in the right efforts. For example, unlocking world modeling, spatial intelligence, and the simulation layer are all part of this important effort. Can we bridge the gap? I do believe World Labs is researching one of the most critical technologies in the physical intelligence field. Obviously, that is where the hope lies. ## AI Safety: Be Wary of Alarmism, Focus on Scientific Foundation You're more cautious about AI safety, skeptical of doomsday talk and over-regulation. Looking across the industry, where do you see real safety work and where is safety theater? Are there people doing it right? Overall, I'm more cautious about any alarmist rhetoric, which frankly makes me boring. I think there's too much hype. Obviously, we need to build the right technology, and we need to set guardrails for it. Whether you use the term responsible, safe, or trustworthy, building the right technology and products that can empower, enhance, and uplift humans without harming them is the goal of all our work, AI or not. Where is it done right? I really hope that every company, every product being built, the people behind it are very mindful of this, thinking about what data we use, what systems we build, what evaluations we conduct, what guardrails we set, how we communicate with users and customers, how we cooperate with regulators, so that at critical moments we are responsible. I do believe a lot of work is happening, not theater. For example, in the pharmaceutical and healthcare industries, companies are integrating AI. I actually came straight from the hospital to your discussion because a family member is about to undergo surgery in an hour. I just saw in the ward how AI is already being used and where it could be used—it's already happening. Doctors use AI to help record medical records, radiologists use AI to assist in reading MRIs and CT scans. I really hope we have more AI to help nurses and families. Last night I got a long radiology report; the first thing I did was send it to AI to help interpret it. So it's all happening, and safety measures are happening too. But more needs to be done in the right way, on a scientific foundation. This should be an ongoing conversation, not what you call theater. ## Educational Reform: AI Must Change Learning We're both mothers, both with teenagers. How do you think AI will change university learning? It must change learning. It must change learning from K to 16. I think this is one of the greatest opportunities for humanity in the next ten years. Because the most valuable resource in the world is human capital. When we have a technology that can answer standardized tests—from Common Core to International Mathematical Olympiad—and do better than the average human, then it's not that humans are bad, but we need to change the education system. We need to change how we evaluate, how we empower teachers to teach, and educate the next generation of students so they can use these tools and powers to do things we never imagined. So do you think our children will still learn? Of course, if we teach them correctly and society prepares them. All children today should not be afraid of AI. They should feel human agency—to lead AI, use it correctly, and use it to make the impact they want on the world. ## Views on AGI: I Don't Engage with the Term, Focus on Scientific Exploration Anthropic's CEO Dario Amodei believes AGI will arrive in 2 to 3 years by scaling the current paradigm. Demis Hassabis says we are at the foothills of the singularity. You've said you don't even engage with the term AGI. Are they wrong? Or is the disagreement just about what we call the goal? I don't engage with the term AGI because the founding fathers of AI as a scientific field had a dream of machines that think and act. It's a scientific exploration, and it's been my lifelong career. I'm still on this exploration path. Now, I'm combining that scientific exploration with building products that improve people's lives. That's the field called AI. Others can call it anything, even an apple. I focus on building a technology that can truly change how people live and work. ## Looking Ahead: Spatial Intelligence Models Will Inspire New Opportunities What will you launch this year that we'll still be talking about next year? I hope we will launch a model for spatial intelligence that will inspire incredible, exciting product opportunities that people have never seen before. *Source: https://youtu.be/pNYVckbCFuk*

Similar Articles

@drfeifei: https://x.com/drfeifei/status/2062247238143996275

X AI KOLs Timeline

Fei-Fei Li and the World Labs team present a functional taxonomy of world models, distinguishing between renderers, physics engines, and other components within the reinforcement learning loop, and arguing that spatial intelligence is AI's next frontier.

World Models Explained: What Every AI Is Missing

Reddit r/ArtificialInteligence

The article explains the concept of world models in detail, comparing them to LLMs, introduces two major camps (pixel prediction and meaning prediction) and representative works such as Dreamer v3, GameNGen, Genie, and JEPA, discusses applications in autonomous driving and robotics, and points out that world models are a key component of physical AI.

LeCun's take on World Models

Reddit r/artificial

Yann LeCun discusses his new lab AMI Labs in Paris, his focus on world models as an alternative to LLMs, and his vision for AI that can plan and act in the physical world, backed by $1B investment and using JEPA architecture.