Tag
World model AI companies like AMI Labs and World Labs are maintaining secrecy about their commercialization plans due to competitive pressures and ease of fundraising in the industry.
VABench introduces a benchmark to evaluate embodied spatial intelligence in models by testing their ability to observe, reason, and act through visual demonstrations and active perception. It shows that active camera control improves task success, but no model completes long-horizon episodes.
This paper introduces SpatialBlock-15k, a synthetic dataset for block-stacking problems, to enhance 3D spatial reasoning in large vision-language models, demonstrating improved performance and generalization to real-world tasks.
Atlas is introduced as a foundation model for spatial intelligence in a video that features atmospheric music and sparse spoken words.
World Labs announced Atlas, an omni world model that natively processes text, images, video, and 3D geometry in a shared spatial context to advance AI capabilities in simulating space, time, and physical interactions.
World Labs introduces Atlas, a multimodal world model that can generate, reconstruct, and simulate 3D environments with camera control and spatial consistency.
LightNav-0 is a compact generalist navigation model that leverages pretrained vision-language models' spatial intelligence to achieve state-of-the-art embodied navigation across diverse tasks and robot embodiments.
This paper introduces a continuous metric field framework trained by a single causal contrastive loss that unifies geometric structure discovery from robot navigation to black hole emergence, demonstrating zero-shot generalization across dimensions.
Stanford HAI discusses the emergence of world models and spatial intelligence in AI, calling for governance frameworks that address capabilities beyond language processing.
World Labs shares early results on building worlds to train robots, advancing spatial intelligence for interaction with physical and virtual environments.
Promotes the eBook 'Python for GIS & Spatial Intelligence' which teaches interactive mapping and spatial analysis using Python.
Google Maps Platform Product Manager Angela Yu discussed the shift from developers to builders in the AI era at NYC Community Day, emphasizing agents, spatial intelligence, and how Google's data empowers builders.
An Ars Technica article explores the promise and limitations of world models as an emerging AI paradigm, contrasting them with LLMs and featuring expert insights on their applications in robotics, research, and asset generation.
WildCity introduces a large-scale multimodal dataset for city-scale urban navigation and spatial representation, collected by autonomous fleets. It provides 18 long trajectories and establishes baselines for reconstruction and closed-loop simulation to advance AI systems that can perceive and reason about city-scale environments.
This paper introduces S-Agent, which uses spatial tool-use to elicit reasoning for spatial intelligence.
Fei-Fei Li's World Labs, focused on spatial intelligence AI, raised $1 billion in funding, signaling a shift from large language models to world models.
This paper proposes Embodied-BenchClaw, an autonomous multi-agent system that automatically constructs embodied spatial intelligence benchmarks from user intent through a five-stage pipeline with process quality control and an extensible Skill Library.
Fei-Fei Li explains that World Labs focuses on building large world models to unlock spatial intelligence, considering this the next frontier after language models, and argues its value from perspectives of evolutionary history, application scenarios, and technology classification, while expressing a pragmatic attitude towards AI safety and the necessity of educational reform.
GeoVR enhances multimodal large language models with 3D awareness by restructuring their semantic latent space through geometric knowledge distillation from 3D foundation models using multiple geometric targets.
Fei-Fei Li and the World Labs team present a functional taxonomy of world models, distinguishing between renderers, physics engines, and other components within the reinforcement learning loop, and arguing that spatial intelligence is AI's next frontier.