Tag
SceneMosaic combines learned image priors with vision-language agents to efficiently generate diverse, physically valid indoor scenes, achieving a 24x speedup and improved physical validity over existing methods.
MIT CSAIL and Toyota Research Institute introduce SceneSmith, a system using AI agents powered by GPT-5.2 to automatically generate rich 3D virtual scenes, providing diverse simulation environments for robot training without extensive real-world testing.
Xiaomi Robotics introduces U0, a 38-billion-parameter multimodal autoregressive model for unified embodied synthesis, treating embodied generation as an extension of image and video generation. It achieves state-of-the-art results on multiple embodied tasks, outperforming GPT-Image-2.0 and improving real-world manipulation success rates.
MAGIC is a system that uses large language models to generate connected, navigable multi-scene game worlds from a single natural-language prompt, addressing cross-scene consistency, in-scene navigability, and transition evaluation. It achieves high precision and recall on a benchmark of 100 multi-scene cases.
SynCity 3000 introduces a framework for generating large, globally coherent 3D scenes by adapting image-to-3D generators as convolutional operators, fine-tuned on synthetic scene data from a new data engine.
SimFoundry is a modular system that automates real-to-sim scene construction from video, generating digital twins and affordance-preserving variations for zero-shot robot policy training, achieving strong transfer to real-world tasks and high simulation-to-real performance prediction.
FLAT is a method that directly decodes explicit triangle splats from compressed video diffusion latents in a single forward pass, improving geometric accuracy while enabling fast rasterization and physics-based interaction.
FLAT proposes a method to decode explicit triangle splats directly from video diffusion latents for geometrically accurate 3D scene generation. It introduces a ray-centered rotation parameterization and a product window function to improve gradient flow, achieving better geometric accuracy than prior feedforward methods while supporting real-time rendering.
Sat3DGen introduces a geometry-first approach for generating street-level 3D scenes from a single satellite image, achieving improved geometric accuracy and photorealism through novel constraints and training strategies. The method demonstrates significant improvements over prior work on the VIGOR-OOD benchmark.
HY-World 2.0 is a multi-modal world model framework that generates high-fidelity 3D Gaussian Splatting scenes from text, images, and videos through specialized modules for panorama generation, trajectory planning, and scene composition, achieving state-of-the-art performance among open-source approaches.
MetaSpatial is a reinforcement learning framework that enhances 3D spatial reasoning in vision-language models, enabling coherent and physically plausible 3D scene generation without hard-coded optimizations.