scene-generation

Tag

Cards List
#scene-generation

SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution

Hugging Face Daily Papers · 2026-09-04 Cached

SceneMosaic combines learned image priors with vision-language agents to efficiently generate diverse, physically valid indoor scenes, achieving a 24x speedup and improved physical validity over existing methods.

0 favorites 0 likes
#scene-generation

AI agents create virtual playgrounds to help robots get crucial training data

MIT News — Artificial Intelligence · 2026-07-13 Cached

MIT CSAIL and Toyota Research Institute introduce SceneSmith, a system using AI agents powered by GPT-5.2 to automatically generate rich 3D virtual scenes, providing diverse simulation environments for robot training without extensive real-world testing.

0 favorites 0 likes
#scene-generation

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

Hugging Face Daily Papers · 2026-07-13 Cached

Xiaomi Robotics introduces U0, a 38-billion-parameter multimodal autoregressive model for unified embodied synthesis, treating embodied generation as an extension of image and video generation. It achieves state-of-the-art results on multiple embodied tasks, outperforming GPT-Image-2.0 and improving real-world manipulation success rates.

0 favorites 0 likes
#scene-generation

MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models

Hugging Face Daily Papers · 2026-07-13 Cached

MAGIC is a system that uses large language models to generate connected, navigable multi-scene game worlds from a single natural-language prompt, addressing cross-scene consistency, in-scene navigability, and transition evaluation. It achieves high precision and recall on a benchmark of 100 multi-scene cases.

0 favorites 0 likes
#scene-generation

SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion

Hugging Face Daily Papers · 2026-07-06 Cached

SynCity 3000 introduces a framework for generating large, globally coherent 3D scenes by adapting image-to-3D generators as convolutional operators, fine-tuned on synthetic scene data from a new data engine.

0 favorites 0 likes
#scene-generation

SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation

Hugging Face Daily Papers · 2026-06-26 Cached

SimFoundry is a modular system that automates real-to-sim scene construction from video, generating digital twins and affordance-preserving variations for zero-shot robot policy training, achieving strong transfer to real-world tasks and high simulation-to-real performance prediction.

0 favorites 0 likes
#scene-generation

Triangle Splats from Video Diffusion Latents (5 minute read)

TLDR AI · 2026-06-25 Cached

FLAT is a method that directly decodes explicit triangle splats from compressed video diffusion latents in a single forward pass, improving geometric accuracy while enabling fast rasterization and physics-based interaction.

0 favorites 0 likes
#scene-generation

FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation

Hugging Face Daily Papers · 2026-06-23 Cached

FLAT proposes a method to decode explicit triangle splats directly from video diffusion latents for geometrically accurate 3D scene generation. It introduces a ray-centered rotation parameterization and a product window function to improve gradient flow, achieving better geometric accuracy than prior feedforward methods while supporting real-time rendering.

0 favorites 0 likes
#scene-generation

Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite Image

Hugging Face Daily Papers · 2026-05-14 Cached

Sat3DGen introduces a geometry-first approach for generating street-level 3D scenes from a single satellite image, achieving improved geometric accuracy and photorealism through novel constraints and training strategies. The method demonstrates significant improvements over prior work on the VIGOR-OOD benchmark.

0 favorites 0 likes
#scene-generation

HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds

Hugging Face Daily Papers · 2026-04-15 Cached

HY-World 2.0 is a multi-modal world model framework that generates high-fidelity 3D Gaussian Splatting scenes from text, images, and videos through specialized modules for panorama generation, trajectory planning, and scene composition, achieving state-of-the-art performance among open-source approaches.

0 favorites 0 likes
#scene-generation

MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse

Papers with Code Trending · 2025-03-24 Cached

MetaSpatial is a reinforcement learning framework that enhances 3D spatial reasoning in vision-language models, enabling coherent and physically plausible 3D scene generation without hard-coded optimizations.

0 favorites 0 likes
← Back to home

Submit Feedback