Tag
SceneActBench is a benchmark for evaluating VLM agents on acting in complete multi-object 3D scenes, using task-specific geometric metrics across five tasks.
China open-sourced a real-time 3D scene reconstruction model that works from a single regular video without LiDAR, achieving 20 FPS on a single GPU and maintaining stability over 10,000+ frames.
LaviGen is a framework that repurposes 3D generative models for autoregressive 3D layout generation, using an adapted 3D diffusion model with dual-guidance self-rollout distillation to achieve 19% higher physical plausibility and 65% faster computation than state-of-the-art methods on the LayoutVLM benchmark.