scene-understanding

Tag

Cards List
#scene-understanding

SceneActBench: Can Agents Act on the 3D Scenes They See?

Hugging Face Daily Papers · 2026-07-24 Cached

SceneActBench is a benchmark for evaluating VLM agents on acting in complete multi-object 3D scenes, using task-specific geometric metrics across five tasks.

0 favorites 0 likes
#scene-understanding

@thesupermanmx: China open-sourced a model that reconstructs any scene in 3D from a regular video, in real-time. one camera. no LiDAR. …

X AI KOLs Timeline · 2026-07-16

China open-sourced a real-time 3D scene reconstruction model that works from a single regular video without LiDAR, achieving 20 FPS on a single GPU and maintaining stability over 10,000+ frames.

0 favorites 0 likes
#scene-understanding

Repurposing 3D Generative Model for Autoregressive Layout Generation

Hugging Face Daily Papers · 2026-04-17 Cached

LaviGen is a framework that repurposes 3D generative models for autoregressive 3D layout generation, using an adapted 3D diffusion model with dual-guidance self-rollout distillation to achieve 19% higher physical plausibility and 65% faster computation than state-of-the-art methods on the LayoutVLM benchmark.

0 favorites 0 likes
← Back to home

Submit Feedback