3d-scene-understanding

Tag

Cards List
#3d-scene-understanding

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

Hugging Face Daily Papers · 2026-08-05 Cached

SmartMage is a unified multimodal LLM that dynamically orchestrates visual and geometric modalities for query-dependent 3D scene understanding, achieving state-of-the-art results across five benchmarks.

0 favorites 0 likes
#3d-scene-understanding

MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs

arXiv cs.AI · 2026-07-13 Cached

MultiView-Bench is a diagnostic benchmark for evaluating vision-language models on their ability to integrate multiple viewpoints into a coherent 3D mental model, revealing systematic failures in 3D spatial reasoning, and introducing ViewNavigator to mitigate these issues.

0 favorites 0 likes
#3d-scene-understanding

CM-EVS: Sparse Panoramic RGB-D-Pose Data for Complete Scene Coverage

Hugging Face Daily Papers · 2026-05-15 Cached

This paper proposes COVER, a training-free method for converting 3D assets into sparse panoramic RGB-D-pose data with complete scene coverage and low redundancy, and introduces the CM-EVS dataset containing 36,373 curated frames from indoor and outdoor scenes.

0 favorites 0 likes
#3d-scene-understanding

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation

Hugging Face Daily Papers · 2026-04-30 Cached

This paper introduces HERMES++, a unified driving world model that integrates 3D scene understanding and future geometry prediction using BEV representation, LLM-enhanced queries, and joint geometric optimization.

0 favorites 0 likes
← Back to home

Submit Feedback