3d-vision

Tag

Cards List
#3d-vision

[R] Would you keep a robot demonstration if hand tracking missed the moment the plug went in?

Reddit r/MachineLearning ↗ · 1h ago Cached

MEgoVista is an offline pipeline from Maniformer that turns unprepared egocentric MEgo View recordings into metric two-hand and head motion in a gravity-aligned world frame, using calibrated stereo for metric scale and validating outputs against independent Chingmu optical motion capture. The work positions itself as a scalable, unconstrained alternative to studio rigs for generating metric hand supervision for robot manipulation learning.

0 favorites 0 likes
#3d-vision

Generative Semantic Scene Completion

Hugging Face Daily Papers ↗ · 2026-08-27 Cached

This paper presents generative semantic scene completion with corrected performance metrics and real-time inference, releasing code, weights, and the PS3 corpus for reproducibility.

0 favorites 0 likes
#3d-vision

Projector Is All You Train

arXiv cs.CL ↗ · 2026-08-21 Cached

This paper demonstrates that training only the projector in multimodal large language models achieves strong performance on 3D tasks, avoids language model drift, and improves training efficiency compared to joint training methods.

0 favorites 0 likes
#3d-vision

Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models

Hugging Face Daily Papers ↗ · 2026-08-11 Cached

Presents Self-Geometry, a plug-and-play test-time adaptation pipeline that enforces explicit multi-view geometric constraints using 2D pixel correspondences to improve geometrically consistent 3D vision foundation models.

0 favorites 0 likes
#3d-vision

3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering

Hugging Face Daily Papers ↗ · 2026-08-02 Cached

3DZip is a three-stage token compression framework for 3D vision-language models that uses voxelization, diversity-guided anchor selection, and spatial constraint merging to reduce tokens while preserving spatial reasoning. It retains 94.7% of original performance with only 128 tokens and achieves 1.92x faster inference on 3DQA benchmarks.

0 favorites 0 likes
#3d-vision

MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors

Hugging Face Daily Papers ↗ · 2026-07-13 Cached

MetaView proposes a diffusion-based monocular novel view synthesis framework that combines implicit geometry priors with metric depth guidance to achieve consistent and controllable rendering under large viewpoint changes from a single image.

0 favorites 0 likes
#3d-vision

A Cookbook of 3D Vision: Data, Learning Paradigms, and Application

Hugging Face Daily Papers ↗ · 2026-06-02 Cached

This paper presents a comprehensive taxonomy of 3D vision research, covering geometric representations, datasets, learning paradigms, and applications in reconstruction, generation, and video modeling.

0 favorites 0 likes
#3d-vision

Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence

Hugging Face Daily Papers ↗ · 2026-05-28 Cached

This paper introduces a post-training framework that leverages 3D priors from SAM3D to improve semantic correspondence in 2D foundation features, addressing issues like left-right confusion and repeated parts. The method uses instance-specific 3D reconstruction without pose annotations or spherical geometry shortcuts.

0 favorites 0 likes
#3d-vision

SpatialBench: Is Your Spatial Foundation Model an All-Round Player?

Hugging Face Daily Papers ↗ · 2026-05-26 Cached

SpatialBench is a comprehensive benchmark for evaluating spatial foundation models across diverse domains and tasks, revealing limitations in current models and introducing DA-Next-5M and DA-Next to advance spatial representation learning.

0 favorites 0 likes
#3d-vision

@rwayne: Regarding how to pursue a PhD, a Zhejiang University expert posted the solution directly on GitHub. It covers the entire research lifecycle, including getting started, topic selection, conducting experiments, advisor meetings, project management, writing, rebuttals, and presentation slides. The `getting_started` file for the 3D Vision direction...

X AI KOLs Timeline ↗ · 2026-05-10

A Zhejiang University researcher shared a comprehensive PhD guide on GitHub, covering the entire research lifecycle from topic selection to rebuttals, specifically tailored for the 3D Vision direction.

0 favorites 0 likes
#3d-vision

facebook/VGGT-Omega

Hugging Face Models Trending ↗ · 2026-03-17 Cached

Meta AI and Oxford VGG released VGGT-Omega, a foundation model for 3D vision, with project page and GitHub repository.

0 favorites 0 likes
← Back to home

Submit Feedback