rgb-d

Tag

Cards List
#rgb-d

Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator

Hugging Face Daily Papers · 2026-07-07 Cached

Image2Sim is a neural simulation framework that creates high-fidelity interactive environments from RGB-D images, enabling scalable training for embodied navigation agents. It generates nearly 20K scenes and over 10 million training samples, showing strong benchmark improvements and effective real-world zero-shot transfer.

0 favorites 0 likes
#rgb-d

RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

Hugging Face Daily Papers · 2026-07-07 Cached

RynnWorld-4D is a generative world model that co-produces future RGB, depth, and optical flow from a single RGB-D image and language instruction using a unified diffusion process, enabling efficient robotic manipulation through inverse dynamics policy learning. It achieves state-of-the-art on real-world bimanual manipulation tasks.

0 favorites 0 likes
#rgb-d

Human Universal Grasping

Hugging Face Daily Papers · 2026-06-15 Cached

A flow-matching model generates diverse human grasps from RGB-D images, enabling zero-shot robotic grasping with improved performance over existing methods. The model, trained on a large egocentric dataset, significantly outperforms state-of-the-art baselines on a new benchmark.

0 favorites 0 likes
#rgb-d

Revisiting Articulated Parts Perception in Robot Manipulation

Hugging Face Daily Papers · 2026-06-06 Cached

This paper introduces Geometric Primary Structure (GPS), a new representation for articulated parts perception in robot manipulation, enabling efficient VR-based annotation and achieving a 73% success rate without fine-tuning.

0 favorites 0 likes
#rgb-d

AFUN: Towards an Affordance Foundation Model for Functionality Understanding

Hugging Face Daily Papers · 2026-06-01 Cached

AFUN proposes an affordance foundation model that predicts functional masks and 3D motion curves from RGB-D observations and language descriptions, enabling generalizable robot manipulation across diverse environments. The model outperforms baselines on multiple benchmarks and can be deployed for real-world tasks without fine-tuning.

0 favorites 0 likes
#rgb-d

CM-EVS: Sparse Panoramic RGB-D-Pose Data for Complete Scene Coverage

Hugging Face Daily Papers · 2026-05-15 Cached

This paper proposes COVER, a training-free method for converting 3D assets into sparse panoramic RGB-D-pose data with complete scene coverage and low redundancy, and introduces the CM-EVS dataset containing 36,373 curated frames from indoor and outdoor scenes.

0 favorites 0 likes
← Back to home

Submit Feedback