robotic-manipulation

Tag

Cards List
#robotic-manipulation

Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

Hugging Face Daily Papers ↗ · 2026-09-17 Cached

Introduces Movement Trend Guidance to enhance 3D diffusion policies in robotic manipulation by providing foresight without explicit trajectories, achieving improved performance on benchmarks like RoboTwin2.0 and LIBERO-40.

0 favorites 0 likes
#robotic-manipulation

Agent as Policy for Robotic Manipulation

Hugging Face Daily Papers ↗ · 2026-09-11 Cached

A general-purpose agent directly controls a physical robot by interpreting visuals, writing executable programs, and revising actions based on physical feedback across diverse manipulation tasks, achieving high success rates without task-specific training.

0 favorites 0 likes
#robotic-manipulation

Memory as Plans: World-Action Modeling with Memory-Grounded Planning

Hugging Face Daily Papers ↗ · 2026-09-10 Cached

MaP-WAM introduces a framework for non-Markovian robotic manipulation by separating memory-grounded planning from execution, using episodic memory to maintain fixed inference latency and achieving state-of-the-art results.

0 favorites 0 likes
#robotic-manipulation

Why is this robot slowly dropping the object?

Lobsters Hottest ↗ · 2026-09-09 Cached

Researchers enable simple robotic grippers to achieve precise in-hand object manipulation through controlled sliding using tactile sensing and advanced friction models, mimicking human dexterity.

0 favorites 0 likes
#robotic-manipulation

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

Hugging Face Daily Papers ↗ · 2026-09-04 Cached

This paper introduces GE-Act 2.0, a world-action model pretrained from scratch to enable scalable zero-shot robotic manipulation with improved success rates across diverse tasks and conditions.

0 favorites 0 likes
#robotic-manipulation

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

Hugging Face Daily Papers ↗ · 2026-08-31 Cached

ZimaBlue introduces a scalable framework for learning generalizable world action models from large-scale egocentric video, substantially improving zero-shot robotic manipulation through a three-stage curriculum and slow-fast architecture.

0 favorites 0 likes
#robotic-manipulation

PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration

Hugging Face Daily Papers ↗ · 2026-08-21 Cached

PhysCaP is a physics-informed code-generation agent that actively explores objects to infer hidden physical properties for efficient robotic manipulation.

0 favorites 0 likes
#robotic-manipulation

GOAG: Generative and Object-Agnostic Grasp Planner for Dexterous Robotic Manipulation

Hugging Face Daily Papers ↗ · 2026-08-20 Cached

GOAG is a deep generative grasp planner that learns a gripper-specific contact surface distribution to sample valid grasps for unseen objects without object-specific training, achieving state-of-the-art results on the MultiDex dataset for dexterous robotic manipulation.

0 favorites 0 likes
#robotic-manipulation

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

Hugging Face Daily Papers ↗ · 2026-08-13 Cached

DreamX-Phi 1.0 is an action-conditioned video world model for robotic manipulation that predicts future observations using geometric attention encoding, depth estimation, and distillation, achieving top results in the WorldArena 2.0 Challenge.

0 favorites 0 likes
#robotic-manipulation

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Hugging Face Daily Papers ↗ · 2026-07-29 Cached

TurboVLA introduces a new Vision-Language-Action paradigm that directly maps vision and language to action, achieving 97.7% success on LIBERO with only 0.2B parameters and real-time inference at 32 Hz on consumer GPUs, significantly reducing computational cost.

0 favorites 0 likes
#robotic-manipulation

TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable Manipulation

Hugging Face Daily Papers ↗ · 2026-07-23 Cached

TableVerse introduces a fully automated Real2Sim pipeline that converts unstructured, in-the-wild images into high-fidelity, simulation-ready tabletop environments with accurate metrics and physical stability, along with a large-scale dataset (TableVerse-100K) for generalizable robotic manipulation.

0 favorites 0 likes
#robotic-manipulation

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

Hugging Face Daily Papers ↗ · 2026-07-08 Cached

LaMem-VLA proposes a latent-memory-native framework that integrates short-term and long-term historical experience directly into Vision-Language-Action reasoning, enabling better performance on long-horizon robotic manipulation tasks.

0 favorites 0 likes
#robotic-manipulation

RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

Hugging Face Daily Papers ↗ · 2026-07-07 Cached

RynnWorld-4D is a generative world model that co-produces future RGB, depth, and optical flow from a single RGB-D image and language instruction using a unified diffusion process, enabling efficient robotic manipulation through inverse dynamics policy learning. It achieves state-of-the-art on real-world bimanual manipulation tasks.

0 favorites 0 likes
#robotic-manipulation

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation

Hugging Face Daily Papers ↗ · 2026-06-26 Cached

PhysisForcing is a training framework that enhances embodied video generation for robotic manipulation by enforcing physical consistency through pixel-level trajectory alignment and semantic-level relational alignment losses in a DiT-based architecture, achieving notable improvements on benchmarks.

0 favorites 0 likes
#robotic-manipulation

Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents

Hugging Face Daily Papers ↗ · 2026-06-22 Cached

Foresight is a failure detection framework for long-horizon robotic manipulation that uses action-conditioned world model latents and functional conformal prediction to monitor trajectories, trained only with final task labels. It demonstrates state-of-the-art performance across simulation and real robot tasks.

0 favorites 0 likes
#robotic-manipulation

EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies

Hugging Face Daily Papers ↗ · 2026-06-18 Cached

EventVLA introduces a sparse visual evidence memory framework for long-horizon robotic manipulation, achieving an average success rate improvement of +40% over state-of-the-art memory-augmented VLAs.

0 favorites 0 likes
#robotic-manipulation

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Hugging Face Daily Papers ↗ · 2026-06-17 Cached

Presents Qwen-RobotManip, a Vision-Language-Action foundation model for robotic manipulation that achieves generalization through unified alignment across representation, motion, and behavior dimensions, enabling large-scale training on diverse data sources. It outperforms prior state-of-the-art models across multiple out-of-distribution benchmarks and demonstrates emergent capabilities like zero-shot instruction following and cross-embodiment transfer.

0 favorites 0 likes
#robotic-manipulation

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation

Hugging Face Daily Papers ↗ · 2026-06-16 Cached

PAIWorld enhances diffusion-transformer world models with geometric awareness and cross-view attention to improve multi-view 3D consistency for robotic manipulation tasks, achieving state-of-the-art results on benchmarks.

0 favorites 0 likes
#robotic-manipulation

WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation

Hugging Face Daily Papers ↗ · 2026-06-11 Cached

WEAVER is a multi-view world model for robotic manipulation that achieves high fidelity, consistency, and efficiency using flow-matching loss, demonstrating superior performance in policy evaluation, improvement, and test-time planning with significant real-world improvements.

0 favorites 0 likes
#robotic-manipulation

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding

Hugging Face Daily Papers ↗ · 2026-06-04 Cached

AffordanceVLA introduces a unified framework using structured affordance forecasting as an intermediate representation to improve perception-action mapping in robotic manipulation, leveraging vision-language models and a Mixture-of-Transformer architecture.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback