robot-manipulation

Tag

Cards List
#robot-manipulation

DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation

Hugging Face Daily Papers · 2026-08-06 Cached

DyPES-VLA is a cross-embodiment VLA model that learns shared dynamics priors via future prediction and uses an embodiment-specific Mixture-of-Experts action head to control robots in their native action spaces, achieving state-of-the-art results on LIBERO, RoboCasa, and RoboTwin benchmarks.

0 favorites 0 likes
#robot-manipulation

World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

Hugging Face Daily Papers · 2026-08-05 Cached

Introduces W2-VLA, a vision-language-action model for fine-grained robot manipulation that models task-conditioned future wrist observations, achieving improved manipulation performance on benchmarks and real-world tasks.

0 favorites 0 likes
#robot-manipulation

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

Hugging Face Daily Papers · 2026-08-05 Cached

BridgeVLA++ is a memory-augmented vision-language-action framework for 3D robot manipulation that builds on BridgeVLA to add spatio-temporal memory, achieving state-of-the-art results on memory-dependent manipulation benchmarks while preserving data efficiency and generalization.

0 favorites 0 likes
#robot-manipulation

DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents

Hugging Face Daily Papers · 2026-08-01 Cached

DreamTraj predicts 6-DoF object trajectories from a single RGB image and a language instruction by decoding internal video diffusion latents, eliminating the need for video, depth, or CAD models at inference. It introduces the MOVEdataset with fine-grained language-to-motion annotations and achieves state-of-the-art performance while running 4.6x faster than generate-then-extract pipelines.

0 favorites 0 likes
#robot-manipulation

πR^2: Reactive Real-time Flow Policies

Hugging Face Daily Papers · 2026-07-28 Cached

πR^2 introduces a reactive real-time flow policy for robot manipulation that splits conditioning into fast and slow channels and uses a latency-adaptive flow schedule, enabling closed-loop replanning ~4x faster than baseline policies and improving success rates by up to 30%.

0 favorites 0 likes
#robot-manipulation

Grabette: an open system to record robot-manipulation data

Hugging Face Blog · 2026-07-21 Cached

Grabette is an open, low-cost system for recording robot manipulation data using a handheld gripper and camera, aiming to build a shared dataset for robot learning.

0 favorites 0 likes
#robot-manipulation

@Haoyu_Xiong_: Success rate has long been the primary metric for evaluating robot manipulation. What about speed? Today, we introduce …

X AI KOLs Following · 2026-07-15 Cached

Introduces B-spline Policy (BSP), which parameterizes actions as continuous B-spline curves instead of discrete fixed-rate action chunks, enabling faster and smoother manipulation on low-cost robot arms.

0 favorites 0 likes
#robot-manipulation

CLAP: Direct VLM-to-VLA Adaptation via Language-Action Grounding

arXiv cs.AI · 2026-07-13 Cached

CLAP proposes a method to convert pretrained vision-language models (VLMs) into vision-language-action models (VLAs) by prepending natural-language action descriptions to action token sequences, preserving semantic capabilities without architectural changes. It achieves 90.8% on LIBERO and improves robustness.

0 favorites 0 likes
#robot-manipulation

Prompt-Driven Exploration

arXiv cs.LG · 2026-07-13 Cached

The paper introduces Prompt-Driven Exploration (PDE), a method that uses a vision-language model to iteratively refine natural language prompts for reinforcement learning policies, enabling global exploration and successful policy learning even from zero-reward starts.

0 favorites 0 likes
#robot-manipulation

InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization

Hugging Face Daily Papers · 2026-07-06 Cached

InternVLA-A1.5 integrates pretrained vision-language models with future prediction in latent space to enable efficient robot manipulation with compositional generalization and long-horizon execution, achieving state-of-the-art results on simulation benchmarks.

0 favorites 0 likes
#robot-manipulation

@almond_robotics: Axol folds a towel

X AI KOLs Following · 2026-07-03 Cached

Almond Robotics' Axol robot demonstrates folding a towel.

0 favorites 0 likes
#robot-manipulation

VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon

Hugging Face Daily Papers · 2026-07-02 Cached

VLA-Corrector introduces a lightweight detect-and-correct inference framework that adaptively adjusts action horizons in Vision-Language-Action policies without retraining, improving robustness and efficiency in robot manipulation tasks.

0 favorites 0 likes
#robot-manipulation

Warp RL: Reshaping Base Policy Distributions for Dynamics Adaptation

arXiv cs.LG · 2026-07-01 Cached

Warp RL replaces additive residual corrections in reinforcement learning with an invertible, state-conditioned transformation of the base policy's action distribution using monotonic rational-quadratic spline flows, enabling adaptation of distribution shape, scale, and geometry under dynamics shifts. It matches or outperforms residual correction in ManiSkill3 manipulation tasks and achieves 30% faster task completion in a real robot peg-insertion task.

0 favorites 0 likes
#robot-manipulation

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance

Hugging Face Daily Papers · 2026-06-30 Cached

3D HAMSTER enhances robot manipulation by using a vision-language model with depth encoding to generate 3D trajectories for point cloud-based control, outperforming 2D-guided baselines.

0 favorites 0 likes
#robot-manipulation

Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)

Hugging Face Daily Papers · 2026-06-25 Cached

This paper describes the prizewinning solution for the LeHome Challenge at ICRA 2026, where a two-armed robot learns to fold various garments using a novel RL approach with a self-contained value function, asynchronous training, and heavy sim-to-real augmentation.

0 favorites 0 likes
#robot-manipulation

Geometric Action Model for Robot Policy Learning

Hugging Face Daily Papers · 2026-06-15 Cached

The Geometric Action Model (GAM) repurposes a pretrained geometric foundation model (GFM) as a unified backbone for language-conditioned robot manipulation, achieving higher accuracy, robustness, and efficiency than existing foundation-model-scale baselines across simulation and real-world benchmarks.

0 favorites 0 likes
#robot-manipulation

APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies

Hugging Face Daily Papers · 2026-06-10 Cached

Researchers propose APT, a two-stage training method that pretrains action experts on vision-action pairs before integrating language conditioning, significantly improving out-of-distribution instruction generalization for Vision-Language-Action policies.

0 favorites 0 likes
#robot-manipulation

AEGIS: A Backup Reflex for Physical AI

arXiv cs.AI · 2026-06-08 Cached

AEGIS uses activation-probe early warning to switch to a stronger policy before failures compound in long-horizon robot manipulation, recovering twice as many failures as budget-matched escalation.

0 favorites 0 likes
#robot-manipulation

AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing

Hugging Face Daily Papers · 2026-06-08 Cached

AHA-WAM is an asynchronous world-action model that uses dual Diffusion Transformers to decouple world prediction from action execution, achieving efficient long-horizon planning and real-time control. It achieves state-of-the-art performance on robotic manipulation tasks with up to 92.8% success on RoboTwin and 78.3% on real-world tasks, while reaching 24.17 Hz closed-loop control.

0 favorites 0 likes
#robot-manipulation

Revisiting Articulated Parts Perception in Robot Manipulation

Hugging Face Daily Papers · 2026-06-06 Cached

This paper introduces Geometric Primary Structure (GPS), a new representation for articulated parts perception in robot manipulation, enabling efficient VR-based annotation and achieving a 73% success rate without fine-tuning.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback