Tag
Introduces Movement Trend Guidance to enhance 3D diffusion policies in robotic manipulation by providing foresight without explicit trajectories, achieving improved performance on benchmarks like RoboTwin2.0 and LIBERO-40.
A general-purpose agent directly controls a physical robot by interpreting visuals, writing executable programs, and revising actions based on physical feedback across diverse manipulation tasks, achieving high success rates without task-specific training.
MaP-WAM introduces a framework for non-Markovian robotic manipulation by separating memory-grounded planning from execution, using episodic memory to maintain fixed inference latency and achieving state-of-the-art results.
Researchers enable simple robotic grippers to achieve precise in-hand object manipulation through controlled sliding using tactile sensing and advanced friction models, mimicking human dexterity.
This paper introduces GE-Act 2.0, a world-action model pretrained from scratch to enable scalable zero-shot robotic manipulation with improved success rates across diverse tasks and conditions.
ZimaBlue introduces a scalable framework for learning generalizable world action models from large-scale egocentric video, substantially improving zero-shot robotic manipulation through a three-stage curriculum and slow-fast architecture.
PhysCaP is a physics-informed code-generation agent that actively explores objects to infer hidden physical properties for efficient robotic manipulation.
GOAG is a deep generative grasp planner that learns a gripper-specific contact surface distribution to sample valid grasps for unseen objects without object-specific training, achieving state-of-the-art results on the MultiDex dataset for dexterous robotic manipulation.
DreamX-Phi 1.0 is an action-conditioned video world model for robotic manipulation that predicts future observations using geometric attention encoding, depth estimation, and distillation, achieving top results in the WorldArena 2.0 Challenge.
TurboVLA introduces a new Vision-Language-Action paradigm that directly maps vision and language to action, achieving 97.7% success on LIBERO with only 0.2B parameters and real-time inference at 32 Hz on consumer GPUs, significantly reducing computational cost.
TableVerse introduces a fully automated Real2Sim pipeline that converts unstructured, in-the-wild images into high-fidelity, simulation-ready tabletop environments with accurate metrics and physical stability, along with a large-scale dataset (TableVerse-100K) for generalizable robotic manipulation.
LaMem-VLA proposes a latent-memory-native framework that integrates short-term and long-term historical experience directly into Vision-Language-Action reasoning, enabling better performance on long-horizon robotic manipulation tasks.
RynnWorld-4D is a generative world model that co-produces future RGB, depth, and optical flow from a single RGB-D image and language instruction using a unified diffusion process, enabling efficient robotic manipulation through inverse dynamics policy learning. It achieves state-of-the-art on real-world bimanual manipulation tasks.
PhysisForcing is a training framework that enhances embodied video generation for robotic manipulation by enforcing physical consistency through pixel-level trajectory alignment and semantic-level relational alignment losses in a DiT-based architecture, achieving notable improvements on benchmarks.
Foresight is a failure detection framework for long-horizon robotic manipulation that uses action-conditioned world model latents and functional conformal prediction to monitor trajectories, trained only with final task labels. It demonstrates state-of-the-art performance across simulation and real robot tasks.
EventVLA introduces a sparse visual evidence memory framework for long-horizon robotic manipulation, achieving an average success rate improvement of +40% over state-of-the-art memory-augmented VLAs.
Presents Qwen-RobotManip, a Vision-Language-Action foundation model for robotic manipulation that achieves generalization through unified alignment across representation, motion, and behavior dimensions, enabling large-scale training on diverse data sources. It outperforms prior state-of-the-art models across multiple out-of-distribution benchmarks and demonstrates emergent capabilities like zero-shot instruction following and cross-embodiment transfer.
PAIWorld enhances diffusion-transformer world models with geometric awareness and cross-view attention to improve multi-view 3D consistency for robotic manipulation tasks, achieving state-of-the-art results on benchmarks.
WEAVER is a multi-view world model for robotic manipulation that achieves high fidelity, consistency, and efficiency using flow-matching loss, demonstrating superior performance in policy evaluation, improvement, and test-time planning with significant real-world improvements.
AffordanceVLA introduces a unified framework using structured affordance forecasting as an intermediate representation to improve perception-action mapping in robotic manipulation, leveraging vision-language models and a Mixture-of-Transformer architecture.