@omarrayyann: Meet FetchMan: a vision-based humanoid policy trained entirely in simulation that transfers zero-shot to diverse real-w…
Summary
FetchMan is a vision-based humanoid policy trained entirely in simulation, capable of zero-shot transfer to diverse real-world scenes and objects for loco-manipulation tasks.
View Cached Full Text
Cached at: 08/23/26, 07:43 AM
Meet FetchMan: a vision-based humanoid policy trained entirely in simulation that transfers zero-shot to diverse real-world scenes and objects.
Simulation has produced impressive locomotion policies that transfer to the real world. We wanted to see how far the same recipe goes for vision-based loco-manipulation. More below
Similar Articles
Niantic Spatial, Flexion, and NVIDIA: Closing the Sim2Real Gap for Humanoids
Niantic Spatial, Flexion, and NVIDIA demonstrate a sim2real pipeline for humanoid robots using digital twins and RL, achieving zero-shot transfer from simulation to real office navigation.
SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation
SimFoundry is a modular system that automates real-to-sim scene construction from video, generating digital twins and affordance-preserving variations for zero-shot robot policy training, achieving strong transfer to real-world tasks and high simulation-to-real performance prediction.
Retrieve, Don't Retrain: Extending Vision Language Action Models to New Tasks at Test Time
This paper introduces a retrieval-augmented vision-language-action policy that eliminates per-task fine-tuning by using pre-trained models with indexed demonstrations, enabling efficient cross-embodiment generalization and task adaptation at test time.
OASIS: From Simulation Data Collection to Real-World Humanoid Loco-Manipulation
OASIS is a simulation-data-driven framework for humanoid loco-manipulation that uses 3D generative models and hierarchical visuomotor policies. It achieves better zero-shot performance than real-robot training by leveraging domain randomization in simulation.
DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation
DeVI introduces a framework that turns text-conditioned synthetic videos into physically plausible dexterous robot control via a hybrid 3D-2D tracking reward, enabling zero-shot generalization to unseen objects.