imitation-learning

Tag

Cards List
#imitation-learning

RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation

Hugging Face Daily Papers ↗ · 2026-07-07 Cached

Introduces digital teleoperation using action-conditioned world models to generate diverse training data for robotics, decoupling data collection from physical hardware. The system achieves real-time generation and enables zero-shot Sim2Real transfer.

0 favorites 0 likes
#imitation-learning

SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models

Hugging Face Daily Papers ↗ · 2026-07-07 Cached

SIEVE is a structure-aware data selection method for vision-language-action imitation learning that identifies reusable visuo-motor primitives and transition interfaces, outperforming full-data training with only 50% of demonstrations and training steps.

0 favorites 0 likes
#imitation-learning

Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation

arXiv cs.AI ↗ · 2026-07-03 Cached

Introduces Φ-Nav, a unified on-policy framework that uses hindsight reasoning to synthetically generate path-level instructions from exploratory trajectories, bridging the semantic supervision gap in Vision-Language Navigation and achieving competitive results on R2R-CE and RxR-CE benchmarks with fewer expert demonstrations.

0 favorites 0 likes
#imitation-learning

Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization

arXiv cs.LG ↗ · 2026-07-02 Cached

Active-GRPO introduces an adaptive imitation and self-improving reasoning framework that dynamically decides when to imitate references and when to reinforce the model's own discoveries for molecular optimization, achieving statistically significant improvements over previous methods on the TOMG-Bench-MolOpt benchmark.

0 favorites 0 likes
#imitation-learning

Behavior Cloning is Not All You Need: The Optimality of On-Policy Distillation for Noisy Expert Feedback

arXiv cs.LG ↗ · 2026-07-01 Cached

This paper proposes a noisy expert model to explain the gap between offline and online imitation learning, showing that offline learning from noisy trajectories requires exponential sample complexity while online on-policy distillation achieves polynomial dependence. The analysis leads to an alternative loss function and experiments confirm the theoretical findings.

0 favorites 0 likes
#imitation-learning

ReGuide: From Test-Time Guidance to Self-Improving Diffusion Policies

arXiv cs.LG ↗ · 2026-06-30 Cached

ReGuide introduces a self-improving framework for diffusion policies that uses test-time guidance to generate corrective rollouts, then fine-tunes the policy on this data, achieving 1.3–7.7× success improvement on Robomimic tasks.

0 favorites 0 likes
#imitation-learning

@oprydai: lerobot so-101 is interesting because it makes robot learning feel like an actual software pipeline instead of random l…

X AI KOLs Following ↗ · 2026-06-26 Cached

Lerobot SO-101 is a framework that transforms robot learning from messy lab magic into a clean, reproducible software pipeline, enabling structured data collection, policy training, and deployment, similar to what PyTorch did for deep learning.

0 favorites 0 likes
#imitation-learning

Do as the Romans Do: Learning Universal Behaviors from Heterogeneous Agents

arXiv cs.LG ↗ · 2026-06-18 Cached

This paper introduces GRID, a social learning method that extracts universal behaviors from heterogeneous agents by decomposing per-agent reward functions into general and specific rewards, enabling a generalist agent that avoids mode-averaging bias.

0 favorites 0 likes
#imitation-learning

Phase-Localized Curation Does Not Help: A Negative Result on Per-Phase Metric Selection for Demonstration Filtering

arXiv cs.LG ↗ · 2026-06-16 Cached

This paper investigates whether per-phase metric selection improves demonstration curation for behavior cloning policies. The authors find that phase-gated curation never outperforms global or uniform metric application, and the dilution of defect signals across phases explains the failure.

0 favorites 0 likes
#imitation-learning

Reinforcement Learning-Guided Retrieval with Soft Fusion for Robust Multimodal Imitation Learning under Missing Modalities

Hugging Face Daily Papers ↗ · 2026-06-13 Cached

RL4IL introduces a reinforcement learning-guided retrieval method that uses soft fusion over frozen demonstration libraries to handle missing sensor modalities in robotic imitation learning at inference time, achieving high success rates under complete camera dropout.

0 favorites 0 likes
#imitation-learning

@IlirAliu_: ETH Zurich just open-sourced their entire 2026 robot learning course. Not a MOOC. The actual course. Slides, lecture re…

X AI KOLs Timeline ↗ · 2026-06-10 Cached

ETH Zurich has open-sourced their entire 2026 robot learning course, including slides, lecture recordings, coding assignments, and a GitHub repository, covering topics from imitation learning to foundation models for robotics, with guest lectures from industry leaders.

0 favorites 0 likes
#imitation-learning

APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies

Hugging Face Daily Papers ↗ · 2026-06-10 Cached

Researchers propose APT, a two-stage training method that pretrains action experts on vision-action pairs before integrating language conditioning, significantly improving out-of-distribution instruction generalization for Vision-Language-Action policies.

0 favorites 0 likes
#imitation-learning

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning

Hugging Face Daily Papers ↗ · 2026-06-09 Cached

QGF is an RL algorithm that improves policies at test time by using a value gradient to guide a pre-trained flow policy, avoiding training-time instability while maintaining competitive performance.

0 favorites 0 likes
#imitation-learning

@RuohanZhang76: Excited to introduce StereoPolicy, led by @EvansXuHan. StereoPolicy is an effective way to add geometric cues to modern…

X AI KOLs Following ↗ · 2026-06-03 Cached

Introduces StereoPolicy, a framework that leverages synchronized stereo image pairs to improve geometric reasoning for robot manipulation policies, avoiding the fragility of RGB-D and point clouds. It integrates with diffusion-based and vision-language-action policies, showing consistent improvements in simulation and real-world tasks.

0 favorites 0 likes
#imitation-learning

Reinforcement Learning from Rich Feedback with Distributional DAgger

Hugging Face Daily Papers ↗ · 2026-06-03 Cached

Introduces DistIL, a method for reinforcement learning from rich feedback that guarantees monotonic policy improvement, outperforming existing methods on science reasoning, coding, and mathematical reasoning.

0 favorites 0 likes
#imitation-learning

Learning Bilevel Policies over Symbolic World Models for Long-Horizon Planning

arXiv cs.AI ↗ · 2026-05-18 Cached

Proposes BISON, a system combining learned low-level neural policies with high-level symbolic planning for long-horizon embodied tasks, showing strong generalization and efficiency.

0 favorites 0 likes
#imitation-learning

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation

Hugging Face Daily Papers ↗ · 2026-05-14 Cached

IntentVLA is a history-conditioned visual-language-action framework that improves robot imitation learning stability by encoding short-horizon intents from visual observations, addressing challenges from partial observability and ambiguous observations. It also introduces AliasBench, an ambiguity-aware benchmark for evaluating such methods.

0 favorites 0 likes
#imitation-learning

Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates

arXiv cs.LG ↗ · 2026-05-13 Cached

This paper introduces Trust Region Inverse Reinforcement Learning (TRIRL), a method that combines monotonic dual improvement with efficient local policy updates to outperform state-of-the-art imitation learning methods. It addresses the trade-off between stability and computational cost in IRL by using trust-region constraints.

0 favorites 0 likes
#imitation-learning

Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding

Hugging Face Daily Papers ↗ · 2026-05-12 Cached

This paper introduces LC-MAPF, a pre-trained model with a learnable communication module for multi-agent pathfinding that improves coordination and outperforms existing learning-based solvers while maintaining scalability.

0 favorites 0 likes
#imitation-learning

Learning to play Minecraft with Video PreTraining

OpenAI Blog ↗ · 2022-06-23 Cached

OpenAI introduced Video PreTraining (VPT), a semi-supervised method that trains neural networks to play Minecraft by learning from 70,000 hours of unlabeled human gameplay video combined with a small labeled dataset. The model learns complex sequential tasks using the native human interface (keyboard and mouse) and demonstrates capabilities like crafting diamond tools and pillar jumping, representing progress toward general computer-using agents.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback