imitation-learning

Tag

Cards List
#imitation-learning

Learning from Mixed-Quality Deployment Experience for Robot Manipulation

arXiv cs.LG ↗ · yesterday Cached

The paper proposes Predictive Action Chunk Learning (PACL) for robot manipulation, which learns from mixed-quality deployment experiences using a predictive critic and diffusion actor to improve policy performance.

0 favorites 0 likes
#imitation-learning

Minimal Recurrent Behavioral Memory for Imitation under Partial Observability

arXiv cs.LG ↗ · 3d ago Cached

This paper characterizes the minimal recurrent behavioral memory required for imitating an expert under partial observability, using information-theoretic measures and experimental validation.

0 favorites 0 likes
#imitation-learning

Imitation Learning for Autonomous Driving in CARLA

arXiv cs.AI ↗ · 2026-09-17 Cached

The paper investigates the closed-loop driving competence of a multimodal behavioral-cloning policy trained on offline expert demonstrations in the CARLA simulator, demonstrating effective autonomous driving without collisions and releasing all artifacts.

0 favorites 0 likes
#imitation-learning

JumpStart Your Policy Learning with Lessons from 160,000 Training Runs

arXiv cs.LG ↗ · 2026-09-15 Cached

This paper presents a large-scale empirical study of offline reinforcement and imitation learning, analyzing over 160,000 training runs to understand the effects of hyperparameters and dataset properties, and introduces JumpStart, a resource suite for reliable policy-learning research.

0 favorites 0 likes
#imitation-learning

Salesforce Finds Better Ways to Co-Evolve Agents and Their Harnesses (9 minute read)

TLDR AI ↗ · 2026-09-11 Cached

This research explores combining harness evolution with model adaptation for AI agents, discovering that direct imitation from experts harms weaker models' performance and proposing an on-policy correction method to improve performance without breaking harness fit for enterprise tasks.

0 favorites 0 likes
#imitation-learning

Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds

Hugging Face Daily Papers ↗ · 2026-09-01 Cached

This paper compares the temporal robustness of expert and imitation-learned policies in dexterous manipulation tasks, finding that imitation-learned policies degrade more sharply with increased execution speeds, primarily due to insertion misalignments.

0 favorites 0 likes
#imitation-learning

Predicting Consequences and Reinforcing Navigation Policies with Latent World Models

arXiv cs.AI ↗ · 2026-08-28 Cached

This paper proposes a Latent World Model (LWM) for robot navigation that predicts action-conditioned latent feature compatibility, enabling policy learning from unlabeled video data and reinforcement learning without additional environment interaction, outperforming existing methods.

0 favorites 0 likes
#imitation-learning

EXIMO: VLM Guided Exploration of VLA Policies

Hugging Face Daily Papers ↗ · 2026-08-20 Cached

EXIMO proposes an efficient algorithm for fine-tuning vision-language-action robot policies using a three-stage process: VLM-guided exploration, imitation learning, and residual reinforcement learning, showing improved sample-efficiency and performance.

0 favorites 0 likes
#imitation-learning

SAGE: SLO-Aware Adaptive Retrieval for Production RAG Systems

arXiv cs.LG ↗ · 2026-08-11 Cached

SAGE is a learned SLO-aware adaptive retrieval policy for production RAG systems that dynamically selects the number of retrieved passages per query, improving SLO compliance and reducing latency/cost with minimal quality loss.

0 favorites 0 likes
#imitation-learning

@svlevine: Action chunking is a mysteriously effective method. Modern large-scale imitation learning basically doesn't work withou…

X AI KOLs Following ↗ · 2026-08-07 Cached

Sergey Levine highlights a new paper investigating why action chunking is so effective in modern large-scale imitation learning for robotics, breaking down the underlying reasons.

0 favorites 0 likes
#imitation-learning

The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents

Hugging Face Daily Papers ↗ · 2026-08-06 Cached

Proposes Gated Hindsight Distillation (GHD), a method that uses future screenshots as privileged information to recover correct reasoning during training of mobile GUI agents, improving task success on AndroidWorld and AndroidLab across two vision-language models.

0 favorites 0 likes
#imitation-learning

@starlitmatcha: Today I read few papers about Action Chunking Transformers (ACT) and Diffusion Policy, which are the implementation of …

X AI KOLs Timeline ↗ · 2026-08-04 Cached

Khushi shares her reading notes on Action Chunking Transformers and Diffusion Policy, explaining how action chunking with generative models like VAEs and diffusion improves imitation learning for robotics, and how they solve inference latency with decoupled planning and execution.

0 favorites 0 likes
#imitation-learning

Mirror Learning

arXiv cs.LG ↗ · 2026-08-03 Cached

This paper proposes 'Mirror Learning', a framework for imitation learning from third-person observation that uses a fine-tuned video diffusion model for perspective transformation and an inverse dynamics model to synthesize pseudo first-person expert data, showing that this mirror data alone can train effective policies and improve behavior cloning.

0 favorites 0 likes
#imitation-learning

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

Hugging Face Daily Papers ↗ · 2026-07-28 Cached

HiFi-UMI introduces a portable data-production system for robot-free UMI data that achieves high trajectory accuracy using stereo-inertial SLAM and wide-angle cameras. Training manipulation policies on this data alone enables zero-shot deployment on real robots, matching or exceeding teleoperation baselines across several model families, and the authors open-source a 2,000-hour high-fidelity dataset.

0 favorites 0 likes
#imitation-learning

Why first person video may matter for robot learning[D]

Reddit r/MachineLearning ↗ · 2026-07-25

The article discusses the potential benefits and challenges of using first-person video for robot learning, highlighting that while direct imitation is limited, the sequence of visual attention may transfer. It references LingBot-VLA 2.0 and calls for controlled evaluations to separate viewpoint effects from data volume.

0 favorites 0 likes
#imitation-learning

I built a no-code tool that learns to play any 2D game by watching you play it

Reddit r/ArtificialInteligence ↗ · 2026-07-20

DeepEpoch is a no-code desktop app that learns to play 2D games by watching users play via behavior cloning, with human-in-the-loop fine-tuning.

0 favorites 0 likes
#imitation-learning

@aimalysheva: latent actions are having a moment, especially in robotics: instead of predicting a robot's actual joint commands or ga…

X AI KOLs Following ↗ · 2026-07-18 Cached

Latent actions are gaining traction in robotics as a way to learn from unlabeled video without action labels. Recent papers from DeepMind and FAIR demonstrate progress from controlled game environments to in-the-wild internet video, promising scalable training for imitation learning.

0 favorites 0 likes
#imitation-learning

Distributionally Robust and Safe Imitation Learning

arXiv cs.LG ↗ · 2026-07-16 Cached

This paper proposes a unified imitation learning framework using Taylor Series Imitation Learning and distributionally robust adaptive control to address both policy-induced and uncertainty-induced distribution shifts, with a UAV case study demonstrating safety under uncertainty.

0 favorites 0 likes
#imitation-learning

I watched a robot keep up with a live air-hockey puck at real speed, and it predicts the play instead of just reacting

Reddit r/ArtificialInteligence ↗ · 2026-07-10

The article discusses the shift from reactive to prediction-based robot control, highlighted by the LingBot-VA 2.0 model which can keep up with fast-moving objects like an air-hockey puck and learn from few demonstrations.

0 favorites 0 likes
#imitation-learning

Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning

arXiv cs.AI ↗ · 2026-07-10 Cached

This paper introduces Feedback Manipulation Regularization (FMR), an algorithm-agnostic method that uses evaluative feedback to improve alignment in imitation learning, achieving up to 98% reduction in misalignment in Safety Gymnasium environments.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback