behavior-cloning

Tag

Cards List
#behavior-cloning

Mirror Learning

arXiv cs.LG · 2026-08-03 Cached

This paper proposes 'Mirror Learning', a framework for imitation learning from third-person observation that uses a fine-tuned video diffusion model for perspective transformation and an inverse dynamics model to synthesize pseudo first-person expert data, showing that this mirror data alone can train effective policies and improve behavior cloning.

0 favorites 0 likes
#behavior-cloning

I built a no-code tool that learns to play any 2D game by watching you play it

Reddit r/ArtificialInteligence · 2026-07-20

DeepEpoch is a no-code desktop app that learns to play 2D games by watching users play via behavior cloning, with human-in-the-loop fine-tuning.

0 favorites 0 likes
#behavior-cloning

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment

Hugging Face Daily Papers · 2026-07-15 Cached

Anchor-Align augments behavioral cloning with vision-language anchoring to preserve pretrained representations and language-action alignment, improving real-robot success rates by over 20% on xArm7 and showing consistent gains in simulation benchmarks.

0 favorites 0 likes
#behavior-cloning

Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

Hugging Face Daily Papers · 2026-07-02 Cached

Task-Agnostic Pretraining (TAP) decomposes VLA training into self-supervised motor skill learning from unlabeled interaction data, then lightweight language grounding, achieving strong performance with minimal expert demonstrations. It matches or outperforms models trained on millions of expert trajectories while being robust to real-world perturbations.

0 favorites 0 likes
#behavior-cloning

Skill-Guided Continuation Distillation for GUI Agents

arXiv cs.AI · 2026-06-18 Cached

The paper proposes Skill-Guided Continuation Distillation (SGCD), an iterative self-improvement framework that uses skill-guided policies to generate supervision for off-trajectory states during closed-loop execution, improving GUI agent success rates on OSWorld-Verified from around 30% to over 50%.

0 favorites 0 likes
#behavior-cloning

Phase-Localized Curation Does Not Help: A Negative Result on Per-Phase Metric Selection for Demonstration Filtering

arXiv cs.LG · 2026-06-16 Cached

This paper investigates whether per-phase metric selection improves demonstration curation for behavior cloning policies. The authors find that phase-gated curation never outperforms global or uniform metric application, and the dilution of defect signals across phases explains the failure.

0 favorites 0 likes
#behavior-cloning

From Noise to Control: Parameterized Diffusion Policies

arXiv cs.AI · 2026-06-02 Cached

This paper introduces Parameterized Diffusion Policy (PDP), a framework that makes diffusion policies controllable by conditioning on low-dimensional latent parameters, enabling smooth behavior interpolation and adaptation without retraining. It demonstrates improved performance on complex multimodal robot tasks in simulation and real-world experiments.

0 favorites 0 likes
#behavior-cloning

PRO-CUA: Process-Reward Optimization for Computer Use Agents

arXiv cs.AI · 2026-05-29 Cached

This paper introduces PRO-CUA, a process-reward optimization framework for training Computer Use Agents (CUAs) using iterative step-level reinforcement learning. The method decouples on-policy environment interaction from policy optimization, enabling dense credit assignment without relying on expert trajectories, and demonstrates effectiveness on live web benchmarks.

0 favorites 0 likes
#behavior-cloning

Frequency-Guided Action Diffusion via Sub-Frequency Manifold Traversal

Hugging Face Daily Papers · 2026-05-27 Cached

This paper introduces the Frequency Guidance Operator (FGO), a method for diffusion policies that smooths action generation by steering noisy samples through intermediate sub-frequency manifolds, improving performance on robotic manipulation tasks.

0 favorites 0 likes
#behavior-cloning

@rohanpaul_ai: WOW, A leaked audio from Meta’s April 30 all-hands. Meta is reportedly using its own engineers’ work traces to train co…

X AI KOLs Following · 2026-05-21 Cached

Leaked audio reveals Meta is using engineers' work traces to train coding AI through behavior cloning, while cutting thousands of jobs.

0 favorites 0 likes
← Back to home

Submit Feedback