policy-learning

Tag

Cards List
#policy-learning

The Sample Complexity of Policy Learning with Mu-Resets

arXiv cs.LG · 7h ago Cached

This paper studies the sample complexity of policy learning under the mu-resets interaction protocol in reinforcement learning, resolving a question about the role of policy realizability and showing horizon dependence is exponential under all-policy concentrability and sqrt-exponential under pushforward concentrability.

0 favorites 0 likes
#policy-learning

EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents

arXiv cs.LG · 4d ago Cached

Introduces EvoHarness-RL, a framework that learns runtime harness policies for long-horizon LLM agents, enabling them to construct and update external state (belief, progress, experience) during task execution. Using Qwen3-8B on ALFWorld, it achieves 96.9% success and reveals harness annealing and evolution dynamics.

0 favorites 0 likes
#policy-learning

Internalizing Academic Writing Workflows for Introduction Generation via Struct-Aware Policy Learning

arXiv cs.CL · 6d ago Cached

This paper proposes StructPO, a structure-aware policy learning framework that internalizes multi-stage academic writing workflows into a single-pass LLM policy using explicit stage tokens and refinement-guided optimization, improving introduction generation quality and efficiency.

0 favorites 0 likes
#policy-learning

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space

Hugging Face Daily Papers · 2026-08-02 Cached

This paper proposes SG-WAM, a self-guided framework for learning geometry-aware action-conditioned world models directly in policy-derived representation space. It achieves state-of-the-art success rates on LIBERO and LIBERO-Plus benchmarks, outperforming strong baselines in real-world evaluations.

0 favorites 0 likes
#policy-learning

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

Hugging Face Daily Papers · 2026-07-28 Cached

HiFi-UMI introduces a portable data-production system for robot-free UMI data that achieves high trajectory accuracy using stereo-inertial SLAM and wide-angle cameras. Training manipulation policies on this data alone enables zero-shot deployment on real robots, matching or exceeding teleoperation baselines across several model families, and the authors open-source a 2,000-hour high-fidelity dataset.

0 favorites 0 likes
#policy-learning

From Agent Failures to Text Policies: What Works and What Breaks

arXiv cs.CL · 2026-07-24 Cached

This paper investigates why text-based optimization (TextGrad) fails for language agents, showing that while frozen agents can follow good policies, they cannot reliably learn and select policies from their own trajectories.

0 favorites 0 likes
#policy-learning

Graph-Constrained Policy Learning for Extreme Clinical Code Prediction

arXiv cs.LG · 2026-07-15 Cached

Proposes a graph-constrained traversal policy that reformulates ICD-10-CM code prediction as a finite-horizon decision process over a pruned code hierarchy, outperforming flat baselines on MIMIC-IV discharge summaries.

0 favorites 0 likes
#policy-learning

BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding

Hugging Face Daily Papers · 2026-06-30 Cached

BlockPilot proposes an instance-adaptive policy that predicts the optimal block size for diffusion-based speculative decoding, achieving significant speedup with minimal overhead.

0 favorites 0 likes
#policy-learning

SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation

Hugging Face Daily Papers · 2026-06-26 Cached

SimFoundry is a modular system that automates real-to-sim scene construction from video, generating digital twins and affordance-preserving variations for zero-shot robot policy training, achieving strong transfer to real-world tasks and high simulation-to-real performance prediction.

0 favorites 0 likes
#policy-learning

World Value Models for Robotic Manipulation

Hugging Face Daily Papers · 2026-06-23 Cached

The paper presents World Value Model (WVM), a generalist robotic value model that combines world models with value estimation to accurately assess task progression and improve robotic policy learning from mixed-quality data, achieving state-of-the-art results on standard benchmarks and a new suboptimal data benchmark.

0 favorites 0 likes
#policy-learning

@omarsar0: // Automating SKILL.md Generation // Increasingly, mining sessions is one of the best ways to improve your agents. Open…

X AI KOLs Following · 2026-06-19 Cached

This paper from MIT and Harvard explores automating SKILL.md generation by mining GUI interaction trajectories, finding that clusters are readable but do not improve policy performance across domains.

0 favorites 0 likes
#policy-learning

PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning

Hugging Face Daily Papers · 2026-06-19 Cached

PoLAR introduces a geometrically structured latent action representation in hyperbolic space that separates transition extent from mode, improving robotic policy learning performance.

0 favorites 0 likes
#policy-learning

@HuggingPapers: Geometric Action Model for Robot Policy Learning Repurposes a geometric foundation model as one backbone for perception…

X AI KOLs Following · 2026-06-16 Cached

Geometric Action Model repurposes a geometric foundation model for robot policy learning, achieving 85.5% on LIBERO-Plus with 6.9 ms inference, 55× faster than baselines.

0 favorites 0 likes
#policy-learning

Geometric Action Model for Robot Policy Learning

Hugging Face Daily Papers · 2026-06-15 Cached

The Geometric Action Model (GAM) repurposes a pretrained geometric foundation model (GFM) as a unified backbone for language-conditioned robot manipulation, achieving higher accuracy, robustness, and efficiency than existing foundation-model-scale baselines across simulation and real-world benchmarks.

0 favorites 0 likes
#policy-learning

Light-WAM: Efficient World Action Models with State-Fusion Action Decoding

Hugging Face Daily Papers · 2026-06-06 Cached

Light-WAM is a lightweight world action model for efficient robot manipulation that uses a compact video backbone and downsampled latent space for future-video supervision, achieving high performance with low inference latency.

0 favorites 0 likes
#policy-learning

DiffAero: A GPU-Accelerated Differentiable Simulation Framework for Efficient Quadrotor Policy Learning

arXiv cs.AI · 2026-06-04 Cached

DiffAero is a GPU-accelerated, fully differentiable simulation framework for quadrotor control policy learning that supports environment- and agent-level parallelism, multiple dynamics models, and customizable sensors. It enables robust flight policy learning in hours on consumer-grade hardware and is released as open-source.

0 favorites 0 likes
#policy-learning

Capability Self-Assessment: Teaching LLMs to Know Their Limits

arXiv cs.AI · 2026-06-02 Cached

This paper introduces Capability Self-Assessment (CSA) for LLMs, formulating it as a policy-learning problem. Experiments show that reinforcement learning effectively teaches models to recognize their own limits and delegate queries they cannot solve, outperforming supervised fine-tuning and generalizing well out-of-distribution.

0 favorites 0 likes
#policy-learning

τ_0-WM: A Unified Video-Action World Model for Robotic Manipulation

Hugging Face Daily Papers · 2026-05-31 Cached

τ_0-WM is a unified video-action world model for robotic manipulation that integrates policy learning, video prediction, and action evaluation using a shared video diffusion backbone. It shows superior performance on challenging long-horizon and fine-grained tasks.

0 favorites 0 likes
#policy-learning

Actionable World Representation

Hugging Face Daily Papers · 2026-05-18 Cached

WorldString is a neural architecture that models object state manifolds from point clouds or RGB-D video streams, serving as a foundational component for physical world models with differentiable structure for policy learning integration.

0 favorites 0 likes
#policy-learning

Learning and Reusing Policy Decompositions for Hierarchical Generalized Planning with LLM Agents

arXiv cs.AI · 2026-05-11 Cached

This paper introduces HCL-GP, a dynamic policy-learning framework that integrates generalized planning and hierarchical task decomposition to enable LLM-based agents to learn and reuse executable policy components, significantly improving performance on the AppWorld benchmark.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback