trajectory-optimization

Tag

Cards List
#trajectory-optimization

Beyond Success and Failure: Length-Aware Contrastive Learning for GUI Agents

arXiv cs.AI · 6h ago Cached

The paper proposes Length-Aware Contrastive Learning for GUI Agents (LACL-GUI), a contrastive reinforcement learning framework that incorporates trajectory-level quality signals to improve agent performance by addressing reward-gradient misalignment.

0 favorites 0 likes
#trajectory-optimization

@ZiyunClaudeWang: What if every meter of robot motion were optimized for reconstruction? Excited to share TRACE, our new work on active 3…

X AI KOLs Timeline · 2026-08-07 Cached

TRACE introduces a novel approach to active 3D reconstruction by optimizing full sensor trajectories for ergodic coverage of scene information, outperforming next-best-view baselines with a 1.5 dB PSNR improvement.

0 favorites 0 likes
#trajectory-optimization

@HaiyuWu1: Locally straightening the latent trajectory can largely reduce the optimization difficulty of world model planning! Sou…

X AI KOLs Following · 2026-07-14 Cached

The paper proposes locally straightening latent trajectories by maximizing cosine similarity between adjacent velocity vectors to reduce optimization difficulty in world model planning, with experiments on PushT dataset achieving 36% success rate for long-horizon planning.

0 favorites 0 likes
#trajectory-optimization

ExTra: Exploratory Trajectory Optimization for Language Model Reinforcement Learning

arXiv cs.LG · 2026-06-25 Cached

ExTra introduces exploratory trajectory optimization for language model reinforcement learning, combining novelty rewards and entropy-guided prefix regeneration to improve both single-sample accuracy and inference-time coverage on mathematical reasoning benchmarks.

0 favorites 0 likes
#trajectory-optimization

CKM-Driven Communication-Aware UAV Intelligent Trajectory Optimization for Urban Inspection

arXiv cs.LG · 2026-06-25 Cached

This paper proposes a CKM-driven framework for multi-UAV trajectory planning in urban inspection, using diffusion models to reconstruct high-fidelity channel quality maps and a graph attention network with soft actor-critic algorithm for communication-aware path planning.

0 favorites 0 likes
#trajectory-optimization

Read the Trace, Steer the Path: Trajectory-Aware Reinforcement Learning for Diffusion Language Models

arXiv cs.CL · 2026-06-04 Cached

This paper introduces CAPR (Cached-Amortized Path Refinement), a reinforcement learning algorithm for diffusion large language models that extracts tree-like supervision signals from the denoising trace without the compute cost of full tree rollouts. CAPR achieves state-of-the-art performance on reasoning benchmarks like GSM8K, Math500, Sudoku, and Countdown at roughly 0.75x the cost of flat rollouts.

0 favorites 0 likes
#trajectory-optimization

On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment

Hugging Face Daily Papers · 2026-05-12 Cached

This paper introduces FATE, an on-policy framework that leverages failure trajectories to enhance the safety and performance of tool-using LLM agents through self-evolution and Pareto-aware optimization.

0 favorites 0 likes
#trajectory-optimization

Plan online, learn offline: Efficient learning and exploration via model-based control

OpenAI Blog · 2018-11-05 Cached

OpenAI proposes POLO (Plan Online, Learn Offline), a framework combining model-based control with value function learning and coordinated exploration to enable efficient learning on complex control tasks like humanoid locomotion and dexterous manipulation with minimal real-world experience.

0 favorites 0 likes
#trajectory-optimization

Prediction and control with temporal segment models

OpenAI Blog · 2017-03-12 Cached

OpenAI introduces a method for learning complex nonlinear system dynamics using deep generative models over temporal segments, enabling stable long-horizon predictions and differentiable trajectory optimization for model-based control.

0 favorites 0 likes
← Back to home

Submit Feedback