language-agents

Tag

Cards List
#language-agents

Memory Depth, Not Memory Access: Selective Parametric Consolidation for Long-Running Language Agents

arXiv cs.AI · 2026-06-26 Cached

This paper introduces the concept of memory depth for long-running language agents, distinguishing it from retrieval-based memory access, and proposes EVAF, a selective parametric consolidation mechanism using surprise- and valence-gated LoRA updates. Experiments across multiple models show EVAF improves goal persistence after context unload with minimal parametric writes.

0 favorites 0 likes
#language-agents

OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

Hugging Face Daily Papers · 2026-06-25 Cached

OPID proposes an on-policy skill distillation framework that extracts dense hindsight supervision from completed trajectories, combining outcome-based RL with token-level self-distillation to improve language agent training efficiency and performance on multi-turn tasks.

0 favorites 0 likes
#language-agents

Learning What to Remember: Observability-Safe Memory Retention via Constrained Optimization for Long-Horizon Language Agents

arXiv cs.AI · 2026-06-10 Cached

This paper formulates memory retention for long-horizon language agents as a constrained stochastic optimization problem, introducing OSL-MR, a framework that enforces observability-safe learning with a Mixed-Score heuristic. Experiments show consistent improvements over existing heuristic baselines under tight memory budgets.

0 favorites 0 likes
#language-agents

What Spatial Memory Must Store: Occlusion as the Test for Language-Agent Memory

arXiv cs.AI · 2026-06-10 Cached

This paper investigates whether spatial geometry improves language-agent memory recall, demonstrating that geometry must lead recall over recency and importance, and that a ray-tracing visibility predicate is crucial for occlusion handling in 3D voxel worlds.

0 favorites 0 likes
#language-agents

Policy-Conditioned Counterfactual Credit for Verifiable Reinforcement Learning of Long-Horizon Language Agents

arXiv cs.LG · 2026-06-05 Cached

Proposes CVT-RL, a constrained policy-gradient algorithm with policy-conditioned counterfactual contribution estimation and verifiable rewards, improving long-horizon language agent reliability and reducing reward hacking.

0 favorites 0 likes
#language-agents

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?

Hugging Face Daily Papers · 2026-06-04 Cached

This paper introduces ArcANE, an automatically constructed benchmark for evaluating role-playing language agents' alignment with character psychological trajectories across narrative phases, showing that conditioning on character arc information improves performance, especially in scenarios beyond the source text.

0 favorites 0 likes
#language-agents

AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents

Hugging Face Daily Papers · 2026-06-02 Cached

A comprehensive evaluation framework for continual learning in language agents is introduced, emphasizing controlled task streams and memory design analysis to better assess reusable experience and learning stability.

0 favorites 0 likes
#language-agents

Policy and World Modeling Co-Training for Language Agents

Hugging Face Daily Papers · 2026-06-01 Cached

This paper introduces PaW, a co-training framework that adds auxiliary world modeling supervision to policy learning during on-policy RL rollouts, improving language agent training without additional computational overhead.

0 favorites 0 likes
#language-agents

LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents

Hugging Face Daily Papers · 2026-05-28 Cached

LiteCoder-Terminal-Gen introduces a zero-dependency synthetic pipeline that generates executable terminal training environments, producing SFT and RL datasets that enable language agents to achieve significant performance gains on Terminal Bench benchmarks.

0 favorites 0 likes
#language-agents

RICE-PO: Turning Retrieval Interactions into Credit Signals for Reasoning Agents

arXiv cs.CL · 2026-05-27 Cached

RICE-PO is a critic-free policy optimization framework that turns retrieval interactions into localized credit signals for training reasoning agents, outperforming prompt-based and group-based RL baselines on BRIGHT and BEIR benchmarks.

0 favorites 0 likes
#language-agents

From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills

Hugging Face Daily Papers · 2026-05-22 Cached

This paper systematically evaluates model-generated skills for language agents across the full lifecycle of experience generation, extraction, and consumption, finding that skills are beneficial on average but exhibit non-trivial negative transfer, leading to a meta-skill that improves skill quality.

0 favorites 0 likes
#language-agents

Auto-Dreamer: Learning Offline Memory Consolidation for Language Agents

arXiv cs.CL · 2026-05-21 Cached

Auto-Dreamer introduces a learned offline memory consolidation method for language agents, decoupling fast memory acquisition from slow cross-session consolidation, and achieving higher performance with smaller memory banks, generalizing to unseen environments.

0 favorites 0 likes
#language-agents

@tom_doerr: AI agents for data analysis, plugins, and web browsing https://github.com/xlang-ai/OpenAgents…

X AI KOLs Timeline · 2026-05-15 Cached

OpenAgents is an open platform for using and hosting language agents in everyday life, featuring agents for data analysis, plugins, and web browsing, with open code and a demo.

0 favorites 0 likes
#language-agents

@tom_doerr: Builds custom AI agents with reinforcement learning https://github.com/agentica-project/rllm…

X AI KOLs Timeline · 2026-05-14 Cached

rLLM is an open-source framework for post-training language agents via reinforcement learning, with notable model releases like DeepSWE-Preview and DeepCoder-14B-Preview achieving state-of-the-art results.

0 favorites 0 likes
#language-agents

Milestone-Guided Policy Learning for Long-Horizon Language Agents

arXiv cs.CL · 2026-05-08 Cached

This paper introduces BEACON, a milestone-guided policy learning framework designed to improve credit assignment and sample efficiency for long-horizon language agents. It demonstrates significant performance improvements over GRPO and GiGPO on benchmarks like ALFWorld, WebShop, and ScienceWorld.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback