language-agents

Tag

Cards List
#language-agents

Heavy-Tailed Memory Traces in Long-Horizon Language Agents

arXiv cs.AI ↗ · yesterday Cached

The paper audits the shape of memory use in long-horizon language agents, finding heavy-tailed (core-tail) retrieval traces, and proposes Core-Tail World Model (CTWM), a rank-based memory controller that cuts prompt tokens by up to 24.48% on LongMemEval while preserving accuracy.

0 favorites 0 likes
#language-agents

NAQD Env: A benchmark for selective withdrawal in language agents

arXiv cs.AI ↗ · 2d ago Cached

NAQD-Env is a synthetic benchmark that evaluates how language agents selectively suspend actions when evidence, permissions, or stop instructions change, finding near-zero withdrawal recall across models despite fine-tuning gains in decision accuracy.

0 favorites 0 likes
#language-agents

RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents

arXiv cs.AI ↗ · 2026-09-17 Cached

RideWay introduces an efficiency-centered benchmark and the Efficiency Utility metric for evaluating tool-using language agents in ridehailing tasks, measuring success-gapped performance based on tool calls and user turns.

0 favorites 0 likes
#language-agents

What Should an Agent Forget? Separating What Is Stored from What Is Used

arXiv cs.AI ↗ · 2026-09-11 Cached

This paper introduces RD-Forget, a training-free framework for persistent language agents that separates stored memory from query-conditioned evidence to handle changing facts while preserving historical information.

0 favorites 0 likes
#language-agents

Procedural Memory Under Change: Reuse and Interference in Controlled Web Tasks

arXiv cs.AI ↗ · 2026-09-11 Cached

This study investigates how procedural memory mismatch in language agents affects web tasks, finding that under controlled conditions, mismatch does not lead to behavioral disruption or errors.

0 favorites 0 likes
#language-agents

ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance

arXiv cs.AI ↗ · 2026-09-11 Cached

ContractEval is a diagnostic framework that evaluates procedural conformance in LLM agents by making active obligations explicit and matching them against evidence, detecting failures missed by traditional judges.

0 favorites 0 likes
#language-agents

Outcome Monitors: Recovery Affordances for Silent Tool Failures

arXiv cs.AI ↗ · 2026-08-21 Cached

This paper introduces Outcome Monitors, a deterministic detector that identifies silent failures in tool calls for language agents and provides recovery receipts, significantly improving task completion rates in benchmarks like ToolMaze and τ-bench.

0 favorites 0 likes
#language-agents

AQuA: Recursively Self-Improving Quantitative Trading Research Agents

arXiv cs.CL ↗ · 2026-08-14 Cached

AQuA is a research system with two independent language-model-driven agents that recursively self-improve in quantitative trading research, achieving strong information coefficients on crypto and US equities while using sealed sandboxes to prevent data leakage.

0 favorites 0 likes
#language-agents

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction

arXiv cs.CL ↗ · 2026-08-13 Cached

This paper introduces DARC, a diagnosis-guided recovery harness that makes agent self-correction selective by profiling failure modes and pruning mismatched interventions before test-time correction, improving performance on ALFWorld, AppWorld, and XBRL Finance.

0 favorites 0 likes
#language-agents

EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents

arXiv cs.AI ↗ · 2026-08-13 Cached

Presents EvoGraph-Mem, a failure-aware editable graph memory framework for long-term language agents that tracks positive/negative evidence and activation states for insights, enabling memory maintenance through utility-aware retrieval and graph-level editing.

0 favorites 0 likes
#language-agents

Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards

arXiv cs.CL ↗ · 2026-08-11 Cached

Search-G1 proposes a representation-based intrinsic reward framework for search-augmented language agents, using intervention-calibrated readouts to balance retrieval necessity and evidence reliance, improving search efficiency without costly annotations.

0 favorites 0 likes
#language-agents

ContextWeave: A Real-World Workflow Benchmark

arXiv cs.AI ↗ · 2026-08-06 Cached

ContextWeave is a new longitudinal benchmark that evaluates whether recalled memory improves downstream agent performance in realistic office-work streams, using privacy-preserved multi-month workflows of 14 participants to create 1,005 executable tasks.

0 favorites 0 likes
#language-agents

Role Steering of Language Models for Social Simulations

arXiv cs.CL ↗ · 2026-08-04 Cached

Introduces an activation-steering screening workflow for role-conditioned LLM agents in social simulations, evaluated on OLMo-3-7B-Instruct across a 275-role inventory and showing role-specific directions outperform assistant-axis control.

0 favorites 0 likes
#language-agents

Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents

Hugging Face Daily Papers ↗ · 2026-08-03 Cached

This paper introduces SIEVE, a search-inspect-fetch strategy that uses Boolean Query Language to make deep-research agents retrieve only relevant document sections, achieving higher accuracy with 20.7–50.6% fewer tokens across multiple benchmark datasets and agent backbones.

0 favorites 0 likes
#language-agents

SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation

Hugging Face Daily Papers ↗ · 2026-08-03 Cached

Introduces SKT, a verified data synthesis pipeline for skill-use training of language model agents, producing 4,000 task packages and 27,164 verified trajectories from 2,000 public skills. Supervised fine-tuning on SKT-generated trajectories consistently improves skill-use performance across models and agent harnesses.

0 favorites 0 likes
#language-agents

Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance

arXiv cs.LG ↗ · 2026-07-31 Cached

This position paper argues that long-horizon benchmark failures must be compared against baselines built from matched short stages, introducing the 'horizon residual' metric to distinguish task size from task difficulty in LLM agent evaluation.

0 favorites 0 likes
#language-agents

From Agent Failures to Text Policies: What Works and What Breaks

arXiv cs.CL ↗ · 2026-07-24 Cached

This paper investigates why text-based optimization (TextGrad) fails for language agents, showing that while frozen agents can follow good policies, they cannot reliably learn and select policies from their own trajectories.

0 favorites 0 likes
#language-agents

RF-Agent: A Practical Framework for Building Language Agents for RFIC Design

arXiv cs.CL ↗ · 2026-07-22 Cached

RF-Agent introduces a textbook-driven knowledge distillation pipeline to create the first RF-domain reasoning dataset and benchmark, demonstrating that domain-specific fine-tuning and semantic retrieval significantly improve LLM reasoning for RF circuit design.

0 favorites 0 likes
#language-agents

New LLM Coordination Benchmark - Benchmarking Open-Ended Multi-Agent Coordination in Language Agents [R]

Reddit r/MachineLearning ↗ · 2026-07-14

Introduces a new benchmark for evaluating multi-agent coordination in LLMs, finding that most models struggle with long-horizon open-ended tasks, but Gemini 3.1 Pro performs comparably to trained MARL agents on the hardest setting.

0 favorites 0 likes
#language-agents

RIMRULE: Improving Tool-Using Language Agents via MDL-Guided Rule Learning

arXiv cs.CL ↗ · 2026-07-09 Cached

RimRule proposes a neuro-symbolic method that distills compact, interpretable rules from failure traces using the Minimum Description Length principle, improving LLM tool-use performance without modifying weights, and demonstrating rule portability across models.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback