Tag
The paper audits the shape of memory use in long-horizon language agents, finding heavy-tailed (core-tail) retrieval traces, and proposes Core-Tail World Model (CTWM), a rank-based memory controller that cuts prompt tokens by up to 24.48% on LongMemEval while preserving accuracy.
NAQD-Env is a synthetic benchmark that evaluates how language agents selectively suspend actions when evidence, permissions, or stop instructions change, finding near-zero withdrawal recall across models despite fine-tuning gains in decision accuracy.
RideWay introduces an efficiency-centered benchmark and the Efficiency Utility metric for evaluating tool-using language agents in ridehailing tasks, measuring success-gapped performance based on tool calls and user turns.
This paper introduces RD-Forget, a training-free framework for persistent language agents that separates stored memory from query-conditioned evidence to handle changing facts while preserving historical information.
This study investigates how procedural memory mismatch in language agents affects web tasks, finding that under controlled conditions, mismatch does not lead to behavioral disruption or errors.
ContractEval is a diagnostic framework that evaluates procedural conformance in LLM agents by making active obligations explicit and matching them against evidence, detecting failures missed by traditional judges.
This paper introduces Outcome Monitors, a deterministic detector that identifies silent failures in tool calls for language agents and provides recovery receipts, significantly improving task completion rates in benchmarks like ToolMaze and τ-bench.
AQuA is a research system with two independent language-model-driven agents that recursively self-improve in quantitative trading research, achieving strong information coefficients on crypto and US equities while using sealed sandboxes to prevent data leakage.
This paper introduces DARC, a diagnosis-guided recovery harness that makes agent self-correction selective by profiling failure modes and pruning mismatched interventions before test-time correction, improving performance on ALFWorld, AppWorld, and XBRL Finance.
Presents EvoGraph-Mem, a failure-aware editable graph memory framework for long-term language agents that tracks positive/negative evidence and activation states for insights, enabling memory maintenance through utility-aware retrieval and graph-level editing.
Search-G1 proposes a representation-based intrinsic reward framework for search-augmented language agents, using intervention-calibrated readouts to balance retrieval necessity and evidence reliance, improving search efficiency without costly annotations.
ContextWeave is a new longitudinal benchmark that evaluates whether recalled memory improves downstream agent performance in realistic office-work streams, using privacy-preserved multi-month workflows of 14 participants to create 1,005 executable tasks.
Introduces an activation-steering screening workflow for role-conditioned LLM agents in social simulations, evaluated on OLMo-3-7B-Instruct across a 275-role inventory and showing role-specific directions outperform assistant-axis control.
This paper introduces SIEVE, a search-inspect-fetch strategy that uses Boolean Query Language to make deep-research agents retrieve only relevant document sections, achieving higher accuracy with 20.7–50.6% fewer tokens across multiple benchmark datasets and agent backbones.
Introduces SKT, a verified data synthesis pipeline for skill-use training of language model agents, producing 4,000 task packages and 27,164 verified trajectories from 2,000 public skills. Supervised fine-tuning on SKT-generated trajectories consistently improves skill-use performance across models and agent harnesses.
This position paper argues that long-horizon benchmark failures must be compared against baselines built from matched short stages, introducing the 'horizon residual' metric to distinguish task size from task difficulty in LLM agent evaluation.
This paper investigates why text-based optimization (TextGrad) fails for language agents, showing that while frozen agents can follow good policies, they cannot reliably learn and select policies from their own trajectories.
RF-Agent introduces a textbook-driven knowledge distillation pipeline to create the first RF-domain reasoning dataset and benchmark, demonstrating that domain-specific fine-tuning and semantic retrieval significantly improve LLM reasoning for RF circuit design.
Introduces a new benchmark for evaluating multi-agent coordination in LLMs, finding that most models struggle with long-horizon open-ended tasks, but Gemini 3.1 Pro performs comparably to trained MARL agents on the hardest setting.
RimRule proposes a neuro-symbolic method that distills compact, interpretable rules from failure traces using the Minimum Description Length principle, improving LLM tool-use performance without modifying weights, and demonstrating rule portability across models.