Tag
Telco-GAIA is a bilingual, multi-modal benchmark for evaluating tool-using agents in the telecom domain, comprising 100 human-verified tasks requiring multi-hop reasoning over heterogeneous sources, with objective scoring via exact string matching.
Introduces DFAH-Bench, a replay benchmark to measure behavioral instability in financial agent decision-making, finding that outcome agreement alone misses significant trajectory divergence.
This paper introduces the concept of 'Verifier Tax' to categorize AI agent outcomes as safe success, unsafe success, or failure, and proposes a two-tier verification architecture for tool-using LLM agents.
Introduces CICL, a decision-aware context layer that selects and compresses evidence for tool-using LLM agents by treating context as a decision-time intervention, using counterfactual-inspired scoring and typed memory cards under a token budget. Experiments on SWE-bench and RepoBench show concrete gains in retrieval accuracy and action criticality.
Aurora is an agentic video editing framework that pairs a tool-augmented vision-language model agent with a diffusion transformer to automatically resolve textual and visual underspecification in user requests, enabling unified video editing tasks like replacement, removal, style transfer, and reference-driven insertion.