tool-using

Tag

Cards List
#tool-using

Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain

arXiv cs.AI · 3d ago Cached

Telco-GAIA is a bilingual, multi-modal benchmark for evaluating tool-using agents in the telecom domain, comprising 100 human-verified tasks requiring multi-hop reasoning over heterogeneous sources, with objective scoring via exact string matching.

0 favorites 0 likes
#tool-using

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making

arXiv cs.AI · 3d ago Cached

Introduces DFAH-Bench, a replay benchmark to measure behavioral instability in financial agent decision-making, finding that outcome agreement alone misses significant trajectory divergence.

0 favorites 0 likes
#tool-using

Can an AI agent complete a task and still fail?

Reddit r/artificial · 2026-06-14

This paper introduces the concept of 'Verifier Tax' to categorize AI agent outcomes as safe success, unsafe success, or failure, and proposes a two-tier verification architecture for tool-using LLM agents.

0 favorites 0 likes
#tool-using

Decision-Aware Memory Cards: Counterfactual-Inspired Context Selection and Compression for Tool-Using LLM Agents

arXiv cs.AI · 2026-06-09 Cached

Introduces CICL, a decision-aware context layer that selects and compresses evidence for tool-using LLM agents by treating context as a decision-time intervention, using counterfactual-inspired scoring and typed memory cards under a token budget. Experiments on SWE-bench and RepoBench show concrete gains in retrieval accuracy and action criticality.

0 favorites 0 likes
#tool-using

Aurora: Unified Video Editing with a Tool-Using Agent

Hugging Face Daily Papers · 2026-05-18 Cached

Aurora is an agentic video editing framework that pairs a tool-augmented vision-language model agent with a diffusion transformer to automatically resolve textual and visual underspecification in user requests, enabling unified video editing tasks like replacement, removal, style transfer, and reference-driven insertion.

0 favorites 0 likes
← Back to home

Submit Feedback