interactive-benchmarks

Tag

Cards List
#interactive-benchmarks

@tomas_hk: Today we’re releasing our methodology for evaluating model routing with interactive benchmarks, which represent agent c…

X AI KOLs Timeline ↗ · 2026-09-01 Cached

Releasing a methodology for evaluating model routing with interactive benchmarks that achieve Pareto-dominance over leading benchmarks, offering higher quality at lower cost.

0 favorites 0 likes
#interactive-benchmarks

Uncertainty Decomposition for Clarification Seeking in LLM Agents

arXiv cs.AI ↗ · 2026-06-20 Cached

This paper proposes a prompt-based uncertainty decomposition method for LLM agents that separates action confidence from request uncertainty, enabling proactive clarification seeking in underspecified tasks. The method is evaluated on new clarification-augmented benchmarks across five LLM backbones, showing significant improvements.

0 favorites 0 likes
← Back to home

Submit Feedback