multi-dimensional-evaluation

Tag

Cards List
#multi-dimensional-evaluation

Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking

arXiv cs.AI ↗ · 2026-07-15 Cached

This paper introduces GenAI Evaluation, a governed and configuration-driven pipeline for scalable multi-dimensional evaluation of retail conversational agents. It achieves high accuracy using LLM-as-a-judge scoring with selective re-evaluation, validated against human-labeled data.

0 favorites 0 likes
#multi-dimensional-evaluation

Search Discipline for Long-Horizon Research Agents

arXiv cs.AI ↗ · 2026-06-11 Cached

This paper identifies a failure mode in long-horizon research agents where optimizing an aggregate metric can select candidates that improve the headline number but break critical subgroups (inversion). It proposes a search-discipline protocol with an external control loop that audits candidates based on disaggregated behavior rather than the score.

0 favorites 0 likes
← Back to home

Submit Feedback