agent-as-a-judge

Tag

Cards List
#agent-as-a-judge

Online Agent-as-a-Judge: Situation-Generating Evaluation for Interactive Agents

arXiv cs.AI · 2026-06-09 Cached

Proposes Online Agent-as-a-Judge, an evaluation framework that uses an in-world evaluator agent to actively generate situations for testing interactive social agents, improving coverage and reliability over passive methods.

0 favorites 0 likes
#agent-as-a-judge

WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models

Hugging Face Daily Papers · 2026-04-20 Cached

WebCompass is a multimodal benchmark for evaluating LLMs on web coding tasks across three input modalities (text, image, video) and three task types (generation, editing, repair). It introduces an Agent-as-a-Judge paradigm that autonomously executes generated websites in a real browser to assess visual fidelity and interactivity.

0 favorites 0 likes
← Back to home

Submit Feedback