interactive-evaluation

Tag

Cards List
#interactive-evaluation

Interactive Evaluation Requires a Design Science

Hugging Face Daily Papers · 2026-05-18 Cached

This position paper argues that interactive AI evaluation should be treated as a design science paradigm, proposing a two-axis taxonomy and reporting standards for assessing dynamic system behavior through trajectories.

0 favorites 0 likes
#interactive-evaluation

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents

arXiv cs.AI · 2026-05-15 Cached

ClawForge is a generator-backed benchmark framework for executable command-line workflows under state conflict, evaluating LLM agents on tasks with pre-existing partial, stale, or conflicting artifacts across 17 scenarios.

0 favorites 0 likes
← Back to home

Submit Feedback