Tag
This post announces Autoresearch Bench, an open-source benchmark for coding agents to autonomously tackle research problems, noting stark differences between models in autoresearch loops.
SWE-Interact has been updated to enable dynamic, user-driven evaluation sessions for software engineering tasks, with native support in Harbor for simulating multi-turn user-agent interactions.
Roboflow Playground is a new developer tool that enables side-by-side comparison of over 30 computer vision models across tasks like object detection and classification, streamlining the evaluation process.
AgentR 3.0 is a hiring evaluation tool designed to address challenges in the AI cheating era, launched on ProductHunt.
The author is seeking 5-10 beta testers for Behave, an AI agent testing/evaluation tool that goes beyond simple answer checking to catch issues like hallucination, premature conclusions, unsafe advice, and failures to self-correct.