Tag
Future AGI is an open-source platform combining evaluations, tracing, simulations, guardrails, and optimization to help teams ship self-improving AI agents, with a nightly release available for early testing.
Miles Brundage highlights excellent work from the AI Security Institute investigating the impact of test-time compute budgets for frontier AI model evaluations, with praise from Noam Brown.
This tweet summarizes key takeaways from a detailed writeup on building complex cybersecurity evaluations for AI agents, covering long-horizon tasks, sourced benchmarks, varying difficulty levels, outcome verification challenges, and QA-based progress testing.
This article argues that AI agents lacking proper evaluations are not yet viable products, emphasizing the need for rigorous testing and benchmarking in AI development.
Summary of three key takeaways from conversations with leading AI researchers at CAISconf, covering the importance of evaluations for AI agents, the trade-offs between industry and academia, and a novel pedagogical RL approach.
This paper demonstrates that allowing attackers to strategically choose when to attack (attack selection) in agentic AI control evaluations significantly reduces measured safety, suggesting that current evaluations may overestimate safety against selective attackers.
Discusses the need for evolving AI evaluation benchmarks through difficulty, quality, and diversity refinement, citing examples like MMLU-Pro, MMLU-Redux, BIG-Bench Extra Hard, RealMath, MathArena, and DatBench.
A developer asks for recommendations for open-source alternatives to LangSmith for tracing, evaluations, and debugging agent workflows, citing restrictive paywalls.
Arize Phoenix announces a free 2-hour evaluations workshop from the AI Engineer: Europe conference, led by head of DevRel Laurie Voss, covering manual data examination and built-in/custom evals.