evaluations

Tag

Cards List
#evaluations

@tom_doerr: Future AGI is an open-source platform that combines evaluations, tracing, and guardrails to help teams ship self-improv…

X AI KOLs Timeline · 2026-08-02 Cached

Future AGI is an open-source platform combining evaluations, tracing, simulations, guardrails, and optimization to help teams ship self-improving AI agents, with a nightly release available for early testing.

0 favorites 0 likes
#evaluations

@Miles_Brundage: We could have this at home, American friends (a government agency that is staffed to do, and allowed to publish, stuff …

X AI KOLs Timeline · 2026-07-03 Cached

Miles Brundage highlights excellent work from the AI Security Institute investigating the impact of test-time compute budgets for frontier AI model evaluations, with praise from Noam Brown.

0 favorites 0 likes
#evaluations

@cwolferesearch: This cybersecurity eval writeup is great for understanding how complex / realistic evals are built. Some key takeaways …

X AI KOLs Timeline · 2026-06-27 Cached

This tweet summarizes key takeaways from a detailed writeup on building complex cybersecurity evaluations for AI agents, covering long-horizon tasks, sourced benchmarks, varying difficulty levels, outcome verification challenges, and QA-based progress testing.

0 favorites 0 likes
#evaluations

If your AI agent has no evals, you probably don’t have a product yet

Reddit r/AI_Agents · 2026-06-26

This article argues that AI agents lacking proper evaluations are not yet viable products, emphasizing the need for rigorous testing and benchmarking in AI development.

0 favorites 0 likes
#evaluations

@andykonwinski: 3 take-aways from chatting w/ top AI researchers last month: - Evals are the “source code” of AI agents (48:35) - BigAI…

X AI KOLs Following · 2026-06-25 Cached

Summary of three key takeaways from conversations with leading AI researchers at CAISconf, covering the importance of evaluations for AI agents, the trade-offs between industry and academia, and a novel pedagogical RL approach.

0 favorites 0 likes
#evaluations

Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety

arXiv cs.AI · 2026-06-08 Cached

This paper demonstrates that allowing attackers to strategically choose when to attack (attack selection) in agentic AI control evaluations significantly reduces measured safety, suggesting that current evaluations may overestimate safety against selective attackers.

0 favorites 0 likes
#evaluations

@cwolferesearch: Evaluations should not be static. We need to evolve evaluation sets / benchmarks over time so that they remain relevant…

X AI KOLs Following · 2026-05-29

Discusses the need for evolving AI evaluation benchmarks through difficulty, quality, and diversity refinement, citing examples like MMLU-Pro, MMLU-Redux, BIG-Bench Extra Hard, RealMath, MathArena, and DatBench.

0 favorites 0 likes
#evaluations

Are there any genuinely good open-source alternatives to LangSmith right now?

Reddit r/AI_Agents · 2026-05-15

A developer asks for recommendations for open-source alternatives to LangSmith for tracing, evaluations, and debugging agent workflows, citing restrictive paywalls.

0 favorites 0 likes
#evaluations

@ArizePhoenix: A comprehensive 2-hour evaluations workshop, for free! At AI Engineer: Europe, head of DevRel Laurie Voss gave this wor…

X AI KOLs Following · 2026-05-14 Cached

Arize Phoenix announces a free 2-hour evaluations workshop from the AI Engineer: Europe conference, led by head of DevRel Laurie Voss, covering manual data examination and built-in/custom evals.

0 favorites 0 likes
← Back to home

Submit Feedback