agent-evals

Tag

Cards List
#agent-evals

@Vtrivedy10: ok so it’s early but @mattpocockuk’s grill-me skill feels like great DX for iteratively building evals/environments wit…

X AI KOLs Timeline · 2026-07-14 Cached

A tweet thread discussing the iterative process of building evaluations and environments for AI agents, emphasizing human-agent collaboration and the importance of data and verifier design.

0 favorites 0 likes
#agent-evals

Production agent evals should test incident replay not just task success

Reddit r/AI_Agents · 2026-07-08

Discusses that production agent evaluations should include failure replay and resume capabilities, not just happy-path task success, emphasizing the need for observability that enables recovery.

0 favorites 0 likes
#agent-evals

@LangChain: How to run agent evals with @harborframework and LangSmith sandboxes, full traces included.

X AI KOLs Following · 2026-07-06 Cached

A guide on running agent evaluations using Harbor framework and LangSmith sandboxes with full trace support.

0 favorites 0 likes
#agent-evals

@BraceSproul: I've been thinking a lot about the two different groups of evals you need in general agents/agents which handle broad t…

X AI KOLs Following · 2026-05-19 Cached

A Twitter thread discussing two distinct evaluation suites needed for general AI agents: a lightweight benchmark eval for quick iteration and a comprehensive test coverage eval for thorough validation across diverse user paths.

0 favorites 0 likes
#agent-evals

@ArizePhoenix: Who judges the evaluators? When you use LLM-as-a-judge, you’re trusting a model to decide whether your agent, workflow,…

X AI KOLs Following · 2026-05-07

The article discusses the challenges of debugging and evaluating LLM judges using Arize Phoenix, which traces evaluator runs via OpenTelemetry to inspect decision logic, costs, and potential biases.

0 favorites 0 likes
← Back to home

Submit Feedback