@LangChain: In 13 minutes, @jeffbarg, Vyshu Khota, and Soroush Khadem walk through how Clay scaled agent evals agents at 300M+ runs…
Summary
Clay scaled agent evaluations to over 300 million runs per month, covering their four-quadrant eval framework and the challenges of closing the production-to-eval loop.
View Cached Full Text
Cached at: 08/30/26, 12:18 PM
In 13 minutes, @jeffbarg, Vyshu Khota, and Soroush Khadem walk through how Clay scaled agent evals agents at 300M+ runs a month.
Topics covered: Their four quadrant eval framework Why closing the production-to-eval loop is the hardest part How a data lake and long context changed what agents can do with data
Similar Articles
@LangChain: .@AdamRLucek on how we use traces to build evals for production agents.
Adam Łucek discusses how LangChain uses trace data to build evaluations for production agents.
@LangChain: What does it look like to automate eval and environment engineering for agents? Join Harrison Chase, Vivek Trivedy, Wil…
LangChain is hosting a webinar on automating evaluation and environment engineering for AI agents, featuring discussions on using repository context and production traces for building and testing evals.
@Vtrivedy10: my fave question, talked about this coding agent Eval+Improvement loop infra + UX in my AIE talk yesterday! biased but …
The speaker discusses the importance of evaluating and improving coding agents, highlighting LangSmith's integration with Harbor to provide a unified stack for running, tracing, and improving agent evaluations in isolated environments.
@LangChain: Improving agents The old way: Manually reading traces, looking for patterns, writing evals, and creating fixes. The bet…
This tweet contrasts the old manual approach to improving AI agents with a new automated method using LangSmith Engine, which cycles through tracing, eval, and fixes.
@LangChain: How to run agent evals with @harborframework and LangSmith sandboxes, full traces included.
A guide on running agent evaluations using Harbor framework and LangSmith sandboxes with full trace support.