Tag
A tweet introducing the eval-engineering skill from langchain-ai/langchain-skills, which uses human feedback to generate aligned environments, harnesses, and tasks for agent evaluation. It explains the workflow and provides installation instructions for the open-source tool.
A six-step guide to building evaluation gates that let AI agents merge changes autonomously, covering judge bias, runtime evals, trajectory grading, and more.
LangChain released an Eval Engineering Skill that automatically generates executable Harbor evaluations by mapping agent repositories and production traces, with an iterative user interview process to refine evals.
Jerry Liu observes that AI development is shifting from manually building workflows to specifying goals, and from prompt engineering to goal and eval engineering.