things i wish i knew before evaluating AI agents in production
Summary
Personal lessons on evaluating AI agents in production, including mapping symptoms to layers, using trajectory evaluation, calibrating LLM judges, converting failures to test cases, and performing adversarial testing.
Similar Articles
Everyone says write evals for your agent. But what should you actually test?
A practical guide on writing effective evaluations for AI agents, focusing on starting from observed failures and using a mix of deterministic checks and LLM judges.
How to go about evaluation and Observability while building AI agents?
The author discusses challenges in evaluating and monitoring AI agents in production, including offline vs online evals, LLM-as-a-judge, tracing, and cost tracking, while citing tools like Langfuse and LangSmith but focusing on underlying processes.
How are people evaluating AI agents after they go into production?
The article discusses methods and challenges for evaluating AI agents in production environments, focusing on quality assurance for real-world conversations beyond pre-defined evaluation sets.
10 things I'd tell anyone starting to build AI agents in production
A practitioner shares ten critical lessons for deploying AI agents in production, emphasizing code-based constraints, context management, and security over relying solely on prompts.
AI Agents Testing before deploying to production
Discusses best practices for testing AI agents before deploying them to production environments.