Tag
A discussion comparing LLM evaluation and observability tools (LangSmith, Weave, Phoenix, Braintrust, Galileo, Opik) for fixing prompt failures and introducing an open-source platform that integrates the full eval-to-fix loop on a single trace.
Introduces how to use Phoenix's Pixie assistant to automatically classify error spans, batch annotate, and generate system prompt repair suggestions based on failure patterns.
Latitude is an open-source AI Agent Monitoring tool that provides issue detection, traces, and evals for LLM-based agents, similar to Sentry for AI.
LangChain's LangSmith enables developers to use tracing as compliance evidence for the EU AI Act, with customizable evaluators for bias, hallucination, toxicity, accuracy, and adversarial inputs.
Langfuse open-sources its LLM engineering platform to offer self-hosted tracing, analytics, and evaluation tools for production AI applications.