Tag
LangChain introduces an Engine that learns how your agent works by reading traces and repositories, then tests for agent-specific weaknesses such as hallucinations and system prompt violations, returning a list of verified issues to fix before they impact users.
The author discusses common challenges in debugging AI agent evaluation failures, seeking insights on efficient investigation methods and the reliability of comparison runs versus other evidence sources.
The article discusses common errors in using AI agent traces and offers advice on proper implementation.
LangChain introduces Trajectories in LangSmith, a chronological view of agent sessions that simplifies debugging by aggregating messages from humans, AI, and tools in order for easier navigation and analysis.
The article explores common challenges in running multiple AI agents, such as debugging, state management, and cost, and invites community input on real-world experiences in production.
Enzo Health uses LangSmith to trace and evaluate its clinical AI pipelines in a HIPAA-compliant healthcare setting, and the company is hiring engineers to advance AI in clinical workflows.
Clay demonstrates integrating LangSmith tracing into their agent harness built on Vercel AI SDK, requiring only one line of code.
Phoenix now supports Google's Agent Development Kit for Java, enabling tracing and enrichment of OpenTelemetry spans without modifying agent code.
LangChain has added a capstone project to their LangSmith Essentials course, providing more opportunities to practice debugging and tracing skills in AI agent engineering.
The author created an open-source build profiler called buildprof to visualize where time is spent during software compilation on Linux, motivated by Bun's compile time improvements.
The article highlights the Python package 'wrapture' as an indispensable tool for monkey patching, testing, and observability, with growing tutorials and potential despite being in alpha.
A user discusses strategies to debug AI agent systems in production where all indicators show success but outcomes are incorrect, seeking community advice on evidence and methods for diagnosis.
ArizePhoenix has released an update that automatically searches for duplicates, drafts GitHub issues with linked traces and spans, and uses a secure browser-based GitHub token for filing.
LangChain Academy introduces a free course on LangSmith Essentials to teach the essentials of the LangSmith platform for agent engineering, including tracing, debugging, and deployments.
Introducing wrapture, a new Python library for monkeypatching, testing, and tracing with OpenTelemetry support, entirely AI-assisted in its development.
Epiq is a distributed, Git-native issue tracker that can replay state on demand to audit and trace agentic workflows, addressing key challenges in multi-agent environments.
LangSmith revealed that PodiumHQ's AI agents were acting rationally based on context, contrary to initial perceptions of being broken. Principal Software Engineer Walker Ward discusses tracing agent reasoning end to end.
The article explores an early concept called agentuptime, which addresses verifying AI agent actions by independently checking outcomes to ensure that an agent's completion claim matches the actual state of external systems.
TraceMotive v0.2.0, a local AI-agent debugging tool, has been released with improvements like persistent storage, one-command startup, and trace comparison features, addressing feedback from the initial version.
OpenObserve is an open-source observability platform built in Rust that supports logs, metrics, distributed tracing, and RUM. Its storage cost is 140x lower than Elasticsearch, it can be deployed as a single file, and it serves as an open-source alternative to Datadog.