Tag
A comprehensive guide to 15 AI agent design patterns for production systems, explaining when to use each pattern and common pitfalls.
The author built a runtime control layer to address the problem of AI agents failing silently in production environments.
A discussion on strategies for managing token budgets when deploying multiple AI agents in production, covering cost and efficiency considerations.
Blog post by Xingyao Wang explaining why OpenHands V1 chose a different architecture from Claude Managed Agents, arguing that reliability comes from implementation details rather than topology.
A curated, open-source learning path for building voice agents, covering from STT to production, with 190+ resources and a 5-week plan.
A discussion on the missing infrastructure required to run AI agents in production, including monitoring, permissions, recovery, and audit trails, questioning whether this will become a new infrastructure category.
Open-sourced a full-stack production starter for voice agents using LiveKit, FastAPI, and React, handling both web and telephony with a single code path, deployable via Docker Compose.
Personal lessons on evaluating AI agents in production, including mapping symptoms to layers, using trajectory evaluation, calibrating LLM judges, converting failures to test cases, and performing adversarial testing.
A practical guide outlining seven prioritized security layers for AI agents before production, including hardening system prompts, adversarial testing, input/output scanning, and multi-turn session tracking, based on findings that 73% of production AI deployments have prompt injection exposure.
A blog post argues that current agent checkpointing is insufficient for production-grade resiliency, highlighting gaps like failure detection, automatic retries, and high availability, and suggests building agents on a highly-available orchestration layer.
Discusses common failure modes of AI agents in enterprise environments, such as over-reliance on long-term memory and stateless tool gating leading to security risks.
A developer shares the challenge of debugging multi-step agents in production, where failures are hard to trace due to complex tool use and confident wrong answers, and asks the community for better monitoring and regression detection approaches.
Adaline 2.0 is an agent self-improvement layer that watches real user interactions, clusters failures by pattern, automatically writes hundreds of tests daily, and generates new agent candidates for approval before deployment.
A thread explaining the four essential layers for building production-grade RAG systems beyond simple chunk-embed-retrieve-generate: intelligent query routing, advanced indexing, multi-type retrieval, and continuous evaluation.
The tweet recommends an article on agent architecture in production, highlighting the use of Traces to diagnose issues and implement an iterative improvement loop.
This paper presents a five-plane reference architecture for runtime governance of production AI agents, addressing security risks from delegated actions. It defines primitives, invariants, and an evaluation framework to ensure safety and utility.
AI agents often fail due to messy environments rather than bad models; improving environment stability makes simple agents perform well.
A discussion about deploying multi-agent AI systems in production, where different agents handle planning, execution, communication, and project management, asking about real-world experiences and bottlenecks.
This article presents a repository that systematically gathers and benchmarks 35 AI agent architectures, helping developers choose effective control structures for production systems.
Discusses a common failure mode in AI agents where the model confidently claims to have performed an action (e.g., sending an email) without actually executing the required tool call, and asks the community how they detect and handle such silent failures in production.