Tag
The article argues that current security tools overlook the risks posed by AI agents operating in production environments, suggesting a misalignment in monitoring strategies.
A statistic reveals that while 89% of agent teams in production have observability, only 52% run evaluations, raising questions about how prompt changes are gated.
Decagon runs 90% of workloads on fine-tuned open-source models for latency and performance, while overall enterprise spending on open-source LLMs has dropped to 11% due to a surge in new use cases using frontier models. The article argues that as use cases mature, they will migrate from closed to open-source models.
A discussion prompt seeking approaches to AI governance for autonomous agents in production environments.
The article discusses common causes of cost spikes in AI workflows, such as retries, repeated tool calls, long-running workflows, and growing context, and asks how teams investigate such issues.
A discussion about strategies for detecting and preventing erroneous tool calls by AI agents before they execute in production environments.
The author reflects on their experience building AI agents and argues that creating production-ready agents should not be overly complex.
A developer argues that most production AI agents lack essential observability like session traces and cost tracking, comparing it to deploying a web app without monitoring. The article questions whether agent observability is an unsolved problem.
A discussion about handling prompt injection attacks in AI agents that read external content like emails and webpages, exploring production-level defenses and the subtle threats beyond obvious patterns.
Vercel's AI Gateway now supports routing rules that allow developers to dynamically rewrite model routes (e.g., from retired models like Claude Fable-5 to Claude Opus-5) without code changes, ensuring production workloads remain resilient.
A real AI agent running on OpenClaw hosts an AMA about production AI agents.
A discussion on where to place guardrails to prevent AI coding agents from making unauthorized changes, exploring friction points at various stages of the deployment workflow.
A developer outlines a structured skill stack for production-grade AI agents, covering planning, tool-use, permissions, recovery, observation, budget, and escalation as explicit contracts rather than loose tool wrappers.
A developer seeks advice on how to control and bound AI agents' actions in production environments, particularly when they interact with real systems like databases and customer data, asking about current practices and whether this is a known headache.
Mycelium is an experimental open-source runtime guard that prevents predictable failures in AI agents before they reach the LLM, aiming to improve reliability in production environments.
REAP is an automated pipeline that curates production-derived benchmarks for coding agents from real developer-agent sessions, using LLM-based classification and stability checks to ensure reliable evaluation without manual labeling.
A developer argues that the harness (critics, scaffolding) around an AI model is more important than the model itself, sharing an example where a 27B model with good critics became usable for coding work.
Fireworks AI announces Serverless 2.0, introducing three serving tiers (Standard, Priority, Fast) to handle traffic congestion without pre-provisioning GPUs, enabling per-request routing for reliability and cost efficiency.
The article argues that granting broad tool permissions to AI agents is an inadequate abstraction for production environments, suggesting more granular control is needed.
This article explores common reasons why AI agents fail shortly after being deployed in production, highlighting pitfalls and lessons learned.