Tag
The article discusses strategies for capping retries in AI agents to balance cost and performance, emphasizing the need to differentiate between retryable errors and cases requiring escalation to more capable models in production.
The article highlights the hidden costs of building AI agents on external models, specifically unannounced behavioral regressions after updates that can disrupt automated workflows, and suggests strategies like version pinning to mitigate risks.
The article argues that many enterprise AI pilots fail because they use generative models for tasks better suited to discriminative models, highlighting differences in mathematical objectives and update mechanisms.
The author argues that multi-agent collaboration in AI is currently a false premise because language models' flaws are amplified in such systems, making them unreliable for production use.
The author discusses the challenge of keeping AI agent stacks current with evolving models and tools, and seeks insights from production teams on benchmarking and update practices.
This article outlines a 12-step roadmap for AI Agent Engineers in 2026, focusing on seven interconnected pillars like context, tools, and memory, with Claude-based workflows to build reliable production agents.
LinkedIn presents a self-evolving agentic customer support system that integrates RAG with evolutionary auto-prompting and modular evaluation, achieving significant gains in production A/B tests including a 9.0-point increase in QA self-serve and 30.6-point improvement in routing accuracy.
This article argues that LLM hallucinations in production are typically a system architecture problem rather than a model problem, and outlines four key guardrails: RAG grounding, live tools/function calling, selective human oversight, and red teaming/adversarial testing.
Uber Eats describes its self-tuning multi-agent AI system for automatically fixing merchant photos, using router, editor, QA agents with centralized logging and an autonomous Diagnoser Agent that rewrites prompts and auto-deploys after passing a golden benchmark.
The article argues that selling AI wrappers (simple interfaces over existing models) is easier than building AI systems that actually work reliably in production, highlighting challenges in deployment.
A technical write-up discusses the shift from agent loops to structured graphs in production AI agent work, backed by references to durable execution engines (Temporal, Restate) and research like AFlow which uses Monte Carlo Tree Search to optimize workflow graphs.
This paper proposes Agentic Context Management (ACM), treating agent memory as a lifecycle problem with five primitives, and presents Maximem Synap, a reference implementation achieving strong benchmark results.
The article discusses the growing dominance of open-weight models, especially from Chinese firms, in production AI workloads, challenging the relevance of frontier models from companies like Anthropic and OpenAI.
A production team migrated their QA agent from GPT-5.3-codex to MiniMax M3, finding that while the new model uses more tokens per task, its lower per-token price led to a 55% median cost reduction. The post also highlights the importance of inference provider selection and hidden reasoning tokens affecting effective pricing.
The article discusses how AI agent systems waste spend in production due to hidden inefficiencies like over-context, inappropriate model selection, and retries, and questions what runtime decisions should govern model calls.
Observations on the shift from addressing AI hallucinations to the more pressing problem of production AI failures, emphasizing the need for system reliability, tracking decisions, and limiting blast radius in enterprise deployments.
The author shares practical learnings from running AI agent loops for a month, emphasizing the importance of loop contracts, state, and logs to make agents autonomous and reliable.
Schneider Electric uses LangChain's LangSmith to run over 60 production AI agents across 100+ countries, serving 160,000 employees with their AI Assistant, demonstrating enterprise-scale LLMOps.
This article argues that common LLM cost advice focusing on token reduction is too shallow, and that the more impactful strategy in production is to route different workflow steps to different models rather than using a single default model.
The article discusses the challenge of building a reliable, long-running multi-agent production system, noting that it currently requires integrating multiple fragmented tools such as CrewAI, Temporal, Browserbase, and Langfuse, and questions whether a more unified runtime exists.