Tag
An analysis of the gap between episodic and procedural memory in LLM agents, citing a new paper (Memp) from Zhejiang University and Alibaba that builds procedural memory from agent trajectories and uses failure signals to revise stored procedures.
At the AI Engineer World's Fair, Phil Schmid gave a talk on why vibe-checking agent skills break in production and how to build reliable automated evals using negative test cases, skill limits, and ablation tests.
A developer reflects on critical safety measures—such as spending caps, rate limits, and fallback models—that should be in place before launching an AI Agent app publicly to avoid hidden costs and unexpected behaviors.
Discusses best practices for testing AI agents before deploying them to production environments.
Netflix shares the design decisions behind its in-house LLM serving stack, including engine selection (vLLM), model packaging, API surface, and deployment strategy, highlighting trade-offs revealed under production load.
Jerry Liu highlights the engineering challenges of productionizing agentic retrieval systems, emphasizing that success depends on careful tuning of chunking, synchronization, reranking, and tool API design rather than novel techniques.
A comparison of three voice AI agents — Retell, Vapi, and Plura AI — evaluating their performance for production use cases.
Kevin Bass advocates for quickly shipping 'vibe code slop' into production, arguing that future AI models will handle architectural improvements, freeing humans from that task.
Toyota's enterprise AI team shared the ToyotaGPT platform at the Interrupt conference. Based on LangGraph and LangSmith, it reduced AI agent delivery from 6 months to 4 days, and has over 50 agents running in production, saving millions of dollars.
A 7-week course with 7.7k stars on GitHub, building a production-grade RAG system from scratch, covering Docker, FastAPI, hybrid search, LangGraph agentic RAG, and a Telegram bot, with hands-on coding throughout.
A discussion asking how developers handle changes in tools, APIs, or model versions that their AI agents depend on in production, including detection, fixes, and costs.
The state of open source AI report by Mozilla highlights that open-weight models have reached parity with closed models on many tasks, while inference costs have dropped 50× in 36 months. The majority of production tokens now route through open models, and the competitive landscape has shifted to the agentic layer above.
A practitioner shares concerns about an upcoming audit revealing undocumented AI agents in production, highlighting governance gaps and risks with customer PII access.
Alvin Sng explains why their team moved away from using client SDKs for Stripe, WorkOS, and Slack, opting instead to call their REST APIs directly via a centralized wrapper. They argue that SDKs hide critical debugging details, are fragile in production, and encourage anti-patterns that are now more easily avoided with AI-assisted coding.
The article discusses the significant challenges enterprises face when deploying AI agents to production, highlighting the lack of standard deployment infrastructure, security concerns, and the need for an orchestration layer to manage agent lifecycles.
The article shares production learnings for reliably generating structured JSON output from LLMs, covering methods like JSON mode, schema validation, and retry loops, achieving 99.5% validity.
Microsoft shares insights from shipping thousands of production AI agents at enterprise scale, covering the engineering challenges of moving from prototype to production, including the agent harness, retrieval-as-a-subagent, agent identity, and rubric-based evaluation loops.
A survey shows that most teams keep agents on a short leash, and data indicates narrow-scope agents succeed 65% of the time vs 16% for broad scope, suggesting the reliability issue may be more about scope than model capability.
An article introducing loop engineering as the 2026 successor to prompt engineering, focusing on designing agent loops rather than hand-writing prompts, with emphasis on the verifier as the bottleneck.
An AI engineer released an open-source project teaching how to build a local RAG system from scratch and a production-grade agentic architecture with LangGraph, hybrid retrieval, caching, and observability.