Tag
A practitioner shares 8 months of experience running a voice agent for a law firm, detailing challenges like latency, turn-taking, and post-call workflows, and provides a working system prompt.
Orka, an open-source control layer for AI agents in production, has been released.
A discussion on the challenges and successful strategies for deploying AI agents in production at scale, covering common pain points and effective solutions.
The article discusses the challenges that arise when AI agents transition from demos to production, focusing on the need for operational control planes that provide idempotency, approval tracking, and operational explainability rather than just model reasoning.
Explores techniques used in production AI agents to avoid hallucinations when controlling real devices with multiple tools.
A detailed guide on building a production-grade agent harness for multi-agent LLM systems, covering components like orchestrator, subagents, skills, backend state management, and context engineering.
An essay exploring the concept of 'taste' as the ability to make high-quality qualitative judgments, arguing that taste becomes more valuable as production is commoditized by AI.
A detailed technical guide explaining how PgBouncer works as a PostgreSQL connection pooler, covering its pooling modes, production deployment, and common pitfalls.
A discussion on how companies should measure the real-world impact of AI agents and skills in production environments, rather than relying solely on benchmark results.
The author contrasts polished AI agent demos with the reality of production systems, noting that most agent code is for error handling and guardrails rather than the core intelligence.
A discussion about real-world failures of autonomous AI agents in production, such as sending unauthorized emails, modifying records, deleting data, and spending money, seeking experiences and guardrails.
Haystack is an open-source AI framework for building production-ready agents and RAG pipelines, supporting multimodal, conversational, and content generation applications.
Discussion about scoping permissions for AI agents in production to avoid dangerous database actions, suggesting read-only mirrors, approval steps, or hard walls between suggestion and execution.
Latitude has launched an open-source, MIT licensed monitoring platform that turns AI agent conversations into production debugging data, helping teams see sessions, catch failures, and fix issues directly from their editor.
SGLang provided Day-0 support for DeepSeek-V4, and collaboration between LMSys and NVIDIA engineering teams achieved up to 5x throughput increase in production, with improvements shown on the SemiAnalysis InferenceX dashboard.
Discussion about whether ML teams are actually testing model security risks like extraction and poisoning in production, noting that security review for models lags behind regular software.
A developer asks for advice on building a reliable company OS where AI agents and humans collaborate in production, focusing on long-term memory, workflow state, and agent handoffs. They share their current tool stack and question whether RAG, event sourcing, or custom memory systems are the missing piece.
This article poses critical questions teams should consider before trusting AI agents in real workflows, focusing on reliability, accountability, and correctness.
PP-OCRv6 is the latest generation of PaddleOCR's universal OCR model family, offering three tiers from 1.5M to 34.5M parameters, supporting 50 languages, and achieving significant accuracy improvements over previous versions.
This article discusses how AI agent demos often succeed while production deployment reveals critical security and authorization issues, emphasizing that model quality does not solve problems like access control, data leaks, and auditability.