Tag
The article discusses the growing disparity between AI agent capabilities and the necessary control mechanisms for production use, highlighting challenges in permissions, escalation, and accountability.
The article explores the key challenges in transitioning LLM application ideas to production in the EU, covering legal compliance, traceability, abuse prevention, and accountability, and seeks concrete examples from practitioners.
This post discusses common challenges with unattended AI agents, such as looping, overspending, and incorrect task completion, and asks how practitioners handle issues like verification, stall detection, and hard limits in production.
The article discusses the need for re-checking permissions and actions immediately before an AI agent executes a task in production, due to potential changes in conditions like record states or approvals.
The author questions why tiny AI models with fewer than 50M parameters or swarms of specialized micro-models are rarely deployed in production, speculating on reasons like tooling biases or the convenience of generalist models.
This paper reports on a production migration for a customer support conversational assistant, replacing a monolithic AI model with an agentic orchestration system to improve precision, reduce hallucinations, and lower serving costs.
Mercury 2.5 is announced as the most capable diffusion language model, with a 40% increase in intelligence over Mercury 2, operating at over 1,100 tokens/sec on NVIDIA GPUs, and optimized for production with low latency and cost.
A post soliciting feedback on the pain points of deploying AI agents in real workplace settings, highlighting issues like security reviews and operational challenges.
The article explains why AI demos are not suitable for production use and outlines key engineering practices needed to build reliable AI systems.
An inquiry into the real-world effectiveness of AI agents, highlighting challenges in production such as inefficiency and error rates, and inviting community experiences.
The author shares three unexpected learnings from running a language-learning product with per-user persistent AI agents, including benefits in memory handling and scalability, and challenges with proactive engagement and security.
The article discusses challenges and asks for community experiences regarding the breakdown of AI agent memory systems after months in production use.
The article warns about the limitations of free-tier vector databases, highlighting issues like data deletion and deployment constraints, and advises choosing based on where your AI agent runs.
The article discusses considerations for integrating vision capabilities like screenshots and documents into AI agent workflows, referencing the DeepSeek-V4-Flash-Vision-Exp experimental API and suggesting evaluation steps before production use.
Two talks and a blog post argue that the feedback loop and harness engineering are more important than model weights for owning AI intelligence in production, highlighting context management and cost considerations.
The article discusses common issues with AI agents in production, such as handling incomplete context, API failures, and state management, emphasizing that system design often outweighs model decisions.
The paper presents a scalable framework that bridges search and CRM workflows using AI-powered Product Research Agents for proactive customer re-engagement in e-commerce, evaluated in a production deployment with improved CTR and sales.
This article explores common misconceptions in agentic AI, highlighting the gap between theoretical assumptions and real-world production challenges, and invites practitioners to share their experiences.
A firsthand report detailing the costs incurred from running AI agents in production for one week, emphasizing that token usage dominates expenses and sharing insights on cost management like integer math for billing.
Evaluation of gpt-4o-mini and gpt-4o on an event classification system showed gpt-4o performed better, but both models had unreliable confidence scores for real-world decision-making.