Tag
LangChain published a new guide on scaling AI agents from isolated experiments to enterprise production, featuring lessons from Schneider Electric, Vodafone, and monday.com across shared agent platforms, LLMOps, observability, multi-agent architectures, security, and data residency.
The article explains the key differences between DevOps, MLOps, and LLMOps, highlighting how each addresses distinct challenges in software development, machine learning, and LLM applications, with a focus on unique monitoring and optimization in LLMOps.
The article discusses the importance of system design for AI agents, covering concepts like Agent Harness, LLMOps, and Evals, and provides a proof-of-concept implementation with plans for future parts.
This arXiv paper presents a unified LLMOps architecture for real-time, enterprise-ready LLM deployments, integrating data ingestion, continual learning, RAG, and feedback loops. It introduces components like AIPO, STAR+FAR, and SAGE to address knowledge staleness, hallucination, and latency-cost trade-offs in regulated sectors.
Schneider Electric uses LangChain's LangSmith to run over 60 production AI agents across 100+ countries, serving 160,000 employees with their AI Assistant, demonstrating enterprise-scale LLMOps.
TensorZero, an open-source LLMOps platform that raised $7.3 million in seed funding, has archived its GitHub repository. The platform provides a unified gateway, observability, evaluation, optimization, and experimentation for LLMs.
RiskKernel is a self-hosted, single Go binary that enforces hard per-run budgets (cost, loop count, wall-clock), kill switches, and human approval gates for AI agents, supporting Anthropic and OpenAI providers with no telemetry.
An open-source interactive playbook for building an Agentic DevOps pipeline, covering observability, test-driven prompt evaluations, guardrails, and cost control for multi-agent systems.