Workshop on Sep 12: shipping LLM systems that actually survive production
摘要
There's a hands-on masterclass on Sep 12 for anyone building agents who wants real engineering discipline instead of finding out something broke from an invoice. Covers: Tracing, token/cost monitoring, latency budgets, and caching, so problems show up in dashboards, not surprise bills Tool-using agents with function calling, validation, guardrails, retries, and fallbacks, so a bad call degrades gracefully instead of compounding Versioned prompts with regression tests, so an edit can't silently d
相似文章
大多数LLM功能在发布时缺乏我们绝不会在常规软件中忽略的工程纪律
文章强调了在发布LLM功能时缺乏工程严谨性,并推广了9月12日的大师课,该课程教授有纪律的评估、测试和生产方法。
低延迟系统中工具制作与自进化LLM代理
本文提出了一种方法,将重复的标准操作流程步骤编译为经过验证、有版本管理的工具,在部署前完成,替代推理时的代码生成。在一个配送中心的报警分类系统中,该方法将p50延迟降低了42%,端到端错误率降低了最多53%。
@Zephyr_hg: https://x.com/Zephyr_hg/status/2062176187384807488
一篇实用指南认为,掌握子代理需要在周末构建四个特定工作流,涵盖分解、上下文打包、验证和成本控制,而不是花费200小时看教程。
生产环境中的高效基准测试:一项演化LLM代理研究
本文探索生产环境中LLM代理的高效重复评估方法,比较自适应测试和固定子集等技术,并提供部署的实用建议。
有人在AI智能体交付的“第二天”阶段碰壁了吗?
一位从业者分享了将AI智能体部署到生产环境中的真实挑战,指出治理、审计和部署防护措施现在已成为瓶颈,而非智能体构建本身,并提到了Lyzr Control Plane和微软参考架构等新兴解决方案。