Workshop on Sep 12: shipping LLM systems that actually survive production

Reddit r/AI_Agents 新闻

摘要

There's a hands-on masterclass on Sep 12 for anyone building agents who wants real engineering discipline instead of finding out something broke from an invoice. Covers: Tracing, token/cost monitoring, latency budgets, and caching, so problems show up in dashboards, not surprise bills Tool-using agents with function calling, validation, guardrails, retries, and fallbacks, so a bad call degrades gracefully instead of compounding Versioned prompts with regression tests, so an edit can't silently d

There's a hands-on masterclass on Sep 12 for anyone building agents who wants real engineering discipline instead of finding out something broke from an invoice. Covers: Tracing, token/cost monitoring, latency budgets, and caching, so problems show up in dashboards, not surprise bills Tool-using agents with function calling, validation, guardrails, retries, and fallbacks, so a bad call degrades gracefully instead of compounding Versioned prompts with regression tests, so an edit can't silently degrade quality A real eval harness combining deterministic checks and LLM-as-judge Bootstrap confidence intervals and paired significance testing for model comparisons Evaluated RAG with retrieval metrics (recall@k, MRR) Led by Bruno Gonçalves, PhD, founder of Data For Science, who trains engineers at Fortune 500 companies on this exact stack. Link in comments.
查看原文

相似文章

低延迟系统中工具制作与自进化LLM代理

arXiv cs.CL

本文提出了一种方法,将重复的标准操作流程步骤编译为经过验证、有版本管理的工具,在部署前完成,替代推理时的代码生成。在一个配送中心的报警分类系统中,该方法将p50延迟降低了42%,端到端错误率降低了最多53%。

有人在AI智能体交付的“第二天”阶段碰壁了吗?

Reddit r/artificial

一位从业者分享了将AI智能体部署到生产环境中的真实挑战,指出治理、审计和部署防护措施现在已成为瓶颈,而非智能体构建本身,并提到了Lyzr Control Plane和微软参考架构等新兴解决方案。