humanlayer/12-factor-agents
摘要
一份名为'12-Factor Agents'的指南,概述了构建可靠LLM应用的原则,受12 Factor App方法论启发。它提供了针对生产级AI代理开发的可行指南。
查看缓存全文
缓存时间: 2026/05/18 12:33
humanlayer/12-factor-agents 来源:https://github.com/humanlayer/12-factor-agents
12-Factor Agents —— 构建可靠 LLM 应用的原则
致敬 12 Factor Apps (https://12factor.net/)。 本项目的源代码公开在 https://github.com/humanlayer/12-factor-agents,欢迎您的反馈和贡献。让我们一起把它搞清楚!
错过了 AI Engineer World’s Fair?可以在这里看演讲 (https://www.youtube.com/watch?v=8kMaTybvDUw)
想了解 Context Engineering?直接跳到第 3 条 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md)
想为
npx/uvx create-12-factor-agent做贡献?请查看讨论帖 (https://github.com/humanlayer/12-factor-agents/discussions/61)
你好,我是 Dex。我研究 AI agent (https://youtu.be/8bIHcttkOTE) 有一阵子了 (https://theouterloop.substack.com),主要关注 (https://humanlayer.dev) 我几乎尝试了市面上所有的 agent 框架,从即插即用的 crew/langchain 到标榜“极简”的 smolagents,再到号称“生产级”的 langraph、griptape 等等。我和很多非常优秀的创始人聊过,无论是 YC 内部还是外部的,他们都在用 AI 构建令人印象深刻的产品。大多数人都选择自己搭建整个技术栈。我在面向客户的生产级 agent 中很少看到框架的身影。我惊讶地发现,市面上很多自称“AI Agent”的产品其实并没有那么智能。其中大部分主要是确定性代码,只是在恰到好处的地方加入了一些 LLM 步骤,使体验变得真正神奇。Agent(至少是好 agent)并不遵循“给你一个提示、一袋工具,循环直到达成目标” (https://www.anthropic.com/engineering/building-effective-agents#agents) 的模式。相反,它们主要由普通的软件构成。
因此,我试图回答:
有哪些原则可以用来构建 LLM 驱动的软件,使其真正好到可以交付给生产客户?
欢迎来到 12-factor agents。正如自戴利以来每位芝加哥市长都持续在各大机场贴满标语一样,我们很高兴你来到这里。
特别感谢 @iantbutler01 (https://github.com/iantbutler01)、@tnm (https://github.com/tnm)、@hellovai (https://www.github.com/hellovai)、@stantonk (https://www.github.com/stantonk)、@balanceiskey (https://www.github.com/balanceiskey)、@AdjectiveAllison (https://www.github.com/AdjectiveAllison)、@pfbyjy (https://www.github.com/pfbyjy)、@a-churchill (https://www.github.com/a-churchill) 以及旧金山 MLOps 社区对本文档的早期反馈。
简版:12 要素
即使 LLM 继续以指数级变得更强大 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-10-small-focused-agents.md#what-if-llms-get-smarter),仍然会有一些核心工程技巧让 LLM 驱动的软件更可靠、更具扩展性且更易于维护。
- 我们如何走到今天:软件简史 (https://github.com/humanlayer/12-factor-agents/blob/main/content/brief-history-of-software.md)
- 第 1 条:自然语言到工具调用 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-01-natural-language-to-tool-calls.md)
- 第 2 条:掌控你的提示词 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-02-own-your-prompts.md)
- 第 3 条:掌控你的上下文窗口 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md)
- 第 4 条:工具只是结构化输出 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-04-tools-are-structured-outputs.md)
- 第 5 条:统一执行状态与业务状态 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-05-unify-execution-state.md)
- 第 6 条:通过简单 API 启动/暂停/恢复 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-06-launch-pause-resume.md)
- 第 7 条:通过工具调用联系人类 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-07-contact-humans-with-tools.md)
- 第 8 条:掌控你的控制流 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-08-own-your-control-flow.md)
- 第 9 条:将错误压缩到上下文窗口中 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-09-compact-errors.md)
- 第 10 条:小型、专注的 Agent (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-10-small-focused-agents.md)
- 第 11 条:从任何地方触发,在用户所在之处满足他们 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-11-trigger-from-anywhere.md)
- 第 12 条:让你的 Agent 成为无状态规约器 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-12-stateless-reducer.md)
可视化导航
| 第 1 条 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-01-natural-language-to-tool-calls.md) | 第 2 条 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-02-own-your-prompts.md) | 第 3 条 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md) |
| 第 4 条 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-04-tools-are-structured-outputs.md) | 第 5 条 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-05-unify-execution-state.md) | 第 6 条 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-06-launch-pause-resume.md) |
| 第 7 条 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-07-contact-humans-with-tools.md) | 第 8 条 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-08-own-your-control-flow.md) | 第 9 条 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-09-compact-errors.md) |
| 第 10 条 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-10-small-focused-agents.md) | 第 11 条 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-11-trigger-from-anywhere.md) | 第 12 条 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-12-stateless-reducer.md) |
我们如何走到今天
想深入了解我的 agent 之旅以及是什么促使我们走到这里,请查看《软件简史》(https://github.com/humanlayer/12-factor-agents/blob/main/content/brief-history-of-software.md) —— 以下是快速总结:
Agent 的承诺
我们会大量讨论有向图(DG)及其无环兄弟 DAG。首先我要指出……好吧……软件本身就是有向图。我们曾经用流程图表示程序,这是有原因的。
010-software-dag
从代码到 DAG
大约 20 年前,DAG 编排器开始流行起来。我们说的是经典工具如 Airflow (https://airflow.apache.org/)、Prefect (https://www.prefect.io/)、一些前辈,以及一些较新的如 Dagster (https://dagster.io/)、Inngest (https://www.inngest.com/)、Windmill (https://www.windmill.dev/)。它们遵循相同的图模式,并额外提供了可观测性、模块化、重试、管理等功能。
015-dag-orchestrators
Agent 的承诺
我不是第一个这么说的人 (https://youtu.be/Dc99-zTMyMg?si=bcT0hIwWij2mR-40&t=73),但我开始学习 agent 时最大的收获是,你可以扔掉 DAG。不再需要软件工程师编码每一个步骤和边缘情况,你可以给 agent 一个目标和一组转换:
025-agent-dag
让 LLM 实时做出决策,找到路径
026-agent-dag-lines
这里的承诺是:你编写更少的软件,只需给 LLM 图的“边缘”,让它找出节点。你可以从错误中恢复,可以写更少的代码,而且你可能会发现 LLM 能找出新颖的问题解决方案。
Agent 即循环
正如我们稍后会看到的,这实际上并不太奏效。让我们再深入一步——对于 agent,你有一个包含 3 个步骤的循环:
- LLM 决定工作流中的下一步,输出结构化 JSON(“工具调用”)
- 确定性代码执行工具调用
- 结果被附加到上下文窗口
- 重复,直到下一步被确定为“完成”
initial_event = {"message": "..."}
context = [initial_event]
while True:
next_step = await llm.determine_next_step(context)
context.append(next_step)
if (next_step.intent === "done"):
return next_step.final_answer
result = await execute_step(next_step)
context.append(result)
我们初始的上下文就是起始事件(可能是用户消息、cron 触发、webhook 等),然后我们让 LLM 选择下一步(工具)或判断是否已经完成。以下是一个多步骤示例:
027-agent-loop-animation (https://github.com/user-attachments/assets/3beb0966-fdb1-4c12-a47f-ed4e8240f8fd)
GIF 版本
027-agent-loop-animation
为什么需要 12-factor agents?
说到底,这种方法并没有我们期望的那么好。在构建 HumanLayer 的过程中,我与至少 100 位 SaaS 构建者(主要是技术创始人)交流过,他们希望让自己的现有产品更加智能化。这个过程通常是这样的:
- 决定构建一个 agent
- 产品设计、UX 映射、确定要解决什么问题
- 想要快速推进,所以选择 $FRAMEWORK 并开始构建
- 达到 70-80% 的质量水平
- 意识到 80% 对于大多数面向客户的功能来说还不够好
- 意识到要突破 80%,必须逆向工程框架、提示词、流程等
- 从头开始重写
随机声明
声明:我不确定在哪里说这个最合适,但这里似乎还可以:这绝无批评任何框架或那些为之工作的聪明人的意思。它们启发了无数奇迹,加速了 AI 生态系统。我希望这篇文章的一个结果是,agent 框架构建者可以从我以及其他人的经历中学习,让框架变得更好,尤其是对那些想要快速推进但又需要深度控制的构建者。
声明 2:我不会谈论 MCP。我相信你能看到它的定位。
声明 3:我主要使用 TypeScript,原因见这里 (https://www.linkedin.com/posts/dexterihorthy_llms-typescript-aiagents-activity-7290858296679313408-Lh9e?utm_source=share&utm_medium=member_desktop&rcm=ACoAAA4oHTkByAiD-wZjnGsMBUL_JT6nyyhOh30),但这些内容也适用于 Python 或你喜欢的任何语言。
无论如何,回到正题…
构建优秀 LLM 应用的设计模式
在深挖了数百个 AI 库并与几十位创始人合作后,我的直觉是这样的:
- 有一些核心要素能让 agent 变得出色
- 全盘投入一个框架并进行基本上是绿地重写的工作可能适得其反
- 有一些核心原则能让 agent 变得出色,如果你引入一个框架,你可能会得到大部分(或全部)原则
- 但是,我所见过的让构建者将高质量 AI 软件交到客户手中的最快方法,是从 agent 构建中提取小型、模块化的概念,并将其融入现有产品中
- 这些来自 agent 的模块化概念可以被大多数有经验的软件工程师(即使没有 AI 背景)定义和应用
我所见过的让构建者将优秀 AI 软件交到客户手中的最快方法,是从 agent 构建中提取小型、模块化的概念,并将其融入现有产品中
12 要素(再次列出)
- 我们如何走到今天:软件简史 (https://github.com/humanlayer/12-factor-agents/blob/main/content/brief-history-of-software.md)
- 第 1 条:自然语言到工具调用 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-01-natural-language-to-tool-calls.md)
- 第 2 条:掌控你的提示词 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-02-own-your-prompts.md)
- 第 3 条:掌控你的上下文窗口 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md)
- 第 4 条:工具只是结构化输出 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-04-tools-are-structured-outputs.md)
- 第 5 条:统一执行状态与业务状态 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-05-unify-execution-state.md)
- 第 6 条:通过简单 API 启动/暂停/恢复 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-06-launch-pause-resume.md)
- 第 7 条:通过工具调用联系人类 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-07-contact-humans-with-tools.md)
- 第 8 条:掌控你的控制流 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-08-own-your-control-flow.md)
- 第 9 条:将错误压缩到上下文窗口中 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-09-compact-errors.md)
- 第 10 条:小型、专注的 Agent (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-10-small-focused-agents.md)
- 第 11 条:从任何地方触发,在用户所在之处满足他们 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-11-trigger-from-anywhere.md)
- 第 12 条:让你的 Agent 成为无状态规约器 (https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-12-stateless-reducer.md)
荣誉提名 / 其他建议
- 第 13 条:预取你可能需要的所有上下文 (https://github.com/humanlayer/12-factor-agents/blob/main/content/appendix-13-pre-fetch.md)
相关资源
- 在这里对本文档做出贡献 (https://github.com/humanlayer/12-factor-agents)
- 我在 2025 年 3 月的 Tool Use 播客 (https://youtu.be/8bIHcttkOTE) 中讨论了很多这些内容
- 我有时在 The Outer Loop (https://theouterloop.substack.com) 写关于这些主题的文章
- 我与 @hellovai (https://github.com/hellovai) 一起举办关于最大化 LLM 性能 (https://github.com/hellovai/ai-that-works/tree/main) 的网络研讨会
- 我们使用此方法论在 got-agents/agents (https://github.com/got-agents/agents) 下构建开源 agent
- 我们忽略了自己的所有建议,构建了一个用于在 kubernetes 中运行分布式 agent 的框架 (https://github.com/humanlayer/kubechain)
- 本文档中的其他链接:
- 12 Factor Apps (https://12factor.net)
- 构建有效 Agent(Anthropic)(https://www.anthropic.com/engineering/building-effective-agents#agents)
- Prompts are Functions
- 库模式:为什么框架是邪恶的 (https://tomasp.net/blog/2015/library-frameworks/)
- 错误的抽象 (https://sandimetz.com/blog/2016/1/20/the-wrong-abstraction)
- Mailcrew Agent (https://github.com/dexhorthy/mailcrew)
- Mailcrew 演示视频 (https://www.youtube.com/watch?v=f_cKnoPC_Oo)
- Chainlit Demo (https://x.com/chainlit_io/status/1858613325921480922)
- 用于 LLM 的 TypeScript (https://www.linkedin.com/posts/dexterihorthy_llms-typescript-aiagents-activity-7290858296679313408-Lh9e)
- Schema Aligned Parsing (https://www.boundaryml.com/blog/schema-aligned-parsing)
- Function Calling vs Structured Outputs vs JSON Mode (https://www.vellum.ai/blog/when-should-i-use-function-calling-structured-outputs-or-json-mode)
- BAML on GitHub (https://github.com/boundaryml/baml)
- OpenAI JSON vs Function Calling (https://docs.llamaindex.ai/en/stable/examples/llm/openai_json_vs_function_calling/)
- Outer Loop Agents (https://theouterloop.substack.com/p/openais-realtime-api-is-a-step-towards)
- Airflow (https://airflow.apache.org/)
- Prefect (https://www.prefect.io/)
- Dagster (https://dagster.io/)
- Inngest (https://www.inngest.com/)
- Windmill (https://www.windmill.dev/)
相似文章
@mdancho84: 构建自主LLM代理的基础知识——一份38页PDF,揭示了构建具有自主性的AI代理的秘密…
一条推文,推广一份关于构建自主LLM代理的38页PDF指南,提供免费资源学习关于自主性AI系统。
@hwchase17: https://x.com/hwchase17/status/2053157547985834227
文章概述了一个系统的“智能体开发生命周期”(构建、测试、部署、监控),以有效创建和管理 AI 智能体,重点介绍了 LangChain、LangGraph 和 CrewAI 等关键框架。
构建智能体
关于构建AI智能体的指南或资源。
LLM Agents Factory: Retrieval of Domain-Specific LLM Agents
The paper presents LLM Agents Factory, a retrieval-based framework that constructs domain-specific LLM agents from a base of over 20K predefined agent profiles, offering a cost-efficient and controllable alternative to dynamic agent generation. Experiments show accuracy comparable to AutoGen with a 120B backbone at substantially lower inference cost.
可信代理AI:故障模式、缓解策略与自主LLM系统的生命周期框架
论文回顾了基于LLM的可信代理AI系统的故障模式和缓解策略,并介绍了可信代理开发生命周期(TADL)框架,用于安全开发。