vectorize-io/hindsight
摘要
Hindsight 是一个智能体记忆系统,专为增强AI智能体的长期学习能力而设计,在记忆基准测试中达到了最先进的精度。
查看缓存全文
缓存时间: 2026/09/24 15:09
vectorize-io/hindsight
来源:https://github.com/vectorize-io/hindsight
Hindsight 文档 (https://hindsight.vectorize.io) • 集成 (https://hindsight.vectorize.io/integrations) • 示例手册 (https://hindsight.vectorize.io/cookbook) • 基准测试 (https://benchmarks.hindsight.vectorize.io/) • 论文 (https://arxiv.org/abs/2512.12818) • Hindsight Cloud (https://ui.hindsight.vectorize.io/signup)
发布 (https://github.com/vectorize-io/hindsight/actions/workflows/release.yml) 版本 (https://pypi.org/project/hindsight-api/) PyPI 下载 (https://pypi.org/project/hindsight-client/) NPM 下载 (https://www.npmjs.com/package/@vectorize-io/hindsight-client) Slack 社区 (https://vectorize.io/slack) 许可证:MIT (https://opensource.org/licenses/MIT)
什么是 Hindsight?
HindsightTM 是一个智能体记忆系统,旨在创建能够随时间学习的更智能的智能体。大多数智能体记忆系统侧重于回忆对话历史。Hindsight 则专注于让智能体学习,而不仅仅是记忆。它消除了 RAG 和知识图谱等替代技术的缺点,并在长期记忆任务中提供了最先进的性能。
目录
- 记忆性能与准确性
- 快速开始
— 服务器 · 客户端 · 支持平台 · 嵌入式 - 将 Hindsight 添加到你的智能体
— LLM 包装器 · 集成 · 编程智能体 · MCP - 核心概念
— 记忆类型 · 保存 / 召回 / 反思 · 观察 · 心智模型与知识页 · 记忆库 - 使用场景
- 生产环境运行
- 资源
记忆性能与准确性
根据基准测试结果,Hindsight 是迄今为止测试过的最准确的智能体记忆系统。它在广泛用于评估各种对话 AI 场景记忆系统性能的 LongMemEval 基准测试中取得了最先进的性能。
截至 2026 年 1 月,Hindsight 及其他智能体记忆解决方案的当前报告性能如下所示:
概览
实时、持续更新的结果(包括每个模型的准确性、延迟和成本)发布在 benchmarks.hindsight.vectorize.io (https://benchmarks.hindsight.vectorize.io/)。
Hindsight 的基准测试数据已由弗吉尼亚理工学院桑哈尼人工智能与数据分析中心 (https://sanghani.cs.vt.edu/) 和《华盛顿邮报》的研究合作伙伴独立复现。其他分数由软件供应商自行报告。
Hindsight 正在被财富 500 强企业和数量不断增长的 AI 初创公司用于生产环境。
🤖 正在使用编程智能体?
安装 Hindsight 文档技能,以便在编码时即时访问文档:npx skills add https://github.com/vectorize-io/hindsight --skill hindsight-docs适用于 Claude Code、Cursor 和其他 AI 编程助手。
快速开始
1. 启动服务器
Docker(推荐)
export OPENAI_API_KEY=sk-xxx
docker run -it --pull always --name hindsight --restart unless-stopped -p 8888:8888 -p 9999:9999 \
-e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
-v hindsight-data:/home/hindsight/.pg0 \
ghcr.io/vectorize-io/hindsight:latest
API:http://localhost:8888
UI:http://localhost:9999
Hindsight 通过 HINDSIGHT_API_LLM_PROVIDER 支持 25+ 个 LLM 提供商——托管服务(openai、anthropic、gemini、groq、bedrock、vertexai、minimax、deepseek、atlas、meta…)、完全本地化(ollama、lmstudio、llamacpp)、任何 OpenAI 兼容端点,以及能够访问其他提供商的网关(litellm、litellmrouter)。
现有订阅也可使用:openai-codex(ChatGPT Plus/Pro)、claude-code(Claude Pro/Max)、cursor(Cursor)和 github-copilot(GitHub Copilot)无需 API 密钥。
查看支持的模型 (https://hindsight.vectorize.io/developer/models)。
Docker(外部 PostgreSQL)
export OPENAI_API_KEY=sk-xxx
export HINDSIGHT_DB_PASSWORD=choose-a-password
cd docker/docker-compose
docker compose up
Oracle AI 数据库也支持企业部署,具有完全相同的功能。详情请参阅存储文档 (https://hindsight.vectorize.io/developer/storage)。
裸机(pip)
pip install hindsight-api
export HINDSIGHT_API_LLM_API_KEY=sk-xxx
hindsight-api
Kubernetes(Helm)
helm install hindsight oci://ghcr.io/vectorize-io/charts/hindsight \
--set api.llm.provider=openai \
--set api.llm.apiKey=sk-xxx \
--set postgresql.enabled=true
托管(无需服务器)
Hindsight Cloud (https://vectorize.io/pricing) 是托管选项:自动扩展的托管基础设施,外加仪表板、备份、团队协作和 99.9% 的正常运行时间 SLA。
采用基于使用量的计费方式,并提供免费额度入门——无固定月费或按席位费用。只需将任何客户端指向 https://api.hindsight.vectorize.io 并使用你的 API 密钥即可,完全跳过部署。
比较自托管、Cloud 和企业版 → (https://vectorize.io/pricing) · 注册 → (https://ui.hindsight.vectorize.io/signup)
所有选项(包括 Windows 和气隙环境)均在安装指南 (https://hindsight.vectorize.io/developer/installation) 中介绍。
2. 连接客户端
pip install hindsight-client -U # Python
npm install @vectorize-io/hindsight-client # Node.js / TypeScript
go get github.com/vectorize-io/hindsight/hindsight-clients/go # Go
curl -fsSL https://hindsight.vectorize.io/get-cli | bash # CLI
Python
from hindsight_client import Hindsight
client = Hindsight(base_url="http://localhost:8888")
# 保存:存储信息
client.retain(bank_id="my-bank", content="Alice works at Google as a software engineer")
# 召回:搜索记忆
client.recall(bank_id="my-bank", query="What does Alice do?")
# 反思:生成考虑处置的回复
client.reflect(bank_id="my-bank", query="Tell me about Alice")
Node.js / TypeScript
const { HindsightClient } = require('@vectorize-io/hindsight-client');
const main = async () => {
const client = new HindsightClient({ baseUrl: 'http://localhost:8888' });
await client.retain('my-bank', 'Alice loves hiking in Yosemite');
const results = await client.recall('my-bank', 'What does Alice like?');
console.log(results);
};
main();
完整参考:Python (https://hindsight.vectorize.io/sdks/python) · Node.js (https://hindsight.vectorize.io/sdks/nodejs) · Go (https://hindsight.vectorize.io/sdks/go) · CLI (https://hindsight.vectorize.io/sdks/cli) · REST API (https://hindsight.vectorize.io/api-reference)
支持平台
| 平台 | Docker | 裸机(pip) | 嵌入式数据库(pg0) |
|---|---|---|---|
| Linux (x86_64, ARM64) | ✅ | ✅ | ✅ |
| macOS (Apple Silicon / arm64) | ✅ | ✅ | ✅ |
| macOS (Intel / x86_64) | ✅ | ⚠️ | ✅ |
| Windows (x86_64) | ✅ | ✅ | ✅ |
⚠️ Intel Mac:请使用 hindsight-all-slim——详情请参阅安装指南 (https://hindsight.vectorize.io/developer/installation#supported-platforms)。
Python 嵌入式(无需服务器)
pip install hindsight-all -U
在 Intel (x86_64) Mac 上,请改为安装 hindsight-all-slim——请参阅支持平台。
import os
from hindsight import HindsightServer, HindsightClient
with HindsightServer(
llm_provider="openai",
llm_model="gpt-5-mini",
llm_api_key=os.environ["OPENAI_API_KEY"]
) as server:
client = HindsightClient(base_url=server.url)
client.retain(bank_id="my-bank", content="Alice works at Google")
results = client.recall(bank_id="my-bank", query="Where does Alice work?")
还提供了 Node.js 等效方案 (https://hindsight.vectorize.io/sdks/hindsight-all-npm) 和守护进程 CLI (https://hindsight.vectorize.io/sdks/embed)。
将 Hindsight 添加到你的智能体
LLM 包装器(仅需两行代码)
将记忆添加到现有智能体最简单的方法是使用 LLM 包装器。将你的 LLM 客户端替换为包装后的客户端——之后每次调用都会自动存储和检索记忆,无需对代码进行其他更改。
pip install hindsight-litellm
from openai import OpenAI
from hindsight_litellm import wrap_openai
# 包装你现有的 LLM 客户端即可完成。
# 默认使用 Hindsight Cloud;对于自托管服务器,请传入 hindsight_api_url。
client = wrap_openai(
OpenAI(),
bank_id="user-123",
hindsight_api_url="http://localhost:8888",
)
# Hindsight 会在调用前召回相关记忆
# 并在调用后保存对话。
response = client.chat.completions.create(
model="gpt-5-mini",
messages=[{"role": "user", "content": "What do you know about me?"}],
)
wrap_anthropic() 对 Anthropic SDK 执行相同操作,并且每个设置(记忆库、召回预算、事实类型、反思而非召回)都可以通过 hindsight_* 关键字参数在每次调用时覆盖。
LiteLLM 位于底层,因此相同的集成覆盖 100+ 个模型。
请参阅 LiteLLM 集成 (https://hindsight.vectorize.io/sdks/integrations/litellm)。
如果你需要明确控制何时存储和召回记忆,请直接使用 SDK 或 REST API。
集成
60+ 个集成——大多数无需更改代码。
| 编程智能体 | Claude Code (https://hindsight.vectorize.io/sdks/integrations/claude-code) · Codex (https://hindsight.vectorize.io/sdks/integrations/codex) · Cursor (https://hindsight.vectorize.io/sdks/integrations/cursor) · GitHub Copilot (https://hindsight.vectorize.io/sdks/integrations/github-copilot) · opencode (https://hindsight.vectorize.io/sdks/integrations/opencode) · Cline (https://hindsight.vectorize.io/sdks/integrations/cline) · Aider (https://hindsight.vectorize.io/sdks/integrations/aider) · Zed (https://hindsight.vectorize.io/sdks/integrations/zed) · Continue (https://hindsight.vectorize.io/sdks/integrations/continue) · Roo Code (https://hindsight.vectorize.io/sdks/integrations/roo-code) · OpenHands (https://hindsight.vectorize.io/sdks/integrations/openhands) |
| 智能体框架 | LangGraph / LangChain (https://hindsight.vectorize.io/sdks/integrations/langgraph) · LlamaIndex (https://hindsight.vectorize.io/sdks/integrations/llamaindex) · CrewAI (https://hindsight.vectorize.io/sdks/integrations/crewai) · Pydantic AI (https://hindsight.vectorize.io/sdks/integrations/pydantic-ai) · OpenAI Agents SDK (https://hindsight.vectorize.io/sdks/integrations/openai-agents) · Google ADK (https://hindsight.vectorize.io/sdks/integrations/google-adk) · Agno (https://hindsight.vectorize.io/sdks/integrations/agno) · Strands (https://hindsight.vectorize.io/sdks/integrations/strands) · AutoGen (https://hindsight.vectorize.io/sdks/integrations/autogen) · Microsoft Agent Framework (https://hindsight.vectorize.io/sdks/integrations/agent-framework) · Vercel AI SDK (https://hindsight.vectorize.io/sdks/integrations/ai-sdk) · Haystack (https://hindsight.vectorize.io/sdks/integrations/haystack) |
| 无代码 / 低代码 | n8n (https://hindsight.vectorize.io/sdks/integrations/n8n) · Zapier (https://hindsight.vectorize.io/sdks/integrations/zapier) · Dify (https://hindsight.vectorize.io/sdks/integrations/dify) · Flowise (https://hindsight.vectorize.io/sdks/integrations/flowise) |
| 应用与工具 | ChatGPT (https://hindsight.vectorize.io/sdks/integrations/chatgpt) · Perplexity (https://hindsight.vectorize.io/sdks/integrations/perplexity) · Obsidian (https://hindsight.vectorize.io/sdks/integrations/obsidian) · Pipecat (https://hindsight.vectorize.io/sdks/integrations/pipecat) · Vapi (https://hindsight.vectorize.io/sdks/integrations/vapi) |
👉 浏览所有集成 (https://hindsight.vectorize.io/integrations)
编程智能体
一个包即可为 CLI 编程智能体提供长期项目记忆:一个自动从 git 历史和过去会话构建的、每个仓库一个的记忆库,在智能体开始工作时注入,加上涵盖架构、约定和进行中工作的精选知识页。
npx @vectorize-io/hindsight-coding-agents install all # 为所有检测到的智能体安装,原生连接
npx @vectorize-io/hindsight-coding-agents install claude-code # 或只安装一个
支持 Claude Code、Codex CLI、Cursor CLI、GitHub Copilot CLI、opencode、Kilo CLI、Cline CLI、Antigravity CLI、Devin CLI、pi、Prime Agent、Grok Build 和 DeepSeek Harness。
摄取是自动的——无需设置命令。
请参阅编程智能体集成 (https://hindsight.vectorize.io/sdks/integrations/coding-agents)。
MCP 服务器
每个服务器都内置一个 Model Context Protocol (https://modelcontextprotocol.io/) 端点,每个记忆库一个,默认启用:
http://localhost:8888/mcp/{bank_id}/
将任何 MCP 客户端指向它,即可将保存、召回和反思作为工具暴露。
请参阅 MCP 服务器文档 (https://hindsight.vectorize.io/developer/mcp-server)。
核心概念
概览
记忆类型
大多数智能体记忆实现依赖于基本的向量搜索,有时使用知识图谱。Hindsight 使用仿生数据结构来组织智能体记忆,其方式更类似于人类记忆的工作方式:
- 世界事实: 关于世界的事实(“炉子会变热”)
- 经验: 智能体自身的经验(“我碰了炉子,真的很疼”)
- 观察: 从许多记忆中形成的、有证据支持的、整合的信念
- 心智模型: 从观察和事实中综合而得的、对智能体世界的习得性理解
记忆存储在记忆库中。当添加记忆时,它们会被推送到世界事实或经验通路,然后表示为实体、关系和时间序列的组合,并辅以稀疏/密集向量表示,以帮助后续召回。
三个核心操作
保存
retain 操作用于将新记忆推送到 Hindsight。它告诉 Hindsight 保留你作为输入传递的信息。
client.retain(
bank_id="my-bank",
content="Alice got promoted to senior engineer",
context="career update",
timestamp="2025-06-15T10:00:00Z",
)
在幕后,保存使用 LLM 提取关键事实、时间数据、实体和关系。这些数据会经过规范化过程,转换为规范实体、时间序列和搜索索引以及元数据。这些表示为在召回和反思操作中准确检索记忆创造了途径。
保存文档 → (https://hindsight.vectorize.io/developer/retain)
召回
recall 操作用于检索记忆。这些记忆可以来自任何记忆类型(世界、经验等)。
client.recall(bank_id="my-bank", query="What does Alice do?")
client.recall(bank_id="my-bank", query="What happened in June?") # 时间相关
召回并行执行 4 种检索策略:
- 语义:向量相似度
- 关键词:BM25 精确匹配
- 图谱:实体/时间/因果链接
- 时间:时间范围过滤
单独的结果被合并,使用倒数排名融合和交叉编码器重排模型按相关性排序,然后根据需要修剪以符合令牌限制。
召回文档 → (https://hindsight.vectorize.io/developer/retrieval)
反思
reflect 操作对现有记忆进行更彻底的分析。这使得智能体能够在记忆之间形成新的联系,并建立对其世界更透彻的理解——或者回答需要深度思考而非简单查找的问题。
client.reflect(bank_id="my-bank", query="What should I know about Alice?")
例如,反思支持以下使用场景:
- 一个 AI 项目
相似文章
rohitg00/agentmemory
agentmemory 是一个开源的持久化记忆层,专为 AI 编程智能体(Claude Code、Cursor、Gemini CLI、Codex CLI 等)设计。它通过知识图谱、置信度评分和混合搜索技术,借助 MCP、Hooks 或 REST API,为智能体提供跨会话的长期记忆能力。该项目基于 iii 引擎构建,无需外部数据库,提供 51 个 MCP 工具。
如何在不耗尽你全部显存的情况下为AI智能体提供长期记忆 (Hillock v0.5)
Hillock v0.5.0 是一个轻量级的神经符号记忆引擎,专为AI智能体设计。它通过SQLite中的结构化三元组和超向量高效管理持久化记忆,旨在为显存有限的本地环境提供服务。
@_avichawla:为你的Agent构建类人记忆(开源)!每个智能体和RAG系统在实时知识方面都面临挑战……
Graphiti是一个开源工具,通过持续演进且具有时间感知的知识图谱,为AI代理构建类人记忆,相比MemGPT,准确率最高提升18.5%,延迟降低90%。
activeloopai/hivemind
Hivemind 是 Activeloop/Deeplake 推出的开源工具,为 AI 编程代理提供自动学习、云支持的共享内存。它捕获痕迹、编码模式并在代理间传播技能,在 LoCoMo 基准测试上实现 25% 的成本节省和 1.7 倍的令牌减少。
@yoheinakajima: https://x.com/yoheinakajima/status/2081741659260477666
该线程探讨了大脑的双重记忆系统(海马体和新皮层)为构建长期运行的AI智能体提供的启示,指出智能体需要快速的情景捕获机制和缓慢的巩固机制来避免灾难性干扰,而不是仅仅依赖带有临时支撑结构的冻结模型。