@CycleDecoded: 搞自动化和爬虫的兄弟可以把之前的方案扔了。 GitHub 上突然爆火了一个开源项目 Index(由 AI 开发者平台 Laminar 团队打造),直接把网页浏览器变成了可调用的 API。这玩意本质上是一个极度丝滑的“AI 浏览器 Agen…
摘要
Index 是一个由 Laminar 团队开发的开源 AI 浏览器 Agent,可将任何网站转化为可调用的 API,支持 Claude、Gemini 等视觉模型,在 WebVoyager 上准确率达 92%。
查看缓存全文
缓存时间: 2026/08/03 19:49
搞自动化和爬虫的兄弟可以把之前的方案扔了。
GitHub 上突然爆火了一个开源项目 Index(由 AI 开发者平台 Laminar 团队打造),直接把网页浏览器变成了可调用的 API。这玩意本质上是一个极度丝滑的“AI 浏览器 Agent”,你在终端输一句人话,它就能像真人在浏览器里点按、抓数据、填表单,甚至跨站点联动办完一整套复杂流程。在 WebVoyager 跑分里直接飙到了 92% 的惊人准确率。
一句命令就能跑:pip install 之后敲 index run 就能直接在终端调用。
能直接用本地 Chrome:带 –local-chrome 参数,自动复用你已登录的账号状态。
顶配视觉推理引擎:原生支持 Claude 3.7 Sonnet、Gemini 2.5 Pro 等顶尖多模态大模型。
全程会话回放与排错:自带开箱即用的可视化录屏与步骤 Trace 监控。
自动抓取并生成表格:一句话搞定“去网站抓数据并新建 Google Sheets 写入”。
开源协议:Apache-2.0 GitHub 传送门:https://github.com/lmnr-ai/index
lmnr-ai/index
Source: https://github.com/lmnr-ai/index
Index
Index is a state-of-the-art open-source browser agent that autonomously executes complex web tasks. It turns any website into an accessible API and can be seamlessly integrated with just a few lines of code.
-
Powered by reasoning LLMs with vision capabilities.
- Gemini 2.5 Pro (really fast and accurate)
- Claude 3.7 Sonnet with extended thinking (reliable and accurate)
- OpenAI o4-mini (depending on the reasoning effort, provides good balance between speed, cost and accuracy)
- Gemini 2.5 Flash (really fast, cheap, and good for less complex tasks)
-
pip install lmnr-indexand use it in your project -
index runto run the agent in the interactive CLI - Supports structured output via Pydantic schemas for reliable data extraction.
- Index is also available as a serverless API.
- You can also try out Index via Chat UI.
- Supports advanced browser agent observability powered by open-source platform Laminar.
prompt: go to ycombinator.com. summarize first 3 companies in the W25 batch and make new spreadsheet in google sheets.
https://github.com/user-attachments/assets/2b46ee20-81b6-4188-92fb-4d97fe0b3d6a
Documentation
Check out full documentation here
Quickstart
Install dependencies
pip install lmnr-index 'lmnr[all]'
# Install playwright
playwright install chromium
Setup model API keys
Setup your model API keys in .env file in your project root:
GEMINI_API_KEY=
ANTHROPIC_API_KEY=
OPENAI_API_KEY=
# Optional, to trace the agent's actions and record browser session
LMNR_PROJECT_API_KEY=
Run Index with code
import asyncio
from index import Agent, GeminiProvider
from pydantic import BaseModel
from lmnr import Laminar
import os
# to trace the agent's actions and record browser session
Laminar.initialize()
# Define Pydantic schema for structured output
class NewsSummary(BaseModel):
title: str
summary: str
async def main():
llm = GeminiProvider(model="gemini-2.5-pro-preview-05-06")
agent = Agent(llm=llm)
# Example of getting structured output
output = await agent.run(
prompt="Navigate to news.ycombinator.com, find a post about AI, extract its title and provide a concise summary.",
output_model=NewsSummary
)
summary = NewsSummary.model_validate(output.result.content)
print(f"Title: {summary.title}")
print(f"Summary: {summary.summary}")
if __name__ == "__main__":
asyncio.run(main())
Run Index with CLI
Index CLI features:
- Browser state persistence between sessions
- Follow-up messages with support for “give human control” action
- Real-time streaming updates
- Beautiful terminal UI using Textual
You can run Index CLI with the following command.
index run
Output will look like this:
Loaded existing browser state
╭───────────────────── Interactive Mode ─────────────────────╮
│ Index Browser Agent Interactive Mode │
│ Type your message and press Enter. The agent will respond. │
│ Press Ctrl+C to exit. │
╰────────────────────────────────────────────────────────────╯
Choose an LLM model:
1. Gemini 2.5 Flash
2. Claude 3.7 Sonnet
3. OpenAI o4-mini
Select model [1/2] (1): 3
Using OpenAI model: o4-mini
Loaded existing browser state
Your message: go to lmnr.ai, summarize pricing page
Agent is working...
Step 1: Opening lmnr.ai
Step 2: Opening Pricing page
Step 3: Scrolling for more pricing details
Step 4: Scrolling back up to view pricing tiers
Step 5: Provided concise summary of the three pricing tiers
Running CLI with a personal Chrome instance
You can use Index with personal Chrome browser instance instead of launching a new browser. Main advantage is that all your existing logged-in sessions will be available.
# Basic usage with default Chrome path
index run --local-chrome
Use Index via API
The easiest way to use Index in production is with serverless API. Index API manages remote browser sessions, agent infrastructure and browser observability. To get started, create a project API key in Laminar.
Install Laminar
pip install lmnr
Use Index via API
from lmnr import Laminar, LaminarClient
# you can also set LMNR_PROJECT_API_KEY environment variable
# Initialize tracing
Laminar.initialize(project_api_key="your_api_key")
# Initialize the client
client = LaminarClient(project_api_key="your_api_key")
for chunk in client.agent.run(
stream=True,
model_provider="gemini",
model="gemini-2.5-pro-preview-05-06",
prompt="Navigate to news.ycombinator.com, find a post about AI, and summarize it"
):
print(chunk)
Browser agent observability
Both code run and API run provide advanced browser observability. To trace Index agent’s actions and record browser session you simply need to initialize Laminar tracing before running the agent.
from lmnr import Laminar
Laminar.initialize(project_api_key="your_api_key")
Then you will get full observability on the agent’s actions synced with the browser session in the Laminar platform. Learn more about browser agent observability in the documentation.
Made with ❤️ by the Laminar team
相似文章
@CycleDecoded: 搞 AI Agent 和自动化爬虫的兄弟们直接看,这玩意简直是“把全网网站给 LLM 当饭吃”的神器。 以前想给 AI Feeding 动态网页或者抓数据,要折腾 Puppeteer、搞动态代理、解 JavaScript,还得调 API …
介绍开源项目 Firecrawl:一个能将任意网址转换为干净 Markdown/JSON、支持 AI 交互和整站爬取的网页数据 API,专为 LLM 和 Agent 设计,已获 2.5w+ GitHub Star。
@GitHub_Daily: 让 AI Agent 自动化操作浏览器或抓数据,经常被各种反爬机制拦截,遇到验证码、人机验证直接卡死。 最近 BrowserAct 团队开源了一个 Skill,专为 AI Agent 设计的浏览器自动化命令行工具。 提供三层反封锁机制,从…
BrowserAct 团队开源了一个专为 AI Agent 设计的浏览器自动化命令行工具,提供三层反封锁机制(指纹伪装、验证码破解、人类接管),支持多浏览器并行、账户隔离,并优化了输出格式以节省Token。
@CycleDecoded: 一个叫 WebRover 的开源 AI Agent,你给它一句话,它自己打开浏览器、识别网页元素、点击、翻页、抓数据、搞定任务,最后把你要的结果整理得明明白白。 开源协议:MIT License 定位:自主式 Web 自动化 AI 智能体…
WebRover 是一个基于 MIT 协议的开源 AI Agent,通过自然语言即可驱动浏览器完成网页自动化、跨站数据抓取与深度研究,支持本地部署。
@Jolyne_AI: 开源 AI 网页自动化工具:Nanobrowser。 OpenAI Operator 的开源替代方案,本地在浏览器里运行,支持多智能体协作。 免费、重视隐私、LLM 选择灵活、代码完全开源,让网页操作更智能、更高效。 GitHub:htt…
Nanobrowser 是一个开源 AI 网页自动化工具,作为 OpenAI Operator 的免费替代方案,在本地浏览器中运行,支持多智能体协作,注重隐私且 LLM 选择灵活。
@GoJun315: 一位 16 岁开发者,开源了一个无头浏览器引擎,专为爬虫和 AI Agent 自动化设计。 项目名叫 Obscura,使用 Rust 构建,已狂揽 14600+ GitHub Star。 与 headless Chrome 对比优势明显:…
一位16岁开发者开源了基于Rust的无头浏览器引擎Obscura,专为爬虫和AI Agent自动化设计,内存占用仅30MB,已获得超14600 GitHub星标。