Tag
本文基于PyCon China 2026的演讲,讨论如何设计AI代理的组成、运行环境和生命周期。
The article critiques human-in-the-loop systems for leading to disengaged approval and advocates for scoped authority in AI agent design, emphasizing evidence and enforcement over routine clicks.
The author investigates design challenges for AI agents to manage tool budgets effectively, focusing on preventing wasteful retries and handling API failures through budget-aware strategies.
The blog post introduces Maka's Design Blog, which discusses practices and designs from an infrastructure architect, particularly relevant for AI agents. It highlights the idea that agents are replayable histories rather than processes, with logs serving as the runtime.
The article discusses how persistent memory in AI agents can amplify prompt injection risks by storing hostile instructions as trusted context, and explores a design to mitigate this while acknowledging limitations and the need for further testing over extended periods.
The article argues that multi-agent systems are often overused in AI applications, suggesting that a single agent with good tools, strict state, and clear stop conditions can be more efficient, easier to debug, and cost-effective for many workflows.
This article discusses the relationship between Graph Engineers and Loop Engineers, emphasizing that a Loop is the smallest Graph, and references research from 'Nature Machine Intelligence' to analyze the applicable scenarios of multi-agent systems.
The article discusses a research paper where LLMs used to write and review a trading feature missed a future-data bug, highlighting the need for structural redesigns in agent systems to prevent such issues.
The article discusses the importance of system design for AI agents, covering concepts like Agent Harness, LLMOps, and Evals, and provides a proof-of-concept implementation with plans for future parts.
Argues that approval gates for AI agents should be based on reversibility rather than fear, suggesting that building undo mechanisms can eliminate the need for many human-in-the-loop checks.
The article discusses a failure mode where LLM agents following prompt-level sequential loop instructions can silently skip items at production scale, and recommends an orchestrator/worker architecture with platform-level batch dispatch to guarantee every item is processed.
The author comments on Deepseek Harness's plugin-based philosophy, arguing that it turns everything (GUI/TUI, etc.) into plugins and opposes hardcoding workflows. They compare it with Pi Agent and Prime-Agent, propose the idea of building a self-evolving Agent in a REPL, and have already started implementing it with Codex.
An article exploring the origins and core components of loop engineering, the practice of designing autonomous loops that prompt AI coding agents instead of prompting them one message at a time, highlighting the importance of objective gates and state.
The author argues that adding an undo button—not new capabilities—unlocked experimentation with their AI agent, suggesting agent design is really about reducing the cost of reverting changes.
This post explains the layered relationship between prompt, context, harness, loop, and graph engineering, emphasizing that each layer builds on the previous one rather than replacing it, and how to identify which layer to debug.
A practical note on AI agent reliability, arguing that production agents need explicit gates for evidence thresholds, retry budgets, and impact assessment rather than relying on memory alone to determine task completion.
This paper presents the first systematic exploration of filesystem-based memory for LLM agents, formalizing roles of management, search, and execution agents around a shared memory store. It finds that organization primarily reduces retrieval cost but does not yet improve answer quality, and that tooling choices affect store shape as much as model selection.
This article discusses a technique for reducing context bloat in AI agents by decoupling the planning phase from execution, improving agent efficiency and accuracy.
Andrew Ng shifts the focus from whether a system is an agent to how much autonomy it has, recommending building agentic workflows with deliberate autonomy levels per task rather than full autonomy.
This article explores the tradeoffs between stateful and stateless agent design for scalable AI systems, providing implementation examples using Groq API and Llama 3.1 8B Instant model.