Tag
The author shares their experience building AI agent stacks with one smart AI component for reasoning and reliable deterministic tools to prevent failures, emphasizing the value of minimizing model discretion in mundane tasks.
Xing4.0-29B-A4B is a next-generation MoE large language model developed by China Telecom, featuring 29B total parameters with 4B active per token, native support for 256K context length, and optimization for Ascend NPU with agent-oriented architecture for complex engineering tasks.
Xing4.0-29B-A4B is an open-source 29B-parameter large language model optimized for agent tasks and Ascend NPU, featuring a MoE architecture and achieving high training efficiency with competitive benchmark results.
An individual experiments with adding an explicit architecture layer to coding agents, building an open-source agent harness to test the idea, and discusses potential tradeoffs in agent design.
The paper introduces Recuris, a recursive memory architecture that improves long-horizon agent success by tracking progress and guiding skill selection through localized, validation-gated updates.
The article explains how using loops for automated checks and graphs for workflow optimization can reduce manual oversight in managing AI agents.
The article discusses the problem of AI agents that fail silently by not escalating when stuck, and suggests implementing explicit checks or circuit breakers to improve reliability in production.
monday.com rebuilt their AI copilot Sidekick by evolving from a single general-purpose agent to a system with specialized subagents and LangSmith Sandboxes, enhancing its capability to handle complex, iterative work in production.
An essay arguing that typical LLM-based agent pipelines built as LLM → tool → action lack proper uncertainty handling, and proposing a belief-state architecture with Bayesian updates and value-of-information policies. The LLM should act as investigator/translator, while the system enforces permissions and calibrated beliefs.
Explores how multi-layered AI agent delegation affects reliability, arguing that the downstream influence of errors matters more than the number of layers.
The author outlines four patterns for how AI agents handle external email replies (human-in-the-middle, no-reply outbound, shared inbox, agent-owned address), asking which approach readers use and whether no-reply agents have incurred real costs. They note their work at Atomic Mail, an email service built for AI agents.
The article presents an agent architecture principle where arguments are looked up along a provenance chain (user_answer, instruction, pre_set_data, measured_data, prior_state) rather than generated, treating unknown fields as a valid state that triggers asking the user.
A senior Anthropic engineer published a 12-page PDF detailing a graph engineering approach for multiagent systems, using knowledge graphs as persistent shared memory to overcome context window limitations.
The article outlines four essential boundaries (identity, intent, policy/execution, system-of-record) to ensure LLMs never have direct access to production systems, preventing irreversible actions. It emphasizes the need for multi-layered permission checks and human approval for risky operations.
Andrew Ng released an 8-page PDF detailing four key agentic workflows: reflection, tool use, planning, and multi-agent collaboration, emphasizing that a weak model with proper architecture can outperform a strong one.
This article breaks down the five infrastructure layers required to run production web agents beyond just a browser, covering warm pools, isolation, identity, observability, and model gateways, and discusses when it makes sense to build vs. buy.
This article analyzes the industry shift from single-loop to graph-based self-improvement architectures in AI agents, explaining why optimizing a single metric often fails and how a network of improvement cycles provides a more robust solution.
The author reflects on building a local AI agent capable of deleting files and concludes that safety gates are more critical than agent autonomy.
This blog post by Ai2 describes the architecture and lessons learned from building Shippy, a maritime AI agent for real-time domain awareness, covering aspects like agent anatomy (soul, skills, config), deterministic tools, sandboxed hosting, and evaluation.
A blog post arguing that using a frontier model only for planning and a cheaper model for execution is not cost-effective because reading—not editing—is the primary cost driver; duplicate reading offsets any savings.