Tag
A Google paper introduces Procedural Graphs, an editable workflow structure for LLM agents that improves performance on long tasks by evolving from execution feedback, outperforming baselines in most benchmarks.
This position paper proposes a Foundation Model Operating System (FMOS) to virtualize foundation model interactions, providing applications with dedicated, trustworthy instances and enabling self-evolving capabilities through orchestration and policy enforcement.
EvoOntology introduces a self-evolving ontology layer for data agents, encapsulated as an MCP server, to bridge the agent-data gap and improve performance on heterogeneous data tasks as shown in benchmarks.
This paper introduces Procedural Graph, a framework that organizes LLM agent actions into structured triplets for improved long-horizon tool use, with self-evolving topology to enhance performance.
Enhanced the TUI of ouroboros, a CLI tool for self-evolving auto-research systems, and deleted 2.79 million lines of vestigial code to improve usability.
CHIME is a credit-aware hierarchical memory framework that separates planning and execution memory banks to improve long-horizon agentic planning by accurately attributing task outcomes and outperforming baselines.
J-Zero is a unified framework for co-evolving Challenger, Solver, and Judge models from zero data, enabling self-improvement in language models across both verifiable and unverifiable domains with performance surpassing baselines.
JIT-Agent is a trainable model that synthesizes adaptive agent harnesses for off-the-shelf LLMs, improving performance across diverse models and tasks.
This paper surveys agent memory in AI agents, addressing the problem of context explosion and proposing a framework for self-evolving agents through various memory types and management strategies.
This article comments on the low barrier to entry for autonomous-evolving code agent tools, drawing an analogy to stand-up comedy performances, pointing out that in the current code agent field, everyone can claim to be skilled at complex tasks.
Introduces ERSkill, a retrieval-centric framework for self-evolving, skill-guided adaptive memory access in LLM agents. It co-evolves retrieval skills and a routing policy, substantially outperforming strong baselines across agent memory benchmarks.
The article introduces MindMemOS, a portable and self-evolving memory operating layer for AI agents that uses a unified entity-property-time structure, with algorithms for memory refinement and skill evolution. It achieves notable accuracy on LOCOMO and PersonaMem benchmarks and improves SpreadsheetBench performance by 9.2 percentage points.
The author comments on Deepseek Harness's plugin-based philosophy, arguing that it turns everything (GUI/TUI, etc.) into plugins and opposes hardcoding workflows. They compare it with Pi Agent and Prime-Agent, propose the idea of building a self-evolving Agent in a REPL, and have already started implementing it with Codex.
MEGA is a self-evolving infrastructure for coding-agent optimization that distills reusable wisdom from sessions, composes it via a typed Wisdom Graph, and uses operational evidence to continuously improve both the agents and the knowledge guiding their optimization.
GeoForge is a training-free, self-evolving framework for Earth-observation reasoning that structures completed trajectories into nonparametric memories to improve LLM agent planning and tool-use without updating the backbone model.
LinkedIn presents a self-evolving agentic customer support system that integrates RAG with evolutionary auto-prompting and modular evaluation, achieving significant gains in production A/B tests including a 9.0-point increase in QA self-serve and 30.6-point improvement in routing accuracy.
Presents NeSy-Spatial, a neuro-symbolic framework that self-evolves spatial reasoning skills by composing tool-use and geometry skills, improving accuracy on spatial reasoning benchmarks.
BigBang-v1 is a self-evolving 36B LLM from Endless Frontier Lab in Shanghai, trained with AI-generated frontier tasks and achieving strong performance with only 10K high-quality examples across science, coding, tool use, and long context.
Introduces EvoHarness-RL, a framework that learns runtime harness policies for long-horizon LLM agents, enabling them to construct and update external state (belief, progress, experience) during task execution. Using Qwen3-8B on ALFWorld, it achieves 96.9% success and reveals harness annealing and evolution dynamics.
A tweet expressing amazement at the concept of dynamic agent orgs—self-evolving multi-agent systems where the graph structure rewrites itself during execution.