Tag
Professor Li Hongyi from National Taiwan University has updated the 2026 spring machine learning course, focusing on AI Agents and model self-evolution in the large model era. This marks a complete transition of the course from classical machine learning to cutting-edge large model technologies.
This paper introduces VQS, a method for self-evolving vision-language models that uses program verification to generate and label questions from images, significantly improving accuracy and model performance over existing baselines.
TimeEvo is a failure-driven self-evolution method for time series agents that diagnoses capability gaps, synthesizes tools, and improves accuracy across tasks and backbones.
This article explores verifying the genuine improvement of AI agent self-evolution through evaluation frameworks like Meta-Harness, emphasizing the separation of powers among agent modification, evaluation, and deployment decisions to prevent false progress.
The paper introduces SELF-INDEX, a framework that enables search indexes to self-evolve autonomously, improving retrieval performance and benefiting downstream applications such as search agents and agent memory systems.
The article proposes a method for continual learning in AI where the model autonomously decides what to learn and updates its weights in real-time, potentially enabling continuous self-evolution towards AGI.
This paper proposes a four-layer architecture for cognitive digital twins that enables self-evolving operational loops through integrated cognitive capabilities and task-oriented feedback mechanisms.
Procedural Graphs presents a framework for LLM agents to self-evolve their execution structures, enhancing adaptability and efficiency in autonomous task completion.
The paper introduces reSolve, a surrogate-guided solve-and-reproduce framework for self-evolving agent skills that achieves 74.9% performance, surpassing human-curated baselines by 14.8 points.
StudyBench introduces a controlled physics benchmark to measure how efficiently self-evolution methods convert training material into transferable problem-solving ability, revealing gaps in guidance and compute.
HarnessDev evaluates LLMs by their ability to build and evolve execution harnesses, revealing significant variations in performance and poor transferability across models.
DiagEvo improves language-model self-evolution by deriving training direction from internal failure history via hierarchical error-cause memory and double-confidence filtering, outperforming baselines that rely on external resources.
ASPIRE introduces a benchmark for self-evolving LLM agents from vague natural-language goals, revealing challenges in goal interpretation and stable weight-level improvement.
EvoUndo introduces a framework for evaluating and ensuring recoverability in self-modifying LLM agents, showing that reliable recovery requires co-designing verification, state grounding, and recovery language expressivity.
This article discusses the importance of trajectory learning in AI agent self-evolution and introduces the Apodex 1.1 system and open-source tool FrontierAgent for executing and evaluating long-duration research tasks.
RubSE is a framework that uses rubric-guided self-evolution to enhance the stability of UI-to-code generation by mitigating visual repair coupling and trajectory collapse.
This paper audits self-evolution mechanisms in financial AI agents, revealing capability improvements alongside security risks such as prompt injection drift and execution-interface mismatches, emphasizing the need for holistic auditing.
GenRouter is a unified routing framework for agentic image generation that adaptively directs prompts to optimal workflows, significantly reducing costs and latency while improving visual alignment through demand profiling and self-evolution.
This paper introduces SHAPER, a self-evolving framework for embodied agents that keeps model parameters frozen and improves performance by evolving reusable skills and context-code harnesses through target-environment rollouts. Evaluated on VLABench and ESI-Bench, it proposes a practical alternative to fine-tuning when training is expensive or unavailable.
This paper introduces Spatial Memory Agent (SMA), a runtime framework that improves frozen vision-language models' spatial reasoning through verifier-guided reflection and reusable memory without parameter updates or external tools, achieving strong results across five benchmarks and four base VLMs.