Tag
KernelZero introduces a co-evolution framework with Proposer and Coder models to enhance GPU kernel generation, achieving superior performance on CUDA and Triton benchmarks.
MACE is a memory-agent co-evolution framework for LLM-based multi-agent systems that uses adaptive memory graphs to organize functional memory units and adapt through execution feedback, improving task performance over baselines.
This research explores combining harness evolution with model adaptation for AI agents, discovering that direct imitation from experts harms weaker models' performance and proposing an on-policy correction method to improve performance without breaking harness fit for enterprise tasks.
J-Zero is a unified framework for co-evolving Challenger, Solver, and Judge models from zero data, enabling self-improvement in language models across both verifiable and unverifiable domains with performance surpassing baselines.
EnvHarness introduces a programmable layer to dynamically reshape static environments for reinforcement learning, improving agent performance through automated targeting of weaknesses with EnvRigger.
This paper proposes model-harness co-evolution as a fundamental principle for recursive self-improvement in AI agents, introducing HELIX, a source-traceable system that improves both execution and learning by generating structured training signals from verified trajectories.
Researchers propose the Red Queen Gödel Machine, a framework for self-improving AI where agents and evaluators evolve together to overcome evaluation ceilings, showing enhanced performance in tasks like scientific paper writing and grading.
This paper introduces PIAC, a framework that improves LLM-based automatic construction of parallel algorithm portfolios by using a potential-gain metric that eliminates the need for reference solutions and by leveraging LLMs to generate diverse instance mutators. It consistently outperforms existing LLM-ACP baselines on TSP and CVRP, achieving up to 19.76% relative improvement.
A survey paper proposing a three-stage taxonomy of co-evolution in agentic systems, covering agent-agent, agent-environment, and meta co-evolution to enable open-ended improvement beyond fixed human-designed paths.
This paper proposes scaffold-mediated post-training, a paradigm where procedural scaffolds co-evolve with LLM parameters through discovery, distillation, and dynamic recompilation. On FeatureBench, automatically discovered skills improve pass rate by 8.1pp, with a 27.7% pass rate after distillation.
This paper introduces CoCoEvolve, a self-supervised method that improves cross-representation understanding across charts, tables, and code by enforcing one-to-one consistency between representations, with training-time and test-time co-evolution objectives.
Echoverse presents a method for generating deep, evolving synthetic environments to train computer-use agents, demonstrating substantial accuracy gains and releasing a benchmark with grounded graders.
Introduces DecoEvo, a score-decoupled co-evolution method for LLM optimization in text space that jointly improves solver and rubric-generator skills without gold rubrics, achieving 2.8–5.0% relative gains over baselines across five benchmarks.
EvoSQL is a co-evolution framework for Text-to-SQL that iteratively improves SQL generation via a generator-critic pair with episodic memory, achieving gains on Spider and BIRD benchmarks.
The paper introduces Skill Self-Play (Skill-SP), a co-evolutionary framework that uses a proposer, solver, and skill controller to bridge structured verification and open-ended exploration, improving LLM performance on tool-use and reasoning benchmarks.
MSCE is a training-free framework that organizes LLM agent experience into three memory levels and converts them into reusable skills with evidence links, outperforming existing memory and skill-augmented baselines.
Introduces EvolvingWorld, an open-schema framework and benchmark for co-evolving role-play agents and world models in interactive literary worlds, enabling long-horizon simulation with persistent character and world state updates.
This paper proposes a method for co-evolving evaluation metrics and skills in self-improving LLM agent systems, demonstrating that metrics can be evolved and that a co-evolution approach recovers most of the performance of a ground-truth-driven oracle across code generation, text-to-SQL, and report generation tasks.
HASE is a reinforcement-learning framework that co-evolves model weights, task solutions, and harness components (guidance and evaluation) in a unified agentic process, enabling a single 8B-parameter model to match the performance of much larger systems on text classification and alpha factor mining tasks.
The Red Queen Gödel Machine enables recursive self-improvement in AI by co-evolving the agent and evaluator, achieving better coding performance with fewer tokens.