Tag
EpiCon presents a shared multimodal memory framework for collective learning among AI agents, enhancing performance across eleven benchmarks without updating host model parameters.
The paper presents ACLArena, a framework for evaluating Agent Continual Learning in multi-stage post-training, analyzing forgetting and generalization mechanisms, and proposing an improved ACL recipe using offline replay and LoRA experts.
ScienceIDE introduces infrastructure for converting scientific code repositories into programmable environments for scientific agents, enabling task generation, execution, and verification, with trained models showing improvements in scientific code repair and general capabilities.
This paper introduces a framework for constructing verified synthetic web environments to improve the training of web agents, demonstrating enhanced performance and transferability across benchmarks.
EnvHarness introduces a programmable layer to dynamically reshape static environments for reinforcement learning, improving agent performance through automated targeting of weaknesses with EnvRigger.
The author shares a deterministic learning harness that lets multi-agent systems improve across episodes without fine-tuning or prompt edits, by promoting successful strategies into persistent playbooks. On the Mini Amusement Park benchmark, reward improved from 12,121 to 483,019, reaching #1 on the leaderboard.
Proposes Experience Distillation, a method that internalizes in-context learning gains from agent interaction histories into model weights without requiring additional environment interaction, achieving significant sample efficiency improvements on software engineering and text-adventure tasks.
Hermes Agent by Nous Research introduces /learn, a command that lets the agent deliberately create skills from documentation, code, or instructions without needing to first fail at the task, turning any source into a reusable skill.