PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
Summary
Introduces PAST-Bench, a benchmark for evaluating whether personal AI agents improve from retained experience across sessions, and Hermes+, an extension with targeted interventions. Finds improvement is real but uneven across capabilities and models.
View Cached Full Text
Cached at: 08/05/26, 05:43 AM
Paper page - PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
Source: https://huggingface.co/papers/2608.04003
Abstract
Recursiveself-improvementrequiresagentstoturnaccumulatedexperienceintobetterfuturebehavior.PersonalAIagentsofferaconcretesettingforstudyingthiscapabilitybecausetheyretainpreferences,taskhistories,toolroutines,andlearnedskillsacrosssessions.Yetwhetherretainedexperienceactuallyimprovesthemovertimehasnotbeensystematicallytested.WeintroducePAST-Bench,abenchmarkdesignedtoisolatethisquestion.Eachagentrunsthroughorderedsequencesoffresh-sessiontasksundermatchedconditionsthatturnretainedexperienceonandoff.Itspans26scenariosand204episodesacrossmemory,proceduralreuse,informationgathering,andupdate.Wereportbothlater-taskgainsandwhetherthosegainsfollowtheintendedsave,retrieve,andupdatepathway.Acrosssevenbasemodelsandfouragentframeworks,improvementisrealbutunevenacrosscapabilities.Agentswiththesameheadlinegaincandiffermarkedlyinwhetherthatgainissupportedbyevidenceoftheintendedpathway.Guidedbythesefindings,wedevelopHermes+,whichextendsHermeswithfivetargetedinterventionsacrossstagesoftheagentloop.Hermes+raisestheaveragegainfromretainedexperienceandprovidesclearerpathwayevidence,withitsstrongestimprovementontasksrequiringoutdatedstatetobereplaced,althoughtheeffectremainscapability-andmodel-dependent.Together,PAST-BenchandHermes+provideanevaluationanddiagnosticfoundationforstudyinghowpersistentagentscanprogressfromretainingexperiencetosystematicallyimprovingthroughit.Code:https://github.com/Gen-Verse/PAST-Bench
View arXiv pageView PDFGitHub2Add to collection
Get this paper in your agent:
hf papers read 2608\.04003
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.04003 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.04003 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.04003 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning
AhaBench is a benchmark suite that evaluates whether language agents improve from prior experience in long-horizon tasks by testing exploration, knowledge transfer, and delayed feedback handling.
@EinsiaAI: 1/ Recursive self-improvement (RSI) depends on agents improving how AI systems are trained —not just tuning hyperparame…
The article presents AI4AI-Bench, a benchmark evaluating AI agents' ability to improve training algorithms, showing low performance scores and high exploration costs across ten research repositories.
AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?
AutoDataBench introduces a benchmark to evaluate if agents can write tasks for data pipelines that meet practical acceptance standards, aiming to enable scalable data synthesis for recursive self-improvement in AI.
FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents
Introduces FinPersona-Bench, a benchmark to measure how well autonomous financial agents maintain their assigned behavioral mandates over time, revealing Mandate Salience Decay (MSD) that worsens with temporal distance and varies by model and agent profile.
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
This paper introduces HealthAgentBench, a suite of 54 realistic healthcare tasks for evaluating frontier AI agents. It finds that even the best agent (Codex GPT-5.5) achieves only ~42% success, highlighting substantial room for improvement.