Tag
Paper2Agent, a multi-agent framework published in Nature, automatically transforms research papers into virtual authors, enabling AI agents to interact with and build upon scientific knowledge for enhanced discovery.
RSIAgent is a training-free multi-agent framework that enables digital agents to adapt to new environments through recursive self-improvement, autonomous memory construction, and broad-then-deep exploration, outperforming closed-source models on benchmarks.
This paper presents a benchmark (DNBench) for evaluating LLMs in database normalization tasks and introduces a multi-agent reasoning framework (MARS) that improves accuracy by 82% over single-prompt methods.
The paper introduces PACE, a dataset for evaluating whether AI models can identify hidden conflicts in user requests by retrieving implicit knowledge base facts, and proposes PaceMaker, a multi-agent framework to enhance conflict-aware decision-making.
The article describes GenOS, a custom multi-agent framework that autonomously evolved Rust algorithms to solve the NP-Hard Reverse Game of Life problem, discovering three optimization paradigms and proving a mathematical limit of 378/400.
HERMES is a scalable multi-agent framework for extracting structured knowledge from ultra-long scientific documents in geoscience, achieving high accuracy and sixfold efficiency improvement over manual methods.
This paper investigates quality issues in LLM-generated answers for hardware description language questions, finding over-answering tendencies like redundancy (65.7%) and verbosity (69.1%), and proposes a multi-agent framework that reduces core answers by 37% and non-core content length by 31% while improving quality scores.
Introduces an automated prompt optimization framework for LLM game agents that decomposes the observation-to-action pipeline into two agents and iteratively refines prompts via an evolutionary loop guided by environment returns. Evaluated on BabyAI tasks, it significantly improves success rates (e.g., from 0% to 72.5% on PutNext) without updating model weights.
PRISM is a multi-agent framework that decouples speech perception, response generation, and speech synthesis to improve empathetic spoken dialogue by integrating prosodic cues with LLM reasoning and external knowledge tools.
Visual Para-Thinker++ proposes a single-policy multi-agent framework for visual reasoning that uses role-conditioned agents (Main, Worker, Summary) and dedicated training methods to reduce hallucinations and improve efficiency, outperforming baselines on hallucination-sensitive benchmarks.
This paper presents AuditFlow, a graph-grounded multi-agent framework that uses executable symbolic environments for structured financial reporting verification, achieving 82.09% audit accuracy on a FinAuditing-derived sample under GPT-5.5.