Tag
This paper presents a benchmark and a multi-agent system to improve scientific diagram generation through multi-turn human feedback, addressing issues like quality drift and forgetting across revisions.
The Station is an open-world multi-agent environment where AI agents autonomously build a scientific literature, leading to novel discoveries on five mathematical problems out of twelve evaluated.
MARS is a multi-agent LLM framework using specialized agents for different algorithmic topics to solve competitive programming problems, achieving a 0.624 pass rate with lower cost and variance than direct prompting.
This paper proposes a multi-agent framework that enables LLM agents to conduct controlled experiments using simulation models for pharmaceutical process design, yielding more specific and actionable recommendations than language-only reasoning.
The project Munder Difflin integrates multiple AI Agents into a virtual office, achieving task allocation, collaboration, and automated management through a manager Agent.
This article provides a practical architecture for building a 24/7 multi-agent system with Grok Bot, detailing steps from individual bots to an automated team that can handle recurring tasks with minimal human intervention.
The author built an AI writing reviewer agent by synthesizing user interviews and behavioral data into a psychological profile that configures the agent to critique drafts from the perspective of a cynical, technical audience.
This paper proposes a multi-agent framework built on CrewAI for automated conversational business intelligence, using five specialized AI agents to process queries, retrieve data, and generate insights, with evaluation showing significant accuracy and quality improvements over baselines.
This paper introduces Sanyu Studio, a multi-agent dialogue system that uses LLMs to construct art-historical narratives for Sanyu oil paintings, showing how AI can support diverse interpretations through user interactions.
The paper introduces GxP-Agent, a multi-agent system that uses a directed acyclic graph (DAG) to encode regulatory processes for reliable clinical trial programming with LLMs, achieving 100% accuracy on benchmarks compared to 0% for baseline approaches.
The article discusses how human-AI collaboration in hypothesis-driven research is advancing mathematical problem-solving and AI development, citing examples like progress on the Riemann Hypothesis and the AIRA-Compose system for architecture search.
SKILL is a self-correcting knowledge-guided iterative large language model agent that unifies multi-agent LLM reasoning and RL-based interaction for logic synthesis optimization, achieving significant improvements over expert flows.
The paper introduces aDSL, a co-designed domain-specific language and multi-agent system that improve LLM-driven 3D program synthesis through relational operators and iterative feedback, enhancing robustness and controllability in 3D content creation.
TeachMateGPT presents a multi-agent framework that improves retrieval-augmented generation for creating pedagogical assessments from science textbooks, achieving higher faithfulness and answer relevancy compared to baseline systems.
This paper introduces AEROBAT, the first multi-agent system to automate behavioral scientific research on AI agents, generating hypotheses, designing and executing controlled experiments, and writing reports. The authors demonstrate its efficacy across 12 target behaviors, finding statistical evidence for 26 hypotheses.
This paper presents IntelliAudit, a retrieval-grounded multi-agent system that uses large language models to evaluate IT audit controls against evidence corpora, generating cited recommendations and remediation guidance. The authors instantiate it on ISO/IEC 27001 and find it useful for audit preparation while emphasizing the need for human oversight.
QFoldAgent is a closed-loop multi-agent framework for quantum-classical protein structure prediction that iteratively optimizes Hamiltonian penalties using VQE and feedback, achieving improved RMSD and structural validity on 5-residue fragments.
This paper proposes a multi-agent embodied conversational system that generates level-appropriate dialogue for English learners using a generate-evaluate-regenerate loop with LLMs and a CEFR classifier. A pilot study with Japanese university students showed improved level appropriateness but no statistically significant reduction in foreign language anxiety.
This paper describes a hybrid multi-agent LLM system for conversational depression screening submitted to the eRisk 2026 challenge, using either a paid GPT-5-nano or open-source Gemma 27B model with algorithmic guidance (dialogue tree, reliability-weighted aggregation, cluster-based imputation) to achieve competitive BDI-II assessment at lower cost.
Pythia is a multi-agent system that autonomously writes and optimizes extraction prompts for clinical concepts without manual prompt engineering or fine-tuning, using a locally hosted open-weights model. It achieves mean sensitivity of 0.76 and specificity of 0.95 on clinical symptom detection, outperforming lexicon-based methods on specificity.