Tag
Raven is an open-source harness for orchestrating multiple AI agents to perform complex tasks, with capabilities for recursive self-improvement.
A project running 400+ LLM agents in a private MMO server, exploring challenges like stale perception, unacknowledged actions, and load shedding using locally served Qwen models.
This paper presents MACBT, a multi-agent cognitive behavioral therapy decision support system with longitudinal memory, trained on Chinese dialogues and shown to improve clinical authenticity and session quality compared to other AI systems.
The paper proposes a multi-agent AI architecture called SOFAI, inspired by Kahneman's dual-system theory, to enhance AI capabilities through metacognition and balancing fast and slow reasoning processes.
Microsoft Research introduces Agensh, a self-organized multi-agent harness that scales to over 1,000 agents, demonstrating improved test-pass rates on coding tasks without a central orchestrator.
The author suggests that heavy AI users may inefficiently allocate model capacity by not optimizing model selection, proposing that work should be routed to the least expensive capable model to save costs and improve efficiency.
This paper introduces a multi-agent system called CB-MAS for document sensitivity classification, which addresses the limitation of fixed input length in transformer models like BERT by using iterative consultation and channel boosting.
Schematize is an open-source multi-agent system that interactively generates and refines information-extraction schemas for legal research, achieving top performance in human evaluations.
This paper presents MedNotes, a multi-agent pipeline for generating source-grounded synthetic clinical notes from longitudinal structured EHR data, achieving high accuracy and improving downstream clinical modeling tasks.
Agensh is a scalable self-organized multi-agent system without a central orchestrator that improves performance on complex tasks by scaling the number of agents, showing significant test-pass rate increases on benchmarks like ProgramBench and pandoc.
FinSkillOps is a multi-agent system for SEC filing question answering that introduces controlled skill management for self-evolution, improving accuracy and reducing errors in financial QA systems.
ScientistTwo is a fully autonomous multi-agent framework that conducts end-to-end scientific research, generating expert-level papers and codebases that outperform human state-of-the-art models and meet acceptance standards at top-tier AI conferences like ICLR and NeurIPS.
This paper proposes a hybrid agentic AI framework for supply chain analytics that uses a coordinator agent and specialized agents to improve decision-making, achieving 90% accuracy and reducing token usage by fourfold.
This paper presents SurgicalRoomAgent, a voice-interactive multi-agent system for smart operating rooms that uses large language models to enable natural language device control, focusing on architecture design and key technologies like KV Cache optimization and progressive prompt disclosure to reduce latency in sterile environments.
This paper introduces a multi-agent AI system for measuring and diagnosing competitive visibility in LLM-mediated e-commerce using Agentic Share-of-Search, with an ablation study showing feasibility.
This paper introduces a human-in-the-loop, multi-agent GeoAI system for Arctic eco-navigation that integrates operational, ecological, and community criteria to plan safer and more socially responsible maritime routes.
The paper introduces Dude, a dual-detection multi-agent system using LLMs to detect discrepancies between research papers and code, addressing limitations of single-agent approaches with improved recall and precision.
NS-Copilot is an LLM-driven multi-agent system that autonomously supports end-to-end workflows for diverse neuroscience analysis tasks, outperforming baselines on key benchmarks.
ConvDeck introduces a multi-agent pipeline for conversational paper-to-slide generation that enables users to iteratively refine presentations through stage-specific feedback at different stages of the process.
EULER is a multi-agent system that explores cross-domain transfers to automatically prove or refute mathematical conjectures, validated with stress tests and evaluated on 120 recent conjectures, producing proofs, refutations, and partial results.