Tag
PhoenixNest-Video introduces an evidence-grounded multimodal agent framework for automated video interview assessment, achieving 91.50% grade-level accuracy on a benchmark by using structured video graphs and reinforcement learning.
HyGRAIL is a cost-aware framework that combines GNN triage with LLM review for scientific hypothesis discovery on knowledge graphs, achieving improved F1 scores and reduced LLM calls on the MatKG benchmark.
SGHA is a fully automated system that uses a local 9B language model to discover research problems from scientific literature by structuring evidence and detecting structural gaps, offering transparency and privacy over proprietary models.
This paper investigates agentic data cleaning without a clean reference, proposing an evidence-grounded framework and evaluating trade-offs across multiple configurations.
This paper presents ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools for thyroid ultrasound, storing outputs as auditable evidence records. Developed on a large multicentre dataset, it achieves strong results in nodule segmentation, benign-malignant classification, and report generation.
OmniScientist is an end-to-end omni-modal AI scientist that performs multidisciplinary research directly from heterogeneous raw evidence using autonomous agents and lifecycle-wide perception. Evaluated on 36 real-data cases, it improves evidence-grounded discovery across diverse scientific modalities.
TumorBoard is a multi-agent decision-support system for longitudinal neuro-oncology that uses a shared longitudinal case state and auditable claim-evidence ledger. It outperforms baselines on a 360-case benchmark, with a safety governor reducing harmful recommendations.
Introduces XL-DocBench, a human-verified benchmark for extra-long document understanding with 1,519 questions across six professional domains, requiring multi-page evidence and structured reasoning, showing current LLMs still struggle with long-context professional documents.
This paper proposes EGTA, an Evidence-Grounded Terminology Adaptation framework for simultaneous speech translation that selectively uses document-specific terminology to improve translation of rare terms, achieving significant gains in named-entity and acronym recall without full-model fine-tuning.
MSCE is a training-free framework that organizes LLM agent experience into three memory levels and converts them into reusable skills with evidence links, outperforming existing memory and skill-augmented baselines.
Introduces Knowledgeless Language Models (KLLMs), pretrained on corpora with anonymized entities to suppress parametric recall and enhance evidence-grounded reasoning, achieving substantial improvements on contextual QA, fact verification, and hallucination detection benchmarks.
This paper introduces ResearchStudio-Idea, a reusable research ideation skill suite based on analysis of ML conference papers. It includes Paper-Search, Scoop-Check, and IdeaSpark, generating evidence-supported research ideas using 15 patterns extracted from 1,947 papers.
This paper introduces Data Journalist Agent (Data2Story), a multi-agent framework that automates data journalism by generating evidence-grounded, multimodal news stories while ensuring transparency and verifiability.