Tag
This paper introduces a process-centric benchmark for evaluating AI-assisted peer review systems, aiming to improve transparency and reliability beyond final decision accuracy.
This paper presents ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools for thyroid ultrasound, storing outputs as auditable evidence records. Developed on a large multicentre dataset, it achieves strong results in nodule segmentation, benign-malignant classification, and report generation.
Proposes Causal-Audit, a framework for explicit and auditable causal reasoning in LLMs using target-aware causal graph construction and path-level evidence aggregation, outperforming existing methods on benchmarks.
FirstResearch introduces a structured framework for LLM scientific discovery agents that generates a Research Question Certificate containing primitive definitions, assumptions, mechanism, falsifiable hypothesis, and failure update rules, making the proposed research question inspectable before execution. Preliminary evaluations using LLM judges show that the certificate-centered approach outperforms baseline systems in audibility and score.