Tag
Parenting & Baby Advisor is an AI agent that delivers evidence-based parenting guidance from pregnancy through young adulthood, adapted to each child's age and family needs.
TwinCheck is an inference-time verification policy that enhances stateful tool agents by using evidence-grounded negative-twin comparisons, significantly improving task success rates in benchmarks like BFCL V4.
This article discusses the challenges in establishing proof and validation for medical AI systems, highlighting issues of trust, scientific rigor, and reliability in healthcare applications.
The article explores whether AI agents should be allowed to express uncertainty, such as saying 'I don't know', in their actions and governance, emphasizing the value of honest communication over forced classification.
GPS-Bench is an evidence-grounded benchmark for governance policy simulation that uses legislative records and public evidence to model actors and outcomes, enabling controlled comparisons of LLM-based methods for policy analysis.
Alternate Historian is an AI agent that turns 'what if' questions into rigorous historical thought experiments, providing evidence-backed analyses of how divergent scenarios could unfold.
FaithMed is a framework that trains LLMs for faithful evidence-based medical reasoning by integrating clinician-designed rubrics with reinforcement learning using step-level process reward assignment, achieving significant improvements over baselines on multiple medical benchmarks.
This perspective paper develops a conceptual and methodological framework for evaluating evidence-licensed claims in AI-assisted research, emphasizing calibration as a mechanism for managing scientific assertion rights and distinguishing between different AI research routes.
PathPocket is a multimodal AI agentic co-pilot for evidence-grounded pathology, utilizing a comprehensive evidence corpus and hypergraph to outperform existing state-of-the-art methods on over 200,000 real-world cases.
Proposes EVIDENT, a framework that integrates Bayesian training and evidence-based ranking for neural architecture selection, demonstrated on subject-specific blood glucose forecasting in type 1 diabetes, systematically selecting low-capacity models that generalize reliably.
This paper proposes an evidence-based model to automatically generate query keywords from query-free summarization datasets, enabling the creation of query-focused summarization datasets. Experimental results show that summaries generated using evidence-based queries achieve competitive ROUGE scores compared to original queries.
DeepER-Med introduces an agentic AI framework for evidence-based medical research with explicit evidence appraisal criteria and a new benchmark dataset (DeepER-MedQA) of 100 expert-curated medical questions, demonstrating superior performance over production platforms with clinical validation on real-world cases.