Tag
This paper introduces a multi-agent peer-reviewed reasoning method where multiple LLMs independently generate chain-of-thought reasoning and then evaluate each other's outputs to select the best answer. The method outperforms single-model reasoning and majority voting on medical QA benchmarks.
This paper investigates few-shot biomedical relation extraction using prompt-based learning with LLMs, comparing pairwise classification and joint generation approaches. The best model achieves micro-F1 of 0.44, outperforming previous few-shot results but remaining below supervised baselines, while macro-F1 surpasses the supervised baseline on rare relation types.
This review examines the role of machine learning across the biomedical Raman spectroscopy pipeline, from preprocessing and signal correction to clinical translation, highlighting barriers and future directions for reliable Raman-AI systems.
Fine-tuning small LLMs (3B-7B) with QLoRA on biomedical claim verification achieves higher F1 than GPT-4o and GPT-5 at 44.5x lower cost, and reveals a structural artifact in SciFact. The study demonstrates robust cross-domain transfer when training on structurally sound data.
This exploratory study evaluates whether augmenting AI agents with a medical research skill package improves the quality of transcriptomic research analysis outputs compared to native AI, using a multi-model human evaluation in an NSCLC biomarker task. Results show a directional but statistically non-significant improvement, highlighting the need for larger, more robust evaluations.
The NIH has added the Hugging Face Hub to its official list of Generalist Repositories for data sharing, allowing NIH-funded researchers to use it in their data sharing plans.
This paper introduces BioConCal, a supervised scorer that uses inference-time panel and candidate features to rank biomedical entity candidates surfaced by LLM panels, significantly improving over raw agreement for curator triage.
This arXiv paper presents a protocol for evaluating ChatGPT's ability to generate and verify biomedical associations using a RAG-enabled, cross-model majority voting workflow to address hallucination and ontology limitations.
BELIEF is a structured evidence modeling and uncertainty-aware fusion framework for biomedical question answering that converts retrieved documents into evidence objects and combines symbolic Dempster-Shafer reasoning with LLM-based inference. Experiments on PubMedQA, MedQA, and MedMCQA show BELIEF achieves state-of-the-art results in the majority of settings.
This paper introduces MHGraphBench, a knowledge-graph-grounded benchmark for evaluating large language models on mental health knowledge, including entity recognition, relation judgment, and multi-hop reasoning. Experiments across 15 LLMs reveal a gap between recognition and judgment capabilities.
This paper systematically evaluates five imbalance handling methods (RUS, ROS, SMOTE, re-weighting, direct F1 optimization) on three biomedical datasets (tabular, text, image) using models of varying complexity. Results show that benefits depend on model complexity and data modality, with ROS, re-weighting, and direct F1 optimization being effective for complex models on unstructured data.
MAML is a novel multi-modal AI model that unifies understanding of chemistry, genetics, and proteins, outperforming specialized models on 11 drug discovery benchmarks, promising to accelerate pharmaceutical research and improve success rates.
This paper introduces a unified benchmark to evaluate the robustness of Graph Neural Networks on noisy, text-derived knowledge graphs and the effectiveness of graph construction methods in the biomedical domain.
BioTool introduces a comprehensive biomedical tool-calling dataset with 34 tools and 7,040 human-verified query-API pairs, enabling fine-tuned LLMs to outperform GPT-5.1 on biomedical tool use and significantly enhance answer quality.
Researchers developed an electronics-free smart contact lens using microfluidics to continuously monitor eye pressure and automatically release glaucoma medication when needed, achieving 94% accuracy with a phone-based imaging app.