Tag
Surg-UniWorld is a unified surgical world model with multimodal control experts, enabling controllable generation of coherent instrument-tissue interaction videos using edge, depth, and optical-flow inputs. It introduces a new benchmark (Cholec80-SurgWAM) and outperforms existing controllable video generation methods.
TotalSegmentator now has an MCP server, enabling AI agents like Codex to run it and answer clinical questions about CT scans, e.g., detecting hepatosplenomegaly or NAFLD/NASH.
Introduces LiNC, a lightweight noise correction method that learns per-sample trust parameters to distinguish clean and noisy labels using a Gaussian Mixture Model, achieving robust accuracy gains on medical imaging datasets under high label noise.
This paper proposes DualIFM, an interpretable-by-design foundation model for retinal fundus images, achieving performance comparable to RETFound with far fewer parameters while providing interpretable predictions.
This paper presents a model-agnostic framework for per-modality failure analysis in multimodal clinical AI, distinguishing loud vs silent failures when a modality is dropped. Validated on planted ground truth and applied to EchoJEPA and HuBERT-ECG embeddings for LVEF prediction, it shows that dropping echo nearly doubles error.
This paper performs a forensic reproducibility audit of a radiology vision-language model benchmark, finding divergences between the intended protocol and released artifacts that invalidate the original claims. The authors propose a benchmark contract to expose such failure classes.
Marion Lepert introduces OpenDerm, an open-source 4-DOF robot for high-resolution skin imaging to enable early melanoma detection at home, arguing that general-purpose home robots could make routine skin screening accessible.
QFedPolyp proposes a federated learning framework for polyp segmentation that uses quantization-aware training to reduce communication costs and achieve faster inference while preserving privacy.
Discussion on which area of healthcare will see the biggest transformation from AI in the next few years, including diagnosis, drug discovery, patient monitoring, and medical imaging.
A systematic review of diffusion-based methods for medical image inpainting, covering architectures, applications, datasets, and evaluation strategies, with a proposed taxonomy and identification of challenges such as lack of standardized benchmarks.
GLI-AL is a new label resource for glioma MRI that unifies anatomy and lesion labels, expanding supervision to include healthy tissues and previously unlabeled abnormalities. It provides 1,251 label sets aligned with BraTS-GLI cases.
ClinFusion is a vision-centric multimodal large language model for holistic medical understanding that unifies 2D and 3D medical image analysis using a cascaded vision encoder. It achieves state-of-the-art results on 20 out of 24 benchmarks and outperforms proprietary models like GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks.
A tutorial from freeCodeCamp teaching how to build a tumor segmentation pipeline for breast ultrasound images using MONAI, emphasizing data profiling before model selection.
This paper introduces PathReportEval, a standardized benchmark and evaluation framework for pathology report generation from whole-slide images, including a new clinically grounded metric called Clinical Report Quality Score (CRQS) that better captures factual correctness than conventional lexical metrics.
FedCC proposes a federated learning framework combining a frozen DINOv2 backbone with a lightweight YOLO detection head and LoRA modules for robust corpus callosum localization in fetal ultrasound images, achieving strong performance with greatly reduced communication cost in a multi-center setting.
Introduces qZACH-ViT, a quantization-aware extension of ZACH-ViT with recursive intrinsic explanations, and Recursive Attribution-Stabilized Optimization (RASO) for stable attribution gradients. Achieves high prediction agreement and speedups on MedMNIST datasets after INT8 conversion.
PnP-CoSMo is a plug-and-play framework for multi-contrast MRI reconstruction that learns content/style models from image data, enabling reconstruction without raw k-space training data. It is generalizable across contrasts and forward operators.
Proposes a decoupled training strategy that adapts normalization layers and uses precomputed features to reduce overhead in transfer learning, achieving competitive accuracy with significantly reduced training time and energy consumption.
LMU researchers developed a Quantum Boltzmann Machine model for pneumonia detection from chest X-rays. Using fewer than 9,000 parameters (vs. millions in classical CNNs), the model achieves 84–86% accuracy, demonstrating potential advantages of quantum machine learning for small medical datasets.
Indian scientists at IIT-Madras have created the most detailed 3D atlas of the human brainstem at cellular resolution, called Anchor, integrating over 500 tissue sections to map more than 200 cell clusters and pathways, bridging whole-brain imaging and cellular pathology.