Tag
AI is not replacing radiologists but is transforming their jobs by enhancing accuracy and efficiency in medical imaging, fostering collaboration between humans and AI systems.
The paper presents a locally deployed multi-agent AI system for structuring radiology reports and performing quality assurance, with radiologist evaluation showing favorable performance.
RadleyAI is creating the first AI-native radiology practice to tackle the radiology crisis, having acquired a practice and reached $4.3M in annualized revenue.
RadFusion is a framework that adds threshold controllability to radiology report generation by fusing a multi-label classifier with a VQA-based generator and an LLM rewrite step, enabling sensitivity-specificity trade-offs and ROC-based validation. Experiments on MIMIC-CXR show improved diagnostic accuracy and clinically adaptable report behavior.
An article discussing why AI forecasting is difficult, emphasizing that capability gains often fail to translate into end-to-end impact due to social, institutional, and tacit bottlenecks, using radiology as a case study.
Microsoft Research introduces CARE-X, a unified chest X-ray vision-language model that combines flexible reasoning, calibrated predictions, and tool-augmented measurement for clinically useful radiology interpretation.
TotalSegmentator now has an MCP server, enabling AI agents like Codex to run it and answer clinical questions about CT scans, e.g., detecting hepatosplenomegaly or NAFLD/NASH.
This paper highlights that VLMs for chest x-ray report generation can score well on benchmarks while erasing clinically meaningful terms and introducing biased language, and proposes a framework to measure these failures.
This paper performs a forensic reproducibility audit of a radiology vision-language model benchmark, finding divergences between the intended protocol and released artifacts that invalidate the original claims. The authors propose a benchmark contract to expose such failure classes.
Muse Spark 1.1 outperforms GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam 2.0, a new visual reasoning benchmark for autonomous AI diagnosis in healthcare, though it still lags behind Fable and human radiologists.
This paper introduces a pipeline that converts free-text chest radiograph reports into multi-label matrices with a single structured annotation pass, enabling reconfiguration of label schemas via dictionary edits without relabeling, saving significant cost and time.
This paper adapts a diffusion language model for interactive radiology report drafting, showing it matches autoregressive models in accuracy while offering unique infill capabilities that allow radiologists to fix report fragments and have the model fill in the text between them.
CONFLUX is a 3D latent diffusion model for chest CT synthesis that achieves high-fidelity volumetric generation with controllable clinical attributes, enhanced by a reinforcement learning post-training stage to improve conditioning reliability. The model and a synthetic dataset of ~200k chest CT volumes are released.
This paper adapts a mixture-of-experts diffusion language model, DiffusionGemma-26B, for interactive radiology report drafting, showing it matches or exceeds autoregressive models in medical VQA with 3.5-4.4x faster decoding and bidirectional infill capabilities.
This paper measures information degradation in AI-rewritten radiology reports, finding that tasks producing cleaner text for multimodal training cause greater cross-modal alignment loss, a phenomenon termed the 'slop paradox'.
Proposes MedExpMem, an experience memory framework that enables medical vision-language models to accumulate and retrieve discriminative diagnostic experience from past cases, improving differential diagnosis accuracy by up to 7.0% on a radiology benchmark.
AnchorDiff proposes a topology-aware masked diffusion framework for radiology report generation, integrating RadGraph-derived clinical anchors and confidence-based rewriting to achieve state-of-the-art results on MIMIC-CXR and MIMIC-RG4 benchmarks.
A new AI model (REDMOD) can detect pancreatic cancer up to three years earlier than human doctors by analyzing CT scans for subtle irregularities, potentially improving early diagnosis and survival rates.