MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models
Summary
MedPMC is an automated framework that transforms medical literature into high-fidelity multimodal data for foundation models, achieving significant improvements across multiple benchmarks and clinical settings.
View Cached Full Text
Cached at: 07/13/26, 03:51 PM
Paper page - MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models
Source: https://huggingface.co/papers/2607.07673 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Medicineisinherentlymultimodal,requiringclinicianstosynthesizeinformationacrossdiversedatastreams.Yetthedevelopmentofmultimodalfoundationmodelsisconstrainedbylimitedaccesstolarge-scale,high-qualityclinicaldata.AlthoughPubMedCentral(PMC)offersacomplementarysourceofexpert-authoredimage-textdata,existingPMC-derivedresourcesremainlimitedinfidelity,reproducibility,andclinicalvalidation.WeintroduceMedPMC,anautomated,continuouslyupdatableframeworkthattransformspermissivelylicensedliteratureintohigh-fidelityinfrastructureformedicalmultimodalmodels.Appliedto6.1millionPMCarticles,MedPMCcurated11millionmedicalimage-textpairs.Componentevaluationsshowedstrongperformanceforinitialscreening(F1=93.2),multi-panelfiguredetection(F1=96.5),figureseparation(mAP=89.8),captionseparationandalignment(F1=81.4;ROUGE-L=85.3),andmedicalfigureclassification(F1=96.5).Manualreviewbyfiveannotators,threewithmedicaltraining,found95.3%ofMedPMCimagesmedicallyrelevant,versus19.7%inapriorPMC-deriveddataset.Across26benchmarksspanning11specialties,aMedPMC-trainedCLIP-stylemodelimprovedaveragezero-shotAUCby7.1percentagepointsoverthestrongestarchitecture-matchedbiomedicalCLIPbaselinedespiteusingfewerthanhalfasmanyimage-textpairs.Asthevisionencoderinamultimodallargelanguagemodel,itimprovedmedicalvisualquestion-answeringby1.9and16.9percentagepointsacrosstwobenchmarks.In10,524YaleNewHavenHealthSystemdermatologyphotographs,itimprovedmorphology-to-imageretrievalRecall@5by11.7percentagepoints.Thesefindingsshowthathigh-fidelityliteraturecurationstrengthensmedicalmultimodalfoundationmodelsacrossbenchmarkandclinicalsettings.Wepubliclyreleasetheframework,corpus,benchmarks,andpretrainedmodels.
View arXiv pageView PDFGitHub1Add to collection
Get this paper in your agent:
hf papers read 2607\.07673
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.07673 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.07673 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.07673 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis
This paper systematically evaluates foundation model representations for multimodal cancer analysis, benchmarking unimodal and multimodal fusion strategies on real-world cohorts, and assessing trustworthiness via conformal prediction.
MEDSYN: Benchmarking Multi-Evidence Synthesis in Complex Clinical Cases for Multimodal Large Language Models
MEDSYN is a multilingual multimodal benchmark for evaluating MLLMs on complex clinical cases with up to 7 distinct visual evidence types per case. The study reveals that while frontier models match human experts on differential diagnosis generation, all MLLMs show significant gaps in final diagnosis selection due to poor synthesis of heterogeneous clinical evidence.
MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation
MedRealMM is a new multimodal benchmark for Chinese online medical consultation, built from real-world patient-doctor interactions, evaluating LLMs on next-response generation with clinical rubrics.
Large Language Models as Unified Multimodal Learners for Clinical Prediction
The paper proposes converting multimodal patient data (text, labs, vitals) into a single natural language sequence and fine-tuning LLMs for clinical prediction, achieving comparable or better performance than specialized fusion architectures across three tasks.
MMIR-TCM: Memory-Integrated Multimodal Inference and Retrieval for TCM Clinical Decision Support
This paper introduces MMIR-TCM, a novel framework that integrates multimodal large language models with memory-augmented segmentation and retrieval-augmented generation to support Traditional Chinese Medicine clinical decision making, along with a new dataset MedTCM and evaluation metric TDEU.