MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models
Summary
MedPMC is an automated framework that transforms medical literature into high-fidelity multimodal data for foundation models, achieving significant improvements across multiple benchmarks and clinical settings.
View Cached Full Text
Cached at: 07/13/26, 03:51 PM
Paper page - MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models
Source: https://huggingface.co/papers/2607.07673 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Medicineisinherentlymultimodal,requiringclinicianstosynthesizeinformationacrossdiversedatastreams.Yetthedevelopmentofmultimodalfoundationmodelsisconstrainedbylimitedaccesstolarge-scale,high-qualityclinicaldata.AlthoughPubMedCentral(PMC)offersacomplementarysourceofexpert-authoredimage-textdata,existingPMC-derivedresourcesremainlimitedinfidelity,reproducibility,andclinicalvalidation.WeintroduceMedPMC,anautomated,continuouslyupdatableframeworkthattransformspermissivelylicensedliteratureintohigh-fidelityinfrastructureformedicalmultimodalmodels.Appliedto6.1millionPMCarticles,MedPMCcurated11millionmedicalimage-textpairs.Componentevaluationsshowedstrongperformanceforinitialscreening(F1=93.2),multi-panelfiguredetection(F1=96.5),figureseparation(mAP=89.8),captionseparationandalignment(F1=81.4;ROUGE-L=85.3),andmedicalfigureclassification(F1=96.5).Manualreviewbyfiveannotators,threewithmedicaltraining,found95.3%ofMedPMCimagesmedicallyrelevant,versus19.7%inapriorPMC-deriveddataset.Across26benchmarksspanning11specialties,aMedPMC-trainedCLIP-stylemodelimprovedaveragezero-shotAUCby7.1percentagepointsoverthestrongestarchitecture-matchedbiomedicalCLIPbaselinedespiteusingfewerthanhalfasmanyimage-textpairs.Asthevisionencoderinamultimodallargelanguagemodel,itimprovedmedicalvisualquestion-answeringby1.9and16.9percentagepointsacrosstwobenchmarks.In10,524YaleNewHavenHealthSystemdermatologyphotographs,itimprovedmorphology-to-imageretrievalRecall@5by11.7percentagepoints.Thesefindingsshowthathigh-fidelityliteraturecurationstrengthensmedicalmultimodalfoundationmodelsacrossbenchmarkandclinicalsettings.Wepubliclyreleasetheframework,corpus,benchmarks,andpretrainedmodels.
View arXiv pageView PDFGitHub1Add to collection
Get this paper in your agent:
hf papers read 2607\.07673
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.07673 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.07673 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.07673 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis
This paper systematically evaluates foundation model representations for multimodal cancer analysis, benchmarking unimodal and multimodal fusion strategies on real-world cohorts, and assessing trustworthiness via conformal prediction.
MEDSYN: Benchmarking Multi-Evidence Synthesis in Complex Clinical Cases for Multimodal Large Language Models
MEDSYN is a multilingual multimodal benchmark for evaluating MLLMs on complex clinical cases with up to 7 distinct visual evidence types per case. The study reveals that while frontier models match human experts on differential diagnosis generation, all MLLMs show significant gaps in final diagnosis selection due to poor synthesis of heterogeneous clinical evidence.
OpenMHC: Accelerating the Science of Wearable Foundation Models
OpenMHC introduces the largest open-access wearable health dataset with over 60 million hours of data and open-source implementations of wearable foundation models, including a unified benchmark for prediction, imputation, and forecasting.
MedMix: Specialization-Consistent Federated Sparse MoEs under Modality Heterogeneity
MedMix is a semantic-alignment framework for federated multimodal sparse Mixture-of-Experts that addresses modality heterogeneity by coordinating routing and expert specialization, achieving improved performance in medical AI datasets.
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
ClinFusion is a vision-centric multimodal large language model for holistic medical understanding that unifies 2D and 3D medical image analysis using a cascaded vision encoder. It achieves state-of-the-art results on 20 out of 24 benchmarks and outperforms proprietary models like GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks.