FedCC: A Low-Resource Federated Adaptation of Foundation Models for Robust Corpus Callosum localization in Fetal Ultrasound Images

arXiv cs.LG Papers

Summary

FedCC proposes a federated learning framework combining a frozen DINOv2 backbone with a lightweight YOLO detection head and LoRA modules for robust corpus callosum localization in fetal ultrasound images, achieving strong performance with greatly reduced communication cost in a multi-center setting.

arXiv:2607.18283v1 Announce Type: new Abstract: Accurate localization of the corpus callosum (CC) in fetal ultrasound (US) images is crucial for the early identification of neurodevelopmental abnormalities. However, this task remains highly challenging due to the intrinsic limitations of US imaging, including low contrast, speckle noise, and the considerable anatomical variability of the CC. We propose FedCC, a federated learning (FL)-based framework for CC localization in fetal US images, specifically designed for realistic multi-center and resource-constrained clinical settings without requiring data sharing. The framework integrates a frozen DINOv2 backbone with a lightweight YOLO-based detection head. To enable parameter-efficient adaptation, Low-Rank Adaptation (LoRA) modules are incorporated, allowing only a small subset of parameters to be optimized and exchanged among clients. This strategy substantially reduces both computational and communication overhead, making the framework suitable for low-resource environments. The proposed approach was evaluated on a multi-center dataset comprising 10,970 ultrasound frames acquired from 58 pregnant women during routine neurosonographic examinations across three clinical sites using heterogeneous imaging devices. The proposed framework achieved strong performance in the federated setting. In particular, the combination of DINOv2 and LoRA under the FedAvg strategy achieved an average mAP@50 of 0.857 and an F1-score of 0.803, outperforming both full fine-tuning and encoder-freezing baselines. Notably, the proposed approach reduced the number of trainable parameters to 2.9M compared with 24.4M in full fine-tuning, corresponding to an approximately 8.5$\times$ reduction in communication cost. These findings represent a promising step toward scalable, privacy-preserving, and clinically deployable AI systems for fetal neurosonography.
Original Article
View Cached Full Text

Cached at: 07/22/26, 08:18 AM

# FedCC: A Low-Resource Federated Adaptation of Foundation Models for Robust Corpus Callosum localization in Fetal Ultrasound Images
Source: [https://arxiv.org/html/2607.18283](https://arxiv.org/html/2607.18283)
Sara MocciaGiuseppe RizzoGianpaolo GrisoliaRicciarda RaffaelliLorenzo VasciaveoFrancesco d’AntonioMaria Chiara Fiorentino[mariachiara\.fiorentino@unich\.it](https://arxiv.org/html/2607.18283v1/mailto:[email protected])

###### Abstract

Background:Accurate localization of the corpus callosum \(CC\) in fetal ultrasound \(US\) images is crucial for the early identification of neurodevelopmental abnormalities\. However, this task remains highly challenging due to the intrinsic limitations of US imaging, including low contrast, speckle noise, and the considerable anatomical variability of the CC\. In addition, substantial inter\-site variability caused by heterogeneous acquisition protocols and imaging devices further complicates automated analysis\. Despite its clinical importance, automatic CC localization in fetal US has received limited attention in the literature, mainly because of these technical challenges and the scarcity of publicly available annotated datasets\.

Methods:We proposeFedCC, a federated learning \(FL\)\-based framework for CC localization in fetal US images, specifically designed for realistic multi\-center and resource\-constrained clinical settings without requiring data sharing\. The framework integrates a frozen DINOv2 backbone with a lightweight YOLO\-based detection head\. To enable parameter\-efficient adaptation, Low\-Rank Adaptation \(LoRA\) modules are incorporated, allowing only a small subset of parameters to be optimized and exchanged among clients\. This strategy substantially reduces both computational and communication overhead, making the framework suitable for low\-resource environments\. The proposed approach was evaluated on a multi\-center dataset comprising 10,970 ultrasound frames acquired from 58 pregnant women during routine neurosonographic examinations across three clinical sites using heterogeneous imaging devices\.

Results & Discussions:The proposed framework achieved strong performance in the federated setting\. In particular, the combination of DINOv2 and LoRA under the FedAvg strategy achieved an average mAP@50 of 0\.857 and an F1\-score of 0\.803, outperforming both full fine\-tuning and encoder\-freezing baselines\. Notably, the proposed approach reduced the number of trainable parameters to 2\.9M compared with 24\.4M in full fine\-tuning, corresponding to an approximately 8\.5×\\timesreduction in communication cost\. Furthermore, the federated framework improved generalization across heterogeneous domains, especially on the most challenging client\. These findings represent a promising step toward scalable, privacy\-preserving, and clinically deployable AI systems for fetal neurosonography, although further validation on larger and more diverse cohorts is still required\.

###### keywords:

Federated Learning , Fetal Ultrasound , Foundation Models , Corpus Callosum

\\affiliation

\[1\]organization=Department of Innovative Technologies in Medicine & Dentistry, Università degli Studi ”G\. D’Annunzio” Chieti \- Pescara, city=Chieti, country=Italy\\affiliation\[2\]organization=Centre for Fetal care and High Risk Pregnancy, Università degli Studi ”G\. D’Annunzio” Chieti \- Pescara, city=Chieti, country=Italy\\affiliation\[3\]organization=Department of Obstetrics and Gynaecology, Sapienza Università di Roma, city=Roma, country=Italy\\affiliation\[4\]organization=Department of Obstetrics and Gynaecology, city=Mantova, country=Italy\\affiliation\[5\]organization=Department of Obstetrics and Gynaecology, Università di Verona, city=Verona, country=Italy\\affiliation\[6\]organization=Department of Obstetrics and Gynaecology, Università di Foggia, city=Foggia, country=Italy

\{graphicalabstract\}![[Uncaptioned image]](https://arxiv.org/html/2607.18283v1/framework.png)

\{highlights\}

A federated learning framework for corpus callosum detection in fetal ultrasound images is proposed\.

A frozen DINOv2 backbone is combined with LoRA adapters and a YOLO\-based detection architecture\.

Only LoRA adapters and detection\-head parameters are shared across federated clients\.

Communication costs are reduced by up to 8\.5×\\timescompared with full model fine\-tuning\.

An average mAP@50 of 0\.857 and an F1\-score of 0\.803 are achieved on a multi\-center fetal ultrasound dataset\.

## 1Introduction

Fetal ultrasound \(US\) is the standard imaging modality for monitoring fetal development and assessing the anatomy of the central nervous system throughout pregnancy\[[44](https://arxiv.org/html/2607.18283#bib.bib2),[33](https://arxiv.org/html/2607.18283#bib.bib3)\]\. During the second and third trimester, neurosonographic examination plays a fundamental role in the early detection of brain abnormalities, enabling timely prenatal counseling and clinical management\[[43](https://arxiv.org/html/2607.18283#bib.bib4)\]\. Among the anatomical structures evaluated in this context, the corpus callosum \(CC\) holds particular clinical relevance\. As the main white matter commissure connecting the two cerebral hemispheres, the CC is essential for interhemispheric communication and normal neurodevelopment\[[22](https://arxiv.org/html/2607.18283#bib.bib7)\], and its prenatal assessment is routinely performed on the mid\-sagittal plane, where it appears as a thin, curved hypoechoic band above the cavum septi pellucidi \(Fig\.[1](https://arxiv.org/html/2607.18283#S1.F1)\)\.

![Refer to caption](https://arxiv.org/html/2607.18283v1/corpus.png)\(a\)Sagittal view of the fetal brain highlighting the corpus callosum \(yellow structure\) and surrounding midline structures\.
![Refer to caption](https://arxiv.org/html/2607.18283v1/cc1.jpg)\(b\)Representative fetal ultrasound frame with the ground\-truth bounding box \(green\) delineating the corpus callosum in the mid\-sagittal plane\.

Figure 1:Anatomical context and localisation target for corpus callosum detection\.Accurate localization of this structure is clinically critical, as abnormalities of the CC, including agenesis, dysgenesis, and hypoplasia, represent some of the most common congenital malformations of the fetal central nervous system, frequently associated with neurodevelopmental impairment, chromosomal abnormalities, and additional cerebral anomalies\[[34](https://arxiv.org/html/2607.18283#bib.bib6),[28](https://arxiv.org/html/2607.18283#bib.bib8)\]\. Early and reliable detection is therefore essential to guide further diagnostic investigations, genetic counseling, and prognosis estimation\[[35](https://arxiv.org/html/2607.18283#bib.bib10)\]\.

Despite its clinical importance, robust localization of the CC in fetal US remains challenging\. The structure is inherently small, and its morphological appearance changes substantially during gestation, particularly between 18 and 32 weeks, as the CC progressively elongates and thickens\[[6](https://arxiv.org/html/2607.18283#bib.bib41)\]\. These developmental changes are compounded by the intrinsic limitations of fetal US imaging, including speckle noise, low contrast, acoustic shadowing, and considerable variability introduced by fetal position, maternal body habitus, and probe orientation\[[30](https://arxiv.org/html/2607.18283#bib.bib14),[11](https://arxiv.org/html/2607.18283#bib.bib5)\]\. In this context, automatic methods for CC localization could meaningfully support clinicians by improving the consistency and efficiency of fetal neurosonographic examinations\. Deep learning \(DL\)\-based approaches have shown considerable promise in automating complex image analysis tasks in fetal US, potentially supporting clinicians by improving consistency, reproducibility, and efficiency across different clinical settings\. Several methods have been proposed for fetal US analysis, including standard plane detection, biometric estimation, and landmark localization\[[11](https://arxiv.org/html/2607.18283#bib.bib5),[40](https://arxiv.org/html/2607.18283#bib.bib11)\]\. However, realizing their full potential requires training on large and diverse datasets that capture the variability of real\-world acquisitions across different centers, scanners, and populations, which is a requirement that may directly conflict with the regulatory and ethical constraints that typically limit direct data sharing between institutions\. As a consequence, most existing methods for fetal US image analysis have been developed using data from a single center and a specific US system\[[3](https://arxiv.org/html/2607.18283#bib.bib9),[41](https://arxiv.org/html/2607.18283#bib.bib15)\], making them inherently prone to overfitting to a particular acquisition setting and poorly generalizable to external clinical environments\.

To address these limitations, recent studies have begun to explore the use of foundation models \(FMs\), which are pretrained on large and heterogeneous datasets and are therefore able to learn more general and transferable representations\[[17](https://arxiv.org/html/2607.18283#bib.bib36),[21](https://arxiv.org/html/2607.18283#bib.bib34)\]\. Compared with conventional models trained from scratch on limited fetal US datasets, FMs have the potential to better handle the variability introduced by different scanners, acquisition protocols, and clinical centers\[[31](https://arxiv.org/html/2607.18283#bib.bib38),[27](https://arxiv.org/html/2607.18283#bib.bib39)\]\. However, adapting foundation models \(FMs\) to task\-specific applications typically relies on fine\-tuning strategies that require updating a large number of parameters, resulting in substantial computational and memory demands, particularly in resource\-constrained environments\. Parameter\-Efficient Fine\-Tuning \(PEFT\) approaches address this limitation by introducing only a small set of trainable parameters while keeping most of the pretrained backbone frozen\. In this way, PEFT significantly reduces computational and communication costs while preserving the rich and generalizable representations learned during pretraining\[[45](https://arxiv.org/html/2607.18283#bib.bib35),[8](https://arxiv.org/html/2607.18283#bib.bib37)\]\.

These characteristics become particularly advantageous in federated learning \(FL\) settings, where multiple institutions collaboratively train a shared model without exchanging sensitive patient data, thereby preserving privacy and complying with data\-sharing regulations\[[38](https://arxiv.org/html/2607.18283#bib.bib71)\]\. When integrated with PEFT, only the lightweight adapter parameters need to be exchanged among participating institutions during the federated optimization process, while the FM backbone remains locally stored and frozen\. This strategy substantially decreases both the number of trainable parameters and the communication overhead across sites, making the deployment and adaptation of large FMs more practical in distributed clinical environments\[[2](https://arxiv.org/html/2607.18283#bib.bib73),[38](https://arxiv.org/html/2607.18283#bib.bib71)\]\.

Moreover, reducing the number of trainable and transmitted parameters also lowers the computational and energy demands associated with model adaptation, potentially decreasing the environmental impact of AI development and deployment\. This perspective is aligned with the European Commission’s guidelines for trustworthy AI, which identify societal and environmental well\-being as a key requirement and encourage the design of AI systems that consider their environmental footprint throughout their lifecycle\[[7](https://arxiv.org/html/2607.18283#bib.bib81)\]\. By requiring only lightweight parameter updates, the proposed framework lowers both the computational and communication burden, thereby facilitating participation in collaborative training initiatives even in resource\-constrained settings\[[38](https://arxiv.org/html/2607.18283#bib.bib71),[37](https://arxiv.org/html/2607.18283#bib.bib12)\]\. Taken together, the integration of FL, FMs, and PEFT represents a principled and promising strategy for developing robust, generalizable systems for automatic CC localization across heterogeneous fetal US datasets\. The main contributions of this work are as follows:

- •FedCC, a FL framework for automatic CC localization that combines a DINOv2\-based FM with a YOLO\-based detection head\. The FM backbone is kept frozen, and only lightweight Low\-Rank Adaptation \(LoRA\) adapter modules\[[13](https://arxiv.org/html/2607.18283#bib.bib1)\]are trained and shared across participating centers during federation, drastically reducing both the number of trainable parameters and the communication overhead\. This design makes the adaptation of large FMs practically feasible in distributed clinical settings while preserving patient privacy\.
- •A systematic evaluation of robustness under multi\-device domain shift: We assessFedCCunder a realistic multi\-center and multi\-device scenario characterized by substantial inter\-client variability in scanner type, spatial resolution, image quality, and acquisition conditions\. Through extensive comparisons against alternative foundation\-model backbones, CNN\-based baselines, centralized training, different adaptation strategies, and FL aggregation methods, we show that the proposed LoRA\-based federated adaptation improves cross\-client generalization and maintains robust performance under heterogeneous acquisition settings\.
- •A novel fetal US dataset for CC detection, acquired during clinical practice across three Italian clinical centers located in Abruzzo, Puglia, and Veneto regions, encompassing images collected with both high\- and low\-resolution US devices\. The dataset captures the substantial variability encountered, and represents a valuable resource for the development and evaluation of generalizable CC detection methods\.

## 2Related Work

Only a limited number of studies have specifically addressed the automatic analysis of the CC in fetal US, reflecting both the inherent difficulty of the task and the scarcity of annotated datasets\. Among the earliest contributions,\[[14](https://arxiv.org/html/2607.18283#bib.bib42)\]proposes a framework for the automatic segmentation of the CC and choroid plexus in fetal US images\. The framework first identifies regions with homogeneous intensity patterns and then characterizes them using descriptors that encode both shape and local intensity information\. These descriptors are subsequently combined with a boosting classifier to segment the target anatomical structures\. The framework was evaluated on two small 2\-D fetal US datasets: 219 midsagittal images for CC segmentation, acquired from fetuses between 21 and 30 gestational weeks, and 120 trans\-thalamic images for choroid plexus segmentation, acquired from healthy fetuses between 20 and 22 gestational weeks\. The framework achieved high segmentation accuracy, with Dice coefficients of0\.81±0\.060\.81\\pm 0\.06for the CC and0\.76±0\.080\.76\\pm 0\.08for the choroid plexus \(CP\)\. However, its reliance on handcrafted features and a conventional classifier may limit its generalizability across heterogeneous acquisition settings\.

More recently, the work in\[[42](https://arxiv.org/html/2607.18283#bib.bib43)\]proposes FB\-ZWUNet, an end\-to\-end CC segmentation network integrating an attention module for feature enhancement, a wavelet Attention Module for optimized feature fusion, and a morphological constraint module for accurate edge and region capture\. The network was trained on the FB\-CC dataset, which consists of 1,336 annotated fetal brain mid\-sagittal US images, from 18 to 32 weeks of gestation, collected using three different US devices, achieving a Dice coefficient of 0\.874 and IoU of 0\.781 on a dedicated fetal brain corpus callosum dataset of 1,336 annotated mid\-sagittal US images\. The work in\[[24](https://arxiv.org/html/2607.18283#bib.bib46)\]proposes CC\-FocusNet, a DL framework that integrates automated region localization with an anatomy\-aware dual\-stream architecture for multi\-view analysis of CC abnormalities in prenatal US, achieving 0\.973 of accuracy in distinguishing normal from absent CC cases on an independent external test set of 93 cases\.

Although the works in\[[42](https://arxiv.org/html/2607.18283#bib.bib43),[24](https://arxiv.org/html/2607.18283#bib.bib46)\]represent important advances in fetal CC analysis, addressing segmentation and abnormality detection, respectively, both were developed and evaluated within relatively constrained acquisition settings\. As a result, their robustness to domain shifts induced by different scanners, imaging protocols, and clinical centers remains unclear\. Moreover, neither study considers privacy\-preserving collaborative training across institutions, which limits their applicability to realistic multi\-center deployment scenarios\.

More recently, FL has gained increasing attention in fetal US analysis, demonstrating its potential across a variety of clinical applications while preserving data privacy across institutions\. Existing studies have explored FL for tasks ranging from fetal standard plane detection through prototype\-based federated denoising approaches\[[9](https://arxiv.org/html/2607.18283#bib.bib80)\], to fetal biometry and multi\-structure abnormality detection with frameworks such as FB\-SeUNet\+\+\[[19](https://arxiv.org/html/2607.18283#bib.bib70)\], as well as prenatal congenital heart disease assessment, including the detection of interrupted aortic arch\[[12](https://arxiv.org/html/2607.18283#bib.bib79)\]\.

Despite these advances, FL applications in fetal neurosonography remain largely unexplored\. In particular, no previous study has investigated CC detection or localization in fetal US within a multi\-center, multi\-device setting\. Furthermore, the potential of FMs adapted through parameter\-efficient fine\-tuning strategies and trained under privacy\-preserving FL paradigms has not yet been examined for this task\. This represents a significant gap in the literature, given the clinical importance of CC assessment and the challenges associated with acquiring sufficiently large and diverse neurosonography datasets\.

![Refer to caption](https://arxiv.org/html/2607.18283v1/framework.png)Figure 2:Overview of the proposed federated learning framework\. Each client trains a local model consisting of a frozen DINOv2 backbonef​\(x;ω\)f\(x;\\omega\), a set of trainable LoRA adaptersαi\\alpha\_\{i\}, and a detection headϕi\\phi\_\{i\}\. Only the trainable parametersθi=\{αi,ϕi\}\\theta\_\{i\}=\\\{\\alpha\_\{i\},\\phi\_\{i\}\\\}are transmitted to the server, where they are aggregated via a server\-side strategy𝒜\\mathcal\{A\}and redistributed at the next communication round\. The backbone weightsω\\omegaremain frozen throughout and are never exchanged, substantially reducing communication overhead while preserving data privacy across clinical sites\.
## 3Materials and methods

In this section, we first present the overall framework \(Sec\.[3\.1](https://arxiv.org/html/2607.18283#S3.SS1)\)\. We then describe the dataset and its federated partitioning \(Sec\.[3\.2](https://arxiv.org/html/2607.18283#S3.SS2)\), followed by the configuration adopted for model training \(Sec\.[3\.3](https://arxiv.org/html/2607.18283#S3.SS3)\)\.

### 3\.1Proposed framework

The proposed framework, shown in Fig\.[2](https://arxiv.org/html/2607.18283#S2.F2), is built around three integrated components: i\) a Vision Transformer \(ViT\) backbone pretrained through self\-supervised learning, ii\) a lightweight single\-scale detection head derived from the YOLOv8 architecture, and iii\) a federated training strategy that coordinates learning across multiple clinical centres without exchanging raw fetal US images\.

Backbone\.DINOv2\[[32](https://arxiv.org/html/2607.18283#bib.bib53)\]is employed as a frozen feature extractor owing to its strong capability to learn transferable representations and its demonstrated robustness in fetal imaging domains\. In particular, previous studies have shown that DINO\-based representations generalise well across different fetal US datasets and acquisition settings\[[1](https://arxiv.org/html/2607.18283#bib.bib47)\]\. Given an input fetal US imagex∈ℝH×Wx\\in\\mathbb\{R\}^\{H\\times W\}, the image is resized and partitioned into non\-overlapping patches, which are linearly projected into a sequence of visual tokens and processed by the transformer encoder:

z=fDINOv2​\(x;ω\),z=f\_\{\\mathrm\{DINOv2\}\}\(x;\\omega\),\(1\)whereω\\omegadenotes the pretrained DINOv2 parameters andzzis the latent representation extracted from the final transformer layer\. To preserve the pretrained knowledge encoded in the backbone, the parametersω\\omegaremain frozen throughout training and are never transmitted during federated optimisation\. Adaptation to the CC localisation task is achieved through LoRA modules inserted into the query and value projection matrices of each attention block\. Rather than directly updating a pretrained weight matrixW∈ℝd×kW\\in\\mathbb\{R\}^\{d\\times k\}, LoRA constrains the update to a low\-rank residual:

W′=W\+Δ​W,Δ​W=B​AW^\{\\prime\}=W\+\\Delta W,\\qquad\\Delta W=BA\(2\)whereA∈ℝr×kA\\in\\mathbb\{R\}^\{r\\times k\}andB∈ℝd×rB\\in\\mathbb\{R\}^\{d\\times r\}are trainable matrices andr≪min⁡\(d,k\)r\\ll\\min\(d,k\)is the LoRA rank\. Therefore, the number of trainable parameters is reduced fromd×kd\\times ktor​\(d\+k\)r\(d\+k\)\.

![Refer to caption](https://arxiv.org/html/2607.18283v1/x1.png)Figure 3:Overview of the proposed architecture for corpus callosum localization in ultrasound images\. The input image is processed by a frozen DINOv2 backbone, in which LoRA layers are introduced to adapt the Transformer \(Trans\) features to the target task\. The resulting image embeddings are then forwarded to a YOLO\-like branch composed of an SPPF and C3f blocks \(\[[39](https://arxiv.org/html/2607.18283#bib.bib54)\]\), followed by the detection head to predict the bounding box\.Detection head\.To preserve a lightweight and computationally efficient single\-stage detection pipeline, the latent representationzzextracted by the frozen DINOv2 backbone is reshaped into a two\-dimensional feature map and processed by a compact YOLO\-like detection branch \(Fig\.[3](https://arxiv.org/html/2607.18283#S3.F3)\)\. This detection branch enables direct bounding\-box regression and confidence prediction, bypassing the need for region proposals or additional refinement stages\. Given that the backbone lacks a native multi\-resolution feature pyramid, a single detection head was preferred to minimize architectural complexity and facilitate federated optimization\. This streamlined approach is further justified by the single\-target nature of the task and the localized, compact appearance of the CC\. The detection branch can be expressed as

h=fdet​\(z;ϕ\),h=f\_\{\\mathrm\{det\}\}\(z;\\phi\),\(3\)whereϕ\\phidenotes the trainable parameters of the detection branch andhhis the refined feature representation used for localisation\.fdet​\(⋅\)f\_\{\\mathrm\{det\}\}\(\\cdot\)consists of a spatial pyramid pooling–fast \(SPPF\) block, followed by a convolutional layer, a C3f block, and a final detection layer, following the design principles of YOLOv8\[[39](https://arxiv.org/html/2607.18283#bib.bib54)\]\. The SPPF, convolutional, and C3f modules were adopted from YOLOv8 because they provide an effective balance between feature refinement and computational efficiency\. In particular, the SPPF block enlarges the receptive field and aggregates contextual information at multiple scales, which is beneficial in fetal ultrasound images where the corpus callosum may appear with variable size and surrounding anatomical context\. The convolutional layer and C3f block further refinezzwhile maintaining a limited number of trainable parameters and efficient gradient propagation\. A detailed description of these modules can be found in\[[39](https://arxiv.org/html/2607.18283#bib.bib54)\]\.

The final detection layer then mapshhto the predicted corpus callosum bounding boxb^\\hat\{b\}and corresponding confidence scorec^\\hat\{c\}:

\(b^,c^\)=g​\(h;ϕ\),\(\\hat\{b\},\\hat\{c\}\)=g\(h;\\phi\),\(4\)
whereg​\(⋅\)g\(\\cdot\)denotes the final prediction layer\. This YOLO\-inspired design was selected because it provides a favourable trade\-off between localisation accuracy and parameter efficiency, which is particularly important in the federated setting, where limiting the number of trainable and shared parameters is essential\.

Federated training\.LetNNclinical centres participate, indexed byi∈\{1,…,N\}i\\in\\\{1,\\dots,N\\\}\. The federally transmitted parameters comprise the LoRA adapters and the detection head weights:

θi=\{αi,ϕi\}=\{Ai,Bi,ϕi\},\\theta\_\{i\}=\\\{\\alpha\_\{i\},\\,\\phi\_\{i\}\\\}=\\\{A\_\{i\},\\,B\_\{i\},\\,\\phi\_\{i\}\\\},\(5\)
while the backboneω\\omegaremains frozen and is never transmitted\. At each communication roundtt, the server broadcasts the current global parametersθ\(t\)\\theta^\{\(t\)\}to all selected centres\. Each centre minimises a local detection objective overEEepochs:

ℒ=λbox​ℒbox\+λcls​ℒcls\+λdfl​ℒdfl\\mathcal\{L\}=\\lambda\_\{\\text\{box\}\}\\mathcal\{L\}\_\{\\text\{box\}\}\+\\lambda\_\{\\text\{cls\}\}\\mathcal\{L\}\_\{\\text\{cls\}\}\+\\lambda\_\{\\text\{dfl\}\}\\mathcal\{L\}\_\{\\text\{dfl\}\}\(6\)whereℒ​box\\mathcal\{L\}\{\\mathrm\{box\}\},ℒ​cls\\mathcal\{L\}\{\\mathrm\{cls\}\}, andℒdfl\\mathcal\{L\}\_\{\\mathrm\{dfl\}\}denote the bounding\-box regression, classification, and distribution focal losses, respectively, as defined in the YOLOv8 architecture\[[18](https://arxiv.org/html/2607.18283#bib.bib55)\]\. Upon completion, the updated parametersθi\(t\)\\theta\_\{i\}^\{\(t\)\}are returned to the server and aggregated via a server\-side strategy𝒜\\mathcal\{A\}:

θ\(t\+1\)=𝒜​\(\{θi\(t\)\}i=1N,\{\|𝒟i\|\}i=1N\)\.\\theta^\{\(t\+1\)\}=\\mathcal\{A\}\\\!\\left\(\\\{\\theta\_\{i\}^\{\(t\)\}\\\}\_\{i=1\}^\{N\},\\,\\\{\|\\mathcal\{D\}\_\{i\}\|\\\}\_\{i=1\}^\{N\}\\right\)\.\(7\)
The updated global parameters are redistributed at the next round\. Because neither raw images, annotations, nor backbone weights are ever exchanged, the approach substantially reduces communication overhead while preserving data privacy across centres\.

### 3\.2Dataset

The dataset, collected under the approval of the local Ethical Committees and with the informed consent of all participants, comprises 10,970 fetal US frames extracted from video acquisitions collected from 58 pregnant women during routine neurosonographic examinations\. The examinations were acquired at a native frame rate of 25 Hz in three Italian clinical centers: Chieti, Foggia, and Verona\. All included cases correspond to normal pregnancies, and no CC malformations were present in the dataset\. Image acquisition is achieved through the anterior fontanelle of the fetal skull, which provides an acoustic window for US\. The CC should be examined in the mid\-sagittal plane of the fetal brain, offering a clear view of all its components: splenium, isthmus, genu, corpus, and rostrum\. It appears as a thin, elongated, C\-shaped structure with two parallel echogenic margins, positioned above the cavum septi pellucidi \(CSP\)—an anechoic, fluid\-filled space that follows the shape of the CC\. Below these structures lies the tela choroidea of the third ventricle, seen as a hyperechoic, double\-concave line descending toward the cerebellar vermis\. The vermis, located at the bottom of the sagittal view, has a hyperechoic shell\-like appearance, with an internal tent\-shaped indentation representing the fourth ventricle\.

Among the collected frames, 4,561 contain a visible CC and were manually annotated by expert clinicians using a single bounding box enclosing the structure\. The remaining 6,409 frames do not contain the CC and were retained as background samples to increase the robustness of the detection task\. All annotations were provided in YOLO format considering a single target class corresponding to the CC\.

To simulate a realistic FL scenario, the dataset was further distributed across three clients, each corresponding to a distinct clinical site and US system\. Client 1 included 28 patients and 4,453 frames acquired using a GE Voluson E10 system; Client 2 included 16 patients and 3,391 frames acquired using Samsung HERA W9 and Samsung HERA W10 systems; Client 3 included 14 patients and 3,126 frames acquired using a Samsung HERA Z20 scanner\. This partitioning introduces a realistic domain shift across clients, driven by differences in acquisition hardware, imaging settings, and local clinical protocols\.

![Refer to caption](https://arxiv.org/html/2607.18283v1/clients1.png)Figure 4:Representative fetal ultrasound frames grouped by client, highlighting inter\-domain variability across the dataset\. The images show differences in field\-of\-view, fetal pose, and apparent image scale, leading to variability in the size and position of the corpus callosum\.Table 1:Dataset statistics across federated clients\. For each site, we report the number of images, the proportion of positive samples \(CC present\), image intensity statistics, spatial resolution, and geometric properties of the CC bounding box \(normalized area, aspect ratio AR, and vertical center position Ctr\-Y\)\. Substantial inter\-client heterogeneity is observed across all dimensions\.Image StatisticsBBox StatisticsClient\#Patients\#ImagesPos\.IntensityStdResolution \(W×\\timesH\)AreaARCtr\-YClient 1284,4530\.4771\.05±26\.3471\.05\\pm 26\.3467\.02±7\.1067\.02\\pm 7\.101278±219×8521278\{\\pm\}219\\times 8520\.042±0\.0220\.042\{\\pm\}0\.0221\.13±0\.371\.13\{\\pm\}0\.370\.49±0\.050\.49\{\\pm\}0\.05Client 2163,3910\.3958\.96±9\.6058\.96\\pm 9\.6070\.70±3\.4470\.70\\pm 3\.441062±348×642±1431062\{\\pm\}348\\times 642\{\\pm\}1430\.034±0\.0110\.034\{\\pm\}0\.0111\.08±0\.261\.08\{\\pm\}0\.260\.55±0\.060\.55\{\\pm\}0\.06Client 3143,1260\.3760\.79±14\.2560\.79\\pm 14\.2568\.35±4\.3468\.35\\pm 4\.341232×9241232\\times 9240\.039±0\.0220\.039\{\\pm\}0\.0221\.20±0\.261\.20\{\\pm\}0\.260\.47±0\.100\.47\{\\pm\}0\.10Total5810,9700\.41—

Inter\-client heterogeneity is substantial and multi\-factorial, as shown in Fig\.[4](https://arxiv.org/html/2607.18283#S3.F4)and quantified in Table[1](https://arxiv.org/html/2607.18283#S3.T1)\. Image appearance differs markedly across sites, with well\-separated intensity distributions confirmed by pairwise Wasserstein distances\. Spatial resolution is fixed within Clients 1 and 3, whereas Client 2 exhibits pronounced variability, reflecting heterogeneous acquisition settings\. Geometric properties of the annotated structures also vary across clients, with systematic differences in the apparent scale and vertical positioning of the CC\. Conversely, bounding\-box aspect ratios remain consistent across sites, suggesting that the morphological appearance of the CC is largely preserved despite inter\-domain variability\.

Beyond these statistical differences, the images are affected by challenges intrinsic to fetal US acquisition, including acoustic shadowing, speckle noise, and variable image contrast\. The apparent position of the CC within the image frame varies considerably depending on probe orientation and patient\-specific factors, as does the degree of zoom, ranging from close\-up views centred on the fetal brain to wider fields of view encompassing surrounding anatomical structures\. In some cases, the CC is only partially visible or poorly delineated, further increasing the difficulty of the localisation task\.

### 3\.3Training settings

Raw US frames were processed through a standardized preprocessing pipeline prior to training\. Scanner\-specific overlay artefacts, including textual information bars and control panels, were removed by zeroing fixed peripheral regions of the image, preventing the model from exploiting cues unrelated to the target anatomy\.

The dataset was divided at the patient level into training, validation, and test subsets for each federated client, preventing data leakage between splits\. Client 1 contributed 2,684 training images \(15 patients\), 644 validation images \(3 patients\), and 1,125 test images \(10 patients\)\. Client 2 contributed 2,031 training images \(9 patients\), 454 validation images \(3 patients\), and 906 test images \(4 patients\), whereas Client 3 contributed 2,165 training images \(9 patients\), 472 validation images \(2 patients\), and 683 test images \(3 patients\)\. The complete distribution of patients and images across clients and dataset partitions is reported in Table[2](https://arxiv.org/html/2607.18283#S3.T2)\.

Table 2:Patient\- and image\-level distribution across federated clients and data splits\.Input images were resized to544×544544\\times 544pixels prior to being fed to the backbone\. All experiments were conducted on NVIDIA A100 GPUs with Automatic Mixed Precision \(AMP\) to reduce memory footprint\. Each federated simulation runs for 20 communication rounds withE=4E=4local epochs per round\. Local optimisation was performed using AdamW with a learning rate of1×10−41\\times 10^\{\-4\}, weight decay of5×10−45\\times 10^\{\-4\}, and a batch size of 64\. The LoRA rank was set tor=8r=8, with scaling factorαLoRA=16\\alpha\_\{\\mathrm\{LoRA\}\}=16\. The detection loss weights were set toλbox=7\.5\\lambda\_\{\\mathrm\{box\}\}=7\.5,λcls=0\.5\\lambda\_\{\\mathrm\{cls\}\}=0\.5, andλdfl=1\.5\\lambda\_\{\\mathrm\{dfl\}\}=1\.5, with an IoU threshold of 0\.6 and a confidence threshold of1×10−31\\times 10^\{\-3\}\. At each communication round, allNN= 3 clients participated in the federated update\.

Table 3:Summary of the experimental configurations\. For each backbone and initialization strategy, all combinations of adaptation strategy and FL aggregation method were evaluated\.

## 4Experimental Protocol

### 4\.1Comparisons

To thoroughly assess the effectiveness of the proposedFedCCframework, we designed a comprehensive comparison involving different backbone architectures, pretraining strategies, adaptation schemes, and federated aggregation methods\. In particular, the experimental analysis was aimed at answering three main questions: \(i\) whether, in a federated setting, sharing only parameter\-efficient adaptation modules is preferable to updating and sharing the entire backbone; \(ii\) whether natural\-image, generic US, or fetal US\-specific pretraining provides the most suitable initialization for fetal CC detection; and \(iii\) whether a lightweight single\-head detector can achieve performance comparable to, or better than, a more complex multi\-head architecture\.

We considered two variants of the DINOv2 backbone\. Both were based on the small architecture \(ViT\-S/14\), which was selected to ensure a fair comparison within the low\-resource federated setting addressed in this work, and to prevent performance differences from being driven by model scaling rather than by the proposed adaptation strategy\. The first, denoted asD​I​N​O​v​2b​a​s​eDINOv2\_\{base\}, was initialized with the original weights pretrained on large\-scale natural image collections\[[32](https://arxiv.org/html/2607.18283#bib.bib53)\]\. The second, referred to asD​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}, used the same architecture but initialized from a version further adapted through self\-supervised learning on a publicly available fetal US dataset derived from\[[5](https://arxiv.org/html/2607.18283#bib.bib56)\]\. Comparing these two variants allows us to isolate the contribution of fetal\-specific pretraining with respect to standard natural\-image initialization\. Finally, all experiments were performed using both a lightweight single\-head detector and its multi\-head counterpart\. This design allows the effect of the detection architecture to be assessed independently of the backbone initialization and federated optimization strategy, providing a clearer understanding of the contribution of each component within the proposed framework\.

To investigate the effect of different FM families, we also considered the Segment Anything Model \(SAM\), which, although originally pretrained on natural images, has recently been successfully adopted in fetal imaging applications\[[8](https://arxiv.org/html/2607.18283#bib.bib37),[46](https://arxiv.org/html/2607.18283#bib.bib57),[23](https://arxiv.org/html/2607.18283#bib.bib58)\]\. In addition, we evaluated UltraSAM, an US\-specific adaptation of SAM explicitly designed to better model the texture, speckle noise, and anatomical properties of sonographic images\[[15](https://arxiv.org/html/2607.18283#bib.bib59)\]\. Finally, we considered UltraFedFM\[[16](https://arxiv.org/html/2607.18283#bib.bib60)\], a recent privacy\-preserving US foundation model pretrained in a federated manner across multiple institutions\. UltraFedFM is particularly relevant in this context, as it combines US\-specific representations with a training paradigm intrinsically aligned with FL\.

For all considered backbones \(D​I​N​O​v​2b​a​s​eDINOv2\_\{base\},D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}, SAM, UltraSAM, and UltraFedFM\), we systematically evaluated three adaptation strategies: \(i\) full fine\-tuning \(Full FT\), in which all backbone parameters were updated and shared across clients; \(ii\) encoder freezing \(F​r​e​e​z​eFreeze\), in which the pretrained encoder remained fixed and only the detection head was optimized; and \(iii\) the proposed LoRA\-based strategy \(p​r​o​p​o​s​e​dproposed\), in which only a small set of trainable low\-rank adapters was introduced and shared among clients\. This comparison allows us to quantify whether parameter\-efficient adaptation is more effective and communication\-efficient than conventional fine\-tuning in a federated setting\.

Each adaptation strategy was combined with two standard FL aggregation methods, namely FedAvg and FedProx\[[29](https://arxiv.org/html/2607.18283#bib.bib66),[25](https://arxiv.org/html/2607.18283#bib.bib67)\]\. FedAvg was adopted as the reference strategy, as it is the most widely used optimization approach in FL\. FedProx was additionally considered to account for the strong heterogeneity expected across fetal US datasets acquired at different institutions, scanners, and gestational ages\. By introducing a proximal regularization term during local optimization, FedProx can mitigate client drift and improve convergence under non\-IID conditions\. Overall, for each backbone and initialization strategy, we evaluated all six possible federated configurations resulting from the combination of the three adaptation schemes and the two aggregation methods\.

As representative convolutional baselines, we additionally included three variants ofY​O​L​O​26YOLO26\(nano, small, and medium\)\[[36](https://arxiv.org/html/2607.18283#bib.bib61)\], fully fine\-tuned from COCO\-pretrained weights\. These models were selected because the YOLO family is widely recognized for providing a favorable trade\-off between detection accuracy and computational efficiency, requiring substantially fewer parameters than large FMs\. Including YOLO26 therefore enables us to assess whether the proposed parameter\-efficient adaptation strategy can outperform lightweight convolutional models while maintaining a similarly efficient parameter footprint\.

### 4\.2Performance settings

The detection performance of the proposed framework was evaluated using four standard object detection metrics: Precision, Recall, F1\-score, and mean Average Precision at an Intersection over Union \(IoU\) threshold of 0\.5 \(mAP@50\)\.

Precision quantifies the fraction of predicted bounding boxes that correctly identify the CC, reflecting the model’s tendency to avoid spurious detections\. Recall measures the fraction of ground\-truth CC instances that are successfully localized, thus characterizing detection sensitivity\. The F1\-score was computed as their harmonic mean to provide a single balanced summary statistic:

Precision=T​PT​P\+F​P,Recall=T​PT​P\+F​N\\displaystyle\\text\{Precision\}=\\frac\{TP\}\{TP\+FP\},\\qquad\\text\{Recall\}=\\frac\{TP\}\{TP\+FN\}\(1\)F1=2⋅Precision⋅RecallPrecision\+Recall\\displaystyle\\text\{F1\}=2\\cdot\\frac\{\\text\{Precision\}\\cdot\\text\{Recall\}\}\{\\text\{Precision\}\+\\text\{Recall\}\}\(2\)whereT​PTP,F​PFP, andF​NFNdenote the number of true positives, false positives, and false negatives, respectively\.

To obtain a threshold\-independent summary of localization performance, the Average Precision \(AP\) was computed as the area under the Precision–Recall curve, and the mAP@50 as its mean across allNNevaluated classes:

AP=∫01p​\(r\)​𝑑r,mAP@50=1N​∑i=1NA​PiIoU=0\.5\\text\{AP\}=\\int\_\{0\}^\{1\}p\(r\)\\,dr,\\qquad\\text\{mAP@50\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}AP\_\{i\}^\{\\,\\text\{IoU\}=0\.5\}\(3\)wherep​\(r\)p\(r\)denotes precision as a function of recall\. Since the detection task involves a single anatomical target \(CC\)N=1N=1and mAP@50 reduces to AP@50\. It is nonetheless reported under the standard mAP notation for consistency with the object detection literature\[[47](https://arxiv.org/html/2607.18283#bib.bib52)\]\.

In addition, since the proposed framework relies on parameter\-efficient adaptation, the different methods were also compared in terms of the number of trainable parameters\.

Table 4:Core comparison study onD​I​N​O​v​2b​a​s​eDINOv2\_\{base\}in the federated setting\. We evaluate adaptation strategy, federated optimization, and initialization\. Metrics are reported per client \(C1, C2, C3\) and averaged \(AVG\)\.Boldindicates the best result per metric\.mAP50PrecisionRecallF1ComponentVariantC1C2C3\\cellcolorlightgrayAVGC1C2C3\\cellcolorlightgrayAVGC1C2C3\\cellcolorlightgrayAVGC1C2C3\\cellcolorlightgrayAVGFederated \- FedProxD​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}Freeze0\.6700\.9610\.4720\.7010\.8360\.8500\.4740\.7200\.5840\.9540\.5590\.6990\.6870\.8990\.5130\.700Full FT0\.7210\.9370\.4670\.7080\.8270\.8590\.5060\.7310\.6600\.9390\.5040\.7010\.7340\.8970\.5050\.712Proposed0\.7410\.9650\.7270\.8110\.8270\.9490\.6540\.8100\.6240\.8820\.7570\.7540\.7110\.9140\.7020\.776D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}Freeze0\.7050\.9610\.5480\.7380\.7130\.9120\.5540\.7260\.6840\.9030\.5930\.7270\.6980\.9070\.5730\.726Full FT0\.7420\.9690\.6010\.7710\.7350\.9350\.6220\.7640\.7020\.9260\.6410\.7560\.7180\.9300\.6310\.760Proposed0\.8140\.9730\.7470\.8450\.7520\.9220\.6830\.7860\.7390\.9010\.6830\.7740\.7520\.9220\.6830\.786Federated \- FedAvgD​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}Freeze0\.7110\.9310\.4330\.6920\.8120\.8540\.4570\.7080\.6600\.9350\.4850\.6940\.7280\.8930\.4710\.697Full FT0\.6880\.9700\.5070\.7220\.8300\.8570\.6320\.7730\.6170\.9730\.4410\.6770\.7080\.9120\.5200\.713Proposed0\.7410\.9780\.8000\.8390\.7050\.8920\.7800\.7920\.6720\.9660\.7310\.7900\.6880\.9270\.7550\.790D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}Freeze0\.6810\.9570\.5070\.7150\.6920\.9050\.5160\.7040\.6570\.8920\.5640\.7040\.6740\.8980\.5390\.704Full FT0\.6300\.9530\.5790\.7210\.7060\.9240\.5690\.7330\.6710\.9140\.6150\.7330\.6880\.9240\.5870\.733Proposed0\.8060\.9770\.7880\.8570\.7610\.9570\.8460\.8550\.7460\.8900\.6460\.7610\.7530\.9220\.7320\.803

Table 5:Impact of detection head design forD​I​N​O​v​2b​a​s​eDINOv2\_\{base\}pretrained\. We compare single\-scale and multi\-scale \(2\-scale, 3\-scale\) detection heads using their best\-performing configurations\. Metrics are reported per client \(C1, C2, C3\) and averaged \(AVG\)\.Boldindicates the best result per metric\.Table 6:Federated performance of the considered FM\-based backbones under FedAvg and FedProx\. Best result for each metric is highlighted in bold\.mAP50PrecisionRecallF1ModelTuningC1C2C3\\cellcolorlightgrayAVGC1C2C3\\cellcolorlightgrayAVGC1C2C3\\cellcolorlightgrayAVGC1C2C3\\cellcolorlightgrayAVGCentralizedYOLO26nFull FT0\.7770\.9640\.1650\.6350\.7520\.9370\.1990\.6300\.6790\.8630\.3010\.6140\.71400\.89860\.24020\.617YOLO26sFull FT0\.7490\.9410\.5610\.7500\.7660\.9510\.6010\.7730\.6600\.8820\.5550\.6990\.7090\.9150\.5770\.734YOLO26mFull FT0\.6220\.9210\.5490\.6970\.7100\.9030\.5860\.7330\.5500\.8560\.5210\.6420\.6200\.8790\.5520\.684D​I​N​O​v​3b​a​s​eDINOv3\_\{base\}Full FT0\.7320\.9550\.6360\.7740\.7760\.8840\.7040\.7880\.6770\.9390\.5520\.7230\.7230\.9110\.6190\.751D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}proposed0\.7240\.9770\.4220\.7080\.7850\.9520\.5120\.7500\.6610\.9430\.4780\.6940\.7180\.9470\.4950\.720D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}proposed0\.7520\.9830\.6080\.7810\.7970\.9560\.6330\.7950\.6450\.9200\.5370\.7010\.7130\.9380\.5810\.744Federated – FedProxYOLO26nFull FT0\.7880\.9590\.6730\.8070\.8540\.9080\.7590\.8400\.6980\.9390\.6010\.7460\.7690\.9230\.6700\.787YOLO26sFull FT0\.8340\.9750\.5180\.7750\.8590\.9410\.5630\.7870\.8010\.9010\.4930\.7320\.8290\.9210\.5250\.758YOLO26mFull FT0\.7910\.9660\.7340\.8300\.8320\.9230\.7630\.8390\.7280\.8750\.7280\.7770\.7770\.8980\.7280\.801SAMproposed0\.6890\.9520\.4480\.6960\.7840\.9030\.4720\.7200\.5900\.8890\.5520\.6770\.6730\.8960\.5080\.692D​I​N​O​v​3b​a​s​eDINOv3\_\{base\}Feat\.Ext0\.7880\.9740\.7110\.8240\.8180\.9080\.6620\.7960\.6660\.9620\.6760\.7680\.7340\.9340\.6690\.779UltraFedFMFull FT0\.6750\.9690\.2430\.6290\.8580\.8450\.3590\.6870\.5900\.9770\.4490\.6720\.6990\.9060\.3990\.668UltraSAMFull FT0\.7960\.9790\.6140\.7960\.8080\.9730\.5420\.7750\.6830\.9010\.6100\.7320\.7410\.9360\.5740\.750D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}proposed0\.7410\.9650\.7270\.8110\.8270\.9490\.6540\.8100\.6240\.8820\.7570\.7540\.7110\.9140\.7020\.776D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}Proposed0\.8140\.9730\.7470\.8450\.7520\.9220\.6830\.7860\.7390\.9010\.6830\.7740\.7520\.9220\.6830\.786Federated – FedAvgYOLO26nFull FT0\.7900\.9720\.5940\.7850\.8770\.9300\.6370\.8150\.6780\.9700\.5660\.7380\.7650\.9490\.5990\.771YOLO26sFull FT0\.8410\.9600\.4580\.7530\.8920\.9130\.4390\.7480\.7310\.9320\.5990\.7540\.8040\.9220\.5070\.744YOLO26mFull FT0\.8020\.9770\.4160\.7320\.8690\.9670\.4570\.7640\.7270\.8900\.5150\.7100\.7910\.92660\.48450\.7342SAMproposed0\.6730\.9550\.4270\.6850\.7780\.8990\.4540\.7100\.5670\.8940\.5380\.6660\.6560\.8960\.4920\.681D​I​N​O​v​3b​a​s​eDINOv3\_\{base\}proposed0\.7730\.9720\.7240\.8230\.7990\.9450\.7070\.8170\.6560\.9090\.6900\.7520\.7210\.9270\.6980\.782UltraFedFMproposed0\.6520\.9650\.1760\.5980\.7680\.9040\.2790\.6500\.5500\.9200\.4710\.6470\.6410\.9120\.3500\.634UltraSAMFull FT0\.8340\.9800\.7420\.8520\.8770\.9600\.6360\.8240\.6780\.9010\.7200\.7670\.7650\.9300\.6760\.790D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}proposed0\.7410\.9780\.8000\.8390\.7050\.8920\.7800\.7920\.6720\.9660\.7310\.7900\.6880\.9270\.7550\.790D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}proposed0\.8060\.9770\.7880\.8570\.7610\.9570\.8460\.8550\.7460\.8900\.6460\.7610\.7530\.9220\.7320\.803

![Refer to caption](https://arxiv.org/html/2607.18283v1/x2.png)Figure 5:Qualitative comparison of corpus callosum localization across three federated clients and six models:D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}\(Centralized\), YOLO26m \(FedProx\),D​I​N​O​v​3b​a​s​eDINOv3\_\{base\}Proposed \(FedProx\),D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}\(FedAvg\), UltraSAM \(FedAvg\), andD​I​N​O​v​2b​a​s​eDINOv2\_\{base\}Proposed \(FedAvg\)\. Ground truth annotations are shown in red; model predictions in green\. Each row pair corresponds to one federated client, reflecting distinct acquisition domains\.![Refer to caption](https://arxiv.org/html/2607.18283v1/figure_publication.png)Figure 6:Trainable parameter count across model families and adaptation strategies\. Each bar reports the number of trainable parameters \(in millions, M\) for a given combination of backbone architecture and fine\-tuning strategy\. Four model families are compared: DINOv2, UltraSAM, UltraFedFM, and YOLO26 \(variants n/s/m\)\. For the first three families, three adaptation strategies are evaluated: Low\-Rank Adaptation \(LoRA\), encoder freezing with trainable decoder \(Freeze\), and full fine\-tuning \(Full FT\)\. Parameter counts are reported as floating\-point millions\. Hatch patterns distinguish model families to ensure readability in greyscale reproduction\.

## 5Results

Table[4](https://arxiv.org/html/2607.18283#S4.T4)reports the comparison between the two DINOv2 initializations,D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}andD​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}, under different adaptation strategies and federated optimization methods\. Across both FedAvg and FedProx,D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}consistently achieves higher average performance compared toD​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}\. Under FedAvg, the best results are obtained with the proposed adaptation strategy applied toD​I​N​O​v​2b​a​s​eDINOv2\_\{base\}, reaching an average mAP50 of 0\.857 and an F1\-score of 0\.803\. Similarly, under FedProx, the same configuration achieves the highest performance, with an average mAP50 of 0\.845 and F1\-score of 0\.786\. ForD​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}, the proposed strategy improves performance compared to both freezing and full FT across most metrics and clients\. Under FedAvg, it achieves an average mAP50 of 0\.839 and F1\-score of 0\.790, while under FedProx it reaches 0\.811 and 0\.776, respectively\. Given the superior performance ofD​I​N​O​v​2b​a​s​eDINOv2\_\{base\}across all settings, Table[5](https://arxiv.org/html/2607.18283#S4.T5)further investigates this configuration by analyzing the impact of the detection head design\. The single\-scale configuration achieves the highest performance across all metrics, with an average mAP50 of 0\.857 and F1\-score of 0\.803\. The 2\-scale and 3\-scale variants show lower results, with average mAP50 values of 0\.757 and 0\.741, respectively\.

Table[6](https://arxiv.org/html/2607.18283#S4.T6)reports the performance of all considered FM\-based backbones, comparing centralized training with federated optimization under both FedProx and FedAvg\. In the centralized setting, YOLO\-based models achieve competitive performance, with YOLO26s reaching the highest average mAP50 \(0\.750\) and F1\-score \(0\.734\)\. Among transformer\-based models,D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}slightly outperformsD​I​N​O​v​2b​a​s​eDINOv2\_\{base\}, achieving an average mAP50 of 0\.781 and F1\-score of 0\.744 when combined with the proposed adaptation strategy\. Under federated optimization, all models show an overall improvement compared to the centralized setting\. In particular, under FedAvg,D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}with the proposed strategy achieves the best overall performance, reaching an average mAP50 of 0\.857 and F1\-score of 0\.803\. Similarly,D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}achieves strong results, with an average mAP50 of 0\.839 and F1\-score of 0\.790\. Among the alternative FMs, UltraSAM and UltraFedFM show competitive performance, especially under FedAvg\. UltraSAM achieves an average F1\-score of 0\.790, comparable toD​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}, while UltraFedFM shows lower performance overall, particularly in terms of mAP50\. SAM exhibits lower performance compared to both DINOv2 variants and UltraSAM in all configurations\. Similarly, YOLO\-based models, despite strong centralized results, do not reach the same performance levels in the federated setting\. Under FedProx, a similar trend is observed\.D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}with the proposed strategy achieves the highest average performance \(mAP50 = 0\.845, F1 = 0\.786\), followed byD​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}\(mAP50 = 0\.811, F1 = 0\.776\)\. Other models, including UltraSAM and YOLO variants, achieve lower but comparable results\.

Figure[5](https://arxiv.org/html/2607.18283#S4.F5)presents qualitative results across the three federated clients for a representative set of models:D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}\(Centralized\), YOLO26m \(FedProx\),D​I​N​O​v​3b​a​s​eDINOv3\_\{base\}with the proposed strategy \(FedProx\),D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}\(FedAvg\) with the proposed strategy, UltraSAM \(FedAvg\), andD​I​N​O​v​2b​a​s​eDINOv2\_\{base\}with the proposed strategy \(FedAvg\)\. Ground truth annotations are shown in red, while model predictions are shown in green\. Client 1 presents several inherently challenging examples in which the CC is only partially visible or exhibits low contrast against the surrounding fetal brain tissue\. However, most of the models exhibits good results\. The proposed LoRA\-based strategies demonstrate comparatively more stable predictions, better aligned with the ground truth even under low\-visibility conditions\. Client 2 represents the least challenging domain in the federation\. The CC is generally well\-delineated and clearly distinguishable from adjacent anatomical structures, resulting in consistent and accurate predictions across most models\. This client\-level advantage is reflected in the quantitative results, where all methods achieve their highest per\-client scores\. Client 3 introduces a distinct type of difficulty: while the CC is frequently present within the field of view, its boundaries tend to merge with surrounding fetal soft tissue, reducing edge contrast and increasing anatomical ambiguity\. This domain shift, driven by differences in scanner characteristics, fetal positioning, or gestational stage, visibly degrades the predictions of several baseline models, which either miss the structure entirely or produce poorly localized outputs\. The proposed federated adaptation methods, by contrast, retain greater localization accuracy, suggesting improved robustness to appearance variability across acquisition domains\.

Figure[6](https://arxiv.org/html/2607.18283#S4.F6)reports the number of trainable parameters across the different model families and adaptation strategies\. For all FMs, the proposed LoRA\-based approach drastically reduces the number of trainable parameters compared with full FT, while remaining close to the freezing configuration\. Specifically, for DINOv2, LoRA requires only 2\.9M trainable parameters, compared with 24\.4M for full FT and 2\.3M when the encoder is frozen\. A similar behavior is observed for UltraSAM, where LoRA involves 3\.3M trainable parameters versus 92\.5M for full FT, and for UltraFedFM, with 4\.7M parameters compared with 89\.3M under the full FT setting\. Overall, the proposed strategy consistently achieves a substantial reduction in trainable parameters across all considered backbones, offering a parameter\-efficient alternative\. In contrast, the number of trainable parameters in the YOLO26 models varies according to model scale, ranging from 2\.4M for the nano variant to 20\.4M for the medium version\.

## 6Discussion

Automatic detection of CC in fetal neurosonography is a clinically relevant yet challenging task, requiring models that are robust to scanner\-dependent variability, effective under limited supervision, and compatible with privacy\-preserving clinical deployment\. The results reported in Tables[4](https://arxiv.org/html/2607.18283#S4.T4),[6](https://arxiv.org/html/2607.18283#S4.T6), and[5](https://arxiv.org/html/2607.18283#S4.T5)indicate that these requirements can be effectively addressed by combining a self\-supervised transformer backbone with parameter\-efficient adaptation in a federated setting\. The comparison betweenD​I​N​O​v​2b​a​s​eDINOv2\_\{base\}andD​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}provides an interesting insight into the role of domain\-specific pretraining in federated fetal US analysis\. Although fetal\-specific initialization could be expected to provide a better starting point for CC detection,D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}consistently achieved higher average performance across both FedAvg and FedProx\. This result suggests that, in this setting, the breadth and diversity of the original DINOv2 pretraining may be more beneficial than a subsequent domain\-specific adaptation performed on a substantially smaller fetal US dataset\. A possible explanation is that large\-scale self\-supervised pretraining on highly diverse natural images allowsD​I​N​O​v​2b​a​s​eDINOv2\_\{base\}to learn general visual representations that remain transferable to fetal US, especially when coupled with a parameter\-efficient adaptation strategy\. Conversely, the fetal\-pretrained variant may have shifted the representation toward the characteristics of the pretraining fetal dataset, potentially reducing its generality when applied to CC data\. This effect may be particularly relevant in FL, where each client can exhibit scanner\-, protocol\-, and population\-dependent appearance differences\. Therefore, while domain\-specific pretraining remains a promising direction, these results indicate that its effectiveness depends not only on domain similarity, but also on the scale, diversity, and representativeness of the pretraining data\.

The results also show that the proposed LoRA\-based adaptation consistently provides the best or near\-best performance for both initializations and aggregation strategies\. This suggests that updating a limited set of trainable parameters can better preserve the transferable representations learned during pretraining while still enabling task\-specific adaptation\. In contrast, full FT does not consistently improve performance and, in some cases, leads to lower results, supporting the use of parameter\-efficient adaptation in low\-resource and heterogeneous federated settings\. The comparison with other FM backbones further highlights differences in performance across architectures\. In particular, SAM consistently achieves lower results across all configurations\. A possible explanation is that SAM is primarily designed as a prompt\-based segmentation model, where performance depends on the availability and quality of input prompts\. In the considered detection setting, where no explicit prompt information is provided, its representations may be less aligned with the task requirements\. In addition, its pretraining on natural images may limit its ability to capture the specific appearance characteristics of fetal US data\. UltraSAM represents the closest FM\-based competitor, achieving performance comparable to the proposed approach under FedAvg\. However, its performance decreases more markedly on the most challenging client \(C3\), where both mAP50 and F1\-score drop compared toD​I​N​O​v​2b​a​s​eDINOv2\_\{base\}\. This behavior may be related to the higher model capacity and the increased number of trainable parameters, which could make the model more sensitive to limited data availability and domain shifts across clients\. Similarly, UltraFedFM does not achieve performance comparable to DINOv2\-based models\. Although it benefits from US\-specific pretraining, its relatively large number of trainable parameters may limit its robustness in the considered federated setting, where each client provides only a subset of the overall data distribution\. This can result in less stable generalization across heterogeneous centers\. Compared with CNN\-based YOLO26 baselines,D​I​N​O​v​2DINOv2exhibits more consistent performance across clients\. This difference is unlikely to be explained by model capacity alone, as YOLO26m has a comparable number of parameters\. Instead, it suggests that the self\-supervised representations learned by DINOv2 are more transferable under scanner\-dependent appearance shifts\. This is consistent with the design of DINOv2, which leverages large\-scale self\-supervised pretraining to learn general\-purpose visual features\[[32](https://arxiv.org/html/2607.18283#bib.bib53)\], and with prior evidence highlighting the impact of scanner, vendor, and protocol variability on model generalization in US imaging\[[11](https://arxiv.org/html/2607.18283#bib.bib5),[40](https://arxiv.org/html/2607.18283#bib.bib11)\]\.

The comparison between FedAvg and FedProx should be interpreted in a model\-dependent manner\. FedProx was introduced to improve federated optimization under statistical and systems heterogeneity by constraining local updates toward the global model\[[25](https://arxiv.org/html/2607.18283#bib.bib67)\]\. In our experiments, this proximal regularization improved some configurations, most notably the CNN\-based YOLO26m baseline, whose average mAP50 increased from 0\.732 under FedAvg to 0\.830 under FedProx\. This suggests that models trained through full fine\-tuning may benefit from an explicit constraint on local optimization when client distributions are heterogeneous\. However, the same trend was not observed for the proposed LoRA\-based DINOv2 configuration\. One possible explanation is that LoRA already constrains local adaptation by limiting trainable updates to a compact low\-rank parameter space\. In this setting, the additional proximal term may provide limited extra regularization and may partially reduce the flexibility required to adapt to the smallest and most shifted client\. This interpretation is consistent with prior observations that client heterogeneity, data imbalance, and client drift can strongly influence federated optimization dynamics\[[29](https://arxiv.org/html/2607.18283#bib.bib66),[20](https://arxiv.org/html/2607.18283#bib.bib78)\]\. From a deployment perspective, the fact that FedAvg provided the best trade\-off for the proposed LoRA\-based configuration is practically relevant\. FedAvg avoids the need to tune the proximal coefficient required by FedProx, reduces local optimization complexity, and simplifies deployment across heterogeneous clinical infrastructures\. This advantage is further reinforced by the communication efficiency of the proposed approach\. The proposed federated configuration outperformed its centralized counterpart trained on pooled data\. Although centralized training is often considered an upper\-bound setting because all data are available during optimization, this assumption does not necessarily hold when the pooled dataset is heterogeneous and imbalanced across domains\. In the present study, centralizedD​I​N​O​v​2b​a​s​eDINOv2\_\{base\}achieved an average mAP50 of 0\.708 and degraded substantially on C3, reaching only 0\.422 mAP50\. By contrast, the federatedD​I​N​O​v​2b​a​s​eDINOv2\_\{base\}\+ LoRA model trained with FedAvg achieved 0\.857 average mAP50 and retained 0\.788 mAP50 on C3\. This suggests that federated optimization may have acted as a form of structured multi\-domain training, where each client contributes domain\-specific updates before aggregation\. In this setting, the repeated local training and aggregation cycle may help the model preserve representations that are useful across scanners, rather than overfitting to the dominant acquisition domain in the pooled dataset\. However, this interpretation should be considered specific to the dataset and partitioning used in this study, and should be further validated on larger multi\-center cohorts\.

Several limitations must be acknowledged\. First, although the dataset is multi\-center and multi\-device, it includes 58 patients from three clinical sites\. Validation on larger and more diverse cohorts, including additional centers, US manufacturers, acquisition protocols, and gestational\-age ranges, is necessary before considering clinical translation\[[10](https://arxiv.org/html/2607.18283#bib.bib77)\]\. Second, all included cases correspond to normal pregnancies, and no CC malformations were present\. Therefore, the current study evaluates localization robustness but does not address abnormality detection or classification\. Third, the framework performs bounding\-box detection rather than segmentation, preventing direct extraction of biometric measurements such as CC length, thickness, area, or subregional morphology\. Fourth, the federated setting was simulated under controlled conditions, with all clients participating at each communication round\. Real cross\-institutional deployments may involve asynchronous updates, client dropout\.

## 7Conclusions

In this work, we proposed a federated and parameter efficient framework for automated detection of CC in fetal neurosonography, addressing the joint challenges of privacy preservation, inter\-scanner domain shift, and communication efficiency\. Our central contribution is the combination of a DINOv2 backbone with LoRA in a federated pipeline, which we evaluated on a multi\-site dataset of 10,970 ultrasound frames from 58 patients acquired with four distinct scanners\.

Our results suggest that LoRA serves a dual role in this context: it reduces the per\-round parameter transmission by8\.5×8\.5\\timesand provides a constrained adaptation space that may improve the stability of federated training under heterogeneous client distributions\.

Future work will extend this framework toward automated CC segmentation for biometric quantification and validation on larger and more diverse fetal US datasets collected across multiple clinical centers\. Expanding the number of participating institutions and increasing dataset heterogeneity will enable a more comprehensive assessment of model robustness and generalizability across different acquisition settings and patient populations\. In addition, a key future direction will be the detection of other anatomical subregions of the CC, such as the rostrum and splenium, as accurate localization of these structures may provide valuable structural and developmental information\[[4](https://arxiv.org/html/2607.18283#bib.bib62),[26](https://arxiv.org/html/2607.18283#bib.bib63)\]\.

## 8Aknowledgements

This work was partially supported by the Italian Fund for Applied Sciences \(FISA\), grant no\. FISA2022\-00696\.

## 9Code Availability

The source code will be made publicly available upon acceptance of the manuscript\.

## 10Supplementary Material

Table 7:Expanded comparison of the experimental configurations associated with the main comparison study\. Metrics are reported per client \(C1, C2, C3\) and averaged \(AVG\)\. Only configurations directly related to the models and settings discussed in the main comparison are included\.mAP50PrecisionRecallF1ModelTuningC1C2C3AVGC1C2C3AVGC1C2C3AVGC1C2C3AVGCentralizedYOLO26nFull FT0\.7770\.9640\.1650\.6350\.7520\.9370\.1990\.6300\.6790\.8630\.3010\.6140\.7140\.8990\.2400\.617YOLO26sFull FT0\.7490\.9410\.5610\.7500\.7660\.9510\.6010\.7730\.6600\.8820\.5550\.6990\.7090\.9150\.5770\.734YOLO26mFull FT0\.6220\.9210\.5490\.6970\.7100\.9030\.5860\.7330\.5500\.8560\.5210\.6420\.6200\.8790\.5520\.684D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}Freeze0\.7130\.9510\.3940\.6860\.7500\.8940\.4630\.7020\.6390\.9350\.3160\.6300\.6900\.9140\.3760\.660D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}Full FT0\.6600\.9600\.2810\.6330\.8060\.9030\.2910\.6670\.5550\.9500\.5980\.7010\.6570\.9260\.3920\.658D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}Proposed0\.7240\.9770\.4220\.7080\.7850\.9520\.5120\.7500\.6610\.9430\.4780\.6940\.7180\.9470\.4950\.720D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}Freeze0\.7820\.9770\.3850\.7150\.7690\.8940\.4940\.7190\.7310\.9620\.4960\.7300\.7490\.9270\.4950\.724D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}Full FT0\.7180\.9160\.4080\.6810\.7640\.9000\.4490\.7040\.6770\.8520\.5000\.6760\.7180\.8750\.4730\.689D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}Proposed0\.7520\.9830\.6080\.7810\.7970\.9560\.6330\.7950\.6450\.9200\.5370\.7010\.7130\.9380\.5810\.744D​I​N​O​v​3b​a​s​eDINOv3\_\{base\}Freeze0\.5790\.9730\.6270\.7260\.6090\.8950\.6350\.7130\.5760\.9430\.6540\.7240\.5920\.9190\.6450\.718D​I​N​O​v​3b​a​s​eDINOv3\_\{base\}Full FT0\.7320\.9550\.6360\.7740\.7760\.8840\.7040\.7880\.6770\.9390\.5520\.7230\.7230\.9110\.6190\.751D​I​N​O​v​3b​a​s​eDINOv3\_\{base\}Proposed0\.6860\.9650\.6200\.7570\.7330\.9010\.5840\.7390\.6760\.9370\.7430\.7850\.7030\.9190\.6540\.759Federated – FedProxYOLO26nFull FT0\.7880\.9590\.6730\.8070\.8540\.9080\.7590\.8400\.6980\.9390\.6010\.7460\.7690\.9230\.6700\.787YOLO26sFull FT0\.8340\.9750\.5180\.7750\.8590\.9410\.5630\.7870\.8010\.9010\.4930\.7320\.8290\.9210\.5250\.758YOLO26mFull FT0\.7910\.9660\.7340\.8300\.8320\.9230\.7630\.8390\.7280\.8750\.7280\.7770\.7770\.8980\.7280\.801SAMProposed0\.6890\.9520\.4480\.6960\.7840\.9030\.4720\.7200\.5900\.8890\.5520\.6770\.6730\.8960\.5080\.692D​I​N​O​v​3b​a​s​eDINOv3\_\{base\}Freeze0\.7880\.9740\.7110\.8240\.8180\.9080\.6620\.7960\.6660\.9620\.6760\.7680\.7340\.9340\.6690\.779D​I​N​O​v​3b​a​s​eDINOv3\_\{base\}Full FT0\.7890\.9650\.6550\.8030\.7970\.9570\.6300\.7940\.7340\.8930\.7500\.7920\.7640\.9240\.6850\.791D​I​N​O​v​3b​a​s​eDINOv3\_\{base\}Proposed0\.7720\.9760\.7020\.8160\.7670\.9490\.7130\.8090\.7090\.9200\.6840\.7710\.7370\.9340\.6980\.790UltraFedFMProposed0\.6950\.9740\.0610\.5760\.7480\.9430\.1320\.6070\.6030\.9010\.1910\.5650\.6680\.9210\.1560\.582UltraFedFMFreeze0\.6790\.9680\.1860\.6110\.7510\.8460\.3190\.6390\.5990\.9830\.4710\.6840\.6660\.9100\.3810\.652UltraFedFMFull FT0\.6750\.9690\.2430\.6290\.8580\.8450\.3590\.6870\.5900\.9770\.4490\.6720\.6990\.9060\.3990\.668UltraSAMFreeze0\.8130\.9570\.5620\.7770\.7960\.8600\.5730\.7430\.7540\.9770\.5370\.7560\.7740\.9150\.5540\.748UltraSAMFull FT0\.7960\.9790\.6140\.7960\.8080\.9730\.5420\.7750\.6830\.9010\.6100\.7320\.7410\.9360\.5740\.750UltraSAMProposed0\.7870\.9620\.5590\.7690\.8310\.8730\.5400\.7480\.7740\.9510\.6180\.7810\.8010\.9100\.5760\.763D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}Freeze0\.6700\.9610\.4720\.7010\.8360\.8500\.4740\.7200\.5840\.9540\.5590\.6990\.6870\.8990\.5130\.700D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}Full FT0\.7210\.9370\.4670\.7080\.8270\.8590\.5060\.7310\.6600\.9390\.5040\.7010\.7340\.8970\.5050\.712D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}Proposed0\.7410\.9650\.7270\.8110\.8270\.9490\.6540\.8100\.6240\.8820\.7570\.7540\.7110\.9140\.7020\.776D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}Freeze0\.7050\.9610\.5480\.7380\.7130\.9120\.5540\.7260\.6840\.9030\.5930\.7270\.6980\.9070\.5730\.726D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}Full FT0\.7420\.9690\.6010\.7710\.7350\.9350\.6220\.7640\.7020\.9260\.6410\.7560\.7180\.9300\.6310\.760D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}Proposed0\.8140\.9730\.7470\.8450\.7520\.9220\.6830\.7860\.7390\.9010\.6830\.7740\.7520\.9220\.6830\.786Federated – FedAvgYOLO26nFull FT0\.7900\.9720\.5940\.7850\.8770\.9300\.6370\.8150\.6780\.9700\.5660\.7380\.7650\.9490\.5990\.771YOLO26sFull FT0\.8410\.9600\.4580\.7530\.8920\.9130\.4390\.7480\.7310\.9320\.5990\.7540\.8040\.9220\.5070\.744YOLO26mFull FT0\.8020\.9770\.4160\.7320\.8690\.9670\.4570\.7640\.7270\.8900\.5150\.7100\.7910\.9270\.4850\.734SAMProposed0\.6730\.9550\.4270\.6850\.7780\.8990\.4540\.7100\.5670\.8940\.5380\.6660\.6560\.8960\.4920\.681D​I​N​O​v​3b​a​s​eDINOv3\_\{base\}Freeze0\.7720\.9750\.6410\.7960\.8160\.9170\.6710\.8010\.6510\.9510\.6900\.7640\.7240\.9340\.6810\.779D​I​N​O​v​3b​a​s​eDINOv3\_\{base\}Full FT0\.7740\.9770\.6960\.8160\.8030\.8740\.6740\.7840\.6970\.9810\.6760\.7850\.7460\.9250\.6750\.782D​I​N​O​v​3b​a​s​eDINOv3\_\{base\}Proposed0\.7730\.9720\.7240\.8230\.7990\.9450\.7070\.8170\.6560\.9090\.6900\.7520\.7210\.9270\.6980\.782UltraFedFMProposed0\.6520\.9650\.1760\.5980\.7680\.9040\.2790\.6500\.5500\.9200\.4710\.6470\.6410\.9120\.3500\.634UltraFedFMFreeze0\.5890\.9640\.3120\.6220\.6510\.8620\.3920\.6350\.5640\.9730\.5290\.6890\.6050\.9140\.4510\.656UltraFedFMFull FT0\.6240\.9540\.2070\.5950\.7140\.8530\.3150\.6270\.5450\.9850\.4490\.6600\.6180\.9140\.3700\.634UltraSAMFreeze0\.8050\.9690\.4910\.7550\.8340\.8570\.4050\.6980\.7520\.9770\.7210\.8170\.7910\.9130\.5180\.741UltraSAMFull FT0\.8340\.9800\.7420\.8520\.8770\.9600\.6360\.8240\.6780\.9010\.7200\.7670\.7650\.9300\.6760\.790UltraSAMProposed0\.7320\.9730\.4830\.7290\.7600\.8930\.5340\.7290\.6570\.9580\.5910\.7350\.7050\.9240\.5610\.730D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}Freeze0\.7110\.9310\.4330\.6920\.8120\.8540\.4570\.7080\.6600\.9350\.4850\.6940\.7280\.8930\.4710\.697D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}Full FT0\.6880\.9700\.5070\.7220\.8300\.8570\.6320\.7730\.6170\.9730\.4410\.6770\.7080\.9120\.5200\.713D​I​N​O​v​2f​e​t​a​lDINOv2\_\{fetal\}Proposed0\.7410\.9780\.8000\.8390\.7050\.8920\.7800\.7920\.6720\.9660\.7310\.7900\.6880\.9270\.7550\.790D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}Freeze0\.6810\.9570\.5070\.7150\.6920\.9050\.5160\.7040\.6570\.8920\.5640\.7040\.6740\.8980\.5390\.704D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}Full FT0\.6300\.9530\.5790\.7210\.7060\.9240\.5690\.7330\.6710\.9140\.6150\.7330\.6880\.9240\.5870\.733D​I​N​O​v​2b​a​s​eDINOv2\_\{base\}Proposed0\.8060\.9770\.7880\.8570\.7610\.9570\.8460\.8550\.7460\.8900\.6460\.7610\.7530\.9220\.7320\.803
## References

- \[1\]J\. Ambsdorf, A\. Munk, S\. Llambias, A\. N\. Christensen, K\. Mikolaj, R\. Balestriero, M\. G\. Tolsgaard, A\. Feragen, and M\. Nielsen\(2025\)General methods make great domain\-specific foundation models: a case\-study on fetal ultrasound\.InInternational Conference on Medical Image Computing and Computer\-Assisted Intervention,pp\. 271–281\.Cited by:[§3\.1](https://arxiv.org/html/2607.18283#S3.SS1.p2.1)\.
- \[2\]J\. Bian, Y\. Peng, L\. Wang, Y\. Huang, and J\. Xu\(2025\)A survey on parameter\-efficient fine\-tuning for foundation models in federated learning\.arXiv preprint arXiv:2504\.21099\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p5.1)\.
- \[3\]X\. P\. Burgos\-Artizzu, D\. Coronado\-Gutiérrez, B\. Valenzuela\-Alcaraz, E\. Bonet\-Carne, E\. Eixarch, F\. Crispi, and E\. Gratacós\(2020\)Evaluation of deep convolutional neural networks for automatic classification of common maternal fetal ultrasound planes\.Scientific Reports10\(1\),pp\. 10200\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p3.1)\.
- \[4\]J\. H\. Cha, J\. Lim, Y\. H\. Jang, J\. K\. Hwang, J\. Y\. Na, J\. Lee, H\. J\. Lee, and J\. Ahn\(2022\)Altered microstructure of the splenium of corpus callosum is associated with neurodevelopmental impairment in preterm infants with necrotizing enterocolitis\.Italian Journal of Pediatrics48\(1\),pp\. 6\.Cited by:[§7](https://arxiv.org/html/2607.18283#S7.p3.1)\.
- \[5\]E\. Conti, R\. Rosati, L\. Federici, A\. Mancini, and M\. C\. Fiorentin\(2025\)Challenging DINOv3 foundation model under low inter\-class variability: a case study on fetal brain ultrasound\.arXiv preprint arXiv:2511\.01915\.Cited by:[§4\.1](https://arxiv.org/html/2607.18283#S4.SS1.p2.2)\.
- \[6\]G\. Egana\-Ugrinovic, S\. Savchev, C\. Bazán\-Arcos, B\. Puerto, E\. Gratacos, and M\. Sanz\-Cortes\(2015\)Neurosonographic assessment of the corpus callosum as imaging biomarker of abnormal neurodevelopment in late\-onset fetal growth restriction\.Fetal Diagnosis and Therapy37\(4\),pp\. 281–288\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p3.1)\.
- \[7\]European Parliament and Council of the European Union\(2024\)Regulation \(EU\) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence \(Artificial Intelligence Act\)\.Note:Official Journal of the European UnionAccessed: 2026\-05\-13External Links:[Link](https://eur-lex.europa.eu/eli/reg/2024/1689/oj)Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p6.1)\.
- \[8\]M\. C\. Fiorentino, L\. Federici, A\. P\. La Camera, and E\. G\. Caiani\(2025\)Adapt or specialize? a comprehensive evaluation of adapted sam versus task\-specific cnns for fetal abdominal segmentation\.Computer Methods and Programs in Biomedicine,pp\. 109178\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p4.1),[§4\.1](https://arxiv.org/html/2607.18283#S4.SS1.p3.1)\.
- \[9\]M\. C\. Fiorentino, G\. Migliorelli, F\. P\. Villani, E\. Frontoni, and S\. Moccia\(2025\)Contrastive prototype federated learning against noisy labels in fetal standard plane detection\.International Journal of Computer Assisted Radiology and Surgery20\(7\),pp\. 1431–1439\.Cited by:[§2](https://arxiv.org/html/2607.18283#S2.p4.1)\.
- \[10\]M\. C\. Fiorentino, S\. Moccia, M\. D\. Cosmo, E\. Frontoni, B\. Giovanola, and S\. Tiribelli\(2025\)Uncovering ethical biases in publicly available fetal ultrasound datasets\.npj Digital Medicine8\(1\),pp\. 355\.Cited by:[§6](https://arxiv.org/html/2607.18283#S6.p4.1)\.
- \[11\]M\. C\. Fiorentino, F\. P\. Villani, M\. Di Cosmo, E\. Frontoni, and S\. Moccia\(2023\)A review on deep\-learning algorithms for fetal ultrasound\-image analysis\.Medical Image Analysis83,pp\. 102629\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p3.1),[§6](https://arxiv.org/html/2607.18283#S6.p2.2)\.
- \[12\]J\. Han, H\. Wang, Y\. Feng, Q\. Yang, J\. Li, H\. Zhang, Y\. He, J\. Liu, T\. Nakamura, Y\. Cao,et al\.\(2026\)Federated learning for prenatal detection of interrupted aortic arch using fetal ultrasound imaging\.Biomedical Signal Processing and Control119,pp\. 109795\.Cited by:[§2](https://arxiv.org/html/2607.18283#S2.p4.1)\.
- \[13\]E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, W\. Chen,et al\.\(2022\)LoRA: low\-rank adaptation of large language models\.Iclr1\(2\),pp\. 3\.Cited by:[1st item](https://arxiv.org/html/2607.18283#S1.I1.i1.p1.1)\.
- \[14\]R\. Huang, A\. Namburete, and A\. Noble\(2018\)Learning to segment key clinical anatomical structures in fetal neurosonography informed by a region\-based descriptor\.Journal of Medical Imaging5\(1\),pp\. 014007–014007\.Cited by:[§2](https://arxiv.org/html/2607.18283#S2.p1.2)\.
- \[15\]T\. Jiang, Y\. Li, W\. Xing, R\. Cao, M\. Yu, Y\. Zhu, Y\. Chen, B\. Li, and D\. Ta\(2025\)UltraSAM: a foundational medical ultrasound segmentation model with limited training data\.Expert Systems with Applications,pp\. 130223\.Cited by:[§4\.1](https://arxiv.org/html/2607.18283#S4.SS1.p3.1)\.
- \[16\]Y\. Jiang, C\. Feng, J\. Ren, J\. Wei, Z\. Zhang, Y\. Hu, Y\. Liu, R\. Sun, X\. Tang, J\. Du,et al\.\(2025\)From pretraining to privacy: federated ultrasound foundation model with self\-supervised learning\.npj Digital Medicine8\(1\),pp\. 714\.Cited by:[§4\.1](https://arxiv.org/html/2607.18283#S4.SS1.p3.1)\.
- \[17\]J\. Jiao, J\. Zhou, X\. Li, M\. Xia, Y\. Huang, L\. Huang, N\. Wang, X\. Zhang, S\. Zhou, Y\. Wang,et al\.\(2024\)Usfm: a universal ultrasound foundation model generalized to tasks and organs towards label efficient image analysis\.Medical Image Analysis96,pp\. 103202\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p4.1)\.
- \[18\]Ultralytics YOLOExternal Links:[Link](https://github.com/ultralytics/ultralytics)Cited by:[§3\.1](https://arxiv.org/html/2607.18283#S3.SS1.p10.5)\.
- \[19\]A\. A\. Judi, P\. Suresh, and T\. A\. B\. Raj\(2026\)FB\-UNet\+\+: federated biometric unet\+\+ model for segmentation and classification network of fetal anomaly detection in prenatal care\.Biomedical Signal Processing and Control112,pp\. 108706\.Cited by:[§2](https://arxiv.org/html/2607.18283#S2.p4.1)\.
- \[20\]S\. P\. Karimireddy, S\. Kale, M\. Mohri, S\. Reddi, S\. Stich, and A\. T\. Suresh\(2020\)Scaffold: stochastic controlled averaging for federated learning\.InInternational Conference on Machine Learning,pp\. 5132–5143\.Cited by:[§6](https://arxiv.org/html/2607.18283#S6.p3.2)\.
- \[21\]W\. Khan, S\. Leem, K\. B\. See, J\. K\. Wong, S\. Zhang, and R\. Fang\(2025\)A comprehensive survey of foundation models in medicine\.IEEE Reviews in Biomedical Engineering\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p4.1)\.
- \[22\]V\. Lanzarone, E\. Eixarch, and A\. Borrell\(2025\)Fetal corpus callosum anomalies: a review of underlying genetic disorders and prenatal testing options\.Journal of Ultrasound in Medicine44\(4\),pp\. 637–652\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p1.1)\.
- \[23\]M\. H\. Le, K\. D\. Pham, T\. Vinh, T\. Nguyen, H\. H\. Huynh, K\. T\. Le, A\. M\. Vu, H\. N\. Luong, U\. Bagci, M\. Xu,et al\.\(2026\)VISCERA\-SAM: adapting segment anything for multi\-visceral fetal abdominal ultrasound segmentation\.InMedical Imaging with Deep Learning,Cited by:[§4\.1](https://arxiv.org/html/2607.18283#S4.SS1.p3.1)\.
- \[24\]M\. Li, S\. Liu, Z\. Zhang, Q\. Li, and X\. Xu\(2026\)Deep learning\-based automated detection of fetal corpus callosum abnormalities in prenatal ultrasound\.Frontiers in Pediatrics14,pp\. 1774586\.Cited by:[§2](https://arxiv.org/html/2607.18283#S2.p2.1),[§2](https://arxiv.org/html/2607.18283#S2.p3.1)\.
- \[25\]T\. Li, A\. K\. Sahu, M\. Zaheer, M\. Sanjabi, A\. Talwalkar, and V\. Smith\(2020\)Federated optimization in heterogeneous networks\.Proceedings of Machine learning and Systems2,pp\. 429–450\.Cited by:[§4\.1](https://arxiv.org/html/2607.18283#S4.SS1.p5.1),[§6](https://arxiv.org/html/2607.18283#S6.p3.2)\.
- \[26\]M\. Lubián\-Gutiérrez, I\. Benavente\-Fernández, Y\. Marín\-Almagro, N\. Jiménez\-Luque, A\. Zuazo\-Ojeda, Y\. Sánchez\-Sandoval, and S\. P\. Lubián\-López\(2024\)Corpus callosum long\-term biometry in very preterm children related to cognitive and motor outcomes\.Pediatric Research96\(2\),pp\. 409–417\.Cited by:[§7](https://arxiv.org/html/2607.18283#S7.p3.1)\.
- \[27\]C\. Ma, J\. Jiao, S\. Liang, J\. Fu, Q\. Wang, Z\. Li, Y\. Wang, and Y\. Guo\(2025\)TinyUSFM: towards compact and efficient ultrasound foundation models\.arXiv preprint arXiv:2510\.19239\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p4.1)\.
- \[28\]K\. K\. Marathu, F\. Vahedifard, M\. Kocak, X\. Liu, J\. O\. Adepoju, R\. M\. Bowker, M\. Supanich, R\. M\. Cosme\-Cruz, and S\. Byrd\(2024\)Fetal MRI analysis of corpus callosal abnormalities: classification, and associated anomalies\.Diagnostics14\(4\),pp\. 430\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p2.1)\.
- \[29\]B\. McMahan, E\. Moore, D\. Ramage, S\. Hampson, and B\. A\. y Arcas\(2017\)Communication\-efficient learning of deep networks from decentralized data\.InArtificial Intelligence and Statistics,pp\. 1273–1282\.Cited by:[§4\.1](https://arxiv.org/html/2607.18283#S4.SS1.p5.1),[§6](https://arxiv.org/html/2607.18283#S6.p3.2)\.
- \[30\]L\. Meng, D\. Zhao, Z\. Yang, and B\. Wang\(2020\)Automatic display of fetal brain planes and automatic measurements of fetal brain parameters by transabdominal three\-dimensional ultrasound\.Journal of Clinical Ultrasound48\(2\),pp\. 82–88\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p3.1)\.
- \[31\]A\. Meyer, A\. Murali, F\. Zarin, D\. Mutter, and N\. Padoy\(2025\)Ultrasam: a foundation model for ultrasound using large open\-access segmentation datasets\.International Journal of Computer Assisted Radiology and Surgery,pp\. 1–10\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p4.1)\.
- \[32\]M\. Oquab, T\. Darcet, T\. Moutakanni, H\. Vo, M\. Szafraniec, V\. Khalidov, P\. Fernandez, D\. Haziza, F\. Massa, A\. El\-Nouby,et al\.\(2023\)DINOv2: learning robust visual features without supervision\.arXiv preprint arXiv:2304\.07193\.Cited by:[§3\.1](https://arxiv.org/html/2607.18283#S3.SS1.p2.1),[§4\.1](https://arxiv.org/html/2607.18283#S4.SS1.p2.2),[§6](https://arxiv.org/html/2607.18283#S6.p2.2)\.
- \[33\]D\. Paladini, G\. Malinger, A\. Monteagudo, G\. Pilu, I\. Timor Tritsch, A\. Toi,et al\.\(2007\)Sonographic examination of the fetal central nervous system: guidelines for performing the ‘basic examination’and the ‘fetal neurosonogram’\.Ultrasound in Obstetrics & Gynecology29\(1\),pp\. 109–116\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p1.1)\.
- \[34\]G\. Pilu and Z\. Alfirevic\(2016\)Fetal central nervous system anomalies\.Fetal Medicine,pp\. 81\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p2.1)\.
- \[35\]S\. Santo, F\. D’antonio, T\. Homfray, P\. Rich, G\. Pilu, A\. Bhide, B\. Thilaganathan, and A\. Papageorghiou\(2012\)Counseling in fetal medicine: agenesis of the corpus callosum\.Ultrasound in Obstetrics & Gynecology40\(5\),pp\. 513–521\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p2.1)\.
- \[36\]R\. Sapkota, R\. H\. Cheppally, A\. Sharda, and M\. Karkee\(2025\)YOLO26: key architectural enhancements and performance benchmarking for real\-time object detection\.arXiv preprint arXiv:2509\.25164\.Cited by:[§4\.1](https://arxiv.org/html/2607.18283#S4.SS1.p6.1)\.
- \[37\]C\. Sendra\-Balcells, V\. M\. Campello, J\. Torrents\-Barrena, Y\. A\. Ahmed, M\. Elattar, B\. Ohene\-Botwe, P\. Nyangulu, W\. Stones, M\. Ammar, L\. N\. Benamer,et al\.\(2023\)Generalisability of fetal ultrasound deep learning models to low\-resource imaging settings in five african countries\.Scientific Reports13\(1\),pp\. 2728\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p6.1)\.
- \[38\]M\. J\. Sheller, B\. Edwards, G\. A\. Reina, J\. Martin, S\. Pati, A\. Kotrotsou, M\. Milchenko, W\. Xu, D\. Marcus, R\. R\. Colen, and S\. Bakas\(2020\)Federated learning in medicine: facilitating multi\-institutional collaborations without sharing patient data\.Scientific Reports10\(1\),pp\. 12598\.External Links:[Document](https://dx.doi.org/10.1038/s41598-020-69250-1)Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p5.1),[§1](https://arxiv.org/html/2607.18283#S1.p6.1)\.
- \[39\]M\. Sohan, T\. Sai Ram, and C\. V\. Rami Reddy\(2024\)A review on YOLOv8 and its advancements\.InInternational Conference on Data Intelligence and Cognitive Informatics,pp\. 529–545\.Cited by:[Figure 3](https://arxiv.org/html/2607.18283#S3.F3),[Figure 3](https://arxiv.org/html/2607.18283#S3.F3.3.2),[§3\.1](https://arxiv.org/html/2607.18283#S3.SS1.p4.4)\.
- \[40\]H\. R\. Torres, P\. Morais, B\. Oliveira, C\. Birdir, M\. Rüdiger, J\. C\. Fonseca, and J\. L\. Vilaça\(2022\)A review of image processing methods for fetal head and brain analysis in ultrasound images\.Computer Methods and Programs in Biomedicine215,pp\. 106629\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p3.1),[§6](https://arxiv.org/html/2607.18283#S6.p2.2)\.
- \[41\]Q\. Wang, J\. Pei, J\. Ouyang, Y\. Chen, J\. Pu, A\. Humayun, D\. Zhao, and B\. Liu\(2023\)A method framework of automatic localization and quantitative segmentation for the cavum septum pellucidum complex and the cerebellar vermis in fetal brain ultrasound images\.Quantitative Imaging in Medicine and Surgery13\(9\),pp\. 6059\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p3.1)\.
- \[42\]Q\. Wang, D\. Zhao, H\. Ma, and B\. Liu\(2025\)FB\-zwunet: a deep learning network for corpus callosum segmentation in fetal brain ultrasound images for prenatal diagnostics\.Biomedical Signal Processing and Control104,pp\. 107499\.Cited by:[§2](https://arxiv.org/html/2607.18283#S2.p2.1),[§2](https://arxiv.org/html/2607.18283#S2.p3.1)\.
- \[43\]X\. Wang, C\. Wang, W\. Yang, Q\. Yao, and L\. Zuo\(2024\)Assessment of the development of the central nervous system in fetuses with fetal growth restriction\.Archives of Gynecology and Obstetrics310\(6\),pp\. 2963–2971\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p1.1)\.
- \[44\]N\. Zhang, H\. Dong, P\. Wang, Z\. Wang, Y\. Wang, and Z\. Guo\(2020\)The value of obstetric ultrasound in screening fetal nervous system malformation\.World Neurosurgery138,pp\. 645–653\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p1.1)\.
- \[45\]X\. Zhang, E\. Z\. Chen, L\. Zhao, X\. Chen, Y\. Liu, B\. Maihe, J\. S\. Duncan, T\. Chen, and S\. Sun\(2025\)Adapting vision foundation models for real\-time ultrasound image segmentation\.InInternational Conference on Medical Image Computing and Computer\-Assisted Intervention,pp\. 24–34\.Cited by:[§1](https://arxiv.org/html/2607.18283#S1.p4.1)\.
- \[46\]Z\. Zhou, Y\. Lu, J\. Bai, V\. M\. Campello, F\. Feng, and K\. Lekadir\(2025\)Segment anything model for fetal head\-pubic symphysis segmentation in intrapartum ultrasound image analysis\.Expert Systems with Applications263,pp\. 125699\.Cited by:[§4\.1](https://arxiv.org/html/2607.18283#S4.SS1.p3.1)\.
- \[47\]Z\. Zou, K\. Chen, Z\. Shi, Y\. Guo, and J\. Ye\(2023\)Object detection in 20 years: a survey\.Proceedings of the IEEE111\(3\),pp\. 257–276\.Cited by:[§4\.2](https://arxiv.org/html/2607.18283#S4.SS2.p3.3)\.

Similar Articles

Accurate and Resource-Efficient Federated Continual Learning

arXiv cs.LG

FedRAN is a resource-aware analytic federated continual learning framework that replaces gradient-based updates with compact random feature statistics, achieving high accuracy with significantly lower communication and computation costs.

Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients

arXiv cs.LG

This paper proposes FedSLM, a parameter-centric framework for federated fine-tuning of foundation models with heterogeneous compressed clients, using SVD-based decomposition and a weak-to-strong elicitation step to handle resource asymmetry. Experiments show it outperforms existing federated baselines while reducing client GPU memory by ~50%.