Encoder-Sharing Hierarchical Federated Multi-Task Learning for VANETs

arXiv cs.LG Papers

Summary

This paper proposes EN-HMTFL, an encoder-sharing hierarchical multi-task federated learning framework for vehicular ad hoc networks that lets vehicles train heterogeneous perception tasks collaboratively via a shared encoder while keeping task-specific decoders local, improving accuracy by up to 24% and reducing communication rounds by up to 28.8%.

arXiv:2609.36157v1 Announce Type: new Abstract: Most federated learning frameworks for vehicular ad hoc networks assume that all vehicles collaboratively train a single model for a common task. This assumption limits their applicability to practical vehicular environments, where vehicles may perform heterogeneous but related perception tasks with different output spaces. This paper proposes encoder-sharing hierarchical multi-task federated learning (EN-HMTFL), which integrates cluster-based hierarchical federated learning with a globally shared encoder and vehicle-local decoders. EN-HMTFL enables vehicles performing different tasks to collaboratively learn a transferable feature representation while preserving their task-specific models locally. Only the encoder is exchanged and aggregated through the hierarchy, whereas raw data and local decoder parameters remain at the vehicles. The proposed framework is evaluated on the MNIST and GTSRB datasets in different vehicular scenarios. Across the evaluated scenarios, EN-HMTFL improves accuracy by up to 24.0% relative to the compared representation-sharing benchmark. In scenarios where EN-HMTFL converges earlier, the reduction reaches up to 69 communication rounds (28.8%).
Original Article
View Cached Full Text

Cached at: 09/30/26, 09:46 AM

# Encoder-Sharing Hierarchical Federated Multi-Task Learning for VANETs
Source: [https://arxiv.org/html/2609.36157](https://arxiv.org/html/2609.36157)
## Encoder\-Sharing Hierarchical Federated Multi\-Task Learning for VANETsThanks:This work was supported by the Scientific and Technological Research Council of Turkey \(TÜBİTAK\) under Grant 119C058 and Ford Otosan\.

M\. Saeid HaghighiFard and Sinem ColeriAffiliation:Department of Electrical and Electronics Engineering, Koç University, Istanbul, Türkiye Email: \{mhaghighifard21, scoleri\}@ku\.edu\.trAffiliation:

###### Abstract

Most federated learning frameworks for vehicular ad hoc networks assume that all vehicles collaboratively train a single model for a common task\. This assumption limits their applicability to practical vehicular environments, where vehicles may perform heterogeneous but related perception tasks with different output spaces\. This paper proposes encoder\-sharing hierarchical multi\-task federated learning \(EN\-HMTFL\), which integrates cluster\-based hierarchical federated learning with a globally shared encoder and vehicle\-local decoders\. EN\-HMTFL enables vehicles performing different tasks to collaboratively learn a transferable feature representation while preserving their task\-specific models locally\. Only the encoder is exchanged and aggregated through the hierarchy, whereas raw data and local decoder parameters remain at the vehicles\. The proposed framework is evaluated on the MNIST and GTSRB datasets in different vehicular scenarios\. Across the evaluated scenarios, EN\-HMTFL improves accuracy by up to 24\.0% relative to the compared representation\-sharing benchmark\. In scenarios where EN\-HMTFL converges earlier, the reduction reaches up to 69 communication rounds \(28\.8%\)\.

###### Index Terms:

vehicular ad hoc networks, hierarchical federated learning, multi\-task learning, shared encoder, local decoder

## IIntroduction

Vehicular Ad hoc Networks \(VANETs\) increasingly employ machine learning \(ML\) for safety\- and mobility\-critical services such as traffic\-sign recognition, object detection, trajectory prediction, traffic\-flow estimation, and driving\-scene understanding\. These services rely on the large volume of observations generated by onboard cameras, radar, LiDAR, and other sensors\[[1](https://arxiv.org/html/2609.36157#bib.bib1),[2](https://arxiv.org/html/2609.36157#bib.bib5),[3](https://arxiv.org/html/2609.36157#bib.bib6)\]\. A conventional centralized learning architecture requires vehicles to upload their locally generated data to a remote processor\. Such an architecture becomes difficult to scale as the number of vehicles and the volume of sensory data increase; it also consumes radio resources and exposes raw observations outside the vehicle\[[4](https://arxiv.org/html/2609.36157#bib.bib2)\]\. Federated learning \(FL\) addresses these limitations by allowing vehicles to train locally and exchange model parameters instead of raw data\[[5](https://arxiv.org/html/2609.36157#bib.bib3),[6](https://arxiv.org/html/2609.36157#bib.bib4)\]\. However, most FL formulations assume that every participating vehicle optimizes the same model for one common learning objective\.

Multi\-task federated learning \(MTFL\) replaces the single\-task assumption by coordinating multiple related learning objectives\. The original federated multi\-task formulation treats clients as related tasks and jointly optimizes their models while accounting for communication and systems constraints\[[7](https://arxiv.org/html/2609.36157#bib.bib7)\]\. FedEM learns a mixture of shared models and adapts the mixture coefficients to each client distribution\[[8](https://arxiv.org/html/2609.36157#bib.bib8)\]\. Ditto jointly learns a global reference model and a personalized model for every client\[[9](https://arxiv.org/html/2609.36157#bib.bib9)\]\. When Ditto is executed independently for each task, it provides within\-task personalization but does not create a common representation that enables heterogeneous tasks to collaborate\. These methods, therefore, do not directly resolve the incompatibility among task\-specific output components\.

A natural solution is to divide each model into a representation component and a task\-dependent component\. FedRep exchanges a common representation while retaining personalized prediction layers\[[10](https://arxiv.org/html/2609.36157#bib.bib10)\]\. M\-Fed similarly adopts an encoder\-decoder architecture to support collaboration across heterogeneous tasks while preserving task\-dependent components locally\[[11](https://arxiv.org/html/2609.36157#bib.bib11)\]\. In this design, the encoder learns transferable features from vehicles performing different tasks, whereas the decoder maps these features to the output space required by each vehicle and task\. However, existing representation\-sharing methods generally rely on a flat client\-server architecture, in which every participating vehicle sends its model update directly to a central server and receives the updated global representation directly from that server in each communication round\. In VANETs, this direct exchange can overload the infrastructure link, scale poorly with the number of vehicles, and fail to exploit the temporary local connectivity among nearby vehicles\.

This paper proposes an encoder\-sharing hierarchical multi\-task federated learning framework for cluster\-based vehicular networks\. Within the proposed framework, vehicles performing heterogeneous but related tasks jointly learn a shared encoder through the hierarchy while retaining their task\-specific decoders locally\. We conduct extensive simulations under different vehicular network scenarios\. Comparisons with two related benchmarks, Ditto and M\-Fed, demonstrate that the proposed algorithm achieves the highest accuracy across the evaluated scenarios and can reduce the number of communication rounds required for convergence, depending on the scenario\.

The remainder of this paper is organized as follows\. Section II presents the system model, Section III describes the proposed encoder\-sharing hierarchical MTFL algorithm, Section IV provides the experimental evaluation, and Section V concludes the paper\.

## IISystem Model

We consider a dynamic vehicular ad hoc network \(VANET\) in which vehicles communicate through vehicle\-to\-vehicle \(V2V\) links and access the Evolved Packet Core \(EPC\) through vehicle\-to\-infrastructure \(V2I\) or vehicle\-to\-network \(V2N\) connectivity\. V2V communication may be supported by IEEE 802\.11p\[[12](https://arxiv.org/html/2609.36157#bib.bib17)\], IEEE 802\.11bd\[[13](https://arxiv.org/html/2609.36157#bib.bib18)\], or LTE\-based device\-to\-device communication\[[14](https://arxiv.org/html/2609.36157#bib.bib19)\], whereas V2I/V2N connectivity is provided through 5G New Radio V2X\[[15](https://arxiv.org/html/2609.36157#bib.bib20)\]\.

The cluster\-based hierarchical federated learning \(CbHFL\) architecture introduced in\[[16](https://arxiv.org/html/2609.36157#bib.bib12)\]serves as the underlying communication and coordination framework\. In hierarchical federated learning \(HFL\), intermediate aggregators are placed between participating clients and the central server so that local model exchanges are first collected and aggregated at an intermediate level before being forwarded to the central entity\. This hierarchy reduces the number of direct client\-server transmissions and localizes frequent model exchanges\.

In CbHFL, vehicles dynamically transition among four states:*INITIAL*\(IN\),*STATE ELECTION*\(SE\),*CLUSTER HEAD*\(CH\), and*CLUSTER MEMBER*\(CM\)\. Each vehicle maintains a Vehicle Information Base \(VIB\) containing its current state, mobility information, neighboring vehicles, candidate CHs, active task set, shared encoder parameters, and task\-specific local decoder parameters\. The VIB is continuously updated through state transitions and periodic “HELLO\_PACKET” exchanges, enabling distributed cluster formation and maintenance under dynamic network conditions\.

After initialization, every vehicle enters the SE state and attempts to associate with a neighboring CH\. When multiple candidate CHs are available, the vehicle associates with the CH that offers the most stable mobility relationship\. If no suitable CH is available, a CH election is performed among neighboring vehicles in the SE state, selecting the vehicle with the lowest average relative speed to its neighbors as the CH; otherwise, the vehicle remains in the SE state until the topology changes\. During network operation, cluster membership is continuously updated: CMs that lose connectivity re\-enter the SE state, whereas neighboring CHs may merge to eliminate redundant clusters and reduce communication toward the EPC\.

The resulting hierarchy consists of three entities: CMs, CHs, and the EPC\. During global roundrr, let𝒞\(r\)\\mathcal\{C\}^\{\(r\)\}denote the set of active clusters and𝒮c\(r\)\\mathcal\{S\}\_\{c\}^\{\(r\)\}denote the set of CMs associated with CHccthat successfully complete local training\. Each CM trains its local multi\-task model using private data and uploads only its shared encoder parameters\. Each CHccaggregates the encoder parameters received from𝒮c\(r\)\\mathcal\{S\}\_\{c\}^\{\(r\)\}into a cluster\-level encoder and forwards it to the EPC, which performs inter\-cluster aggregation to construct the global shared encoder\. The updated encoder is then disseminated through the reverse hierarchy\. Since task\-specific decoders remain local throughout the learning process, they are never exchanged, aggregated, or involved in cluster formation\.

The proposed algorithm has three successive components\. First, each CM jointly trains the shared encoder with the decoder for every locally supported task\. Second, each CH averages only the encoders received from its CMs and forwards the cluster encoder to the EPC\. Third, the EPC averages the cluster encoders, redistributes the resulting global encoder, and checks convergence using the EPC\-level accuracy\. The decoder parameters do not participate in either hierarchical average\.

Algorithm 1Local Training at Cluster Member \(CM\) with Shared Encoder and Local DecodersInput :Global encoder

𝜽EPCenc,\(r\)\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\)\}\_\{\\mathrm\{EPC\}\}; task set

𝒯i\\mathcal\{T\}\_\{i\}; local datasets

\{𝒟i,t\}t∈𝒯i\\\{\\mathcal\{D\}\_\{i,t\}\\\}\_\{t\\in\\mathcal\{T\}\_\{i\}\}; local decoders

\{𝜽i,tdec\}t∈𝒯i\\\{\\boldsymbol\{\\theta\}^\{\\mathrm\{dec\}\}\_\{i,t\}\\\}\_\{t\\in\\mathcal\{T\}\_\{i\}\}; local epochs

EiE\_\{i\}
Output :Updated encoder

𝜽ienc,\(r\)\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\)\}\_\{i\}; updated local decoders retained at vehicle

ii
1

𝜽ienc←𝜽EPCenc,\(r\)\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\}\}\_\{i\}\\leftarrow\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\)\}\_\{\\mathrm\{EPC\}\};

2Retain

\{𝜽i,tdec\}t∈𝒯i\\\{\\boldsymbol\{\\theta\}^\{\\mathrm\{dec\}\}\_\{i,t\}\\\}\_\{t\\in\\mathcal\{T\}\_\{i\}\}from the preceding round;

3for*e=1,…,Eie=1,\\ldots,E\_\{i\}*do

4

Li←0L\_\{i\}\\leftarrow 0;

5foreach*t∈𝒯it\\in\\mathcal\{T\}\_\{i\}*do

6Sample mini\-batch

ℬi,t⊂𝒟i,t\\mathcal\{B\}\_\{i,t\}\\subset\\mathcal\{D\}\_\{i,t\};

7

𝐳i,t←Enc⁡\(xi,t,𝜽ienc\)\\mathbf\{z\}\_\{i,t\}\\leftarrow\\mathrm\{Enc\}\(x\_\{i,t\};\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\}\}\_\{i\}\);

8

y^i,t←Deci,t​\(𝐳i,t,𝜽i,tdec\)\\widehat\{y\}\_\{i,t\}\\leftarrow\\mathrm\{Dec\}\_\{i,t\}\(\\mathbf\{z\}\_\{i,t\};\\boldsymbol\{\\theta\}^\{\\mathrm\{dec\}\}\_\{i,t\}\);

9

Li,t←ℓt​\(y^i,t,yi,t\)L\_\{i,t\}\\leftarrow\\ell\_\{t\}\(\\widehat\{y\}\_\{i,t\},y\_\{i,t\}\);

10

Li←Li\+Li,tL\_\{i\}\\leftarrow L\_\{i\}\+L\_\{i,t\};

11

𝜽ienc←𝜽ienc−ηenc​∇𝜽iencLi\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\}\}\_\{i\}\\leftarrow\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\}\}\_\{i\}\-\\eta\_\{\\mathrm\{enc\}\}\\nabla\_\{\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\}\}\_\{i\}\}L\_\{i\};

12foreach*t∈𝒯it\\in\\mathcal\{T\}\_\{i\}*do

13

𝜽i,tdec←𝜽i,tdec−ηt​∇𝜽i,tdecLi,t\\boldsymbol\{\\theta\}^\{\\mathrm\{dec\}\}\_\{i,t\}\\leftarrow\\boldsymbol\{\\theta\}^\{\\mathrm\{dec\}\}\_\{i,t\}\-\\eta\_\{t\}\\nabla\_\{\\boldsymbol\{\\theta\}^\{\\mathrm\{dec\}\}\_\{i,t\}\}L\_\{i,t\};

14

𝜽ienc,\(r\)←𝜽ienc\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\)\}\_\{i\}\\leftarrow\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\}\}\_\{i\}and send it to the associated CH;

15Keep all

\{𝜽i,tdec\}t∈𝒯i\\\{\\boldsymbol\{\\theta\}^\{\\mathrm\{dec\}\}\_\{i,t\}\\\}\_\{t\\in\\mathcal\{T\}\_\{i\}\}at vehicle

ii;

Algorithm[1](https://arxiv.org/html/2609.36157#algorithm1)describes the local multi\-task learning procedure executed at each CM\. Let𝒯\\mathcal\{T\}denote the set of learning tasks considered by the system, and let𝒯i⊆𝒯\\mathcal\{T\}\_\{i\}\\subseteq\\mathcal\{T\}denote the nonempty subset of tasks supported by vehicleii\. Each taskt∈𝒯t\\in\\mathcal\{T\}has its own input space𝒳t\\mathcal\{X\}\_\{t\}, output space𝒴t\\mathcal\{Y\}\_\{t\}, and loss functionℓt\\ell\_\{t\}\. For every supported taskt∈𝒯it\\in\\mathcal\{T\}\_\{i\}, vehicleiistores a private non\-IID dataset

𝒟i,t=\{\(xi,t\(n\),yi,t\(n\)\)\}n=1Ni,t,\\mathcal\{D\}\_\{i,t\}=\\\{\(x\_\{i,t\}^\{\(n\)\},y\_\{i,t\}^\{\(n\)\}\)\\\}\_\{n=1\}^\{N\_\{i,t\}\},\(1\)whereNi,tN\_\{i,t\}denotes the number of local samples of taskttavailable at vehicleii\. The dataset is never transmitted to another vehicle, a CH, or the EPC\. Vehicleiimaintains a shared encoderEnc⁡\(⋅,𝜽ienc\)\\mathrm\{Enc\}\(\\cdot;\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\}\}\_\{i\}\)and a task\-specific decoderDeci,t​\(⋅,𝜽i,tdec\)\\mathrm\{Dec\}\_\{i,t\}\(\\cdot;\\boldsymbol\{\\theta\}^\{\\mathrm\{dec\}\}\_\{i,t\}\)for everyt∈𝒯it\\in\\mathcal\{T\}\_\{i\}\. Here,𝜽ienc\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\}\}\_\{i\}denotes the shared encoder parameters at vehicleii, and𝜽i,tdec\\boldsymbol\{\\theta\}^\{\\mathrm\{dec\}\}\_\{i,t\}denotes the local decoder parameters for taskttat vehicleii\. For a task\-specific inputxi,tx\_\{i,t\}, the encoder first produces the latent representation

𝐳i,t\\displaystyle\\mathbf\{z\}\_\{i,t\}=Enc⁡\(xi,t,𝜽ienc\),\\displaystyle=\\mathrm\{Enc\}\(x\_\{i,t\};\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\}\}\_\{i\}\),\(2\)y^i,t\\displaystyle\\widehat\{y\}\_\{i,t\}=Deci,t​\(𝐳i,t,𝜽i,tdec\),\\displaystyle=\\mathrm\{Dec\}\_\{i,t\}\(\\mathbf\{z\}\_\{i,t\};\\boldsymbol\{\\theta\}^\{\\mathrm\{dec\}\}\_\{i,t\}\),\(3\)where𝐳i,t\\mathbf\{z\}\_\{i,t\}is the latent representation generated for taskttat vehicleii, andy^i,t\\widehat\{y\}\_\{i,t\}is the corresponding predicted output produced by the task\-specific decoder, which maps the shared latent representation to the output space of tasktt\. The shared encoder has the same architecture and parameter dimensions across vehicles, enabling its parameters to be aggregated hierarchically\. The encoder parameters are the only trainable parameters exchanged through the hierarchical architecture, whereas decoder parameters are both vehicle\- and task\-specific and always remain local\. The corresponding task loss and local multi\-task objective are defined as

Li,t\\displaystyle L\_\{i,t\}=ℓt​\(y^i,t,yi,t\),t∈𝒯i,\\displaystyle=\\ell\_\{t\}\(\\widehat\{y\}\_\{i,t\},y\_\{i,t\}\),\\qquad t\\in\\mathcal\{T\}\_\{i\},\(4\)Li\\displaystyle L\_\{i\}=∑t∈𝒯iLi,t\.\\displaystyle=\\sum\_\{t\\in\\mathcal\{T\}\_\{i\}\}L\_\{i,t\}\.\(5\)
At the beginning of each communication roundrr, every participating CM replaces its local encoder𝜽ienc\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\}\}\_\{i\}with the latest global encoder𝜽EPCenc,\(r\)\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\)\}\_\{\\mathrm\{EPC\}\}received from the EPC while preserving all decoder parameters learned in previous rounds \(Lines 1–2\)\. For each local epoch, the CM initializes the total loss and iterates over every task in its assigned task set; consequently, a vehicle supporting multiple tasks jointly trains all corresponding decoders during the same communication round \(Lines 3–5\)\. For each task, the CM samples a mini\-batch from its private dataset, computes the latent representation using the shared encoder, applies the corresponding decoder, evaluates the task loss, and accumulates it into the local multi\-task objective \(Lines 6–10\)\. The shared encoder is then updated using the gradient of the complete multi\-task objective, integrating knowledge learned from different tasks into a common representation, whereas each decoder is updated only with the loss of its corresponding task, preventing interference between task\-specific output spaces \(Lines 11–13\)\. After local optimization, only the updated encoder parameters𝜽ienc,\(r\)\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\)\}\_\{i\}are transmitted to the associated CH, while all decoder parameters and private datasets remain local to the vehicle \(Lines 14–15\)\.

Algorithm 2Intra\-Cluster Encoder Aggregation at Cluster Head \(CH\)Input :Encoder updates

\{𝜽ienc,\(r\)\}i∈𝒮c\(r\)\\\{\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\)\}\_\{i\}\\\}\_\{i\\in\\mathcal\{S\}\_\{c\}^\{\(r\)\}\}from the active CMs of cluster

cc
Output :Cluster encoder

𝜽cenc,\(r\)\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\)\}\_\{c\}; disseminated EPC encoder

1Collect

𝜽ienc,\(r\)\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\)\}\_\{i\}from every

i∈𝒮c\(r\)i\\in\\mathcal\{S\}\_\{c\}^\{\(r\)\};

2

𝜽cenc,\(r\)←1\|𝒮c\(r\)\|​∑i∈𝒮c\(r\)𝜽ienc,\(r\)\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\)\}\_\{c\}\\leftarrow\\frac\{1\}\{\|\\mathcal\{S\}\_\{c\}^\{\(r\)\}\|\}\\sum\_\{i\\in\\mathcal\{S\}\_\{c\}^\{\(r\)\}\}\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\)\}\_\{i\};

3Send

𝜽cenc,\(r\)\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\)\}\_\{c\}to the EPC;

4Receive

𝜽EPCenc,\(r\+1\)\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\+1\)\}\_\{\\mathrm\{EPC\}\}from the EPC;

5Broadcast

𝜽EPCenc,\(r\+1\)\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\+1\)\}\_\{\\mathrm\{EPC\}\}to all associated CMs;

Algorithm[2](https://arxiv.org/html/2609.36157#algorithm2)first collects one encoder update from each active CM in the current cluster \(Line 1\)\. It then applies the average for clustercc\. Decoder parameters and task identities are not inputs to this operation \(Line 2\)\. The CH sends the resulting cluster encoder to the EPC rather than forwarding every individual CM encoder over the infrastructure link \(Line 3\)\. After the EPC completes the next global update, the CH receives the new global encoder and broadcasts it to its CMs for the next round \(Lines 4–5\)\.

Algorithm 3Global Encoder Aggregation at EPCInput :Initial global encoder

𝜽EPCenc,\(0\)\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(0\)\}\_\{\\mathrm\{EPC\}\}
Output :Global encoder sequence

1for*r=0,…,Rmax−1r=0,\\ldots,R\_\{\\max\}\-1*do

2Collect

\{𝜽cenc,\(r\)\}c∈𝒞\(r\)\\\{\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\)\}\_\{c\}\\\}\_\{c\\in\\mathcal\{C\}^\{\(r\)\}\}from the active CHs;

3

𝜽EPCenc,\(r\+1\)←1\|𝒞\(r\)\|​∑c∈𝒞\(r\)𝜽cenc,\(r\)\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\+1\)\}\_\{\\mathrm\{EPC\}\}\\leftarrow\\frac\{1\}\{\|\\mathcal\{C\}^\{\(r\)\}\|\}\\sum\_\{c\\in\\mathcal\{C\}^\{\(r\)\}\}\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\)\}\_\{c\};

4Broadcast

𝜽EPCenc,\(r\+1\)\\boldsymbol\{\\theta\}^\{\\mathrm\{enc\},\(r\+1\)\}\_\{\\mathrm\{EPC\}\}to all active CHs;

For algorithm[3](https://arxiv.org/html/2609.36157#algorithm3), in every global round, the EPC collects one encoder from each active CH, computes the average, and sends the new global encoder back to the CHs\.

## IVExperimental Evaluation

We compare the proposed encoder\-sharing hierarchical multi\-task federated learning algorithm, termed EN\-HMTFL, with Ditto and M\-Fed, as these benchmarks capture the two main learning paradigms most relevant to our design\. Ditto is adopted as a personalized federated learning benchmark and adapted to the underlying vehicular setting by executing independent Ditto processes in parallel for the tasks under consideration\[[9](https://arxiv.org/html/2609.36157#bib.bib9)\]\. Each process learns a task\-specific global reference model together with personalized local models for the vehicles assigned to that task\. Hence, Ditto provides personalization within each task, but it does not enable knowledge transfer or representation sharing across heterogeneous tasks\. M\-Fed is selected as the closest representation\-sharing benchmark because it adopts an encoder\-decoder architecture and enables cross\-task knowledge transfer through a global encoder\[[11](https://arxiv.org/html/2609.36157#bib.bib11)\]\. However, M\-Fed does not exchange only encoder parameters\. Clients upload both their encoders and task\-specific decoders, and the server first aggregates complete models belonging to the same task to construct task\-global models\. It then extracts and aggregates the encoder components of these task\-global models to obtain a cross\-task global encoder\. In contrast, EN\-HMTFL communicates and aggregates only encoder parameters through the hierarchy, while all decoders remain strictly local and are never transmitted or aggregated\.

### IV\-ASimulation Setup

The mobility and communication environment is generated using SUMO, while Kafka supports communication events and model updates via streaming\[[17](https://arxiv.org/html/2609.36157#bib.bib15),[18](https://arxiv.org/html/2609.36157#bib.bib16)\]\. Learning is implemented in PyTorch and scikit\-learn\. IEEE 802\.11p with the Winner\+ B1 channel model is used for V2V communication, and the infrastructure link follows the Friis\-based 5G NR model\[[19](https://arxiv.org/html/2609.36157#bib.bib13),[20](https://arxiv.org/html/2609.36157#bib.bib14)\]\. The evaluated scenarios include a 50\-vehicle setting with a 100 m transmission range and a 20\-vehicle setting with transmission ranges of 100 and 500 m\. Vehicles are randomly assigned to perform MNIST, GTSRB, or both tasks, and each vehicle trains using its own private non\-IID dataset\.

For GTSRB, a lightweight CNN extracts a 32\-dimensional feature vector from each image, while MNIST samples are flattened into 64\-dimensional vectors\. These features are mapped by a three\-layer shared encoder to a 128\-dimensional latent representation\. Each vehicle maintains a local reconstruction decoder and a task\-specific classification head; only the encoder parameters are exchanged and aggregated\. The transmitted encoder accounts for approximately 46\.9% of the MNIST model parameters and 45\.7% of the GTSRB model parameters\.

Training uses SGD with momentum0\.90\.9, an initial learning rate of0\.010\.01, and a decay factor of0\.950\.95every 10 communication rounds\. Each CM performs 10 local epochs per round\. Mean\-squared error is used for reconstruction, and cross\-entropy for classification\. EPC accuracy is evaluated over 250 rounds, and convergence is declared when the change in accuracy remains below0\.0050\.005for three consecutive rounds\. The reported convergence round is the first to complete this three\-round streak\.

\(a\)MNIST\(b\)GTSRB
Fig\. 1:EPC accuracy versus communication round for 50 vehicles, a 100 m transmission range: \(a\) MNIST and \(b\) GTSRB\.
### IV\-BPerformance evaluation

Fig\.[1](https://arxiv.org/html/2609.36157#S4.F1)compares EN\-HMTFL with Ditto and M\-Fed for 50 vehicles and a transmission range of 100 m\. On MNIST, EN\-HMTFL achieves a late\-round mean EPC accuracy of69\.29±0\.22%69\.29\\pm 0\.22\\%, outperforming Ditto and M\-Fed, which attain61\.89±0\.67%61\.89\\pm 0\.67\\%and49\.81±1\.06%49\.81\\pm 1\.06\\%, respectively\. This corresponds to gains of 7\.41 percentage points \(12\.0% relative\) over Ditto and 19\.49 percentage points \(39\.1% relative\) over M\-Fed\. EN\-HMTFL satisfies the convergence criterion at round 211, only two rounds after Ditto and 12 rounds before M\-Fed\. More importantly, its late\-round standard deviation is substantially smaller than those of both benchmarks, indicating that the proposed method not only achieves higher accuracy but also maintains a more stable operating regime\. Therefore, the slight two\-round delay relative to Ditto is negligible compared with the considerable improvement in sustained accuracy and stability\.

The performance gap becomes more pronounced on GTSRB\. EN\-HMTFL achieves a late\-round mean EPC accuracy of76\.39±7\.79%76\.39\\pm 7\.79\\%, compared with29\.92±0\.41%29\.92\\pm 0\.41\\%for Ditto and61\.58±0\.86%61\.58\\pm 0\.86\\%for M\-Fed\. Accordingly, EN\-HMTFL improves the accuracy by 46\.47 percentage points \(155\.3% relative\) over Ditto and by 14\.81 percentage points \(24\.0% relative\) over M\-Fed\. Although Ditto meets the stopping criterion earlier, at round 211, this earlier stabilization occurs at a markedly inferior accuracy level and therefore does not indicate a better learned solution\. EN\-HMTFL converges at round 233, five rounds earlier than M\-Fed, and reaches a final plotted accuracy of 85\.74%, whereas Ditto and M\-Fed reach only 30\.32% and 62\.25%, respectively\. The larger temporal variation of EN\-HMTFL on GTSRB is caused by a temporary late\-round degradation; however, the method subsequently recovers and remains clearly superior in both its final and sustained accuracy\.

Overall, EN\-HMTFL achieves the highest late\-round EPC accuracy in the 50\-vehicle scenario, while its convergence round is comparable to the benchmarks and depends on the task\. Unlike Ditto, it enables cross\-task representation sharing, and unlike M\-Fed, it exploits localized CbHFL aggregation\. This combination improves knowledge transfer and yields higher EPC accuracy, particularly for the more heterogeneous GTSRB task\.

\(a\)100 m, MNIST\(b\)100 m, GTSRB\(c\)500 m, MNIST\(d\)500 m, GTSRB
Fig\. 2:EPC accuracy versus communication round for 20 vehicles: \(a\)–\(b\) 100 m and \(c\)–\(d\) 500 m transmission ranges\.Fig\.[2](https://arxiv.org/html/2609.36157#S4.F2)evaluates the effect of transmission range for 20 vehicles\. For MNIST at 100 m, EN\-HMTFL achieves a late\-round mean EPC accuracy of73\.56±0\.30%73\.56\\pm 0\.30\\%, compared with72\.07±0\.37%72\.07\\pm 0\.37\\%for Ditto and63\.71±0\.38%63\.71\\pm 0\.38\\%for M\-Fed\. This corresponds to improvements of 1\.49 percentage points over Ditto and 9\.85 percentage points over M\-Fed, equivalent to relative gains of 2\.1% and 15\.5%, respectively\. EN\-HMTFL also converges at round 177, which is 51 rounds earlier than Ditto and 35 rounds earlier than M\-Fed\. This improvement results from combining cross\-task encoder sharing with localized CH\-level aggregation, which promotes transferable representation learning while limiting the propagation of heterogeneous updates under short\-range connectivity\.

At 500 m, EN\-HMTFL remains the most accurate method for MNIST, reaching72\.82±0\.33%72\.82\\pm 0\.33\\%\. It exceeds Ditto by 0\.75 percentage points and M\-Fed by 9\.10 percentage points, corresponding to relative gains of 1\.0% and 14\.3%, respectively\. EN\-HMTFL and Ditto satisfy the convergence criterion at the same round, whereas M\-Fed converges 16 rounds earlier\. However, M\-Fed stabilizes at an accuracy that is 9\.10 percentage points lower, showing that earlier convergence does not necessarily indicate a better learned model\. Compared with the 100 m case, the wider transmission range reduces the EN\-HMTFL accuracy by 0\.75 percentage points and delays convergence by 51 rounds\. This behavior is attributed to the denser neighborhood, which improves reachability but also introduces a broader and more heterogeneous set of encoder updates into uniform aggregation\.

TABLE I:EPC Accuracy and Convergence for the different ScenariosA similar trend is observed for GTSRB\. At 100 m, EN\-HMTFL achieves81\.81±0\.07%81\.81\\pm 0\.07\\%, outperforming Ditto and M\-Fed by 3\.62 and 3\.25 percentage points, respectively\. These correspond to relative improvements of 4\.6% over Ditto and 4\.1% over M\-Fed\. EN\-HMTFL converges at round 171, which is 52 rounds earlier than Ditto and 69 rounds earlier than M\-Fed\. The very small late\-round standard deviation further confirms that the proposed method reaches a stable high\-accuracy regime\. At 500 m, EN\-HMTFL still achieves the highest mean accuracy,78\.93±0\.10%78\.93\\pm 0\.10\\%, while converging 18 rounds earlier than Ditto and 35 rounds earlier than M\-Fed\. Although the accuracy margins narrow to 0\.74 percentage points over Ditto and 0\.38 percentage points over M\-Fed, EN\-HMTFL remains superior in both accuracy and convergence speed\. The stronger gains at 100 m indicate that hierarchical encoder aggregation is most effective when localized clusters reduce update heterogeneity, whereas the wider range weakens this advantage by mixing updates from a more diverse set of vehicles\.

Table[I](https://arxiv.org/html/2609.36157#S4.T1)summarizes these accuracy and convergence results\. Overall, EN\-HMTFL consistently achieves the highest late\-round EPC accuracy across both tasks and both transmission ranges\. Its convergence behavior is scenario\-dependent: EN\-HMTFL converges earlier than both benchmarks in the 20\-vehicle GTSRB scenarios and in the 20\-vehicle MNIST scenario at 100 m, matches Ditto at 500 m for MNIST, and converges slightly later than Ditto in the 50\-vehicle scenario\. These results show that the principal advantage of EN\-HMTFL is its consistently higher accuracy, while convergence gains are obtained in several, but not all, evaluated scenarios\.

## VConclusion

This paper introduced EN\-HMTFL, a hierarchical multi\-task federated learning framework for dynamic VANETs in which vehicles exchange a shared encoder through a cluster\-based CM–CH–EPC hierarchy while keeping task\-specific decoders local\. This design enables knowledge transfer across heterogeneous yet related tasks without aggregating incompatible task\-specific model components\. Across the evaluated vehicle\-density and transmission\-range settings, EN\-HMTFL consistently achieved the highest late\-round EPC accuracy, with relative accuracy gains of up to 24\.0% over M\-Fed\. Its convergence behavior varied across scenarios: EN\-HMTFL converged earlier than the compared methods in several settings, achieving a maximum reduction of 69 communication rounds \(28\.8%\), whereas in other settings it converged at the same round as, or later than, one benchmark\. These results demonstrate the accuracy benefit of hierarchical encoder sharing while showing that the convergence advantage depends on the vehicular scenario\. Future work will extend the framework to more complex perception and multimodal tasks and investigate task\-aware clustering and reliability\-aware aggregation\.

## References

- \[1\]\(2024\)Federated learning in intelligent transportation systems: recent applications and open problems\.IEEE Transactions on Intelligent Transportation Systems25\(5\),pp\. 3259–3285\.External Links:[Document](https://dx.doi.org/10.1109/TITS.2023.3324962)Cited by:[§I](https://arxiv.org/html/2609.36157#S1.p1.1)\.
- \[2\]S\. Wang, C\. Li, D\. W\. K\. Ng, Y\. C\. Eldar, H\. V\. Poor, Q\. Hao, and C\. Xu\(2023\)Federated deep learning meets autonomous vehicle perception: design and verification\.IEEE Network37\(3\),pp\. 16–25\.External Links:[Document](https://dx.doi.org/10.1109/MNET.104.2100403)Cited by:[§I](https://arxiv.org/html/2609.36157#S1.p1.1)\.
- \[3\]V\. P\. Chellapandi, L\. Yuan, C\. G\. Brinton, S\. H\. Żak, and Z\. Wang\(2024\)Federated learning for connected and automated vehicles: a survey of existing approaches and challenges\.IEEE Transactions on Intelligent Vehicles9\(1\),pp\. 119–137\.External Links:[Document](https://dx.doi.org/10.1109/TIV.2023.3332675)Cited by:[§I](https://arxiv.org/html/2609.36157#S1.p1.1)\.
- \[4\]W\. Y\. B\. Lim, N\. C\. Luong, D\. T\. Hoang, Y\. Jiao, Y\. Liang, Q\. Yang, D\. Niyato, and C\. Miao\(2020\)Federated learning in mobile edge networks: a comprehensive survey\.IEEE Communications Surveys & Tutorials22\(3\),pp\. 2031–2063\.External Links:[Document](https://dx.doi.org/10.1109/COMST.2020.2986024)Cited by:[§I](https://arxiv.org/html/2609.36157#S1.p1.1)\.
- \[5\]B\. McMahan, E\. Moore, D\. Ramage, S\. Hampson, and B\. A\. y Arcas\(2017\)Communication\-efficient learning of deep networks from decentralized data\.InProceedings of the 20th International Conference on Artificial Intelligence and Statistics \(AISTATS\),Vol\.54,pp\. 1273–1282\.Cited by:[§I](https://arxiv.org/html/2609.36157#S1.p1.1)\.
- \[6\]A\. M\. Elbir, B\. Soner, S\. Coleri, D\. Gündüz, and M\. Bennis\(2022\)Federated learning in vehicular networks\.InProceedings of the IEEE International Mediterranean Conference on Communications and Networking \(MeditCom\),pp\. 72–77\.External Links:[Document](https://dx.doi.org/10.1109/MeditCom55741.2022.9928621)Cited by:[§I](https://arxiv.org/html/2609.36157#S1.p1.1)\.
- \[7\]V\. Smith, C\. Chiang, M\. Sanjabi, and A\. S\. Talwalkar\(2017\)Federated multi\-task learning\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.30,pp\. 4424–4434\.Cited by:[§I](https://arxiv.org/html/2609.36157#S1.p2.1)\.
- \[8\]O\. Marfoq, G\. Neglia, A\. Bellet, L\. Kameni, and R\. Vidal\(2021\)Federated multi\-task learning under a mixture of distributions\.InAdvances in Neural Information Processing Systems,M\. Ranzato, A\. Beygelzimer, Y\. Dauphin, P\.S\. Liang, and J\. W\. Vaughan \(Eds\.\),Vol\.34,pp\. 15434–15447\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2021/file/82599a4ec94aca066873c99b4c741ed8-Paper.pdf)Cited by:[§I](https://arxiv.org/html/2609.36157#S1.p2.1)\.
- \[9\]T\. Li, S\. Hu, A\. Beirami, and V\. Smith\(2021\)Ditto: fair and robust federated learning through personalization\.InProceedings of the 38th International Conference on Machine Learning \(ICML\),Vol\.139,pp\. 6357–6368\.Cited by:[§I](https://arxiv.org/html/2609.36157#S1.p2.1),[§IV](https://arxiv.org/html/2609.36157#S4.p1.1)\.
- \[10\]L\. Collins, H\. Hassani, A\. Mokhtari, and S\. Shakkottai\(2021\)Exploiting shared representations for personalized federated learning\.InProceedings of the 38th International Conference on Machine Learning \(ICML\),Vol\.139,pp\. 2089–2099\.Cited by:[§I](https://arxiv.org/html/2609.36157#S1.p3.1)\.
- \[11\]J\. Zhou, W\. Bao, J\. Wang, D\. Zhang, X\. Zhang, and Y\. Zhang\(2025\)Multi\-task federated learning with encoder–decoder structure: enabling collaborative learning across different tasks\.International Journal of Machine Learning and Cybernetics16,pp\. 10403 – 10420\.External Links:[Link](https://api.semanticscholar.org/CorpusID:277780856)Cited by:[§I](https://arxiv.org/html/2609.36157#S1.p3.1),[§IV](https://arxiv.org/html/2609.36157#S4.p1.1)\.
- \[12\]J\. B\. Kenney\(2011\)Dedicated short\-range communications \(DSRC\) standards in the united states\.Proceedings of the IEEE99\(7\),pp\. 1162–1182\.Cited by:[§II](https://arxiv.org/html/2609.36157#S2.p1.1)\.
- \[13\]M\. Noor\-A\-Rahim, Z\. Liu, H\. Lee, M\. O\. Khyam, J\. He, D\. Pesch, K\. Moessner, W\. Saad, and H\. V\. Poor\(2022\)6G for vehicle\-to\-everything \(V2X\) communications: enabling technologies, challenges, and opportunities\.Proceedings of the IEEE110\(6\),pp\. 712–734\.Cited by:[§II](https://arxiv.org/html/2609.36157#S2.p1.1)\.
- \[14\]H\. Seo, K\.\-D\. Lee, S\. Yasukawa, Y\. Peng, and P\. Sartori\(2016\)LTE evolution for vehicle\-to\-everything services\.IEEE Communications Magazine54\(6\),pp\. 22–28\.Cited by:[§II](https://arxiv.org/html/2609.36157#S2.p1.1)\.
- \[15\]H\. Bagheri, M\. Noor\-A\-Rahim, Z\. Liu, H\. Lee, D\. Pesch, K\. Moessner, and P\. Xiao\(2021\)5G NR\-V2X: toward connected and cooperative autonomous driving\.IEEE Communications Standards Magazine5\(1\),pp\. 48–54\.Cited by:[§II](https://arxiv.org/html/2609.36157#S2.p1.1)\.
- \[16\]M\. S\. HaghighiFard and S\. Coleri\(2025\)Hierarchical federated learning in multi\-hop cluster\-based VANETs\.IEEE Transactions on Vehicular Technology74\(10\),pp\. 15371–15385\.External Links:[Document](https://dx.doi.org/10.1109/TVT.2025.3569179)Cited by:[§II](https://arxiv.org/html/2609.36157#S2.p2.1)\.
- \[17\]Eclipse SUMOSimulation of urban mobility \(SUMO\)\.Note:Online:https://eclipse\.dev/sumo/Cited by:[§IV\-A](https://arxiv.org/html/2609.36157#S4.SS1.p1.1)\.
- \[18\]Apache Software FoundationApache kafka\.Note:Online:https://kafka\.apache\.org/Cited by:[§IV\-A](https://arxiv.org/html/2609.36157#S4.SS1.p1.1)\.
- \[19\]M\. Sepulcre, M\. Gonzalez\-Martín, J\. Gozalvez, R\. Molina\-Masegosa, and B\. Coll\-Perales\(2022\)Analytical models of the performance of IEEE 802\.11p vehicle\-to\-vehicle communications\.IEEE Transactions on Vehicular Technology71\(1\),pp\. 713–724\.External Links:[Document](https://dx.doi.org/10.1109/TVT.2021.3124708)Cited by:[§IV\-A](https://arxiv.org/html/2609.36157#S4.SS1.p1.1)\.
- \[20\]3rd Generation Partnership Project \(3GPP\)\(2020\)Study on channel model for frequencies from 0\.5 to 100 GHz\.Technical reportTechnical ReportTR 38\.901,3GPP\.Cited by:[§IV\-A](https://arxiv.org/html/2609.36157#S4.SS1.p1.1)\.

Similar Articles

Federated Foundation Models over Vehicular Networks

arXiv cs.LG

This paper presents a vision for integrating multi-modal multi-task federated foundation models (M3T FedFMs) into vehicular networks, discussing training principles, use cases, challenges, and a case study on the Waymo Open Dataset.