FlatLand: Personalized Graph Federated Learning via Tailored Lorentz Space
Summary
The paper proposes FlatLand, a personalized federated learning method that uses tailored Lorentz space in hyperbolic geometry to handle heterogeneous graph structures among clients, improving performance in privacy-preserving collaborative training.
View Cached Full Text
Cached at: 08/24/26, 04:36 AM
# Personalized Graph Federated Learning via Tailored Lorentz Space Source: [https://arxiv.org/html/2608.21096](https://arxiv.org/html/2608.21096) Jiahong LiuAffiliation:Department of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong SAR, ChinaCorrespondence to:[jiahong\.liu21@gmail\.com](mailto:[email protected])Xinyu FuAffiliation:Huawei Technologies Co\., Ltd\., Hong Kong SAR, ChinaMenglin YangAffiliation:AI Thrust, The Hong Kong University of Science and Technology \(Guangzhou\), Guangzhou, Guangdong, ChinaWeixi ZhangAffiliation:Huawei Technologies Co\., Ltd\., Hong Kong SAR, ChinaRex YingAffiliation:Department of Computer Science, Yale University, New Haven, CT, USAIrwin KingAffiliation:Department of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong SAR, ChinaCorrespondence to:[king@cse\.cuhk\.edu\.hk](mailto:[email protected]) ###### Abstract Federated learning enables privacy\-preserving collaborative training, but highly heterogeneous client data remain challenging, especially in graph federated learning where clients possess structurally diverse graphs\. Existing personalized federated learning \(PFL\) methods ignore the intrinsic geometric properties of diverse graph structures\. We propose𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}, a novel personalizedFederatedlearning method that embeds different clients’ data intailoredLorentz space of hyperbolic geometry\. Our key insight is that hyperbolic geometry naturally accommodates the intrinsic negative curvature prevalent in real\-world graphs, while the time\-like dimension in Lorentz space provides a principled way to encode client\-specific heterogeneity\. We develop a parameter decoupling strategy that separates heterogeneous information \(captured in time\-like parameters\) from common knowledge \(preserved in space\-like parameters\), enabling direct aggregation without requiring client similarity estimation and extra calculation modules\. Empirical results on diverse federated graph learning tasks demonstrate that𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}achieves superior performance, particularly in low\-dimensional settings\. Code is available in our[GitHub repository](https://github.com/HUBERILT/FlatLand_ICML)\. ###### Keywords: Federated Learning, Graph Neural Networks, Hyperbolic Geometry, Personalization ## 1Introduction Figure 1:Example: \(a\) KDE of degree distributions from three CiteSeer clients\([11](https://arxiv.org/html/2608.21096#bib.bib10)\), and \(b\) their respective 2D Lorentz Spaces with different Lorentz scale parametersKK\.Federated learning \(FL\) has emerged as a paradigm that enables collaborative machine learning across multiple clients while preserving data privacy\. Traditional FL struggles with data heterogeneity, as one model cannot satisfy diverse local requirements\([57](https://arxiv.org/html/2608.21096#bib.bib63)\)\. This challenge is magnified in graph federated learning, where complex topology yields pronounced structural heterogeneity across clients\([61](https://arxiv.org/html/2608.21096#bib.bib59);[58](https://arxiv.org/html/2608.21096#bib.bib50)\)\. In severe cases, federated learning may even underperform local training\([5](https://arxiv.org/html/2608.21096#bib.bib51)\)\. To address heterogeneity, Personalized federated learning \(PFL\) resolves this by sharing common model knowledge and allowing for client\-specific adaptations\. Current PFL approaches for graph data primarily address heterogeneity through three main strategies during aggregation: \(1\)parameter disentanglement: splitting models into shared and personalized components\([45](https://arxiv.org/html/2608.21096#bib.bib28);[58](https://arxiv.org/html/2608.21096#bib.bib50)\); \(2\)client similarity estimation: analyzing weights or gradients to evaluate client similarities\([61](https://arxiv.org/html/2608.21096#bib.bib59)\); and \(3\)auxiliary module calculation: incorporating additional modules to distinguish between globally beneficial and client\-specific parameters\([5](https://arxiv.org/html/2608.21096#bib.bib51)\)\. Despite their effectiveness, existing methods for PFL are confined to Euclidean space, implicitly assuming a uniform flat geometry across all client data distributions\. This assumption necessitates the design of intricate mechanisms to address heterogeneity\. Simple parameter disentanglement often fails, while more advanced techniques, such as client similarity estimation or auxiliary module integration, achieve better performance but incur significant computational overhead\. Therefore, we revisit PFL through a geometric lens using Ricci curvature\([16](https://arxiv.org/html/2608.21096#bib.bib14);[56](https://arxiv.org/html/2608.21096#bib.bib15)\), which characterizes the intrinsic properties of graph structures: its sign indicates hyperbolic \(negative\), flat \(zero\), or spherical \(positive\) geometry, while its magnitude captures how much the space curves\. Observations\.Our empirical analysis across multiple real\-world datasets reveals two critical observations \([Figure 3](https://arxiv.org/html/2608.21096#S5.F3)and[Figure 2](https://arxiv.org/html/2608.21096#S1.F2)\): \(1\) client graphs predominantly exhibit*negative*Ricci curvature, indicating inherent hyperbolic structure, and \(2\) curvature values*vary substantially*across different clients, revealing intrinsic geometric heterogeneity that extends beyond simple statistical differences\. These findings suggest that the assumption of unified Euclidean geometry in existing methods is fundamentally misaligned with the true geometric nature of graph data\([48](https://arxiv.org/html/2608.21096#bib.bib18);[1](https://arxiv.org/html/2608.21096#bib.bib16);[28](https://arxiv.org/html/2608.21096#bib.bib8);[58](https://arxiv.org/html/2608.21096#bib.bib50);[22](https://arxiv.org/html/2608.21096#bib.bib66)\), resulting in suboptimal representations and complicating heterogeneity modeling\. Figure 2:Averaged Forman\-Ricci curvature across datasets \(Cora, ogbn\-arxiv, and Amazon\-Photo\)\. Higher bars indicate more pronounced non\-Euclidean characteristics in these datasets\.To move beyond a single Euclidean geometry, we theoretically establish the advantages ofLorentz geometryfor PFL \([Section 4\.1](https://arxiv.org/html/2608.21096#S4.SS1)\), showing it serves as a natural testbed for two reasons:First, enhanced representational power: Lorentz space enables low\-distortion representation of the estimated inherent graph properties\([48](https://arxiv.org/html/2608.21096#bib.bib18);[50](https://arxiv.org/html/2608.21096#bib.bib19);[4](https://arxiv.org/html/2608.21096#bib.bib7);[63](https://arxiv.org/html/2608.21096#bib.bib23)\)\. When client graphs exhibit varying Ricci curvature, assigning each client an appropriate hyperbolic curvature supports a more faithful modeling \([Theorem 4\.1](https://arxiv.org/html/2608.21096#S4.Thmtheorem1)\)\. See[Figure 1](https://arxiv.org/html/2608.21096#S1.F1)\(a\), distributions are long\-tailed with varying skewness\. In particular, Client 1 is ‘steeper’ and benefits from a Lorentz space with larger curvature magnitude \(smallerKK\), which offers a ‘roomier’ embedding environment where tail nodes can be separated\.Second,natural heterogeneity encoding: The additional time\-like dimension in Lorentz space provides a carrier to capture intrinsic geometric heterogeneity across clients \([Theorem 4\.3](https://arxiv.org/html/2608.21096#S4.Thmtheorem3)\)\. In the example \([Figure 1](https://arxiv.org/html/2608.21096#S1.F1)\(b\)\)111For convenience, all origins of Lorentz spaces in the figure are shown as the same, but in reality, their origins are not in the same location\., heterogeneous properties such as “how significant is the imbalance between tail nodes and head nodes?” can be naturally distinguished in Lorentz space through the time\-like dimensionxtx\_\{t\}, while common information \(e\.g\., “the star is a tail node”\) remains preserved in the space\-like dimensions𝐱s\\mathbf\{x\}\_\{s\}\("Flatland"\) as a shared node representation\. Based on the insights, we propose𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}, an exploratory PFL framework that embeds client data in tailored Lorentz spaces to faithfully capture their intrinsic geometry \([Section 5](https://arxiv.org/html/2608.21096#S5)\)\. Yet such embedding alone cannot resolve the heterogeneity challenge in parameter aggregation\. Therefore, we further leverage the nature of the time\-like dimension and develop a theoretically groundedparameter decoupling strategythat designates heterogeneity\-related parameters as personalized, while aggregating only those carrying shared information\. This design effectively mitigates heterogeneity without auxiliary modules or client\-similarity estimation and preserves the validity of Lorentz geometry\. To the best of our knowledge, this is the first work to bridge hyperbolic geometry and PFL for addressing client heterogeneity in a principled manner, providing a newsuccinctandeffectiveperspective that leverages geometric properties\. Experimental results demonstrate that𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}achieves superior performance than its Euclidean counterpart, particularly in low\-dimensional settings that are crucial for communication\-efficient federated learning\. ## 2Related Work ##### Personalized federated learning on graphs\. Personalized federated learning \(PFL\) addresses statistical heterogeneity by learning client\-adapted models instead of a single global model\([26](https://arxiv.org/html/2608.21096#bib.bib32);[15](https://arxiv.org/html/2608.21096#bib.bib33);[25](https://arxiv.org/html/2608.21096#bib.bib34);[10](https://arxiv.org/html/2608.21096#bib.bib47);[7](https://arxiv.org/html/2608.21096#bib.bib49)\)\. For graph data, existing personalized federated graph learning methods commonly cluster clients by gradients\([61](https://arxiv.org/html/2608.21096#bib.bib59)\), introduce additional personalized modules\([58](https://arxiv.org/html/2608.21096#bib.bib50)\), or estimate client similarities for customized aggregation\([5](https://arxiv.org/html/2608.21096#bib.bib51)\)\. These approaches are effective but usually operate in Euclidean spaces and rely on client\-similarity estimation or auxiliary components to handle heterogeneity\. Such designs do not explicitly model the scale\-free and hierarchical structures widely observed in real\-world graphs\([1](https://arxiv.org/html/2608.21096#bib.bib16);[30](https://arxiv.org/html/2608.21096#bib.bib17)\), motivating a geometry\-aware treatment of personalized graph FL\. ##### Hyperbolic federated learning\. Hyperbolic representations provide a natural geometry for hierarchical and power\-law data\([48](https://arxiv.org/html/2608.21096#bib.bib18);[6](https://arxiv.org/html/2608.21096#bib.bib9)\)\. Recent federated methods leverage hyperbolic distances for knowledge distillation\([2](https://arxiv.org/html/2608.21096#bib.bib54)\), hyperbolic prototypes for non\-IID learning\([38](https://arxiv.org/html/2608.21096#bib.bib53)\), or hyperbolic GNNs inside a FedAvg\-style graph FL pipeline\([14](https://arxiv.org/html/2608.21096#bib.bib52)\)\. However, these methods do not personalize the underlying client geometry or separate client\-specific geometric information from shared knowledge during aggregation\.𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}differs by assigning each client a tailored Lorentz space and by decoupling time\-like personalized parameters from space\-like shared parameters, enabling direct aggregation without client clustering or extra similarity estimation\. A more detailed discussion is provided in[AppendixA](https://arxiv.org/html/2608.21096#A1)\. ## 3Preliminaries ##### Lorentz Model of Hyperbolic Geometry\. The Lorentz model, also known as the hyperboloid model, is a representation ofRiemannianhyperbolic space embedded in flat Minkowski ambient spaceℝd\+1\\mathbb\{R\}^\{d\+1\}\([48](https://arxiv.org/html/2608.21096#bib.bib18);[9](https://arxiv.org/html/2608.21096#bib.bib27)\)\. Given add\-dimensional Lorentz manifoldℒKd\\mathcal\{L\}\_\{K\}^\{d\}with a constant negative curvature−1/K\(K\>0\)\-1/K\(K\>0\), suppose a point𝐱∈ℒKd\\mathbf\{x\}\\in\\mathcal\{L\}\_\{K\}^\{d\}, which has the form𝐱=\[xt𝐱s\]⊤∈ℝd\+1\\mathbf\{x\}=\\begin\{bmatrix\}x\_\{t\}&\\mathbf\{x\}\_\{s\}\\end\{bmatrix\}^\{\\top\}\\in\\mathbb\{R\}^\{d\+1\}, where the first dimensionxt∈ℝx\_\{t\}\\in\\mathbb\{R\}is calledtime\-likedimension and others𝐱s∈ℝd\\mathbf\{x\}\_\{s\}\\in\\mathbb\{R\}^\{d\}arespace\-likedimensions\. It satisfies the following conditions:⟨𝐱,𝐱⟩ℒ=−K\\langle\\mathbf\{x\},\\mathbf\{x\}\\rangle\_\{\\mathcal\{L\}\}=\-Kandxt\>0x\_\{t\}\>0, where⟨𝐱,𝐲⟩ℒ=−xtyt\+𝐱s⊤𝐲s\\langle\\mathbf\{x\},\\mathbf\{y\}\\rangle\_\{\\mathcal\{L\}\}=\-x\_\{t\}y\_\{t\}\+\\mathbf\{x\}\_\{s\}^\{\\top\}\\mathbf\{y\}\_\{s\}is the Lorentzian inner product\. Here,KKis the positive Lorentz scale parameter; differentKKvalues induce different hyperbolic geometries, with sectional curvature−1/K\-1/K\. Formal definitions are shown in[AppendixB\.1](https://arxiv.org/html/2608.21096#A2.SS1)\. Typically, inputs reside in Euclidean space and need to be mapped into hyperbolic space\. The way of projecting the data𝐯E∈ℝd\\mathbf\{v\}^\{E\}\\in\\mathbb\{R\}^\{d\}in Euclidean to Lorentz space𝐱∈ℒKd\\mathbf\{x\}\\in\\mathcal\{L\}\_\{K\}^\{d\}can be simplified as222For clarity, all Lorentz space embeddings are denoted by⋅H\\cdot^\{H\}\. Specifically, if the Lorentz scale parameter is known asKK, it is denoted by⋅K\\cdot^\{K\}\. In contrast, Euclidean space embeddings are denoted by⋅E\\cdot^\{E\}\. All vectors𝐱\\mathbf\{x\}, if not superscripted, are assumed to be in Lorentz space\. 𝐱K\\displaystyle\\mathbf\{x\}^\{K\}=exp𝐨K\(𝐯E\)=exp𝐨K\(\[0,𝐯E\]\)=\[xt,𝐱s\]⊤,\\displaystyle=\\exp\_\{\\mathbf\{o\}\}^\{K\}\\\!\\left\(\\mathbf\{v\}^\{E\}\\right\)=\\exp\_\{\\mathbf\{o\}\}^\{K\}\\\!\\left\(\\left\[0,\\mathbf\{v\}^\{E\}\\right\]\\right\)=\\left\[x\_\{t\},\\mathbf\{x\}\_\{s\}\\right\]^\{\\top\},\(1\)xt\\displaystyle x\_\{t\}=Kcosh\(‖𝐯E‖2K\),\\displaystyle=\\sqrt\{K\}\\cosh\\\!\\left\(\\frac\{\\\|\\mathbf\{v\}^\{E\}\\\|\_\{2\}\}\{\\sqrt\{K\}\}\\right\),𝐱s\\displaystyle\\mathbf\{x\}\_\{s\}=Ksinh\(‖𝐯E‖2K\)𝐯E‖𝐯E‖2\.\\displaystyle=\\sqrt\{K\}\\sinh\\\!\\left\(\\frac\{\\\|\\mathbf\{v\}^\{E\}\\\|\_\{2\}\}\{\\sqrt\{K\}\}\\right\)\\frac\{\\mathbf\{v\}^\{E\}\}\{\\\|\\mathbf\{v\}^\{E\}\\\|\_\{2\}\}\.DifferentKKmaps𝐯E\\mathbf\{v\}^\{E\}to different Lorentz surfaces\. ##### Fully Lorentz Neural Networks\. Fully Lorentz network\([9](https://arxiv.org/html/2608.21096#bib.bib27)\)has been shown to be ideal for PFL due to their reduced need for space projections, enhancing computational efficiency\. These networks also incorporate Lorentz transformations \(boosts and rotations\), improving data heterogeneity handling and parameter interpretability \([AppendixB\.3](https://arxiv.org/html/2608.21096#A2.SS3)\)\. Given an input vector𝐱∈ℒKn\\mathbf\{x\}\\in\\mathcal\{L\}\_\{K\}^\{n\}, and a linear layer matrix𝐖∈ℝm×\(n\+1\)\\mathbf\{W\}\\in\\mathbb\{R\}^\{m\\times\(n\+1\)\}to optimize, the fully Lorentz linear layer can be denoted asLT\\mathrm\{LT\}in a general form asLT\(𝐱,f,𝐖\):=\(‖f\(𝐖𝐱\)‖2\+K,f\(𝐖𝐱\)\)T,\\mathrm\{LT\}\(\\mathbf\{x\};f;\\mathbf\{W\}\):=\\left\(\{\\sqrt\{\\\|f\(\\mathbf\{Wx\}\)\\\|^\{2\}\+K\}\},\{f\(\\mathbf\{Wx\}\)\}\\right\)^\{T\},whereffis a function like activation, dropout, and bias\. ##### Problem Statement\. Given clients𝒞=\{1,2,…,C\}\\mathcal\{C\}=\\\{1,2,\\ldots,C\\\}, each with a dataset𝒟c=\(𝐱ic,yic\)i=1Nc\\mathcal\{D\}\_\{c\}=\{\(\\mathbf\{x\}\_\{i\}^\{c\},y\_\{i\}^\{c\}\)\}\_\{i=1\}^\{N\_\{c\}\}and distributionpc\(𝐱,y\)p\_\{c\}\(\\mathbf\{x\},y\), Personalized Federated Learning \(PFL\) encounters heterogeneity ifpi\(𝐱,y\)≠pj\(𝐱,y\)p\_\{i\}\(\\mathbf\{x\},y\)\\neq p\_\{j\}\(\\mathbf\{x\},y\)for any client pairi≠ji\\neq j, which degrades performance\. In PFL, the goal is to optimize personalized modelsfc\(⋅,𝜽c,𝜽s\)f\_\{c\}\(\\cdot;\\bm\{\\theta\}\_\{c\},\\bm\{\\theta\}\_\{s\}\)for each client using specific and shared parameters𝜽c\\bm\{\\theta\}\_\{c\},𝜽s\\bm\{\\theta\}\_\{s\}: min𝜽c\|c=1C,𝜽s\\displaystyle\\min\_\{\{\\bm\{\\theta\}\_\{c\}\}\|\_\{c=1\}^\{C\},\\bm\{\\theta\}\_\{s\}\}∑c=1C𝔼\(𝐱,y\)∼pc\(𝐱,y\)\[ℒc\(f\(𝐱,𝜽c,𝜽s\),y\)\]\\displaystyle\\sum\_\{c=1\}^\{C\}\\mathbb\{E\}\_\{\(\\mathbf\{x\},y\)\\sim p\_\{c\}\(\\mathbf\{x\},y\)\}\[\\mathcal\{L\}\_\{c\}\(f\(\\mathbf\{x\};\\bm\{\\theta\}\_\{c\},\\bm\{\\theta\}\_\{s\}\),y\)\]\(2\)\+λΩ\(𝜽c\|c=1C,𝜽s\)\.\\displaystyle\+\\lambda\\Omega\(\{\\bm\{\\theta\}\_\{c\}\}\|\_\{c=1\}^\{C\},\\bm\{\\theta\}\_\{s\}\)\. This function merges local lossℒc\\mathcal\{L\}\_\{c\}with regularizationΩ\\Omega, balanced by hyperparameterλ\\lambda\. Our goalsare\(1\)toeffectivelyrepresent the inherent properties of each local client data;\(2\)tosuccinctlyreflect heterogeneity among client data and facilitate the communication of shared information without requiring additional computations\. ## 4Motivation and Insights We investigate PFL for graph data from a geometric perspective via Ricci curvature \([AppendixB\.2](https://arxiv.org/html/2608.21096#A2.SS2)\), which distinguishes hyperbolic \(negative\), flat \(zero\), and spherical \(positive\) geometries and quantifies curvature strength through its magnitude\. This provides a principled measure of the intrinsic geometry of client graphs, enabling us to analyze structural differences beyond conventional statistical heterogeneity\.[Figure 2](https://arxiv.org/html/2608.21096#S1.F2)and[AppendixC\.1](https://arxiv.org/html/2608.21096#A3.SS1)show the empirical results on real\-world datasets across clients, from which we identifytwo consistent patterns: \(1\) client graphs mostly have*negative*curvature, evidencing hyperbolic structure, and \(2\) curvaturevalues vary considerablyacross clients, indicating geometric heterogeneity beyond statistical differences\. These findings confirm that assuming a unified Euclidean geometry leads to distorted representations and ineffective heterogeneity modeling\([48](https://arxiv.org/html/2608.21096#bib.bib18);[50](https://arxiv.org/html/2608.21096#bib.bib19);[4](https://arxiv.org/html/2608.21096#bib.bib7)\), underscoring the necessity of exploring solutions beyond Euclidean space\. ### 4\.1Motivation: Why Lorentz Space for PFL? In this section, we bridge non\-Euclidean geometry and PFL, and theoretically claim that the Lorentz geometry of hyperbolic space is particularly suitable, as it faithfully captures the non\-Euclidean properties of client data and aligns with the goals outlined in[Section 3](https://arxiv.org/html/2608.21096#S3)\. Why the Lorentz \(hyperboloid\) model specifically?While any model of hyperbolic space is geometrically equivalent, the Lorentz model provides unique structural advantages for PFL: \(1\)Time–space decomposition: The natural split into time\-like and space\-like coordinates directly supports our parameter decoupling strategy, where time\-like parameters capture heterogeneity and space\-like parameters are safely aggregated\. \(2\)Block\-structured isometries: Lorentz transformations have a linear\-algebraic form that guarantees aggregated parameters remain on the hyperbolic manifold \([Proposition6\.1](https://arxiv.org/html/2608.21096#S6.Thmtheorem1)\)\. \(3\)Closed\-form curvature\-aware gradients: The exponential map yields analytically transparent, curvature\-dependent gradient weighting\. Other hyperbolic models \(e\.g\., Poincaré ball\) lack an intrinsically privileged direction for such clean decomposition\. For Goal[\(1\)](https://arxiv.org/html/2608.21096#S3.I1.i1)\.Prevalent non\-Euclidean heterogeneity can be captured by hyperbolic curvature\. The observednegativecurvature shows that hyperbolic space naturally fits client graphs with such properties\([64](https://arxiv.org/html/2608.21096#bib.bib20)\), and its curvature can be adjusted to accommodatevariedclient distributions\([30](https://arxiv.org/html/2608.21096#bib.bib17)\)\. Next, we theorize the use of hyperbolic geometry in PFL\. ###### Theorem 4\.1\(Necessity of tailored curvature\)\. Let\{Gc\}c=1C\\\{G\_\{c\}\\\}\_\{c=1\}^\{C\}be client graphs with average Forman\-Ricci curvaturesR¯c=Ric¯\(Gc\)\\bar\{R\}\_\{c\}=\\overline\{\\mathrm\{Ric\}\}\(G\_\{c\}\), and letℒKd\\mathcal\{L\}^\{d\}\_\{K\}denote thedd\-dimensional hyperbolic space with Lorentz scale parameterK\>0K\>0and constant sectional curvature−1/K<0\-1/K<0\. For each clientcc, letεc∗\(K\)\\varepsilon\_\{c\}^\{\*\}\(K\)be the minimal edge distortion of any\(1\+ε\)\(1\+\\varepsilon\)\-bi\-Lipschitz embeddingfc:Gc→ℒKdf\_\{c\}:G\_\{c\}\\to\\mathcal\{L\}^\{d\}\_\{K\}\. Then the following holds: max1≤c≤Cεc∗\(K\)≥cd2max1≤i<j≤C\|R¯i−R¯j\|,\\max\_\{1\\leq c\\leq C\}\\ \\varepsilon\_\{c\}^\{\*\}\(K\)\\;\\;\\geq\\;\\;\\frac\{c\_\{d\}\}\{2\}\\,\\max\_\{1\\leq i<j\\leq C\}\\,\|\\bar\{R\}\_\{i\}\-\\bar\{R\}\_\{j\}\|,\(3\)wherecd\>0c\_\{d\}\>0is a dimension\-dependent constant \(Proof in[AppendixD\.1](https://arxiv.org/html/2608.21096#A4.SS1)\)\. For Goal[\(2\)](https://arxiv.org/html/2608.21096#S3.I1.i2)\.Strong correlation between heterogeneity and hyperbolic time\-like dimension\. ###### Theorem 4\.3\. LetC∈\{1,…,m\}C\\in\\\{1,\\dots,m\\\}, each clientcchave Lorentz scale parameterKc\>0K\_\{c\}\>0withVar\(KC\)\>0\\mathrm\{Var\}\(K\_\{C\}\)\>0, and𝐱=\[xt𝐱s\]⊤∈𝕃Kcd\\mathbf\{x\}=\[x\_\{t\}\\;\\mathbf\{x\}\_\{s\}\]^\{\\top\}\\in\\mathbb\{L\}^\{d\}\_\{K\_\{c\}\}admit hyperbolic polar coordinates\(ρ,𝐮\)\(\\rho,\\mathbf\{u\}\)withxt=Kccosh\(ρ/Kc\)x\_\{t\}=\\sqrt\{K\_\{c\}\}\\cosh\(\\rho/\\sqrt\{K\_\{c\}\}\),𝐱s=Kcsinh\(ρ/Kc\)𝐮\\mathbf\{x\}\_\{s\}=\\sqrt\{K\_\{c\}\}\\sinh\(\\rho/\\sqrt\{K\_\{c\}\}\)\\,\\mathbf\{u\},𝐮:=𝐱s/‖𝐱s‖\\mathbf\{u\}:=\\mathbf\{x\}\_\{s\}/\\\|\\mathbf\{x\}\_\{s\}\\\|\. Then \(1\)I\(𝐮,C\)=0I\(\\mathbf\{u\};C\)=0; \(2\)I\(xt,C\)\>0I\(x\_\{t\};C\)\>0if the pushforward measures ofTK\(ρ\):=Kcosh\(ρ/K\)T\_\{K\}\(\\rho\):=\\sqrt\{K\}\\cosh\(\\rho/\\sqrt\{K\}\)differ acrossKK; \(3\)I\(\(xt,𝐱s\);C∣Kc\)=I\(xt;C∣Kc\)I\\big\(\(x\_\{t\},\\mathbf\{x\}\_\{s\}\);C\\mid K\_\{c\}\\big\)=I\(x\_\{t\};C\\mid K\_\{c\}\)\(Proof in[AppendixD\.2](https://arxiv.org/html/2608.21096#A4.SS2)\)\. ### 4\.2Insights: Introduce a Higher Dimension \(time dimension\) to"Flatland" In the above case,"Flatland"captures the common feature of a cylinder or a sphere, while a higher dimension \(the third dimension\) highlights the differences between the objects\. Analogous to our setting, informally speaking, by introducing an additionaltime\-likedimension, we can imagine each client’s data residing in a unique Lorentz space \(a curved world in a higher\-dimensional space\), where the curvature reflects the distinct distributions \(objects\)\."Flatland"333Our method is named after Edwin Abbott’s book ”Flatland: A Romance of Many Dimensions”, highlighting our insights of exploring a new perspective that maps various data distributions onto different Lorentz surfaces of hyperbolic geometry\.,ℝd\\mathbb\{R\}^\{d\}\(flat\), serves as a metaphor for a platform where common information \(circle\) is exchanged and integrated\. ![[Uncaptioned image]](https://arxiv.org/html/2608.21096v1/Figures/icon/attention-to-detail.png)In "Flatland", a two\-dimensional flat plane, the same shapes may represent the projections of various three\-dimensional objects\. For instance, a circle could be the projection of either a cylinder or a sphere from a higher dimension\. ## 5The𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}Framework We propose𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}, using tailored Lorentz spaces for each client and a well\-designed parameter decoupling strategy to mitigate heterogeneity\. The main steps are outlined in[Figure 3](https://arxiv.org/html/2608.21096#S5.F3)and[Algorithm 2](https://arxiv.org/html/2608.21096#alg2)\. Our method is succinct, directly built upon FedAvg, and requires no additional clustering computations or auxiliary modules; further details are provided in[AppendixC\.2](https://arxiv.org/html/2608.21096#A3.SS2)\. Figure 3:The𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}framework\. It comprises three stages: \(1\) Initialization of the client\-specific Lorentz scale \([Section 5\.1](https://arxiv.org/html/2608.21096#S5.SS1)\), personalized, and shared parameters; \(2\) Local updates within client\-specific Lorentz spaces \([AppendixC\.2\.3](https://arxiv.org/html/2608.21096#A3.SS2.SSS3)\); and \(3\) Server updates, aggregating only shared parameters while preserving personalization by keeping personalized parameters locally \([Section 5\.2](https://arxiv.org/html/2608.21096#S5.SS2)\)\.### 5\.1Learnable Curvature Initialization Each client is assigned a learnable Lorentz scale parameterKcK\_\{c\}\(sectional curvature−1/Kc\-1/K\_\{c\}\), rather than a fixed pre\-computed estimate\. Client heterogeneity is not solely structural: node features, label distributions, and task\-specific local optimization can also affect the preferred geometry, making any single pre\-training statistic insufficient\. We use Forman\-Ricci curvature only as a lightweight structure\-aware initialization prior\. Given client graphGcG\_\{c\}, its average graph curvature initializes a raw scalar, which is mapped to a positiveKcK\_\{c\}by sigmoid reparameterization and updated during local training\. This follows prior work treating curvature as a trainable geometric quantity rather than a fixed scaling constant\([19](https://arxiv.org/html/2608.21096#bib.bib11);[65](https://arxiv.org/html/2608.21096#bib.bib5);[18](https://arxiv.org/html/2608.21096#bib.bib62)\)\. Details are provided in[AppendixC\.2\.2](https://arxiv.org/html/2608.21096#A3.SS2.SSS2)\. ### 5\.2Parameter Decoupling Strategy Since each client has its own Lorentz space under the Lorentz model in[Equation \(1\)](https://arxiv.org/html/2608.21096#S3.E1), the Lorentz scale parameterKcK\_\{c\}is kept client\-specific\. We further split the transformation parameters𝐌^\\hat\{\\mathbf\{M\}\}of the Lorentz linear layer into personalized parameters𝜽c\\bm\{\\theta\}\_\{c\}and shared parameters𝜽s\\bm\{\\theta\}\_\{s\}\. The split follows the Lorentz coordinates:𝜽s\\bm\{\\theta\}\_\{s\}carries transferable information inspace\-likedimensions, while𝜽c\\bm\{\\theta\}\_\{c\}captures client\-specific heterogeneity through thetime\-likedimension\. For clarity, we derive the split on a Lorentz linear layer and omit the auxiliary functionsff\. Given input𝐱\(l\)=\[xt\(l\)𝐱s\(l\)\]⊤∈ℒKn\\mathbf\{x\}^\{\(l\)\}=\\begin\{bmatrix\}x\_\{t\}^\{\(l\)\}&\\mathbf\{x\}\_\{s\}^\{\(l\)\}\\end\{bmatrix\}^\{\\top\}\\in\\mathcal\{L\}\_\{K\}^\{n\}at layerll, wherext\(l\)∈ℝx\_\{t\}^\{\(l\)\}\\in\\mathbb\{R\}and𝐱s\(l\)∈ℝn\\mathbf\{x\}\_\{s\}^\{\(l\)\}\\in\\mathbb\{R\}^\{n\}\. We decompose the learnable matrix𝐌^\(l\)\\hat\{\\mathbf\{M\}\}^\{\(l\)\}as\[𝐦\(l\)𝐌\(l\)\]∈ℝm×\(n\+1\)\\begin\{bmatrix\}\\mathbf\{m\}^\{\(l\)\}&\\mathbf\{M\}^\{\(l\)\}\\end\{bmatrix\}\\in\\mathbb\{R\}^\{m\\times\(n\+1\)\}, with𝐦\(l\)∈ℝm×1\\mathbf\{m\}^\{\(l\)\}\\in\\mathbb\{R\}^\{m\\times 1\}and𝐌\(l\)∈ℝm×n\\mathbf\{M\}^\{\(l\)\}\\in\\mathbb\{R\}^\{m\\times n\}\. Here,𝐦\(l\)\\mathbf\{m\}^\{\(l\)\}controls the time\-like contribution, whereas𝐌\(l\)\\mathbf\{M\}^\{\(l\)\}operates on the transferable space\-like dimensions\. Let𝐳\(l\)=𝐦\(l\)xt\(l\)\+𝐌\(l\)𝐱s\(l\)∈ℝm\\mathbf\{z\}^\{\(l\)\}=\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\+\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}\\in\\mathbb\{R\}^\{m\}\. The output𝐱\(l\+1\)\\mathbf\{x\}^\{\(l\+1\)\}is 𝐱\(l\+1\)=LT\(𝐱\(l\),𝐌^\(l\)\)=\(‖𝐳\(l\)‖2\+K,𝐳\(l\)\)T\.\\mathbf\{x\}^\{\(l\+1\)\}=\\mathrm\{LT\}\(\\mathbf\{x\}^\{\(l\)\};\\hat\{\\mathbf\{M\}\}^\{\(l\)\}\)=\\left\(\\sqrt\{\\\|\\mathbf\{z\}^\{\(l\)\}\\\|^\{2\}\+K\},\\mathbf\{z\}^\{\(l\)\}\\right\)^\{T\}\.\(4\) ##### Decoupled aggregation\. Guided by the derivation in[AppendixC\.3](https://arxiv.org/html/2608.21096#A3.SS3), we federate only the parameters associated with space\-like dimensions\. The reason is twofold\. First,[Theorem 4\.3](https://arxiv.org/html/2608.21096#S4.Thmtheorem3)indicates that client\-specific heterogeneity is expressed through the time\-like coordinate\. Since𝐦\(l\)\\mathbf\{m\}^\{\(l\)\}directly weightsxt\(l\)x\_\{t\}^\{\(l\)\}in𝐳\(l\)\\mathbf\{z\}^\{\(l\)\}, averaging it across clients would mix scale\-sensitive quantities from different Lorentz spaces\. Second,KcK\_\{c\}determines the Lorentz surface where clientccembeds and updates its data, so it must remain local\. After local training, clientccuploads only\{𝐌c\(l\)\}l=1L\\\{\\mathbf\{M\}^\{\(l\)\}\_\{c\}\\\}\_\{l=1\}^\{L\}, and the server averages these space\-like parameters\. The averaged𝐌\(l\)\\mathbf\{M\}^\{\(l\)\}is then combined locally with𝐦\(l\)\\mathbf\{m\}^\{\(l\)\}andKcK\_\{c\}\. This does not assume identical space\-like features across clients; rather,𝐌\\mathbf\{M\}acts on transferable space\-like coordinates, while𝐦\(l\)\\mathbf\{m\}^\{\(l\)\}andKcK\_\{c\}remain personalized\. By[Proposition6\.1](https://arxiv.org/html/2608.21096#S6.Thmtheorem1), replacing𝐌\\mathbf\{M\}with its aggregated counterpart still maps outputs to the same client\-specific Lorentz spaceℒKc\\mathcal\{L\}\_\{K\_\{c\}\}\. For the transformation parameters of a modelℳ\\mathcal\{M\}withLLlayers:𝜽c=⋃l=1L\{𝐦\(l\)\},𝜽s=⋃l=1L\{𝐌\(l\)\},\\bm\{\\theta\}\_\{c\}=\\bigcup\_\{l=1\}^\{L\}\\\{\\mathbf\{m\}^\{\(l\)\}\\\},\\qquad\\bm\{\\theta\}\_\{s\}=\\bigcup\_\{l=1\}^\{L\}\\\{\\mathbf\{M\}^\{\(l\)\}\\\},where𝜽c\\bm\{\\theta\}\_\{c\}and𝜽s\\bm\{\\theta\}\_\{s\}denote thepersonalizedandsharedtransformation parameter sets, respectively\. ## 6Analysis This section analyzes three properties of𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}:Correctness, showing that client representations remain valid in Lorentz space during federated communication;Convergence, comparing its convergence behavior with FedAvg; andEfficiency, quantifying computational overhead\. Table 1:Comparison of node classification performance across real\-world datasets with varying numbers of clients\. The results, presented as mean and standard deviation, are based on five separate trials\. Performances that are statistically significant \(p<0\.05p<0\.05\) are highlighted in bold\.Table 2:Comparison of node classification performance across heterophilic datasets with varying numbers of clients\. The results, presented as mean and standard deviation, are based on five separate trials\. Performances that are statistically significant \(p<0\.05p<0\.05\) are highlighted in bold\.##### Correctness\. Although a fully Lorentz neural network ensures that representations remain in hyperbolic space during local training, as guaranteed by[LemmaD\.12](https://arxiv.org/html/2608.21096#A4.Thmtheorem12), we further need to verify that our proposed decoupling strategy also preserves this property after server\-side aggregation\. ###### Proposition 6\.1\. Let𝐌^=\[𝐦𝐌\]\\hat\{\\mathbf\{M\}\}=\\begin\{bmatrix\}\\mathbf\{m\}&\\mathbf\{M\}\\end\{bmatrix\}, where𝐌^∈ℝm×\(n\+1\)\\hat\{\\mathbf\{M\}\}\\in\\mathbb\{R\}^\{m\\times\(n\+1\)\},𝐦∈ℝm×1\\mathbf\{m\}\\in\\mathbb\{R\}^\{m\\times 1\}, and𝐌∈ℝm×n\\mathbf\{M\}\\in\\mathbb\{R\}^\{m\\times n\}\. LetΦ\(𝐌^,𝐍\)=\[𝐦𝐍\]\\Phi\\left\(\\hat\{\\mathbf\{M\}\},\\mathbf\{N\}\\right\)=\\begin\{bmatrix\}\\mathbf\{m\}&\\mathbf\{N\}\\end\{bmatrix\}, where𝐍∈ℝm×n\\mathbf\{N\}\\in\\mathbb\{R\}^\{m\\times n\}is the aggregated shared parameter matrix obtained from\{𝐌i\}i=1C\\\{\\mathbf\{M\}\_\{i\}\\\}\_\{i=1\}^\{C\}using the proposed decoupling strategy\. For all𝐱∈ℒKn\\mathbf\{x\}\\in\\mathcal\{L\}\_\{K\}^\{n\}, we haveLT\(𝐱,Φ\(𝐌^,𝐍\)\)∈ℒKm\.\\mathrm\{LT\}\\left\(\\mathbf\{x\};\\Phi\\left\(\\hat\{\\mathbf\{M\}\},\\mathbf\{N\}\\right\)\\right\)\\in\\mathcal\{L\}\_\{K\}^\{m\}\. [Proposition6\.1](https://arxiv.org/html/2608.21096#S6.Thmtheorem1)\(refer to the proof in[AppendixD\.3](https://arxiv.org/html/2608.21096#A4.SS3)\) implies that, even after the aggregation of shared parameters on the server, the transformation of any client vector𝐱∈ℒKn\\mathbf\{x\}\\in\\mathcal\{L\}\_\{K\}^\{n\}by this updated matrix will still yield results in the Lorentz spaceℒKm\\mathcal\{L\}\_\{K\}^\{m\}with the same Lorentz scale parameter, and hence the same sectional curvature\. ##### Convergence\. We analyze in[AppendixD\.4](https://arxiv.org/html/2608.21096#A4.SS4)whether our method hinders the convergence rate\. The results show that, under the standard FedAvg scheme, our method does not affect the convergence rate, which remains at𝒪\(1/T\)\\mathcal\{O\}\(1/T\)\. The key reason is that𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}changes only the geometric parameterization of local updates\. Server\-side aggregation still averages shared space\-like parameters as in FedAvg, without estimating client similarity\. ##### Efficiency\. We provide the time complexity analysis in[AppendixC\.4](https://arxiv.org/html/2608.21096#A3.SS4)\.𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}introduces minimal operations, like theO\(1\)O\(1\)exponential map and curvature estimation, which can be mitigated by pre\-computation\. These minimal costs are offset by reduced communication overhead and enhanced representation in Lorentz space, making𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}efficient for practical personalized FL\. Figure 4:Performance of CiteSeer \(20 clients\) with varying dimensions for node classification scenario\.Figure 5:Ablation study of𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}on the Cora dataset\. In summary,𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}preserves correctness by keeping representations in Lorentz space after aggregation, achieves the same convergence rate as FedAvg, and incurs only minimal overhead comparable to FedAvg\. Besides, we further justify the rationale of our method from the perspective of Lorentz transformations in[AppendixD\.5](https://arxiv.org/html/2608.21096#A4.SS5)\. ## 7Experiments In this section, we validate the effectiveness of𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}through experiments on*node classification*and*graph classification*on a series of benchmark datasets\. The experiments are designed to address the following research questions\.RQ1\.Can𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}outperform personalized and hyperbolic FL baselines?RQ2\.Can𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}still perform well in low\-dimensional settings?RQ3\.Can𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}maintain high performance under partial client participation in FL?RQ4\.Are the proposed novel components really beneficial? ### 7\.1Experimental Setup ##### Datasets and Baselines\. The details about datasets are listed in[AppendixE\.1](https://arxiv.org/html/2608.21096#A5.SS1)\. Implementation details are shown in[AppendixE\.2](https://arxiv.org/html/2608.21096#A5.SS2)\. To assess𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}and demonstrate its superiority, we compare it with the following baselines: \(1\) Local: clients train their models locally without any communication\. Local \(EE\) refers to self\-training in the Euclidean model, while Local \(LL\) refers to training in the Lorentz model; \(2\) FedAvg\([45](https://arxiv.org/html/2608.21096#bib.bib28)\)and \(3\) FedProx\([33](https://arxiv.org/html/2608.21096#bib.bib29)\): the most popular FL baselines; \(4\) FedPer\([3](https://arxiv.org/html/2608.21096#bib.bib42)\): a PFL baseline with personalized model layers; \(5\) GCFL\([61](https://arxiv.org/html/2608.21096#bib.bib59)\): a PFGL baseline with client clustering and cluster\-wise model aggregation; \(6\) FedGNN\([60](https://arxiv.org/html/2608.21096#bib.bib30)\)and \(7\) FedSage\+\([67](https://arxiv.org/html/2608.21096#bib.bib31)\): two FGL baselines; \(8\) FED\-PUB\([5](https://arxiv.org/html/2608.21096#bib.bib51)\): a PFGL baseline with personalized model aggregation and local weight masking; \(9\) FedGTA\([37](https://arxiv.org/html/2608.21096#bib.bib68)\)introduces a personalized optimization strategy that leverages topology\-aware local smoothing confidence and mixed neighbor features; \(10\) AdaFGL\([36](https://arxiv.org/html/2608.21096#bib.bib69)\)addresses structural non\-IID challenges by introducing a decoupled, two\-stage personalized learning strategy ; \(11\) FedHGCN\([14](https://arxiv.org/html/2608.21096#bib.bib52)\): a hyperbolic FGL baseline that omits explicit client\-geometry modeling\. Together, these baselines cover local, standard FL, personalized FL, graph FL, and hyperbolic graph FL settings; further selection rationale is provided in[AppendixE\.3](https://arxiv.org/html/2608.21096#A5.SS3)\. ### 7\.2Main Experimental Results \(RQ1\) Table 3:Performance on graph classification tasks\. The results are reported as mean±\\pmstandard deviation over five runs\. Bold indicates statistical significance \(p<0\.05p<0\.05\)\.##### Node Classification\. We tackle node classification onhighly heterogeneous homophilic and heterophilic datasets, with non\-overlapping node partitions for each client, which many previous methods are not designed to address\. This challenge highlights our method’s ability to handle heterogeneity that previous approaches could not address\. Tables[2](https://arxiv.org/html/2608.21096#S6.T2)and[2](https://arxiv.org/html/2608.21096#S6.T2)show that our proposed𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}achieves the best or competitive performance in most settings, with particularly clear gains on CiteSeer, ogbn\-arxiv, and Minesweeper\. \(1\) Local \(LL\) often surpasses Local \(EE\), suggesting that hyperbolic space can better represent most datasets, with the difference being particularly pronounced in heterophilic graphs\. \(2\) Vanilla Euclidean FL baselines such as FedAvg, FedProx, and FedGNN often underperform local training under strong heterogeneity\. Among stronger Euclidean/PFGL baselines, FED\-PUB is often the strongest on homophilic node datasets, while FedGTA is particularly competitive on Photo and FedSage\+, FED\-PUB, and AdaFGL remain competitive on some heterophilic datasets\. \(3\) FedHGCN, despite operating in hyperbolic space, underperforms in many heterogeneous settings because it does not explicitly model client\-specific geometry, akin to FedAvg vs Local \(EE\) in Euclidean space\. Due to the quadratic time and space complexity of FedHGCN’s node selection module, it can easily encounter out\-of\-memory \(OOM\) issues with large datasets like ogbn\-arxiv\. In contrast, our experiments show that𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}effectively mitigates heterogeneity and yields substantial improvements on both highly heterogeneous homophilic datasets \(e\.g\., CiteSeer\) and heterophilic datasets \(e\.g\., Minesweeper and Roman\-Empire\)\. ##### Graph Classification\. [Table 3](https://arxiv.org/html/2608.21096#S7.T3)shows the results of the graph classification task, which is conducted with multiple datasets from one or more domains owned by different clients in each task/setting\. In the single\-dataset CHEM setting, Local \(LL\) outperforms Local \(EE\) due to inherent hyperbolic characteristics better captured by hyperbolic geometry\. However, in multiple\-dataset settings like BIO\-CHEM\-SN, Local \(LL\) fails to surpass Local \(EE\), potentially because not all datasets exhibit prominent hyperbolicfeatures\.With our proposedfederated graph learning approach,𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}can significantly enhance the performance of the Lorentzian model, outperforming the Euclidean baselines, and demonstrating its effectiveness\. Figure 6:Performance comparison between FedAvg, FedPer, FedHGCN, and𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}under different client participation rates on Cora with 50 clients\.Table 4:Ablation study results about the necessity of using Lorentz space to do parameter decoupling\. ### 7\.3Varying Embedding Dimensions \(RQ2\) Lower embedding and hidden dimensions reduce the parameter transmission cost in federated learning, as fewer parameters are communicated between the server and clients during training\. Considering the representational power of hyperbolic spaces in lower dimensions\([6](https://arxiv.org/html/2608.21096#bib.bib9)\), we reduced the embedding dimension from 64 to 4 to evaluate𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}’s ability to mitigate data heterogeneity using compact representations\.[Figure 5](https://arxiv.org/html/2608.21096#S6.F5)shows the results on CiteSeer \(20 clients\), with similar trends observed across datasets\. Dimensionality reduction from 64 to 4 had a relatively small impact on the hyperbolic methods \(𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}and FedHGCN\) compared to their Euclidean counterparts\. Notably, while FedHGCN underperformed Euclidean methods at higher dimensions, it outperformed them when the dimension was reduced to 16\.𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}consistently outperformed all other methods in different embedding dimensions, and its performance advantage over the baselines became increasingly significant as the dimensionality was reduced\. ### 7\.4Partial Client Participation \(RQ3\) We further evaluate whether𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}remains effective when only a subset of clients participates in each communication round\. This setting is common in practical FL systems, where coordinating all clients simultaneously can be difficult\. We conduct experiments on Cora with 50 clients, a larger client\-pool configuration for graph FL\([14](https://arxiv.org/html/2608.21096#bib.bib52)\), and compare FedAvg, FedPer, FedHGCN, and𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}under different participation rates\. [Figure 6](https://arxiv.org/html/2608.21096#S7.F6)shows that𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}is robust to partial participation\. Even with only 10% client participation \(5 clients\),𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}achieves 81\.82% accuracy, while FedAvg reaches only 18\.14%\. Across all participation rates,𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}consistently outperforms all baselines, whereas FedAvg remains low and fluctuates because its aggregated model is sensitive to the sampled clients\. These results suggest that separating personalized time\-like parameters from shared space\-like parameters reduces the damage caused by missing or inconsistent client\-specific updates\. ### 7\.5Ablation Study \(RQ4\) To analyze the contribution of each component, we conduct ablation studies on the proposed design choices\. A unified summary of these ablations is provided in[AppendixE\.4](https://arxiv.org/html/2608.21096#A5.SS4)\. Figure 7:Client\-level performance comparison on Cora \(left\) and CiteSeer \(right\) with 10 clients\. Each group corresponds to one client and reports accuracy of𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}w/o DS,𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}, and Local \(LL\)\.𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}improves most clients, while removing DS causes client\-wise degradation, supporting the role of time\-like parameter decoupling\.The Benefits of Adaptive Curvature\.The "w/o TS" \(without tailored curvature\) in[Figure 5](https://arxiv.org/html/2608.21096#S6.F5)refers to setting a constant Lorentz scale parameterK=1K=1for all clients instead of employing tailored curvature settings\. The results indicate that using a fixed hyperbolic space with a constant Lorentz scale parameter yields inferior performance compared with adaptive tailored curvature\. In contrast, learning client\-specific Lorentz scale parameters allows the full𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}model to surpass the Local \(LL\) baseline, showing that adaptive geometry is important for effective federated transfer\. The Benefits of Time\-like Parameter Decoupling\.The "w/o DS" in[Figure 5](https://arxiv.org/html/2608.21096#S6.F5)refers to no parameter decoupling strategy \(DS\), which exhibits significant fluctuations across rounds because the aggregation process incorporates heterogeneous information, adversely affecting the results\. This highlights the effectiveness of our proposed decoupling strategy and validates that the time\-like dimension can effectively capture heterogeneous information\. Moreover, we analyze the benefits of DS for each client’s performance\. As shown in[Figure 7](https://arxiv.org/html/2608.21096#S7.F7), with client IDs on the x\-axis, Flatland outperforms the local method for the vast majority of clients, notably improving performance for clients with inherently poorer results, like c\_8 in the CiteSeer dataset\. This underscoresthe necessity of federated settings for hyperbolic models\. Without our proposed DS, performance deteriorates significantly \(e\.g\., c\_7 in CiteSeer\), furthervalidating our hypothesis that the time\-like parameter encapsulates crucial heterogeneity information\. The Necessity of Lorentz Space\.We conducted experiments to further evaluate the necessity of using Lorentz space\.[Table 4](https://arxiv.org/html/2608.21096#S7.T4)presents the results of an ablation study on the Lorentz transformation\. FlatLand \(EE\) represents our proposed method with a parameter decoupling strategy implemented using a Euclidean backbone\. Without Lorentz geometry, FlatLand \(EE\) underperforms because the time\-like parameter loses its geometric meaning\. It even falls short of FedPer in most cases, which uses the classifier layer for personalization\. These results validate our hypothesis and underscore the importance of hyperbolic representation\. Robustness of Curvature Initialization\.[AppendixE\.5](https://arxiv.org/html/2608.21096#A5.SS5)shows that𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}remains stable across Ricci, Ollivier, constant, and MLP initialization strategies, indicating that learning client\-specific curvature matters more than the particular initialization rule\. This suggests that curvature initialization mainly provides a warm start, while the learnableKcK\_\{c\}can adapt to each client’s geometry during training\. ## 8Conclusions This paper introduces𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}, an exploratory PFL approach that uses hyperbolic geometry to model heterogeneous client graph distributions\. By assigning clients tailored Lorentz spaces and learning client\-specific Lorentz scale parameters,𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}preserves local geometric structure while avoiding explicit client similarity estimation\. Its parameter decoupling strategy separates personalized time\-like parameters from shared space\-like parameters, allowing the server to aggregate common information while reducing interference from heterogeneous updates\. Experiments on node\- and graph\-level benchmarks show that this geometric design improves personalization, particularly when compact representations are required, and highlights non\-Euclidean geometry as a promising direction for federated personalization in heterogeneous settings\. ##### Limitation and future work\. Hyperbolic geometry is not universally optimal for all graph structures, since some clients may be closer to Euclidean or positively curved geometries\. Future work can extend𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}to mixed\-curvature spaces and evaluate richer Lorentz backbones and broader data modalities beyond graph benchmarks\. ## Acknowledgments Jiahong Liu and Irwin King acknowledge partial support from RGC of Hong Kong SAR, China \(CUHK 2300246, RGC C1043\-24G; CUHK 14203425, RGC GRF 2151317\)\. Menglin Yang acknowledges partial support from the General Program of Guangdong Provincial Natural Science Foundation \(No\. 2026A1515012118\)\. ## Impact Statement This work presents a novel geometric approach to personalized federated learning that enhances privacy\-preserving collaborative machine learning\. The proposed𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}framework inherently supports privacy by design through local training without requiring raw data sharing\. Compared to many PFL methods that additionally share similarity matrices, clustering assignments, or other client\-level statistics,𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}is more privacy\-preserving: only the shared \(space\-like\) parameters are sent to the server for aggregation, while personalized parameters and Lorentz scale parameters remain local to each client\. Standard privacy\-enhancing techniques such as secure aggregation or differential privacy can be applied on top of𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}in the same way as for standard FL methods\. We evaluate our approach on standard academic benchmarks without sensitive personal information\. ## References - Albert and Barabási \(2002\)R\. Albert and A\. BarabásiStatistical mechanics of complex networks\.Reviews of Modern Physics74\(1\),pp\. 47\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2608.21096#S1.p3.1),[§2](https://arxiv.org/html/2608.21096#S2.SS0.SSS0.Px1.p1.1)\. - Anet al\.\(2023\)X\. An, L\. Shen, H\. Hu, and Y\. LuoFederated learning with manifold regularization and normalized update reaggregation\.Advances in Neural Information Processing Systems36\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px3.p1.1),[§E\.3](https://arxiv.org/html/2608.21096#A5.SS3.SSS0.Px4.p1.1),[§2](https://arxiv.org/html/2608.21096#S2.SS0.SSS0.Px2.p1.1)\. - Arivazhaganet al\.\(2019\)M\. G\. Arivazhagan, V\. Aggarwal, A\. K\. Singh, and S\. ChoudharyFederated learning with personalization layers\.arXiv preprint arXiv:1912\.00818\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1),[§E\.3](https://arxiv.org/html/2608.21096#A5.SS3.SSS0.Px1.p1.1),[§7\.1](https://arxiv.org/html/2608.21096#S7.SS1.SSS0.Px1.p1.1)\. - Atighet al\.\(2022\)M\. G\. Atigh, J\. Schoep, E\. Acar, N\. Van Noord, and P\. MettesHyperbolic image segmentation\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 4453–4462\.Cited by:[§1](https://arxiv.org/html/2608.21096#S1.p4.1),[§4](https://arxiv.org/html/2608.21096#S4.p1.1)\. - Baeket al\.\(2023\)J\. Baek, W\. Jeong, J\. Jin, J\. Yoon, and S\. J\. HwangPersonalized subgraph federated learning\.InInternational Conference on Machine Learning,pp\. 1396–1415\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px2.p1.1),[§E\.1](https://arxiv.org/html/2608.21096#A5.SS1.p1.1),[§1](https://arxiv.org/html/2608.21096#S1.p1.1),[§2](https://arxiv.org/html/2608.21096#S2.SS0.SSS0.Px1.p1.1),[§7\.1](https://arxiv.org/html/2608.21096#S7.SS1.SSS0.Px1.p1.1)\. - Chamiet al\.\(2019\)I\. Chami, Z\. Ying, C\. Ré, and J\. LeskovecHyperbolic graph convolutional neural networks\.Advances in Neural Information Processing Systems32\.Cited by:[§2](https://arxiv.org/html/2608.21096#S2.SS0.SSS0.Px2.p1.1),[§7\.3](https://arxiv.org/html/2608.21096#S7.SS3.p1.1)\. - Chen and Chao \(2022\)H\. Chen and W\. ChaoOn bridging generic and personalized federated learning for image classification\.InInternational Conference on Learning Representations,Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2608.21096#S2.SS0.SSS0.Px1.p1.1)\. - Chenet al\.\(2024\)J\. Chen, R\. Lei, and Z\. WeiPolyGCL: graph contrastive learning via learnable spectral polynomial filters\.InThe Twelfth International Conference on Learning Representations,Cited by:[§E\.2](https://arxiv.org/html/2608.21096#A5.SS2.SSS0.Px3.p1.1)\. - Chenet al\.\(2021\)W\. Chen, X\. Han, Y\. Lin, H\. Zhao, Z\. Liu, P\. Li, M\. Sun, and J\. ZhouFully hyperbolic neural networks\.arXiv preprint arXiv:2105\.14686\.Cited by:[§B\.1](https://arxiv.org/html/2608.21096#A2.SS1.p2.1),[§B\.2](https://arxiv.org/html/2608.21096#A2.SS2.p2.1),[§B\.3](https://arxiv.org/html/2608.21096#A2.SS3.p1.1),[§D\.5](https://arxiv.org/html/2608.21096#A4.SS5.p1.1),[§E\.2](https://arxiv.org/html/2608.21096#A5.SS2.SSS0.Px1.p1.1),[§E\.2](https://arxiv.org/html/2608.21096#A5.SS2.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2608.21096#S3.SS0.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2608.21096#S3.SS0.SSS0.Px2.p1.1)\. - Collinset al\.\(2021\)L\. Collins, H\. Hassani, A\. Mokhtari, and S\. ShakkottaiExploiting shared representations for personalized federated learning\.InInternational Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.139,pp\. 2089–2099\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2608.21096#S2.SS0.SSS0.Px1.p1.1)\. - Daviset al\.\(2011\)R\. A\. Davis, K\. Lii, and D\. N\. PolitisRemarks on some nonparametric estimates of a density function\.Selected Works of Murray Rosenblatt,pp\. 95–100\.Cited by:[Figure 1](https://arxiv.org/html/2608.21096#S1.F1),[Figure 1](https://arxiv.org/html/2608.21096#S1.F1.4)\. - Denget al\.\(2020\)Y\. Deng, M\. M\. Kamani, and M\. MahdaviAdaptive personalized federated learning\.arXiv preprint arXiv:2003\.13461\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1)\. - Dinhet al\.\(2020\)C\. T\. Dinh, N\. H\. Tran, and T\. D\. NguyenPersonalized federated learning with moreau envelopes\.InAdvances in Neural Information Processing Systems,Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1)\. - Duet al\.\(2024\)H\. Du, C\. Liu, H\. Liu, X\. Ding, and H\. HuoAn efficient federated learning framework for graph learning in hyperbolic space\.Knowledge\-Based Systems289,pp\. 111438\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px3.p1.1),[§E\.2](https://arxiv.org/html/2608.21096#A5.SS2.SSS0.Px2.p1.1),[§E\.3](https://arxiv.org/html/2608.21096#A5.SS3.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2608.21096#S2.SS0.SSS0.Px2.p1.1),[§7\.1](https://arxiv.org/html/2608.21096#S7.SS1.SSS0.Px1.p1.1),[§7\.4](https://arxiv.org/html/2608.21096#S7.SS4.p1.1)\. - Fallahet al\.\(2020\)A\. Fallah, A\. Mokhtari, and A\. E\. OzdaglarPersonalized federated learning with theoretical guarantees: A model\-agnostic meta\-learning approach\.InAdvances in Neural Information Processing Systems,Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2608.21096#S2.SS0.SSS0.Px1.p1.1)\. - Forman \(2003\)FormanBochner’s method for cell complexes and combinatorial ricci curvature\.Discrete & Computational Geometry29,pp\. 323–374\.Cited by:[§B\.2](https://arxiv.org/html/2608.21096#A2.SS2.p3.1),[§1](https://arxiv.org/html/2608.21096#S1.p2.1)\. - Fuet al\.\(2026\)Y\. Fu, J\. Li, J\. Liu, Q\. Xing, Q\. Wang, and I\. KingHC\-GLAD: dual hyperbolic contrastive learning for unsupervised graph\-level anomaly detection\.Neural Networks202,pp\. 109009\.External Links:[Document](https://dx.doi.org/10.1016/j.neunet.2026.109009)Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px2.p1.1)\. - Gaoet al\.\(2023\)Z\. Gao, Y\. Wu, M\. Harandi, and Y\. JiaCurvature\-Adaptive Meta\-Learning for Fast Adaptation to Manifold Data\.IEEE Transactions on Pattern Analysis and Machine Intelligence45\(2\),pp\. 1545–1562\(en\)\.External Links:ISSN 0162\-8828, 2160\-9292, 1939\-3539,[Document](https://dx.doi.org/10.1109/TPAMI.2022.3164894)Cited by:[§5\.1](https://arxiv.org/html/2608.21096#S5.SS1.p2.1)\. - Gaoet al\.\(2021\)Z\. Gao, Y\. Wu, Y\. Jia, and M\. HarandiCurvature generation in curved spaces for few\-shot learning\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 8691–8700\.Cited by:[§5\.1](https://arxiv.org/html/2608.21096#S5.SS1.p2.1)\. - Gray \(2003\)A\. GrayTubes\.Vol\.221,Springer Science & Business Media\.Cited by:[Lemma D\.2](https://arxiv.org/html/2608.21096#A4.Thmtheorem2)\. - Hanzely and Richtárik \(2020\)F\. Hanzely and P\. RichtárikFederated learning of a mixture of global and local models\.arXiv preprint arXiv:2002\.05516\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1)\. - Heet al\.\(2025\)N\. He, J\. Liu, B\. Zhang, N\. Bui, A\. Maatouk, I\. King, M\. Yang, M\. Weber, and R\. YingPosition: beyond euclidean – foundation models should embrace non\-euclidean geometries\.InProceedings of the Fourth Learning on Graphs Conference,Proceedings of Machine Learning Research, Vol\.269\.Cited by:[§1](https://arxiv.org/html/2608.21096#S1.p3.1)\. - Huet al\.\(2020\)W\. Hu, M\. Fey, M\. Zitnik, Y\. Dong, H\. Ren, B\. Liu, M\. Catasta, and J\. LeskovecOpen graph benchmark: datasets for machine learning on graphs\.InAdvances in Neural Information Processing Systems,Cited by:[§E\.1](https://arxiv.org/html/2608.21096#A5.SS1.p1.1)\. - Huanget al\.\(2021\)Y\. Huang, L\. Chu, Z\. Zhou, L\. Wang, J\. Liu, J\. Pei, and Y\. ZhangPersonalized cross\-silo federated learning on non\-iid data\.InProceedings of the AAAI Conference on Artificial Intelligence,pp\. 7865–7873\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1)\. - Jianget al\.\(2019\)Y\. Jiang, J\. Konečný, K\. Rush, and S\. KannanImproving federated learning personalization via model agnostic meta learning\.arXiv preprint arXiv:1909\.12488\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2608.21096#S2.SS0.SSS0.Px1.p1.1)\. - Kairouzet al\.\(2021\)P\. Kairouz, H\. B\. McMahan, B\. Avent, A\. Bellet, M\. Bennis, A\. N\. Bhagoji, K\. A\. Bonawitz, Z\. Charles, G\. Cormode, R\. Cummings, R\. G\. L\. D’Oliveira, H\. Eichner, S\. E\. Rouayheb, D\. Evans, J\. Gardner, Z\. Garrett, A\. Gascón, B\. Ghazi, P\. B\. Gibbons, M\. Gruteser, Z\. Harchaoui, C\. He, L\. He, Z\. Huo, B\. Hutchinson, J\. Hsu, M\. Jaggi, T\. Javidi, G\. Joshi, M\. Khodak, J\. Konečný, A\. Korolova, F\. Koushanfar, S\. Koyejo, T\. Lepoint, Y\. Liu, P\. Mittal, M\. Mohri, R\. Nock, A\. Özgür, R\. Pagh, H\. Qi, D\. Ramage, R\. Raskar, M\. Raykova, D\. Song, W\. Song, S\. U\. Stich, Z\. Sun, A\. T\. Suresh, F\. Tramèr, P\. Vepakomma, J\. Wang, L\. Xiong, Z\. Xu, Q\. Yang, F\. X\. Yu, H\. Yu, and S\. ZhaoAdvances and open problems in federated learning\.Foundations and Trends in Machine Learning14\(1\-2\),pp\. 1–210\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2608.21096#S2.SS0.SSS0.Px1.p1.1)\. - Karypis and Kumar \(1995\)G\. Karypis and V\. KumarMETIS – unstructured graph partitioning and sparse matrix ordering system, version 2\.0\.Cited by:[§E\.1](https://arxiv.org/html/2608.21096#A5.SS1.p1.1)\. - Khrulkovet al\.\(2020\)V\. Khrulkov, L\. Mirvakhabova, E\. Ustinova, I\. Oseledets, and V\. LempitskyHyperbolic image embeddings\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 6418–6428\.Cited by:[§1](https://arxiv.org/html/2608.21096#S1.p3.1)\. - Kipf and Welling \(2017\)T\. N\. Kipf and M\. WellingSemi\-supervised classification with graph convolutional networks\.InInternational Conference on Learning Representations,Cited by:[§E\.2](https://arxiv.org/html/2608.21096#A5.SS2.SSS0.Px2.p1.1)\. - Krioukovet al\.\(2010\)D\. Krioukov, F\. Papadopoulos, M\. Kitsak, A\. Vahdat, and M\. BogunáHyperbolic geometry of complex networks\.Physical Review E82\(3\),pp\. 036106\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px2.p1.1),[Assumption D\.6](https://arxiv.org/html/2608.21096#A4.Thmtheorem6),[§2](https://arxiv.org/html/2608.21096#S2.SS0.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2608.21096#S4.SS1.p4.1)\. - Lee \(2018\)J\. M\. LeeIntroduction to riemannian manifolds\.Vol\.2,Springer\.Cited by:[§B\.2](https://arxiv.org/html/2608.21096#A2.SS2.p2.1)\. - Liet al\.\(2021a\)T\. Li, S\. Hu, A\. Beirami, and V\. SmithDitto: fair and robust federated learning through personalization\.InInternational Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.139,pp\. 6357–6368\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1)\. - Liet al\.\(2020a\)T\. Li, A\. K\. Sahu, M\. Zaheer, M\. Sanjabi, A\. Talwalkar, and V\. SmithFederated optimization in heterogeneous networks\.InProceedings of Machine Learning and Systems,Cited by:[§D\.4](https://arxiv.org/html/2608.21096#A4.SS4.p2.1),[§D\.4](https://arxiv.org/html/2608.21096#A4.SS4.p22.1),[§D\.4](https://arxiv.org/html/2608.21096#A4.SS4.p24.1),[§E\.3](https://arxiv.org/html/2608.21096#A5.SS3.SSS0.Px1.p1.1),[§7\.1](https://arxiv.org/html/2608.21096#S7.SS1.SSS0.Px1.p1.1)\. - Liet al\.\(2020b\)X\. Li, K\. Huang, W\. Yang, S\. Wang, and Z\. ZhangOn the convergence of fedavg on non\-iid data\.InInternational Conference on Learning Representations,Cited by:[§D\.4](https://arxiv.org/html/2608.21096#A4.SS4.p1.1),[§D\.4](https://arxiv.org/html/2608.21096#A4.SS4.p11.1)\. - Liet al\.\(2021b\)X\. Li, D\. Zhan, Y\. Shao, B\. Li, and S\. SongFedPHP: federated personalization with inherited private models\.InMachine Learning and Knowledge Discovery in Databases,Lecture Notes in Computer Science, Vol\.12975,pp\. 587–602\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1)\. - Liet al\.\(2024\)X\. Li, Z\. Wu, W\. Zhang, H\. Sun, R\. Li, and G\. WangAdaFGL: a new paradigm for federated node classification with topology heterogeneity\.In2024 IEEE 40th International Conference on Data Engineering,pp\. 2517–2530\.External Links:[Document](https://dx.doi.org/10.1109/ICDE60146.2024.00198)Cited by:[§7\.1](https://arxiv.org/html/2608.21096#S7.SS1.SSS0.Px1.p1.1)\. - Liet al\.\(2023\)X\. Li, Z\. Wu, W\. Zhang, Y\. Zhu, R\. Li, and G\. WangFedGTA: topology\-aware averaging for federated graph learning\.Proceedings of the VLDB Endowment17\(1\),pp\. 41–50\.External Links:[Document](https://dx.doi.org/10.14778/3617838.3617842)Cited by:[§7\.1](https://arxiv.org/html/2608.21096#S7.SS1.SSS0.Px1.p1.1)\. - Liaoet al\.\(2023\)X\. Liao, W\. Liu, C\. Chen, P\. Zhou, H\. Zhu, Y\. Tan, J\. Wang, and Y\. QiHyperFed: hyperbolic prototypes exploration with consistent aggregation for non\-iid data in federated learning\.InProceedings of the Thirty\-Second International Joint Conference on Artificial Intelligence,pp\. 3957–3965\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px3.p1.1),[§E\.3](https://arxiv.org/html/2608.21096#A5.SS3.SSS0.Px4.p1.1),[§2](https://arxiv.org/html/2608.21096#S2.SS0.SSS0.Px2.p1.1)\. - Liuet al\.\(2022\)J\. Liu, P\. Fournier\-Viger, M\. Zhou, G\. He, and M\. NouiouaCSPM: discovering compressing stars in attributed graphs\.Information Sciences611,pp\. 126–158\.External Links:[Document](https://dx.doi.org/10.1016/j.ins.2022.08.008)Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px2.p1.1)\. - Liuet al\.\(2025\)J\. Liu, Z\. Qiu, Z\. Li, Q\. Dai, W\. Yu, J\. Zhu, M\. Hu, M\. Yang, T\. Chua, and I\. KingA survey of personalized large language models: progress and future directions\.arXiv preprint arXiv:2502\.11528\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1)\. - Liuet al\.\(2021\)J\. Liu, M\. Yang, M\. Zhou, S\. Feng, and P\. Fournier\-VigerEnhancing hyperbolic graph embeddings via contrastive learning\.InWorkshop on Self\-Supervised Learning: Theory and Practice at Advances in Neural Information Processing Systems 2021,Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px2.p1.1)\. - Liuet al\.\(2026\)J\. Liu, W\. Yu, Q\. Dai, Z\. Li, J\. Zhu, M\. Yang, T\. Chua, and I\. KingPerFit: exploring personalization shifts in representation space of LLMs\.InThe Fourteenth International Conference on Learning Representations,Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1)\. - Luo and Wu \(2022\)J\. Luo and S\. WuAdapt to adaptation: learning personalization for cross\-silo federated learning\.InProceedings of the International Joint Conference on Artificial Intelligence,pp\. 2166–2173\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1)\. - Mansouret al\.\(2020\)Y\. Mansour, M\. Mohri, J\. Ro, and A\. T\. SureshThree approaches for personalization with applications to federated learning\.arXiv preprint arXiv:2002\.10619\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1)\. - McMahanet al\.\(2017\)B\. McMahan, E\. Moore, D\. Ramage, S\. Hampson, and B\. A\. y ArcasCommunication\-efficient learning of deep networks from decentralized data\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 1273–1282\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1),[§B\.4](https://arxiv.org/html/2608.21096#A2.SS4.p1.1),[§E\.3](https://arxiv.org/html/2608.21096#A5.SS3.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.21096#S1.p1.1),[§7\.1](https://arxiv.org/html/2608.21096#S7.SS1.SSS0.Px1.p1.1)\. - Moretti \(2002\)V\. MorettiThe interplay of the polar decomposition theorem and the lorentz group\.arXiv preprint math\-ph/0211047\.Cited by:[§B\.3](https://arxiv.org/html/2608.21096#A2.SS3.p1.1)\. - Morriset al\.\(2020\)C\. Morris, N\. M\. Kriege, F\. Bause, K\. Kersting, P\. Mutzel, and M\. NeumannTUDataset: A collection of benchmark datasets for learning with graphs\.arXiv preprint arXiv:2007\.08663\.Cited by:[§E\.1](https://arxiv.org/html/2608.21096#A5.SS1.p2.1)\. - Nickel and Kiela \(2018\)M\. Nickel and D\. KielaLearning continuous hierarchies in the lorentz model of hyperbolic geometry\.InInternational Conference on Machine Learning,pp\. 3779–3788\.Cited by:[§B\.1](https://arxiv.org/html/2608.21096#A2.SS1.p1.1),[§B\.1](https://arxiv.org/html/2608.21096#A2.SS1.p2.1),[§1](https://arxiv.org/html/2608.21096#S1.p3.1),[§1](https://arxiv.org/html/2608.21096#S1.p4.1),[§2](https://arxiv.org/html/2608.21096#S2.SS0.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2608.21096#S3.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2608.21096#S4.p1.1)\. - Ollivier \(2009\)Y\. OllivierRicci curvature of markov chains on metric spaces\.Journal of Functional Analysis256\(3\),pp\. 810–864\.Cited by:[§B\.2](https://arxiv.org/html/2608.21096#A2.SS2.p3.1)\. - Penget al\.\(2021\)W\. Peng, T\. Varanka, A\. Mostafa, H\. Shi, and G\. ZhaoHyperbolic deep neural networks: a survey\.IEEE Transactions on Pattern Analysis and Machine Intelligence44\(12\),pp\. 10023–10044\.Cited by:[§1](https://arxiv.org/html/2608.21096#S1.p4.1),[§4](https://arxiv.org/html/2608.21096#S4.p1.1)\. - Reddiet al\.\(2020\)S\. Reddi, Z\. Charles, M\. Zaheer, Z\. Garrett, K\. Rush, J\. Konečnỳ, S\. Kumar, and H\. B\. McMahanAdaptive federated optimization\.arXiv preprint arXiv:2003\.00295\.Cited by:[§D\.4](https://arxiv.org/html/2608.21096#A4.SS4.p11.1)\. - Senet al\.\(2008\)P\. Sen, G\. Namata, M\. Bilgic, L\. Getoor, B\. Gallagher, and T\. Eliassi\-RadCollective classification in network data\.AI Magazine29\(3\),pp\. 93–106\.Cited by:[§E\.1](https://arxiv.org/html/2608.21096#A5.SS1.p1.1)\. - Shchuret al\.\(2018\)O\. Shchur, M\. Mumme, A\. Bojchevski, and S\. GünnemannPitfalls of graph neural network evaluation\.arXiv preprint arXiv:1811\.05868\.Cited by:[§E\.1](https://arxiv.org/html/2608.21096#A5.SS1.p1.1)\. - Southernet al\.\(2023\)J\. Southern, J\. Wayland, M\. Bronstein, and B\. RieckCurvature filtrations for graph generative model evaluation\.Advances in Neural Information Processing Systems36\.Cited by:[§B\.2](https://arxiv.org/html/2608.21096#A2.SS2.p6.1)\. - Sunet al\.\(2021\)B\. Sun, H\. Huo, Y\. Yang, and B\. BaiPartialFed: cross\-domain personalized federated learning via partial initialization\.InAdvances in Neural Information Processing Systems,pp\. 23309–23320\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1)\. - Sunet al\.\(2024\)L\. Sun, J\. Ye, J\. Zhang, Y\. Yang, M\. Liu, F\. Wang, and P\. S\. YuContrastive sequential interaction network learning on co\-evolving riemannian spaces\.International Journal of Machine Learning and Cybernetics15\(4\),pp\. 1397–1413\.Cited by:[§B\.2](https://arxiv.org/html/2608.21096#A2.SS2.p1.1),[§1](https://arxiv.org/html/2608.21096#S1.p2.1)\. - Tanet al\.\(2022\)A\. Z\. Tan, H\. Yu, L\. Cui, and Q\. YangTowards personalized federated learning\.IEEE Transactions on Neural Networks and Learning Systems34\(12\),pp\. 9587–9603\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.21096#S1.p1.1)\. - Tanet al\.\(2023\)Y\. Tan, Y\. Liu, G\. Long, J\. Jiang, Q\. Lu, and C\. ZhangFederated learning on non\-iid graphs via structural knowledge sharing\.InProceedings of the AAAI Conference on Artificial Intelligence,pp\. 9953–9961\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2608.21096#S1.p1.1),[§1](https://arxiv.org/html/2608.21096#S1.p3.1),[§2](https://arxiv.org/html/2608.21096#S2.SS0.SSS0.Px1.p1.1)\. - Wanget al\.\(2019\)K\. Wang, R\. Mathews, C\. Kiddon, H\. Eichner, F\. Beaufays, and D\. RamageFederated evaluation of on\-device personalization\.arXiv preprint arXiv:1910\.10252\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1)\. - Wuet al\.\(2021\)C\. Wu, F\. Wu, Y\. Cao, Y\. Huang, and X\. XieFedGNN: federated graph neural network for privacy\-preserving recommendation\.arXiv preprint arXiv:2102\.04925\.Cited by:[§7\.1](https://arxiv.org/html/2608.21096#S7.SS1.SSS0.Px1.p1.1)\. - Xieet al\.\(2021\)H\. Xie, J\. Ma, L\. Xiong, and C\. YangFederated graph classification over non\-iid graphs\.Advances in Neural Information Processing Systems34,pp\. 18839–18852\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px2.p1.1),[§E\.1](https://arxiv.org/html/2608.21096#A5.SS1.p2.1),[§E\.3](https://arxiv.org/html/2608.21096#A5.SS3.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.21096#S1.p1.1),[§2](https://arxiv.org/html/2608.21096#S2.SS0.SSS0.Px1.p1.1),[§7\.1](https://arxiv.org/html/2608.21096#S7.SS1.SSS0.Px1.p1.1)\. - Xuet al\.\(2018\)K\. Xu, W\. Hu, J\. Leskovec, and S\. JegelkaHow powerful are graph neural networks?\.arXiv preprint arXiv:1810\.00826\.Cited by:[§E\.2](https://arxiv.org/html/2608.21096#A5.SS2.SSS0.Px2.p1.1)\. - Yanget al\.\(2024\)M\. Yang, H\. Verma, D\. C\. Zhang, J\. Liu, I\. King, and R\. YingHypformer: exploring efficient transformer fully in hyperbolic space\.InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,pp\. 3770–3781\.External Links:[Document](https://dx.doi.org/10.1145/3637528.3672039)Cited by:[§1](https://arxiv.org/html/2608.21096#S1.p4.1)\. - Yanget al\.\(2022\)M\. Yang, M\. Zhou, Z\. Li, J\. Liu, L\. Pan, H\. Xiong, and I\. KingHyperbolic graph neural networks: a review of methods and applications\.arXiv preprint arXiv:2202\.13852\.Cited by:[§4\.1](https://arxiv.org/html/2608.21096#S4.SS1.p4.1)\. - Yeet al\.\(2019\)Z\. Ye, K\. S\. Liu, T\. Ma, J\. Gao, and C\. ChenCurvature graph network\.InInternational Conference on Learning Representations,Cited by:[§B\.2](https://arxiv.org/html/2608.21096#A2.SS2.p1.1),[§5\.1](https://arxiv.org/html/2608.21096#S5.SS1.p2.1)\. - Zhanget al\.\(2023\)J\. Zhang, Y\. Hua, H\. Wang, T\. Song, Z\. Xue, R\. Ma, and H\. GuanFedALA: adaptive local aggregation for personalized federated learning\.InProceedings of the AAAI Conference on Artificial Intelligence,pp\. 11237–11244\.Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1)\. - Zhanget al\.\(2021a\)K\. Zhang, C\. Yang, X\. Li, L\. Sun, and S\. YiuSubgraph federated learning with missing neighbor generation\.InAdvances in Neural Information Processing Systems,pp\. 6671–6682\.Cited by:[§7\.1](https://arxiv.org/html/2608.21096#S7.SS1.SSS0.Px1.p1.1)\. - Zhanget al\.\(2021b\)M\. Zhang, K\. Sapra, S\. Fidler, S\. Yeung, and J\. M\. ÁlvarezPersonalized federated learning with first order model optimization\.InInternational Conference on Learning Representations,Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px1.p1.1)\. - Zhanget al\.\(2021c\)Y\. Zhang, X\. Wang, C\. Shi, N\. Liu, and G\. SongLorentzian graph convolutional networks\.InProceedings of the Web Conference 2021,pp\. 1249–1261\.Cited by:[§B\.1](https://arxiv.org/html/2608.21096#A2.SS1.p2.1),[§B\.3](https://arxiv.org/html/2608.21096#A2.SS3.p1.1)\. - Zhanget al\.\(2025\)Y\. Zhang, H\. Zhu, M\. Yang, J\. Liu, R\. Ying, I\. King, and P\. KoniuszUnderstanding and mitigating hyperbolic dimensional collapse in graph contrastive learning\.InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\.1,pp\. 1984–1995\.External Links:[Document](https://dx.doi.org/10.1145/3690624.3709249)Cited by:[Appendix A](https://arxiv.org/html/2608.21096#A1.SS0.SSS0.Px2.p1.1)\. Appendix ###### Contents 1. [1Introduction](https://arxiv.org/html/2608.21096#S1) 2. [2Related Work](https://arxiv.org/html/2608.21096#S2) 3. [3Preliminaries](https://arxiv.org/html/2608.21096#S3) 4. [4Motivation and Insights](https://arxiv.org/html/2608.21096#S4)1. [4\.1Motivation: Why Lorentz Space for PFL?](https://arxiv.org/html/2608.21096#S4.SS1) 2. [4\.2Insights: Introduce a Higher Dimension \(time dimension\) to"Flatland"](https://arxiv.org/html/2608.21096#S4.SS2) 5. [5The𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}Framework](https://arxiv.org/html/2608.21096#S5)1. [5\.1Learnable Curvature Initialization](https://arxiv.org/html/2608.21096#S5.SS1) 2. [5\.2Parameter Decoupling Strategy](https://arxiv.org/html/2608.21096#S5.SS2) 6. [6Analysis](https://arxiv.org/html/2608.21096#S6) 7. [7Experiments](https://arxiv.org/html/2608.21096#S7)1. [7\.1Experimental Setup](https://arxiv.org/html/2608.21096#S7.SS1) 2. [7\.2Main Experimental Results \(RQ1\)](https://arxiv.org/html/2608.21096#S7.SS2) 3. [7\.3Varying Embedding Dimensions \(RQ2\)](https://arxiv.org/html/2608.21096#S7.SS3) 4. [7\.4Partial Client Participation \(RQ3\)](https://arxiv.org/html/2608.21096#S7.SS4) 5. [7\.5Ablation Study \(RQ4\)](https://arxiv.org/html/2608.21096#S7.SS5) 8. [8Conclusions](https://arxiv.org/html/2608.21096#S8) 9. [References](https://arxiv.org/html/2608.21096#bib) 10. [ARelated Work](https://arxiv.org/html/2608.21096#A1) 11. [BPreliminaries](https://arxiv.org/html/2608.21096#A2)1. [B\.1Lorentz \(Hyperboloid\) Model: Formal Definitions](https://arxiv.org/html/2608.21096#A2.SS1) 2. [B\.2Forman\-Ricci Curvature](https://arxiv.org/html/2608.21096#A2.SS2) 3. [B\.3Lorentz Transformations \(Hyperbolic Isometries\)](https://arxiv.org/html/2608.21096#A2.SS3) 4. [B\.4The FedAvg Algorithm](https://arxiv.org/html/2608.21096#A2.SS4) 12. [CMethodology Supplementary](https://arxiv.org/html/2608.21096#A3)1. [C\.1Statistics of Forman\-Ricci Curvature in Other Datasets](https://arxiv.org/html/2608.21096#A3.SS1) 2. [C\.2The FlatLand Algorithm](https://arxiv.org/html/2608.21096#A3.SS2)1. [C\.2\.1Overall Process](https://arxiv.org/html/2608.21096#A3.SS2.SSS1) 2. [C\.2\.2Curvature Initialization and Learning Details](https://arxiv.org/html/2608.21096#A3.SS2.SSS2) 3. [C\.2\.3Local Training Procedure](https://arxiv.org/html/2608.21096#A3.SS2.SSS3) 3. [C\.3Derivation of Parameters Disentanglement](https://arxiv.org/html/2608.21096#A3.SS3) 4. [C\.4Time and Space Complexity Compared with FedAvg](https://arxiv.org/html/2608.21096#A3.SS4) 13. [DTheoretical Analysis Supplementary](https://arxiv.org/html/2608.21096#A4)1. [D\.1Proof for Theorem](https://arxiv.org/html/2608.21096#A4.SS1) 2. [D\.2Proof for Theorem](https://arxiv.org/html/2608.21096#A4.SS2) 3. [D\.3Proof for Proposition](https://arxiv.org/html/2608.21096#A4.SS3) 4. [D\.4Convergence Analysis](https://arxiv.org/html/2608.21096#A4.SS4) 5. [D\.5Perspectives on Lorentz Transformations](https://arxiv.org/html/2608.21096#A4.SS5) 14. [EExperimental Supplementary](https://arxiv.org/html/2608.21096#A5)1. [E\.1Datasets](https://arxiv.org/html/2608.21096#A5.SS1) 2. [E\.2Implementation Details](https://arxiv.org/html/2608.21096#A5.SS2) 3. [E\.3Baseline Selection and Comparison](https://arxiv.org/html/2608.21096#A5.SS3) 4. [E\.4Unified Summary of Ablations](https://arxiv.org/html/2608.21096#A5.SS4) 5. [E\.5Impact of Curvature Initialization](https://arxiv.org/html/2608.21096#A5.SS5) 6. [E\.6Convergence Curves](https://arxiv.org/html/2608.21096#A5.SS6) 7. [E\.7Curvature Sensitivity and Interpretation](https://arxiv.org/html/2608.21096#A5.SS7) ## Appendix ARelated Work ##### Personalized Federated Learning Under statistical heterogeneity\([26](https://arxiv.org/html/2608.21096#bib.bib32)\), conventional FL frameworks such as FedAvg\([45](https://arxiv.org/html/2608.21096#bib.bib28)\)often fail to obtain a single global model that generalizes well to every client \(the basic framework is shown in[AppendixB\.4](https://arxiv.org/html/2608.21096#A2.SS4)\)\. Motivated by this, researchers have proposed personalized FL \(PFL\) to train customized local models\([57](https://arxiv.org/html/2608.21096#bib.bib63);[40](https://arxiv.org/html/2608.21096#bib.bib64);[42](https://arxiv.org/html/2608.21096#bib.bib65)\)\. Generally speaking, existing PFL techniques can be categorized into the following three groups: \(1\) techniques that personalize client models via local fine\-tuning\([15](https://arxiv.org/html/2608.21096#bib.bib33);[25](https://arxiv.org/html/2608.21096#bib.bib34);[59](https://arxiv.org/html/2608.21096#bib.bib35)\), \(2\) techniques that personalize client models via customized model aggregation\([24](https://arxiv.org/html/2608.21096#bib.bib36);[35](https://arxiv.org/html/2608.21096#bib.bib40);[43](https://arxiv.org/html/2608.21096#bib.bib41);[55](https://arxiv.org/html/2608.21096#bib.bib39);[66](https://arxiv.org/html/2608.21096#bib.bib38);[68](https://arxiv.org/html/2608.21096#bib.bib37)\), and \(3\) techniques that personalize client models by creating localized models or layers\([3](https://arxiv.org/html/2608.21096#bib.bib42);[7](https://arxiv.org/html/2608.21096#bib.bib49);[10](https://arxiv.org/html/2608.21096#bib.bib47);[12](https://arxiv.org/html/2608.21096#bib.bib46);[13](https://arxiv.org/html/2608.21096#bib.bib43);[21](https://arxiv.org/html/2608.21096#bib.bib48);[32](https://arxiv.org/html/2608.21096#bib.bib44);[44](https://arxiv.org/html/2608.21096#bib.bib45)\)\. However, these PFL methods typically encode data samples in Euclidean spaces, which makes it difficult to capture scale\-free properties and implicit hierarchical structure in client data\. ##### Personalized Federated Graph Learning When applied to graph data, personalized federated graph learning \(PFGL\) can intuitively exhibit the problem mentioned above\. For example,\([61](https://arxiv.org/html/2608.21096#bib.bib59)\)clusters clients based on gradients to aggregate models with similar data distributions\. Another method\([58](https://arxiv.org/html/2608.21096#bib.bib50)\)introduces additional personalized models to capture client\-specific knowledge of graph structure\.\([5](https://arxiv.org/html/2608.21096#bib.bib51)\)calculates client\-client similarities to apply personalized model aggregation with local weight masking\. These methods learn node representations in Euclidean spaces, which cannot naturally model the power\-law degree distributions that widely exist in real\-world graph data\([1](https://arxiv.org/html/2608.21096#bib.bib16);[30](https://arxiv.org/html/2608.21096#bib.bib17)\)\. More broadly, graph\-structure mining and hyperbolic graph representation learning show that structural patterns and non\-Euclidean geometry are central to graph modeling\([39](https://arxiv.org/html/2608.21096#bib.bib21);[41](https://arxiv.org/html/2608.21096#bib.bib22);[70](https://arxiv.org/html/2608.21096#bib.bib24);[17](https://arxiv.org/html/2608.21096#bib.bib25)\)\. Additionally, the client clustering procedure and additional model components introduce computational overhead that may not be feasible in real\-world scenarios with strict privacy constraints or limited resources\. ##### Hyperbolic Federated Learning Only a few studies have considered incorporating hyperbolic spaces into federated settings\.\([2](https://arxiv.org/html/2608.21096#bib.bib54)\)leverages hyperbolic distances to distill knowledge from the global model to the local model, to mitigate model inconsistency caused by data heterogeneity\.\([38](https://arxiv.org/html/2608.21096#bib.bib53)\)applies hyperbolic prototype learning to capture the hierarchical structure among data samples\. As the work most similar to our𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}, FedHGCN\([14](https://arxiv.org/html/2608.21096#bib.bib52)\)is a simple combination of FedAvg and hyperbolic graph neural networks along with a node selection process\. Although these methods can benefit from the hyperbolic space to capture the hierarchical structure in the data, they do not have the personalization capability to adaptively model client data spaces with different curvatures\. This may lead to suboptimal results when there is severe data heterogeneity\. Therefore, our goal is to design a novel FL framework that can encode client data in hyperbolic spaces with adaptive curvatures using personalization techniques\. ## Appendix BPreliminaries ### B\.1Lorentz \(Hyperboloid\) Model: Formal Definitions Hyperbolic space is non\-Euclidean geometry with a constant negative curvature\. The curvature of hyperbolic space is a measure of how the geometry of the space deviates from the flatness of Euclidean space\. The Lorentz manifold, also known as the hyperboloid model, is one of the most commonly used mathematical representations of hyperbolic space\. Its greater stability for numerical optimization makes it a popular choice for hyperbolic geometry methods\([48](https://arxiv.org/html/2608.21096#bib.bib18)\)\. Clarification on terminology\.Throughout this paper, “Lorentz space” or “Lorentz model” refers specifically to the hyperboloid model ofRiemannianhyperbolic space, not to Lorentzian spacetime in the sense of pseudo\-Riemannian geometry or general relativity\. We use the Lorentzian inner product purely as a mathematical tool for representing hyperbolic geometry, consistent with prior hyperbolic learning literature\([48](https://arxiv.org/html/2608.21096#bib.bib18);[9](https://arxiv.org/html/2608.21096#bib.bib27);[69](https://arxiv.org/html/2608.21096#bib.bib26)\)\. All constructions take place in flat Minkowski ambient spaceℝd\+1\\mathbb\{R\}^\{d\+1\}, where the Lorentz groupO\(1,d\)O\(1,d\)acts as global isometries\. We do not assume curved Lorentzian manifolds \(e\.g\., anti\-de Sitter space\), and notions such as light cones, causality, or speed limits from physics are not invoked or interpreted in our framework\. ###### Definition B\.1\(Lorentz Manifold\)\. Add\-dimensional Lorentz manifoldℒKd\\mathcal\{L\}\_\{K\}^\{d\}with constant negative curvature−1/K\-1/K\(K\>0K\>0\) is defined as the hyperboloidℍKd=\{𝐱∈ℝd\+1:⟨𝐱,𝐱⟩ℒ=−K,x0\>0\}\\mathbb\{H\}\_\{K\}^\{d\}=\\left\\\{\\mathbf\{x\}\\in\\mathbb\{R\}^\{d\+1\}:\\langle\\mathbf\{x\},\\mathbf\{x\}\\rangle\_\{\\mathcal\{L\}\}=\-K,x\_\{0\}\>0\\right\\\}endowed with the Riemannian metric induced by the Lorentzian inner product\. ###### Definition B\.2\(Lorentzian Inner Product\)\. The inner product⟨𝐱,𝐲⟩ℒ\\langle\\mathbf\{x\},\\mathbf\{y\}\\rangle\_\{\\mathcal\{L\}\}for𝐱,𝐲∈ℝd\+1\\mathbf\{x\},\\mathbf\{y\}\\in\\mathbb\{R\}^\{d\+1\}is defined as⟨𝐱,𝐲⟩ℒ=−x0y0\+∑i=1dxiyi\\langle\\mathbf\{x\},\\mathbf\{y\}\\rangle\_\{\\mathcal\{L\}\}=\-x\_\{0\}y\_\{0\}\+\\sum\_\{i=1\}^\{d\}x\_\{i\}y\_\{i\}\. Based on the constraint⟨𝐱,𝐱⟩ℒ=−K\\langle\\mathbf\{x\},\\mathbf\{x\}\\rangle\_\{\\mathcal\{L\}\}=\-K, it holds for any point𝐱=\(x0,𝐱′\)∈ℝd\+1\\mathbf\{x\}=\\left\(x\_\{0\},\\mathbf\{x\}^\{\\prime\}\\right\)\\in\\mathbb\{R\}^\{d\+1\}that𝐱∈ℒKd⇔x0=‖𝐱′‖22\+K\\mathbf\{x\}\\in\\mathcal\{L\}^\{d\}\_\{K\}\\Leftrightarrow x\_\{0\}=\\sqrt\{\\left\\\|\\mathbf\{x\}^\{\\prime\}\\right\\\|\_\{2\}^\{2\}\+K\}\. Here,KKis the positive Lorentz scale parameter, and the corresponding sectional curvature is−1/K\-1/K\. Next, the corresponding Lorentzian distance function for two points𝐱,𝐲∈ℒKd\\mathbf\{x\},\\mathbf\{y\}\\in\\mathcal\{L\}^\{d\}\_\{K\}is provided as dℒK\(𝐱,𝐲\)=Karcosh\(−⟨𝐱,𝐲⟩ℒ/K\)\.d\_\{\\mathcal\{L\}\}^\{K\}\(\\mathbf\{x\},\\mathbf\{y\}\)=\\sqrt\{K\}\\mbox\{arcosh\}\(\-\\langle\\mathbf\{x\},\\mathbf\{y\}\\rangle\_\{\\mathcal\{L\}\}/K\)\.\(5\) ###### Definition B\.3\(Tangent Space\)\. For a point𝐱∈ℒKd\\mathbf\{x\}\\in\\mathcal\{L\}^\{d\}\_\{K\}, the tangent space𝒯𝐱ℒKd\\mathcal\{T\}\_\{\\mathbf\{x\}\}\\mathcal\{L\}^\{d\}\_\{K\}consists of all vectors orthogonal to𝐱\\mathbf\{x\}, where orthogonality is defined with respect to the Lorentzian inner product \([DefinitionB\.2](https://arxiv.org/html/2608.21096#A2.Thmtheorem2)\)\. Hence,𝒯𝐱ℒKd=\{𝐯:⟨𝐱,𝐯⟩ℒ=0\}\.\\mathcal\{T\}\_\{\\mathbf\{x\}\}\\mathcal\{L\}^\{d\}\_\{K\}=\\left\\\{\\mathbf\{v\}:\\langle\\mathbf\{x\},\\mathbf\{v\}\\rangle\_\{\\mathcal\{L\}\}=0\\right\\\}\. ###### Definition B\.4\(Exponential and Logarithmic Maps\)\. Let𝐯∈𝒯xℒKd\\mathbf\{v\}\\in\\mathcal\{T\}\_\{x\}\\mathcal\{L\}^\{d\}\_\{K\}\. The exponential mapexp𝐱K:𝒯𝐱ℒKd→ℒKd\\exp^\{K\}\_\{\\mathbf\{x\}\}:\\mathcal\{T\}\_\{\\mathbf\{x\}\}\\mathcal\{L\}^\{d\}\_\{K\}\\rightarrow\\mathcal\{L\}^\{d\}\_\{K\}and logarithmic maplog𝐱K:ℒKd→𝒯𝐱ℒKd\\log^\{K\}\_\{\\mathbf\{x\}\}:\\mathcal\{L\}^\{d\}\_\{K\}\\rightarrow\\mathcal\{T\}\_\{\\mathbf\{x\}\}\\mathcal\{L\}^\{d\}\_\{K\}are defined as exp𝐱K\(𝐯\)=cosh\(‖𝐯‖ℒK\)𝐱\+Ksinh\(‖𝐯‖ℒK\)𝐯‖𝐯‖ℒ,\\exp^\{K\}\_\{\\mathbf\{x\}\}\(\\mathbf\{v\}\)=\\cosh\\left\(\\frac\{\\\|\\mathbf\{v\}\\\|\_\{\\mathcal\{L\}\}\}\{\\sqrt\{K\}\}\\right\)\\mathbf\{x\}\+\\sqrt\{K\}\\sinh\\left\(\\frac\{\\\|\\mathbf\{v\}\\\|\_\{\\mathcal\{L\}\}\}\{\\sqrt\{K\}\}\\right\)\\frac\{\\mathbf\{v\}\}\{\\\|\\mathbf\{v\}\\\|\_\{\\mathcal\{L\}\}\},log𝐱K\(𝐲\)=dℒK\(𝐱,𝐲\)𝐲\+1K⟨𝐱,𝐲⟩ℒ𝐱‖𝐲\+1K⟨𝐱,𝐲⟩ℒ𝐱‖ℒ,\\log\_\{\\mathbf\{x\}\}^\{K\}\(\\mathbf\{y\}\)=d\_\{\\mathcal\{L\}\}^\{K\}\(\\mathbf\{x\},\\mathbf\{y\}\)\\frac\{\\mathbf\{y\}\+\\frac\{1\}\{K\}\\langle\\mathbf\{x\},\\mathbf\{y\}\\rangle\_\{\\mathcal\{L\}\}\\mathbf\{x\}\}\{\\left\\\|\\mathbf\{y\}\+\\frac\{1\}\{K\}\\langle\\mathbf\{x\},\\mathbf\{y\}\\rangle\_\{\\mathcal\{L\}\}\\mathbf\{x\}\\right\\\|\_\{\\mathcal\{L\}\}\},where‖𝐯‖ℒ=⟨𝐯,𝐯⟩ℒ\\\|\\mathbf\{v\}\\\|\_\{\\mathcal\{L\}\}=\\sqrt\{\\langle\\mathbf\{v\},\\mathbf\{v\}\\rangle\_\{\\mathcal\{L\}\}\}denotes the norm of𝐯\\mathbf\{v\}in𝒯𝐱ℒKd\\mathcal\{T\}\_\{\\mathbf\{x\}\}\\mathcal\{L\}^\{d\}\_\{K\}\. Particularly, for the sake of calculation, the origin of Lorentz manifold𝐨=\(K,0,0,…,0\)∈ℒKd\\mathbf\{o\}=\(\\sqrt\{K\},0,0,\\ldots,0\)\\in\\mathcal\{L\}^\{d\}\_\{K\}is chosen as the reference point for the exponential and logarithmic maps, which can be simplified as exp𝐨K\(𝐯\)\\displaystyle\\exp\_\{\\mathbf\{o\}\}^\{K\}\\left\(\\mathbf\{v\}\\right\)=exp𝐨K\(\[0,𝐯E\]\)\\displaystyle=\\exp\_\{\\mathbf\{o\}\}^\{K\}\\left\(\\left\[0,\\mathbf\{v\}^\{E\}\\right\]\\right\)\(6\)=\(Kcosh\(‖𝐯E‖2K\)⏟time\-like dimension,Ksinh\(‖𝐯E‖2K\)𝐯E‖𝐯E‖2⏟space\-like dimension\),\\displaystyle=\\left\(\\underbrace\{\\sqrt\{K\}\\cosh\\left\(\\frac\{\\\|\\mathbf\{v\}^\{E\}\\\|\_\{2\}\}\{\\sqrt\{K\}\}\\right\)\}\_\{\\text\{time\-like dimension\}\},\\underbrace\{\\sqrt\{K\}\\sinh\\left\(\\frac\{\\\|\\mathbf\{v\}^\{E\}\\\|\_\{2\}\}\{\\sqrt\{K\}\}\\right\)\\frac\{\\mathbf\{v\}^\{E\}\}\{\\\|\\mathbf\{v\}^\{E\}\\\|\_\{2\}\}\}\_\{\\text\{space\-like dimension\}\}\\right\),where the\(,\)\(,\)denotes concatenation and the⋅E\\cdot^\{E\}denotes the embedding in Euclidean space \. ### B\.2Forman\-Ricci Curvature Curvature is a metric used in Riemannian geometry that expresses how far a curved line deviates from a straight line, or how much a surface deviates from planarity\. In this context, knowledge of the local and global geometrical features depends on an understanding of sectional curvature and Ricci curvature, respectively\([56](https://arxiv.org/html/2608.21096#bib.bib15);[65](https://arxiv.org/html/2608.21096#bib.bib5)\)\. Sectional Curvature\.This type of curvature is determined at any given point on a manifold by examining all possible two\-dimensional subspaces that intersect at that point\. It provides a more straightforward representation than the Riemann curvature tensor\([31](https://arxiv.org/html/2608.21096#bib.bib3)\)\. Recent studies\([9](https://arxiv.org/html/2608.21096#bib.bib27)\)often treat sectional curvature uniformly across the manifold, simplifying it to a singular constant value\. Ricci Curvature\.Ricci curvature averages the sectional curvatures at a specific point\. In graph theory, various discrete versions of Ricci curvature have been developed, such as Ollivier\-Ricci curvature\([49](https://arxiv.org/html/2608.21096#bib.bib4)\)and Forman\-Ricci curvature\([16](https://arxiv.org/html/2608.21096#bib.bib14)\)\. The Ricci curvature on graphs is intended to assess how the local structure around a graph edge deviates from that of a grid graph\. Notably, the Ollivier approach provides a rougher estimate of Ricci curvature, whereas the Forman method is more combinatorial and computationally efficient\. For a weighted graphG=\(V,E,w\)G=\(V,E,w\), the overall Forman\-Ricci curvatureRic¯\(G\)\\overline\{\\mathrm\{Ric\}\}\(G\)can be calculated as follows: Ric¯\(G\)=1\|E\|∑\(i,j\)∈ERic\(i,j\),\\overline\{\\mathrm\{Ric\}\}\(G\)=\\frac\{1\}\{\|E\|\}\\sum\_\{\(i,j\)\\in E\}\\mathrm\{Ric\}\(i,j\), where\|E\|\|E\|represents the cardinality of the edge setEE\(i\.e\., the total number of edges\), andRic\(i,j\)\\mathrm\{Ric\}\(i,j\)is the Forman\-Ricci curvature of the edge\(i,j\)\(i,j\), computed as\([54](https://arxiv.org/html/2608.21096#bib.bib1)\) Ric\(i,j\)=:we\(wiwe\+wjwe−∑el∼iwiwewel−∑el∼jwjwewel\)\\mathrm\{Ric\}\(i,j\)=:w\_\{e\}\\left\(\\frac\{w\_\{i\}\}\{w\_\{e\}\}\+\\frac\{w\_\{j\}\}\{w\_\{e\}\}\-\\sum\_\{e\_\{l\}\\sim i\}\\frac\{w\_\{i\}\}\{\\sqrt\{w\_\{e\}w\_\{e\_\{l\}\}\}\}\-\\sum\_\{e\_\{l\}\\sim j\}\\frac\{w\_\{j\}\}\{\\sqrt\{w\_\{e\}w\_\{e\_\{l\}\}\}\}\\right\)wherewew\_\{e\}denotes the weight of the edgeee, i\.e\.,\(i,j\)\(i,j\), andwiw\_\{i\}andwjw\_\{j\}are the weights of verticesiiandjj, respectively\. The sums overel∼ke\_\{l\}\\sim krun over all edgesele\_\{l\}incident on vertexkkexcludingee\. Specifically, when vertex and edge weights are set to11, the curvature is Ric\(i,j\):=4−di−dj\+3\|\#Δ\|,\\mathrm\{Ric\}\(i,j\):=4\-d\_\{i\}\-d\_\{j\}\+3\|\\\#\\Delta\|,wheredid\_\{i\}is the degree of nodeiiand\|\#Δ\|\|\\\#\\Delta\|is the number of 3\-cycles \(i\.e\. triangles\) containing the adjacent nodes\. In our experiments, Forman\-Ricci curvature is computed once before training using theGraphRicciCurvaturepackage\. We do not introduce additional approximations beyond the package implementation; therefore, this curvature\-estimation step is a preprocessing cost and is not part of the per\-round federated training loop\. Therefore, the overall Forman\-Ricci curvature of the graph is the weighted average of the curvature values of all edges\. ### B\.3Lorentz Transformations \(Hyperbolic Isometries\) In the context of hyperbolic geometry, Lorentz transformations are isometries of the hyperboloid model that preserve the Lorentzian inner product\. While the terminology originates from special relativity, in our setting these transformations serve purely as mathematical tools for manipulating hyperbolic embeddings\([9](https://arxiv.org/html/2608.21096#bib.bib27);[69](https://arxiv.org/html/2608.21096#bib.bib26)\)\. They can be decomposed into a combination of aLorentz Boostand aLorentz Rotation\([46](https://arxiv.org/html/2608.21096#bib.bib6)\)\. The Lorentz boost is parameterized by a vectorv∈ℝnv\\in\\mathbb\{R\}^\{n\}with‖v‖<1\\\|v\\\|<1and represented by the matrixBB\. The Lorentz rotation matrixRRrotates space\-like coordinates and is a special orthogonal matrix, i\.e\.,R⊤R=IR^\{\\top\}R=Ianddet\(R\)=1\\det\(R\)=1\. ###### Definition B\.5\(Lorentz Boost\)\. A Lorentz boost is a standard isometry of the hyperboloid model\. Given a vector𝐯∈ℝn\\mathbf\{v\}\\in\\mathbb\{R\}^\{n\}with∥𝐯∥<1\\lVert\\mathbf\{v\}\\rVert<1and the factorγ=11−∥𝐯∥2\\gamma=\\frac\{1\}\{\\sqrt\{1\-\\lVert\\mathbf\{v\}\\rVert^\{2\}\}\}, the Lorentz boost matrix is defined as: 𝐁=\[γ−γ𝐯⊤−γ𝐯𝐈\+γ21\+γ𝐯𝐯⊤\]\\mathbf\{B\}=\\begin\{bmatrix\}\\gamma&\-\\gamma\\mathbf\{v\}^\{\\top\}\\\\ \-\\gamma\\mathbf\{v\}&\\mathbf\{I\}\+\\frac\{\\gamma^\{2\}\}\{1\+\\gamma\}\\mathbf\{v\}\\mathbf\{v\}^\{\\top\}\\end\{bmatrix\}\(7\) where𝐈\\mathbf\{I\}is then×nn\\times nidentity matrix\. In our use, this matrix is interpreted only as a hyperbolic isometry acting on embeddings\. ###### Definition B\.6\(Lorentz Rotation\)\. A Lorentz rotation describes a rotation of the spatial coordinates\. The Lorentz rotation matrix is defined as: 𝐑=\[1𝟎⊤𝟎𝐑~\]\\mathbf\{R\}=\\begin\{bmatrix\}1&\\mathbf\{0\}^\{\\top\}\\\\ \\mathbf\{0\}&\\tilde\{\\mathbf\{R\}\}\\end\{bmatrix\}\(8\) where𝐑~∈SO\(n\)\\tilde\{\\mathbf\{R\}\}\\in\\text\{SO\}\(n\)is a special orthogonal matrix satisfying𝐑~⊤𝐑~=𝐈\\tilde\{\\mathbf\{R\}\}^\{\\top\}\\tilde\{\\mathbf\{R\}\}=\\mathbf\{I\}anddet\(𝐑~\)=1\\det\(\\tilde\{\\mathbf\{R\}\}\)=1\. A Lorentz rotation changes the orientation of the space\-like coordinates while leaving the time\-like coordinate unchanged\. Both the Lorentz boost and the Lorentz rotation are linear transformations defined directly in the Lorentz model\. For any point𝐱∈ℒKn\\mathbf\{x\}\\in\\mathcal\{L\}^\{n\}\_\{K\}, we have𝐁𝐱∈ℒKn\\mathbf\{B\}\\mathbf\{x\}\\in\\mathcal\{L\}^\{n\}\_\{K\}and𝐑𝐱∈ℒKn\\mathbf\{R\}\\mathbf\{x\}\\in\\mathcal\{L\}^\{n\}\_\{K\}\. ### B\.4The FedAvg Algorithm Federated Learning \(FL\) is a distributed learning approach that enables the training of machine learning models using data residing on local devices\. A cornerstone algorithm within the FL paradigm is the FedAvg algorithm\([45](https://arxiv.org/html/2608.21096#bib.bib28)\)\. FedAvg is particularly effective for scenarios where data is decentralized and not identically distributed across participants\. Algorithm 1FedAvg0:Model parameters 𝜽\\bm\{\\theta\}, learning rate η\\eta, and client dataset 𝒟c\\mathcal\{D\}\_\{c\}for each client c∈𝒞c\\in\\mathcal\{C\} 0:Aggregated model parameters 𝜽\\bm\{\\theta\} 1:Initialize model parameters 𝜽\(0\)\\bm\{\\theta\}^\{\(0\)\} 2:foreach communication round rrdo 3:foreach client ccin 𝒞\\mathcal\{C\}do 4:Client ccreceives global model parameters 𝜽\(r\)\\bm\{\\theta\}^\{\(r\)\} 5:forlocal epochs eedo 6:Compute gradients ∇ℒ=∇𝜽\(r\)∑\(𝐱,𝐲\)∈𝒟cℒc\(f\(𝐱;𝜽\(r\)\),y\)\\nabla\\mathcal\{L\}=\\nabla\_\{\\bm\{\\theta\}^\{\(r\)\}\}\\sum\_\{\(\\mathbf\{x\},\\mathbf\{y\}\)\\in\\mathcal\{D\}\_\{c\}\}\\mathcal\{L\}\_\{c\}\(f\(\\mathbf\{x\};\\bm\{\\theta\}^\{\(r\)\}\),y\) 7:endfor 8:Update local model 𝜽\(r\+1\)←𝜽\(r\)−η∇ℒ\\bm\{\\theta\}^\{\(r\+1\)\}\\leftarrow\\bm\{\\theta\}^\{\(r\)\}\-\\eta\\nabla\\mathcal\{L\} 9:Send 𝜽\(r\+1\)\\bm\{\\theta\}^\{\(r\+1\)\}to the server 10:endfor 11: N=∑c∈𝒞\|𝒟c\|N=\\sum\_\{c\\in\\mathcal\{C\}\}\|\\mathcal\{D\}\_\{c\}\| 12:Server aggregates models 𝜽\(r\+1\)←∑c∈𝒞\|𝒟c\|N𝜽c\(r\+1\)\\bm\{\\theta\}^\{\(r\+1\)\}\\leftarrow\\sum\_\{c\\in\\mathcal\{C\}\}\\frac\{\|\\mathcal\{D\}\_\{c\}\|\}\{N\}\\bm\{\\theta\}\_\{c\}^\{\(r\+1\)\} 13:endfor ## Appendix CMethodology Supplementary ### C\.1Statistics of Forman\-Ricci Curvature in Other Datasets We compute the Forman\-Ricci curvature \([AppendixB\.2](https://arxiv.org/html/2608.21096#A2.SS2)\) for each client in the Cora, Photo, and ogbn\-arxiv datasets, with 10 clients per dataset\. The CiteSeer client statistics are shown in the initialization stage of[Figure 3](https://arxiv.org/html/2608.21096#S5.F3)\. Figure 8:Averaged Forman\-Ricci curvature across datasets \(Cora, ogbn\-arxiv, and Amazon\-Photo\)\. Higher bars indicate more pronounced non\-Euclidean characteristics in these datasets\.Figure 9:Averaged Forman\-Ricci curvature across heterophilic datasets \(Questions and Tolokers\)\. Higher bars indicate more pronounced non\-Euclidean characteristics in these datasets\. ### C\.2The𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}Algorithm This section introduces the supplementary details of our𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}with pseudocode shown in[Algorithm 2](https://arxiv.org/html/2608.21096#alg2)\. #### C\.2\.1Overall Process 1. S1Initialization\.At the initial communication roundr=0r=0, the parameters that need to be initialized can be divided into three groups: 1. \(1\)Lorentz scale parameters ofCCclients\{K1,K2,…,KC\}\\\{K\_\{1\},K\_\{2\},\\ldots,K\_\{C\}\\\}; \([Section 5\.1](https://arxiv.org/html/2608.21096#S5.SS1)\) 2. \(2\)Personalized parameters ofCCclients\{𝜽1,𝜽2,…,𝜽C\}\\\{\\bm\{\\theta\}\_\{1\},\\bm\{\\theta\}\_\{2\},\\ldots,\\bm\{\\theta\}\_\{C\}\\\}; \([Section 5\.2](https://arxiv.org/html/2608.21096#S5.SS2)\) 3. \(3\)Shared parameters𝜽¯s\\overline\{\\bm\{\\theta\}\}\_\{s\}of the central server\. All the parameters of clientiiat round00can be written as𝚯i\(0\)=\(Ki,𝜽i\(0\),𝜽¯s\(0\)\)\\bm\{\\Theta\}\_\{i\}^\{\(0\)\}=\\left\(K\_\{i\};\\bm\{\\theta\}\_\{i\}^\{\(0\)\};\\overline\{\\bm\{\\theta\}\}\_\{s\}^\{\(0\)\}\\right\)and server parameters as𝜽¯s\(0\)\\overline\{\\bm\{\\theta\}\}\_\{s\}^\{\(0\)\}\. 2. S2Local updates\.Given the learning rateη\\etafor the roundrr, each local client model performs training on the data𝒟i\\mathcal\{D\}\_\{i\}to minimize the task lossℒ\(𝒟i,𝚯i\(r\)\)\\mathcal\{L\}\(\\mathcal\{D\}\_\{i\};\\bm\{\\Theta\}\_\{i\}^\{\(r\)\}\)and then updates the parameters𝚯i\(r\+1\)←𝚯i\(r\)−η∇ℒ\\bm\{\\Theta\}\_\{i\}^\{\(r\+1\)\}\\leftarrow\\bm\{\\Theta\}\_\{i\}^\{\(r\)\}\-\\eta\\nabla\\mathcal\{L\}\. \([AppendixC\.2\.3](https://arxiv.org/html/2608.21096#A3.SS2.SSS3)\) 3. S3Server updates\.After local training, only the shared parameters𝜽sc\(r\+1\)\\bm\{\\theta\}\_\{s\_\{c\}\}^\{\(r\+1\)\}are updated on the server for each clientcc\. These are then aggregated using FedAvg:𝜽¯s\(r\+1\)←∑c=1CNcN𝜽sc\(r\+1\)\\overline\{\\bm\{\\theta\}\}\_\{s\}^\{\(r\+1\)\}\\leftarrow\\sum\_\{c=1\}^\{C\}\\frac\{N\_\{c\}\}\{N\}\\bm\{\\theta\}\_\{s\_\{c\}\}^\{\(r\+1\)\}, whereN=∑c=1CNcN=\\sum\_\{c=1\}^\{C\}N\_\{c\}\. The aggregated parameters𝜽¯s\(r\+1\)\\overline\{\\bm\{\\theta\}\}\_\{s\}^\{\(r\+1\)\}are subsequently distributed to the clients for the next round\. #### C\.2\.2Curvature Initialization and Learning Details We use graph curvature only as a structure\-aware initialization\. Specifically, given a weighted graphGc=\(V,E,w\)G\_\{c\}=\(V,E,w\)on clientcc, we adopt Forman\-Ricci curvature \([AppendixB\.2](https://arxiv.org/html/2608.21096#A2.SS2)\) as a lightweight topological prior\. The overall graph curvature is computed as Ric¯\(Gc\)=1\|E\|∑\(x,y\)∈ERic\(x,y\),\\overline\{\\mathrm\{Ric\}\}\(G\_\{c\}\)=\\frac\{1\}\{\|E\|\}\\sum\_\{\(x,y\)\\in E\}\\mathrm\{Ric\}\(x,y\),\(9\)whereVVdenotes graph nodes,\|E\|\|E\|is the number of edges, and\(x,y\)\(x,y\)denotes an edge between nodesxxandyy\. Since the Lorentz scale parameterKcK\_\{c\}is positive while graph Ricci curvature can be negative, we do not directly setKc=Ric¯\(Gc\)K\_\{c\}=\\overline\{\\mathrm\{Ric\}\}\(G\_\{c\}\)\. Instead, a raw learnable scalar is initialized from the normalized signed magnitude ofRic¯\(Gc\)\\overline\{\\mathrm\{Ric\}\}\(G\_\{c\}\)and mapped to a positive Lorentz scale parameter through a sigmoid reparameterization\. The resultingKcK\_\{c\}is then optimized during local training\. This makes the Ricci statistic a topology\-aware warm start rather than a fixed pre\-estimated curvature\. Empirical comparisons with other initialization strategies are reported in[AppendixE\.5](https://arxiv.org/html/2608.21096#A5.SS5)\. #### C\.2\.3Local Training Procedure Given the Lorentz scale parameterKc\(r\)K\_\{c\}^\{\(r\)\}at roundrr, we directly project the client input𝐱iE∈𝒟c\\mathbf\{x\}\_\{i\}^\{E\}\\in\\mathcal\{D\}\_\{c\}into its corresponding Lorentz space via the exponential map𝐱Kc=exp𝐨Kc\(𝐱E\)\\mathbf\{x\}^\{K\_\{c\}\}=\\exp\_\{\\mathbf\{o\}\}^\{K\_\{c\}\}\(\\mathbf\{x\}^\{E\}\), as shown in[Equation \(1\)](https://arxiv.org/html/2608.21096#S3.E1)\. Afterward, the training data are fed into the Lorentz modelℳ\\mathcal\{M\}, and the model output isf\(𝐱Kc,𝜽c,𝜽sc\)f\(\\mathbf\{x\}^\{K\_\{c\}\};\\bm\{\\theta\}\_\{c\},\\bm\{\\theta\}\_\{s\_\{c\}\}\)\. At clientcc, the objective function isminKc\>0,𝜽c,𝜽scℒc\(f\(𝐱Kc,𝜽c,𝜽sc\),y\)\+λ‖𝜽sc−𝜽¯s‖22,\\min\_\{K\_\{c\}\>0,\\bm\{\\theta\}\_\{c\},\\bm\{\\theta\}\_\{s\_\{c\}\}\}\\mathcal\{L\}\_\{c\}\(f\(\\mathbf\{x\}^\{K\_\{c\}\};\\bm\{\\theta\}\_\{c\},\\bm\{\\theta\}\_\{s\_\{c\}\}\),y\)\+\\lambda\\\|\\bm\{\\theta\}\_\{s\_\{c\}\}\-\\overline\{\\bm\{\\theta\}\}\_\{s\}\\\|\_\{2\}^\{2\},whereλ\\lambdais a hyperparameter,𝜽sc\\bm\{\\theta\}\_\{s\_\{c\}\}denotes clientcc’s local copy of the shared block, and‖𝜽sc−𝜽¯s‖22\\\|\\bm\{\\theta\}\_\{s\_\{c\}\}\-\\overline\{\\bm\{\\theta\}\}\_\{s\}\\\|\_\{2\}^\{2\}is the regularization term that prevents the locally updated model𝜽sc\\bm\{\\theta\}\_\{s\_\{c\}\}from deviating too far from the server\-shared parameters𝜽¯s\\overline\{\\bm\{\\theta\}\}\_\{s\}\. The optimization overKcK\_\{c\}is implemented through the raw\-scalar reparameterization described above, whileKcK\_\{c\}remains local and is not uploaded for server aggregation\. Algorithm 2𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}0:Personalized parameters 𝜽c\(0\),Kc\(0\)\\bm\{\\theta\}\_\{c\}^\{\(0\)\},K\_\{c\}^\{\(0\)\}and dataset 𝒟c\\mathcal\{D\}\_\{c\}, for each client c∈𝒞c\\in\\mathcal\{C\}; Shared parameters 𝜽¯s\(0\)\\overline\{\\bm\{\\theta\}\}\_\{s\}^\{\(0\)\}; Learning rate η\\eta 0:Client model parameters 𝚯c=\(Kc,𝜽c,𝜽¯s\)\\bm\{\\Theta\}\_\{c\}=\\left\(K\_\{c\};\\bm\{\\theta\}\_\{c\};\\overline\{\\bm\{\\theta\}\}\_\{s\}\\right\), for each client c∈𝒞c\\in\\mathcal\{C\}; Shared parameters 𝜽¯s\\overline\{\\bm\{\\theta\}\}\_\{s\} 1:Initialize model parameters: 𝜽¯s\(0\)and𝚯c\(0\)=\(Kc\(0\),𝜽c\(0\),𝜽¯s\(0\)\)\\overline\{\\bm\{\\theta\}\}\_\{s\}^\{\(0\)\}\\text\{ and \}\\bm\{\\Theta\}\_\{c\}^\{\(0\)\}=\\left\(K\_\{c\}^\{\(0\)\};\\bm\{\\theta\}\_\{c\}^\{\(0\)\};\\overline\{\\bm\{\\theta\}\}\_\{s\}^\{\(0\)\}\\right\), for c∈𝒞c\\in\\mathcal\{C\} 2:foreach communication round rrdo 3:foreach client ccin 𝒞\\mathcal\{C\}do 4: 𝐱=exp𝐨Kc\(r\)\(𝐱\)\\mathbf\{x\}=\\exp\_\{\\mathbf\{o\}\}^\{K\_\{c\}^\{\(r\)\}\}\{\(\\mathbf\{x\}\)\}, for 𝐱∈𝒟c\\mathbf\{x\}\\in\\mathcal\{D\}\_\{c\} 5:Client ccreceives global model parameters 𝜽¯s\(r\)\\overline\{\\bm\{\\theta\}\}\_\{s\}^\{\(r\)\} 6: 𝚯c\(r\)=\(Kc\(r\),𝜽c\(r\),𝜽¯s\(r\)\)\\bm\{\\Theta\}\_\{c\}^\{\(r\)\}=\\left\(K\_\{c\}^\{\(r\)\};\\bm\{\\theta\}\_\{c\}^\{\(r\)\};\\overline\{\\bm\{\\theta\}\}\_\{s\}^\{\(r\)\}\\right\) 7:forlocal epochs eedo 8:Compute gradients ∇ℒ=∇𝚯c\(r\)∑\(𝐱,𝐲\)∈𝒟cℒc\(f\(𝐱;𝚯c\(r\)\),y\)\\nabla\\mathcal\{L\}=\\nabla\_\{\\bm\{\\Theta\}^\{\(r\)\}\_\{c\}\}\\sum\_\{\(\\mathbf\{x\},\\mathbf\{y\}\)\\in\\mathcal\{D\}\_\{c\}\}\\mathcal\{L\}\_\{c\}\(f\(\\mathbf\{x\};\\bm\{\\Theta\}^\{\(r\)\}\_\{c\}\),y\) 9:endfor 10:Update local model 𝚯\(r\+1\)c←𝚯c\(r\)−η∇ℒ\\bm\{\\Theta\}^\{\(r\+1\)\}\_\{c\}\\leftarrow\\bm\{\\Theta\}\_\{c\}^\{\(r\)\}\-\\eta\\nabla\\mathcal\{L\} 11:Send 𝜽sc\(r\+1\)∈𝚯c\(r\+1\)\\bm\{\\theta\}\_\{s\_\{c\}\}^\{\(r\+1\)\}\\in\\bm\{\\Theta\}\_\{c\}^\{\(r\+1\)\}to the server 12:endfor 13: N=∑c∈𝒞\|𝒟c\|N=\\sum\_\{c\\in\\mathcal\{C\}\}\|\\mathcal\{D\}\_\{c\}\| 14:Server aggregates models 𝜽¯s\(r\+1\)←∑c∈𝒞\|𝒟c\|N𝜽sc\(r\+1\)\\overline\{\\bm\{\\theta\}\}\_\{s\}^\{\(r\+1\)\}\\leftarrow\\sum\_\{c\\in\\mathcal\{C\}\}\\frac\{\|\\mathcal\{D\}\_\{c\}\|\}\{N\}\\bm\{\\theta\}\_\{s\_\{c\}\}^\{\(r\+1\)\} 15:endfor ### C\.3Derivation of Parameters Disentanglement The reformulated Lorentz neural network in layerllis shown as 𝐱\(l\+1\)\\displaystyle\\mathbf\{x\}^\{\(l\+1\)\}=LT\(𝐱\(l\),𝐌^\(l\)\)\\displaystyle=\\mathrm\{LT\}\(\\mathbf\{x\}^\{\(l\)\};\\hat\{\\mathbf\{M\}\}^\{\(l\)\}\)\(10\)=\(‖𝐦\(l\)xt\(l\)\+𝐌\(l\)𝐱s\(l\)‖2\+K⏟time\-like dimensionxt\(l\+1\),𝐦\(l\)xt\(l\)\+𝐌\(l\)𝐱s\(l\)⏟space\-like dimensions𝐱s\(l\+1\)\)T\.\\displaystyle=\\left\(\\underbrace\{\\sqrt\{\\\|\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\+\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}\\\|^\{2\}\+K\}\}\_\{\\begin\{subarray\}\{c\}\\text\{time\-like dimension \}\\\\ x\_\{t\}^\{\(l\+1\)\}\\end\{subarray\}\},\\underbrace\{\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\+\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}\}\_\{\\begin\{subarray\}\{c\}\\text\{space\-like dimensions \}\\\\ \\mathbf\{x\}\_\{s\}^\{\(l\+1\)\}\\end\{subarray\}\}\\right\)^\{T\}\.\(11\) For the lossℒc\(f\(𝐱,𝜽c,𝜽s\),y\)\\mathcal\{L\}\_\{c\}\(f\(\\mathbf\{x\};\\bm\{\\theta\}\_\{c\},\\bm\{\\theta\}\_\{s\}\),y\)of clientcc, the relevant partial derivatives can be calculated as follows: ### Time\-like Dimensionxt\(l\+1\)x\_\{t\}^\{\(l\+1\)\} First, we compute the partial derivative ofxt\(l\+1\)x\_\{t\}^\{\(l\+1\)\}with respect to the matrix𝐌\(l\)\\mathbf\{M\}^\{\(l\)\}and the vector block𝐦\(l\)\\mathbf\{m\}^\{\(l\)\}\. Using the chain rule: ∂xt\(l\+1\)∂𝐌\(l\)=∂∂𝐌‖𝐦\(l\)xt\(l\)\+𝐌\(l\)𝐱s\(l\)‖2\+K;\\frac\{\\partial x\_\{t\}^\{\(l\+1\)\}\}\{\\partial\\mathbf\{M\}^\{\(l\)\}\}=\\frac\{\\partial\}\{\\partial\\mathbf\{M\}\}\\sqrt\{\\\|\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\+\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}\\\|^\{2\}\+K\};∂xt\(l\+1\)∂𝐦\(l\)=∂∂𝐦‖𝐦\(l\)xt\(l\)\+𝐌\(l\)𝐱s\(l\)‖2\+K\.\\frac\{\\partial x\_\{t\}^\{\(l\+1\)\}\}\{\\partial\\mathbf\{m\}^\{\(l\)\}\}=\\frac\{\\partial\}\{\\partial\\mathbf\{m\}\}\\sqrt\{\\\|\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\+\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}\\\|^\{2\}\+K\}\.Applying the chain rule, we get: ∂xt\(l\+1\)∂𝐌\(l\)\\displaystyle\\frac\{\\partial x\_\{t\}^\{\(l\+1\)\}\}\{\\partial\\mathbf\{M\}^\{\(l\)\}\}=12\(‖𝐦\(l\)xt\(l\)\+𝐌\(l\)𝐱s\(l\)‖2\+K\)−12\\displaystyle=\\frac\{1\}\{2\}\\left\(\\\|\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\+\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}\\\|^\{2\}\+K\\right\)^\{\-\\frac\{1\}\{2\}\}\(12\)⋅2\(𝐦\(l\)xt\(l\)\+𝐌\(l\)𝐱s\(l\)\)⋅∂\(𝐌\(l\)𝐱s\(l\)\)∂𝐌\(l\)\\displaystyle\\cdot 2\(\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\+\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}\)\\cdot\\frac\{\\partial\(\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}\)\}\{\\partial\\mathbf\{M\}^\{\(l\)\}\}=𝐦\(l\)xt\(l\)\+𝐌\(l\)𝐱s\(l\)‖𝐦\(l\)xt\(l\)\+𝐌\(l\)𝐱s\(l\)‖2\+K\\displaystyle=\\frac\{\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\+\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}\}\{\\sqrt\{\\\|\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\+\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}\\\|^\{2\}\+K\}\}⋅∂\(𝐌\(l\)𝐱s\(l\)\)∂𝐌\(l\)\\displaystyle\\cdot\\frac\{\\partial\(\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}\)\}\{\\partial\\mathbf\{M\}^\{\(l\)\}\} ∂xt\(l\+1\)∂𝐦\(l\)\\displaystyle\\frac\{\\partial x\_\{t\}^\{\(l\+1\)\}\}\{\\partial\\mathbf\{m\}^\{\(l\)\}\}=12\(‖𝐦\(l\)xt\(l\)\+𝐌\(l\)𝐱s\(l\)‖2\+K\)−12\\displaystyle=\\frac\{1\}\{2\}\\left\(\\\|\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\+\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}\\\|^\{2\}\+K\\right\)^\{\-\\frac\{1\}\{2\}\}\(13\)⋅2\(𝐦\(l\)xt\(l\)\+𝐌\(l\)𝐱s\(l\)\)⋅∂\(𝐦\(l\)xt\(l\)\)∂𝐦\(l\)\\displaystyle\\cdot 2\(\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\+\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}\)\\cdot\\frac\{\\partial\(\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\)\}\{\\partial\\mathbf\{m\}^\{\(l\)\}\}=𝐦\(l\)xt\(l\)\+𝐌\(l\)𝐱s\(l\)‖𝐦\(l\)xt\(l\)\+𝐌\(l\)𝐱s\(l\)‖2\+K⋅xt\(l\)\\displaystyle=\\frac\{\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\+\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}\}\{\\sqrt\{\\\|\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\+\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}\\\|^\{2\}\+K\}\}\\cdot x\_\{t\}^\{\(l\)\} ### Space\-like Dimension𝐱s\(l\+1\)\\mathbf\{x\}\_\{s\}^\{\(l\+1\)\} Assume that the update rule for the space\-like vector𝐱s\(l\+1\)\\mathbf\{x\}\_\{s\}^\{\(l\+1\)\}is given by the following formula: 𝐱s\(l\+1\)=𝐦\(l\)xt\(l\)\+𝐌\(l\)𝐱s\(l\)\\mathbf\{x\}\_\{s\}^\{\(l\+1\)\}=\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\+\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\} Similarly, we have ∂𝐱s\(l\+1\)∂𝐌\(l\)=∂\(𝐌\(l\)𝐱s\(l\)\)∂𝐌\(l\),∂𝐱s\(l\+1\)∂𝐦\(l\)=∂\(𝐦\(l\)xt\(l\)\)∂𝐦\(l\)\.\\frac\{\\partial\\mathbf\{x\}\_\{s\}^\{\(l\+1\)\}\}\{\\partial\\mathbf\{M\}^\{\(l\)\}\}=\\frac\{\\partial\\left\(\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}\\right\)\}\{\\partial\\mathbf\{M\}^\{\(l\)\}\},\\quad\\frac\{\\partial\\mathbf\{x\}\_\{s\}^\{\(l\+1\)\}\}\{\\partial\\mathbf\{m\}^\{\(l\)\}\}=\\frac\{\\partial\\left\(\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\\right\)\}\{\\partial\\mathbf\{m\}^\{\(l\)\}\}\.\(14\) "Flatland"denotes the space\-like dimensions1:n1\{:\}n, which serve as the platform where common information is exchanged and integrated\. In the space\-like update𝐱s\(l\+1\)=𝐦\(l\)xt\(l\)\+𝐌\(l\)𝐱s\(l\)\\mathbf\{x\}\_\{s\}^\{\(l\+1\)\}=\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\+\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}, the update path of𝐌\(l\)\\mathbf\{M\}^\{\(l\)\}is tied to𝐱s\(l\)\\mathbf\{x\}\_\{s\}^\{\(l\)\}, whereas the update path of𝐦\(l\)\\mathbf\{m\}^\{\(l\)\}is tied toxt\(l\)x\_\{t\}^\{\(l\)\}\. For better illustration, here, we let𝐱\(l\)∈ℒKn\\mathbf\{x\}^\{\(l\)\}\\in\\mathcal\{L\}\_\{K\}^\{n\},𝐱\(l\+1\)∈ℒKn\\mathbf\{x\}^\{\(l\+1\)\}\\in\\mathcal\{L\}\_\{K\}^\{n\}, and𝐌^\(l\)∈ℝn×\(n\+1\)\\hat\{\\mathbf\{M\}\}^\{\(l\)\}\\in\\mathbb\{R\}^\{n\\times\(n\+1\)\}\. The introduced"Flatland"ℝn\\mathbb\{R\}^\{n\}is defined as a manifold spanning dimensions11tonn\. This construct serves as a metaphorical platform for the exchange and integration of common information, andxtx\_\{t\}serves as the heterogeneous information\. Consider the same transformation of a space\-like vector𝐱s\(l\)\\mathbf\{x\}\_\{s\}^\{\(l\)\}to𝐱s\(l\+1\)\\mathbf\{x\}\_\{s\}^\{\(l\+1\)\}in different clients, formulated as 𝐱s\(l\)→\(𝐌\(l\)𝐱s\(l\)\+𝐦\(l\)xt\(l\)\),\\mathbf\{x\}\_\{s\}^\{\(l\)\}\\rightarrow\\left\(\\mathbf\{M\}^\{\(l\)\}\\mathbf\{x\}\_\{s\}^\{\(l\)\}\+\\mathbf\{m\}^\{\(l\)\}x\_\{t\}^\{\(l\)\}\\right\), ### C\.4Time and Space Complexity Compared with FedAvg We analyze the computational complexity of FlatLand compared to FedAvg, which gives insight into the scalability\. ##### Local Update\. The additional operations in𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}’s local update phase compared with FedAvg \- curvature estimation \(Section[5\.1](https://arxiv.org/html/2608.21096#S5.SS1)\), exponential map \(line 4 in Algorithm[2](https://arxiv.org/html/2608.21096#alg2), Equation[6](https://arxiv.org/html/2608.21096#A2.E6)\)\. Notably, the curvature estimation can bepre\-computedsince each client’s data distribution corresponds to a constant curvature value\. For exponential map, the transformation only requiresa singlenon\-linear mapping operation based on the norm of input samples with the time complexity ofO\(1\)O\(1\)\. These norms can also bepre\-computed and cached\.Therefore, while𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}introduces these additional steps compared to FedAvg, their practical computational overhead is limited due to pre\-computation opportunities and constant\-time operations\. ##### Aggregation\. 𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}and FedAvg have the same aggregation time complexity when the hidden embedding dimension is the same\. Though𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}introduces extra time\-like space parameters, it only aggregates shared parameters𝜽s\\bm\{\\theta\}\_\{s\}while maintaining personalized parameters\. The overhead of the shared parameters is the same\. Moreover,𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}can perform better in low dimensionality \(Section[7\.3](https://arxiv.org/html/2608.21096#S7.SS3)\), which potentially reduces practical communication costs\. ##### Space Requirements and Storage\. 𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}requires extraO\(d\+1\)O\(d\+1\)storage per client compared to FedAvg due to the additional time\-like dimension and Lorentz scale parameter, whereddis the hidden dimension\. Since typicallyddis small, the increase in storage is small\. Moreover,𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}demonstrates superior performance even in low\-dimensional settings compared with the Euclidean counterparts, which further limits the practical storage overhead\. This analysis suggests that𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}can balance the trade\-off between computational overhead and model effectiveness, showing the scalability for the increase in clients\. While it introduces additional operations in local computations, these overheads are limited and offer significant optimization opportunities through pre\-computation and caching strategies\. The method compensates for these minimal costs through reduced communication overhead and enhanced representation capabilities in the Lorentz space, making it a practical and efficient choice for personalized federated learning applications\. ## Appendix DTheoretical Analysis Supplementary ### D\.1Proof for Theorem[4\.1](https://arxiv.org/html/2608.21096#S4.Thmtheorem1) ###### Lemma D\.1\. LetG=\(V,E\)G=\(V,E\)be a connected, simple, unweighted graph with maximum degreeΔ<∞\\Delta<\\inftyand average Forman–Ricci curvatureR¯=Ric¯\(G\)\\bar\{R\}=\\overline\{\\mathrm\{Ric\}\}\(G\)\(edge\-average\), where the unweighted edge curvature follows[AppendixB\.2](https://arxiv.org/html/2608.21096#A2.SS2)\. Define the averaged ball countsVG\(r\):=1\|V\|∑v∈V\|BrG\(v\)\|V\_\{G\}\(r\):=\\frac\{1\}\{\|V\|\}\\sum\_\{v\\in V\}\|B\_\{r\}^\{G\}\(v\)\|forr=0,1,2r=0,1,2\. Then Δ2VG\(1\):=VG\(2\)−2VG\(1\)\+VG\(0\)≥1−\|E\|\|V\|R¯\.\\Delta^\{2\}V\_\{G\}\(1\)\\;:=\\;V\_\{G\}\(2\)\-2V\_\{G\}\(1\)\+V\_\{G\}\(0\)\\ \\geq\\ 1\-\\frac\{\|E\|\}\{\|V\|\}\\,\\bar\{R\}\.\(15\) ###### Proof\. For the unweighted specialization in[AppendixB\.2](https://arxiv.org/html/2608.21096#A2.SS2), the Forman–Ricci curvature on an edgee=\(u,v\)e=\(u,v\)isRic\(u,v\)=4−deg\(u\)−deg\(v\)\+3tuv\\mathrm\{Ric\}\(u,v\)=4\-\\deg\(u\)\-\\deg\(v\)\+3t\_\{uv\}, wheretuvt\_\{uv\}is the number of triangles containingee\. Hencedeg\(u\)\+deg\(v\)−2=2\+3tuv−Ric\(u,v\)\\deg\(u\)\+\\deg\(v\)\-2=2\+3t\_\{uv\}\-\\mathrm\{Ric\}\(u,v\)\. Counting non\-backtracking two\-step walks yields 1\|V\|∑v∈V∑u∼v\(deg\(u\)−1\)\\displaystyle\\frac\{1\}\{\|V\|\}\\sum\_\{v\\in V\}\\sum\_\{u\\sim v\}\(\\deg\(u\)\-1\)=1\|V\|∑\(u,v\)∈E\(deg\(u\)\+deg\(v\)−2\)\\displaystyle=\\frac\{1\}\{\|V\|\}\\sum\_\{\(u,v\)\\in E\}\\bigl\(\\deg\(u\)\+\\deg\(v\)\-2\\bigr\)=1\|V\|\(2\|E\|\+3∑e∈Ete−∑e∈ERic\(e\)\)\\displaystyle=\\frac\{1\}\{\|V\|\}\\Bigl\(2\|E\|\+3\\sum\_\{e\\in E\}t\_\{e\}\-\\sum\_\{e\\in E\}\\mathrm\{Ric\}\(e\)\\Bigr\)≥1\|V\|\(2\|E\|−∑e∈ERic\(e\)\)\.\\displaystyle\\geq\\frac\{1\}\{\|V\|\}\\Bigl\(2\|E\|\-\\sum\_\{e\\in E\}\\mathrm\{Ric\}\(e\)\\Bigr\)\.UsingVG\(0\)=1V\_\{G\}\(0\)=1,VG\(1\)=1\+1\|V\|∑vdeg\(v\)=1\+2\|E\|\|V\|V\_\{G\}\(1\)=1\+\\frac\{1\}\{\|V\|\}\\sum\_\{v\}\\deg\(v\)=1\+\\frac\{2\|E\|\}\{\|V\|\}, and adding the two\-step term gives[Equation \(15\)](https://arxiv.org/html/2608.21096#A4.E15)\. The triangle term contributes a nonnegative correction, so dropping it yields a weaker bound in the same form; whente=0t\_\{e\}=0for all edges, this reduces to the simplified no\-triangle case\. ∎ ###### Lemma D\.2\(\([20](https://arxiv.org/html/2608.21096#bib.bib67)\)\)\. LetℒKd\\mathcal\{L\}\_\{K\}^\{d\}be thedd\-dimensional hyperbolic space with Lorentz scale parameterK\>0K\>0and constant sectional curvature−1/K<0\-1/K<0, andVolℒ\(Bρ\)\\mathrm\{Vol\}\_\{\\mathcal\{L\}\}\(B\_\{\\rho\}\)the volume of a radius\-ρ\\rhoball\. For smallρ\\rho, Volℒ\(Bρ\)=ωdρd\(1\+αdd−1Kρ2\+O\(ρ4\)\),αd=d6\(d\+2\)\.\\mathrm\{Vol\}\_\{\\mathcal\{L\}\}\(B\_\{\\rho\}\)=\\omega\_\{d\}\\,\\rho^\{d\}\\Bigl\(1\+\\alpha\_\{d\}\\,\\frac\{d\-1\}\{K\}\\,\\rho^\{2\}\+O\(\\rho^\{4\}\)\\Bigr\),\\quad\\alpha\_\{d\}=\\frac\{d\}\{6\(d\+2\)\}\.\(16\)Taking the discrete second difference inρ∈\{0,1,2\}\\rho\\in\\\{0,1,2\\\}gives Δ2Volℒ\(1\)=Volℒ\(B2\)−2Volℒ\(B1\)\+Volℒ\(B0\)=Cdd−1K\+O\(1\),\\Delta^\{2\}\\mathrm\{Vol\}\_\{\\mathcal\{L\}\}\(1\)=\\mathrm\{Vol\}\_\{\\mathcal\{L\}\}\(B\_\{2\}\)\-2\\mathrm\{Vol\}\_\{\\mathcal\{L\}\}\(B\_\{1\}\)\+\\mathrm\{Vol\}\_\{\\mathcal\{L\}\}\(B\_\{0\}\)=C\_\{d\}\\,\\frac\{d\-1\}\{K\}\+O\(1\),\(17\)whereCd:=ωdαdΔ2\[ρd\+2\]ρ=1\>0C\_\{d\}:=\\omega\_\{d\}\\,\\alpha\_\{d\}\\,\\Delta^\{2\}\[\\rho^\{d\+2\}\]\_\{\\rho=1\}\>0\. ###### Lemma D\.3\(Local bi\-Lipschitz sandwich\)\. Letf:V→ℒKdf:V\\to\\mathcal\{L\}\_\{K\}^\{d\}be\(1\+ε\)\(1\+\\varepsilon\)\-bi\-Lipschitz on graph balls of radius22, i\.e\.,dℒ\(f\(x\),f\(y\)\)∈\[\(1\+ε\)−1dG\(x,y\),\(1\+ε\)dG\(x,y\)\]d\_\{\\mathcal\{L\}\}\\bigl\(f\(x\),f\(y\)\\bigr\)\\in\[\(1\+\\varepsilon\)^\{\-1\}d\_\{G\}\(x,y\),\(1\+\\varepsilon\)d\_\{G\}\(x,y\)\]wheneverdG\(x,y\)≤2d\_\{G\}\(x,y\)\\leq 2\. Then there exist constantsAd,Δ,Bd,Δ\>0A\_\{d,\\Delta\},B\_\{d,\\Delta\}\>0such that Ad,ΔVolℒ\(B\(1−ε\)ρ\)≤VG\(ρ\)≤Bd,ΔVolℒ\(B\(1\+ε\)ρ\),ρ∈\{0,1,2\}\.A\_\{d,\\Delta\}\\,\\mathrm\{Vol\}\_\{\\mathcal\{L\}\}\\bigl\(B\_\{\(1\-\\varepsilon\)\\rho\}\\bigr\)\\ \\leq\\ V\_\{G\}\(\\rho\)\\ \\leq\\ B\_\{d,\\Delta\}\\,\\mathrm\{Vol\}\_\{\\mathcal\{L\}\}\\bigl\(B\_\{\(1\+\\varepsilon\)\\rho\}\\bigr\),\\qquad\\rho\\in\\\{0,1,2\\\}\.\(18\)Consequently, taking discrete second differences and Taylor\-expanding atε=0\\varepsilon=0, Δ2VG\(1\)=Γd,Δd−1K±Λd,Δε\+O\(ε2\),\\Delta^\{2\}V\_\{G\}\(1\)=\\Gamma\_\{d,\\Delta\}\\,\\frac\{d\-1\}\{K\}\\ \\pm\\ \\Lambda\_\{d,\\Delta\}\\,\\varepsilon\\ \+\\ O\(\\varepsilon^\{2\}\),\(19\)for someΓd,Δ,Λd,Δ\>0\\Gamma\_\{d,\\Delta\},\\Lambda\_\{d,\\Delta\}\>0\. ###### Proof\. Inclusionsf\(BρG\(v\)\)⊂B\(1\+ε\)ρℒ\(f\(v\)\)f\(B\_\{\\rho\}^\{G\}\(v\)\)\\subset B\_\{\(1\+\\varepsilon\)\\rho\}^\{\\mathcal\{L\}\}\(f\(v\)\)andB\(1−ε\)ρℒ\(f\(v\)\)⊂f\(BρG\(v\)\)B\_\{\(1\-\\varepsilon\)\\rho\}^\{\\mathcal\{L\}\}\(f\(v\)\)\\subset f\(B\_\{\\rho\}^\{G\}\(v\)\)follow from the bi\-Lipschitz bounds\. Degree boundΔ\\Deltagives packing/covering constants relating point counts and volumes, yielding the sandwich; applyΔ2\\Delta^\{2\}inρ\\rhoand a first\-order Taylor expansion inε\\varepsilon\. ∎ ###### Lemma D\.4\(Curvature mismatch⇒\\Rightarrowlocal distortion\)\. Under the conditions of[LemmaD\.1](https://arxiv.org/html/2608.21096#A4.Thmtheorem1)–[LemmaD\.3](https://arxiv.org/html/2608.21096#A4.Thmtheorem3), there exist constantscd\>0c\_\{d\}\>0andCd,Δ\>0C\_\{d,\\Delta\}\>0such that any\(1\+ε\)\(1\+\\varepsilon\)\-bi\-Lipschitzffon radius\-22balls satisfies ε≥cd\|R¯\+\(d−1\)/K\|−Cd,Δε2\.\\varepsilon\\ \\geq\\ c\_\{d\}\\bigl\|\\bar\{R\}\+\(d\-1\)/K\\bigr\|\-C\_\{d,\\Delta\}\\,\\varepsilon^\{2\}\.\(20\)In particular, forε∈\(0,ε0\(d,Δ\)\)\\varepsilon\\in\(0,\\varepsilon\_\{0\}\(d,\\Delta\)\), ε≥12cd\|R¯\+\(d−1\)/K\|\.\\varepsilon\\ \\geq\\ \\tfrac\{1\}\{2\}c\_\{d\}\\bigl\|\\bar\{R\}\+\(d\-1\)/K\\bigr\|\.\(21\) ###### Proof\. Combine[Equation \(15\)](https://arxiv.org/html/2608.21096#A4.E15)and[Equation \(19\)](https://arxiv.org/html/2608.21096#A4.E19); rearrange to isolate\|R¯\+\(d−1\)/K\|\|\\bar\{R\}\+\(d\-1\)/K\|in terms ofε\\varepsilon, absorbing constants intocd,Cd,Δc\_\{d\},C\_\{d,\\Delta\}\. For sufficiently smallε\\varepsilon, the quadratic term is dominated, giving[Equation \(21\)](https://arxiv.org/html/2608.21096#A4.E21)\. ∎ ###### Theorem D\.5\(Recall of[Theorem 4\.1](https://arxiv.org/html/2608.21096#S4.Thmtheorem1)\)\. Let\{Gc\}c=1C\\\{G\_\{c\}\\\}\_\{c=1\}^\{C\}be client graphs with average Forman\-Ricci curvaturesR¯c=Ric¯\(Gc\)\\bar\{R\}\_\{c\}=\\overline\{\\mathrm\{Ric\}\}\(G\_\{c\}\), and letℒKd\\mathcal\{L\}^\{d\}\_\{K\}denote thedd\-dimensional hyperbolic space with Lorentz scale parameterK\>0K\>0and constant sectional curvature−1/K<0\-1/K<0\. For each clientcc, letεc∗\(K\)\\varepsilon\_\{c\}^\{\*\}\(K\)be the minimal edge distortion of any\(1\+ε\)\(1\+\\varepsilon\)\-bi\-Lipschitz embeddingfc:Gc→ℒKdf\_\{c\}:G\_\{c\}\\to\\mathcal\{L\}^\{d\}\_\{K\}\. Then the following holds: max1≤c≤Cεc∗\(K\)≥cd2max1≤i<j≤C\|R¯i−R¯j\|,\\max\_\{1\\leq c\\leq C\}\\ \\varepsilon\_\{c\}^\{\*\}\(K\)\\;\\;\\geq\\;\\;\\frac\{c\_\{d\}\}\{2\}\\,\\max\_\{1\\leq i<j\\leq C\}\\,\|\\bar\{R\}\_\{i\}\-\\bar\{R\}\_\{j\}\|,\(22\)wherecd\>0c\_\{d\}\>0is a dimension\-dependent constant\. ###### Proof\. For clientcc,[LemmaD\.4](https://arxiv.org/html/2608.21096#A4.Thmtheorem4)applied toGcG\_\{c\}gives \(for small optimal distortion\) εc∗\(K\)≥cd\|R¯c\+\(d−1\)/K\|\.\\varepsilon\_\{c\}^\{\*\}\(K\)\\ \\geq\\ c\_\{d\}\\,\\bigl\|\\bar\{R\}\_\{c\}\+\(d\-1\)/K\\bigr\|\.\(23\)For anyi≠ji\\neq jsetai:=R¯i\+\(d−1\)/Ka\_\{i\}:=\\bar\{R\}\_\{i\}\+\(d\-1\)/K,aj:=R¯j\+\(d−1\)/Ka\_\{j\}:=\\bar\{R\}\_\{j\}\+\(d\-1\)/K\. Thenai−aj=R¯i−R¯ja\_\{i\}\-a\_\{j\}=\\bar\{R\}\_\{i\}\-\\bar\{R\}\_\{j\}, somax\{\|ai\|,\|aj\|\}≥12\|ai−aj\|=12\|R¯i−R¯j\|\\max\\\{\|a\_\{i\}\|,\|a\_\{j\}\|\\\}\\geq\\tfrac\{1\}\{2\}\|a\_\{i\}\-a\_\{j\}\|=\\tfrac\{1\}\{2\}\|\\bar\{R\}\_\{i\}\-\\bar\{R\}\_\{j\}\|\. Taking the maximum over all pairs\(i,j\)\(i,j\)yields max1≤c≤C\|R¯c\+\(d−1\)/K\|≥12max1≤i<j≤C\|R¯i−R¯j\|\.\\max\_\{1\\leq c\\leq C\}\\bigl\|\\,\\bar\{R\}\_\{c\}\+\(d\-1\)/K\\,\\bigr\|\\ \\geq\\ \\frac\{1\}\{2\}\\max\_\{1\\leq i<j\\leq C\}\|\\bar\{R\}\_\{i\}\-\\bar\{R\}\_\{j\}\|\.\(24\)Combining[Equation \(23\)](https://arxiv.org/html/2608.21096#A4.E23)and[Equation \(24\)](https://arxiv.org/html/2608.21096#A4.E24)gives max1≤c≤Cεc∗\(K\)≥cdmaxc\|R¯c\+\(d−1\)/K\|≥cd2max1≤i<j≤C\|R¯i−R¯j\|,\\max\_\{1\\leq c\\leq C\}\\ \\varepsilon\_\{c\}^\{\*\}\(K\)\\ \\geq\\ c\_\{d\}\\max\_\{c\}\\bigl\|\\,\\bar\{R\}\_\{c\}\+\(d\-1\)/K\\,\\bigr\|\\ \\geq\\ \\frac\{c\_\{d\}\}\{2\}\\,\\max\_\{1\\leq i<j\\leq C\}\|\\bar\{R\}\_\{i\}\-\\bar\{R\}\_\{j\}\|,which is the desired inequality\. If\{R¯c\}\\\{\\bar\{R\}\_\{c\}\\\}are not all equal, the right\-hand side is strictly positive, hence a single Lorentz scale parameterKKcannot make all clients’ distortions simultaneously small\. ∎ ### D\.2Proof for Theorem[4\.3](https://arxiv.org/html/2608.21096#S4.Thmtheorem3) For each clientccwith Lorentz scale parameterKc\>0K\_\{c\}\>0, consider𝐱∈𝕃Kcd\\mathbf\{x\}\\in\\mathbb\{L\}\_\{K\_\{c\}\}^\{d\}expressed in hyperbolic polar coordinates\(ρ,𝐮\)\(\\rho,\\mathbf\{u\}\), whereρ≥0\\rho\\geq 0is the radial coordinate and𝐮∈Ω:=𝕊d−1\\mathbf\{u\}\\in\\Omega:=\\mathbb\{S\}^\{d\-1\}is the angular coordinate \(unit direction vector\)\. We denote byp\(ρ,𝐮∣Kc\)p\(\\rho,\\mathbf\{u\}\\mid K\_\{c\}\)the joint distribution of\(ρ,𝐮\)\(\\rho,\\mathbf\{u\}\)givenKcK\_\{c\}\. ###### Assumption D\.6\(Isotropy at the basepoint\([30](https://arxiv.org/html/2608.21096#bib.bib17)\)\)\. For each clientccwith Lorentz scale parameterKcK\_\{c\},p\(ρ,𝐮∣Kc\)p\(\\rho,\\mathbf\{u\}\\mid K\_\{c\}\)isGoG\_\{o\}–isotropic at the basepointoo, i\.e\.p\(g⋅𝐱∣Kc\)=p\(𝐱∣Kc\)∀g∈Go≃SO\(d\)\.p\(g\\\!\\cdot\\\!\\mathbf\{x\}\\mid K\_\{c\}\)=p\(\\mathbf\{x\}\\mid K\_\{c\}\)\\quad\\forall g\\in G\_\{o\}\\simeq SO\(d\)\.Equivalently, the joint law factorizes as p\(ρ,𝐮∣Kc\)=pρ\(ρ∣Kc\)pΩ\(𝐮\),p\(\\rho,\\mathbf\{u\}\\mid K\_\{c\}\)=p\_\{\\rho\}\(\\rho\\mid K\_\{c\}\)\\,p\_\{\\Omega\}\(\\mathbf\{u\}\),wherepρ\(ρ∣Kc\)p\_\{\\rho\}\(\\rho\\mid K\_\{c\}\)is the radial density depending onKcK\_\{c\}, andpΩ\(𝐮\)p\_\{\\Omega\}\(\\mathbf\{u\}\)is theSO\(d\)SO\(d\)–invariant angular density onΩ\\Omega, independent ofKcK\_\{c\}\(in particular, uniform on𝕊d−1\\mathbb\{S\}^\{d\-1\}\)\. ###### Lemma D\.7\. Under Assumption[D\.6](https://arxiv.org/html/2608.21096#A4.Thmtheorem6), the angular component𝐮\\mathbf\{u\}is independent of the client identityCC\. In particular,I\(𝐮,C\)=0I\(\\mathbf\{u\};C\)=0\. ###### Proof\. SinceKcK\_\{c\}is a deterministic function ofCC, for anyccwe compute p\(𝐮∣C=c\)\\displaystyle p\(\\mathbf\{u\}\\mid C=c\)=∫p\(ρ,𝐮∣C=c\)𝑑ρ\\displaystyle=\\int p\(\\rho,\\mathbf\{u\}\\mid C=c\)\\,d\\rho=∫p\(ρ,𝐮∣Kc\)𝑑ρ\\displaystyle=\\int p\(\\rho,\\mathbf\{u\}\\mid K\_\{c\}\)\\,d\\rho=∫pρ\(ρ∣Kc\)pΩ\(𝐮\)𝑑ρ\\displaystyle=\\int p\_\{\\rho\}\(\\rho\\mid K\_\{c\}\)\\,p\_\{\\Omega\}\(\\mathbf\{u\}\)\\,d\\rho\(Assumption[D\.6](https://arxiv.org/html/2608.21096#A4.Thmtheorem6)\)=pΩ\(𝐮\)∫pρ\(ρ∣Kc\)𝑑ρ\\displaystyle=p\_\{\\Omega\}\(\\mathbf\{u\}\)\\int p\_\{\\rho\}\(\\rho\\mid K\_\{c\}\)\\,d\\rho=pΩ\(𝐮\)\\displaystyle=p\_\{\\Omega\}\(\\mathbf\{u\}\)\(25\) Therefore the marginal satisfies p\(𝐮\)=∑cp\(C=c\)p\(𝐮∣C=c\)=∑cp\(C=c\)pΩ\(𝐮\)=pΩ\(𝐮\)\.p\(\\mathbf\{u\}\)=\\sum\_\{c\}p\(C=c\)\\,p\(\\mathbf\{u\}\\mid C=c\)=\\sum\_\{c\}p\(C=c\)\\,p\_\{\\Omega\}\(\\mathbf\{u\}\)=p\_\{\\Omega\}\(\\mathbf\{u\}\)\. Substituting into the KL formulation of mutual information, I\(𝐮,C\)\\displaystyle I\(\\mathbf\{u\};C\)=∑cp\(C=c\)DKL\(p\(𝐮∣C=c\)∥p\(𝐮\)\)\\displaystyle=\\sum\_\{c\}p\(C=c\)\\,D\_\{\\mathrm\{KL\}\}\\\!\\big\(p\(\\mathbf\{u\}\\mid C=c\)\\,\\\|\\,p\(\\mathbf\{u\}\)\\big\)=∑cp\(C=c\)∫Ωp\(𝐮∣C=c\)logp\(𝐮∣C=c\)p\(𝐮\)𝑑σ\(𝐮\)\\displaystyle=\\sum\_\{c\}p\(C=c\)\\,\\int\_\{\\Omega\}p\(\\mathbf\{u\}\\mid C=c\)\\,\\log\\frac\{p\(\\mathbf\{u\}\\mid C=c\)\}\{p\(\\mathbf\{u\}\)\}\\,d\\sigma\(\\mathbf\{u\}\)=∑cp\(C=c\)∫ΩpΩ\(𝐮\)logpΩ\(𝐮\)pΩ\(𝐮\)𝑑σ\(𝐮\)\\displaystyle=\\sum\_\{c\}p\(C=c\)\\,\\int\_\{\\Omega\}p\_\{\\Omega\}\(\\mathbf\{u\}\)\\,\\log\\frac\{p\_\{\\Omega\}\(\\mathbf\{u\}\)\}\{p\_\{\\Omega\}\(\\mathbf\{u\}\)\}\\,d\\sigma\(\\mathbf\{u\}\)=∑cp\(C=c\)DKL\(pΩ∥pΩ\)=0\.\\displaystyle=\\sum\_\{c\}p\(C=c\)\\,D\_\{\\mathrm\{KL\}\}\(p\_\{\\Omega\}\\,\\\|\\,p\_\{\\Omega\}\)=0\.\(26\) Thus the angular component carries no client information, which completes the proof\. ∎ ###### Lemma D\.8\. Letxt=TKc\(ρ\):=Kccosh\(ρ/Kc\)x\_\{t\}=T\_\{K\_\{c\}\}\(\\rho\):=\\sqrt\{K\_\{c\}\}\\cosh\(\\rho/\\sqrt\{K\_\{c\}\}\)withρ∼pρ\(⋅∣Kc\)\\rho\\sim p\_\{\\rho\}\(\\cdot\\mid K\_\{c\}\)\. Here\(TKc\)\#pρ\(⋅∣Kc\)\(T\_\{K\_\{c\}\}\)\_\{\\\#\}p\_\{\\rho\}\(\\cdot\\mid K\_\{c\}\)denotes the pushforward law ofpρ\(⋅∣Kc\)p\_\{\\rho\}\(\\cdot\\mid K\_\{c\}\)throughTKcT\_\{K\_\{c\}\}, i\.e\. the distribution ofxtx\_\{t\}\. If there existc1≠c2c\_\{1\}\\neq c\_\{2\}withℙ\(C=ci\)\>0\\mathbb\{P\}\(C=c\_\{i\}\)\>0such that\(TKc1\)\#pρ\(⋅∣Kc1\)≠\(TKc2\)\#pρ\(⋅∣Kc2\)\(T\_\{K\_\{c\_\{1\}\}\}\)\_\{\\\#\}p\_\{\\rho\}\(\\cdot\\mid K\_\{c\_\{1\}\}\)\\neq\(T\_\{K\_\{c\_\{2\}\}\}\)\_\{\\\#\}p\_\{\\rho\}\(\\cdot\\mid K\_\{c\_\{2\}\}\), thenxtx\_\{t\}is not independent ofCCand henceI\(xt,C\)\>0I\(x\_\{t\};C\)\>0\. ###### Proof\. Conditioned onC=cC=c, the Lorentz scale parameter is fixed toKcK\_\{c\}, and the law ofxtx\_\{t\}is the pushforward ofpρ\(⋅∣Kc\)p\_\{\\rho\}\(\\cdot\\mid K\_\{c\}\)underTKcT\_\{K\_\{c\}\}: for every Borel setA⊂ℝA\\subset\\mathbb\{R\}, ℙ\(xt∈A∣C=c\)\\displaystyle\\mathbb\{P\}\(x\_\{t\}\\in A\\mid C=c\)=ℙ\(TKc\(ρ\)∈A∣Kc\)\\displaystyle=\\mathbb\{P\}\\big\(T\_\{K\_\{c\}\}\(\\rho\)\\in A\\mid K\_\{c\}\\big\)=\[\(TKc\)\#pρ\(⋅∣Kc\)\]\(A\)\.\\displaystyle=\\big\[\(T\_\{K\_\{c\}\}\)\_\{\\\#\}p\_\{\\rho\}\(\\cdot\\mid K\_\{c\}\)\\big\]\(A\)\.\(27\)By the assumption of differing pushforward laws, there existc1≠c2c\_\{1\}\\neq c\_\{2\}withPr\(C=ci\)\>0\\Pr\(C=c\_\{i\}\)\>0such that p\(xt∣C=c1\)≠p\(xt∣C=c2\)\.p\(x\_\{t\}\\mid C=c\_\{1\}\)\\;\\neq\\;p\(x\_\{t\}\\mid C=c\_\{2\}\)\. Letp\(xt\)=∑cp\(C=c\)p\(xt∣C=c\)p\(x\_\{t\}\)=\\sum\_\{c\}p\(C=c\)\\,p\(x\_\{t\}\\mid C=c\)denote the marginal ofxtx\_\{t\}\. Using the KL expansion of mutual information, I\(xt,C\)\\displaystyle I\(x\_\{t\};C\)=∑cp\(C=c\)DKL\(p\(xt∣C=c\)∥p\(xt\)\)\\displaystyle=\\sum\_\{c\}p\(C=c\)\\,D\_\{\\mathrm\{KL\}\}\\\!\\big\(p\(x\_\{t\}\\mid C=c\)\\,\\\|\\,p\(x\_\{t\}\)\\big\)=∑cp\(C=c\)∫ℝp\(xt∣C=c\)logp\(xt∣C=c\)p\(xt\)dxt\.\\displaystyle=\\sum\_\{c\}p\(C=c\)\\,\\int\_\{\\mathbb\{R\}\}p\(x\_\{t\}\\mid C=c\)\\,\\log\\frac\{p\(x\_\{t\}\\mid C=c\)\}\{p\(x\_\{t\}\)\}\\,dx\_\{t\}\.\(28\)Ifp\(xt∣C=c1\)≠p\(xt∣C=c2\)p\(x\_\{t\}\\mid C=c\_\{1\}\)\\neq p\(x\_\{t\}\\mid C=c\_\{2\}\)and bothp\(C=ci\)\>0p\(C=c\_\{i\}\)\>0, then at least one ofp\(xt∣C=ci\)p\(x\_\{t\}\\mid C=c\_\{i\}\)differs from the mixturep\(xt\)p\(x\_\{t\}\); by Gibbs’ inequality, the corresponding KL term is strictly positive: DKL\(p\(xt∣C=ci\)∥p\(xt\)\)\>0for somei∈\{1,2\}\.D\_\{\\mathrm\{KL\}\}\\\!\\big\(p\(x\_\{t\}\\mid C=c\_\{i\}\)\\,\\\|\\,p\(x\_\{t\}\)\\big\)\\;\>\\;0\\quad\\text\{for some \}i\\in\\\{1,2\\\}\.Since all KL terms are nonnegative,[Equation \(28\)](https://arxiv.org/html/2608.21096#A4.E28)yieldsI\(xt,C\)\>0I\(x\_\{t\};C\)\>0\. Equivalently, from[Equation \(27\)](https://arxiv.org/html/2608.21096#A4.E27), there exists a Borel setAAsuch thatℙ\(xt∈A∣C=c1\)≠ℙ\(xt∈A∣C=c2\)\\mathbb\{P\}\(x\_\{t\}\\in A\\mid C=c\_\{1\}\)\\neq\\mathbb\{P\}\(x\_\{t\}\\in A\\mid C=c\_\{2\}\), which already rules out independence and thus forcesI\(xt,C\)\>0I\(x\_\{t\};C\)\>0\. This completes the proof\. ∎ ###### Lemma D\.9\. Under[AssumptionD\.6](https://arxiv.org/html/2608.21096#A4.Thmtheorem6), we haveI\(\(xt,𝐱s\);C∣Kc\)=I\(xt;C∣Kc\)\.I\\big\(\(x\_\{t\},\\mathbf\{x\}\_\{s\}\);C\\mid K\_\{c\}\\big\)=I\(x\_\{t\};C\\mid K\_\{c\}\)\. ###### Proof\. By[AssumptionD\.6](https://arxiv.org/html/2608.21096#A4.Thmtheorem6),𝐮\\mathbf\{u\}is independent of\(Kc,ρ\)\(K\_\{c\},\\rho\), hence of any measurable function thereof; in particular𝐮⟂\(xt,Kc,C\)\\mathbf\{u\}\\perp\(x\_\{t\},K\_\{c\},C\)\. Using the chain rule, I\(𝐱s;C∣xt,Kc\)\\displaystyle I\(\\mathbf\{x\}\_\{s\};C\\mid x\_\{t\},K\_\{c\}\)=I\(rs,𝐮;C∣xt,Kc\)\\displaystyle=I\(r\_\{s\},\\mathbf\{u\};C\\mid x\_\{t\},K\_\{c\}\)=I\(𝐮;C∣xt,Kc\)\+I\(rs;C∣xt,Kc,𝐮\)\\displaystyle=I\(\\mathbf\{u\};C\\mid x\_\{t\},K\_\{c\}\)\+I\(r\_\{s\};C\\mid x\_\{t\},K\_\{c\},\\mathbf\{u\}\)=0\+0\(since𝐮⟂\(xt,Kc,C\)andrsis deterministic given\(xt,Kc\)\)\\displaystyle=0\+0\\qquad\\text\{\(since $\\mathbf\{u\}\\perp\(x\_\{t\},K\_\{c\},C\)$ and $r\_\{s\}$ is deterministic given $\(x\_\{t\},K\_\{c\}\)$\)\}=0\.\\displaystyle=0\. Finally, apply the chain rule conditioned onKcK\_\{c\}: I\(\(xt,𝐱s\);C∣Kc\)=I\(xt;C∣Kc\)\+I\(𝐱s;C∣xt,Kc\)=I\(xt;C∣Kc\),I\\big\(\(x\_\{t\},\\mathbf\{x\}\_\{s\}\);C\\mid K\_\{c\}\\big\)=I\(x\_\{t\};C\\mid K\_\{c\}\)\+I\(\\mathbf\{x\}\_\{s\};C\\mid x\_\{t\},K\_\{c\}\)=I\(x\_\{t\};C\\mid K\_\{c\}\),which proves the stated conclusion\. ∎ ###### Theorem D\.11\(Recall of[Theorem 4\.3](https://arxiv.org/html/2608.21096#S4.Thmtheorem3)\)\. LetC∈\{1,…,m\}C\\in\\\{1,\\dots,m\\\}, each clientcchave Lorentz scale parameterKc\>0K\_\{c\}\>0withVar\(KC\)\>0\\mathrm\{Var\}\(K\_\{C\}\)\>0, and𝐱=\[xt𝐱s\]⊤∈𝕃Kcd\\mathbf\{x\}=\[x\_\{t\}\\;\\mathbf\{x\}\_\{s\}\]^\{\\top\}\\in\\mathbb\{L\}^\{d\}\_\{K\_\{c\}\}admit hyperbolic polar coordinates\(ρ,𝐮\)\(\\rho,\\mathbf\{u\}\)withxt=Kccosh\(ρ/Kc\)x\_\{t\}=\\sqrt\{K\_\{c\}\}\\cosh\(\\rho/\\sqrt\{K\_\{c\}\}\),𝐱s=Kcsinh\(ρ/Kc\)𝐮\\mathbf\{x\}\_\{s\}=\\sqrt\{K\_\{c\}\}\\sinh\(\\rho/\\sqrt\{K\_\{c\}\}\)\\,\\mathbf\{u\},𝐮:=𝐱s/‖𝐱s‖\\mathbf\{u\}:=\\mathbf\{x\}\_\{s\}/\\\|\\mathbf\{x\}\_\{s\}\\\|\. Then \(1\)I\(𝐮,C\)=0I\(\\mathbf\{u\};C\)=0; \(2\)I\(xt,C\)\>0I\(x\_\{t\};C\)\>0if the pushforward measures ofTK\(ρ\):=Kcosh\(ρ/K\)T\_\{K\}\(\\rho\):=\\sqrt\{K\}\\cosh\(\\rho/\\sqrt\{K\}\)differ acrossKK; \(3\)I\(\(xt,𝐱s\);C∣Kc\)=I\(xt;C∣Kc\)I\\big\(\(x\_\{t\},\\mathbf\{x\}\_\{s\}\);C\\mid K\_\{c\}\\big\)=I\(x\_\{t\};C\\mid K\_\{c\}\)\. ###### Proof\. The first claim follows from[LemmaD\.7](https://arxiv.org/html/2608.21096#A4.Thmtheorem7), the second from[LemmaD\.8](https://arxiv.org/html/2608.21096#A4.Thmtheorem8), and the third from[LemmaD\.9](https://arxiv.org/html/2608.21096#A4.Thmtheorem9)\. ∎ ##### Scope of the theoretical claims\. The mutual\-information statement above should be read under the stated isotropy and scale\-conditioned assumptions, rather than as a claim that the space\-like coordinates are strictly client\-invariant for every trained network or every graph distribution\. In practice, residual curvature\-dependent scaling can still appear in the space\-like magnitude\. Our theory is intended to justify the main design principle that the time\-like component provides a natural carrier for client\-specific geometric variation, while the empirical ablations validate the resulting parameter\-decoupling strategy\. ### D\.3Proof for Proposition[6\.1](https://arxiv.org/html/2608.21096#S6.Thmtheorem1) ###### Lemma D\.12\. LetℒKn\\mathcal\{L\}\_\{K\}^\{n\}denote thenn\-dimensional Lorentz space with constant sectional curvature−1/K\-1/K\. For any𝐱∈ℒKn\\mathbf\{x\}\\in\\mathcal\{L\}\_\{K\}^\{n\}and any transformation matrix𝐖∈ℝm×\(n\+1\)\\mathbf\{W\}\\in\\mathbb\{R\}^\{m\\times\(n\+1\)\}, the Lorentz transformationLT\\mathrm\{LT\}preserves the Lorentz structure, i\.e\.,LT\(𝐱,𝐖\)∈ℒKm\.\\mathrm\{LT\}\(\\mathbf\{x\};\\mathbf\{W\}\)\\in\\mathcal\{L\}\_\{K\}^\{m\}\. ###### Proof\. Let𝐱∈ℒKn\\mathbf\{x\}\\in\\mathcal\{L\}\_\{K\}^\{n\}\. By the definition of the Lorentz transformationLT\\mathrm\{LT\}, we compute the Lorentzian inner product: ⟨LT\(𝐱,𝐖\),LT\(𝐱,𝐖\)⟩ℒ=−K\.\\left\\langle\\mathrm\{LT\}\(\\mathbf\{x\};\\mathbf\{W\}\),\\mathrm\{LT\}\(\\mathbf\{x\};\\mathbf\{W\}\)\\right\\rangle\_\{\\mathcal\{L\}\}=\-K\.Since this condition characterizes membership in the Lorentz spaceℒKm\\mathcal\{L\}\_\{K\}^\{m\}, it follows thatLT\(𝐱,𝐖\)∈ℒKm\\mathrm\{LT\}\(\\mathbf\{x\};\\mathbf\{W\}\)\\in\\mathcal\{L\}\_\{K\}^\{m\}, which proves that the Lorentz transformation preserves membership in the target Lorentz space and completes the proof\. ∎ ###### Proposition D\.13\(Recall of Proposition[6\.1](https://arxiv.org/html/2608.21096#S6.Thmtheorem1)\)\. Let𝐌^=\[𝐦𝐌\]\\hat\{\\mathbf\{M\}\}=\\begin\{bmatrix\}\\mathbf\{m\}&\\mathbf\{M\}\\end\{bmatrix\}, where𝐌^∈ℝm×\(n\+1\)\\hat\{\\mathbf\{M\}\}\\in\\mathbb\{R\}^\{m\\times\(n\+1\)\},𝐦∈ℝm×1\\mathbf\{m\}\\in\\mathbb\{R\}^\{m\\times 1\}, and𝐌∈ℝm×n\\mathbf\{M\}\\in\\mathbb\{R\}^\{m\\times n\}\. LetΦ\(𝐌^,𝐍\)=\[𝐦𝐍\]\\Phi\\left\(\\hat\{\\mathbf\{M\}\},\\mathbf\{N\}\\right\)=\\begin\{bmatrix\}\\mathbf\{m\}&\\mathbf\{N\}\\end\{bmatrix\}, where𝐍∈ℝm×n\\mathbf\{N\}\\in\\mathbb\{R\}^\{m\\times n\}is the aggregated shared block obtained from\{𝐌i\}i=1C\\\{\\mathbf\{M\}\_\{i\}\\\}\_\{i=1\}^\{C\}using the proposed decoupling strategy\. For all𝐱∈ℒKn\\mathbf\{x\}\\in\\mathcal\{L\}\_\{K\}^\{n\}, we haveLT\(𝐱,Φ\(𝐌^,𝐍\)\)∈ℒKm\.\\mathrm\{LT\}\\left\(\\mathbf\{x\};\\Phi\\left\(\\hat\{\\mathbf\{M\}\},\\mathbf\{N\}\\right\)\\right\)\\in\\mathcal\{L\}\_\{K\}^\{m\}\. ###### Proof\. Let𝐱=\[xt𝐱s\]∈ℒKn\\mathbf\{x\}=\\begin\{bmatrix\}x\_\{t\}\\\\ \\mathbf\{x\}\_\{s\}\\end\{bmatrix\}\\in\\mathcal\{L\}\_\{K\}^\{n\}, wherext∈ℝ,𝐱s∈ℝnx\_\{t\}\\in\\mathbb\{R\},\\mathbf\{x\}\_\{s\}\\in\\mathbb\{R\}^\{n\}\. According to[Equation \(4\)](https://arxiv.org/html/2608.21096#S5.E4), we have: LT\(𝐱,Φ\(𝐌^,𝐍\)\)=\[‖𝐦xt\+𝐍𝐱s‖2\+K𝐦xt\+𝐍𝐱s\]\\mathrm\{LT\}\\left\(\\mathbf\{x\};\\Phi\(\\hat\{\\mathbf\{M\}\},\\mathbf\{N\}\)\\right\)=\\begin\{bmatrix\}\\sqrt\{\\\|\\mathbf\{m\}x\_\{t\}\+\\mathbf\{N\}\\mathbf\{x\}\_\{s\}\\\|^\{2\}\+K\}\\\\ \\mathbf\{m\}x\_\{t\}\+\\mathbf\{N\}\\mathbf\{x\}\_\{s\}\\end\{bmatrix\} We need to prove thatLT\(𝐱,Φ\(𝐌^,𝐍\)\)∈ℒKm\\mathrm\{LT\}\(\\mathbf\{x\};\\Phi\(\\hat\{\\mathbf\{M\}\},\\mathbf\{N\}\)\)\\in\\mathcal\{L\}\_\{K\}^\{m\}, i\.e\., to prove that it satisfies the definition condition of the Lorentz manifold⟨⋅,⋅⟩ℒ=−K\\langle\\cdot,\\cdot\\rangle\_\{\\mathcal\{L\}\}=\-K: ⟨LT\(𝐱,Φ\(𝐌^,𝐍\)\),LT\(𝐱,Φ\(𝐌^,𝐍\)\)⟩ℒ\\displaystyle\\left\\langle\\mathrm\{LT\}\\left\(\\mathbf\{x\};\\Phi\(\\hat\{\\mathbf\{M\}\},\\mathbf\{N\}\)\\right\),\\mathrm\{LT\}\\left\(\\mathbf\{x\};\\Phi\(\\hat\{\\mathbf\{M\}\},\\mathbf\{N\}\)\\right\)\\right\\rangle\_\{\\mathcal\{L\}\}=\\displaystyle=⟨\[‖𝐦xt\+𝐍𝐱s‖2\+K𝐦xt\+𝐍𝐱s\],\[‖𝐦xt\+𝐍𝐱s‖2\+K𝐦xt\+𝐍𝐱s\]⟩ℒ\([DefinitionB\.2](https://arxiv.org/html/2608.21096#A2.Thmtheorem2)\)\\displaystyle\\left\\langle\\begin\{bmatrix\}\\sqrt\{\\\|\\mathbf\{m\}x\_\{t\}\+\\mathbf\{N\}\\mathbf\{x\}\_\{s\}\\\|^\{2\}\+K\}\\\\ \\mathbf\{m\}x\_\{t\}\+\\mathbf\{N\}\\mathbf\{x\}\_\{s\}\\end\{bmatrix\},\\begin\{bmatrix\}\\sqrt\{\\\|\\mathbf\{m\}x\_\{t\}\+\\mathbf\{N\}\\mathbf\{x\}\_\{s\}\\\|^\{2\}\+K\}\\\\ \\mathbf\{m\}x\_\{t\}\+\\mathbf\{N\}\\mathbf\{x\}\_\{s\}\\end\{bmatrix\}\\right\\rangle\_\{\\mathcal\{L\}\}\\quad\\text\{\(\\hyperref@@ii\[def: inner product\]\{Definition~\\ref\*\{def: inner product\}\}\)\}=\\displaystyle=−\(‖𝐦xt\+𝐍𝐱s‖2\+K\)2\+‖𝐦xt\+𝐍𝐱s‖2\\displaystyle\-\\left\(\\sqrt\{\\\|\\mathbf\{m\}x\_\{t\}\+\\mathbf\{N\}\\mathbf\{x\}\_\{s\}\\\|^\{2\}\+K\}\\right\)^\{2\}\+\\\|\\mathbf\{m\}x\_\{t\}\+\\mathbf\{N\}\\mathbf\{x\}\_\{s\}\\\|^\{2\}=\\displaystyle=−K\\displaystyle\-K Therefore, we have proved thatLT\(𝐱,Φ\(𝐌^,𝐍\)\)∈ℒKm\\mathrm\{LT\}\\left\(\\mathbf\{x\};\\Phi\(\\hat\{\\mathbf\{M\}\},\\mathbf\{N\}\)\\right\)\\in\\mathcal\{L\}\_\{K\}^\{m\}\. ∎ ### D\.4Convergence Analysis FedAvg converges to the global optimum at a rate ofO\(1T\)O\(\\frac\{1\}\{T\}\)for strongly convex and smooth functions and non\-iid data\. When the learning rate is sufficiently small, the effect ofEEsteps of local updates is similar to a step update with a larger learning rate\([34](https://arxiv.org/html/2608.21096#bib.bib60)\)\. In this section, we demonstrate that𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}achieves a convergence rate ofO\(1T\)O\(\\frac\{1\}\{T\}\)without regularization, which is consistent with FedAvg\. Furthermore, when incorporating regularization similar to FedProx\([33](https://arxiv.org/html/2608.21096#bib.bib29)\), the convergence rate can be bounded by a constant that reflects the degree of data heterogeneity, analogous to FedProx’s theoretical guarantees\. This analysis confirms that our special geometric enhanced decoupling strategy maintains the overall convergence properties while addressing the challenges of heterogeneous data distribution\. To simplify the analysis, we consider each client conducts full batch gradient descent with one step\. At clientcc, the objective function can be generally written as minKc\>0,𝜽c,𝜽scℒc\(f\(𝐱Kc,𝜽c,𝜽sc\),y\)\+λ‖𝜽sc−𝜽¯s‖22,\\min\_\{K\_\{c\}\>0,\\bm\{\\theta\}\_\{c\},\\bm\{\\theta\}\_\{s\_\{c\}\}\}\\mathcal\{L\}\_\{c\}\(f\(\\mathbf\{x\}^\{K\_\{c\}\};\\bm\{\\theta\}\_\{c\},\\bm\{\\theta\}\_\{s\_\{c\}\}\),y\)\+\\lambda\\\|\\bm\{\\theta\}\_\{s\_\{c\}\}\-\\overline\{\\bm\{\\theta\}\}\_\{s\}\\\|\_\{2\}^\{2\},\(29\) whereλ\\lambdais a hyperparameter,y∈𝒴y\\in\\mathcal\{Y\},𝜽sc\\bm\{\\theta\}\_\{s\_\{c\}\}denotes clientcc’s local copy of the shared block, and‖𝜽sc−𝜽¯s‖22\\\|\\bm\{\\theta\}\_\{s\_\{c\}\}\-\\overline\{\\bm\{\\theta\}\}\_\{s\}\\\|\_\{2\}^\{2\}is the regularization term that prevents the locally updated model𝜽sc\\bm\{\\theta\}\_\{s\_\{c\}\}from deviating too far from the server\-shared parameters𝜽¯s\\overline\{\\bm\{\\theta\}\}\_\{s\}\. The optimization overKcK\_\{c\}is implemented through the raw\-scalar reparameterization described in[AppendixC\.2\.2](https://arxiv.org/html/2608.21096#A3.SS2.SSS2), andKcK\_\{c\}is kept local rather than aggregated by the server\. Letℓc=ℒc\(f\(𝐱Kc,𝜽c,𝜽sc\),y\)\\ell\_\{c\}=\\mathcal\{L\}\_\{c\}\(f\(\\mathbf\{x\}^\{K\_\{c\}\};\\bm\{\\theta\}\_\{c\},\\bm\{\\theta\}\_\{s\_\{c\}\}\),y\)\. Then the global loss is taken as an average of the loss of each client:ℓ=∑c∈𝒞pcℓc\\ell=\\sum\_\{c\\in\\mathcal\{C\}\}p\_\{c\}\\ell\_\{c\}, wherepc≥0p\_\{c\}\\geq 0and∑cpc=1\\sum\_\{c\}p\_\{c\}=1\. The local update is performed using vanilla gradient descent with a local learning rateη\\etain each client, and𝚯c\(r\)∈ℰ\\bm\{\\Theta\}\_\{c\}\(r\)\\in\\mathcal\{E\}represents the client\-local parameters\(Kc\(r\),𝜽c\(r\),𝜽sc\(r\)\)\(K\_\{c\}^\{\(r\)\},\\bm\{\\theta\}\_\{c\}^\{\(r\)\},\\bm\{\\theta\}\_\{s\_\{c\}\}^\{\(r\)\}\)in the roundrr\. Then, for global roundrr, Δ𝚯c\(r\)=𝚯c\(r\+1\)−𝚯c\(r\)=−η\(∇ℓc\(𝚯\(r\)\)\+2λ\(𝜽sc\(r\)−𝜽^s\(r\)\)\)\.\\Delta\\bm\{\\Theta\}\_\{c\}^\{\(r\)\}=\\bm\{\\Theta\}\_\{c\}^\{\(r\+1\)\}\-\\bm\{\\Theta\}\_\{c\}^\{\(r\)\}=\-\\eta\\left\(\\nabla\\ell\_\{c\}\(\\bm\{\\Theta\}^\{\(r\)\}\)\+2\\lambda\\left\(\\bm\{\\theta\}\_\{s\_\{c\}\}^\{\(r\)\}\-\\bm\{\\hat\{\\theta\}\}\_\{s\}^\{\(r\)\}\\right\)\\right\)\.Here the regularization gradient is applied only to the shared block𝜽sc\\bm\{\\theta\}\_\{s\_\{c\}\}; its components onKcK\_\{c\}and𝜽c\\bm\{\\theta\}\_\{c\}are zero\. To better calculate the difference between personalized parameters and shared parameters in the Lorentz linear block, we decompose the Lorentz\-layer part of𝚯c\(r\)\\bm\{\\Theta\}\_\{c\}^\{\(r\)\}as 𝚯LT,c\(r\)=𝜽c\(r\)\+𝜽sc\(r\),\\bm\{\\Theta\}\_\{\\mathrm\{LT\},c\}^\{\(r\)\}=\\bm\{\\theta\}\_\{c\}^\{\(r\)\}\+\\bm\{\\theta\}\_\{s\_\{c\}\}^\{\(r\)\},where𝜽c\(r\)=\[𝐦\(r\)𝟎m×n\]\\bm\{\\theta\}\_\{c\}^\{\(r\)\}=\[\\mathbf\{m\}^\{\(r\)\}\\quad\\mathbf\{0\}\_\{m\\times n\}\]and𝜽sc\(r\)=\[𝟎m×1𝐌\(r\)\]\\bm\{\\theta\}^\{\(r\)\}\_\{s\_\{c\}\}=\[\\mathbf\{0\}\_\{m\\times 1\}\\quad\\mathbf\{M\}^\{\(r\)\}\]\. Specifically, the global aggregation procedure is conducted by taking the average of local updates of shared parameters𝜽𝒔\\bm\{\\theta\_\{s\}\}of all\|𝒞\|\|\\mathcal\{C\}\|clients\. According to 𝜽s\(r\+1\)=𝜽¯s\(r\)=∑c∈𝒞\|𝒟c\|N𝜽sc\(r\)=∑c∈𝒞pc𝜽sc\(r\)\\bm\{\\theta\}\_\{s\}^\{\(r\+1\)\}=\\bm\{\\bar\{\\theta\}\}\_\{s\}^\{\(r\)\}=\\sum\_\{c\\in\\mathcal\{C\}\}\\frac\{\|\\mathcal\{D\}\_\{c\}\|\}\{N\}\\bm\{\\theta\}\_\{s\_\{c\}\}^\{\(r\)\}=\\sum\_\{c\\in\\mathcal\{C\}\}p\_\{c\}\\bm\{\\theta\}\_\{s\_\{c\}\}^\{\(r\)\} We next introduce standard non\-convex optimization assumptions commonly used in federated analyses\([34](https://arxiv.org/html/2608.21096#bib.bib60);[51](https://arxiv.org/html/2608.21096#bib.bib61)\), which provide the basis for the following descent bound\. ###### Assumption D\.14\(L\-smoothness\)\. ∀c∈𝒞ℓc\\forall\_\{c\\in\\mathcal\{C\}\}\\ell\_\{c\}areLL\-smooth: for all𝚯1∈𝔼\\bm\{\\Theta\}\_\{1\}\\in\\mathbb\{E\}and𝚯2∈𝔼\\bm\{\\Theta\}\_\{2\}\\in\\mathbb\{E\}, ℓc\(𝚯1\)≤ℓc\(𝚯2\)\+\(𝚯1−𝚯2\)T∇ℓc\(𝚯2\)\+L2∥𝚯1−𝚯2∥22\.\\ell\_\{c\}\(\\bm\{\\Theta\}\_\{1\}\)\\leq\\ell\_\{c\}\(\\bm\{\\Theta\}\_\{2\}\)\+\(\\bm\{\\Theta\}\_\{1\}\-\\bm\{\\Theta\}\_\{2\}\)^\{T\}\\nabla\\ell\_\{c\}\(\\bm\{\\Theta\}\_\{2\}\)\+\\frac\{L\}\{2\}\\\|\\bm\{\\Theta\}\_\{1\}\-\\bm\{\\Theta\}\_\{2\}\\\|\_\{2\}^\{2\}\. ###### Assumption D\.15\(Bounded Gradients\)\. The functionℓc\(𝚯\)\\ell\_\{c\}\(\\bm\{\\Theta\}\)haveGG\-bounded gradients, i\.e\., for anyc∈𝒞c\\in\\mathcal\{C\},𝚯∈ℝd\\bm\{\\Theta\}\\in\\mathbb\{R\}^\{d\}we have∥∇ℓc\(𝚯\)∥≤G\\lVert\\nabla\\ell\_\{c\}\(\\bm\{\\Theta\}\)\\rVert\\leq G\. ###### Lemma D\.16\(Smooth Descent Lemma\)\. Letℓ:ℰ→ℝ\\ell:\\mathcal\{E\}\\rightarrow\\mathbb\{R\}be an L\-smooth function\. Then for any𝚯\(r\),𝚯\(r\+1\)∈𝔼\\bm\{\\Theta\}^\{\(r\)\},\\bm\{\\Theta\}^\{\(r\+1\)\}\\in\\mathbb\{E\}, the following inequality holds: ℓ\(𝚯\(r\+1\)\)≤ℓ\(𝚯\(r\)\)\+⟨∇ℓ\(𝚯\(r\)\),Δ𝚯\(r\)⟩\+L2‖Δ𝚯\(r\)‖2\.\\ell\(\\bm\{\\Theta\}^\{\(r\+1\)\}\)\\leq\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\+\\langle\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\),\\Delta\\bm\{\\Theta\}^\{\(r\)\}\\rangle\+\\frac\{L\}\{2\}\\\|\\Delta\\bm\{\\Theta\}^\{\(r\)\}\\\|^\{2\}\. Letδ\(r\)=2λ∑c∈𝒞\|𝒟c\|N\(𝜽sc−𝜽¯𝒔\)\\delta^\{\(r\)\}=2\\lambda\\sum\_\{c\\in\\mathcal\{C\}\}\\frac\{\|\\mathcal\{D\}\_\{c\}\|\}\{N\}\\left\(\\bm\{\\theta\}\_\{s\_\{c\}\}\-\\bm\{\\bar\{\\theta\}\_\{s\}\}\\right\)\. Based on Lemma 1, we have ℓ\(𝚯\(r\+1\)\)\\displaystyle\\ell\(\\bm\{\\Theta\}^\{\(r\+1\)\}\)≤ℓ\(𝚯\(r\)\)\+⟨∇ℓ\(𝚯\(r\)\),Δ𝚯\(r\)⟩\+L2‖Δ𝚯\(r\)‖2\\displaystyle\\leq\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\+\\langle\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\),\\Delta\\bm\{\\Theta\}^\{\(r\)\}\\rangle\+\\frac\{L\}\{2\}\\\|\\Delta\\bm\{\\Theta\}^\{\(r\)\}\\\|^\{2\}=ℓ\(𝚯\(r\)\)\+⟨∇ℓ\(𝚯\(r\)\),−η\(∇ℓ\(𝚯\(r\)\)\+δ\(r\)\)⟩\\displaystyle=\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\+\\left\\langle\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\),\-\\eta\\left\(\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\+\\delta^\{\(r\)\}\\right\)\\right\\rangle\+Lη22‖∇ℓ\(𝚯\(r\)\)\+δ\(r\)‖2\\displaystyle\+\\frac\{L\\eta^\{2\}\}\{2\}\\\|\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\+\\delta^\{\(r\)\}\\\|^\{2\}=ℓ\(𝚯\(r\)\)−η⟨∇ℓ\(𝚯\(r\)\),∇ℓ\(𝚯\(r\)\)\+δ\(r\)⟩\\displaystyle=\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\-\\eta\\left\\langle\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\),\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\+\\delta^\{\(r\)\}\\right\\rangle\+Lη22‖∇ℓ\(𝚯\(r\)\)\+δ\(r\)‖2\\displaystyle\+\\frac\{L\\eta^\{2\}\}\{2\}\\\|\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\+\\delta^\{\(r\)\}\\\|^\{2\}=ℓ\(𝚯\(r\)\)−η‖∇ℓ\(𝚯\(r\)\)‖2−η⟨∇ℓ\(𝚯\(r\)\),δ\(r\)⟩\\displaystyle=\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\-\\eta\\\|\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\\\|^\{2\}\-\\eta\\left\\langle\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\),\\delta^\{\(r\)\}\\right\\rangle\+Lη22‖∇ℓ\(𝚯\(r\)\)‖2\+Lη2⟨∇ℓ\(𝚯\(r\),δ\(r\)\)⟩\+Lη22‖δ\(r\)‖2\\displaystyle\+\\frac\{L\\eta^\{2\}\}\{2\}\\\|\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\\\|^\{2\}\+L\\eta^\{2\}\\langle\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\},\\delta^\{\(r\)\}\)\\rangle\+\\frac\{L\\eta^\{2\}\}\{2\}\\\|\\delta^\{\(r\)\}\\\|^\{2\}=ℓ\(𝚯\(r\)\)\+\(Lη22−η\)‖∇ℓ\(𝚯\(r\)\)‖2\+Lη22‖δ\(r\)‖2\\displaystyle=\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\+\(\\frac\{L\\eta^\{2\}\}\{2\}\-\\eta\)\\\|\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\\\|^\{2\}\+\\frac\{L\\eta^\{2\}\}\{2\}\\\|\\delta^\{\(r\)\}\\\|^\{2\}\+\(Lη2−η\)⟨∇ℓ\(𝚯\(r\)\),δ\(r\)⟩\\displaystyle\+\(L\\eta^\{2\}\-\\eta\)\\left\\langle\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\),\\delta^\{\(r\)\}\\right\\rangle=ℓ\(𝚯\(r\)\)\+\(Lη22−η\)‖∇ℓ\(𝚯\(r\)\)‖2\+Lη22‖δ\(r\)‖2\\displaystyle=\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\+\(\\frac\{L\\eta^\{2\}\}\{2\}\-\\eta\)\\\|\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\\\|^\{2\}\+\\frac\{L\\eta^\{2\}\}\{2\}\\\|\\delta^\{\(r\)\}\\\|^\{2\}\+Lη2−η2\(‖∇ℓ\(𝚯\(r\)\)‖2\+‖δ\(r\)‖2−‖∇ℓ\(𝚯\(r\)\)\+δ\(r\)‖2\)\\displaystyle\+\\frac\{L\\eta^\{2\}\-\\eta\}\{2\}\\left\(\\\|\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\\\|^\{2\}\+\\\|\\delta^\{\(r\)\}\\\|^\{2\}\-\\\|\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\+\\delta^\{\(r\)\}\\\|^\{2\}\\right\)=ℓ\(𝚯\(r\)\)\+\(Lη2−3η2\)‖∇ℓ\(𝚯\(r\)\)‖2\+\(Lη2−η2\)‖δ\(r\)‖2\\displaystyle=\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\+\(L\\eta^\{2\}\-\\frac\{3\\eta\}\{2\}\)\\\|\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\\\|^\{2\}\+\(L\\eta^\{2\}\-\\frac\{\\eta\}\{2\}\)\\\|\\delta^\{\(r\)\}\\\|^\{2\}−Lη2−η2‖∇ℓ\(𝚯\(r\)\)\+δ\(r\)‖2\\displaystyle\-\\frac\{L\\eta^\{2\}\-\\eta\}\{2\}\\\|\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\+\\delta^\{\(r\)\}\\\|^\{2\}\(30\) We selectη=1L\\eta=\\frac\{1\}\{L\}, so we have ℓ\(𝚯\(r\+1\)\)≤ℓ\(𝚯\(r\)\)−12L‖∇ℓ\(𝚯\(r\)\)‖2\+12L‖δ\(r\)‖2\\ell\(\\bm\{\\Theta\}^\{\(r\+1\)\}\)\\leq\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\-\\frac\{1\}\{2L\}\\\|\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\\\|^\{2\}\+\\frac\{1\}\{2L\}\\\|\\delta^\{\(r\)\}\\\|^\{2\}\(31\) Rearranging the above inequality gives ‖∇ℓ\(𝚯\(r\)\)‖2≤2L\(ℓ\(𝚯\(r\)\)−ℓ\(𝚯\(r\+1\)\)\)\+‖δ\(r\)‖2\\\|\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\\\|^\{2\}\\leq 2L\\left\(\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\-\\ell\(\\bm\{\\Theta\}^\{\(r\+1\)\}\)\\right\)\+\\\|\\delta^\{\(r\)\}\\\|^\{2\}\(32\) Then, summingrrfrom11toTT, we have minr∈\[T\]‖∇ℓ\(𝚯\(r\)\)‖2≤2L\(ℓ\(𝚯\(1\)\)−ℓ\(𝚯\(T\+1\)\)\)T\+1T∑r∈\[T\]‖δ\(r\)‖2\\min\_\{r\\in\[T\]\}\\\|\\nabla\\ell\(\\bm\{\\Theta\}^\{\(r\)\}\)\\\|^\{2\}\\leq\\frac\{2L\\left\(\\ell\(\\bm\{\\Theta\}^\{\(1\)\}\)\-\\ell\(\\bm\{\\Theta\}^\{\(T\+1\)\}\)\\right\)\}\{T\}\+\\frac\{1\}\{T\}\\sum\_\{r\\in\[T\]\}\\\|\\delta^\{\(r\)\}\\\|^\{2\}\(33\) ###### Definition D\.17\(BB\-local dissimilarity\)\. The local functionsℓc\\ell\_\{c\}areBB\-locally dissimilar at𝚯\\bm\{\\Theta\}if 𝔼c\[‖∇ℓc\(𝚯\)‖2\]≤‖∇ℓ\(𝚯\)‖2B2\.\\mathbb\{E\}\_\{c\}\[\\\|\\nabla\\ell\_\{c\}\(\\bm\{\\Theta\}\)\\\|^\{2\}\]\\leq\\\|\\nabla\\ell\(\\bm\{\\Theta\}\)\\\|^\{2\}B^\{2\}\.We further defineB\(𝚯\)=𝔼c\[‖∇ℓc\(𝚯\)‖2\]‖∇ℓ\(𝚯\)‖2B\(\\bm\{\\Theta\}\)=\\sqrt\{\\frac\{\\mathbb\{E\}\_\{c\}\[\\\|\\nabla\\ell\_\{c\}\(\\bm\{\\Theta\}\)\\\|^\{2\}\]\}\{\\\|\\nabla\\ell\(\\bm\{\\Theta\}\)\\\|^\{2\}\}\}for‖∇ℓ\(𝚯\)‖≠0\\\|\\nabla\\ell\(\\bm\{\\Theta\}\)\\\|\\neq 0\. ###### Definition D\.18\(γ\\gamma\-inexact solution\)\. For a functionh\(w,w0\)=F\(w\)\+λ‖w−w0‖2,andγ∈\[0,1\],h\(w;w\_\{0\}\)=F\(w\)\+\\lambda\\\|w\-w\_\{0\}\\\|^\{2\},\\quad\\text\{and\}\\quad\\gamma\\in\[0,1\],we sayw∗w^\{\*\}is aγ\\gamma\-inexact solution ofminwh\(w,w0\)\\min\_\{w\}h\(w;w\_\{0\}\)if‖∇h\(w∗,w0\)‖≤γ‖∇h\(w0,w0\)‖,\\\|\\nabla h\(w^\{\*\};w\_\{0\}\)\\\|\\leq\\gamma\\\|\\nabla h\(w\_\{0\};w\_\{0\}\)\\\|,where∇h\(w,w0\)=∇F\(w\)\+μ\(w−w0\),\\nabla h\(w;w\_\{0\}\)=\\nabla F\(w\)\+\\mu\(w\-w\_\{0\}\),where,μ=2λ\\mu=2\\lambda\. Note that smallerγ\\gammacorresponds to higher accuracy\. Using the notion ofγ\\gamma\-inexactness for each local client, we can defineec\(r\)e\_\{c\}^\{\(r\)\}such that ∇ℓc\(𝚯c\(r\+1\)\)\+μ\(𝜽^s\(r\)−𝜽sc\(r\)\)\+μ\(𝜽c\(r\+1\)−𝜽c\(r\)\)−ec\(r\)=0,‖ec\(r\)‖≤γ‖∇ℓc\(𝚯c\(r\)\)‖\.\\begin\{array\}\[\]\{l\}\\nabla\\ell\_\{c\}\\left\(\\bm\{\\Theta\}\_\{c\}^\{\(r\+1\)\}\\right\)\+\\mu\\left\(\\bm\{\\hat\{\\theta\}\}\_\{s\}^\{\(r\)\}\-\\bm\{\\theta\}\_\{s\_\{c\}\}^\{\(r\)\}\\right\)\+\\mu\\left\(\\bm\{\\theta\}\_\{c\}^\{\(r\+1\)\}\-\\bm\{\\theta\}\_\{c\}^\{\(r\)\}\\right\)\-e\_\{c\}^\{\(r\)\}=0,\\\\ \\\|e\_\{c\}^\{\(r\)\}\\\|\\leq\\gamma\\\|\\nabla\\ell\_\{c\}\\left\(\\bm\{\\Theta\}\_\{c\}^\{\(r\)\}\\right\)\\\|\.\\end\{array\}\(34\) Then we have 𝜽s\(r\+1\)−𝜽s\(r\)=−1μ𝔼c\[∇ℓc\(𝚯c\(r\)\)\]\+1μ𝔼c\[ec\(r\)\]−𝔼c\[Δ𝜽c\(r\)\],\\bm\{\\theta\}\_\{s\}^\{\(r\+1\)\}\-\\bm\{\\theta\}\_\{s\}^\{\(r\)\}=\\frac\{\-1\}\{\\mu\}\\mathbb\{E\}\_\{c\}\\left\[\\nabla\\ell\_\{c\}\\left\(\\bm\{\\Theta\}\_\{c\}^\{\(r\)\}\\right\)\\right\]\+\\frac\{1\}\{\\mu\}\\mathbb\{E\}\_\{c\}\[e\_\{c\}^\{\(r\)\}\]\-\\mathbb\{E\}\_\{c\}\\left\[\\Delta\\bm\{\\theta\}\_\{c\}^\{\(r\)\}\\right\],\(35\) According to\([33](https://arxiv.org/html/2608.21096#bib.bib29)\)and triangle inequality, when a regularization is incorporated, \(λ\>0\\lambda\>0\), we have 14λ2‖δ\(r\)‖2\\displaystyle\\frac\{1\}\{4\\lambda^\{2\}\}\\\|\\delta^\{\(r\)\}\\\|^\{2\}≤\(𝔼c\[‖𝜽s\(r\+1\)−𝜽sc\(r\)‖\]\)2\\displaystyle\\leq\\left\(\\mathbb\{E\}\_\{c\}\\left\[\\\|\\bm\{\\theta\}\_\{s\}^\{\(r\+1\)\}\-\\bm\{\\theta\}\_\{s\_\{c\}\}^\{\(r\)\}\\\|\\right\]\\right\)^\{2\}≤\(1\+γμ¯\)2\(𝔼c\[‖∇ℓc\(𝚯c\(r\)\)−Δ𝜽c\(r\)‖\]\)2\\displaystyle\\leq\\left\(\\frac\{1\+\\gamma\}\{\\bar\{\\mu\}\}\\right\)^\{2\}\\left\(\\mathbb\{E\}\_\{c\}\\left\[\\\|\\nabla\\ell\_\{c\}\\left\(\\bm\{\\Theta\}\_\{c\}^\{\(r\)\}\\right\)\-\\Delta\\bm\{\\theta\}\_\{c\}^\{\(r\)\}\\\|\\right\]\\right\)^\{2\}≤\(1\+γμ¯\)2\(𝔼c\[‖∇ℓc\(𝚯c\(r\)\)−Δ𝜽c\(r\)‖2\]\)\\displaystyle\\leq\\left\(\\frac\{1\+\\gamma\}\{\\bar\{\\mu\}\}\\right\)^\{2\}\\left\(\\mathbb\{E\}\_\{c\}\\left\[\\\|\\nabla\\ell\_\{c\}\\left\(\\bm\{\\Theta\}\_\{c\}^\{\(r\)\}\\right\)\-\\Delta\\bm\{\\theta\}\_\{c\}^\{\(r\)\}\\\|^\{2\}\\right\]\\right\)≤B2\(1\+γ\)2μ¯2𝔼\[‖∇ℓc\(𝚯c\(r\)\)‖2\]\+C,\\displaystyle\\leq\\frac\{B^\{2\}\(1\+\\gamma\)^\{2\}\}\{\\bar\{\\mu\}^\{2\}\}\\mathbb\{E\}\\left\[\\\|\\nabla\\ell\_\{c\}\\left\(\\bm\{\\Theta\}\_\{c\}^\{\(r\)\}\\right\)\\\|^\{2\}\\right\]\+C, Based on the assumption of the bounded gradients \(Assumption[D\.15](https://arxiv.org/html/2608.21096#A4.Thmtheorem15)\), we find that theδ\(r\)\\delta^\{\(r\)\}is also bounded\. Specifically,C=\(1\+γμ¯\)2𝔼c\[‖Δ𝜽c‖2\]≈\(1\+γμ¯\)2𝔼\[‖ΔMc‖2\]C=\\left\(\\frac\{1\+\\gamma\}\{\\bar\{\\mu\}\}\\right\)^\{2\}\\mathbb\{E\}\_\{c\}\[\\\|\\Delta\\bm\{\\theta\}\_\{c\}\\\|^\{2\}\]\\approx\\left\(\\frac\{1\+\\gamma\}\{\\bar\{\\mu\}\}\\right\)^\{2\}\\mathbb\{E\}\[\\\|\\Delta M\_\{c\}\\\|^\{2\}\]\.∥δ\(r\)∥2\\lVert\\delta^\{\(r\)\}\\rVert^\{2\}measures the degree of data heterogeneity\. Overall, whenλ=0\\lambda=0, the termδ\(r\)=0\\delta^\{\(r\)\}=0, eliminating the impact of data heterogeneity and resulting in a convergence rate ofO\(1T\)O\\left\(\\frac\{1\}\{T\}\\right\), consistent with FedAvg\. And when incorporating regularization \(λ\>0\\lambda\>0\), we establish that‖δ\(r\)‖2\\left\\\|\\delta^\{\(r\)\}\\right\\\|^\{2\}is bounded, analogous to the theoretical guarantees provided by FedProx\([33](https://arxiv.org/html/2608.21096#bib.bib29)\)\. This analysis suggests that𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}can balance the trade\-off between computational overhead and model effectiveness as the number of clients increases\. While it introduces additional operations in local computation, these overheads are limited and offer optimization opportunities through pre\-computation and caching\. The method compensates for these costs through reduced communication overhead and enhanced representation capability in Lorentz space, making it a practical choice for personalized federated learning\. ### D\.5Perspectives on Lorentz Transformations Lorentz Boosts and Lorentz Rotations \([AppendixB\.3](https://arxiv.org/html/2608.21096#A2.SS3)\) are understood as transformations that are encapsulated byLT\(𝐱,𝐌^\)\\mathrm\{LT\}\\left\(\\mathbf\{x\};\\hat\{\\mathbf\{M\}\}\\right\)when the dimension is unchanged\([9](https://arxiv.org/html/2608.21096#bib.bib27)\)\. We can easily prove that the Lorentz transformations are still covered byLT\(⋅,Φ\(𝐌^,𝐍\)\)\\mathrm\{LT\}\\left\(\\cdot;\\Phi\\left\(\\hat\{\\mathbf\{M\}\},\\mathbf\{N\}\\right\)\\right\), where the corresponding Lorentz\-linear block is𝐌^=\[𝐦𝐌\]∈ℝn×\(n\+1\)\\hat\{\\mathbf\{M\}\}=\[\\mathbf\{m\}\\ \\mathbf\{M\}\]\\in\\mathbb\{R\}^\{n\\times\(n\+1\)\}and the replacement block satisfies𝐍∈ℝn×n\\mathbf\{N\}\\in\\mathbb\{R\}^\{n\\times n\}\. For any data point𝐱∈𝒟c\\mathbf\{x\}\\in\\mathcal\{D\}\_\{c\}, transformationsLT\(𝐱,𝐌^\)\\mathrm\{LT\}\\left\(\\mathbf\{x\};\\hat\{\\mathbf\{M\}\}\\right\)andLT\(𝐱,Φ\(𝐌^,𝐍\)\)\\mathrm\{LT\}\\left\(\\mathbf\{x\};\\Phi\\left\(\\hat\{\\mathbf\{M\}\},\\mathbf\{N\}\\right\)\\right\)map𝐱\\mathbf\{x\}to new Lorentz\-model coordinates while preserving the Lorentzian norm constraint \([Proposition6\.1](https://arxiv.org/html/2608.21096#S6.Thmtheorem1)\)\. Thus, replacing the shared block after aggregation keeps the representation on the same client\-specific hyperboloid\. Clients with different Lorentz scale parametersKcK\_\{c\}therefore remain in distinct Lorentz spaces, reflecting differences in their underlying data distributions\. Moreover, according to the definition of Lorentz rotation in[Equation \(8\)](https://arxiv.org/html/2608.21096#A2.E8), the server updates only𝐌\\mathbf\{M\}while leaving the time\-like component local\. This operation is a relaxation of a Lorentz rotation, consistent with our "Flatland" assumption that aggregates only space\-like information\. ## Appendix EExperimental Supplementary ### E\.1Datasets For federated node classification, we adopt four benchmark datasets constructed by\([5](https://arxiv.org/html/2608.21096#bib.bib51)\): Cora, CiteSeer, ogbn\-arxiv, and Photo\([52](https://arxiv.org/html/2608.21096#bib.bib55);[23](https://arxiv.org/html/2608.21096#bib.bib56);[53](https://arxiv.org/html/2608.21096#bib.bib57)\)\. Cora, CiteSeer, and ogbn\-arxiv are citation graphs\. Photo is a product graph\. Each graph dataset is divided into a certain number of disjoint subgraphs using the METIS graph partitioning algorithm\([27](https://arxiv.org/html/2608.21096#bib.bib12)\), where each subgraph belongs to an FL client\. Statistics of datasets are summarized in[Table 6](https://arxiv.org/html/2608.21096#A5.T6)\. For federated graph classification, we consider the non\-IID settings proposed by\([61](https://arxiv.org/html/2608.21096#bib.bib59)\)\. In total, there are 13 graph classification datasets from three different domains, including small molecules \(MUTAG, BZR, COX2, DHFR, PTC\_MR, AIDS, NCI1\) denoted as CHEM, bioinformatics \(ENZYMES, DD, PROTEINS\) denoted as BIO, and social networks \(COLLAB, IMDB\-BINARY, IMDB\-MULTI\)\([47](https://arxiv.org/html/2608.21096#bib.bib58)\)denoted as SN\. To simulate data heterogeneity, two non\-IID settings are constructed: \(1\) a cross\-dataset setting based on the small molecule datasets \(CHEM\), \(2\) a cross\-domain setting based on all datasets \(BIO\-CHEM\-SN\)\. In each setting, one dataset corresponds to one FL client\. Statistics of datasets are summarized in[Table 5](https://arxiv.org/html/2608.21096#A5.T5)and[Table 7](https://arxiv.org/html/2608.21096#A5.T7)\. Table 5:Statistics of graph classification datasets\. We report the \(average\) number of graphs, nodes, edges, classes, and node features of each dataset\.Table 6:Statistics of homophilic node classification datasets\. We report the \(average\) number of nodes, edges, classes, clustering coefficient, and heterogeneity for different numbers of clients\.Table 7:Statistics of heterophilic node classification datasets\. We report the \(average\) number of nodes, edges, and classes\. ### E\.2Implementation Details ##### Implementation of learnable curvature\. For each clientcc, we optimize a raw scalarκc\\kappa\_\{c\}and map it to a positive Lorentz scale parameterKc=sigmoid\(κc\)\+0\.5K\_\{c\}=\\operatorname\{sigmoid\}\(\\kappa\_\{c\}\)\+0\.5\. This keepsKcK\_\{c\}in the effective range\[0\.5,1\.5\]\[0\.5,1\.5\], so the corresponding sectional curvature−1/Kc\-1/K\_\{c\}remains negative while numerical optimization stays stable\([9](https://arxiv.org/html/2608.21096#bib.bib27)\)\. The reparameterization allows the corresponding curvature to adapt to heterogeneous client data during local training\. ##### Implementation of node classification and graph classification tasks\. For node classification, we use a 2\-layer GCN\([29](https://arxiv.org/html/2608.21096#bib.bib13)\)for Euclidean models, a 2\-layer LGCN\([9](https://arxiv.org/html/2608.21096#bib.bib27)\)for𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}, and HGCN with node selection for FedHGCN\([14](https://arxiv.org/html/2608.21096#bib.bib52)\)\. LGCN serves as the backbone for our graph learning framework, combining Lorentz linear layers with graph aggregation operations, similar to how Euclidean counterparts such as GCN and GIN integrate linear layers with graph aggregation\. Each layer applies a Lorentz transformation followed by neighbor aggregation using the adjacency matrix to obtain node representations\. We conduct 100 rounds for Cora/CiteSeer and 200 rounds for larger datasets such as Photo and ogbn\-arxiv, with 1–3 local epochs and 128\-dimensional hidden layers\. For graph classification, we use a 3\-layer GIN\([62](https://arxiv.org/html/2608.21096#bib.bib2)\)as the Euclidean encoder and the same 3\-layer hyperbolic encoders as in node classification for hyperbolic models, with 1 local epoch and 200 rounds\. The learning rate is chosen from\{0\.01,0\.001\}\\\{0\.01,0\.001\\\}, and the weight decay is10−510^\{\-5\}\. We optimize with Adam and report node\- or graph\-level accuracy averaged across clients\. All experiments are implemented in Python 3\.10 and PyTorch and run on an RTX A6000 GPU\. Each client is allocated a worker; one node\-classification local epoch takes roughly one second per round\. ##### Backbone structure\. Inspired by recent GNN models\([8](https://arxiv.org/html/2608.21096#bib.bib70)\)that highlight the benefits of incorporating high\-pass information for heterophilic graphs, all hyperbolic baselines in our framework combine information from both low\-pass \(adjacency\-based\) and high\-pass \(Laplacian\-based\) operations through a learnable gating mechanism\. For a fair comparison, we also applied this design to the Euclidean baselines, but it did not yield improvements over the hyperbolic counterparts, suggesting that hyperbolic models benefit more than Euclidean baselines from jointly using low\-pass and high\-pass information for heterophilic graph representation\. ### E\.3Baseline Selection and Comparison We provide clarification on our baseline selection to ensure fair and comprehensive comparison: ##### Euclidean baselines\. We compare against representative methods from each major PFL category: \(1\)basic FL: FedAvg\([45](https://arxiv.org/html/2608.21096#bib.bib28)\); \(2\)regularization\-based: FedProx\([33](https://arxiv.org/html/2608.21096#bib.bib29)\); \(3\)parameter decoupling: FedPer\([3](https://arxiv.org/html/2608.21096#bib.bib42)\); and \(4\)client clustering: GCFL\([61](https://arxiv.org/html/2608.21096#bib.bib59)\)\. These baselines span the spectrum of PFL approaches and use the same GNN backbone \(GCN/GIN\) as our method for fair comparison\. ##### Hyperbolic baseline\. FedHGCN\([14](https://arxiv.org/html/2608.21096#bib.bib52)\)is the most relevant hyperbolic FL baseline, combining FedAvg with hyperbolic GNNs\. Unlike𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}, FedHGCN uses a fixed global curvature and lacks personalization capability, making it a direct comparison point for demonstrating the benefits of tailored curvature and parameter decoupling\. ##### Local training baselines\. We include Local \(EE\) and Local \(LL\) to show the performance of purely local training in Euclidean and Lorentz spaces respectively, without any federated aggregation\. This helps isolate the contribution of federated learning from the geometric representation\. ##### Why not other hyperbolic models? Methods like HyperFed\([38](https://arxiv.org/html/2608.21096#bib.bib53)\)and FedMRUR\([2](https://arxiv.org/html/2608.21096#bib.bib54)\)focus on different aspects \(prototype learning and knowledge distillation\) rather than personalized geometric modeling\. Our comparison therefore focuses on methods that directly address the heterogeneity challenge through model architecture or aggregation designs that explicitly handle client heterogeneity\. ### E\.4Unified Summary of Ablations The main text and appendices evaluate the components of𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}from different angles\. We collect these results in[Table 8](https://arxiv.org/html/2608.21096#A5.T8)to clarify which design choice each ablation isolates\. Table 8:Unified ablation summary on Cora and CiteSeer with 20 clients\. Results are reported as mean accuracy\.GroupVariantCoraCiteSeerSourceFull model𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}82\.4972\.24Sec\.[7\.2](https://arxiv.org/html/2608.21096#S7.SS2),[Table 2](https://arxiv.org/html/2608.21096#S6.T2)RepresentationLocal \(EE\)80\.3065\.98Sec\.[7\.2](https://arxiv.org/html/2608.21096#S7.SS2),[Table 2](https://arxiv.org/html/2608.21096#S6.T2)RepresentationLocal \(LL\)80\.4669\.52Sec\.[7\.2](https://arxiv.org/html/2608.21096#S7.SS2),[Table 2](https://arxiv.org/html/2608.21096#S6.T2)Decoupling𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}\(EE\)76\.2366\.29Sec\.[7\.5](https://arxiv.org/html/2608.21096#S7.SS5),[Table 4](https://arxiv.org/html/2608.21096#S7.T4)Decouplingw/o DS71\.7663\.08Sec\.[7\.5](https://arxiv.org/html/2608.21096#S7.SS5),[Figure 5](https://arxiv.org/html/2608.21096#S6.F5)Tailored curvaturew/o TS78\.8367\.93Sec\.[7\.5](https://arxiv.org/html/2608.21096#S7.SS5),[Figure 5](https://arxiv.org/html/2608.21096#S6.F5)Curvature initializationRicci82\.4972\.24App\.[E\.5](https://arxiv.org/html/2608.21096#A5.SS5),[Table 9](https://arxiv.org/html/2608.21096#A5.T9)Curvature initializationConstant81\.9171\.89App\.[E\.5](https://arxiv.org/html/2608.21096#A5.SS5),[Table 9](https://arxiv.org/html/2608.21096#A5.T9)Curvature initializationOllivier82\.5172\.21App\.[E\.5](https://arxiv.org/html/2608.21096#A5.SS5),[Table 9](https://arxiv.org/html/2608.21096#A5.T9)Curvature initializationMLP82\.3372\.59App\.[E\.5](https://arxiv.org/html/2608.21096#A5.SS5),[Table 9](https://arxiv.org/html/2608.21096#A5.T9)Overall, the gains do not come from a single factor\. Local \(LL\) versus Local \(EE\) shows that Lorentz representations are beneficial for many graph clients, while the gap between𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}and w/o DS highlights the importance of separating personalized time\-like parameters from shared space\-like parameters\. The drop of w/o TS further indicates that client\-specific curvature is useful, whereas the similar performance across curvature initializers suggests that learnable curvature matters more than a particular initialization rule\. ### E\.5Impact of Curvature Initialization We further investigate the impact of different curvature initialization strategies on model performance\. Specifically, we compare four initialization methods: \(1\)Forman\-Ricci curvature, \(2\)Ollivier\-Ricci curvature, \(3\)constant\-K=1K=1, and \(4\) anMLP\-basedestimator that updates the Lorentz scale parameter with an MLP layer\. Table[9](https://arxiv.org/html/2608.21096#A5.T9)reports the corresponding results for CiteSeer under the 10\-client and 20\-client settings\. Table 9:Performance with different curvature initialization methods on CiteSeer\. Results are reported as mean±\\pmstandard deviation over five runs\.As shown in Table[9](https://arxiv.org/html/2608.21096#A5.T9), the choice of initialization method has only a marginal effect on performance\. This indicates that our method is robust to the initialization choice, and that the specific initializer isnot criticalto the effectiveness of our method\. What truly matters is that each client is assigned alearnable curvature, which can be adapted during training to better fit its local data distribution \(see[Figure 5](https://arxiv.org/html/2608.21096#S6.F5)\)\. ### E\.6Convergence Curves Figure 10:Convergence curves of𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}and strong baselines\.The convergence curves are shown in[Figure 10](https://arxiv.org/html/2608.21096#A5.F10)\. As the figures demonstrate, our proposed method achieves competitive convergence speed while reaching higher final performance\. This is consistent with the theoretical discussion in[AppendixD\.4](https://arxiv.org/html/2608.21096#A4.SS4): the decoupling strategy does not introduce additional client\-similarity estimation or clustering optimization, and the server\-side aggregation remains as simple as averaging the shared space\-like parameters\. Therefore,𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}can retain FedAvg\-level convergence behavior while better handling heterogeneous local geometries\. ### E\.7Curvature Sensitivity and Interpretation To further clarify why curvature should be client\-specific and learnable, we summarize the relevant empirical observations: ##### Client\-wise preferred curvature\. Different sources of client heterogeneity, including graph\-structural variation \(e\.g\., degree distribution, clustering patterns, and homophily\) and label\-distribution shift, may induce different preferred geometric biases across clients\. Empirically, we observe that different clients can achieve their best validation performance under different curvature values, rather than sharing a single globally optimal curvature or consistently preferring the Euclidean limit\. This suggests that client\-specific curvature acts as a compact geometric proxy for heterogeneous client conditions, instead of merely serving as a numerical scaling factor\. This also explains why directly fixing client curvature from a structure\-only Ricci estimate is insufficient\. The estimated Ricci curvature provides an informative initialization signal, but the curvature that benefits downstream prediction is also shaped by label\-distribution shift, feature geometry, and task\-specific optimization\. Therefore,𝖥𝗅𝖺𝗍𝖫𝖺𝗇𝖽\\mathsf\{FlatLand\}treats curvature as a learnable client\-specific parameter: the structural estimate serves as a prior, while training can adapt it toward the effective geometry preferred by each client\. ##### Fixed vs\. learnable curvature\. As shown in[Table 9](https://arxiv.org/html/2608.21096#A5.T9)and the ablation study in the main text, fixing curvature to a constant value consistently underperforms learnable curvature, with performance gaps of 1–3% across datasets\. This demonstrates the importance of client\-specific geometric adaptation\.
Similar Articles
Federated Learning
The article explains the concept of Federated Learning as a privacy-preserving machine learning technique that trains models on local devices rather than central servers. It details the process of encrypted parameter updates and aggregation to mitigate data leakage risks while maintaining model performance.
Towards Federated Long-Tailed Graph Learning: An Energy-Guided Dual Decoupling Approach
This paper introduces FedEPD, a framework for federated graph learning under long-tailed data distributions. It uses an energy-guided dual decoupling approach to separate topological purification from semantic recalibration, achieving state-of-the-art performance on benchmarks with up to 4.97% accuracy improvement.
Towards the Next Frontier of LLMs, Training on Private Data: A Cross-Domain Benchmark for Federated Fine-Tuning
This paper presents a cross-domain benchmark for federated fine-tuning of large language models on private data, evaluating LoRA, QLoRA, and IA3 strategies on healthcare and finance datasets. Results show federated fine-tuning approaches centralized performance and outperforms isolated learning, supporting its viability for adapting LLMs when data cannot be shared.
COSMOS: Model-Agnostic Personalized Federated Learning with Clustered Server Models and Pseudo-Label-Only Communication
This paper introduces COSMOS, a model-agnostic personalized federated learning framework that uses clustered server models and pseudo-label-only communication. It provides theoretical analysis showing exponential personalization risk contraction and demonstrates superior performance over existing baselines in heterogeneous environments.
Federated continual learning: A comprehensive survey on lifelong and privacy-preserving learning over distributed and non-stationary data
This paper provides a comprehensive survey of Federated Continual Learning (FCL), an emerging field that combines Federated Learning and Continual Learning to enable lifelong, adaptive, and privacy-preserving learning over distributed and non-stationary data. It proposes a taxonomy, reviews applications, metrics, and open challenges.