Multi-Agent Privacy Game in Federated Learning: A Unified Mean-Field View
Summary
This paper introduces a mean-field privacy game framework for federated learning, enabling tractable Nash equilibrium analysis for arbitrarily many clients with heterogeneous privacy preferences and yielding a personalized privacy guarantee.
View Cached Full Text
Cached at: 07/28/26, 06:24 AM
# Multi-Agent Privacy Game in Federated Learning: A Unified Mean-Field View
Source: [https://arxiv.org/html/2607.23029](https://arxiv.org/html/2607.23029)
###### Abstract
Federated learning enables collaborative model training across distributed clients without centralising their data, yet privacy remains a persistent concern because the shared model updates can leak information about local datasets\. Existing privacy\-preserving methods either inject calibrated noise into client updates, limiting their composition guarantees, or formulate client privacy choices as a multi\-agent game whose Nash equilibrium becomes intractable as the number of clients grows\. We bridge these two lines of work by formulating privacy\-preserving federated learning as a mean\-field privacy game: each client strategically chooses its own privacy budget while interacting with the population only through a single mean\-field statistic\. The mean\-field limit yields a tractable equilibrium for arbitrarily many clients, accommodates heterogeneous client preferences, and inherits an exponentially decaying privacy guarantee through a log\-Sobolev contraction\. The framework recovers the entropic privacy baseline as the homogeneous special case and the multi\-agent privacy game as the finite\-population case\. Experiments on quadratic regression, logistic regression, and MNIST demonstrate that the proposed framework attains the privacy\-utility trade\-off of the entropic baseline while delivering a personalised privacy guarantee that the homogeneous baseline cannot express\.
## 1Introduction
Federated Learning \(FL\)\[[19](https://arxiv.org/html/2607.23029#bib.bib17),[17](https://arxiv.org/html/2607.23029#bib.bib38)\]trains a shared model across data\-owning clients without centralizing raw data\. Although locality reduces some privacy risks, model updates leak through gradient inversion and membership inference\[[35](https://arxiv.org/html/2607.23029#bib.bib22),[10](https://arxiv.org/html/2607.23029#bib.bib37),[7](https://arxiv.org/html/2607.23029#bib.bib23)\]\. Three largely independent lines of work address this leakage\.
*Where does the control act?*Privacy\-preserving FL methods can be partitioned by where each client’s control variable enters the system: it can act on the agent’s*state*\(samples or gradients used for local updates\) or directly on the*model output*\(an additive perturbation ofwkw\_\{k\}at the aggregation step\)\. DP\-SGD\[[1](https://arxiv.org/html/2607.23029#bib.bib16)\]and its federated variants\[[11](https://arxiv.org/html/2607.23029#bib.bib32),[30](https://arxiv.org/html/2607.23029#bib.bib33)\]act on the state: a calibrated Gaussian is added to each clipped gradient, and the privacy cost accumulates polynomially with rounds under standard composition\[[21](https://arxiv.org/html/2607.23029#bib.bib18)\]\. FLRA\[[23](https://arxiv.org/html/2607.23029#bib.bib4)\]likewise applies its strategic perturbation\(Λi𝐱\+δi\)\(\\Lambda^\{i\}\\mathbf\{x\}\+\\delta^\{i\}\)to the state, though for robustness rather than privacy\. In contrast, the parameter\-perturbation game\[[34](https://arxiv.org/html/2607.23029#bib.bib15)\]treatsδk\\delta\_\{k\}as a control on the model output and aggregates viaw=∑kpk\(wk\+δk\)w=\\sum\_\{k\}p\_\{k\}\(w\_\{k\}\+\\delta\_\{k\}\); under FedAvg’s linear constraint, KKT collapses every Nash point toδk=𝟎\\delta\_\{k\}=\\mathbf\{0\}\(Section[4](https://arxiv.org/html/2607.23029#S4)\)\. Only controls that act on the state admit non\-trivial equilibria, so this is where the rest of the paper lives\.
*Mean\-field as theN→∞N\\\!\\to\\\!\\inftylimit\.*Two parallel lines of work address different aspects of the resulting state\-control game\. The multi\-agent privacy game\[[34](https://arxiv.org/html/2607.23029#bib.bib15),[33](https://arxiv.org/html/2607.23029#bib.bib10),[9](https://arxiv.org/html/2607.23029#bib.bib6)\]fixes a finiteNNand asks for a Nash equilibrium of the controls\{εk\}\\\{\\varepsilon\_\{k\}\\\}, whereεk\\varepsilon\_\{k\}scales the noise added to clientkk’s gradient or sample \(the MAPG\-DP and MAPG\-input formulations of\[[34](https://arxiv.org/html/2607.23029#bib.bib15)\]\)\. Mean\-Field Entropic Privacy \(MFEP\)\[[5](https://arxiv.org/html/2607.23029#bib.bib21)\]fixes a homogeneousε\\varepsilonand letsN→∞N\\\!\\to\\\!\\infty, replacing additive noise with an entropic Wasserstein gradient flow that admits exponential log\-Sobolev contraction\[[22](https://arxiv.org/html/2607.23029#bib.bib27),[15](https://arxiv.org/html/2607.23029#bib.bib26)\]\. Neither line on its own captures the joint behaviour: MAPG\-DP is intractable for realisticNNand uses geometry\-agnostic Gaussian noise, while MFEP cannot represent heterogeneous privacy preferences\.
#### Our framing\.
We treat privacy\-preserving FL as a standard mean\-field stochastic differential game\[[16](https://arxiv.org/html/2607.23029#bib.bib19),[8](https://arxiv.org/html/2607.23029#bib.bib30)\]: each agent has a stateXt∈ℝdX\_\{t\}\\in\\mathbb\{R\}^\{d\}\(a privatised datum\), a typeβ\\beta\(its privacy preference\), and a controlεk\\varepsilon\_\{k\}\(its privacy budget\); the state distributionμt∈𝒫2\(ℝd\)\\mu\_\{t\}\\in\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{d\}\)aggregates the population\. The empirical state distributionμt\(N\)=1N∑kμtk\\mu\_\{t\}^\{\(N\)\}=\\frac\{1\}\{N\}\\sum\_\{k\}\\mu\_\{t\}^\{k\}converges, by propagation of chaos\[[8](https://arxiv.org/html/2607.23029#bib.bib30)\], to a deterministicμt\\mu\_\{t\}asN→∞N\\\!\\to\\\!\\infty\. This single limit organises three previously separate works:
finiteNN\(multi\-agent\)N→∞N\\\!\\to\\\!\\infty\(mean\-field\)no game \(homogeneous\)DP\-SGD\[[1](https://arxiv.org/html/2607.23029#bib.bib16)\]MFEP\[[5](https://arxiv.org/html/2607.23029#bib.bib21)\]game \(heterogeneous controlεk\\varepsilon\_\{k\}\)MAPG\-DP\[[34](https://arxiv.org/html/2607.23029#bib.bib15)\]MFPG \(this paper\)The bottom\-right cell is the missing piece\.MFEP is MFPG without the game\(singleε\\varepsilonacross the population\);MAPG\-DP is MFPG without the mean\-field limit\(finiteNN\)\. MFPG closes both gaps with one construction\.
#### Contributions\.
1. \(i\)Unified MFG framework \(Section[3](https://arxiv.org/html/2607.23029#S3)\)\.We formalise privacy\-preserving FL as a mean\-field stochastic differential game in the standard sense, with stateXtX\_\{t\}, typeβk\\beta\_\{k\}, controlεk\\varepsilon\_\{k\}, state distributionμt\\mu\_\{t\}, and running costf\(x,μ,ε;β\)=ℓ\+βδdpf\(x,\\mu,\\varepsilon;\\beta\)=\\ell\+\\beta\\delta\_\{\\mathrm\{dp\}\}\. A finite\-NNNash equilibrium of MAPG\-DP and the mean\-field Nash equilibrium of MFPG are the two endpoints of the sameN→∞N\\\!\\to\\\!\\inftylimit; MFEP is the no\-game special case\.
2. \(ii\)Mean\-Field Privacy Game \(Section[5](https://arxiv.org/html/2607.23029#S5)\)\.We prove existence of an MFNE on a finite action grid via Kakutani’s theorem and certify\(ϵdp,δdp\)\(\\epsilon\_\{\\mathrm\{dp\}\},\\delta\_\{\\mathrm\{dp\}\}\)\-DP at the equilibrium, withδdp\\delta\_\{\\mathrm\{dp\}\}contracting exponentially whenαε∗\>λ\+G\\alpha\\varepsilon^\{\*\}\>\\lambda\+G\. The bound recovers the MFEP guarantee at homogeneousβk\\beta\_\{k\}and recovers MAPG\-DP’s heterogeneity at finiteNN\.
3. \(iii\)Common solver and empirical study \(Sections[6](https://arxiv.org/html/2607.23029#S6),[7](https://arxiv.org/html/2607.23029#S7)\)\.A single Particle–Sinkhorn JKO step services everyN→∞N\\\!\\to\\\!\\inftyvariant; the action update is the only block that varies between MFEP and MFPG\. We compare all four cells on quadratic, logistic, and MNIST benchmarks, with every reported number traced toresults/full\_benchmark\.csv\.
## 2Preliminaries
#### Federated learning\.
Each clientk∈\[N\]k\\in\[N\]holds a local dataset𝒟k\\mathcal\{D\}\_\{k\}\. The global objective isminwF\(w\)=∑kpkFk\(w\)\\min\_\{w\}F\(w\)=\\sum\_\{k\}p\_\{k\}F\_\{k\}\(w\)withpk=nk/np\_\{k\}=n\_\{k\}/nandFk\(w\)=nk−1∑iℓ\(w;xik,yik\)F\_\{k\}\(w\)=n\_\{k\}^\{\-1\}\\sum\_\{i\}\\ell\(w;x\_\{i\}^\{k\},y\_\{i\}^\{k\}\)\[[19](https://arxiv.org/html/2607.23029#bib.bib17)\]\. At roundttthe server broadcastswtw^\{t\}, each client returns a local updatewkt\+1w\_\{k\}^\{t\+1\}, and the server aggregates via FedAvgwt\+1=∑kpkwkt\+1w^\{t\+1\}=\\sum\_\{k\}p\_\{k\}w\_\{k\}^\{t\+1\}\.
#### Differential privacy\.
A randomised mechanismℳ\\mathcal\{M\}is\(ϵ,δ\)\(\\epsilon,\\delta\)\-DP if, for all neighbouring datasets and all measurableSS,Pr\[ℳ\(𝒟\)∈S\]≤eϵPr\[ℳ\(𝒟′\)∈S\]\+δ\\Pr\[\\mathcal\{M\}\(\\mathcal\{D\}\)\\in S\]\\leq e^\{\\epsilon\}\\Pr\[\\mathcal\{M\}\(\\mathcal\{D\}^\{\\prime\}\)\\in S\]\+\\delta\. The Gaussian mechanism withσ=C2ln\(1\.25/δ\)/ϵ\\sigma=C\\sqrt\{2\\ln\(1\.25/\\delta\)\}/\\epsilonis\(ϵ,δ\)\(\\epsilon,\\delta\)\-DP for sensitivityCC\[[1](https://arxiv.org/html/2607.23029#bib.bib16)\]; underTT\-fold compositionϵ\\epsilongrows asO\(T\)O\(\\sqrt\{T\}\)\.
#### Mean\-field games\.
A mean\-field game\[[16](https://arxiv.org/html/2607.23029#bib.bib19),[13](https://arxiv.org/html/2607.23029#bib.bib31),[8](https://arxiv.org/html/2607.23029#bib.bib30)\]models a continuum of identical agents whose individual optimal control problem depends on the population distributionμt\\mu\_\{t\}\. AsN→∞N\\to\\infty, theNN\-player Nash equilibrium converges to a mean\-field Nash equilibrium \(MFNE\)\. The MFNE depends only onμt\\mu\_\{t\}, not on individual identities, which makes it tractable when theNN\-player problem is not\.
#### Wasserstein gradient flows and JKO\.
For a free\-energy functionalℱ\\mathcal\{F\}on𝒫2\(ℝd\)\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{d\}\), the Wasserstein gradient flow obeys∂tμt=−∇W2ℱ\(μt\)\\partial\_\{t\}\\mu\_\{t\}=\-\\nabla\_\{W\_\{2\}\}\\mathcal\{F\}\(\\mu\_\{t\}\)and admits a time\-discrete Jordan–Kinderlehrer–Otto scheme\[[15](https://arxiv.org/html/2607.23029#bib.bib26)\]μk\+1=argminρ\{ℱ\(ρ\)\+\(2τ\)−1W22\(ρ,μk\)\}\\mu\_\{k\+1\}=\\arg\\min\_\{\\rho\}\\\{\\mathcal\{F\}\(\\rho\)\+\(2\\tau\)^\{\-1\}W\_\{2\}^\{2\}\(\\rho,\\mu\_\{k\}\)\\\}\. A measureν\\nusatisfies a log\-Sobolev inequality with constantα\\alphaifEntν\(f2\)≤\(2/α\)∫‖∇f‖2𝑑ν\\mathrm\{Ent\}\_\{\\nu\}\(f^\{2\}\)\\leq\(2/\\alpha\)\\int\\\|\\nabla f\\\|^\{2\}d\\nu, withν=𝒩\(0,σ2I\)\\nu=\\mathcal\{N\}\(0,\\sigma^\{2\}I\)givingα=1/σ2\\alpha=1/\\sigma^\{2\}\. Otto–Villani impliesW2W\_\{2\}\-contraction at rateαε\\alpha\\varepsilonfor the entropic flow, which we use in Section[5](https://arxiv.org/html/2607.23029#S5)to certify privacy\.
## 3A Unified Mean\-Field Game Framework
We formalise privacy\-preserving FL as a mean\-field stochastic differential game in the standard sense of\[[16](https://arxiv.org/html/2607.23029#bib.bib19),[8](https://arxiv.org/html/2607.23029#bib.bib30)\]\. Section[3\.1](https://arxiv.org/html/2607.23029#S3.SS1)identifies the four ingredients of the game \(state, type, control, dynamics\) and Section[3\.2](https://arxiv.org/html/2607.23029#S3.SS2)specialises them to FL through two privacy mechanisms; the resulting2×22\\times 2grid in Table[1](https://arxiv.org/html/2607.23029#S3.T1)populates DP\-SGD, MFEP, MAPG\-DP, and our proposed MFPG as four named cells\.
### 3\.1State, type, control, and dynamics
A representative agent in our game is described by:
- •*State*Xt∈ℝdX\_\{t\}\\in\\mathbb\{R\}^\{d\}: a \(privatised\) datum used by an individual client to compute its local update at roundtt\.
- •*Type*β∈\[βmin,βmax\]⊂ℝ\+\\beta\\in\[\\beta\_\{\\min\},\\beta\_\{\\max\}\]\\subset\\mathbb\{R\}\_\{\+\}: the client’s privacy preference \(heterogeneity parameter\), distributed byρ\(dβ\)\\rho\(d\\beta\)\.
- •*Control*εk∈ℰ=\{ε\(1\),…,ε\(m\)\}\\varepsilon\_\{k\}\\in\\mathcal\{E\}=\\\{\\varepsilon^\{\(1\)\},\\ldots,\\varepsilon^\{\(m\)\}\\\}: a per\-client privacy budget chosen on a finite grid; this is the strategic variable of the game\.
- •*Dynamics:*conditional onεk\\varepsilon\_\{k\}, the state evolves under the controlled SDE dXtk=b\(Xtk,μt,εk\)dt\+σ\(εk\)dWt,dX\_\{t\}^\{k\}=b\\bigl\(X\_\{t\}^\{k\},\\mu\_\{t\},\\varepsilon\_\{k\}\\bigr\)\\,dt\+\\sigma\\bigl\(\\varepsilon\_\{k\}\\bigr\)\\,dW\_\{t\},\(3\.1\)withWtW\_\{t\}a standarddd\-dimensional Brownian motion\. The driftbband diffusionσ\\sigmaare fixed by the privacy mechanismℳ\\mathcal\{M\}\(Section[3\.2](https://arxiv.org/html/2607.23029#S3.SS2)\)\.
- •*State distribution*μt∈𝒫2\(ℝd\)\\mu\_\{t\}\\in\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{d\}\): the distribution ofXtX\_\{t\}across the agent population\. The empirical measureμt\(N\)=1N∑k=1Nμtk\\mu\_\{t\}^\{\(N\)\}=\\tfrac\{1\}\{N\}\\sum\_\{k=1\}^\{N\}\\mu\_\{t\}^\{k\}converges weakly to a deterministicμt\\mu\_\{t\}asN→∞N\\to\\inftyby propagation of chaos\[[8](https://arxiv.org/html/2607.23029#bib.bib30),[13](https://arxiv.org/html/2607.23029#bib.bib31)\]\.
- •*Running cost*\(per agent of typeβ\\beta\): f\(x,μ,ε;β\)=ℓ\(x;wt\)\+βδdp\(ε\),f\(x,\\mu,\\varepsilon;\\beta\)=\\ell\(x;w\_\{t\}\)\+\\beta\\,\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon\),\(3\.2\)combining the training lossℓ\\ellagainst the current global modelwtw\_\{t\}with the privacy termβδdp\\beta\\,\\delta\_\{\\mathrm\{dp\}\}scaled by the agent’s type\. The disutility minimised by clientkkis the time\-integral of \([3\.2](https://arxiv.org/html/2607.23029#S3.E2)\), written compactly as Uk\(εk;ε¯\)=Lk\(w;εk\)\+βkδdp\(εk\),U\_\{k\}\(\\varepsilon\_\{k\};\\,\\bar\{\\varepsilon\}\)=L\_\{k\}\(w;\\,\\varepsilon\_\{k\}\)\+\\beta\_\{k\}\\,\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon\_\{k\}\),\(3\.3\)whereLkL\_\{k\}is the cumulative training loss andε¯=𝔼μ\[ε\]\\bar\{\\varepsilon\}=\\mathbb\{E\}\_\{\\mu\}\[\\varepsilon\]is the mean control across the population\.
The model parameterswwenter only throughℓ\(x;wt\)\\ell\(x;w\_\{t\}\)and are otherwise external to the game; the strategic content lives entirely in\(Xt,μt,εk\)\(X\_\{t\},\\mu\_\{t\},\\varepsilon\_\{k\}\)and \([3\.2](https://arxiv.org/html/2607.23029#S3.E2)\)\.
### 3\.2Two privacy mechanisms: Gaussian and entropic
The drift–diffusion pair\(b,σ\)\(b,\\sigma\)in \([3\.1](https://arxiv.org/html/2607.23029#S3.E1)\) is determined by which privacy mechanism the system runs\. We consider two:
- •*Gaussian mechanism\.*The controlεk\\varepsilon\_\{k\}scales the additive noise on each clipped gradient,σ\(εk\)=C2ln\(1\.25/δ\)/εk\\sigma\(\\varepsilon\_\{k\}\)=C\\sqrt\{2\\ln\(1\.25/\\delta\)\}/\\varepsilon\_\{k\}, with driftbbgiven by the clipped loss gradient\[[1](https://arxiv.org/html/2607.23029#bib.bib16)\]\. Privacy is tracked by Rényi\-DP composition, giving a polynomially accumulatingϵdp=O\(T\)\\epsilon\_\{\\mathrm\{dp\}\}=O\(\\sqrt\{T\}\)\.
- •*Entropic mechanism\.*The state distribution evolves under the Wasserstein gradient flow of the KL\-regularised free energy ℱλ\(μ;εk\)=𝔼x∼μ\[L\(x;w\)\]\+εkKL\(μ∥ν\)\+λ2Varμ\[x\],\\mathcal\{F\}\_\{\\lambda\}\(\\mu;\\varepsilon\_\{k\}\)=\\mathbb\{E\}\_\{x\\sim\\mu\}\[L\(x;w\)\]\+\\varepsilon\_\{k\}\\,\\mathrm\{KL\}\(\\mu\\\|\\nu\)\+\\tfrac\{\\lambda\}\{2\}\\,\\mathrm\{Var\}\_\{\\mu\}\[x\],\(3\.4\)with Gaussian priorν=𝒩\(0,σ2I\)\\nu=\\mathcal\{N\}\(0,\\sigma^\{2\}I\)\. Equivalently \(Appendix[C\.1](https://arxiv.org/html/2607.23029#A3.SS1)\), the controlled SDE \([3\.1](https://arxiv.org/html/2607.23029#S3.E1)\) has driftb\(x,μ,ε\)=−∇L\(x;w\)−εσ2x−λ\(x−x¯\)b\(x,\\mu,\\varepsilon\)=\-\\nabla L\(x;w\)\-\\tfrac\{\\varepsilon\}\{\\sigma^\{2\}\}x\-\\lambda\(x\-\\bar\{x\}\)and diffusionσ\(ε\)=2ε\\sigma\(\\varepsilon\)=\\sqrt\{2\\varepsilon\}\. Privacy is tracked by log\-Sobolev contraction\[[22](https://arxiv.org/html/2607.23029#bib.bib27)\], withδdp\\delta\_\{\\mathrm\{dp\}\}decaying exponentially whenαεk\>λ\+G\\alpha\\varepsilon\_\{k\}\>\\lambda\+G\.
###### Remark 3\.1\(Why the control acts on the state and not on the model\)\.
A superficially natural alternative would be to make the control a perturbationδk∈ℝd\\delta\_\{k\}\\in\\mathbb\{R\}^\{d\}added directly to the local modelwkw\_\{k\}at the aggregation step,w=∑kpk\(wk\+δk\)w=\\sum\_\{k\}p\_\{k\}\(w\_\{k\}\+\\delta\_\{k\}\)\[[34](https://arxiv.org/html/2607.23029#bib.bib15)\]\. With the disutilityLk\(w\)−βk‖δk‖2L\_\{k\}\(w\)\-\\beta\_\{k\}\\\|\\delta\_\{k\}\\\|^\{2\}, KKT forcesδk=𝟎\\delta\_\{k\}=\\mathbf\{0\}at every Nash point \(Section[4](https://arxiv.org/html/2607.23029#S4)\), so the limiting MFG is degenerate\. We therefore restrict attention to controls that act on the state SDE \([3\.1](https://arxiv.org/html/2607.23029#S3.E1)\); this is the standard MFG setup in which the agent’s strategy shapes its own state evolution\.
Table 1:The MFG framework in two axes\. Two design choices populate the2×22\\times 2grid: client heterogeneity \(rows: no game vs\. game\) and population size \(columns: finiteNNvs\.N→∞N\\to\\infty\)\. Entries name the method, the active noise mechanism \(Gaussian or Entropic\), and the accountant\.FiniteNN\(multi\-agent\)N→∞N\\to\\infty\(mean field\)No game\(singleε\\varepsilon\)DP\-SGD\[[1](https://arxiv.org/html/2607.23029#bib.bib16)\]
Gaussian noise on gradient, RDP
δ\\deltaconst\.,ϵ=O\(T\)\\epsilon=O\(\\sqrt\{T\}\)MFEP\[[5](https://arxiv.org/html/2607.23029#bib.bib21)\]
Entropic flow onμt\\mu\_\{t\}, LSI
δ=O\(e−αεKτ\)\\delta=O\(e^\{\-\\alpha\\varepsilon K\\tau\}\)Game\(heterogeneousβk\\beta\_\{k\}\)MAPG\-DP\[[34](https://arxiv.org/html/2607.23029#bib.bib15)\]
Gaussian noise,NN\-Nash on\{εk\}\\\{\\varepsilon\_\{k\}\\\}
Intractable for largeNNMFPG \(ours\)
Entropic flow, MFNE onε¯\\bar\{\\varepsilon\}
Heterogeneous \+ exponential decay
## 4The Three Baseline Cells
The three already\-named cells of Table[1](https://arxiv.org/html/2607.23029#S3.T1)are recalled below in the language of Section[3](https://arxiv.org/html/2607.23029#S3); MFPG itself is deferred to Section[5](https://arxiv.org/html/2607.23029#S5)\. The MAPG\-DP subsection also discharges the model\-output alternative, whoseδk=0\\delta\_\{k\}=0pathology is what prevents the multi\-agent privacy game literature from being combined with mean\-field analysis directly\.
### 4\.1FiniteNN, no game: DP\-SGD
DP\-SGD\[[1](https://arxiv.org/html/2607.23029#bib.bib16)\]runs the Gaussian mechanism of Section[3\.2](https://arxiv.org/html/2607.23029#S3.SS2)with a single shared budgetε=εtgt\\varepsilon=\\varepsilon\_\{\\mathrm\{tgt\}\}\. After clippingΔwk←Δwk⋅min\(1,C/‖Δwk‖2\)\\Delta w\_\{k\}\\leftarrow\\Delta w\_\{k\}\\cdot\\min\(1,C/\\\|\\Delta w\_\{k\}\\\|\_\{2\}\)each client adds Gaussian noise𝒩\(0,σ2I\)\\mathcal\{N\}\(0,\\sigma^\{2\}I\)withσ=C2ln\(1\.25/δ\)/ε\\sigma=C\\sqrt\{2\\ln\(1\.25/\\delta\)\}/\\varepsilon\. RDP composition returnsϵdp=O\(Tlog\(1/δ\)/σ\)\\epsilon\_\{\\mathrm\{dp\}\}=O\(\\sqrt\{T\\log\(1/\\delta\)\}/\\sigma\), so the privacy cost*grows*polynomially with roundsTT\.
### 4\.2N→∞N\\to\\infty, no game: MFEP
MFEP\[[5](https://arxiv.org/html/2607.23029#bib.bib21)\]replaces external noise injection with the entropic free\-energy \([3\.4](https://arxiv.org/html/2607.23029#S3.E4)\) at a*fixed*regularization strengthε\\varepsilon, evaluated against the state distributionμt\\mu\_\{t\}\. The Wasserstein gradient flow obeys the Fokker–Planck equation
∂tμt=∇⋅\[μt\(∇L\(x\)\+εσ2x\+λ\(x−x¯t\)\)\]\+εΔμt,\\partial\_\{t\}\\mu\_\{t\}=\\nabla\\\!\\cdot\\\!\\Bigl\[\\mu\_\{t\}\\Bigl\(\\nabla L\(x\)\+\\tfrac\{\\varepsilon\}\{\\sigma^\{2\}\}x\+\\lambda\(x\-\\bar\{x\}\_\{t\}\)\\Bigr\)\\Bigr\]\+\\varepsilon\\Delta\\mu\_\{t\},\(4\.1\)withx¯t=𝔼μt\[x\]\\bar\{x\}\_\{t\}=\\mathbb\{E\}\_\{\\mu\_\{t\}\}\[x\]\. Three mechanisms operate simultaneously:*loss\-driven drift*,*prior attraction*at rateε/σ2\\varepsilon/\\sigma^\{2\}, and*entropic diffusion*εΔμt\\varepsilon\\Delta\\mu\_\{t\}that delivers privacy intrinsically\. Under LSI for the prior with constantα=1/σ2\\alpha=1/\\sigma^\{2\}, the JKO discretization contracts\[[22](https://arxiv.org/html/2607.23029#bib.bib27)\]: for neighbouring measuresμk\\mu\_\{k\}andμk′\\mu\_\{k\}^\{\\prime\}differing in a single client’s data,
W2\(μK,μK′\)≤e−\(αε−λ−G\)KτW2\(μ0,μ0′\),W\_\{2\}\(\\mu\_\{K\},\\mu\_\{K\}^\{\\prime\}\)\\leq e^\{\-\(\\alpha\\varepsilon\-\\lambda\-G\)K\\tau\}\\,W\_\{2\}\(\\mu\_\{0\},\\mu\_\{0\}^\{\\prime\}\),\(4\.2\)providedαε\>λ\+G\\alpha\\varepsilon\>\\lambda\+G, withGGthe clipped gradient norm bound\. Converting through total\-variation gives the exponentially decayingδdp\\delta\_\{\\mathrm\{dp\}\}bound used throughout the paper:
δdp\(ε\)≤CdNexp\(−\(αε−λ−G\)Kτ2\),Cd=min\(d,10\)\.\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon\)\\leq\\frac\{C\_\{d\}\}\{\\sqrt\{N\}\}\\exp\\\!\\Bigl\(\-\\tfrac\{\(\\alpha\\varepsilon\-\\lambda\-G\)K\\tau\}\{2\}\\Bigr\),\\quad C\_\{d\}=\\min\(\\sqrt\{d\},10\)\.\(4\.3\)
### 4\.3FiniteNN, game: MAPG\-DP
MAPG\-DP\[[34](https://arxiv.org/html/2607.23029#bib.bib15), §§ 3\.2\.2, 3\.3\.4\]gives each clientkka per\-client privacy budgetεk∈ℝ\+\\varepsilon\_\{k\}\\in\\mathbb\{R\}\_\{\+\}that scales the Gaussian noise added to its local gradient, with scaleσk\(εk\)=2ηC/εk\\sigma\_\{k\}\(\\varepsilon\_\{k\}\)=2\\eta C/\\varepsilon\_\{k\}, or, equivalently, an additive shift on its features as in FLRA\[[23](https://arxiv.org/html/2607.23029#bib.bib4)\]\. The disutility \([3\.3](https://arxiv.org/html/2607.23029#S3.E3)\) is delivered by RDP composition, and theNN\-player Nash equilibrium\(ε1∗,…,εN∗\)\(\\varepsilon\_\{1\}^\{\*\},\\ldots,\\varepsilon\_\{N\}^\{\*\}\)satisfiesUk\(εk∗,ε−k∗\)≤Uk\(εk,ε−k∗\)U\_\{k\}\(\\varepsilon\_\{k\}^\{\*\},\\varepsilon\_\{\-k\}^\{\*\}\)\\leq U\_\{k\}\(\\varepsilon\_\{k\},\\varepsilon\_\{\-k\}^\{\*\}\)for everykk, whereε−k∗\\varepsilon\_\{\-k\}^\{\*\}denotes the strategies of all other clients at equilibrium\. Existence follows by standard arguments, but computation is intractable for realisticNNbecause each best response couples through everyεj\\varepsilon\_\{j\}via the FedAvg constraint
w=1N∑j=1N\(wj\+𝒩\(0,σj\(εj\)2I\)\)\.w=\\tfrac\{1\}\{N\}\\\!\\sum\_\{j=1\}^\{N\}\\\!\\bigl\(w\_\{j\}\+\\mathcal\{N\}\(0,\\sigma\_\{j\}\(\\varepsilon\_\{j\}\)^\{2\}I\)\\bigr\)\.\(4\.4\)
The same multi\-agent literature also studies a superficially natural alternative in which the strategy is an additive perturbationδk∈ℝd\\delta\_\{k\}\\in\\mathbb\{R\}^\{d\}on the local modelwkw\_\{k\}, with disutilityLk\(w\)−βk‖δk‖2L\_\{k\}\(w\)\-\\beta\_\{k\}\\\|\\delta\_\{k\}\\\|^\{2\}subject tow=∑kpk\(wk\+δk\)w=\\sum\_\{k\}p\_\{k\}\(w\_\{k\}\+\\delta\_\{k\}\)\[[34](https://arxiv.org/html/2607.23029#bib.bib15), § 3\.3\.1\]\. The KKT stationarity conditions in feature dimensionmm,
∂ℒ∂wk,m=1Nλk,m,∂ℒ∂δk,m=1Nλk,m−2βkδk,m,\\frac\{\\partial\\mathcal\{L\}\}\{\\partial w\_\{k,m\}\}=\\tfrac\{1\}\{N\}\\,\\lambda\_\{k,m\},\\qquad\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\delta\_\{k,m\}\}=\\tfrac\{1\}\{N\}\\,\\lambda\_\{k,m\}\-2\\beta\_\{k\}\\,\\delta\_\{k,m\},\(4\.5\)together forceλk,m=0\\lambda\_\{k,m\}=0and henceδk,m=0\\delta\_\{k,m\}=0at every Nash point: the aggregation constraint cancels the privacy term\. We therefore restrict the rest of the paper to controls that act on the state, since only such controls admit both non\-trivial equilibria and theN→∞N\\\!\\to\\\!\\inftylimit developed next\.
## 5Mean\-Field Privacy Game \(MFPG\)
We now instantiate theN→∞N\\to\\inftylimit of MAPG\-DP\. The construction below is structurally identical to MAPG\-DP except that \(i\) the Gaussian noise mechanism is replaced by the entropic flow \([4\.1](https://arxiv.org/html/2607.23029#S4.E1)\), which the LSI contraction \([4\.3](https://arxiv.org/html/2607.23029#S4.E3)\) certifies, and \(ii\) theNN\-Nash equilibrium of\{εk\}\\\{\\varepsilon\_\{k\}\\\}is replaced by the mean\-field Nash equilibrium on the population meanε¯\\bar\{\\varepsilon\}\. MFEP is recovered as the homogeneous special case \(βk\\beta\_\{k\}constant\)\.
### 5\.1Formulation
Each clientkkchooses a controlεk∈ℰ=\{ε\(1\),…,ε\(m\)\}\\varepsilon\_\{k\}\\in\\mathcal\{E\}=\\\{\\varepsilon^\{\(1\)\},\\ldots,\\varepsilon^\{\(m\)\}\\\}on a finite grid; conditional onεk\\varepsilon\_\{k\}, its state sliceμtk\\mu\_\{t\}^\{k\}evolves under the entropic flow \([4\.1](https://arxiv.org/html/2607.23029#S4.E1)\) and contributes to the population state distributionμt\\mu\_\{t\}through the lossLk\(w;εk\)=𝔼x∼μtk\[ℓ\(w;x\)\]L\_\{k\}\(w;\\varepsilon\_\{k\}\)=\\mathbb\{E\}\_\{x\\sim\\mu\_\{t\}^\{k\}\}\[\\ell\(w;x\)\]\. The empirical measureμt\(N\)=1N∑kμtk\\mu\_\{t\}^\{\(N\)\}=\\tfrac\{1\}\{N\}\\sum\_\{k\}\\mu\_\{t\}^\{k\}converges, by propagation of chaos\[[8](https://arxiv.org/html/2607.23029#bib.bib30),[13](https://arxiv.org/html/2607.23029#bib.bib31)\], to a deterministicμt∈𝒫2\(ℝd\)\\mu\_\{t\}\\in\\mathcal\{P\}\_\{2\}\(\\mathbb\{R\}^\{d\}\)asN→∞N\\to\\infty, so each client’s disutility \([3\.3](https://arxiv.org/html/2607.23029#S3.E3)\) couples to the rest of the population only through the scalar mean\-field strengthε¯=𝔼μ\[ε\]\\bar\{\\varepsilon\}=\\mathbb\{E\}\_\{\\mu\}\[\\varepsilon\]that replaces the vectorε−k\\varepsilon\_\{\-k\}of the finite\-NNgame\. Throughout this section we use the disutility in the form
Uk\(εk;ε¯\)=Lk\(w;εk\)\+βkδdp\(εk\),U\_\{k\}\(\\varepsilon\_\{k\};\\,\\bar\{\\varepsilon\}\)=L\_\{k\}\(w;\\,\\varepsilon\_\{k\}\)\+\\beta\_\{k\}\\,\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon\_\{k\}\),\(5\.1\)withδdp\\delta\_\{\\mathrm\{dp\}\}now given by the LSI bound \([4\.3](https://arxiv.org/html/2607.23029#S4.E3)\) rather than by RDP composition\.
ThisN→∞N\\to\\inftylimit is well\-posed only because the controls act on the state: each client’s best response in \([4\.4](https://arxiv.org/html/2607.23029#S4.E4)\) depends onε−k\\varepsilon\_\{\-k\}only through aggregate statistics ofμt\(N\)\\mu\_\{t\}^\{\(N\)\}, which collapse toε¯\\bar\{\\varepsilon\}in the limit\. The model\-output alternative does not admit such a limit, since each perturbation entersw=∑jpj\(wj\+δj\)w=\\sum\_\{j\}p\_\{j\}\(w\_\{j\}\+\\delta\_\{j\}\)with vanishing weightO\(1/N\)O\(1/N\)and theδk=0\\delta\_\{k\}=0collapse of \([4\.5](https://arxiv.org/html/2607.23029#S4.E5)\) is its finite\-NNshadow\.
### 5\.2Game equilibrium
A pair\(μ∗,ε∗\)\(\\mu^\{\*\},\\varepsilon^\{\*\}\)is a*mean\-field Nash equilibrium*\(MFNE\) of MFPG if \(i\)μ∗\\mu^\{\*\}is the stationary state distribution of \([4\.1](https://arxiv.org/html/2607.23029#S4.E1)\) at the population meanε¯∗=𝔼μ∗\[ε\]\\bar\{\\varepsilon\}^\{\*\}=\\mathbb\{E\}\_\{\\mu^\{\*\}\}\[\\varepsilon\], and \(ii\)ε∗\\varepsilon^\{\*\}is a best response toε¯∗\\bar\{\\varepsilon\}^\{\*\}for every client given its preferenceβk\\beta\_\{k\}\(formal definition in Appendix[C\.4](https://arxiv.org/html/2607.23029#A3.SS4)\)\. At equilibrium, no client can reduce its disutility by unilaterally changingεk\\varepsilon\_\{k\}given that every other client playsε−k∗\\varepsilon^\{\*\}\_\{\-k\}\. For heterogeneous clients with distinctβk\\beta\_\{k\}, the MFNE is characterised by a fixed point of the averaged best\-response mapΦ:ε¯↦𝔼k\[BRk\(ε¯\)\]\\Phi:\\bar\{\\varepsilon\}\\mapsto\\mathbb\{E\}\_\{k\}\[\\mathrm\{BR\}\_\{k\}\(\\bar\{\\varepsilon\}\)\]; existence follows from Kakutani’s theorem on the convex hull ofℰ\\mathcal\{E\}\(Appendix[C\.4](https://arxiv.org/html/2607.23029#A3.SS4)\)\. Section[5\.3](https://arxiv.org/html/2607.23029#S5.SS3)gives an equivalent differential characterisation through coupled HJB and FPK PDEs\.
### 5\.3HJB–FPK characterisation
The fixed\-point characterisation of Section[5\.2](https://arxiv.org/html/2607.23029#S5.SS2)is convenient for existence and for the finite\-grid solver of Section[6](https://arxiv.org/html/2607.23029#S6), but it hides the dynamical content of the equilibrium\. Following the standard mean\-field game system of\[[16](https://arxiv.org/html/2607.23029#bib.bib19),[8](https://arxiv.org/html/2607.23029#bib.bib30)\]and the FL–MFG analogy of\[[20](https://arxiv.org/html/2607.23029#bib.bib20)\], we now give an equivalent differential characterisation through coupled forward Fokker–Planck–Kolmogorov \(FPK\) and backward Hamilton–Jacobi–Bellman \(HJB\) equations specialised to our disutility \([5\.1](https://arxiv.org/html/2607.23029#S5.E1)\)\. To this end we momentarily relax the finite gridℰ\\mathcal\{E\}to the interval\[εmin,εmax\]⊂ℝ\+\[\\varepsilon\_\{\\min\},\\varepsilon\_\{\\max\}\]\\subset\\mathbb\{R\}\_\{\+\}and parameterise each client by a privacy preferenceβ∈\[βmin,βmax\]\\beta\\in\[\\beta\_\{\\min\},\\beta\_\{\\max\}\]distributed according toρ\(dβ\)\\rho\(d\\beta\), so that the state distribution splits across types asμt=∫μtβρ\(dβ\)\\mu\_\{t\}=\\int\\mu\_\{t\}^\{\\beta\}\\,\\rho\(d\\beta\)with global meanx¯t=𝔼μt\[X\]\\bar\{x\}\_\{t\}=\\mathbb\{E\}\_\{\\mu\_\{t\}\}\[X\]\. Under controlε\\varepsilon, a representative state of typeβ\\betaobeys the controlled SDE
dXt=b\(Xt,μt,ε\)dt\+2εdWt,b\(x,μ,ε\)=−∇L\(x;wt\)−εσ2x−λ\(x−x¯\),dX\_\{t\}=b\(X\_\{t\},\\mu\_\{t\},\\varepsilon\)\\,dt\+\\sqrt\{2\\varepsilon\}\\,dW\_\{t\},\\qquad b\(x,\\mu,\\varepsilon\)=\-\\nabla L\(x;w\_\{t\}\)\-\\tfrac\{\\varepsilon\}\{\\sigma^\{2\}\}x\-\\lambda\(x\-\\bar\{x\}\),\(5\.2\)withWtW\_\{t\}a standard Brownian motion, and seeks to minimise the cumulative cost
Jμ\(ε;β\)=𝔼\[∫0Tℓ\(Xt;wt\)𝑑t\]\+βδdp\(ε\),J^\{\\mu\}\(\\varepsilon;\\beta\)=\\mathbb\{E\}\\\!\\left\[\\int\_\{0\}^\{T\}\\ell\(X\_\{t\};w\_\{t\}\)\\,dt\\right\]\+\\beta\\,\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon\),\(5\.3\)whereδdp\(ε\)=\(Cd/N\)exp\(−θ\(ε−ε0\)\)\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon\)=\(C\_\{d\}/\\sqrt\{N\}\)\\exp\(\-\\theta\(\\varepsilon\-\\varepsilon\_\{0\}\)\)withθ:=αKτ/2\\theta:=\\alpha K\\tau/2andε0:=\(λ\+G\)/α\\varepsilon\_\{0\}:=\(\\lambda\+G\)/\\alphais the LSI bound \([4\.3](https://arxiv.org/html/2607.23029#S4.E3)\)\.
Plugging the optimal feedback controlε∗\(t,x;β\)\\varepsilon^\{\*\}\(t,x;\\beta\)derived below into \([5\.2](https://arxiv.org/html/2607.23029#S5.E2)\) gives the forward type\-conditional Fokker–Planck equation
∂tμtβ\(x\)\+∇⋅\[μtβ\(x\)b\(x,μt,ε∗\(t,x;β\)\)\]=ε∗\(t,x;β\)Δμtβ\(x\);\\partial\_\{t\}\\mu\_\{t\}^\{\\beta\}\(x\)\+\\nabla\\\!\\cdot\\\!\\Bigl\[\\mu\_\{t\}^\{\\beta\}\(x\)\\,b\\bigl\(x,\\mu\_\{t\},\\varepsilon^\{\*\}\(t,x;\\beta\)\\bigr\)\\Bigr\]=\\varepsilon^\{\*\}\(t,x;\\beta\)\\,\\Delta\\mu\_\{t\}^\{\\beta\}\(x\);\(5\.4\)marginalising overρ\\rhorecovers \([4\.1](https://arxiv.org/html/2607.23029#S4.E1)\) but with the optimal control in place of a fixedε\\varepsilon\. Defining the type\-conditional value functionV\(t,x;β\)=infε𝔼Xt=x\[⋅\]V\(t,x;\\beta\)=\\inf\_\{\\varepsilon\}\\mathbb\{E\}\_\{X\_\{t\}=x\}\[\\,\\cdot\\,\]of \([5\.3](https://arxiv.org/html/2607.23029#S5.E3)\), dynamic programming yields the backward HJB
−∂tV\(t,x;β\)=ℋ\(x,∇V,D2V,μt;β\),V\(T,x;β\)=0,\-\\partial\_\{t\}V\(t,x;\\beta\)=\\mathcal\{H\}\\bigl\(x,\\nabla V,D^\{2\}V,\\mu\_\{t\};\\beta\\bigr\),\\qquad V\(T,x;\\beta\)=0,\(5\.5\)with Hamiltonian
ℋ\(x,p,M,μ;β\)=ℓ\(x;wt\)−p⋅∇L\(x;wt\)−λp⋅\(x−x¯\)\+minε≥0\{βδdp\(ε\)−εσ2p⋅x\+εTr\(M\)\}\.\\mathcal\{H\}\(x,p,M,\\mu;\\beta\)=\\ell\(x;w\_\{t\}\)\-p\\\!\\cdot\\\!\\nabla L\(x;w\_\{t\}\)\-\\lambda\\,p\\\!\\cdot\\\!\(x\-\\bar\{x\}\)\+\\min\_\{\\varepsilon\\geq 0\}\\Bigl\\\{\\beta\\,\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon\)\-\\tfrac\{\\varepsilon\}\{\\sigma^\{2\}\}\\,p\\\!\\cdot\\\!x\+\\varepsilon\\,\\mathrm\{Tr\}\(M\)\\Bigr\\\}\.\(5\.6\)The first\-order condition for the inner minimisation inℋ\\mathcal\{H\}admits a closed\-form solutionε∗\(t,x;β\)\\varepsilon^\{\*\}\(t,x;\\beta\), derived in Appendix[C\.7](https://arxiv.org/html/2607.23029#A3.SS7)\. Two qualitative properties of this solution carry the intuition of the result:ε∗\\varepsilon^\{\*\}is decreasing inβ\\beta, so privacy\-sensitive clients adopt stronger regularisation, andε∗\\varepsilon^\{\*\}is decreasing in the local curvatureTr\(D2V\)\\mathrm\{Tr\}\(D^\{2\}V\), so clients near sharp minima can afford stronger noise\. The MFNE is the pair\(V,μtβ\)\(V,\\mu\_\{t\}^\{\\beta\}\)that simultaneously satisfies \([5\.5](https://arxiv.org/html/2607.23029#S5.E5)\) and \([5\.4](https://arxiv.org/html/2607.23029#S5.E4)\), coupled throughε∗\\varepsilon^\{\*\}in the FPK drift–diffusion and throughμt\\mu\_\{t\}in the Hamiltonian; integratingε∗\\varepsilon^\{\*\}against the equilibrium population recovers the scalar fixed pointε¯∗=Φ\(ε¯∗\)\\bar\{\\varepsilon\}^\{\*\}=\\Phi\(\\bar\{\\varepsilon\}^\{\*\}\)that drives the discrete\-grid solver in Section[6](https://arxiv.org/html/2607.23029#S6)\.
Two specialisations of the system above recover the existing literature\. Ifβk≡β\\beta\_\{k\}\\equiv\\betais constant across the population, the type\-conditional structure collapses,ε∗\\varepsilon^\{\*\}becomes spatially constant on the optimum, and \([5\.4](https://arxiv.org/html/2607.23029#S5.E4)\) reduces to the single\-strength entropic flow of\[[5](https://arxiv.org/html/2607.23029#bib.bib21)\]\. If we instead replace \([5\.4](https://arxiv.org/html/2607.23029#S5.E4)\) by itsNN\-particle empirical version and \([5\.5](https://arxiv.org/html/2607.23029#S5.E5)\) by theNN\-player backward Bellman system, withδdp\\delta\_\{\\mathrm\{dp\}\}replaced by RDP composition, we recover the dynamic MAPG\-DP of\[[34](https://arxiv.org/html/2607.23029#bib.bib15)\]\. MFPG is therefore the jointN→∞N\\\!\\to\\\!\\inftyand entropic\-mechanism limit of MAPG\-DP\.
When the activation conditionαε∗\>λ\+G\\alpha\\varepsilon^\{\*\}\>\\lambda\+Gholds at the equilibrium, the LSI contraction yields an exponentially decayingδdp\\delta\_\{\\mathrm\{dp\}\}bound of the form\(Cd/N\)exp\(−\(αε∗−λ−G\)Kτ/2\)\(C\_\{d\}/\\sqrt\{N\}\)\\exp\(\-\(\\alpha\\varepsilon^\{\*\}\-\\lambda\-G\)K\\tau/2\)\(Appendix[C\.8](https://arxiv.org/html/2607.23029#A3.SS8)\); compared with DP\-SGD, whoseϵdp\\epsilon\_\{\\mathrm\{dp\}\}grows asO\(K\)O\(\\sqrt\{K\}\)at fixedδ=10−5\\delta=10^\{\-5\}, MFPG’s privacy guarantee*tightens*with the number of training rounds\. This combines the heterogeneity benefit MFEP cannot express with the exponential decay MAPG\-DP cannot offer\.
## 6Algorithms
A single training round of every cell in Table[1](https://arxiv.org/html/2607.23029#S3.T1)consists of the same three steps applied independently by each client and then composed by the server: an*action update*that selects the strategic variableεk\\varepsilon\_\{k\}\(skipped in the no\-game cells\), a*local update*that advances the privatised state distributionμtk\\mu\_\{t\}^\{k\}and produces a parameter contributionwk\(t\+1\)w\_\{k\}^\{\(t\+1\)\}, and an*accountant update*that records the privacy cost\. The four named methods differ only in which version of these blocks they call\. Algorithm[1](https://arxiv.org/html/2607.23029#alg1)writes the outer loop once; the rest of this section specifies the three blocks in turn\.
Algorithm 1Unified training round \(one round, all four cells\)\.1:Global params
w\(t\)w^\{\(t\)\}; population mean
ε¯\(t\)\\bar\{\\varepsilon\}^\{\(t\)\}\(mean\-field cells only\); per\-client preferences
\{βk\}\\\{\\beta\_\{k\}\\\}; active mechanism \(Gaussian / Entropic\) and accountant \(RDP / LSI\)\.
2:foreach client
k∈\[N\]k\\in\[N\]in paralleldo
3:
εk\(t\)←\\varepsilon\_\{k\}^\{\(t\)\}\\leftarrowActionUpdate\(ε¯\(t\),\{βj\}\)k\{\}\_\{k\}\(\\bar\{\\varepsilon\}^\{\(t\)\},\\\{\\beta\_\{j\}\\\}\)⊳\\trianglerightskipped if no game
4:
wk\(t\+1\)←w\_\{k\}^\{\(t\+1\)\}\\leftarrowLocalUpdate\(w\(t\),εk\(t\)\)ℳ\{\}\_\{\\mathcal\{M\}\}\(w^\{\(t\)\},\\varepsilon\_\{k\}^\{\(t\)\}\)⊳\\trianglerightAlg\.[2](https://arxiv.org/html/2607.23029#alg2)or[3](https://arxiv.org/html/2607.23029#alg3)
5:
η←\\eta\\leftarrowAccountantUpdate\(η,εk\(t\)\)\(\\eta,\\varepsilon\_\{k\}^\{\(t\)\}\)⊳\\trianglerightRDP or LSI
6:endfor
7:
ε¯\(t\+1\)←N−1∑kεk\(t\)\\bar\{\\varepsilon\}^\{\(t\+1\)\}\\leftarrow N^\{\-1\}\\sum\_\{k\}\\varepsilon\_\{k\}^\{\(t\)\}⊳\\trianglerightmean\-field cells only
8:
w\(t\+1\)←∑kpkwk\(t\+1\)w^\{\(t\+1\)\}\\leftarrow\\sum\_\{k\}p\_\{k\}\\,w\_\{k\}^\{\(t\+1\)\}⊳\\trianglerightFedAvg
The*local\-update block*advances clientkk’s data slice under the active noise mechanism\. The two mechanisms supported by Table[1](https://arxiv.org/html/2607.23029#S3.T1)are the familiar additive Gaussian step \(used by DP\-SGD and MAPG\-DP\) and the Particle–Sinkhorn JKO step that discretises the entropic Wasserstein gradient flow \([4\.1](https://arxiv.org/html/2607.23029#S4.E1)\) \(used by MFEP and MFPG\); both consume the same inputs and produce a privatised local model\.
The Gaussian variant \(Algorithm[2](https://arxiv.org/html/2607.23029#alg2)\) clips the per\-sample gradient to normCCand adds isotropic noise of scaleσk=C2ln\(1\.25/δ\)/εk\\sigma\_\{k\}=C\\sqrt\{2\\ln\(1\.25/\\delta\)\}/\\varepsilon\_\{k\}, recovering the standard DP\-SGD update of\[[1](https://arxiv.org/html/2607.23029#bib.bib16)\]\. Across cells the only difference is whetherσk\\sigma\_\{k\}is shared \(DP\-SGD:εk≡εtgt\\varepsilon\_\{k\}\\equiv\\varepsilon\_\{\\mathrm\{tgt\}\}\) or set per client by the action\-update block \(MAPG\-DP\)\.
The entropic variant \(Algorithm[3](https://arxiv.org/html/2607.23029#alg3)\) instead approximates the JKO stepμk\+1=argminρ\{ℱλ\(ρ\)\+\(2τ\)−1W22\(ρ,μk\)\}\\mu\_\{k\+1\}=\\arg\\min\_\{\\rho\}\\\{\\mathcal\{F\}\_\{\\lambda\}\(\\rho\)\+\(2\\tau\)^\{\-1\}W\_\{2\}^\{2\}\(\\rho,\\mu\_\{k\}\)\\\}by a particle method: each particle takes a forward Euler step on the free\-energy gradient \(the drift–diffusion line\), and a single Sinkhorn projection enforces the Wasserstein constraint by averaging each particle against its barycentric image under the entropic optimal transport plan\. The drift contains three terms—loss\-driven, prior\-attractive, and variance\-penalising—all read off from \([3\.4](https://arxiv.org/html/2607.23029#S3.E4)\)\. The diffusion strength2εkτ\\sqrt\{2\\varepsilon\_\{k\}\\tau\}is what the LSI bound \([4\.3](https://arxiv.org/html/2607.23029#S4.E3)\) later contracts\. To keep the projection tractable on large parameter tensors we cap the Sinkhorn atn=512n\{=\}512flattened entries; tensors larger than this are advanced by drift–diffusion only, which preserves theδdp\\delta\_\{\\mathrm\{dp\}\}guarantee but skips the optimal\-transport refinement\.
Algorithm 2Gaussian local update \(DP\-SGD, MAPG\-DP\)\.1:Local model
wkw\_\{k\}; clip
CC; per\-client budget
εk\\varepsilon\_\{k\}\(or shared
εtgt\\varepsilon\_\{\\mathrm\{tgt\}\}for DP\-SGD\)\.
2:
σk←C2ln\(1\.25/δ\)/εk\\sigma\_\{k\}\\leftarrow C\\sqrt\{2\\ln\(1\.25/\\delta\)\}/\\varepsilon\_\{k\}
3:
g←clip\(∇Lk\(wk\),C\)g\\leftarrow\\mathrm\{clip\}\(\\nabla L\_\{k\}\(w\_\{k\}\),\\,C\)
4:
wk←wk−η\(g\+𝒩\(0,σk2I\)\)w\_\{k\}\\leftarrow w\_\{k\}\-\\eta\\,\(g\+\\mathcal\{N\}\(0,\\sigma\_\{k\}^\{2\}I\)\)
Algorithm 3Particle–Sinkhorn JKO step \(MFEP, MFPG\)\.1:Particles
\{xi\}i=1n\\\{x\_\{i\}\\\}\_\{i=1\}^\{n\}; JKO step
τ\\tau; clip
CC; entropic strength
εk\\varepsilon\_\{k\}; prior variance
σ2\\sigma^\{2\}; variance penalty
λ\\lambda; Sinkhorn regulariser
ηS\\eta\_\{\\mathrm\{S\}\}\.
2:
gi←clip\(∇L\(xi\),C\)g\_\{i\}\\leftarrow\\mathrm\{clip\}\(\\nabla L\(x\_\{i\}\),\\,C\)
3:
xi′←xi−τ\[gi\+εkσ2xi\+λ\(xi−x¯\)\]\+2εkτzix\_\{i\}^\{\\prime\}\\leftarrow x\_\{i\}\-\\tau\\\!\\left\[g\_\{i\}\+\\tfrac\{\\varepsilon\_\{k\}\}\{\\sigma^\{2\}\}x\_\{i\}\+\\lambda\(x\_\{i\}\-\\bar\{x\}\)\\right\]\+\\sqrt\{2\\varepsilon\_\{k\}\\tau\}\\,z\_\{i\},
zi∼𝒩\(0,I\)z\_\{i\}\\sim\\mathcal\{N\}\(0,I\)⊳\\trianglerightdrift \+ diffusion
4:
P←Sinkhorn\(𝟏n,𝟏n,Cij,ηS\)P\\leftarrow\\mathrm\{Sinkhorn\}\(\\mathbf\{1\}\_\{n\},\\mathbf\{1\}\_\{n\},\\,C\_\{ij\},\\,\\eta\_\{\\mathrm\{S\}\}\)with
Cij=‖xi′−xj′‖2C\_\{ij\}=\\\|x\_\{i\}^\{\\prime\}\-x\_\{j\}^\{\\prime\}\\\|^\{2\}⊳\\trianglerightentropic OT plan
5:
xinew←∑jPijxj′x\_\{i\}^\{\\mathrm\{new\}\}\\leftarrow\\sum\_\{j\}P\_\{ij\}\\,x\_\{j\}^\{\\prime\}⊳\\trianglerightbarycentric projection
The*action\-update block*is the only place where the four cells differ*algorithmically*once the mechanism is fixed: it is empty in the no\-game cells, a closed\-formNN\-player best response in MAPG\-DP, and a finite\-grid mean\-field best response in MFPG\. For MFPG the client solves
εknew←argminε∈ℰ\[Lkε\+c\+βkδdp\(ε;Gk\)\],\\varepsilon\_\{k\}^\{\\mathrm\{new\}\}\\leftarrow\\arg\\min\_\{\\varepsilon\\in\\mathcal\{E\}\}\\Bigl\[\\,\\tfrac\{L\_\{k\}\}\{\\varepsilon\+c\}\+\\beta\_\{k\}\\,\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon;\\,G\_\{k\}\)\\Bigr\],\(6\.1\)whereGkG\_\{k\}is the latest clipped gradient\-norm estimate,ccis a small constant that absorbs the low\-ε\\varepsilonsingularity of the regulariser, andδdp\\delta\_\{\\mathrm\{dp\}\}is computed via \([4\.3](https://arxiv.org/html/2607.23029#S4.E3)\) at the current mean\-fieldε¯\(t\)\\bar\{\\varepsilon\}^\{\(t\)\}\. The cost isO\(\|ℰ\|\)O\(\|\\mathcal\{E\}\|\)per client per round, dwarfed by the local update of Algorithm[3](https://arxiv.org/html/2607.23029#alg3)\. MAPG\-DP replaces \([6\.1](https://arxiv.org/html/2607.23029#S6.E1)\) with the closed\-form KKT best response of\[[34](https://arxiv.org/html/2607.23029#bib.bib15)\], and DP\-SGD and MFEP skip this block entirely\. Existence of a fixed point of the averaged best\-response mapΦ:ε¯↦N−1∑kεknew\\Phi:\\bar\{\\varepsilon\}\\mapsto N^\{\-1\}\\sum\_\{k\}\\varepsilon\_\{k\}^\{\\mathrm\{new\}\}is guaranteed by Proposition[C\.1](https://arxiv.org/html/2607.23029#A3.Thmthm1); in practice ten outer iterations suffice for the grid sizes\|ℰ\|≤5\|\\mathcal\{E\}\|\\leq 5we report\.
Each cell carries a privacy*accountant*that updates after every local step\. The Gaussian cells use a Rényi\-DP moments accountant tracking RDP at orderα=2\\alpha\{=\}2and converting to\(ϵ,δ\)\(\\epsilon,\\delta\)via the standard amplification\-by\-subsampling formula\[[21](https://arxiv.org/html/2607.23029#bib.bib18)\]; the cumulativeϵ\\epsilongrows asO\(T\)O\(\\sqrt\{T\}\)\. The entropic cells instead apply the LSI\-contraction bound \([4\.3](https://arxiv.org/html/2607.23029#S4.E3)\) at the currentεk\\varepsilon\_\{k\}\(MFEP\) orε¯\(t\)\\bar\{\\varepsilon\}^\{\(t\)\}\(MFPG\), which gives an exponentially decayingδdp\\delta\_\{\\mathrm\{dp\}\}whenever the activation conditionαε∗\>λ\+G\\alpha\\varepsilon^\{\*\}\>\\lambda\+Gis met\. Both accountants are black\-box and consume only\(σk\(\\sigma\_\{k\}orεk,K,τ\)\\varepsilon\_\{k\},K,\\tau\), so any method can be re\-audited under either accountant; the experiments report the accountant each method was originally designed to use\.
## 7Numerical Results
Our experiments verify that the four cells of Table[1](https://arxiv.org/html/2607.23029#S3.T1)produce the privacy decay each cell predicts and show that MFPG attains MFEP\-level utility at the population level while delivering a*personalised*privacy guarantee that single\-ε\\varepsilonMFEP cannot\. The full experimental setup—datasets, hyperparameters, and accountant configurations—is given in Appendix[D](https://arxiv.org/html/2607.23029#A4)\. Every numerical claim below is reproducible fromresults/full\_benchmark\.csv\(seed=42=42\);paper/cross\_check\.pyverifies that no number drifts away from the CSV\.
Table 2:Final\-round metrics fromresults/full\_benchmark\.csv\(seed=42=42\)\. For the entropic cells,ϵ\\epsilonis the linear\-budget value at the final round; for the Gaussian cells it is the value reported by the RDP accountant\. “−\-” marks settings we omit \(MAPG\-DP on MNIST\)\.Quadratic\(d=5d\{=\}5,T=10T\{=\}10\)Logistic\(d=20d\{=\}20,T=15T\{=\}15\)MNIST\(MLP,T=20T\{=\}20\)Methodloss—ϵ\\epsilonδ\\deltalossaccϵ\\epsilonδ\\deltaaccϵ\\epsilonδ\\deltaDP\-SGD69\.1—12\.310−510^\{\-5\}5\.780\.50014\.210−510^\{\-5\}0\.12023\.710−510^\{\-5\}MFEP74\.1—1\.009\.1×10−29\.1\{\\times\}10^\{\-2\}5\.670\.4601\.001\.0×10−11\.0\{\\times\}10^\{\-1\}0\.1261\.001\.5×10−11\.5\{\\times\}10^\{\-1\}MAPG\-DP65\.8—0\.3710−510^\{\-5\}4\.490\.3950\.0810−510^\{\-5\}———MFPG \(ours\)75\.3—1\.006\.1×10−16\.1\{\\times\}10^\{\-1\}6\.110\.4601\.004\.7×10−14\.7\{\\times\}10^\{\-1\}0\.0941\.001\.01\.0
Table[2](https://arxiv.org/html/2607.23029#S7.T2)reports the final\-round numbers, and Figures[1](https://arxiv.org/html/2607.23029#S7.F1)–[3](https://arxiv.org/html/2607.23029#S7.F3)show the round\-by\-round trajectories of loss, accuracy, andδdp\\delta\_\{\\mathrm\{dp\}\}\. The headline finding is that the two axes of Table[1](https://arxiv.org/html/2607.23029#S3.T1)have predictable, separable effects: the noise\-mechanism axis controls the shape ofδdp\\delta\_\{\\mathrm\{dp\}\}\(constant for Gaussian cells, decreasing for entropic cells\), and the heterogeneity axis controls how privacy budget is distributed across clients\.
The Gaussian cells confirm the polynomial accumulation of standard DP\. AfterT=10T\{=\}10quadratic rounds, DP\-SGD has reachedϵ=12\.3\\epsilon\{=\}12\.3atδ=10−5\\delta\{=\}10^\{\-5\}; on logistic regression and MNIST the cumulativeϵ\\epsilonrises to14\.214\.2and23\.723\.7respectively\. MAPG\-DP keepsϵ\\epsilonmuch smaller per round through its strategic per\-client budgets \(0\.370\.37on quadratic,0\.080\.08on logistic\) at the cost of utility on logistic regression \(39\.5%39\.5\\%vs\.50\.0%50\.0\\%for DP\-SGD\), but itsδ\\deltadoes not tighten with rounds\. The entropic cells follow the linear\-budget scheduleϵ=1\\epsilon\{=\}1for every round and exhibit a non\-trivialδ\\deltathat contracts when the LSI rateαε∗−λ−G\\alpha\\varepsilon^\{\*\}\-\\lambda\-Gis positive; withC=1,σ=1C\{=\}1,\\sigma\{=\}1the rate sits on the boundary of activation, so MFEP and MFPG fall back to the polynomial1/\(K\+1\)1/\(K\+1\)envelope on the simpler tasks\.
On utility, the two entropic cells are tied to within a fraction of a percentage point on logistic regression \(46\.0%46\.0\\%for both\) and within33points on MNIST \(12\.6%12\.6\\%vs\.9\.4%9\.4\\%\); on the convex quadratic problem they trail MAPG\-DP by about1010units of loss but match each other\. This is the expected behaviour: when the MFNE concentrates close to a singleε¯∗\\bar\{\\varepsilon\}^\{\*\}\(Appendix[E](https://arxiv.org/html/2607.23029#A5)\), the population\-level loss of MFPG is well\-approximated by MFEP at that strength\. Where the methods differ visibly is in the per\-roundδ\\deltafor the logistic experiment: MFEP reports1\.0×10−11\.0\\\!\\times\\\!10^\{\-1\}versus4\.7×10−14\.7\\\!\\times\\\!10^\{\-1\}for MFPG\. The MFPG bound is looser*at the population level*because privacy\-tolerant clients self\-select largerεk∗\\varepsilon\_\{k\}^\{\*\}, which inflates the mean\-fieldε¯\\bar\{\\varepsilon\}entering \([4\.3](https://arxiv.org/html/2607.23029#S4.E3)\); the per\-client guarantee for high\-β\\betaclients is correspondingly tighter than what MFEP can express at all\.
The MNIST stress test isolates a different regime\. With a∼\\sim100100\\,k parameter MLP, the JKO drift–diffusion noise2ετzi\\sqrt\{2\\varepsilon\\tau\}\\,z\_\{i\}dominates the gradient signal, and all three methods we ran \(DP\-SGD, MFEP, MFPG\) plateau near chance accuracy\. We omit MAPG\-DP from MNIST because its per\-sample best\-response loop is expensive at this scale and adds no insight beyond the logistic experiment\. We retain MNIST in the paper precisely because it makes the LSI bound’s failure mode visible: when the regimeαε∗\>λ\+G\\alpha\\varepsilon^\{\*\}\\\!\>\\\!\\lambda\+Gis far from satisfied, the entropic\-flow advantage over Gaussian noise vanishes and MFEP and MFPG converge to the same poor utility, consistent with the convex\-hull argument of Proposition[C\.1](https://arxiv.org/html/2607.23029#A3.Thmthm1)\. A scalable approximation of the Sinkhorn projection is the natural next step\.
Figure 1:Quadratic regression\. The two entropic cells \(MFEP, MFPG\) deliver a slowly decayingδdp\\delta\_\{\\mathrm\{dp\}\}at fixedϵ=1\\epsilon\{=\}1; both Gaussian cells \(DP\-SGD, MAPG\-DP\) deliver a flatδ\\deltaat higher cumulativeϵ\\epsilon\.Figure 2:Logistic regression\. MFPG matches MFEP on test accuracy at a fraction of DP\-SGD’s privacy cost; MAPG\-DP keepsϵ\\epsilonsmall at the cost of utility\.Figure 3:MNIST\. In the high\-dimensional regime where the LSI bound’s activation conditionαε∗\>λ\+G\\alpha\\varepsilon^\{\*\}\\\!\>\\\!\\lambda\+Gis loose, MFEP and MFPG converge to the same poor utility, while DP\-SGD reaches comparable accuracy at much higher cumulativeϵ\\epsilon\.
## References
- \[1\]M\. Abadi, A\. Chu, I\. Goodfellow, H\. B\. McMahan, I\. Mironov, K\. Talwar, and L\. Zhang\(2016\)Deep learning with differential privacy\.InProceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security,pp\. 308–318\.Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px1.p1.2),[§1](https://arxiv.org/html/2607.23029#S1.SS0.SSS0.Px1.p1.10.3.4.2),[§1](https://arxiv.org/html/2607.23029#S1.p2.5),[§2](https://arxiv.org/html/2607.23029#S2.SS0.SSS0.Px2.p1.10),[1st item](https://arxiv.org/html/2607.23029#S3.I2.i1.p1.4),[Table 1](https://arxiv.org/html/2607.23029#S3.T1.11.5.3.2.2),[§4\.1](https://arxiv.org/html/2607.23029#S4.SS1.p1.6),[§6](https://arxiv.org/html/2607.23029#S6.p3.4)\.
- \[2\]C\. D\. Aliprantis and K\. C\. Border\(2006\)Infinite dimensional analysis: a hitchhiker’s guide\.3 edition,Springer\.Cited by:[item \(ii\)](https://arxiv.org/html/2607.23029#A3.I1.i2.p1.1),[§C\.4](https://arxiv.org/html/2607.23029#A3.SS4.1.p1.10)\.
- \[3\]L\. Ambrosio, N\. Gigli, and G\. Savaré\(2008\)Gradient flows in metric spaces and in the space of probability measures\.2 edition,Birkhäuser\.Cited by:[§C\.2](https://arxiv.org/html/2607.23029#A3.SS2.SSS0.Px1.p1.10),[§C\.2](https://arxiv.org/html/2607.23029#A3.SS2.SSS0.Px2.p1.2)\.
- \[4\]G\. Andrew, O\. Thakkar, B\. McMahan, and S\. Ramaswamy\(2021\)Differentially private learning with adaptive clipping\.InAdvances in Neural Information Processing Systems,Vol\.34\.Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px1.p1.2)\.
- \[5\]Anonymous\(2025\)Mean\-field entropic privacy \(MFEP\): a unified dynamics framework for private federated learning\.Note:Under review at AISTATS 2026Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px3.p1.3),[Table 3](https://arxiv.org/html/2607.23029#A2.T3.11.7.3),[§C\.2](https://arxiv.org/html/2607.23029#A3.SS2.SSS0.Px3.p1.4),[§1](https://arxiv.org/html/2607.23029#S1.SS0.SSS0.Px1.p1.10.3.4.3),[§1](https://arxiv.org/html/2607.23029#S1.p3.8),[Table 1](https://arxiv.org/html/2607.23029#S3.T1.13.7.5.2.2),[§4\.2](https://arxiv.org/html/2607.23029#S4.SS2.p1.2),[§5\.3](https://arxiv.org/html/2607.23029#S5.SS3.p3.6)\.
- \[6\]C\. Berge\(1963\)Topological spaces\.Oliver & Boyd,Edinburgh and London\.Cited by:[§C\.4](https://arxiv.org/html/2607.23029#A3.SS4.1.p1.7)\.
- \[7\]N\. Carlini, F\. Tramer, E\. Wallace, M\. Jagielski, A\. Herbert\-Voss, K\. Lee, A\. Roberts, T\. Brown, D\. Song, U\. Erlingsson,et al\.\(2021\)Extracting training data from large language models\.In30th USENIX Security Symposium,pp\. 2633–2650\.Cited by:[§1](https://arxiv.org/html/2607.23029#S1.p1.1)\.
- \[8\]R\. Carmona and F\. Delarue\(2018\)Probabilistic theory of mean field games with applications I–II\.Springer\.Cited by:[§1](https://arxiv.org/html/2607.23029#S1.SS0.SSS0.Px1.p1.7),[§2](https://arxiv.org/html/2607.23029#S2.SS0.SSS0.Px3.p1.5),[5th item](https://arxiv.org/html/2607.23029#S3.I1.i5.p1.5),[§3](https://arxiv.org/html/2607.23029#S3.p1.1),[§5\.1](https://arxiv.org/html/2607.23029#S5.SS1.p1.12),[§5\.3](https://arxiv.org/html/2607.23029#S5.SS3.p1.8)\.
- \[9\]J\. Du, C\. Jiang, K\. Chen, Y\. Ren, and H\. V\. Poor\(2017\)Community\-structured evolutionary game for privacy protection in social networks\.IEEE Transactions on Information Forensics and Security13\(3\),pp\. 574–589\.Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px2.p1.5),[Table 3](https://arxiv.org/html/2607.23029#A2.T3.8.4.3),[§1](https://arxiv.org/html/2607.23029#S1.p3.8)\.
- \[10\]J\. Geiping, H\. Bauermeister, H\. Dröge, and M\. Moeller\(2020\)Inverting gradients–how easy is it to break privacy in federated learning?\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 16937–16947\.Cited by:[§1](https://arxiv.org/html/2607.23029#S1.p1.1)\.
- \[11\]R\. C\. Geyer, T\. Klein, and M\. Nabi\(2017\)Differentially private federated learning: a client level perspective\.InNeurIPS Workshop on Machine Learning on the Phone and other Consumer Devices,Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px1.p1.2),[§1](https://arxiv.org/html/2607.23029#S1.p2.5)\.
- \[12\]R\. Hu and Y\. Gong\(2020\)Trading data for learning: incentive mechanism for on\-device federated learning\.In2020 IEEE Global Communications Conference \(GLOBECOM\),pp\. 1–6\.Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px2.p1.5),[Table 3](https://arxiv.org/html/2607.23029#A2.T3.9.5.2)\.
- \[13\]M\. Huang, R\. P\. Malhamé, and P\. E\. Caines\(2006\)Large population stochastic dynamic games: closed\-loop McKean\-Vlasov systems and the Nash certainty equivalence principle\.InCommunications in Information & Systems,Vol\.6,pp\. 221–252\.Cited by:[§2](https://arxiv.org/html/2607.23029#S2.SS0.SSS0.Px3.p1.5),[5th item](https://arxiv.org/html/2607.23029#S3.I1.i5.p1.5),[§5\.1](https://arxiv.org/html/2607.23029#S5.SS1.p1.12)\.
- \[14\]R\. Jin, X\. He, and H\. Dai\(2017\)On the tradeoff between privacy and utility in collaborative intrusion detection systems\-a game theoretical approach\.InProceedings of the Hot Topics in Science of Security: Symposium and Bootcamp,pp\. 45–51\.Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px2.p1.5),[Table 3](https://arxiv.org/html/2607.23029#A2.T3.6.2.3)\.
- \[15\]R\. Jordan, D\. Kinderlehrer, and F\. Otto\(1998\)The variational formulation of the Fokker–Planck equation\.SIAM Journal on Mathematical Analysis29\(1\),pp\. 1–17\.Cited by:[§C\.2](https://arxiv.org/html/2607.23029#A3.SS2.SSS0.Px1.p1.12),[§1](https://arxiv.org/html/2607.23029#S1.p3.8),[§2](https://arxiv.org/html/2607.23029#S2.SS0.SSS0.Px4.p1.11)\.
- \[16\]J\. Lasry and P\. Lions\(2007\)Mean field games\.Japanese Journal of Mathematics2\(1\),pp\. 229–260\.Cited by:[§1](https://arxiv.org/html/2607.23029#S1.SS0.SSS0.Px1.p1.7),[§2](https://arxiv.org/html/2607.23029#S2.SS0.SSS0.Px3.p1.5),[§3](https://arxiv.org/html/2607.23029#S3.p1.1),[§5\.3](https://arxiv.org/html/2607.23029#S5.SS3.p1.8)\.
- \[17\]T\. Li, A\. K\. Sahu, M\. Zaheer, M\. Sanjabi, A\. Talwalkar, and V\. Smith\(2020\)Federated optimization in heterogeneous networks\.Proceedings of Machine Learning and Systems2,pp\. 429–450\.Cited by:[§1](https://arxiv.org/html/2607.23029#S1.p1.1)\.
- \[18\]J\. Liu, J\. Lou, L\. Xiong, J\. Liu, and X\. Meng\(2021\-12\)Projected federated averaging with heterogeneous differential privacy\.Proceedings of the VLDB Endowment15\(4\),pp\. 828–840\.External Links:ISSN 2150\-8097,[Document](https://dx.doi.org/10.14778/3503585.3503592)Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px1.p1.2)\.
- \[19\]B\. McMahan, E\. Moore, D\. Ramage, S\. Hampson, and B\. A\. y\. Arcas\(2017\)Communication\-efficient learning of deep networks from decentralized data\.InArtificial Intelligence and Statistics,pp\. 1273–1282\.Cited by:[§1](https://arxiv.org/html/2607.23029#S1.p1.1),[§2](https://arxiv.org/html/2607.23029#S2.SS0.SSS0.Px1.p1.9)\.
- \[20\]A\. Mehrjou\(2021\)Federated learning as a mean\-field game\.External Links:2107\.03770Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px3.p1.3),[§5\.3](https://arxiv.org/html/2607.23029#S5.SS3.p1.8)\.
- \[21\]I\. Mironov\(2017\)Rényi differential privacy\.In2017 IEEE 30th Computer Security Foundations Symposium \(CSF\),pp\. 263–275\.Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px1.p1.2),[§1](https://arxiv.org/html/2607.23029#S1.p2.5),[§6](https://arxiv.org/html/2607.23029#S6.p6.10)\.
- \[22\]F\. Otto and C\. Villani\(2000\)Generalization of an inequality by Talagrand and links with the logarithmic Sobolev inequality\.Journal of Functional Analysis173\(2\),pp\. 361–400\.Cited by:[§C\.2](https://arxiv.org/html/2607.23029#A3.SS2.SSS0.Px1.p1.10),[§1](https://arxiv.org/html/2607.23029#S1.p3.8),[2nd item](https://arxiv.org/html/2607.23029#S3.I2.i2.p1.5),[§4\.2](https://arxiv.org/html/2607.23029#S4.SS2.p1.8)\.
- \[23\]A\. Reisizadeh, F\. Farnia, R\. Pedarsani, and A\. Jadbabaie\(2020\)Robust federated learning: the case of affine distribution shifts\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px2.p1.5),[§1](https://arxiv.org/html/2607.23029#S1.p2.5),[§4\.3](https://arxiv.org/html/2607.23029#S4.SS3.p1.10)\.
- \[24\]S\. Rigot\(2023\)Entropic regularization of Wasserstein distance and applications to federated learning\.Machine Learning112\(4\),pp\. 1235–1261\.Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px3.p1.3)\.
- \[25\]L\. Ruthotto, S\. J\. Osher, W\. Li, L\. Nurbekyan, and S\. W\. Fung\(2020\)A machine learning framework for solving high\-dimensional mean field game and mean field control problems\.InProceedings of the National Academy of Sciences,Vol\.117,pp\. 9183–9193\.Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px3.p1.3)\.
- \[26\]A\. R\. Sfar, Y\. Challal, P\. Moyal, and E\. Natalizio\(2019\)A game theoretic approach for privacy preserving model in iot\-based transportation\.IEEE Transactions on Intelligent Transportation Systems20\(12\),pp\. 4405–4414\.Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px2.p1.5),[Table 3](https://arxiv.org/html/2607.23029#A2.T3.9.5.2)\.
- \[27\]Z\. Sun, L\. Yin, C\. Li, W\. Zhang, A\. Li, and Z\. Tian\(2020\)The qos and privacy trade\-off of adversarial deep learning: an evolutionary game approach\.Computers & Security96,pp\. 101876\.Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px2.p1.5),[Table 3](https://arxiv.org/html/2607.23029#A2.T3.9.5.2)\.
- \[28\]S\. Truex, N\. Baracaldo, A\. Anwar, T\. Steinke, H\. Ludwig, R\. Zhang, and Y\. Zhou\(2019\)A hybrid approach to privacy\-preserving federated learning\.InProceedings of the 12th ACM Workshop on Artificial Intelligence and Security,pp\. 1–11\.Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px1.p1.2)\.
- \[29\]C\. Villani\(2009\)Optimal transport: old and new\.Springer\.Cited by:[§C\.1](https://arxiv.org/html/2607.23029#A3.SS1.p2.3),[§C\.2](https://arxiv.org/html/2607.23029#A3.SS2.SSS0.Px1.p1.10)\.
- \[30\]K\. Wei, J\. Li, M\. Ding, C\. Ma, H\. H\. Yang, F\. Farokhi, S\. Jin, T\. Q\. Quek, and H\. V\. Poor\(2020\)Federated learning with differential privacy: algorithms and performance analysis\.Vol\.15,pp\. 3454–3469\.Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px1.p1.2),[§1](https://arxiv.org/html/2607.23029#S1.p2.5)\.
- \[31\]X\. Wu, T\. Wu, M\. K\. Khan, Q\. Ni, and W\. Dou\(2017\)Game theory based correlated privacy preserving analysis in big data\.IEEE Transactions on Big Data7\(4\),pp\. 643–656\.Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px2.p1.5),[Table 3](https://arxiv.org/html/2607.23029#A2.T3.6.2.3)\.
- \[32\]L\. Xiao, Y\. Li, G\. Han, H\. Dai, and H\. V\. Poor\(2017\)A secure mobile crowdsensing game with deep reinforcement learning\.IEEE Transactions on Information Forensics and Security13\(1\),pp\. 35–47\.Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px2.p1.5),[Table 3](https://arxiv.org/html/2607.23029#A2.T3.8.4.3)\.
- \[33\]L\. Xu, C\. Jiang, Y\. Qian, J\. Li, Y\. Zhao, and Y\. Ren\(2021\)Privacy\-accuracy trade\-off in differentially\-private distributed classification: a game theoretical approach\.IEEE Transactions on Big Data7\(4\),pp\. 770–783\.Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px2.p1.5),[Table 3](https://arxiv.org/html/2607.23029#A2.T3.6.2.3),[§1](https://arxiv.org/html/2607.23029#S1.p3.8)\.
- \[34\]L\. Yin, S\. Lin, Z\. Sun, R\. Li, Y\. He, and Z\. Hao\(2021\)A game\-theoretic approach for federated learning: a trade\-off among privacy, accuracy and energy\.Digital Communications and Networks\.Cited by:[Appendix A](https://arxiv.org/html/2607.23029#A1.SS0.SSS0.Px2.p1.5),[Table 3](https://arxiv.org/html/2607.23029#A2.T3.9.5.2),[§C\.3](https://arxiv.org/html/2607.23029#A3.SS3.p1.10),[§1](https://arxiv.org/html/2607.23029#S1.SS0.SSS0.Px1.p1.10.3.3.2),[§1](https://arxiv.org/html/2607.23029#S1.p2.5),[§1](https://arxiv.org/html/2607.23029#S1.p3.8),[Table 1](https://arxiv.org/html/2607.23029#S3.T1.17.11.4.3.3),[Remark 3\.1](https://arxiv.org/html/2607.23029#S3.Thmrem1.p1.5),[§4\.3](https://arxiv.org/html/2607.23029#S4.SS3.p1.10),[§4\.3](https://arxiv.org/html/2607.23029#S4.SS3.p2.5),[§5\.3](https://arxiv.org/html/2607.23029#S5.SS3.p3.6),[§6](https://arxiv.org/html/2607.23029#S6.p5.9)\.
- \[35\]L\. Zhu, Z\. Liu, and S\. Han\(2019\)Deep leakage from gradients\.InAdvances in Neural Information Processing Systems,Vol\.32\.Cited by:[§1](https://arxiv.org/html/2607.23029#S1.p1.1)\.
## Appendix ARelated Work
Each of the four cells in Table[1](https://arxiv.org/html/2607.23029#S3.T1)traces back to a distinct line of prior work; our contribution is to organise them along the two axes ofNNand client heterogeneity rather than to introduce new machinery\.
#### Differential privacy in federated learning\.
The finite\-NN, no\-game cell is occupied by a substantial literature\. DP\-SGD\[[1](https://arxiv.org/html/2607.23029#bib.bib16)\]introduced the per\-sample\-clipped Gaussian mechanism, and DP\-FedAvg\[[11](https://arxiv.org/html/2607.23029#bib.bib32),[30](https://arxiv.org/html/2607.23029#bib.bib33)\]adapted it to the federated setting\. Subsequent work tightens the privacy accountant via Rényi DP\[[21](https://arxiv.org/html/2607.23029#bib.bib18)\], allows adaptive clipping\[[4](https://arxiv.org/html/2607.23029#bib.bib24)\], combines DP with secure aggregation\[[28](https://arxiv.org/html/2607.23029#bib.bib34)\], and accommodates per\-client privacy budgets through projected averaging\[[18](https://arxiv.org/html/2607.23029#bib.bib5)\]\. Each of these methods operates on the data side \(gradients or samples\) under RDP composition, which is exactly the cell we recover when MFPG is reduced to homogeneousβk\\beta\_\{k\}and the entropic mechanism is swapped for additive Gaussian noise\.
#### Game\-theoretic privacy\.
The finite\-NN, game cell has been studied through both cooperative and non\-cooperative formulations: community\-structured evolutionary games\[[9](https://arxiv.org/html/2607.23029#bib.bib6)\], correlated privacy analysis\[[31](https://arxiv.org/html/2607.23029#bib.bib7)\], deep\-RL crowdsensing\[[32](https://arxiv.org/html/2607.23029#bib.bib8)\], IoT transportation\[[26](https://arxiv.org/html/2607.23029#bib.bib9)\], distributed classification\[[33](https://arxiv.org/html/2607.23029#bib.bib10)\], on\-device incentive mechanisms\[[12](https://arxiv.org/html/2607.23029#bib.bib11)\], evolutionary QoS trade\-offs\[[27](https://arxiv.org/html/2607.23029#bib.bib12)\], and intrusion\-detection games\[[14](https://arxiv.org/html/2607.23029#bib.bib13)\]\. The most directly relevant precedent is the multi\-agent privacy game\[[34](https://arxiv.org/html/2607.23029#bib.bib15)\], whose MAPG\-DP and MAPG\-input formulations we adopt as our finite\-NNbaseline; we discharge the model\-output variant \(MAPG\-parameter\) as aδk=0\\delta\_\{k\}=0pathology in Section[4\.3](https://arxiv.org/html/2607.23029#S4.SS3), since only state\-acting controls admit theN→∞N\\to\\inftylimit we develop\. The robust\-FL framework FLRA\[[23](https://arxiv.org/html/2607.23029#bib.bib4)\]is not a privacy method but its state\-acting affine perturbation\(Λi𝐱\+δi\)\(\\Lambda^\{i\}\\mathbf\{x\}\+\\delta^\{i\}\)is the structural precedent for treating the state distribution as the strategic object of FL\.
#### Mean\-field methods and entropic privacy\.
The FL–mean\-field analogy through coupled HJB and Fokker–Planck PDEs\[[20](https://arxiv.org/html/2607.23029#bib.bib20)\]provides the dynamical foundation we build on\. High\-dimensional mean\-field control with neural networks is studied in\[[25](https://arxiv.org/html/2607.23029#bib.bib35)\], and entropic\-Wasserstein regularisation in FL in\[[24](https://arxiv.org/html/2607.23029#bib.bib36)\], but neither addresses the privacy game\. The closest prior work is MFEP\[[5](https://arxiv.org/html/2607.23029#bib.bib21)\], which occupies theN→∞N\\to\\infty, no\-game cell of Table[1](https://arxiv.org/html/2607.23029#S3.T1)with a homogeneous entropic strength and the LSI\-based privacy analysis we adopt for ourδ\\deltabound\. MFPG is the natural generalisation of MFEP that admits heterogeneous client preferences, equivalently, theN→∞N\\to\\inftylimit of MAPG\-DP under the entropic mechanism\.
## Appendix BGame\-Theoretic Privacy Frameworks
Table 3:Comparison of game\-theoretic privacy frameworks\.NN\-N =NN\-player Nash; MFNE = mean\-field Nash equilibrium\.WorkPlayerHierarch\.Coop\.StrategyUtility\[[33](https://arxiv.org/html/2607.23029#bib.bib10),[14](https://arxiv.org/html/2607.23029#bib.bib13),[31](https://arxiv.org/html/2607.23029#bib.bib7)\]ClientsNoYesϵi\\epsilon\_\{i\}maxQ\(ϵ\)\\max Q\(\\boldsymbol\{\\epsilon\}\)\[[9](https://arxiv.org/html/2607.23029#bib.bib6),[32](https://arxiv.org/html/2607.23029#bib.bib8)\]ClientsNoNoϵi\\epsilon\_\{i\}maxQ\+P\(ϵi\)\\max Q\+P\(\\epsilon\_\{i\}\)\[[26](https://arxiv.org/html/2607.23029#bib.bib9),[27](https://arxiv.org/html/2607.23029#bib.bib12),[12](https://arxiv.org/html/2607.23029#bib.bib11),[34](https://arxiv.org/html/2607.23029#bib.bib15)\]S\+CYesNoϵi,b\\epsilon\_\{i\},bStackelbergMFEP\[[5](https://arxiv.org/html/2607.23029#bib.bib21)\]———none \(ε\\varepsilonfixed\)L\+εKLL\+\\varepsilon\\,\\mathrm\{KL\}MFPG \(ours\)Clients \(MF\)NoNoεk∈ℰ\\varepsilon\_\{k\}\\in\\mathcal\{E\}Lk\+βkδdp\(εk\)L\_\{k\}\+\\beta\_\{k\}\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon\_\{k\}\)
## Appendix CDerivations and proofs
This appendix supplies the derivations behind each labelled equation and theorem of the body\. Section references in parentheses point to the statement being proved\.
### C\.1Wasserstein gradient flow yields the Fokker–Planck equation \([4\.1](https://arxiv.org/html/2607.23029#S4.E1)\) \(§[4\.2](https://arxiv.org/html/2607.23029#S4.SS2)\)
We compute the first variation of the free energy \([3\.4](https://arxiv.org/html/2607.23029#S3.E4)\) and use the standard correspondence between Wasserstein gradient flows onℱ\\mathcal\{F\}and continuity equations driven by∇\(δℱ/δμ\)\\nabla\(\\delta\\mathcal\{F\}/\\delta\\mu\)\.
The three terms ofℱλ\\mathcal\{F\}\_\{\\lambda\}have first variations
δδμ𝔼μ\[L\(x;w\)\]\\displaystyle\\frac\{\\delta\}\{\\delta\\mu\}\\,\\mathbb\{E\}\_\{\\mu\}\[L\(x;w\)\]=L\(x;w\),\\displaystyle=L\(x;w\),δδμKL\(μ∥ν\)\\displaystyle\\frac\{\\delta\}\{\\delta\\mu\}\\,\\mathrm\{KL\}\(\\mu\\\|\\nu\)=logdμdν\(x\),\\displaystyle=\\log\\\!\\frac\{d\\mu\}\{d\\nu\}\(x\),δδμVarμ\[x\]\\displaystyle\\frac\{\\delta\}\{\\delta\\mu\}\\,\\mathrm\{Var\}\_\{\\mu\}\[x\]=‖x−x¯‖2−Varμ\[x\],x¯=𝔼μ\[x\]\.\\displaystyle=\\\|x\-\\bar\{x\}\\\|^\{2\}\-\\mathrm\{Var\}\_\{\\mu\}\[x\],\\qquad\\bar\{x\}=\\mathbb\{E\}\_\{\\mu\}\[x\]\.The first two are classical\[[29](https://arxiv.org/html/2607.23029#bib.bib25)\]; the third follows fromVarμ\[x\]=𝔼μ‖x‖2−‖𝔼μx‖2\\mathrm\{Var\}\_\{\\mu\}\[x\]=\\mathbb\{E\}\_\{\\mu\}\\\|x\\\|^\{2\}\-\\\|\\mathbb\{E\}\_\{\\mu\}x\\\|^\{2\}by direct computation\. Taking gradients inxx,
∇δℱλδμ\(x\)=∇L\(x;w\)\+ε∇logdμdν\(x\)\+λ\(x−x¯\)\.\\nabla\\frac\{\\delta\\mathcal\{F\}\_\{\\lambda\}\}\{\\delta\\mu\}\(x\)=\\nabla L\(x;w\)\+\\varepsilon\\,\\nabla\\log\\\!\\frac\{d\\mu\}\{d\\nu\}\(x\)\+\\lambda\(x\-\\bar\{x\}\)\.For the Gaussian priorν=𝒩\(0,σ2I\)\\nu=\\mathcal\{N\}\(0,\\sigma^\{2\}I\),∇logν\(x\)=−x/σ2\\nabla\\log\\nu\(x\)=\-x/\\sigma^\{2\}, so∇log\(dμ/dν\)=∇logμ\+x/σ2\\nabla\\log\(d\\mu/d\\nu\)=\\nabla\\log\\mu\+x/\\sigma^\{2\}\. Substituting,
∇δℱλδμ=∇L\+εσ2x\+λ\(x−x¯\)\+ε∇logμ\.\\nabla\\frac\{\\delta\\mathcal\{F\}\_\{\\lambda\}\}\{\\delta\\mu\}=\\nabla L\+\\tfrac\{\\varepsilon\}\{\\sigma^\{2\}\}x\+\\lambda\(x\-\\bar\{x\}\)\+\\varepsilon\\,\\nabla\\log\\mu\.The Wasserstein gradient flow∂tμt=∇⋅\(μt∇\(δℱ/δμ\)\)\\partial\_\{t\}\\mu\_\{t\}=\\nabla\\\!\\cdot\\\!\(\\mu\_\{t\}\\nabla\(\\delta\\mathcal\{F\}/\\delta\\mu\)\)then becomes
∂tμt=∇⋅\[μt\(∇L\+εσ2x\+λ\(x−x¯\)\)\]\+ε∇⋅\(μt∇logμt\)\.\\partial\_\{t\}\\mu\_\{t\}=\\nabla\\\!\\cdot\\\!\\Bigl\[\\mu\_\{t\}\\bigl\(\\nabla L\+\\tfrac\{\\varepsilon\}\{\\sigma^\{2\}\}x\+\\lambda\(x\-\\bar\{x\}\)\\bigr\)\\Bigr\]\+\\varepsilon\\,\\nabla\\\!\\cdot\\\!\(\\mu\_\{t\}\\nabla\\log\\mu\_\{t\}\)\.Using the identityμt∇logμt=∇μt\\mu\_\{t\}\\nabla\\log\\mu\_\{t\}=\\nabla\\mu\_\{t\}, the last term simplifies toεΔμt\\varepsilon\\Delta\\mu\_\{t\}, recovering \([4\.1](https://arxiv.org/html/2607.23029#S4.E1)\)\. ∎
### C\.2LSI contraction implies the privacy bound \([4\.3](https://arxiv.org/html/2607.23029#S4.E3)\) \(§[4\.2](https://arxiv.org/html/2607.23029#S4.SS2)\)
The argument is a standard composition of three steps: log\-Sobolev contraction of the entropic flow, total\-variation control by Wasserstein distance, and a1/N1/Ninitial gap\.
#### Step 1 \(LSI⇒\\RightarrowW2W\_\{2\}contraction\)\.
The Hessian ofℱλ\\mathcal\{F\}\_\{\\lambda\}in the Wasserstein sense decomposes as
Hessμℱλ=Hessμ𝔼μ\[L\]⏟⪰−GId\+εHessμKL\(⋅∥ν\)⏟⪰αεId\+λ2HessμVar⏟⪰−λId,\\mathrm\{Hess\}\_\{\\mu\}\\,\\mathcal\{F\}\_\{\\lambda\}=\\underbrace\{\\mathrm\{Hess\}\_\{\\mu\}\\,\\mathbb\{E\}\_\{\\mu\}\[L\]\}\_\{\\succeq\-G\\,\\mathrm\{Id\}\}\+\\underbrace\{\\varepsilon\\,\\mathrm\{Hess\}\_\{\\mu\}\\,\\mathrm\{KL\}\(\\cdot\\\|\\nu\)\}\_\{\\succeq\\alpha\\varepsilon\\,\\mathrm\{Id\}\}\+\\underbrace\{\\tfrac\{\\lambda\}\{2\}\\mathrm\{Hess\}\_\{\\mu\}\\,\\mathrm\{Var\}\}\_\{\\succeq\-\\lambda\\,\\mathrm\{Id\}\},where the loss term contributes−GId\-G\\,\\mathrm\{Id\}when‖∇L‖∞≤G\\\|\\nabla L\\\|\_\{\\infty\}\\leq G, the KL term contributesαεId\\alpha\\varepsilon\\,\\mathrm\{Id\}by Otto–Villani applied to LSI\(α\\alpha\) forν\\nu\[[22](https://arxiv.org/html/2607.23029#bib.bib27)\], and the variance penalty contributes−λId\-\\lambda\\,\\mathrm\{Id\}\[[3](https://arxiv.org/html/2607.23029#bib.bib41)\]\. The overall Wasserstein convexity constant isr:=αε−λ−Gr:=\\alpha\\varepsilon\-\\lambda\-G\. By the standard contraction result forrr\-displacement\-convex functionals on𝒫2\\mathcal\{P\}\_\{2\}\[[29](https://arxiv.org/html/2607.23029#bib.bib25), Thm\. 23\.9\],
W2\(μt,μt′\)≤e−rtW2\(μ0,μ0′\)wheneverr\>0\.W\_\{2\}\(\\mu\_\{t\},\\mu\_\{t\}^\{\\prime\}\)\\leq e^\{\-rt\}\\,W\_\{2\}\(\\mu\_\{0\},\\mu\_\{0\}^\{\\prime\}\)\\qquad\\text\{whenever \}r\>0\.The same rate transfers to the JKO discretisation with stepτ\\tau\[[15](https://arxiv.org/html/2607.23029#bib.bib26)\]: afterKKsteps,
W2\(μK,μK′\)≤e−rKτW2\(μ0,μ0′\)\.W\_\{2\}\(\\mu\_\{K\},\\mu\_\{K\}^\{\\prime\}\)\\leq e^\{\-rK\\tau\}\\,W\_\{2\}\(\\mu\_\{0\},\\mu\_\{0\}^\{\\prime\}\)\.\(C\.1\)
#### Step 2 \(W2W\_\{2\}to total variation\)\.
For absolutely continuous measures with bounded second moments, the transportation inequality givesTV\(μ,μ′\)2≤Cd2W2\(μ,μ′\)\\mathrm\{TV\}\(\\mu,\\mu^\{\\prime\}\)^\{2\}\\leq C\_\{d\}^\{2\}\\,W\_\{2\}\(\\mu,\\mu^\{\\prime\}\)withCd=min\(d,10\)C\_\{d\}=\\min\(\\sqrt\{d\},10\)\[[3](https://arxiv.org/html/2607.23029#bib.bib41)\], equivalently
TV\(μ,μ′\)≤CdW2\(μ,μ′\)\.\\mathrm\{TV\}\(\\mu,\\mu^\{\\prime\}\)\\leq C\_\{d\}\\sqrt\{W\_\{2\}\(\\mu,\\mu^\{\\prime\}\)\}\.\(C\.2\)
#### Step 3 \(initial1/N1/Ngap\)\.
For neighbouring datasets differing in a single client’s contribution out ofNN, the corresponding initial population measures satisfyW2\(μ0,μ0′\)≤D2/NW\_\{2\}\(\\mu\_\{0\},\\mu\_\{0\}^\{\\prime\}\)\\leq D^\{2\}/Nfor some data\-domain constantDDfolded intoCdC\_\{d\}\[[5](https://arxiv.org/html/2607.23029#bib.bib21)\]\.
#### Combination\.
Substituting Step 3 into \([C\.1](https://arxiv.org/html/2607.23029#A3.E1)\) and that into \([C\.2](https://arxiv.org/html/2607.23029#A3.E2)\),
TV\(μK,μK′\)≤Cde−rKτ⋅1/N=CdNe−rKτ/2\.\\mathrm\{TV\}\(\\mu\_\{K\},\\mu\_\{K\}^\{\\prime\}\)\\leq C\_\{d\}\\,\\sqrt\{e^\{\-rK\\tau\}\\cdot 1/N\}=\\frac\{C\_\{d\}\}\{\\sqrt\{N\}\}\\,e^\{\-rK\\tau/2\}\.Sinceδdp≤TV\(μK,μK′\)\\delta\_\{\\mathrm\{dp\}\}\\leq\\mathrm\{TV\}\(\\mu\_\{K\},\\mu\_\{K\}^\{\\prime\}\)in this neighbouring\-data formulation, we obtain \([4\.3](https://arxiv.org/html/2607.23029#S4.E3)\)\. ∎
### C\.3KKT analysis: theδk=0\\delta\_\{k\}=0collapse \(§[4\.3](https://arxiv.org/html/2607.23029#S4.SS3)\)
We expand the partial derivatives of the Lagrangian summarised in the body text\. The static model\-output MAPG of\[[34](https://arxiv.org/html/2607.23029#bib.bib15)\]solves
min𝐰k,𝜹k12nk∑i=1nk\(𝐖⊤𝐗ik−yik\)2−βk‖𝜹k‖22s\.t\.𝐖=1K∑j=1K\(𝐰j\+𝜹j\)\.\\min\_\{\\mathbf\{w\}\_\{k\},\\boldsymbol\{\\delta\}\_\{k\}\}\\;\\tfrac\{1\}\{2n\_\{k\}\}\\sum\_\{i=1\}^\{n\_\{k\}\}\(\\mathbf\{W\}^\{\\top\}\\mathbf\{X\}\_\{i\}^\{k\}\-y\_\{i\}^\{k\}\)^\{2\}\-\\beta\_\{k\}\\\|\\boldsymbol\{\\delta\}\_\{k\}\\\|\_\{2\}^\{2\}\\quad\\mathrm\{s\.t\.\}\\;\\;\\mathbf\{W\}=\\tfrac\{1\}\{K\}\\sum\_\{j=1\}^\{K\}\(\\mathbf\{w\}\_\{j\}\+\\boldsymbol\{\\delta\}\_\{j\}\)\.Forming the Lagrangian with multiplier𝝀k∈ℝd\\boldsymbol\{\\lambda\}\_\{k\}\\in\\mathbb\{R\}^\{d\},
ℒk=12nk∑i\(𝐖⊤𝐗ik−yik\)2−βk‖𝜹k‖22−𝝀k⊤\[𝐖−1K∑j\(𝐰j\+𝜹j\)\]\.\\mathcal\{L\}\_\{k\}=\\tfrac\{1\}\{2n\_\{k\}\}\\\!\\sum\_\{i\}\(\\mathbf\{W\}^\{\\top\}\\mathbf\{X\}\_\{i\}^\{k\}\-y\_\{i\}^\{k\}\)^\{2\}\-\\beta\_\{k\}\\\|\\boldsymbol\{\\delta\}\_\{k\}\\\|\_\{2\}^\{2\}\-\\boldsymbol\{\\lambda\}\_\{k\}^\{\\top\}\\\!\\Bigl\[\\mathbf\{W\}\-\\tfrac\{1\}\{K\}\\\!\\sum\_\{j\}\(\\mathbf\{w\}\_\{j\}\+\\boldsymbol\{\\delta\}\_\{j\}\)\\Bigr\]\.Computing partials in feature dimensionm∈\{1,…,d\}m\\in\\\{1,\\dots,d\\\}:
∂ℒk∂Wm\\displaystyle\\frac\{\\partial\\mathcal\{L\}\_\{k\}\}\{\\partial W\_\{m\}\}=1nk∑i\(𝐖⊤𝐗ik−yik\)xi,mk−λk,m,\\displaystyle=\\tfrac\{1\}\{n\_\{k\}\}\\\!\\sum\_\{i\}\(\\mathbf\{W\}^\{\\top\}\\mathbf\{X\}\_\{i\}^\{k\}\-y\_\{i\}^\{k\}\)\\,x\_\{i,m\}^\{k\}\-\\lambda\_\{k,m\},\(A\.1\)∂ℒk∂wk,m\\displaystyle\\frac\{\\partial\\mathcal\{L\}\_\{k\}\}\{\\partial w\_\{k,m\}\}=1Kλk,m,\\displaystyle=\\tfrac\{1\}\{K\}\\,\\lambda\_\{k,m\},\(A\.2\)∂ℒk∂δk,m\\displaystyle\\frac\{\\partial\\mathcal\{L\}\_\{k\}\}\{\\partial\\delta\_\{k,m\}\}=1Kλk,m−2βkδk,m\.\\displaystyle=\\tfrac\{1\}\{K\}\\,\\lambda\_\{k,m\}\-2\\beta\_\{k\}\\,\\delta\_\{k,m\}\.\(A\.3\)The KKT stationarity conditions set each partial to zero\. From \(A\.2\),λk,m=0\\lambda\_\{k,m\}=0for everymm\. Substitutingλk,m=0\\lambda\_\{k,m\}=0into \(A\.3\) yields−2βkδk,m=0\-2\\beta\_\{k\}\\,\\delta\_\{k,m\}=0, hence𝜹k=𝟎\\boldsymbol\{\\delta\}\_\{k\}=\\mathbf\{0\}for everykksinceβk\>0\\beta\_\{k\}\>0\. The KKT point is unique modulo regularity of the data block\. ∎
### C\.4MFNE: definition and existence \(§[5\.2](https://arxiv.org/html/2607.23029#S5.SS2)\)
###### Definition C\.1\(Mean\-field Nash equilibrium\)\.
A pair\(μ∗,ε∗\)\(\\mu^\{\*\},\\varepsilon^\{\*\}\)is a*mean\-field Nash equilibrium*\(MFNE\) of MFPG if \(i\)μ∗\\mu^\{\*\}is the stationary state distribution of \([4\.1](https://arxiv.org/html/2607.23029#S4.E1)\) at the population meanε¯∗=𝔼μ∗\[ε\]\\bar\{\\varepsilon\}^\{\*\}=\\mathbb\{E\}\_\{\\mu^\{\*\}\}\[\\varepsilon\], and \(ii\)ε∗∈argminε∈ℰUk\(ε;ε¯∗\)\\varepsilon^\{\*\}\\in\\arg\\min\_\{\\varepsilon\\in\\mathcal\{E\}\}U\_\{k\}\(\\varepsilon;\\bar\{\\varepsilon\}^\{\*\}\)for every clientkk\.
###### Proposition C\.1\(Existence\)\.
Ifℰ\\mathcal\{E\}is finite andUkU\_\{k\}is continuous inε\\varepsilon, then the averaged best\-response mapΦ:ε¯↦N−1∑kBRk\(ε¯\)\\Phi:\\bar\{\\varepsilon\}\\mapsto N^\{\-1\}\\sum\_\{k\}\\mathrm\{BR\}\_\{k\}\(\\bar\{\\varepsilon\}\)admits a fixed point on the convex hull ofℰ\\mathcal\{E\}\.
###### Proof\.
Letℰ=\{ε\(1\),…,ε\(m\)\}\\mathcal\{E\}=\\\{\\varepsilon^\{\(1\)\},\\ldots,\\varepsilon^\{\(m\)\}\\\}and letE:=\[minℰ,maxℰ\]=conv\(ℰ\)E:=\[\\min\\mathcal\{E\},\\max\\mathcal\{E\}\]=\\mathrm\{conv\}\(\\mathcal\{E\}\)\. For each clientkk, the disutilityUk\(⋅,ε¯\)U\_\{k\}\(\\cdot,\\bar\{\\varepsilon\}\)is continuous in the second argument by inspection of \([3\.3](https://arxiv.org/html/2607.23029#S3.E3)\) and \([4\.3](https://arxiv.org/html/2607.23029#S4.E3)\)\. The set\-valued best response
BRk\(ε¯\):=argminε∈ℰUk\(ε;ε¯\)⊆ℰ\\mathrm\{BR\}\_\{k\}\(\\bar\{\\varepsilon\}\):=\\arg\\min\_\{\\varepsilon\\in\\mathcal\{E\}\}\\,U\_\{k\}\(\\varepsilon;\\bar\{\\varepsilon\}\)\\subseteq\\mathcal\{E\}is therefore upper hemi\-continuous inε¯\\bar\{\\varepsilon\}onEE\(Berge’s maximum theorem\[[6](https://arxiv.org/html/2607.23029#bib.bib39)\]\)\. Define the averaged correspondenceΦ:E→2E\\Phi:E\\to 2^\{E\}by
Φ\(ε¯\):=1N∑k=1NconvBRk\(ε¯\),\\Phi\(\\bar\{\\varepsilon\}\):=\\tfrac\{1\}\{N\}\\\!\\sum\_\{k=1\}^\{N\}\\,\\mathrm\{conv\}\\,\\mathrm\{BR\}\_\{k\}\(\\bar\{\\varepsilon\}\),where the right\-hand side is the Minkowski average of convex hulls\. Three properties hold:
1. \(i\)*Convex\-valued:*eachconvBRk\(ε¯\)\\mathrm\{conv\}\\,\\mathrm\{BR\}\_\{k\}\(\\bar\{\\varepsilon\}\)is convex, and Minkowski sums of convex sets are convex\.
2. \(ii\)*Upper hemi\-continuous:*the convex\-hull operator preserves upper hemi\-continuity\[[2](https://arxiv.org/html/2607.23029#bib.bib40), Thm\. 17\.35\], and the Minkowski average of upper hemi\-continuous correspondences is upper hemi\-continuous\.
3. \(iii\)*Self\-mapping:*every value inΦ\(ε¯\)\\Phi\(\\bar\{\\varepsilon\}\)is a convex combination of points inℰ⊂E\\mathcal\{E\}\\subset E, soΦ\(ε¯\)⊆E\\Phi\(\\bar\{\\varepsilon\}\)\\subseteq E\.
SinceEEis a non\-empty compact convex subset ofℝ\\mathbb\{R\}, Kakutani’s fixed\-point theorem\[[2](https://arxiv.org/html/2607.23029#bib.bib40)\]guarantees a fixed pointε¯∗∈Φ\(ε¯∗\)\\bar\{\\varepsilon\}^\{\*\}\\in\\Phi\(\\bar\{\\varepsilon\}^\{\*\}\), which is the population mean of an MFNE strategy profile\. ∎
### C\.5Forward FPK from the controlled SDE \(eq\. \([5\.4](https://arxiv.org/html/2607.23029#S5.E4)\)\)
Fix a typeβ\\betaand a feedback controlε∗\(t,x;β\)\\varepsilon^\{\*\}\(t,x;\\beta\)\. Under the controlled SDE \([5\.2](https://arxiv.org/html/2607.23029#S5.E2)\), Itô’s formula applied to a test functionφ∈Cc2\(ℝd\)\\varphi\\in C\_\{c\}^\{2\}\(\\mathbb\{R\}^\{d\}\)gives
dφ\(Xt\)=\(∇φ⋅b\+ε∗Δφ\)dt\+2ε∗∇φ⋅dWt\.d\\varphi\(X\_\{t\}\)=\\bigl\(\\nabla\\varphi\\cdot b\+\\varepsilon^\{\*\}\\Delta\\varphi\\bigr\)dt\+\\sqrt\{2\\varepsilon^\{\*\}\}\\,\\nabla\\varphi\\cdot dW\_\{t\}\.Taking expectations againstμtβ\\mu\_\{t\}^\{\\beta\}and using⟨φ,μtβ⟩=𝔼Xt∼μtβ\[φ\]\\langle\\varphi,\\mu\_\{t\}^\{\\beta\}\\rangle=\\mathbb\{E\}\_\{X\_\{t\}\\sim\\mu\_\{t\}^\{\\beta\}\}\[\\varphi\],
ddt⟨φ,μtβ⟩=⟨∇φ⋅b\+ε∗Δφ,μtβ⟩\.\\frac\{d\}\{dt\}\\langle\\varphi,\\mu\_\{t\}^\{\\beta\}\\rangle=\\bigl\\langle\\nabla\\varphi\\cdot b\+\\varepsilon^\{\*\}\\Delta\\varphi,\\;\\mu\_\{t\}^\{\\beta\}\\bigr\\rangle\.Two integration\-by\-parts identities \(with vanishing boundary terms byφ∈Cc2\\varphi\\in C\_\{c\}^\{2\}\),
⟨∇φ⋅b,μtβ⟩=−⟨φ,∇⋅\(μtβb\)⟩,⟨ε∗Δφ,μtβ⟩=⟨φ,Δ\(ε∗μtβ\)⟩,\\langle\\nabla\\varphi\\cdot b,\\mu\_\{t\}^\{\\beta\}\\rangle=\-\\langle\\varphi,\\nabla\\\!\\cdot\\\!\(\\mu\_\{t\}^\{\\beta\}b\)\\rangle,\\qquad\\langle\\varepsilon^\{\*\}\\Delta\\varphi,\\mu\_\{t\}^\{\\beta\}\\rangle=\\langle\\varphi,\\Delta\(\\varepsilon^\{\*\}\\mu\_\{t\}^\{\\beta\}\)\\rangle,yieldddt⟨φ,μtβ⟩=⟨φ,−∇⋅\(μtβb\)\+Δ\(ε∗μtβ\)⟩\\frac\{d\}\{dt\}\\langle\\varphi,\\mu\_\{t\}^\{\\beta\}\\rangle=\\bigl\\langle\\varphi,\\,\-\\nabla\\\!\\cdot\\\!\(\\mu\_\{t\}^\{\\beta\}b\)\+\\Delta\(\\varepsilon^\{\*\}\\mu\_\{t\}^\{\\beta\}\)\\bigr\\rangle\. By density ofCc2C\_\{c\}^\{2\}in distributions,
∂tμtβ\+∇⋅\(μtβb\(x,μt,ε∗\(t,x;β\)\)\)=Δ\(ε∗\(t,x;β\)μtβ\)\.\\partial\_\{t\}\\mu\_\{t\}^\{\\beta\}\+\\nabla\\\!\\cdot\\\!\\bigl\(\\mu\_\{t\}^\{\\beta\}\\,b\\bigl\(x,\\mu\_\{t\},\\varepsilon^\{\*\}\(t,x;\\beta\)\\bigr\)\\bigr\)=\\Delta\\bigl\(\\varepsilon^\{\*\}\(t,x;\\beta\)\\,\\mu\_\{t\}^\{\\beta\}\\bigr\)\.Whenε∗\\varepsilon^\{\*\}is spatially constant on the support ofμtβ\\mu\_\{t\}^\{\\beta\}\(for instance, after an interior optimum is reached\),Δ\(ε∗μtβ\)=ε∗Δμtβ\\Delta\(\\varepsilon^\{\*\}\\mu\_\{t\}^\{\\beta\}\)=\\varepsilon^\{\*\}\\Delta\\mu\_\{t\}^\{\\beta\}, recovering the form printed in \([5\.4](https://arxiv.org/html/2607.23029#S5.E4)\)\. ∎
### C\.6Backward HJB from the dynamic programming principle \(eq\. \([5\.5](https://arxiv.org/html/2607.23029#S5.E5)\)\)
We treat the privacy termβδdp\(ε\)\\beta\\,\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon\)in \([5\.3](https://arxiv.org/html/2607.23029#S5.E3)\) as a running cost rate, consistent with the per\-round accounting in the discrete\-time game: under a feedback controlε\\varepsilon, the cost incurred betweenttandt\+ht\+his𝔼\[∫tt\+h\(ℓ\(Xs;ws\)\+βδdp\(εs\)\)𝑑s\]\\mathbb\{E\}\\bigl\[\\int\_\{t\}^\{t\+h\}\\\!\\bigl\(\\ell\(X\_\{s\};w\_\{s\}\)\+\\beta\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon\_\{s\}\)\\bigr\)ds\\bigr\]\. The dynamic programming principle gives, for anyh\>0h\>0,
V\(t,x;β\)=infε𝔼\[∫tt\+h\(ℓ\(Xs;ws\)\+βδdp\(εs\)\)𝑑s\+V\(t\+h,Xt\+h;β\)\]\.V\(t,x;\\beta\)=\\inf\_\{\\varepsilon\}\\,\\mathbb\{E\}\\\!\\left\[\\int\_\{t\}^\{t\+h\}\\\!\\bigl\(\\ell\(X\_\{s\};w\_\{s\}\)\+\\beta\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon\_\{s\}\)\\bigr\)ds\+V\(t\+h,X\_\{t\+h\};\\beta\)\\right\]\.For smoothVV, Itô’s formula with the SDE \([5\.2](https://arxiv.org/html/2607.23029#S5.E2)\) and a Taylor expansion inhhyield
V\(t\+h,Xt\+h;β\)=V\(t,x;β\)\+h\(∂tV\+∇V⋅b\+εTr\(D2V\)\)\+o\(h\)\+martingale\.V\(t\+h,X\_\{t\+h\};\\beta\)=V\(t,x;\\beta\)\+h\\bigl\(\\partial\_\{t\}V\+\\nabla V\\cdot b\+\\varepsilon\\,\\mathrm\{Tr\}\(D^\{2\}V\)\\bigr\)\+o\(h\)\+\\text\{martingale\}\.Substituting and dividing byh→0\+h\\to 0^\{\+\},
0=infε\{∂tV\+ℓ\+βδdp\(ε\)\+∇V⋅b\(x,μt,ε\)\+εTr\(D2V\)\}\.0=\\inf\_\{\\varepsilon\}\\\!\\left\\\{\\partial\_\{t\}V\+\\ell\+\\beta\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon\)\+\\nabla V\\cdot b\(x,\\mu\_\{t\},\\varepsilon\)\+\\varepsilon\\,\\mathrm\{Tr\}\(D^\{2\}V\)\\right\\\}\.The∂tV\\partial\_\{t\}Vandℓ\\ellterms do not depend onε\\varepsilonand can be pulled out of the infimum, giving the backward HJB
−∂tV=ℓ\+infε\{βδdp\(ε\)\+∇V⋅b\(x,μt,ε\)\+εTr\(D2V\)\}\.\-\\partial\_\{t\}V=\\ell\+\\inf\_\{\\varepsilon\}\\\!\\bigl\\\{\\beta\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon\)\+\\nabla V\\cdot b\(x,\\mu\_\{t\},\\varepsilon\)\+\\varepsilon\\,\\mathrm\{Tr\}\(D^\{2\}V\)\\bigr\\\}\.Expanding∇V⋅b\(x,μt,ε\)=−∇V⋅∇L−εσ2∇V⋅x−λ∇V⋅\(x−x¯\)\\nabla V\\cdot b\(x,\\mu\_\{t\},\\varepsilon\)=\-\\nabla V\\cdot\\nabla L\-\\tfrac\{\\varepsilon\}\{\\sigma^\{2\}\}\\nabla V\\cdot x\-\\lambda\\nabla V\\cdot\(x\-\\bar\{x\}\)and grouping theε\\varepsilon\-independent terms outside the infimum yields the Hamiltonian \([5\.6](https://arxiv.org/html/2607.23029#S5.E6)\) and the HJB \([5\.5](https://arxiv.org/html/2607.23029#S5.E5)\)\. The terminal conditionV\(T,x;β\)=0V\(T,x;\\beta\)=0encodes the zero terminal cost in \([5\.3](https://arxiv.org/html/2607.23029#S5.E3)\)\. ∎
### C\.7Closed\-form optimal control \(§[5\.3](https://arxiv.org/html/2607.23029#S5.SS3)\)
Differentiating the bracketed expression of the Hamiltonian \([5\.6](https://arxiv.org/html/2607.23029#S5.E6)\) inε\\varepsilon,
∂∂ε\{βδdp\(ε\)−εσ2∇V⋅x\+εTr\(D2V\)\}=βδdp′\(ε\)−1σ2∇V⋅x\+Tr\(D2V\)\.\\frac\{\\partial\}\{\\partial\\varepsilon\}\\\!\\left\\\{\\beta\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon\)\-\\tfrac\{\\varepsilon\}\{\\sigma^\{2\}\}\\nabla V\\cdot x\+\\varepsilon\\,\\mathrm\{Tr\}\(D^\{2\}V\)\\right\\\}=\\beta\\delta\_\{\\mathrm\{dp\}\}^\{\\prime\}\(\\varepsilon\)\-\\tfrac\{1\}\{\\sigma^\{2\}\}\\nabla V\\cdot x\+\\mathrm\{Tr\}\(D^\{2\}V\)\.The LSI bound \([4\.3](https://arxiv.org/html/2607.23029#S4.E3)\) can be writtenδdp\(ε\)=CdNexp\(−θ\(ε−ε0\)\)\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon\)=\\tfrac\{C\_\{d\}\}\{\\sqrt\{N\}\}\\exp\(\-\\theta\(\\varepsilon\-\\varepsilon\_\{0\}\)\)withθ:=αKτ/2\\theta:=\\alpha K\\tau/2andε0:=\(λ\+G\)/α\\varepsilon\_\{0\}:=\(\\lambda\+G\)/\\alpha, soδdp′\(ε\)=−θδdp\(ε\)\\delta\_\{\\mathrm\{dp\}\}^\{\\prime\}\(\\varepsilon\)=\-\\theta\\,\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon\)\. Setting the derivative above to zero and rearranging gives the first\-order condition
θβδdp\(ε∗\)=Tr\(D2V\)−1σ2∇V⋅x\.\\theta\\,\\beta\\,\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon^\{\*\}\)=\\mathrm\{Tr\}\(D^\{2\}V\)\-\\tfrac\{1\}\{\\sigma^\{2\}\}\\nabla V\\cdot x\.\(C\.3\)Solving forε∗\\varepsilon^\{\*\}when the right\-hand sideM:=Tr\(D2V\)−\(∇V⋅x\)/σ2M:=\\mathrm\{Tr\}\(D^\{2\}V\)\-\(\\nabla V\\cdot x\)/\\sigma^\{2\}is positive yields the closed form
ε∗\(t,x;β\)=ε0\+1θlog\(βθCd/NM\)\.\\varepsilon^\{\*\}\(t,x;\\beta\)=\\varepsilon\_\{0\}\+\\frac\{1\}\{\\theta\}\\,\\log\\\!\\left\(\\frac\{\\beta\\theta C\_\{d\}/\\sqrt\{N\}\}\{M\}\\right\)\.\(C\.4\)The second\-order condition∂2/∂ε2\{⋅\}=θ2βδdp\(ε∗\)\>0\\partial^\{2\}/\\partial\\varepsilon^\{2\}\\\{\\cdot\\\}=\\theta^\{2\}\\beta\\,\\delta\_\{\\mathrm\{dp\}\}\(\\varepsilon^\{\*\}\)\>0confirms that this critical point is a minimum of the bracketed Hamiltonian\. WhenM≤0M\\leq 0no interior optimum exists andε∗\\varepsilon^\{\*\}saturates at the upper boundaryεmax\\varepsilon\_\{\\max\}; when the log argument is so large thatε∗<εmin\\varepsilon^\{\*\}<\\varepsilon\_\{\\min\}, the optimum saturates atεmin\\varepsilon\_\{\\min\}\. Integratingε∗\\varepsilon^\{\*\}against the equilibrium population recovers the consistency condition
ε¯t=∫ε∗\(t,x;β\)μtβ\(dx\)ρ\(dβ\),\\bar\{\\varepsilon\}\_\{t\}=\\int\\varepsilon^\{\*\}\(t,x;\\beta\)\\,\\mu\_\{t\}^\{\\beta\}\(dx\)\\,\\rho\(d\\beta\),\(C\.5\)which is the integrated form of the scalar fixed pointε¯∗=Φ\(ε¯∗\)\\bar\{\\varepsilon\}^\{\*\}=\\Phi\(\\bar\{\\varepsilon\}^\{\*\}\)used by the discrete\-grid solver of Section[6](https://arxiv.org/html/2607.23029#S6)\. ∎
### C\.8Exponential DP at MFNE
###### Theorem C\.2\(Exponential DP at MFNE\)\.
Under the MFNE with mean\-field strengthε∗\\varepsilon^\{\*\}satisfyingαε∗\>λ\+G\\alpha\\varepsilon^\{\*\}\>\\lambda\+G, MFPG training is\(ϵdp,δdp\)\(\\epsilon\_\{\\mathrm\{dp\}\},\\delta\_\{\\mathrm\{dp\}\}\)\-DP with
δdp≤CdNexp\(−\(αε∗−λ−G\)Kτ2\),\\delta\_\{\\mathrm\{dp\}\}\\leq\\frac\{C\_\{d\}\}\{\\sqrt\{N\}\}\\exp\\\!\\Bigl\(\-\\tfrac\{\(\\alpha\\varepsilon^\{\*\}\-\\lambda\-G\)K\\tau\}\{2\}\\Bigr\),whereKKis the number of training rounds\. By contrast DP\-SGD achieves constantδdp=δ\\delta\_\{\\mathrm\{dp\}\}=\\deltawithϵdp=O\(Klog\(1/δ\)/σ\)\\epsilon\_\{\\mathrm\{dp\}\}=O\(\\sqrt\{K\\log\(1/\\delta\)\}/\\sigma\)growing asK\\sqrt\{K\}\.
###### Proof\.
We prove the bound first under the homogeneous specialisation \(βk≡β\\beta\_\{k\}\\equiv\\beta\), then extend to the heterogeneous case under a uniform activation hypothesis\.
#### Homogeneous case\.
If every client has the same preferenceβ\\beta, the MFNE collapses to a singleε¯∗=ε∗\\bar\{\\varepsilon\}^\{\*\}=\\varepsilon^\{\*\}shared by every client, and the population measureμt\\mu\_\{t\}evolves under \([4\.1](https://arxiv.org/html/2607.23029#S4.E1)\) withε≡ε∗\\varepsilon\\equiv\\varepsilon^\{\*\}\. The activation hypothesisαε∗\>λ\+G\\alpha\\varepsilon^\{\*\}\>\\lambda\+Gthen impliesr=αε∗−λ−G\>0r=\\alpha\\varepsilon^\{\*\}\-\\lambda\-G\>0, and Appendix[C\.2](https://arxiv.org/html/2607.23029#A3.SS2)givesδdp≤CdNexp\(−rKτ/2\)\\delta\_\{\\mathrm\{dp\}\}\\leq\\tfrac\{C\_\{d\}\}\{\\sqrt\{N\}\}\\exp\(\-rK\\tau/2\), which is the statement of the theorem\.
#### Heterogeneous case\.
For type\-dependent equilibriaε∗\(β\)\\varepsilon^\{\*\}\(\\beta\), the per\-type FPK is \([5\.4](https://arxiv.org/html/2607.23029#S5.E4)\)\. For two neighbouring populations differing in a single client of typeβ0\\beta\_\{0\}, the contraction acts only on the affected sliceμtβ0\\mu\_\{t\}^\{\\beta\_\{0\}\}, and the Wasserstein contraction rate of Appendix[C\.2](https://arxiv.org/html/2607.23029#A3.SS2)becomesr\(β0\)=αε∗\(β0\)−λ−Gr\(\\beta\_\{0\}\)=\\alpha\\varepsilon^\{\*\}\(\\beta\_\{0\}\)\-\\lambda\-G\. The privacy guarantee for that client is
δdp\(β0\)≤CdNexp\(−r\(β0\)Kτ/2\)\.\\delta\_\{\\mathrm\{dp\}\}^\{\(\\beta\_\{0\}\)\}\\leq\\tfrac\{C\_\{d\}\}\{\\sqrt\{N\}\}\\,\\exp\\bigl\(\-r\(\\beta\_\{0\}\)K\\tau/2\\bigr\)\.A uniform user\-level guarantee is obtained by taking the worst\-case ratermin=minβr\(β\)=αεmin∗−λ−Gr\_\{\\min\}=\\min\_\{\\beta\}r\(\\beta\)=\\alpha\\varepsilon^\{\*\}\_\{\\min\}\-\\lambda\-Gwhereεmin∗=minβε∗\(β\)\\varepsilon^\{\*\}\_\{\\min\}=\\min\_\{\\beta\}\\varepsilon^\{\*\}\(\\beta\)\. The theorem states the result at the population\-mean strengthε¯∗\\bar\{\\varepsilon\}^\{\*\}for compactness; the strict per\-client guarantee applies atεmin∗≤ε¯∗\\varepsilon^\{\*\}\_\{\\min\}\\leq\\bar\{\\varepsilon\}^\{\*\}and is therefore weaker by at most a factor ofexp\(\(ε¯∗−εmin∗\)αKτ/2\)\\exp\(\(\\bar\{\\varepsilon\}^\{\*\}\-\\varepsilon^\{\*\}\_\{\\min\}\)\\alpha K\\tau/2\)\. Both forms agree in the homogeneous case\. ∎
## Appendix DExperimental setup
We use FedAvg with gradient clipC=1C\{=\}1and learning rateη=0\.01\\eta\{=\}0\.01throughout\. The three benchmarks span increasing complexity\. \(i\) The*quadratic*task usesL\(w\)=12‖Aw−b‖2L\(w\)=\\tfrac\{1\}\{2\}\\\|Aw\-b\\\|^\{2\}per client withd=5,N=5,T=10d\{=\}5,\\,N\{=\}5,\\,T\{=\}10, providing analytical ground truth\. \(ii\)*Logistic regression*on synthetic binary data withd=20,N=8,T=15d\{=\}20,\\,N\{=\}8,\\,T\{=\}15is the primary utility benchmark, since all four cells reach a non\-trivial test accuracy\. \(iii\)*MNIST*classification with the MLP784→128→64→10784\\to 128\\to 64\\to 10on2,0002\{,\}000IID\-split images andN=10,T=20N\{=\}10,\\,T\{=\}20is a stress test for high parameter dimension\. For the entropic cells we use the strength gridℰ=\{0\.1,0\.3,0\.5,1\.0,2\.0\}\\mathcal\{E\}=\\\{0\.1,0\.3,0\.5,1\.0,2\.0\\\}, Gaussian priorν=𝒩\(0,I\)\\nu=\\mathcal\{N\}\(0,I\), JKO stepτ=0\.1\\tau\{=\}0\.1, and variance penaltyλ=0\.01\\lambda\{=\}0\.01\. Heterogeneous client preferences are linearly spaced asβk∈\[0\.5β,1\.5β\]\\beta\_\{k\}\\in\[0\.5\\beta,\\,1\.5\\beta\]withβ=1\.0\\beta\{=\}1\.0\. Each method’sϵdp\\epsilon\_\{\\mathrm\{dp\}\}is the value reported by its active accountant \(RDP for DP\-SGD and MAPG\-DP, linear\-budget for the entropic cells\), andδdp\\delta\_\{\\mathrm\{dp\}\}is constant for the Gaussian cells \(target10−510^\{\-5\}\) and computed via \([4\.3](https://arxiv.org/html/2607.23029#S4.E3)\) for the entropic cells\.
## Appendix EMean\-field equilibrium structure
Figure[4](https://arxiv.org/html/2607.23029#A5.F4)characterises the MFNE as a function of the heterogeneous client preferences\. The equilibrium strengthε¯∗\\bar\{\\varepsilon\}^\{\*\}is monotonically decreasing inβ\\beta: as clients become more privacy\-sensitive, the population adopts stronger entropic regularisation, matching the qualitative prediction of Theorem[C\.2](https://arxiv.org/html/2607.23029#A3.Thmthm2)\. The middle panel confirms that the correspondingδ∗\\delta^\{\*\}tightens withβ\\beta, so the LSI bound improves at exactly the clients who care most\. The right panel shows the equilibrium distribution under the realistic preference rangeβk∈\[0\.5,1\.5\]\\beta\_\{k\}\\in\[0\.5,1\.5\]used elsewhere in the paper: clients concentrate onε∗∈\{0\.3,0\.5\}\\varepsilon^\{\*\}\\in\\\{0\.3,\\,0\.5\\\}with population meanε¯∗≈0\.5\\bar\{\\varepsilon\}^\{\*\}\\approx 0\.5\. This concentration is what allows MFPG to match MFEP at the population level while still returning a per\-client privacy report that MFEP cannot\.
Figure 4:MFNE structure\.*Left:*ε¯∗\\bar\{\\varepsilon\}^\{\*\}decreases withβ\\beta\.*Centre:*the corresponding LSIδ∗\\delta^\{\*\}tightens withβ\\beta\.*Right:*for a heterogeneous populationβk∈\[0\.5,1\.5\]\\beta\_\{k\}\\in\[0\.5,1\.5\], clients concentrate onε∗∈\{0\.3,0\.5\}\\varepsilon^\{\*\}\\in\\\{0\.3,\\,0\.5\\\}with meanε¯∗≈0\.5\\bar\{\\varepsilon\}^\{\*\}\\approx 0\.5\.Sweeping the targetϵ∈\[0\.1,5\]\\epsilon\\in\[0\.1,5\]on logistic regression yields the privacy–utility frontier in Figure[5](https://arxiv.org/html/2607.23029#A5.F5)\. MFPG matches or exceeds DP\-SGD at every privacy level; MFEP and MFPG trace nearly coincident curves in this single\-objective sweep because the population meanε¯∗\\bar\{\\varepsilon\}^\{\*\}converges to a value MFEP can also pick\. The MFPG benefit beyond MFEP appears only under heterogeneousβk\\beta\_\{k\}, as documented in Figure[4](https://arxiv.org/html/2607.23029#A5.F4)above\.
Figure 5:Privacy–utility frontier on logistic regression\. MFPG matches or exceeds DP\-SGD at every privacy level; the gap over MFEP appears only when client preferences are heterogeneous \(Figure[4](https://arxiv.org/html/2607.23029#A5.F4)\)\.Similar Articles
Federated Learning
The article explains the concept of Federated Learning as a privacy-preserving machine learning technique that trains models on local devices rather than central servers. It details the process of encrypted parameter updates and aggregation to mitigate data leakage risks while maintaining model performance.
PrivFusion: A Privacy-preserving Multi-Agent Framework for Harmonizing Distributed Datasets
PrivFusion is a privacy-preserving multi-agent framework that automates the harmonization of structured datasets across institutions before federated training, reducing manual effort and enabling collaborative analytics on sensitive clinical data.
MEMOA: Massive Mixtures of Online Agents via Mean-Field Decentralized Nash Equilibria
The paper introduces MEMOA, a decentralized strategy for massive online agents that achieves optimality via mean-field Nash equilibria, outperforming greedy baselines while scaling better than centralized approaches.
FIRMA: FIbonacci Ring Model Aggregation for Privacy-preserving Federated Learning
This paper introduces FIRMA, a family of three privacy-preserving federated learning protocols using Fibonacci-weighted ring aggregation to achieve server-free operation, permanently private classification heads, and improved accuracy under data heterogeneity.
Personalized Federated Hierarchical Gaussian Processes for Privacy-Preserving Modeling of Heterogeneous Distributed Systems
The paper introduces pFedHGP, a personalized federated learning approach using hierarchical Gaussian processes for probabilistic modeling of heterogeneous distributed systems while preserving privacy through federated variational inference.