对称正定流形上函数数据的几何特征学习

arXiv cs.LG 论文

摘要

本文介绍了 MatFAE,一种用于从对称正定 (SPD) 流形上的轨迹中学习表示的功能神经网络,可应用于神经影像学及其他科学领域。

arXiv:2609.30487v1 Announce Type: new Abstract: We here develop a functional neural network, termed MatFAE, for learning trajectories on the Riemannian manifold of symmetric positive definite (SPD) matrices. MatFAE features intrinsic layers that map manifold-valued functions to Euclidean vector-valued functions, followed by a functional layer that projects them into a finite-dimensional Euclidean space. Unlike most neural networks for discrete-time sequences, MatFAE treats each sequence as a continuous function and can therefore encode trajectory dynamics (e.g., first-order derivatives) in its latent representations. Additionally, the morphology of the functional weights in the functional layer offers interpretability by revealing the regions of the input functional data that contribute most to the latent representations. We justify the design principles and properties of each intrinsic layer and detail how matrix factorization is handled during backpropagation. We apply MatFAE to a range of fMRI datasets, demonstrating its ability to efficiently learn informative representations from high-dimensional SPD trajectories and its practical value for real-world neuroimaging analysis.
查看原文
查看缓存全文

缓存时间: 2026/09/29 09:38

# Geometric Feature Learning for Functional Data Valued on the Symmetric Positive Definite Manifold
Source: [https://arxiv.org/html/2609.30487](https://arxiv.org/html/2609.30487)
Samuel V\. Singh & Mimi ZhangAffiliation:School of Computer Science and StatisticsAffiliation:Trinity College DublinAffiliation:Dublin, IrelandAffiliation:\{ssingh9, mimi\.zhang\}@tcd\.ie

###### Abstract

We here develop a functional neural network, termedMatFAE, for learning trajectories on the Riemannian manifold of symmetric positive definite \(SPD\) matrices\.MatFAEfeatures intrinsic layers that map manifold\-valued functions to Euclidean vector\-valued functions, followed by a functional layer that projects them into a finite\-dimensional Euclidean space\. Unlike most neural networks for discrete\-time sequences,MatFAEtreats each sequence as a continuous function and can therefore encode trajectory dynamics \(e\.g\., first\-order derivatives\) in its latent representations\. Additionally, the morphology of the functional weights in the functional layer offers interpretability by revealing the regions of the input functional data that contribute most to the latent representations\. We justify the design principles and properties of each intrinsic layer and detail how matrix factorization is handled during backpropagation\. We applyMatFAEto a range of fMRI datasets, demonstrating its ability to efficiently learn informative representations from high\-dimensional SPD trajectories and its practical value for real\-world neuroimaging analysis\.

## 1Introduction

Functional data taking values in the Riemannian manifold of symmetric positive\-definite \(SPD\) matrices are increasingly common in scientific and engineering applications\. In neuroimaging, time\-varying functional connectivity from fMRI is typically represented as trajectories of SPD covariance or correlation matrices\[[Lurie et al\., 2020](https://arxiv.org/html/2609.30487#bib.bib19)\]\. In DT\-MRI, each voxel is associated with a3×33\\times 3SPD diffusion tensor that, when sampled along white matter tracts, yields SPD\-valued functions over a one\-dimensional spatial domain\[[Yuan et al\., 2013](https://arxiv.org/html/2609.30487#bib.bib18)\]\. In subsurface hydrology and geomechanics, anisotropic hydraulic conductivity \(or permeability\) is commonly modelled by a3×33\\times 3SPD tensor describing directional flow of groundwater or heat; profiles measured along boreholes or evolving in time during injection, pumping, or thermal loading likewise form SPD\-valued functional data\[[Sanchez\-Vila et al\., 2006](https://arxiv.org/html/2609.30487#bib.bib20)\]\. Other application fields include dynamic texture/action recognition, where each frame or temporal segment is encoded by an SPD descriptor[Guo et al\. \[2013\]](https://arxiv.org/html/2609.30487#bib.bib45), and structural health monitoring, where covariance matrices of vibration or strain signals from sensor networks are tracked over time\.

A natural representation\-learning strategy for SPD\-matrix trajectories is to interpret each matrix as a graph and model each trajectory as a sequence of graphs, to which one can apply temporal attention mechanisms or recurrent architectures to capture dependencies over time\[[Chakraborty et al\., 2018](https://arxiv.org/html/2609.30487#bib.bib21),[Kim et al\., 2021](https://arxiv.org/html/2609.30487#bib.bib17)\]\. Such sequence\-based approaches, however, treat the functional data as discrete\-time snapshots and may fail to capture structure in the underlying continuous\-time dynamics\. We hypothesise that first\-order manifold behaviour, which captures how trajectories evolve in direction and speed, encodes additional discriminative and mechanistic information\. A functional data perspective on SPD\-valued trajectories provides a principled way to capture both the trajectories and their derivatives on the manifold, and to embed this richer geometric and dynamical information into latent representations for downstream tasks such as clustering\. A detailed review of related work is given in Appendix[A](https://arxiv.org/html/2609.30487#A1)\.

Leveraging recent advances in Riemannian neural networks\[[Huang and Gool, 2017](https://arxiv.org/html/2609.30487#bib.bib2),[Chakraborty et al\., 2022](https://arxiv.org/html/2609.30487#bib.bib13),[Brooks et al\., 2019](https://arxiv.org/html/2609.30487#bib.bib22)\]and functional data analysis\[[Akeweje and Zhang, 2024](https://arxiv.org/html/2609.30487#bib.bib23),[Singh et al\., 2025](https://arxiv.org/html/2609.30487#bib.bib1)\], we introduceMatFAE, a neural architecture for representation learning with functional data taking values on the manifold of SPD matrices\. To the best of our knowledge,MatFAEis the first neural network explicitly designed for this setting\. Its geometric layers are intrinsically Riemannian, respecting the non\-Euclidean structure of SPD\-valued trajectories and inducing a principled geometric bias that yields more efficient and structured representations\.

TheMatFAEarchitecture is modular\. In Section[2](https://arxiv.org/html/2609.30487#S2), we explain the dimension\-reduction module, which maps high\-dimensional SPD\-valued functional data to lower\-dimensional Euclidean functional data\. In Section[3](https://arxiv.org/html/2609.30487#S3), we present the autoencoder module, which learns latent representations of the vector\-valued Euclidean functional data and reconstructs the original SPD\-valued functional data\. Section[4](https://arxiv.org/html/2609.30487#S4)outlines the computational challenges in training the network, and Section[5](https://arxiv.org/html/2609.30487#S5)reports extensive numerical studies\. All proofs are collected in the appendix\.

## 2Mapping manifold\-valued functions to Euclidean functions

Let𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}denote the manifold ofm×mm\\times mSPD matrices:𝒮\+m=\{𝐗∈ℝm×m:𝐗=𝐗⊤,𝐗≻0\}\\mathcal\{S\}^\{m\}\_\{\+\}=\\\{\\mathbf\{X\}\\in\\mathbb\{R\}^\{m\\times m\}:\\mathbf\{X\}=\\mathbf\{X\}^\{\\top\},\\mathbf\{X\}\\succ 0\\\}\. Given a matrix\-valued trajectory𝐗⁡\(t\)∈𝒮\+m\\mathbf\{X\}\(t\)\\in\\mathcal\{S\}^\{m\}\_\{\+\}fort∈\[t0,t1\]t\\in\[t\_\{0\},t\_\{1\}\], we below develop a geometric networkℱ\\mathcal\{F\}that maps𝐗⁡\(t\)\\mathbf\{X\}\(t\)to app\-dimensional Euclidean function𝒚⁡\(t\)=ℱ⁡\(𝐗⁡\(t\)\)∈ℝp\\boldsymbol\{y\}\(t\)=\\mathcal\{F\}\(\\mathbf\{X\}\(t\)\)\\in\\mathbb\{R\}^\{p\}\. We employ the Log\-Euclidean \(LE\) metric\[[Arsigny et al\., 2007](https://arxiv.org/html/2609.30487#bib.bib6)\]and the Log\-Cholesky \(LC\) metric\[[Lin, 2019](https://arxiv.org/html/2609.30487#bib.bib5)\], both having the bounded\-determinant property: for a collection of SPD matrices, the determinant of the Fréchet mean under the metric lies within the range of the input determinants\. This property is crucial for the average\-pooling layers in our network\. Although the affine\-invariant metric\[[Pennec et al\., 2006](https://arxiv.org/html/2609.30487#bib.bib9)\]also satisfies the bounded\-determinant property, there is no general closed\-form expression for the Fréchet mean, making its use substantially more costly and complicating backpropagation\. To avoid confusion, we write exp\(𝐔\\mathbf\{U\}\) for the matrix exponential and log\(𝐔\\mathbf\{U\}\) for the principal matrix logarithm\. By contrast,Exp𝐗​\(𝐔\)\\text\{Exp\}\_\{\\mathbf\{X\}\}\(\\mathbf\{U\}\)denotes the Riemannian exponential at𝐗\\mathbf\{X\}andLog𝐗​\(𝐙\)\\text\{Log\}\_\{\\mathbf\{X\}\}\(\\mathbf\{Z\}\)denotes the Riemannian logarithm at𝐗\\mathbf\{X\}\.

All trainable weights inℱ\\mathcal\{F\}are time\-invariant; thus, when no ambiguity arises, we expound the architecture at a fixedttand write𝐗\\mathbf\{X\}for𝐗⁡\(t\)\\mathbf\{X\}\(t\)\. The manifold𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}is an open cone inSym​\(m\)=\{𝐔∈ℝm×m:𝐔=𝐔⊤\}\\text\{Sym\}\(m\)=\\\{\\mathbf\{U\}\\in\\mathbb\{R\}^\{m\\times m\}:\\mathbf\{U\}=\\mathbf\{U\}^\{\\top\}\\\}, and hence the tangent space𝒯𝐗​𝒮\+m\\mathcal\{T\}\_\{\\mathbf\{X\}\}\\mathcal\{S\}^\{m\}\_\{\+\}at any point𝐗∈𝒮\+m\\mathbf\{X\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}identifies withSym​\(m\)\\text\{Sym\}\(m\):𝒯𝐗​𝒮\+m=Sym​\(m\)\\mathcal\{T\}\_\{\\mathbf\{X\}\}\\mathcal\{S\}^\{m\}\_\{\+\}=\\text\{Sym\}\(m\), independently of𝐗\\mathbf\{X\}\. Let𝔏⁡\(𝐗\)\\mathfrak\{L\}\(\\mathbf\{X\}\)be the lower triangular matrix of the Cholesky decomposition of𝐗\\mathbf\{X\}:𝐗=𝔏⁡\(𝐗\)​𝔏​\(𝐗\)⊤\\mathbf\{X\}=\\mathfrak\{L\}\(\\mathbf\{X\}\)\\mathfrak\{L\}\(\\mathbf\{X\}\)^\{\\top\}, and⌊𝔏⁡\(𝐗\)⌋\\lfloor\\mathfrak\{L\}\(\\mathbf\{X\}\)\\rfloorthe strictly lower triangular matrix of𝔏⁡\(𝐗\)\\mathfrak\{L\}\(\\mathbf\{X\}\)\. Define the log\-Cholesky coordinate mapΦ:𝒮\+m→ℝm⁡\(m\+1\)/2\\Phi:\\mathcal\{S\}^\{m\}\_\{\+\}\\rightarrow\\mathbb\{R\}^\{m\(m\+1\)/2\},Φ⁡\(𝐗\)=\(vnz​\(𝔏⁡\(𝐗\)\),log⁡\(vd​\(𝔏⁡\(𝐗\)\)\)\)⊤\\Phi\(\\mathbf\{X\}\)=\(\\text\{vnz\}\(\\mathfrak\{L\}\(\\mathbf\{X\}\)\),\\log\(\\text\{vd\}\(\\mathfrak\{L\}\(\\mathbf\{X\}\)\)\)\)^\{\\top\}, wherevnz​\(𝔏​\(𝐗\)\)\\text\{vnz\}\(\\mathfrak\{L\}\(\\mathbf\{X\}\)\)is the vector of the non\-zero entries of⌊𝔏⁡\(𝐗\)⌋\\lfloor\\mathfrak\{L\}\(\\mathbf\{X\}\)\\rfloor, andvd​\(𝔏​\(𝐗\)\)\\text\{vd\}\(\\mathfrak\{L\}\(\\mathbf\{X\}\)\)is the vector of the diagonal entries of𝔏⁡\(𝐗\)\\mathfrak\{L\}\(\\mathbf\{X\}\)\. The LC metric makesΦ\\Phia global Riemannian isometry: equippingΦ⁡\(𝐗\)\\Phi\(\\mathbf\{X\}\)with the standard Euclidean inner product and pulling it back byΦ\\Phidefines the LC metric:

⟨𝐔,𝐕⟩𝐗LC=⟨\(D𝐗​Φ\)​\(𝐔\),\(D𝐗​Φ\)​\(𝐕\)⟩2,𝐔,𝐕∈𝒯𝐗​𝒮\+m=Sym​\(m\),\\langle\\mathbf\{U\},\\mathbf\{V\}\\rangle\_\{\\mathbf\{X\}\}^\{\\text\{LC\}\}=\\langle\(D\_\{\\mathbf\{X\}\}\\Phi\)\(\\mathbf\{U\}\),~\(D\_\{\\mathbf\{X\}\}\\Phi\)\(\\mathbf\{V\}\)\\rangle\_\{2\},~~~\\mathbf\{U\},\\mathbf\{V\}\\in\\mathcal\{T\}\_\{\\mathbf\{X\}\}\\mathcal\{S\}^\{m\}\_\{\+\}=\\text\{Sym\}\(m\),where\(D𝐗​Φ\)​\(𝐔\)=dd​s\|s=0​Φ​\(𝐗\+s​𝐔\)\(D\_\{\\mathbf\{X\}\}\\Phi\)\(\\mathbf\{U\}\)=\\frac\{d\}\{ds\}\|\_\{s=0\}\\Phi\(\\mathbf\{X\}\+s\\mathbf\{U\}\)is the Fréchet derivative ofΦ\\Phiat the point𝐗\\mathbf\{X\}to the tangent direction𝐔\\mathbf\{U\}\. Likewise, the LE metric is the pullback of the Frobenius metric onSym​\(m\)\\text\{Sym\}\(m\)via the principal matrix logarithmlog:𝒮\+m→Sym​\(m\)\\log:\\mathcal\{S\}^\{m\}\_\{\+\}\\rightarrow\\text\{Sym\}\(m\); for𝐗∈𝒮\+m\\mathbf\{X\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}and𝐔,𝐕∈𝒯𝐗​𝒮\+m\\mathbf\{U\},\\mathbf\{V\}\\in\\mathcal\{T\}\_\{\\mathbf\{X\}\}\\mathcal\{S\}^\{m\}\_\{\+\}, the LE inner product at𝐗\\mathbf\{X\}is

⟨𝐔,𝐕⟩𝐗LE=⟨\(D𝐗​log\)​\(𝐔\),\(D𝐗​log\)​\(𝐕\)⟩F,\\langle\\mathbf\{U\},\\mathbf\{V\}\\rangle\_\{\\mathbf\{X\}\}^\{\\text\{LE\}\}=\\langle\(D\_\{\\mathbf\{X\}\}\\log\)\(\\mathbf\{U\}\),~\(D\_\{\\mathbf\{X\}\}\\log\)\(\\mathbf\{V\}\)\\rangle\_\{F\},where\(D𝐗​log\)​\(𝐔\)=dd​s\|s=0​log⁡\(𝐗\+s​𝐔\)\(D\_\{\\mathbf\{X\}\}\\log\)\(\\mathbf\{U\}\)=\\frac\{d\}\{ds\}\|\_\{s=0\}\\log\(\\mathbf\{X\}\+s\\mathbf\{U\}\)is the Fréchet derivative oflog\\logat the point𝐗\\mathbf\{X\}to the tangent direction𝐔\\mathbf\{U\}\. Both Riemannian manifolds\(𝒮\+m,⟨⋅,⋅⟩LC\)\(\\mathcal\{S\}^\{m\}\_\{\+\},\\langle\\cdot,\\cdot\\rangle^\{\\text\{LC\}\}\)and\(𝒮\+m,⟨⋅,⋅⟩LE\)\(\\mathcal\{S\}^\{m\}\_\{\+\},\\langle\\cdot,\\cdot\\rangle^\{\\text\{LE\}\}\)are flat \(zero sectional curvature\), geodesically complete, and simply connected; hence each is a Cartan\-Hadamard manifold\. Further properties of the LC and LE metrics are provided in Appendix[B](https://arxiv.org/html/2609.30487#A2)\.

According to Theorem 1 in[Singh et al\. \[2025\]](https://arxiv.org/html/2609.30487#bib.bib1), on Cartan\-Hadamard manifolds we can linearize the manifold\-valued functions by mapping them into Euclidean space via the \(globally defined\) Riemannian logarithm map\. The linearized functions can then be processed by the FAEclust functional network, which guarantees universal approximation\. However, a direct mapping of𝐗⁡\(t\)\\mathbf\{X\}\(t\)ontoSym​\(m\)\\text\{Sym\}\(m\)produces𝒚⁡\(t\)\\boldsymbol\{y\}\(t\)with dimensionm⁡\(m\+1\)/2m\(m\+1\)/2, which imposes a heavy training cost for FAEclust\. Therefore, we here develop a modular geometric network that compresses the trajectory𝐗⁡\(t\)\\mathbf\{X\}\(t\)to a lower\-order SPD trajectory𝐗4​\(t\)∈𝒮\+m2\\mathbf\{X\}\_\{4\}\(t\)\\in\\mathcal\{S\}^\{m\_\{2\}\}\_\{\+\}withm2<mm\_\{2\}<m, and then applies the Riemannian logarithm to map the trajectory𝐗4​\(t\)\\mathbf\{X\}\_\{4\}\(t\)to the Euclidean spaceSym​\(m2\)\\text\{Sym\}\(m\_\{2\}\)for subsequent processing by any functional neural network\. This dimensionality reduction substantially decreases computational burden while preserving the relevant geometry\.

An overview of the network architecture is provided in[Figure 1](https://arxiv.org/html/2609.30487#S2.F1)\. The two BiMap layers perform dimensionality reduction, the Pooling layer aggregates SPD outputs across heads, and the Activation layer appropriately regularize the SPD matrix\. The Logarithm layer performs Riemannian logarithm, mapping each head𝐗4h​\(t\)\\mathbf\{X\}\_\{4\}^\{h\}\(t\)to𝐘h​\(t\)\\mathbf\{Y\}^\{h\}\(t\):𝐘h​\(t\)=Log𝐈LC​\(𝐗4h​\(t\)\)\\mathbf\{Y\}^\{h\}\(t\)=\\text\{Log\}^\{\\text\{LC\}\}\_\{\\mathbf\{I\}\}\(\\mathbf\{X\}\_\{4\}^\{h\}\(t\)\)or𝐘h​\(t\)=Log𝐈LE​\(𝐗4h​\(t\)\)\\mathbf\{Y\}^\{h\}\(t\)=\\text\{Log\}^\{\\text\{LE\}\}\_\{\\mathbf\{I\}\}\(\\mathbf\{X\}\_\{4\}^\{h\}\(t\)\), where𝐈\\mathbf\{I\}is the identity matrix\. The explicit expressions forLog𝐈LC​\(⋅\)\\text\{Log\}^\{\\text\{LC\}\}\_\{\\mathbf\{I\}\}\(\\cdot\)andLog𝐈LE​\(⋅\)\\text\{Log\}^\{\\text\{LE\}\}\_\{\\mathbf\{I\}\}\(\\cdot\)are provided in Eqs\. \([6](https://arxiv.org/html/2609.30487#A2.E6)\) and \([10](https://arxiv.org/html/2609.30487#A2.E10)\), respectively\. In the Concat layer, we vectorize the lower\-triangular entries of𝐘h​\(t\)\\mathbf\{Y\}^\{h\}\(t\)and concatenate them across heads to obtain a vector\-valued Euclidean function𝒚⁡\(t\)\\boldsymbol\{y\}\(t\)\. We now provide a rigorous justification for the BiMap, Pooling, and Activation layers and detail their properties\.

![Refer to caption](https://arxiv.org/html/2609.30487v1/fSPDnet.png)Figure 1:The two BiMap layers reduce the matrix dimension frommmtom1\(<m\)m\_\{1\}\(<m\), and further fromm1m\_\{1\}tom2\(<m1\)m\_\{2\}\(<m\_\{1\}\); that is,𝐗1h​\(t\)∈𝒮\+m1\\mathbf\{X\}\_\{1\}^\{h\}\(t\)\\in\\mathcal\{S\}^\{m\_\{1\}\}\_\{\+\}and𝐗4h​\(t\)∈𝒮\+m2\\mathbf\{X\}\_\{4\}^\{h\}\(t\)\\in\\mathcal\{S\}^\{m\_\{2\}\}\_\{\+\}for1≤h≤H1\\leq h\\leq H\. The Pooling, Activation and Logarithm layers are geometric layers that depend on the Riemannian metric on𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}\. The Pooling layer performs manifold\-average pooling of the SPD matrices\{𝐗1h​\(t\):1≤h≤H\}\\\{\\mathbf\{X\}\_\{1\}^\{h\}\(t\):1\\leq h\\leq H\\\}; the Activation layer applies an activation on𝐗2​\(t\)\\mathbf\{X\}\_\{2\}\(t\), and the Logarithm layer maps each𝐗4h​\(t\)\\mathbf\{X\}\_\{4\}^\{h\}\(t\)to the tangent space via the Riemannian logarithmic mapLog𝐈LC​\(⋅\)\\text\{Log\}^\{\\text\{LC\}\}\_\{\\mathbf\{I\}\}\(\\cdot\)orLog𝐈LE​\(⋅\)\\text\{Log\}^\{\\text\{LE\}\}\_\{\\mathbf\{I\}\}\(\\cdot\)\. Finally, we vectorize the lower\-triangular entries of𝐘h​\(t\)\\mathbf\{Y\}^\{h\}\(t\)and concatenate them across heads to obtain a vector\-valued function𝒚⁡\(t\)\\boldsymbol\{y\}\(t\)\.### 2\.1Multi\-head BiMap

[Huang and Gool \[2017\]](https://arxiv.org/html/2609.30487#bib.bib2)introduced SPDNet, a deep architecture that operates directly on SPD matrices and preserves positive\-definiteness throughout\. Its core linear block \(i\.e\., the BiMap layer\) applies a congruence transformation:𝐗k=𝐖k​𝐗k−1​𝐖k⊤\\mathbf\{X\}\_\{k\}=\\mathbf\{W\}\_\{k\}\\mathbf\{X\}\_\{k\-1\}\\mathbf\{W\}\_\{k\}^\{\\top\}, where𝐗k−1∈𝒮\+dk−1\\mathbf\{X\}\_\{k\-1\}\\in\\mathcal\{S\}^\{d\_\{k\-1\}\}\_\{\+\}, and𝐖k∈ℝdk×dk−1​\(dk<dk−1\)\\mathbf\{W\}\_\{k\}\\in\\mathbb\{R\}^\{d\_\{k\}\\times d\_\{k\-1\}\}~\(d\_\{k\}<d\_\{k\-1\}\)is a full row rank weight matrix with orthonormal rows\. The bilinear congruence ensures that the output remains SPD, i\.e\.,𝐗k∈𝒮\+dk\\mathbf\{X\}\_\{k\}\\in\\mathcal\{S\}^\{d\_\{k\}\}\_\{\+\}, while reducing dimension fromdk−1d\_\{k\-1\}todkd\_\{k\}\. Stacking multiple BiMap layers yields progressively more compact and task\-discriminative SPD representations\.

For the𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}\-valued function𝐗⁡\(t\)\\mathbf\{X\}\(t\), a natural extension of the bilinear mapping is𝐗1​\(t\)=𝐖1​𝐗​\(t\)​𝐖1⊤\\mathbf\{X\}\_\{1\}\(t\)=\\mathbf\{W\}\_\{1\}\\mathbf\{X\}\(t\)\\mathbf\{W\}\_\{1\}^\{\\top\}\. However, time\-varying SPD signals almost never commute across time \(i\.e\.,𝐗⁡\(t1\)​𝐗​\(t2\)≠𝐗⁡\(t2\)​𝐗​\(t1\)\\mathbf\{X\}\(t\_\{1\}\)\\mathbf\{X\}\(t\_\{2\}\)\\neq\\mathbf\{X\}\(t\_\{2\}\)\\mathbf\{X\}\(t\_\{1\}\)\), and therefore no single global projection𝐖1\\mathbf\{W\}\_\{1\}can align the covariance geometry uniformly overtt\. When regimes switch and eigenvalues cross, causing the spectral subspaces to drift, any fixed𝐖1\\mathbf\{W\}\_\{1\}will inevitably overfit some intervals and underfit others\. We therefore introduce a multi\-head BiMap layer withHHparallel bilinear maps:𝐗1h​\(t\)=𝐖1h​𝐗​\(t\)​\(𝐖1h\)⊤\\mathbf\{X\}\_\{1\}^\{h\}\(t\)=\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(t\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\},h=1,…,Hh=1,\\ldots,H, each with its own full row rank𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}\.

When𝐗⁡\(t\)\\mathbf\{X\}\(t\)for differentttvalues do not share a common eigenbasis, different heads approximate different \(temporally\) local joint structures, improving conditioning and reducing bias compared with a single global projection\. Different heads naturally specialize to complementary, time\-dependent relationships \(e\.g\., distinct temporal regimes, frequency bands, or sensor groups\)\. Their SPD outputs are then fused by an SPD\-preserving pooling operator in the Pooling layer\. This design also mirrors 1\-D CNNs along the temporal axis: each𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}serves as a shared temporal filter, and multiple heads enhance expressive power by capturing diverse temporal patterns while retaining weight sharing\.

### 2\.2Pooling

The Pooling layer aggregates the SPD inputs\{𝐗1h:1≤h≤H\}\\\{\\mathbf\{X\}\_\{1\}^\{h\}:1\\leq h\\leq H\\\}\. A natural choice is the Fréchet mean on𝒮\+m1\\mathcal\{S\}^\{m\_\{1\}\}\_\{\+\}under either the LC or LE geometry:

𝐗2=arg​min𝐗∈𝒮\+m1∑h=1Hdg\(𝐗,𝐗1h\)2,\\mathbf\{X\}\_\{2\}=\\argmin\_\{\\mathbf\{X\}\\in\\mathcal\{S\}^\{m\_\{1\}\}\_\{\+\}\}\\sum\_\{h=1\}^\{H\}d\_\{g\}\(\\mathbf\{X\},\\mathbf\{X\}\_\{1\}^\{h\}\)^\{2\},wheredgd\_\{g\}is the geodesic distance\. Under either the LC metric or the LE metric, the Fréchet mean has a closed form \(Eqs\. \([7](https://arxiv.org/html/2609.30487#A2.E7)\) and \([11](https://arxiv.org/html/2609.30487#A2.E11)\)\)\. Conceptually, this operation is the manifold analogue of global average pooling: it compressesHHinputs into one SPD representative while strictly preserving positive\-definiteness\. A caveat is that the congruence mapping𝐖1h​𝐗​\(𝐖1h\)⊤\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}does not commute with the matrix logarithm/exponential in the LE metric: in general,

log⁡\(𝐖1h​𝐗​\(𝐖1h\)⊤\)≠𝐖1h​log⁡\(𝐗\)​\(𝐖1h\)⊤​and​exp⁡\(𝐖1h​𝐗​\(𝐖1h\)⊤\)≠𝐖1h​exp⁡\(𝐗\)​\(𝐖1h\)⊤\.\\log\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\\neq\\mathbf\{W\}\_\{1\}^\{h\}\\log\(\\mathbf\{X\}\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}~\\text\{ and \}~\\exp\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\\neq\\mathbf\{W\}\_\{1\}^\{h\}\\exp\(\\mathbf\{X\}\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\.Likewise, for the LC metric, the Cholesky factor𝔏⁡\(𝐖1h​𝐗​\(𝐖1h\)⊤\)\\mathfrak\{L\}\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)depends nonlinearly on both𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}and𝐗\\mathbf\{X\}\. Consequently, the Fréchet mean computed in the Pooling layer is neither invariant nor equivariant to dimension\-reducing congruence actions applied to𝐗⁡\(t\)\\mathbf\{X\}\(t\)\. We detail this lack of invariance/equivariance below\.

A canonical family of symmetries of𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}is given by the congruence action of GL\(mm\): any invertible matrix𝐔∈GL​\(m\)\\mathbf\{U\}\\in\\text\{GL\}\(m\)defines the symmetry𝐗↦𝐔𝐗𝐔⊤\\mathbf\{X\}\\mapsto\\mathbf\{U\}\\mathbf\{X\}\\mathbf\{U\}^\{\\top\}, acting on all points𝐗∈𝒮\+m\\mathbf\{X\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}\[[López et al\., 2021](https://arxiv.org/html/2609.30487#bib.bib10)\]\. \(1\) When𝐔\\mathbf\{U\}is an element of𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}, the symmetry𝐗↦𝐔𝐗𝐔⊤\\mathbf\{X\}\\mapsto\\mathbf\{U\}\\mathbf\{X\}\\mathbf\{U\}^\{\\top\}is a generalization of the Euclideantranslation, fixing no points of𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}\. \(2\) When𝐔∈O​\(m\)\\mathbf\{U\}\\in\\text\{O\}\(m\)is an orthogonal matrix, the symmetry𝐗↦𝐔𝐗𝐔⊤\\mathbf\{X\}\\mapsto\\mathbf\{U\}\\mathbf\{X\}\\mathbf\{U\}^\{\\top\}is conjugation by𝐔\\mathbf\{U\}, and hence fixes the base point𝐈=𝐔𝐈𝐔⊤=𝐔𝐔⊤=𝐈\\mathbf\{I\}=\\mathbf\{U\}\\mathbf\{I\}\\mathbf\{U\}^\{\\top\}=\\mathbf\{U\}\\mathbf\{U\}^\{\\top\}=\\mathbf\{I\}\. The stabilizer of𝐈\\mathbf\{I\}is preciselyO​\(m\)\\text\{O\}\(m\): elements withdet​\(𝐔\)=1\\text\{det\}\(\\mathbf\{U\}\)=1are𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}\-rotations, and elements withdet​\(𝐔\)=−1\\text\{det\}\(\\mathbf\{U\}\)=\-1are𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}\-reflections\.

###### Proposition 1\.

The LE Fréchet mean𝐗2\\mathbf\{X\}\_\{2\}is invariant or equivariant under orthogonal congruence only under pathological conditions\. The LC Fréchet mean is neither invariant nor equivariant under orthogonal congruence\. Moreover, under the LC and LE geometries, the Fréchet mean𝐗2\\mathbf\{X\}\_\{2\}is generally neither invariant nor equivariant under actions of the isometry group on𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}\.

#### Max pooling on𝒮\+m1\\mathcal\{S\}^\{m\_\{1\}\}\_\{\+\}

Unlike Euclidean space, the manifold𝒮\+m1\\mathcal\{S\}^\{m\_\{1\}\}\_\{\+\}admits neither canonical global coordinates nor a natural total order\. Therefore, a “max” operation is not intrinsically defined: any elementwise maximum defined in a particular chart \(e\.g\.,log\\logorΦ\\Phi\) is chart\-dependent and generally fails to be equivariant under manifold isometries\. An intrinsic alternative is to define a scalar scores:𝒮\+m1→ℝs:\\mathcal\{S\}^\{m\_\{1\}\}\_\{\+\}\\rightarrow\\mathbb\{R\}and select the head𝐗1h\\mathbf\{X\}\_\{1\}^\{h\}with the highest scores⁡\(𝐗1h\)s\(\\mathbf\{X\}\_\{1\}^\{h\}\)\. However, because features pass through the congruence map𝐗↦𝐖1h​𝐗​\(𝐖1h\)⊤\\mathbf\{X\}\\mapsto\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}, no non\-trivial choice ofssyields invariance/equivariance to general matrix congruence; this follows from the same obstruction discussed in Appendix[C](https://arxiv.org/html/2609.30487#A3)\. Equivariance is attainable only in pathological cases \(e\.g\., orthogonal conjugation with spectral scores, or specific metric\-based selections with matched transformations of any reference\)\. Moreover, themax\\maxoperation is non\-smooth and therefore less amenable to optimization\. For these reasons,MatFAEadopts average \(Fréchet\) pooling\.

Although the BiMap’s rectangular congruence deprives the Pooling layer of invariance or equivariance guarantees, it remains a highly efficient dimension\-reduction operator: it preserves SPD by construction and needs only two matrix\-matrix products, completely avoiding eigendecomposition\.

### 2\.3Activation

Conventional deep architectures interleave pointwise nonlinearitiesa:ℝ→ℝa:\\mathbb\{R\}\\rightarrow\\mathbb\{R\}between linear layers \(e\.g\., ReLU\)\. Such activations are non\-linear and, in most cases, non\-expansive with respect to the Euclidean norm \(i\.e\.,‖a⁡\(𝒙\)−a⁡\(𝒚\)‖2≤L​‖𝒙−𝒚‖2\\\|a\(\\boldsymbol\{x\}\)\-a\(\\boldsymbol\{y\}\)\\\|\_\{2\}\\leq L\\\|\\boldsymbol\{x\}\-\\boldsymbol\{y\}\\\|\_\{2\}withL≤1L\\leq 1\); certain choices \(e\.g\., the logistic sigmoid\) are strictly contractive\. Non\-expansiveness helps control the network’s overall Lipschitz constant, improving numerical stability, robustness to perturbations, and generalization by biasing toward smoother mappings\. The non\-linearities between layers prevent the deep neural networks from collapsing to a single fully connected layer\. Below we prove that the Fréchet\-mean pooling operator is, in general, neither contractive nor non\-expansive \(Proposition[2](https://arxiv.org/html/2609.30487#ThmPro2)and Proposition[3](https://arxiv.org/html/2609.30487#ThmPro3)\)\. However, the Fréchet\-mean pooling operator itself is non\-linear \(Proposition[4](https://arxiv.org/html/2609.30487#ThmPro4)\)\. Therefore, a repeated block “BiMap→\\rightarrowPooling→\\rightarrowBiMap→\\rightarrowPooling” will not collapse to a single “BiMap→\\rightarrowPooling” block\.

###### Proposition 2\.

Fix any spectral band0<α≤β<∞0<\\alpha\\leq\\beta<\\inftyand consider the subset𝒳α,β=\{𝐗∈𝒮\+m:α​𝐈⪯𝐗⪯β​𝐈\}\\mathcal\{X\}\_\{\\alpha,\\beta\}=\\\{\\mathbf\{X\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}:\\alpha\\mathbf\{I\}\\preceq\\mathbf\{X\}\\preceq\\beta\\mathbf\{I\}\\\}\. Under the LE metric, the mappingF:𝒮\+m↦𝒮\+m1F:\\mathcal\{S\}^\{m\}\_\{\+\}\\mapsto\\mathcal\{S\}^\{m\_\{1\}\}\_\{\+\},

F⁡\(𝐗,\{𝐖1h\}h=1H\)=𝔼LE​\(𝐖11​𝐗​\(𝐖11\)⊤,…,𝐖1H​𝐗​\(𝐖1H\)⊤\)=exp⁡\(1H​∑h=1Hlog⁡\(𝐖1h​𝐗​\(𝐖1h\)⊤\)CLOSE,F\(\\mathbf\{X\};\\\{\\mathbf\{W\}\_\{1\}^\{h\}\\\}\_\{h=1\}^\{H\}\)=\\mathbb\{E\}\_\{\\text\{LE\}\}\(\\mathbf\{W\}\_\{1\}^\{1\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{1\}\)^\{\\top\},\\ldots,\\mathbf\{W\}\_\{1\}^\{H\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{H\}\)^\{\\top\}\)=\\exp\(\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\log\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\),is Lipschitz on𝒳α,β\\mathcal\{X\}\_\{\\alpha,\\beta\}with constantL≤β/αL\\leq\\beta/\\alpha\. In particular,FFis neither contractive nor non\-expansive\.

###### Proposition 3\.

Fix any spectral band0<α≤β<∞0<\\alpha\\leq\\beta<\\inftyand consider the subset𝒳α,β=\{𝐗∈𝒮\+m:α​𝐈⪯𝐗⪯β​𝐈\}\\mathcal\{X\}\_\{\\alpha,\\beta\}=\\\{\\mathbf\{X\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}:\\alpha\\mathbf\{I\}\\preceq\\mathbf\{X\}\\preceq\\beta\\mathbf\{I\}\\\}\. Under the LC metric, the mappingF:𝒮\+m↦𝒮\+m1F:\\mathcal\{S\}^\{m\}\_\{\+\}\\mapsto\\mathcal\{S\}^\{m\_\{1\}\}\_\{\+\},

F⁡\(𝐗,\{𝐖1h\}h=1H\)=𝔼LC​\(𝐖11​𝐗​\(𝐖11\)⊤,…,𝐖1H​𝐗​\(𝐖1H\)⊤\)=Φ−1​\(1H​∑h=1HΦ⁡\(𝐖1h​𝐗​\(𝐖1h\)⊤\)CLOSE,F\(\\mathbf\{X\};\\\{\\mathbf\{W\}\_\{1\}^\{h\}\\\}\_\{h=1\}^\{H\}\)=\\mathbb\{E\}\_\{\\text\{LC\}\}\(\\mathbf\{W\}\_\{1\}^\{1\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{1\}\)^\{\\top\},\\ldots,\\mathbf\{W\}\_\{1\}^\{H\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{H\}\)^\{\\top\}\)=\\Phi^\{\-1\}\(\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\Phi\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\),is Lipschitz with constantL≤2​β3/2α​max⁡\{1,α−1\}L\\leq\\frac\{2\\beta^\{3/2\}\}\{\\alpha\}\\max\\\{1,\\alpha^\{\-1\}\\\}\. In particular,FFis neither contractive nor non\-expansive\.

###### Proposition 4\.

The pooling map is itself nonlinear\. Without any non\-linear activation after the Pooling layer, the repeated block “BiMap→\\rightarrowPooling→\\rightarrowBiMap→\\rightarrowPooling” will not collapse to a single “BiMap→\\rightarrowPooling” block\.

Proposition[4](https://arxiv.org/html/2609.30487#ThmPro4)implies that, if contractivity is not required, the Activation layer can be omitted\. The architecture in Figure[1](https://arxiv.org/html/2609.30487#S2.F1)then reduces to “BiMap→\\rightarrowPooling→\\rightarrowBiMap→\\rightarrowLogarithm→\\rightarrowConcat”\. However, the Fréchet\-mean pooling is only Lipschitz, neither contractive nor non\-expensive \(Propositions[2](https://arxiv.org/html/2609.30487#ThmPro2)and[3](https://arxiv.org/html/2609.30487#ThmPro3)\)\. Therefore, when stability or spectral control is desired, inserting a contractive activation after pooling will enforce a uniform Lipschitz bound, curb eigenvalue inflation, and improve numerical stability\.

The ReEig layer introduced in[Huang and Gool \[2017\]](https://arxiv.org/html/2609.30487#bib.bib2)performs a spectral rectification by eigendecomposing𝐗2=𝐐​𝚲​𝐐⊤\\mathbf\{X\}\_\{2\}=\\mathbf\{Q\}\\mathbf\{\\Lambda\}\\mathbf\{Q\}^\{\\top\}, replacing each eigenvalueλi\\lambda\_\{i\}withmax⁡\{λi,ϵ\}\\max\\\{\\lambda\_\{i\},\\epsilon\\\}, and reconstructing𝐗3=𝐐​max⁡\{𝚲,ϵ​𝐈\}​𝐐⊤\\mathbf\{X\}\_\{3\}=\\mathbf\{Q\}\\max\\\{\\mathbf\{\\Lambda\},\\epsilon\\mathbf\{I\}\\\}\\mathbf\{Q\}^\{\\top\}\. However, the ReEig activation operator is not scale\-invariant \(positively homogeneous\); forc\>0c\>0, we will havec​𝐗3≠𝐐​max⁡\{c​𝚲,ϵ​𝐈\}​𝐐⊤c\\mathbf\{X\}\_\{3\}\\neq\\mathbf\{Q\}\\max\\\{c\\mathbf\{\\Lambda\},\\epsilon\\mathbf\{I\}\\\}\\mathbf\{Q\}^\{\\top\}, whenever a scaled eigenvalue crosses the threshold\. We can modify the ReEig operation to be scale\-invariant:

𝐗3=s⁡\(𝐗2\)​𝐐​exp⁡\(a⁡\(log⁡\(𝚲\)−log⁡\(s⁡\(𝐗2\)\)​𝐈\)\)​𝐐⊤,\\mathbf\{X\}\_\{3\}=s\(\\mathbf\{X\}\_\{2\}\)\\mathbf\{Q\}\\exp\(a\(\\log\(\\mathbf\{\\Lambda\}\)\-\\log\(s\(\\mathbf\{X\}\_\{2\}\)\)\\mathbf\{I\}\)\)\\mathbf\{Q\}^\{\\top\},wherea⁡\(⋅\)a\(\\cdot\)is elementwise activation, ands⁡\(𝐗2\)s\(\\mathbf\{X\}\_\{2\}\)is any 1\-homogeneous scale functional, e\.g\.,s⁡\(𝐗2\)=1m1​trace​\(𝐗2\)s\(\\mathbf\{X\}\_\{2\}\)=\\frac\{1\}\{m\_\{1\}\}\\text\{trace\}\(\\mathbf\{X\}\_\{2\}\)\. The map is positively homogeneous:log⁡\(c​𝐗2\)=𝐐⁡\[log⁡\(𝚲\)\+log⁡\(c\)​𝐈\]​𝐐⊤\\log\(c\\mathbf\{X\}\_\{2\}\)=\\mathbf\{Q\}\[\\log\(\\mathbf\{\\Lambda\}\)\+\\log\(c\)\\mathbf\{I\}\]\\mathbf\{Q\}^\{\\top\}ands⁡\(c​𝐗2\)=c×s⁡\(𝐗2\)s\(c\\mathbf\{X\}\_\{2\}\)=c\\times s\(\\mathbf\{X\}\_\{2\}\)\. However, embeddings⁡\(𝐗2\)s\(\\mathbf\{X\}\_\{2\}\)inside the spectral nonlinearity couples the scale and eigenvalue paths in backpropagation, requiring extra spectral Fréchet operations and raising the computational burden \(≈𝒪⁡\(m13\)\\approx\\mathcal\{O\}\(m\_\{1\}^\{3\}\)\)\. Another type of activation, introduced by[Zhang et al\. \[2018\]](https://arxiv.org/html/2609.30487#bib.bib12), directly applies the scalar functionexp⁡\(⋅\)\\exp\(\\cdot\),sinh⁡\(⋅\)\\sinh\(\\cdot\)orcosh⁡\(⋅\)\\cosh\(\\cdot\)elementwise to𝐗2\\mathbf\{X\}\_\{2\}\. However,exp⁡\(⋅\)\\exp\(\\cdot\)andcosh⁡\(⋅\)\\cosh\(\\cdot\)grow rapidly; their derivatives can trigger exploding activations/gradients and inflate condition numbers\. Moreover, elementwise maps provide no direct spectral control and can mix scales, obscuring links to the underlying process\.

Motivated by the need to control the Lipschitz bound and by the fact that our pooling operator is already non\-linear, we develop an intrinsic geodesic\-shrinkage activation on𝒮\+m1\\mathcal\{S\}^\{m\_\{1\}\}\_\{\+\}, which preserves scale behavior while yielding tractable derivatives\.

#### Geodesic shrinkage

Letϕ\\phidenote either the principal matrix logarithmlog\\logor the log\-Cholesky coordinate mapΦ\\Phi\. The Pooling layer computes the Fréchet mean inϕ\\phi\-coordinates via

F⁡\(𝐗,\{𝐖1h\}h=1H\)=ϕ−1​\(1H​∑h=1Hϕ⁡\(𝐖1h​𝐗​\(𝐖1h\)⊤\)\)\.F\(\\mathbf\{X\};\\\{\\mathbf\{W\}\_\{1\}^\{h\}\\\}\_\{h=1\}^\{H\}\)=\\phi^\{\-1\}\(\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\phi\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\)\.We want to regularize the Fréchet mean𝐗2=F⁡\(𝐗,\{𝐖1h\}h=1H\)\\mathbf\{X\}\_\{2\}=F\(\\mathbf\{X\};\\\{\\mathbf\{W\}\_\{1\}^\{h\}\\\}\_\{h=1\}^\{H\}\)by moving𝐗2\\mathbf\{X\}\_\{2\}back toward the identity matrix𝐈\\mathbf\{I\}along the unique geodesic joining𝐈\\mathbf\{I\}to𝐗2\\mathbf\{X\}\_\{2\}, stopping at a proportion0<α<10<\\alpha<1of the total geodesic length\. Becauseϕ\\phimaps𝐈\\mathbf\{I\}to the origin of the coordinate space:ϕ⁡\(𝐈\)=𝟎\\phi\(\\mathbf\{I\}\)=\\mathbf\{0\}, geodesic distances from𝐈\\mathbf\{I\}reduce to Euclidean norms:dLE​\(𝐈,𝐗2\)=‖log⁡\(𝐗2\)‖Fd\_\{\\text\{LE\}\}\(\\mathbf\{I\},\\mathbf\{X\}\_\{2\}\)=\\\|\\log\(\\mathbf\{X\}\_\{2\}\)\\\|\_\{F\}under the LE metric, anddLC​\(𝐈,𝐗2\)=‖Φ⁡\(𝐗2\)‖2d\_\{\\text\{LC\}\}\(\\mathbf\{I\},\\mathbf\{X\}\_\{2\}\)=\\\|\\Phi\(\\mathbf\{X\}\_\{2\}\)\\\|\_\{2\}under the LC metric\. Define the activation output𝐗3\\mathbf\{X\}\_\{3\}as the point at distanceα​dg​\(𝐈,𝐗2\)\\alpha d\_\{g\}\(\\mathbf\{I\},\\mathbf\{X\}\_\{2\}\)from𝐈\\mathbf\{I\}along that geodesic\. Since geodesics are straight lines inϕ\\phi\-coordinates, this construction has a particularly simple coordinate description:

ϕ⁡\(𝐗3\)=\(1−α\)​ϕ​\(𝐈\)\+α​ϕ​\(𝐗2\)=α​ϕ​\(𝐗2\),\\phi\(\\mathbf\{X\}\_\{3\}\)=\(1\-\\alpha\)\\phi\(\\mathbf\{I\}\)\+\\alpha\\phi\(\\mathbf\{X\}\_\{2\}\)=\\alpha\\phi\(\\mathbf\{X\}\_\{2\}\),or, equivalently,𝐗3=ϕ−1​\(α​ϕ​\(𝐗2\)CLOSE\\mathbf\{X\}\_\{3\}=\\phi^\{\-1\}\(\\alpha\\phi\(\\mathbf\{X\}\_\{2\}\)\)\. In words: we shrink the coordinate vector toward zero \(the image of𝐈\\mathbf\{I\}\) by a factorα\\alpha, and then map back to the manifold\. This yieldsdg​\(𝐈,𝐗3\)=α​dg​\(𝐈,𝐗2\)d\_\{g\}\(\\mathbf\{I\},\\mathbf\{X\}\_\{3\}\)=\\alpha d\_\{g\}\(\\mathbf\{I\},\\mathbf\{X\}\_\{2\}\)\. Hence the operation contracts distances from the identity exactly byα\\alpha\. The effect is directly analogous to weight decay in Euclidean networks, but carried out in intrinsic coordinates of the manifold\.

To keep overhead negligible, rather than “pool, then activate,” we combine the Pooling layer and the Activation layer into one layer, namely, the “Activated Pooling” layer in Figure[1](https://arxiv.org/html/2609.30487#S2.F1)\. The activated\-pooling computes the average in theϕ\\phi\-chart, shrinks the average byα\\alpha, and maps the shrinked average back to the manifold:

𝐗3=ϕ−1​\(αH​∑h=1Hϕ⁡\(𝐖1h​𝐗​\(𝐖1h\)⊤\)\)\.\\mathbf\{X\}\_\{3\}=\\phi^\{\-1\}\(\\frac\{\\alpha\}\{H\}\\sum\_\{h=1\}^\{H\}\\phi\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\)\.This eliminates the intermediate𝐗2\\mathbf\{X\}\_\{2\}\. The only extra computation relative to plain pooling is a scalar multiplication byα\\alphain theϕ\\phi\-domain\.

### 2\.4Concatenation

Substituting the intermediate layers, the composite mapping from the input𝐗⁡\(t\)\\mathbf\{X\}\(t\)to𝐘h​\(t\)\\mathbf\{Y\}^\{h\}\(t\)is:

𝐘h​\(t\)=Log𝐈g​\(𝐖4h​ϕ−1​\(αH​∑k=1Hϕ⁡\(𝐖1k​𝐗​\(t\)​\(𝐖1k\)⊤\)\)​\(𝐖4h\)⊤\)\.\\mathbf\{Y\}^\{h\}\(t\)=\\text\{Log\}^\{g\}\_\{\\mathbf\{I\}\}\\left\(\\mathbf\{W\}\_\{4\}^\{h\}\\phi^\{\-1\}\(\\frac\{\\alpha\}\{H\}\\sum\_\{k=1\}^\{H\}\\phi\\left\(\\mathbf\{W\}\_\{1\}^\{k\}\\mathbf\{X\}\(t\)\(\\mathbf\{W\}\_\{1\}^\{k\}\)^\{\\top\}\)\\right\)\(\\mathbf\{W\}\_\{4\}^\{h\}\)^\{\\top\}\\right\)\.\(1\)For each headh=1,…,Hh=1,\\ldots,Hand timett, the Concatenation layer maps a symmetric matrix𝐘h​\(t\)∈Sym​\(m2\)\\mathbf\{Y\}^\{h\}\(t\)\\in\\text\{Sym\}\(m\_\{2\}\)to a vectorvl​\(𝐘h​\(t\)\)∈ℝm2​\(m2\+1\)/2\\text\{vl\}\(\\mathbf\{Y\}^\{h\}\(t\)\)\\in\\mathbb\{R\}^\{m\_\{2\}\(m\_\{2\}\+1\)/2\}, and then concatenates the vectors across heads\. We equipSym​\(m2\)\\text\{Sym\}\(m\_\{2\}\)with the Frobenius inner product andℝm2​\(m2\+1\)/2\\mathbb\{R\}^\{m\_\{2\}\(m\_\{2\}\+1\)/2\}with the Euclidean inner product and requirevl​\(⋅\)\\text\{vl\}\(\\cdot\)to be an isometry:⟨𝐘1,𝐘2⟩F=⟨vl​\(𝐘1\),vl​\(𝐘2\)⟩2\\langle\\mathbf\{Y\}\_\{1\},\\mathbf\{Y\}\_\{2\}\\rangle\_\{F\}=\\langle\\text\{vl\}\(\\mathbf\{Y\}\_\{1\}\),\\text\{vl\}\(\\mathbf\{Y\}\_\{2\}\)\\rangle\_\{2\}, for all𝐘1,𝐘2\\mathbf\{Y\}\_\{1\},\\mathbf\{Y\}\_\{2\}\. Isometry guarantees norm preservation and, crucially for backpropagation, it makes the adjoint equal the inverse\. Fix any indexingκ:\{\(i,j\):1≤j≤i≤m2\}↦\{1,…,m2​\(m2\+1\)/2\}\\kappa:\\\{\(i,j\):1\\leq j\\leq i\\leq m\_\{2\}\\\}\\mapsto\\\{1,\\ldots,m\_\{2\}\(m\_\{2\}\+1\)/2\\\}\(e\.g\., lexicographic\)\. Define

\[vl​\(𝐘\)\]κ⁡\(i,i\)=Yi,i,\[vl​\(𝐘\)\]κ⁡\(i,j\)=2​Yi,j,for​i\>j\.\[\\text\{vl\}\(\\mathbf\{Y\}\)\]\_\{\\kappa\(i,i\)\}=Y\_\{i,i\},~~~\[\\text\{vl\}\(\\mathbf\{Y\}\)\]\_\{\\kappa\(i,j\)\}=\\sqrt\{2\}Y\_\{i,j\},\\text\{ for \}i\>j\.It is easy to prove thatvl​\(⋅\)\\text\{vl\}\(\\cdot\)is an isometry, and the adjoint equals the inverse\. Finally, the output of the network is𝒚​\(t\)⊤=\[vl​\(𝐘1​\(t\)\)⊤,…,vl​\(𝐘H​\(t\)\)⊤\]\\boldsymbol\{y\}\(t\)^\{\\top\}=\[\\text\{vl\}\(\\mathbf\{Y\}^\{1\}\(t\)\)^\{\\top\},\\ldots,\\text\{vl\}\(\\mathbf\{Y\}^\{H\}\(t\)\)^\{\\top\}\]\.

## 3Functional autoencoder with geometry\-aware layers

Figure 2:The FAEclust architecture generalizes the multilayer perceptron \(MLP\) autoencoder by appending functional layers at the encoder input and the decoder output\. In the above figure, the plum block denotes the MLP autoencoder, and the blue components denote functional nodes/edges\.The matrix\-to\-vector network in Figure[1](https://arxiv.org/html/2609.30487#S2.F1)is modular: it can be paired with any network that accept vector\-valued functions\. We here adapt the FAEclust architecture\[[Singh et al\., 2025](https://arxiv.org/html/2609.30487#bib.bib1)\]to learn a latent embedding of the Euclidean vector\-valued function𝒚⁡\(t\)\\boldsymbol\{y\}\(t\)and, via a geometry\-aware decoder, to reconstruct the original SPD trajectory𝐗⁡\(t\)\\mathbf\{X\}\(t\)\. FAEclust comprises an encoderℰ\\mathcal\{E\}and a decoder𝒟\\mathcal\{D\}\(see Figure[2](https://arxiv.org/html/2609.30487#S3.F2)\)\. Letℋ⁡\(\[t0,t1\],ℝp\)\\mathcal\{H\}\(\[t\_\{0\},t\_\{1\}\];\\mathbb\{R\}^\{p\}\)denote the separable Hilbert space of square\-integrable functions that are defined on\[t0,t1\]\[t\_\{0\},t\_\{1\}\]and taking values inℝp\\mathbb\{R\}^\{p\}\. Given app\-variate function𝒚∈ℋ⁡\(\[t0,t1\],ℝp\)\\boldsymbol\{y\}\\in\\mathcal\{H\}\(\[t\_\{0\},t\_\{1\}\],\\mathbb\{R\}^\{p\}\), the encoder maps the entire trajectory to a latent multivariate data point𝒙=ℰ⁡\(𝒚\)∈ℝs\\boldsymbol\{x\}=\\mathcal\{E\}\(\\boldsymbol\{y\}\)\\in\\mathbb\{R\}^\{s\}; the decoder then reconstructs a functional trajectory𝒚^=𝒟⁡\(𝒙\)∈ℋ⁡\(\[t0,t1\],ℝp\)\\hat\{\\boldsymbol\{y\}\}=\\mathcal\{D\}\(\\boldsymbol\{x\}\)\\in\\mathcal\{H\}\(\[t\_\{0\},t\_\{1\}\],\\mathbb\{R\}^\{p\}\)\. Cluster analysis is performed on the embedded data in the latent space\. Unlike a conventional MLP autoencoder, FAEclust introduces four functional layers: one placed at the encoder entrance and three at the decoder exit, which act on functions rather than vectors\. The feedforward equation for the functional layer in the encoder is given by

𝒙\(1\)=a⁡\(∫t0t1𝑾\(1\)​\(t\)​𝒚​\(t\)​𝑑t\+𝒃\),\\boldsymbol\{x\}^\{\(1\)\}=a\(\\int\_\{t\_\{0\}\}^\{t\_\{1\}\}\\boldsymbol\{W\}^\{\(1\)\}\(t\)\\boldsymbol\{y\}\(t\)dt\+\\boldsymbol\{b\}\),\(2\)where𝑾\(1\)​\(t\)\\boldsymbol\{W\}^\{\(1\)\}\(t\)is a functional weight matrix,𝒃\\boldsymbol\{b\}is a bias vector, andaais an activation function\.𝒙\(1\)\\boldsymbol\{x\}^\{\(1\)\}is then fed into the MLP autoencoder\. The equations for the three functional layers in the decoder are:

𝒚^\(1\)​\(t\)\\displaystyle\\hat\{\\boldsymbol\{y\}\}^\{\(1\)\}\(t\)=a⁡\(𝓦\(1\)​\(t\)​𝒙^\+𝒃1​\(t\)\),\\displaystyle=a\(\\boldsymbol\{\\mathcal\{W\}\}^\{\(1\)\}\(t\)\\hat\{\\boldsymbol\{x\}\}\+\\boldsymbol\{b\}\_\{1\}\(t\)\),𝒚^\(2\)​\(t\)\\displaystyle\\hat\{\\boldsymbol\{y\}\}^\{\(2\)\}\(t\)=a⁡\(𝓦\(2\)​\(t\)​𝒚^\(1\)​\(t\)\+𝒃2​\(t\)\),\\displaystyle=a\(\\boldsymbol\{\\mathcal\{W\}\}^\{\(2\)\}\(t\)\\hat\{\\boldsymbol\{y\}\}^\{\(1\)\}\(t\)\+\\boldsymbol\{b\}\_\{2\}\(t\)\),𝒚^​\(t\)\\displaystyle\\hat\{\\boldsymbol\{y\}\}\(t\)=𝓦\(3\)​\(t\)​𝒚^\(2\)​\(t\),\\displaystyle=\\boldsymbol\{\\mathcal\{W\}\}^\{\(3\)\}\(t\)\\hat\{\\boldsymbol\{y\}\}^\{\(2\)\}\(t\),where𝒙^\\hat\{\\boldsymbol\{x\}\}is the output of the MLP autoencoder,\{𝓦\(l\)​\(t\),l=1,2,3\}\\\{\\boldsymbol\{\\mathcal\{W\}\}^\{\(l\)\}\(t\),l=1,2,3\\\}are functional weight matrices, and\{𝒃1​\(t\),𝒃2​\(t\)\}\\\{\\boldsymbol\{b\}\_\{1\}\(t\),\\boldsymbol\{b\}\_\{2\}\(t\)\\\}are functional biases\. For the output layer, the activation function is linear, and there is no functional bias\.

For a matrix𝐔∈Sym​\(m\)\\mathbf\{U\}\\in\\text\{Sym\}\(m\), let𝒱⁡\(𝐔\)=\(vec​\(⌊𝐔⌋\),vec​\(diag​\(𝐔\)\)\)∈ℝm⁡\(m\+1\)/2\\mathcal\{V\}\(\\mathbf\{U\}\)=\(\\text\{vec\}\(\\lfloor\\mathbf\{U\}\\rfloor\),\\text\{vec\}\(\\text\{diag\}\(\\mathbf\{U\}\)\)\)\\in\\mathbb\{R\}^\{m\(m\+1\)/2\}denote the vector of the lower triangular entries of𝐔\\mathbf\{U\}\. Note that𝒱\\mathcal\{V\}is different from the log\-Cholesky coordinate mapΦ\\Phi\. The inverse𝒱−1\\mathcal\{V\}^\{\-1\}maps a vector𝒖∈ℝm⁡\(m\+1\)/2\\boldsymbol\{u\}\\in\\mathbb\{R\}^\{m\(m\+1\)/2\}back to anm×mm\\times mmatrix inSym​\(m\)\\text\{Sym\}\(m\):𝐔=𝒱−1​\(𝒱​\(𝐔\)\)\\mathbf\{U\}=\\mathcal\{V\}^\{\-1\}\(\\mathcal\{V\}\(\\mathbf\{U\}\)\)\. We can directly apply FAEclust and train the network to output𝒚^​\(t\)∈ℝm⁡\(m\+1\)/2\\hat\{\\boldsymbol\{y\}\}\(t\)\\in\\mathbb\{R\}^\{m\(m\+1\)/2\}that approximates𝒱⁡\(Log𝐈LC​\(𝐗⁡\(t\)\)\)\\mathcal\{V\}\(\\text\{Log\}^\{\\text\{LC\}\}\_\{\\mathbf\{I\}\}\(\\mathbf\{X\}\(t\)\)\)or𝒱⁡\(Log𝐈LE​\(𝐗⁡\(t\)\)\)\\mathcal\{V\}\(\\text\{Log\}^\{\\text\{LE\}\}\_\{\\mathbf\{I\}\}\(\\mathbf\{X\}\(t\)\)\)\. In practice, this is computationally prohibitive: each functional weight matrix𝓦\(l\)​\(t\)\\boldsymbol\{\\mathcal\{W\}\}^\{\(l\)\}\(t\)will have size aboutm2×m2m^\{2\}\\times m^\{2\}, yielding𝒪⁡\(m4\)\\mathcal\{O\}\(m^\{4\}\)parameters per layer and a heavy training burden\. We cut computation by decoding in a low\-dimensional tangent space, then lifting via conjugation with thin, column\-orthonormal matrices, giving an injective isometry for the Frobenius inner product\. Let𝒚^\(1\)​\(t\)∈ℝp1​\(p1\+1\)/2\\hat\{\\boldsymbol\{y\}\}^\{\(1\)\}\(t\)\\in\\mathbb\{R\}^\{p\_\{1\}\(p\_\{1\}\+1\)/2\}, wherep1<mp\_\{1\}<m, and𝒱−1​\(𝒚^\(1\)​\(t\)\)∈Sym​\(p1\)\\mathcal\{V\}^\{\-1\}\(\\hat\{\\boldsymbol\{y\}\}^\{\(1\)\}\(t\)\)\\in\\text\{Sym\}\(p\_\{1\}\)is a trajectory in the tangent space of𝒮\+p1\\mathcal\{S\}^\{p\_\{1\}\}\_\{\+\}\. Let𝐌1∈ℝp2×p1\\mathbf\{M\}\_\{1\}\\in\\mathbb\{R\}^\{p\_\{2\}\\times p\_\{1\}\}and𝐌2∈ℝm×p2\\mathbf\{M\}\_\{2\}\\in\\mathbb\{R\}^\{m\\times p\_\{2\}\}be two thin matrices of orthonormal columns\. The modified functional layers are

𝒚^\(1\)​\(t\)\\displaystyle\\hat\{\\boldsymbol\{y\}\}^\{\(1\)\}\(t\)=\\displaystyle=a⁡\(𝓦\(1\)​\(t\)​𝒙^\+𝒃1​\(t\)\),\\displaystyle a\(\\boldsymbol\{\\mathcal\{W\}\}^\{\(1\)\}\(t\)\\hat\{\\boldsymbol\{x\}\}\+\\boldsymbol\{b\}\_\{1\}\(t\)\),𝒚^\(2\)​\(t\)\\displaystyle\\hat\{\\boldsymbol\{y\}\}^\{\(2\)\}\(t\)=\\displaystyle=a⁡\(𝓦\(2\)​\(t\)​𝒱​\(𝐌1​𝒱−1​\(𝒚^\(1\)​\(t\)\)​𝐌1⊤\)\+𝒃2​\(t\)\),\\displaystyle a\(\\boldsymbol\{\\mathcal\{W\}\}^\{\(2\)\}\(t\)\{\\color\[rgb\]\{0\.5,0\.5,0\.5\}\\mathcal\{V\}\(\\mathbf\{M\}\_\{1\}\\mathcal\{V\}^\{\-1\}\(\\hat\{\\boldsymbol\{y\}\}^\{\(1\)\}\(t\)\)\\mathbf\{M\}\_\{1\}^\{\\top\}\)\}\+\\boldsymbol\{b\}\_\{2\}\(t\)\),\(3\)𝐗^​\(t\)\\displaystyle\\hat\{\\mathbf\{X\}\}\(t\)=\\displaystyle=Exp𝐈g​\(𝐌2​𝒱−1​\(𝒚^\(2\)​\(t\)\)​𝐌2⊤\),\\displaystyle\\text\{Exp\}^\{g\}\_\{\\mathbf\{I\}\}\(\{\\color\[rgb\]\{0\.5,0\.5,0\.5\}\\mathbf\{M\}\_\{2\}\\mathcal\{V\}^\{\-1\}\(\\hat\{\\boldsymbol\{y\}\}^\{\(2\)\}\(t\)\)\\mathbf\{M\}\_\{2\}^\{\\top\}\}\),\(4\)where𝐗^​\(t\)∈𝒮\+m\\hat\{\\mathbf\{X\}\}\(t\)\\in\\mathcal\{S\}^\{m\}\_\{\+\}is the reconstructed SPD trajectory, andExp𝐈g\\text\{Exp\}^\{g\}\_\{\\mathbf\{I\}\}is the Riemannian exponential at𝐈\\mathbf\{I\}under the chosen metricgg\(i\.e\., LC or LE\)\. In Eq\. equation[3](https://arxiv.org/html/2609.30487#S3.E3), we lift fromSym​\(p1\)\\text\{Sym\}\(p\_\{1\}\)toSym​\(p2\)\\text\{Sym\}\(p\_\{2\}\)by conjugation before applying the nonlinearity in vectorized form, and in Eq\. equation[4](https://arxiv.org/html/2609.30487#S3.E4)we lift again toSym​\(m\)\\text\{Sym\}\(m\), and map back to the manifold\. In line with the output layer of FAEclust, Eq\. equation[4](https://arxiv.org/html/2609.30487#S3.E4)only involves a bilinear mapping \(in the tangent space\)\.

## 4Network training

For a Riemannian metricggon𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}\(i\.e\., the LC or LE metric\), defineℋg\(\[t0,t1\];𝒮\+m\)=\{𝑿:\[t0,t1\]→𝒮\+mmeasurable\|∫t0t1dg\(𝑿\(t\),𝐈\)2dt<∞\}\\mathcal\{H\}\_\{g\}\(\[t\_\{0\},t\_\{1\}\];\\mathcal\{S\}^\{m\}\_\{\+\}\)=\\\{\{\\bm\{\\mathsfit\{X\}\}\}:\[t\_\{0\},t\_\{1\}\]\\rightarrow\\mathcal\{S\}^\{m\}\_\{\+\}\\text\{ measurable \}\|\\int\_\{t\_\{0\}\}^\{t\_\{1\}\}d\_\{g\}\(\{\\bm\{\\mathsfit\{X\}\}\}\(t\),\\mathbf\{I\}\)^\{2\}dt<\\infty\\\}\. The matrix\-to\-vector module is the mappingℱ\\mathcal\{F\}:𝑿∈ℋg​\(\[t0,t1\],𝒮\+m\)↦𝒚=ℱ⁡\(𝑿\)∈ℋ⁡\(\[t0,t1\],ℝp\)\{\\bm\{\\mathsfit\{X\}\}\}\\in\\mathcal\{H\}\_\{g\}\(\[t\_\{0\},t\_\{1\}\];\\mathcal\{S\}^\{m\}\_\{\+\}\)\\mapsto\\boldsymbol\{y\}=\\mathcal\{F\}\(\{\\bm\{\\mathsfit\{X\}\}\}\)\\in\\mathcal\{H\}\(\[t\_\{0\},t\_\{1\}\];\\mathbb\{R\}^\{p\}\), and theMatFAEnetwork is the mapping𝒟∘ℰ∘ℱ\\mathcal\{D\}\\circ\\mathcal\{E\}\\circ\\mathcal\{F\}:𝑿∈ℋg​\(\[t0,t1\],𝒮\+m\)↦𝑿^=𝒟⁡\(ℰ⁡\(ℱ⁡\(𝑿\)\)\)∈ℋg​\(\[t0,t1\],𝒮\+m\)\{\\bm\{\\mathsfit\{X\}\}\}\\in\\mathcal\{H\}\_\{g\}\(\[t\_\{0\},t\_\{1\}\];\\mathcal\{S\}^\{m\}\_\{\+\}\)\\mapsto\\hat\{\{\\bm\{\\mathsfit\{X\}\}\}\}=\\mathcal\{D\}\(\\mathcal\{E\}\(\\mathcal\{F\}\(\{\\bm\{\\mathsfit\{X\}\}\}\)\)\)\\in\\mathcal\{H\}\_\{g\}\(\[t\_\{0\},t\_\{1\}\];\\mathcal\{S\}^\{m\}\_\{\+\}\)\. Given a set of𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}\-valued trajectories\{𝑿i∈ℋg\(\[t0,t1\];𝒮\+m\)\}i=1n\\\{\{\\bm\{\\mathsfit\{X\}\}\}\_\{i\}\\in\\mathcal\{H\}\_\{g\}\(\[t\_\{0\},t\_\{1\}\];\\mathcal\{S\}^\{m\}\_\{\+\}\)\\\}\_\{i=1\}^\{n\},MatFAEproduces reconstructions\{𝑿^i∈ℋg\(\[t0,t1\];𝒮\+m\)\}i=1n\\\{\\hat\{\{\\bm\{\\mathsfit\{X\}\}\}\}\_\{i\}\\in\\mathcal\{H\}\_\{g\}\(\[t\_\{0\},t\_\{1\}\];\\mathcal\{S\}^\{m\}\_\{\+\}\)\\\}\_\{i=1\}^\{n\}\. Under the LE metric, the reconstruction loss is

ℒr=∑i=1n∫t0t1dLE​\(𝑿i​\(t\),𝑿^i​\(t\)\)2​𝑑t=∑i=1n∫t0t1‖log⁡\(𝑿i​\(t\)\)−𝐌2​𝒱−1​\(𝒚^i\(2\)​\(t\)\)​𝐌2⊤‖F2​𝑑t\.\\mathcal\{L\}\_\{r\}=\\sum\_\{i=1\}^\{n\}\\int\_\{t\_\{0\}\}^\{t\_\{1\}\}d\_\{\\text\{LE\}\}\(\{\\bm\{\\mathsfit\{X\}\}\}\_\{i\}\(t\),\\hat\{\{\\bm\{\\mathsfit\{X\}\}\}\}\_\{i\}\(t\)\)^\{2\}dt=\\sum\_\{i=1\}^\{n\}\\int\_\{t\_\{0\}\}^\{t\_\{1\}\}\\\|\\log\(\{\\bm\{\\mathsfit\{X\}\}\}\_\{i\}\(t\)\)\-\\mathbf\{M\}\_\{2\}\\mathcal\{V\}^\{\-1\}\(\\hat\{\\boldsymbol\{y\}\}^\{\(2\)\}\_\{i\}\(t\)\)\\mathbf\{M\}\_\{2\}^\{\\top\}\\\|\_\{F\}^\{2\}dt\.Under the LC metric, the reconstruction lossℒr\\mathcal\{L\}\_\{r\}is

∑i=1n∫t0t1dLC​\(𝑿i​\(t\),𝑿^i​\(t\)\)2​𝑑t\\displaystyle\\sum\_\{i=1\}^\{n\}\\int\_\{t\_\{0\}\}^\{t\_\{1\}\}d\_\{\\text\{LC\}\}\(\{\\bm\{\\mathsfit\{X\}\}\}\_\{i\}\(t\),\\hat\{\{\\bm\{\\mathsfit\{X\}\}\}\}\_\{i\}\(t\)\)^\{2\}dt=∑i=1n∫t0t1‖⌊𝔏⁡\(𝑿i​\(t\)\)⌋−⌊𝐌2​𝒱−1​\(𝒚^i\(2\)​\(t\)\)​𝐌2⊤⌋‖F2​𝑑t\\displaystyle=\\sum\_\{i=1\}^\{n\}\\int\_\{t\_\{0\}\}^\{t\_\{1\}\}\\\|\\lfloor\\mathfrak\{L\}\(\{\\bm\{\\mathsfit\{X\}\}\}\_\{i\}\(t\)\)\\rfloor\-\\lfloor\\mathbf\{M\}\_\{2\}\\mathcal\{V\}^\{\-1\}\(\\hat\{\\boldsymbol\{y\}\}\_\{i\}^\{\(2\)\}\(t\)\)\\mathbf\{M\}\_\{2\}^\{\\top\}\\rfloor\\\|\_\{F\}^\{2\}dt\+∑i=1n∫t0t1∥log\(diag\(𝔏\(𝑿i\(t\)\)\)\)−12diag\(𝐌2𝒱−1\(𝒚^i\(2\)\(t\)\)𝐌2⊤\)∥F2dt\.\\displaystyle\+\\sum\_\{i=1\}^\{n\}\\int\_\{t\_\{0\}\}^\{t\_\{1\}\}\\\|\\log\(\\text\{diag\}\(\\mathfrak\{L\}\(\{\\bm\{\\mathsfit\{X\}\}\}\_\{i\}\(t\)\)\)\)\-\\frac\{1\}\{2\}\\text\{diag\}\(\\mathbf\{M\}\_\{2\}\\mathcal\{V\}^\{\-1\}\(\\hat\{\\boldsymbol\{y\}\}\_\{i\}^\{\(2\)\}\(t\)\)\\mathbf\{M\}\_\{2\}^\{\\top\}\)\\\|\_\{F\}^\{2\}dt\.
The training objective function ofMatFAEcomprises three components: \(1\)ℒr\\mathcal\{L\}\_\{r\}, the reconstruction loss; \(2\)ℒw\\mathcal\{L\}\_\{w\}, a regularization term on the functional weights and functional biases\{𝑾\(1\)\(t\)\\\{\\boldsymbol\{W\}^\{\(1\)\}\(t\),𝓦\(1\)\(t\),𝓦\(2\)\(t\),𝓦\(3\)\(t\),𝒃1\(t\),𝒃2\(t\)\}\\boldsymbol\{\\mathcal\{W\}\}^\{\(1\)\}\(t\),\\boldsymbol\{\\mathcal\{W\}\}^\{\(2\)\}\(t\),\\boldsymbol\{\\mathcal\{W\}\}^\{\(3\)\}\(t\),\\boldsymbol\{b\}\_\{1\}\(t\),\\boldsymbol\{b\}\_\{2\}\(t\)\\\}; and \(3\)ℒdiv\\mathcal\{L\}\_\{\\mathrm\{div\}\}, a row\-space decorrelation penalty on the congruence weight matrices\{𝐖1h\}h=1H\\\{\\mathbf\{W\}\_\{1\}^\{h\}\\\}\_\{h=1\}^\{H\}and\{𝐖4h\}h=1H\\\{\\mathbf\{W\}\_\{4\}^\{h\}\\\}\_\{h=1\}^\{H\}, The detailed formulation ofℒw\\mathcal\{L\}\_\{w\}is given in\[[Singh et al\., 2025](https://arxiv.org/html/2609.30487#bib.bib1)\], while the formulation ofℒdiv\\mathcal\{L\}\_\{\\mathrm\{div\}\}is provided in Appendix[F](https://arxiv.org/html/2609.30487#A6)\.

#### Riemannian gradient

MatFAEinvolves a few matrix parameters that are constrained to Stiefel manifolds\. We here briefly outline the gradient\-descent procedure for optimizing them\. LetSt​\(m,p\)=\{𝐌∈ℝm×p:𝐌⊤​𝐌=𝐈\}\\text\{St\}\(m,p\)=\\\{\\mathbf\{M\}\\in\\mathbb\{R\}^\{m\\times p\}:\\mathbf\{M\}^\{\\top\}\\mathbf\{M\}=\\mathbf\{I\}\\\}denote the Stiefel manifold ofm×pm\\times pmatrices with orthonormal columns \(i\.e\.,m\>pm\>p\)\. For any smooth mappingf:St​\(m,p\)→ℝf:\\text\{St\}\(m,p\)\\rightarrow\\mathbb\{R\}, let𝐃=∇𝐌f​\(𝐌\)\\mathbf\{D\}=\\nabla\_\{\\mathbf\{M\}\}f\(\\mathbf\{M\}\)denote the normal Euclidean gradient\. Then project𝐃\\mathbf\{D\}onto the tangent space to get the Riemannian gradient:Δ=𝐃−12​𝐌​\(𝐌⊤​𝐃\+𝐃⊤​𝐌\)\\Delta=\\mathbf\{D\}\-\\frac\{1\}\{2\}\\mathbf\{M\}\(\\mathbf\{M\}^\{\\top\}\\mathbf\{D\}\+\\mathbf\{D\}^\{\\top\}\\mathbf\{M\}\)\. For a given step sizeη\>0\\eta\>0, compute a thin QR factorization of𝐌\+η⁡\(−Δ\)\\mathbf\{M\}\+\\eta\(\-\\Delta\):𝐌−η​Δ=𝐐𝐑\\mathbf\{M\}\-\\eta\\Delta=\\mathbf\{Q\}\\mathbf\{R\}; then the updated𝐌\\mathbf\{M\}is𝐐​diag​\(sign​\(diag​\(𝐑\)\)\)\\mathbf\{Q\}\\text\{diag\}\(\\text\{sign\}\(\\text\{diag\}\(\\mathbf\{R\}\)\)\), wheresign​\(⋅\)\\text\{sign\}\(\\cdot\)is the sign function, andsign​\(0\)=1\\text\{sign\}\(0\)=1\.

#### Matrix factorization in backpropagation

The matrix\-to\-vector mappingℱ\\mathcal\{F\}involves eigen\-decomposition or Cholesky factorization \(e\.g\.,𝐗3​\(t\)=ϕ−1​\(αH​∑h=1Hϕ⁡\(𝐖1h​𝐗​\(t\)​\(𝐖1h\)⊤\)\)\\mathbf\{X\}\_\{3\}\(t\)=\\phi^\{\-1\}\(\\frac\{\\alpha\}\{H\}\\sum\_\{h=1\}^\{H\}\\phi\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(t\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\)for geodesic\-shrinkage activated pooling\)\. Computing the partial derivatives required for backpropagation is therefore quite demanding\. Appendix[E](https://arxiv.org/html/2609.30487#A5)provides the detailed derivations, as well as the functional gradients required for backpropagation from the encoderℰ\\mathcal\{E\}to the mappingℱ\\mathcal\{F\}\.

## 5Experiments

### 5\.1Cluster analysis

MatFAEis available as a fully documented Python package\. We here compareMatFAEwith TSRVF\[[Dai et al\., 2020](https://arxiv.org/html/2609.30487#bib.bib14)\], GeoAtt\[[Dan et al\., 2022](https://arxiv.org/html/2609.30487#bib.bib15)\], SPDNet\[[Huang and Gool, 2017](https://arxiv.org/html/2609.30487#bib.bib2)\], PGA\[[Fletcher et al\., 2004](https://arxiv.org/html/2609.30487#bib.bib16)\], MLP autoencoder \(DC\), and k\-means applied directly to the pairwise distance matrix \(PDM\) induced by the LE metric\. Appendix[G\.1](https://arxiv.org/html/2609.30487#A7.SS1)briefly explains each benchmark method and reports the configuration settings used for all algorithms\. For methods that produce vector representations, clustering is performed using k\-means\. The number of clusters is selected by maximizing the silhouette score overk∈\{2,3,4,5\}k\\in\\\{2,3,4,5\\\}\. All simulated and real datasets are available in theMatFAEGitHub repository\.

#### Simulation study

We consider ten simulation scenarios, and the data\-generating mechanism for each is provided in detail in Appendix[G\.3](https://arxiv.org/html/2609.30487#A7.SS3)\. For each scenario, we generate 100 independent datasets and apply all benchmark methods to each dataset\. Clustering performance is then summarized by the mean and standard deviation of the AMI and ARI across the 100 replications, reported in Tables[1](https://arxiv.org/html/2609.30487#S5.T1)and[4](https://arxiv.org/html/2609.30487#A7.T4), respectively\. The benchmark is structured as an ablation ladder, with each synthetic dataset designed to isolate a single type of structural signal\. Accordingly, failure on a given rung indicates that the method lacks the feature\-extraction mechanism needed to capture that particular structure\.

Table[1](https://arxiv.org/html/2609.30487#S5.T1)shows that MatFAE performs best overall across the ten simulation scenarios\. It achieves the highest mean AMI in six scenarios \(A, C, D, E, F, and I\) and ties for the best in H and J\. Its main advantage appears in the more difficult settings, especially dwell time \(C\), transition frequency \(D\), and smoothness \(E\), where competing methods show only modest or near\-chance performance\. The gain in E is particularly pronounced: MatFAE reaches 0\.725, whereas the next\-best method, PGA, achieves only 0\.163\. Scenario I is the most challenging for all methods, with all AMI values below 0\.16\. Even in this case, however, MatFAE achieves the highest mean score, suggesting limited but detectable sensitivity to the multi\-scale structure\.

Table 1:AMI scores for the ten simulation scenarios\. The table reports the mean \(top row\) and standard deviation \(bottom row\) of the scores over 100 repetitions\.
#### Real datasets

We used six real resting\-state fMRI datasets from two public neuroimaging repositories: OpenNeuro and the International Neuroimaging Data\-sharing Initiative\. Detailed dataset profiles, preprocessing procedures, and additional analyses are provided in Appendix[G\.4](https://arxiv.org/html/2609.30487#A7.SS4)\. Using the COBRE dataset, Appendix[G\.6](https://arxiv.org/html/2609.30487#A7.SS6)provides a step\-by\-step demonstration of MatMAE’s intrinsic interpretability\. Specifically, it shows how the temporal profiles of the functional weights𝑾\(1\)​\(t\)\\boldsymbol\{W\}^\{\(1\)\}\(t\)in the encoder’s functional layer identify the portions of the SPD trajectory that contribute most strongly to the latent representations\.

Table 2:AMI scores for the six real resting\-state fMRI datasets\.Clustering performance is reported using AMI and ARI in Tables[2](https://arxiv.org/html/2609.30487#S5.T2)and[5](https://arxiv.org/html/2609.30487#A7.T5), respectively\.MatFAEachieves the best performance across all six cohorts, demonstrating its consistent advantage over the competing methods\. At the same time, the absolute AMI values reflect the intrinsic difficulty of unsupervised diagnostic discovery from resting\-state fMRI\. ADHD\-200 and ABIDE\-I are heterogeneous multi\-site cohorts, CNP contains small patient subgroups across four DSM categories, and resting\-state cognition is itself an unconstrained latent state\. Even supervised models on these datasets typically improve only moderately over chance, with reported balanced accuracies of approximately 0\.60\-0\.70 for ABIDE\-I\[[Heinsfeld et al\., 2018](https://arxiv.org/html/2609.30487#bib.bib59),[Saponaro et al\., 2022](https://arxiv.org/html/2609.30487#bib.bib61)\], 0\.72\-0\.81 for COBRE\[[Zeng et al\., 2018](https://arxiv.org/html/2609.30487#bib.bib60)\], and 0\.55\-0\.66 for ADHD\-200\[[Qiu et al\., 2024](https://arxiv.org/html/2609.30487#bib.bib62)\]\. These results are about0\.100\.10\-0\.250\.25above chance, indicating the intrinsic difficulty of prediction from resting\-state fMRI\.

### 5\.2Classification analysis

We further evaluate the encoder blocks of MatFAE in a supervised setting, denoted MatFAE\-C, by removing the decoder and appending a classification head to the latent representation\. We compare MatFAE\-C against three recent classifiers for temporal SPD/graph\-structured data on the same six real datasets: SPD\-SRU\[[Chakraborty et al\., 2018](https://arxiv.org/html/2609.30487#bib.bib21)\], which extends recurrent neural networks to operate directly on sequences of SPD matrices; STAGIN\[[Kim et al\., 2021](https://arxiv.org/html/2609.30487#bib.bib17)\], which learns dynamic brain\-connectome representations via spatio\-temporal graph attention; and BW\-Norm\[[Wang et al\., 2025](https://arxiv.org/html/2609.30487#bib.bib39)\], a Riemannian batch\-normalization method under the Bures\-Wasserstein metric that processes SPD trajectories frame by frame\. Full configuration details for all methods are reported in Appendix[G\.2](https://arxiv.org/html/2609.30487#A7.SS2)\. Classification results are reported in Table[3](https://arxiv.org/html/2609.30487#S5.T3)\. The folds are not site\-stratified and no ComBat harmonisation is applied; therefore, the reported results should be interpreted as within\-distribution classification performance rather than evidence of out\-of\-site generalisation\.

Table 3:Balanced classification accuracy \(mean±\\pmstd over five\-fold CV\) for MatFAE\-C and three classification baselines\. Chance\-level accuracy is shown in parentheses next to each dataset name\.The supervised variantMatFAE\-Cachieves the highest balanced classification accuracy on all six cohorts\. Although the CNP accuracy appears lower in absolute terms, CNP is a four\-class problem, for which chance\-level balanced accuracy is 0\.25 rather than 0\.50; the CNP result therefore remains substantially above chance, and in fact represents the largest relative margin over the best baseline of any cohort \(0\.453 vs\. 0\.338 for BW\-norm, a34%34\\%relative improvement\)\. Statistical significance analyses of the reported improvements are provided in Appendix[G\.5](https://arxiv.org/html/2609.30487#A7.SS5)\. Overall, the performance ofMatFAE\-Cis comparable to conservative resting\-state fMRI classification results reported on multisite psychiatric datasets such as ABIDE\-I and ADHD\-200, where accuracies around 0\.60\-0\.70 are commonly observed under realistic validation settings\. However, it is below some highly optimized supervised studies on single binary datasets, such as COBRE schizophrenia classification, where substantially higher accuracies have been reported using task\-specific feature engineering, feature selection, multimodal inputs, or deep supervised architectures\. These comparisons should therefore be interpreted cautiously, as published results differ substantially in diagnostic task, number of classes, preprocessing, site handling, validation protocol, and whether external generalisation is evaluated\. The purpose ofMatFAE\-Cis not to claim state\-of\-the\-art supervised diagnostic classification, but to demonstrate that the representation learned by theMatFAEencoder contains diagnostically relevant information\. In neuroscience, many neurological and psychiatric conditions are heterogeneous and may contain clinically meaningful subtypes that are not known in advance\.MatFAEcan therefore be used to identify latent subgroups and support symptom\-driven, transdiagnostic characterizations of psychiatric conditions\.

## 6Conclusion

We introducedMatFAE, a geometry\-aware neural architecture for representation learning from functional data valued on the SPD manifold\. Beyond combining SPD\-aware dimensionality reduction with functional feature extraction,MatFAEcontributes three theoretically grounded mechanisms: a regularized multi\-head BiMap layer that captures heterogeneous, time\-varying joint structure inaccessible to a single global projection; a geodesic\-shrinkage activation, derived from our analysis of the non\-contractive, nonlinear properties of Fréchet\-mean pooling, which generalizes to Riemannian manifolds beyond the SPD setting; and an intrinsic interpretability procedure that traces latent representations back to the original SPD\-valued trajectories\. Together, these components allowMatFAEto capture both the geometric structure and the temporal morphology of SPD\-valued functional data within a single, lightweight, non\-recurrent architecture\.

Across ten simulation scenarios constructed as an ablation ladder,MatFAEachieved the best or tied\-best performance on eight of ten rungs, with especially large gains on rungs requiring sensitivity to trajectory smoothness, dwell time, and transition frequency\. On six real resting\-state fMRI cohorts,MatFAEachieved the best AMI and ARI on all six, and its supervised variant, MatFAE\-C, achieved balanced classification accuracies of 0\.61\-0\.74, comparable to or exceeding published results for diagnosis prediction on these cohorts\. The consistency ofMatFAE’s advantage across both idealized synthetic settings and heterogeneous real data, despite the latter’s inherently lower achievable signal, suggests that the architecture captures generalizable structure in SPD\-valued dynamics rather than overfitting to a particular data\-generating process\.

## Acknowledgments

This work was conducted with the financial support of the Research Ireland Centre for Research Training in Digitally\-Enhanced Reality \(d\-real\) under Grant No\. 18/CRT/6224\.

## References

- E\. Akeweje and M\. ZhangLearning mixtures of Gaussian processes through random projection\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 720–739\.Cited by:[§1](https://arxiv.org/html/2609.30487#S1.p3.1)\.
- Allenet al\.\(2014\)E\. A\. Allen, E\. Damaraju, S\. M\. Plis, E\. B\. Erhardt, T\. Eichele, and V\. D\. CalhounTracking whole\-brain connectivity dynamics in the resting state\.Cerebral Cortex24\(3\),pp\. 663–676\.External Links:[Document](https://dx.doi.org/10.1093/cercor/bhs352)Cited by:[item G \( = q 30 , = k 2 \)](https://arxiv.org/html/2609.30487#A7.I3.ix7.p1.1)\.
- Arsignyet al\.\(2006\)V\. Arsigny, P\. Fillard, X\. Pennec, and N\. AyacheLog\-euclidean metrics for fast and simple calculus on diffusion tensors\.Magnetic Resonance in Medicine56\(2\),pp\. 411–421\.External Links:[Document](https://dx.doi.org/10.1002/mrm.20965)Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px2.p1.1)\.
- Arsignyet al\.\(2007\)V\. Arsigny, P\. Fillard, X\. Pennec, and N\. AyacheGeometric means in a novel vector space structure on symmetric positive‐definite matrices\.SIAM Journal on Matrix Analysis and Applications29\(1\),pp\. 328–347\.External Links:[Document](https://dx.doi.org/10.1137/050637996)Cited by:[item H \( = q 30 , = k 2 \)](https://arxiv.org/html/2609.30487#A7.I3.ix8.p1.1),[§2](https://arxiv.org/html/2609.30487#S2.p1.1)\.
- Balasubramanianet al\.\(2025\)K\. Balasubramanian, H\. Müller, and B\. K\. SriperumbudurFunctional linear and single\-index models: a unified approach via gaussian stein identity\.Bernoulli31\(2\),pp\. 973 – 1006\.External Links:[Document](https://dx.doi.org/10.3150/24-BEJ1755)Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px1.p2.1)\.
- Bellecet al\.\(2017\)P\. Bellec, C\. Chu, F\. Chouinard\-Decorte, Y\. Benhajali, D\. S\. Margulies, and R\. C\. CraddockThe neuro bureau adhd\-200 preprocessed repository\.NeuroImage144\(Part B\),pp\. 275–286\.External Links:[Document](https://dx.doi.org/10.1016/j.neuroimage.2016.06.034)Cited by:[item ADHD\-200 \( = n 138 , = q 57 , = m 100 , = k 2 \) \[ , \]](https://arxiv.org/html/2609.30487#A7.I4.ix3),[item ADHD\-200 \( = n 138 , = q 57 , = m 100 , = k 2 \) \[ , \]](https://arxiv.org/html/2609.30487#A7.I4.ix3.1)\.
- Bilderet al\.\(2020\)R\. Bilder, R\. Poldrack, T\. Cannon, E\. London, N\. Freimer, E\. Congdon, K\. Karlsgodt, and F\. Sabb”UCLA consortium for neuropsychiatric phenomics la5c study”\.OpenNeuro\.External Links:[Document](https://dx.doi.org/10.18112/openneuro.ds000030.v1.0.0)Cited by:[item CNP \( = n 257 , = q 57 , = m 100 , = k 4 \) \[ , \]](https://arxiv.org/html/2609.30487#A7.I4.ix1),[item CNP \( = n 257 , = q 57 , = m 100 , = k 4 \) \[ , \]](https://arxiv.org/html/2609.30487#A7.I4.ix1.1)\.
- Bonetet al\.\(2023\)C\. Bonet, B\. Malézieux, A\. Rakotomamonjy, L\. Drumetz, T\. Moreau, M\. Kowalski, and N\. CourtySliced\-Wasserstein on symmetric positive definite matrices for M/EEG signals\.InProceedings of the 40th International Conference on Machine Learning,A\. Krause, E\. Brunskill, K\. Cho, B\. Engelhardt, S\. Sabato, and J\. Scarlett \(Eds\.\),Proceedings of Machine Learning Research, Vol\.202,pp\. 2777–2805\.Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px2.p2.1)\.
- Brookset al\.\(2019\)D\. Brooks, O\. Schwander, F\. Barbaresco, J\. Schneider, and M\. CordRiemannian batch normalization for SPD neural networks\.InAdvances in Neural Information Processing Systems,Vol\.32,pp\.\.Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.30487#S1.p3.1)\.
- Caiet al\.\(2022\)X\. Cai, L\. Xue, and J\. CaoVARIABLE selection for multiple function\-on\-function linear regression\.Statistica Sinica32\(3\),pp\. 1435 – 1465\.External Links:[Document](https://dx.doi.org/10.5705/ss.202020.0473)Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px1.p2.1)\.
- Campet al\.\(2024\)C\. C\. Camp, S\. Noble, D\. Scheinost, A\. Stringaris, and D\. M\. NielsonTest\-retest reliability of functional connectivity in adolescents with depression\.Biological Psychiatry: Cognitive Neuroscience and Neuroimaging9\(1\),pp\. 21–29\.External Links:[Document](https://dx.doi.org/10.1016/j.bpsc.2023.09.002)Cited by:[item CAT\-D \( = n 118 , = q 49 , = m 100 , = k 2 \) \[ , \]](https://arxiv.org/html/2609.30487#A7.I4.ix6),[item CAT\-D \( = n 118 , = q 49 , = m 100 , = k 2 \) \[ , \]](https://arxiv.org/html/2609.30487#A7.I4.ix6.1)\.
- Chakrabortyet al\.\(2022\)R\. Chakraborty, J\. Bouza, J\. H\. Manton, and B\. C\. VemuriManifoldNet: a deep neural network for manifold\-valued data with applications\.IEEE Transactions on Pattern Analysis and Machine Intelligence44\(2\),pp\. 799–810\.External Links:[Document](https://dx.doi.org/10.1109/TPAMI.2020.3003846)Cited by:[§1](https://arxiv.org/html/2609.30487#S1.p3.1)\.
- Chakrabortyet al\.\(2018\)R\. Chakraborty, C\. Yang, X\. Zhen, M\. Banerjee, D\. Archer, D\. Vaillancourt, V\. Singh, and B\. VemuriA statistical recurrent model on the manifold of symmetric positive definite matrices\.InAdvances in Neural Information Processing Systems,Vol\.31,pp\.\.Cited by:[§1](https://arxiv.org/html/2609.30487#S1.p2.1),[§5\.2](https://arxiv.org/html/2609.30487#S5.SS2.p1.1)\.
- Chenet al\.\(2023\)Z\. Chen, T\. Xu, X\. Wu, R\. Wang, Z\. Huang, and J\. KittlerRiemannian local mechanism for spd neural networks\.InProceedings of the Thirty\-Seventh AAAI Conference on Artificial Intelligence and Thirty\-Fifth Conference on Innovative Applications of Artificial Intelligence and Thirteenth Symposium on Educational Advances in Artificial Intelligence,External Links:[Document](https://dx.doi.org/10.1609/aaai.v37i6.25867)Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px2.p1.1)\.
- Chiouet al\.\(2014\)J\. Chiou, Y\. Chen, and Y\. YangMULTIVARIATE functional principal component analysis: a normalization approach\.Statistica Sinica24\(4\),pp\. 1571–1596\.Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px1.p2.1)\.
- Chopraet al\.\(2025\)S\. Chopra, C\. V\. Cocuzza, C\. Lawhead, J\. A\. Ricard, L\. Labache, L\. M\. Patrick, P\. Kumar, A\. Rubenstein, J\. Moses, L\. Chen, C\. Blankenbaker, B\. Gillis, L\. T\. Germine, I\. Harpaz\-Rotem, B\. T\. T\. Yeo, J\. T\. Baker, and A\. J\. HolmesThe Transdiagnostic Connectome Project: an open dataset for studying brain\-behavior relationships in psychiatry\.Scientific Data12\(1\),pp\. 923\.Note:OpenNeurods005237External Links:[Document](https://dx.doi.org/10.1038/s41597-025-04895-z)Cited by:[item TCP \( = n 241 , = q 80 , = m 100 , = k 2 \) \[ , \]](https://arxiv.org/html/2609.30487#A7.I4.ix5),[item TCP \( = n 241 , = q 80 , = m 100 , = k 2 \) \[ , \]](https://arxiv.org/html/2609.30487#A7.I4.ix5.1)\.
- Daiet al\.\(2020\)M\. Dai, Z\. Zhang, and A\. SrivastavaAnalyzing dynamical brain functional connectivity as trajectories on space of covariance matrices\.IEEE Transactions on Medical Imaging39\(3\),pp\. 611–620\.External Links:[Document](https://dx.doi.org/10.1109/TMI.2019.2931708)Cited by:[item TSRVF](https://arxiv.org/html/2609.30487#A7.I1.ix1.p1.1),[§5\.1](https://arxiv.org/html/2609.30487#S5.SS1.p1.1)\.
- Danet al\.\(2022\)T\. Dan, Z\. Huang, H\. Cai, P\. J\. Laurienti, and G\. WuLearning brain dynamics of evolving manifold functional mri data using geometric\-attention neural network\.IEEE Transactions on Medical Imaging41\(10\),pp\. 2752–2763\.External Links:[Document](https://dx.doi.org/10.1109/TMI.2022.3169640)Cited by:[§5\.1](https://arxiv.org/html/2609.30487#S5.SS1.p1.1)\.
- Dette and Tang \(2024\)H\. Dette and J\. TangStatistical inference for function\-on\-function linear regression\.Bernoulli30\(1\),pp\. 304 – 331\.External Links:[Document](https://dx.doi.org/10.3150/23-BEJ1598)Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px1.p2.1)\.
- Di Martinoet al\.\(2014\)A\. Di Martino, C\.\-G\. Yan, Q\. Li, E\. Denio, F\. X\. Castellanos, K\. Alaerts, J\. S\. Anderson, M\. Assaf, S\. Y\. Bookheimer, M\. Dapretto, B\. Deen, S\. Delmonte, I\. Dinstein, B\. Ertl\-Wagner, D\. A\. Fair, L\. Gallagher, D\. P\. Kennedy, C\. L\. Keown, C\. Keysers, J\. E\. Lainhart, C\. Lord, B\. Luna, V\. Menon, N\. J\. Minshew, C\. S\. Monk, S\. Mueller, R\.\-A\. Müller, M\. B\. Nebel, J\. T\. Nigg, K\. O’Hearn, K\. A\. Pelphrey, S\. J\. Peltier, J\. D\. Rudie, S\. Sunaert, M\. Thioux, J\. M\. Tyszka, L\. Q\. Uddin, J\. S\. Verhoeven, N\. Wenderoth, J\. L\. Wiggins, S\. H\. Mostofsky, and M\. P\. MilhamThe autism brain imaging data exchange: towards a large\-scale evaluation of the intrinsic brain architecture in autism\.Molecular Psychiatry19\(6\),pp\. 659–667\.External Links:[Document](https://dx.doi.org/10.1038/mp.2013.78)Cited by:[item ABIDE\-I \( = n 846 , = q 57 , = m 101 , = k 2 \)](https://arxiv.org/html/2609.30487#A7.I4.ix4.p1.1)\.
- Estebanet al\.\(2019\)O\. Esteban, C\. J\. Markiewicz, R\. W\. Blair, C\. A\. Moodie, A\. I\. Isik, A\. Erramuzpe, J\. D\. Kent, M\. Goncalves, E\. DuPre, M\. Snyder, H\. Oya, S\. S\. Ghosh, J\. Wright, J\. Durnez, R\. A\. Poldrack, and K\. J\. GorgolewskifMRIPrep: a robust preprocessing pipeline for functional MRI\.Nature Methods16\(1\),pp\. 111–116\.External Links:[Document](https://dx.doi.org/10.1038/s41592-018-0235-4)Cited by:[item CAT\-D \( = n 118 , = q 49 , = m 100 , = k 2 \) \[ , \]](https://arxiv.org/html/2609.30487#A7.I4.ix6.p2.1)\.
- Fan and Pall \(1957\)K\. Fan and G\. PallImbedding conditions for Hermitian and normal matrices\.Canadian Journal of Mathematics9,pp\. 298 – 304\.External Links:[Document](https://dx.doi.org/10.4153/CJM-1957-036-1)Cited by:[§D\.1](https://arxiv.org/html/2609.30487#A4.SS1.p2.5),[§D\.3](https://arxiv.org/html/2609.30487#A4.SS3.p2.1)\.
- Fineet al\.\(1998\)S\. Fine, Y\. Singer, and N\. TishbyThe hierarchical hidden markov model: analysis and applications\.Machine Learning32\(1\),pp\. 41 – 62\.External Links:[Document](https://dx.doi.org/10.1023/A%3A1007469218079)Cited by:[item I \( = q 60 , = k 2 \)](https://arxiv.org/html/2609.30487#A7.I3.ix9.p1.1)\.
- Fletcheret al\.\(2004\)P\.T\. Fletcher, C\. Lu, S\.M\. Pizer, and S\. JoshiPrincipal geodesic analysis for the study of nonlinear statistics of shape\.IEEE Transactions on Medical Imaging23\(8\),pp\. 995–1005\.External Links:[Document](https://dx.doi.org/10.1109/TMI.2004.831793)Cited by:[item PGA](https://arxiv.org/html/2609.30487#A7.I1.ix4.p1.1),[§5\.1](https://arxiv.org/html/2609.30487#S5.SS1.p1.1)\.
- Glasseret al\.\(2016\)M\. F\. Glasser, T\. S\. Coalson, E\. C\. Robinson, C\. D\. Hacker, J\. Harwell, E\. Yacoub, K\. Ugurbil, J\. Andersson, C\. F\. Beckmann, M\. Jenkinson, S\. M\. Smith, and D\. C\. Van EssenA multi\-modal parcellation of human cerebral cortex\.Nature536\(7615\),pp\. 171–178\.External Links:[Document](https://dx.doi.org/10.1038/nature18933)Cited by:[item TCP \( = n 241 , = q 80 , = m 100 , = k 2 \) \[ , \]](https://arxiv.org/html/2609.30487#A7.I4.ix5.p1.1)\.
- Glover \(1999\)G\. H\. GloverDeconvolution of impulse response in event\-related bold fmri1\.NeuroImage9\(4\),pp\. 416–429\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1006/nimg.1998.0419)Cited by:[item G \( = q 30 , = k 2 \)](https://arxiv.org/html/2609.30487#A7.I3.ix7.p1.1)\.
- Gorgolewskiet al\.\(2017\)K\. J\. Gorgolewski, J\. Durnez, and R\. A\. PoldrackPreprocessed Consortium for Neuropsychiatric Phenomics dataset\.F1000Research6,pp\. 1262\.External Links:[Document](https://dx.doi.org/10.12688/f1000research.11964.2)Cited by:[item CNP \( = n 257 , = q 57 , = m 100 , = k 4 \) \[ , \]](https://arxiv.org/html/2609.30487#A7.I4.ix1.p1.1)\.
- Guoet al\.\(2013\)K\. Guo, P\. Ishwar, and J\. KonradAction recognition from video using feature covariance matrices\.IEEE Transactions on Image Processing22\(6\),pp\. 2479–2494\.External Links:[Document](https://dx.doi.org/10.1109/TIP.2013.2252622)Cited by:[§1](https://arxiv.org/html/2609.30487#S1.p1.1)\.
- Hall and Horowitz \(2007\)P\. Hall and J\. L\. HorowitzMethodology and convergence rates for functional linear regression\.Annals of Statistics35\(1\),pp\. 70 – 91\.External Links:[Document](https://dx.doi.org/10.1214/009053606000000957)Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px1.p2.1)\.
- Heinrichset al\.\(2023\)F\. Heinrichs, M\. Heim, and C\. WeberFunctional neural networks: shift invariant models for functional data with applications to EEG classification\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\. 12866–12881\.Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px1.p3.1)\.
- Heinsfeldet al\.\(2018\)A\. S\. Heinsfeld, A\. R\. Franco, R\. C\. Craddock, A\. Buchweitz, and F\. MeneguzziIdentification of autism spectrum disorder using deep learning and the ABIDE dataset\.NeuroImage: Clinical17,pp\. 16–23\.External Links:[Document](https://dx.doi.org/10.1016/j.nicl.2017.08.017)Cited by:[§5\.1](https://arxiv.org/html/2609.30487#S5.SS1.SSS0.Px2.p2.1)\.
- Huang and Gool \(2017\)Z\. Huang and L\. V\. GoolA Riemannian network for SPD matrix learning\.InProceedings of the Thirty\-First AAAI Conference on Artificial Intelligence,pp\. 2036–2042\.Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.30487#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.30487#S2.SS1.p1.1),[§2\.3](https://arxiv.org/html/2609.30487#S2.SS3.p3.1),[§5\.1](https://arxiv.org/html/2609.30487#S5.SS1.p1.1)\.
- Huanget al\.\(2015\)Z\. Huang, R\. Wang, S\. Shan, X\. Li, and X\. ChenLog\-euclidean metric learning on symmetric positive definite manifold with application to image set classification\.InProceedings of the 32nd International Conference on International Conference on Machine Learning \- Volume 37,pp\. 720–729\.Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px2.p1.1)\.
- Janget al\.\(2023\)C\. Jang, Y\. Lee, Y\. Noh, and F\. C\. ParkGeometrically regularized autoencoders for non\-euclidean data\.InThe Eleventh International Conference on Learning Representations,Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px2.p1.1)\.
- Kimet al\.\(2021\)B\. Kim, J\. C\. Ye, and J\. KimLearning dynamic graph representation of brain connectome with spatio\-temporal attention\.InAdvances in Neural Information Processing Systems,M\. Ranzato, A\. Beygelzimer, Y\. Dauphin, P\.S\. Liang, and J\. W\. Vaughan \(Eds\.\),Vol\.34,pp\. 4314–4327\.Cited by:[§1](https://arxiv.org/html/2609.30487#S1.p2.1),[§5\.2](https://arxiv.org/html/2609.30487#S5.SS2.p1.1)\.
- Lin \(2019\)Z\. LinRiemannian geometry of symmetric positive definite matrices via Cholesky decomposition\.SIAM Journal on Matrix Analysis and Applications40\(4\),pp\. 1353–1370\.External Links:[Document](https://dx.doi.org/10.1137/18M1221084)Cited by:[§B\.1](https://arxiv.org/html/2609.30487#A2.SS1.SSS0.Px2.p1.3),[§B\.1](https://arxiv.org/html/2609.30487#A2.SS1.p1.1),[§2](https://arxiv.org/html/2609.30487#S2.p1.1)\.
- Liuet al\.\(2019\)H\. Liu, J\. Li, Y\. Wu, and R\. JiLearning neural bag\-of\-matrix\-summarization with riemannian network\.InProceedings of the Thirty\-Third AAAI Conference on Artificial Intelligence and Thirty\-First Innovative Applications of Artificial Intelligence Conference and Ninth AAAI Symposium on Educational Advances in Artificial Intelligence,External Links:[Document](https://dx.doi.org/10.1609/aaai.v33i01.33018746)Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px2.p1.1)\.
- Lópezet al\.\(2021\)F\. López, B\. Pozzetti, S\. Trettel, M\. Strube, and A\. WienhardVector\-valued distance and gyrocalculus on the space of symmetric positive definite matrices\.InProceedings of the 35th International Conference on Neural Information Processing Systems,NIPS ’21\.Cited by:[§2\.2](https://arxiv.org/html/2609.30487#S2.SS2.p2.1)\.
- Lurieet al\.\(2020\)D\. J\. Lurie, D\. Kessler, D\. S\. Bassett, R\. F\. Betzel, M\. Breakspear, S\. Keilholz, A\. Kucyi, R\. Liégeois, M\. A\. Lindquist, A\. R\. McIntosh, R\. A\. Poldrack, J\. M\. Shine, W\. H\. Thompson, N\. Z\. Bielczyk, L\. Douw, D\. Kraft, R\. L\. Miller, M\. Muthuraman, L\. Pasquini, A\. Razi, D\. Vidaurre, H\. Xie, and V\. D\. CalhounQuestions and controversies in the study of time\-varying functional connectivity in resting fMRI\.Network Neuroscience4\(1\),pp\. 30 – 69\.External Links:[Document](https://dx.doi.org/10.1162/netn%5Fa%5F00116)Cited by:[§1](https://arxiv.org/html/2609.30487#S1.p1.1)\.
- Nguyenet al\.\(2024\)X\. S\. Nguyen, Y\. Yang, and A\. HistaceMatrix manifold neural networks\+\+\.InInternational Conference on Learning Representations,B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(Eds\.\),Vol\.2024,pp\. 40448–40478\.Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px2.p1.1)\.
- Patileaet al\.\(2016\)V\. Patilea, C\. Sánchez\-Sellero, and M\. SaumardTesting the predictor effect on a functional response\.Journal of the American Statistical Association111\(516\),pp\. 1684–1695\.External Links:[Document](https://dx.doi.org/10.1080/01621459.2015.1110031)Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px1.p2.1)\.
- Pennecet al\.\(2006\)X\. Pennec, P\. Fillard, and N\. AyacheA Riemannian framework for tensor computing\.International Journal of Computer Vision66\(1\),pp\. 41 – 66\.External Links:[Document](https://dx.doi.org/10.1007/s11263-005-3222-z)Cited by:[item H \( = q 30 , = k 2 \)](https://arxiv.org/html/2609.30487#A7.I3.ix8.p1.1),[§2](https://arxiv.org/html/2609.30487#S2.p1.1)\.
- Qiuet al\.\(2024\)B\. Qiu, Q\. Wang, X\. Li, W\. Li, W\. Shao, and M\. WangAdaptive spatial\-temporal neural network for ADHD identification using functional fMRI\.Frontiers in Neuroscience18,pp\. 1394234\.External Links:[Document](https://dx.doi.org/10.3389/fnins.2024.1394234)Cited by:[§5\.1](https://arxiv.org/html/2609.30487#S5.SS1.SSS0.Px2.p2.1)\.
- Rossiet al\.\(2002\)F\. Rossi, B\. Conan\-Guez, and F\. FleuretFunctional data analysis with multi layer perceptrons\.In2002 International Joint Conference on Neural Networks \(IJCNN\),Vol\.3,pp\. 2843–2848 vol\.3\.External Links:[Document](https://dx.doi.org/10.1109/IJCNN.2002.1007599)Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px1.p3.1)\.
- Sadeghiet al\.\(2022\)N\. Sadeghi, P\. Q\. Fors, L\. Eisner, J\. Taigman, K\. Qi, L\. S\. Gorham, C\. C\. Camp, G\. O’Callaghan, D\. Rodriguez, J\. McGuire, E\. M\. Garth, C\. Engel, M\. Davis, K\. E\. Towbin, A\. Stringaris, and D\. M\. NielsonMood and behaviors of adolescents with depression in a longitudinal study before and during the COVID\-19 pandemic\.Journal of the American Academy of Child & Adolescent Psychiatry61\(11\),pp\. 1341–1350\.Note:OpenNeurods004627External Links:[Document](https://dx.doi.org/10.1016/j.jaac.2022.04.004)Cited by:[item CAT\-D \( = n 118 , = q 49 , = m 100 , = k 2 \) \[ , \]](https://arxiv.org/html/2609.30487#A7.I4.ix6),[item CAT\-D \( = n 118 , = q 49 , = m 100 , = k 2 \) \[ , \]](https://arxiv.org/html/2609.30487#A7.I4.ix6.1)\.
- Salimi\-Khorshidiet al\.\(2014\)G\. Salimi\-Khorshidi, G\. Douaud, C\. F\. Beckmann, M\. F\. Glasser, L\. Griffanti, and S\. M\. SmithAutomatic denoising of functional MRI data: combining independent component analysis and hierarchical fusion of classifiers\.NeuroImage90,pp\. 449–468\.External Links:[Document](https://dx.doi.org/10.1016/j.neuroimage.2013.11.046)Cited by:[item TCP \( = n 241 , = q 80 , = m 100 , = k 2 \) \[ , \]](https://arxiv.org/html/2609.30487#A7.I4.ix5.p1.1)\.
- Saloet al\.\(2023\)T\. Salo, T\. Yarkoni, T\. E\. Nichols, J\. Poline, M\. Bilgel, K\. L\. Bottenhorn, S\. B\. Eickhoff, D\. Jarecka, J\. D\. Kent, A\. Kimbler, D\. M\. Nielson, K\. M\. Oudyk, J\. A\. Peraza, A\. Pérez, P\. C\. Reeders, J\. A\. Yanes, and A\. R\. LairdNiMARE: Neuroimaging Meta\-Analysis Research Environment\.Aperture Neuro3,pp\. 1–32\.External Links:[Document](https://dx.doi.org/10.52294/001c.87681)Cited by:[item J \( = q 50 , = k 2 \) :](https://arxiv.org/html/2609.30487#A7.I3.ix10.p1.1)\.
- Sanchez\-Vilaet al\.\(2006\)X\. Sanchez\-Vila, A\. Guadagnini, and J\. CarreraRepresentative hydraulic conductivities in saturated groundwater flow\.Reviews of Geophysics44\(3\),pp\.\.External Links:[Document](https://dx.doi.org/10.1029/2005RG000169)Cited by:[§1](https://arxiv.org/html/2609.30487#S1.p1.1)\.
- Saponaroet al\.\(2022\)S\. Saponaro, A\. Giuliano, R\. Bellotti, A\. Lombardi, S\. Tangaro, P\. Oliva, S\. Calderoni, and A\. ReticoMulti\-site harmonization of MRI data uncovers machine\-learning discrimination capability in barely separable populations: an example from the ABIDE dataset\.NeuroImage: Clinical35,pp\. 103082\.External Links:[Document](https://dx.doi.org/10.1016/j.nicl.2022.103082)Cited by:[§5\.1](https://arxiv.org/html/2609.30487#S5.SS1.SSS0.Px2.p2.1)\.
- Singhet al\.\(2025\)S\. Singh, S\. Coyle, and M\. ZhangShape\-informed clustering of multi\-dimensional functional data via deep functional autoencoders\.InAdvances in Neural Information Processing Systems,Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px1.p3.1),[§1](https://arxiv.org/html/2609.30487#S1.p3.1),[§2](https://arxiv.org/html/2609.30487#S2.p3.1),[§3](https://arxiv.org/html/2609.30487#S3.p1.1),[§4](https://arxiv.org/html/2609.30487#S4.p2.1)\.
- The ADHD\-200 Consortium \(2012\)The ADHD\-200 ConsortiumThe adhd\-200 consortium: a model to advance the translational potential of neuroimaging in clinical neuroscience\.Frontiers in Systems Neuroscience6,pp\. 62\.External Links:[Document](https://dx.doi.org/10.3389/fnsys.2012.00062)Cited by:[item ADHD\-200 \( = n 138 , = q 57 , = m 100 , = k 2 \) \[ , \]](https://arxiv.org/html/2609.30487#A7.I4.ix3),[item ADHD\-200 \( = n 138 , = q 57 , = m 100 , = k 2 \) \[ , \]](https://arxiv.org/html/2609.30487#A7.I4.ix3.1)\.
- Thindet al\.\(2022\)B\. Thind, K\. Multani, and J\. CaoDeep learning with functional inputs\.Journal of Computational and Graphical Statistics0\(0\),pp\. 1–10\.External Links:[Document](https://dx.doi.org/10.1080/10618600.2022.2097914)Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px1.p3.1)\.
- Tuzelet al\.\(2008\)O\. Tuzel, F\. Porikli, and P\. MeerPedestrian detection via classification on riemannian manifolds\.IEEE Transactions on Pattern Analysis and Machine Intelligence30\(10\),pp\. 1713–1727\.External Links:[Document](https://dx.doi.org/10.1109/TPAMI.2008.75)Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px2.p1.1)\.
- Vidaurreet al\.\(2017\)D\. Vidaurre, S\. M\. Smith, and M\. W\. WoolrichBrain network dynamics are hierarchically organized in time\.Proceedings of the National Academy of Sciences114\(48\),pp\. 12827–12832\.External Links:[Document](https://dx.doi.org/10.1073/pnas.1705120114)Cited by:[item C \( = q 120 , = k 3 \)](https://arxiv.org/html/2609.30487#A7.I3.ix3.p1.1)\.
- Wanget al\.\(2016\)J\. Wang, J\. Chiou, and H\. MüllerFunctional data analysis\.Annual Review of Statistics and Its Application3\(Volume 3, 2016\),pp\. 257–295\.External Links:[Document](https://dx.doi.org/10.1146/annurev-statistics-041715-033624)Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px1.p2.1)\.
- Wanget al\.\(2025\)R\. Wang, S\. Jin, Z\. Chen, X\. Luo, and X\. WuLearning to normalize on the spd manifold under bures\-wasserstein geometry\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 8289–8298\.Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px2.p1.1),[§5\.2](https://arxiv.org/html/2609.30487#S5.SS2.p1.1)\.
- Yarkoniet al\.\(2011\)T\. Yarkoni, R\. A\. Poldrack, T\. E\. Nichols, D\. C\. Van Essen, and T\. D\. WagerLarge\-scale automated synthesis of human functional neuroimaging data\.Nature Methods8\(8\),pp\. 665–670\.External Links:[Document](https://dx.doi.org/10.1038/nmeth.1635)Cited by:[item J \( = q 50 , = k 2 \) :](https://arxiv.org/html/2609.30487#A7.I3.ix10.p1.1)\.
- Yeonet al\.\(2023\)H\. Yeon, X\. Dai, and D\. J\. NordmanBootstrap inference in functional linear regression models with scalar response\.Bernoulli29\(4\),pp\. 2599 – 2626\.External Links:[Document](https://dx.doi.org/10.3150/22-BEJ1554)Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px1.p2.1)\.
- Yuanet al\.\(2013\)Y\. Yuan, H\. Zhu, M\. Styner, J\. H\. Gilmore, and J\.S\. MarronVarying coefficient model for modeling diffusion tensors along white matter tracts\.Annals of Applied Statistics7\(1\),pp\. 102 – 125\.External Links:[Document](https://dx.doi.org/10.1214/12-AOAS574)Cited by:[§1](https://arxiv.org/html/2609.30487#S1.p1.1)\.
- Zenget al\.\(2018\)L\. Zeng, H\. Wang, P\. Hu, B\. Yang, W\. Pu, H\. Shen, X\. Chen, Z\. Liu, H\. Yin, Q\. Tan, K\. Wang, and D\. HuMulti\-site diagnostic classification of schizophrenia using discriminant deep learning with functional connectivity MRI\.EBioMedicine30,pp\. 74–85\.External Links:[Document](https://dx.doi.org/10.1016/j.ebiom.2018.03.017)Cited by:[§5\.1](https://arxiv.org/html/2609.30487#S5.SS1.SSS0.Px2.p2.1)\.
- Zhang and Parnell \(2023\)M\. Zhang and A\. ParnellReview of clustering methods for functional data\.ACM Transactions on Knowledge Discovery from Data17\(7\),pp\. 34\.External Links:[Document](https://dx.doi.org/10.1145/3581789)Cited by:[Appendix A](https://arxiv.org/html/2609.30487#A1.SS0.SSS0.Px1.p2.1)\.
- Zhanget al\.\(2018\)T\. Zhang, W\. Zheng, Z\. Cui, and C\. LiDeep manifold\-to\-manifold transforming network\.In2018 25th IEEE International Conference on Image Processing \(ICIP\),Vol\.,pp\. 4098–4102\.External Links:[Document](https://dx.doi.org/10.1109/ICIP.2018.8451626)Cited by:[§2\.3](https://arxiv.org/html/2609.30487#S2.SS3.p3.2)\.

## Appendix ARelated work

#### Functional data analysis

Functional data analysis studies data whose basic observational units are functions \(e\.g\., curves over time or space\) and, more broadly, other complex objects such as images and shapes\. In functional data analysis, each subject in a sample is represented by one or multiple functions observed over a continuum \(such as time, wavelength, or spatial location\)\. Let\{𝒚i:i=1,…,n\}\\\{\\boldsymbol\{y\}\_\{i\}:i=1,\\ldots,n\\\}be a set ofnnindependentpp\-dimensional sample functions:𝒚i​\(t\)=\(yi1​\(t\),…,yip​\(t\)\)⊤∈ℝp\\boldsymbol\{y\}\_\{i\}\(t\)=\(y\_\{i\}^\{1\}\(t\),\\ldots,y\_\{i\}^\{p\}\(t\)\)^\{\\top\}\\in\\mathbb\{R\}^\{p\}for anyt∈\[t0,t1\]t\\in\[t\_\{0\},t\_\{1\}\]; the superscript⊤\\topis the transpose operator\. Each sample function𝒚i\\boldsymbol\{y\}\_\{i\}is a random realization of app\-dimensional random functionY→=\(Y1,…,Yp\)⊤\\vec\{Y\}=\(Y^\{1\},\\ldots,Y^\{p\}\)^\{\\top\}\. Ford=1,…,pd=1,\\ldots,p, thennone\-dimensional sample functions\{y1d,…,ynd\}\\\{y^\{d\}\_\{1\},\\ldots,y^\{d\}\_\{n\}\\\}are independent realizations of the component random functionYd=Yd​\(t\)=Yd​\(t,ω\)Y^\{d\}=Y^\{d\}\(t\)=Y^\{d\}\(t,\\omega\), defined on a probability space\(Ω,ℱ,Pr\)\(\\Omega,\\mathcal\{F\},\\Pr\)and taking values inℋ⁡\(\[t0,t1\],ℝ\)\\mathcal\{H\}\(\[t\_\{0\},t\_\{1\}\],\\mathbb\{R\}\)\. Here,ℋ⁡\(\[t0,t1\],ℝ\)\\mathcal\{H\}\(\[t\_\{0\},t\_\{1\}\],\\mathbb\{R\}\)is the separable Hilbert space of all square\-integrable measurable functions that are defined on\[t0,t1\]\[t\_\{0\},t\_\{1\}\]and taking values inℝ\\mathbb\{R\}\. That is,YdY^\{d\}is a measurable map fromΩ\\Omegatoℋ⁡\(\[t0,t1\],ℝ\)\\mathcal\{H\}\(\[t\_\{0\},t\_\{1\}\],\\mathbb\{R\}\)\.

Functional regression is commonly classified by whether the response and the covariates are functional or finite\-dimensional \(vector/scalar\)\. This yields three main settings: \(i\) functional responses with functional covariates[Dette and Tang \[2024\]](https://arxiv.org/html/2609.30487#bib.bib25),[Cai et al\. \[2022\]](https://arxiv.org/html/2609.30487#bib.bib28), \(ii\) vector \(including scalar\) responses with functional covariates[Hall and Horowitz \[2007\]](https://arxiv.org/html/2609.30487#bib.bib24),[Yeon et al\. \[2023\]](https://arxiv.org/html/2609.30487#bib.bib27),[Balasubramanian et al\. \[2025\]](https://arxiv.org/html/2609.30487#bib.bib26), and \(iii\) functional responses with vector covariates[Patilea et al\. \[2016\]](https://arxiv.org/html/2609.30487#bib.bib29)\. Among these, the most extensively studied case is \(ii\), particularly scalar\-on\-function regression, where the response is a scalar and the predictor is a function\. For an overview of statistical tools in functional data analysis, see[Wang et al\. \[2016\]](https://arxiv.org/html/2609.30487#bib.bib30)\. Beyond regression, clustering is a core task in functional data analysis\.[Zhang and Parnell \[2023\]](https://arxiv.org/html/2609.30487#bib.bib31)provided a comprehensive review of clustering methods for functional data and observed that approaches for clustering multi\-dimensional functional data generally follow one of two paradigms: \(1\) first applying multivariate functional principal component analysis[Chiou et al\. \[2014\]](https://arxiv.org/html/2609.30487#bib.bib32)and then clustering the resulting score vectors using standard multivariate clustering methods; or \(2\) defining an appropriate \(dis\)similarity measure for multi\-dimensional functional data and clustering directly using distance\- or similarity\-based algorithms\.

More recently, functional neural networks have been proposed to capture nonlinear relationships that may be difficult to represent with classical functional data analysis models[Rossi et al\. \[2002\]](https://arxiv.org/html/2609.30487#bib.bib33),[Thind et al\. \[2022\]](https://arxiv.org/html/2609.30487#bib.bib34),[Heinrichs et al\. \[2023\]](https://arxiv.org/html/2609.30487#bib.bib35)\.[Singh et al\. \[2025\]](https://arxiv.org/html/2609.30487#bib.bib1)provided the first proof that functional neural network has the universal approximation property\. A key distinction of these architectures is the inclusion of functional layers, in which the learnable weights are themselves functions rather than scalar parameters, enabling the architectures operate directly on functional inputs\. Although one can treat the sequence of discrete evaluations of eachyid​\(t\)y^\{d\}\_\{i\}\(t\)over a grid as a time series, enabling the use of recurrent neural networks or Transformer\-based models, such approaches do not naturally encode shape \(morphological\) characteristics or derivative\-based \(first\-order\) information of the underlying functions\. By contrast, functional neural networks offer a principled framework for learning representations that respect the functional nature of the data rather than relying solely on discrete measurements\.

#### Riemannian neural networks

Deep neural networks for SPD matrices must respect the fact that the data live on a Riemannian manifold\. Over the past decade, two main modeling paradigms have emerged: tangent\-space learning and intrinsic manifold deep networks, with more recent work extending these ideas to attention, normalization, and alternative geometries\. In the tangent\-space learning paradigm, we map an SPD matrix to a vector space using a Riemannian logarithm map, then apply standard neural architectures on the mapped data[Arsigny et al\. \[2006\]](https://arxiv.org/html/2609.30487#bib.bib36),[Tuzel et al\. \[2008\]](https://arxiv.org/html/2609.30487#bib.bib37),[Huang et al\. \[2015\]](https://arxiv.org/html/2609.30487#bib.bib38)\. This approach is computationally attractive because it avoids repeated eigen/SVD constraints inside the network while keeping a principled geometric link to the manifold\. In the intrinsic deep learning paradigm, every operation preserves positive definiteness\. The canonical example is SPDNet[Huang and Gool \[2017\]](https://arxiv.org/html/2609.30487#bib.bib2), which introduced a stack of SPD\-preserving layers: bilinear mapping \(BiMap\) for dimension reduction on SPD, eigenvalue rectification \(ReEig\) as a nonlinearity, and log\-eigenvalue \(LogEig\) to move to a Euclidean space only near the output for standard classifiers\. This architecture family has been extended in multiple directions: regularization on manifolds[Jang et al\. \[2023\]](https://arxiv.org/html/2609.30487#bib.bib43), Riemannian batch normalization on the SPD manifold[Brooks et al\. \[2019\]](https://arxiv.org/html/2609.30487#bib.bib22),[Wang et al\. \[2025\]](https://arxiv.org/html/2609.30487#bib.bib39), a multi\-channel mechanism to capture multi\-scale features[Chen et al\. \[2023\]](https://arxiv.org/html/2609.30487#bib.bib40), and replacing LogEig with a learnable SPD encoder[Liu et al\. \[2019\]](https://arxiv.org/html/2609.30487#bib.bib41)\. Another emerging direction is to broaden the layer toolbox \(e\.g\., convolutional layers\) to make SPD networks more closely aligned with standard deep learning building blocks, while still ensuring valid manifold backpropagation[Nguyen et al\. \[2024\]](https://arxiv.org/html/2609.30487#bib.bib42)\.

To the best of our knowledge, there is currently no deep neural network architecture specifically designed to model SPD\-valued trajectories\{𝑿i∈ℋg\(\[t0,t1\];𝒮\+m\)\}i=1n\\\{\{\\bm\{\\mathsfit\{X\}\}\}\_\{i\}\\in\\mathcal\{H\}\_\{g\}\(\[t\_\{0\},t\_\{1\}\];\\mathcal\{S\}^\{m\}\_\{\+\}\)\\\}\_\{i=1\}^\{n\}, where each observation𝑿i\{\\bm\{\\mathsfit\{X\}\}\}\_\{i\}is a function of time, and its value𝑿i​\(t\)\{\\bm\{\\mathsfit\{X\}\}\}\_\{i\}\(t\)is an SPD matrix\. Such SPD\-valued functional data arise naturally when second\-order statistics evolve continuously over time, and they are increasingly encountered in modern data\-rich scientific domains, most notably neuroscience\. A representative example is magnetoencephalography \(MEG\) or electroencephalography \(EEG\) analysis\. Let𝒙i​\(t\)∈ℝm\\boldsymbol\{x\}\_\{i\}\(t\)\\in\\mathbb\{R\}^\{m\}denote the multichannel sensor signal of subjectiiat timett\. Instead of analyzing raw signals, it is common to compute time\-resolved covariance or cross\-spectral matrices𝑿i​\(t\)\{\\bm\{\\mathsfit\{X\}\}\}\_\{i\}\(t\)over short, possibly overlapping, temporal windows\. As the window slides continuously, this procedure yields a smooth trajectory of SPD matrices over time, capturing the dynamic evolution of brain networks during cognitive tasks, motor imagery, or resting\-state activity[Bonet et al\. \[2023\]](https://arxiv.org/html/2609.30487#bib.bib44)\. In this work, we developMatFAEto bridge this methodological gap by unifying recent advances in Riemannian deep learning and functional data analysis for end\-to\-end modeling of SPD\-valued trajectories\.

## Appendix BProperties of the two Riemannian manifolds

### B\.1The Log\-Cholesky metric

Letℒm\\mathcal\{L\}^\{m\}denote the set ofm×mm\\times mlower triangular matrices, andℒ\+m⊂ℒm\\mathcal\{L\}^\{m\}\_\{\+\}\\subset\\mathcal\{L\}^\{m\}the subset of lower triangular matrices with positive diagonal elements\. Bothℒ\+m\\mathcal\{L\}^\{m\}\_\{\+\}and𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}are smooth manifolds and therefore have differential geometry properties\. The tangent space𝒯𝐋​ℒ\+m\\mathcal\{T\}\_\{\\mathbf\{L\}\}\\mathcal\{L\}^\{m\}\_\{\+\}ofℒ\+m\\mathcal\{L\}^\{m\}\_\{\+\}at any point𝐋∈ℒ\+m\\mathbf\{L\}\\in\\mathcal\{L\}^\{m\}\_\{\+\}isℒm\\mathcal\{L\}^\{m\}\. For any𝐗∈𝒮\+m\\mathbf\{X\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}, there is a unique lower triangular matrix with positive diagonal elements𝐋=𝔏⁡\(𝐗\)∈ℒ\+m\\mathbf\{L\}=\\mathfrak\{L\}\(\\mathbf\{X\}\)\\in\\mathcal\{L\}^\{m\}\_\{\+\}that satisfies𝐗=𝐋𝐋⊤\\mathbf\{X\}=\\mathbf\{L\}\\mathbf\{L\}^\{\\top\}\. Furthermore, the map𝔏:𝒮\+m→ℒ\+m\\mathfrak\{L\}:\\mathcal\{S\}^\{m\}\_\{\+\}\\rightarrow\\mathcal\{L\}^\{m\}\_\{\+\}is a diffeomorphism\[[Lin, 2019](https://arxiv.org/html/2609.30487#bib.bib5), Proposition 2\]\. The differentialD𝐗​𝔏:𝒯𝐗​𝒮\+m↦𝒯𝔏⁡\(𝐗\)​ℒ\+mD\_\{\\mathbf\{X\}\}\\mathfrak\{L\}:\\mathcal\{T\}\_\{\\mathbf\{X\}\}\\mathcal\{S\}^\{m\}\_\{\+\}\\mapsto\\mathcal\{T\}\_\{\\mathfrak\{L\}\(\\mathbf\{X\}\)\}\\mathcal\{L\}^\{m\}\_\{\+\}of𝔏\\mathfrak\{L\}at𝐗\\mathbf\{X\}is given by

\(D𝐗​𝔏\)​\(𝐔\)=𝔏⁡\(𝐗\)​\(𝔏​\(𝐗\)−1​𝐔​𝔏​\(𝐗\)−⁣⊤\)12,𝐔∈𝒯𝐗​𝒮\+m,\(D\_\{\\mathbf\{X\}\}\\mathfrak\{L\}\)\(\\mathbf\{U\}\)=\\mathfrak\{L\}\(\\mathbf\{X\}\)\(\\mathfrak\{L\}\(\\mathbf\{X\}\)^\{\-1\}\\mathbf\{U\}\\mathfrak\{L\}\(\\mathbf\{X\}\)^\{\-\\top\}\)\_\{\\frac\{1\}\{2\}\},~~\\mathbf\{U\}\\in\\mathcal\{T\}\_\{\\mathbf\{X\}\}\\mathcal\{S\}^\{m\}\_\{\+\},where\(𝐕\)12=⌊𝐕⌋\+12​diag​\(𝐕\)\(\\mathbf\{V\}\)\_\{\\frac\{1\}\{2\}\}=\\lfloor\\mathbf\{V\}\\rfloor\+\\frac\{1\}\{2\}\\text\{diag\}\(\\mathbf\{V\}\)is the lower triangular part of𝐕\\mathbf\{V\}with the diagonal elements halved, anddiag​\(𝐕\)\\text\{diag\}\(\\mathbf\{V\}\)is the matrix of the main diagonal elements of𝐕\\mathbf\{V\}\. Let𝔏−1:ℒ\+m↦𝒮\+m\\mathfrak\{L\}^\{\-1\}:\\mathcal\{L\}^\{m\}\_\{\+\}\\mapsto\\mathcal\{S\}^\{m\}\_\{\+\}denote the inverse of the map𝔏\\mathfrak\{L\}:𝔏−1​\(𝐋\)=𝐋𝐋⊤\\mathfrak\{L\}^\{\-1\}\(\\mathbf\{L\}\)=\\mathbf\{L\}\\mathbf\{L\}^\{\\top\}\. The differentialD𝐋​𝔏−1:𝒯𝐋​ℒ\+m↦𝒯𝐋𝐋⊤​𝒮\+mD\_\{\\mathbf\{L\}\}\\mathfrak\{L\}^\{\-1\}:\\mathcal\{T\}\_\{\\mathbf\{L\}\}\\mathcal\{L\}^\{m\}\_\{\+\}\\mapsto\\mathcal\{T\}\_\{\\mathbf\{L\}\\mathbf\{L\}^\{\\top\}\}\\mathcal\{S\}^\{m\}\_\{\+\}of𝔏−1\\mathfrak\{L\}^\{\-1\}at𝐋\\mathbf\{L\}is given by

\(D𝐋​𝔏−1\)​\(𝐔\)=𝐋𝐔⊤\+𝐔𝐋⊤,𝐔∈𝒯𝐋​ℒ\+m\.\(D\_\{\\mathbf\{L\}\}\\mathfrak\{L\}^\{\-1\}\)\(\\mathbf\{U\}\)=\\mathbf\{L\}\\mathbf\{U\}^\{\\top\}\+\\mathbf\{U\}\\mathbf\{L\}^\{\\top\},~~\\mathbf\{U\}\\in\\mathcal\{T\}\_\{\\mathbf\{L\}\}\\mathcal\{L\}^\{m\}\_\{\+\}\.Because𝔏\\mathfrak\{L\}is a diffeomorphism, the Riemannian metric for𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}can be induced \(via pullback\) from the Riemannian metric forℒ\+m\\mathcal\{L\}^\{m\}\_\{\+\}\.

For the manifoldℒ\+m\\mathcal\{L\}^\{m\}\_\{\+\}, the Riemannian metric at a point𝐋∈ℒ\+m\\mathbf\{L\}\\in\\mathcal\{L\}^\{m\}\_\{\+\}is

⟨𝐔,𝐕⟩𝐋ℒ\+m=⟨⌊𝐔⌋,⌊𝐕⌋⟩F\+⟨diag​\(𝐋\)−1​diag​\(𝐔\),diag​\(𝐋\)−1​diag​\(𝐕\)⟩F,𝐔,𝐕∈𝒯𝐋​ℒ\+m,\\langle\\mathbf\{U\},\\mathbf\{V\}\\rangle\_\{\\mathbf\{L\}\}^\{\\mathcal\{L\}^\{m\}\_\{\+\}\}=\\langle\\lfloor\\mathbf\{U\}\\rfloor,\\lfloor\\mathbf\{V\}\\rfloor\\rangle\_\{F\}\+\\langle\\text\{diag\}\(\\mathbf\{L\}\)^\{\-1\}\\text\{diag\}\(\\mathbf\{U\}\),\\text\{diag\}\(\\mathbf\{L\}\)^\{\-1\}\\text\{diag\}\(\\mathbf\{V\}\)\\rangle\_\{F\},~~\\mathbf\{U\},\\mathbf\{V\}\\in\\mathcal\{T\}\_\{\\mathbf\{L\}\}\\mathcal\{L\}^\{m\}\_\{\+\},where⟨⋅,⋅⟩F\\langle\\cdot,\\cdot\\rangle\_\{F\}is the Frobenius inner product, with induced norm∥⋅∥F\\\|\\cdot\\\|\_\{F\}\. The Riemannian exponential map at𝐋\\mathbf\{L\}is

Exp𝐋​\(𝐔\)=⌊𝐋⌋\+⌊𝐔⌋\+diag​\(𝐋\)​exp⁡\(diag​\(𝐔\)​diag​\(𝐋\)−1\),𝐔∈𝒯𝐋​ℒ\+m,\\text\{Exp\}\_\{\\mathbf\{L\}\}\(\\mathbf\{U\}\)=\\lfloor\\mathbf\{L\}\\rfloor\+\\lfloor\\mathbf\{U\}\\rfloor\+\\text\{diag\}\(\\mathbf\{L\}\)\\exp\(\\text\{diag\}\(\\mathbf\{U\}\)\\text\{diag\}\(\\mathbf\{L\}\)^\{\-1\}\),~~\\mathbf\{U\}\\in\\mathcal\{T\}\_\{\\mathbf\{L\}\}\\mathcal\{L\}^\{m\}\_\{\+\},and the Riemannian logarithmic map at𝐋\\mathbf\{L\}is

Log𝐋​\(𝐊\)=⌊𝐊⌋−⌊𝐋⌋\+diag​\(𝐋\)​log⁡\(diag​\(𝐋\)−1​diag​\(𝐊\)\),𝐊∈ℒ\+m\.\\text\{Log\}\_\{\\mathbf\{L\}\}\(\\mathbf\{K\}\)=\\lfloor\\mathbf\{K\}\\rfloor\-\\lfloor\\mathbf\{L\}\\rfloor\+\\text\{diag\}\(\\mathbf\{L\}\)\\log\(\\text\{diag\}\(\\mathbf\{L\}\)^\{\-1\}\\text\{diag\}\(\\mathbf\{K\}\)\),~~\\mathbf\{K\}\\in\\mathcal\{L\}^\{m\}\_\{\+\}\.The geodesic distance between two𝐋\\mathbf\{L\}and𝐊∈ℒ\+m\\mathbf\{K\}\\in\\mathcal\{L\}^\{m\}\_\{\+\}is

dℒ\+m​\(𝐋,𝐊\)=‖⌊𝐋⌋−⌊𝐊⌋‖F2\+‖log⁡\(diag​\(𝐋\)\)−log⁡\(diag​\(𝐊\)\)‖F2\.d\_\{\\mathcal\{L\}^\{m\}\_\{\+\}\}\(\\mathbf\{L\},\\mathbf\{K\}\)=\\sqrt\{\\\|\\lfloor\\mathbf\{L\}\\rfloor\-\\lfloor\\mathbf\{K\}\\rfloor\\\|\_\{F\}^\{2\}\+\\\|\\log\(\\text\{diag\}\(\\mathbf\{L\}\)\)\-\\log\(\\text\{diag\}\(\\mathbf\{K\}\)\)\\\|\_\{F\}^\{2\}\}\.
#### Riemannian geometry

Given the Riemannian metric⟨⋅,⋅⟩⋅ℒ\+m\\langle\\cdot,\\cdot\\rangle\_\{\\cdot\}^\{\\mathcal\{L\}^\{m\}\_\{\+\}\}forℒ\+m\\mathcal\{L\}^\{m\}\_\{\+\}, the manifold map𝔏:𝒮\+m↦ℒ\+m\\mathfrak\{L\}:\\mathcal\{S\}^\{m\}\_\{\+\}\\mapsto\\mathcal\{L\}^\{m\}\_\{\+\}induces the Riemannian metric for𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}, called the Log\-Cholesky \(LC\) metric:

⟨𝐔,𝐕⟩𝐗LC=⟨\(D𝐗​𝔏\)​\(𝐔\),\(D𝐗​𝔏\)​\(𝐕\)⟩𝔏⁡\(𝐗\)ℒ\+m,𝐔,𝐕∈𝒯𝐗​𝒮\+m\.\\langle\\mathbf\{U\},\\mathbf\{V\}\\rangle\_\{\\mathbf\{X\}\}^\{\\text\{LC\}\}=\\langle\(D\_\{\\mathbf\{X\}\}\\mathfrak\{L\}\)\(\\mathbf\{U\}\),\(D\_\{\\mathbf\{X\}\}\\mathfrak\{L\}\)\(\\mathbf\{V\}\)\\rangle\_\{\\mathfrak\{L\}\(\\mathbf\{X\}\)\}^\{\\mathcal\{L\}^\{m\}\_\{\+\}\},~~\\mathbf\{U\},\\mathbf\{V\}\\in\\mathcal\{T\}\_\{\\mathbf\{X\}\}\\mathcal\{S\}^\{m\}\_\{\+\}\.The Riemannian exponential map at𝐗∈𝒮\+m\\mathbf\{X\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}is

Exp𝐗LC​\(𝐕\)=\[Exp𝔏⁡\(𝐗\)​\(\(D𝐗​𝔏\)​\(𝐕\)\)\]​\[Exp𝔏⁡\(𝐗\)​\(\(D𝐗​𝔏\)​\(𝐕\)\)\]⊤,𝐕∈𝒯𝐗​𝒮\+m\.\\text\{Exp\}^\{\\text\{LC\}\}\_\{\\mathbf\{X\}\}\(\\mathbf\{V\}\)=\[\\text\{Exp\}\_\{\\mathfrak\{L\}\(\\mathbf\{X\}\)\}\(\(D\_\{\\mathbf\{X\}\}\\mathfrak\{L\}\)\(\\mathbf\{V\}\)\)\]\[\\text\{Exp\}\_\{\\mathfrak\{L\}\(\\mathbf\{X\}\)\}\(\(D\_\{\\mathbf\{X\}\}\\mathfrak\{L\}\)\(\\mathbf\{V\}\)\)\]^\{\\top\},~~\\mathbf\{V\}\\in\\mathcal\{T\}\_\{\\mathbf\{X\}\}\\mathcal\{S\}^\{m\}\_\{\+\}\.The Riemannian logarithmic map at𝐗\\mathbf\{X\}is

Log𝐗LC​\(𝐙\)=\(D𝔏⁡\(𝐗\)​𝔏−1\)​\(Log𝔏⁡\(𝐗\)​\(𝔏⁡\(𝐙\)\)\),𝐙∈𝒮\+m\.\\text\{Log\}^\{\\text\{LC\}\}\_\{\\mathbf\{X\}\}\(\\mathbf\{Z\}\)=\(D\_\{\\mathfrak\{L\}\(\\mathbf\{X\}\)\}\\mathfrak\{L\}^\{\-1\}\)\(\\text\{Log\}\_\{\\mathfrak\{L\}\(\\mathbf\{X\}\)\}\(\\mathfrak\{L\}\(\\mathbf\{Z\}\)\)\),~~\\mathbf\{Z\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}\.\(5\)When the reference point is the identity matrix𝐈\\mathbf\{I\}, the Riemannian logarithmic map of𝐙\\mathbf\{Z\}is

Log𝐈LC​\(𝐙\)=⌊𝔏⁡\(𝐙\)⌋\+⌊𝔏⁡\(𝐙\)⌋⊤\+2​log⁡\(diag​\(𝔏⁡\(𝐙\)\)\)\.\\text\{Log\}^\{\\text\{LC\}\}\_\{\\mathbf\{I\}\}\(\\mathbf\{Z\}\)=\\lfloor\\mathfrak\{L\}\(\\mathbf\{Z\}\)\\rfloor\+\\lfloor\\mathfrak\{L\}\(\\mathbf\{Z\}\)\\rfloor^\{\\top\}\+2\\log\(\\text\{diag\}\(\\mathfrak\{L\}\(\\mathbf\{Z\}\)\)\)\.\(6\)The map𝔏\\mathfrak\{L\}is an isometry between\(𝒮\+m,⟨⋅,⋅⟩⋅LC\)\(\\mathcal\{S\}^\{m\}\_\{\+\},\\langle\\cdot,\\cdot\\rangle\_\{\\cdot\}^\{\\text\{LC\}\}\)and\(ℒ\+m,⟨⋅,⋅⟩⋅ℒ\+m\)\(\\mathcal\{L\}^\{m\}\_\{\+\},\\langle\\cdot,\\cdot\\rangle\_\{\\cdot\}^\{\\mathcal\{L\}^\{m\}\_\{\+\}\}\), and therefore the geodesic distance between𝐗,𝐙∈𝒮\+m\\mathbf\{X\},\\mathbf\{Z\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}is simplydLC​\(𝐗,𝐙\)=dℒ\+m​\(𝔏⁡\(𝐗\),𝔏⁡\(𝐙\)\)d\_\{\\text\{LC\}\}\(\\mathbf\{X\},\\mathbf\{Z\}\)=d\_\{\\mathcal\{L\}^\{m\}\_\{\+\}\}\(\\mathfrak\{L\}\(\\mathbf\{X\}\),\\mathfrak\{L\}\(\\mathbf\{Z\}\)\)\.

#### Fréchet mean

Given a set of SPD matrices\{𝐗i∈𝒮\+m:i=1,…,n\}\\\{\\mathbf\{X\}\_\{i\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}:i=1,\\ldots,n\\\}, under mild conditions, the Fréchet mean𝔼LC\(𝐗1:n\)\\mathbb\{E\}\_\{\\text\{LC\}\}\(\\mathbf\{X\}\_\{1:n\}\)exits and is unique:

𝔼LC\(𝐗1:n\)=arg​min𝐗∈𝒮\+m∑i=1ndLC\(𝐗,𝐗i\)2=𝐋n𝐋n⊤,\\mathbb\{E\}\_\{\\text\{LC\}\}\(\\mathbf\{X\}\_\{1:n\}\)=\\argmin\_\{\\mathbf\{X\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}\}\\sum\_\{i=1\}^\{n\}d\_\{\\text\{LC\}\}\(\\mathbf\{X\},\\mathbf\{X\}\_\{i\}\)^\{2\}=\\mathbf\{L\}\_\{n\}\\mathbf\{L\}\_\{n\}^\{\\top\},where

𝐋n=1n​∑i=1n⌊𝔏⁡\(𝐗i\)⌋\+exp⁡\(1n​∑i=1nlog⁡\(diag​\(𝔏⁡\(𝐗i\)\)\)\)\.\\mathbf\{L\}\_\{n\}=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\lfloor\\mathfrak\{L\}\(\\mathbf\{X\}\_\{i\}\)\\rfloor\+\\exp\(\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\log\(\\text\{diag\}\(\\mathfrak\{L\}\(\\mathbf\{X\}\_\{i\}\)\)\)\)\.\(7\)The sectional curvature of\(𝒮\+m,⟨⋅,⋅⟩⋅LC\)\(\\mathcal\{S\}^\{m\}\_\{\+\},\\langle\\cdot,\\cdot\\rangle\_\{\\cdot\}^\{\\text\{LC\}\}\)is constantly zero\[[Lin, 2019](https://arxiv.org/html/2609.30487#bib.bib5), Proposition 8\], and therefore\(𝒮\+m,⟨⋅,⋅⟩⋅LC\)\(\\mathcal\{S\}^\{m\}\_\{\+\},\\langle\\cdot,\\cdot\\rangle\_\{\\cdot\}^\{\\text\{LC\}\}\)is a Cartan\-Hadamard manifold\.

### B\.2The Log\-Euclidean metric

Every real symmetric matrix𝐔∈Sym​\(m\)\\mathbf\{U\}\\in\\text\{Sym\}\(m\)admits an orthogonal eigen\-decomposition:𝐔=𝐐​diag​\(λ1,…,λm\)​𝐐⊤\\mathbf\{U\}=\\mathbf\{Q\}\\text\{diag\}\(\\lambda\_\{1\},\\ldots,\\lambda\_\{m\}\)\\mathbf\{Q\}^\{\\top\}, and therefore taking the matrix exponential can be done by simply applying its scalar version to eigenvalues:

exp⁡\(𝐔\)=𝐐​diag​\(exp⁡\(λ1\),…,exp⁡\(λm\)\)​𝐐⊤\.\\exp\(\\mathbf\{U\}\)=\\mathbf\{Q\}\\text\{diag\}\(\\exp\(\\lambda\_\{1\}\),\\ldots,\\exp\(\\lambda\_\{m\}\)\)\\mathbf\{Q\}^\{\\top\}\.Furthermore, for an SPD matrix𝐗∈𝒮\+m⊂Sym​\(m\)\\mathbf\{X\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}\\subset\\text\{Sym\}\(m\), the \(principal\) matrix logarithm exists, is unique, and also admits the eigen\-mapping form:

log⁡\(𝐗\)=𝐐​diag​\(log⁡\(γ1\),…,log⁡\(γm\)\)​𝐐⊤,\\log\(\\mathbf\{X\}\)=\\mathbf\{Q\}\\text\{diag\}\(\\log\(\\gamma\_\{1\}\),\\ldots,\\log\(\\gamma\_\{m\}\)\)\\mathbf\{Q\}^\{\\top\},where𝐗=𝐐​diag​\(γ1,…,γm\)​𝐐⊤\\mathbf\{X\}=\\mathbf\{Q\}\\text\{diag\}\(\\gamma\_\{1\},\\ldots,\\gamma\_\{m\}\)\\mathbf\{Q\}^\{\\top\}is the eigen\-decomposition\. Bothexp:Sym​\(m\)↦𝒮\+m\\exp:\\text\{Sym\}\(m\)\\mapsto\\mathcal\{S\}^\{m\}\_\{\+\}and its inverselog:𝒮\+m↦Sym​\(m\)\\log:\\mathcal\{S\}^\{m\}\_\{\+\}\\mapsto\\text\{Sym\}\(m\)are diffeomorphisms\. Let\(D𝐗​log\)​\(𝐔\)\(D\_\{\\mathbf\{X\}\}\\log\)\(\\mathbf\{U\}\)be the differential \(Fréchet derivative\) of the matrix logarithm at the reference point𝐗\\mathbf\{X\}, evaluated in the direction𝐔\\mathbf\{U\}, and\(D𝐔​exp\)​\(𝐕\)\(D\_\{\\mathbf\{U\}\}\\exp\)\(\\mathbf\{V\}\)the differential of the matrix exponential\.

Let𝐗=𝐐​𝚲​𝐐⊤\\mathbf\{X\}=\\mathbf\{Q\}\\mathbf\{\\Lambda\}\\mathbf\{Q\}^\{\\top\}with𝚲=diag​\(γ1,…,γm\)\\mathbf\{\\Lambda\}=\\text\{diag\}\(\\gamma\_\{1\},\\ldots,\\gamma\_\{m\}\)\. The differentialD𝐗​log:𝒯𝐗​𝒮\+m↦𝒯log⁡\(𝐗\)​Sym​\(m\)D\_\{\\mathbf\{X\}\}\\log:\\mathcal\{T\}\_\{\\mathbf\{X\}\}\\mathcal\{S\}^\{m\}\_\{\+\}\\mapsto\\mathcal\{T\}\_\{\\log\(\\mathbf\{X\}\)\}\\text\{Sym\}\(m\)of the matrix logarithm at𝐗\\mathbf\{X\}is given by

\(D𝐗​log\)​\(𝐔\)=𝐐⁡\(𝚲^⊙𝐔~\)​𝐐⊤,𝐔∈𝒯𝐗​𝒮\+m,\(D\_\{\\mathbf\{X\}\}\\log\)\(\\mathbf\{U\}\)=\\mathbf\{Q\}\(\\hat\{\\mathbf\{\\Lambda\}\}\\odot\\tilde\{\\mathbf\{U\}\}\)\\mathbf\{Q\}^\{\\top\},~~\\mathbf\{U\}\\in\\mathcal\{T\}\_\{\\mathbf\{X\}\}\\mathcal\{S\}^\{m\}\_\{\+\},\(8\)where⊙\\odotis the Hadamard product,𝐔~=𝐐⊤​𝐔𝐐\\tilde\{\\mathbf\{U\}\}=\\mathbf\{Q\}^\{\\top\}\\mathbf\{U\}\\mathbf\{Q\}, and the Loewner matrix𝚲^\\hat\{\\mathbf\{\\Lambda\}\}has diagonal entriesΛ^i,i=γi−1\\hat\{\\Lambda\}\_\{i,i\}=\\gamma\_\{i\}^\{\-1\}and off\-diagonal entriesΛ^i,j=log⁡\(γi\)−log⁡\(γj\)γi−γj\\hat\{\\Lambda\}\_\{i,j\}=\\frac\{\\log\(\\gamma\_\{i\}\)\-\\log\(\\gamma\_\{j\}\)\}\{\\gamma\_\{i\}\-\\gamma\_\{j\}\}\.

Let𝐔=𝐐​𝚲​𝐐⊤\\mathbf\{U\}=\\mathbf\{Q\}\\mathbf\{\\Lambda\}\\mathbf\{Q\}^\{\\top\}with𝚲=diag​\(λ1,…,λm\)\\mathbf\{\\Lambda\}=\\text\{diag\}\(\\lambda\_\{1\},\\ldots,\\lambda\_\{m\}\)\. The differentialD𝐔​exp:𝒯𝐔​Sym​\(m\)↦𝒯exp⁡\(𝐔\)​𝒮\+mD\_\{\\mathbf\{U\}\}\\exp:\\mathcal\{T\}\_\{\\mathbf\{U\}\}\\text\{Sym\}\(m\)\\mapsto\\mathcal\{T\}\_\{\\exp\(\\mathbf\{U\}\)\}\\mathcal\{S\}^\{m\}\_\{\+\}of matrix exponential map at𝐔\\mathbf\{U\}is given by

\(D𝐔​exp\)​\(𝐊\)=𝐐⁡\(𝚲^⊙𝐊~\)​𝐐⊤,𝐕∈𝒯𝐔​Sym​\(m\),\(D\_\{\\mathbf\{U\}\}\\exp\)\(\\mathbf\{K\}\)=\\mathbf\{Q\}\(\\hat\{\\mathbf\{\\Lambda\}\}\\odot\\tilde\{\\mathbf\{K\}\}\)\\mathbf\{Q\}^\{\\top\},~~\\mathbf\{V\}\\in\\mathcal\{T\}\_\{\\mathbf\{U\}\}\\text\{Sym\}\(m\),\(9\)where𝐊~=𝐐⊤​𝐊𝐐\\tilde\{\\mathbf\{K\}\}=\\mathbf\{Q\}^\{\\top\}\\mathbf\{K\}\\mathbf\{Q\}, and the Loewner matrix𝚲^\\hat\{\\mathbf\{\\Lambda\}\}has diagonal entriesΛ^i,i=exp⁡\(λi\)\\hat\{\\Lambda\}\_\{i,i\}=\\exp\(\\lambda\_\{i\}\)and off\-diagonal entriesΛ^i,j=exp⁡\(λi\)−exp⁡\(λj\)λi−λj\\hat\{\\Lambda\}\_\{i,j\}=\\frac\{\\exp\(\\lambda\_\{i\}\)\-\\exp\(\\lambda\_\{j\}\)\}\{\\lambda\_\{i\}\-\\lambda\_\{j\}\}\.

#### Riemannian geometry

The LE Riemannian metric for𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}at𝐗∈𝒮\+m\\mathbf\{X\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}is

⟨𝐔,𝐕⟩𝐗LE=⟨\(D𝐗​log\)​\(𝐔\),\(D𝐗​log\)​\(𝐕\)⟩F,𝐔,𝐕∈𝒯𝐗​𝒮\+m\.\\langle\\mathbf\{U\},\\mathbf\{V\}\\rangle\_\{\\mathbf\{X\}\}^\{\\text\{LE\}\}=\\langle\(D\_\{\\mathbf\{X\}\}\\log\)\(\\mathbf\{U\}\),\(D\_\{\\mathbf\{X\}\}\\log\)\(\\mathbf\{V\}\)\\rangle\_\{F\},~~\\mathbf\{U\},\\mathbf\{V\}\\in\\mathcal\{T\}\_\{\\mathbf\{X\}\}\\mathcal\{S\}^\{m\}\_\{\+\}\.The Riemannian exponential map at𝐗∈𝒮\+m\\mathbf\{X\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}is

Exp𝐗LE​\(𝐕\)=exp⁡\(log⁡\(𝐗\)\+\(D𝐗​log\)​\(𝐕\)\),𝐕∈𝒯𝐗​𝒮\+m\.\\text\{Exp\}^\{\\text\{LE\}\}\_\{\\mathbf\{X\}\}\(\\mathbf\{V\}\)=\\exp\(\\log\(\\mathbf\{X\}\)\+\(D\_\{\\mathbf\{X\}\}\\log\)\(\\mathbf\{V\}\)\),~~\\mathbf\{V\}\\in\\mathcal\{T\}\_\{\\mathbf\{X\}\}\\mathcal\{S\}^\{m\}\_\{\+\}\.The Riemannian logarithmic map at𝐗\\mathbf\{X\}is

Log𝐗LE​\(𝐙\)=\(Dlog⁡\(𝐗\)​exp\)​\(log⁡\(𝐙\)−log⁡\(𝐗\)\),𝐙∈𝒮\+m\.\\text\{Log\}^\{\\text\{LE\}\}\_\{\\mathbf\{X\}\}\(\\mathbf\{Z\}\)=\(D\_\{\\log\(\\mathbf\{X\}\)\}\\exp\)\(\\log\(\\mathbf\{Z\}\)\-\\log\(\\mathbf\{X\}\)\),~~\\mathbf\{Z\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}\.\(10\)When the reference point is the identity matrix𝐈\\mathbf\{I\}, the matrix logarithmlog⁡\(𝐈\)\\log\(\\mathbf\{I\}\)is the zero matrix, and hence the differentialDlog⁡\(𝐗\)​expD\_\{\\log\(\\mathbf\{X\}\)\}\\expis the identity linear map\. Therefore, we haveLog𝐈LE​\(𝐙\)=log⁡\(𝐙\)\\text\{Log\}^\{\\text\{LE\}\}\_\{\\mathbf\{I\}\}\(\\mathbf\{Z\}\)=\\log\(\\mathbf\{Z\}\)\. The matrix exponential mapexp\\expis an isometry between\(𝒮\+m,⟨⋅,⋅⟩⋅LE\)\(\\mathcal\{S\}^\{m\}\_\{\+\},\\langle\\cdot,\\cdot\\rangle\_\{\\cdot\}^\{\\text\{LE\}\}\)and\(Sym​\(m\),⟨⋅,⋅⟩F\)\(\\text\{Sym\}\(m\),\\langle\\cdot,\\cdot\\rangle\_\{F\}\), and therefore the geodesic distance between𝐗,𝐙∈𝒮\+m\\mathbf\{X\},\\mathbf\{Z\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}is simplydLE​\(𝐗,𝐙\)=‖log⁡\(𝐗\)−log⁡\(𝐙\)‖Fd\_\{\\text\{LE\}\}\(\\mathbf\{X\},\\mathbf\{Z\}\)=\\\|\\log\(\\mathbf\{X\}\)\-\\log\(\\mathbf\{Z\}\)\\\|\_\{F\}\.

#### Fréchet mean

Given a set of SPD matrices\{𝐗i∈𝒮\+m:i=1,…,n\}\\\{\\mathbf\{X\}\_\{i\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}:i=1,\\ldots,n\\\}, the Fréchet mean𝔼LE\(𝐗1:n\)\\mathbb\{E\}\_\{\\text\{LE\}\}\(\\mathbf\{X\}\_\{1:n\}\)exits, is unique, and is simply the arithmetic mean in the domain of matrix logarithms:

𝔼LE\(𝐗1:n\)=arg​min𝐗∈𝒮\+m∑i=1ndLE\(𝐗,𝐗i\)2=exp\(1n∑i=1nlog\(𝐗i\)\)\.\\mathbb\{E\}\_\{\\text\{LE\}\}\(\\mathbf\{X\}\_\{1:n\}\)=\\argmin\_\{\\mathbf\{X\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}\}\\sum\_\{i=1\}^\{n\}d\_\{\\text\{LE\}\}\(\\mathbf\{X\},\\mathbf\{X\}\_\{i\}\)^\{2\}=\\exp\(\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\log\(\\mathbf\{X\}\_\{i\}\)\)\.\(11\)Under the LE metric,\(𝒮\+m,⟨⋅,⋅⟩⋅LE\)\(\\mathcal\{S\}^\{m\}\_\{\+\},\\langle\\cdot,\\cdot\\rangle\_\{\\cdot\}^\{\\text\{LE\}\}\)is globally isometric via the matrix logarithm to the Euclidean space\(Sym​\(m\),⟨⋅,⋅⟩F\)\(\\text\{Sym\}\(m\),\\langle\\cdot,\\cdot\\rangle\_\{F\}\), and hence\(𝒮\+m,⟨⋅,⋅⟩⋅LE\)\(\\mathcal\{S\}^\{m\}\_\{\+\},\\langle\\cdot,\\cdot\\rangle\_\{\\cdot\}^\{\\text\{LE\}\}\)is a Cartan\-Hadamard manifold\. Moreover, the sectional curvature of\(𝒮\+m,⟨⋅,⋅⟩⋅LE\)\(\\mathcal\{S\}^\{m\}\_\{\+\},\\langle\\cdot,\\cdot\\rangle\_\{\\cdot\}^\{\\text\{LE\}\}\)is constantly zero\.

## Appendix CProof of Proposition[1](https://arxiv.org/html/2609.30487#ThmPro1)

### C\.1Congruence

After conjugation of the input matrix𝐗\\mathbf\{X\}, the output of the multi\-head BiMap layer is\{𝐖1h𝐔𝐗𝐔⊤\(𝐖1h\)⊤:h=1,…,H\}\\\{\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{U\}\\mathbf\{X\}\\mathbf\{U\}^\{\\top\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}:h=1,\\ldots,H\\\}, where𝐔∈O​\(m\)\\mathbf\{U\}\\in\\text\{O\}\(m\)is an orthogonal matrix\. Each weight matrix𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}is full row rank with orthonormal rows\. Let𝐗~2\\tilde\{\\mathbf\{X\}\}\_\{2\}denote the Fréchet mean after congruence on the input𝐗\\mathbf\{X\}\. If there exists an orthogonal matrix𝐐∈O​\(m1\)\\mathbf\{Q\}\\in\\text\{O\}\(m\_\{1\}\)that𝐖1h​𝐔=𝐐𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{U\}=\\mathbf\{Q\}\\mathbf\{W\}\_\{1\}^\{h\}for all1≤h≤H1\\leq h\\leq H, then

𝐖1h​𝐔𝐗𝐔⊤​\(𝐖1h\)⊤=𝐐𝐖1h​𝐗​\(𝐖1h\)⊤​𝐐⊤=𝐐𝐗1h​𝐐⊤\.\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{U\}\\mathbf\{X\}\\mathbf\{U\}^\{\\top\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}=\\mathbf\{Q\}\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\\mathbf\{Q\}^\{\\top\}=\\mathbf\{Q\}\\mathbf\{X\}\_\{1\}^\{h\}\\mathbf\{Q\}^\{\\top\}\.\(12\)Under the LE Riemainnian metric, we havelog⁡\(𝐐𝐗1h​𝐐⊤\)=𝐐​log⁡\(𝐗1h\)​𝐐⊤\\log\(\\mathbf\{Q\}\\mathbf\{X\}\_\{1\}^\{h\}\\mathbf\{Q\}^\{\\top\}\)=\\mathbf\{Q\}\\log\(\\mathbf\{X\}\_\{1\}^\{h\}\)\\mathbf\{Q\}^\{\\top\}, and that the LE geodesic distance is invariant under orthogonal conjugation:

dLE​\(𝐗2,𝐗1h\)=dLE​\(𝐐𝐗2​𝐐⊤,𝐐𝐗1h​𝐐⊤\)\.d\_\{\\text\{LE\}\}\(\\mathbf\{X\}\_\{2\},\\mathbf\{X\}\_\{1\}^\{h\}\)=d\_\{\\text\{LE\}\}\(\\mathbf\{Q\}\\mathbf\{X\}\_\{2\}\\mathbf\{Q\}^\{\\top\},\\mathbf\{Q\}\\mathbf\{X\}\_\{1\}^\{h\}\\mathbf\{Q\}^\{\\top\}\)\.Therefore, after matrix congruence, the LE Fréchet mean is𝐗~2=𝐐𝐗2​𝐐⊤\\tilde\{\\mathbf\{X\}\}\_\{2\}=\\mathbf\{Q\}\\mathbf\{X\}\_\{2\}\\mathbf\{Q\}^\{\\top\}\.

Eq\. \([12](https://arxiv.org/html/2609.30487#A3.E12)\) implies that conjugating the input by𝐔\\mathbf\{U\}is equivalent to conjugating every head’s output by𝐐\\mathbf\{Q\}\. LetIm​\(\(𝐖1h\)⊤\)⊂ℝm1\\text\{Im\}\(\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\\subset\\mathbb\{R\}^\{m\_\{1\}\}denote the subspace spanned by the rows of𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}\. The relation𝐖1h​𝐔=𝐐𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{U\}=\\mathbf\{Q\}\\mathbf\{W\}\_\{1\}^\{h\}forces𝐔​Im​\(\(𝐖1h\)⊤\)=Im​\(\(𝐖1h\)⊤\)\\mathbf\{U\}~\\text\{Im\}\(\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)=\\text\{Im\}\(\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)and𝐖1h​𝐔​\(𝐖1h\)⊤=𝐐\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{U\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}=\\mathbf\{Q\}; that is, each head’s row\-space is𝐔\\mathbf\{U\}\-invariant, and the “compressed” action of𝐔\\mathbf\{U\}onto that subspace is the same𝐐\\mathbf\{Q\}for all heads\. With one head, such a𝐐\\mathbf\{Q\}always exists\. With multiple heads, asking for the same𝐐\\mathbf\{Q\}across allhhis strong: either the subspacesIm​\(\(𝐖1h\)⊤\)\\text\{Im\}\(\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)are all the same and𝐔\\mathbf\{U\}acts on that subspace as𝐐\\mathbf\{Q\}, or𝐔\\mathbf\{U\}acts trivially on everyIm​\(\(𝐖1h\)⊤\)\\text\{Im\}\(\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\.

If each head rotates by its own𝐐h∈O​\(m1\)\\mathbf\{Q\}\_\{h\}\\in\\text\{O\}\(m\_\{1\}\):𝐖1h​𝐔=𝐐h​𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{U\}=\\mathbf\{Q\}\_\{h\}\\mathbf\{W\}\_\{1\}^\{h\}, assume additionally that the candidate Fréchet mean𝐗~2\\tilde\{\\mathbf\{X\}\}\_\{2\}commutes with every𝐐h\\mathbf\{Q\}\_\{h\}:𝐗~2​𝐐h=𝐐h​𝐗~2\\tilde\{\\mathbf\{X\}\}\_\{2\}\\mathbf\{Q\}\_\{h\}=\\mathbf\{Q\}\_\{h\}\\tilde\{\\mathbf\{X\}\}\_\{2\}\. Then under the LE metric, we have

∑h=1H‖log⁡\(𝐗~2\)−log⁡\(𝐐h​𝐗1h​𝐐h⊤\)‖F2\\displaystyle\\sum\_\{h=1\}^\{H\}\\\|\\log\(\\tilde\{\\mathbf\{X\}\}\_\{2\}\)\-\\log\(\\mathbf\{Q\}\_\{h\}\\mathbf\{X\}\_\{1\}^\{h\}\\mathbf\{Q\}\_\{h\}^\{\\top\}\)\\\|\_\{F\}^\{2\}=∑h=1H‖log⁡\(𝐗~2\)−𝐐h​log⁡\(𝐗1h\)​𝐐h⊤‖F2\\displaystyle=\\sum\_\{h=1\}^\{H\}\\\|\\log\(\\tilde\{\\mathbf\{X\}\}\_\{2\}\)\-\\mathbf\{Q\}\_\{h\}\\log\(\\mathbf\{X\}\_\{1\}^\{h\}\)\\mathbf\{Q\}\_\{h\}^\{\\top\}\\\|\_\{F\}^\{2\}=∑h=1H‖log⁡\(𝐗~2\)−log⁡\(𝐗1h\)‖F2\.\\displaystyle=\\sum\_\{h=1\}^\{H\}\\\|\\log\(\\tilde\{\\mathbf\{X\}\}\_\{2\}\)\-\\log\(\\mathbf\{X\}\_\{1\}^\{h\}\)\\\|\_\{F\}^\{2\}\.Therefore, the LE Fréchet mean is𝐗~2=𝐗2\\tilde\{\\mathbf\{X\}\}\_\{2\}=\\mathbf\{X\}\_\{2\}, invariant under general congruence\.

The commutation requirement𝐗~2​𝐐h=𝐐h​𝐗~2\\tilde\{\\mathbf\{X\}\}\_\{2\}\\mathbf\{Q\}\_\{h\}=\\mathbf\{Q\}\_\{h\}\\tilde\{\\mathbf\{X\}\}\_\{2\}implies that𝐗~2\\tilde\{\\mathbf\{X\}\}\_\{2\}is invariant under all the head\-wise rotations:𝐗~2=𝐐h​𝐗~2​𝐐h⊤\\tilde\{\\mathbf\{X\}\}\_\{2\}=\\mathbf\{Q\}\_\{h\}\\tilde\{\\mathbf\{X\}\}\_\{2\}\\mathbf\{Q\}\_\{h\}^\{\\top\}, for everyhh; that is,𝐗~2\\tilde\{\\mathbf\{X\}\}\_\{2\}shares a common orthonormal eigenbasis with all𝐐h\\mathbf\{Q\}\_\{h\}up to blocks: the space splits into invariant blocks for the group⟨𝐐h⟩\\langle\\mathbf\{Q\}\_\{h\}\\rangle, and on each block𝐗~2\\tilde\{\\mathbf\{X\}\}\_\{2\}is a scalar multiple of the identity\. In group\-theory terms,𝐗~2\\tilde\{\\mathbf\{X\}\}\_\{2\}lies in the centralizer of the subgroup generated by\{𝐐h:1≤h≤H\}\\\{\\mathbf\{Q\}\_\{h\}:1\\leq h\\leq H\\\}\.

For the LC Riemannian metric, because the LC geometry depends on Cholesky factors of the matrices𝐗1h=𝐖1h​𝐗​\(𝐖1h\)⊤\\mathbf\{X\}\_\{1\}^\{h\}=\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\(not on𝐗\\mathbf\{X\}directly\), the inner congruence by𝐔∈GL​\(m\)\\mathbf\{U\}\\in\\text\{GL\}\(m\)before compression does not translate to a uniform congruence on the𝐗1h\\mathbf\{X\}\_\{1\}^\{h\}\. In particular, the LC Fréchet mean is defined by mapping each SPD𝐗1h\\mathbf\{X\}\_\{1\}^\{h\}to its Cholesky factor⌊𝔏⁡\(𝐗1h\)⌋\+log⁡\(diag​\(𝔏⁡\(𝐗1h\)\)\)\\lfloor\\mathfrak\{L\}\(\\mathbf\{X\}\_\{1\}^\{h\}\)\\rfloor\+\\log\(\\text\{diag\}\(\\mathfrak\{L\}\(\\mathbf\{X\}\_\{1\}^\{h\}\)\)\), averaging in that triangular space, then mapping back\. This construction is not affine\-invariant and not even orthogonally equivariant in general\. Therefore, the LC Fréchet mean changes in a way that is not a simple transform of𝐗2\\mathbf\{X\}\_\{2\}\.

The affine\-invariant \(AI\) Fréchet mean has congruence equivariance:

𝔼AI​\(𝐐𝐗1​𝐐⊤,…,𝐐𝐗n​𝐐⊤\)=𝐐​𝔼AI​\(𝐗1,…,𝐗n\)​𝐐⊤\.\\mathbb\{E\}\_\{\\text\{AI\}\}\(\\mathbf\{Q\}\\mathbf\{X\}\_\{1\}\\mathbf\{Q\}^\{\\top\},\\ldots,\\mathbf\{Q\}\\mathbf\{X\}\_\{n\}\\mathbf\{Q\}^\{\\top\}\)=\\mathbf\{Q\}\\mathbb\{E\}\_\{\\text\{AI\}\}\(\\mathbf\{X\}\_\{1\},\\ldots,\\mathbf\{X\}\_\{n\}\)\\mathbf\{Q\}^\{\\top\}\.Therefore, the AI Fréchet mean is equivariant under matrix congruence, only if there exists a single𝐐∈GL​\(m1\)\\mathbf\{Q\}\\in\\text\{GL\}\(m\_\{1\}\)that𝐖1h​𝐔=𝐐𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{U\}=\\mathbf\{Q\}\\mathbf\{W\}\_\{1\}^\{h\}for all1≤h≤H1\\leq h\\leq H\.

### C\.2Isometry

A diffeomorphismψ:𝒮\+m→𝒮\+m\\psi:\\mathcal\{S\}^\{m\}\_\{\+\}\\rightarrow\\mathcal\{S\}^\{m\}\_\{\+\}is an isometry ifdg​\(ψ⁡\(𝐗\),ψ⁡\(𝐙\)\)=dg​\(𝐗,𝐙\)d\_\{g\}\(\\psi\(\\mathbf\{X\}\),\\psi\(\\mathbf\{Z\}\)\)=d\_\{g\}\(\\mathbf\{X\},\\mathbf\{Z\}\)for all𝐗,𝐙∈𝒮\+m\\mathbf\{X\},\\mathbf\{Z\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}\. Here,dgd\_\{g\}is the geodesic distance induced by the LC or LE Riemannian metric\. The isometries of𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}form a group under composition, denoted byI⁡\(𝒮\+m\)I\(\\mathcal\{S\}^\{m\}\_\{\+\}\)\. LetF:𝒮\+m→𝒮\+m1F:\\mathcal\{S\}^\{m\}\_\{\+\}\\rightarrow\\mathcal\{S\}^\{m\_\{1\}\}\_\{\+\}be the function defined by𝐗↦𝔼g​\(𝐖11​𝐗​\(𝐖11\)⊤,…,𝐖1H​𝐗​\(𝐖1H\)⊤\)\\mathbf\{X\}\\mapsto\\mathbb\{E\}\_\{g\}\(\\mathbf\{W\}\_\{1\}^\{1\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{1\}\)^\{\\top\},\\ldots,\\mathbf\{W\}\_\{1\}^\{H\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{H\}\)^\{\\top\}\), where𝔼g\\mathbb\{E\}\_\{g\}is either the LC or the LE Fréchet mean\.

Given any two points𝐗,𝐙∈𝒮\+m\\mathbf\{X\},\\mathbf\{Z\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}, the LE geodesic betweenF⁡\(𝐗\)F\(\\mathbf\{X\}\)andF⁡\(𝐙\)F\(\\mathbf\{Z\}\)is

dLE​\(F​\(𝐗\),F​\(𝐙\)\)\\displaystyle d\_\{\\text\{LE\}\}\(F\(\\mathbf\{X\}\),F\(\\mathbf\{Z\}\)\)=‖log⁡\(F⁡\(𝐗\)\)−log⁡\(F⁡\(𝐙\)\)‖F\\displaystyle=\\\|\\log\(F\(\\mathbf\{X\}\)\)\-\\log\(F\(\\mathbf\{Z\}\)\)\\\|\_\{F\}=‖1H​∑h=1Hlog⁡\(𝐖1h​𝐗​\(𝐖1h\)⊤\)−1H​∑h=1Hlog⁡\(𝐖1h​𝐙​\(𝐖1h\)⊤\)‖F,\\displaystyle=\\\|\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\log\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\-\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\log\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{Z\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\\\|\_\{F\},and the LE geodesic betweenF⁡\(ψ⁡\(𝐗\)\)F\(\\psi\(\\mathbf\{X\}\)\)andF⁡\(ψ⁡\(𝐙\)\)F\(\\psi\(\\mathbf\{Z\}\)\)is

dLE​\(F⁡\(ψ⁡\(𝐗\)\),F⁡\(ψ⁡\(𝐙\)\)\)=‖1H​∑h=1Hlog⁡\(𝐖1h​ψ​\(𝐗\)​\(𝐖1h\)⊤\)−1H​∑h=1Hlog⁡\(𝐖1h​ψ​\(𝐙\)​\(𝐖1h\)⊤\)‖F,d\_\{\\text\{LE\}\}\(F\(\\psi\(\\mathbf\{X\}\)\),F\(\\psi\(\\mathbf\{Z\}\)\)\)=\\\|\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\log\(\\mathbf\{W\}\_\{1\}^\{h\}\\psi\(\\mathbf\{X\}\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\-\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\log\(\\mathbf\{W\}\_\{1\}^\{h\}\\psi\(\\mathbf\{Z\}\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\\\|\_\{F\},Any LE\-isometry is exactly a Euclidean isometry in the log domain:ψ⁡\(𝐗\)=exp⁡\(𝒪⁡\(log⁡\(𝐗\)\)\+𝐔\)\\psi\(\\mathbf\{X\}\)=\\exp\(\\mathcal\{O\}\(\\log\(\\mathbf\{X\}\)\)\+\\mathbf\{U\}\), where𝒪:Sym​\(m\)→Sym​\(m\)\\mathcal\{O\}:\\text\{Sym\}\(m\)\\rightarrow\\text\{Sym\}\(m\)is linear orthogonal w\.r\.t\. the Frobenius inner product, and𝐔∈Sym​\(m\)\\mathbf\{U\}\\in\\text\{Sym\}\(m\)is a fixed symmetric matrix \(a translation in the log domain\)\. The crucial observation is that for semi\-orthogonal𝐖1h\\mathbf\{W\}\_\{1\}^\{h\},log⁡\(𝐖1h​ψ​\(𝐗\)​\(𝐖1h\)⊤\)≠𝐖1h​log⁡\(ψ⁡\(𝐗\)\)​\(𝐖1h\)⊤\\log\(\\mathbf\{W\}\_\{1\}^\{h\}\\psi\(\\mathbf\{X\}\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\\neq\\mathbf\{W\}\_\{1\}^\{h\}\\log\(\\psi\(\\mathbf\{X\}\)\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}in general\. Hence the orthogonal action𝒪\\mathcal\{O\}and translation𝐔\\mathbf\{U\}in the log domain do not factor out through the𝐖1h​\(⋅\)​\(𝐖1h\)⊤\\mathbf\{W\}\_\{1\}^\{h\}\(\\cdot\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}compression, so the two distances above generally differ\. Therefore, the mapFFis not invariant \(nor equivariant\) to the action of the LE\-isometry group\.

Likewise, any LC\-isometry is exactly a Euclidean isometry in the LC chart domain:ψ⁡\(𝐗\)=Φ−1​\(𝐑​Φ​\(𝐗\)\+𝐮\)\\psi\(\\mathbf\{X\}\)=\\Phi^\{\-1\}\(\\mathbf\{R\}\\Phi\(\\mathbf\{X\}\)\+\\mathbf\{u\}\), where𝐑∈O⁡\(m⁡\(m\+1\)/2\)\\mathbf\{R\}\\in O\(m\(m\+1\)/2\),𝐮∈ℝm⁡\(m\+1\)/2\\mathbf\{u\}\\in\\mathbb\{R\}^\{m\(m\+1\)/2\}, andΦ\\Phiis the log\-Cholesky coordinate map for the LC metric\. In general,𝔏⁡\(𝐖1h​ψ​\(𝐗\)​\(𝐖1h\)⊤\)≠𝐖1h​𝔏​\(ψ⁡\(𝐗\)\)\\mathfrak\{L\}\(\\mathbf\{W\}\_\{1\}^\{h\}\\psi\(\\mathbf\{X\}\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\\neq\\mathbf\{W\}\_\{1\}^\{h\}\\mathfrak\{L\}\(\\psi\(\\mathbf\{X\}\)\), even for orthogonal𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}\. In fact, the Cholesky factor𝔏⁡\(𝐖1h​ψ​\(𝐗\)​\(𝐖1h\)⊤\)\\mathfrak\{L\}\(\\mathbf\{W\}\_\{1\}^\{h\}\\psi\(\\mathbf\{X\}\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)depends nonlinearly on bothψ⁡\(𝐗\)\\psi\(\\mathbf\{X\}\)and𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}via a re\-triangularization step\. Consequently, the linear orthogonal action𝐑\\mathbf\{R\}and translation𝐮\\mathbf\{u\}in the LC chart do not commute with the compression𝐖1h​ψ​\(𝐗\)​\(𝐖1h\)⊤\\mathbf\{W\}\_\{1\}^\{h\}\\psi\(\\mathbf\{X\}\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\. Therefore,FFis not invariant \(nor equivariant\) under the action of the LC\-isometry group\.

## Appendix DProof of Propositions[2](https://arxiv.org/html/2609.30487#ThmPro2)\-[4](https://arxiv.org/html/2609.30487#ThmPro4)

### D\.1Proof of Proposition[2](https://arxiv.org/html/2609.30487#ThmPro2)

Under the LE metric, we haveF⁡\(𝐗\)=exp⁡\(1H​∑h=1Hlog⁡\(𝐖1h​𝐗​\(𝐖1h\)⊤\)\)F\(\\mathbf\{X\}\)=\\exp\(\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\log\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\), and

dLE​\(F​\(𝐗\),F​\(𝐙\)\)\\displaystyle d\_\{\\text\{LE\}\}\(F\(\\mathbf\{X\}\),F\(\\mathbf\{Z\}\)\)=‖1H​∑h=1H\[log⁡\(𝐖1h​𝐗​\(𝐖1h\)⊤\)−log⁡\(𝐖1h​𝐙​\(𝐖1h\)⊤\)\]‖F\\displaystyle=\\\|\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\left\[\\log\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\-\\log\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{Z\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\\right\]\\\|\_\{F\}≤1H​∑h=1H‖log⁡\(𝐖1h​𝐗​\(𝐖1h\)⊤\)−log⁡\(𝐖1h​𝐙​\(𝐖1h\)⊤\)‖F\.\\displaystyle\\leq\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\\|\\log\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\-\\log\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{Z\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\\\|\_\{F\}\.Letγ⁡\(t\)=exp⁡\(\(1−t\)​log⁡\(𝐙\)\+t​log⁡\(𝐗\)\)\\gamma\(t\)=\\exp\(\(1\-t\)\\log\(\\mathbf\{Z\}\)\+t\\log\(\\mathbf\{X\}\)\)be the LE geodesic from𝐙\\mathbf\{Z\}to𝐗\\mathbf\{X\}, and definefh​\(t\)=log⁡\(𝐖1h​γ​\(t\)​\(𝐖1h\)⊤\)f\_\{h\}\(t\)=\\log\(\\mathbf\{W\}\_\{1\}^\{h\}\\gamma\(t\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\. The functionfh​\(t\)f\_\{h\}\(t\)is absolutely continuous, and therefore we have

‖log⁡\(𝐖1h​𝐗​\(𝐖1h\)⊤\)−log⁡\(𝐖1h​𝐙​\(𝐖1h\)⊤\)‖F=‖fh​\(1\)−fh​\(0\)‖F=‖∫01fh′​\(t\)​𝑑t‖F≤∫01‖fh′​\(t\)‖F​𝑑t\.\\\|\\log\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\-\\log\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{Z\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\\\|\_\{F\}=\\\|f\_\{h\}\(1\)\-f\_\{h\}\(0\)\\\|\_\{F\}=\\\|\\int\_\{0\}^\{1\}f\_\{h\}^\{\\prime\}\(t\)dt\\\|\_\{F\}\\leq\\int\_\{0\}^\{1\}\\\|f\_\{h\}^\{\\prime\}\(t\)\\\|\_\{F\}dt\.We below prove that

‖fh′​\(t\)‖F≤κ⁡\(γ⁡\(t\)\)​‖log⁡\(𝐗\)−log⁡\(𝐙\)‖F=κ⁡\(γ⁡\(t\)\)​dLE​\(𝐗,𝐙\),\\\|f\_\{h\}^\{\\prime\}\(t\)\\\|\_\{F\}\\leq\\kappa\(\\gamma\(t\)\)\\\|\\log\(\\mathbf\{X\}\)\-\\log\(\\mathbf\{Z\}\)\\\|\_\{F\}=\\kappa\(\\gamma\(t\)\)d\_\{\\text\{LE\}\}\(\\mathbf\{X\},\\mathbf\{Z\}\),\(13\)whereκ⁡\(γ⁡\(t\)\)=λmax​\(γ⁡\(t\)\)/λmin​\(γ⁡\(t\)\)\\kappa\(\\gamma\(t\)\)=\\lambda\_\{\\text\{max\}\}\(\\gamma\(t\)\)/\\lambda\_\{\\text\{min\}\}\(\\gamma\(t\)\)is the condition number ofγ⁡\(t\)\\gamma\(t\)\. When𝐗,𝐙∈𝒳α,β\\mathbf\{X\},\\mathbf\{Z\}\\in\\mathcal\{X\}\_\{\\alpha,\\beta\}, we haveγ⁡\(t\)∈𝒳α,β\\gamma\(t\)\\in\\mathcal\{X\}\_\{\\alpha,\\beta\}and henceκ⁡\(γ⁡\(t\)\)≤β/α\\kappa\(\\gamma\(t\)\)\\leq\\beta/\\alpha\. Then it directly follows from Eq\. \([13](https://arxiv.org/html/2609.30487#A4.E13)\) thatFFis Lipschitz on𝒳α,β\\mathcal\{X\}\_\{\\alpha,\\beta\}with constantL≤β/αL\\leq\\beta/\\alpha\.

Write𝐔⁡\(t\)=log⁡\(γ⁡\(t\)\)=\(1−t\)​log⁡\(𝐙\)\+t​log⁡\(𝐗\)\\mathbf\{U\}\(t\)=\\log\(\\gamma\(t\)\)=\(1\-t\)\\log\(\\mathbf\{Z\}\)\+t\\log\(\\mathbf\{X\}\)and𝐕=log⁡\(𝐗\)−log⁡\(𝐙\)\\mathbf\{V\}=\\log\(\\mathbf\{X\}\)\-\\log\(\\mathbf\{Z\}\)\. We haveγ′​\(t\)=\(D𝐔⁡\(t\)​exp\)​\(𝐕\)\\gamma^\{\\prime\}\(t\)=\(D\_\{\\mathbf\{U\}\(t\)\}\\exp\)\(\\mathbf\{V\}\), and the chain rule gives

fh′​\(t\)=\(D𝐖1h​γ​\(t\)​\(𝐖1h\)⊤​log\)​\(𝐖1h​\(D𝐔⁡\(t\)​exp\)​\(𝐕\)​\(𝐖1h\)⊤\)\.f\_\{h\}^\{\\prime\}\(t\)=\(D\_\{\\mathbf\{W\}\_\{1\}^\{h\}\\gamma\(t\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\}\\log\)\\left\(\\mathbf\{W\}\_\{1\}^\{h\}\(D\_\{\\mathbf\{U\}\(t\)\}\\exp\)\(\\mathbf\{V\}\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\\right\)\.The differentialD𝐗​log:𝒯𝐗​𝒮\+m↦𝒯log⁡\(𝐗\)​Sym​\(m\)D\_\{\\mathbf\{X\}\}\\log:\\mathcal\{T\}\_\{\\mathbf\{X\}\}\\mathcal\{S\}^\{m\}\_\{\+\}\\mapsto\\mathcal\{T\}\_\{\\log\(\\mathbf\{X\}\)\}\\text\{Sym\}\(m\)of the matrix logarithm at𝐗\\mathbf\{X\}has the integral form

\(D𝐗​log\)​\(𝐊\)=∫0∞\(𝐗\+s​𝐈\)−1​𝐊​\(𝐗\+s​𝐈\)−1​𝑑s,𝐊∈𝒯𝐗​𝒮\+m,\(D\_\{\\mathbf\{X\}\}\\log\)\(\\mathbf\{K\}\)=\\int\_\{0\}^\{\\infty\}\(\\mathbf\{X\}\+s\\mathbf\{I\}\)^\{\-1\}\\mathbf\{K\}\(\\mathbf\{X\}\+s\\mathbf\{I\}\)^\{\-1\}ds,~~\\mathbf\{K\}\\in\\mathcal\{T\}\_\{\\mathbf\{X\}\}\\mathcal\{S\}^\{m\}\_\{\+\},Taking Frobenius norms and using‖\(𝐗\+s​𝐈\)−1‖2=1λmin​\(𝐗\)\+s\\\|\(\\mathbf\{X\}\+s\\mathbf\{I\}\)^\{\-1\}\\\|\_\{2\}=\\frac\{1\}\{\\lambda\_\{\\text\{min\}\}\(\\mathbf\{X\}\)\+s\}, we have

‖\(D𝐗​log\)​\(𝐊\)‖F≤∫0∞‖\(𝐗\+s​𝐈\)−1‖22​‖𝐊‖F​𝑑s=‖𝐊‖F​∫0∞d​s\(λmin​\(𝐗\)\+s\)2=‖𝐊‖Fλmin​\(𝐗\)\.\\\|\(D\_\{\\mathbf\{X\}\}\\log\)\(\\mathbf\{K\}\)\\\|\_\{F\}\\leq\\int\_\{0\}^\{\\infty\}\\\|\(\\mathbf\{X\}\+s\\mathbf\{I\}\)^\{\-1\}\\\|\_\{2\}^\{2\}\\\|\\mathbf\{K\}\\\|\_\{F\}ds=\\\|\\mathbf\{K\}\\\|\_\{F\}\\int\_\{0\}^\{\\infty\}\\frac\{ds\}\{\(\\lambda\_\{\\text\{min\}\}\(\\mathbf\{X\}\)\+s\)^\{2\}\}=\\frac\{\\\|\\mathbf\{K\}\\\|\_\{F\}\}\{\\lambda\_\{\\text\{min\}\}\(\\mathbf\{X\}\)\}\.Therefore, we have

‖fh′​\(t\)‖F≤‖𝐖1h​\(D𝐔⁡\(t\)​exp\)​\(𝐕\)​\(𝐖1h\)⊤‖Fλmin​\(𝐖1h​γ​\(t\)​\(𝐖1h\)⊤\)≤‖\(D𝐔⁡\(t\)​exp\)​\(𝐕\)‖Fλmin​\(𝐖1h​γ​\(t\)​\(𝐖1h\)⊤\),\\\|f\_\{h\}^\{\\prime\}\(t\)\\\|\_\{F\}\\leq\\frac\{\\\|\\mathbf\{W\}\_\{1\}^\{h\}\(D\_\{\\mathbf\{U\}\(t\)\}\\exp\)\(\\mathbf\{V\}\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\\\|\_\{F\}\}\{\\lambda\_\{\\text\{min\}\}\(\\mathbf\{W\}\_\{1\}^\{h\}\\gamma\(t\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\}\\leq\\frac\{\\\|\(D\_\{\\mathbf\{U\}\(t\)\}\\exp\)\(\\mathbf\{V\}\)\\\|\_\{F\}\}\{\\lambda\_\{\\text\{min\}\}\(\\mathbf\{W\}\_\{1\}^\{h\}\\gamma\(t\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\},where we have utilized the property that‖𝐗𝐘𝐙‖F≤‖𝐗‖2​‖𝐘‖F​‖𝐙‖2\\\|\\mathbf\{X\}\\mathbf\{Y\}\\mathbf\{Z\}\\\|\_\{F\}\\leq\\\|\\mathbf\{X\}\\\|\_\{2\}\\\|\\mathbf\{Y\}\\\|\_\{F\}\\\|\\mathbf\{Z\}\\\|\_\{2\}for any matrices\{𝐗,𝐘,𝐙\}\\\{\\mathbf\{X\},\\mathbf\{Y\},\\mathbf\{Z\}\\\}, and‖𝐗‖2=1\\\|\\mathbf\{X\}\\\|\_\{2\}=1when𝐗\\mathbf\{X\}has orthonormal rows\. Moreover, because𝐖1h​γ​\(t\)​\(𝐖1h\)⊤\\mathbf\{W\}\_\{1\}^\{h\}\\gamma\(t\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}is a compression ofγ⁡\(t\)\\gamma\(t\)by an isometry, eigenvalues interlace\[[Fan and Pall, 1957](https://arxiv.org/html/2609.30487#bib.bib11)\]:λmin​\(𝐖1h​γ​\(t\)​\(𝐖1h\)⊤\)≥λmin​\(γ⁡\(t\)\)\\lambda\_\{\\text\{min\}\}\(\\mathbf\{W\}\_\{1\}^\{h\}\\gamma\(t\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\\geq\\lambda\_\{\\text\{min\}\}\(\\gamma\(t\)\)\.

The differentialD𝐔​exp:𝒯𝐔​Sym​\(m\)↦𝒯exp⁡\(𝐔\)​𝒮\+mD\_\{\\mathbf\{U\}\}\\exp:\\mathcal\{T\}\_\{\\mathbf\{U\}\}\\text\{Sym\}\(m\)\\mapsto\\mathcal\{T\}\_\{\\exp\(\\mathbf\{U\}\)\}\\mathcal\{S\}^\{m\}\_\{\+\}of matrix exponential map at𝐔\\mathbf\{U\}has the integral form

\(D𝐔​exp\)​\(𝐕\)=∫01exp⁡\(\(1−s\)​𝐔\)​𝐕​exp⁡\(s​𝐔\)​𝑑s,𝐕∈𝒯𝐗​𝒮\+m\.\(D\_\{\\mathbf\{U\}\}\\exp\)\(\\mathbf\{V\}\)=\\int\_\{0\}^\{1\}\\exp\(\(1\-s\)\\mathbf\{U\}\)\\mathbf\{V\}\\exp\(s\\mathbf\{U\}\)ds,~~\\mathbf\{V\}\\in\\mathcal\{T\}\_\{\\mathbf\{X\}\}\\mathcal\{S\}^\{m\}\_\{\+\}\.Taking Frobenius norms, we have

‖\(D𝐔​exp\)​\(𝐕\)‖F≤∫01‖exp⁡\(\(1−s\)​𝐔\)‖2​‖𝐕‖F​‖exp⁡\(s​𝐔\)‖2​𝑑s=‖𝐕‖F​‖exp⁡\(𝐔\)‖2\.\\\|\(D\_\{\\mathbf\{U\}\}\\exp\)\(\\mathbf\{V\}\)\\\|\_\{F\}\\leq\\int\_\{0\}^\{1\}\\\|\\exp\(\(1\-s\)\\mathbf\{U\}\)\\\|\_\{2\}\\\|\\mathbf\{V\}\\\|\_\{F\}\\\|\\exp\(s\\mathbf\{U\}\)\\\|\_\{2\}ds=\\\|\\mathbf\{V\}\\\|\_\{F\}\\\|\\exp\(\\mathbf\{U\}\)\\\|\_\{2\}\.When𝐔=𝐔⁡\(t\)\\mathbf\{U\}=\\mathbf\{U\}\(t\), we have‖exp⁡\(𝐔⁡\(t\)\)‖2=‖γ⁡\(t\)‖2=λmax​\(γ⁡\(t\)\)\\\|\\exp\(\\mathbf\{U\}\(t\)\)\\\|\_\{2\}=\\\|\\gamma\(t\)\\\|\_\{2\}=\\lambda\_\{\\text\{max\}\}\(\\gamma\(t\)\), and substituting𝐕\\mathbf\{V\}withlog⁡\(𝐗\)−log⁡\(𝐙\)\\log\(\\mathbf\{X\}\)\-\\log\(\\mathbf\{Z\}\), we have

‖fh′​\(t\)‖F≤λmax​\(γ⁡\(t\)\)​‖log⁡\(𝐗\)−log⁡\(𝐙\)‖Fλmin​\(𝐖1h​γ​\(t\)​\(𝐖1h\)⊤\)≤λmax​\(γ​\(t\)\)λmin​\(γ​\(t\)\)​‖log⁡\(𝐗\)−log⁡\(𝐙\)‖F\.\\\|f\_\{h\}^\{\\prime\}\(t\)\\\|\_\{F\}\\leq\\frac\{\\lambda\_\{\\text\{max\}\}\(\\gamma\(t\)\)\\\|\\log\(\\mathbf\{X\}\)\-\\log\(\\mathbf\{Z\}\)\\\|\_\{F\}\}\{\\lambda\_\{\\text\{min\}\}\(\\mathbf\{W\}\_\{1\}^\{h\}\\gamma\(t\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\}\\leq\\frac\{\\lambda\_\{\\text\{max\}\}\(\\gamma\(t\)\)\}\{\\lambda\_\{\\text\{min\}\}\(\\gamma\(t\)\)\}\\\|\\log\(\\mathbf\{X\}\)\-\\log\(\\mathbf\{Z\}\)\\\|\_\{F\}\.

### D\.2Proof of Proposition[3](https://arxiv.org/html/2609.30487#ThmPro3)

The log\-Cholesky coordinate mapΦ:𝒮\+m→ℝm⁡\(m\+1\)/2\\Phi:\\mathcal\{S\}^\{m\}\_\{\+\}\\rightarrow\\mathbb\{R\}^\{m\(m\+1\)/2\},Φ⁡\(𝐗\)=\(vnz​\(𝔏⁡\(𝐗\)\),log⁡\(vd​\(𝔏⁡\(𝐗\)\)\)\)\\Phi\(\\mathbf\{X\}\)=\(\\text\{vnz\}\(\\mathfrak\{L\}\(\\mathbf\{X\}\)\),\\log\(\\text\{vd\}\(\\mathfrak\{L\}\(\\mathbf\{X\}\)\)\)\)is a global coordinate map, wherevnz​\(𝔏​\(𝐗\)\)\\text\{vnz\}\(\\mathfrak\{L\}\(\\mathbf\{X\}\)\)is the vector of the non\-zero entries of⌊𝔏⁡\(𝐗\)⌋\\lfloor\\mathfrak\{L\}\(\\mathbf\{X\}\)\\rfloor, andvd​\(𝔏​\(𝐗\)\)\\text\{vd\}\(\\mathfrak\{L\}\(\\mathbf\{X\}\)\)is the vector of the diagonal entries of𝔏⁡\(𝐗\)\\mathfrak\{L\}\(\\mathbf\{X\}\)\. The LC metric makes𝒮\+m\\mathcal\{S\}^\{m\}\_\{\+\}flat via the chart mapΦ\\Phi, and the LC geodesic distance is simply the Euclidean norm in the log\-Cholesky coordinate:dLC​\(𝐗,𝐙\)=‖Φ⁡\(𝐗\)−Φ⁡\(𝐙\)‖2d\_\{\\text\{LC\}\}\(\\mathbf\{X\},\\mathbf\{Z\}\)=\\\|\\Phi\(\\mathbf\{X\}\)\-\\Phi\(\\mathbf\{Z\}\)\\\|\_\{2\}\. For any𝐗,𝐙∈𝒮\+m\\mathbf\{X\},\\mathbf\{Z\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}, we have

dLC​\(F​\(𝐗\),F​\(𝐙\)\)\\displaystyle d\_\{\\text\{LC\}\}\(F\(\\mathbf\{X\}\),F\(\\mathbf\{Z\}\)\)=‖1H​∑h=1H\[Φ⁡\(𝐖1h​𝐗​\(𝐖1h\)⊤\)−Φ⁡\(𝐖1h​𝐙​\(𝐖1h\)⊤\)\]‖2\\displaystyle=\\\|\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\left\[\\Phi\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\-\\Phi\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{Z\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\\right\]\\\|\_\{2\}≤1H​∑h=1H‖Φ⁡\(𝐖1h​𝐗​\(𝐖1h\)⊤\)−Φ⁡\(𝐖1h​𝐙​\(𝐖1h\)⊤\)‖2\.\\displaystyle\\leq\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\\|\\Phi\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\-\\Phi\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{Z\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\\\|\_\{2\}\.In the proof below, we drop the indices and write𝐖\\mathbf\{W\}for𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}\. We have𝐖𝐗𝐖⊤=\(𝐖​𝔏​\(𝐗\)\)​\(𝐖​𝔏​\(𝐗\)\)⊤\\mathbf\{W\}\\mathbf\{X\}\\mathbf\{W\}^\{\\top\}=\(\\mathbf\{W\}\\mathfrak\{L\}\(\\mathbf\{X\}\)\)\(\\mathbf\{W\}\\mathfrak\{L\}\(\\mathbf\{X\}\)\)^\{\\top\}, and a thin QR decomposition:\(𝐖​𝔏​\(𝐗\)\)⊤=𝐐𝐗​𝐑𝐗,\(\\mathbf\{W\}\\mathfrak\{L\}\(\\mathbf\{X\}\)\)^\{\\top\}=\\mathbf\{Q\}\_\{\\mathbf\{X\}\}\\mathbf\{R\}\_\{\\mathbf\{X\}\},where𝐐𝐗∈ℝm×m1\\mathbf\{Q\}\_\{\\mathbf\{X\}\}\\in\\mathbb\{R\}^\{m\\times m\_\{1\}\}has orthonormal columns, and𝐑𝐗∈ℝm1×m1\\mathbf\{R\}\_\{\\mathbf\{X\}\}\\in\\mathbb\{R\}^\{m\_\{1\}\\times m\_\{1\}\}is upper\-triangular with positive diagonal\. The Cholesky factor of𝐖𝐗𝐖⊤\\mathbf\{W\}\\mathbf\{X\}\\mathbf\{W\}^\{\\top\}is exactly𝐑𝐗⊤\\mathbf\{R\}\_\{\\mathbf\{X\}\}^\{\\top\}:𝔏⁡\(𝐖𝐗𝐖⊤\)=𝐑𝐗⊤\\mathfrak\{L\}\(\\mathbf\{W\}\\mathbf\{X\}\\mathbf\{W\}^\{\\top\}\)=\\mathbf\{R\}\_\{\\mathbf\{X\}\}^\{\\top\}\. Therefore, we have

‖Φ⁡\(𝐖𝐗𝐖⊤\)−Φ⁡\(𝐖𝐙𝐖⊤\)‖22=‖⌊𝐑𝐗⊤⌋−⌊𝐑𝐙⊤⌋‖F2\+‖log⁡\(diag​\(𝐑𝐗⊤\)\)−log⁡\(diag​\(𝐑𝐙⊤\)\)‖F2\.\\\|\\Phi\(\\mathbf\{W\}\\mathbf\{X\}\\mathbf\{W\}^\{\\top\}\)\-\\Phi\(\\mathbf\{W\}\\mathbf\{Z\}\\mathbf\{W\}^\{\\top\}\)\\\|\_\{2\}^\{2\}=\\\|\\lfloor\\mathbf\{R\}\_\{\\mathbf\{X\}\}^\{\\top\}\\rfloor\-\\lfloor\\mathbf\{R\}\_\{\\mathbf\{Z\}\}^\{\\top\}\\rfloor\\\|\_\{F\}^\{2\}\+\\\|\\log\(\\text\{diag\}\(\\mathbf\{R\}\_\{\\mathbf\{X\}\}^\{\\top\}\)\)\-\\log\(\\text\{diag\}\(\\mathbf\{R\}\_\{\\mathbf\{Z\}\}^\{\\top\}\)\)\\\|\_\{F\}^\{2\}\.
Fix any spectral band0<α≤β<∞0<\\alpha\\leq\\beta<\\infty, and consider the subset𝒳α,β=\{𝐗∈𝒮\+m:α​𝐈⪯𝐗⪯β​𝐈\}\\mathcal\{X\}\_\{\\alpha,\\beta\}=\\\{\\mathbf\{X\}\\in\\mathcal\{S\}^\{m\}\_\{\+\}:\\alpha\\mathbf\{I\}\\preceq\\mathbf\{X\}\\preceq\\beta\\mathbf\{I\}\\\}\. Then we have

‖Φ⁡\(𝐖𝐗𝐖⊤\)−Φ⁡\(𝐖𝐙𝐖⊤\)‖22\\displaystyle\\\|\\Phi\(\\mathbf\{W\}\\mathbf\{X\}\\mathbf\{W\}^\{\\top\}\)\-\\Phi\(\\mathbf\{W\}\\mathbf\{Z\}\\mathbf\{W\}^\{\\top\}\)\\\|\_\{2\}^\{2\}≤‖⌊𝐑𝐗⊤⌋−⌊𝐑𝐙⊤⌋‖F2\+α−1​‖diag​\(𝐑𝐗⊤\)−diag​\(𝐑𝐙⊤\)‖F2\\displaystyle\\leq\\\|\\lfloor\\mathbf\{R\}\_\{\\mathbf\{X\}\}^\{\\top\}\\rfloor\-\\lfloor\\mathbf\{R\}\_\{\\mathbf\{Z\}\}^\{\\top\}\\rfloor\\\|\_\{F\}^\{2\}\+\\alpha^\{\-1\}\\\|\\text\{diag\}\(\\mathbf\{R\}\_\{\\mathbf\{X\}\}^\{\\top\}\)\-\\text\{diag\}\(\\mathbf\{R\}\_\{\\mathbf\{Z\}\}^\{\\top\}\)\\\|\_\{F\}^\{2\}≤max⁡\{1,α−1\}​‖𝐑𝐗⊤−𝐑𝐙⊤‖F2\.\\displaystyle\\leq\\max\\\{1,\\alpha^\{\-1\}\\\}\\\|\\mathbf\{R\}\_\{\\mathbf\{X\}\}^\{\\top\}\-\\mathbf\{R\}\_\{\\mathbf\{Z\}\}^\{\\top\}\\\|\_\{F\}^\{2\}\.
Define a line in them×mm\\times mmatrix space:𝐘⁡\(t\)=\(1−t\)​𝐙\+t​𝐗\\mathbf\{Y\}\(t\)=\(1\-t\)\\mathbf\{Z\}\+t\\mathbf\{X\}, for0≤t≤10\\leq t\\leq 1, and let𝐀⁡\(t\)=𝐖𝐘⁡\(t\)​𝐖⊤\\mathbf\{A\}\(t\)=\\mathbf\{W\}\\mathbf\{Y\}\(t\)\\mathbf\{W\}^\{\\top\}\. It can be readily proved that𝐀⁡\(t\)∈𝒳α,β\\mathbf\{A\}\(t\)\\in\\mathcal\{X\}\_\{\\alpha,\\beta\}, for any0≤t≤10\\leq t\\leq 1\. The derivative inttis𝐀˙​\(t\)=𝐖⁡\(𝐗−𝐙\)​𝐖⊤\\dot\{\\mathbf\{A\}\}\(t\)=\\mathbf\{W\}\(\\mathbf\{X\}\-\\mathbf\{Z\}\)\\mathbf\{W\}^\{\\top\}\.

Let𝐋⁡\(t\)\\mathbf\{L\}\(t\)be the Cholesky factor of𝐀⁡\(t\)\\mathbf\{A\}\(t\):𝔏⁡\(𝐀⁡\(t\)\)=𝐋⁡\(t\)\\mathfrak\{L\}\(\\mathbf\{A\}\(t\)\)=\\mathbf\{L\}\(t\)\. In particular,𝐋⁡\(0\)=𝐑𝐙⊤\\mathbf\{L\}\(0\)=\\mathbf\{R\}\_\{\\mathbf\{Z\}\}^\{\\top\}and𝐋⁡\(1\)=𝐑𝐗⊤\\mathbf\{L\}\(1\)=\\mathbf\{R\}\_\{\\mathbf\{X\}\}^\{\\top\}\. Differentiate the Cholesky map𝐀⁡\(t\)=𝐋⁡\(t\)​𝐋​\(t\)⊤\\mathbf\{A\}\(t\)=\\mathbf\{L\}\(t\)\\mathbf\{L\}\(t\)^\{\\top\}:𝐀˙​\(t\)=𝐋˙​\(t\)​𝐋​\(t\)⊤\+𝐋⁡\(t\)​𝐋˙​\(t\)⊤\\dot\{\\mathbf\{A\}\}\(t\)=\\dot\{\\mathbf\{L\}\}\(t\)\\mathbf\{L\}\(t\)^\{\\top\}\+\\mathbf\{L\}\(t\)\\dot\{\\mathbf\{L\}\}\(t\)^\{\\top\}, and then we have

𝐋​\(t\)−1​𝐀˙​\(t\)​𝐋​\(t\)−⁣⊤=𝐋​\(t\)−1​𝐋˙​\(t\)\+𝐋˙​\(t\)⊤​𝐋​\(t\)−⁣⊤=𝛀\+𝛀⊤,\\mathbf\{L\}\(t\)^\{\-1\}\\dot\{\\mathbf\{A\}\}\(t\)\\mathbf\{L\}\(t\)^\{\-\\top\}=\\mathbf\{L\}\(t\)^\{\-1\}\\dot\{\\mathbf\{L\}\}\(t\)\+\\dot\{\\mathbf\{L\}\}\(t\)^\{\\top\}\\mathbf\{L\}\(t\)^\{\-\\top\}=\\mathbf\{\\Omega\}\+\\mathbf\{\\Omega\}^\{\\top\},where𝛀=𝐋​\(t\)−1​𝐋˙​\(t\)\\mathbf\{\\Omega\}=\\mathbf\{L\}\(t\)^\{\-1\}\\dot\{\\mathbf\{L\}\}\(t\)is lower\-triangular\. Therefore, we have‖𝛀‖F≤‖𝛀\+𝛀⊤‖F≤‖𝐋​\(t\)−1‖22​‖𝐀˙​\(t\)‖F\.\\\|\\mathbf\{\\Omega\}\\\|\_\{F\}\\leq\\\|\\mathbf\{\\Omega\}\+\\mathbf\{\\Omega\}^\{\\top\}\\\|\_\{F\}\\leq\\\|\\mathbf\{L\}\(t\)^\{\-1\}\\\|\_\{2\}^\{2\}\\\|\\dot\{\\mathbf\{A\}\}\(t\)\\\|\_\{F\}\.Finally𝐋˙​\(t\)=𝐋​\(t\)​𝛀\\dot\{\\mathbf\{L\}\}\(t\)=\\mathbf\{L\}\(t\)\\mathbf\{\\Omega\}gives

‖𝐋˙​\(t\)‖F≤‖𝐋⁡\(t\)‖2​‖𝛀‖F≤‖𝐋⁡\(t\)‖2​‖𝐋​\(t\)−1‖22​‖𝐀˙​\(t\)‖F\.\\\|\\dot\{\\mathbf\{L\}\}\(t\)\\\|\_\{F\}\\leq\\\|\\mathbf\{L\}\(t\)\\\|\_\{2\}\\\|\\mathbf\{\\Omega\}\\\|\_\{F\}\\leq\\\|\\mathbf\{L\}\(t\)\\\|\_\{2\}\\\|\\mathbf\{L\}\(t\)^\{\-1\}\\\|\_\{2\}^\{2\}\\\|\\dot\{\\mathbf\{A\}\}\(t\)\\\|\_\{F\}\.The squared singular values of𝐋⁡\(t\)\\mathbf\{L\}\(t\)are the eigen values of𝐀⁡\(t\)\\mathbf\{A\}\(t\), and therefore‖𝐋⁡\(t\)‖2≤β\\\|\\mathbf\{L\}\(t\)\\\|\_\{2\}\\leq\\sqrt\{\\beta\}and‖𝐋​\(t\)−1‖2≤1α\\\|\\mathbf\{L\}\(t\)^\{\-1\}\\\|\_\{2\}\\leq\\frac\{1\}\{\\sqrt\{\\alpha\}\}\. Finally, we have‖𝐋˙​\(t\)‖F≤βα​‖𝐀˙​\(t\)‖F\.\\\|\\dot\{\\mathbf\{L\}\}\(t\)\\\|\_\{F\}\\leq\\frac\{\\sqrt\{\\beta\}\}\{\\alpha\}\\\|\\dot\{\\mathbf\{A\}\}\(t\)\\\|\_\{F\}\.

Integrating along the path, we have

‖𝐑𝐗⊤−𝐑𝐙⊤‖F≤∫01‖𝐋˙​\(t\)‖F​𝑑t≤βα​∫01‖𝐀˙​\(t\)‖F​𝑑t=βα​‖𝐖⁡\(𝐗−𝐙\)​𝐖⊤‖F\.\\\|\\mathbf\{R\}\_\{\\mathbf\{X\}\}^\{\\top\}\-\\mathbf\{R\}\_\{\\mathbf\{Z\}\}^\{\\top\}\\\|\_\{F\}\\leq\\int\_\{0\}^\{1\}\\\|\\dot\{\\mathbf\{L\}\}\(t\)\\\|\_\{F\}dt\\leq\\frac\{\\sqrt\{\\beta\}\}\{\\alpha\}\\int\_\{0\}^\{1\}\\\|\\dot\{\\mathbf\{A\}\}\(t\)\\\|\_\{F\}dt=\\frac\{\\sqrt\{\\beta\}\}\{\\alpha\}\\\|\\mathbf\{W\}\(\\mathbf\{X\}\-\\mathbf\{Z\}\)\\mathbf\{W\}^\{\\top\}\\\|\_\{F\}\.Given that𝐖𝐖⊤=𝐈\\mathbf\{W\}\\mathbf\{W\}^\{\\top\}=\\mathbf\{I\}, we have‖𝐖⁡\(𝐗−𝐙\)​𝐖⊤‖F≤‖𝐖‖2​‖𝐗−𝐙‖F​‖𝐖⊤‖2=‖𝐗−𝐙‖F\\\|\\mathbf\{W\}\(\\mathbf\{X\}\-\\mathbf\{Z\}\)\\mathbf\{W\}^\{\\top\}\\\|\_\{F\}\\leq\\\|\\mathbf\{W\}\\\|\_\{2\}\\\|\\mathbf\{X\}\-\\mathbf\{Z\}\\\|\_\{F\}\\\|\\mathbf\{W\}^\{\\top\}\\\|\_\{2\}=\\\|\\mathbf\{X\}\-\\mathbf\{Z\}\\\|\_\{F\}\. Write𝐗−𝐙=𝔏⁡\(𝐗\)​\(𝔏⁡\(𝐗\)−𝔏⁡\(𝐙\)\)⊤\+\(𝔏⁡\(𝐗\)−𝔏⁡\(𝐙\)\)​𝔏​\(𝐙\)⊤\\mathbf\{X\}\-\\mathbf\{Z\}=\\mathfrak\{L\}\(\\mathbf\{X\}\)\(\\mathfrak\{L\}\(\\mathbf\{X\}\)\-\\mathfrak\{L\}\(\\mathbf\{Z\}\)\)^\{\\top\}\+\(\\mathfrak\{L\}\(\\mathbf\{X\}\)\-\\mathfrak\{L\}\(\\mathbf\{Z\}\)\)\\mathfrak\{L\}\(\\mathbf\{Z\}\)^\{\\top\}, we have

‖𝐗−𝐙‖F≤\(‖𝔏⁡\(𝐗\)‖2\+‖𝔏⁡\(𝐙\)‖2\)​‖𝔏⁡\(𝐗\)−𝔏⁡\(𝐙\)‖F≤2​β​‖𝔏⁡\(𝐗\)−𝔏⁡\(𝐙\)‖F\.\\\|\\mathbf\{X\}\-\\mathbf\{Z\}\\\|\_\{F\}\\leq\(\\\|\\mathfrak\{L\}\(\\mathbf\{X\}\)\\\|\_\{2\}\+\\\|\\mathfrak\{L\}\(\\mathbf\{Z\}\)\\\|\_\{2\}\)\\\|\\mathfrak\{L\}\(\\mathbf\{X\}\)\-\\mathfrak\{L\}\(\\mathbf\{Z\}\)\\\|\_\{F\}\\leq 2\\sqrt\{\\beta\}\\\|\\mathfrak\{L\}\(\\mathbf\{X\}\)\-\\mathfrak\{L\}\(\\mathbf\{Z\}\)\\\|\_\{F\}\.Applying the mean\-value theorem on the diagonal entries, we have

\|\[𝔏⁡\(𝐗\)\]i,i−\[𝔏⁡\(𝐙\)\]i,i\|\\displaystyle\|\[\\mathfrak\{L\}\(\\mathbf\{X\}\)\]\_\{i,i\}\-\[\\mathfrak\{L\}\(\\mathbf\{Z\}\)\]\_\{i,i\}\|≤max⁡\{\[𝔏⁡\(𝐗\)\]i,i,\[𝔏⁡\(𝐙\)\]i,i\}​\|log⁡\(\[𝔏⁡\(𝐗\)\]i,i\)−log⁡\(\[𝔏⁡\(𝐙\)\]i,i\)\|\\displaystyle\\leq\\max\\\{\[\\mathfrak\{L\}\(\\mathbf\{X\}\)\]\_\{i,i\},\[\\mathfrak\{L\}\(\\mathbf\{Z\}\)\]\_\{i,i\}\\\}\|\\log\(\[\\mathfrak\{L\}\(\\mathbf\{X\}\)\]\_\{i,i\}\)\-\\log\(\[\\mathfrak\{L\}\(\\mathbf\{Z\}\)\]\_\{i,i\}\)\|≤β​\|log⁡\(\[𝔏⁡\(𝐗\)\]i,i\)−log⁡\(\[𝔏⁡\(𝐙\)\]i,i\)\|,\\displaystyle\\leq\\sqrt\{\\beta\}\|\\log\(\[\\mathfrak\{L\}\(\\mathbf\{X\}\)\]\_\{i,i\}\)\-\\log\(\[\\mathfrak\{L\}\(\\mathbf\{Z\}\)\]\_\{i,i\}\)\|,and therefore

‖𝔏⁡\(𝐗\)−𝔏⁡\(𝐙\)‖F2\\displaystyle\\\|\\mathfrak\{L\}\(\\mathbf\{X\}\)\-\\mathfrak\{L\}\(\\mathbf\{Z\}\)\\\|\_\{F\}^\{2\}≤‖⌊𝔏⁡\(𝐗\)⌋−⌊𝔏⁡\(𝐙\)⌋‖F2\+β​‖log⁡\(diag​\(𝔏⁡\(𝐗\)\)\)−log⁡\(diag​\(𝔏⁡\(𝐙\)\)\)‖F2\\displaystyle\\leq\\\|\\lfloor\\mathfrak\{L\}\(\\mathbf\{X\}\)\\rfloor\-\\lfloor\\mathfrak\{L\}\(\\mathbf\{Z\}\)\\rfloor\\\|\_\{F\}^\{2\}\+\\beta\\\|\\log\(\\text\{diag\}\(\\mathfrak\{L\}\(\\mathbf\{X\}\)\)\)\-\\log\(\\text\{diag\}\(\\mathfrak\{L\}\(\\mathbf\{Z\}\)\)\)\\\|\_\{F\}^\{2\}≤β​‖Φ⁡\(𝐗\)−Φ⁡\(𝐙\)‖22\\displaystyle\\leq\\beta\\\|\\Phi\(\\mathbf\{X\}\)\-\\Phi\(\\mathbf\{Z\}\)\\\|\_\{2\}^\{2\}Putting everything together, we have

‖Φ⁡\(𝐖𝐗𝐖⊤\)−Φ⁡\(𝐖𝐙𝐖⊤\)‖2≤max⁡\{1,α−1\}​2​β3/2α​‖Φ⁡\(𝐗\)−Φ⁡\(𝐙\)‖2,\\\|\\Phi\(\\mathbf\{W\}\\mathbf\{X\}\\mathbf\{W\}^\{\\top\}\)\-\\Phi\(\\mathbf\{W\}\\mathbf\{Z\}\\mathbf\{W\}^\{\\top\}\)\\\|\_\{2\}\\leq\\max\\\{1,\\alpha^\{\-1\}\\\}\\frac\{2\\beta^\{3/2\}\}\{\\alpha\}\\\|\\Phi\(\\mathbf\{X\}\)\-\\Phi\(\\mathbf\{Z\}\)\\\|\_\{2\},for any𝐖\\mathbf\{W\}of orthonormal rows\.

### D\.3Proof of Proposition[4](https://arxiv.org/html/2609.30487#ThmPro4)

Letϕ\\phibe either the matrix logarithm maplog\\logor the log\-Cholesky coordinate mapΦ\\Phi\. The unified formula for the Fréchet mean in the Pooling layer is

F⁡\(𝐗,\{𝐖1h\}h=1H\)=ϕ−1​\(1H​∑h=1Hϕ⁡\(𝐖1h​𝐗​\(𝐖1h\)⊤\)\)\.F\(\\mathbf\{X\};\\\{\\mathbf\{W\}\_\{1\}^\{h\}\\\}\_\{h=1\}^\{H\}\)=\\phi^\{\-1\}\(\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\phi\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\)\.Write𝐙=F⁡\(𝐗,\{𝐖1h\}h=1H\)\\mathbf\{Z\}=F\(\\mathbf\{X\};\\\{\\mathbf\{W\}\_\{1\}^\{h\}\\\}\_\{h=1\}^\{H\}\)\. If the functionFFis “linear”, then there exists another weight matrix𝐖∈ℝm1×m\\mathbf\{W\}\\in\\mathbb\{R\}^\{m\_\{1\}\\times m\}with𝐖𝐖⊤=𝐈\\mathbf\{W\}\\mathbf\{W\}^\{\\top\}=\\mathbf\{I\}such that𝐙=F⁡\(𝐗,\{𝐖1h\}h=1H\)=𝐖𝐗𝐖⊤\\mathbf\{Z\}=F\(\\mathbf\{X\};\\\{\\mathbf\{W\}\_\{1\}^\{h\}\\\}\_\{h=1\}^\{H\}\)=\\mathbf\{W\}\\mathbf\{X\}\\mathbf\{W\}^\{\\top\}\.

Letλ1​\(𝐗\)≥…≥λm​\(𝐗\)\\lambda\_\{1\}\(\\mathbf\{X\}\)\\geq\.\.\.\\geq\\lambda\_\{m\}\(\\mathbf\{X\}\)be the eigenvalues of𝐗\\mathbf\{X\}, andμ1​\(𝐙\)≥…≥μm1​\(𝐙\)\\mu\_\{1\}\(\\mathbf\{Z\}\)\\geq\.\.\.\\geq\\mu\_\{m\_\{1\}\}\(\\mathbf\{Z\}\)be the eigenvalues of𝐙\\mathbf\{Z\}\. According to[Fan and Pall \[1957\]](https://arxiv.org/html/2609.30487#bib.bib11), if and only if the eigenvalues of𝐙\\mathbf\{Z\}satisfy the interlacing inequalities:λj​\(𝐗\)≥μj​\(𝐙\)≥λj\+\(m−m1\)​\(𝐗\)\\lambda\_\{j\}\(\\mathbf\{X\}\)\\geq\\mu\_\{j\}\(\\mathbf\{Z\}\)\\geq\\lambda\_\{j\+\(m\-m\_\{1\}\)\}\(\\mathbf\{X\}\), for anyj=1,…,m1j=1,\\ldots,m\_\{1\}, then there exists an𝐖∈ℝm1×m\\mathbf\{W\}\\in\\mathbb\{R\}^\{m\_\{1\}\\times m\}with𝐖𝐖⊤=𝐈\\mathbf\{W\}\\mathbf\{W\}^\{\\top\}=\\mathbf\{I\}such that𝐙=𝐖𝐗𝐖⊤\\mathbf\{Z\}=\\mathbf\{W\}\\mathbf\{X\}\\mathbf\{W\}^\{\\top\}\.

Under the LE metric, with the eigenvalue decomposition𝐖1h​𝐗​\(𝐖1h\)⊤=𝐐1h​𝚲1h​\(𝐐1h\)⊤\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}=\\mathbf\{Q\}\_\{1\}^\{h\}\\mathbf\{\\Lambda\}\_\{1\}^\{h\}\(\\mathbf\{Q\}\_\{1\}^\{h\}\)^\{\\top\}, the Fréchet mean is

𝐙=exp⁡\(1H​∑j=1H𝐐1h​log⁡\(𝚲1h\)​\(𝐐1h\)⊤\)\.\\mathbf\{Z\}=\\exp\(\\frac\{1\}\{H\}\\sum\_\{j=1\}^\{H\}\\mathbf\{Q\}\_\{1\}^\{h\}\\log\(\\mathbf\{\\Lambda\}\_\{1\}^\{h\}\)\(\\mathbf\{Q\}\_\{1\}^\{h\}\)^\{\\top\}\)\.Note that, in general, the heads\{𝐖1h𝐗\(𝐖1h\)⊤:h=1,…,H\}\\\{\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}:h=1,\\ldots,H\\\}do not have a common orthonormal eigenbasis\. Therefore, although each compression𝐖1h​𝐗​\(𝐖1h\)⊤\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}individually interlaces with𝐗\\mathbf\{X\}, the eigenvalues of the average𝐙\\mathbf\{Z\}in general do not satisfy the interlacing inequalities\.

The key point is thatΩ:=\{log⁡\(𝐖𝐗𝐖⊤\):𝐖𝐖⊤=𝐈\}\\Omega:=\\\{\\log\(\\mathbf\{W\}\\mathbf\{X\}\\mathbf\{W\}^\{\\top\}\):\\mathbf\{W\}\\mathbf\{W\}^\{\\top\}=\\mathbf\{I\}\\\}is not a convex set form1≥2m\_\{1\}\\geq 2, and hence taking average of\{log\(𝐖1h𝐗\(𝐖1h\)⊤\):h=1,…,H\}\\\{\\log\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\):h=1,\\ldots,H\\\}can leave the setΩ\\Omega\. Therefore, in general there is no𝐖\\mathbf\{W\}with𝐖𝐖⊤=𝐈\\mathbf\{W\}\\mathbf\{W\}^\{\\top\}=\\mathbf\{I\}such thatOPENlog⁡\(𝐖𝐗𝐖⊤\)=1H​∑h=1Hlog⁡\(𝐖1h​𝐗​\(𝐖1h\)⊤\)\)\\log\(\\mathbf\{W\}\\mathbf\{X\}\\mathbf\{W\}^\{\\top\}\)=\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\log\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\)\. Sinceexp\\expandlog\\logare inverse diffeomorphisms between𝒮\+m1\\mathcal\{S\}^\{m\_\{1\}\}\_\{\+\}and Sym\(m1m\_\{1\}\), this implies there is no such𝐖\\mathbf\{W\}withF⁡\(𝐗,\{𝐖1h\}h=1H\)=𝐖𝐗𝐖⊤F\(\\mathbf\{X\};\\\{\\mathbf\{W\}\_\{1\}^\{h\}\\\}\_\{h=1\}^\{H\}\)=\\mathbf\{W\}\\mathbf\{X\}\\mathbf\{W\}^\{\\top\}\. Consequently, the spectrum of the average1H​∑j=1H𝐐1h​log⁡\(𝚲1h\)​\(𝐐1h\)⊤\\frac\{1\}\{H\}\\sum\_\{j=1\}^\{H\}\\mathbf\{Q\}\_\{1\}^\{h\}\\log\(\\mathbf\{\\Lambda\}\_\{1\}^\{h\}\)\(\\mathbf\{Q\}\_\{1\}^\{h\}\)^\{\\top\}need not satisfy the Fan\-Pall interlacing inequalities relative to𝐗\\mathbf\{X\}\.

An identical argument holds for the LC metric by replacinglog\\logwithΦ\\Phiand noting that the set\{Φ⁡\(𝐖𝐗𝐖⊤\):𝐖𝐖⊤=𝐈\}\\\{\\Phi\(\\mathbf\{W\}\\mathbf\{X\}\\mathbf\{W\}^\{\\top\}\):\\mathbf\{W\}\\mathbf\{W\}^\{\\top\}=\\mathbf\{I\}\\\}is likewise non\-convex\.

## Appendix EBackpropagation in the matrix\-to\-vector module

We absorb each BiMap layer into its succeeding operation, yielding two composite mapsfa​p​\(⋅\)f^\{ap\}\(\\cdot\)andfl​o​g​\(⋅\)f^\{log\}\(\\cdot\)\. In particular,fa​p​\(⋅\)f^\{ap\}\(\\cdot\)directly maps𝐗⁡\(t\)\\mathbf\{X\}\(t\)to𝐗3​\(t\)\\mathbf\{X\}\_\{3\}\(t\):

𝐗3​\(t\)=fa​p​\(𝐗⁡\(t\),\{𝐖1h\}h=1H\)=ϕ−1​\(αH​∑h=1Hϕ⁡\(𝐖1h​𝐗​\(t\)​\(𝐖1h\)⊤\)\),\\mathbf\{X\}\_\{3\}\(t\)=f^\{ap\}\(\\mathbf\{X\}\(t\);\\\{\\mathbf\{W\}\_\{1\}^\{h\}\\\}\_\{h=1\}^\{H\}\)=\\phi^\{\-1\}\(\\frac\{\\alpha\}\{H\}\\sum\_\{h=1\}^\{H\}\\phi\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(t\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\),thenfl​o​g​\(⋅\)f^\{log\}\(\\cdot\)directly maps𝐗3​\(t\)\\mathbf\{X\}\_\{3\}\(t\)to𝐘h​\(t\)\\mathbf\{Y\}^\{h\}\(t\):

𝐘h\(t\)=fl​o​g\(𝐗3\(t\);𝐖4h\)=Log𝐈g\(𝐖4h𝐗3\(t\)\(𝐖4h\)⊤\),h=1,…,H\.\\mathbf\{Y\}^\{h\}\(t\)=f^\{log\}\(\\mathbf\{X\}\_\{3\}\(t\);\\mathbf\{W\}\_\{4\}^\{h\}\)=\\text\{Log\}\_\{\\mathbf\{I\}\}^\{g\}\(\\mathbf\{W\}\_\{4\}^\{h\}\\mathbf\{X\}\_\{3\}\(t\)\(\\mathbf\{W\}\_\{4\}^\{h\}\)^\{\\top\}\),~~h=1,\\ldots,H\.
Note that bothfa​p​\(⋅\)f^\{ap\}\(\\cdot\)andfl​o​g​\(⋅\)f^\{log\}\(\\cdot\)act pointwise in time: they are applied independently at eachttto the matrix𝐗⁡\(t\)\\mathbf\{X\}\(t\)or𝐗3​\(t\)\\mathbf\{X\}\_\{3\}\(t\), rather than being functionals of the entire trajectory𝑿∈ℋg​\(\[t0,t1\],𝒮\+m\)\{\\bm\{\\mathsfit\{X\}\}\}\\in\\mathcal\{H\}\_\{g\}\(\[t\_\{0\},t\_\{1\}\];\\mathcal\{S\}^\{m\}\_\{\+\}\)\. Consequently, the derivatives with respect to the inputs, namely∂fa​p∂𝐗⁡\(t\)\\frac\{\\partial f^\{ap\}\}\{\\partial\\mathbf\{X\}\(t\)\}and∂fl​o​g∂𝐗3​\(t\)\\frac\{\\partial f^\{log\}\}\{\\partial\\mathbf\{X\}\_\{3\}\(t\)\}, are the usual matrix Jacobians evaluated at timett\(as in the static, single\-matrix case\), rather than functional/variational derivatives with respect to the path\.

However, the functional layer in the encoderℰ\\mathcal\{E\}, namely Eq\. equation[2](https://arxiv.org/html/2609.30487#S3.E2), is a functional of the input vector\-valued function\. We write𝒛=∫t0t1𝑾\(1\)​\(t\)​𝒚​\(t\)​𝑑t\+𝒃\\boldsymbol\{z\}=\\int\_\{t\_\{0\}\}^\{t\_\{1\}\}\\boldsymbol\{W\}^\{\(1\)\}\(t\)\\boldsymbol\{y\}\(t\)dt\+\\boldsymbol\{b\}and𝒙\(1\)=a⁡\(𝒛\)\\boldsymbol\{x\}^\{\(1\)\}=a\(\\boldsymbol\{z\}\)\. For a perturbation𝒚→𝒚\+ϵ​𝜼\\boldsymbol\{y\}\\rightarrow\\boldsymbol\{y\}\+\\epsilon\\boldsymbol\{\\eta\}, the first variation is

δ​𝒛=∫t0t1𝑾\(1\)​\(t\)​𝜼​\(t\)​𝑑t,δ​𝒙\(1\)=diag​\(a′​\(𝒛\)\)​δ​𝒛=∫t0t1\[diag​\(a′​\(𝒛\)\)​𝑾\(1\)​\(t\)\]​𝜼​\(t\)​𝑑t,\\delta\\boldsymbol\{z\}=\\int\_\{t\_\{0\}\}^\{t\_\{1\}\}\\boldsymbol\{W\}^\{\(1\)\}\(t\)\\boldsymbol\{\\eta\}\(t\)dt,~~~\\delta\\boldsymbol\{x\}^\{\(1\)\}=\\text\{diag\}\(a^\{\\prime\}\(\\boldsymbol\{z\}\)\)\\delta\\boldsymbol\{z\}=\\int\_\{t\_\{0\}\}^\{t\_\{1\}\}\[\\text\{diag\}\(a^\{\\prime\}\(\\boldsymbol\{z\}\)\)\\boldsymbol\{W\}^\{\(1\)\}\(t\)\]\\boldsymbol\{\\eta\}\(t\)dt,wherediag​\(a′​\(𝒛\)\)\\text\{diag\}\(a^\{\\prime\}\(\\boldsymbol\{z\}\)\)is the Jacobian of the element\-wise activationaaat𝒛\\boldsymbol\{z\}\. Hence, the functional derivative density with respect to the whole function𝒚\\boldsymbol\{y\}is the matrix kernel∂𝒙\(1\)∂𝒚​\(s\)=diag​\(a′​\(𝒛\)\)​𝑾\(1\)​\(s\)∈ℝq1×p\\frac\{\\partial\\boldsymbol\{x\}^\{\(1\)\}\}\{\\partial\\boldsymbol\{y\}\}\(s\)=\\text\{diag\}\(a^\{\\prime\}\(\\boldsymbol\{z\}\)\)\\boldsymbol\{W\}^\{\(1\)\}\(s\)\\in\\mathbb\{R\}^\{q\_\{1\}\\times p\}fors∈\[t0,t1\]s\\in\[t\_\{0\},t\_\{1\}\], whereq1q\_\{1\}is the dimension of𝒙\(1\)\\boldsymbol\{x\}^\{\(1\)\}\.

Recall that𝒚​\(t\)⊤=\[vl​\(𝐘1​\(t\)\)⊤,…,vl​\(𝐘H​\(t\)\)⊤\]\\boldsymbol\{y\}\(t\)^\{\\top\}=\[\\text\{vl\}\(\\mathbf\{Y\}^\{1\}\(t\)\)^\{\\top\},\\ldots,\\text\{vl\}\(\\mathbf\{Y\}^\{H\}\(t\)\)^\{\\top\}\]and vl is an isometry\. Now split the matrix kernel intoHHblocks:diag​\(a′​\(𝒛\)\)​𝑾\(1\)​\(s\)=\[𝐊1​\(s\),…,𝐊H​\(s\)\]\\text\{diag\}\(a^\{\\prime\}\(\\boldsymbol\{z\}\)\)\\boldsymbol\{W\}^\{\(1\)\}\(s\)=\[\\mathbf\{K\}\_\{1\}\(s\),\\ldots,\\mathbf\{K\}\_\{H\}\(s\)\], where𝐊h​\(s\)∈ℝq1×m2​\(m2\+1\)/2\\mathbf\{K\}\_\{h\}\(s\)\\in\\mathbb\{R\}^\{q\_\{1\}\\times m\_\{2\}\(m\_\{2\}\+1\)/2\}andp=H×m2​\(m2\+1\)/2p=H\\times m\_\{2\}\(m\_\{2\}\+1\)/2\. Then the functional derivative density w\.r\.t\. the whole𝐘h\\mathbf\{Y\}^\{h\}is

∂xr\(1\)∂𝐘h\(s\)=vl−1\(\[𝐊h\(s\)\]r,:⊤\),forr=1,…,q1\.\\frac\{\\partial x^\{\(1\)\}\_\{r\}\}\{\\partial\\mathbf\{Y\}^\{h\}\}\(s\)=\\text\{vl\}^\{\-1\}\(\[\\mathbf\{K\}\_\{h\}\(s\)\]\_\{r,:\}^\{\\top\}\),\\text\{ for \}r=1,\\ldots,q\_\{1\}\.
### E\.1Logarithm layer

For each head1≤h≤H1\\leq h\\leq H, we have𝐘h​\(t\)=Log𝐈g​\(𝐖4h​𝐗3​\(t\)​\(𝐖4h\)⊤\)\\mathbf\{Y\}^\{h\}\(t\)=\\text\{Log\}\_\{\\mathbf\{I\}\}^\{g\}\(\\mathbf\{W\}\_\{4\}^\{h\}\\mathbf\{X\}\_\{3\}\(t\)\(\\mathbf\{W\}\_\{4\}^\{h\}\)^\{\\top\}\), where\(𝐖4h\)⊤∈St​\(m1,m2\)\(\\mathbf\{W\}\_\{4\}^\{h\}\)^\{\\top\}\\in\\text\{St\}\(m\_\{1\},m\_\{2\}\)\. To simplify notation, we drop the head indexhhand write𝐗4​\(t\)=𝐖4​𝐗3​\(t\)​𝐖4⊤\\mathbf\{X\}\_\{4\}\(t\)=\\mathbf\{W\}\_\{4\}\\mathbf\{X\}\_\{3\}\(t\)\\mathbf\{W\}\_\{4\}^\{\\top\}\.

#### LE metric

Whenggis the LE metric, the Riemannian logarithm at the identity is exactly the matrix logarithm:Log𝐈g​\(𝐗4​\(t\)\)=log⁡\(𝐗4​\(t\)\)\\text\{Log\}\_\{\\mathbf\{I\}\}^\{g\}\(\\mathbf\{X\}\_\{4\}\(t\)\)=\\log\(\\mathbf\{X\}\_\{4\}\(t\)\)\. The Fréchet derivative of log at𝐗4​\(t\)\\mathbf\{X\}\_\{4\}\(t\)applied to a direction𝐔\\mathbf\{U\}is

\(D𝐗4​\(t\)​log\)​\(𝐔\)=∫0∞\[𝐗4​\(t\)\+s​𝐈\]−1​𝐔​\[𝐗4​\(t\)\+s​𝐈\]−1​𝑑s,\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(\\mathbf\{U\}\)=\\int\_\{0\}^\{\\infty\}\[\\mathbf\{X\}\_\{4\}\(t\)\+s\\mathbf\{I\}\]^\{\-1\}\\mathbf\{U\}\[\\mathbf\{X\}\_\{4\}\(t\)\+s\\mathbf\{I\}\]^\{\-1\}ds,which is self\-adjoint w\.r\.t\. the Frobenius inner product\. If𝐔\\mathbf\{U\}is symmetric \(resp\. skew\), then\(D𝐗4​\(t\)​log\)​\(𝐔\)\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(\\mathbf\{U\}\)is symmetric \(resp\. skew\)\. The spectral form for the Fréchet derivative\(D𝐗4​\(t\)​log\)​\(𝐔\)\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(\\mathbf\{U\}\)is given in Eq\. equation[8](https://arxiv.org/html/2609.30487#A2.E8)\. Then, by the chain ruleD𝐖4​𝐘​\(t\)=D𝐗4​\(t\)​log∘D𝐖4​𝐗4​\(t\)D\_\{\\mathbf\{W\}\_\{4\}\}\\mathbf\{Y\}\(t\)=D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\\circ D\_\{\\mathbf\{W\}\_\{4\}\}\\mathbf\{X\}\_\{4\}\(t\), the directional derivative with respect to𝐖4\\mathbf\{W\}\_\{4\}in directionΔ​𝐖\\Delta\\mathbf\{W\}is

\(D𝐖4​𝐘​\(t\)\)​\(Δ​𝐖\)=\(D𝐗4​\(t\)​log\)​\(Δ​𝐖𝐗3​\(t\)​𝐖4⊤\+𝐖4​𝐗3​\(t\)​Δ​𝐖⊤\)\.\(D\_\{\\mathbf\{W\}\_\{4\}\}\\mathbf\{Y\}\(t\)\)\(\\Delta\\mathbf\{W\}\)=\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(\\Delta\\mathbf\{W\}\\mathbf\{X\}\_\{3\}\(t\)\\mathbf\{W\}\_\{4\}^\{\\top\}\+\\mathbf\{W\}\_\{4\}\\mathbf\{X\}\_\{3\}\(t\)\\Delta\\mathbf\{W\}^\{\\top\}\)\.
Let𝐕⁡\(t\)\\mathbf\{V\}\(t\)be an “upstream gradient”, and writef⁡\(t\)=⟨𝐕⁡\(t\),𝐘⁡\(t\)⟩Ff\(t\)=\\langle\\mathbf\{V\}\(t\),\\mathbf\{Y\}\(t\)\\rangle\_\{F\}\. The linear mapping\(D𝐗4​\(t\)​log\)​\(⋅\)\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(\\cdot\)is self\-adjoint, and therefore we have

d​f​\(t\)=⟨𝐕⁡\(t\),d​𝐘​\(t\)⟩F\\displaystyle df\(t\)=\\langle\\mathbf\{V\}\(t\),d\\mathbf\{Y\}\(t\)\\rangle\_\{F\}=⟨𝐕⁡\(t\),\(D𝐗4​\(t\)​log\)​\(d​𝐗4​\(t\)\)⟩F\\displaystyle=\\langle\\mathbf\{V\}\(t\),\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(d\\mathbf\{X\}\_\{4\}\(t\)\)\\rangle\_\{F\}=⟨𝐕s​\(t\),\(D𝐗4​\(t\)​log\)​\(d​𝐗4​\(t\)\)⟩F=⟨\(D𝐗4​\(t\)​log\)​\(𝐕s​\(t\)\),d​𝐗4​\(t\)⟩F,\\displaystyle=\\langle\\mathbf\{V\}\_\{s\}\(t\),\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(d\\mathbf\{X\}\_\{4\}\(t\)\)\\rangle\_\{F\}=\\langle\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(\\mathbf\{V\}\_\{s\}\(t\)\),d\\mathbf\{X\}\_\{4\}\(t\)\\rangle\_\{F\},whered​𝐗4​\(t\)=\(d​𝐖4\)​𝐗3​\(t\)​𝐖4⊤\+𝐖4​𝐗3​\(t\)​\(d​𝐖4\)⊤d\\mathbf\{X\}\_\{4\}\(t\)=\(d\\mathbf\{W\}\_\{4\}\)\\mathbf\{X\}\_\{3\}\(t\)\\mathbf\{W\}\_\{4\}^\{\\top\}\+\\mathbf\{W\}\_\{4\}\\mathbf\{X\}\_\{3\}\(t\)\(d\\mathbf\{W\}\_\{4\}\)^\{\\top\}, and𝐕s​\(t\)=12​\(𝐕⁡\(t\)\+𝐕​\(t\)⊤\)\\mathbf\{V\}\_\{s\}\(t\)=\\frac\{1\}\{2\}\(\\mathbf\{V\}\(t\)\+\\mathbf\{V\}\(t\)^\{\\top\}\)\. Then we obtain

d​f​\(t\)\\displaystyle df\(t\)=tr​\(\(D𝐗4​\(t\)​log\)​\(𝐕s​\(t\)\)​\(d​𝐖4\)​𝐗3​\(t\)​𝐖4⊤\)\+tr​\(\(D𝐗4​\(t\)​log\)​\(𝐕s​\(t\)\)​𝐖4​𝐗3​\(t\)​\(d​𝐖4\)⊤\)\\displaystyle=\\text\{tr\}\(\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(\\mathbf\{V\}\_\{s\}\(t\)\)\(d\\mathbf\{W\}\_\{4\}\)\\mathbf\{X\}\_\{3\}\(t\)\\mathbf\{W\}\_\{4\}^\{\\top\}\)\+\\text\{tr\}\(\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(\\mathbf\{V\}\_\{s\}\(t\)\)\\mathbf\{W\}\_\{4\}\\mathbf\{X\}\_\{3\}\(t\)\(d\\mathbf\{W\}\_\{4\}\)^\{\\top\}\)=tr​\(𝐗3​\(t\)​𝐖4⊤​\(D𝐗4​\(t\)​log\)​\(𝐕s​\(t\)\)​\(d​𝐖4\)\)\+tr​\(\(\(D𝐗4​\(t\)​log\)​\(𝐕s​\(t\)\)​𝐖4​𝐗3​\(t\)\)⊤​\(d​𝐖4\)\)\\displaystyle=\\text\{tr\}\(\\mathbf\{X\}\_\{3\}\(t\)\\mathbf\{W\}\_\{4\}^\{\\top\}\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(\\mathbf\{V\}\_\{s\}\(t\)\)\(d\\mathbf\{W\}\_\{4\}\)\)\+\\text\{tr\}\(\(\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(\\mathbf\{V\}\_\{s\}\(t\)\)\\mathbf\{W\}\_\{4\}\\mathbf\{X\}\_\{3\}\(t\)\)^\{\\top\}\(d\\mathbf\{W\}\_\{4\}\)\)=⟨2​\(D𝐗4​\(t\)​log\)​\(𝐕s​\(t\)\)​𝐖4​𝐗3​\(t\),d​𝐖4⟩F\.\\displaystyle=\\langle 2\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(\\mathbf\{V\}\_\{s\}\(t\)\)\\mathbf\{W\}\_\{4\}\\mathbf\{X\}\_\{3\}\(t\),d\\mathbf\{W\}\_\{4\}\\rangle\_\{F\}\.Then theEuclideangradient w\.r\.t\.𝐖4\\mathbf\{W\}\_\{4\}is

∇𝐖4f​\(t\)=2​\(D𝐗4​\(t\)​log\)​\(𝐕s​\(t\)\)​𝐖4​𝐗3​\(t\)\.\\nabla\_\{\\mathbf\{W\}\_\{4\}\}f\(t\)=2\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(\\mathbf\{V\}\_\{s\}\(t\)\)\\mathbf\{W\}\_\{4\}\\mathbf\{X\}\_\{3\}\(t\)\.\(14\)Then we can project∇𝐖4f​\(t\)\\nabla\_\{\\mathbf\{W\}\_\{4\}\}f\(t\)onto the tangent space to get the Riemannian gradient\.

When taking directional derivative with respect to𝐗3​\(t\)\\mathbf\{X\}\_\{3\}\(t\), we have

d​f​\(t\)=⟨𝐕⁡\(t\),\(D𝐗4​\(t\)​log\)​\(d​𝐗4​\(t\)\)⟩F\\displaystyle df\(t\)=\\langle\\mathbf\{V\}\(t\),\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(d\\mathbf\{X\}\_\{4\}\(t\)\)\\rangle\_\{F\}=⟨\(D𝐗4​\(t\)​log\)​\(𝐕s​\(t\)\),d​𝐗4​\(t\)⟩F\\displaystyle=\\langle\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(\\mathbf\{V\}\_\{s\}\(t\)\),d\\mathbf\{X\}\_\{4\}\(t\)\\rangle\_\{F\}=⟨\(D𝐗4​\(t\)​log\)​\(𝐕s​\(t\)\),𝐖4​d​𝐗3​\(t\)​𝐖4⊤⟩F\\displaystyle=\\langle\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(\\mathbf\{V\}\_\{s\}\(t\)\),\\mathbf\{W\}\_\{4\}d\\mathbf\{X\}\_\{3\}\(t\)\\mathbf\{W\}\_\{4\}^\{\\top\}\\rangle\_\{F\}=⟨𝐖4⊤​\(D𝐗4​\(t\)​log\)​\(𝐕s​\(t\)\)​𝐖4,d​𝐗3​\(t\)⟩F\.\\displaystyle=\\langle\\mathbf\{W\}\_\{4\}^\{\\top\}\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(\\mathbf\{V\}\_\{s\}\(t\)\)\\mathbf\{W\}\_\{4\},d\\mathbf\{X\}\_\{3\}\(t\)\\rangle\_\{F\}\.The Euclidean gradient w\.r\.t\. the input𝐗3​\(t\)\\mathbf\{X\}\_\{3\}\(t\)is

∇𝐗3​\(t\)f​\(t\)=𝐖4⊤​\(D𝐗4​\(t\)​log\)​\(𝐕s​\(t\)\)​𝐖4\.\\nabla\_\{\\mathbf\{X\}\_\{3\}\(t\)\}f\(t\)=\\mathbf\{W\}\_\{4\}^\{\\top\}\(D\_\{\\mathbf\{X\}\_\{4\}\(t\)\}\\log\)\(\\mathbf\{V\}\_\{s\}\(t\)\)\\mathbf\{W\}\_\{4\}\.\(15\)

#### LC metric

Let𝐋⁡\(t\)=ℒ⁡\(𝐗4​\(t\)\)\\mathbf\{L\}\(t\)=\\mathcal\{L\}\(\\mathbf\{X\}\_\{4\}\(t\)\)denote the Cholesky factor of𝐗4​\(t\)=𝐖4​𝐗3​\(t\)​𝐖4⊤\\mathbf\{X\}\_\{4\}\(t\)=\\mathbf\{W\}\_\{4\}\\mathbf\{X\}\_\{3\}\(t\)\\mathbf\{W\}\_\{4\}^\{\\top\}\. Then𝐘⁡\(t\)=Log𝐈LC​\(𝐗4​\(t\)\)=2​log⁡\(diag​\(𝐋⁡\(t\)\)\)\+⌊𝐋⁡\(t\)⌋\+⌊𝐋⁡\(t\)⌋⊤\\mathbf\{Y\}\(t\)=\\text\{Log\}\_\{\\mathbf\{I\}\}^\{\\text\{LC\}\}\(\\mathbf\{X\}\_\{4\}\(t\)\)=2\\log\(\\text\{diag\}\(\\mathbf\{L\}\(t\)\)\)\+\\lfloor\\mathbf\{L\}\(t\)\\rfloor\+\\lfloor\\mathbf\{L\}\(t\)\\rfloor^\{\\top\}, and hence

d​𝐘​\(t\)=2​diag​\(d​L1,1​\(t\)L1,1​\(t\),⋯,d​Lm2,m2​\(t\)Lm2,m2​\(t\)\)\+⌊d​𝐋​\(t\)⌋\+⌊d​𝐋​\(t\)⌋⊤\.d\\mathbf\{Y\}\(t\)=2\\text\{diag\}\(\\frac\{dL\_\{1,1\}\(t\)\}\{L\_\{1,1\}\(t\)\},\\cdots,\\frac\{dL\_\{m\_\{2\},m\_\{2\}\}\(t\)\}\{L\_\{m\_\{2\},m\_\{2\}\}\(t\)\}\)\+\\lfloor d\\mathbf\{L\}\(t\)\\rfloor\+\\lfloor d\\mathbf\{L\}\(t\)\\rfloor^\{\\top\}\.Let𝐕⁡\(t\)\\mathbf\{V\}\(t\)be an “upstream gradient”, and writef⁡\(t\)=⟨𝐕⁡\(t\),𝐘⁡\(t\)⟩Ff\(t\)=\\langle\\mathbf\{V\}\(t\),\\mathbf\{Y\}\(t\)\\rangle\_\{F\}\. Then the differential is

d​f​\(t\)=⟨𝐕⁡\(t\),d​𝐘​\(t\)⟩F\\displaystyle df\(t\)=\\langle\\mathbf\{V\}\(t\),d\\mathbf\{Y\}\(t\)\\rangle\_\{F\}=2​∑i=1m2Vi,i​\(t\)​d​Li,i​\(t\)Li,i​\(t\)\+∑i\>j\[Vi,j​\(t\)\+Vj,i​\(t\)\]​d​Li,j​\(t\)\\displaystyle=2\\sum\_\{i=1\}^\{m\_\{2\}\}V\_\{i,i\}\(t\)\\frac\{dL\_\{i,i\}\(t\)\}\{L\_\{i,i\}\(t\)\}\+\\sum\_\{i\>j\}\[V\_\{i,j\}\(t\)\+V\_\{j,i\}\(t\)\]dL\_\{i,j\}\(t\)=⟨𝐇⁡\(t\),d​𝐋​\(t\)⟩F,\\displaystyle=\\langle\\mathbf\{H\}\(t\),d\\mathbf\{L\}\(t\)\\rangle\_\{F\},where𝐇⁡\(t\)\\mathbf\{H\}\(t\)is a lower\-triangular matrix:Hi,i​\(t\)=2​Vi,i​\(t\)Li,i​\(t\)H\_\{i,i\}\(t\)=\\frac\{2V\_\{i,i\}\(t\)\}\{L\_\{i,i\}\(t\)\},Hi,j​\(t\)=Vi,j​\(t\)\+Vj,i​\(t\)H\_\{i,j\}\(t\)=V\_\{i,j\}\(t\)\+V\_\{j,i\}\(t\)\(i\>ji\>j\), andHi,j​\(t\)=0H\_\{i,j\}\(t\)=0\(i<ji<j\)\.

The differential of the Cholesky factor satisfiesd​𝐋​\(t\)=𝐋⁡\(t\)​Ψ​\(𝐋​\(t\)−1​d​𝐗4​\(t\)​𝐋​\(t\)−⁣⊤\)d\\mathbf\{L\}\(t\)=\\mathbf\{L\}\(t\)\\Psi\(\\mathbf\{L\}\(t\)^\{\-1\}d\\mathbf\{X\}\_\{4\}\(t\)\\mathbf\{L\}\(t\)^\{\-\\top\}\), whereΨ\\Psiis the lower\-symmetrizer \(keep strictly lower part and take half of the diagonal; zero upper part\)\. SinceΨ\\Psiis self\-adjoint, the adjoint mapping gives

⟨𝐇⁡\(t\),d​𝐋​\(t\)⟩F\\displaystyle\\langle\\mathbf\{H\}\(t\),d\\mathbf\{L\}\(t\)\\rangle\_\{F\}=⟨𝐋​\(t\)⊤​𝐇​\(t\),Ψ⁡\(𝐋​\(t\)−1​d​𝐗4​\(t\)​𝐋​\(t\)−⁣⊤\)⟩F\\displaystyle=\\langle\\mathbf\{L\}\(t\)^\{\\top\}\\mathbf\{H\}\(t\),\\Psi\(\\mathbf\{L\}\(t\)^\{\-1\}d\\mathbf\{X\}\_\{4\}\(t\)\\mathbf\{L\}\(t\)^\{\-\\top\}\)\\rangle\_\{F\}=⟨Ψ⁡\(𝐋​\(t\)⊤​𝐇​\(t\)\),𝐋​\(t\)−1​d​𝐗4​\(t\)​𝐋​\(t\)−⁣⊤⟩F\\displaystyle=\\langle\\Psi\(\\mathbf\{L\}\(t\)^\{\\top\}\\mathbf\{H\}\(t\)\),\\mathbf\{L\}\(t\)^\{\-1\}d\\mathbf\{X\}\_\{4\}\(t\)\\mathbf\{L\}\(t\)^\{\-\\top\}\\rangle\_\{F\}=⟨𝐋​\(t\)−⁣⊤​Ψ​\(𝐋​\(t\)⊤​𝐇​\(t\)\)⊤​𝐋​\(t\)−1,d​𝐗4​\(t\)⟩F\\displaystyle=\\langle\\mathbf\{L\}\(t\)^\{\-\\top\}\\Psi\(\\mathbf\{L\}\(t\)^\{\\top\}\\mathbf\{H\}\(t\)\)^\{\\top\}\\mathbf\{L\}\(t\)^\{\-1\},d\\mathbf\{X\}\_\{4\}\(t\)\\rangle\_\{F\}Becaused​𝐗4​\(t\)d\\mathbf\{X\}\_\{4\}\(t\)is symmetric, we have⟨𝐇⁡\(t\),d​𝐋​\(t\)⟩F=⟨𝐔⁡\(t\),d​𝐗4​\(t\)⟩F\\langle\\mathbf\{H\}\(t\),d\\mathbf\{L\}\(t\)\\rangle\_\{F\}=\\langle\\mathbf\{U\}\(t\),d\\mathbf\{X\}\_\{4\}\(t\)\\rangle\_\{F\}, where𝐔⁡\(t\)=12​𝐋​\(t\)−⁣⊤​\(Ψ⁡\(𝐋​\(t\)⊤​𝐇​\(t\)\)\+Ψ​\(𝐋​\(t\)⊤​𝐇​\(t\)\)⊤\)​𝐋​\(t\)−1\\mathbf\{U\}\(t\)=\\frac\{1\}\{2\}\\mathbf\{L\}\(t\)^\{\-\\top\}\(\\Psi\(\\mathbf\{L\}\(t\)^\{\\top\}\\mathbf\{H\}\(t\)\)\+\\Psi\(\\mathbf\{L\}\(t\)^\{\\top\}\\mathbf\{H\}\(t\)\)^\{\\top\}\)\\mathbf\{L\}\(t\)^\{\-1\}\.

When taking directional derivative with respect to𝐖4\\mathbf\{W\}\_\{4\}, we have

d​f​\(t\)=⟨𝐔⁡\(t\),d​𝐗4​\(t\)⟩F\\displaystyle df\(t\)=\\langle\\mathbf\{U\}\(t\),d\\mathbf\{X\}\_\{4\}\(t\)\\rangle\_\{F\}=⟨𝐔⁡\(t\),\(d​𝐖4\)​𝐗3​\(t\)​𝐖4⊤\+𝐖4​𝐗3​\(t\)​\(d​𝐖4\)⊤⟩F\\displaystyle=\\langle\\mathbf\{U\}\(t\),\(d\\mathbf\{W\}\_\{4\}\)\\mathbf\{X\}\_\{3\}\(t\)\\mathbf\{W\}\_\{4\}^\{\\top\}\+\\mathbf\{W\}\_\{4\}\\mathbf\{X\}\_\{3\}\(t\)\(d\\mathbf\{W\}\_\{4\}\)^\{\\top\}\\rangle\_\{F\}=⟨2​𝐔​\(t\)​𝐖4​𝐗3​\(t\),d​𝐖4⟩F\.\\displaystyle=\\langle 2\\mathbf\{U\}\(t\)\\mathbf\{W\}\_\{4\}\\mathbf\{X\}\_\{3\}\(t\),d\\mathbf\{W\}\_\{4\}\\rangle\_\{F\}\.Therefore, theEuclideangradient w\.r\.t\.𝐖4\\mathbf\{W\}\_\{4\}is

∇𝐖4f​\(t\)=2​𝐔​\(t\)​𝐖4​𝐗3​\(t\)\.\\nabla\_\{\\mathbf\{W\}\_\{4\}\}f\(t\)=2\\mathbf\{U\}\(t\)\\mathbf\{W\}\_\{4\}\\mathbf\{X\}\_\{3\}\(t\)\.\(16\)When taking directional derivative with respect to𝐗3​\(t\)\\mathbf\{X\}\_\{3\}\(t\), we have

d​f​\(t\)=⟨𝐔⁡\(t\),d​𝐗4​\(t\)⟩F=⟨𝐔⁡\(t\),𝐖4​d​𝐗3​\(t\)​𝐖4⊤⟩F=⟨𝐖4⊤​𝐔​\(t\)​𝐖4,d​𝐗3​\(t\)⟩F\.\\displaystyle df\(t\)=\\langle\\mathbf\{U\}\(t\),d\\mathbf\{X\}\_\{4\}\(t\)\\rangle\_\{F\}=\\langle\\mathbf\{U\}\(t\),\\mathbf\{W\}\_\{4\}d\\mathbf\{X\}\_\{3\}\(t\)\\mathbf\{W\}\_\{4\}^\{\\top\}\\rangle\_\{F\}=\\langle\\mathbf\{W\}\_\{4\}^\{\\top\}\\mathbf\{U\}\(t\)\\mathbf\{W\}\_\{4\},d\\mathbf\{X\}\_\{3\}\(t\)\\rangle\_\{F\}\.Therefore, the Euclidean gradient w\.r\.t\.𝐗3​\(t\)\\mathbf\{X\}\_\{3\}\(t\)is

∇𝐗3​\(t\)f​\(t\)=𝐖4⊤​𝐔​\(t\)​𝐖4\.\\nabla\_\{\\mathbf\{X\}\_\{3\}\(t\)\}f\(t\)=\\mathbf\{W\}\_\{4\}^\{\\top\}\\mathbf\{U\}\(t\)\\mathbf\{W\}\_\{4\}\.\(17\)

### E\.2Activated\-pooling layer

For the geodesic\-shrinkage activation, we have

𝐗3​\(t\)=ϕ−1​\(αH​∑h=1Hϕ⁡\(𝐖1h​𝐗​\(t\)​\(𝐖1h\)⊤\)\),\\mathbf\{X\}\_\{3\}\(t\)=\\phi^\{\-1\}\(\\frac\{\\alpha\}\{H\}\\sum\_\{h=1\}^\{H\}\\phi\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(t\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)\),whereϕ\\phiis the principal matrix logarithm log for the LE metric, and the log\-Cholesky coordinate mapΦ\\Phifor the LC metric\. Let𝐀⁡\(t\)=αH​∑h=1Hϕ⁡\(𝐖1h​𝐗​\(t\)​\(𝐖1h\)⊤\)\\mathbf\{A\}\(t\)=\\frac\{\\alpha\}\{H\}\\sum\_\{h=1\}^\{H\}\\phi\(\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(t\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\)and𝐗1h​\(t\)=𝐖1h​𝐗​\(t\)​\(𝐖1h\)⊤\\mathbf\{X\}\_\{1\}^\{h\}\(t\)=\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(t\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\. By the chain rule, for anyhh, we have

\(D𝐖1h​𝐗3​\(t\)\)​\(Δ​𝐖\)=\(D𝐀⁡\(t\)​ϕ−1\)​\(αH​\(D𝐗1h​\(t\)​ϕ\)​\(Δ​𝐖𝐗​\(t\)​\(𝐖1h\)⊤\+𝐖1h​𝐗​\(t\)​Δ​𝐖⊤\)\)\.\(D\_\{\\mathbf\{W\}\_\{1\}^\{h\}\}\\mathbf\{X\}\_\{3\}\(t\)\)\(\\Delta\\mathbf\{W\}\)=\(D\_\{\\mathbf\{A\}\(t\)\}\\phi^\{\-1\}\)\(\\frac\{\\alpha\}\{H\}\(D\_\{\\mathbf\{X\}\_\{1\}^\{h\}\(t\)\}\\phi\)\(\\Delta\\mathbf\{W\}\\mathbf\{X\}\(t\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\+\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(t\)\\Delta\\mathbf\{W\}^\{\\top\}\)\)\.
#### LE metric

Under the LE metric, we have

\(D𝐖1h​𝐗3​\(t\)\)​\(Δ​𝐖\)=\(D𝐀⁡\(t\)​exp\)​\(αH​\(D𝐗1h​\(t\)​log\)​\(Δ​𝐖𝐗​\(t\)​\(𝐖1h\)⊤\+𝐖1h​𝐗​\(t\)​Δ​𝐖⊤\)\)\.\(D\_\{\\mathbf\{W\}\_\{1\}^\{h\}\}\\mathbf\{X\}\_\{3\}\(t\)\)\(\\Delta\\mathbf\{W\}\)=\(D\_\{\\mathbf\{A\}\(t\)\}\\exp\)\(\\frac\{\\alpha\}\{H\}\(D\_\{\\mathbf\{X\}\_\{1\}^\{h\}\(t\)\}\\log\)\(\\Delta\\mathbf\{W\}\\mathbf\{X\}\(t\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\+\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(t\)\\Delta\\mathbf\{W\}^\{\\top\}\)\)\.The spectral form for the Fréchet derivative\(D𝐗​log\)​\(𝐔\)\(D\_\{\\mathbf\{X\}\}\\log\)\(\\mathbf\{U\}\)is given in Eq\. equation[8](https://arxiv.org/html/2609.30487#A2.E8), and for\(D𝐔​exp\)​\(𝐊\)\(D\_\{\\mathbf\{U\}\}\\exp\)\(\\mathbf\{K\}\)given in Eq\. equation[9](https://arxiv.org/html/2609.30487#A2.E9)\.

Let𝐕⁡\(t\)\\mathbf\{V\}\(t\)be an “upstream gradient”, and writef⁡\(t\)=⟨𝐕⁡\(t\),𝐗3​\(t\)⟩Ff\(t\)=\\langle\\mathbf\{V\}\(t\),\\mathbf\{X\}\_\{3\}\(t\)\\rangle\_\{F\}\. The mapping\(D𝐔​exp\)​\(⋅\)\(D\_\{\\mathbf\{U\}\}\\exp\)\(\\cdot\)is self\-adjoint, and therefore we have

d​f​\(t\)=⟨𝐕⁡\(t\),d​𝐗3​\(t\)⟩F=⟨𝐕⁡\(t\),\(D𝐀⁡\(t\)​exp\)​\(d​𝐀​\(t\)\)⟩F=⟨\(D𝐀⁡\(t\)​exp\)​\(𝐕s​\(t\)\),d​𝐀​\(t\)⟩F,df\(t\)=\\langle\\mathbf\{V\}\(t\),d\\mathbf\{X\}\_\{3\}\(t\)\\rangle\_\{F\}=\\langle\\mathbf\{V\}\(t\),\(D\_\{\\mathbf\{A\}\(t\)\}\\exp\)\(d\\mathbf\{A\}\(t\)\)\\rangle\_\{F\}=\\langle\(D\_\{\\mathbf\{A\}\(t\)\}\\exp\)\(\\mathbf\{V\}\_\{s\}\(t\)\),d\\mathbf\{A\}\(t\)\\rangle\_\{F\},where again𝐕s​\(t\)=12​\(𝐕⁡\(t\)\+𝐕​\(t\)⊤\)\\mathbf\{V\}\_\{s\}\(t\)=\\frac\{1\}\{2\}\(\\mathbf\{V\}\(t\)\+\\mathbf\{V\}\(t\)^\{\\top\}\)\. Then we obtain

d​f​\(t\)=⟨\(D𝐀⁡\(t\)​exp\)​\(𝐕s​\(t\)\),αH​\(D𝐗1h​\(t\)​log\)​\(d​𝐗1h\)⟩F=αH​⟨\(D𝐗1h​\(t\)​log\)​\(\(D𝐀⁡\(t\)​exp\)​\(𝐕s​\(t\)\)\),d​𝐗1h⟩F\.\\displaystyle df\(t\)=\\langle\(D\_\{\\mathbf\{A\}\(t\)\}\\exp\)\(\\mathbf\{V\}\_\{s\}\(t\)\),\\frac\{\\alpha\}\{H\}\(D\_\{\\mathbf\{X\}\_\{1\}^\{h\}\(t\)\}\\log\)\(d\\mathbf\{X\}\_\{1\}^\{h\}\)\\rangle\_\{F\}=\\frac\{\\alpha\}\{H\}\\langle\(D\_\{\\mathbf\{X\}\_\{1\}^\{h\}\(t\)\}\\log\)\(\(D\_\{\\mathbf\{A\}\(t\)\}\\exp\)\(\\mathbf\{V\}\_\{s\}\(t\)\)\),d\\mathbf\{X\}\_\{1\}^\{h\}\\rangle\_\{F\}\.Finally, we obtaind​f​\(t\)=⟨2​αH​\(D𝐗1h​\(t\)​log\)​\(\(D𝐀⁡\(t\)​exp\)​\(𝐕s​\(t\)\)\)​𝐖1h​𝐗​\(t\),d​𝐖1h⟩Fdf\(t\)=\\langle\\frac\{2\\alpha\}\{H\}\(D\_\{\\mathbf\{X\}\_\{1\}^\{h\}\(t\)\}\\log\)\(\(D\_\{\\mathbf\{A\}\(t\)\}\\exp\)\(\\mathbf\{V\}\_\{s\}\(t\)\)\)\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(t\),d\\mathbf\{W\}\_\{1\}^\{h\}\\rangle\_\{F\}, and therefore theEuclideangradient w\.r\.t\.𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}is

∇𝐖1hf​\(t\)=2​αH​\(D𝐗1h​\(t\)​log\)​\(\(D𝐀⁡\(t\)​exp\)​\(𝐕s​\(t\)\)\)​𝐖1h​𝐗​\(t\)\.\\nabla\_\{\\mathbf\{W\}\_\{1\}^\{h\}\}f\(t\)=\\frac\{2\\alpha\}\{H\}\(D\_\{\\mathbf\{X\}\_\{1\}^\{h\}\(t\)\}\\log\)\(\(D\_\{\\mathbf\{A\}\(t\)\}\\exp\)\(\\mathbf\{V\}\_\{s\}\(t\)\)\)\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(t\)\.\(18\)

#### LC metric

Under the LC metric, for a perturbationΔ​𝐖\\Delta\\mathbf\{W\}in𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}, we haved​𝐗1h​\(t\)=Δ​𝐖𝐗​\(t\)​\(𝐖1h\)⊤\+𝐖1h​𝐗​\(t\)​Δ​𝐖⊤d\\mathbf\{X\}\_\{1\}^\{h\}\(t\)=\\Delta\\mathbf\{W\}\\mathbf\{X\}\(t\)\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\+\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(t\)\\Delta\\mathbf\{W\}^\{\\top\},d​𝐀​\(t\)=αH​∑h=1H\(D𝐗1h​\(t\)​Φ\)​\(d​𝐗1h​\(t\)\)d\\mathbf\{A\}\(t\)=\\frac\{\\alpha\}\{H\}\\sum\_\{h=1\}^\{H\}\(D\_\{\\mathbf\{X\}\_\{1\}^\{h\}\(t\)\}\\Phi\)\(d\\mathbf\{X\}\_\{1\}^\{h\}\(t\)\)andd​𝐗3​\(t\)=\(D𝐀⁡\(t\)​Φ−1\)​\(d​𝐀​\(t\)\)d\\mathbf\{X\}\_\{3\}\(t\)=\(D\_\{\\mathbf\{A\}\(t\)\}\\Phi^\{\-1\}\)\(d\\mathbf\{A\}\(t\)\)\. Let𝐕⁡\(t\)\\mathbf\{V\}\(t\)be an “upstream gradient”, and writef⁡\(t\)=⟨𝐕⁡\(t\),𝐗3​\(t\)⟩Ff\(t\)=\\langle\\mathbf\{V\}\(t\),\\mathbf\{X\}\_\{3\}\(t\)\\rangle\_\{F\}\. Let\(D𝐀⁡\(t\)​Φ−1\)∗\(D\_\{\\mathbf\{A\}\(t\)\}\\Phi^\{\-1\}\)^\{\*\}be the adjoint operator of\(D𝐀⁡\(t\)​Φ−1\)\(D\_\{\\mathbf\{A\}\(t\)\}\\Phi^\{\-1\}\), and\(D𝐗1h​\(t\)​Φ\)∗\(D\_\{\\mathbf\{X\}\_\{1\}^\{h\}\(t\)\}\\Phi\)^\{\*\}the adjoint operator of\(D𝐗1h​\(t\)​Φ\)\(D\_\{\\mathbf\{X\}\_\{1\}^\{h\}\(t\)\}\\Phi\)\. Then we have

d​f​\(t\)\\displaystyle df\(t\)=αH​⟨\(D𝐗1h​\(t\)​Φ\)∗​\(\(D𝐀⁡\(t\)​Φ−1\)∗​\(𝐕s​\(t\)\)\),d​𝐗1h⟩F\\displaystyle=\\frac\{\\alpha\}\{H\}\\langle\(D\_\{\\mathbf\{X\}\_\{1\}^\{h\}\(t\)\}\\Phi\)^\{\*\}\(\(D\_\{\\mathbf\{A\}\(t\)\}\\Phi^\{\-1\}\)^\{\*\}\(\\mathbf\{V\}\_\{s\}\(t\)\)\),d\\mathbf\{X\}\_\{1\}^\{h\}\\rangle\_\{F\}=⟨2​αH​\(D𝐗1h​\(t\)​Φ\)∗​\(\(D𝐀⁡\(t\)​Φ−1\)∗​\(𝐕s​\(t\)\)\)​𝐖1h​𝐗​\(t\),d​𝐖1h⟩F,\\displaystyle=\\langle\\frac\{2\\alpha\}\{H\}\(D\_\{\\mathbf\{X\}\_\{1\}^\{h\}\(t\)\}\\Phi\)^\{\*\}\(\(D\_\{\\mathbf\{A\}\(t\)\}\\Phi^\{\-1\}\)^\{\*\}\(\\mathbf\{V\}\_\{s\}\(t\)\)\)\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(t\),d\\mathbf\{W\}\_\{1\}^\{h\}\\rangle\_\{F\},where𝐕s​\(t\)=12​\(𝐕⁡\(t\)\+𝐕​\(t\)⊤\)\\mathbf\{V\}\_\{s\}\(t\)=\\frac\{1\}\{2\}\(\\mathbf\{V\}\(t\)\+\\mathbf\{V\}\(t\)^\{\\top\}\)\. Then theEuclideangradient w\.r\.t\.𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}is

∇𝐖1hf​\(t\)=2​αH​\(D𝐗1h​\(t\)​Φ\)∗​\(\(D𝐀⁡\(t\)​Φ−1\)∗​\(𝐕s​\(t\)\)\)​𝐖1h​𝐗​\(t\)\.\\nabla\_\{\\mathbf\{W\}\_\{1\}^\{h\}\}f\(t\)=\\frac\{2\\alpha\}\{H\}\(D\_\{\\mathbf\{X\}\_\{1\}^\{h\}\(t\)\}\\Phi\)^\{\*\}\(\(D\_\{\\mathbf\{A\}\(t\)\}\\Phi^\{\-1\}\)^\{\*\}\(\\mathbf\{V\}\_\{s\}\(t\)\)\)\\mathbf\{W\}\_\{1\}^\{h\}\\mathbf\{X\}\(t\)\.\(19\)We below give the explicit adjoints\(D𝐀⁡\(t\)​Φ−1\)∗\(D\_\{\\mathbf\{A\}\(t\)\}\\Phi^\{\-1\}\)^\{\*\}and\(D𝐗1h​\(t\)​Φ\)∗\(D\_\{\\mathbf\{X\}\_\{1\}^\{h\}\(t\)\}\\Phi\)^\{\*\}for the log\-Cholesky coordinates\.

Let𝐋⁡\(t\)=ℒ⁡\(𝐗3​\(t\)\)\\mathbf\{L\}\(t\)=\\mathcal\{L\}\(\\mathbf\{X\}\_\{3\}\(t\)\)denote the Cholesky factor of𝐗3​\(t\)\\mathbf\{X\}\_\{3\}\(t\)\. Define𝐁⁡\(t\)=𝐋​\(t\)⊤​𝐕s​\(t\)\\mathbf\{B\}\(t\)=\\mathbf\{L\}\(t\)^\{\\top\}\\mathbf\{V\}\_\{s\}\(t\); The adjoint\(D𝐀⁡\(t\)​Φ−1\)∗​\(𝐕s​\(t\)\)∈ℝm1​\(m1\+1\)/2\(D\_\{\\mathbf\{A\}\(t\)\}\\Phi^\{\-1\}\)^\{\*\}\(\\mathbf\{V\}\_\{s\}\(t\)\)\\in\\mathbb\{R\}^\{m\_\{1\}\(m\_\{1\}\+1\)/2\}is the vector of the lower\-diagonal entries of the following matrix:

2​diag​\(𝐋⁡\(t\)\)​diag​\(𝐁⁡\(t\)\)\+⌊𝐁⁡\(t\)⌋\+⌊𝐁​\(t\)⊤⌋\.2\\text\{diag\}\(\\mathbf\{L\}\(t\)\)\\text\{diag\}\(\\mathbf\{B\}\(t\)\)\+\\lfloor\\mathbf\{B\}\(t\)\\rfloor\+\\lfloor\\mathbf\{B\}\(t\)^\{\\top\}\\rfloor\.
Let𝐋h​\(t\)=ℒ⁡\(𝐗1h​\(t\)\)\\mathbf\{L\}^\{h\}\(t\)=\\mathcal\{L\}\(\\mathbf\{X\}\_\{1\}^\{h\}\(t\)\)denote the Cholesky factor of𝐗1h​\(t\)\\mathbf\{X\}\_\{1\}^\{h\}\(t\)forh=1,…,Hh=1,\\ldots,H\. Given a coordinate covector𝒈⁡\(t\)∈ℝm1​\(m1\+1\)/2\\boldsymbol\{g\}\(t\)\\in\\mathbb\{R\}^\{m\_\{1\}\(m\_\{1\}\+1\)/2\}, build a lower\-triangular𝐇h​\(t\)\\mathbf\{H\}^\{h\}\(t\)by

\[𝐇h​\(t\)\]i,i=gκ⁡\(i,i\)​\(t\)\[𝐋h​\(t\)\]i,i,\[𝐇h​\(t\)\]i,j=gκ⁡\(i,j\)​\(t\)​\(i\>j\),\[𝐇h​\(t\)\]i,j=0​\(i<j\),\[\\mathbf\{H\}^\{h\}\(t\)\]\_\{i,i\}=\\frac\{g\_\{\\kappa\(i,i\)\}\(t\)\}\{\[\\mathbf\{L\}^\{h\}\(t\)\]\_\{i,i\}\},~~\[\\mathbf\{H\}^\{h\}\(t\)\]\_\{i,j\}=g\_\{\\kappa\(i,j\)\}\(t\)~\(i\>j\),~~\[\\mathbf\{H\}^\{h\}\(t\)\]\_\{i,j\}=0~\(i<j\),whereκ\\kappais the indexing function\. LetΨ\\Psibe the lower\-symmetrizer\. Then

\(D𝐗1h​\(t\)​Φ\)∗​\(𝒈⁡\(t\)\)=12​𝐋h​\(t\)−⁣⊤​\(Ψ⁡\(𝐋h​\(t\)⊤​𝐇h​\(t\)\)\+Ψ​\(𝐋h​\(t\)⊤​𝐇h​\(t\)\)⊤\)​𝐋h​\(t\)−1\.\(D\_\{\\mathbf\{X\}\_\{1\}^\{h\}\(t\)\}\\Phi\)^\{\*\}\(\\boldsymbol\{g\}\(t\)\)=\\frac\{1\}\{2\}\\mathbf\{L\}^\{h\}\(t\)^\{\-\\top\}\(\\Psi\(\\mathbf\{L\}^\{h\}\(t\)^\{\\top\}\\mathbf\{H\}^\{h\}\(t\)\)\+\\Psi\(\\mathbf\{L\}^\{h\}\(t\)^\{\\top\}\\mathbf\{H\}^\{h\}\(t\)\)^\{\\top\}\)\\mathbf\{L\}^\{h\}\(t\)^\{\-1\}\.

## Appendix FCongruence matrix regularization

Without additional regularization on the congruence weight matrices\{𝐖1h\}h=1H\\\{\\mathbf\{W\}\_\{1\}^\{h\}\\\}\_\{h=1\}^\{H\}and\{𝐖4h\}h=1H\\\{\\mathbf\{W\}\_\{4\}^\{h\}\\\}\_\{h=1\}^\{H\}, different heads are likely to focus on the same dominant patterns in the functional data\. To encourage the heads to learn complementary features, we introduce a row\-space decorrelation penalty\.

For the matrices\{𝐖1h\}h=1H\\\{\\mathbf\{W\}\_\{1\}^\{h\}\\\}\_\{h=1\}^\{H\}, a natural choice is to penalize the overlap between their row spaces\. Specifically, for each pair1≤h<k≤H1\\leq h<k\\leq H, we consider

‖𝐖1h​\(𝐖1k\)⊤‖F2,\\\|\\mathbf\{W\}\_\{1\}^\{h\}\(\\mathbf\{W\}\_\{1\}^\{k\}\)^\{\\top\}\\\|\_\{F\}^\{2\},which quantifies the degree of alignment between the row spaces of𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}and𝐖1k\\mathbf\{W\}\_\{1\}^\{k\}\. Summing over all pairs gives the total pairwise\-overlap penalty:

ℒpair=∑1≤h<k≤H‖𝐖1h​\(𝐖1k\)⊤‖F2\.\\mathcal\{L\}\_\{\\mathrm\{pair\}\}=\\sum\_\{1\\leq h<k\\leq H\}\\\|\\mathbf\{W\}\_\{1\}^\{h\}\(\\mathbf\{W\}\_\{1\}^\{k\}\)^\{\\top\}\\\|\_\{F\}^\{2\}\.To reduce the computational cost, we exploit the fact that each matrix𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}is constrained to be full row rank with orthonormal rows, that is,𝐖1h​\(𝐖1h\)⊤=𝐈m1\.\\mathbf\{W\}\_\{1\}^\{h\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}=\\mathbf\{I\}\_\{m\_\{1\}\}\.For each headhh, we define

𝐏h=\(𝐖1h\)⊤​𝐖1h∈ℝm×m,\\mathbf\{P\}\_\{h\}=\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\\mathbf\{W\}\_\{1\}^\{h\}\\in\\mathbb\{R\}^\{m\\times m\},which is the orthogonal projector onto the row space of𝐖1h\\mathbf\{W\}\_\{1\}^\{h\}\. Using‖𝐖1h​\(𝐖1k\)⊤‖F2=tr⁡\(𝐏h​𝐏k\)\\\|\\mathbf\{W\}\_\{1\}^\{h\}\(\\mathbf\{W\}\_\{1\}^\{k\}\)^\{\\top\}\\\|\_\{F\}^\{2\}=\\operatorname\{tr\}\(\\mathbf\{P\}\_\{h\}\\mathbf\{P\}\_\{k\}\), we can rewrite the total pairwise\-overlap penalty as

ℒpair=∑1≤h<k≤Htr⁡\(𝐏h​𝐏k\)=12​\(‖∑h=1H𝐏h‖F2−∑h=1H‖𝐏h‖F2\)\.\\mathcal\{L\}\_\{\\mathrm\{pair\}\}=\\sum\_\{1\\leq h<k\\leq H\}\\operatorname\{tr\}\(\\mathbf\{P\}\_\{h\}\\mathbf\{P\}\_\{k\}\)=\\frac\{1\}\{2\}\\left\(\\\|\\sum\_\{h=1\}^\{H\}\\mathbf\{P\}\_\{h\}\\\|\_\{F\}^\{2\}\-\\sum\_\{h=1\}^\{H\}\\\|\\mathbf\{P\}\_\{h\}\\\|\_\{F\}^\{2\}\\right\)\.Moreover, for each projector𝐏h\\mathbf\{P\}\_\{h\},

‖𝐏h‖F2=tr⁡\(𝐏h2\)=tr⁡\(𝐏h\)=m1\.\\\|\\mathbf\{P\}\_\{h\}\\\|\_\{F\}^\{2\}=\\operatorname\{tr\}\(\\mathbf\{P\}\_\{h\}^\{2\}\)=\\operatorname\{tr\}\(\\mathbf\{P\}\_\{h\}\)=m\_\{1\}\.Hence,

ℒpair=12​\(‖∑h=1H𝐏h‖F2−H​m1\)\.\\mathcal\{L\}\_\{\\mathrm\{pair\}\}=\\frac\{1\}\{2\}\\left\(\\\|\\sum\_\{h=1\}^\{H\}\\mathbf\{P\}\_\{h\}\\\|\_\{F\}^\{2\}\-Hm\_\{1\}\\right\)\.Therefore, minimizingℒpair\\mathcal\{L\}\_\{\\mathrm\{pair\}\}is equivalent to minimizing‖∑h=1H𝐏h‖F2\\\|\\sum\_\{h=1\}^\{H\}\\mathbf\{P\}\_\{h\}\\\|\_\{F\}^\{2\}\.

WhenH​m1\>mHm\_\{1\}\>m, the row spaces of all heads cannot be mutually orthogonal\. In this overcomplete regime, a more appropriate objective is to distribute the row spaces as uniformly as possible, which corresponds to encouraging the sum of the projection matrices to be close to an isotropic target:

∑h=1H𝐏h≈H​m1m​𝐈m\.\\sum\_\{h=1\}^\{H\}\\mathbf\{P\}\_\{h\}\\approx\\frac\{Hm\_\{1\}\}\{m\}\\mathbf\{I\}\_\{m\}\.Accordingly, we define the following diversity regularizer:

ℒdiv=λ​‖∑h=1H\(𝐖1h\)⊤​𝐖1h−H​m1m​𝐈m‖F2,\\mathcal\{L\}\_\{\\mathrm\{div\}\}=\\lambda\\left\\\|\\sum\_\{h=1\}^\{H\}\(\\mathbf\{W\}\_\{1\}^\{h\}\)^\{\\top\}\\mathbf\{W\}\_\{1\}^\{h\}\-\\frac\{Hm\_\{1\}\}\{m\}\\mathbf\{I\}\_\{m\}\\right\\\|\_\{F\}^\{2\},whereλ\>0\\lambda\>0is a tuning parameter\. The same regularization is also applied to\{𝐖4h\}h=1H\\\{\\mathbf\{W\}\_\{4\}^\{h\}\\\}\_\{h=1\}^\{H\}:

ℒdiv=λ​‖∑h=1H\(𝐖4h\)⊤​𝐖4h−H​m2m1​𝐈m1‖F2\.\\mathcal\{L\}\_\{\\mathrm\{div\}\}=\\lambda\\left\\\|\\sum\_\{h=1\}^\{H\}\(\\mathbf\{W\}\_\{4\}^\{h\}\)^\{\\top\}\\mathbf\{W\}\_\{4\}^\{h\}\-\\frac\{Hm\_\{2\}\}\{m\_\{1\}\}\\mathbf\{I\}\_\{m\_\{1\}\}\\right\\\|\_\{F\}^\{2\}\.

## Appendix GSupplementary experimental details and results

### G\.1Details on the clustering benchmark methods and their configuration

TSRVFIn\[[Dai et al\., 2020](https://arxiv.org/html/2609.30487#bib.bib14)\], trajectory dissimilarity is measured in the Transported Square\-Root Vector Field \(TSRVF\) space, with temporal reparameterization used to factor out inter\-subject differences in execution rate\. The authors adopted a particular Riemannian metric on the SPD manifold, which provides closed\-form expressions for the key geometric operations they need \(e\.g\., parallel transport\)\. When the matrix dimensionmmis large \(e\.g\.,m\>100m\>100\), efficient comparison of covariance trajectories becomes computationally impractical\. Therefore,[Dai et al\. \[2020\]](https://arxiv.org/html/2609.30487#bib.bib14)apply dimension reduction before trajectory comparison\. We adopt the same approach here and set the reduced dimension to 8\. The other hyperparameter is the number of optimization iterations which is set to 30\.

GeoAttThe network feeds each SPD matrix𝐗⁡\(t\)\\mathbf\{X\}\(t\)into a manifold\-aware CNN to produce a lower\-dimensional SPD matrix𝐙⁡\(t\)\\mathbf\{Z\}\(t\), and the sequence\{𝐙⁡\(t\)\}\\\{\\mathbf\{Z\}\(t\)\\\}is then processed by a manifold\-aware GRU\. We take the hidden state from the GRU as the latent embedding of the whole trajectory\{𝐗⁡\(t\)\}\\\{\\mathbf\{X\}\(t\)\\\}\. The hyperparameters and their values are: epochs=100\{\}=100, latent dimension=8\{\}=8, Adam learning rate=0\.001\{\}=0\.001\(Euclidean parameters\), Stiefel manifold learning rate=0\.005\{\}=0\.005\(BiMap weights\), and batch size=16\{\}=16\.

SPDNetSPDNet learns a latent vector embedding𝐱⁡\(t\)\\mathbf\{x\}\(t\)for each input SPD matrix𝐗⁡\(t\)\\mathbf\{X\}\(t\)\. Given a sequence of SPD matrices\{𝐗⁡\(t1\),𝐗⁡\(t2\),𝐗⁡\(t3\),…\}\\\{\\mathbf\{X\}\(t\_\{1\}\),\\mathbf\{X\}\(t\_\{2\}\),\\mathbf\{X\}\(t\_\{3\}\),\\ldots\\\}, we average the corresponding latent embeddings\{𝐱⁡\(t1\),𝐱⁡\(t2\),𝐱⁡\(t3\),…\}\\\{\\mathbf\{x\}\(t\_\{1\}\),\\mathbf\{x\}\(t\_\{2\}\),\\mathbf\{x\}\(t\_\{3\}\),\\ldots\\\}, and use this mean vector as the latent representation of the entire SPD matrix\-valued function\. The hyperparameter settings are as follows: training epochs=100\{\}=100, latent dimension=8\{\}=8, Adam learning rate=0\.001\{\}=0\.001\(Euclidean parameters\), Stiefel manifold learning rate=0\.005\{\}=0\.005\(BiMap weights\), and batch size=16\{\}=16\. The intermediate BiMap layer dimensions are determined automatically bym1=max⁡\(4,⌊m/2⌋\)m\_\{1\}=\\max\(4,\\lfloor m/2\\rfloor\)andm2=max⁡\(3,⌊m1/2⌋\)m\_\{2\}=\\max\(3,\\lfloor m\_\{1\}/2\\rfloor\)\.

PGAThe authors in\[[Fletcher et al\., 2004](https://arxiv.org/html/2609.30487#bib.bib16)\]introduced principal geodesic analysis \(PGA\), which is essentially a nonlinear generalization of PCA to manifold\-valued data\. In practice, PGA is implemented by mapping the manifold\-valued data to the tangent space at the intrinsic mean via the Logarithm map and then performing standard PCA on the resulting tangent vectors\. In our approach, we perform standard functional PCA in the tangent space, and then apply k\-means on the principal component score vectors\. We adopt the log\-Euclidean metric and retain 10 principal components\.

PDMWe compute the pairwise distances between functions using the affine\-invariant metric and then apply k\-medoids on the resulting distance matrix\. The number of random initializations for the clustering method is 10\.

DCDC is the standard MLP autoencoder for deep clustering \(DC\)\. To adapt it to functional data, we make the input layer containm⁡\(m\+1\)/2m\(m\+1\)/2nodes, where each node receives a T\-dimensional input vector \(corresponding to oneXi​j​\(t\)X\_\{ij\}\(t\)observed over T time points\)\. The hyperparameter settings are as follows: training epochs=100\{\}=100, latent dimension=8\{\}=8, Adam learning rate=0\.001\{\}=0\.001, and batch size=16\{\}=16\. The encoder consists of 1D convolutional layers with channel sizestri→64→32\\mathrm\{tri\}\\to 64\\to 32, followed by a fully connected layer that produces the latent representation\.

MatFAEThe algorthm repository documentation provides a clear explanation of all hyperparameters and their default values\.

For all four deep learning models, including GeoAtt, SPDNet, DC, and MatFAE, we used Bayesian optimization implemented in Optuna for hyperparameter tuning\. In all the experiments below, MatFAE training and competitive methods computations were performed on a 2\.5 GHz Intel Xeon Gold 6548Y\+ CPU \(32 cores\) and an NVIDIA L40S GPU with 46 GB of VRAM\.

### G\.2Details on the classification benchmark methods and their configuration

Shared protocol for all SPD\-native methods \(SPD\-SRU, BW\-norm, MatFAE\-C\): internal 80/20 split with early stopping on validation balanced accuracy \(patience 25\), class\-weighted cross\-entropy loss, Adam optimizer on Euclidean weights combined with a manual Stiefel\-manifold step \(learning rate5×10−35\\times 10^\{\-3\}\) on BiMap/orthogonal weights, gradient clipping at 5\.0, and an SPD floor ofε=10−4\\varepsilon=10^\{\-4\}\. STAGIN, which does not operate on the SPD manifold, uses the same early\-stopping and loss protocol but standard Euclidean Adam throughout \(no Stiefel step or SPD floor\)\. All architecture designs and default hyperparameters \(learning rate, weight decay, batch size, dropout, epoch budget\) otherwise follow each method’s original publication, as detailed below\.

SPD\-SRUEach frame is reduced to an 8×\\times8 SPD matrix via BiMap, then passed through one SPD\-SRU recurrent cell with five time scalesα=\[0\.01,0\.25,0\.5,0\.9,0\.99\]\\alpha=\[0\.01,0\.25,0\.5,0\.9,0\.99\]\(multi\-scale log\-Euclidean\-weighted running SPD means, orthogonal recurrence, learnable scale/output/blend gates\), followed by LogEig→\\tovech→\\tohead\. Batch size 40, weight decay 0,≤\\leq200 epochs \(early\-stopped\), dropout 0\.1\.

STAGINPer\-frame correlation graphs are thresholded to a top\-30%\\%binary adjacency and passed through a 4\-layer GIN \(sum aggregation, hidden dimension 128\) with per\-layer SERO readouts summed, sinusoidal temporal positional encoding, one\-head temporal self\-attention, and a summed CLS token over time; a linear head produces the final prediction\. Learning rate5×10−45\\times 10^\{\-4\}, weight decay10−510^\{\-5\}, batch size 3, 100 epochs, dropout 0\.5\.

BW\-normStandard SPDNet\-style BiMap layer→\\toBW batch normalization \(running BW barycenter, momentum 0\.1,ε=10−5\\varepsilon=10^\{\-5\}, geodesic running mean; centering in the BW log\-chart, scaling by BW variance, with a learnable SPD bias\)→\\toReEig→\\toLogEig→\\tolinear head; per\-frame outputs are averaged before the final head\. Learning rate2\.5×10−32\.5\\times 10^\{\-3\}, weight decay5×10−25\\times 10^\{\-2\}, batch size 30,≤\\leq200 epochs, dropout 0\.1\.

MatFAE\-CWe retain theMatFAEencoder backbone, remove the decoder, and append a linear classification head to the latent representation𝒙=ℰ⁡\(ℱ⁡\(𝐗\)\)\\boldsymbol\{x\}=\\mathcal\{E\}\(\\mathcal\{F\}\(\\mathbf\{X\}\)\)\. The objective combines class\-weighted cross\-entropy with the same regularization terms used in the unsupervised model, namely congruence matrix regularization, the orthogonality penalty on functional weights, and the roughness penalty on functional weights; class weights are computed within each training fold usingsklearn\.utils\.class\_weight\.compute\_class\_weight\. Evaluation uses a class\-stratified outer five\-fold split\. Euclidean parameters use Adam with learning rate10−310^\{\-3\}and weight decay10−510^\{\-5\}, batch size1616, latent dimension1616, and dropout0\.10\.1before the classification head\. Unless otherwise stated, all backbone hyperparameters\(H,m1,m2,p1,p2,α,αmin,basis​K\)\(H,m\_\{1\},m\_\{2\},p\_\{1\},p\_\{2\},\\alpha,\\alpha\_\{\\min\},\\mathrm\{basis~\}K\), encoder and MLP widths, and loss weights are set to their default values\.

### G\.3Details on the simulation scenarios and additional experimental results

Each synthetic dataset consists of 100 functional trajectories, each defined on the manifold𝒮\+48\\mathcal\{S\}^\{48\}\_\{\+\}\. The matrix dimension is determined by the 48 cortical regions in the Harvard\-Oxford atlas\. Below, letqqdenote the number of time points at which the function𝐗⁡\(t\)\\mathbf\{X\}\(t\)is evaluated over the interval\[t0,t1\]\[t\_\{0\},t\_\{1\}\], andkkthe true number of clusters\.

The benchmark is organized as an ablation ladder, where each synthetic dataset isolates a single type of structural signal\. Consequently, failure on a given rung suggests that the method does not possess the feature\-extraction mechanism required to capture that specific structure\.

A\(q=40,k=2\)\(q=40,k=2\): The data generation process is as follows\. First, we define two well\-separated SPD states, A and B\. For Group 0, each trajectory starts in state A and switches to state B halfway through, producing an A→\\toB pattern\. For Group 1, each trajectory starts in state B and switches to state A halfway through, producing a B→\\toA pattern\. To introduce between\-subject variability, we add random jitter of up to±3\\pm 3windows to the transition point, together with tangent\-space noise\.*Rung: temporal ordering\.*

B\(q=40,k=2\)\(q=40,k=2\): We define three SPD states, A, B, and C\. Every trajectory begins and ends in state A\. The two groups differ in the intermediate state they visit: Group 0 passes through state B, while Group 1 passes through state C\.*Rung: static identity\.*

C\(q=120,k=3\)\(q=120,k=3\): We follow the HMM\-based dFC construction of\[[Vidaurre et al\., 2017](https://arxiv.org/html/2609.30487#bib.bib3)\]\. In particular, we consider three brain states and assign each group a different Markov transition matrix, such that the corresponding stationary dwell\-time ratios are 75/15/10, 50/30/20, and 25/35/40, respectively\.*Rung: dwell time\.*

D\(q=60,k=2\)\(q=60,k=2\): Both groups switch between the same two states, A and B, following piecewise\-constant trajectories\. Group 0 has a switching period of approximatelyq/4q/4, while Group 1 has a shorter switching period of approximatelyq/12q/12\. As a result, both groups share the same visited\-state set and stationary proportions, but differ in how frequently they transition between states\.*Rung: transition frequency\.*

E\(q=30,k=2\)\(q=30,k=2\): Each trajectory is generated by sampling a symmetric\-matrix\-valued Gaussian process in the tangent space with squared\-exponential kernel

κℓ​\(t,t′\)=exp⁡\(−\(t−t′\)22​ℓ2\),\\kappa\_\{\\ell\}\(t,t^\{\\prime\}\)=\\exp\\\!\\left\(\-\\frac\{\(t\-t^\{\\prime\}\)^\{2\}\}\{2\\ell^\{2\}\}\\right\),and then mapping it to the SPD manifold through𝐗⁡\(t\)=exp⁡\(𝐘⁡\(t\)\)\\mathbf\{X\}\(t\)=\\exp\(\\mathbf\{Y\}\(t\)\)\. Group 0 usesℓ=12\\ell=12, while Group 1 usesℓ=2\\ell=2\. The cluster means are asymptotically the same, so the groups are distinguished only by trajectory smoothness\.*Rung: smoothness\.*

F\(q=20,k=3\)\(q=20,k=3\)For each time point, we sample𝐗⁡\(t\)∼𝒲m​\(𝐕,vc\)/vc\\mathbf\{X\}\(t\)\\sim\\mathcal\{W\}\_\{m\}\(\\mathbf\{V\},v\_\{c\}\)/v\_\{c\}, with common scale matrix𝐕=𝐈\\mathbf\{V\}=\\mathbf\{I\}and group\-specific degrees of freedomvc∈\{150,75,50\}v\_\{c\}\\in\\\{150,75,50\\\}\. Because the mean is the same for all groups, the groups differ only by their concentration, as reflected in quantities such as the log\-determinant and trace\.*Rung: Wishart concentration\.*

G\(q=30,k=2\)\(q=30,k=2\): We generate group\-specific neural source signals and convolve them with the canonical double\-gamma hemodynamic response function \(HRF\)\[[Glover, 1999](https://arxiv.org/html/2609.30487#bib.bib8)\]\. We then construct the functional connectivity trajectory𝐗⁡\(t\)\\mathbf\{X\}\(t\)by applying sliding\-window Pearson correlation to the simulated signals\[[Allen et al\., 2014](https://arxiv.org/html/2609.30487#bib.bib7)\]\.*Rung: real HRF\-blurred functional connectivity\.*

H\(q=30,k=2\)\(q=30,k=2\): Following\[[Arsigny et al\., 2007](https://arxiv.org/html/2609.30487#bib.bib6),[Pennec et al\., 2006](https://arxiv.org/html/2609.30487#bib.bib9)\], we consider two groups that share the same endpoints,𝐗0\\mathbf\{X\}\_\{0\}and𝐗1\\mathbf\{X\}\_\{1\}\. The geodesic connecting these two points on the SPD manifold is given by

γ\(t\)=𝐗01/2exp\(tlog\(𝐗0−1/2𝐗1𝐗0−1/2\)\)𝐗01/2\.\\gamma\(t\)=\\mathbf\{X\}\_\{0\}^\{1/2\}\\exp\(t\\log\(\\mathbf\{X\}\_\{0\}^\{\-1/2\}\\mathbf\{X\}\_\{1\}\\mathbf\{X\}\_\{0\}^\{\-1/2\}\)\)\\mathbf\{X\}\_\{0\}^\{1/2\}\.For Group 0, the trajectory follows the geodesic under the identity warp:τ↦τ\\tau\\mapsto\\tau\. For Group 1, the trajectory follows the same geodesic under a triangle\-bounce warp:τ↦2​τ\\tau\\mapsto 2\\tauforτ≤1/2\\tau\\leq 1/2, andτ↦2​\(1−τ\)\\tau\\mapsto 2\(1\-\\tau\)forτ\>1/2\\tau\>1/2\. Since the two warps have the same mean \(1/21/2\) and variance \(1/121/12\), the Jensen gap is zero\. Thus, the groups are distinguished only by temporal progression along the geodesic\.*Rung: trajectory direction\.*

I\(q=60,k=2\)\(q=60,k=2\): We generate the data using a two\-level hierarchical HMM\[[Fine et al\., 1998](https://arxiv.org/html/2609.30487#bib.bib4)\]\. The outer state, taking values in\{O1,O2\}\\\{O\_\{1\},O\_\{2\}\\\}, evolves slowly over time, while the inner state, taking values in\{I1,I2,I3\}\\\{I\_\{1\},I\_\{2\},I\_\{3\}\\\}, evolves conditional on the current outer state\. The two groups differ in the coupling structure between the outer and inner states\.*Rung: multi\-scale dynamics\.*

J\(q=50,k=2\)\(q=50,k=2\):The two groups are anchored to meta\-analytic brain maps from the Neurosynth database\[[Yarkoni et al\., 2011](https://arxiv.org/html/2609.30487#bib.bib51)\], accessed throughNiMARE\[[Salo et al\., 2023](https://arxiv.org/html/2609.30487#bib.bib52)\]\. We selected six cognitive terms and divided them into two concept families: an attentional\-control family \(*attention*,*executive*,*inhibition*\) and a memory family \(*memory*,*recall*,*encoding*\)\. For each termtt, we identified all Neurosynth studies whose abstract\-level TF\-IDF loading for that term exceeded10−310^\{\-3\}\. We then computed a multilevel kernel density analysis map by placing a 10 mm spherical kernel at each reported peak coordinate and summing the resulting kernels across studies\. Each term\-specific density map was parcellated into the 48 Harvard\-Oxford cortical regions, producing a regional score vector𝐬⁡\(t\)∈ℝ48\\mathbf\{s\}\(t\)\\in\\mathbb\{R\}^\{48\}\. This vector was mean\-centred and unit\-normalised to obtain𝐮⁡\(t\)\\mathbf\{u\}\(t\), and then converted into an SPD anchor matrix using𝐗⁡\(t\)=ϵ​𝐈\+𝐮⁡\(t\)​𝐮​\(t\)⊤,\\mathbf\{X\}\(t\)=\\epsilon\\mathbf\{I\}\+\\mathbf\{u\}\(t\)\\mathbf\{u\}\(t\)^\{\\top\},whereϵ\\epsilonwas chosen so that𝐗⁡\(t\)\\mathbf\{X\}\(t\)had condition number 7\.5\. For each synthetic trajectory, the sequence was divided into three approximately equal temporal segments\. Each group cycled through the three Neurosynth\-derived anchors from its own concept family, with a random±3\\pm 3\-window jitter applied to each segment boundary\. Additional log\-Euclidean tangent\-space perturbations were added at both the subject and window levels\.

Table 4:ARI scores for the ten simulation scenarios\. The table reports the mean \(top row\) and standard deviation \(bottom row\) of the scores over 100 repetitions\.Table[4](https://arxiv.org/html/2609.30487#A7.T4)shows a pattern broadly consistent with the ARI results\. MatFAE achieves the highest mean ARI in seven scenarios \(A, C, D, E, F, H and J\), with especially large advantages in the more challenging settings C and E\. The gain in the smoothness scenario E is particularly striking: MatFAE reaches 0\.808, whereas all competing methods remain near zero\. It also performs best in F and ties for perfect recovery in H and J\.

The baseline methods again show more specialized behavior\. TSRVF and DC perform very well in the temporally driven scenarios A and H, while PGA is strongest in the static identity and real\-HRF scenarios B, J and G\. Scenario I remains the most difficult overall, with all methods achieving low ARI values; unlike the AMI results, TSRVF attains the highest mean score there, although the margin over MatFAE is small\. Overall, the ARI results reinforce the conclusion that MatFAE is the most robust method across the benchmark, with its clearest advantage appearing in the harder scenarios\.

### G\.4Details on the real datasets and additional experimental results

All six resting\-state fMRI cohorts were obtained from publicly available preprocessed releases\. We retained the preprocessing provided by each consortium, including motion correction, T1 coregistration, spatial normalization to MNI space, and, where applicable, upstream confound regression\. To improve comparability across datasets, we then applied a common post\-hoc processing pipeline where the required data were available\.

Unless stated otherwise, the post\-hoc pipeline included nuisance regression using the full Friston\-24 motion model, 5\-6 aCompCor components, and volume\-level scrubbing for frames with framewise displacement greater than0\.5​mm0\.5\\,\\mathrm\{mm\}\. Scrubbing was implemented either as one\-hot spike regressors or through sample\-mask exclusion, depending on the confound information provided with each release\. We then applied temporal band\-pass filtering at0\.010\.01\-0\.1​Hz0\.1\\,\\mathrm\{Hz\}, linear and quadratic detrending, and per\-ROIzz\-scoring\.

When native NIfTI images were available, we parcellated the data using the 100\-parcel, 7\-network Schaefer atlas in MNI2​mm2\\,\\mathrm\{mm\}space\. When only pre\-extracted ROI time series were provided, we used the atlas distributed with the corresponding release\. Time\-varying functional connectivity was estimated using Ledoit\-Wolf shrinkage covariance matrices computed within30​s30\\,\\mathrm\{s\}sliding windows with a4​s4\\,\\mathrm\{s\}step\. This produced one SPD\-valued covariance trajectory for each subject\.

Dataset\-specific deviations from this common pipeline are described below\. Here,nndenotes the number of retained subjects,qqthe number of sliding windows,mmthe final ROI dimension after applying a common non\-zero\-variance mask, andkkthe number of diagnostic classes\.

CNP\(n=257,q=57,m=100,k=4\)\(n=257,\\ q=57,\\ m=100,\\ k=4\)\[[Bilder et al\., 2020](https://arxiv.org/html/2609.30487#bib.bib46)\]: The Consortium for Neuropsychiatric Phenomics \(CNP\) dataset was obtained from OpenNeuro under accession numberds000030\. The cohort comprises healthy controls\(n=120\)\(n=120\), individuals with schizophrenia\(n=48\)\(n=48\), bipolar disorder\(n=49\)\(n=49\), and ADHD\(n=40\)\(n=40\)\. Each participant has one resting\-state fMRI run with TR=2​s=2\\,\\mathrm\{s\}and 152 volumes\. We used the fMRIPrep\-v0\.4\.4 preprocessed release\[[Gorgolewski et al\., 2017](https://arxiv.org/html/2609.30487#bib.bib47)\], which provides motion\-corrected BOLD images spatially normalized to the MNI152NLin2009cAsym template \(the asymmetric 2009c nonlinear MNI152 standard space\) at 2 mm resolution\. The release also includes a 24\-column confound TSV file containing 6 motion parameters, 6 aCompCor components, 6 tCompCor components, white\-matter signal, global signal, framewise displacement, and DVARS\.

We applied a standard preprocessing pipeline consisting of Friston\-24 motion regression, 6 aCompCor components, FD\>0\.5​mm\>0\.5\\,\\mathrm\{mm\}spike regression, and temporal band\-pass filtering\. The data were then parcellated in MNI space using the 100\-parcel Schaefer atlas\. Ledoit\-Wolf covariance matrices were computed over 30 s sliding windows, corresponding to 15 TRs, with a step size of 2 TRs\. Four subjects were excluded because more than 50% of their volumes exceeded the FD threshold\.

COBRE\(n=144,q=68,m=100,k=2\)\(n=144,\\ q=68,\\ m=100,\\ k=2\): The COBRE \(RRID:SCR\_010482\) dataset was obtained from the International Neuroimaging Data\-sharing Initiative\. The cohort includes individuals with schizophrenia \(n=70n=70\) and healthy controls \(n=74n=74\), giving a total sample size of 144 participants\. Each resting\-state scan was acquired with TR=2​s=2\\,\\mathrm\{s\}and contains 150 volumes\. The data were preprocessed upstream by Bellec and colleagues using the NIAK pipeline \(v0\.17\)\. This preprocessing produced BOLD images in MNI152NLin2009a space at 6 mm isotropic resolution, together with per\-subject confound TSV files\. These confound files include 6 motion parameters, framewise displacement, a scrub flag, 6 slow\-drift discrete cosine transform basis functions, white\-matter and ventricle mean signals, and 5 aCompCor components\.

Importantly, these confound variables are provided with the data but were not regressed out in the released BOLD time series\. We therefore applied our common nuisance\-regression and filtering pipeline, including Friston\-24 motion regressors, 5 aCompCor components, white\-matter and ventricle signals, one\-hot spike regressors for volumes marked with scrub = 1, and temporal band\-pass filtering\. The resulting time series were parcellated in MNI space using the 100\-parcel Schaefer atlas\. Time\-varying functional connectivity was then estimated by computing Ledoit\-Wolf covariance matrices within 30 s sliding windows, corresponding to 15 TRs, with a step size of 2 TRs\.

ADHD\-200\(n=138,q=57,m=100,k=2\)\(n=138,\\ q=57,\\ m=100,\\ k=2\)\[[The ADHD\-200 Consortium, 2012](https://arxiv.org/html/2609.30487#bib.bib48),[Bellec et al\., 2017](https://arxiv.org/html/2609.30487#bib.bib49)\]: The ADHD\-200 dataset is a multi\-site cohort released by the ADHD\-200 Consortium and preprocessed by the Preprocessed Connectomes Project\. We used the CPAC\-preprocessed version distributed ats3://fcp\-indi/data/Projects/ADHD200/Outputs/cpac/, with the configurationpc10\.linear1\.wm0\.global0\.motion1\.quadratic1\.gm0\.compcor1\.csf0\. The cohort includes typically developing controls\(n=88\)\(n=88\)and individuals with ADHD\(n=50\)\(n=50\)\. Data were acquired across 10 sites, with site\-specific TR values of1\.51\.5,1\.961\.96,2\.02\.0, or2\.5​s2\.5\\,\\mathrm\{s\}\. In this preprocessing stream, 6 motion parameters, 5 aCompCor components, and linear and quadratic trends were already regressed from the BOLD time series\. ROI time series were then extracted using the 100\-parcel Schaefer atlas\.

Because only the extracted ROI time series are available, rather than the raw BOLD images or confound TSV files, our post\-hoc preprocessing was limited to temporal band\-pass filtering using the site\-specific TR\. We then estimated time\-varying functional connectivity by computing Ledoit\-Wolf covariance matrices within sliding windows\. Additional Friston\-24 motion regression or volume scrubbing could not be applied retrospectively\. Subjects were excluded if their scan duration satisfiedq⋅TR<240​sq\\cdot\\mathrm\{TR\}<240\\,\\mathrm\{s\}, or if more than 5% of parcels had zero variance\. This led to the exclusion of one outlier,sub\-0015038, and allowed the full 100\-ROI Schaefer atlas to be retained across the cohort\.

ABIDE\-I\(n=846,q=57,m=101,k=2\)\(n=846,\\ q=57,\\ m=101,\\ k=2\): The Autism Brain Imaging Data Exchange I \(ABIDE I\) dataset\[[Di Martino et al\., 2014](https://arxiv.org/html/2609.30487#bib.bib50)\]is a multi\-site resting\-state fMRI cohort comprising healthy controls\(n=455\)\(n=455\)and individuals with autism spectrum disorder \(ASD;n=391n=391\)\. Data were acquired across 17 sites, with TR values ranging from1\.51\.5to3\.0​s3\.0\\,\\mathrm\{s\}\. We used the CPAC\-preprocessednofilt\_noglobalrelease from the Preprocessed Connectomes Project, distributed ats3://fcp\-indi/data/Projects/ABID\_Initiative/Outputs/cpac/nofilt\_noglobal/\. In this preprocessing stream, 24 Friston motion parameters, 5 aCompCor components, and linear and quadratic trends were regressed from the BOLD data\. As indicated by thenofiltlabel, temporal band\-pass filtering was not applied upstream; volume scrubbing was also not performed\.

Because Schaefer\-atlas time series are not provided in the PCP release for ABIDE I, we used the shipped Harvard\-Oxford ROI time series directly\. Our post\-hoc preprocessing therefore consisted of temporal band\-pass filtering using the site\-specific TR, followed by estimation of time\-varying functional connectivity using Ledoit\-Wolf covariance matrices within sliding windows\. After applying a common non\-zero\-variance mask across all subjects, 101 of the 111 Harvard\-Oxford parcels were retained\.

TCP\(n=241,q=80,m=100,k=2\)\(n=241,\\ q=80,\\ m=100,\\ k=2\)\[[Chopra et al\., 2025](https://arxiv.org/html/2609.30487#bib.bib53)\]: The Transdiagnostic Connectome Project \(TCP\) dataset was obtained from OpenNeuro under accession numberds005237\. The cohort includes healthy general\-population controls\(n=92\)\(n=92\)and transdiagnostic patients\(n=149\)\(n=149\), aged 18\-70 years, recruited at two acquisition sites: Yale University and McLean Hospital\. Each participant has four resting\-state fMRI runs, acquired with anterior–posterior and posterior–anterior phase\-encoding directions \(task\-restAP\_run\-01/02andtask\-restPA\_run\-01/02\)\. Each run was acquired with TR=0\.8​s=0\.8\\,\\mathrm\{s\}, multiband factor 8, and 488 volumes\. The OpenNeuro release provides time series after the consortium’s HCP\-style minimal preprocessing pipeline, including motion and distortion correction, T1 coregistration, spatial normalization to MNI space, CIFTI grayordinate projection, ICA\-FIX automated noise classification and regression\[[Salimi\-Khorshidi et al\., 2014](https://arxiv.org/html/2609.30487#bib.bib55)\], global signal regression, and1/2000​Hz1/2000\\,\\mathrm\{Hz\}high\-pass detrending\. The released time series are parcellated using the 488\-region CAB\-NP atlas, comprising 360 Glasser cortical parcels\[[Glasser et al\., 2016](https://arxiv.org/html/2609.30487#bib.bib54)\]and 128 subcortical CIFTI grayordinate parcels\.

We retained the first anterior\-posterior resting\-state run\(task\-restAP\_run\-01\)\(\\texttt\{task\-restAP\\\_run\-01\}\)for each subject\. Because the released data were already denoised and provided as parcellated time series, our post\-hoc preprocessing was limited to temporal band\-pass filtering at0\.010\.01\-0\.1​Hz0\.1\\,\\mathrm\{Hz\}, followed by estimation of time\-varying functional connectivity using Ledoit\-Wolf covariance matrices\. Covariance matrices were computed within 30 s sliding windows, corresponding to 38 TRs, with a step size of 5 TRs\. To improve comparability with CNP, COBRE, and ADHD\-200, we mapped the 360 Glasser cortical parcels to the 100\-parcel Schaefer atlas using a row\-stochastic voxel\-overlap mapping computed in MNI152NLin2009cAsym2​mm2\\,\\mathrm\{mm\}space; the 128 subcortical parcels were not used\. Subjects were excluded if more than 50% of volumes exceeded the framewise\-displacement threshold of0\.5​mm0\.5\\,\\mathrm\{mm\}, based on the releasedmotion\_FD/files\.

CAT\-D\(n=118,q=49,m=100,k=2\)\(n=118,\\ q=49,\\ m=100,\\ k=2\)\[[Sadeghi et al\., 2022](https://arxiv.org/html/2609.30487#bib.bib56),[Camp et al\., 2024](https://arxiv.org/html/2609.30487#bib.bib57)\]: The NIMH Characterization and Treatment of Adolescent Depression \(CAT\-D\) dataset was obtained from OpenNeuro under accession numberds004627\. The cohort includes healthy volunteers\(n=50\)\(n=50\)and adolescents with major depressive disorder\(n=68\)\(n=68\), all scanned at a single site on a GE Discovery MR750 3T scanner\. Each participant has a multi\-echo resting\-state fMRI run with TR=2\.5​s=2\.5\\,\\mathrm\{s\}and three echoes, with TE approximately 14\.4, 28\.8, and 43\.2ms\\mathrm\{ms\}\. We retained only the second echo for parity with the other real datasets and restricted the analysis to the baseline session \(ses\-v1\), so that each subject contributed a single resting\-state trajectory\.

Unlike the other datasets, CAT\-D does not have a publicly available upstream\-preprocessed release\. We therefore preprocessed the data using fMRIPrep v23\.2\.0\[[Esteban et al\., 2019](https://arxiv.org/html/2609.30487#bib.bib58)\], which produced motion\-corrected BOLD images spatially normalized to the MNI152NLin2009cAsym template at2​mm2\\,\\mathrm\{mm\}resolution, together with per\-subject confound TSV files\. We then applied our post\-hoc preprocessing pipeline, including Friston\-24 motion regression, 6 anatomical aCompCor components, white\-matter and CSF mean signals, volume censoring based on FD\>0\.5​mm\>0\.5\\,\\mathrm\{mm\}and standardized DVARS\>1\.5\>1\.5, and temporal band\-pass filtering at0\.010\.01\-0\.1​Hz0\.1\\,\\mathrm\{Hz\}\. The resulting time series were parcellated in MNI space using the 100\-parcel Schaefer atlas\. Time\-varying functional connectivity was estimated by computing Ledoit\-Wolf covariance matrices within 30 s sliding windows, corresponding to 12 TRs, with a step size of 2 TRs\.

Table 5:The ARI scores for the six real datasets\.MatFAEachieves the highest ARI on all six cohorts, indicating a consistent advantage over the competing methods\. The absolute ARI values, however, remain modest, which reflects the difficulty of unsupervised diagnostic discovery from resting\-state fMRI\. Several patterns are worth noting\. First, the deep\-learning baselines, especially GeoAtt and SPDNet, perform close to chance on several cohorts, including CNP, ABIDE\-I, TCP, and CAT\-D\. This suggests that parameter\-heavy SPD encoders may be vulnerable to weak disease signal, limited sample size, site effects, and preprocessing variability\. Second, the strongest competing baselines differ across cohorts: TSRVF is the strongest baseline on CNP and COBRE, PDM and DC are strongest on ADHD\-200, and PGA is strongest on CAT\-D\. Nevertheless,MatFAEremains consistently ahead of these cohort\-specific competitors, with the largest margins observed on COBRE and CNP\. This suggests that the combination of geometric layers and functional feature extraction provides a more stable representation than relying only on temporal alignment, tangent\-space summaries, or static geometric structure\. Read together with the simulation study, the real\-data results support the same conclusion\. Individual baselines are competitive in specific regimes, such as TSRVF and DC for temporal ordering, PGA for static identity or HRF\-blurred connectivity, and GeoAtt for Wishart concentration\. By contrast,MatFAEperforms strongly across multiple regimes and is particularly effective on dwell time, transition frequency, and smoothness, which are relevant to resting\-state dynamic connectivity\. The real fMRI results therefore show thatMatFAEis more discriminative than the tested alternatives, while the low absolute ARI values also highlight that unsupervised recovery of DSM diagnostic labels from resting\-state dynamic connectivity remains a challenging open problem\.

### G\.5Statistical significance tests and ablation study

We test the statistical significance for the reported improvements on both simulated and real data\. For each synthetic rung, we test whether MatFAE \(clustering\) and MatFAE\-C \(classification\) perform significantly above chance\. Specifically, we apply a one\-sided one\-samplett\-test to the scores across random seeds, using a chance level of 0 for adjusted ARI and 0\.5 for balanced accuracy on the binary\-class rungs \(0\.33 for the three\-class rung\)\. We report both thepp\-value and effect size \(Cohen’sdd\)\. The clustering analysis uses 100 seeds per rung, while the classification analysis uses 20 seeds \(5×\\times3 stratified CV\)\. For the six real datasets, we apply the same procedure: For clustering, we set the number of clusters to the known number and test whether the mean ARI across 20 training seeds exceeds the chance level of 0\. For classification, we test whether the mean balanced accuracy across the 15 outer test folds of the5×35\\times 3stratified cross\-validation exceeds the chance level of1/K1/K, whereKKis the number of classes\. This corresponds to 0\.25 for the four\-class CNP dataset and 0\.5 for the remaining binary datasets\. The test results are given in Table[6](https://arxiv.org/html/2609.30487#A7.T6)\.

Table 6:Statistical significance of performance on the synthetic rungs and real datasets\.We conduct a component\-wise ablation, removing one module at a time\. Each model variant is evaluated on three representative simulation rungs \(E, G, and J\) and the mean ARI over 100 independent repetitions is reported\.

Table 7:Component\-wise ablation results for MatFAE on three simulation rungs\. Each row indicates a component removed or simplified relative to the full model\. Values are mean ARIs over 100 independent repetitions\.Across the three rungs, the most pronounced and recurring performance degradation occurs when the multi\-head BiMap is reduced to a single head, the matrix\-to\-vector moduleℱ\\mathcal\{F\}is replaced, or the functional layers are removed\. These results indicate that all three components are critical to the performance of MatFAE\.

We conduct a one\-factor\-at\-a\-time sensitivity analysis on the same three simulation rungs: E, G, and J\. Each hyperparameter is varied while all others are held fixed\. Table[8](https://arxiv.org/html/2609.30487#A7.T8)reports the mean ARI over 100 independent repetitions\.

Table 8:One\-factor\-at\-a\-time hyperparameter sensitivity analysis on three simulation rungs\. The number of heads \(HH\), latent dimension \(ss\), basis size \(KK\), intermediate SPD dimension \(m1m\_\{1\}\), and geodesic shrinkage parameter \(α\\alpha\) are varied individually while the remaining hyperparameters are held fixed\. Values are mean ARIs over 100 independent repetitions\.The results reveal the following trends:

Number of heads H:Performance generally improves as additional heads capture distinct localized temporal patterns, before plateauing once the relevant temporal regimes are adequately represented\.

Latent dimensionss:Performance is relatively stable over a moderate range of latent dimensions\. Small values can restrict representational capacity, whereas larger values provide no consistent benefit and may introduce redundant features\.

Basis sizeKK:Performance improves until the basis is sufficiently expressive, after which additional basis functions yield little or no benefit\. Thus,KKacts primarily as a minimum capacity requirement rather than following a larger\-is\-better relationship\.

Intermediate SPD dimensionm1m\_\{1\}:Moderate values generally provide sufficient representational capacity\. Increasingm1m\_\{1\}beyond this range produces only marginal improvements, indicating limited sensitivity once an adequate dimension is reached\.

Geodesic shrinkageα\\alpha:Intermediate values generally perform best\. Because the pooling operation already introduces nonlinearity, geodesic shrinkage primarily acts as a regularizer; excessive or insufficient shrinkage can reduce performance\.

### G\.6Interpretability ofMatFAE

Compared with existing complex deep learning models used in neuroscience,MatFAEcombines a lightweight architecture with intrinsic interpretability\. In particular, the temporal profile of the functional weights𝑾\(1\)​\(t\)\\boldsymbol\{W\}^\{\(1\)\}\(t\)in the encoder functional layer indicates which slices of the SPD trajectory contribute most strongly to the latent representations\. We illustrate this interpretability using the COBRE dataset by comparing statistics of𝒚⁡\(t\)\\boldsymbol\{y\}\(t\)with the temporal profile of𝑾\(1\)​\(t\)\\boldsymbol\{W\}^\{\(1\)\}\(t\)learned byMatFAEunder the configuration reported in Appendix[G\.4](https://arxiv.org/html/2609.30487#A7.SS4)\.

#### Profile and peak window

Letg⁡\(t\)=‖𝑾\(1\)​\(t\)‖Fg\(t\)=\\\|\\boldsymbol\{W\}^\{\(1\)\}\(t\)\\\|\_\{F\}denote the Frobenius norm of the functional weight at time indextt, and lete⁡\(t\)=∑k=1KBk​\(t\)2e\(t\)=\\sqrt\{\\sum\_\{k=1\}^\{K\}B\_\{k\}\(t\)^\{2\}\}denote the pointwiseℓ2\\ell\_\{2\}envelope of the B\-spline basis\. Because clamped open\-uniform B\-splines satisfye⁡\(0\)=e⁡\(T−1\)=1e\(0\)=e\(T\-1\)=1, while the interior values are below unity, the raw normg⁡\(t\)g\(t\)contains a deterministic boundary effect induced by the basis itself\. To isolate the data\-driven temporal preference of the encoder, we therefore use the basis\-corrected profileg~​\(t\)=g​\(t\)/e​\(t\)\\tilde\{g\}\(t\)=g\(t\)/e\(t\)\. The peak window\[t1,t2\]\[t\_\{1\},t\_\{2\}\]is selected as the contiguous interior interval of width⌊0\.20​T⌋\\lfloor 0\.20\\,T\\rfloorcentred atarg⁡maxt​g~​\(t\)\\arg\\max\_\{t\}\\tilde\{g\}\(t\), with two end\-windows excluded on each side\.

#### Cohort divergence in𝒚⁡\(t\)\\boldsymbol\{y\}\(t\)

For each subjectii, we compute𝒚i​\(t\)∈ℝp\\boldsymbol\{y\}\_\{i\}\(t\)\\in\\mathbb\{R\}^\{p\}and define two pointwise cohort\-divergence summaries: theℓ2\\ell\_\{2\}norm of the cohort\-mean difference,‖𝒚¯P​\(t\)−𝒚¯C​\(t\)‖2\\\|\\bar\{\\boldsymbol\{y\}\}\_\{P\}\(t\)\-\\bar\{\\boldsymbol\{y\}\}\_\{C\}\(t\)\\\|\_\{2\}, and the mean absolute Welch two\-samplett\-statistic across theppcoordinates,meanp​\|tWelch​\(t,p\)\|\\mathrm\{mean\}\_\{p\}\\,\|t\_\{\\mathrm\{Welch\}\}\(t,p\)\|\. These two quantities identify the time indices at which the patient and control cohorts are most separable in𝒚⁡\(t\)\\boldsymbol\{y\}\(t\)\.[Figure 3](https://arxiv.org/html/2609.30487#A7.F3)supports the design claim that𝑾\(1\)​\(t\)\\boldsymbol\{W\}^\{\(1\)\}\(t\)concentrates on the same temporal interval in which the geometry\-module output𝒚⁡\(t\)\\boldsymbol\{y\}\(t\)shows the strongest cohort separation\. Using the basis\-corrected profile is important because clamped B\-spline endpoints inflateg⁡\(t\)g\(t\)att∈\{0,T−1\}t\\in\\\{0,T\-1\\\}by a factor of approximately1/0\.681/0\.68for purely geometric reasons; these endpoint spikes therefore do not reflect learned temporal content\.

Figure 3:The encoder functional weight localises the cohort\-divergence window ofy⁡\(t\)\\boldsymbol\{y\}\(t\)\.\(A\)Cohort\-mean‖𝒚⁡\(t\)‖2\\\|\\boldsymbol\{y\}\(t\)\\\|\_\{2\}for healthy controls \(n=74n=74\) and schizophrenia patients \(n=70n=70\), with within\-cohort SEM bands\.\(B\)Two pointwise cohort\-divergence summaries on𝒚⁡\(t\)\\boldsymbol\{y\}\(t\), namely‖𝒚¯P​\(t\)−𝒚¯C​\(t\)‖2\\\|\\bar\{\\boldsymbol\{y\}\}\_\{P\}\(t\)\-\\bar\{\\boldsymbol\{y\}\}\_\{C\}\(t\)\\\|\_\{2\}andmeanp​\|tWelch​\(t,p\)\|\\mathrm\{mean\}\_\{p\}\\,\|t\_\{\\mathrm\{Welch\}\}\(t,p\)\|, both min\-max normalised\.\(C\)Basis\-corrected encoder weight profileg~​\(t\)=‖𝑾\(1\)​\(t\)‖F/∑kBk​\(t\)2\\tilde\{g\}\(t\)=\\\|\\boldsymbol\{W\}^\{\(1\)\}\(t\)\\\|\_\{F\}/\\sqrt\{\\sum\_\{k\}B\_\{k\}\(t\)^\{2\}\}, with the selected peak window shaded\.\(D\)Min\-max overlay of \(B\) and \(C\)\. The maximum ofg~\\tilde\{g\}att=35t=35lies55\-77windows away from the cohort\-divergence maxima att=28t=28formeanp​\|tWelch​\(t,p\)\|\\mathrm\{mean\}\_\{p\}\\,\|t\_\{\\mathrm\{Welch\}\}\(t,p\)\|andt=30t=30for‖𝒚¯P​\(t\)−𝒚¯C​\(t\)‖2\\\|\\bar\{\\boldsymbol\{y\}\}\_\{P\}\(t\)\-\\bar\{\\boldsymbol\{y\}\}\_\{C\}\(t\)\\\|\_\{2\}\. All three peaks fall within the encoder peak window\[t1,t2\]=\[28,41\]\[t\_\{1\},t\_\{2\}\]=\[28,41\]\.
#### Connectivity contrast restricted to the peak window

We next provide a static functional\-connectivity visualisation of the contrast thatMatFAEis expected to capture\.[Figure 4](https://arxiv.org/html/2609.30487#A7.F4)shows the edges for which the schizophrenia cohort has higher mean Fisher\-zzcorrelation than the healthy\-control cohort\. This visualisation identifies the spatial structure of the static signal and helps explain why COBRE is the cohort on whichMatFAEachieves its largest gain\. To construct this static visualisation, each subject’s SPD trajectory\{𝑿i​\(t\)\}t=1q\\\{\{\\bm\{\\mathsfit\{X\}\}\}\_\{i\}\(t\)\\\}\_\{t=1\}^\{q\}, withq=68q=68sliding\-window covariance matrices andm=100m=100Schaefer\-100 7\-network parcels, is collapsed to a single SPD matrix by the Log\-Euclidean Fréchet mean,

𝐗¯i=Exp𝐈LE​\(1q​∑t=1qLog𝐈LE​\(𝑿i​\(t\)\)\)=exp⁡\(1q​∑t=1qlog⁡\(𝑿i​\(t\)\)\)\.\\bar\{\\mathbf\{X\}\}\_\{i\}\\;=\\;\\mathrm\{Exp\}^\{\\mathrm\{LE\}\}\_\{\\mathbf\{I\}\}\\\!\\Big\(\\frac\{1\}\{q\}\\sum\_\{t=1\}^\{q\}\\mathrm\{Log\}^\{\\mathrm\{LE\}\}\_\{\\mathbf\{I\}\}\(\{\\bm\{\\mathsfit\{X\}\}\}\_\{i\}\(t\)\)\\Big\)\\;=\\;\\exp\\\!\\Big\(\\frac\{1\}\{q\}\\sum\_\{t=1\}^\{q\}\\log\\,\(\{\\bm\{\\mathsfit\{X\}\}\}\_\{i\}\(t\)\)\\Big\)\.\(20\)The static covariance matrix𝐗¯i\\bar\{\\mathbf\{X\}\}\_\{i\}is then converted to a correlation matrix𝐑i\\mathbf\{R\}\_\{i\}and Fisher\-zztransformed entry\-wise\. For each off\-diagonal edge\(j,k\)\(j,k\), a Welch two\-samplett\-statistic is computed between the patient \(n=70n=70\) and control \(n=74n=74\) cohorts using\{atanh⁡\(\[𝐑i\]j,k\)\}i\\\{\\mathrm\{atanh\}\(\[\\mathbf\{R\}\_\{i\}\]\_\{j,k\}\)\\\}\_\{i\}\. The top6060edges with the largest Welch\|t\|\|t\|in each direction are retained\. Edges are displayed in panel A on a multi\-view glass brain, using parcel centroids in MNI 2 mm space, and in panel B on a 100\-node chord plot grouped by Yeo\-7 network\. Within each hemisphere, networks are ordered asVis→SomMot→DorsAttn→SalVentAttn→Limbic→Cont→Default\\text\{Vis\}\\to\\text\{SomMot\}\\to\\text\{DorsAttn\}\\to\\text\{SalVentAttn\}\\to\\text\{Limbic\}\\to\\text\{Cont\}\\to\\text\{Default\}, so that both*Vis*arcs are placed at 12 o’clock and both*Default*arcs at 6 o’clock\. The right hemisphere fills the right half of the circle clockwise from 12 o’clock, while the left hemisphere fills the left half counter\-clockwise, mirroring the standard radiological axial view\. Edge thickness and opacity scale linearly across the top\-60 Welch\|t\|\|t\|range, with weaker edges drawn first so that the strongest edges appear on top\.

Figure 4:Increased connectivity in schizophrenia on COBRE\.Top\-60 edges by Welch\|t\|\|t\|where the schizophrenia cohort \(n=70n=70\) has higher Fisher\-zzcorrelation than the healthy\-control cohort \(n=74n=74\)\.\(A\)Multi\-view glass brain, including left lateral, right lateral, coronal, and axial views, with parcel centroids in MNI 2 mm space\. Nodes are coloured by Yeo\-7 network and sized by the total Welch\|t\|\|t\|at each node\.\(B\)Chord plot with all 100 Schaefer\-100 parcels on the rim, grouped by Yeo\-7 network\. The right hemisphere fills the right half of the circle clockwise from 12 o’clock, while the left hemisphere fills the left half counter\-clockwise\. Both Visual arcs are located at 12 o’clock and both Default\-Mode arcs at 6 o’clock, with a cyclic gap at the top and an internal hemisphere gap at the bottom\. Edge thickness and opacity scale linearly across the top\-60 Welch\|t\|\|t\|range, with the weakest edges drawn first so that the strongest edges appear on top\.To test whether the encoder\-identified peak window contains the strongest static functional\-connectivity contrast, we recompute the LE\-Fréchet mean in Eq\. equation[20](https://arxiv.org/html/2609.30487#A7.E20)using onlyt∈\[t1,t2\]=\[28,41\]t\\in\[t\_\{1\},t\_\{2\}\]=\[28,41\]\. We then repeat the Welch/Fisher\-zzedge\-level test and retain the top 60 edges by\|t\|\|t\|for which the patient cohort exhibits higher correlation than the control cohort\.[Figure 5](https://arxiv.org/html/2609.30487#A7.F5)shows that the canonical schizophrenia hyperconnectivity pattern is preserved when the average is restricted to the encoder\-flagged window\. The contrast remains dominated by connections among the Default\-Mode, Salience/Ventral\-Attention, and Frontoparietal\-Control networks, with retained\|t\|∈\[3\.9,5\.1\]\|t\|\\in\[3\.9,5\.1\], compared with\[4\.0,5\.4\]\[4\.0,5\.4\]under all\-window pooling\. The modest reduction in the upper\|t\|\|t\|bound is consistent with the peak\-window analysis using only14/68≈21%14/68\\approx 21\\%of the temporal data, while still preserving the main spatial structure of the contrast\.

Figure 5:Increased connectivity restricted to the encoder peak window\.Top\-60 edges by Welch\|t\|\|t\|where the schizophrenia cohort has higher Fisher\-zzcorrelation than the healthy\-control cohort, computed from per\-subject LE\-Fréchet means restricted tot∈\[28,41\]t\\in\[28,41\], the window highlighted by𝑾\(1\)​\(t\)\\boldsymbol\{W\}^\{\(1\)\}\(t\)in[Figure 3](https://arxiv.org/html/2609.30487#A7.F3)\. The same Default\-Mode / Salience\-Ventral\-Attention / Frontoparietal\-Control hyperconnectivity pattern is recovered, with retained\|t\|\|t\|values ranging from 3\.9 to 5\.1\. This suggests that the encoder\-highlighted window overlaps with the period in which the cohort\-discriminative static signal is most concentrated\.Taken together, Figures[3](https://arxiv.org/html/2609.30487#A7.F3)\-[5](https://arxiv.org/html/2609.30487#A7.F5)support the interpretability claim\. The functional weight profile𝑾\(1\)​\(t\)\\boldsymbol\{W\}^\{\(1\)\}\(t\)acts as an internally derived temporal saliency map, and its peak window aligns with the interval in which the underlying clinical contrast is most pronounced\.

## Appendix HLimitations

Individual\-level diagnosis and patient/control classification from resting\-state fMRI remain challenging because disease\-related signals are often weak, distributed, and strongly confounded by scanner/site differences, preprocessing choices, parcellation schemes, and the unknown mental state of participants during scanning\. Although we applied a common post\-hoc processing pipeline whenever possible, the real datasets were not fully homogeneous in preprocessing or parcellation\. Three cohorts used the Schaefer\-100 atlas, whereas ABIDE\-I used a 101\-region Harvard\-Oxford parcellation\. Therefore, performance differences across cohorts may partly reflect differences in ROI definition, rather than only differences in disease phenotype or model behaviour\.

Several additional sources of heterogeneity may affect the estimated dynamic functional connectivity trajectories\. ADHD\-200 and ABIDE\-I are multi\-site cohorts with TR values ranging from1\.51\.5to3\.0​s3\.0\\,\\mathrm\{s\}\. We used site\-specific band\-pass filtering, but did not explicitly harmonise scanner, acquisition, or site effects; residual site\-related variability may therefore remain\. Upstream preprocessing also differed across consortia, including fMRIPrep for CNP, NIAK for COBRE, and CPAC for ADHD\-200 and ABIDE\-I\. In particular, for ADHD\-200 and ABIDE\-I, only pre\-extracted ROI time series were available, and the raw BOLD images and confound files were not redistributed\. Consequently, retrospective Friston\-24 motion regression and frame\-wise scrubbing could not be applied, making nuisance\-control procedures less directly comparable across datasets\.

Finally, the two BiMap layers in the matrix\-to\-vector module containH⁡\(m1​m\+m2​m1\)\+1H\(m\_\{1\}m\+m\_\{2\}m\_\{1\}\)\+1trainable entries, independent of the time\-series length\. The geometry\-aware decoder introduces two additional matrices,𝐌1\\mathbf\{M\}\_\{1\}and𝐌2\\mathbf\{M\}\_\{2\}, containingp2​p1\+m​p2p\_\{2\}p\_\{1\}\+mp\_\{2\}trainable entries\. All functional weights in FAEclust are represented using basis expansions, so only their expansion coefficients are trainable\. The main parameter bottleneck is𝒲\(2\)​\(t\)\\mathcal\{W\}^\{\(2\)\}\(t\)in the decoder\. WithKKbasis functions, the matrix containsO⁡\(K​p24\)O\(Kp\_\{2\}^\{4\}\)trainable coefficients\. This limits scalability to very high\-dimensional parcellations, such as Schaefer\-400 or Schaefer\-1000, and motivates future work on more efficient geometric layers, low\-rank approximations, or sparse SPD representations\.

相似文章

SAEs 能否捕捉神经几何?(6分钟阅读)

TLDR AI

本文探讨了稀疏自动编码器(SAEs)如何捕捉弯曲的神经几何,揭示了SAE特征表示流形的三种不同方式,并提出了一个无监督流程来揭示神经表征中的几何结构。