Simultaneous inference of environmental and interaction forces in collective dynamics
Summary
This paper proposes a nonparametric variational learning framework to simultaneously infer interaction kernels and environmental forces in collective dynamics, validated on benchmark models with a model-selection procedure.
View Cached Full Text
Cached at: 08/27/26, 09:35 AM
# Simultaneous inference of environmental and interaction forces in collective dynamics Source: [https://arxiv.org/html/2608.25181](https://arxiv.org/html/2608.25181) ###### Abstract Collective dynamics arise in a wide range of physical, biological, and engineering applications\. Examples include cell migration, swarm robotics, social dynamics, and animal behavior\. A defining characteristic of these systems is the emergence of large\-scale coordination from local interactions among agents; a fundamental scientific question is thus to understand the local interactions that give rise to the observed emergent dynamics\. We are interested in methods for learning interactions generally, which can describe a wide class of physical systems exhibiting collective dynamics defined by an interaction kernel, without making any a priori assumptions regarding the analytical form of this kernel \(i\.e\. it is nonparametric\)\. The advantage of this kernel\-based approach is that it incorporates the underlying physics of the model \(i\.e\. collective dynamics\), which more general equation\-learning approaches may ignore, potentially limiting their effectiveness with regards to model accuracy and predictions\. In this work, we extend existing variational learning approaches to collective systems with both interaction kernels and environmental/intra\-agent forces\. The proposed framework simultaneously infers the interaction kernel non\-parametrically while learning the environmental force using either semi\-parametric or fully nonparametric representations\. The proposed methodology is validated on several benchmark models exhibiting synchronization, alignment, attraction\-repulsion, and external environmental forces\. In addition, we introduce a model\-selection procedure based on our nonparametric learning framework to identify models that optimally explain a given set of trajectory observations\. By exploiting the feature\-identification capability of the learned models, the proposed procedure can distinguish among different collective dynamics frameworks and recover mechanistic interaction mechanisms directly from trajectory data\. 11footnotetext:Department of Mathematics, Clarkson University, Potsdam, NY, United States22footnotetext:Department of Mathematics, University of Houston, Houston, TX, United States\*\*footnotetext:Corresponding author,[jgreene@clarkson\.edu](mailto:[email protected])## 1Introduction An important problem in the study of collective dynamics of interacting agents is how one may quantify and predict their behavior via mathematical models\. Such models may be constructed from first principles, utilizing \(for example\) Newton’s equations of motion and other well\-known physical theories, or from phenomenological considerations based on experimental observations; we note also that such models may be constructed by combining aspects of both approaches\. Many such models of collective dynamics exist in a wide variety of disciplines, including physics, biology, ecology, neurobiology, social sciences, and economics\[[1](https://arxiv.org/html/2608.25181#bib.bib7),[2](https://arxiv.org/html/2608.25181#bib.bib8),[3](https://arxiv.org/html/2608.25181#bib.bib5),[4](https://arxiv.org/html/2608.25181#bib.bib9)\]\. Furthermore, once a mathematical model is known and validated, the emergence of complex behavior may be analyzed and understood theoretically and precisely; again, such models provide mechanistic insight into the causality of emergence, which may be both scientifically and practically important\. Indeed, inter\-agent dynamics have had applications across biology, engineering, and the social sciences, providing mathematical frameworks for understanding collective cell migration, tissue morphogenesis and tumor invasion\[[5](https://arxiv.org/html/2608.25181#bib.bib20),[6](https://arxiv.org/html/2608.25181#bib.bib74)\], microbial swarming\[[7](https://arxiv.org/html/2608.25181#bib.bib17),[8](https://arxiv.org/html/2608.25181#bib.bib18),[9](https://arxiv.org/html/2608.25181#bib.bib19)\], animal flocking\[[10](https://arxiv.org/html/2608.25181#bib.bib11),[11](https://arxiv.org/html/2608.25181#bib.bib26),[12](https://arxiv.org/html/2608.25181#bib.bib27),[13](https://arxiv.org/html/2608.25181#bib.bib28),[14](https://arxiv.org/html/2608.25181#bib.bib12),[15](https://arxiv.org/html/2608.25181#bib.bib33),[16](https://arxiv.org/html/2608.25181#bib.bib31)\], pedestrian and traffic flow\[[17](https://arxiv.org/html/2608.25181#bib.bib24)\], opinion formation\[[18](https://arxiv.org/html/2608.25181#bib.bib10),[19](https://arxiv.org/html/2608.25181#bib.bib39)\], swarm robotics\[[20](https://arxiv.org/html/2608.25181#bib.bib21),[21](https://arxiv.org/html/2608.25181#bib.bib15)\], and the control of distributed autonomous systems\[[22](https://arxiv.org/html/2608.25181#bib.bib72),[23](https://arxiv.org/html/2608.25181#bib.bib73)\]\. Such mechanistic insight will thus be challenging without prior knowledge of the governing equations\. A learning\-based approach that first discovers the governing equations from observations of collective dynamics can address these challenges in predictive modeling, since predictions are then generated from the learned mechanistic model rather than from direct extrapolation of dynamical data\[[24](https://arxiv.org/html/2608.25181#bib.bib58)\], the latter of which generally fails in predictive capacity outside of training data\[[25](https://arxiv.org/html/2608.25181#bib.bib75),[26](https://arxiv.org/html/2608.25181#bib.bib1)\]\. Research on the data\-driven discovery of governing equations of dynamical systems can be traced back to the early works of Lagrange, Laplace, and Gauss\[[27](https://arxiv.org/html/2608.25181#bib.bib59)\]\. The work presented here focuses on discovering the governing equations of collective dynamics base on observed temporal agent trajectory data, which is often available with modern experimental techniques\[[17](https://arxiv.org/html/2608.25181#bib.bib24),[28](https://arxiv.org/html/2608.25181#bib.bib22),[29](https://arxiv.org/html/2608.25181#bib.bib65),[30](https://arxiv.org/html/2608.25181#bib.bib57),[31](https://arxiv.org/html/2608.25181#bib.bib76),[32](https://arxiv.org/html/2608.25181#bib.bib77)\]\. We note that all models considered in this work are defined via ordinary differential equations \(ODEs\), and as such take the general form x˙=F\(x\),x\(t0\)=x0∈ℝD,t∈\[t0,T\],\\displaystyle\\dot\{x\}=F\(x\),\\quad x\(t\_\{0\}\)=x\_\{0\}\\in\\mathbb\{R\}^\{D\},\\quad t\\in\[t\_\{0\},T\],\(1\)for someD∈ℕD\\in\\mathbb\{N\}\. For the general ODE system \([1](https://arxiv.org/html/2608.25181#S1.E1)\), the discovery of the governing equations is thus represented as follows: inferFFfrom observations ofx\(t\)x\(t\)at various timestt\. Many system identification methods have been developed to solve this problem, including classical regression\-inspired techniques\[[33](https://arxiv.org/html/2608.25181#bib.bib52),[34](https://arxiv.org/html/2608.25181#bib.bib53),[35](https://arxiv.org/html/2608.25181#bib.bib54)\], sparse\-identification methods\[[36](https://arxiv.org/html/2608.25181#bib.bib55),[26](https://arxiv.org/html/2608.25181#bib.bib1),[37](https://arxiv.org/html/2608.25181#bib.bib56),[38](https://arxiv.org/html/2608.25181#bib.bib60),[39](https://arxiv.org/html/2608.25181#bib.bib61)\], and statistical mechanics formulations\[[30](https://arxiv.org/html/2608.25181#bib.bib57)\]\. Physics\-informed neural networks \(PINNs\)\[[40](https://arxiv.org/html/2608.25181#bib.bib62)\], neural ODEs\[[41](https://arxiv.org/html/2608.25181#bib.bib2)\], physics\-informed neural vector fields\[[42](https://arxiv.org/html/2608.25181#bib.bib16),[43](https://arxiv.org/html/2608.25181#bib.bib4)\]have also been utilized\. Several challenges arise when attempting to learn governing equations for collective systems from trajectory data\. If one treats the dynamics as an arbitrary vector field on the full state space, then the resulting regression problem is high\-dimensional and computationally demanding, with a dimension that grows with the number of agents\. This is particularly problematic in biological collective dynamics, where systems may consist of many organisms or cells, each evolving inℝ2\\mathbb\{R\}^\{2\}orℝ3\\mathbb\{R\}^\{3\}\. In addition, the relevant functional forms of the governing mechanisms are often not known a priori, making it difficult to construct an appropriate candidate library\. Library\-based methods may therefore be susceptible to model misspecification when important interaction, environmental, or self\-propulsion terms are omitted\[[44](https://arxiv.org/html/2608.25181#bib.bib3),[45](https://arxiv.org/html/2608.25181#bib.bib78)\]\. For a more detailed discussion of these issues in the context of learning collective dynamics, we refer the interested reader to\[[46](https://arxiv.org/html/2608.25181#bib.bib6),[47](https://arxiv.org/html/2608.25181#bib.bib64)\]\. The previously discussed methods are applicable for general dynamical systems of the form \([1](https://arxiv.org/html/2608.25181#S1.E1)\) \(we note that extensions to partial differential equations and stochastic differential equations also exist\[[48](https://arxiv.org/html/2608.25181#bib.bib83),[49](https://arxiv.org/html/2608.25181#bib.bib82),[39](https://arxiv.org/html/2608.25181#bib.bib61),[50](https://arxiv.org/html/2608.25181#bib.bib84),[51](https://arxiv.org/html/2608.25181#bib.bib85)\]\)\. However, these methods generally do not take into account the specific form of the vector fieldFF, or any symmetry that may be a priori known in the application of interest\. For agents interacting collectively,FFtypically assumes a structured form, and one can use this known structure to improve inference\. Indeed, the authors in\[[46](https://arxiv.org/html/2608.25181#bib.bib6),[52](https://arxiv.org/html/2608.25181#bib.bib49),[53](https://arxiv.org/html/2608.25181#bib.bib50),[24](https://arxiv.org/html/2608.25181#bib.bib58),[54](https://arxiv.org/html/2608.25181#bib.bib51)\]have designed variational\-based learning methods for collective systems that can be applied to a diverse range of fields; these methods avoid the problems faced by regression\-based dynamical methods, and are the basis of the work presented here\. More specifically, this general framework assumes that the generally high\-dimensional dynamical system \([1](https://arxiv.org/html/2608.25181#S1.E1)\) is defined by a low\-dimensional interaction kernelϕ:ℝ≥0→ℝ\\phi:\\mathbb\{R\}\_\{\\geq 0\}\\to\\mathbb\{R\}, which quantifies the effect of local pairwise interactions between agents\. The simplest example of such a system is that the interactions directly influence dynamics via pairwise averaging: x˙i\\displaystyle\\dot\{x\}\_\{i\}=1N∑j=1j≠iNϕ\(\|xj−xi\|\)\(xj−xi\)\.\\displaystyle=\\frac\{1\}\{N\}\\sum\_\{\\begin\{subarray\}\{c\}j=1\\\\ j\\neq i\\end\{subarray\}\}^\{N\}\\phi\(\|x\_\{j\}\-x\_\{i\}\|\)\(x\_\{j\}\-x\_\{i\}\)\.\(2\)Here,xix\_\{i\}denotes the state inℝd\\mathbb\{R\}^\{d\}of theithi^\{\\text\{th\}\}agent, andϕ\(\|xj−xi\|\)\\phi\(\|x\_\{j\}\-x\_\{i\}\|\)denotes the strength of the interaction of agentjjon agentii:ϕ\>0\\phi\>0models attraction, whileϕ<0\\phi<0models repulsion\. The functionϕ=ϕ\(r\)\\phi=\\phi\(r\)isassumedto be a function of only the distance between agents, where\|⋅\|\|\\cdot\|denotes the standard Euclidean distance inℝd\\mathbb\{R\}^\{d\}\. Many models in science and engineering take the form \([2](https://arxiv.org/html/2608.25181#S1.E2)\), including\[[18](https://arxiv.org/html/2608.25181#bib.bib10),[10](https://arxiv.org/html/2608.25181#bib.bib11)\]\. Note that the problem of inferringF:ℝD→ℝDF:\\mathbb\{R\}^\{D\}\\to\\mathbb\{R\}^\{D\}is thus reduced to inferringϕ:ℝ≥0→ℝ\\phi:\\mathbb\{R\}\_\{\\geq 0\}\\to\\mathbb\{R\}, as we are utilizing physical knowledge of the phenomena to reduce the dimensionality of the inverse problem\. Of course, other collective systems exist beyond those of the form \([2](https://arxiv.org/html/2608.25181#S1.E2)\) \(e\.g\. second\-order systems\), and we emphasize that the learning methods generally apply; indeed, such variational methods will perform effectively for systems when a significant degree of dimensionality reduction in the candidate vector fields exists and is known a priori\. We note that system \([2](https://arxiv.org/html/2608.25181#S1.E2)\) is defined entirely by interaction forces between agents; there are no intra\-agent forces assumed\. However, in many applications, agents experience both intra\- and inter\-agent dynamics, the former of which includes self\-propulsion, friction, and other environmental forces\. Thus there remains a need for a unified learning framework that simultaneously identifies both the interactions between agents \(ϕ\\phi\) and the intra\-agent environmental forces via trajectory data\. It is the goal of this work to introduce an extended mechanistic learning approach to capture both interactions and the environmental forces simultaneously by generalizing the variational framework\[[46](https://arxiv.org/html/2608.25181#bib.bib6)\]\. Mathematically, we are thus generalizing the approach for system \([2](https://arxiv.org/html/2608.25181#S1.E2)\) to \(for example\) first\-order equations of the form x˙i\\displaystyle\\dot\{x\}\_\{i\}=f\(xi\)\+1N∑j=1j≠iNϕ\(\|xj−xi\|\)\(xj−xi\),\\displaystyle=f\(x\_\{i\}\)\+\\frac\{1\}\{N\}\\sum\_\{\\begin\{subarray\}\{c\}j=1\\ j\\neq i\\end\{subarray\}\}^\{N\}\\phi\(\|x\_\{j\}\-x\_\{i\}\|\)\(x\_\{j\}\-x\_\{i\}\),\(3\)where the environmental forces are represented by the functionf:ℝd→ℝdf:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{d\}; extensions to more general collective systems will be discussed in the manuscript\. The extension \([3](https://arxiv.org/html/2608.25181#S1.E3)\) thus includes both the interaction\-driven dynamics \(ϕ\\phi\) with agent\-level intrinsic effects \(ff\) while preserving the pairwise interaction structure\. Note that we are now interested in inferring bothf:ℝd→ℝdf:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{d\}andϕ:ℝ≥0→ℝ\\phi:\\mathbb\{R\}\_\{\\geq 0\}\\to\\mathbb\{R\}, which still represents a significant reduction in dimension when compared to \([1](https://arxiv.org/html/2608.25181#S1.E1)\), whereF:ℝD→ℝDF:\\mathbb\{R\}^\{D\}\\to\\mathbb\{R\}^\{D\}withD=dND=dNin the case of collective dynamics, as generallyd≪dNd\\ll dN\. Equations such as \([3](https://arxiv.org/html/2608.25181#S1.E3)\) can be utilized to describe a broad class of phenomena, including attraction, repulsion, milling, flocking, swarming, clustering, synchronization, and alignment in agent\-based systems, and are observed in a diverse range of processes such as the phototaxis of bacteria\[[55](https://arxiv.org/html/2608.25181#bib.bib32)\], self\-propelled particles and the coordinated movement of fish schools\[[56](https://arxiv.org/html/2608.25181#bib.bib30)\], swarm robotics\[[57](https://arxiv.org/html/2608.25181#bib.bib87),[58](https://arxiv.org/html/2608.25181#bib.bib86)\], chemotaxis with social interactions\[[59](https://arxiv.org/html/2608.25181#bib.bib88)\], crowd dynamics with destination guidance\[[60](https://arxiv.org/html/2608.25181#bib.bib23)\], and firefly synchronization\[[61](https://arxiv.org/html/2608.25181#bib.bib41)\]\. The main contributions of this work are as follows\. We introduce two approaches \(semi\-parametric and fully non\-parametric\) and demonstrate that they can both be applied within a unified first\- and second\-order model framework to simultaneously recover the environmental forceffand the interaction kernelϕ\\phi\. We note that the parametric formulation may be utilized in the case when a functional form for the intra\-agent dynamics is known, and can be naturally combined to non\-parametrically identify the interaction kernelϕ\\phi\. We next quantify the performance of the learning methods by comparing the mechanistic estimates for the forces utilizing a dynamics\-driven weightedL2L^\{2\}\-distance, as well as by predicting trajectories from both trained and untrained initial conditions\. We validate the performance of the proposed approaches through quantifying the dependence on training data using using statistics computed over different training sets, evaluating the sample\-complexity \(the amount of training data required to learn the model\), and robustness analyses under observational noise on a number of benchmark models of collective dynamics\. Numerical evidence further suggests that these estimators scale favorably with respect to prediction on a significantly larger number of agents than trained on, which is experimentally relevant\. Lastly, we demonstrate how the generalized framework can be used for model selection to identify the “active" forces in a given dataset so as to determine the mechanistic equations that likely describe the observed dynamics\. The structure of the manuscript is as follows\. In Section[2](https://arxiv.org/html/2608.25181#S2), we describe in detail the unified learning approaches which capture both the environmental intra\-agent dynamics as well as the interaction kernel under the semi\-parametric and fully non\-parametric learning approaches\. In Section[3](https://arxiv.org/html/2608.25181#S3), we define performance measures for the learning approaches via a number of metrics, including feature recovery, governing\-equation consistency, dynamical accuracy, statistical reliability, sample complexity, and robustness to observational noise\. Section[4](https://arxiv.org/html/2608.25181#S4)introduces a framework which can be applied to sampled trajectory data to determine the most likely mechanistic model of collective dynamics; we note that this identifies the prescence of intra\-agent forcesff\. Section[5](https://arxiv.org/html/2608.25181#S5)presents a detailed study of three fundamental dynamical systems that vary across order, features, and agent characteristics for both learning approaches, together with an extended study of model selection using an additional three dynamical systems that vary across feature complexity\. Lastly, we conclude the paper with a discussion of limitations and possible future directions\. ### 1\.1Related Methods We briefly discuss results in the variational approach on systems \([2](https://arxiv.org/html/2608.25181#S1.E2)\)\. In\[[46](https://arxiv.org/html/2608.25181#bib.bib6)\]and\[[52](https://arxiv.org/html/2608.25181#bib.bib49)\], the non\-parametric variational approach is proposed to learn the interaction kernelϕ\\phiof the framework \([2](https://arxiv.org/html/2608.25181#S1.E2)\)\. Specifically,\[[52](https://arxiv.org/html/2608.25181#bib.bib49)\]considers a first\-order model of homogeneous agents and studies the convergence to its mean\-field limit, together with the inference of the mean\-field interaction kernel from observations of trajectories of systems with a finite, yet increasing, number of agents\. In\[[46](https://arxiv.org/html/2608.25181#bib.bib6)\], the authors extend the approach to the setting where the number of agents is fixed, but the number of observations increases, showing that the non\-parametric estimators for the interaction kernel converge at the near\-optimal rate for one\-dimensional regression, independent of the dimensionddof the state space\. It further generalizes the estimators to first\- and second\-order heterogeneous agent systems with one\-dimensional interaction kernels based on pairwise distances, providing substantial numerical evidence of the performance of these generalizations\. The work of\[[53](https://arxiv.org/html/2608.25181#bib.bib50)\]analyzes in detail the estimators for first\-order heterogeneous agent\-based systems, generalizing the theoretical results of\[[46](https://arxiv.org/html/2608.25181#bib.bib6)\]to that setting while sharpening some of the constructions\. The work of\[[24](https://arxiv.org/html/2608.25181#bib.bib58)\]extends the application of these approaches to rather general classes of agent\-based systems driven by first\- and second\-order dynamics, with interaction kernels depending not only on pairwise distances but also on other pairwise quantities that depend on the states of the agents\. The overview in\[[24](https://arxiv.org/html/2608.25181#bib.bib58)\]provides a summary of these learning methods under different scenarios, together with possible future extensions and novel approaches related to the framework introduced in\[[46](https://arxiv.org/html/2608.25181#bib.bib6)\]\. ## 2Methods for learning collective systems Here we discuss the methods utilized to simultaneously infer both the interaction and environmental dynamics in general models of collective dynamics\. For notational simplicity, we restrict our attention to first\-order systems; an analogous description of the methods for second\-order systems is provided in Appendix[A](https://arxiv.org/html/2608.25181#A1)\. Section[2\.1](https://arxiv.org/html/2608.25181#S2.SS1)provides a detailed description of the fully non\-parametric procedure, while the semi\-parametric approach is described in Section[2\.2](https://arxiv.org/html/2608.25181#S2.SS2)\. ### 2\.1Variational methods for learning collective and environmental forces: fully non\-parametric approach \(first\-order\) In this section, we describe our algorithm for inferring both the environmental and interaction forces in first\-order models of collective dynamics from observed trajectory data\. Specifically, we assume that the dynamics of a system ofNNparticles are described by the following system of first\-order ordinary differential equations \(ODEs\), fori=1,2,…,Ni=1,2,\\ldots,N: x˙i\\displaystyle\\dot\{x\}\_\{i\}=f\(xi\)\+1N∑j=1j≠iNϕ\(\|xj−xi\|\)\(xj−xi\)\.\\displaystyle=f\(x\_\{i\}\)\+\\frac\{1\}\{N\}\\sum\_\{\\begin\{subarray\}\{c\}j=1\\\\ j\\neq i\\end\{subarray\}\}^\{N\}\\phi\(\|x\_\{j\}\-x\_\{i\}\|\)\(x\_\{j\}\-x\_\{i\}\)\.\(4\)Herexi=xi\(t\)∈ℝdx\_\{i\}=x\_\{i\}\(t\)\\in\\mathbb\{R\}^\{d\}denotes the state of theithi^\{\\text\{th\}\}agent at timett\. The functionf:ℝd→ℝdf:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{d\}models the environmental forces on the agents, which for simplicity we assume to be identical for all agents in the system; typical examples include frictional forces and self\-propulsion, which may be due to both the external environment and/or intrinsic agent dynamics of the specific model considered\. We also note that in certain applications, the domain offfmay be a proper subset ofΩ⊆ℝd\\Omega\\subseteq\\mathbb\{R\}^\{d\}, which will not change the proposed algorithm in any significant manner\. The functionϕ:ℝ≥0→ℝ\\phi:\\mathbb\{R\}\_\{\\geq 0\}\\to\\mathbb\{R\}encodes the interaction strength between agents, which we assume, again for simplicity of presentation, to be a function only of the distance\|⋅\|\|\\cdot\|between agents\. As discussed previously and presented in Section[1](https://arxiv.org/html/2608.25181#S1), the general form of system \([4](https://arxiv.org/html/2608.25181#S2.E4)\) has been utilized in many physical, biological, engineering, and social science contexts\. We note that this is a modeling assumption, and that the proposed algorithm can be readily adapted to more general interaction kernels\. Here\|⋅\|\|\\cdot\|denotes the standard Euclidean norm onℝd\\mathbb\{R\}^\{d\}, but other norms or metrics may be utilized as appropriate for the scientific application considered\. Our goal is to obtain estimators forffandϕ\\phiin \([4](https://arxiv.org/html/2608.25181#S2.E4)\) from observed trajectory data\. In the following subsections, we describe both the form of the observed data which we use to obtain the estimators, as well as the estimation algorithm\. We note that the subsequent descriptions are for systems of the form \([4](https://arxiv.org/html/2608.25181#S2.E4)\); the analogous procedure for second\-order systems can be found in Appendix[A](https://arxiv.org/html/2608.25181#A1) #### 2\.1\.1Trajectory data In this subsection, we formalize the form and notation for the trajectory data which will be utilized to construct the estimators forffandϕ\\phiin \([4](https://arxiv.org/html/2608.25181#S2.E4)\)\. We begin by rewriting \([4](https://arxiv.org/html/2608.25181#S2.E4)\) as an explicit dynamical system inℝdN\\mathbb\{R\}^\{dN\}\(NNagents each taking values inℝd\\mathbb\{R\}^\{d\}\)\. Denote the full state of the system at timettvia x\(t\)\\displaystyle x\(t\):=\(x1\(t\)x2\(t\)xN\(t\)\)∈ℝdN\.\\displaystyle:=\\begin\{pmatrix\}x\_\{1\}\(t\)\\\\ x\_\{2\}\(t\)\\\\ \\vdots\\\\ x\_\{N\}\(t\)\\end\{pmatrix\}\\in\\mathbb\{R\}^\{dN\}\.\(5\)Thus, by defining Ff\(x\)\\displaystyle F\_\{f\}\(x\):=\(f\(x1\)f\(x2\)f\(xN\)\)∈ℝdNFϕ\(x\):=1N\(∑j=1j≠1Nϕ\(\|xj−x1\|\)\(xj−x1\)∑j=1j≠2Nϕ\(\|xj−x2\|\)\(xj−x2\)∑j=1j≠NNϕ\(\|xj−xN\|\)\(xj−xN\),\)∈ℝdN\\displaystyle:=\\begin\{pmatrix\}f\(x\_\{1\}\)\\\\ f\(x\_\{2\}\)\\\\ \\vdots\\\\ f\(x\_\{N\}\)\\end\{pmatrix\}\\in\\mathbb\{R\}^\{dN\}\\quad\\quad F\_\{\\phi\}\(x\):=\\frac\{1\}\{N\}\\begin\{pmatrix\}\\sum\\limits\_\{\\begin\{subarray\}\{c\}j=1\\\\ j\\neq 1\\end\{subarray\}\}^\{N\}\\phi\(\|x\_\{j\}\-x\_\{1\}\|\)\(x\_\{j\}\-x\_\{1\}\)\\\\ \\sum\\limits\_\{\\begin\{subarray\}\{c\}j=1\\\\ j\\neq 2\\end\{subarray\}\}^\{N\}\\phi\(\|x\_\{j\}\-x\_\{2\}\|\)\(x\_\{j\}\-x\_\{2\}\)\\\\ \\vdots\\\\ \\sum\\limits\_\{\\begin\{subarray\}\{c\}j=1\\\\ j\\neq N\\end\{subarray\}\}^\{N\}\\phi\(\|x\_\{j\}\-x\_\{N\}\|\)\(x\_\{j\}\-x\_\{N\}\),\\end\{pmatrix\}\\in\\mathbb\{R\}^\{dN\}\(6\)we see that we can realize \([4](https://arxiv.org/html/2608.25181#S2.E4)\) in the form x˙\\displaystyle\\dot\{x\}=Ff\(x\)\+Fϕ\(x\)\.\\displaystyle=F\_\{f\}\(x\)\+F\_\{\\phi\}\(x\)\.\(7\)We denotexi∈ℝdx\_\{i\}\\in\\mathbb\{R\}^\{d\}as theithi^\{\\text\{th\}\}component ofx∈ℝdx\\in\\mathbb\{R\}^\{d\}, and will utilize this notation to precisely define the form of the observation data and estimation algorithm below\. Assume thatMMreplicates of initial conditions are independently and identically \(i\.i\.d\.\) sampled from a fixed but generally unknown probability distributionμ0\\mu\_\{0\}onℝdN\\mathbb\{R\}^\{dN\}; note that each replicate corresponds to a set of initial conditions inℝd\\mathbb\{R\}^\{d\}for theNNagents\. Denote these initial conditions via X0\(m\)∈ℝdN,\\displaystyle X^\{\(m\)\}\_\{0\}\\in\\mathbb\{R\}^\{dN\},\(8\)form=1,2,…,Mm=1,2,\\ldots,M\. For each suchmm, denote the solution at timettof the corresponding initial\-value problem \(IVP\) \{x˙=Ff\(x\)\+Fϕ\(x\)x\(0\)=X0\(m\)\\displaystyle\\begin\{cases\}\\dot\{x\}&=F\_\{f\}\(x\)\+F\_\{\\phi\}\(x\)\\\\ x\(0\)&=X^\{\(m\)\}\_\{0\}\\end\{cases\}\(9\)asx\(m\)\(t\)x^\{\(m\)\}\(t\)\. We assume that the experimental data consists of observations at discrete time points0=t1<t2<⋯<tL=T0=t\_\{1\}<t\_\{2\}<\\cdots<t\_\{L\}=T, which we denote as Xℓ\(m\)\\displaystyle X^\{\(m\)\}\_\{\\ell\}:=x\(m\)\(tℓ\)\\displaystyle:=x^\{\(m\)\}\(t\_\{\\ell\}\)\(10\)As before,Xℓ\(m\)∈ℝdNX^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{dN\}, and thus the set\(Xℓ\(m\)\)m,ℓ=1M,L\(X^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}is generally the set of observed trajectory data which will be utilized to obtain the estimators forffandϕ\\phi\. Note that by construction, \(Xℓ\(m\)\)i\\displaystyle\(X^\{\(m\)\}\_\{\\ell\}\)\_\{i\}=xi\(m\)\(tℓ\)∈ℝd,\\displaystyle=x^\{\(m\)\}\_\{i\}\(t\_\{\\ell\}\)\\in\\mathbb\{R\}^\{d\},\(11\)i\.e\. that theithi^\{\\text\{th\}\}component \(with respect to the natural decomposition ofXℓ\(m\)X^\{\(m\)\}\_\{\\ell\}via \([5](https://arxiv.org/html/2608.25181#S2.E5)\)\) ofXℓ\(m\)X^\{\(m\)\}\_\{\\ell\}is the state of agentiiat timetℓt\_\{\\ell\}of themthm^\{\\text\{th\}\}replicate\. The learning algorithm introduced in Section[2\.1\.2](https://arxiv.org/html/2608.25181#S2.SS1.SSS2)will also require observations of the velocityv:=x˙v:=\\dot\{x\}in \([7](https://arxiv.org/html/2608.25181#S2.E7)\)\. Specifically, we define Vℓ\(m\)\\displaystyle V^\{\(m\)\}\_\{\\ell\}:=x˙\(m\)\(tℓ\),\\displaystyle:=\\dot\{x\}^\{\(m\)\}\(t\_\{\\ell\}\),\(12\)which may be approximated from\(Xℓ\(m\)\)m,ℓ=1M,L\(X^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}\. In practice, we may obtain this approximation by first applying a low\-pass filterℋ\\mathcal\{H\}to the trajectory data, and thus estimate the necessary velocities via forward differences of the filtered data: Vℓ\(m\)\\displaystyle V^\{\(m\)\}\_\{\\ell\}≈1Δtℓ\(ℋ\(Xℓ\+1\(m\)\)−ℋ\(Xℓ\(m\)\)\),\\displaystyle\\approx\\frac\{1\}\{\\Delta t\_\{\\ell\}\}\(\\mathcal\{H\}\(X^\{\(m\)\}\_\{\\ell\+1\}\)\-\\mathcal\{H\}\(X^\{\(m\)\}\_\{\\ell\}\)\),\(13\)whereΔtℓ:=tℓ\+1−tℓ\\Delta t\_\{\\ell\}:=t\_\{\\ell\+1\}\-t\_\{\\ell\}forℓ=1,2,…,L\\ell=1,2,\\ldots,L; see also the discussion in Section[3\.5\.4](https://arxiv.org/html/2608.25181#S3.SS5.SSS4)\. Taken together, the observed data utilized to estimateϕ\\phiandfftakes the form\(Xℓ\(m\),Vℓ\(m\)\)m,ℓ=1M,L\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}, whereXℓ\(m\),Vℓ\(m\)∈ℝdN\.X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{dN\}\.Since the interaction kernelϕ\\phidepends on on pairwise distancesrrbetween agents, it is natural to also introduce notation for the observed pairwise distances from the trajectory data\. Thus, we defineRℓ\(m\)∈ℝN×NR^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{N\\times N\}via \(Rℓ\(m\)\)ij:=\|\(Xℓ\(m\)\)i−\(Xℓ\(m\)\)j\|∈ℝ,\\displaystyle\(R^\{\(m\)\}\_\{\\ell\}\)\_\{ij\}:=\|\(X^\{\(m\)\}\_\{\\ell\}\)\_\{i\}\-\(X^\{\(m\)\}\_\{\\ell\}\)\_\{j\}\|\\in\\mathbb\{R\},\(14\)i\.e\.\(Rℓ\(m\)\)ij\(R^\{\(m\)\}\_\{\\ell\}\)\_\{ij\}measures the distance between agentsiiandjjat timetℓt\_\{\\ell\}of replicate \(initial condition\)mm\. Prior to formalizing the estimation procedure, we note that the formulation \([7](https://arxiv.org/html/2608.25181#S2.E7)\) induces a natural inner product \(and thus norm\) onℝdN\\mathbb\{R\}^\{dN\}via the standard Euclidean norm\|⋅\|\|\\cdot\|onℝd\\mathbb\{R\}^\{d\}and theNN\-agent symmetry of \([4](https://arxiv.org/html/2608.25181#S2.E4)\)\. Specifically, we define, forX,Y∈ℝdN,X,Y\\in\\mathbb\{R\}^\{dN\}, ⟨X,Y⟩ℝdN:=1N∑i=1N⟨Xi,Yi⟩=1NXTY,\\displaystyle\\begin\{split\}\\left\\langle X,Y\\right\\rangle\_\{\\mathbb\{R\}^\{dN\}\}&:=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\left\\langle X\_\{i\},Y\_\{i\}\\right\\rangle=\\frac\{1\}\{N\}X^\{T\}Y,\\end\{split\}\(15\)whereXTX^\{T\}denotes the transpose ofXX,X=\(X1,X2,…,XN\),Y=\(Y1,Y2,…,YN\)X=\(X\_\{1\},X\_\{2\},\\ldots,X\_\{N\}\),Y=\(Y\_\{1\},Y\_\{2\},\\ldots,Y\_\{N\}\)andXi,Yi∈ℝdX\_\{i\},Y\_\{i\}\\in\\mathbb\{R\}^\{d\}fori=1,2,…,Ni=1,2,\\ldots,N, and⟨⋅,⋅⟩\\left\\langle\\cdot,\\cdot\\right\\rangledenotes the standard Euclidean inner product onℝd\\mathbb\{R\}^\{d\}\. We then define the associated norm onℝdN\\mathbb\{R\}^\{dN\}as ‖⋅‖ℝdN\\displaystyle\\left\\lVert\\cdot\\right\\rVert\_\{\\mathbb\{R\}^\{dN\}\}:=⟨⋅,⋅⟩ℝdN\.\\displaystyle:=\\sqrt\{\\left\\langle\\cdot,\\cdot\\right\\rangle\_\{\\mathbb\{R\}^\{dN\}\}\}\.\(16\)This norm and inner product will be utilized in the error functional defined in Section[2\.1\.2](https://arxiv.org/html/2608.25181#S2.SS1.SSS2)\. #### 2\.1\.2Estimation algorithm The central idea in the learning algorithm is that we assume the unknown functionsffandϕ\\phican be well\-approximated within finite\-dimensional hypothesis spaces; our goal then is to calculate these approximations, which will then serve as our desired estimators\. Specifically, we fix two finite\-dimensional functional vector spacesℋf⊆ℱ\(ℝd,ℝd\)\\mathcal\{H\}\_\{f\}\\subseteq\\mathcal\{F\}\(\\mathbb\{R\}^\{d\},\\mathbb\{R\}^\{d\}\)andℋϕ⊆ℱ\(ℝ≥0,ℝ\)\\mathcal\{H\}\_\{\\phi\}\\subseteq\\mathcal\{F\}\(\\mathbb\{R\}\_\{\\geq 0\},\\mathbb\{R\}\), whereℱ\(X,Y\)\\mathcal\{F\}\(X,Y\)denotes the set of all functions between setsXXandYY; we refer to these spaces as hypothesis spaces\. A discussion of hypothesis space selection is provided in Section[2\.1\.4](https://arxiv.org/html/2608.25181#S2.SS1.SSS4)\. We then define the following error functionalℰℋf,ℋϕ\\mathcal\{E\}\_\{\\mathcal\{H\}\_\{f\},\\mathcal\{H\}\_\{\\phi\}\}over the hypothesis spacesℋf\\mathcal\{H\}\_\{f\}andℋϕ\\mathcal\{H\}\_\{\\phi\}as ℰℋf,ℋϕ\(f~,ϕ~\)\\displaystyle\\mathcal\{E\}\_\{\\mathcal\{H\}\_\{f\},\\mathcal\{H\}\_\{\\phi\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\):=1ML∑m,ℓ=1M,L‖Vℓ\(m\)−\(Ff~\(Xℓ\(m\)\)\+Fϕ~\(Xℓ\(m\)\)\)‖ℝdN2,\\displaystyle:=\\frac\{1\}\{ML\}\\sum\_\{m,\\ell=1\}^\{M,L\}\\left\\lVert V^\{\(m\)\}\_\{\\ell\}\-\\Big\(F\_\{\\tilde\{f\}\}\(X^\{\(m\)\}\_\{\\ell\}\)\+F\_\{\\tilde\{\\phi\}\}\(X^\{\(m\)\}\_\{\\ell\}\)\\Big\)\\right\\rVert\_\{\\mathbb\{R\}^\{dN\}\}^\{2\},\(17\)where\(Xℓ\(m\),Vℓ\(m\)\)m,ℓ=1M,L\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}denotes the \(fixed\) trajectory data as introduced in Section[2\.1\.1](https://arxiv.org/html/2608.25181#S2.SS1.SSS1), andf~∈ℋf,ϕ~∈ℋϕ\\tilde\{f\}\\in\\mathcal\{H\}\_\{f\},\\tilde\{\\phi\}\\in\\mathcal\{H\}\_\{\\phi\}, withFf~F\_\{\\tilde\{f\}\}andFϕ~F\_\{\\tilde\{\\phi\}\}defined as in \([6](https://arxiv.org/html/2608.25181#S2.E6)\) and \([7](https://arxiv.org/html/2608.25181#S2.E7)\)\. Note that the notational choice off~\\tilde\{f\}andϕ~\\tilde\{\\phi\}reflects our assumption thatffandϕ\\phiin \([4](https://arxiv.org/html/2608.25181#S2.E4)\) represent thetrueenvironmental and inter\-agent forces, respectively, and our goal is to construct estimatorsf^\\hat\{f\}andϕ^\\hat\{\\phi\}via the error functionalℰℋf,ℋϕ\.\\mathcal\{E\}\_\{\\mathcal\{H\}\_\{f\},\\mathcal\{H\}\_\{\\phi\}\}\.Intuitively,ℰℋf,ℋϕ\(f~,ϕ~\)\\mathcal\{E\}\_\{\\mathcal\{H\}\_\{f\},\\mathcal\{H\}\_\{\\phi\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\)measures the average error \(over time and replicates\) of the sum\-of\-squares error of observing the data\(Xℓ\(m\),Vℓ\(m\)\)m,ℓ=1M,L\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}when predicting with the vector fieldsf~\\tilde\{f\}andϕ~\\tilde\{\\phi\}in \([7](https://arxiv.org/html/2608.25181#S2.E7)\)\. Indeed, by the law of large numbers,ℰℋf,ℋϕ\(f~,ϕ~\)\\mathcal\{E\}\_\{\\mathcal\{H\}\_\{f\},\\mathcal\{H\}\_\{\\phi\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\)is an estimator for a time\-discretized version of 𝔼X\(0\)∼μ0\[1T∫0T∥X˙\(t\)−\(Ff~\(X\(t\)\)\+Fϕ~\(X\(t\)\)\)∥ℝdN2𝑑t\]\.\\displaystyle\\mathbb\{E\}\_\{X\(0\)\\sim\\mu\_\{0\}\}\\left\[\\frac\{1\}\{T\}\\int\_\{0\}^\{T\}\\lVert\\dot\{X\}\(t\)\-\\left\(F\_\{\\tilde\{f\}\}\(X\(t\)\)\+F\_\{\\tilde\{\\phi\}\}\(X\(t\)\)\\right\)\\rVert\_\{\\mathbb\{R\}^\{dN\}\}^\{2\}\\,\\mathrm\{d\}t\\right\]\.\(18\)As in standard likelihood approaches, our goal is to minimize \([17](https://arxiv.org/html/2608.25181#S2.E17)\) overℋf\\mathcal\{H\}\_\{f\}andℋϕ\\mathcal\{H\}\_\{\\phi\}\. Asℋf\\mathcal\{H\}\_\{f\}andℋϕ\\mathcal\{H\}\_\{\\phi\}are finite dimensional, let\(hk\)k=1nf\(h\_\{k\}\)\_\{k=1\}^\{n\_\{f\}\}and\(ψk\)k=1nϕ\(\\psi\_\{k\}\)\_\{k=1\}^\{n\_\{\\phi\}\}be respective ordered bases\. Thus, for anyf~∈ℋf\\tilde\{f\}\\in\\mathcal\{H\}\_\{f\}andϕ~∈ℋϕ\\tilde\{\\phi\}\\in\\mathcal\{H\}\_\{\\phi\}, there existsα=\(α1,α2,…,αnf\)∈ℝnf\\alpha=\(\\alpha\_\{1\},\\alpha\_\{2\},\\ldots,\\alpha\_\{n\_\{f\}\}\)\\in\\mathbb\{R\}^\{n\_\{f\}\}andβ=\(β1,β2,…,βnϕ\)∈ℝnϕ\\beta=\(\\beta\_\{1\},\\beta\_\{2\},\\ldots,\\beta\_\{n\_\{\\phi\}\}\)\\in\\mathbb\{R\}^\{n\_\{\\phi\}\}such that f~=∑k=1nfαkhkandϕ~=∑k=1nϕβkψk\.\\displaystyle\\begin\{split\}\\tilde\{f\}&=\\sum\_\{k=1\}^\{n\_\{f\}\}\\alpha\_\{k\}h\_\{k\}\\quad\\text\{and\}\\quad\\tilde\{\\phi\}=\\sum\_\{k=1\}^\{n\_\{\\phi\}\}\\beta\_\{k\}\\psi\_\{k\}\.\\end\{split\}\(19\)As in \([6](https://arxiv.org/html/2608.25181#S2.E6)\), for each1≤k≤nf1\\leq k\\leq n\_\{f\}, writingx=\(x1,x2,…,xN\)∈ℝdNx=\(x\_\{1\},x\_\{2\},\\ldots,x\_\{N\}\)\\in\\mathbb\{R\}^\{dN\}withxi∈ℝdx\_\{i\}\\in\\mathbb\{R\}^\{d\}, we define hk\(x\)\\displaystyle h\_\{k\}\(x\):=\(hk\(x1\)hk\(x2\)⋯hk\(xN\)\)\\displaystyle:=\\begin\{pmatrix\}h\_\{k\}\(x\_\{1\}\)\\\\ h\_\{k\}\(x\_\{2\}\)\\\\ \\cdots\\\\ h\_\{k\}\(x\_\{N\}\)\\end\{pmatrix\}\(20\)and the matrixHℓ\(m\)∈ℝdN×nfH^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{dN\\times n\_\{f\}\} Hℓ\(m\)\\displaystyle H^\{\(m\)\}\_\{\\ell\}:=\(h1\(Xℓ\(m\)\)h2\(Xℓ\(m\)\)⋯hnf\(Xℓ\(m\)\)\)\.\\displaystyle:=\\Big\(\\,h\_\{1\}\(X^\{\(m\)\}\_\{\\ell\}\)\\quad h\_\{2\}\(X^\{\(m\)\}\_\{\\ell\}\)\\quad\\cdots\\quad h\_\{n\_\{f\}\}\(X^\{\(m\)\}\_\{\\ell\}\)\\,\\Big\)\.\(21\)We note that for1≤m≤M1\\leq m\\leq Mand1≤ℓ≤L1\\leq\\ell\\leq L, each matrixHℓ\(m\)H^\{\(m\)\}\_\{\\ell\}is constant in the estimation algorithm, as a function only of the trajectory data\(Xℓ\(m\)\)m,ℓ=1M,L\(X^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}and the selected basis\(hk\)k=1nf\(h\_\{k\}\)\_\{k=1\}^\{n\_\{f\}\}forℋf\\mathcal\{H\}\_\{f\}\. Thus, expandingf~∈ℋf\\tilde\{f\}\\in\\mathcal\{H\}\_\{f\}as in \([19](https://arxiv.org/html/2608.25181#S2.E19)\),Ff~\(Xℓ\(m\)\)F\_\{\\tilde\{f\}\}\(X^\{\(m\)\}\_\{\\ell\}\)in \([17](https://arxiv.org/html/2608.25181#S2.E17)\) takes the form of matrix\-vector multiplication: Ff~\(Xℓ\(m\)\)\\displaystyle F\_\{\\tilde\{f\}\}\(X^\{\(m\)\}\_\{\\ell\}\)=Hℓ\(m\)α\.\\displaystyle=H^\{\(m\)\}\_\{\\ell\}\\alpha\.\(22\)Similarly, we define ψk\(x\)\\displaystyle\\psi\_\{k\}\(x\):=\(ψk\(x1\)ψk\(x2\)⋯ψk\(xN\)\)\\displaystyle:=\\begin\{pmatrix\}\\psi\_\{k\}\(x\_\{1\}\)\\\\ \\psi\_\{k\}\(x\_\{2\}\)\\\\ \\cdots\\\\ \\psi\_\{k\}\(x\_\{N\}\)\\end\{pmatrix\}\(23\)and Fψk\(x\):=1N\(∑j=1j≠1Nψk\(\|xj−x1\|\)\(xj−x1\)∑j=1j≠2Nψk\(\|xj−x2\|\)\(xj−x2\)∑j=1j≠NNψk\(\|xj−xN\|\)\(xj−xN\),\)\\displaystyle F\_\{\\psi\_\{k\}\}\(x\):=\\frac\{1\}\{N\}\\begin\{pmatrix\}\\sum\\limits\_\{\\begin\{subarray\}\{c\}j=1\\\\ j\\neq 1\\end\{subarray\}\}^\{N\}\\psi\_\{k\}\(\|x\_\{j\}\-x\_\{1\}\|\)\(x\_\{j\}\-x\_\{1\}\)\\\\ \\sum\\limits\_\{\\begin\{subarray\}\{c\}j=1\\\\ j\\neq 2\\end\{subarray\}\}^\{N\}\\psi\_\{k\}\(\|x\_\{j\}\-x\_\{2\}\|\)\(x\_\{j\}\-x\_\{2\}\)\\\\ \\vdots\\\\ \\sum\\limits\_\{\\begin\{subarray\}\{c\}j=1\\\\ j\\neq N\\end\{subarray\}\}^\{N\}\\psi\_\{k\}\(\|x\_\{j\}\-x\_\{N\}\|\)\(x\_\{j\}\-x\_\{N\}\),\\end\{pmatrix\}\(24\)for each1≤k≤nϕ1\\leq k\\leq n\_\{\\phi\}as in \([6](https://arxiv.org/html/2608.25181#S2.E6)\), from which we further define, for1≤m≤M1\\leq m\\leq Mand1≤ℓ≤L1\\leq\\ell\\leq Lthe matrixΨℓ\(m\)∈ℝdN×nϕ\\Psi^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{dN\\times\{n\_\{\\phi\}\}\} Ψℓ\(m\)\\displaystyle\\Psi^\{\(m\)\}\_\{\\ell\}:=\(Fψ1\(Xℓ\(m\)\)Fψ2\(Xℓ\(m\)\)⋯Fψnϕ\(Xℓ\(m\)\)\)\.\\displaystyle:=\\Big\(\\,F\_\{\\psi\_\{1\}\}\(X^\{\(m\)\}\_\{\\ell\}\)\\quad F\_\{\\psi\_\{2\}\}\(X^\{\(m\)\}\_\{\\ell\}\)\\quad\\cdots\\quad F\_\{\\psi\_\{n\_\{\\phi\}\}\}\(X^\{\(m\)\}\_\{\\ell\}\)\\,\\Big\)\.\(25\)Recall again that for1≤m≤M1\\leq m\\leq Mand1≤ℓ≤L1\\leq\\ell\\leq L, each matrixΨℓ\(m\)\\Psi^\{\(m\)\}\_\{\\ell\}is constant in the estimation algorithm, as a function only of the trajectory data\(Xℓ\(m\)\)m,ℓ=1M,L\(X^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}and the selected basis\(ψk\)k=1nf\(\\psi\_\{k\}\)\_\{k=1\}^\{n\_\{f\}\}forℋϕ\\mathcal\{H\}\_\{\\phi\}\. Expandingϕ~∈ℋϕ\\tilde\{\\phi\}\\in\\mathcal\{H\}\_\{\\phi\}as in \([19](https://arxiv.org/html/2608.25181#S2.E19)\), the termFϕ~\(Xℓ\(m\)\)F\_\{\\tilde\{\\phi\}\}\(X^\{\(m\)\}\_\{\\ell\}\)in \([17](https://arxiv.org/html/2608.25181#S2.E17)\) thus also takes the form of matrix\-vector multiplication: Fϕ~\(Xℓ\(m\)\)\\displaystyle F\_\{\\tilde\{\\phi\}\}\(X^\{\(m\)\}\_\{\\ell\}\)=Ψℓ\(m\)β\.\\displaystyle=\\Psi^\{\(m\)\}\_\{\\ell\}\\beta\.\(26\) Using \([22](https://arxiv.org/html/2608.25181#S2.E22)\) and \([26](https://arxiv.org/html/2608.25181#S2.E26)\), we observe that minimizing \([17](https://arxiv.org/html/2608.25181#S2.E17)\) overℋf\\mathcal\{H\}\_\{f\}andℋϕ\\mathcal\{H\}\_\{\\phi\}is equivalent to minimizing ℰ\(α,β\)\\displaystyle\\mathcal\{E\}\(\\alpha,\\beta\):=1ML∑m,ℓ=1M,L‖Vℓ\(m\)−\(Hℓ\(m\)α\+Ψℓ\(m\)β\)‖ℝdN2,\\displaystyle:=\\frac\{1\}\{ML\}\\sum\_\{m,\\ell=1\}^\{M,L\}\\left\\lVert V^\{\(m\)\}\_\{\\ell\}\-\\Big\(H^\{\(m\)\}\_\{\\ell\}\\alpha\+\\Psi^\{\(m\)\}\_\{\\ell\}\\beta\\Big\)\\right\\rVert\_\{\\mathbb\{R\}^\{dN\}\}^\{2\},\(27\)over\(α,β\)∈ℝnf\+nϕ\(\\alpha,\\beta\)\\in\\mathbb\{R\}^\{n\_\{f\}\+n\_\{\\phi\}\}\. As \([27](https://arxiv.org/html/2608.25181#S2.E27)\) is quadratic in\(α,β\)\(\\alpha,\\beta\)\(recall that the norm‖⋅‖ℝdN\\left\\lVert\\cdot\\right\\rVert\_\{\\mathbb\{R\}^\{dN\}\}is defined in \([15](https://arxiv.org/html/2608.25181#S2.E15)\) \- \([16](https://arxiv.org/html/2608.25181#S2.E16)\) via the standard Euclidean inner product onℝd\\mathbb\{R\}^\{d\}\), a solution\(α^,β^\)\(\\hat\{\\alpha\},\\hat\{\\beta\}\)always exists: \(α^,β^\)\\displaystyle\(\\hat\{\\alpha\},\\hat\{\\beta\}\):=argminα∈ℝnf,β∈ℝnϕℰ\(α,β\)\.\\displaystyle:=\\argmin\_\{\\alpha\\in\\mathbb\{R\}^\{n\_\{f\}\},\\,\\beta\\in\\mathbb\{R\}^\{n\_\{\\phi\}\}\}\\mathcal\{E\}\(\\alpha,\\beta\)\.\(28\) #### 2\.1\.3Normal equations We provide the system of linear equations \(i\.e\. the normal equations\) which characterize solutions of \([28](https://arxiv.org/html/2608.25181#S2.E28)\)\. For simplicity, defineθ∈ℝnf\+nϕ\\theta\\in\\mathbb\{R\}^\{n\_\{f\}\+n\_\{\\phi\}\}as ordered set of all coefficients to be determined, θ\\displaystyle\\theta:=\(αβ\),\\displaystyle:=\\begin\{pmatrix\}\\alpha\\\\ \\beta\\end\{pmatrix\},\(29\)and the matricesCℓ\(m\)∈ℝdN×\(nf\+nϕ\)C^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{dN\\times\{\(n\_\{f\}\+n\_\{\\phi\}\)\}\} Cℓ\(m\)\\displaystyle C^\{\(m\)\}\_\{\\ell\}:=\(Hℓ\(m\)Ψℓ\(m\)\)\\displaystyle:=\(\\,H^\{\(m\)\}\_\{\\ell\}\\quad\\Psi^\{\(m\)\}\_\{\\ell\}\\,\)\(30\)form=1,2,…,Mm=1,2,\\ldots,Mandℓ=1,2,…,L\\ell=1,2,\\ldots,L\. With respect to this notation, equation \([27](https://arxiv.org/html/2608.25181#S2.E27)\) takes the form ℰ\(θ\)\\displaystyle\\mathcal\{E\}\(\\theta\):=1ML∑m,ℓ=1M,L‖Vℓ\(m\)−Cℓ\(m\)θ‖ℝdN2\.\\displaystyle:=\\frac\{1\}\{ML\}\\sum\_\{m,\\ell=1\}^\{M,L\}\\left\\lVert V^\{\(m\)\}\_\{\\ell\}\-C^\{\(m\)\}\_\{\\ell\}\\theta\\right\\rVert\_\{\\mathbb\{R\}^\{dN\}\}^\{2\}\.\(31\)Thus, we minimize \([31](https://arxiv.org/html/2608.25181#S2.E31)\) with respect toθ=\(α,β\)∈ℝnf\+nϕ\\theta=\(\\alpha,\\beta\)\\in\\mathbb\{R\}^\{n\_\{f\}\+n\_\{\\phi\}\}, and obtain the estimator coefficients \(α^,β^\)\\displaystyle\(\\hat\{\\alpha\},\\hat\{\\beta\}\)=θ^:=argminθ∈ℝnf\+nϕℰ\(θ\)\.\\displaystyle=\\hat\{\\theta\}:=\\argmin\_\{\\theta\\in\\mathbb\{R\}^\{n\_\{f\}\+n\_\{\\phi\}\}\}\\mathcal\{E\}\(\\theta\)\.\(32\)Since‖⋅‖ℝdN\\left\\lVert\\cdot\\right\\rVert\_\{\\mathbb\{R\}^\{dN\}\}is derived from the standard inner product onℝdN\\mathbb\{R\}^\{dN\}via \([15](https://arxiv.org/html/2608.25181#S2.E15)\), solving the minimization problem \([32](https://arxiv.org/html/2608.25181#S2.E32)\) corresponds precisely to a least\-squares problem with multiple observations \(one observation for eachm=1,2,…,Mm=1,2,\\ldots,Mandℓ=1,2,…,L\\ell=1,2,\\ldots,L\)\. Thus, as a minimizingθ^\\hat\{\\theta\}necessarily satisfies ∇θℰ\(θ^\)\\displaystyle\\nabla\_\{\\theta\}\\mathcal\{E\}\(\\hat\{\\theta\}\)=0,\\displaystyle=0,\(33\)and as ∇θ⟨Vℓ\(m\)−Cℓ\(m\)θ,Vℓ\(m\)−Cℓ\(m\)θ⟩ℝdN\\displaystyle\\nabla\_\{\\theta\}\\left\\langle V^\{\(m\)\}\_\{\\ell\}\-C^\{\(m\)\}\_\{\\ell\}\\theta,V^\{\(m\)\}\_\{\\ell\}\-C^\{\(m\)\}\_\{\\ell\}\\theta\\right\\rangle\_\{\\mathbb\{R\}^\{dN\}\}=2N\(\(Cℓ\(m\)\)TCℓ\(m\)θ−\(Cℓ\(m\)\)TVℓ\(m\)\),\\displaystyle=\\frac\{2\}\{N\}\\Big\(\(C^\{\(m\)\}\_\{\\ell\}\)^\{T\}C^\{\(m\)\}\_\{\\ell\}\\theta\-\(C^\{\(m\)\}\_\{\\ell\}\)^\{T\}V^\{\(m\)\}\_\{\\ell\}\\Big\),\(34\)a minimizingθ^\\hat\{\\theta\}must satisfy the linear system \(∑m,ℓ=1M,L\(Cℓ\(m\)\)TCℓ\(m\)\)θ^\\displaystyle\\left\(\\sum\_\{m,\\ell=1\}^\{M,L\}\(C^\{\(m\)\}\_\{\\ell\}\)^\{T\}C^\{\(m\)\}\_\{\\ell\}\\right\)\\hat\{\\theta\}=∑m,ℓ=1M,L\(Cℓ\(m\)\)TVℓ\(m\)\.\\displaystyle=\\sum\_\{m,\\ell=1\}^\{M,L\}\(C^\{\(m\)\}\_\{\\ell\}\)^\{T\}V^\{\(m\)\}\_\{\\ell\}\.\(35\)Defining A:=∑m,ℓ=1M,L\(C\(m\)ℓ\)TC\(m\)ℓ∈ℝ\(nf\+nϕ\)×\(nf\+nϕ\)andb:=∑m,ℓ=1M,L\(C\(m\)ℓ\)TV\(m\)ℓ∈ℝnf\+nϕ,\\displaystyle\\begin\{split\}A&:=\\sum\_\{m,\\ell=1\}^\{M,L\}\(C^\{\(m\)\}\_\{\\ell\}\)^\{T\}C^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{\(n\_\{f\}\+n\_\{\\phi\}\)\\times\(n\_\{f\}\+n\_\{\\phi\}\)\}\\quad\\text\{and\}\\quad b:=\\sum\_\{m,\\ell=1\}^\{M,L\}\(C^\{\(m\)\}\_\{\\ell\}\)^\{T\}V^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{n\_\{f\}\+n\_\{\\phi\}\},\\end\{split\}\(36\)we see that \([35](https://arxiv.org/html/2608.25181#S2.E35)\) is equivalent to solving Aθ^\\displaystyle A\\hat\{\\theta\}=b\.\\displaystyle=b\.\(37\)We lastly note that sinceCℓ\(m\)C^\{\(m\)\}\_\{\\ell\}takes the form \([30](https://arxiv.org/html/2608.25181#S2.E30)\), the components of\(Cℓ\(m\)\)TCℓ\(m\)\(C^\{\(m\)\}\_\{\\ell\}\)^\{T\}C^\{\(m\)\}\_\{\\ell\}and\(Cℓ\(m\)\)TVℓ\(m\)\(C^\{\(m\)\}\_\{\\ell\}\)^\{T\}V^\{\(m\)\}\_\{\\ell\}take the form \(Cℓ\(m\)\)TCℓ\(m\)=\(\(Hℓ\(m\)\)THℓ\(m\)\(Hℓ\(m\)\)TΨℓ\(m\)\(Ψℓ\(m\)\)THℓ\(m\)\(Ψℓ\(m\)\)TΨℓ\(m\)\)and\(Cℓ\(m\)\)TVℓ\(m\)=\(\(Hℓ\(m\)\)TVℓ\(m\)\(Ψℓ\(m\)\)TVℓ\(m\)\)\.\\displaystyle\\begin\{split\}\(C^\{\(m\)\}\_\{\\ell\}\)^\{T\}C^\{\(m\)\}\_\{\\ell\}=\\begin\{pmatrix\}\(H^\{\(m\)\}\_\{\\ell\}\)^\{T\}H^\{\(m\)\}\_\{\\ell\}&\(H^\{\(m\)\}\_\{\\ell\}\)^\{T\}\\Psi^\{\(m\)\}\_\{\\ell\}\\\\ \(\\Psi^\{\(m\)\}\_\{\\ell\}\)^\{T\}H^\{\(m\)\}\_\{\\ell\}&\(\\Psi^\{\(m\)\}\_\{\\ell\}\)^\{T\}\\Psi^\{\(m\)\}\_\{\\ell\}\\end\{pmatrix\}\\quad\\text\{and\}\\quad\(C^\{\(m\)\}\_\{\\ell\}\)^\{T\}V^\{\(m\)\}\_\{\\ell\}=\\begin\{pmatrix\}\(H^\{\(m\)\}\_\{\\ell\}\)^\{T\}V^\{\(m\)\}\_\{\\ell\}\\\\ \(\\Psi^\{\(m\)\}\_\{\\ell\}\)^\{T\}V^\{\(m\)\}\_\{\\ell\}\\end\{pmatrix\}\.\\end\{split\}\(38\) #### 2\.1\.4Hypothesis space selection As discussed in Section[2\.1\.2](https://arxiv.org/html/2608.25181#S2.SS1.SSS2), estimators are determined within fixed hypothesis spacesℋϕ\\mathcal\{H\}\_\{\\phi\}andℋf\\mathcal\{H\}\_\{f\}by minimizing the error functional \([17](https://arxiv.org/html/2608.25181#S2.E17)\)\. If the hypothesis spaces are finite\-dimensional, minimizing \([17](https://arxiv.org/html/2608.25181#S2.E17)\) is equivalent to minimizing \([27](https://arxiv.org/html/2608.25181#S2.E27)\), i\.e\. to solving a least\-squares problem, and in this section, we describe typical constructions for the spacesℋϕ\\mathcal\{H\}\_\{\\phi\}andℋf\\mathcal\{H\}\_\{f\}\. Motivated by the Weierstrass approximation theorem, we seek relatively simple basis functions with localized support; intuitively, altering a single coefficient in the estimation will thus only affect the approximation locally, making the estimation procedure robust to both outlier data and noise\. Consider firstℋϕ\\mathcal\{H\}\_\{\\phi\}, the hypothesis space approximating the \(true\) interaction kernelϕ\\phi\. Define Rmin:=minm,ℓ,1≤i<j≤N\(Rℓ\(m\)\)ijandRmax:=maxm,ℓ,1≤i<j≤N\(Rℓ\(m\)\)ij\\displaystyle R\_\{\\min\}:=\\min\_\{m,\\ell,1\\leq i<j\\leq N\}\(R^\{\(m\)\}\_\{\\ell\}\)\_\{ij\}\\quad\\text\{and\}\\quad R\_\{\\max\}:=\\max\_\{m,\\ell,1\\leq i<j\\leq N\}\(R^\{\(m\)\}\_\{\\ell\}\)\_\{ij\}to be the minimum and maximum, respectively, of pairwise distances obtained from the observed trajectory data as discussed in Section[2\.1\.1](https://arxiv.org/html/2608.25181#S2.SS1.SSS1)\. Thus, all observed pairwise distances are within the interval\[Rmin,Rmax\]\[R\_\{\\min\},R\_\{\\max\}\], which serves as the effective domain on whichϕ\\phican be learned from the data\. We then partition\[Rmin,Rmax\]\[R\_\{\\min\},R\_\{\\max\}\]into sub\-intervals based on the observed data; this is analogous to selecting bins in a histogram or choosing knots for a spline approximation: the partition should be fine enough to resolve variation inϕ\\phi, but coarse enough to ensure that each interval contains sufficient observations\. In practice, this partition may be chosen using standard data\-dependent histogram/bin\-width rules, such as Scott’s rule or the Freedman\-Diaconis rule, or by placing breakpoints at quantiles of the empirical pairwise distance distributionρ^R\\hat\{\\rho\}\_\{R\}discussed below in Section[3\.1](https://arxiv.org/html/2608.25181#S3.SS1)\[[62](https://arxiv.org/html/2608.25181#bib.bib79),[63](https://arxiv.org/html/2608.25181#bib.bib80),[64](https://arxiv.org/html/2608.25181#bib.bib81)\]\. For each element of the partition of\[Rmin,Rmax\]\[R\_\{\\min\},R\_\{\\max\}\], we then define a set of functions whose support is precisely that element of the partition; standard choices are indicator functions for the partition, or localized low\-degree polynomials\. For example, if the partition element is of the formIp:=\[rp,rp\+1\)⊆\[Rmin,Rmax\]I\_\{p\}:=\[r\_\{p\},r\_\{p\+1\}\)\\subseteq\[R\_\{\\min\},R\_\{\\max\}\], then we may define the basis of localized polynomials toIpI\_\{p\}of degree at most two as\{𝟏Ip\(r\),r𝟏Ip\(r\),r2𝟏Ip\(r\)\}\\\{\\mathbf\{1\}\_\{I\_\{p\}\}\(r\),r\\mathbf\{1\}\_\{I\_\{p\}\}\(r\),r^\{2\}\\mathbf\{1\}\_\{I\_\{p\}\}\(r\)\\\}\. We then extend this basis naturally to all partitions of\[Rmin,Rmax\]\[R\_\{\\min\},R\_\{\\max\}\]to obtain the basis\(ψk\)k=1nϕ\(\\psi\_\{k\}\)\_\{k=1\}^\{n\_\{\\phi\}\}, withℋϕ\\mathcal\{H\}\_\{\\phi\}defined as the span of this set of functions\. Note that if there arePPpartitions of\[Rmin,Rmax\]\[R\_\{\\min\},R\_\{\\max\}\]andnPn\_\{P\}number of localized basis functions, thennϕ=P×nPn\_\{\\phi\}=P\\times n\_\{P\}\. As localized polynomials are in general discontinuous, the resulting approximation is also discontinuous, as a linear combination of discontinuous functions \(see \([19](https://arxiv.org/html/2608.25181#S2.E19)\)\)\. If it known that the interaction kernel is continuous, it may be natural to select a set of continuous basis functions; an example of such a choice is B\-splines, which are utilized frequently in this work\. We refer the interested reader to the supplementary material of\[[46](https://arxiv.org/html/2608.25181#bib.bib6)\]for a detailed discussion regarding the choice of basis functions\. The same localization principle applies to the construction of the hypothesis spaceℋf\\mathcal\{H\}\_\{f\}for variables associated with the environmental forcing termff\. For the first\-order system \([4](https://arxiv.org/html/2608.25181#S2.E4)\),f=f\(x\)f=f\(x\), wherex∈ℝdx\\in\\mathbb\{R\}^\{d\}; note the same principle applies for second\-order systems, but in general in this casef=f\(x,v\)f=f\(x,v\), i\.e\.\(x,v\)∈ℝd×ℝd\(x,v\)\\in\\mathbb\{R\}^\{d\}\\times\\mathbb\{R\}^\{d\}\. Specifically, for first\-order systems, the data of observed states\(Xℓ\(m\)\)ij∈ℝd\(X^\{\(m\)\}\_\{\\ell\}\)\_\{ij\}\\in\\mathbb\{R\}^\{d\}are partitioned intodd\-dimensional hyper\-rectangles\. For example, whend=2d=2, the state data is partitioned into rectangles of the form\[x1,p,x1,p\+1\)×\[x2,p,x2,p\+1\)\[x\_\{1,p\},x\_\{1,p\+1\}\)\\times\[x\_\{2,p\},x\_\{2,p\+1\}\)\. Localized basis functions are constructed on these hyper\-rectangular partition elements, with each function defining a specifichkh\_\{k\}, and the hypothesis spaceℋf\\mathcal\{H\}\_\{f\}is thus defined as the span of this collection of localized basis functions\. Asf∈ℱ\(ℝd,ℝd\)f\\in\\mathcal\{F\}\(\\mathbb\{R\}^\{d\},\\mathbb\{R\}^\{d\}\), each basis functionhk∈ℱ\(ℝd,ℝd\)h\_\{k\}\\in\\mathcal\{F\}\(\\mathbb\{R\}^\{d\},\\mathbb\{R\}^\{d\}\)of the hypothesis spaceℋf\\mathcal\{H\}\_\{f\}isℝd\\mathbb\{R\}^\{d\}valued, i\.e\. is a function of the formhk=\(hk1,hk2,…,hkd\)h\_\{k\}=\(h^\{1\}\_\{k\},h^\{2\}\_\{k\},\\ldots,h^\{d\}\_\{k\}\), wherehkj:ℝd→ℝh^\{j\}\_\{k\}:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}\. Thus, eachhkbh^\{b\}\_\{k\}is a localized scalar\-valued function defined on a partition element ofℝd\\mathbb\{R\}^\{d\}; again, we may take localized multivariate \(withddbeing the number of variates\) polynomials of small degree, or the tensor product of B\-splines\. In practice, we thus construct a localized scalar basis, which we assume identical for each partition, and then take the \(pairwise\) tensor product of this basis with the standard basis ofℝd\\mathbb\{R\}^\{d\}to extend this to a set of functions inℱ\(ℝd,ℝd\)\\mathcal\{F\}\(\\mathbb\{R\}^\{d\},\\mathbb\{R\}^\{d\}\), which defines\(hk\)k=1nf\(h\_\{k\}\)\_\{k=1\}^\{n\_\{f\}\}and thusℋf\\mathcal\{H\}\_\{f\}as the span of these functions\. #### 2\.1\.5Algorithm Algorithm 1Algorithm for learningϕ\\phi\(interaction\) andff\(environmental\) forces from trajectory data of first\-order systems of the form \([4](https://arxiv.org/html/2608.25181#S2.E4)\)\.Input: \(Xℓ\(m\),Vℓ\(m\)\)m,ℓ=1M,L\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\} Output:Estimators: ϕ^\\hat\{\\phi\}, f^\\hat\{f\} Construct pairwise quantities and estimate the observed supports Generate interaction distances \(Rℓ\(m\)\)m,ℓ=1M,L\(R^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}Find the maximum and minimum interaction radii Rmax,Rmin∈ℝR\_\{\\max\},R\_\{\\min\}\\in\\mathbb\{R\} Find the component\-wrise maximum and minimum values of state variable Xmax,Xmin∈ℝdX\_\{\\max\},X\_\{\\min\}\\in\\mathbb\{R\}^\{d\} Observed supports \[Rmin,Rmax\],\[Xmin,Xmax\]\[R\_\{\\min\},R\_\{\\max\}\],\[X\_\{\\min\},X\_\{\\max\}\] Construct localized basis functions Construct the kernel basis \(ψk\)k=1nϕ\(\\psi\_\{k\}\)\_\{k=1\}^\{n\_\{\\phi\}\} Construct \(hk\)k=1nf\(h\_\{k\}\)\_\{k=1\}^\{n\_\{f\}\} Create Cℓ\(m\)C^\{\(m\)\}\_\{\\ell\}by column\-wise concatenating \(Hℓ\(m\)Ψℓ\(m\)\)\(\\,H^\{\(m\)\}\_\{\\ell\}\\quad\\Psi^\{\(m\)\}\_\{\\ell\}\\,\)as in \([30](https://arxiv.org/html/2608.25181#S2.E30)\) Assemble Aθ^=bA\\hat\{\\theta\}=b\(in parallel for ℓ,m\\ell,m\), using \([35](https://arxiv.org/html/2608.25181#S2.E35)\), \([36](https://arxiv.org/html/2608.25181#S2.E36)\), \([37](https://arxiv.org/html/2608.25181#S2.E37)\), \([29](https://arxiv.org/html/2608.25181#S2.E29)\) Solve for θ^\\hat\{\\theta\} Assemble f^\\hat\{f\}and ϕ^\\hat\{\\phi\}as in \([19](https://arxiv.org/html/2608.25181#S2.E19)\) #### 2\.1\.6Computational complexity The proposed learning framework, described in Sections[2\.1\.2](https://arxiv.org/html/2608.25181#S2.SS1.SSS2)and[2\.1\.3](https://arxiv.org/html/2608.25181#S2.SS1.SSS3), is naturally parallelizable across both replicates and observation times\. In general, the data utilized in training is𝒪\(MLD\)\\mathcal\{O\}\(MLD\), whereD=dND=dNfor first\-order systems andD=2dND=2dNfor second\-order systems, as there existMMreplicates \(initial conditions\) ofLLtime observations, with each observation an element ofℝD\\mathbb\{R\}^\{D\}\. However, the assembly of the least\-squares system \([36](https://arxiv.org/html/2608.25181#S2.E36)\) is performed independently across both replicates and times \(i\.e\. eachCℓ\(m\)C^\{\(m\)\}\_\{\\ell\}may be constructed independently across different computational cores\)\. For each replicatemmand observation timetℓt\_\{\\ell\}, all pairwise interaction distancesRℓ\(m\)R^\{\(m\)\}\_\{\\ell\}must be computed to constructΨℓ\(m\)\\Psi^\{\(m\)\}\_\{\\ell\}, which requires𝒪\(N2\)\\mathcal\{O\}\(N^\{2\}\)operations, whereNNis the number of agents, and thus requires \(across allMMreplicates andLLobservation times\)𝒪\(MLN2\)\\mathcal\{O\}\(MLN^\{2\}\)operations\. Letn:=nϕ\+nfn:=n\_\{\\phi\}\+n\_\{f\}be the number of basis functions \(corresponding toℋϕ\\mathcal\{H\}\_\{\\phi\}andℋf\\mathcal\{H\}\_\{f\}\)\. DistancesRℓ\(m\)R^\{\(m\)\}\_\{\\ell\}and stateXℓ\(m\)X^\{\(m\)\}\_\{\\ell\}must be evaluated on thennbasis functions, which yields an additional total complexity of𝒪\(MLn\)\\mathcal\{O\}\(MLn\)\. Note that we are assuming that the complexity of evaluation of each basis function isO\(1\)O\(1\); regardless, this cost is fixed when a set of basis functions forℋϕ\\mathcal\{H\}\_\{\\phi\}andℋf\\mathcal\{H\}\_\{f\}is selected\. Thus, the construction ofAAandbbin \([36](https://arxiv.org/html/2608.25181#S2.E36)\) requires𝒪\(MLN2\+MLn\)\\mathcal\{O\}\(MLN^\{2\}\+MLn\)operations, with the pairwise\-distance computations dominating in practice \(i\.e\.n≪N2n\\ll N^\{2\}\), so that we assume𝒪\(MLN2\+MLn\)=𝒪\(MLN2\)\\mathcal\{O\}\(MLN^\{2\}\+MLn\)=\\mathcal\{O\}\(MLN^\{2\}\)\. Solving the resulting linear system viaQRQRfactorization or singular\-value decomposition requires𝒪\(n3\)\\mathcal\{O\}\(n^\{3\}\)operations\. Therefore, the overall computational complexity of the learning algorithm is𝒪\(MLN2\+n3\)\\mathcal\{O\}\(MLN^\{2\}\+n^\{3\}\)\. Note that since the basis dimensionnnis typically several orders of magnitude smaller than the number of observations, the dominant computational cost arises from matrix construction rather than solving the resulting linear systemAθ^=bA\\hat\{\\theta\}=b\. New trajectories may also be incorporated incrementally by updating the accumulated matrix/vector in the normal equations, thus making the framework naturally amenable to online\-learning scenarios\. ### 2\.2Variational methods for learning collective and environmental forces: semi\-parametric approach \(first\-order\) In contrast to the fully non\-parametric learning approach described in Section[2\.1](https://arxiv.org/html/2608.25181#S2.SS1), in certain physical scenarios we may have knowledge of the functional form of the environmental forces acting on the collective system\. More precisely, in this case we assume that there exists a prescribed functional form for the environmental force termffin \([4](https://arxiv.org/html/2608.25181#S2.E4)\), which we generally write f=f\(x,p\),\\displaystyle f=f\(x;p\),\(39\)wherep∈ℝnpp\\in\\mathbb\{R\}^\{n\_\{p\}\}is a vector of parameters which are needed to fully specifyff\. The task of estimatingffthus reduces to estimating the parameter vectorppfrom trajectory data\. We continue to estimate the interaction kernelϕ\\phinon\-parametrically as described in Section[2\.1](https://arxiv.org/html/2608.25181#S2.SS1), and thus we term this approachsemi\-parametric\. We assume the same training data notation as described in Section[2\.1\.1](https://arxiv.org/html/2608.25181#S2.SS1.SSS1), and all results presented comparing the fully non\-parametric and semi\-parametric approaches will utilize identical samples from initial conditions\. Analogously to the algorithm presented in Section[2\.1\.2](https://arxiv.org/html/2608.25181#S2.SS1.SSS2), we define an error functional of the form ℰℋϕ\(p,ϕ~\)\\displaystyle\\mathcal\{E\}\_\{\\mathcal\{H\}\_\{\\phi\}\}\(p,\\tilde\{\\phi\}\):=1ML∑m,ℓ=1M,L‖Vℓ\(m\)−\(Ff\(Xℓ\(m\),p\)\+Fϕ~\(Xℓ\(m\)\)\)‖ℝdN2,\\displaystyle:=\\frac\{1\}\{ML\}\\sum\_\{m,\\ell=1\}^\{M,L\}\\left\\lVert V^\{\(m\)\}\_\{\\ell\}\-\\Big\(F\_\{f\}\(X^\{\(m\)\}\_\{\\ell\};p\)\+F\_\{\\tilde\{\\phi\}\}\(X^\{\(m\)\}\_\{\\ell\}\)\\Big\)\\right\\rVert\_\{\\mathbb\{R\}^\{dN\}\}^\{2\},\(40\)whereFf\(x,p\):=\(f\(x1,p\),f\(x2,p\),…,f\(xN,p\)\)∈ℝdNF\_\{f\}\(x;p\):=\(f\(x\_\{1\};p\),f\(x\_\{2\};p\),\\ldots,f\(x\_\{N\};p\)\)\\in\\mathbb\{R\}^\{dN\}forx∈ℝdNx\\in\\mathbb\{R\}^\{dN\}as in \([6](https://arxiv.org/html/2608.25181#S2.E6)\)\. As before,ℋϕ\\mathcal\{H\}\_\{\\phi\}is a fixed hypothesis space, and once a basis\(ψk\)k=1nϕ\(\\psi\_\{k\}\)\_\{k=1\}^\{n\_\{\\phi\}\}is chosen, the error functional \([40](https://arxiv.org/html/2608.25181#S2.E40)\) takes the form ℰ\(p,β\)\\displaystyle\\mathcal\{E\}\(p,\\beta\):=1ML∑m,ℓ=1M,L‖Vℓ\(m\)−\(Ff\(Xℓ\(m\),p\)\+Ψℓ\(m\)β\)‖ℝdN2,\\displaystyle:=\\frac\{1\}\{ML\}\\sum\_\{m,\\ell=1\}^\{M,L\}\\left\\lVert V^\{\(m\)\}\_\{\\ell\}\-\\Big\(F\_\{f\}\(X^\{\(m\)\}\_\{\\ell\};p\)\+\\Psi^\{\(m\)\}\_\{\\ell\}\\beta\\Big\)\\right\\rVert\_\{\\mathbb\{R\}^\{dN\}\}^\{2\},\(41\)whereϕ~=∑k=1nϕβkψk\\tilde\{\\phi\}=\\sum\_\{k=1\}^\{n\_\{\\phi\}\}\\beta\_\{k\}\\psi\_\{k\}andΨℓ\(m\)\\Psi^\{\(m\)\}\_\{\\ell\}is given by \([25](https://arxiv.org/html/2608.25181#S2.E25)\)\. We then minimizeℰ\(p,β\)\\mathcal\{E\}\(p,\\beta\)with respect to\(p,β\)∈ℝnp×ℝnϕ\(p,\\beta\)\\in\\mathbb\{R\}^\{n\_\{p\}\}\\times\\mathbb\{R\}^\{n\_\{\\phi\}\}\(note that if certain constraints are known for the parameterspp, then we minimize over a subset ofℝnp\\mathbb\{R\}^\{n\_\{p\}\}\)\. In general, this is non\-convex optimization problem, as the vector fieldffmay depend nonlinearly onpp\(e\.gf\(x,p\)=x/\(x\+p\)f\(x;p\)=x/\(x\+p\)\), and existence, uniqueness, and computation of minimizers \(global and local\) may be non\-trivial\[[65](https://arxiv.org/html/2608.25181#bib.bib89),[66](https://arxiv.org/html/2608.25181#bib.bib90)\]\. We note that the structure of \([41](https://arxiv.org/html/2608.25181#S2.E41)\) can be exploited computationally\. Although the objective is generally non\-quadratic inpp, it is quadratic in the coefficientsβ\\beta, and a “variable projection” or separable nonlinear least\-squares formulation reduces the dimension of the nonlinear optimization problem and is often preferable to optimizing over\(p,β\)\(p,\\beta\)simultaneously\[[67](https://arxiv.org/html/2608.25181#bib.bib92),[68](https://arxiv.org/html/2608.25181#bib.bib94),[69](https://arxiv.org/html/2608.25181#bib.bib93)\]\. The reduced problem may then be solved using standard methods for nonlinear least squares, such as Gauss–Newton, Levenberg–Marquardt, or trust\-region methods, typically with multiple initial guesses when non\-convexity is expected\[[70](https://arxiv.org/html/2608.25181#bib.bib91),[66](https://arxiv.org/html/2608.25181#bib.bib90),[71](https://arxiv.org/html/2608.25181#bib.bib95)\]\. We note that ifffis linear in the parameterspp\(i\.e\.f=∑k=1npηkpkf=\\sum\_\{k=1\}^\{n\_\{p\}\}\\eta\_\{k\}p\_\{k\}, whereηk:ℝd→ℝd\\eta\_\{k\}:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{d\}are known and fixed\), then \([41](https://arxiv.org/html/2608.25181#S2.E41)\) is indeed quadratic in\(p,β\)\(p,\\beta\), and we can constructHℓ\(m\)H^\{\(m\)\}\_\{\\ell\}analogously as in \([21](https://arxiv.org/html/2608.25181#S2.E21)\) to obtain a set of normal equations \([37](https://arxiv.org/html/2608.25181#S2.E37)\) for the estimators\(p^,β^\)\(\\hat\{p\},\\hat\{\\beta\}\)\. We note that all the systems considered in this work are linear in parameters, so that we obtain that we obtain a linear least\-squares problem even in the semi\-parametric case, and hence do not discuss issues related to non\-convex optimization\. See Sections[5\.2](https://arxiv.org/html/2608.25181#S5.SS2)and[5\.6](https://arxiv.org/html/2608.25181#S5.SS6)for details on the models considered in this manuscript\. ## 3Evaluation of estimators In this section we describe various methods to evaluate the effectiveness of the learning methods presented in Sections[2](https://arxiv.org/html/2608.25181#S2)and Appendix[A](https://arxiv.org/html/2608.25181#A1)\. Our goal is to assess the estimated model with respect to both model features \(i\.e\. interaction and environmental mechanisms\) and trajectory predictions\. The proposed methodology quantifies the performance of the inference in various capacities, including feature recovery, consistency of governing equations, trajectory accuracy, statistical variability with respect to observation data, scalability with respect to the number of agents, and robustness to observational noise\. Collectively, these metrics provide a comprehensive assessment of the identifiability, predictive capacity, consistency, and practical applicability of the proposed variational learning framework\. Throughout this section, we define metrics for the first\-order system \([4](https://arxiv.org/html/2608.25181#S2.E4)\), and discuss natural extensions to the second\-order system \([85](https://arxiv.org/html/2608.25181#A1.E85)\)\. We will denote byffandϕ\\phithe true \(but unknown\) environmental and interaction forces, and byf~\\tilde\{f\}andϕ~\\tilde\{\\phi\}their respective estimators\. Note that we define following evaluation metrics with respect toanypair of functionsf~\\tilde\{f\}andϕ~\\tilde\{\\phi\}, but results in Section[5](https://arxiv.org/html/2608.25181#S5)will generally be shown for the estimators defined by the variational approaches, as discussed in Sections[2\.1](https://arxiv.org/html/2608.25181#S2.SS1),[2\.2](https://arxiv.org/html/2608.25181#S2.SS2), and Appendix[A](https://arxiv.org/html/2608.25181#A1); e\.g\. forf^\\hat\{f\}andϕ^\\hat\{\\phi\}defined by f^=∑k=1nfα^khkandϕ^=∑k=1nϕβ^kψk,\\displaystyle\\begin\{split\}\\hat\{f\}&=\\sum\_\{k=1\}^\{n\_\{f\}\}\\hat\{\\alpha\}\_\{k\}h\_\{k\}\\quad\\text\{and\}\\quad\\hat\{\\phi\}=\\sum\_\{k=1\}^\{n\_\{\\phi\}\}\\hat\{\\beta\}\_\{k\}\\psi\_\{k\},\\end\{split\}\(42\)where\(α^,β^\)\(\\hat\{\\alpha\},\\hat\{\\beta\}\)are given by \([28](https://arxiv.org/html/2608.25181#S2.E28)\) for the fully non\-parametric approach, or by minimization of \([41](https://arxiv.org/html/2608.25181#S2.E41)\) for the semi\-parametric approach\. ### 3\.1Induced measures on data The randomnessμ0\\mu\_\{0\}in the initial conditions of \([9](https://arxiv.org/html/2608.25181#S2.E9)\) implies that the observed trajectory data \(statesxxand velocitiesx˙\\dot\{x\}\) are random variables, and thus that there exists distributions both on the state spacexxand thus also the pairwise distancesrrinduced by the solutions of the \(generally nonlinear\) IVP \([9](https://arxiv.org/html/2608.25181#S2.E9)\)\. As our goal is to estimatef=f\(x\)f=f\(x\)andϕ=ϕ\(r\)\\phi=\\phi\(r\), it is natural to measure the accuracy of these estimators with respect to these distributions\. Indeed, one can generally only obtain accurate estimates on regions of state and pairwise distance space that are explored by the solutions of \([9](https://arxiv.org/html/2608.25181#S2.E9)\); regions that are highly explored are thus prioritized when determining accuracy as they contain a larger number of observations\. Hence, following the constructions in\[[46](https://arxiv.org/html/2608.25181#bib.bib6)\], we consider the expected empirical measures for continuous\-time observations and their counterparts for discrete\-time observations, the latter of which are utilized numerically\. Consider the IVP x˙=Ff\(x\)\+Fϕ\(x\),x\(0\)=x0,\\displaystyle\\begin\{split\}\\dot\{x\}&=F\_\{f\}\(x\)\+F\_\{\\phi\}\(x\),\\quad x\(0\)=x\_\{0\},\\end\{split\}\(43\)wherex0∼μ0x\_\{0\}\\sim\\mu\_\{0\}andμ0\\mu\_\{0\}is a fixed distribution onℝdN\\mathbb\{R\}^\{dN\}\. Note that this is the same notation as introduced previously in \([9](https://arxiv.org/html/2608.25181#S2.E9)\), but here are emphasizing the randomness in the solutionx=x\(t,x0\)=x\(t,x0\(ω\)\)x=x\(t,x\_\{0\}\)=x\(t,x\_\{0\}\(\\omega\)\)induced by samplingx0x\_\{0\}fromμ0\\mu\_\{0\}; the notationx0\(ω\)x\_\{0\}\(\\omega\)denotes the dependence ofx0x\_\{0\}on given sampleω\\omega\. Assume that the solution exists on an interval\[0,T\]\[0,T\]\. Since the solutionxxconsists ofNNcomponents inℝd\\mathbb\{R\}^\{d\}\(one for each agent; see \([5](https://arxiv.org/html/2608.25181#S2.E5)\)\), we obtain\(N2\)\\binom\{N\}\{2\}unique pairwise distance trajectories, rij\(t,ω\)\\displaystyle r\_\{ij\}\(t,\\omega\):=\|xj\(t,x0\(ω\)\)−xi\(t,x0\(ω\)\)\|,\\displaystyle:=\|x\_\{j\}\(t,x\_\{0\}\(\\omega\)\)\-x\_\{i\}\(t,x\_\{0\}\(\\omega\)\)\|,\(44\)where1≤i,j≤N1\\leq i,j\\leq N\. The distribution of\(rij\)i,j=1N\(r\_\{ij\}\)\_\{i,j=1\}^\{N\}naturally induces a measure onℝ≥0\\mathbb\{R\}\_\{\\geq 0\}of average pairwise distances: ρR\(A\):=1\(N2\)T∫0T𝔼x0∼μ0\[∑1≤i<j≤N𝟏A\(rij\(t,ω\)\)\]𝑑t\\displaystyle\\rho\_\{R\}\(A\):=\\frac\{1\}\{\\binom\{N\}\{2\}\\,T\}\\int\_\{0\}^\{T\}\\mathbb\{E\}\_\{x\_\{0\}\\sim\\mu\_\{0\}\}\\left\[\\sum\_\{1\\leq i<j\\leq N\}\\mathbf\{1\}\_\{A\}\(r\_\{ij\}\(t,\\omega\)\)\\right\]\\,\\mathrm\{d\}t\(45\)for any Borel setA⊆ℝ≥0A\\subseteq\\mathbb\{R\}\_\{\\geq 0\}, where𝟏A\\mathbf\{1\}\_\{A\}denotes the indicator function ofAA\. Intuitively,ρR\(A\)\\rho\_\{R\}\(A\)measures the time\-averaged expected fraction of observed trajectories whose pairwise distance lies inAA\. Assume now that we have measurement data as discussed in Section[2\.1\.1](https://arxiv.org/html/2608.25181#S2.SS1.SSS1), i\.e\.MMsamples of initial conditionsm=1,2,…,Mm=1,2,\\ldots,Mevaluated on time interval\[0,T\]\[0,T\]at pointstℓt\_\{\\ell\}forℓ=1,2,…,L\\ell=1,2,\\ldots,L\. Simultaneously discretizing both the integration interval\[0,T\]\[0,T\]\(viaΔt:=T/L\\Delta t:=T/L\) and utilizing the law of large numbers to estimate the expectation overx0∼μ0x\_\{0\}\\sim\\mu\_\{0\}in \([45](https://arxiv.org/html/2608.25181#S3.E45)\), we obtain the empirical version of \([45](https://arxiv.org/html/2608.25181#S3.E45)\): ρ^R\(A\)\\displaystyle\\hat\{\\rho\}\_\{R\}\(A\):=1\(N2\)ML∑m,ℓ=1M,L\(∑1≤i<j≤N𝟏A\(\(Rℓ\(m\)\)ij\)\)\.\\displaystyle:=\\frac\{1\}\{\\binom\{N\}\{2\}\\,ML\}\\sum\_\{m,\\ell=1\}^\{M,L\}\\left\(\\sum\_\{1\\leq i<j\\leq N\}\\mathbf\{1\}\_\{A\}\(\(R^\{\(m\)\}\_\{\\ell\}\)\_\{ij\}\)\\right\)\.\(46\)Recall from \([14](https://arxiv.org/html/2608.25181#S2.E14)\) that\(Rℓ\(m\)\)ij\(R^\{\(m\)\}\_\{\\ell\}\)\_\{ij\}denotes the observed distance between agentsiiandjjof sample \(replicate\)mmat timetℓt\_\{\\ell\}\. As in \([45](https://arxiv.org/html/2608.25181#S3.E45)\),AAdenotes a Borel subset ofℝ≥0\\mathbb\{R\}\_\{\\geq 0\}\. Note that in typical examples, the true measureρR\\rho\_\{R\}is unobservable, so that for all computations demonstrated in the manuscript, the approximationρ^R\\hat\{\\rho\}\_\{R\}will be employed\. We analogously define an induced measure on the state variable for determining the accuracy of the environmental forceffin \([43](https://arxiv.org/html/2608.25181#S3.E43)\); note thatf=f\(x\)f=f\(x\)\. Recalling the notationx\(t,x0\(ω\)\)=\(xi\(t,x0\(ω\)\)\)i=1Nx\(t,x\_\{0\}\(\\omega\)\)=\(x\_\{i\}\(t,x\_\{0\}\(\\omega\)\)\)\_\{i=1\}^\{N\}for a solution of \([43](https://arxiv.org/html/2608.25181#S3.E43)\), we define the induced average state measure ρX\(B\)\\displaystyle\\rho\_\{X\}\(B\):=1TN∫0T𝔼x0∼μ0\[∑i=1N𝟏B\(xi\(t,x0\(ω\)\)\)\]𝑑t,\\displaystyle:=\\frac\{1\}\{TN\}\\int\_\{0\}^\{T\}\\mathbb\{E\}\_\{x\_\{0\}\\sim\\mu\_\{0\}\}\\left\[\\sum\_\{i=1\}^\{N\}\\mathbf\{1\}\_\{B\}\(x\_\{i\}\(t,x\_\{0\}\(\\omega\)\)\)\\right\]\\,\\mathrm\{d\}t,\(47\)whereBBis a Borel set inℝd\\mathbb\{R\}^\{d\}, and its empirical approximation ρ^X\(B\)\\displaystyle\\hat\{\\rho\}\_\{X\}\(B\):=1MLN∑m,ℓ,i=1M,L,N𝟏B\(\(Xℓ\(m\)\)i\)\.\\displaystyle:=\\frac\{1\}\{MLN\}\\sum\_\{m,\\ell,i=1\}^\{M,L,N\}\\mathbf\{1\}\_\{B\}\(\(X^\{\(m\)\}\_\{\\ell\}\)\_\{i\}\)\.\(48\)As withρR\(A\)\\rho\_\{R\}\(A\),ρX\(B\)\\rho\_\{X\}\(B\)is time\-averaged expected fraction of observed trajectories whose state lies inBB\. Here\(Xℓ\(m\)\)i∈ℝd\(X^\{\(m\)\}\_\{\\ell\}\)\_\{i\}\\in\\mathbb\{R\}^\{d\}denotes the observation of agentiiat timetℓt\_\{\\ell\}of sample \(replicate\)mm; see equation \([11](https://arxiv.org/html/2608.25181#S2.E11)\)\. For second\-order systems \(as discussed in Appendix[A](https://arxiv.org/html/2608.25181#A1)\), the environmental forcingffmay depend on both position and velocity \(f=f\(x,v\)f=f\(x,v\); see \([85](https://arxiv.org/html/2608.25181#A1.E85)\)\), in which case we writez:=\(x,v\)z:=\(x,v\), and we extend the measureρZ\\rho\_\{Z\}to Borel sets inℝd×ℝd\\mathbb\{R\}^\{d\}\\times\\mathbb\{R\}^\{d\}in the natural way\. As forρR\\rho\_\{R\}, in practice computations are performed via the empirical approximationρ^X\\hat\{\\rho\}\_\{X\}\(orρ^Z\\hat\{\\rho\}\_\{Z\}, for second\-order systems\), but for notational simplicity, we typically define expressions involving induced measures viaρR\\rho\_\{R\}andρX\\rho\_\{X\}\. Specifically, as there does not exist a universal criterion for how largeMMneeds to be for accurate estimation ofρR\\rho\_\{R\}andρX\\rho\_\{X\}\(viaρ^R\\hat\{\\rho\}\_\{R\}andρ^X\\hat\{\\rho\}\_\{X\}, respectively\), we approximateρR\\rho\_\{R\}andρX\\rho\_\{X\}using datasets generated with large values ofMM\(denoted byMρM\_\{\\rho\}\), and compare them with the empirical estimatorsρ^R\\hat\{\\rho\}\_\{R\}andρ^X\\hat\{\\rho\}\_\{X\}obtained from the training data \(denoted byMtrainM\_\{\\text\{train\}\}\); see Table[22](https://arxiv.org/html/2608.25181#A5.T22)for parameter values utilized for the various example systems studied in this manuscript\. The approximation errors of the learned estimators are then evaluated in the weighted least\-squares spacesL2\(ρR\)L^\{2\}\(\\rho\_\{R\}\)andL2\(ρX\)L^\{2\}\(\\rho\_\{X\}\)via∥ϕ~−ϕ∥L2\(ρR\)\\lVert\\tilde\{\\phi\}\-\\phi\\rVert\_\{L^\{2\}\(\\rho\_\{R\}\)\}, where ∥ψ∥L2\(ρR\):=\(∫ℝ≥0\|ψ\(r\)\|2r2dρR\(r\)\)1/2and∥h∥L2\(ρX\):=\(∫ℝd\|h\(x\)\|2dρX\(x\)\)1/2\\displaystyle\\begin\{split\}\\lVert\\psi\\rVert\_\{L^\{2\}\(\\rho\_\{R\}\)\}&:=\\left\(\\int\_\{\\mathbb\{R\}\\geq 0\}\|\\psi\(r\)\|^\{2\}r^\{2\}\\,\\mathrm\{d\}\\rho\_\{R\}\(r\)\\right\)^\{1/2\}\\quad\\text\{and\}\\quad\\lVert h\\rVert\_\{L^\{2\}\(\\rho\_\{X\}\)\}:=\\left\(\\int\_\{\\mathbb\{R\}^\{d\}\}\|h\(x\)\|^\{2\}\\,\\mathrm\{d\}\\rho\_\{X\}\(x\)\\right\)^\{1/2\}\\end\{split\}\(49\)Note the factor ofr2r^\{2\}in∥⋅∥L2\(ρR\)\\lVert\\cdot\\rVert\_\{L^\{2\}\(\\rho\_\{R\}\)\}, which arises due to the assumed form of the collective force in \([4](https://arxiv.org/html/2608.25181#S2.E4)\)\. An extended discussion of methods for quantifying errors is also provided in Section[3\.2](https://arxiv.org/html/2608.25181#S3.SS2)\. As previously discussed, the weightedL2L^\{2\}norms in \([49](https://arxiv.org/html/2608.25181#S3.E49)\) are estimated via the empirical measuresρ^R\\hat\{\\rho\}\_\{R\}andρ^X\\hat\{\\rho\}\_\{X\}, and for all calculations, these approximations are utilized\. In Appendix[B](https://arxiv.org/html/2608.25181#A2)\(see, for example, Figures[18](https://arxiv.org/html/2608.25181#A2.F18),[20](https://arxiv.org/html/2608.25181#A2.F20), and[22](https://arxiv.org/html/2608.25181#A2.F22)\), the observed distributions obtained from the training data closely align with the empirical approximations ofρR\\rho\_\{R\}andρX\\rho\_\{X\}, which suggests that the training data adequately capture the regions of the domain explored by the dynamics, making the use ofL2\(ρ^R\)L^\{2\}\(\\hat\{\\rho\}\_\{R\}\)andL2\(ρ^X\)L^\{2\}\(\\hat\{\\rho\}\_\{X\}\)reasonable for assessing estimation errors in simulations\. ### 3\.2Feature recovery metrics Feature recovery metrics quantify the accuracy with which the underlying interaction kernels and environmental forces are recovered by the variational \(non\-parametric and semi\-parametric\) estimators as discussed in Section[2](https://arxiv.org/html/2608.25181#S2)and Appendix[A](https://arxiv.org/html/2608.25181#A1)\. As these quantities are known only in synthetic experiments, the metrics in this section are primarily used to validate the ability of the learning algorithms to recover the true interaction and environmental forces, and cannot generally be utilized in experimental settings\. Nevertheless, they provide evidence that the introduced algorithms are able to mechanistically infer dynamics in collective systems\. As discussed in Section[3\.1](https://arxiv.org/html/2608.25181#S3.SS1), we bias the accuracy of estimation to regions highly explored by the dynamics of systems \([4](https://arxiv.org/html/2608.25181#S2.E4)\) or \([85](https://arxiv.org/html/2608.25181#A1.E85)\) via the metricsρR\\rho\_\{R\}\(see \([45](https://arxiv.org/html/2608.25181#S3.E45)\)\) andρX\\rho\_\{X\}\(see \([48](https://arxiv.org/html/2608.25181#S3.E48)\)\)\. We thus define the interaction force and environmental kernel errors as Eϕ\(ϕ~\):=‖ϕ~−ϕ‖L2\(ρR\)\\displaystyle E\_\{\\phi\}\(\\tilde\{\\phi\}\):=\\left\\lVert\\tilde\{\\phi\}\-\\phi\\right\\rVert\_\{L^\{2\}\(\\rho\_\{R\}\)\}\(50\)Ef\(f~\):=‖f~−f‖L2\(ρX\),\\displaystyle E\_\{f\}\(\\tilde\{f\}\):=\\left\\lVert\\tilde\{f\}\-f\\right\\rVert\_\{L^\{2\}\(\\rho\_\{X\}\)\},\(51\)respectively\. To facilitate comparisons across systems with different scales, we also report relative versions of the above errors, defined as Eϕrel\(ϕ~\):=Eϕ\(ϕ~\)‖ϕ‖L2\(ρR\)andEfrel\(f~\):=Ef\(f~\)‖f‖L2\(ρX\)\.\\displaystyle\\begin\{split\}E\_\{\\phi\}^\{\\mathrm\{rel\}\}\(\\tilde\{\\phi\}\)&:=\\frac\{E\_\{\\phi\}\(\\tilde\{\\phi\}\)\}\{\\left\\lVert\\phi\\right\\rVert\_\{L^\{2\}\(\\rho\_\{R\}\)\}\}\\quad\\text\{and\}\\quad E\_\{f\}^\{\\mathrm\{rel\}\}\(\\tilde\{f\}\):=\\frac\{E\_\{f\}\(\\tilde\{f\}\)\}\{\\left\\lVert f\\right\\rVert\_\{L^\{2\}\(\\rho\_\{X\}\)\}\}\.\\end\{split\}\(52\)Note that this error is theoretically independent of the observed data, but in practice is estimated viaρ^R\\hat\{\\rho\}\_\{R\}andρ^X\\hat\{\\rho\}\_\{X\}\. ### 3\.3Residual error While feature recovery metrics in Section[3\.2](https://arxiv.org/html/2608.25181#S3.SS2)assess recovery of the underlying interaction mechanisms, residual errors evaluate how well the estimator satisfies the observed data, i\.e\. it measures the discrepancy between the derivatives derived from the observation data \(velocity for a first\-order system, acceleration for a second\-order system\) and the predicted derivatives obtained from estimators \(e\.g\.ϕ~\\tilde\{\\phi\}andf~\\tilde\{f\}\)\. Unlike the previous mechanistic errors, residual errors are a function of the observed trajectory data\. For notational simplicity, consider the first\-order system \([4](https://arxiv.org/html/2608.25181#S2.E4)\)\. As discussed in Section[2\.1\.1](https://arxiv.org/html/2608.25181#S2.SS1.SSS1), the observed position and velocity data takes the form\(Xℓ\(m\),Vℓ\(m\)\)m,ℓ=1M,L\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}\. For estimatorsϕ~\\tilde\{\\phi\}andf~\\tilde\{f\}, the residual error is defined by Eres\(f~,ϕ~\):=\(1ML∑m,ℓ=1M,L‖Vℓ\(m\)−\(Ff~\(Xℓ\(m\)\)\+Fϕ~\(Xℓ\(m\)\)\)‖ℝdN2\)1/2,\\displaystyle E\_\{\\mathrm\{res\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\):=\\left\(\\frac\{1\}\{ML\}\\sum\_\{m,\\ell=1\}^\{M,L\}\\left\\lVert V^\{\(m\)\}\_\{\\ell\}\-\\left\(F\_\{\\tilde\{f\}\}\(X^\{\(m\)\}\_\{\\ell\}\)\+F\_\{\\tilde\{\\phi\}\}\(X^\{\(m\)\}\_\{\\ell\}\)\\right\)\\right\\rVert\_\{\\mathbb\{R\}^\{dN\}\}^\{2\}\\right\)^\{1/2\},\(53\)where‖⋅‖ℝdN\\left\\lVert\\cdot\\right\\rVert\_\{\\mathbb\{R\}^\{dN\}\}denotes the induced Euclidean norm onℝdN\\mathbb\{R\}^\{dN\}as defined by \([16](https://arxiv.org/html/2608.25181#S2.E16)\)\. Note the close correspondence between the residual error \([53](https://arxiv.org/html/2608.25181#S3.E53)\) and the error functionalℰℋf,ℋϕ\\mathcal\{E\}\_\{\\mathcal\{H\}\_\{f\},\\mathcal\{H\}\_\{\\phi\}\}defined in \([17](https://arxiv.org/html/2608.25181#S2.E17)\); indeed, the goal of the variational approach is precisely to minimize the residual error over the hypothesis spacesℋf\\mathcal\{H\}\_\{f\}andℋϕ\\mathcal\{H\}\_\{\\phi\}\. For second\-order systems \([85](https://arxiv.org/html/2608.25181#A1.E85)\), the observed trajectory data takes the form of position, velocity, and acceleration\(Xℓ\(m\),Vℓ\(m\),Aℓ\(m\)\)m,ℓ=1M,L\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\},A^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}as discussed in Section[A\.1](https://arxiv.org/html/2608.25181#A1.SS1), and the goal is to minimize the error functional \([96](https://arxiv.org/html/2608.25181#A1.E96)\) over hypothesis spacesℋf\\mathcal\{H\}\_\{f\}andℋϕ\\mathcal\{H\}\_\{\\phi\}\. Thus, we define the corresponding residual error \(and relative residual error\) analogously, where \(for example\) we replace velocitiesVℓ\(m\)V^\{\(m\)\}\_\{\\ell\}with accelerationsAℓ\(m\)A^\{\(m\)\}\_\{\\ell\}andFf~\(Xℓ\(m\)\)F\_\{\\tilde\{f\}\}\(X^\{\(m\)\}\_\{\\ell\}\)withFf~\(Xℓ\(m\),Vℓ\(m\)\)F\_\{\\tilde\{f\}\}\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)in \([53](https://arxiv.org/html/2608.25181#S3.E53)\)\. As with the mechanistic errors of Section[3\.2](https://arxiv.org/html/2608.25181#S3.SS2), we also define a relative residual error measuring how well the observed trajectory data are approximated by the candidate functionsϕ~\\tilde\{\\phi\}andf~\\tilde\{f\}\. Since the magnitude of the residual depends on the natural scale of the observed velocities, we normalize by a characteristic velocity scale\. Specifically, we define the residual scaleSresS\_\{\\mathrm\{res\}\}as the root mean square \(RMS\) of the observed velocity data\(Vℓ\(m\)\)m,ℓ=1M,L\(V^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}: Sres\\displaystyle S\_\{\\mathrm\{res\}\}=\(1ML∑m=1M∑ℓ=1L‖Vℓ\(m\)‖ℝdN2\)1/2\.\\displaystyle=\\left\(\\frac\{1\}\{ML\}\\sum\_\{m=1\}^\{M\}\\sum\_\{\\ell=1\}^\{L\}\\left\\lVert V^\{\(m\)\}\_\{\\ell\}\\right\\rVert\_\{\\mathbb\{R\}^\{dN\}\}^\{2\}\\right\)^\{1/2\}\.\(54\)This normalization makes the residual error dimensionless and allows errors to be compared across data sets, parameter regimes, and candidate models with different velocity scales\. In particular, the normalized residual measures the typical discrepancy between the observed and predicted velocities relative to the typical magnitude of the observed velocities\. We therefore define Eresrel\(f~,ϕ~\)\\displaystyle E^\{\\mathrm\{rel\}\}\_\{\\mathrm\{res\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\):=Eres\(f~,ϕ~\)Sres\.\\displaystyle:=\\frac\{E\_\{\\mathrm\{res\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\)\}\{S\_\{\\mathrm\{res\}\}\}\.\(55\)As discussed previously, for second\-order systemsSresS\_\{\\mathrm\{res\}\}is defined with respect to the observed acceleration data \(OPENAℓ\(m\)\)m,ℓ=1M,LA^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\} In Section[4](https://arxiv.org/html/2608.25181#S4), we utilize a regularized form of the relative residual error in \([55](https://arxiv.org/html/2608.25181#S3.E55)\), where a small constantϵ=10−12\\epsilon=10^\{\-12\}is included in the denominator of \([55](https://arxiv.org/html/2608.25181#S3.E55)\) to improve numerical stability when scales of the observed velocities \(accelerations for second\-order systems\) are small: Eres,ϵrel\(f~,ϕ~\)=Eres\(f~,ϕ~\)Sres\+ϵ\.\\displaystyle E\_\{\\mathrm\{res\},\\epsilon\}^\{\\mathrm\{rel\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\)=\\frac\{E\_\{\\mathrm\{res\}\}\(\\widetilde\{f\},\\widetilde\{\\phi\}\)\}\{S\_\{\\mathrm\{res\}\}\+\\epsilon\}\.\(56\) ### 3\.4Trajectory metrics A small residual error implies that estimators predict the observed velocity \(first\-order\) or acceleration \(second\-order\) well\. However, small residual errors do not necessarily guarantee accurate trajectory estimation\. In this section we introduce trajectory\-based metrics to assess accuracy of trajectory data utilized in both training and predictions\. Specifically, if we have observation data on an interval\[0,Tf\]\[0,T\_\{f\}\], we obtain estimators as described in Sections[2](https://arxiv.org/html/2608.25181#S2)and Appendix[A](https://arxiv.org/html/2608.25181#A1)utilizing data restricted to\[0,T\]⊆\[0,Tf\]\[0,T\]\\subseteq\[0,T\_\{f\}\]\(training data\), while the remaining observations in\(T,Tf\]\(T,T\_\{f\}\]are reserved for trajectory extrapolation \(testing data\) via predictions from the obtained estimators\. As in Section[2\.1\.1](https://arxiv.org/html/2608.25181#S2.SS1.SSS1), we define the observed trajectory data from replicatemmat timetℓt\_\{\\ell\}asXℓ\(m\)∈ℝdNX^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{dN\}\. For estimatorsf~\\tilde\{f\}andϕ~\\tilde\{\\phi\}, definex~=\(x~1,x~2,…,x~N\)∈ℝdN\\tilde\{x\}=\(\\tilde\{x\}\_\{1\},\\tilde\{x\}\_\{2\},\\ldots,\\tilde\{x\}\_\{N\}\)\\in\\mathbb\{R\}^\{dN\}as the corresponding solution of the IVP \{x˙=Ff~\(x\)\+Fϕ~\(x\)x\(0\)=X0\(m\)\.\\displaystyle\\begin\{cases\}\\dot\{x\}&=F\_\{\\tilde\{f\}\}\(x\)\+F\_\{\\tilde\{\\phi\}\}\(x\)\\\\ x\(0\)&=X^\{\(m\)\}\_\{0\}\.\\end\{cases\}\(57\)DefineX~ℓ\(m\)∈ℝdN\\tilde\{X\}^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{dN\}as the solutionx~\\tilde\{x\}evaluated at timetℓt\_\{\\ell\}of each replicatemm: \(X~ℓ\(m\)\)i\\displaystyle\(\\tilde\{X\}^\{\(m\)\}\_\{\\ell\}\)\_\{i\}:=x~i\(m\)\(tℓ\)\.\\displaystyle:=\\tilde\{x\}\_\{i\}^\{\(m\)\}\(t\_\{\\ell\}\)\.\(58\)We then define the pointwise agent\-averaged trajectory error of replicate \(initial condition\)m=1,2,…,Mm=1,2,\\ldots,Mat timetℓt\_\{\\ell\}as eℓ\(m\)\(f~,ϕ~\)\\displaystyle e^\{\(m\)\}\_\{\\ell\}\(\\tilde\{f\},\\tilde\{\\phi\}\):=‖Xℓ\(m\)−X~ℓ\(m\)‖ℝdN\.\\displaystyle:=\\left\\lVert X^\{\(m\)\}\_\{\\ell\}\-\\tilde\{X\}^\{\(m\)\}\_\{\\ell\}\\right\\rVert\_\{\\mathbb\{R\}^\{dN\}\}\.\(59\)Note that for each samplemmwe obtain an averaged \(over agents\) time series of trajectory errors\. To quantify goodness\-of\-fit for trajectory data over an entire time interval, we define the following trajectory errors: Erecon\(m\)\(f~,ϕ~\)\\displaystyle E\_\{\\mathrm\{recon\}\}^\{\(m\)\}\(\\tilde\{f\},\\tilde\{\\phi\}\):=maxtℓ∈\[0,T\]eℓ\(m\)\(f~,ϕ~\)\\displaystyle:=\\max\_\{t\_\{\\ell\}\\in\[0,T\]\}e^\{\(m\)\}\_\{\\ell\}\(\\tilde\{f\},\\tilde\{\\phi\}\)\(60\)Epred\(m\)\(f~,ϕ~\)\\displaystyle E\_\{\\mathrm\{pred\}\}^\{\(m\)\}\(\\tilde\{f\},\\tilde\{\\phi\}\):=maxtℓ∈\(T,Tf\]eℓ\(m\)\(f~,ϕ~\)\\displaystyle:=\\max\_\{t\_\{\\ell\}\\in\(T,T\_\{f\}\]\}e^\{\(m\)\}\_\{\\ell\}\(\\tilde\{f\},\\tilde\{\\phi\}\)\(61\)Etraj\(m\)\(f~,ϕ~\)\\displaystyle E\_\{\\mathrm\{traj\}\}^\{\(m\)\}\(\\tilde\{f\},\\tilde\{\\phi\}\):=maxtℓ∈\[0,Tf\]eℓ\(m\)\(f~,ϕ~\)\\displaystyle:=\\max\_\{t\_\{\\ell\}\\in\[0,T\_\{f\}\]\}e^\{\(m\)\}\_\{\\ell\}\(\\tilde\{f\},\\tilde\{\\phi\}\)\(62\)That is,Erecon\(m\),Epred\(m\)E\_\{\\mathrm\{recon\}\}^\{\(m\)\},E\_\{\\mathrm\{pred\}\}^\{\(m\)\}, andEtraj\(m\)E\_\{\\mathrm\{traj\}\}^\{\(m\)\}measure trajectory accuracy with respect to the training data\[0,T\]\[0,T\], the testing \(prediction\) data\(T,Tf\]\(T,T\_\{f\}\], and the entire trajectory\[0,Tf\]\[0,T\_\{f\}\], respectively\. AsErecon\(m\)E\_\{\\mathrm\{recon\}\}^\{\(m\)\}measures the ability of the estimators to reconstruct the training data, we refer to this as the reconstruction error of samplemm\. The trajectory errorEtraj\(m\)E\_\{\\mathrm\{traj\}\}^\{\(m\)\}combines reconstruction and prediction performance into a single quantity, and will serve as the primary trajectory\-based metric used in the model selection procedure discussed in Section[4](https://arxiv.org/html/2608.25181#S4)\. We have defined trajectory\-based metrics for each sampled initial conditionmm\. To assess performance over an entire observed trajectory data set, we aggregate these quantities across allMMrealizations\. Specifically, we define the sample means E¯recon\(f~,ϕ~\):=1M∑m=1MErecon\(m\)\(f~,ϕ~\)E¯pred\(f~,ϕ~\):=1M∑m=1MEpred\(m\)\(f~,ϕ~\)E¯traj\(f~,ϕ~\):=1M∑m=1MEtraj\(m\)\(f~,ϕ~\),\\displaystyle\\begin\{split\}\\bar\{E\}\_\{\\mathrm\{recon\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\)&:=\\frac\{1\}\{M\}\\sum\_\{m=1\}^\{M\}E\_\{\\mathrm\{recon\}\}^\{\(m\)\}\(\\tilde\{f\},\\tilde\{\\phi\}\)\\\\ \\bar\{E\}\_\{\\mathrm\{pred\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\)&:=\\frac\{1\}\{M\}\\sum\_\{m=1\}^\{M\}E\_\{\\mathrm\{pred\}\}^\{\(m\)\}\(\\tilde\{f\},\\tilde\{\\phi\}\)\\\\ \\bar\{E\}\_\{\\mathrm\{traj\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\)&:=\\frac\{1\}\{M\}\\sum\_\{m=1\}^\{M\}E\_\{\\mathrm\{traj\}\}^\{\(m\)\}\(\\tilde\{f\},\\tilde\{\\phi\}\),\\end\{split\}\(63\)and the corresponding sample standard deviations σErecon\(f~,ϕ~\):=\(1M−1∑m=1M\(Erecon\(m\)\(f~,ϕ~\)−E¯recon\(f~,ϕ~\)\)2\)1/2σEpred\(f~,ϕ~\):=\(1M−1∑m=1M\(Epred\(m\)\(f~,ϕ~\)−E¯pred\(f~,ϕ~\)\)2\)1/2σEtraj\(f~,ϕ~\):=\(1M−1∑m=1M\(Etraj\(m\)\(f~,ϕ~\)−E¯traj\(f~,ϕ~\)\)2\)1/2\.\\displaystyle\\begin\{split\}\\sigma\_\{E\_\{\\mathrm\{recon\}\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\)&:=\\left\(\\frac\{1\}\{M\-1\}\\sum\_\{m=1\}^\{M\}\\left\(E\_\{\\mathrm\{recon\}\}^\{\(m\)\}\(\\tilde\{f\},\\tilde\{\\phi\}\)\-\\bar\{E\}\_\{\\mathrm\{recon\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\)\\right\)^\{2\}\\right\)^\{1/2\}\\\\ \\sigma\_\{E\_\{\\mathrm\{pred\}\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\)&:=\\left\(\\frac\{1\}\{M\-1\}\\sum\_\{m=1\}^\{M\}\\left\(E\_\{\\mathrm\{pred\}\}^\{\(m\)\}\(\\tilde\{f\},\\tilde\{\\phi\}\)\-\\bar\{E\}\_\{\\mathrm\{pred\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\)\\right\)^\{2\}\\right\)^\{1/2\}\\\\ \\sigma\_\{E\_\{\\mathrm\{traj\}\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\)&:=\\left\(\\frac\{1\}\{M\-1\}\\sum\_\{m=1\}^\{M\}\\left\(E\_\{\\mathrm\{traj\}\}^\{\(m\)\}\(\\tilde\{f\},\\tilde\{\\phi\}\)\-\\bar\{E\}\_\{\\mathrm\{traj\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\)\\right\)^\{2\}\\right\)^\{1/2\}\.\\end\{split\}\(64\)The averaged errors summarize the sampled\-averaged accuracy of the estimated mechanisms, while the corresponding standard deviations quantify variability across trajectory realizations\. When evaluated on prediction data \(e\.g\. on observation times in\(T,Tf\]\(T,T\_\{f\}\]\), these quantities additionally provide an assessment of predictive performance\. Additionally, as opposed to the maximum pointwise trajectory error, we also compute the RMS error, which characterizes an average measure of fidelity of trajectory data\. Compared with the maximum errors \([63](https://arxiv.org/html/2608.25181#S3.E63)\), this metric is less dominated by outliers, and is utilized primarily in Section[4](https://arxiv.org/html/2608.25181#S4), where the corresponding relative RMS errors \(defined in \([67](https://arxiv.org/html/2608.25181#S3.E67)\) below\) is utilized to determine model frameworks which are incompatible with the observed trajectory data\. We define trajectory RMS error as E∗RMS\(f~,ϕ~\):=\(1ML∑m,ℓ=1M,Leℓ\(m\)\(f~,ϕ~\)\)1/2,\\displaystyle E\_\{\*\}^\{\\mathrm\{RMS\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\):=\\left\(\\frac\{1\}\{ML\}\\sum\_\{m,\\ell=1\}^\{M,L\}e\_\{\\ell\}^\{\(m\)\}\(\\tilde\{f\},\\tilde\{\\phi\}\)\\right\)^\{1/2\},\(65\)where the subscript “∗\{\*\}” indicates the time interval that the error is evaluated, i\.e\.∗∈\{recon,pred,traj\}\{\*\}\\in\\\{\\mathrm\{recon\},\\mathrm\{pred\},\\mathrm\{traj\}\\\}as discussed in \([60](https://arxiv.org/html/2608.25181#S3.E60)\) \- \([62](https://arxiv.org/html/2608.25181#S3.E62)\)\. As in \([54](https://arxiv.org/html/2608.25181#S3.E54)\), the trajectory scaleSXS\_\{X\}is defined as the RMS of the trajectory data, SX:=\(1ML∑m,ℓ=1M,L‖Xℓ\(m\)‖ℝdN2\)1/2,\\displaystyle S\_\{X\}:=\\left\(\\frac\{1\}\{ML\}\\sum\_\{m,\\ell=1\}^\{M,L\}\\left\\lVert X^\{\(m\)\}\_\{\\ell\}\\right\\rVert\_\{\\mathbb\{R\}^\{dN\}\}^\{2\}\\right\)^\{1/2\},\(66\)and the corresponding relative reconstruction error is defined as E∗rel:=E∗RMS\(f~,ϕ~\)SX\+ϵ,\\displaystyle E\_\{\*\}^\{\\mathrm\{rel\}\}:=\\frac\{E\_\{\*\}^\{\\mathrm\{RMS\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\)\}\{S\_\{X\}\+\\epsilon\},\(67\)where again∗∈\{recon,pred,traj\}\{\*\}\\in\\\{\\mathrm\{recon\},\\mathrm\{pred\},\\mathrm\{traj\}\\\}\. Note that as opposed to the case for residual errors, the definitions for trajectory errors for first\- and second\-order systems are identical\. ### 3\.5Metrics for consistency Beyond accuracy, we quantify the consistency of the proposed variational approach with respect to the trajectory data utilized to estimate environmental and interaction forces in models of collective dynamics\. We assess consistency via uncertainty quantification on observation \(training\) trajectory data, scalability with respect to the number of agents \(NN\), accuracy with respect to the number of sample \(MM\) and time \(LL\) measurements, and robustness to observational noise\. #### 3\.5\.1Uncertainty quantification To evaluate variation with respect to sample data\(Xℓ\(m\),Vℓ\(m\)\)m,ℓ=1M,L\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}, the estimation algorithm is repeated independently a total ofTrT\_\{r\}times usingMMsamples generated from the same underlying distributionμ0\\mu\_\{0\}\. That is, we considerTrT\_\{r\}replicates ofMMsamples, and for each trial inTrT\_\{r\}, all evaluation metrics introduced in Sections[3\.2](https://arxiv.org/html/2608.25181#S3.SS2)\-[3\.4](https://arxiv.org/html/2608.25181#S3.SS4)are computed, and trial means and trial standard deviations of the error metrics are subsequently reported\. Small trial\-to\-trial variability indicates that the learning procedure is relatively insensitive to sampling fluctuations in observed trajectory data, and therefore produces reliable estimates of the underlying environmental and interaction mechanisms\. #### 3\.5\.2Scalability We evaluate the estimators ability to scale from a small to large number of agents by obtaining estimators for systems containingNNagents, and subsequently evaluating the performance of those estimators on the same system with4N4Nagents\. The objective of this experiment is to determine whether the recovered interaction mechanisms generalize across system sizes and continue to reproduce the collective dynamics of larger populations without the need to re\-estimate\. Scientifically this is also useful, as it is often more experimentally tractable to measure trajectories on small samples\. Note that in this numerical experiment, we continue to report reconstruction error metrics for trajectories \(e\.gE¯recon\\bar\{E\}\_\{\\mathrm\{recon\}\}\), but emphasize that we do not train on the4N4N\-agent system on any time interval; this is a purely predictive experiment, and in this case such errors simply measure the predictive trajectory error on the early time window\[0,T\]\[0,T\]\. #### 3\.5\.3Dependence on training data We investigate how estimation accuracy improves for the mechanismsffandϕ\\phias a function of the amount of training data\. The following two types of experiments are considered: 1. 1\.Increase the number of samplesMM\. 2. 2\.Increase the temporal resolutionLLof the observations\. For each experiment, the training data is expanded while all other experimental settings remain fixed, and we compute how the relative feature recovery metrics \([52](https://arxiv.org/html/2608.25181#S3.E52)\) vary withMMandLL\. This resulting convergence analysis provide empirical evidence regarding the amount of data required to accurately recover the underlying interaction mechanisms\. #### 3\.5\.4Robustness to noise in observation data Robustness refers to the ability of the learned model to maintain accuracy in the presence of noisy observations\. To assess robustness, controlled perturbations are introduced into the observed trajectories\(Xℓ\(m\),Vℓ\(m\)\)m,ℓ=1M,L\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}for first\-order systems \([4](https://arxiv.org/html/2608.25181#S2.E4)\); for second\-order systems \([85](https://arxiv.org/html/2608.25181#A1.E85)\), perturbations are also added to the acceleration data\(Aℓ\(m\)\)m,ℓ=1M,L\(A^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}\. In this manuscript, we consider multiplicative noise\. In the simplest case, for data utilized in the estimation algorithm \(i\.e\. fortℓ∈\[0,T\]t\_\{\\ell\}\\in\[0,T\]as described in Section[3\.4](https://arxiv.org/html/2608.25181#S3.SS4)\), the observation data takes the perturbed form xi\(m\),noisy:=xi\(m\)\(tℓ\)\(1\+ηi,ℓ\(m\)\)andvi\(m\),noisy:=vi\(m\)\(tℓ\)\(1\+η¯i,ℓ\(m\)\)\\displaystyle\\begin\{split\}x\_\{i\}^\{\(m\),\\text\{noisy\}\}&:=x\_\{i\}^\{\(m\)\}\(t\_\{\\ell\}\)\\left\(1\+\\eta\_\{i,\\ell\}^\{\(m\)\}\\right\)\\quad\\text\{and\}\\quad v\_\{i\}^\{\(m\),\\text\{noisy\}\}:=v\_\{i\}^\{\(m\)\}\(t\_\{\\ell\}\)\\left\(1\+\\bar\{\\eta\}\_\{i,\\ell\}^\{\(m\)\}\\right\)\\end\{split\}\(68\)where ηi,ℓ\(m\),η¯i,ℓ\(m\)\\displaystyle\\eta\_\{i,\\ell\}^\{\(m\)\},\\bar\{\\eta\}\_\{i,\\ell\}^\{\(m\)\}∼i\.i\.d\.Unif\[−ζ,ζ\]\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{i\.i\.d\.\}\}\}\{\{\\sim\}\}\\text\{Unif\}\[\-\\zeta,\\zeta\]\(69\)and\(\(Xℓ\(m\)\)i,\(Vℓ\(m\)\)i\)=\(xi\(m\)\(tℓ\),vi\(m\)\(tℓ\)\)\(\(X^\{\(m\)\}\_\{\\ell\}\)\_\{i\},\(V^\{\(m\)\}\_\{\\ell\}\)\_\{i\}\)=\(x\_\{i\}^\{\(m\)\}\(t\_\{\\ell\}\),v\_\{i\}^\{\(m\)\}\(t\_\{\\ell\}\)\)\. That is, noise is applied independently from a uniform distribution to all observed state variables; thus we assume in \([68](https://arxiv.org/html/2608.25181#S3.E68)\) that noise affects position and velocity independently\. Numerical experiments are performed over the range ζ∈\{0,0\.01,0\.03,0\.05,0\.10\},\\displaystyle\\zeta\\in\\\{0,\\;0\.01,\\;0\.03,\\;0\.05,\\;0\.10\\\},\(70\)and for each noise level, we evaluateEϕE\_\{\\phi\},EfE\_\{f\}, and also the trajectory\-based metrics of performance\. The above assumption of independent noise for position, velocity, and acceleration data is generally unrealistic, as experimentally only position measurements are typically available, from which velocity and acceleration must be constructed\. Since the proposed learning framework requires velocity and, for second\-order systems, acceleration observations, these quantities must be estimated from measured positions\. A straightforward estimation approach is to apply finite\-difference schemes; however, numerical differentiation amplifies measurement noise and often produces poor derivative estimates\[[72](https://arxiv.org/html/2608.25181#bib.bib96)\]\. Several smoothing and differentiation techniques have been proposed for trajectory pre\-processing, including moving\-average windows\[[73](https://arxiv.org/html/2608.25181#bib.bib66)\], Butterworth filters\[[74](https://arxiv.org/html/2608.25181#bib.bib67)\], Savitzky–Golay filters\[[75](https://arxiv.org/html/2608.25181#bib.bib68)\], Whittaker smoothers\[[76](https://arxiv.org/html/2608.25181#bib.bib69)\], and Kalman filtering with Rauch\-Tung\-Striebel \(RTS\) smoothing\[[77](https://arxiv.org/html/2608.25181#bib.bib70),[78](https://arxiv.org/html/2608.25181#bib.bib71)\]\. We note that these techniques have been applied to trajectory data from animal movement, unmanned aerial vehicle \(UAV\) tracking, biomechanics, ecological monitoring, and other systems which require analysis of motion\. Each method has its own advantages and limitations depending on the characteristics of the data\. However, since noise filtering is not the primary focus of this work, we do not attempt a comprehensive comparison of these approaches\. Instead, we evaluate two practical approaches for pre\-processing noisy trajectory data: a moving\-average smoother and a Kalman filter with an RTS smoother\. Both methods reduce the effects of measurement noise before applying the learning algorithm, but they accomplish this in different ways\. The moving\-average filter smooths the position measurements over a local time window, after which velocities and accelerations are obtained through numerical differentiation\. In contrast, the Kalman smoother estimates positions, velocities, and, when appropriate, accelerations simultaneously through a state\-space model, thereby eliminating the need for numerical differentiation\. Both methods will be applied to estimate velocity data for noisy trajectory position data for a first\-order system in Section[5\.5](https://arxiv.org/html/2608.25181#S5.SS5)\. ## 4Model selection The framework introduced in Section[2\.1](https://arxiv.org/html/2608.25181#S2.SS1)and Appendix[A](https://arxiv.org/html/2608.25181#A1)assumes that the precise form of the dynamical system is known a priori, e\.g\. that there exists both an interaction kernelϕ\\phiand an environmental forceff\. In practice however, the exact form of the governing dynamics is often unknown\. For example, a system may involve interaction forces, alignment effects, environmental influences, and/or combinations of these mechanisms\. Therefore, it is natural to consider a more general class of models of collective dynamics and determine, directly from the trajectory data, precisely which mechanistic forces are influencing the dynamics\. Ideally, this selection procedure should be data\-driven: rather than manually choosing the model structure in advance, the algorithm should identify a parsimonious candidate model whose inferred forces accurately reproduce the observed dynamics\. It is the goal of this section to outline such a selection procedure\. For first\-order systems, we consider the generalized framework x˙i\\displaystyle\\dot\{x\}\_\{i\}=f\(xi\)\+1N∑j=1j≠iNϕE\(\|xj−xi\|\)\(xj−xi\),i=1,…,N\\displaystyle=f\(x\_\{i\}\)\+\\frac\{1\}\{N\}\\sum\_\{\\begin\{subarray\}\{c\}j=1\\\\ j\\neq i\\end\{subarray\}\}^\{N\}\\phi^\{E\}\(\|x\_\{j\}\-x\_\{i\}\|\)\(x\_\{j\}\-x\_\{i\}\),\\quad i=1,\\dots,N\(71\)where the environmental forceffand interaction kernelϕE\\phi^\{E\}may be identically zero\. Here we label the interaction kernelϕ=ϕE\\phi=\\phi^\{E\}to be consistent with \([72](https://arxiv.org/html/2608.25181#S4.E72)\)\. Analogously, for second\-order systems, we consider the generalized framework \(fori=1,…,Ni=1,\\dots,N\) \{x˙i=viv˙i=f\(xi,vi\)\+1N∑j=1j≠iNϕE\(\|xj−xi\|\)\(xj−xi\)\+1N∑j=1j≠iNϕA\(\|xj−xi\|\)\(vj−vi\),\\begin\{cases\}\\dot\{x\}\_\{i\}&=v\_\{i\}\\\\ \\dot\{v\}\_\{i\}&=f\(x\_\{i\},v\_\{i\}\)\+\\frac\{1\}\{N\}\\sum\_\{\\begin\{subarray\}\{c\}j=1\\\\ j\\neq i\\end\{subarray\}\}^\{N\}\\phi^\{E\}\(\|x\_\{j\}\-x\_\{i\}\|\)\(x\_\{j\}\-x\_\{i\}\)\+\\frac\{1\}\{N\}\\sum\_\{\\begin\{subarray\}\{c\}j=1\\\\ j\\neq i\\end\{subarray\}\}^\{N\}\\phi^\{A\}\(\|x\_\{j\}\-x\_\{i\}\|\)\(v\_\{j\}\-v\_\{i\}\),\\end\{cases\}\(72\)where the environmental forceff, the energy\-based interaction kernelϕE\\phi^\{E\}, and the alignment interaction kernelϕA\\phi^\{A\}may be identically zero\. The previously introduced variational learning procedure naturally extends to these generalized formulations\. Under the assumption that the data is generated from dynamics corresponding to either the full generalized model or to one of its sub\-models \(e\.g\. ifϕA≡0\\phi^\{A\}\\equiv 0in \([72](https://arxiv.org/html/2608.25181#S4.E72)\)\), the learned coefficients of the estimators with respect to fixed hypothesis spaces provide information about the likelihood of specific interaction mechanisms from the observed data\. Relatedly, we are also interested in determining whether the data supports a first\-order or second\-order model description\. In this section, we discuss a systematic framework for determining which mechanisms are likely present in data generated from an inter\-agent particle system\. We further note that although similar methodologies can be applied to extended systems of collective dynamics \(or more generally to systems exhibiting a high degree of symmetry\), we restrict our attention to the above frameworks throughout this manuscript\. We begin by enumerating a collection of candidate frameworks, referred to as*candidates*and denoted by the labelcc, each corresponding to a specific combination of both order and mechanisms\. As described below, each candidate model is utilized to construct estimators, from which each are evaluated and ranked according to predefined selection criteria\. The framework with the strongest support from the data \(i\.e\. highest ranked\) is selected as the most likely framework which describes the available data\. We assign to each candidate a complexity rank based on its dynamical order and the number of active features, as summarized in Table[1](https://arxiv.org/html/2608.25181#S4.T1); generally first\-order models are less complex when compared to second\-order models, as are models with fewer terms comprising their vector fields\. Thus, complexity is based on the notion of parsimony, i\.e\. our goal is to select the simplest model that is able to describe the data\. For every case, the corresponding model yields estimators as described in Sections[2\.1](https://arxiv.org/html/2608.25181#S2.SS1)\(first\-order\) and Appendix[A](https://arxiv.org/html/2608.25181#A1)\(second\-order\), and its performance on the available trajectory data is evaluated utilizing the metrics described in Section[3\.4](https://arxiv.org/html/2608.25181#S3.SS4)\. Candidate \(OPENc\)c\)Candidate ModelComplexityF1F\_\{1\}x˙=FϕE\\dot\{x\}=F\_\{\\phi^\{E\}\}1F2F\_\{2\}x˙=Ff\+FϕE\\dot\{x\}=F\_\{f\}\+F\_\{\\phi^\{E\}\}2S1S\_\{1\}x¨=FϕE\\ddot\{x\}=F\_\{\\phi^\{E\}\}3S2S\_\{2\}x¨=FϕA\\ddot\{x\}=F\_\{\\phi^\{A\}\}3S3S\_\{3\}x¨=Ff\+FϕE\\ddot\{x\}=F\_\{f\}\+F\_\{\\phi^\{E\}\}4S4S\_\{4\}x¨=Ff\+FϕA\\ddot\{x\}=F\_\{f\}\+F\_\{\\phi^\{A\}\}4S5S\_\{5\}x¨=FϕE\+FϕA\\ddot\{x\}=F\_\{\\phi^\{E\}\}\+F\_\{\\phi^\{A\}\}5S6S\_\{6\}x¨=Ff\+FϕE\+FϕA\\ddot\{x\}=F\_\{f\}\+F\_\{\\phi^\{E\}\}\+F\_\{\\phi^\{A\}\}6Table 1:Hierarchy of candidate models used for model selection\. In the Candidate column,FFdenote a first\-order framework, whileSSdenote the second\-order framework\. The notation utilized for vector fields inℝdN\\mathbb\{R\}^\{dN\}is as described in Sections[2\.1\.1](https://arxiv.org/html/2608.25181#S2.SS1.SSS1)and Appendix[A\.1](https://arxiv.org/html/2608.25181#A1.SS1)\(e\.g\. equation[7](https://arxiv.org/html/2608.25181#S2.E7)\)\. We also suppressx˙i=vi\\dot\{x\}\_\{i\}=v\_\{i\}for second\-order systems for notational simplicity\.Before selecting the most plausible framework, each candidate model must first demonstrate that it is capable of explaining the observed data\. As is standard in machine\-learning contexts, we partition the available observation trajectory data into training, validation, and testing sets\. That is, ifMMsamples are available, we partition them into three subsets, where M\\displaystyle M=Mtrain\+Mval\+Mtest,\\displaystyle=M\_\{\\mathrm\{train\}\}\+M\_\{\\mathrm\{val\}\}\+M\_\{\\mathrm\{test\}\},\(73\)each of whose respective purposes are summarized in Table[2](https://arxiv.org/html/2608.25181#S4.T2)\. Candidate models are fit to the trajectory data using the training subset and subsequently compared via the validation subset\. Once the candidate that best represents the observed dynamics has been selected, the test set is used only for its final evaluation by comparing the predicted trajectories with previously unseen test trajectories; thus the test subset does not participate in either model fitting or model selection and is not utilized in the model selection algorithm\. Testing data is however utilized when we fit a given model to data, as in Section[5\.2](https://arxiv.org/html/2608.25181#S5.SS2)\(see for example Figure[2](https://arxiv.org/html/2608.25181#S5.F2)\)\. In general, the language for partitioning data is as follows: training and testing are utilized when discussing apre\-determinedmodel framework, while training and validating apply for model selection\. Candidate frameworks that cannot adequately reproduce the observed dynamics are removed prior to final model selection\. In particular, poor candidates often exhibit large residual errors, inaccurate reconstruction of the training trajectories, or even fail to generate numerically integrable trajectories\. Motivated by these observations, we introduce a framework compatibility gate that filters out inadmissible candidates prior to ranking\. Specifically, on theMtrainM\_\{\\mathrm\{train\}\}training replicates, define a trajectory reconstruction toleranceτrecon\\tau\_\{\\mathrm\{recon\}\}and a residual toleranceτres\\tau\_\{\\mathrm\{res\}\}\. A candidate frameworkccis then consideredadmissibleif both of the following conditions are satisfied: Ereconrel\(c\)<τreconandEres,ϵrel\(c\)<τres\.\\displaystyle\\begin\{split\}E\_\{\\mathrm\{recon\}\}^\{\\mathrm\{rel\}\}\(c\)&<\\tau\_\{\\mathrm\{recon\}\}\\quad\\text\{and\}\\quad E\_\{\\mathrm\{res\},\\epsilon\}^\{\\mathrm\{rel\}\}\(c\)<\\tau\_\{\\mathrm\{res\}\}\.\\end\{split\}\(74\)whereEreconrelE\_\{\\mathrm\{recon\}\}^\{\\mathrm\{rel\}\}is defined in \([67](https://arxiv.org/html/2608.25181#S3.E67)\) \(∗=recon\*=\\mathrm\{recon\}\), andEres,ϵrelE\_\{\\mathrm\{res\},\\epsilon\}^\{\\mathrm\{rel\}\}is defined in \([56](https://arxiv.org/html/2608.25181#S3.E56)\)\. Intuitively, this requires that the candidate modelccmust approximate well both the training trajectory and velocity \(first\-order\) or acceleration \(second\-order\) data\. In numerical simulations, the above thresholds are fixed atτrecon=τres=0\.25\\tau\_\{\\mathrm\{recon\}\}=\\tau\_\{\\mathrm\{res\}\}=0\.25, and if at least one condition in \([74](https://arxiv.org/html/2608.25181#S4.E74)\) is violated by candidatecc, it is deemedinadmissibleand not considered further in the model selection procedure\. In addition, a candidate framework is rejected if its learned dynamics cannot be integrated successfully or if the condition number of the associated regression matrixAAin \([36](https://arxiv.org/html/2608.25181#S2.E36)\) exceeds a prescribed numerical stability threshold \(101210^\{12\}\)\. The corresponding evaluation metrics are reported asN/A\. Denote by𝒞\\mathcal\{C\}the set of admissible candidate frameworks: 𝒞\\displaystyle\\mathcal\{C\}:=\{c\|Ereconrel\(c\)<τrecon,Eres,ϵrel\(c\)<τres\}\\displaystyle:=\\\{c\\,\|\\,E\_\{\\mathrm\{recon\}\}^\{\\mathrm\{rel\}\}\(c\)<\\tau\_\{\\mathrm\{recon\}\},\\,E\_\{\\mathrm\{res\},\\epsilon\}^\{\\mathrm\{rel\}\}\(c\)<\\tau\_\{\\mathrm\{res\}\}\\\}\(75\)If𝒞=∅\\mathcal\{C\}=\\varnothing, then no first\- or second\-order models of the form \([71](https://arxiv.org/html/2608.25181#S4.E71)\) or \([72](https://arxiv.org/html/2608.25181#S4.E72)\) are selected to describe the trajectory data\. Data setFunctionTraining \(MtrainM\_\{\\mathrm\{train\}\}samples\)Obtain estimators for mechanistic forces for each modelValidation \(MvalM\_\{\\mathrm\{val\}\}samples\)Compare candidate models and tune hyperparameters to select candidate modelTesting \(MtestM\_\{\\mathrm\{test\}\}samples\)Measure final evaluation of models and report performanceTable 2:Description of training, validation, and testing data sets for model selection algorithm\. Note that the full set of initial condition replicatesMMis partitioned between the three classifications\.By construction, all candidates in𝒞\\mathcal\{C\}possess small trajectory and residual error with respect to the training replicates\. The primary metric utilized to further evaluate admissible candidates is validation on trajectory data\. That is, candidate frameworksc∈𝒞c\\in\\mathcal\{C\}are further selected if their mean trajectory error on the validation replicates is below a threshold\. More precisely, define Emin:=minc∈𝒞E¯traj\(c\)andτtraj:=max\{Emin\+εabs,Emin\(1\+εrel\)\}\\displaystyle\\begin\{split\}E\_\{\\min\}&:=\\min\_\{c\\in\\mathcal\{C\}\}\\bar\{E\}\_\{\\mathrm\{traj\}\}\(c\)\\quad\\text\{and\}\\quad\\tau\_\{\\mathrm\{traj\}\}:=\\max\\left\\\{E\_\{\\min\}\+\\varepsilon\_\{\\mathrm\{abs\}\},E\_\{\\min\}\(1\+\\varepsilon\_\{\\mathrm\{rel\}\}\)\\right\\\}\\end\{split\}\(76\)ErrorEminE\_\{\\min\}thus denotes the minimum average trajectory error \(defined by \([63](https://arxiv.org/html/2608.25181#S3.E63)\)\) over all admissible candidates on the validation replicates \(recall that there areMvalM\_\{\\mathrm\{val\}\}of them\), andτtraj\\tau\_\{\\mathrm\{traj\}\}measures both an absolute and a relative tolerance with respect toEminE\_\{\\min\}\. In numerical experiments, we fixεabs=10−3\\varepsilon\_\{\\mathrm\{abs\}\}=10^\{\-3\}andεrel=0\.05\\varepsilon\_\{\\mathrm\{rel\}\}=0\.05\. Denote that𝒞val\\mathcal\{C\}\_\{\\mathrm\{val\}\}the subset of admissible candidates that posses a mean trajectory error \(on validation replicates\) belowτtraj\\tau\_\{\\mathrm\{traj\}\}: 𝒞val\\displaystyle\\mathcal\{C\}\_\{\\mathrm\{val\}\}:=\{c∈𝒞\|E¯traj\(c\)≤τtraj\}\\displaystyle:=\\left\\\{c\\in\\mathcal\{C\}\\,\|\\,\\bar\{E\}\_\{\\mathrm\{traj\}\}\(c\)\\leq\\tau\_\{\\mathrm\{traj\}\}\\right\\\}\(77\) A model framework is selected from𝒞val\\mathcal\{C\}\_\{\\mathrm\{val\}\}as follows\. If\|𝒞val\|=1\|\\mathcal\{C\}\_\{\\mathrm\{val\}\}\|=1, then the uniquec∈𝒞valc\\in\\mathcal\{C\}\_\{\\mathrm\{val\}\}is selected as the most likely model framework to describe the trajectory data\. If\|𝒞val\|\>1\|\\mathcal\{C\}\_\{\\mathrm\{val\}\}\|\>1, i\.e\. multiple models exhibit statistically insignificant differences in mean validation trajectory error, then preference is given to the framework with lower complexity as defined in Table[1](https://arxiv.org/html/2608.25181#S4.T1), following the principle of parsimony\[[79](https://arxiv.org/html/2608.25181#bib.bib63)\]\. If multiple models in𝒞val\\mathcal\{C\}\_\{\\mathrm\{val\}\}exhibit the same complexity, the framework with the smallest residual errorEresE\_\{\\text\{res\}\}\([53](https://arxiv.org/html/2608.25181#S3.E53)\) on theMtrainM\_\{\\mathrm\{train\}\}replicates is selected\. For reference, Table[3](https://arxiv.org/html/2608.25181#S4.T3)summarizes the metrics and partitions of the trajectory data utilized throughout the selection procedure\. The complete algorithm is provided in Algorithm[2](https://arxiv.org/html/2608.25181#algorithm2)\. MetricDataTime horizonFunctionEreconrelE\_\{\\mathrm\{recon\}\}^\{\\mathrm\{rel\}\}MtrainM\_\{\\mathrm\{train\}\}\[0,T\]\[0,T\]Obtain admissible frameworks \(trajectory data\)Eres,ϵrelE\_\{\\mathrm\{res\},\\epsilon\}^\{\\mathrm\{rel\}\}MtrainM\_\{\\mathrm\{train\}\}\[0,T\]\[0,T\]Obtain admissible frameworks \(residual data\)E¯traj\\bar\{E\}\_\{\\mathrm\{traj\}\}MvalM\_\{\\mathrm\{val\}\}\[0,Tf\]\[0,T\_\{f\}\]Primary validation metric \(trajectories\)EresE\_\{\\mathrm\{res\}\}MtrainM\_\{\\mathrm\{train\}\}\[0,T\]\[0,T\]Residual consistencyTable 3:Summary of the evaluation metrics, the datasets on which they are computed, their associated time horizons, and their intended purposes\.Algorithm 2Algorithm for identifying model framework from trajectory dataInput:Training trajectories MtrainM\_\{\\mathrm\{train\}\}on \[0,T\]\[0,T\], validation trajectories MvalM\_\{\\mathrm\{val\}\}on \(T,Tf\]\(T,T\_\{f\}\], test trajectories MtestM\_\{\\mathrm\{test\}\}on \[0,Tf\]\[0,T\_\{f\}\] Output:Selected candidate model c∗c^\{\*\}and its learned interaction laws, or a declaration that the dataset is incompatible with the proposed framework Using MtrainM\_\{\\mathrm\{train\}\}on \[t0,T\]\[t\_\{0\},T\]: Construct pairwise quantities and estimate the observed supports Interaction distances \(Rℓ\(m\)\)m,ℓ=1M,L\(R^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\} Observed supports \[Rmin,Rmax\]\[R\_\{\\min\},R\_\{\\max\}\]and \[Zmin,Zmax\]\[Z\_\{\\min\},Z\_\{\\max\}\], where Z=XZ=Xfor first\-order models and Z=VZ=Vor X,VX,Vfor second\-order models Construct localized basis functions \(ψkE\)k=1nϕ\(\\psi\_\{k\}^\{E\}\)\_\{k=1\}^\{n\_\{\\phi\}\}, \(ψkA\)k=1nϕ\(\\psi\_\{k\}^\{A\}\)\_\{k=1\}^\{n\_\{\\phi\}\}, and \(hk\)k=1nf\(h\_\{k\}\)\_\{k=1\}^\{n\_\{f\}\} foreach*candidatec∈𝒞c\\in\\mathcal\{C\}*do Assemble the least\-squares system Acθ^c=b→c\.A\_\{c\}\\hat\{\\theta\}\_\{c\}=\\vec\{b\}\_\{c\}\.Solve for θ^c\\hat\{\\theta\}\_\{c\} Recover the learned interaction laws: ϕ^E,ϕ^A,f^\\hat\{\\phi\}^\{E\},\\hat\{\\phi\}^\{A\},\\hat\{f\}associated with cc Compute EresE\_\{\\mathrm\{res\}\}\([53](https://arxiv.org/html/2608.25181#S3.E53)\) and E¯recon\\bar\{E\}\_\{\\mathrm\{recon\}\}using MtrainM\_\{\\mathrm\{train\}\}\([63](https://arxiv.org/html/2608.25181#S3.E63)\) Compute E¯traj\\bar\{E\}\_\{\\mathrm\{traj\}\}using MvalM\_\{\\mathrm\{val\}\}\([63](https://arxiv.org/html/2608.25181#S3.E63)\) Assign the candidate status according to compatibility and numerical checks end foreach Reject inadmissible candidates using Eres,ϵrelE\_\{\\mathrm\{res\},\\epsilon\}^\{\\mathrm\{rel\}\}\([56](https://arxiv.org/html/2608.25181#S3.E56)\) and EreconrelE\_\{\\mathrm\{recon\}\}^\{\\mathrm\{rel\}\}\([67](https://arxiv.org/html/2608.25181#S3.E67)\) and inability to integrate numerically if*at least one admissible candidate remains*then Identify EminE\_\{\\min\}via \([76](https://arxiv.org/html/2608.25181#S4.E76)\) Form the set of tied candidates 𝒞val\\mathcal\{C\}\_\{\\mathrm\{val\}\}as in \([77](https://arxiv.org/html/2608.25181#S4.E77)\) Rank candidates in 𝒞val\\mathcal\{C\}\_\{\\mathrm\{val\}\}according to 1. 1\.lower model complexity \(Table[1](https://arxiv.org/html/2608.25181#S4.T1)\), 2. 2\.lower residual errorEresE\_\{\\mathrm\{res\}\}\. Select the highest\-ranked candidate c∗c^\{\*\} else Declare that no candidate within the proposed framework adequately explains the observed dataset end if ## 5Results ### 5\.1Implementation details In this section, we describe the implementation details common to the numerical experiments presented throughout the manuscript\. The models considered here span a range of collective behaviors, including clustering, flocking, milling, and synchronization\. We study both first\- and second\-order systems, as well as models with and without environmental forcing\. The examples are organized by the types of mechanisms present in the dynamics, with the presentation progressing generally from simpler to more complex model systems\. For each system, we generate replicates via the distributionμ0\\mu\_\{0\}of initial conditions for training\(Mtrain\)\(M\_\{\\mathrm\{train\}\}\), testing\(Mtest\)\(M\_\{\\mathrm\{test\}\}\), and validation\(Mval\)\(M\_\{\\mathrm\{val\}\}\); recall that validation is only utilized for results relating to model selection as discussed in Section[4](https://arxiv.org/html/2608.25181#S4)\. Initial conditions are sampled independently from the prescribed distributionμ0\\mu\_\{0\}, and the governing equations are numerically solved from00toTfT\_\{f\}\.111Simulation data are generated using the Pythonsolve\_ivpfunction from thescipy\.integratepackage with the BDF and Radau methods\. Default relative and absolute tolerances are used throughout\.To obtain approximations for the induced empirical measures on interaction distances and state \(ρ^R\\hat\{\\rho\}\_\{R\}andρ^X\\hat\{\\rho\}\_\{X\}, respectively\), we also performMρM\_\{\\rho\}independent samples ofμ0\\mu\_\{0\}, as discussed in Section[3\.1](https://arxiv.org/html/2608.25181#S3.SS1)\. Estimators are obtained on theMtrainM\_\{\\mathrm\{train\}\}samples restricted to the observation interval \(\[0,T\]\[0,T\]\)\. The same observation data is also used to construct the induced probability measuresρ^R\\hat\{\\rho\}\_\{R\}andρ^X\\hat\{\\rho\}\_\{X\}, which approximate the underlying empirical probability measuresρR\\rho\_\{R\}andρX\\rho\_\{X\}; again recall that the latter are constructed utilizing theMρM\_\{\\rho\}samples\. To quantify the variation of the estimators that arises via the distributionμ0\\mu\_\{0\}, we generateTr=10T\_\{r\}=10independent training replicates, with each training replicate containingMtrainM\_\{\\mathrm\{train\}\}samples \(so a total ofMtrain×TrM\_\{\\mathrm\{train\}\}\\times T\_\{r\}samples\), together with a test sample \(of sizeMtestM\_\{\\mathrm\{test\}\}\), a validation sample \(of sizeMvalM\_\{\\mathrm\{val\}\}\), and an induced measure\-estimating sample \(of sizeMρM\_\{\\rho\}\)\. Means and standard deviations will be constructed from theTrT\_\{r\}training replicates\. The general trajectory data takes the form\(Xℓ\(m\),Vℓ\(m\),Aℓ\(m\)\)m,ℓ=1M,L\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\},A^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}as described in Section[2\.1\.1](https://arxiv.org/html/2608.25181#S2.SS1.SSS1)and Appendix[A\.1](https://arxiv.org/html/2608.25181#A1.SS1)\. From this data, we construct the interaction\-distance data\(Rℓ\(m\)\)m,ℓ=1M,L\(R^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}, which contains the relative pairwise information required for estimating interaction kernels\. Initial condition distributions utilized are provided in Table[24](https://arxiv.org/html/2608.25181#A5.T24)in Appendix[E\.3](https://arxiv.org/html/2608.25181#A5.SS3)\. The hypothesis spaces are constructed directly from the observed trajectory data as discussed in Section[2\.1\.4](https://arxiv.org/html/2608.25181#S2.SS1.SSS4)\. The interaction kernelϕ\\phiis estimated on the interval\[Rmin,Rmax\]\[R\_\{\\min\},R\_\{\\max\}\], while environmental forces are learned on hyper\-rectangles derived from component\-wise ranges\[Xmin,Xmax\]\[X\_\{\\min\},X\_\{\\max\}\]and\[Vmin,Vmax\]\[V\_\{\\min\},V\_\{\\max\}\]obtained from\(Xℓ\(m\),Vℓ\(m\)\)m,ℓ=1M,L\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}\. Throughout the manuscript, we employ piecewise\-linear B\-spline basis functions; the number of basis functions and other hyper\-parameters are reported in Appendix[E\.2](https://arxiv.org/html/2608.25181#A5.SS2)\. For all model systems, we report the evaluation metrics introduced in Section[3](https://arxiv.org/html/2608.25181#S3), including feature recovery errors \(Section[3\.2](https://arxiv.org/html/2608.25181#S3.SS2)\), residual errors \(Section[3\.3](https://arxiv.org/html/2608.25181#S3.SS3)\), trajectory errors \(Section[3\.4](https://arxiv.org/html/2608.25181#S3.SS4)\), and consistency measures \(Section[3\.5](https://arxiv.org/html/2608.25181#S3.SS5)\)\. The presentation of results is organized as follows\. We begin by introducing each mathematical model discuss its relevance to the proposed learning framework presented in Sections[2\.1](https://arxiv.org/html/2608.25181#S2.SS1)and Appendix[A](https://arxiv.org/html/2608.25181#A1)\. We then present the estimated interaction kernels and environmental forces, together with trajectory comparisons and evaluation metrics\. For the model selection experiments, we present the a summary of candidate frameworks, tables demonstrating model selection, and supporting visualizations\. Further details relating to data generation may be found in Table[22](https://arxiv.org/html/2608.25181#A5.T22), algorithm parameters in Table[23](https://arxiv.org/html/2608.25181#A5.T23), and estimates from the semi\-parametric approach in Table[25](https://arxiv.org/html/2608.25181#A5.T25)\. ### 5\.2Semi\-parametric versus fully nonparametric variational learning In this section, we compare estimated interaction kernelsϕ^\\hat\{\\phi\}and environmental forcesf^\\hat\{f\}together with generated trajectory dynamics from both the semi\-parametric \(SPSP\) and fully nonparametric \(NPNP\) variants of the proposed learning framework\. Comparisons are conducted through three representative systems: the Kuramoto model of synchronization \(Section[5\.2\.1](https://arxiv.org/html/2608.25181#S5.SS2.SSS1)\), a model self\-propelled particles \(SPP\) \(Section[5\.2\.2](https://arxiv.org/html/2608.25181#S5.SS2.SSS2)\), and a model of phototaxis \(Section[5\.2\.3](https://arxiv.org/html/2608.25181#S5.SS2.SSS3)\)\. For each system, we compare the ability of the two methods to recover the mechanistic forces, their consistency with respect to the error functional utilized in estimation, and the predicted trajectory accuracy obtained by the two learning approaches\. The primary objective is to assess the advantages and limitations of incorporating prior parametric knowledge into the environmental force model relative to the fully nonparametric formulation\. #### 5\.2\.1Kuramoto model The Kuramoto model, initially introduced in\[[61](https://arxiv.org/html/2608.25181#bib.bib41)\], describes synchronization phenomena in large populations of coupled oscillators\. The model was initially motivated by collective behavior observed in chemical and biological oscillators\[[80](https://arxiv.org/html/2608.25181#bib.bib42)\]and has since found applications in neuroscience\[[81](https://arxiv.org/html/2608.25181#bib.bib43)\], electrical power\-grid dynamics\[[82](https://arxiv.org/html/2608.25181#bib.bib44),[83](https://arxiv.org/html/2608.25181#bib.bib45)\], and oscillatory combustion and flame systems\[[84](https://arxiv.org/html/2608.25181#bib.bib47),[85](https://arxiv.org/html/2608.25181#bib.bib46)\]\. For a review of applications of the Kuramoto model to computation in swarming and synchronization, we refer the interested reader to\[[86](https://arxiv.org/html/2608.25181#bib.bib48)\], while a historical review of the development of synchronization models can be found in\[[87](https://arxiv.org/html/2608.25181#bib.bib14)\]\. The Kuramoto model is a first\-order interacting particle system whose state variable is phaseθi∈ℝ\\theta\_\{i\}\\in\\mathbb\{R\}\. The governing equations are given by θ˙i=ωi\+1N∑j=1NKijsin\(θj−θi\),\\displaystyle\\dot\{\\theta\}\_\{i\}=\\omega\_\{i\}\+\\frac\{1\}\{N\}\\sum\_\{j=1\}^\{N\}K\_\{ij\}\\sin\(\\theta\_\{j\}\-\\theta\_\{i\}\),\(78\)whereNNdenotes the number of oscillators,θi\\theta\_\{i\}represents the phase of theithi^\{\\mathrm\{th\}\}oscillator,ωi=ωi\(θi\)\\omega\_\{i\}=\\omega\_\{i\}\(\\theta\_\{i\}\)denotes its intrinsic \(generally state dependent\) frequency, andKijK\_\{ij\}is the coupling strength between oscillatorsiiandjj\. In numerical simulations presented below, the coupling strength is assumed to be homogeneous between agents, so thatKij=KK\_\{ij\}=K, and all oscillators share the same functional form of intrinsic frequencyωi=ω\(θi\)\\omega\_\{i\}=\\omega\(\\theta\_\{i\}\)\. Under these assumptions, the model can be written in the form of a first\-order system \([4](https://arxiv.org/html/2608.25181#S2.E4)\), wherexi=θi,f\(x\)=ω\(x\)x\_\{i\}=\\theta\_\{i\},f\(x\)=\\omega\(x\), andϕ\(r\)=Ksin\(r\)/r\\phi\(r\)=K\\sin\(r\)/r\. For all experiments reported in this section,Kij=2K\_\{ij\}=2andω\(x\)=x/100\\omega\(x\)=x/100\. For the semi\-parametric approach, we note that the environmental force term possesses a single parameterp=1/100=0\.01p=1/100=0\.01\. We begin by comparing estimators for the interaction kernelϕ\\phiand the environmental forcefffor both the semi\-parametric \(SP\) and fully non\-parametric \(NP\) version of the algorithm; results are visualized in Figure[1](https://arxiv.org/html/2608.25181#S5.F1)\. Note that in the semi\-parametric algorithm, we estimatep^=0\.0100000000000000141\\hat\{p\}=0\.0100000000000000141\(see Table[25](https://arxiv.org/html/2608.25181#A5.T25)\)\. We also observe that the interaction kernelϕ\\phiis estimated almost identically by both approaches, with relative errors provided in Table[4](https://arxiv.org/html/2608.25181#S5.T4)on the order of10−210^\{\-2\}, suggesting that the inference ofϕ\\phiis relatively insensitive to the approach, given the data utilized\. For the environmental force, the semi\-parametric model recoversffessentially to machine precision\. The fully non\-parametric formulation nevertheless approximates the environmental forceffaccurately, with a relative errorEfrel≈2\.0×102E\_\{f\}^\{\\mathrm\{rel\}\}\\approx 2\.0\\times 10^\{2\}\. Indeed, despite the large difference inEfrelE^\{\\mathrm\{rel\}\}\_\{f\}between methods, there is visually no discernible difference between the true and predicted trajectories, which are generated from bothMtrainM\_\{\\mathrm\{train\}\}andMtestM\_\{\\mathrm\{test\}\}on the full time horizon\[0,Tf\]\[0,T\_\{f\}\], as observed in Figure[2](https://arxiv.org/html/2608.25181#S5.F2)\. The trajectory errors are also quantified via the metrics discussed in Section[3\.4](https://arxiv.org/html/2608.25181#S3.SS4), and are provided in Table[5](https://arxiv.org/html/2608.25181#S5.T5)\. Note that the semi\-parametric and fully non\-parametric methods produce both small and similar errors with respect to trajectories \(both on training and prediction data\), suggesting that the two approaches provide statistically equivalent performance for the Kuramoto system on the utilized data\. \(a\)Estimation of interaction kernelϕ\\phi\(b\)Estimation of environmental forceff Figure 1:\(Kuramoto model\) Comparison of the mechanistic feature recovery ability of the fully non\-parametric \(NP\) and semi\-parametric \(SP\) estimation algorithms for the Kuramoto model \([78](https://arxiv.org/html/2608.25181#S5.E78)\)\. In Subfigure[1\(a\)](https://arxiv.org/html/2608.25181#S5.F1.sf1)we visualize the true interaction kernelϕ\\phi\(blue\) together with the mean recovered interaction kernelsϕ^¯\\bar\{\\hat\{\\phi\}\}obtained via the NP \(yellow\) and SP \(orange\) approaches\. Subfigure[1\(b\)](https://arxiv.org/html/2608.25181#S5.F1.sf2)provides the corresponding true environmental forceffand its recovered mean estimatesf^¯\\bar\{\\hat\{f\}\}\. Mean estimates are obtained overTr=10T\_\{r\}=10independent learning trials, and the corresponding standard deviations are also plotted\. The background density in Subfigure[1\(a\)](https://arxiv.org/html/2608.25181#S5.F1.sf1)\(right vertical axis\) corresponds to the empirical distributionρ^Δθ\\hat\{\\rho\}\_\{\\Delta\\theta\}of pairwise phase differences, while in Subfigure[1\(b\)](https://arxiv.org/html/2608.25181#S5.F1.sf2)this corresponds to the empirical distributionρ^θ\\hat\{\\rho\}\_\{\\theta\}of oscillator phases observed in the training data\. The corresponding relative feature recovery errors, computed via equation \([52](https://arxiv.org/html/2608.25181#S3.E52)\), are reported in Table[4](https://arxiv.org/html/2608.25181#S5.T4)\. Both formulations recover the interaction kernel with comparable accuracy, while the semi\-parametric formulation achieves nearly exact recovery of the environmental force\. The training data consists ofM=20M=20trajectories withN=10N=10oscillators in dimensiond=1d=1, observed over the interval\[0,T\]=\[0,0\.5\]\[0,T\]=\[0,0\.5\]withL=51L=51time points\.MetricNon\-parametricSemi\-ParametricEresrelE^\{\\mathrm\{rel\}\}\_\{\\mathrm\{res\}\}9\.133 297 54×10−039\.133\\,297\\,54\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.069 999 32×10−031\.069\\,999\\,32\\text\{\\times\}\{10\}^\{\-03\}9\.213 417 49×10−039\.213\\,417\\,49\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.074 338 81×10−031\.074\\,338\\,81\\text\{\\times\}\{10\}^\{\-03\}EϕrelE^\{\\mathrm\{rel\}\}\_\{\\phi\}1\.151 949 00×10−021\.151\\,949\\,00\\text\{\\times\}\{10\}^\{\-02\}±\\pm6\.567 552 90×10−046\.567\\,552\\,90\\text\{\\times\}\{10\}^\{\-04\}1\.150 350 07×10−021\.150\\,350\\,07\\text\{\\times\}\{10\}^\{\-02\}±\\pm6\.505 249 25×10−046\.505\\,249\\,25\\text\{\\times\}\{10\}^\{\-04\}EfrelE^\{\\mathrm\{rel\}\}\_\{f\}1\.992 020 20×10−021\.992\\,020\\,20\\text\{\\times\}\{10\}^\{\-02\}±\\pm5\.681 235 00×10−035\.681\\,235\\,00\\text\{\\times\}\{10\}^\{\-03\}1\.712 688 22×10−141\.712\\,688\\,22\\text\{\\times\}\{10\}^\{\-14\}±\\pm1\.184 873 90×10−141\.184\\,873\\,90\\text\{\\times\}\{10\}^\{\-14\}Table 4:\(Kuramoto model\) Comparison of relative residual error \([55](https://arxiv.org/html/2608.25181#S3.E55)\) and feature recovery metrics \([52](https://arxiv.org/html/2608.25181#S3.E52)\) for the fully non\-parametric and semi\-parametric estimation algorithms for the Kuramoto model \([78](https://arxiv.org/html/2608.25181#S5.E78)\)\. The table reports the mean error±\\pmone standard deviation overTr=10T\_\{r\}=10independent learning trials\. For additional details on algorithm parameters, we refer the reader to the caption of Figure[1](https://arxiv.org/html/2608.25181#S5.F1)\.\(a\)Semi\-parametric trajectories onMtrainM\_\{\\mathrm\{train\}\}training data\(b\)Fully non\-parametric trajectories onMtrainM\_\{\\mathrm\{train\}\}training data\(c\)Semi\-parametric trajectories onMtestM\_\{\\mathrm\{test\}\}testing data\(d\)Fully non\-parametric trajectories onMtestM\_\{\\mathrm\{test\}\}testing data Figure 2:\(Kuramoto model\) Comparison of the true phase trajectoriesθi\\theta\_\{i\}\(blue\) and the estimated phase trajectoriesθ^i\\hat\{\\theta\}\_\{i\}\(yellow\) for both the semi\-parametric \(left column\) and non\-parametric \(right column\) methods\. For trajectory errors, see Table[5](https://arxiv.org/html/2608.25181#S5.T5)\. The top row displays trajectories initialized from theMtrainM\_\{\\mathrm\{train\}\}training trajectories, whereas the bottom rows displays trajectories generated fromMtestM\_\{\\mathrm\{test\}\}previously unseen testing initial conditions\. Note that trajectories are plotted over the full time horizon\[0,Tf\]\[0,T\_\{f\}\], with the end of the training interval timeTTindicated on the horizontal axis\. Recall that models are trained using observations collected only over the interval\[0,T\]=\[0,0\.5\]\[0,T\]=\[0,0\.5\], and are evaluated beyond the fitting horizon to assess both reconstruction and predictive performance\. Both method variants closely reproduce the true dynamics on the training trajectories and generalize well to unseen initial conditions, with no visual differences between the semi\-parametric and non\-parametric models\. For additional details on algorithm parameters, we refer the reader to the caption of Figure[1](https://arxiv.org/html/2608.25181#S5.F1)\.Data setStatisticFully non\-parametricSemi\-parametricMtrainM\_\{\\mathrm\{train\}\}E¯recon\\bar\{E\}\_\{\\mathrm\{recon\}\}1\.486 115 21×10−031\.486\\,115\\,21\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.290 008 01×10−041\.290\\,008\\,01\\text\{\\times\}\{10\}^\{\-04\}1\.515 174 72×10−031\.515\\,174\\,72\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.326 100 95×10−041\.326\\,100\\,95\\text\{\\times\}\{10\}^\{\-04\}σErecon\\sigma\_\{E\_\{\\mathrm\{recon\}\}\}7\.059 050 63×10−047\.059\\,050\\,63\\text\{\\times\}\{10\}^\{\-04\}±\\pm1\.149 350 14×10−041\.149\\,350\\,14\\text\{\\times\}\{10\}^\{\-04\}7\.258 467 47×10−047\.258\\,467\\,47\\text\{\\times\}\{10\}^\{\-04\}±\\pm1\.074 815 46×10−041\.074\\,815\\,46\\text\{\\times\}\{10\}^\{\-04\}E¯pred\\bar\{E\}\_\{\\mathrm\{pred\}\}5\.457 046 59×10−035\.457\\,046\\,59\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.454 816 20×10−031\.454\\,816\\,20\\text\{\\times\}\{10\}^\{\-03\}5\.459 184 78×10−035\.459\\,184\\,78\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.409 017 06×10−031\.409\\,017\\,06\\text\{\\times\}\{10\}^\{\-03\}σEpred\\sigma\_\{E\_\{\\mathrm\{pred\}\}\}5\.665 476 87×10−035\.665\\,476\\,87\\text\{\\times\}\{10\}^\{\-03\}±\\pm2\.895 039 32×10−032\.895\\,039\\,32\\text\{\\times\}\{10\}^\{\-03\}6\.020 934 52×10−036\.020\\,934\\,52\\text\{\\times\}\{10\}^\{\-03\}±\\pm2\.947 784 88×10−032\.947\\,784\\,88\\text\{\\times\}\{10\}^\{\-03\}MtestM\_\{\\mathrm\{test\}\}E¯recon\\bar\{E\}\_\{\\mathrm\{recon\}\}1\.635 224 03×10−031\.635\\,224\\,03\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.994 618 41×10−041\.994\\,618\\,41\\text\{\\times\}\{10\}^\{\-04\}1\.598 611 49×10−031\.598\\,611\\,49\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.293 745 35×10−041\.293\\,745\\,35\\text\{\\times\}\{10\}^\{\-04\}σErecon\\sigma\_\{E\_\{\\mathrm\{recon\}\}\}9\.230 112 42×10−049\.230\\,112\\,42\\text\{\\times\}\{10\}^\{\-04\}±\\pm2\.374 169 98×10−042\.374\\,169\\,98\\text\{\\times\}\{10\}^\{\-04\}9\.094 150 34×10−049\.094\\,150\\,34\\text\{\\times\}\{10\}^\{\-04\}±\\pm2\.159 352 77×10−042\.159\\,352\\,77\\text\{\\times\}\{10\}^\{\-04\}E¯pred\\bar\{E\}\_\{\\mathrm\{pred\}\}5\.106 476 67×10−035\.106\\,476\\,67\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.436 456 17×10−031\.436\\,456\\,17\\text\{\\times\}\{10\}^\{\-03\}4\.785 536 71×10−034\.785\\,536\\,71\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.239 186 08×10−031\.239\\,186\\,08\\text\{\\times\}\{10\}^\{\-03\}σEpred\\sigma\_\{E\_\{\\mathrm\{pred\}\}\}3\.901 501 89×10−033\.901\\,501\\,89\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.104 169 66×10−031\.104\\,169\\,66\\text\{\\times\}\{10\}^\{\-03\}3\.714 524 13×10−033\.714\\,524\\,13\\text\{\\times\}\{10\}^\{\-03\}±\\pm8\.680 604 98×10−048\.680\\,604\\,98\\text\{\\times\}\{10\}^\{\-04\}Table 5:\(Kuramoto model\) Comparison of trajectory\-level errors for the fully non\-parametric and semi\-parametric estimation algorithms for Kuramoto model \([78](https://arxiv.org/html/2608.25181#S5.E78)\)\. For each learning trial, the trajectory statisticsE¯recon\\bar\{E\}\_\{\\mathrm\{recon\}\},E¯pred\\bar\{E\}\_\{\\mathrm\{pred\}\},σErecon\\sigma\_\{E\_\{\\mathrm\{recon\}\}\}, andσEpred\\sigma\_\{E\_\{\\mathrm\{pred\}\}\}are computed according to \([63](https://arxiv.org/html/2608.25181#S3.E63)\) \- \([64](https://arxiv.org/html/2608.25181#S3.E64)\)\. The table reports the mean±\\pmone standard deviation of these quantities computed onTr=10T\_\{r\}=10independent learning trials\. The first section corresponds to trajectories generated fromMtrainM\_\{\\mathrm\{train\}\}training initial conditions, while the second section corresponds toMtestM\_\{\\mathrm\{test\}\}unseen predicted initial conditions sampled fromμ0\\mu\_\{0\}\. Note that across all trajectory measures, error estimates are both small and similar for both methods\. For additional details on algorithm parameters, we refer the reader to the caption of Figure[1](https://arxiv.org/html/2608.25181#S5.F1)\. #### 5\.2\.2Self\-propelled particle \(SPP\) model Understanding how organisms move in cohesive groups has been an extensively studied mathematical modeling problem\. Indeed, it is known in nature that such systems can exhibit various types of behavior, including flocking, where agents converge to a common velocity, milling, where agents rotate around a common center or axis, and swarming, which often appears as a transitional state between flocking and milling\. In this section we consider a model of self\-propelled particles \(SPPs\), which was first introduced in\[[56](https://arxiv.org/html/2608.25181#bib.bib30),[88](https://arxiv.org/html/2608.25181#bib.bib13)\]and further studied in\[[15](https://arxiv.org/html/2608.25181#bib.bib33),[16](https://arxiv.org/html/2608.25181#bib.bib31)\], and can exhibit phenomenon such as clustering, milling, and flocking\. These models are motivated by biological aggregation phenomena and have been widely used to study collective motion\. For additional background on SPP models, we refer the reader to\[[89](https://arxiv.org/html/2608.25181#bib.bib34),[90](https://arxiv.org/html/2608.25181#bib.bib35),[91](https://arxiv.org/html/2608.25181#bib.bib36)\]; mathematical analysis of the dynamics of such systems may be found in\[[92](https://arxiv.org/html/2608.25181#bib.bib37),[93](https://arxiv.org/html/2608.25181#bib.bib38),[15](https://arxiv.org/html/2608.25181#bib.bib33)\]\. The governing equations of the SPP system considered here take the form x˙i=viv˙i=\(α−β\|vi\|2\)vi−∇xiU\(xi\)\\displaystyle\\begin\{split\}\\dot\{x\}\_\{i\}&=v\_\{i\}\\\\ \\dot\{v\}\_\{i\}&=\(\\alpha\-\\beta\|v\_\{i\}\|^\{2\}\)v\_\{i\}\-\\nabla\_\{x\_\{i\}\}U\(x\_\{i\}\)\\end\{split\}\(79\)where the generalized Morse interaction potentialUUis given by U\(xi\)\\displaystyle U\(x\_\{i\}\)=∑j≠iNCre−\|xi−xj\|/lr−Cae−\|xi−xj\|/la,\\displaystyle=\\sum\_\{j\\neq i\}^\{N\}C\_\{r\}e^\{\-\|x\_\{i\}\-x\_\{j\}\|/l\_\{r\}\}\-C\_\{a\}e^\{\-\|x\_\{i\}\-x\_\{j\}\|/l\_\{a\}\},\(80\)and combines the effects of attracting and repulsion between agents\. Hereα\\alphascales the strength of self\-propulsion,β\\betamodels a nonlinear damping effect,lal\_\{a\}andlrl\_\{r\}are the attractive and repulsive interaction ranges, respectively, andCaC\_\{a\}andCrC\_\{r\}denote the corresponding interaction strengths\. Note that the system \([79](https://arxiv.org/html/2608.25181#S5.E79)\) is second\-order\. For numerical experiments, we utilize the following parameters \(as reported in\[[88](https://arxiv.org/html/2608.25181#bib.bib13)\]\):Cr=0\.6,lr=0\.5,Ca=1,la=1,α=1,β=0\.5C\_\{r\}=0\.6,l\_\{r\}=0\.5,C\_\{a\}=1,l\_\{a\}=1,\\alpha=1,\\beta=0\.5\. The remaining simulation and learning parameters are reported in Tables[22](https://arxiv.org/html/2608.25181#A5.T22),[23](https://arxiv.org/html/2608.25181#A5.T23), and[24](https://arxiv.org/html/2608.25181#A5.T24)of Appendix[E](https://arxiv.org/html/2608.25181#A5)\. Equations \([79](https://arxiv.org/html/2608.25181#S5.E79)\) take the general form of a second\-order collective system \([85](https://arxiv.org/html/2608.25181#A1.E85)\), where f\(xi,vi\)=\(α−β\|vi\|2\)viandϕ\(r\)=Nr\(−Crlre−r/lr\+Calae−r/la\)\\displaystyle\\begin\{split\}f\(x\_\{i\},v\_\{i\}\)&=\(\\alpha\-\\beta\|v\_\{i\}\|^\{2\}\)v\_\{i\}\\quad\\text\{and\}\\quad\\phi\(r\)=\\frac\{N\}\{r\}\\left\(\\frac\{\-C\_\{r\}\}\{l\_\{r\}\}e^\{\-r/l\_\{r\}\}\+\\frac\{C\_\{a\}\}\{l\_\{a\}\}e^\{\-r/l\_\{a\}\}\\right\)\\end\{split\}\(81\)Note that the number of agentsNNappears inϕ\\phidue to the fact that that the learning framework averages interaction forces for each agent \(1/N1/Nin \([85](https://arxiv.org/html/2608.25181#A1.E85)\)\), whereas the standard SPP system \([79](https://arxiv.org/html/2608.25181#S5.E79)\) does not\. Figure[3\(a\)](https://arxiv.org/html/2608.25181#S5.F3.sf1)provides a plot of the obtained estimators for the interaction kernelϕ\\phifrom both methods, and demonstrates that they both closely approximateϕ\\phiover the region where the empirical pairwise distance distributionρ^R\\hat\{\\rho\}\_\{R\}is well sampled by the observed trajectory data, i\.e\. where the empirical pairwise distance distributionρ^R\\hat\{\\rho\}\_\{R\}has non\-negligible mass\. Note that when the interaction distance is close to00, the kernelϕ\\phiin \([81](https://arxiv.org/html/2608.25181#S5.E81)\) is singular, and as a result, there exist a limited number of trajectories with small interaction distancesrr\. As we utilize a uniformly\-spaced basis of splines to approximateϕ\\phi, the reduced information content near00leads to visible deviations between the true kernelϕ\\phiand its estimatorsϕ^\\hat\{\\phi\}\. Similarly, Figure[3\(b\)](https://arxiv.org/html/2608.25181#S5.F3.sf2)plots both the true environmental self\-propelled forceff\(left column\) together with estimators from the fully non\-parametric \(middle column\) and semi\-parametric \(right column\) approaches; note that the top row corresponds to the first component off:ℝ2→ℝ2f:\\mathbb\{R\}^\{2\}\\to\\mathbb\{R\}^\{2\}in \([81](https://arxiv.org/html/2608.25181#S5.E81)\) while the bottom row corresponds to the second component\. The approximate empirical distributionρ^V\\hat\{\\rho\}\_\{V\}is also included on the surface plot offfto highlight the regions of velocity explored by the training trajectory data\. Again, in both cases, we observe qualitatively close agreement between estimators and the components offf\. Table[6](https://arxiv.org/html/2608.25181#S5.T6)quantifies the accuracy the inferring mechanisms, demonstrating that the learning approaches recoverϕ\\phiwith nearly identical accuracy, yieldingEϕrel≈0\.12E^\{\\mathrm\{rel\}\}\_\{\\phi\}\\approx 0\.12\. In contrast, the recovery offfimproves substantially under the semi\-parametric formulation, with the relative error reduced by nearly two orders of magnitude\. The truex\(t\)x\(t\)and estimatedx^\(t\)\\hat\{x\}\(t\)trajectories are shown in Figure[4](https://arxiv.org/html/2608.25181#S5.F4)on both training \(Figure[4\(a\)](https://arxiv.org/html/2608.25181#S5.F4.sf1)\) and testing \(Figure[4\(b\)](https://arxiv.org/html/2608.25181#S5.F4.sf2)\) samples\. Quantitative trajectory errors reported in Table[7](https://arxiv.org/html/2608.25181#S5.T7)remain small \(≈10−2\\approx 10^\{\-2\}\), demonstrating that the learned models preserve the long\-time collective behavior of the system\. \(a\)Estimation of interaction kernelϕ\\phi\(b\)Estimation of environmental forceff Figure 3:\(SPP model\) Comparison of mechanistic feature recovery ability of the fully non\-parametric \(NP\) and semi\-parametric \(SP\) estimation algorithms for the SPP model \([79](https://arxiv.org/html/2608.25181#S5.E79)\)\. In Subfigure[3\(a\)](https://arxiv.org/html/2608.25181#S5.F3.sf1)we visualize the true interaction kernelϕ\\phi\(blue\) with the mean recovered kernelsϕ^¯\\bar\{\\hat\{\\phi\}\}obtained via the NP \(yellow\) and SP \(orange\) formulations\. Subfigure[3\(b\)](https://arxiv.org/html/2608.25181#S5.F3.sf2)provides the true environmental forceffwith the corresponding mean learned estimatesf^¯\\bar\{\\hat\{f\}\}\. Asf:ℝ2→ℝ2f:\\mathbb\{R\}^\{2\}\\to\\mathbb\{R\}^\{2\}, the top and bottom rows correspond to the first and second components offf, respectively\. Within each row, the true force \(left column\), the fully non\-parametric mean estimate \(middle column\), and the semi\-parametric mean estimate \(right column\) are shown, together with one standard deviation, which are obtained overTr=10T\_\{r\}=10independent learning trials\. The background density in Subfigure[3\(a\)](https://arxiv.org/html/2608.25181#S5.F3.sf1)\(right vertical axis\) corresponds to the empirical pairwise distance distributionρ^R\\hat\{\\rho\}\_\{R\}, while the density visualized in Subfigure[3\(b\)](https://arxiv.org/html/2608.25181#S5.F3.sf2)\(left column\) corresponds to the empirical velocity distributionρ^V\\hat\{\\rho\}\_\{V\}induced by the training data\. The corresponding relative feature recovery errors are reported in Table[6](https://arxiv.org/html/2608.25181#S5.T6)\. The training data consists ofM=20M=20trajectories withN=10N=10agents in dimensiond=2d=2, observed over\[0,T\]=\[0\.0,0\.5\]\[0,T\]=\[0\.0,0\.5\]withL=51L=51time points\.MetricNon\-parametricSemi\-ParametricEresrelE^\{\\mathrm\{rel\}\}\_\{\\mathrm\{res\}\}3\.250 636 15×10−023\.250\\,636\\,15\\text\{\\times\}\{10\}^\{\-02\}±\\pm3\.804 678 17×10−043\.804\\,678\\,17\\text\{\\times\}\{10\}^\{\-04\}1\.464 225 90×10−021\.464\\,225\\,90\\text\{\\times\}\{10\}^\{\-02\}±\\pm7\.416 366 49×10−047\.416\\,366\\,49\\text\{\\times\}\{10\}^\{\-04\}EϕrelE^\{\\mathrm\{rel\}\}\_\{\\phi\}1\.199 870 60×10−011\.199\\,870\\,60\\text\{\\times\}\{10\}^\{\-01\}±\\pm1\.993 494 02×10−031\.993\\,494\\,02\\text\{\\times\}\{10\}^\{\-03\}1\.197 529 15×10−011\.197\\,529\\,15\\text\{\\times\}\{10\}^\{\-01\}±\\pm1\.841 751 39×10−031\.841\\,751\\,39\\text\{\\times\}\{10\}^\{\-03\}EfrelE^\{\\mathrm\{rel\}\}\_\{f\}2\.974 167 42×10−022\.974\\,167\\,42\\text\{\\times\}\{10\}^\{\-02\}±\\pm3\.449 852 01×10−043\.449\\,852\\,01\\text\{\\times\}\{10\}^\{\-04\}6\.480 458 97×10−046\.480\\,458\\,97\\text\{\\times\}\{10\}^\{\-04\}±\\pm1\.344 640 77×10−041\.344\\,640\\,77\\text\{\\times\}\{10\}^\{\-04\}Table 6:\(SPP model\) Comparison of relative residual error \([55](https://arxiv.org/html/2608.25181#S3.E55)\) and feature recovery metrics \([52](https://arxiv.org/html/2608.25181#S3.E52)\) for the fully non\-parametric and semi\-parametric estimation algorithms for the SPP model \([79](https://arxiv.org/html/2608.25181#S5.E79)\)\. The table reports the mean error±\\pmone standard deviation of these quantities computed overTr=10T\_\{r\}=10independent learning trials\. For additional details on algorithm parameters, we refer the reader to the caption of Figure[3](https://arxiv.org/html/2608.25181#S5.F3)\.\(a\)MtrainM\_\{\\mathrm\{train\}\}training data utilized to obtain estimators \(b\)MtestM\_\{\\mathrm\{test\}\}test data utilized for evaluation Figure 4:\(SPP model\) Comparison of the true trajectoriesxx\(left column\) and the estimated trajectoriesx^\\hat\{x\}for both the semi\-parametric \(middle column\) and fully non\-parametric \(right column\) methods in the phase plane\. For trajectory errors, see Table[7](https://arxiv.org/html/2608.25181#S5.T7)\. The top row displays trajectories initialized from theMtrainM\_\{\\mathrm\{train\}\}training initial\-conditions, whereas the bottom row displays trajectories generated fromMtestM\_\{\\mathrm\{test\}\}previously unseen testing initial conditions\. Note that trajectories are plotted over the full time horizon\[0,Tf\]\[0,T\_\{f\}\], with the color gradient representing the temporal evolution of the dynamics, from yellow att=0t=0to orange att=Tft=T\_\{f\}\. The black segment indicates the training interval\[0,T\]\[0,T\], which is utilized for model fitting; recall that no fitting is performed in the testing data in Figure[4\(b\)](https://arxiv.org/html/2608.25181#S5.F4.sf2)\. Here\[0,T\]=\[0\.0,0\.5\]\[0,T\]=\[0\.0,0\.5\], with testing evaluated beyond the fitting horizon\[0\.0,0\.5\]\[0\.0,0\.5\]untilTf=2\.0T\_\{f\}=2\.0to assess both reconstruction and prediction in Figure[4\(b\)](https://arxiv.org/html/2608.25181#S5.F4.sf2)\. Both method variants closely reproduce the true dynamics on the training trajectories and generalize well to unseen initial conditions\. For additional details on algorithm parameters, we refer the reader to the caption of Figure[3](https://arxiv.org/html/2608.25181#S5.F3)\.Data setStatisticFully non\-parametricSemi\-parametricMtrainM\_\{\\mathrm\{train\}\}E¯recon\\bar\{E\}\_\{\\mathrm\{recon\}\}4\.291 656 12×10−034\.291\\,656\\,12\\text\{\\times\}\{10\}^\{\-03\}±\\pm7\.329 969 06×10−057\.329\\,969\\,06\\text\{\\times\}\{10\}^\{\-05\}2\.762 636 52×10−032\.762\\,636\\,52\\text\{\\times\}\{10\}^\{\-03\}±\\pm8\.642 627 33×10−058\.642\\,627\\,33\\text\{\\times\}\{10\}^\{\-05\}σErecon\\sigma\_\{E\_\{\\mathrm\{recon\}\}\}1\.018 780 56×10−031\.018\\,780\\,56\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.047 847 99×10−041\.047\\,847\\,99\\text\{\\times\}\{10\}^\{\-04\}9\.363 603 47×10−049\.363\\,603\\,47\\text\{\\times\}\{10\}^\{\-04\}±\\pm1\.232 918 49×10−041\.232\\,918\\,49\\text\{\\times\}\{10\}^\{\-04\}E¯pred\\bar\{E\}\_\{\\mathrm\{pred\}\}4\.969 235 82×10−024\.969\\,235\\,82\\text\{\\times\}\{10\}^\{\-02\}±\\pm2\.224 500 17×10−032\.224\\,500\\,17\\text\{\\times\}\{10\}^\{\-03\}2\.484 601 58×10−022\.484\\,601\\,58\\text\{\\times\}\{10\}^\{\-02\}±\\pm1\.481 024 68×10−031\.481\\,024\\,68\\text\{\\times\}\{10\}^\{\-03\}σEpred\\sigma\_\{E\_\{\\mathrm\{pred\}\}\}9\.785 216 05×10−039\.785\\,216\\,05\\text\{\\times\}\{10\}^\{\-03\}±\\pm9\.877 813 09×10−049\.877\\,813\\,09\\text\{\\times\}\{10\}^\{\-04\}9\.042 348 73×10−039\.042\\,348\\,73\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.314 571 31×10−031\.314\\,571\\,31\\text\{\\times\}\{10\}^\{\-03\}MtestM\_\{\\mathrm\{test\}\}E¯recon\\bar\{E\}\_\{\\mathrm\{recon\}\}4\.256 561 62×10−034\.256\\,561\\,62\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.291 396 91×10−041\.291\\,396\\,91\\text\{\\times\}\{10\}^\{\-04\}2\.767 011 67×10−032\.767\\,011\\,67\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.018 333 24×10−041\.018\\,333\\,24\\text\{\\times\}\{10\}^\{\-04\}σErecon\\sigma\_\{E\_\{\\mathrm\{recon\}\}\}1\.133 254 76×10−031\.133\\,254\\,76\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.405 147 45×10−041\.405\\,147\\,45\\text\{\\times\}\{10\}^\{\-04\}8\.268 659 24×10−048\.268\\,659\\,24\\text\{\\times\}\{10\}^\{\-04\}±\\pm9\.900 607 36×10−059\.900\\,607\\,36\\text\{\\times\}\{10\}^\{\-05\}E¯pred\\bar\{E\}\_\{\\mathrm\{pred\}\}4\.939 430 26×10−024\.939\\,430\\,26\\text\{\\times\}\{10\}^\{\-02\}±\\pm2\.915 931 63×10−032\.915\\,931\\,63\\text\{\\times\}\{10\}^\{\-03\}2\.473 505 08×10−022\.473\\,505\\,08\\text\{\\times\}\{10\}^\{\-02\}±\\pm8\.050 993 88×10−048\.050\\,993\\,88\\text\{\\times\}\{10\}^\{\-04\}σEpred\\sigma\_\{E\_\{\\mathrm\{pred\}\}\}8\.626 941 95×10−038\.626\\,941\\,95\\text\{\\times\}\{10\}^\{\-03\}±\\pm2\.362 227 74×10−032\.362\\,227\\,74\\text\{\\times\}\{10\}^\{\-03\}7\.037 997 59×10−037\.037\\,997\\,59\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.442 514 04×10−031\.442\\,514\\,04\\text\{\\times\}\{10\}^\{\-03\}Table 7:\(SPP model\) Comparison of trajectory\-level errors for the fully non\-parametric and semi\-parametric estimation algorithms for the SPP model \([79](https://arxiv.org/html/2608.25181#S5.E79)\)\. For each learning trial, the trajectory statisticsE¯recon\\bar\{E\}\_\{\\mathrm\{recon\}\},E¯pred\\bar\{E\}\_\{\\mathrm\{pred\}\},σErecon\\sigma\_\{E\_\{\\mathrm\{recon\}\}\}, andσEpred\\sigma\_\{E\_\{\\mathrm\{pred\}\}\}are computed according to \([63](https://arxiv.org/html/2608.25181#S3.E63)\) \- \([64](https://arxiv.org/html/2608.25181#S3.E64)\)\. The table reports the mean±\\pmone standard deviation of these quantities computed overTr=10T\_\{r\}=10independent learning trials\. The first section corresponds to trajectories generated fromMtrainM\_\{\\mathrm\{train\}\}training initial conditions, while the second section corresponds toMtestM\_\{\\mathrm\{test\}\}unseen initial conditions sampled fromμ0\\mu\_\{0\}\. Note that across all trajectory measures, error estimates are both small and similar for both methods\. For additional details on algorithm parameters, we refer the reader to the caption of Figure[3](https://arxiv.org/html/2608.25181#S5.F3)\.One limitation of the semi\-parametric algorithm is that its performance depends strongly on the correctness of the prescribed functional form for the environmental forceff\. If the trueffwhich generates the data cannot be realized in the assumed form, the algorithm will generally fail to recover the underlying features and predict trajectories\. Figure[5](https://arxiv.org/html/2608.25181#S5.F5)illustrates such a case for the SPP system\. Numerically, we assume the parametric form of the environmental force takes the formf\(v,p\)=p1v\+p2‖v‖1vf\(v;p\)=p\_\{1\}v\+p\_\{2\}\|\|v\|\|\_\{1\}v, wherep=\(p1,p2\)p=\(p\_\{1\},p\_\{2\}\)is the vector of parameters\. Note that the true environmental force is given by \([81](https://arxiv.org/html/2608.25181#S5.E81)\), which cannot be represented by the assumed form for anyp∈ℝ2p\\in\\mathbb\{R\}^\{2\}\. We then solve the semi\-parametric estimation algorithm on training data generated from \([81](https://arxiv.org/html/2608.25181#S5.E81)\), and the resulting estimators are provided in Figure[5](https://arxiv.org/html/2608.25181#S5.F5)\. Note that we obtain a marginally accurate \(non\-parametric\) estimator for the interaction kernelϕ\\phi\(Figure[5\(b\)](https://arxiv.org/html/2608.25181#S5.F5.sf2)\), but both the environmental force \(Figure[5\(c\)](https://arxiv.org/html/2608.25181#S5.F5.sf3)\) and trajectories \(Figure[5\(a\)](https://arxiv.org/html/2608.25181#S5.F5.sf1)\) are estimated poorly\. This therefore highlights the necessity of being confident in regards to prior knowledge of parametric forms in estimation, and demonstrates the utility of a fully non\-parametric approach in many experimental scenarios, when the exact structure of forces is rarely available\. \(a\)True \(left column\) and estimated \(right column\) trajectoriesxx \(b\)Interaction kernelϕ\\phi \(c\)True \(left column\) and estimated \(right column\) environmental forceff Figure 5:\(Misspecified parametric model, SPP\) Results for the SPP system \([79](https://arxiv.org/html/2608.25181#S5.E79)\) when the semi\-parametric algorithm is supplied with an incorrect functional form for the environmental forceff\. Hereffis assumed of the formf\(v,p\)=p1v\+p2‖v‖1vf\(v;p\)=p\_\{1\}v\+p\_\{2\}\|\|v\|\|\_\{1\}v, while the data is generated from the model \([79](https://arxiv.org/html/2608.25181#S5.E79)\)\. Note that the algorithm is unable to recover the correct mechanistic features and dynamics\. Subfigure[5\(a\)](https://arxiv.org/html/2608.25181#S5.F5.sf1)plots the resulting phase portraits for one initial condition from theMtrainM\_\{\\mathrm\{train\}\}training data; observe that even on the training trajectories, the learned dynamics deviate noticeably from the true dynamics within the predictive interval\[T,Tf\]\[T,T\_\{f\}\]\. Subfigure[5\(b\)](https://arxiv.org/html/2608.25181#S5.F5.sf2)compares the true interaction kernelϕ\\phi\(blue solid line\) with the learned kernelϕ^\\hat\{\\phi\}obtained from the misspecified parametric model \(red dashed line\)\. Subfigure[5\(c\)](https://arxiv.org/html/2608.25181#S5.F5.sf3)compares the true \(left column\) and estimated \(right column\) environmental forcesff, with the upper and lower panels corresponding to the first and second components, respectively\. For additional details on algorithm parameters, we refer the reader to the caption of Figure[3](https://arxiv.org/html/2608.25181#S5.F3)\. #### 5\.2\.3Phototaxis model We next consider a model of phototaxis, describing cell motion guided by an external light stimulus\. Phototaxis is a fundamental behavioral mechanism in many microorganisms, allowing individuals to bias their movement toward or away from light\[[94](https://arxiv.org/html/2608.25181#bib.bib40)\]\. In this work, we adapt the particle\-level model introduced by\[[55](https://arxiv.org/html/2608.25181#bib.bib32)\], in which the motion of each bacterium is also influenced by neighboring agents through a Cucker\-Smale velocity alignment term \(see also Section[5\.6\.1](https://arxiv.org/html/2608.25181#S5.SS6.SSS1)\)\. Specifically, the model takes the general second\-order form \([72](https://arxiv.org/html/2608.25181#S4.E72)\), where f\(x,v\)=I0\(U∞el−v\),ϕA\(r\)=1\(1\+r2\)β,andϕE\(r\)≡0\.\\displaystyle\\begin\{split\}f\(x,v\)&=I\_\{0\}\(U\_\{\\infty\}e\_\{l\}\-v\),\\quad\\phi^\{A\}\(r\)=\\frac\{1\}\{\(1\+r^\{2\}\)^\{\\beta\}\},\\quad\\text\{and\}\\quad\\phi^\{E\}\(r\)\\equiv 0\.\\end\{split\}\(82\)HereI0I\_\{0\}describe the constant intensity of the light source,U∞U\_\{\\infty\}denotes the terminal speed of the particles, andel∈ℝ2e\_\{l\}\\in\\mathbb\{R\}^\{2\}denotes the normalized direction of the light source\. For numerical simulations we fixI0=1I\_\{0\}=1,U∞=0\.1U\_\{\\infty\}=0\.1,el=\(−1/2,1/2\)e\_\{l\}=\(\-1/\\sqrt\{2\},1/\\sqrt\{2\}\), andβ=0\.1\\beta=0\.1\. We emphasize that this example differs significantly from the SPP model in Section[5\.2\.2](https://arxiv.org/html/2608.25181#S5.SS2.SSS2), as the SPP systems combines a nonlinear environmental force with energy\-based interactions, while the model of phototaxis \([82](https://arxiv.org/html/2608.25181#S5.E82)\) describes velocity alignment together with a directional environmental force\. Inferred mechanisms are plotted in Figure[6](https://arxiv.org/html/2608.25181#S5.F6)\. We observe a slight visible deviation between the true interaction kernelϕ\\phiand the estimated kernelsϕ^\\hat\{\\phi\}near the boundaries of the learning interval in both approaches in Figure[6\(a\)](https://arxiv.org/html/2608.25181#S5.F6.sf1), which is consistent with the previous observations for both the SPP and Kuramoto systems, where estimation accuracy deteriorates in regions that are poorly sampled by the trajectories, as measured by the empirical distributionρ^R\\hat\{\\rho\}\_\{R\}\. In regards to the environmental force termff, both methods capture the correct direction of the light source, as well as the overall magnitude of influence \(see Figure[6\(b\)](https://arxiv.org/html/2608.25181#S5.F6.sf2)\)\. Table[8](https://arxiv.org/html/2608.25181#S5.T8)quantifies the estimation of mechanisms, where we note that the semi\-parametric formulation achieves approximately one order of magnitude smaller relative environmental force errorEfrelE\_\{f\}^\{\\mathrm\{rel\}\}than that obtained for the non\-parametric formulation\. This highlights the benefit of incorporating the correct parametric structure into inference when such prior knowledge is available\. Despite the differences observed in feature recovery, both formulations produce highly accurate trajectory reconstructions and predictions \(Figure[7](https://arxiv.org/html/2608.25181#S5.F7)\)\. On both the training and testing data, the reconstruction errors on the fitting interval and the prediction errors on the extrapolation interval remain on the order of10−610^\{\-6\}and10−510^\{\-5\}\. Precise value are provided in Table[9](https://arxiv.org/html/2608.25181#S5.T9)\. \(a\)Estimation of interaction kernelϕ\\phi\(b\)Estimation of environmental forceff Figure 6:\(Phototaxis model\) Comparison of mechanistic feature recovery ability of the fully non\-parametric \(NP\) and semi\-parametric \(SP\) estimation algorithms for the phototaxis model \([82](https://arxiv.org/html/2608.25181#S5.E82)\)\. In Subfigure[6\(a\)](https://arxiv.org/html/2608.25181#S5.F6.sf1)we visualize the true interaction kernelϕ\\phi\(blue\) together with the mean recovered kernelsϕ^¯\\bar\{\\hat\{\\phi\}\}obtained via the NP \(yellow\) and P \(orange\) approaches\. Subfigure[6\(b\)](https://arxiv.org/html/2608.25181#S5.F6.sf2)provides the corresponding true environmental forceffand its recovered mean estimatesf^¯\\bar\{\\hat\{f\}\}\. Asf:ℝ2→ℝ2f:\\mathbb\{R\}^\{2\}\\to\\mathbb\{R\}^\{2\}, the top and bottom rows correspond to first and second components offf, respectively\. Within each row, the true force \(left column\), the fully non\-parametric mean estimate \(middle column\), and semi\-parametric mean estimate \(right column\) are shown, together with one standard deviation, which are obtained overTr=10T\_\{r\}=10independent learning trials\. The background density in Subfigure[6\(a\)](https://arxiv.org/html/2608.25181#S5.F6.sf1)\(right vertical axis\) corresponds to the empirical pairwise distance distributionρ^R\\hat\{\\rho\}\_\{R\}, while the density visualized in Subfigure[6\(b\)](https://arxiv.org/html/2608.25181#S5.F6.sf2)\(left column\) corresponds to the empirical velocity distributionρ^V\\hat\{\\rho\}\_\{V\}induced by the training data\. The corresponding relative feature recovery errors are reported in Table[8](https://arxiv.org/html/2608.25181#S5.T8)\. The training data consists ofM=20M=20trajectories withN=10N=10agents in dimensiond=2d=2, observed over\[0,T\]=\[0,0\.5\]\[0,T\]=\[0,0\.5\]withL=51L=51time points\.MetricNon\-parametricSemi\-ParametricEresrelE^\{\\mathrm\{rel\}\}\_\{\\mathrm\{res\}\}6\.726 628 52×10−056\.726\\,628\\,52\\text\{\\times\}\{10\}^\{\-05\}±\\pm6\.398 876 10×10−066\.398\\,876\\,10\\text\{\\times\}\{10\}^\{\-06\}7\.291 303 63×10−057\.291\\,303\\,63\\text\{\\times\}\{10\}^\{\-05\}±\\pm7\.334 376 82×10−067\.334\\,376\\,82\\text\{\\times\}\{10\}^\{\-06\}EϕrelE^\{\\mathrm\{rel\}\}\_\{\\phi\}1\.985 512 11×10−041\.985\\,512\\,11\\text\{\\times\}\{10\}^\{\-04\}±\\pm2\.292 855 28×10−052\.292\\,855\\,28\\text\{\\times\}\{10\}^\{\-05\}1\.980 978 68×10−041\.980\\,978\\,68\\text\{\\times\}\{10\}^\{\-04\}±\\pm2\.288 845 60×10−052\.288\\,845\\,60\\text\{\\times\}\{10\}^\{\-05\}EfrelE^\{\\mathrm\{rel\}\}\_\{f\}5\.533 440 49×10−055\.533\\,440\\,49\\text\{\\times\}\{10\}^\{\-05\}±\\pm8\.754 940 29×10−068\.754\\,940\\,29\\text\{\\times\}\{10\}^\{\-06\}3\.738 354 94×10−063\.738\\,354\\,94\\text\{\\times\}\{10\}^\{\-06\}±\\pm2\.418 218 40×10−062\.418\\,218\\,40\\text\{\\times\}\{10\}^\{\-06\}Table 8:\(Phototaxis model\) Comparison of relative residual error \([55](https://arxiv.org/html/2608.25181#S3.E55)\) and feature recovery metrics \([52](https://arxiv.org/html/2608.25181#S3.E52)\) for the fully non\-parametric and semi\-parametric estimation algorithms for the phototaxis model \([82](https://arxiv.org/html/2608.25181#S5.E82)\)\. The table reports the mean error±\\pmone standard deviation of these quantities computed overTr=10T\_\{r\}=10independent learning trials\. For additional details on algorithm parameters, we refer the reader to the caption of Figure[6](https://arxiv.org/html/2608.25181#S5.F6)\.\(a\)MtrainM\_\{\\mathrm\{train\}\}training data utilized to obtain estimators \(b\)MtestM\_\{\\mathrm\{test\}\}test data utilized for evaluation Figure 7:\(Phototaxis model\) Comparison of true trajectoriesxx\(left column\) and the estimated trajectoriesx^\\hat\{x\}for both the semi\-parametric \(middle column\) and full non\-parametric \(right column\) methods in the phase plane\. For trajectory errors, see Table[9](https://arxiv.org/html/2608.25181#S5.T9)\. The top row displays trajectories initialized from theMtrainM\_\{\\mathrm\{train\}\}training initial\-conditions, whereas the bottom row displays trajectories generated fromMtestM\_\{\\mathrm\{test\}\}previously unseen testing initial conditions\. Note that trajectories are plotted over the full time horizon\[0,Tf\]\[0,T\_\{f\}\], with the color gradient representing the temporal evolution of the dynamics, from yellow att=0t=0to orange att=Tft=T\_\{f\}\. The black segment indicates the training interval\[0,T\]\[0,T\], which is utilized for model fitting; recall that no fitting is performed in the testing data in Figure[4\(b\)](https://arxiv.org/html/2608.25181#S5.F4.sf2)\. Here\[0,T\]=\[0,0\.5\]\[0,T\]=\[0,0\.5\], with testing evaluated beyond the fitting horizon\[0\.0,0\.5\]\[0\.0,0\.5\]untilTf=2\.0T\_\{f\}=2\.0to assess both reconstruction and prediction in Figure[7\(b\)](https://arxiv.org/html/2608.25181#S5.F7.sf2)\. Both method variants closely reproduce the true dynamics on the training trajectories and generalize well to unseen initial conditions\. For additional details on algorithm parameters, we refer the reader to the caption of Figure[6](https://arxiv.org/html/2608.25181#S5.F6)\.Data setStatisticNon\-parametricSemi\-ParametricMtrainM\_\{\\mathrm\{train\}\}E¯recon\\bar\{E\}\_\{\\mathrm\{recon\}\}4\.027 552 51×10−064\.027\\,552\\,51\\text\{\\times\}\{10\}^\{\-06\}±\\pm6\.549 230 71×10−076\.549\\,230\\,71\\text\{\\times\}\{10\}^\{\-07\}4\.496 428 52×10−064\.496\\,428\\,52\\text\{\\times\}\{10\}^\{\-06\}±\\pm7\.364 670 99×10−077\.364\\,670\\,99\\text\{\\times\}\{10\}^\{\-07\}σErecon\\sigma\_\{E\_\{\\mathrm\{recon\}\}\}1\.009 259 16×10−061\.009\\,259\\,16\\text\{\\times\}\{10\}^\{\-06\}±\\pm3\.055 190 48×10−073\.055\\,190\\,48\\text\{\\times\}\{10\}^\{\-07\}1\.268 869 62×10−061\.268\\,869\\,62\\text\{\\times\}\{10\}^\{\-06\}±\\pm3\.899 608 97×10−073\.899\\,608\\,97\\text\{\\times\}\{10\}^\{\-07\}E¯pred\\bar\{E\}\_\{\\mathrm\{pred\}\}1\.367 294 68×10−051\.367\\,294\\,68\\text\{\\times\}\{10\}^\{\-05\}±\\pm2\.262 020 00×10−062\.262\\,020\\,00\\text\{\\times\}\{10\}^\{\-06\}1\.178 536 84×10−051\.178\\,536\\,84\\text\{\\times\}\{10\}^\{\-05\}±\\pm1\.560 367 45×10−061\.560\\,367\\,45\\text\{\\times\}\{10\}^\{\-06\}σEpred\\sigma\_\{E\_\{\\mathrm\{pred\}\}\}2\.768 822 80×10−062\.768\\,822\\,80\\text\{\\times\}\{10\}^\{\-06\}±\\pm7\.388 936 98×10−077\.388\\,936\\,98\\text\{\\times\}\{10\}^\{\-07\}3\.210 520 55×10−063\.210\\,520\\,55\\text\{\\times\}\{10\}^\{\-06\}±\\pm6\.862 167 86×10−076\.862\\,167\\,86\\text\{\\times\}\{10\}^\{\-07\}MtestM\_\{\\mathrm\{test\}\}E¯recon\\bar\{E\}\_\{\\mathrm\{recon\}\}5\.067 807 84×10−065\.067\\,807\\,84\\text\{\\times\}\{10\}^\{\-06\}±\\pm6\.953 979 44×10−076\.953\\,979\\,44\\text\{\\times\}\{10\}^\{\-07\}4\.523 697 17×10−064\.523\\,697\\,17\\text\{\\times\}\{10\}^\{\-06\}±\\pm5\.193 499 62×10−075\.193\\,499\\,62\\text\{\\times\}\{10\}^\{\-07\}σErecon\\sigma\_\{E\_\{\\mathrm\{recon\}\}\}1\.163 612 56×10−061\.163\\,612\\,56\\text\{\\times\}\{10\}^\{\-06\}±\\pm3\.301 187 82×10−073\.301\\,187\\,82\\text\{\\times\}\{10\}^\{\-07\}1\.012 524 41×10−061\.012\\,524\\,41\\text\{\\times\}\{10\}^\{\-06\}±\\pm3\.038 581 17×10−073\.038\\,581\\,17\\text\{\\times\}\{10\}^\{\-07\}E¯pred\\bar\{E\}\_\{\\mathrm\{pred\}\}1\.543 800 95×10−051\.543\\,800\\,95\\text\{\\times\}\{10\}^\{\-05\}±\\pm2\.754 897 36×10−062\.754\\,897\\,36\\text\{\\times\}\{10\}^\{\-06\}1\.219 029 71×10−051\.219\\,029\\,71\\text\{\\times\}\{10\}^\{\-05\}±\\pm1\.618 797 05×10−061\.618\\,797\\,05\\text\{\\times\}\{10\}^\{\-06\}σEpred\\sigma\_\{E\_\{\\mathrm\{pred\}\}\}2\.643 051 27×10−062\.643\\,051\\,27\\text\{\\times\}\{10\}^\{\-06\}±\\pm6\.207 071 66×10−076\.207\\,071\\,66\\text\{\\times\}\{10\}^\{\-07\}2\.593 189 84×10−062\.593\\,189\\,84\\text\{\\times\}\{10\}^\{\-06\}±\\pm4\.526 170 56×10−074\.526\\,170\\,56\\text\{\\times\}\{10\}^\{\-07\}Table 9:\(Phototaxis model\) Comparison of trajectory\-level errors for the fully non\-parametric and semi\-parametric estimation algorithms for the phototaxis model \([82](https://arxiv.org/html/2608.25181#S5.E82)\)\. For each learning trial, the trajectory statisticsE¯recon\\bar\{E\}\_\{\\mathrm\{recon\}\},E¯pred\\bar\{E\}\_\{\\mathrm\{pred\}\},σErecon\\sigma\_\{E\_\{\\mathrm\{recon\}\}\}, andσEpred\\sigma\_\{E\_\{\\mathrm\{pred\}\}\}are computed according to \([63](https://arxiv.org/html/2608.25181#S3.E63)\) \- \([64](https://arxiv.org/html/2608.25181#S3.E64)\)\. The table reports the mean±\\pmone standard deviation of these quantities computed overTr=10T\_\{r\}=10independent learning trials\. The first section corresponds to trajectories generated fromMtrainM\_\{\\mathrm\{train\}\}training initial conditions, while the second section corresponds toMtestM\_\{\\mathrm\{test\}\}unseen initial conditions sampled fromμ0\\mu\_\{0\}\. Note that across all trajectory measures, error estimates are both small and similar for both methods\. For additional details on algorithm parameters, we refer the reader to the caption of Figure[6](https://arxiv.org/html/2608.25181#S5.F6)\. ### 5\.3Scalability For the remainder of the manuscript, we evaluate the non\-parametric algorithm in several different scenarios, as described in Section[3\.5](https://arxiv.org/html/2608.25181#S3.SS5)\. In practice, prior knowledge of the exact structure of the environmental forceffis rarely available, making the fully non\-parametric formulation more broadly applicable\. The results from Section[5\.2](https://arxiv.org/html/2608.25181#S5.SS2)have established the reliability and consistency of the fully non\-parametric algorithm under repeated trials, so that for the analyses presented in this section, models are trained utilizing one training sample \(i\.e\.Tr=1T\_\{r\}=1\)\. We will later investigate the variability of obtained estimators as a function of the amount of training data in Section[5\.4](https://arxiv.org/html/2608.25181#S5.SS4), where repeated trials are necessary\. To assess whether the learned mechanisms generalize across agent size, we train the algorithm using systems containingNNagents and evaluate the learned models on a testing data set consisting of4N4Nagents of sizeMtest4NM\_\{\\mathrm\{test\}\}^\{4N\}as described in Section[3\.5\.2](https://arxiv.org/html/2608.25181#S3.SS5.SSS2)\(for simplicity we fixMtest4N=MtestN=MtestM\_\{\\mathrm\{test\}\}^\{4N\}=M\_\{\\mathrm\{test\}\}^\{N\}=M\_\{\\mathrm\{test\}\}\)\. Figure[8](https://arxiv.org/html/2608.25181#S5.F8)provides a qualitative comparison of the resulting trajectories for the fully non\-parametric formulation, while Table[10](https://arxiv.org/html/2608.25181#S5.T10)summarizes the corresponding trajectory\-based metrics for both the semi\-parametric and fully non\-parametric formulations for the Kuramoto \(Figure[8\(a\)](https://arxiv.org/html/2608.25181#S5.F8.sf1)\), SPP \(Figure[8\(b\)](https://arxiv.org/html/2608.25181#S5.F8.sf2)\), and phototaxis \(Figure[8\(c\)](https://arxiv.org/html/2608.25181#S5.F8.sf3)\) models\. Across all systems considered, the learned interaction mechanisms generalize consistently to larger populations, and both formulations successfully reproduce the qualitative collective behavior observed in the reference trajectories\. By examining Table[10](https://arxiv.org/html/2608.25181#S5.T10), we observe that the semi\-parametric formulation generally yields slightly smaller mean prediction errors when compared to the fully non\-parametric formulation, but the differences are modest and generally of the same order\. These results thus suggest that the dominant interaction mechanisms are accurately captured by both approaches and transfer well across scales in system size\. Figure[9](https://arxiv.org/html/2608.25181#S5.F9)further demonstrates the ability of the algorithm to predict over a time horizon five times longer than the training interval while simultaneously scaling the number of agents by a factor of ten\. Although the training trajectories contain relatively few agents and do not exhibit the collective milling behavior within the observed training interval, the learned governing mechanisms are sufficiently accurate to reproduce and preserve this emergent behavior when applied to the larger system and evolved into the future\. The predicted agent\-level trajectories may not agree exactly with the reference trajectories on an agent\-by\-agent basis; nevertheless, the learned model preserves the dominant macroscopic collective pattern\. This is a significant strength of the proposed approach, indicating that sufficiently broad sampling of the position and velocity spaces during training enables the algorithm to recover the underlying governing laws, even when the associated emergent phenomenon is not directly observed in the training data\. Data setStatisticNon\-parametricSemi\-parametricKuramotoE¯recon\\bar\{E\}\_\{\\mathrm\{recon\}\}8\.639 226 09×10−048\.639\\,226\\,09\\text\{\\times\}\{10\}^\{\-04\}±\\pm1\.436 067 53×10−041\.436\\,067\\,53\\text\{\\times\}\{10\}^\{\-04\}7\.933 406 60×10−047\.933\\,406\\,60\\text\{\\times\}\{10\}^\{\-04\}±\\pm9\.072 559 42×10−059\.072\\,559\\,42\\text\{\\times\}\{10\}^\{\-05\}σErecon\\sigma\_\{E\_\{\\mathrm\{recon\}\}\}2\.033 986 40×10−042\.033\\,986\\,40\\text\{\\times\}\{10\}^\{\-04\}±\\pm2\.378 269 91×10−052\.378\\,269\\,91\\text\{\\times\}\{10\}^\{\-05\}2\.029 332 07×10−042\.029\\,332\\,07\\text\{\\times\}\{10\}^\{\-04\}±\\pm2\.688 566 76×10−052\.688\\,566\\,76\\text\{\\times\}\{10\}^\{\-05\}E¯pred\\bar\{E\}\_\{\\mathrm\{pred\}\}3\.360 057 73×10−033\.360\\,057\\,73\\text\{\\times\}\{10\}^\{\-03\}±\\pm1\.260 603 89×10−031\.260\\,603\\,89\\text\{\\times\}\{10\}^\{\-03\}2\.721 694 99×10−032\.721\\,694\\,99\\text\{\\times\}\{10\}^\{\-03\}±\\pm9\.891 757 78×10−049\.891\\,757\\,78\\text\{\\times\}\{10\}^\{\-04\}σEpred\\sigma\_\{E\_\{\\mathrm\{pred\}\}\}1\.502 954 18×10−031\.502\\,954\\,18\\text\{\\times\}\{10\}^\{\-03\}±\\pm7\.714 561 34×10−047\.714\\,561\\,34\\text\{\\times\}\{10\}^\{\-04\}1\.372 298 20×10−031\.372\\,298\\,20\\text\{\\times\}\{10\}^\{\-03\}±\\pm6\.667 325 58×10−046\.667\\,325\\,58\\text\{\\times\}\{10\}^\{\-04\}SPPE¯recon\\bar\{E\}\_\{\\mathrm\{recon\}\}7\.886 463 53×10−037\.886\\,463\\,53\\text\{\\times\}\{10\}^\{\-03\}±\\pm2\.475 329 39×10−042\.475\\,329\\,39\\text\{\\times\}\{10\}^\{\-04\}0\.006 884 921 442 267 940\.006\\,884\\,921\\,442\\,267\\,94±\\pm0\.000 253 715 904 779 536 20\.000\\,253\\,715\\,904\\,779\\,536\\,2σErecon\\sigma\_\{E\_\{\\mathrm\{recon\}\}\}1\.737 702 05×10−031\.737\\,702\\,05\\text\{\\times\}\{10\}^\{\-03\}±\\pm4\.804 516 38×10−044\.804\\,516\\,38\\text\{\\times\}\{10\}^\{\-04\}0\.002 113 167 912 756 090 80\.002\\,113\\,167\\,912\\,756\\,090\\,8±\\pm0\.000 543 637 710 861 935 60\.000\\,543\\,637\\,710\\,861\\,935\\,6E¯pred\\bar\{E\}\_\{\\mathrm\{pred\}\}5\.167 463 73×10−025\.167\\,463\\,73\\text\{\\times\}\{10\}^\{\-02\}±\\pm3\.886 650 70×10−033\.886\\,650\\,70\\text\{\\times\}\{10\}^\{\-03\}0\.043 793 787 512 908 120\.043\\,793\\,787\\,512\\,908\\,12±\\pm0\.002 212 161 820 939 125 50\.002\\,212\\,161\\,820\\,939\\,125\\,5σEpred\\sigma\_\{E\_\{\\mathrm\{pred\}\}\}1\.663 482 58×10−021\.663\\,482\\,58\\text\{\\times\}\{10\}^\{\-02\}±\\pm3\.382 845 86×10−033\.382\\,845\\,86\\text\{\\times\}\{10\}^\{\-03\}0\.012 995 345 087 749 8230\.012\\,995\\,345\\,087\\,749\\,823±\\pm0\.002 150 118 775 131 2650\.002\\,150\\,118\\,775\\,131\\,265PhototaxisE¯recon\\bar\{E\}\_\{\\mathrm\{recon\}\}3\.240 223 51×10−063\.240\\,223\\,51\\text\{\\times\}\{10\}^\{\-06\}±\\pm5\.188 340 78×10−075\.188\\,340\\,78\\text\{\\times\}\{10\}^\{\-07\}2\.482 792 67×10−062\.482\\,792\\,67\\text\{\\times\}\{10\}^\{\-06\}±\\pm3\.978 207 70×10−073\.978\\,207\\,70\\text\{\\times\}\{10\}^\{\-07\}σErecon\\sigma\_\{E\_\{\\mathrm\{recon\}\}\}3\.660 011 79×10−073\.660\\,011\\,79\\text\{\\times\}\{10\}^\{\-07\}±\\pm1\.281 899 76×10−071\.281\\,899\\,76\\text\{\\times\}\{10\}^\{\-07\}2\.353 839 58×10−072\.353\\,839\\,58\\text\{\\times\}\{10\}^\{\-07\}±\\pm6\.024 472 96×10−086\.024\\,472\\,96\\text\{\\times\}\{10\}^\{\-08\}E¯pred\\bar\{E\}\_\{\\mathrm\{pred\}\}1\.188 221 88×10−051\.188\\,221\\,88\\text\{\\times\}\{10\}^\{\-05\}±\\pm2\.075 817 43×10−062\.075\\,817\\,43\\text\{\\times\}\{10\}^\{\-06\}7\.507 728 67×10−067\.507\\,728\\,67\\text\{\\times\}\{10\}^\{\-06\}±\\pm6\.925 477 19×10−076\.925\\,477\\,19\\text\{\\times\}\{10\}^\{\-07\}σEpred\\sigma\_\{E\_\{\\mathrm\{pred\}\}\}1\.751 672 58×10−061\.751\\,672\\,58\\text\{\\times\}\{10\}^\{\-06\}±\\pm9\.416 354 18×10−079\.416\\,354\\,18\\text\{\\times\}\{10\}^\{\-07\}1\.671 711 23×10−061\.671\\,711\\,23\\text\{\\times\}\{10\}^\{\-06\}±\\pm1\.223 205 93×10−061\.223\\,205\\,93\\text\{\\times\}\{10\}^\{\-06\}Table 10:\(Scaling number of agents\) Comparison of trajectory level errors for the fully non\-parametric and semi\-parametric estimation algorithms for the Kuramoto \([78](https://arxiv.org/html/2608.25181#S5.E78)\), SPP \([79](https://arxiv.org/html/2608.25181#S5.E79)\), and phototaxis \([82](https://arxiv.org/html/2608.25181#S5.E82)\) models\. For each learning trial, the trajectory statisticsE¯recon\\bar\{E\}\_\{\\mathrm\{recon\}\},E¯pred\\bar\{E\}\_\{\\mathrm\{pred\}\},σErecon\\sigma\_\{E\_\{\\mathrm\{recon\}\}\}, andσEpred\\sigma\_\{E\_\{\\mathrm\{pred\}\}\}are computed according to \([63](https://arxiv.org/html/2608.25181#S3.E63)\) \- \([64](https://arxiv.org/html/2608.25181#S3.E64)\)\. The table reports the mean±\\pmone standard deviation of these quantities computedTr=10T\_\{r\}=10independent learning trials\. Predicted trajectories are made withMtest4NM\_\{\\mathrm\{test\}\}^\{4N\}untrained initial conditions generated for4N4Nagents\. In numerical simulations,Mtest4N=MtestM\_\{\\mathrm\{test\}\}^\{4N\}=M\_\{\\mathrm\{test\}\}\. Further information regarding learning and model parameters may be found in Appendix[E](https://arxiv.org/html/2608.25181#A5)\.\(a\)Kuramoto model trajectories \(b\)SPP model trajectories \(c\)Phototaxis model trajectories Figure 8:\(Scaling number of agents\) The true and predicted trajectories for the Kuramoto \(Figure[8\(a\)](https://arxiv.org/html/2608.25181#S5.F8.sf1)\), SPP \(Figure[8\(b\)](https://arxiv.org/html/2608.25181#S5.F8.sf2)\), and phototaxis \(Figure[8\(c\)](https://arxiv.org/html/2608.25181#S5.F8.sf3)\) systems obtained via the fully non\-parametric method, when obtaining by the estimators by training on a smaller number of agents\. Trajectory predictions are provided onNnew=4NN\_\{\\mathrm\{new\}\}=4Nagents, whenN=10N=10agents were utilized in the inference\. Trajectory errors are reported in Table[10](https://arxiv.org/html/2608.25181#S5.T10)in the semi\-parametric column\. Further information regarding learning and model parameters may be found in Appendix[E](https://arxiv.org/html/2608.25181#A5)\.Figure 9:\(Scaling number of agents: SPP milling model\) The true and predicted trajectories for the SPP milling system obtained using the learned model, with milling parametersCr=1,lr=0\.5,Ca=0\.5,la=2,α=1\.6C\_\{r\}=1,l\_\{r\}=0\.5,C\_\{a\}=0\.5,l\_\{a\}=2,\\alpha=1\.6, andβ=0\.5\\beta=0\.5\. The top row shows the trajectory predictions over the training time interval\[0,T\]=\[0,4\]\[0,T\]=\[0,4\]forN=10N=10agents\. The bottom row shows the trajectory predictions for a scaled system withNnew=100N\_\{\\mathrm\{new\}\}=100agents, evolved over a time horizon five times longer than the training interval, withTf=20T\_\{f\}=20\. Further information regarding the learning parameters, system parameters, and initial conditions may be found in Appendix[E](https://arxiv.org/html/2608.25181#A5)\. ### 5\.4Dependence on training data For the three systems previously considered \(Kuramoto, SPP, and phototaxis\), we are interested in the ability of the fully non\-parametric framework to recover features as either the number of independent trajectory replicatesMMor the number of temporal observationsLLincreases\. This is experimentally relevant, since in applications one would like to determine how many independent replicates are required, and how frequently each trajectory must be sampled in time, to accurately recover the governing mechanistic features\. While one expects accuracy to improve as eitherMMorLLincreases, the relative importance of these two sources of data is not obvious a priori\. We therefore investigate this dependence quantitatively, as described in Section[3\.5\.3](https://arxiv.org/html/2608.25181#S3.SS5.SSS3)\. Figure[10](https://arxiv.org/html/2608.25181#S5.F10)demonstrates that increasing either the number of trajectoriesMMor the temporal resolutionLLleads to more accurate recovery of the underlying interaction mechanisms\. Note that accuracy is measured with respect to the mechanistic feature recovery metrics \([52](https://arxiv.org/html/2608.25181#S3.E52)\), and is averaged overTr=10T\_\{r\}=10learning trials; mean errors are reported with respect to theseTrT\_\{r\}trials\. Interestingly, the results also indicate that increasing the temporal resolutionLLalone cannot fully compensate for a lack of replicate diversityMM\. In particular, when only a single replicate \(M=1M=1\) exists, increasingLLdoes not appear to have a significant effect on the ability to recoverϕ\\phi\(Figure[10\(a\)](https://arxiv.org/html/2608.25181#S5.F10.sf1)\) andff\(Figure[10\(b\)](https://arxiv.org/html/2608.25181#S5.F10.sf2)\), suggesting that a single trajectory does not contain sufficient information to fully identify the interaction mechanisms governing the systems\. Indeed, while increasingLLsamples the same region of state space, increasingMMintroduces new initial conditions and therefore explores a larger portion of state space, which is thus leveraged in the estimation algorithm\. Consequently, trajectory diversity has a critical role in accurately recovering both the interaction kernels and environmental forces, with temporal resolution less crucial with regards to inference\. Recall also that the number of agentsNNis generally large, so that each new replicate introduces a new collection ofNNinitial agent states which are utilized in estimation\. We note that if one is estimating \(for example\) velocitiesVℓ\(m\)V^\{\(m\)\}\_\{\\ell\}from statesXℓ\(m\)X^\{\(m\)\}\_\{\\ell\}as discussed in Section[2\.1\.1](https://arxiv.org/html/2608.25181#S2.SS1.SSS1), increasing the temporal resolutionLLmay improve accuracy in numerically differentiatingXℓ\(m\)X^\{\(m\)\}\_\{\\ell\}\. As we are primarily interested in understanding the errors of the variational inference algorithms, we do not attempt to quantify the accumulation of such errors in this manuscript, where we have assumed exact velocity \(and acceleration, for second–order systems\) data\. One exception to this is presented in Section[5\.5](https://arxiv.org/html/2608.25181#S5.SS5), where we consider the case of reconstruction when utilizing imperfect trajectory data \(see Figure[12](https://arxiv.org/html/2608.25181#S5.F12)\)\. \(a\)Mean error for interaction kernelϕ\\phi \(b\)Mean error for environmental forceff Figure 10:Mean mechanistic relative feature recovery errors \(log scale\) as functions of temporal resolutionLLand independent replicate sizeMMcomputed overTr=10T\_\{r\}=10independent learning trials\. The relatively large errors observed forM=1M=1replicate, even when increasing the number of time samplesLL, suggests that a single trajectory does not contain sufficient information to accurately recover the underlying dynamics\. The above figure, for both interaction kernelϕ\\phi\(Figure[10\(a\)](https://arxiv.org/html/2608.25181#S5.F10.sf1)\) and environmental forceff\(Figure[10\(b\)](https://arxiv.org/html/2608.25181#S5.F10.sf2)\), suggests that the number of replicatesMMis the dominant sampling parameter which controls the accuracy of the estimation algorithm\. Further information regarding learning and model parameters may be found in Appendix[E](https://arxiv.org/html/2608.25181#A5), as well as in the figure captions of Section[5\.2](https://arxiv.org/html/2608.25181#S5.SS2)\. ### 5\.5Robustness to noise in observation data As discussed in Section[3\.5\.4](https://arxiv.org/html/2608.25181#S3.SS5.SSS4), trajectory data arising in applications is typically subject to experimental noise, and we are interested in measuring the ability of the fully non\-parametric variation approach to successfully infer mechanisms in the presence such noise\. We begin by assuming that the observed trajectory, velocity, and acceleration data are independency perturbed at each time sampletℓt\_\{\\ell\}by a uniform random variable of relative sizeζ\\zeta\(see equations \([68](https://arxiv.org/html/2608.25181#S3.E68)\) and \([69](https://arxiv.org/html/2608.25181#S3.E69)\)\); we subsequently varyζ\\zetaand quantify the inference\. Results for relative noise levelsζ∈\[0,0\.1\]\\zeta\\in\[0,0\.1\]are provided in Figure[11](https://arxiv.org/html/2608.25181#S5.F11)for the model of phototaxis \([82](https://arxiv.org/html/2608.25181#S5.E82)\), where we generally observe robustness of the algorithm with respect to observation noise\. Figures[11\(a\)](https://arxiv.org/html/2608.25181#S5.F11.sf1)and[11\(d\)](https://arxiv.org/html/2608.25181#S5.F11.sf4)illustrates feature recovery forffandϕ\\phi, respectively, while trajectory predictions for noise levels ofζ=0\.01\\zeta=0\.01\(Figure[11\(b\)](https://arxiv.org/html/2608.25181#S5.F11.sf2)\) andζ=0\.1\\zeta=0\.1\(Figure[11\(c\)](https://arxiv.org/html/2608.25181#S5.F11.sf3)\) are also provided\. We observe that as the noise levelζ\\zetaincreases, the learned interaction kernel and environmental force gradually deviate from their true values, with the largest deviation occurring primarily in regions where the empirical sampling measuresρ^R\\hat\{\\rho\}\_\{R\}andρ^V\\hat\{\\rho\}\_\{V\}are small, i\.e\. in regions where the available trajectory information is limited\. In contrast, regions that are well supported by the observational data remain largely unaffected by the perturbations\. Similar to previous numerical experiments, training is performed on the time interval\[0,T\]\[0,T\], while testing \(prediction\) is evaluated on\(T,Tf\]\(T,T\_\{f\}\]\.\. Note that even in the presence of10%10\\%relative noise levels \(Figure[11\(c\)](https://arxiv.org/html/2608.25181#S5.F11.sf3)\), the variational algorithm is still able to accurately predict trajectories well outside of the training interval\. Similar behavior is observed across all systems considered in Section[5\.2](https://arxiv.org/html/2608.25181#S5.SS2); the corresponding results are reported in Appendix[C\.2](https://arxiv.org/html/2608.25181#A3.SS2)in Figures[23](https://arxiv.org/html/2608.25181#A3.F23)and[24](https://arxiv.org/html/2608.25181#A3.F24)\. \(a\)Interaction kernelϕ\\phi\(b\)Trajectory data \(ζ=0\.01\\zeta=0\.01\) \(c\)Trajectory data \(ζ=0\.1\\zeta=0\.1\) \(d\)Environmental forceff Figure 11:\(Noise robustness, phototaxis model\) Mechanistic feature recovery under multiplicative observational noise with noise levelsζ\\zeta\. Subfigure[11\(a\)](https://arxiv.org/html/2608.25181#S5.F11.sf1)plots the estimated interaction kernelsϕ\\phifor different noise levels, while Subfigure[11\(d\)](https://arxiv.org/html/2608.25181#S5.F11.sf4)plots the corresponding estimated environmental forcesff\. The estimated empirical measuresρ^R\\hat\{\\rho\}\_\{R\}andρ^V\\hat\{\\rho\}\_\{V\}induced by the training data are also included in each subfigure\. Subfigures[11\(b\)](https://arxiv.org/html/2608.25181#S5.F11.sf2)and[11\(c\)](https://arxiv.org/html/2608.25181#S5.F11.sf3)compare the true, noisy, and learned trajectories forζ=0\.01\\zeta=0\.01andζ=0\.1\\zeta=0\.1, respectively\. True \(blue\), noisy \(grey\), and estimated \(yellow to orange, varying by time\) trajectory data are all plotted; recall that training occurs on the noisy \(grey\) data\. Training occurs on\[0,T\]\[0,T\], and testing \(predictions\) are provided on\(T,Tf\]\(T,T\_\{f\}\]\. We observe that asζ\\zetaincreases, feature recovery gradually deteriorates, primarily in regions with limited trajectory data \(small measures\), while the learned trajectories remain close to the true trajectories from low to moderate noise levels\. Further information regarding learning and model parameters may be found in Appendix[E](https://arxiv.org/html/2608.25181#A5), as well as in the figure captions of Section[5\.2](https://arxiv.org/html/2608.25181#S5.SS2)\.We also provide results for an example where we estimate velocities from noisy trajectory data for the Kuramoto model \([78](https://arxiv.org/html/2608.25181#S5.E78)\), which are provided in Figure[12](https://arxiv.org/html/2608.25181#S5.F12)\. Figure[12\(a\)](https://arxiv.org/html/2608.25181#S5.F12.sf1)provides a sample trajectory \(underlying and noised\), together with both sliding window and Kalman filters, as discussed in Section[3\.5\.4](https://arxiv.org/html/2608.25181#S3.SS5.SSS4)\. In Figure[12\(b\)](https://arxiv.org/html/2608.25181#S5.F12.sf2)we plot the corresponding numerically differentiated velocities for the same representative trajectory, which are utilized as observation dataVℓ\(m\)V^\{\(m\)\}\_\{\\ell\}in the learning algorithm; results are provided in Figures[12\(c\)](https://arxiv.org/html/2608.25181#S5.F12.sf3)\(interaction kernel\),[12\(d\)](https://arxiv.org/html/2608.25181#S5.F12.sf4)\(environmental force\), and[12\(e\)](https://arxiv.org/html/2608.25181#S5.F12.sf5)\(estimated trajectories\)\. The estimation provided here is included simply as a “proof of concept" to demonstrate that mechanism recovery is possible even when estimating velocity from noisy trajectory data, and a comprehensive investigation of such techniques is beyond the scope of this work\. We do note that the optimal choice of numerical differentiation method depends on the characteristics of the underlying dynamics and the level of observational noise, and that for the Kuramoto system presented, the Kalman filter with the Rauch\-Tung\-Striebel \(RTS\) smoother appears to provide the most accurate feature recovery and trajectory reconstruction among the methods considered\. As expected, forward differencing on noisy data yields highly inaccurate results, especially in regards to estimation of the environmental force and trajectory prediction\. \(a\)Trajectoryθ\\theta\(representative agent\)\(b\)Velocityθ˙\\dot\{\\theta\}\(representative agent\)\(c\)Interaction kernelϕ\\phi\(d\)Environmental forceff\(e\)Trajectory data Figure 12:\(Velocity estimation with noisy data, Kuramoto model\) Comparison of filtering methods for trajectory data with multiplicative observational noise in the Kuramoto model \([78](https://arxiv.org/html/2608.25181#S5.E78)\)\. Here the noise level is fixed atζ=0\.01\\zeta=0\.01\(see \([68](https://arxiv.org/html/2608.25181#S3.E68)\) and \([69](https://arxiv.org/html/2608.25181#S3.E69)\)\)\. Subfigures[12\(a\)](https://arxiv.org/html/2608.25181#S5.F12.sf1)and[12\(b\)](https://arxiv.org/html/2608.25181#S5.F12.sf2)provide plots for the noisy observations of stateθ\\thetaand their derivativesθ˙\\dot\{\\theta\}, respectively, for a representative trajectory, with true trajectories \(black\), noisy observations \(blue\), a sliding window \(SW\) filter \(orange\), and a Kalman filter with the Rauch–Tung–Striebel \(KF\-RTS\) smoother \(green\) included\. Subfigure[12\(c\)](https://arxiv.org/html/2608.25181#S5.F12.sf3)plots the estimated interaction kernelsϕ\\phi, while Subfigure[12\(d\)](https://arxiv.org/html/2608.25181#S5.F12.sf4)plot the recovered environmental forcesffobtained from the true, noisy, and filtered observations\. The estimated empirical measuresρ^Δθ\\hat\{\\rho\}\_\{\\Delta\\theta\}andρ^θ\\hat\{\\rho\}\_\{\\theta\}induced by the training data are also included in each subfigure\. Subfigure[12\(e\)](https://arxiv.org/html/2608.25181#S5.F12.sf5)compares the true, noisy, smoothed, and learned trajectories for a representative training trajectory \(i\.e\. a single agent\)\. Further information regarding learning and model parameters may be found in Appendix[E](https://arxiv.org/html/2608.25181#A5), as well as in the caption of Figure[1](https://arxiv.org/html/2608.25181#S5.F1)\. ### 5\.6Model selection In this section, we demonstrate the ability of the proposed model selection framework introduced in Section[4](https://arxiv.org/html/2608.25181#S4)to identify the mechanisms governing a collective system directly from observational data\. The objective is not only to estimate the unknown interaction features but also to determine which candidate framework best explains the observed trajectory data\. We consider three second\-order systems of various complexity: 1\) the pure Cucker\-Smale model \(Section[5\.6\.1](https://arxiv.org/html/2608.25181#S5.SS6.SSS1)\), which describes velocity alignment, 2\) the previously introduced phototaxis model \(Section[5\.6\.2](https://arxiv.org/html/2608.25181#S5.SS6.SSS2)\), which combines alignment with environmental forcing, and 3\) a self\-propelled particle model with Cucker\-Smale dynamics \(Section[5\.6\.3](https://arxiv.org/html/2608.25181#S5.SS6.SSS3)\), which combines features from both the SPP and Cucker\-Smale systems\. We also consider a first\-order model of opinion dynamics \(Section[5\.6\.4](https://arxiv.org/html/2608.25181#S5.SS6.SSS4)\)\. Taken together, these examples allow us to assess the ability of the framework to distinguish among competing model structures and correctly identify the most likely “active" mechanisms given the observed data\. #### 5\.6\.1Cucker\-Smale model We first consider flocking, one of the simplest forms of emergence in collective dynamics\. When flocking, agents asymptotically align their velocities\. The mathematical study of flocking has attracted considerable attention; see for example\[[10](https://arxiv.org/html/2608.25181#bib.bib11),[11](https://arxiv.org/html/2608.25181#bib.bib26),[12](https://arxiv.org/html/2608.25181#bib.bib27),[13](https://arxiv.org/html/2608.25181#bib.bib28),[14](https://arxiv.org/html/2608.25181#bib.bib12)\]and the references therein\. As a representative flocking system, we consider the Cucker\-Smale model\[[10](https://arxiv.org/html/2608.25181#bib.bib11),[95](https://arxiv.org/html/2608.25181#bib.bib29)\], which is a second\-order interacting particle system describing self\-organization driven entirely by a phenomenological interaction velocity alignment kernel\. The Cucker\-Smale model fits naturally within the general second\-order framework \([72](https://arxiv.org/html/2608.25181#S4.E72)\) and is therefore an ideal benchmark for testing whether the model selection procedure can correctly identify a system whose dynamics are governed solely by alignment interactions, as there are no environmental forces in the simplest form of the model\. Indeed, the mechanisms defining the Cucker\-Smale model are given by f\(x,v\)≡0ϕE\(r\)≡0ϕA\(r\)=γ\(1\+r2\)β\\displaystyle\\begin\{split\}f\(x,v\)&\\equiv 0\\\\ \\phi^\{E\}\(r\)&\\equiv 0\\\\ \\phi^\{A\}\(r\)&=\\frac\{\\gamma\}\{\(1\+r^\{2\}\)^\{\\beta\}\}\\end\{split\}\(83\)whereγ,β\\gamma,\\betaare positive constants\. Depending on the choice ofβ\\beta, the Cucker\-Smale system \([83](https://arxiv.org/html/2608.25181#S5.E83)\) exhibits different flocking regimes\[[10](https://arxiv.org/html/2608.25181#bib.bib11)\]\. For example, whenβ<0\.5\\beta<0\.5, flocking occurs for all initial conditions \(unconditional flocking\), while forβ=0\.5\\beta=0\.5, flocking is dependent on the initial velocity configuration, and forβ\>0\.5\\beta\>0\.5, flocking depends on both the initial positions and velocities of the agents\. For numerical experiments reported in this section, we selectγ=1\.0\\gamma=1\.0andβ=0\.1\\beta=0\.1, so that the dynamics exhibit unconditional flocking\. As summarized in Table[11](https://arxiv.org/html/2608.25181#S5.T11), the Cucker Smale model contains only alignment interactions and therefore provides a test case for evaluating whether the model selection procedure can correctly identify a pure alignment system\. The remaining simulation and learning parameters used in this experiment are reported in Tables[E\.2](https://arxiv.org/html/2608.25181#A5.SS2)and[E\.1](https://arxiv.org/html/2608.25181#A5.SS1)of Appendix[E](https://arxiv.org/html/2608.25181#A5)\. We estimate all three mechanisms \(ff,ϕE\\phi^\{E\}, andϕA\\phi^\{A\}\) in the most general second\-order framework \([72](https://arxiv.org/html/2608.25181#S4.E72)\); results are provided in Figure[13](https://arxiv.org/html/2608.25181#S5.F13)\. We observe that the fully non\-parametric variational algorithm exhibits negligible training variability across repeated trials and correctly identifies the dominant alignment interactionϕA\\phi^\{A\}while simultaneously obtaining near zero estimates for inactive energyϕE\\phi^\{E\}and environmental forceffterms\. Indeed, when measuring the weightedL2L^\{2\}norms \([49](https://arxiv.org/html/2608.25181#S3.E49)\), we obtain∥f^∥L2\(ρ^R\),∥ϕ^E∥L2\(ρ^V\)=O\(10−4\)\\lVert\\hat\{f\}\\rVert\_\{L^\{2\}\(\\hat\{\\rho\}\_\{R\}\)\},\\lVert\\hat\{\\phi\}^\{E\}\\rVert\_\{L^\{2\}\(\\hat\{\\rho\}\_\{V\}\)\}=O\(10^\{\-4\}\), while∥ϕ^A∥L2\(ρ^R\)=O\(10−1\)\\lVert\\hat\{\\phi\}^\{A\}\\rVert\_\{L^\{2\}\(\\hat\{\\rho\}\_\{R\}\)\}=O\(10^\{\-1\}\)\. We also provide quantitative measures of the model selection framework in Table[12](https://arxiv.org/html/2608.25181#S5.T12)for the Cucker\-Smale system\. CandidatesF2−S1F\_\{2\}\-S\_\{1\}are rejected by the framework compatibility gate as their relative residual errors satisfyEresrelϵ\>0\.25E\_\{\\mathrm\{res\}\}^\{\\mathrm\{rel\_\{\\epsilon\}\}\}\>0\.25\. The remaining candidatesS2−S6S\_\{2\}\-S\_\{6\}are thus admissible and are retained for ranking\. All admissible candidates satisfy the primary criterion by achieving \(within toleranceτtraj\\tau\_\{\\mathrm\{traj\}\}, as defined in \([77](https://arxiv.org/html/2608.25181#S4.E77)\)\) the same minimum validation trajectory errorEminE\_\{\\min\}\. We therefore move to the secondary criterion of framework complexity\. The minimum complexity is achieved by candidateS2S\_\{2\}, corresponding to the second\-order alignment\-only framework, which is precisely the true underlying model \([83](https://arxiv.org/html/2608.25181#S5.E83)\) which generated the data\. Thus, the framework selection algorithm correctly identified the pure Cucker\-Smale alignment model\. OrderϕE\\phi^\{E\}ϕA\\phi^\{A\}ff2\-✓\-Table 11:Mechanisms of the Cucker\-Smale system \([83](https://arxiv.org/html/2608.25181#S5.E83)\)\. The symbol “\-" indicates that the mechanism is not present, while “✓" indicates that the mechanism is present\.\(a\)Energy\-based interaction kernelϕE\\phi^\{E\}\(b\)Alignment kernelϕA\\phi^\{A\} \(c\)Environmental forceff Figure 13:\(Model selection, Cucker\-Smale\) Mechanisms present in the Cucker\-Smale system \([83](https://arxiv.org/html/2608.25181#S5.E83)\) within the generalized framework \([72](https://arxiv.org/html/2608.25181#S4.E72)\)\. The model selection framework estimates three quantities: an energy\-based interaction kernelϕ^E\\hat\{\\phi\}^\{E\}, an alignment interaction kernelϕ^A\\hat\{\\phi\}^\{A\}, and an environmental forcef^\\hat\{f\}\. From observations of trajectories, the algorithm correctly identify the system as alignment\-dominated\. Estimation is performed overTr=10T\_\{r\}=10independent learning trials, and the inferred alignment kernel \(Figure[13\(b\)](https://arxiv.org/html/2608.25181#S5.F13.sf2)\) has weightedL2L^\{2\}norm‖ϕ^A‖L2\(ρ^R\)=\\\|\\hat\{\\phi\}^\{A\}\\\|\_\{L^\{2\}\(\\hat\{\\rho\}\_\{R\}\)\}=0\.742 1530\.742\\,153±\\pm0\.003 4680\.003\\,468, which is several orders of magnitude larger than the inferred energy kernel‖ϕ^E‖L2\(ρ^R\)=\\\|\\hat\{\\phi\}^\{E\}\\\|\_\{L^\{2\}\(\\hat\{\\rho\}\_\{R\}\)\}=1\.9985×10−041\.9985\\text\{\\times\}\{10\}^\{\-04\}±\\pm1\.0221×10−041\.0221\\text\{\\times\}\{10\}^\{\-04\}\(Figure[13\(a\)](https://arxiv.org/html/2608.25181#S5.F13.sf1)\) and environmental force‖f^‖L2\(ρ^V\)=\\\|\\hat\{f\}\\\|\_\{L^\{2\}\(\\hat\{\\rho\}\_\{V\}\)\}=0\.000 5280\.000\\,528±\\pm0\.000 0700\.000\\,070\(Figure[13\(c\)](https://arxiv.org/html/2608.25181#S5.F13.sf3)\)\. Note the scale of the vertical axes\. Further information regarding learning and model parameters may be found in Appendix[E](https://arxiv.org/html/2608.25181#A5)\.CandidateOrderEresrelϵE\_\{\\mathrm\{res\}\}^\{\\mathrm\{rel\}\_\{\\epsilon\}\}EreconrelE\_\{\\mathrm\{recon\}\}^\{\\mathrm\{rel\}\}E¯traj\(Mval\)\\bar\{E\}\_\{\\mathrm\{traj\}\}\(M\_\{\\mathrm\{val\}\}\)ComplexityEresE\_\{\\mathrm\{res\}\}RankSelectedS2S\_\{2\}20\.0006260\.0000040\.00025930\.0005091YesS4S\_\{4\}20\.0007080\.0000040\.00043740\.0005752NoS5S\_\{5\}20\.0006240\.0000040\.00025550\.0005073NoS6S\_\{6\}20\.0007170\.0000040\.00046460\.0005834NoF2F\_\{2\}10\.5937260\.0032660\.73206120\.4824665NoS3S\_\{3\}20\.5937260\.0032660\.73206140\.4824666NoF1F\_\{1\}10\.9950700\.0129911\.58364110\.8086017NoS1S\_\{1\}20\.9950700\.0129911\.58364130\.8086018NoTable 12:\(Model selection, Cucker\-Smale\) Summary of the candidate models considered during model selection\. For each candidate, we report the system order, framework compatibility diagnostics, validation trajectory error, model complexity, residual error, ranking, and whether the candidate was ultimately selected based on the observed trajectory data\. In this case, the framework correctly identifies the Cucker\-Smale flocking model \([83](https://arxiv.org/html/2608.25181#S5.E83)\)\. #### 5\.6\.2Phototaxis model We perform model selection on the phototaxis system introduced in Section[5\.2\.3](https://arxiv.org/html/2608.25181#S5.SS2.SSS3), with governing equations \([82](https://arxiv.org/html/2608.25181#S5.E82)\) with respect to framework \([72](https://arxiv.org/html/2608.25181#S4.E72)\)\. Phototaxis thus combines alignment with environmental forcing, the latter of which is induced by an external light source\. Therefore, this model provides a natural test case for determining whether model selection can simultaneously identify collective alignment and environmental forcing while rejecting inactive energetic interactions; these mechanisms are summarized in Table[13](https://arxiv.org/html/2608.25181#S5.T13)\. That is, the selection should identify mechanismsϕA\\phi^\{A\}andffwhile rejectingϕE\\phi^\{E\}\. The simulation and learning parameters utilized in this numerical experiment are reported in Tables[E\.2](https://arxiv.org/html/2608.25181#A5.SS2)and[E\.1](https://arxiv.org/html/2608.25181#A5.SS1), with the corresponding values ofϕA\\phi^\{A\}andffspecified in Section[5\.2\.3](https://arxiv.org/html/2608.25181#S5.SS2.SSS3)\. OrderϕE\\phi^\{E\}ϕA\\phi^\{A\}ff2\-✓✓Table 13:Mechanisms of the phototaxis system \([82](https://arxiv.org/html/2608.25181#S5.E82)\)\. The symbol “\-" indicates that the mechanism is not present, while “✓" indicates that the mechanism is present\.We estimate all three mechanisms \(ff,ϕE\\phi^\{E\}, andϕA\\phi^\{A\}\) in the most general second\-order framework \([72](https://arxiv.org/html/2608.25181#S4.E72)\); results are provided in Figure[14](https://arxiv.org/html/2608.25181#S5.F14)\. As in Section[5\.6\.1](https://arxiv.org/html/2608.25181#S5.SS6.SSS1), the model selection algorithm correctly identifies the alignment interactionϕA\\phi^\{A\}and environmental forceffas active features, while obtaining near zero values for the energy interaction kernelϕE\\phi^\{E\}\. Indeed, when measuring the weightedL2L^\{2\}norms \([49](https://arxiv.org/html/2608.25181#S3.E49)\), we obtain∥ϕ^E∥L2\(ρ^R\)=O\(10−4\)\\lVert\\hat\{\\phi\}^\{E\}\\rVert\_\{L^\{2\}\(\\hat\{\\rho\}\_\{R\}\)\}=O\(10^\{\-4\}\), while∥ϕ^A∥L2\(ρ^R\),∥f^∥L2\(ρ^V\),=O\(10−1\)\\lVert\\hat\{\\phi\}^\{A\}\\rVert\_\{L^\{2\}\(\\hat\{\\rho\}\_\{R\}\)\},\\lVert\\hat\{f\}\\rVert\_\{L^\{2\}\(\\hat\{\\rho\}\_\{V\}\)\},=O\(10^\{\-1\}\)\. We also provide quantitative measures on the model selection framework in Table[14](https://arxiv.org/html/2608.25181#S5.T14)\. CandidatesF1−S1F\_\{1\}\-S\_\{1\}are rejected by the framework compatibility gate as their relative residual errors satisfyEresrelϵ\>0\.25E\_\{\\mathrm\{res\}\}^\{\\mathrm\{rel\_\{\\epsilon\}\}\}\>0\.25, while the remaining candidatesS4−S5S\_\{4\}\-S\_\{5\}are thus admissible and are retained for ranking\. Among these admissible candidates, the minimum validation trajectory errorEminE\_\{\\min\}is achieved byS4−S6S\_\{4\}\-S\_\{6\}candidates, resulting in a tie according to the criterion defined in \([77](https://arxiv.org/html/2608.25181#S4.E77)\)\. The selection procedure therefore proceeds to the secondary metric, i\.e\. framework complexity\. The minimum complexity is achieved by candidateS4S\_\{4\}, corresponding to the second\-order alignment\-plus\-environmental force framework, which coincides with the true data\-generation model, demonstrating that the proposed model selection procedure can correctly identify both active mechanisms in the phototaxis model\. \(a\)Energy\-based interaction kernelϕE\\phi^\{E\}\(b\)Alignment interaction kernelϕA\\phi^\{A\} \(c\)Environmental forceff Figure 14:\(Model selection, phototaxis\) Mechanisms present in the phototaxis system \([82](https://arxiv.org/html/2608.25181#S5.E82)\) within the generalized framework \([72](https://arxiv.org/html/2608.25181#S4.E72)\)\. The model selection framework estimates three quantities: an energy\-based interaction kernelϕ^E\\hat\{\\phi\}^\{E\}, an alignment interaction kernelϕ^A\\hat\{\\phi\}^\{A\}, and an environmental forcef^\\hat\{f\}\. From observation of trajectories, the algorithm correctly identify the system as being governed by alignment and environmental force interactions\. Estimation is performed overTr=10T\_\{r\}=10independent learning trials, and the inferred learned alignment kernel \(Figure[14\(b\)](https://arxiv.org/html/2608.25181#S5.F14.sf2)\) has weightedL2L^\{2\}norm‖ϕ^A‖L2\(ρ^R\)=\\\|\\hat\{\\phi\}^\{A\}\\\|\_\{L^\{2\}\(\\hat\{\\rho\}\_\{R\}\)\}=0\.931 0350\.931\\,035±\\pm0\.002 0160\.002\\,016and the estimated environmental force has weightedL2L^\{2\}norm‖f^‖L2\(ρ^V\)=\\\|\\hat\{f\}\\\|\_\{L^\{2\}\(\\hat\{\\rho\}\_\{V\}\)\}=0\.600 5040\.600\\,504±\\pm0\.006 2110\.006\\,211\(Figure[14\(c\)](https://arxiv.org/html/2608.25181#S5.F14.sf3)\), both of which are several orders of magnitude larger than the inferred energy kernel‖ϕ^E‖L2\(ρ^R\)=\\\|\\hat\{\\phi\}^\{E\}\\\|\_\{L^\{2\}\(\\hat\{\\rho\}\_\{R\}\)\}=7\.7845×10−047\.7845\\text\{\\times\}\{10\}^\{\-04\}±\\pm6\.9402×10−046\.9402\\text\{\\times\}\{10\}^\{\-04\}\(Figure[14\(a\)](https://arxiv.org/html/2608.25181#S5.F14.sf1)\)\. Note the scale of the vertical axes\. These results indicate that the model\-selection framework correctly identifies the active mechanisms while suppressing the inactive energy interaction\. Further information regarding learning and model parameters may be found in Appendix[E](https://arxiv.org/html/2608.25181#A5)\.CandidateOrderEresrelϵE\_\{\\mathrm\{res\}\}^\{\\mathrm\{rel\}\_\{\\epsilon\}\}EreconrelE\_\{\\mathrm\{recon\}\}^\{\\mathrm\{rel\}\}E¯traj\(Mval\)\\bar\{E\}\_\{\\mathrm\{traj\}\}\(M\_\{\\mathrm\{val\}\}\)ComplexityEresE\_\{\\mathrm\{res\}\}RankSelectedS4S\_\{4\}20\.0000670\.0000020\.00001640\.0000401YesS6S\_\{6\}20\.0000680\.0000020\.00001760\.0000412NoF2F\_\{2\}10\.2164970\.0090560\.10212420\.1303413NoS3S\_\{3\}20\.2164970\.0090560\.10212440\.1303414NoS2S\_\{2\}20\.2241090\.0156540\.27298130\.1349245NoS5S\_\{5\}20\.2241090\.0156540\.27298150\.1349246NoF1F\_\{1\}11\.0413320\.0886410\.95598010\.6269317NoS1S\_\{1\}21\.0413320\.0886410\.95598030\.6269318NoTable 14:\(Model selection, phototaxis\) Summary of the candidate models considered during model selection\. For each candidate, we report the system order, framework compatibility diagnostics, validation trajectory error, model complexity, residual error, ranking, and whether the candidate was ultimately selected based on the observed trajectory data\. In this case, the framework correctly identifies the phototaxis model \([82](https://arxiv.org/html/2608.25181#S5.E82)\)\. #### 5\.6\.3Self\-propelled particle model with velocity alignment \(SPPCS\) We next consider an extension of the self\-propelled particle \(SPP\) model introduced in Section[5\.2\.2](https://arxiv.org/html/2608.25181#S5.SS2.SSS2)that incorporates a Cucker\-Smale\-type velocity alignment interaction; for notational convenience, we abbreviate this model as SPPCS\. In contrast to the preceding examples, this system includes two interaction kernels, corresponding to attraction\-repulsion and velocity alignment, together with an environmental/self\-propulsion force\. It therefore provides a useful benchmark for assessing whether the proposed framework can distinguish and recover multiple mechanistic force components from trajectory data\. The precise environmental force and energy kernel are thus given by \([81](https://arxiv.org/html/2608.25181#S5.E81)\), while the alignment kernelϕA\\phi^\{A\}takes the form as in \([83](https://arxiv.org/html/2608.25181#S5.E83)\), and the ODE system is given as in the general second\-order framework \([72](https://arxiv.org/html/2608.25181#S4.E72)\)\. Note that this systems contains all three mechanisms described in the generalized framework: an environmental force, an alignment\-based interaction, and an energy\-based interaction\. Thus, this model provides a more challenging selection benchmark example considered, as the correct framework requires all mechanisms to be identified simultaneously\. For trajectory generation, we utilize parameter values as in Sections[5\.6\.1](https://arxiv.org/html/2608.25181#S5.SS6.SSS1)and[5\.2\.2](https://arxiv.org/html/2608.25181#S5.SS2.SSS2), and the remaining simulation and learning parameters are reported in Tables[E\.2](https://arxiv.org/html/2608.25181#A5.SS2)and[E\.1](https://arxiv.org/html/2608.25181#A5.SS1)\. The mechanisms of the SPPCS are summarized in Table[15](https://arxiv.org/html/2608.25181#S5.T15), and the selection algorithm should identify non\-zero mechanismsϕE\\phi^\{E\},ϕA\\phi^\{A\}, andff\. OrderϕE\\phi^\{E\}ϕA\\phi^\{A\}ff2✓✓✓Table 15:Mechanisms of the SPPCS system\. The symbol “\-" indicates that the mechanism is not present, while “✓" indicates that the mechanism is present\.We estimate all three mechanisms \(ff,ϕE\\phi^\{E\}, andϕA\\phi^\{A\}\) in the most general second\-order framework \([72](https://arxiv.org/html/2608.25181#S4.E72)\); results are provided in Figure[15](https://arxiv.org/html/2608.25181#S5.F15)\. The algorithm correctly identifies all three model mechanisms, with all three weightedL2L^\{2\}norms \([49](https://arxiv.org/html/2608.25181#S3.E49)\) of the same order \(O\(10−1\)O\(10^\{\-1\}\)\)\. We also provide quantitative measures on the model selection framework in Table[16](https://arxiv.org/html/2608.25181#S5.T16)\. We observe that candidatesS2−S1S\_\{2\}\-S\_\{1\}are rejected by the framework compatibility gate as their relative residual errors satisfyEresrelϵ\>0\.25E\_\{\\mathrm\{res\}\}^\{\\mathrm\{rel\_\{\\epsilon\}\}\}\>0\.25\. In addition, candidatesS2S\_\{2\}andS5S\_\{5\}fail to produce vector fields that can be numerically integrated, and as candidate frameworks must generate physically meaningful trajectories before compatibility metrics can be evaluated, these candidates are rejected prior to the framework compatibility stage and removed from further consideration\. The remaining candidatesS3−S6S\_\{3\}\-S\_\{6\}satisfy the admissibility criteria and are retained for ranking\. Among these admissible candidates, the minimum validation trajectory errorEminE\_\{\\min\}is achieved uniquely by candidateS6S\_\{6\}\(within toleranceτtraj\\tau\_\{\\mathrm\{traj\}\}in \([76](https://arxiv.org/html/2608.25181#S4.E76)\)\), corresponding to the second\-order energy\-alignment\-environmental force framework, which is precisely the true generating model\. This thus demonstrates that the proposed model selection procedure can correctly identify all active mechanisms when they are simultaneously present in the observation trajectory data\. \(a\)Energy\-based interaction kernelϕE\\phi^\{E\}\(b\)Alignment kernelϕA\\phi^\{A\} \(c\)Environmental forceff Figure 15:\(Model selection, SPPCS\) Mechanisms present in the SPPCS system within the generalized framework \([72](https://arxiv.org/html/2608.25181#S4.E72)\)\. The model selection framework estimates three quantities: an energy\-based interaction kernelϕ^E\\hat\{\\phi\}^\{E\}, an alignment interaction kernelϕ^A\\hat\{\\phi\}^\{A\}, and an environmental forcef^\\hat\{f\}\. From observations of trajectories, the algorithm correctly identify the system as being governed by all three mechanisms\. Estimation is performed overTr=10T\_\{r\}=10independent learning trials, and the inferred alignment kernel has weightedL2L^\{2\}norm‖ϕ^A‖L2\(ρ^R\)=\\\|\\hat\{\\phi\}^\{A\}\\\|\_\{L^\{2\}\(\\hat\{\\rho\}\_\{R\}\)\}=0\.344 5570\.344\\,557±\\pm0\.031 9420\.031\\,942\(Figure[15\(b\)](https://arxiv.org/html/2608.25181#S5.F15.sf2)\), the inferred environmental force has weightedL2L^\{2\}norm‖f^‖L2\(ρ^V\)=\\\|\\hat\{f\}\\\|\_\{L^\{2\}\(\\hat\{\\rho\}\_\{V\}\)\}=3\.062 9653\.062\\,965±\\pm0\.047 2180\.047\\,218\(Figure[15\(c\)](https://arxiv.org/html/2608.25181#S5.F15.sf3)\), and the inferred energy kernel has weightedL2L^\{2\}norm‖ϕ^E‖L2\(ρR\)=\\\|\\hat\{\\phi\}^\{E\}\\\|\_\{L^\{2\}\(\\rho\_\{R\}\)\}=6\.5314×10−016\.5314\\text\{\\times\}\{10\}^\{\-01\}±\\pm6\.0227×10−026\.0227\\text\{\\times\}\{10\}^\{\-02\}\(Figure[15\(a\)](https://arxiv.org/html/2608.25181#S5.F15.sf1)\)\. Note the scale of the vertical axes\. Further information regarding learning and model parameters may be found in Appendix[E](https://arxiv.org/html/2608.25181#A5)\.CandidateOrderEresrelϵE\_\{\\mathrm\{res\}\}^\{\\mathrm\{rel\_\{\\epsilon\}\}\}EreconrelE\_\{\\mathrm\{recon\}\}^\{\\mathrm\{rel\}\}E¯traj\(Mval\)\\bar\{E\}\_\{\\mathrm\{traj\}\}\(M\_\{\\mathrm\{val\}\}\)ComplexityEresE\_\{\\mathrm\{res\}\}RankSelectedS6S\_\{6\}20\.0631760\.0006080\.07223160\.1052551YesS4S\_\{4\}20\.1273280\.0034800\.35443040\.2121382NoF2F\_\{2\}10\.2370080\.0038000\.39650520\.3948733NoS3S\_\{3\}20\.2370080\.0038000\.39650540\.3948734NoS2S\_\{2\}2N/AN/AN/A3N/A5NoS5S\_\{5\}2N/AN/AN/A5N/A6NoF1F\_\{1\}10\.9904630\.0805562\.17668611\.6501837NoS1S\_\{1\}20\.9904630\.0805562\.17668631\.6501838NoTable 16:\(Model selection, SPPCS\) Summary of the candidate models considered during model selection\. For each candidate, we report the system order, framework compatibility diagnostics, validation trajectory error, model complexity, residual error, ranking, and whether the candidate was ultimately selected\. In this case, the framework correctly identifies the SPPCS model with all three mechanisms active\. N/A values indicate errors in trajectory reconstruction due to numerical instability of obtained estimators\. #### 5\.6\.4Opinion Dynamics Here we consider an opinion dynamics model describing how individual opinions evolve through local interactions and may eventually converge toward consensus\. Such models have been widely studied in sociology, control theory, and collective behavior; see, for example,\[[18](https://arxiv.org/html/2608.25181#bib.bib10),[96](https://arxiv.org/html/2608.25181#bib.bib25)\]and the references therein\. We study a first\-order model in which each agent’s state represents its opinion, and the interaction kernel determines how strongly one agent influences another as a function of their opinion difference\. The interaction kernel for this model is defined as ϕE\(r\)=\{1,0≤r<1/2,0\.1,1/2≤r≤1,0,r\>1,\\displaystyle\\phi^\{E\}\(r\)=\\begin\{cases\}1,&0\\leq r<1/\\sqrt\{2\},\\\\ 0\.1,&1/\\sqrt\{2\}\\leq r\\leq 1,\\\\ 0,&r\>1,\\end\{cases\}\(84\)with the governing equations then given by the general first\-order system \([71](https://arxiv.org/html/2608.25181#S4.E71)\)\. As summarized in Table[17](https://arxiv.org/html/2608.25181#S5.T17), the opinion dynamics model \([84](https://arxiv.org/html/2608.25181#S5.E84)\) contains only an energy\-based interaction kernel, with no environmental force terms\. Therefore, the correct framework should identify a first\-order pure interaction model withoutff\. The simulation and learning parameters utilized in this experiment are reported in Tables[22](https://arxiv.org/html/2608.25181#A5.T22)and[23](https://arxiv.org/html/2608.25181#A5.T23)\. OrderϕE\\phi^\{E\}ϕA\\phi^\{A\}ff1✓––Table 17:Mechanisms of the phototaxis system \([84](https://arxiv.org/html/2608.25181#S5.E84)\)\. The symbol “\-" indicates that the mechanism is not present, while “✓" indicates that the mechanism is present\.We estimate both mechanisms \(ffandϕE\\phi^\{E\}\) in the general first\-order framework \([71](https://arxiv.org/html/2608.25181#S4.E71)\), with results provided in Figure[16](https://arxiv.org/html/2608.25181#S5.F16)\. Observe that the algorithm correctly identifies a non\-zero interaction kernelϕ^E\\hat\{\\phi\}^\{E\}\(Figure[16\(a\)](https://arxiv.org/html/2608.25181#S5.F16.sf1)\), while obtaining a near zero estimatef^\\hat\{f\}\(Figure[16\(b\)](https://arxiv.org/html/2608.25181#S5.F16.sf2)\)\. The algorithm also exhibits negligible variability across repeated trials, and thus consistently identifies the energy interaction as the only active feature while suppressing the inactive environmental force term \(Figure[16](https://arxiv.org/html/2608.25181#S5.F16)\)\. The model selection framework also confirms the first\-order pure interaction model as the most likely candidate; details are provided in Table[18](https://arxiv.org/html/2608.25181#S5.T18)\. Specifically, we note that all second\-order \(SS\) candidates are rejected by the framework compatibility gate as their relative residual errors satisfyEresrel\>0\.25E\_\{\\mathrm\{res\}\}^\{\\mathrm\{rel\}\}\>0\.25or are undefined due to numerical instability\. The remaining admissible candidates belong to the first\-order \(FF\) category and all satisfy the primary criterion by achieving \(within toleranceτtraj\\tau\_\{\\mathrm\{traj\}\}in \([76](https://arxiv.org/html/2608.25181#S4.E76)\)\) the same minimum validation trajectory errorEminE\_\{\\min\}, according to the criterion defined in \([77](https://arxiv.org/html/2608.25181#S4.E77)\)\. The selection procedure therefore proceeds to the secondary criterion, i\.e\. framework complexity\. The lowest complexity is achieved by candidateF1F\_\{1\}, corresponding to the first\-order energy\-only framework, which is precisely the model utilized to generate the trajectory data\. Thus, even when the candidate library includes more general second\-order frameworks, the procedure correctly identifies that the observed dynamics are best explained by a first\-order interaction model\. \(a\)Interaction kernelϕE\\phi^\{E\}\(b\)Environmental forceff Figure 16:\(Model selection, opinion dynamics\) Mechanisms present for the opinion dynamics system \([84](https://arxiv.org/html/2608.25181#S5.E84)\) within the generalized first\-order framework \([71](https://arxiv.org/html/2608.25181#S4.E71)\)\. The model selection framework estimates an energy interaction kernelϕ^E\\hat\{\\phi\}^\{E\}and an environmental forcef^\\hat\{f\}from observations of trajectories, and correctly identifies a first\-order model dominated by interaction forces\. Estimation is performed overTr=10T\_\{r\}=10independent learning trials, the inferred interaction kernel has weightedL2L^\{2\}‖ϕ^E‖L2\(ρ^R\)=\\\|\\hat\{\\phi\}^\{E\}\\\|\_\{L^\{2\}\(\\hat\{\\rho\}\_\{R\}\)\}=3\.5858×10−013\.5858\\text\{\\times\}\{10\}^\{\-01\}±\\pm1\.7518×10−021\.7518\\text\{\\times\}\{10\}^\{\-02\}\(Figure[16\(a\)](https://arxiv.org/html/2608.25181#S5.F16.sf1)\), while the inferred environmental force has weightedL2L^\{2\}norm‖f^‖L2\(ρ^V\)=\\\|\\hat\{f\}\\\|\_\{L^\{2\}\(\\hat\{\\rho\}\_\{V\}\)\}=0\.0000×10000\.0000\\text\{\\times\}\{10\}^\{00\}±\\pm0\.0000×10000\.0000\\text\{\\times\}\{10\}^\{00\}\(Figure[16\(b\)](https://arxiv.org/html/2608.25181#S5.F16.sf2)\)\. Note the scale of the vertical axes\. Further information regarding learning and model parameters may be found in Appendix[E](https://arxiv.org/html/2608.25181#A5)\.CandidateOrderEresrelϵE\_\{\\mathrm\{res\}\}^\{\\mathrm\{rel\_\{\\epsilon\}\}\}EreconrelE\_\{\\mathrm\{recon\}\}^\{\\mathrm\{rel\}\}E¯traj\(Mval\)\\bar\{E\}\_\{\\mathrm\{traj\}\}\(M\_\{\\mathrm\{val\}\}\)ComplexityEresE\_\{\\mathrm\{res\}\}RankSelectedF1F\_\{1\}10\.1868860\.0062270\.01643010\.0052281YesF2F\_\{2\}10\.1865800\.0063510\.01586720\.0052202NoS2S\_\{2\}20\.9304540\.0207640\.09966230\.0370023NoS5S\_\{5\}20\.9275060\.0420130\.14710150\.0368854NoS1S\_\{1\}20\.9831270\.0348450\.33578530\.0390975NoS3S\_\{3\}2N/AN/AN/A4N/A6NoS4S\_\{4\}20\.9304540\.0207640\.09966240\.0370027NoS6S\_\{6\}20\.9275060\.0420690\.14709860\.0368858NoTable 18:\(Model selection, opinion dynamics\) Summary of the candidate models considered during model selection\. For each candidate, we report the system order, framework compatibility diagnostics, validation trajectory error, model complexity, residual error, ranking, and whether the candidate was ultimately selected\. In this case, the framework correctly identifies the opinion dynamics model described by interactions only\. N/A values indicate errors in trajectory reconstruction due to numerical instability of obtained estimators\. ## 6Discussion and conclusions In this work, we have extended variational learning methods to collective dynamical systems with both inter\-agent interactions and environmental or intra\-agent forces\. The central idea is to exploit the structural form of collective dynamics: although the full system evolves on a high\-dimensional state space, its vector field is often generated by a small number of low\-dimensional functions, such as interaction kernels and environmental force laws\. By learning these functions directly, rather than treating the dynamics as an arbitrary high\-dimensional vector field, the proposed framework provides mechanistic estimators that are both computationally tractable and interpretable\. We introduced two extensions of the variational learning framework\. The first is a semi\-parametric formulation, in which the environmental force is assumed to have a known functional form depending on unknown parameters\. The second is a fully non\-parametric formulation, in which both the environmental force and the interaction kernels are learned from trajectory data\. When prior knowledge of the environmental force is available, the semi\-parametric formulation enables direct estimation of the corresponding parameters\. When such information is not available, the fully non\-parametric formulation provides a flexible alternative for recovering environmental forces directly from observations\. Across the benchmark systems considered in this manuscript, the proposed methods accurately recover the governing mechanisms and produce learned models that predict collective behavior beyond the training time horizon\. These examples include systems with qualitatively different mechanisms, including synchronization, attraction/repulsion dynamics, velocity alignment, and externally driven motion\. The combination of low feature\-recovery errors, small residual errors, accurate trajectory prediction, stability across repeated learning trials, favorable dependence on size of available training data, and robustness to observational noise suggests that the learned models capture the essential mechanistic features of the underlying dynamics\. The results also highlight the importance of the empirical sampling measures induced by the observed trajectories\. Feature recovery is most accurate on regions of state space and pairwise\-distance space that are well sampled by the data, while recovery may deteriorate in regions that are rarely visited\. This reflects an inherent limitation of trajectory\-based inference: one can only expect to learn mechanisms accurately on the portions of the state space explored by the observed dynamics\. Similarly, the accuracy of the semi\-parametric formulation depends on the correctness of the assumed functional form; when this structure is accurate, parameter recovery can be substantially improved, but model mis\-specification can lead to biased or misleading estimates\. We also developed a model selection procedure for determining which mechanistic components are active in the observed dynamics\. This procedure successfully identifies the correct governing structure in the benchmark systems considered, including models with environmental forces, energy\-based interactions, alignment interactions, and combinations of these mechanisms\. Thus, the framework can be used not only for force estimation, but also as a tool for mechanistic model discovery when the relevant dynamical structure is not known a priori\. More broadly, the approach developed here is not limited to the particular collective systems studied in this manuscript\. The essential requirement is that the governing dynamics possess exploitable structure or symmetry, so that the high\-dimensional vector field can be represented in terms of a small number of lower\-dimensional functions or feature maps\. Interaction kernels provide one example of such a reduction: permutation symmetry and pairwise dependence allow the dynamics of many agents to be encoded through functions of pairwise distances, relative velocities, or other low\-dimensional variables\. Similar ideas may be applicable to other structured dynamical systems in which physical principles, invariances, conservation laws, or known mechanistic features reduce the effective dimension of the learning problem\. Several directions remain for future work\. The present study focuses primarily on homogeneous\-agent systems and relatively low\-dimensional interaction laws\. Future work will investigate heterogeneous agents, higher\-dimensional feature dependencies, stochastic dynamics, and more complex environmental couplings\. On the computational side, further study of basis construction, regularization, and hyperparameter selection may improve both accuracy and efficiency\. On the theoretical side, important questions remain concerning identifiability, convergence, approximation properties, and the relationship between the empirical sampling measures and recoverability of the underlying mechanisms\. ## References - \[1\]J\. A\. Carrillo, Y\. Choi, and S\. P\. Perez\(2017\)A review on attractive–repulsive hydrodynamics for consensus in collective behavior\.Active Particles, Volume 1: Advances in Theory, Models, and Applications,pp\. 259–298\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1)\. - \[2\]T\. Kolokolnikov, H\. Sun, D\. Uminsky, and A\. L\. Bertozzi\(2011\)Stability of ring patterns arising from two\-dimensional particle interactions\.Physical Review E—Statistical, Nonlinear, and Soft Matter Physics84\(1\),pp\. 015203\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1)\. - \[3\]T\. Vicsek and A\. Zafeiris\(2012\)Collective motion\.Physics reports517\(3\-4\),pp\. 71–140\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1)\. - \[4\]Y\. Shoham and K\. Leyton\-Brown\(2008\)Multiagent systems: algorithmic, game\-theoretic, and logical foundations\.Cambridge University Press\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1)\. - \[5\]P\. Friedl and D\. Gilmour\(2009\)Collective cell migration in morphogenesis, regeneration and cancer\.Nature reviews Molecular cell biology10\(7\),pp\. 445–457\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1)\. - \[6\]P\. Rørth\(2009\)Collective cell migration\.Annual review of cell and developmental25,pp\. 407–429\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1)\. - \[7\]A\. Czirók and T\. Vicsek\(2000\)Collective behavior of interacting self\-propelled particles\.Physica A: Statistical Mechanics and its Applications281\(1\-4\),pp\. 17–29\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1)\. - \[8\]A\. Sokolov, I\. S\. Aranson, J\. O\. Kessler, and R\. E\. Goldstein\(2007\)Concentration dependence of the collective dynamics of swimming bacteria\.Physical review letters98\(15\),pp\. 158102\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1)\. - \[9\]H\. Zhang, A\. Be’er, E\. Florin, and H\. L\. Swinney\(2010\)Collective motion and density fluctuations in bacterial colonies\.Proceedings of the National Academy of Sciences107\(31\),pp\. 13626–13630\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1)\. - \[10\]F\. Cucker and S\. Smale\(2007\)Emergent behavior in flocks\.IEEE Transactions on automatic control52\(5\),pp\. 852–862\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1),[§1](https://arxiv.org/html/2608.25181#S1.p5.2),[§5\.6\.1](https://arxiv.org/html/2608.25181#S5.SS6.SSS1.p1.1),[§5\.6\.1](https://arxiv.org/html/2608.25181#S5.SS6.SSS1.p1.2)\. - \[11\]F\. Cucker and J\. Dong\(2010\)Avoiding collisions in flocks\.IEEE transactions on automatic control55\(5\),pp\. 1238–1243\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1),[§5\.6\.1](https://arxiv.org/html/2608.25181#S5.SS6.SSS1.p1.1)\. - \[12\]F\. Cucker and J\. Dong\(2011\)A general collision\-avoiding flocking framework\.IEEE Transactions on Automatic Control56\(5\),pp\. 1124–1129\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1),[§5\.6\.1](https://arxiv.org/html/2608.25181#S5.SS6.SSS1.p1.1)\. - \[13\]F\. Cucker and J\. Dong\(2013\)A conditional, collision\-avoiding, model for swarming\.Discrete and Continuous Dynamical Systems34\(3\),pp\. 1009–1020\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1),[§5\.6\.1](https://arxiv.org/html/2608.25181#S5.SS6.SSS1.p1.1)\. - \[14\]T\. Vicsek, A\. Czirók, E\. Ben\-Jacob, I\. Cohen, and O\. Shochet\(1995\)Novel type of phase transition in a system of self\-driven particles\.Physical review letters75\(6\),pp\. 1226\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1),[§5\.6\.1](https://arxiv.org/html/2608.25181#S5.SS6.SSS1.p1.1)\. - \[15\]G\. Albi, D\. Balagué, José, and J\. von Brecht\(2014\)Stability analysis of flock and mill rings for second order models in swarming\.SIAM Journal on Applied Mathematics74\(3\),pp\. 794–818\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1),[§5\.2\.2](https://arxiv.org/html/2608.25181#S5.SS2.SSS2.p1.1)\. - \[16\]Y\. Chuang, T\. Chou, and M\. R\. D’Orsogna\(2016\)Swarming in viscous fluids: three\-dimensional patterns in swimmer\-and force\-induced flows\.Physical Review E93\(4\),pp\. 043112\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1),[§5\.2\.2](https://arxiv.org/html/2608.25181#S5.SS2.SSS2.p1.1)\. - \[17\]M\. Moussaïd, D\. Helbing, and G\. Theraulaz\(2011\)How simple rules determine pedestrian behavior and crowd disasters\.Proceedings of the National Academy of Sciences108\(17\),pp\. 6884–6888\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1),[§1](https://arxiv.org/html/2608.25181#S1.p3.1)\. - \[18\]U\. Krauseet al\.\(2000\)A discrete nonlinear and non\-autonomous model of consensus formation\.Communications in difference equations2000,pp\. 227–236\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1),[§1](https://arxiv.org/html/2608.25181#S1.p5.2),[§5\.6\.4](https://arxiv.org/html/2608.25181#S5.SS6.SSS4.p1.1)\. - \[19\]I\. D\. Couzin, J\. Krause, N\. R\. Franks, and S\. A\. Levin\(2005\)Effective leadership and decision\-making in animal groups on the move\.Nature433\(7025\),pp\. 513–516\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1)\. - \[20\]L\. Li, M\. Nagy, G\. Amichay, R\. Wu, W\. Wang, O\. Deussen, D\. Rus, and I\. D\. Couzin\(2025\)Reverse engineering the control law for schooling in zebrafish using virtual reality\.Science Robotics10\(101\),pp\. eadq6784\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1)\. - \[21\]Z\. Chen, H\. Ding, P\. S\. Kollipara, J\. Li, and Y\. Zheng\(2024\)Synchronous and fully steerable active particle systems for enhanced mimicking of collective motion in nature\.Advanced Materials36\(7\),pp\. 2304759\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1)\. - \[22\]W\. Ren and R\. W\. Beard\(2008\)Distributed consensus in multi\-vehicle cooperative control: theory and applications\.Springer\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1)\. - \[23\]R\. Olfati\-Saber, J\. A\. Fax, and R\. M\. Murray\(2007\)Consensus and cooperation in networked multi\-agent systems\.Proceedings of the IEEE95\(1\),pp\. 215–233\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p1.1)\. - \[24\]M\. Zhong, J\. Miller, and M\. Maggioni\(2020\)Data\-driven discovery of emergent behaviors in collective dynamics\.Physica D: Nonlinear Phenomena411,pp\. 132542\.Cited by:[§1\.1](https://arxiv.org/html/2608.25181#S1.SS1.p1.1),[§1](https://arxiv.org/html/2608.25181#S1.p2.1),[§1](https://arxiv.org/html/2608.25181#S1.p5.1)\. - \[25\]S\. Sahoo, C\. Lampert, and G\. Martius\(2018\)Learning equations for extrapolation and control\.InInternational conference on machine learning,pp\. 4442–4450\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p2.1)\. - \[26\]S\. L\. Brunton, J\. L\. Proctor, and J\. N\. Kutz\(2016\)Discovering governing equations from data by sparse identification of nonlinear dynamical systems\.Proceedings of the national academy of sciences113\(15\),pp\. 3932–3937\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p2.1),[§1](https://arxiv.org/html/2608.25181#S1.p3.2)\. - \[27\]S\. M\. Stigler\(1990\)The history of statistics: the measurement of uncertainty before 1900\.Harvard University Press\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p3.1)\. - \[28\]C\. Becco, N\. Vandewalle, J\. Delcourt, and P\. Poncin\(2006\)Experimental evidences of a structural and dynamical transition in fish school\.Physica A: Statistical Mechanics and its Applications367,pp\. 487–493\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p3.1)\. - \[29\]D\. Pham, M\. Hansen, F\. Dhellemmes, J\. Krause, and P\. Bideau\(2024\)Watching swarm dynamics from above: a framework for advanced object tracking in drone videos\.arXiv preprint arXiv:2406\.07680\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p3.1)\. - \[30\]W\. Bialek, A\. Cavagna, I\. Giardina, T\. Mora, E\. Silvestri, M\. Viale, and A\. M\. Walczak\(2012\)Statistical mechanics for natural flocks of birds\.Proceedings of the National Academy of Sciences109\(13\),pp\. 4786–4791\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p3.1),[§1](https://arxiv.org/html/2608.25181#S1.p3.2)\. - \[31\]E\. G\. de Lamo, M\. C\. Miguel, and R\. Pastor\-Satorras\(2025\)Data\-driven stochastic modeling of schooling fish: from collective dynamics to individual fluctuations\.arXiv preprint arXiv:2509\.08630\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p3.1)\. - \[32\]H\. Yang, F\. Meyer, S\. Huang, L\. Yang, C\. Lungu, M\. A\. Olayioye, M\. J\. Buehler, and M\. Guo\(2024\)Learning collective cell migratory dynamics from a static snapshot with graph neural networks\.PRX Life2\(4\),pp\. 043010\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p3.1)\. - \[33\]I\. Jeliazkov\(2013\)Nonparametric vector autoregressions: specification, estimation, and inference\.InVAR Models in Macroeconomics – New Developments and Applications: Essays in Honor of Christopher A\. Sims,External Links:ISBN 978\-1\-78190\-752\-8Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p3.2)\. - \[34\]R\. DeVore, G\. Kerkyacharian, D\. Picard, and V\. Temlyakov\(2006\)Approximation methods for supervised learning\.Foundations of Computational Mathematics6\(1\),pp\. 3–58\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p3.2)\. - \[35\]P\. Binev, A\. Cohen, W\. Dahmen, R\. DeVore, V\. Temlyakov, and P\. Bartlett\(2005\)Universal algorithms for learning theory part i: piecewise constant functions\.\.Journal of Machine Learning Research6\(9\)\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p3.2)\. - \[36\]H\. Schaeffer, R\. Caflisch, C\. D\. Hauck, and S\. Osher\(2013\)Sparse dynamics for partial differential equations\.Proceedings of the National Academy of Sciences110\(17\),pp\. 6634–6639\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p3.2)\. - \[37\]G\. Tran and R\. Ward\(2017\)Exact recovery of chaotic systems from highly corrupted data\.Multiscale Modeling & Simulation15\(3\),pp\. 1108–1129\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p3.2)\. - \[38\]D\. A\. Messenger and D\. M\. Bortz\(2022\)Learning mean\-field equations from particle data using wsindy\.Physica D: Nonlinear Phenomena439,pp\. 133406\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p3.2)\. - \[39\]D\. A\. Messenger, G\. E\. Wheeler, X\. Liu, and D\. M\. Bortz\(2022\)Learning anisotropic interaction rules from individual trajectories in a heterogeneous cellular population\.Journal of the Royal Society Interface19\(195\)\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p3.2),[§1](https://arxiv.org/html/2608.25181#S1.p5.1)\. - \[40\]M\. Raissi, P\. Perdikaris, and G\. E\. Karniadakis\(2019\)Physics\-informed neural networks: a deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations\.Journal of Computational physics378,pp\. 686–707\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p3.2)\. - \[41\]R\. T\. Chen, Y\. Rubanova, J\. Bettencourt, and D\. K\. Duvenaud\(2018\)Neural ordinary differential equations\.Advances in neural information processing systems31\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p3.2)\. - \[42\]F\. Djeumou, C\. Neary, E\. Goubault, S\. Putot, and U\. Topcu\(2022\)Neural networks with physics\-informed architectures and constraints for dynamical systems modeling\.InLearning for Dynamics and Control Conference,pp\. 263–277\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p3.2)\. - \[43\]R\. Yu and R\. Wang\(2024\)Learning dynamical systems from data: an introduction to physics\-guided deep learning\.Proceedings of the National Academy of Sciences121\(27\),pp\. e2311808121\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p3.2)\. - \[44\]S\. L\. Brunton and J\. N\. Kutz\(2022\)Data\-driven science and engineering: machine learning, dynamical systems, and control\.Cambridge University Press\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p4.1)\. - \[45\]K\. Champion, B\. Lusch, J\. N\. Kutz, and S\. L\. Brunton\(2019\)Data\-driven discovery of coordinates and governing equations\.Proceedings of the National Academy of Sciences116\(45\),pp\. 22445–22451\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p4.1)\. - \[46\]F\. Lu, M\. Zhong, S\. Tang, and M\. Maggioni\(2019\)Nonparametric inference of interaction laws in systems of agents from trajectory data\.Proceedings of the National Academy of Sciences116\(29\),pp\. 14424–14433\.Cited by:[§1\.1](https://arxiv.org/html/2608.25181#S1.SS1.p1.1),[§1](https://arxiv.org/html/2608.25181#S1.p4.1),[§1](https://arxiv.org/html/2608.25181#S1.p5.1),[§1](https://arxiv.org/html/2608.25181#S1.p6.1),[§2\.1\.4](https://arxiv.org/html/2608.25181#S2.SS1.SSS4.p2.2),[§3\.1](https://arxiv.org/html/2608.25181#S3.SS1.p1.1)\. - \[47\]P\. Gelß, S\. Klus, J\. Eisert, and C\. Schütte\(2019\)Multidimensional approximation of nonlinear dynamical systems\.Journal of Computational and Nonlinear Dynamics14\(6\),pp\. 061006\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p4.1)\. - \[48\]D\. A\. Messenger and D\. M\. Bortz\(2021\)Weak sindy: galerkin\-based data\-driven model selection\.Multiscale Modeling & Simulation19\(3\),pp\. 1474–1497\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p5.1)\. - \[49\]S\. H\. Rudy, S\. L\. Brunton, J\. L\. Proctor, and J\. N\. Kutz\(2017\)Data\-driven discovery of partial differential equations\.Science advances3\(4\),pp\. e1602614\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p5.1)\. - \[50\]L\. Boninsegna, F\. Nüske, and C\. Clementi\(2018\)Sparse learning of stochastic dynamical equations\.The Journal of chemical physics148\(24\)\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p5.1)\. - \[51\]Y\. Huang, Y\. Mabrouk, G\. Gompper, and B\. Sabass\(2022\)Sparse inference and active learning of stochastic differential equations from data\.Scientific Reports12\(1\),pp\. 21691\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p5.1)\. - \[52\]M\. Bongini, M\. Fornasier, M\. Hansen, and M\. Maggioni\(2017\)Inferring interaction rules from observations of evolutive systems i: the variational approach\.Mathematical Models and Methods in Applied Sciences27\(05\),pp\. 909–951\.Cited by:[§1\.1](https://arxiv.org/html/2608.25181#S1.SS1.p1.1),[§1](https://arxiv.org/html/2608.25181#S1.p5.1)\. - \[53\]F\. Lu, M\. Maggioni, and S\. Tang\(2021\)Learning interaction kernels in heterogeneous systems of agents from multiple trajectories\.Journal of Machine Learning Research22,pp\. 1–67\.Cited by:[§1\.1](https://arxiv.org/html/2608.25181#S1.SS1.p1.1),[§1](https://arxiv.org/html/2608.25181#S1.p5.1)\. - \[54\]J\. Feng and M\. Zhong\(2024\)Learning collective behaviors from observation\.InExplorations in the Mathematics of Data Science: The Inaugural Volume of the Center for Approximation and Mathematical Data Analytics,pp\. 101–132\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p5.1)\. - \[55\]S\. Ha and D\. Levy\(2009\)Particle, kinetic and fluid models for phototaxis\.Discrete Contin\. Dyn\. Syst\. Ser\. B12\(1\),pp\. 77–108\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p6.2),[§5\.2\.3](https://arxiv.org/html/2608.25181#S5.SS2.SSS3.p1.1)\. - \[56\]M\. R\. D’Orsogna, Y\. Chuang, A\. L\. Bertozzi, and L\. S\. Chayes\(2006\)Self\-propelled particles with soft\-core interactions: patterns, stability, and collapse\.Physical review letters96\(10\),pp\. 104302\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p6.2),[§5\.2\.2](https://arxiv.org/html/2608.25181#S5.SS2.SSS2.p1.1)\. - \[57\]R\. Olfati\-Saber\(2006\)Flocking for multi\-agent dynamic systems: algorithms and theory\.IEEE Transactions on automatic control51\(3\),pp\. 401–420\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p6.2)\. - \[58\]J\. A\. Fax and R\. M\. Murray\(2004\)Information flow and cooperative control of vehicle formations\.IEEE transactions on automatic control49\(9\),pp\. 1465–1476\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p6.2)\. - \[59\]M\. Burger, V\. Capasso, and D\. Morale\(2007\)On an aggregation model with long and short range interactions\.Nonlinear Analysis: Real World Applications8\(3\),pp\. 939–958\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p6.2)\. - \[60\]D\. Helbing, I\. Farkas, and T\. Vicsek\(2000\)Simulating dynamical features of escape panic\.Nature407\(6803\),pp\. 487–490\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p6.2)\. - \[61\]Y\. Kuramoto\(1975\)International symposium on mathematical problems in theoretical physics\.Lecture notes in Physics30,pp\. 420\.Cited by:[§1](https://arxiv.org/html/2608.25181#S1.p6.2),[§5\.2\.1](https://arxiv.org/html/2608.25181#S5.SS2.SSS1.p1.1)\. - \[62\]D\. W\. Scott\(1979\)On optimal and data\-based histograms\.Biometrika66\(3\),pp\. 605–610\.Cited by:[§2\.1\.4](https://arxiv.org/html/2608.25181#S2.SS1.SSS4.p2.2)\. - \[63\]D\. Freedman and P\. Diaconis\(1981\)On the histogram as a density estimator: l 2 theory\.Zeitschrift für Wahrscheinlichkeitstheorie und verwandte Gebiete57\(4\),pp\. 453–476\.Cited by:[§2\.1\.4](https://arxiv.org/html/2608.25181#S2.SS1.SSS4.p2.2)\. - \[64\]M\. Wand\(1997\)Statistical computing and graphics\.The American Statistician51\(1\),pp\. 59–64\.Cited by:[§2\.1\.4](https://arxiv.org/html/2608.25181#S2.SS1.SSS4.p2.2)\. - \[65\]A\. Tarantola\(2005\)Inverse problem theory and methods for model parameter estimation\.SIAM\.Cited by:[§2\.2](https://arxiv.org/html/2608.25181#S2.SS2.p2.3)\. - \[66\]J\. Nocedal and S\. J\. Wright\(2006\)Numerical optimization\.Springer\.Cited by:[§2\.2](https://arxiv.org/html/2608.25181#S2.SS2.p2.3)\. - \[67\]G\. H\. Golub and V\. Pereyra\(1973\)The differentiation of pseudo\-inverses and nonlinear least squares problems whose variables separate\.SIAM Journal on numerical analysis10\(2\),pp\. 413–432\.Cited by:[§2\.2](https://arxiv.org/html/2608.25181#S2.SS2.p2.3)\. - \[68\]G\. Golub and V\. Pereyra\(2003\)Separable nonlinear least squares: the variable projection method and its applications\.Inverse problems19\(2\),pp\. R1–R26\.Cited by:[§2\.2](https://arxiv.org/html/2608.25181#S2.SS2.p2.3)\. - \[69\]D\. P\. O’Leary and B\. W\. Rust\(2013\)Variable projection for nonlinear least squares problems\.Computational Optimization and Applications54\(3\),pp\. 579–593\.Cited by:[§2\.2](https://arxiv.org/html/2608.25181#S2.SS2.p2.3)\. - \[70\]J\. J\. Moré\(2006\)The levenberg\-marquardt algorithm: implementation and theory\.InNumerical analysis: proceedings of the biennial Conference held at Dundee, June 28–July 1, 1977,pp\. 105–116\.Cited by:[§2\.2](https://arxiv.org/html/2608.25181#S2.SS2.p2.3)\. - \[71\]Å\. Björck\(2024\)Numerical methods for least squares problems\.SIAM\.Cited by:[§2\.2](https://arxiv.org/html/2608.25181#S2.SS2.p2.3)\. - \[72\]F\. Van Breugel, J\. N\. Kutz, and B\. W\. Brunton\(2020\)Numerical differentiation of noisy data: a unifying multi\-objective optimization framework\.IEEE Access8,pp\. 196865–196877\.Cited by:[§3\.5\.4](https://arxiv.org/html/2608.25181#S3.SS5.SSS4.p3.1)\. - \[73\]S\. W\. Smithet al\.\(1997\)The scientist and engineer’s guide to digital signal processing\.California Technical Pub\. San Diego\.Cited by:[§3\.5\.4](https://arxiv.org/html/2608.25181#S3.SS5.SSS4.p3.1)\. - \[74\]S\. Butterworthet al\.\(1930\)On the theory of filter amplifiers\.Wireless Engineer7\(6\),pp\. 536–541\.Cited by:[§3\.5\.4](https://arxiv.org/html/2608.25181#S3.SS5.SSS4.p3.1)\. - \[75\]A\. Savitzky and M\. J\. Golay\(1964\)Smoothing and differentiation of data by simplified least squares procedures\.\.Analytical chemistry36\(8\),pp\. 1627–1639\.Cited by:[§3\.5\.4](https://arxiv.org/html/2608.25181#S3.SS5.SSS4.p3.1)\. - \[76\]P\. H\. Eilers\(2003\)A perfect smoother\.Analytical chemistry75\(14\),pp\. 3631–3636\.Cited by:[§3\.5\.4](https://arxiv.org/html/2608.25181#S3.SS5.SSS4.p3.1)\. - \[77\]H\. E\. Rauch, F\. Tung, and C\. T\. Striebel\(1965\)Maximum likelihood estimates of linear dynamic systems\.AIAA journal3\(8\),pp\. 1445–1450\.Cited by:[§3\.5\.4](https://arxiv.org/html/2608.25181#S3.SS5.SSS4.p3.1)\. - \[78\]S\. Jiang, J\. Shi, and S\. Moura\(2024\)A new framework for nonlinear kalman filters\.measurements1,pp\. 2\.Cited by:[§3\.5\.4](https://arxiv.org/html/2608.25181#S3.SS5.SSS4.p3.1)\. - \[79\]M\. Gori, A\. Betti, and S\. Melacci\(2023\)Machine learning: a constraint\-based approach\.Elsevier\.Cited by:[§4](https://arxiv.org/html/2608.25181#S4.p7.1)\. - \[80\]J\. A\. Acebrón, L\. L\. Bonilla, C\. J\. Pérez Vicente, F\. Ritort, and R\. Spigler\(2005\)The kuramoto model: a simple paradigm for synchronization phenomena\.Reviews of modern physics77\(1\),pp\. 137–185\.Cited by:[§5\.2\.1](https://arxiv.org/html/2608.25181#S5.SS2.SSS1.p1.1)\. - \[81\]M\. Breakspear, S\. Heitmann, and A\. Daffertshofer\(2010\)Generative models of cortical oscillations: neurobiological implications of the kuramoto model\.Frontiers in human neuroscience4,pp\. 190\.Cited by:[§5\.2\.1](https://arxiv.org/html/2608.25181#S5.SS2.SSS1.p1.1)\. - \[82\]G\. Filatrella, A\. H\. Nielsen, and N\. F\. Pedersen\(2008\)Analysis of a power grid using a kuramoto\-like model\.The European Physical Journal B61\(4\),pp\. 485–491\.Cited by:[§5\.2\.1](https://arxiv.org/html/2608.25181#S5.SS2.SSS1.p1.1)\. - \[83\]F\. Dörfler, M\. Chertkov, and F\. Bullo\(2013\)Synchronization in complex oscillator networks and smart grids\.Proceedings of the National Academy of Sciences110\(6\),pp\. 2005–2010\.Cited by:[§5\.2\.1](https://arxiv.org/html/2608.25181#S5.SS2.SSS1.p1.1)\. - \[84\]H\. Susanto and P\. Matthews\(2011\)Variational approximations to homoclinic snaking\.Physical Review E—Statistical, Nonlinear, and Soft Matter Physics83\(3\),pp\. 035201\.Cited by:[§5\.2\.1](https://arxiv.org/html/2608.25181#S5.SS2.SSS1.p1.1)\. - \[85\]I\. Z\. Kiss, Y\. Zhai, and J\. L\. Hudson\(2002\)Emerging coherence in a population of chemical oscillators\.Science296\(5573\),pp\. 1676–1678\.Cited by:[§5\.2\.1](https://arxiv.org/html/2608.25181#S5.SS2.SSS1.p1.1)\. - \[86\]K\. O’Keeffe and C\. Bettstetter\(2019\)A review of swarmalators and their potential in bio\-inspired computing\.Micro\-and Nanotechnology Sensors, Systems, and Applications XI10982,pp\. 383–394\.Cited by:[§5\.2\.1](https://arxiv.org/html/2608.25181#S5.SS2.SSS1.p1.1)\. - \[87\]S\. H\. Strogatz\(2000\)From kuramoto to crawford: exploring the onset of synchronization in populations of coupled oscillators\.Physica D: Nonlinear Phenomena143\(1\-4\),pp\. 1–20\.Cited by:[§5\.2\.1](https://arxiv.org/html/2608.25181#S5.SS2.SSS1.p1.1)\. - \[88\]Y\. Chuang, M\. R\. D’orsogna, D\. Marthaler, A\. L\. Bertozzi, and L\. S\. Chayes\(2007\)State transitions and the continuum limit for a 2d interacting, self\-propelled particle system\.Physica D: Nonlinear Phenomena232\(1\),pp\. 33–47\.Cited by:[§5\.2\.2](https://arxiv.org/html/2608.25181#S5.SS2.SSS2.p1.1),[§5\.2\.2](https://arxiv.org/html/2608.25181#S5.SS2.SSS2.p3.1)\. - \[89\]R\. Lukeman, Y\. Li, and L\. Edelstein\-Keshet\(2009\)A conceptual model for milling formations in biological aggregates\.Bulletin of mathematical biology71\(2\),pp\. 352–382\.Cited by:[§5\.2\.2](https://arxiv.org/html/2608.25181#S5.SS2.SSS2.p1.1)\. - \[90\]N\. Abaid and M\. Porfiri\(2010\)Fish in a ring: spatio\-temporal pattern formation in one\-dimensional animal groups\.Journal of The Royal Society Interface7\(51\),pp\. 1441–1453\.Cited by:[§5\.2\.2](https://arxiv.org/html/2608.25181#S5.SS2.SSS2.p1.1)\. - \[91\]A\. J\. Bernoff and C\. M\. Topaz\(2011\)A primer of swarm equilibria\.SIAM Journal on Applied Dynamical Systems10\(1\),pp\. 212–250\.Cited by:[§5\.2\.2](https://arxiv.org/html/2608.25181#S5.SS2.SSS2.p1.1)\. - \[92\]J\. A\. Carrillo, M\. R\. D’Orsogna, and V\. Panferov\(2009\)Double milling in self\-propelled swarms from kinetic theory\.Kinet\. Relat\. Models2\(2\),pp\. 363–378\.Cited by:[§5\.2\.2](https://arxiv.org/html/2608.25181#S5.SS2.SSS2.p1.1)\. - \[93\]P\. Degond, J\. Liu, S\. Motsch, and V\. Panferov\(2011\)Hydrodynamic models of self\-organized dynamics: derivation and existence theory\.arXiv preprint arXiv:1108\.3160\.Cited by:[§5\.2\.2](https://arxiv.org/html/2608.25181#S5.SS2.SSS2.p1.1)\. - \[94\]D\. Bhaya, A\. Takahashi, and A\. R\. Grossman\(2001\)Light regulation of type iv pilus\-dependent motility by chemosensor\-like elements in synechocystis pcc6803\.Proceedings of the National Academy of Sciences98\(13\),pp\. 7540–7545\.Cited by:[§5\.2\.3](https://arxiv.org/html/2608.25181#S5.SS2.SSS3.p1.1)\. - \[95\]M\. Agueh, R\. Illner, and A\. Richardson\(2011\)Analysis and simulations of a refined flocking and swarming model of cucker\-smale type\.Kinetic and Related Models4\(1\),pp\. 1–16\.Cited by:[§5\.6\.1](https://arxiv.org/html/2608.25181#S5.SS6.SSS1.p1.1)\. - \[96\]S\. Motsch and E\. Tadmor\(2014\)Heterophilious dynamics enhances consensus\.SIAM review56\(4\),pp\. 577–621\.Cited by:[§5\.6\.4](https://arxiv.org/html/2608.25181#S5.SS6.SSS4.p1.1)\. ## Appendix AVariational methods for learning collective and environmental forces: fully non\-parametric approach \(second\-order\) In this section, we describe our algorithm for inferring both the environmental and interaction forces in second\-order models of collective dynamics from observed trajectory data\. Specifically, we assume that the dynamics of a system ofNNparticles are described by the following system of second\-order ordinary differential equations \(ODEs\), fori=1,2,…,Ni=1,2,\\ldots,N: xi˙=viv˙i=f\(xi,vi\)\+1N∑j=1j≠iNϕ\(\|xj−xi\|\)\(xj−xi\)\.\\displaystyle\\begin\{split\}\\dot\{x\_\{i\}\}&=v\_\{i\}\\\\ \\dot\{v\}\_\{i\}&=f\(x\_\{i\},v\_\{i\}\)\+\\frac\{1\}\{N\}\\sum\_\{\\begin\{subarray\}\{c\}j=1\\\\ j\\neq i\\end\{subarray\}\}^\{N\}\\phi\(\|x\_\{j\}\-x\_\{i\}\|\)\(x\_\{j\}\-x\_\{i\}\)\.\\end\{split\}\(85\)Most definitions and notations remain the same as in Section[2\.1](https://arxiv.org/html/2608.25181#S2.SS1)\. Herexi=xi\(t\)∈ℝdx\_\{i\}=x\_\{i\}\(t\)\\in\\mathbb\{R\}^\{d\}denotes the state of theithi^\{\\text\{th\}\}agent at timettandvi=x˙i\(t\)v\_\{i\}=\\dot\{x\}\_\{i\}\(t\)denotes the velocity of theithi^\{\\text\{th\}\}agent at timett\. The functionf:ℝ2d→ℝdf:\\mathbb\{R\}^\{2d\}\\to\\mathbb\{R\}^\{d\}models the environmental forces on the agents, which for simplicity we assume to be identical for all agents in the system\. For all the systems considered in this study, we found that the environmental forceffexhibits no dependence on the spatial variablexx\. Therefore, we simplifyf\(xi,vi\)f\(x\_\{i\},v\_\{i\}\)tof\(vi\)f\(v\_\{i\}\), and the corresponding mapping reduces fromf:ℝ2d→ℝdf:\\mathbb\{R\}^\{2d\}\\to\\mathbb\{R\}^\{d\}tof:ℝd→ℝdf:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{d\}\. ### A\.1Trajectory data For trajectory data, which will be utilized to construct the estimators forffandϕ\\phiin \([85](https://arxiv.org/html/2608.25181#A1.E85)\), we will now use both state space and velocity of the system\. Hence the full state of the system with velocity at timettdenote by x\(t\)\\displaystyle x\(t\):=\(x1\(t\)x2\(t\)xN\(t\)\)∈ℝdN,v\(t\):=\(v1\(t\)v2\(t\)vN\(t\)\)∈ℝdN\.\\displaystyle:=\\begin\{pmatrix\}x\_\{1\}\(t\)\\\\ x\_\{2\}\(t\)\\\\ \\vdots\\\\ x\_\{N\}\(t\)\\end\{pmatrix\}\\in\\mathbb\{R\}^\{dN\},\\qquad v\(t\):=\\begin\{pmatrix\}v\_\{1\}\(t\)\\\\ v\_\{2\}\(t\)\\\\ \\vdots\\\\ v\_\{N\}\(t\)\\end\{pmatrix\}\\in\\mathbb\{R\}^\{dN\}\.\(86\)Thus,FfF\_\{f\}becomes, Ff\(x,v\)\\displaystyle F\_\{f\}\(x,v\):=\(f\(x1,v1\)f\(x2,v2\)f\(xN,,vN\)\)∈ℝdN\\displaystyle:=\\begin\{pmatrix\}f\(x\_\{1\},v\_\{1\}\)\\\\ f\(x\_\{2\},v\_\{2\}\)\\\\ \\vdots\\\\ f\(x\_\{N\},,v\_\{N\}\)\\end\{pmatrix\}\\in\\mathbb\{R\}^\{dN\}\\quad\\quad\(87\) withFϕ\(x\)F\_\{\\phi\}\(x\)remains same as in Section[2\.1\.1](https://arxiv.org/html/2608.25181#S2.SS1.SSS1)\. Hence, the general form of \([85](https://arxiv.org/html/2608.25181#A1.E85)\) can be denote as v˙\\displaystyle\\dot\{v\}=Ff\(x,v\)\+Fϕ\(x\)\.\\displaystyle=F\_\{f\}\(x,v\)\+F\_\{\\phi\}\(x\)\.\(88\)We denotexi∈ℝdx\_\{i\}\\in\\mathbb\{R\}^\{d\},vi∈ℝdv\_\{i\}\\in\\mathbb\{R\}^\{d\}as theithi^\{\\text\{th\}\}components ofx,v∈ℝdx,v\\in\\mathbb\{R\}^\{d\}, and will utilize this notation to precisely define the form of the observation data and estimation algorithm below\. Now the initial conditions denote via X0\(m\),V0\(m\)∈ℝdN,\\displaystyle X^\{\(m\)\}\_\{0\},V^\{\(m\)\}\_\{0\}\\in\\mathbb\{R\}^\{dN\},\(89\)form=1,2,…,Mm=1,2,\\ldots,M\. For each suchmm, denote the solution at the timettof the corresponding initial\-value problem \(IVP\) \{x˙=vv˙=Ff\(x,v\)\+Fϕ\(x\)\(x\(0\),v\(0\)\)=\(X0\(m\),V0\(m\)\)\\displaystyle\\begin\{cases\}\\dot\{x\}&=v\\\\ \\dot\{v\}&=F\_\{f\}\(x,v\)\+F\_\{\\phi\}\(x\)\\\\ \(x\(0\),v\(0\)\)&=\(X^\{\(m\)\}\_\{0\},V^\{\(m\)\}\_\{0\}\)\\end\{cases\}\(90\)as\(x\(m\)\(t\),\(v\(m\)\(t\)\)CLOSE\(x^\{\(m\)\}\(t\),\(v^\{\(m\)\}\(t\)\)\. We assume that the experimental data consists of observations at discrete time points0=t1<t2<⋯<tL=T0=t\_\{1\}<t\_\{2\}<\\cdots<t\_\{L\}=T, which we denote as Xℓ\(m\)\\displaystyle X^\{\(m\)\}\_\{\\ell\}:=x\(m\)\(tℓ\)\\displaystyle:=x^\{\(m\)\}\(t\_\{\\ell\}\)\(91\)Vℓ\(m\)\\displaystyle V^\{\(m\)\}\_\{\\ell\}:=x˙\(m\)\(tℓ\),\\displaystyle:=\\dot\{x\}^\{\(m\)\}\(t\_\{\\ell\}\),\(92\)As before,Xℓ\(m\),Vℓ\(m\)∈ℝdNX^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{dN\}, and thus the set\(Xℓ\(m\),Vℓ\(m\)\)m,ℓ=1M,L\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}is generally the set of observed trajectory data which will be utilized to obtain the estimators forffandϕ\\phi\. Note that by construction, \(Xℓ\(m\)\)i\\displaystyle\(X^\{\(m\)\}\_\{\\ell\}\)\_\{i\}=xi\(m\)\(tℓ\)∈ℝd,\\displaystyle=x^\{\(m\)\}\_\{i\}\(t\_\{\\ell\}\)\\in\\mathbb\{R\}^\{d\},\(93\)\(Vℓ\(m\)\)i\\displaystyle\(V^\{\(m\)\}\_\{\\ell\}\)\_\{i\}=vi\(m\)\(tℓ\)∈ℝd,\\displaystyle=v^\{\(m\)\}\_\{i\}\(t\_\{\\ell\}\)\\in\\mathbb\{R\}^\{d\},\(94\)i\.e\. that theithi^\{\\text\{th\}\}component \(with respect to the natural decomposition ofXℓ\(m\),Vℓ\(m\)X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}via \([86](https://arxiv.org/html/2608.25181#A1.E86)\)\) ofXℓ\(m\),Vℓ\(m\)X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}is the state of agentiiat timetℓt\_\{\\ell\}of themthm^\{\\text\{th\}\}replicate and the velocity of agentiiat timetℓt\_\{\\ell\}of themthm^\{\\text\{th\}\}replicate respectively\. The learning algorithm introduced in Section[A\.2](https://arxiv.org/html/2608.25181#A1.SS2)will also require observations of thex¨=v˙\\ddot\{x\}=\\dot\{v\}in \([88](https://arxiv.org/html/2608.25181#A1.E88)\)\. Specifically, we define Aℓ\(m\)\\displaystyle A^\{\(m\)\}\_\{\\ell\}:=v˙\(m\)\(tℓ\),\\displaystyle:=\\dot\{v\}^\{\(m\)\}\(t\_\{\\ell\}\),\(95\)which can be approximated from\(Vℓ\(m\)\)m,ℓ=1M,L\(V^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}as discussed in Section[2\.1\.1](https://arxiv.org/html/2608.25181#S2.SS1.SSS1)for the velocities of the first\-order system\. ### A\.2Estimation algorithm Analogously to the algorithm presented in Section[2\.1\.2](https://arxiv.org/html/2608.25181#S2.SS1.SSS2), we define an error functional of the form ℰℋf,ℋϕ\(f^,ϕ^\)\\displaystyle\\mathcal\{E\}\_\{\\mathcal\{H\}\_\{f\},\\mathcal\{H\}\_\{\\phi\}\}\(\\hat\{f\},\\hat\{\\phi\}\):=1ML∑m,ℓ=1M,L‖Aℓ\(m\)−\(Ff^\(Xℓ\(m\),Vℓ\(m\)\)\+Fϕ^\(Xℓ\(m\)\)\)‖ℝdN2,\\displaystyle:=\\frac\{1\}\{ML\}\\sum\_\{m,\\ell=1\}^\{M,L\}\\left\\lVert A^\{\(m\)\}\_\{\\ell\}\-\\Big\(F\_\{\\hat\{f\}\}\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)\+F\_\{\\hat\{\\phi\}\}\(X^\{\(m\)\}\_\{\\ell\}\)\\Big\)\\right\\rVert\_\{\\mathbb\{R\}^\{dN\}\}^\{2\},\(96\)where\(Xℓ\(m\),Vℓ\(m\),Aℓ\(m\)\)m,ℓ=1M,L\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\},A^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}denotes the \(fixed\) trajectory data as introduced in Section[A\.1](https://arxiv.org/html/2608.25181#A1.SS1) For the second\-order systemshkh\_\{k\}utilize both\(x,v\)\(x,v\)as in \([87](https://arxiv.org/html/2608.25181#A1.E87)\)\. Hence, for each1≤k≤nf1\\leq k\\leq n\_\{f\}, writingx=\(x1,x2,…,xN\)∈ℝdNx=\(x\_\{1\},x\_\{2\},\\ldots,x\_\{N\}\)\\in\\mathbb\{R\}^\{dN\}withxi∈ℝdx\_\{i\}\\in\\mathbb\{R\}^\{d\}andv=\(v1,v2,…,vN\)∈ℝdNv=\(v\_\{1\},v\_\{2\},\\ldots,v\_\{N\}\)\\in\\mathbb\{R\}^\{dN\}withvi∈ℝdv\_\{i\}\\in\\mathbb\{R\}^\{d\}, we define hk\(x,v\)\\displaystyle h\_\{k\}\(x,v\):=\(hk\(x1,v1\)hk\(x2,v2\)⋯hk\(xN,vN\)\)\\displaystyle:=\\begin\{pmatrix\}h\_\{k\}\(x\_\{1\},v\_\{1\}\)\\\\ h\_\{k\}\(x\_\{2\},v\_\{2\}\)\\\\ \\cdots\\\\ h\_\{k\}\(x\_\{N\},v\_\{N\}\)\\end\{pmatrix\}\(97\)and the matrixHℓ\(m\)∈ℝdN×nfH^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{dN\\times n\_\{f\}\} Hℓ\(m\)\\displaystyle H^\{\(m\)\}\_\{\\ell\}:=\(h1\(Xℓ\(m\),Vℓ\(m\)\)h2\(Xℓ\(m\),Vℓ\(m\)\)⋯hnf\(Xℓ\(m\),Vℓ\(m\)\)\)\.\\displaystyle:=\\Big\(\\,h\_\{1\}\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)\\quad h\_\{2\}\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)\\quad\\cdots\\quad h\_\{n\_\{f\}\}\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)\\,\\Big\)\.\(98\) Thus, expandingf~∈ℋf\\tilde\{f\}\\in\\mathcal\{H\}\_\{f\}as in \([19](https://arxiv.org/html/2608.25181#S2.E19)\),Ff~\(Xℓ\(m\),Vℓ\(m\)\)F\_\{\\tilde\{f\}\}\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)in \([96](https://arxiv.org/html/2608.25181#A1.E96)\) takes the form of matrix\-vector multiplication: Ff~\(Xℓ\(m\),Vℓ\(m\)\)\\displaystyle F\_\{\\tilde\{f\}\}\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)=Hℓ\(m\)α\.\\displaystyle=H^\{\(m\)\}\_\{\\ell\}\\alpha\.\(99\) Using \([99](https://arxiv.org/html/2608.25181#A1.E99)\) and \([26](https://arxiv.org/html/2608.25181#S2.E26)\), we observe that minimizing \([96](https://arxiv.org/html/2608.25181#A1.E96)\) overℋf\\mathcal\{H\}\_\{f\}andℋϕ\\mathcal\{H\}\_\{\\phi\}is equivalent to minimizing ℰ\(α,β\)\\displaystyle\\mathcal\{E\}\(\\alpha,\\beta\):=1ML∑m,ℓ=1M,L‖Aℓ\(m\)−\(Hℓ\(m\)α\+Ψℓ\(m\)β\)‖ℝdN2,\\displaystyle:=\\frac\{1\}\{ML\}\\sum\_\{m,\\ell=1\}^\{M,L\}\\left\\lVert A^\{\(m\)\}\_\{\\ell\}\-\\Big\(H^\{\(m\)\}\_\{\\ell\}\\alpha\+\\Psi^\{\(m\)\}\_\{\\ell\}\\beta\\Big\)\\right\\rVert\_\{\\mathbb\{R\}^\{dN\}\}^\{2\},\(100\)over\(α,β\)∈ℝnf\+nϕ\(\\alpha,\\beta\)\\in\\mathbb\{R\}^\{n\_\{f\}\+n\_\{\\phi\}\}\. As \([100](https://arxiv.org/html/2608.25181#A1.E100)\) is quadratic in\(α,β\)\(\\alpha,\\beta\)\(recall that the norm‖⋅‖ℝdN\\left\\lVert\\cdot\\right\\rVert\_\{\\mathbb\{R\}^\{dN\}\}is defined in \([15](https://arxiv.org/html/2608.25181#S2.E15)\) \- \([16](https://arxiv.org/html/2608.25181#S2.E16)\) via the standard Euclidean inner product onℝd\\mathbb\{R\}^\{d\}\), a solution\(α^,β^\)\(\\hat\{\\alpha\},\\hat\{\\beta\}\)always exists: \(α^,β^\)\\displaystyle\(\\hat\{\\alpha\},\\hat\{\\beta\}\):=argminα∈ℝnf,β∈ℝnϕℰ\(α,β\)\.\\displaystyle:=\\argmin\_\{\\alpha\\in\\mathbb\{R\}^\{n\_\{f\}\},\\,\\beta\\in\\mathbb\{R\}^\{n\_\{\\phi\}\}\}\\mathcal\{E\}\(\\alpha,\\beta\)\.\(101\) The normal equations which characterize the solutions of \([101](https://arxiv.org/html/2608.25181#A1.E101)\) and the solution approach is same as in section[2\.1\.3](https://arxiv.org/html/2608.25181#S2.SS1.SSS3)\. ### A\.3Algorithm pseudocode Algorithm 3Algorithm for learningϕ\\phi\(interaction\) andff\(environmental\) forces from trajectory data of second\-order systems of the form \([85](https://arxiv.org/html/2608.25181#A1.E85)\)\.Input: \(Xℓ\(m\),Vℓ\(m\)Aℓ\(m\)\)m,ℓ=1M,L\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}A^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\} Output:Estimators: ϕ^\\hat\{\\phi\}, f^\\hat\{f\} Construct pairwise quantities and estimate the observed supports Generate interaction distances \(Rℓ\(m\)\)m,ℓ=1M,L\(R^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell=1\}^\{M,L\}Find the maximum and minimum interaction radii Rmin,Rmax∈ℝR\_\{\\min\},R\_\{\\max\}\\in\\mathbb\{R\} Find the maximum and minimum values of state variable and velocity variable Xmin,Xmax,Vmin,Vmax∈ℝdX\_\{\\min\},X\_\{\\max\},V\_\{\\min\},V\_\{\\max\}\\in\\mathbb\{R\}^\{d\} Observed supports \[Rmin,Rmax\],∏k=1d\[Xk,min,,Xk,max\]×∏k=1d\[Vk,min,Vk,max\]\[R\_\{\\min\},R\_\{\\max\}\],\\prod\_\{k=1\}^\{d\}\[X\_\{k,\\min\},,X\_\{k,\\max\}\]\\times\\prod\_\{k=1\}^\{d\}\[V\_\{k,\\min\},V\_\{k,\\max\}\] Construct localized basis functions Construct the kernel basis \(ψk\)k=1nϕ\(\\psi\_\{k\}\)\_\{k=1\}^\{n\_\{\\phi\}\} Construct \(hk\)k=1nf\(h\_\{k\}\)\_\{k=1\}^\{n\_\{f\}\} Create Cℓ\(m\)C^\{\(m\)\}\_\{\\ell\}by column\-wise concatenating \(Hℓ\(m\)Ψℓ\(m\)\)\(\\,H^\{\(m\)\}\_\{\\ell\}\\quad\\Psi^\{\(m\)\}\_\{\\ell\}\\,\)as in \([30](https://arxiv.org/html/2608.25181#S2.E30)\) Assemble Aθ^=b→A\\hat\{\\theta\}=\\vec\{b\}\(in parallel for ℓ,m\\ell,m\), using \([35](https://arxiv.org/html/2608.25181#S2.E35)\), \([36](https://arxiv.org/html/2608.25181#S2.E36)\), \([37](https://arxiv.org/html/2608.25181#S2.E37)\), \([29](https://arxiv.org/html/2608.25181#S2.E29)\) Solve for θ^\\hat\{\\theta\} Assemble ϕ^\\hat\{\\phi\}as in \([19](https://arxiv.org/html/2608.25181#S2.E19)\) Assemble f^=∑k=1nfαkhk\\hat\{f\}=\\sum\_\{k=1\}^\{n\_\{f\}\}\\alpha\_\{k\}h\_\{k\} ## Appendix BHypothesis spaces and empirical measures Here we provide details on both the hypothesis spaces \(Section[2\.1\.4](https://arxiv.org/html/2608.25181#S2.SS1.SSS4)\) and trajectory\-induced empirical measures \(Section[3\.1](https://arxiv.org/html/2608.25181#S3.SS1)\) for the model systems considered in Section[5\.2](https://arxiv.org/html/2608.25181#S5.SS2)\. ### B\.1Kuramoto system The following visualizations illustrate the basis functions for the hypothesis spaces \(Figure[17](https://arxiv.org/html/2608.25181#A2.F17)\) and induced measures \(Figure[18](https://arxiv.org/html/2608.25181#A2.F18)\) utilized for the Kuramoto model \(Section[5\.2\.1](https://arxiv.org/html/2608.25181#S5.SS2.SSS1)\), as described in Sections[2\.1\.4](https://arxiv.org/html/2608.25181#S2.SS1.SSS4)and[3\.1](https://arxiv.org/html/2608.25181#S3.SS1), respectively\. These quantities define the hypothesis spaces and empirical measures employed by the estimation algorithms\. \(a\)Basis functions definingℋϕ\\mathcal\{H\}\_\{\\phi\}\(b\)Basis functions definingℋf\\mathcal\{H\}\_\{f\} Figure 17:\(Kuramoto model\) Linear B\-spline basis functions used to construct hypothesis spaces\. Subfigure[17\(a\)](https://arxiv.org/html/2608.25181#A2.F17.sf1)plots the interaction kernel basis functions\(ψk\)k=1nϕ\(\\psi\_\{k\}\)\_\{k=1\}^\{n\_\{\\phi\}\}inHϕH\_\{\\phi\}withnϕ=10n\_\{\\phi\}=10\. Subfigure[17\(b\)](https://arxiv.org/html/2608.25181#A2.F17.sf2)plots the environmental force basis functions\(hk\)k=1nf\(h\_\{k\}\)\_\{k=1\}^\{n\_\{f\}\}inHfH\_\{f\}withnf=10n\_\{f\}=10\.\(a\)Induced measure onΔθ\\Delta\\theta\(b\)Induced measure onθ\\theta Figure 18:\(Kuramoto model\) Comparison of the empirical measures obtained from theMtrainM\_\{\\mathrm\{train\}\}training data and the reference probability measures fromMρM\_\{\\rho\}samples, as discussed in Section[3\.1](https://arxiv.org/html/2608.25181#S3.SS1)\. See Table[22](https://arxiv.org/html/2608.25181#A5.T22)for parameters utilized\. Subfigure[18\(a\)](https://arxiv.org/html/2608.25181#A2.F18.sf1)plotsρ^Δθ\\hat\{\\rho\}\_\{\\Delta\\theta\}versusρΔθ\\rho\_\{\\Delta\\theta\}, and Subfigure[18\(b\)](https://arxiv.org/html/2608.25181#A2.F18.sf2)plotsρ^θ\\hat\{\\rho\}\_\{\\theta\}versusρθ\\rho\_\{\\theta\}\. The close agreement suggests that the training trajectory data provides an accurate estimate for the regions explored by the dynamics\. ### B\.2Phototaxis system The following visualizations illustrate the basis functions for the hypothesis spaces \(Figure[19](https://arxiv.org/html/2608.25181#A2.F19)\) and induced measures \(Figure[20](https://arxiv.org/html/2608.25181#A2.F20)\) utilized for the phototaxis model \(Section[5\.2\.3](https://arxiv.org/html/2608.25181#S5.SS2.SSS3)\) as described in Sections[2\.1\.4](https://arxiv.org/html/2608.25181#S2.SS1.SSS4)and[3\.1](https://arxiv.org/html/2608.25181#S3.SS1), respectively\. These quantities define the hypothesis spaces and empirical measures employed by the learning algorithm\. \(a\)Basis functions definingℋϕ\\mathcal\{H\}\_\{\\phi\}\(b\)Basis functions definingℋf\\mathcal\{H\}\_\{f\} Figure 19:\(Phototaxis model\) Linear B\-spline basis functions used to construct hypothesis spaces\. Subfigure[19\(a\)](https://arxiv.org/html/2608.25181#A2.F19.sf1)plots the interaction kernel basis functions\(ψk\)k=1nϕ\(\\psi\_\{k\}\)\_\{k=1\}^\{n\_\{\\phi\}\}inHϕH\_\{\\phi\}withnϕ=8n\_\{\\phi\}=8\. Subfigure[19\(b\)](https://arxiv.org/html/2608.25181#A2.F19.sf2)plots the environmental force basis functions\(hk\)k=1nf\(h\_\{k\}\)\_\{k=1\}^\{n\_\{f\}\}inHfH\_\{f\}, where the velocity space is discretized usingnb=10n\_\{b\}=10basis functions in each dimension, resulting innf=10×10=100n\_\{f\}=10\\times 10=100tensor product basis functions forℋf\\mathcal\{H\}\_\{f\}\.\(a\)Induced measure onRR\(b\)Induced measure onVV Figure 20:\(Phototaxis model\) Comparison of the empirical measures obtained from theMtrainM\_\{\\mathrm\{train\}\}training data and the reference measure fromMρM\_\{\\rho\}samples, as discussed in Section[3\.1](https://arxiv.org/html/2608.25181#S3.SS1)\. See Table[22](https://arxiv.org/html/2608.25181#A5.T22)for parameters utilized\. Subfigure[20\(a\)](https://arxiv.org/html/2608.25181#A2.F20.sf1)plotsρ^R\\hat\{\\rho\}\_\{R\}versusρR\\rho\_\{R\}, and Subfigure[20\(b\)](https://arxiv.org/html/2608.25181#A2.F20.sf2)plotsρ^V\\hat\{\\rho\}\_\{V\}versusρV\\rho\_\{V\}\. The close agreement suggests that the training trajectory data provides an accurate estimate for the regions explored by the dynamics\. ### B\.3Self\-propelled particle \(SPP\) system The following visualizations illustrate the basis functions for the hypothesis spaces \(Figure[21](https://arxiv.org/html/2608.25181#A2.F21)\) and induced measures \(Figure[22](https://arxiv.org/html/2608.25181#A2.F22)\) utilized for the self\-propelled particle \(SPP\) model \(Section[5\.2\.2](https://arxiv.org/html/2608.25181#S5.SS2.SSS2)\) as described in Sections[2\.1\.4](https://arxiv.org/html/2608.25181#S2.SS1.SSS4)and[3\.1](https://arxiv.org/html/2608.25181#S3.SS1), respectively\. These quantities define the hypothesis spaces and empirical measures employed by the learning algorithm\. \(a\)Basis functions definingℋϕ\\mathcal\{H\}\_\{\\phi\}\(b\)Basis functions definingℋf\\mathcal\{H\}\_\{f\} Figure 21:\(SPP model\) Linear B\-spline basis functions used to construct the hypothesis spaces\. Subfigure[21\(a\)](https://arxiv.org/html/2608.25181#A2.F21.sf1)plots the interaction kernel basis functions\(ψk\)k=1nϕ\(\\psi\_\{k\}\)\_\{k=1\}^\{n\_\{\\phi\}\}inHϕH\_\{\\phi\}withnϕ=8n\_\{\\phi\}=8\. Subfigure[19\(b\)](https://arxiv.org/html/2608.25181#A2.F19.sf2)plots the environmental force basis functions\(hk\)k=1nf\(h\_\{k\}\)\_\{k=1\}^\{n\_\{f\}\}inHfH\_\{f\}\. The velocity space is discretized usingnb=10n\_\{b\}=10basis functions in each dimension, resulting innf=10×10=100n\_\{f\}=10\\times 10=100tensor product basis functions forℋf\\mathcal\{H\}\_\{f\}\.\(a\)Induced measure onRR\(b\)Induced measure onVV Figure 22:\(SPP\) Comparison of the empirical measures obtained from theMtrainM\_\{\\mathrm\{train\}\}training data and the reference measure fromMρM\_\{\\rho\}samples, as discussed in Section[3\.1](https://arxiv.org/html/2608.25181#S3.SS1)\. See Table[22](https://arxiv.org/html/2608.25181#A5.T22)for parameters utilized\. Subfigure[22\(a\)](https://arxiv.org/html/2608.25181#A2.F22.sf1)plotsρ^R\\hat\{\\rho\}\_\{R\}versusρR\\rho\_\{R\}, and Subfigure[22\(b\)](https://arxiv.org/html/2608.25181#A2.F22.sf2)plotsρ^V\\hat\{\\rho\}\_\{V\}versusρV\\rho\_\{V\}\. The close agreement suggests that the training trajectory data provides an accurate estimate for the regions explored by the dynamics\. ## Appendix CMetric evaluations We provide tables quantifying estimation accuracy as described in Section[3](https://arxiv.org/html/2608.25181#S3), as well as the effect of measurement noise on reconstruction\. ### C\.1Dependence on training data Table 19:\(Kuramoto model\) Kuramoto relative interaction kernel and environmental force errors for differentMM\(number of replicates\) andLL\(number of time samples\)\. Each entry reports the mean±\\pmsample standard deviation overTr=10T\_\{r\}=10independent learning trials\.MLEϕrelE\_\{\\phi\}^\{\\mathrm\{rel\}\}EfrelE\_\{f\}^\{\\mathrm\{rel\}\}1511\.3881×1000±2\.2499×1000$1\.3881\\text\{\\times\}\{10\}^\{00\}$\\pm$2\.2499\\text\{\\times\}\{10\}^\{00\}$3\.6814×1001±6\.4836×1001$3\.6814\\text\{\\times\}\{10\}^\{01\}$\\pm$6\.4836\\text\{\\times\}\{10\}^\{01\}$11011\.8696×1000±2\.0156×1000$1\.8696\\text\{\\times\}\{10\}^\{00\}$\\pm$2\.0156\\text\{\\times\}\{10\}^\{00\}$4\.8435×1001±5\.6091×1001$4\.8435\\text\{\\times\}\{10\}^\{01\}$\\pm$5\.6091\\text\{\\times\}\{10\}^\{01\}$12012\.8650×1000±3\.9872×1000$2\.8650\\text\{\\times\}\{10\}^\{00\}$\\pm$3\.9872\\text\{\\times\}\{10\}^\{00\}$6\.4122×1001±8\.9776×1001$6\.4122\\text\{\\times\}\{10\}^\{01\}$\\pm$8\.9776\\text\{\\times\}\{10\}^\{01\}$14011\.7684×1000±2\.3126×1000$1\.7684\\text\{\\times\}\{10\}^\{00\}$\\pm$2\.3126\\text\{\\times\}\{10\}^\{00\}$4\.2537×1001±5\.9298×1001$4\.2537\\text\{\\times\}\{10\}^\{01\}$\\pm$5\.9298\\text\{\\times\}\{10\}^\{01\}$18012\.6533×1000±3\.9945×1000$2\.6533\\text\{\\times\}\{10\}^\{00\}$\\pm$3\.9945\\text\{\\times\}\{10\}^\{00\}$7\.0362×1001±1\.1683×1002$7\.0362\\text\{\\times\}\{10\}^\{01\}$\\pm$1\.1683\\text\{\\times\}\{10\}^\{02\}$4511\.2803×10−02±2\.1729×10−03$1\.2803\\text\{\\times\}\{10\}^\{\-02\}$\\pm$2\.1729\\text\{\\times\}\{10\}^\{\-03\}$8\.1026×10−02±6\.6889×10−02$8\.1026\\text\{\\times\}\{10\}^\{\-02\}$\\pm$6\.6889\\text\{\\times\}\{10\}^\{\-02\}$41011\.1935×10−02±2\.5299×10−03$1\.1935\\text\{\\times\}\{10\}^\{\-02\}$\\pm$2\.5299\\text\{\\times\}\{10\}^\{\-03\}$7\.9669×10−02±5\.4188×10−02$7\.9669\\text\{\\times\}\{10\}^\{\-02\}$\\pm$5\.4188\\text\{\\times\}\{10\}^\{\-02\}$42011\.2520×10−02±1\.2240×10−03$1\.2520\\text\{\\times\}\{10\}^\{\-02\}$\\pm$1\.2240\\text\{\\times\}\{10\}^\{\-03\}$7\.9739×10−02±2\.9116×10−02$7\.9739\\text\{\\times\}\{10\}^\{\-02\}$\\pm$2\.9116\\text\{\\times\}\{10\}^\{\-02\}$44011\.2743×10−02±2\.0112×10−03$1\.2743\\text\{\\times\}\{10\}^\{\-02\}$\\pm$2\.0112\\text\{\\times\}\{10\}^\{\-03\}$7\.7073×10−02±4\.9390×10−02$7\.7073\\text\{\\times\}\{10\}^\{\-02\}$\\pm$4\.9390\\text\{\\times\}\{10\}^\{\-02\}$48011\.4557×10−02±3\.2833×10−03$1\.4557\\text\{\\times\}\{10\}^\{\-02\}$\\pm$3\.2833\\text\{\\times\}\{10\}^\{\-03\}$1\.4730×10−01±1\.3601×10−01$1\.4730\\text\{\\times\}\{10\}^\{\-01\}$\\pm$1\.3601\\text\{\\times\}\{10\}^\{\-01\}$16511\.1944×10−02±6\.2502×10−04$1\.1944\\text\{\\times\}\{10\}^\{\-02\}$\\pm$6\.2502\\text\{\\times\}\{10\}^\{\-04\}$2\.6885×10−02±7\.5355×10−03$2\.6885\\text\{\\times\}\{10\}^\{\-02\}$\\pm$7\.5355\\text\{\\times\}\{10\}^\{\-03\}$161011\.1872×10−02±7\.0650×10−04$1\.1872\\text\{\\times\}\{10\}^\{\-02\}$\\pm$7\.0650\\text\{\\times\}\{10\}^\{\-04\}$2\.7195×10−02±6\.0489×10−03$2\.7195\\text\{\\times\}\{10\}^\{\-02\}$\\pm$6\.0489\\text\{\\times\}\{10\}^\{\-03\}$162011\.1717×10−02±5\.9480×10−04$1\.1717\\text\{\\times\}\{10\}^\{\-02\}$\\pm$5\.9480\\text\{\\times\}\{10\}^\{\-04\}$2\.4590×10−02±1\.0248×10−02$2\.4590\\text\{\\times\}\{10\}^\{\-02\}$\\pm$1\.0248\\text\{\\times\}\{10\}^\{\-02\}$164011\.1988×10−02±3\.4998×10−04$1\.1988\\text\{\\times\}\{10\}^\{\-02\}$\\pm$3\.4998\\text\{\\times\}\{10\}^\{\-04\}$2\.0456×10−02±7\.1757×10−03$2\.0456\\text\{\\times\}\{10\}^\{\-02\}$\\pm$7\.1757\\text\{\\times\}\{10\}^\{\-03\}$168011\.1814×10−02±9\.0989×10−04$1\.1814\\text\{\\times\}\{10\}^\{\-02\}$\\pm$9\.0989\\text\{\\times\}\{10\}^\{\-04\}$2\.7961×10−02±9\.1794×10−03$2\.7961\\text\{\\times\}\{10\}^\{\-02\}$\\pm$9\.1794\\text\{\\times\}\{10\}^\{\-03\}$64511\.1989×10−02±3\.8711×10−04$1\.1989\\text\{\\times\}\{10\}^\{\-02\}$\\pm$3\.8711\\text\{\\times\}\{10\}^\{\-04\}$1\.3728×10−02±5\.0689×10−03$1\.3728\\text\{\\times\}\{10\}^\{\-02\}$\\pm$5\.0689\\text\{\\times\}\{10\}^\{\-03\}$641011\.2009×10−02±5\.0137×10−04$1\.2009\\text\{\\times\}\{10\}^\{\-02\}$\\pm$5\.0137\\text\{\\times\}\{10\}^\{\-04\}$1\.2694×10−02±3\.9811×10−03$1\.2694\\text\{\\times\}\{10\}^\{\-02\}$\\pm$3\.9811\\text\{\\times\}\{10\}^\{\-03\}$642011\.2031×10−02±4\.0502×10−04$1\.2031\\text\{\\times\}\{10\}^\{\-02\}$\\pm$4\.0502\\text\{\\times\}\{10\}^\{\-04\}$1\.1495×10−02±4\.8004×10−03$1\.1495\\text\{\\times\}\{10\}^\{\-02\}$\\pm$4\.8004\\text\{\\times\}\{10\}^\{\-03\}$644011\.2195×10−02±3\.7997×10−04$1\.2195\\text\{\\times\}\{10\}^\{\-02\}$\\pm$3\.7997\\text\{\\times\}\{10\}^\{\-04\}$1\.3889×10−02±5\.5731×10−03$1\.3889\\text\{\\times\}\{10\}^\{\-02\}$\\pm$5\.5731\\text\{\\times\}\{10\}^\{\-03\}$648011\.2176×10−02±3\.7529×10−04$1\.2176\\text\{\\times\}\{10\}^\{\-02\}$\\pm$3\.7529\\text\{\\times\}\{10\}^\{\-04\}$1\.3101×10−02±2\.5753×10−03$1\.3101\\text\{\\times\}\{10\}^\{\-02\}$\\pm$2\.5753\\text\{\\times\}\{10\}^\{\-03\}$256511\.2236×10−02±1\.8152×10−04$1\.2236\\text\{\\times\}\{10\}^\{\-02\}$\\pm$1\.8152\\text\{\\times\}\{10\}^\{\-04\}$7\.1149×10−03±1\.6387×10−03$7\.1149\\text\{\\times\}\{10\}^\{\-03\}$\\pm$1\.6387\\text\{\\times\}\{10\}^\{\-03\}$2561011\.2238×10−02±2\.0142×10−04$1\.2238\\text\{\\times\}\{10\}^\{\-02\}$\\pm$2\.0142\\text\{\\times\}\{10\}^\{\-04\}$6\.9159×10−03±2\.1138×10−03$6\.9159\\text\{\\times\}\{10\}^\{\-03\}$\\pm$2\.1138\\text\{\\times\}\{10\}^\{\-03\}$2562011\.2295×10−02±1\.7811×10−04$1\.2295\\text\{\\times\}\{10\}^\{\-02\}$\\pm$1\.7811\\text\{\\times\}\{10\}^\{\-04\}$6\.6019×10−03±2\.4540×10−03$6\.6019\\text\{\\times\}\{10\}^\{\-03\}$\\pm$2\.4540\\text\{\\times\}\{10\}^\{\-03\}$2564011\.2287×10−02±1\.3753×10−04$1\.2287\\text\{\\times\}\{10\}^\{\-02\}$\\pm$1\.3753\\text\{\\times\}\{10\}^\{\-04\}$6\.6174×10−03±2\.2207×10−03$6\.6174\\text\{\\times\}\{10\}^\{\-03\}$\\pm$2\.2207\\text\{\\times\}\{10\}^\{\-03\}$2568011\.2283×10−02±1\.3878×10−04$1\.2283\\text\{\\times\}\{10\}^\{\-02\}$\\pm$1\.3878\\text\{\\times\}\{10\}^\{\-04\}$6\.1803×10−03±1\.9047×10−03$6\.1803\\text\{\\times\}\{10\}^\{\-03\}$\\pm$1\.9047\\text\{\\times\}\{10\}^\{\-03\}$Table 20:\(Phototaxis model\) Phototaxis model relative interaction kernel and environmental force errors for differentMM\(number of replicates\) andLL\(number of time samples\)\. Each entry reports the mean±\\pmsample standard deviation overTr=10T\_\{r\}=10independent learning trials\.MLEϕrelE\_\{\\phi\}^\{\\mathrm\{rel\}\}EfrelE\_\{f\}^\{\\mathrm\{rel\}\}1516\.1893×10−03±6\.2809×10−03$6\.1893\\text\{\\times\}\{10\}^\{\-03\}$\\pm$6\.2809\\text\{\\times\}\{10\}^\{\-03\}$6\.1180×10−02±1\.3356×10−02$6\.1180\\text\{\\times\}\{10\}^\{\-02\}$\\pm$1\.3356\\text\{\\times\}\{10\}^\{\-02\}$11014\.4905×10−03±4\.3087×10−03$4\.4905\\text\{\\times\}\{10\}^\{\-03\}$\\pm$4\.3087\\text\{\\times\}\{10\}^\{\-03\}$1\.1115×10−01±8\.8252×10−02$1\.1115\\text\{\\times\}\{10\}^\{\-01\}$\\pm$8\.8252\\text\{\\times\}\{10\}^\{\-02\}$12011\.3141×10−02±2\.9193×10−02$1\.3141\\text\{\\times\}\{10\}^\{\-02\}$\\pm$2\.9193\\text\{\\times\}\{10\}^\{\-02\}$8\.5239×10−02±3\.5008×10−02$8\.5239\\text\{\\times\}\{10\}^\{\-02\}$\\pm$3\.5008\\text\{\\times\}\{10\}^\{\-02\}$14015\.4909×10−03±3\.4038×10−03$5\.4909\\text\{\\times\}\{10\}^\{\-03\}$\\pm$3\.4038\\text\{\\times\}\{10\}^\{\-03\}$8\.5608×10−02±3\.6644×10−02$8\.5608\\text\{\\times\}\{10\}^\{\-02\}$\\pm$3\.6644\\text\{\\times\}\{10\}^\{\-02\}$18018\.6348×10−03±1\.2248×10−02$8\.6348\\text\{\\times\}\{10\}^\{\-03\}$\\pm$1\.2248\\text\{\\times\}\{10\}^\{\-02\}$7\.7069×10−02±1\.3662×10−02$7\.7069\\text\{\\times\}\{10\}^\{\-02\}$\\pm$1\.3662\\text\{\\times\}\{10\}^\{\-02\}$4511\.9382×10−04±3\.1394×10−05$1\.9382\\text\{\\times\}\{10\}^\{\-04\}$\\pm$3\.1394\\text\{\\times\}\{10\}^\{\-05\}$1\.2214×10−02±9\.3570×10−03$1\.2214\\text\{\\times\}\{10\}^\{\-02\}$\\pm$9\.3570\\text\{\\times\}\{10\}^\{\-03\}$41011\.8060×10−04±2\.3614×10−05$1\.8060\\text\{\\times\}\{10\}^\{\-04\}$\\pm$2\.3614\\text\{\\times\}\{10\}^\{\-05\}$1\.0234×10−02±6\.3196×10−03$1\.0234\\text\{\\times\}\{10\}^\{\-02\}$\\pm$6\.3196\\text\{\\times\}\{10\}^\{\-03\}$42011\.8843×10−04±3\.3100×10−05$1\.8843\\text\{\\times\}\{10\}^\{\-04\}$\\pm$3\.3100\\text\{\\times\}\{10\}^\{\-05\}$9\.0533×10−03±7\.2027×10−03$9\.0533\\text\{\\times\}\{10\}^\{\-03\}$\\pm$7\.2027\\text\{\\times\}\{10\}^\{\-03\}$44012\.1438×10−04±3\.3710×10−05$2\.1438\\text\{\\times\}\{10\}^\{\-04\}$\\pm$3\.3710\\text\{\\times\}\{10\}^\{\-05\}$1\.5331×10−02±9\.6510×10−03$1\.5331\\text\{\\times\}\{10\}^\{\-02\}$\\pm$9\.6510\\text\{\\times\}\{10\}^\{\-03\}$48011\.9470×10−04±4\.7980×10−05$1\.9470\\text\{\\times\}\{10\}^\{\-04\}$\\pm$4\.7980\\text\{\\times\}\{10\}^\{\-05\}$8\.2751×10−03±6\.2136×10−03$8\.2751\\text\{\\times\}\{10\}^\{\-03\}$\\pm$6\.2136\\text\{\\times\}\{10\}^\{\-03\}$16511\.9942×10−04±3\.0756×10−05$1\.9942\\text\{\\times\}\{10\}^\{\-04\}$\\pm$3\.0756\\text\{\\times\}\{10\}^\{\-05\}$6\.0096×10−05±1\.0329×10−05$6\.0096\\text\{\\times\}\{10\}^\{\-05\}$\\pm$1\.0329\\text\{\\times\}\{10\}^\{\-05\}$161011\.9766×10−04±2\.4678×10−05$1\.9766\\text\{\\times\}\{10\}^\{\-04\}$\\pm$2\.4678\\text\{\\times\}\{10\}^\{\-05\}$6\.0569×10−05±8\.5808×10−06$6\.0569\\text\{\\times\}\{10\}^\{\-05\}$\\pm$8\.5808\\text\{\\times\}\{10\}^\{\-06\}$162012\.1672×10−04±3\.2822×10−05$2\.1672\\text\{\\times\}\{10\}^\{\-04\}$\\pm$3\.2822\\text\{\\times\}\{10\}^\{\-05\}$6\.8149×10−05±1\.3025×10−05$6\.8149\\text\{\\times\}\{10\}^\{\-05\}$\\pm$1\.3025\\text\{\\times\}\{10\}^\{\-05\}$164012\.0376×10−04±3\.0521×10−05$2\.0376\\text\{\\times\}\{10\}^\{\-04\}$\\pm$3\.0521\\text\{\\times\}\{10\}^\{\-05\}$5\.9792×10−05±1\.2887×10−05$5\.9792\\text\{\\times\}\{10\}^\{\-05\}$\\pm$1\.2887\\text\{\\times\}\{10\}^\{\-05\}$168012\.0988×10−04±3\.5117×10−05$2\.0988\\text\{\\times\}\{10\}^\{\-04\}$\\pm$3\.5117\\text\{\\times\}\{10\}^\{\-05\}$6\.7264×10−05±1\.6062×10−05$6\.7264\\text\{\\times\}\{10\}^\{\-05\}$\\pm$1\.6062\\text\{\\times\}\{10\}^\{\-05\}$64512\.2848×10−04±1\.9651×10−05$2\.2848\\text\{\\times\}\{10\}^\{\-04\}$\\pm$1\.9651\\text\{\\times\}\{10\}^\{\-05\}$3\.3921×10−05±3\.6466×10−06$3\.3921\\text\{\\times\}\{10\}^\{\-05\}$\\pm$3\.6466\\text\{\\times\}\{10\}^\{\-06\}$641012\.2739×10−04±2\.5986×10−05$2\.2739\\text\{\\times\}\{10\}^\{\-04\}$\\pm$2\.5986\\text\{\\times\}\{10\}^\{\-05\}$3\.5154×10−05±4\.1384×10−06$3\.5154\\text\{\\times\}\{10\}^\{\-05\}$\\pm$4\.1384\\text\{\\times\}\{10\}^\{\-06\}$642012\.3161×10−04±2\.3412×10−05$2\.3161\\text\{\\times\}\{10\}^\{\-04\}$\\pm$2\.3412\\text\{\\times\}\{10\}^\{\-05\}$3\.5122×10−05±4\.0520×10−06$3\.5122\\text\{\\times\}\{10\}^\{\-05\}$\\pm$4\.0520\\text\{\\times\}\{10\}^\{\-06\}$644012\.2763×10−04±1\.5265×10−05$2\.2763\\text\{\\times\}\{10\}^\{\-04\}$\\pm$1\.5265\\text\{\\times\}\{10\}^\{\-05\}$3\.4044×10−05±4\.3860×10−06$3\.4044\\text\{\\times\}\{10\}^\{\-05\}$\\pm$4\.3860\\text\{\\times\}\{10\}^\{\-06\}$648012\.3308×10−04±2\.2930×10−05$2\.3308\\text\{\\times\}\{10\}^\{\-04\}$\\pm$2\.2930\\text\{\\times\}\{10\}^\{\-05\}$3\.6565×10−05±5\.4556×10−06$3\.6565\\text\{\\times\}\{10\}^\{\-05\}$\\pm$5\.4556\\text\{\\times\}\{10\}^\{\-06\}$256512\.4231×10−04±1\.2736×10−05$2\.4231\\text\{\\times\}\{10\}^\{\-04\}$\\pm$1\.2736\\text\{\\times\}\{10\}^\{\-05\}$1\.8466×10−05±1\.5881×10−06$1\.8466\\text\{\\times\}\{10\}^\{\-05\}$\\pm$1\.5881\\text\{\\times\}\{10\}^\{\-06\}$2561012\.5664×10−04±1\.7316×10−05$2\.5664\\text\{\\times\}\{10\}^\{\-04\}$\\pm$1\.7316\\text\{\\times\}\{10\}^\{\-05\}$2\.0089×10−05±1\.9047×10−06$2\.0089\\text\{\\times\}\{10\}^\{\-05\}$\\pm$1\.9047\\text\{\\times\}\{10\}^\{\-06\}$2562012\.4479×10−04±2\.5320×10−05$2\.4479\\text\{\\times\}\{10\}^\{\-04\}$\\pm$2\.5320\\text\{\\times\}\{10\}^\{\-05\}$1\.9124×10−05±2\.1326×10−06$1\.9124\\text\{\\times\}\{10\}^\{\-05\}$\\pm$2\.1326\\text\{\\times\}\{10\}^\{\-06\}$2564012\.4944×10−04±1\.7857×10−05$2\.4944\\text\{\\times\}\{10\}^\{\-04\}$\\pm$1\.7857\\text\{\\times\}\{10\}^\{\-05\}$1\.9511×10−05±2\.1490×10−06$1\.9511\\text\{\\times\}\{10\}^\{\-05\}$\\pm$2\.1490\\text\{\\times\}\{10\}^\{\-06\}$2568012\.5230×10−04±2\.0580×10−05$2\.5230\\text\{\\times\}\{10\}^\{\-04\}$\\pm$2\.0580\\text\{\\times\}\{10\}^\{\-05\}$1\.9472×10−05±2\.4619×10−06$1\.9472\\text\{\\times\}\{10\}^\{\-05\}$\\pm$2\.4619\\text\{\\times\}\{10\}^\{\-06\}$Table 21:\(SPP model\) SPP model relative interaction kernel and environmental force errors for differentMM\(number of replicates\) andLL\(number of time samples\)\. Each entry reports the mean±\\pmsample standard deviation overTr=10T\_\{r\}=10independent learning trials\.MLEϕrelE\_\{\\phi\}^\{\\mathrm\{rel\}\}EfrelE\_\{f\}^\{\\mathrm\{rel\}\}1511\.1381×1000±1\.2043×1000$1\.1381\\text\{\\times\}\{10\}^\{00\}$\\pm$1\.2043\\text\{\\times\}\{10\}^\{00\}$2\.7564×1000±3\.3192×1000$2\.7564\\text\{\\times\}\{10\}^\{00\}$\\pm$3\.3192\\text\{\\times\}\{10\}^\{00\}$11011\.4278×1000±1\.7454×1000$1\.4278\\text\{\\times\}\{10\}^\{00\}$\\pm$1\.7454\\text\{\\times\}\{10\}^\{00\}$1\.3453×1000±1\.0674×1000$1\.3453\\text\{\\times\}\{10\}^\{00\}$\\pm$1\.0674\\text\{\\times\}\{10\}^\{00\}$12015\.3246×10−01±5\.6657×10−01$5\.3246\\text\{\\times\}\{10\}^\{\-01\}$\\pm$5\.6657\\text\{\\times\}\{10\}^\{\-01\}$2\.1987×1000±2\.1506×1000$2\.1987\\text\{\\times\}\{10\}^\{00\}$\\pm$2\.1506\\text\{\\times\}\{10\}^\{00\}$14016\.1445×10−01±5\.5824×10−01$6\.1445\\text\{\\times\}\{10\}^\{\-01\}$\\pm$5\.5824\\text\{\\times\}\{10\}^\{\-01\}$1\.0916×1000±1\.5755×1000$1\.0916\\text\{\\times\}\{10\}^\{00\}$\\pm$1\.5755\\text\{\\times\}\{10\}^\{00\}$18018\.0262×10−01±3\.6660×10−01$8\.0262\\text\{\\times\}\{10\}^\{\-01\}$\\pm$3\.6660\\text\{\\times\}\{10\}^\{\-01\}$1\.1428×1000±5\.8255×10−01$1\.1428\\text\{\\times\}\{10\}^\{00\}$\\pm$5\.8255\\text\{\\times\}\{10\}^\{\-01\}$4511\.3722×10−01±3\.5001×10−02$1\.3722\\text\{\\times\}\{10\}^\{\-01\}$\\pm$3\.5001\\text\{\\times\}\{10\}^\{\-02\}$1\.9606×1000±4\.5055×1000$1\.9606\\text\{\\times\}\{10\}^\{00\}$\\pm$4\.5055\\text\{\\times\}\{10\}^\{00\}$41011\.2154×10−01±2\.2052×10−02$1\.2154\\text\{\\times\}\{10\}^\{\-01\}$\\pm$2\.2052\\text\{\\times\}\{10\}^\{\-02\}$4\.1079×10−01±7\.3962×10−01$4\.1079\\text\{\\times\}\{10\}^\{\-01\}$\\pm$7\.3962\\text\{\\times\}\{10\}^\{\-01\}$42011\.2518×10−01±2\.2447×10−02$1\.2518\\text\{\\times\}\{10\}^\{\-01\}$\\pm$2\.2447\\text\{\\times\}\{10\}^\{\-02\}$4\.0368×10−01±7\.0996×10−01$4\.0368\\text\{\\times\}\{10\}^\{\-01\}$\\pm$7\.0996\\text\{\\times\}\{10\}^\{\-01\}$44011\.2693×10−01±3\.5854×10−02$1\.2693\\text\{\\times\}\{10\}^\{\-01\}$\\pm$3\.5854\\text\{\\times\}\{10\}^\{\-02\}$3\.9452×10−01±7\.8274×10−01$3\.9452\\text\{\\times\}\{10\}^\{\-01\}$\\pm$7\.8274\\text\{\\times\}\{10\}^\{\-01\}$48011\.1670×10−01±2\.5842×10−02$1\.1670\\text\{\\times\}\{10\}^\{\-01\}$\\pm$2\.5842\\text\{\\times\}\{10\}^\{\-02\}$2\.4197×10−01±2\.3849×10−01$2\.4197\\text\{\\times\}\{10\}^\{\-01\}$\\pm$2\.3849\\text\{\\times\}\{10\}^\{\-01\}$16511\.1221×10−01±8\.3761×10−03$1\.1221\\text\{\\times\}\{10\}^\{\-01\}$\\pm$8\.3761\\text\{\\times\}\{10\}^\{\-03\}$2\.9490×10−02±5\.8072×10−04$2\.9490\\text\{\\times\}\{10\}^\{\-02\}$\\pm$5\.8072\\text\{\\times\}\{10\}^\{\-04\}$161011\.1085×10−01±7\.1729×10−03$1\.1085\\text\{\\times\}\{10\}^\{\-01\}$\\pm$7\.1729\\text\{\\times\}\{10\}^\{\-03\}$2\.9833×10−02±1\.3988×10−03$2\.9833\\text\{\\times\}\{10\}^\{\-02\}$\\pm$1\.3988\\text\{\\times\}\{10\}^\{\-03\}$162011\.1122×10−01±6\.9548×10−03$1\.1122\\text\{\\times\}\{10\}^\{\-01\}$\\pm$6\.9548\\text\{\\times\}\{10\}^\{\-03\}$2\.9854×10−02±6\.3740×10−04$2\.9854\\text\{\\times\}\{10\}^\{\-02\}$\\pm$6\.3740\\text\{\\times\}\{10\}^\{\-04\}$164011\.0592×10−01±1\.1698×10−02$1\.0592\\text\{\\times\}\{10\}^\{\-01\}$\\pm$1\.1698\\text\{\\times\}\{10\}^\{\-02\}$2\.9147×10−02±5\.4161×10−04$2\.9147\\text\{\\times\}\{10\}^\{\-02\}$\\pm$5\.4161\\text\{\\times\}\{10\}^\{\-04\}$168011\.0950×10−01±7\.2616×10−03$1\.0950\\text\{\\times\}\{10\}^\{\-01\}$\\pm$7\.2616\\text\{\\times\}\{10\}^\{\-03\}$2\.9614×10−02±9\.9649×10−04$2\.9614\\text\{\\times\}\{10\}^\{\-02\}$\\pm$9\.9649\\text\{\\times\}\{10\}^\{\-04\}$64511\.1433×10−01±3\.7577×10−03$1\.1433\\text\{\\times\}\{10\}^\{\-01\}$\\pm$3\.7577\\text\{\\times\}\{10\}^\{\-03\}$2\.8864×10−02±2\.7296×10−04$2\.8864\\text\{\\times\}\{10\}^\{\-02\}$\\pm$2\.7296\\text\{\\times\}\{10\}^\{\-04\}$641011\.1151×10−01±4\.6203×10−03$1\.1151\\text\{\\times\}\{10\}^\{\-01\}$\\pm$4\.6203\\text\{\\times\}\{10\}^\{\-03\}$2\.9080×10−02±2\.3488×10−04$2\.9080\\text\{\\times\}\{10\}^\{\-02\}$\\pm$2\.3488\\text\{\\times\}\{10\}^\{\-04\}$642011\.1438×10−01±3\.0293×10−03$1\.1438\\text\{\\times\}\{10\}^\{\-01\}$\\pm$3\.0293\\text\{\\times\}\{10\}^\{\-03\}$2\.9111×10−02±2\.0774×10−04$2\.9111\\text\{\\times\}\{10\}^\{\-02\}$\\pm$2\.0774\\text\{\\times\}\{10\}^\{\-04\}$644011\.1024×10−01±4\.2098×10−03$1\.1024\\text\{\\times\}\{10\}^\{\-01\}$\\pm$4\.2098\\text\{\\times\}\{10\}^\{\-03\}$2\.9117×10−02±2\.4098×10−04$2\.9117\\text\{\\times\}\{10\}^\{\-02\}$\\pm$2\.4098\\text\{\\times\}\{10\}^\{\-04\}$648011\.1347×10−01±5\.2101×10−03$1\.1347\\text\{\\times\}\{10\}^\{\-01\}$\\pm$5\.2101\\text\{\\times\}\{10\}^\{\-03\}$2\.9282×10−02±2\.7405×10−04$2\.9282\\text\{\\times\}\{10\}^\{\-02\}$\\pm$2\.7405\\text\{\\times\}\{10\}^\{\-04\}$256511\.1533×10−01±2\.2326×10−03$1\.1533\\text\{\\times\}\{10\}^\{\-01\}$\\pm$2\.2326\\text\{\\times\}\{10\}^\{\-03\}$2\.9606×10−02±1\.3604×10−04$2\.9606\\text\{\\times\}\{10\}^\{\-02\}$\\pm$1\.3604\\text\{\\times\}\{10\}^\{\-04\}$2561011\.1210×10−01±2\.4077×10−03$1\.1210\\text\{\\times\}\{10\}^\{\-01\}$\\pm$2\.4077\\text\{\\times\}\{10\}^\{\-03\}$2\.9824×10−02±1\.2639×10−04$2\.9824\\text\{\\times\}\{10\}^\{\-02\}$\\pm$1\.2639\\text\{\\times\}\{10\}^\{\-04\}$2562011\.1398×10−01±2\.2106×10−03$1\.1398\\text\{\\times\}\{10\}^\{\-01\}$\\pm$2\.2106\\text\{\\times\}\{10\}^\{\-03\}$2\.9924×10−02±1\.2837×10−04$2\.9924\\text\{\\times\}\{10\}^\{\-02\}$\\pm$1\.2837\\text\{\\times\}\{10\}^\{\-04\}$2564011\.1318×10−01±1\.9296×10−03$1\.1318\\text\{\\times\}\{10\}^\{\-01\}$\\pm$1\.9296\\text\{\\times\}\{10\}^\{\-03\}$2\.9917×10−02±1\.1030×10−04$2\.9917\\text\{\\times\}\{10\}^\{\-02\}$\\pm$1\.1030\\text\{\\times\}\{10\}^\{\-04\}$2568011\.1401×10−01±1\.4986×10−03$1\.1401\\text\{\\times\}\{10\}^\{\-01\}$\\pm$1\.4986\\text\{\\times\}\{10\}^\{\-03\}$2\.9998×10−02±1\.7770×10−04$2\.9998\\text\{\\times\}\{10\}^\{\-02\}$\\pm$1\.7770\\text\{\\times\}\{10\}^\{\-04\}$ ### C\.2Noise robustness The following figures illustrate the robustness of the variational framework with respect to different observational noise levels for the models discussed in Section[5\.2](https://arxiv.org/html/2608.25181#S5.SS2)\. Figure[23](https://arxiv.org/html/2608.25181#A3.F23)visualizes the effect of noise on recovery of mechanisms, while Figure[24](https://arxiv.org/html/2608.25181#A3.F24)shows the corresponding impact on trajectory reconstruction and prediction\. \(a\)ϕ:\\phi:Kuramoto\(b\)ϕ:\\phi:SPP\(c\)ϕ:\\phi:phototaxis\(d\)f:f:Kuramoto\(e\)f:f:SPP \(f\)ff:phototaxis Figure 23:\(Noise robustness, mechanism recovery\) Each subfigure illustrates feature learning under multiplicative observational noise added to the training data at different noise levelsζ\\zeta, as described in Section[3\.5\.4](https://arxiv.org/html/2608.25181#S3.SS5.SSS4)\. Subfigures[23\(a\)](https://arxiv.org/html/2608.25181#A3.F23.sf1)–[23\(c\)](https://arxiv.org/html/2608.25181#A3.F23.sf3)show the learned interaction kernelsϕ\\phifor increasing values ofζ\\zeta\. Subfigures[23\(d\)](https://arxiv.org/html/2608.25181#A3.F23.sf4)–[23\(f\)](https://arxiv.org/html/2608.25181#A3.F23.sf6)show the corresponding learned environmental forceffunder the same noise levels\. For the SPP and phototaxis systems, the recovery offfis shown component\-wise, where the upper and lower rows correspond to the first and second components, respectively\.The background shading indicates the empirical training measure associated with each feature:ρ^R\\hat\{\\rho\}\_\{R\}for interaction kernels in all systems except Kuramoto, whereρ^Δθ\\hat\{\\rho\}\_\{\\Delta\\theta\}is used;ρ^V\\hat\{\\rho\}\_\{V\}for the environmental force in second\-order systems; andρ^θ\\hat\{\\rho\}\_\{\\theta\}for the environmental variable in the Kuramoto system\. We notice that, as the noise levelζ\\zetaincreases, the learned features gradually deviate from the true features, yet overall recovery is gnerally accurate and estimated features are minimally affected, primarily in regions with small induced measure \(i\.e\. a small amount of trajectory data\)\.          Figure 24:\(Noise robustness, trajectory recovery and prediction\) Each row represents the Kuramoto, Phototaxis, and SPP systems, respectively\. The subfigures show one of the observed trajectories before and after being perturbed by each noise levelζ\\zetadescribed in Section[3\.5\.4](https://arxiv.org/html/2608.25181#S3.SS5.SSS4), over the full time horizon\[0,Tf\]\[0,T\_\{f\}\]\. For the Kuramoto system, the solid blue line represents the true trajectory, the dashed line represents the predicted trajectory learned from the noisy observations, and the semitransparent gray dots denote the noisy trajectory used as the training data \(together with noisy observations of the derivative\)\. For the SPP and phototaxis systems, the solid blue lines represent the true trajectories, the semitransparent gray dotted lines represent the noisy trajectories used as the training data \(together with noisy observations of the derivative\), and the color\-gradient lines represent the predicted trajectories learned from the noisy observations, with yellow corresponding tot0t\_\{0\}and orange corresponding toTfT\_\{f\}\.We observe that the learned trajectories remain close to the true trajectories at lower noise levels\. As the noise level increases, gradual deviations in the predicted dynamics become apparent, particularly for the SPP system\. However, the overall trajectory recovery remains stable, indicating a degree of robustness of the learned models to observational noise\. ## Appendix DLimitations ### D\.1Effect of initial conditions on inferrence The distribution of the initial conditions,μ0\\mu\_\{0\}, has a significant impact on the performance of the learning algorithm\. In particular, training datasets whose initial conditions span a larger region of the state space provide more informative observations, resulting in more accurate recovery of the environmental force\. Conversely, when the initial\-condition distribution is narrowly concentrated, portions of the state space remain poorly sampled, leading to reduced learning accuracy\. Figure[25](https://arxiv.org/html/2608.25181#A4.F25)illustrates this effect by comparing two training datasets generated from different initial\-condition distributions\. \(a\)ϕ\\phi\(b\)ff Figure 25:\(Initial condition distribution, SPP system\) Comparison of the variational learning algorithm under two different initial condition distributionsμ0\\mu\_\{0\}\. Profile 1 \(blue\) is trained using a narrower initial\-condition distribution, while Profile 2 \(orange\) is trained using a distribution with larger variance\. The true kernel is shown in black\. Subfigure[25\(a\)](https://arxiv.org/html/2608.25181#A4.F25.sf1)compares the recovered interaction kernelϕ\\phi\. Both profiles recover the overall kernel accurately; however, Profile 2 fails to capture the narrow, steep peak near the origin\. Subfigure[25\(b\)](https://arxiv.org/html/2608.25181#A4.F25.sf2)compares the recovered environmental forceff\. Profile 1 provides a poor approximation, whereas Profile 2 accurately recovers the true environmental force, demonstrating that a broader initial condition distribution improves recovery of the environmental dynamics\. ### D\.2Basis functions The variational algorithm is generally sensitive to the number of basis functions used to represent both the interaction kernel and the environmental force\. Figure[26](https://arxiv.org/html/2608.25181#A4.F26)illustrates the effect of varying the number of B\-spline basis functions while fixing all other hyperparameters\. For each value of the basis dimension being varied, the reported error is averaged with resect to all tested values of the other basis dimension\. As expected, using too few basis functions leads to underfitting and consequently larger feature recovery errors\. Increasing the basis dimension improves the approximation quality and reduces the recovery error until the performance saturates\. Beyond this, additional basis functions provides minimal improvement and may even degrade performance, as the least\-squares equation \([37](https://arxiv.org/html/2608.25181#S2.E37)\) becomes increasingly ill\-conditioned for a fixed amount of training data\. For the SPP system considered here, both the interaction kernel and the environmental force achieve their lowest average recovery error when approximately1010basis functions are used\. This choice therefore provides a good balance between approximation accuracy and numerical stability\. \(a\)ϕ\\phi\(b\)ff Figure 26:Effect of the number of B\-spline basis functions on the recovery accuracy of the interaction kernel and environmental force for the SPP system\. In each panel, one basis dimension is varied while the recovery error is averaged over all tested values of the other basis dimension\. The solid curves represent the mean relative weightedL2L^\{2\}error, and the shaded regions indicate one standard deviation across the corresponding configurations\. Subfigure[26\(a\)](https://arxiv.org/html/2608.25181#A4.F26.sf1)shows the interaction kernel error,EϕrelE\_\{\\phi\}^\{\\mathrm\{rel\}\}, as the number of kernel basis functions,nϕn\_\{\\phi\}, varies\. Subfigure[26\(b\)](https://arxiv.org/html/2608.25181#A4.F26.sf2)shows the environmental\-force error,EfrelE\_\{f\}^\{\\mathrm\{rel\}\}, as the number of environmental force basis functions,nfn\_\{f\}, varies\. In both cases, the lowest average recovery error is achieved with approximately ten basis functions\.These results indicate that the choice of basis dimension is an important hyperparameter in the nonparametric learning framework\. Consequently, adaptive basis construction and automatic hyperparameter selection may further improve both learning accuracy and computational efficiency\. ## Appendix ETables of parameters ### E\.1Data generation parameters ParameterKuramotoPhototaxisSPPSPPCS2DSPPCSOPprofile\-\-clumpsmilling\-clumps\-System orderss1222221State dimensiondd1222221Number of agentsNN10101010101010training trajectoriesMtrainM\_\{\\mathrm\{train\}\}20205050205020testing trajectoriesMtestM\_\{\\mathrm\{test\}\}10101010101010Reference trajectoriesMρM\_\{\\rho\}2000200020002000200020002000Initial timet0t\_\{0\}000\.250\.25000000000training horizonTT0\.50\.50\.750\.750\.50\.5440\.50\.50\.50\.51010Prediction horizonTfT\_\{f\}222\.252\.2522202022222020Time stepdtdt0\.010\.010\.010\.010\.010\.010\.010\.010\.010\.010\.010\.010\.050\.05training snapshotsLfitL\_\{\\mathrm\{fit\}\}51515151515140140151515151201201Full snapshotsLL20120120120120120120012001201201201201401401Table 22:Data\-generation parameters used in the numerical experiments\. Training data are used to estimate the interaction kernels and environmental forces, testing data are used for trajectory validation, and theMρM\_\{\\rho\}data set is used to construct reference probability densities\. ### E\.2Learning algorithm parameters ParameterKuramotoPhototaxisSPPCS2DSPPCSOPInteraction typeEnergyAlignmentEnergyAlignmentEnergy\+\+AlignmentEnergyRadial basis functionsnϕn\_\{\\phi\}10810101050Energy basis functionsnϕEn\_\{\\phi^\{E\}\}–10–101050Alignment basis functionsnϕAn\_\{\\phi^\{A\}\}–10–101050Environment basis parameternbn\_\{b\}101010101010Environment basis size per statenfd=\(nb\)dn\_\{f^\{d\}\}=\(n\_\{b\}\)^\{d\}1010010010010010Environment basisnf=d∗nfdn\_\{f\}=d\*n\_\{f^\{d\}\}1020020020020010Table 23:Algorithm parameters utilized in the variational least\-squares framework\. The interaction kernels are approximated using radial B\-spline bases, while the environment force is represented using a tensor product of B\-splines\. For first\-order systems the environment variable is the state, whereas for second\-order systems it is generally the velocity\. ### E\.3System initial conditions SystemInitial conditionu0u\_\{0\}Kuramotoθ0\(m,i\)∼𝒰\(0,2π\)\\theta\_\{0\}^\{\(m,i\)\}\{\\sim\}\\mathcal\{U\}\(0,2\\pi\)Phototaxis\[x0\(m,i\)∼𝒰\(\[0,2\]2\),v0\(m,i\)∼μvPhottaxis\]\\Big\[x\_\{0\}^\{\(m,i\)\}\{\\sim\}\\mathcal\{U\}\(\[0,2\]^\{2\}\),\\;v\_\{0\}^\{\(m,i\)\}\\sim\\mu\_\{v\}^\{\\mathrm\{Phottaxis\}\}\\Big\]SPP\[x0\(m,i\)∼μxSPP,v0\(m,i\)∼μvSPP\]\\Big\[x\_\{0\}^\{\(m,i\)\}\\sim\\mu\_\{x\}^\{\\mathrm\{SPP\}\},\\;v\_\{0\}^\{\(m,i\)\}\\sim\\mu\_\{v\}^\{\\mathrm\{SPP\}\}\\Big\]CS2D\[x0\(m,i\)∼𝒰\(0\.1,10\),v0\(m,i\)∼𝒰\(0\.1,5\)\]\\Big\[\\,x\_\{0\}^\{\(m,i\)\}\\sim\\mathcal\{U\}\(0\.1,10\),v\_\{0\}^\{\(m,i\)\}\\sim\\mathcal\{U\}\(0\.1,5\)\\Big\]SPPCS\[x0\(m,i\)∼μxSPP,v0\(m,i\)∼μvSPP\]\\Big\[x\_\{0\}^\{\(m,i\)\}\\sim\\mu\_\{x\}^\{\\mathrm\{SPP\}\},\\;v\_\{0\}^\{\(m,i\)\}\\sim\\mu\_\{v\}^\{\\mathrm\{SPP\}\}\\Big\]OPx0\(m,i\)∼𝒰\(0,10\)x\_\{0\}^\{\(m,i\)\}\{\\sim\}\\mathcal\{U\}\(0,10\)Table 24:System\-specific initial conditions used in numerical experiments\. For the first\-order Kuramoto system,u0=θ0u\_\{0\}=\\theta\_\{0\}\. For second\-order systems,u0=\(x0,v0\)u\_\{0\}=\(x\_\{0\},v\_\{0\}\)\.μvPhototaxis\\mu\_\{v\}^\{\\mathrm\{Phototaxis\}\}denotes the perturbed tensor\-grid measure described in[E\.3](https://arxiv.org/html/2608.25181#A5.SS3)The spatial measuresμxSPP\\mu\_\{x\}^\{\\mathrm\{SPP\}\}andμxSPPCS\\mu\_\{x\}^\{\\mathrm\{SPPCS\}\}are cluster\-based sampling distributions\. The velocity measuresμvSPP\\mu\_\{v\}^\{\\mathrm\{SPP\}\}andμvSPPCS\\mu\_\{v\}^\{\\mathrm\{SPPCS\}\}are perturbed tensor\-grid distributions on\[−22,22\]2\[\-2\\sqrt\{2\},2\\sqrt\{2\}\]^\{2\}\.##### Phototaxis For the Phototaxis system, the velocity initial conditions are generated from a perturbed tensor grid on\[−1,1\]2\[\-1,1\]^\{2\}\. Specifically, a uniform tensor grid is constructed on the square\[−1,1\]2\[\-1,1\]^\{2\}, its points are randomly assigned to agents, and independent Gaussian perturbations are added\. This procedure produces approximately uniform coverage of the velocity domain and improves the conditioning of the tensor\-product basis used to learn the environmental forceff\. ##### SPP, SPPCS The spatial measureμxSPP\\mu\_\{x\}^\{\\mathrm\{SPP\}\}is a cluster\-based sampling distribution obtained from a mixture of uniformly populated spatial clumps\. The velocity measureμvSPP\\mu\_\{v\}^\{\\mathrm\{SPP\}\}is generated from a perturbed tensor grid on the square\[−2veq,2veq\]2\[\-2v\_\{\\mathrm\{eq\}\},2v\_\{\\mathrm\{eq\}\}\]^\{2\}, whereveq=α/βv\_\{\\mathrm\{eq\}\}=\\sqrt\{\\alpha/\\beta\}is the equilibrium speed of the self\-propulsion model\. This construction provides broad coverage of the velocity domain used for learning the environment forceff\. ### E\.4Estimates from semi\-parametric approach SystemComponentTermcoefftruecoeff\_\{\\mathrm\{true\}\}𝔼\(coeff^\)±σ\(coeff^\)\\mathbb\{E\}\(\\hat\{coeff\}\)\\pm\\sigma\(\\hat\{coeff\}\)Kuramoto11θ\\theta0\.010\.010\.009 999 999 999 999 9±2\.087 722 866 931 242 4×10−16$0\.009\\,999\\,999\\,999\\,999\\,9$\\pm$2\.087\\,722\\,866\\,931\\,242\\,4\\text\{\\times\}\{10\}^\{\-16\}$Phototaxis1111I0U∞eℓ\(1\)=−0\.12I\_\{0\}U\_\{\\infty\}e\_\{\\ell\}^\{\(1\)\}=\-\\frac\{0\.1\}\{\\sqrt\{2\}\}−0\.070 710 681 492 353 88±1\.423 730 942 286 253 9×10−07$\-0\.070\\,710\\,681\\,492\\,353\\,88$\\pm$1\.423\\,730\\,942\\,286\\,253\\,9\\text\{\\times\}\{10\}^\{\-07\}$v1v\_\{1\}I0=1\.0I\_\{0\}=1\.01\.000 000 319 453 960 8±3\.538 386 027 474 474×10−06$1\.000\\,000\\,319\\,453\\,960\\,8$\\pm$3\.538\\,386\\,027\\,474\\,474\\text\{\\times\}\{10\}^\{\-06\}$v2v\_\{2\}00−2\.967 810 805 171 508×10−07±2\.676 347 740 402 615×10−06$\-2\.967\\,810\\,805\\,171\\,508\\text\{\\times\}\{10\}^\{\-07\}$\\pm$2\.676\\,347\\,740\\,402\\,615\\text\{\\times\}\{10\}^\{\-06\}$2211I0U∞eℓ\(2\)=0\.12I\_\{0\}U\_\{\\infty\}e\_\{\\ell\}^\{\(2\)\}=\\frac\{0\.1\}\{\\sqrt\{2\}\}0\.070 710 663 656 462 13±5\.863 941 042 682 641×10−08$0\.070\\,710\\,663\\,656\\,462\\,13$\\pm$5\.863\\,941\\,042\\,682\\,641\\text\{\\times\}\{10\}^\{\-08\}$v1v\_\{1\}00−4\.931 084 525 950 616×10−07±2\.709 497 747 258 749×10−06$\-4\.931\\,084\\,525\\,950\\,616\\text\{\\times\}\{10\}^\{\-07\}$\\pm$2\.709\\,497\\,747\\,258\\,749\\text\{\\times\}\{10\}^\{\-06\}$v2v\_\{2\}I0=1\.0I\_\{0\}=1\.00\.999 999 968 739 424±4\.026 680 654 827 744 6×10−06$0\.999\\,999\\,968\\,739\\,424$\\pm$4\.026\\,680\\,654\\,827\\,744\\,6\\text\{\\times\}\{10\}^\{\-06\}$SPP11v1v\_\{1\}α=1\\alpha=11\.000 232 739 385 091 6±0\.000 794 694 792 782 155 5$1\.000\\,232\\,739\\,385\\,091\\,6$\\pm$0\.000\\,794\\,694\\,792\\,782\\,155\\,5$v2v\_\{2\}000\.000 568 144 895 427 883±0\.001 043 046 722 979 779$0\.000\\,568\\,144\\,895\\,427\\,883$\\pm$0\.001\\,043\\,046\\,722\\,979\\,779$‖v‖2v1\\\|v\\\|^\{2\}v\_\{1\}β=0\.5\\beta=0\.50\.499 996 394 287 417 54±0\.000 113 508 445 827 097 8$0\.499\\,996\\,394\\,287\\,417\\,54$\\pm$0\.000\\,113\\,508\\,445\\,827\\,097\\,8$‖v‖2v2\\\|v\\\|^\{2\}v\_\{2\}00−8\.231 290 447 634 012×10−05±0\.000 177 452 175 135 829 3$\-8\.231\\,290\\,447\\,634\\,012\\text\{\\times\}\{10\}^\{\-05\}$\\pm$0\.000\\,177\\,452\\,175\\,135\\,829\\,3$22v1v\_\{1\}006\.237 543 358 736 741×10−05±0\.001 619 508 850 430 651 4$6\.237\\,543\\,358\\,736\\,741\\text\{\\times\}\{10\}^\{\-05\}$\\pm$0\.001\\,619\\,508\\,850\\,430\\,651\\,4$v2v\_\{2\}α=1\\alpha=10\.999 745 922 986 700 8±0\.001 191 278 938 192 407$0\.999\\,745\\,922\\,986\\,700\\,8$\\pm$0\.001\\,191\\,278\\,938\\,192\\,407$‖v‖2v1\\\|v\\\|^\{2\}v\_\{1\}004\.227 453 433 844 442 4×10−05±0\.000 305 343 387 616 725 63$4\.227\\,453\\,433\\,844\\,442\\,4\\text\{\\times\}\{10\}^\{\-05\}$\\pm$0\.000\\,305\\,343\\,387\\,616\\,725\\,63$‖v‖2v2\\\|v\\\|^\{2\}v\_\{2\}β=0\.5\\beta=0\.50\.499 938 858 006 321 65±0\.000 197 572 613 434 238 76$0\.499\\,938\\,858\\,006\\,321\\,65$\\pm$0\.000\\,197\\,572\\,613\\,434\\,238\\,76$ Table 25:Recovery of the parametric coefficients for the candidate basis functions\. For each coefficient, we report the mean and standard deviation overTr=10T\_\{r\}=10independent trials\. Coefficients corresponding to non\-present mechanisms are expected to be close to zero\. ### E\.5Index of notation ### E\.6Index of notation Table 26:Parameters and their definitions\.ParameterDefinitionDynamical systemNNNumber of agents in the systemddDimension of the state space of each agentxi\(t\)∈ℝdx\_\{i\}\(t\)\\in\\mathbb\{R\}^\{d\}State of theii\-th agent at timettx\(t\)∈ℝdNx\(t\)\\in\\mathbb\{R\}^\{dN\}Full stacked state of the system at timettf:ℝd→ℝdf:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{d\}Environmental force function \(assumed identical for all agents\)Ω⊆ℝd\\Omega\\subseteq\\mathbb\{R\}^\{d\}Proper subset ofℝd\\mathbb\{R\}^\{d\}serving as the domain offfin applications whereffis not defined on all ofℝd\\mathbb\{R\}^\{d\}ϕ:ℝ≥0→ℝ\\phi:\\mathbb\{R\}\_\{\\geq 0\}\\to\\mathbb\{R\}Interaction kernel encoding pairwise interaction strength as a function of distanceFf\(x\)∈ℝdNF\_\{f\}\(x\)\\in\\mathbb\{R\}^\{dN\}Stacked environmental force:\(f\(x1\),…,f\(xN\)\)\(f\(x\_\{1\}\),\\ldots,f\(x\_\{N\}\)\)Fϕ\(x\)∈ℝdNF\_\{\\phi\}\(x\)\\in\\mathbb\{R\}^\{dN\}Stacked interaction force:1N∑j≠iϕ\(\|xj−xi\|\)\(xj−xi\)\\frac\{1\}\{N\}\\sum\_\{j\\neq i\}\\phi\(\|x\_\{j\}\-x\_\{i\}\|\)\(x\_\{j\}\-x\_\{i\}\)for each agentiivi\(t\)∈ℝdv\_\{i\}\(t\)\\in\\mathbb\{R\}^\{d\}Velocity of theii\-th agent at timett\(second\-order systems\)v\(t\)∈ℝdNv\(t\)\\in\\mathbb\{R\}^\{dN\}Full stacked velocity of the system at timett:\(v1\(t\),…,vN\(t\)\)\(v\_\{1\}\(t\),\\ldots,v\_\{N\}\(t\)\)Ff\(x,v\)∈ℝdNF\_\{f\}\(x,v\)\\in\\mathbb\{R\}^\{dN\}Stacked environmental force for second\-order systems:\(f\(x1,v1\),…,f\(xN,vN\)\)\(f\(x\_\{1\},v\_\{1\}\),\\ldots,f\(x\_\{N\},v\_\{N\}\)\), reducing toFf\(v\)F\_\{f\}\(v\)whenffhas no spatial dependenceℱ\(X,Y\)\\mathcal\{F\}\(X,Y\)Set of all functions from setXXto setYY\(used to specify thatℋf,ℋϕ\\mathcal\{H\}\_\{f\},\\mathcal\{H\}\_\{\\phi\}are finite\-dimensional subspaces of function spaces\)Trajectory dataμ0\\mu\_\{0\}Distribution onℝdN\\mathbb\{R\}^\{dN\}from which i\.i\.d\. initial conditions are sampledMMNumber of trajectory replicates \(initial condition samples\)LLNumber of discrete observation time points\{tℓ\}ℓ=1L\\\{t\_\{\\ell\}\\\}\_\{\\ell=1\}^\{L\},0=t1<⋯<tL=T0=t\_\{1\}<\\cdots<t\_\{L\}=TDiscrete observation times over the training interval\[0,T\]\[0,T\]Δtℓ:=tℓ\+1−tℓ\\Delta t\_\{\\ell\}:=t\_\{\\ell\+1\}\-t\_\{\\ell\}Time step between consecutive observationsX0\(m\)∈ℝdNX^\{\(m\)\}\_\{0\}\\in\\mathbb\{R\}^\{dN\}Initial condition of themm\-th replicateV0\(m\)∈ℝdNV^\{\(m\)\}\_\{0\}\\in\\mathbb\{R\}^\{dN\}Initial velocity condition of themm\-th replicate \(second\-order systems\)Xℓ\(m\)∈ℝdNX^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{dN\}Observed state of themm\-th replicate at timetℓt\_\{\\ell\}Vℓ\(m\)∈ℝdNV^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{dN\}Observed \(or approximated\) velocity of themm\-th replicate at timetℓt\_\{\\ell\}Aℓ\(m\)∈ℝdNA^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{dN\}Observed \(or approximated\) acceleration of themm\-th replicate at timetℓt\_\{\\ell\}\(second\-order systems only\)\(Rℓ\(m\)\)ij∈ℝ\(R^\{\(m\)\}\_\{\\ell\}\)\_\{ij\}\\in\\mathbb\{R\}Pairwise distance\|\(Xℓ\(m\)\)i−\(Xℓ\(m\)\)j\|\|\(X^\{\(m\)\}\_\{\\ell\}\)\_\{i\}\-\(X^\{\(m\)\}\_\{\\ell\}\)\_\{j\}\|between agentsiiandjjat observation\(m,ℓ\)\(m,\\ell\);Rℓ\(m\)∈ℝN×NR^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{N\\times N\}ℋ\\mathcal\{H\}Low\-pass filter applied to trajectory data for velocity \(or acceleration\) estimationInner product and norm onℝdN\\mathbb\{R\}^\{dN\}⟨X,Y⟩ℝdN\\langle X,Y\\rangle\_\{\\mathbb\{R\}^\{dN\}\}1N∑i=1N⟨Xi,Yi⟩\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\langle X\_\{i\},Y\_\{i\}\\rangle, theNN\-agent\-normalized inner product onℝdN\\mathbb\{R\}^\{dN\}∥⋅∥ℝdN\\\|\\cdot\\\|\_\{\\mathbb\{R\}^\{dN\}\}Norm onℝdN\\mathbb\{R\}^\{dN\}induced by⟨⋅,⋅⟩ℝdN\\langle\\cdot,\\cdot\\rangle\_\{\\mathbb\{R\}^\{dN\}\}Hypothesis spaces and basesℋf⊆ℱ\(ℝd,ℝd\)\\mathcal\{H\}\_\{f\}\\subseteq\\mathcal\{F\}\(\\mathbb\{R\}^\{d\},\\mathbb\{R\}^\{d\}\)Finite\-dimensional hypothesis space for the environmental forceffℋϕ⊆ℱ\(ℝ≥0,ℝ\)\\mathcal\{H\}\_\{\\phi\}\\subseteq\\mathcal\{F\}\(\\mathbb\{R\}\_\{\\geq 0\},\\mathbb\{R\}\)Finite\-dimensional hypothesis space for the interaction kernelϕ\\phinfn\_\{f\}Dimension ofℋf\\mathcal\{H\}\_\{f\}nϕn\_\{\\phi\}Dimension ofℋϕ\\mathcal\{H\}\_\{\\phi\}\(hk\)k=1nf\(h\_\{k\}\)\_\{k=1\}^\{n\_\{f\}\}Ordered basis forℋf\\mathcal\{H\}\_\{f\}; eachhk=\(hk1,…,hkd\)h\_\{k\}=\(h\_\{k\}^\{1\},\\ldots,h\_\{k\}^\{d\}\)isℝd\\mathbb\{R\}^\{d\}\-valued\(ψk\)k=1nϕ\(\\psi\_\{k\}\)\_\{k=1\}^\{n\_\{\\phi\}\}Ordered basis forℋϕ\\mathcal\{H\}\_\{\\phi\}α∈ℝnf\\alpha\\in\\mathbb\{R\}^\{n\_\{f\}\}Coefficient vector forf^\\hat\{f\}in the basis\(hk\)\(h\_\{k\}\)β∈ℝnϕ\\beta\\in\\mathbb\{R\}^\{n\_\{\\phi\}\}Coefficient vector forϕ^\\hat\{\\phi\}in the basis\(ψk\)\(\\psi\_\{k\}\)θ=\(α,β\)∈ℝnf\+nϕ\\theta=\(\\alpha,\\beta\)\\in\\mathbb\{R\}^\{n\_\{f\}\+n\_\{\\phi\}\}Combined coefficient vector for bothf^\\hat\{f\}andϕ^\\hat\{\\phi\}f~∈ℋf,ϕ~∈ℋϕ\\tilde\{f\}\\in\\mathcal\{H\}\_\{f\},\\ \\tilde\{\\phi\}\\in\\mathcal\{H\}\_\{\\phi\}An arbitrary \(not necessarily optimal\) candidate pair of environmental force and interaction kernel used to define the error functionalℰℋf,ℋϕ\\mathcal\{E\}\_\{\\mathcal\{H\}\_\{f\},\\mathcal\{H\}\_\{\\phi\}\}and the evaluation metrics of Section[3](https://arxiv.org/html/2608.25181#S3); contrasted with the fitted estimatorsf^,ϕ^\\hat\{f\},\\hat\{\\phi\}f^∈ℋf\\hat\{f\}\\in\\mathcal\{H\}\_\{f\}Estimator of the environmental forceff, defined as∑kα^khk\\sum\_\{k\}\\hat\{\\alpha\}\_\{k\}h\_\{k\}ϕ^∈ℋϕ\\hat\{\\phi\}\\in\\mathcal\{H\}\_\{\\phi\}Estimator of the interaction kernelϕ\\phi, defined as∑kβ^kψk\\sum\_\{k\}\\hat\{\\beta\}\_\{k\}\\psi\_\{k\}Matrices and normal equationsHℓ\(m\)∈ℝdN×nfH^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{dN\\times n\_\{f\}\}Matrix\(h1\(Xℓ\(m\)\)⋯hnf\(Xℓ\(m\)\)\)\\bigl\(h\_\{1\}\(X^\{\(m\)\}\_\{\\ell\}\)\\;\\cdots\\;h\_\{n\_\{f\}\}\(X^\{\(m\)\}\_\{\\ell\}\)\\bigr\)of stackedff\-basis termsHℓ\(m\)H^\{\(m\)\}\_\{\\ell\}\(second\-order case\)∈ℝdN×nf\\in\\mathbb\{R\}^\{dN\\times n\_\{f\}\}Matrix\(h1\(Xℓ\(m\),Vℓ\(m\)\)⋯hnf\(Xℓ\(m\),Vℓ\(m\)\)\)\\bigl\(h\_\{1\}\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)\\;\\cdots\\;h\_\{n\_\{f\}\}\(X^\{\(m\)\}\_\{\\ell\},V^\{\(m\)\}\_\{\\ell\}\)\\bigr\)of stackedff\-basis terms evaluated on both state and velocityΨℓ\(m\)∈ℝdN×nϕ\\Psi^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{dN\\times n\_\{\\phi\}\}Matrix\(Fψ1\(Xℓ\(m\)\)⋯Fψnϕ\(Xℓ\(m\)\)\)\\bigl\(F\_\{\\psi\_\{1\}\}\(X^\{\(m\)\}\_\{\\ell\}\)\\;\\cdots\\;F\_\{\\psi\_\{n\_\{\\phi\}\}\}\(X^\{\(m\)\}\_\{\\ell\}\)\\bigr\)of stacked interaction\-basis termsCℓ\(m\)∈ℝdN×\(nf\+nϕ\)C^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{dN\\times\(n\_\{f\}\+n\_\{\\phi\}\)\}Concatenated matrix\(Hℓ\(m\)Ψℓ\(m\)\)\(H^\{\(m\)\}\_\{\\ell\}\\;\\;\\Psi^\{\(m\)\}\_\{\\ell\}\)ℰ\(θ\)\\mathcal\{E\}\(\\theta\)Least\-squares error functional1ML∑m,ℓ‖Vℓ\(m\)−Cℓ\(m\)θ‖ℝdN2\\frac\{1\}\{ML\}\\sum\_\{m,\\ell\}\\\|V^\{\(m\)\}\_\{\\ell\}\-C^\{\(m\)\}\_\{\\ell\}\\theta\\\|\_\{\\mathbb\{R\}^\{dN\}\}^\{2\}ℰℋf,ℋϕ\(f^,ϕ^\)\\mathcal\{E\}\_\{\\mathcal\{H\}\_\{f\},\\mathcal\{H\}\_\{\\phi\}\}\(\\hat\{f\},\\hat\{\\phi\}\)Second\-order error functional over hypothesis spaces, equal toℰ\(α,β\)\\mathcal\{E\}\(\\alpha,\\beta\)oncef^,ϕ^\\hat\{f\},\\hat\{\\phi\}are expanded in their bases \(same object asℰ\(θ\)\\mathcal\{E\}\(\\theta\), withθ=\(α,β\)\\theta=\(\\alpha,\\beta\)\)θ^=\(α^,β^\)\\hat\{\\theta\}=\(\\hat\{\\alpha\},\\hat\{\\beta\}\)Minimizer ofℰ\(θ\)\\mathcal\{E\}\(\\theta\); estimated coefficient vectorA∈ℝ\(nf\+nϕ\)×\(nf\+nϕ\)A\\in\\mathbb\{R\}^\{\(n\_\{f\}\+n\_\{\\phi\}\)\\times\(n\_\{f\}\+n\_\{\\phi\}\)\}Normal equation matrix∑m,ℓ\(Cℓ\(m\)\)TCℓ\(m\)\\displaystyle\\sum\_\{m,\\ell\}\(C^\{\(m\)\}\_\{\\ell\}\)^\{T\}C^\{\(m\)\}\_\{\\ell\}b∈ℝnf\+nϕb\\in\\mathbb\{R\}^\{n\_\{f\}\+n\_\{\\phi\}\}Normal equation right\-hand side∑m,ℓ\(Cℓ\(m\)\)TVℓ\(m\)\\displaystyle\\sum\_\{m,\\ell\}\(C^\{\(m\)\}\_\{\\ell\}\)^\{T\}V^\{\(m\)\}\_\{\\ell\}Hypothesis space constructionRmin,RmaxR\_\{\\min\},\\,R\_\{\\max\}Min/max observed pairwise distances; define the effective domain\[Rmin,Rmax\]\[R\_\{\\min\},R\_\{\\max\}\]ofℋϕ\\mathcal\{H\}\_\{\\phi\}PPNumber of sub\-intervals in the data\-dependent partition of\[Rmin,Rmax\]\[R\_\{\\min\},R\_\{\\max\}\]used to buildℋϕ\\mathcal\{H\}\_\{\\phi\}Ip:=\[rp,rp\+1\)I\_\{p\}:=\[r\_\{p\},r\_\{p\+1\}\)pp\-th partition element \(sub\-interval\) of\[Rmin,Rmax\]\[R\_\{\\min\},R\_\{\\max\}\],p=1,…,Pp=1,\\ldots,PnPn\_\{P\}Number of localized basis functions \(e\.g\. degree\-≤2\\leq 2polynomials or B\-splines\) supported on each partition element;nϕ=P×nPn\_\{\\phi\}=P\\times n\_\{P\}Zmin,Zmax∈ℝdZ\_\{\\min\},\\,Z\_\{\\max\}\\in\\mathbb\{R\}^\{d\}Component\-wise min/max observed state values; define the hyper\-rectangular partition domain ofℋf\\mathcal\{H\}\_\{f\}Xmin,Xmax∈ℝdX\_\{\\min\},\\,X\_\{\\max\}\\in\\mathbb\{R\}^\{d\}Component\-wise min/max observed state values \(second\-order notation forZmin,ZmaxZ\_\{\\min\},Z\_\{\\max\}restricted to the state variable\)Vmin,Vmax∈ℝdV\_\{\\min\},\\,V\_\{\\max\}\\in\\mathbb\{R\}^\{d\}Component\-wise min/max observed velocity values; together withXmin,XmaxX\_\{\\min\},X\_\{\\max\}define the domain ofℋf\\mathcal\{H\}\_\{f\}for second\-order systemsXk,min,Xk,max∈ℝX\_\{k,\\min\},\\,X\_\{k,\\max\}\\in\\mathbb\{R\}Min/max of thekk\-th component of the state variable,k=1,…,dk=1,\\ldots,dVk,min,Vk,max∈ℝV\_\{k,\\min\},\\,V\_\{k,\\max\}\\in\\mathbb\{R\}Min/max of thekk\-th component of the velocity variable,k=1,…,dk=1,\\ldots,dhkj:ℝd→ℝh\_\{k\}^\{j\}:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}Scalar\-valued localized component function \(j=1,…,dj=1,\\ldots,d\) used to build theℝd\\mathbb\{R\}^\{d\}\-valued basis functionhk=\(hk1,…,hkd\)h\_\{k\}=\(h\_\{k\}^\{1\},\\ldots,h\_\{k\}^\{d\}\)ofℋf\\mathcal\{H\}\_\{f\}, via tensor product with the standard basis ofℝd\\mathbb\{R\}^\{d\}ψkE,ψkA\\psi\_\{k\}^\{E\},\\ \\psi\_\{k\}^\{A\}Localized basis functions for the energy\-based kernelϕE\\phi^\{E\}and alignment\-based kernelϕA\\phi^\{A\}, respectively, used in the model\-selection algorithm \(playing the role of the genericψk\\psi\_\{k\}for each kernel type\)Computational complexityDDDimension of the stacked state space:D=dND=dN\(first\-order systems\),D=2dND=2dN\(second\-order systems\)n:=nf\+nϕn:=n\_\{f\}\+n\_\{\\phi\}Total number of basis functions acrossℋf\\mathcal\{H\}\_\{f\}andℋϕ\\mathcal\{H\}\_\{\\phi\}Semi\-parametric approachp∈ℝnpp\\in\\mathbb\{R\}^\{n\_\{p\}\}Parameter vector fully specifying the prescribed functional form off\(x,p\)f\(x;p\)npn\_\{p\}Number of unknown parameters in the parametric environmental forceηk:ℝd→ℝd\\eta\_\{k\}:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{d\},k=1,…,npk=1,\\dots,n\_\{p\}Known feature functions whenffis linear inpp, i\.e\.f\(x,p\)=∑kpkηk\(x\)f\(x;p\)=\\sum\_\{k\}p\_\{k\}\\,\\eta\_\{k\}\(x\)Hℓ\(m\)H^\{\(m\)\}\_\{\\ell\}\(semi\-parametric case\)∈ℝdN×np\\in\\mathbb\{R\}^\{dN\\times n\_\{p\}\}Matrix of stacked parametric feature termsηk\(Xℓ\(m\)\)\\eta\_\{k\}\(X^\{\(m\)\}\_\{\\ell\}\), constructed analogously to the non\-parametricHℓ\(m\)H^\{\(m\)\}\_\{\\ell\}with basis\(ηk\)k=1np\(\\eta\_\{k\}\)\_\{k=1\}^\{n\_\{p\}\}in place of\(hk\)k=1nf\(h\_\{k\}\)\_\{k=1\}^\{n\_\{f\}\}Induced measures on dataZZInput variable offf:Z=XZ=X\(first\-order systems\),Z=\(X,V\)Z=\(X,V\)\(second\-order systems\)ω\\omegaSample outcome indexing a random drawx0\(ω\)∼μ0x\_\{0\}\(\\omega\)\\sim\\mu\_\{0\}of the initial conditionrij\(t,ω\)r\_\{ij\}\(t,\\omega\)Pairwise distance trajectory\|xj\(t,x0\(ω\)\)−xi\(t,x0\(ω\)\)\|\|x\_\{j\}\(t,x\_\{0\}\(\\omega\)\)\-x\_\{i\}\(t,x\_\{0\}\(\\omega\)\)\|between agentsi,ji,jA⊆ℝ≥0A\\subseteq\\mathbb\{R\}\_\{\\geq 0\}Borel set used to define the pairwise\-distance measureρR\(A\)\\rho\_\{R\}\(A\)B⊆ℝdB\\subseteq\\mathbb\{R\}^\{d\}Borel set used to define the state measureρX\(B\)\\rho\_\{X\}\(B\)ρR\\rho\_\{R\}Continuous\-time expected measure of observed pairwise distances over\[0,T\]\[0,T\]andμ0\\mu\_\{0\}ρ^R\\hat\{\\rho\}\_\{R\}Empirical approximation ofρR\\rho\_\{R\}from the observed distances\(Rℓ\(m\)\)m,ℓ\(R^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell\}ρZ\\rho\_\{Z\}Continuous\-time expected measure of the state variableZZover\[0,T\]\[0,T\]andμ0\\mu\_\{0\}ρ^Z\\hat\{\\rho\}\_\{Z\}Empirical approximation ofρZ\\rho\_\{Z\}from the observed data\(Xℓ\(m\)\)m,ℓ\(X^\{\(m\)\}\_\{\\ell\}\)\_\{m,\\ell\}Δt:=T/L\\Delta t:=T/LUniform discretization step of the training interval\[0,T\]\[0,T\]used when forming the empirical measuresρ^R,ρ^X\\hat\{\\rho\}\_\{R\},\\hat\{\\rho\}\_\{X\}\(distinct from the per\-stepΔtℓ\\Delta t\_\{\\ell\}used elsewhere\)MρM\_\{\\rho\}Number of replicates used to construct accurate empirical approximations ofρR\\rho\_\{R\}andρZ\\rho\_\{Z\}‖ψ‖L2\(ρR\)\\\|\\psi\\\|\_\{L^\{2\}\(\\rho\_\{R\}\)\},‖h‖L2\(ρX\)\\\|h\\\|\_\{L^\{2\}\(\\rho\_\{X\}\)\}WeightedL2L^\{2\}norms \(with weightr2r^\{2\}forρR\\rho\_\{R\}\) used to measure feature\-recovery errorFeature recovery metricsEϕ\(ϕ~\)E\_\{\\phi\}\(\\tilde\{\\phi\}\)Interaction kernel error‖ϕ~−ϕ‖L2\(ρR\)\\\|\\tilde\{\\phi\}\-\\phi\\\|\_\{L^\{2\}\(\\rho\_\{R\}\)\}\(weighted byr2r^\{2\}\)Ef\(f~\)E\_\{f\}\(\\tilde\{f\}\)Environmental force error‖f~−f‖L2\(ρZ\)\\\|\\tilde\{f\}\-f\\\|\_\{L^\{2\}\(\\rho\_\{Z\}\)\}Eϕrel,EfrelE\_\{\\phi\}^\{\\mathrm\{rel\}\},\\,E\_\{f\}^\{\\mathrm\{rel\}\}Relative versions ofEϕE\_\{\\phi\}andEfE\_\{f\}, normalized by‖ϕ‖L2\(ρR\)\\\|\\phi\\\|\_\{L^\{2\}\(\\rho\_\{R\}\)\}and‖f‖L2\(ρZ\)\\\|f\\\|\_\{L^\{2\}\(\\rho\_\{Z\}\)\}respectivelyResidual errorSresS\_\{\\mathrm\{res\}\}RMS of the observed velocity \(first\-order\) or acceleration \(second\-order\) data; normalization scaleEres\(f~,ϕ~\)E\_\{\\mathrm\{res\}\}\(\\tilde\{f\},\\tilde\{\\phi\}\)Residual error: RMS discrepancy between observed and predicted velocities \(or accelerations\)EresrelE\_\{\\mathrm\{res\}\}^\{\\mathrm\{rel\}\}Relative residual errorEres/SresE\_\{\\mathrm\{res\}\}/S\_\{\\mathrm\{res\}\}Eres,ϵrelE\_\{\\mathrm\{res\},\\epsilon\}^\{\\mathrm\{rel\}\}Regularized relative residual errorEres/\(Sres\+ϵ\)E\_\{\\mathrm\{res\}\}/\(S\_\{\\mathrm\{res\}\}\+\\epsilon\), used in the model\-selection compatibility gateTrajectory metrics\[0,T\]\[0,T\]Training \(fitting\) time horizon\(T,Tf\]\(T,T\_\{f\}\]Prediction \(extrapolation\) time horizon beyond the training windowx~\(t\)∈ℝdN\\tilde\{x\}\(t\)\\in\\mathbb\{R\}^\{dN\}Continuous\-time solution of the IVP driven by estimatorsf~,ϕ~\\tilde\{f\},\\tilde\{\\phi\}\(i\.e\.x~˙=Ff~\(x~\)\+Fϕ~\(x~\)\\dot\{\\tilde\{x\}\}=F\_\{\\tilde\{f\}\}\(\\tilde\{x\}\)\+F\_\{\\tilde\{\\phi\}\}\(\\tilde\{x\}\),x~\(0\)=X0\(m\)\\tilde\{x\}\(0\)=X\_\{0\}^\{\(m\)\}\), of whichX~ℓ\(m\)\\tilde\{X\}^\{\(m\)\}\_\{\\ell\}is the value at timetℓt\_\{\\ell\}X~ℓ\(m\)∈ℝdN\\tilde\{X\}^\{\(m\)\}\_\{\\ell\}\\in\\mathbb\{R\}^\{dN\}State of replicatemmat timetℓt\_\{\\ell\}simulated from the learned model\(f^,ϕ^\)\(\\hat\{f\},\\hat\{\\phi\}\)eℓ\(m\)\(f~,ϕ~\)e^\{\(m\)\}\_\{\\ell\}\(\\tilde\{f\},\\tilde\{\\phi\}\)Pointwise trajectory error‖Xℓ\(m\)−X~ℓ\(m\)‖ℝdN\\\|X^\{\(m\)\}\_\{\\ell\}\-\\tilde\{X\}^\{\(m\)\}\_\{\\ell\}\\\|\_\{\\mathbb\{R\}^\{dN\}\}of replicatemmat timetℓt\_\{\\ell\}Erecon\(m\)E\_\{\\mathrm\{recon\}\}^\{\(m\)\}Reconstruction error:maxtℓ∈\[0,T\]eℓ\(m\)\\max\_\{t\_\{\\ell\}\\in\[0,T\]\}e^\{\(m\)\}\_\{\\ell\}Epred\(m\)E\_\{\\mathrm\{pred\}\}^\{\(m\)\}Prediction error:maxtℓ∈\(T,Tf\]eℓ\(m\)\\max\_\{t\_\{\\ell\}\\in\(T,T\_\{f\}\]\}e^\{\(m\)\}\_\{\\ell\}Etraj\(m\)E\_\{\\mathrm\{traj\}\}^\{\(m\)\}Full trajectory error:maxtℓ∈\[0,Tf\]eℓ\(m\)\\max\_\{t\_\{\\ell\}\\in\[0,T\_\{f\}\]\}e^\{\(m\)\}\_\{\\ell\}E¯recon,E¯pred,E¯traj\\bar\{E\}\_\{\\mathrm\{recon\}\},\\bar\{E\}\_\{\\mathrm\{pred\}\},\\bar\{E\}\_\{\\mathrm\{traj\}\}Sample\-mean reconstruction, prediction, and full trajectory errors overMMreplicatesσErecon,σEpred,σEtraj\\sigma\_\{E\_\{\\mathrm\{recon\}\}\},\\sigma\_\{E\_\{\\mathrm\{pred\}\}\},\\sigma\_\{E\_\{\\mathrm\{traj\}\}\}Sample standard deviations of the corresponding trajectory\-error metricsE∗RMSE\_\{\*\}^\{\\mathrm\{RMS\}\}RMS trajectory error over the time interval indexed by∗∈\{recon,pred,traj\}\{\*\}\\in\\\{\\mathrm\{recon\},\\mathrm\{pred\},\\mathrm\{traj\}\\\}SXS\_\{X\}RMS of observed trajectory data; normalization scale for relative trajectory errorsE∗relE\_\{\*\}^\{\\mathrm\{rel\}\}Relative RMS trajectory errorE∗RMS/\(SX\+ϵ\)E\_\{\*\}^\{\\mathrm\{RMS\}\}/\(S\_\{X\}\+\\epsilon\)Consistency and robustness assessmentTrT\_\{r\}Number of independent repeated trials used to assess reliabilityηi,ℓ\(m\),η¯i,ℓ\(m\)∼i\.i\.d\.Unif\[−ζ,ζ\]\\eta\_\{i,\\ell\}^\{\(m\)\},\\bar\{\\eta\}\_\{i,\\ell\}^\{\(m\)\}\\overset\{\\mathrm\{i\.i\.d\.\}\}\{\\sim\}\\mathrm\{Unif\}\[\-\\zeta,\\zeta\]Multiplicative noise applied independently to observed positions and velocitiesxi\(m\),noisy,vi\(m\),noisy∈ℝdx\_\{i\}^\{\(m\),\\mathrm\{noisy\}\},\\ v\_\{i\}^\{\(m\),\\mathrm\{noisy\}\}\\in\\mathbb\{R\}^\{d\}Multiplicatively perturbed state and velocity of agentii, replicatemm, at timetℓt\_\{\\ell\}, used to assess robustness to observational noiseζ\\zetaNoise level parameter controlling the magnitude of multiplicative perturbationϵ\\epsilonSmall regularization constant \(10−1210^\{\-12\}\) preventing division by zero in normalized error ratiosModel selectionϕE\\phi^\{E\}Energy \(position\-based\) interaction kernel in the generalized frameworkϕA\\phi^\{A\}Alignment \(velocity\-based\) interaction kernel \(second\-order systems\)ccA candidate dynamical framework \(combination of dynamical order and active force terms\)F1,F2F\_\{1\},F\_\{2\}First\-order candidate frameworks \(interaction\-only, and interaction\+environmental\) in Table[1](https://arxiv.org/html/2608.25181#S4.T1)S1,…,S6S\_\{1\},\\ldots,S\_\{6\}Second\-order candidate frameworks, distinguished by which off,ϕE,ϕAf,\\phi^\{E\},\\phi^\{A\}are active, in Table[1](https://arxiv.org/html/2608.25181#S4.T1)Ac,θ^c,bcA\_\{c\},\\ \\hat\{\\theta\}\_\{c\},\\ b\_\{c\}Normal\-equation matrix, coefficient vector, and right\-hand side \(A,θ^,bA,\\hat\{\\theta\},bas in \([36](https://arxiv.org/html/2608.25181#S2.E36)\)\) constructed for a specific candidate modelccin the model\-selection algorithm𝒞\\mathcal\{C\}Set of admissible candidate frameworks passing the compatibility gate𝒞val\\mathcal\{C\}\_\{\\mathrm\{val\}\}Subset of admissible candidates whose mean validation trajectory error is belowτtraj\\tau\_\{\\mathrm\{traj\}\}Mtrain,Mval,MtestM\_\{\\mathrm\{train\}\},M\_\{\\mathrm\{val\}\},M\_\{\\mathrm\{test\}\}Training, validation, and test trajectory splits;M=Mtrain\+Mval\+MtestM=M\_\{\\mathrm\{train\}\}\+M\_\{\\mathrm\{val\}\}\+M\_\{\\mathrm\{test\}\}τrecon\\tau\_\{\\mathrm\{recon\}\}Admissibility threshold for the relative reconstruction error \(set to0\.250\.25\)τres\\tau\_\{\\mathrm\{res\}\}Admissibility threshold for the relative residual error \(set to0\.250\.25\)EminE\_\{\\min\}Minimum mean validation trajectory error among admissible candidates:minc∈𝒞E¯traj\(c\)\\min\_\{c\\in\\mathcal\{C\}\}\\bar\{E\}\_\{\\mathrm\{traj\}\}\(c\)τtraj\\tau\_\{\\mathrm\{traj\}\}Validation trajectory threshold:max\{Emin\+εabs,Emin\(1\+εrel\)\}\\max\\\!\\bigl\\\{E\_\{\\min\}\+\\varepsilon\_\{\\mathrm\{abs\}\},\\,E\_\{\\min\}\(1\+\\varepsilon\_\{\\mathrm\{rel\}\}\)\\bigr\\\}εabs,εrel\\varepsilon\_\{\\mathrm\{abs\}\},\\varepsilon\_\{\\mathrm\{rel\}\}Absolute and relative tolerances defining𝒞val\\mathcal\{C\}\_\{\\mathrm\{val\}\}\(set to10−310^\{\-3\}and0\.050\.05\)c∗c^\{\*\}Selected \(highest\-ranked\) candidate framework
Similar Articles
Learning Dynamical Systems from Multiple Sparse Datasets: A Hierarchical Bayesian Modeling Approach
Proposes a hierarchical Bayesian framework for meta-learning in dynamical systems from multiple sparse, noisy datasets, using gradient-based MCMC with an embedded ODE solver for efficient posterior inference of shared and dataset-specific parameters.
Information-Theoretic Decomposition for Multimodal Interaction Learning
This paper presents an information-theoretic analysis of multimodal learning, revealing the need to capture sample-specific interactions, and proposes DMIL, a paradigm that explicitly models and learns from these interactions via variational decomposition and fine-tuning, achieving superior performance.
Training, learning and inference: unified dynamics of neural systems
This paper proposes a unified dynamical framework for training, learning, and inference in neural systems.
World models of environment, agent and joint agent-environment systems
This paper proposes a framework for world models in reinforcement learning by distinguishing between environment, agent, and joint system channels, using computational mechanics to define canonical predictive models and analyzing their complexity under coupling.
Variational Parameter Calibration with Physics-Aware Latent-Space Surrogates
This paper introduces a physics-aware autoencoder-based latent-space framework for reduced-order forward modeling and variational parameter estimation in parametric dynamical systems, demonstrated on computational fluid dynamics benchmarks. The method enables differentiable surrogate-based inverse modeling and shows improved calibration robustness under realistic noisy or partial observations.