A Semiparametric Framework for Stochastic Fundamental Diagram Modeling

arXiv cs.LG Papers

Summary

This paper proposes a semiparametric framework for stochastic fundamental diagram modeling that combines physically-constrained functional forms with neural network structures to capture complex traffic flow patterns and uncertainty, demonstrating superior performance on real-world datasets.

arXiv:2607.15907v1 Announce Type: new Abstract: The stochastic fundamental diagram (SFD) provides a probabilistic description of the relationship between traffic density and flow or speed, enabling uncertainty-aware traffic modeling. However, existing stochastic models frequently struggle to accommodate rigorous physical constraints while retaining sufficient flexibility to capture complex nonlinear patterns. To address this, we propose a novel semiparametric SFD modeling framework by leveraging specially designed functional forms. These functions intrinsically satisfy physical constraints defined on the moments of the conditional flow distribution given traffic density while incorporating neural-network-based structures to capture complex empirical patterns. We derive a system of moment-matching equations to convert physical constraints into the parameterization of the conditional distribution, proving that a unique solution exists for the location-scale family of distributions, thereby guaranteeing model well-posedness. Furthermore, we demonstrate that the framework can be extended to non-location-scale distributions, including those requiring additional boundary constraints. Empirical evaluations on a real-world dataset reveal that our approach consistently outperforms representative baselines, delivering superior probabilistic accuracy and robust uncertainty quantification, particularly in congested regimes. Overall, the proposed framework provides a theoretically grounded and flexible foundation for stochastic traffic flow modeling.
Original Article
View Cached Full Text

Cached at: 07/20/26, 09:31 AM

# A Semiparametric Framework for Stochastic Fundamental Diagram Modeling
Source: [https://arxiv.org/html/2607.15907](https://arxiv.org/html/2607.15907)
###### Abstract

The stochastic fundamental diagram \(SFD\) provides a probabilistic description of the relationship between traffic density and flow or speed, enabling uncertainty\-aware traffic modeling\. However, existing stochastic models frequently struggle to accommodate rigorous physical constraints while retaining sufficient flexibility to capture complex nonlinear patterns\. To address this, we propose a novel semiparametric SFD modeling framework by leveraging specially designed functional forms\. These functions intrinsically satisfy physical constraints defined on the moments of the conditional flow distribution given traffic density while incorporating neural\-network\-based structures to capture complex empirical patterns\. We derive a system of moment\-matching equations to convert physical constraints into the parameterization of the conditional distribution, proving that a unique solution exists for the location–scale family of distributions, thereby guaranteeing model well\-posedness\. Furthermore, we demonstrate that the framework can be extended to non\-location–scale distributions, including those requiring additional boundary constraints\. Empirical evaluations on the real\-world dataset reveal that our approach consistently outperforms representative baselines, delivering superior probabilistic accuracy and robust uncertainty quantification, particularly in congested regimes\. Overall, the proposed framework provides a theoretically grounded and flexible foundation for stochastic traffic flow modeling\.

###### keywords:

Stochastic fundamental diagram; Semiparametric modeling; Traffic flow modeling; Heteroscedastic uncertainty; Neural network\.

††journal:Transportation Research Part B: Methodological\\affiliation

\[add1\] organization=KTH Royal Institute of Technology, city=Stockholm, country=Sweden

\\affiliation

\[add3\] organization=Swedish Transport Administration, city=Borlänge, country=Sweden

## 1Introduction

The fundamental diagram \(FD\) describes the intrinsic relationship among three macroscopic traffic variables, i\.e\., traffic densityρ\\rho, flow rateqq, and space\-mean speedvv, under steady\-state conditions\. Since traffic flow is the product of speed and density, the fundamental diagram is typically represented by either a speed\-density or flow\-density relationship\. Studies on the fundamental diagram provide critical insights on the vehicle movement patterns on a road segment, enabling the development of models for traffic simulation, congestion detection, and traffic management strategies\.

Classical FD models typically assume a deterministic relationship between traffic density and speed, where each density corresponds to a uniquely speed value\. The seminal work of the Greenshields model\(Greenshieldset al\.,[1935](https://arxiv.org/html/2607.15907#bib.bib1)\)establishes a linear speed\-density relationship, followed by several refinements such as the Greenberg model\(Greenberg,[1959](https://arxiv.org/html/2607.15907#bib.bib2)\)and the Newell model\(Newell,[1961](https://arxiv.org/html/2607.15907#bib.bib3)\)\. Although these classical models have provided valuable insights into traffic flow theory, they inherently overlook the stochastic fluctuations and heterogeneous driver behaviors that characterize traffic in real world\. This limitation has motivated the development of stochastic approaches to FD modeling, which aims to capture inherent uncertainty present in real\-world traffic observations\.

To better capture the random characteristics inherent in road traffic flow, the stochastic fundamental diagram \(SFD\) has been proposed\. Instead of mapping density to a single deterministic flow value, the SFD characterizes a probability distribution of flow or speed for a given density\. Existing SFD models, however, are often closely tied to deterministic FD formulations\. One common approach treats the parameters of a deterministic model, such as free\-flow speed and jam density, as random variables that follow a prescribed distribution\(Jabariet al\.,[2014](https://arxiv.org/html/2607.15907#bib.bib15); Zhou and Zhu,[2020](https://arxiv.org/html/2607.15907#bib.bib22); Baiet al\.,[2021](https://arxiv.org/html/2607.15907#bib.bib14); Chenget al\.,[2024](https://arxiv.org/html/2607.15907#bib.bib6)\)\. Another prevalent approach is to integrate deterministic FD structures with stochastic processes to represent random traffic dynamics\(Niet al\.,[2018](https://arxiv.org/html/2607.15907#bib.bib16); Leiet al\.,[2024](https://arxiv.org/html/2607.15907#bib.bib7); Chenget al\.,[2022](https://arxiv.org/html/2607.15907#bib.bib35)\)\.

While conceptually well\-defined, these approaches inherit limitations from the deterministic models on which they are built\. In particular, widely used single\-regime deterministic models often have limited capacity to represent complex traffic states and transitions across the full range of observed densities\(Ni,[2015](https://arxiv.org/html/2607.15907#bib.bib5)\)\. Consequently, relying on predetermined deterministic forms in certain traffic regimes may introduce systematic bias that undermines modeling accuracy\. In addition, there is no consensus on the appropriate deterministic formulation to adopt in SFD modeling\(Leiet al\.,[2024](https://arxiv.org/html/2607.15907#bib.bib7); Chenget al\.,[2022](https://arxiv.org/html/2607.15907#bib.bib35); Wanget al\.,[2021](https://arxiv.org/html/2607.15907#bib.bib18)\)\. This lack of agreement stems from the different underlying assumptions adopted by existing models, making it difficult to establish clear and unified guidelines for model selection\.

These limitations motivate the development of a modeling framework with enhanced nonlinear modeling capabilities as well as strict adherence to traffic flow physics\. In this study, we adopt a semiparametric modeling strategy that integrates physically motivated constraints with data\-driven components\. Rather than relying on a predefined deterministic fundamental diagram, the proposed approach enforces rigorously\-defined physical properties while allowing functional relationships to be learned from data\. Specifically, a semiparametric formulation is constructed in which physically constrained structures ensure consistency with traffic flow theory, while neural network \(NN\) based correction terms are introduced to capture complex patterns observed in empirical traffic data\. In addition, traffic flow conditioned on density is modeled through a probabilistic distribution capable of representing asymmetric variability commonly observed in real traffic conditions\. The distribution parameters are modeled as semiparametric functions of density, enabling the model to represent the full conditional distribution of traffic flow across different density regimes\.

The main contributions of this paper are summarized as follows:

- 1\.A semiparametric modeling framework is proposed for estimating the stochastic flow\-density relationship in the traffic fundamental diagram, which enforces physical constraints while flexibly capturing complex patterns in empirical data\.
- 2\.A physically constrained parameterization strategy is developed for selected probabilistic distributions; the strategy is proven to be valid for the location\-scale family and demonstrates strong extensibility to other complex distributions\.
- 3\.The proposed approach is validated using extensive experiments on real\-world traffic data, with appropriate evaluation workflow to assess modeling accuracy and uncertainty quantification\.

## 2Related Works

### 2\.1Deterministic Fundamental Diagrams

Traditionally, deterministic fundamental diagrams characterize traffic flow primarily through the speed–density relationship\. The Greenshields model\(Greenshieldset al\.,[1935](https://arxiv.org/html/2607.15907#bib.bib1)\), which assumes a linear relationship between speed and density, was the first to formalize this concept\. Since then, extensive research has sought to improve the accuracy and interpretability of deterministic fundamental diagram models\. Based on their functional forms, these models are commonly classified into three major categories: logarithmic, exponential, and polynomial formulations\. The detailed mathematical expressions for representative models in each category are summarized in Table[1](https://arxiv.org/html/2607.15907#S2.T1)\.

Table 1:Three Categories of Deterministic Models
### 2\.2Stochastic Fundamental Diagrams

The most common approach to modeling SFD is to extend existing deterministic models by introducing random parameters\. In this framework, selected parameters of deterministic models are treated as random variables to capture variability in steady\-state traffic conditions\. For example,Jabariet al\.\([2014](https://arxiv.org/html/2607.15907#bib.bib15)\)andZhou and Zhu \([2020](https://arxiv.org/html/2607.15907#bib.bib22)\)modeled microscopic parameters such as headway and speed adaptation time as random variables and derived the macroscopic fundamental diagram relations under steady\-state assumptions\. Similarly,Quet al\.\([2017](https://arxiv.org/html/2607.15907#bib.bib17)\)andWanget al\.\([2021](https://arxiv.org/html/2607.15907#bib.bib18)\)estimated model parameters at different distribution quantiles, resulting in discrete stochastic representations of the fundamental diagram\. model with random variables and obtained the stochastic speed\-density relationship\. Building on this approach,Ahmedet al\.\([2021](https://arxiv.org/html/2607.15907#bib.bib21)\)further analyzed traffic flow heterogeneity arising from different vehicle types\. In addition to randomizing existing model parameters,Baiet al\.\([2021](https://arxiv.org/html/2607.15907#bib.bib14)\)introduced additional random variables to explicitly represent traffic heterogeneity\.Chenget al\.\([2024](https://arxiv.org/html/2607.15907#bib.bib6)\)investigated analytical properties of the model parameter when making other parameters as random variables\.

In addition to random parameter formulations, another commonly adopted approach is to treat a deterministic fundamental diagram as the expected value and introduce stochasticity through an explicit random process\. In this framework, the deterministic model describes the mean traffic behavior while stochastic variability is modeled separately\. For example,Niet al\.\([2018](https://arxiv.org/html/2607.15907#bib.bib16)\)applied a complex deterministic model\(Niet al\.,[2016](https://arxiv.org/html/2607.15907#bib.bib30)\)for the expected values and characterized the stochasticity using the Maxwell–Boltzmann distribution inspired by ideal gas theory\. Similarly,Liuet al\.\([2023](https://arxiv.org/html/2607.15907#bib.bib31)\)extend the deterministic fundamental diagrams using Gaussian process regression to incorporate additional traffic variables\. Along the same line,Leiet al\.\([2024](https://arxiv.org/html/2607.15907#bib.bib7)\)andChenget al\.\([2022](https://arxiv.org/html/2607.15907#bib.bib35)\)adopted the deterministic models as the expectation function and introduced stochasticity with Gaussian process regression\.

Some studies depart entirely from deterministic fundamental diagram formulations\. Instead, they derive stochastic traffic relationships directly from traffic flow dynamics or alternative stochastic representations\. For example, several studies model traffic evolution based on the continuity equation to describe macroscopic traffic dynamics without relying on predefined fundamental diagrams\(Shiet al\.,[2021](https://arxiv.org/html/2607.15907#bib.bib34); Stormet al\.,[2022](https://arxiv.org/html/2607.15907#bib.bib28); Ngoduy,[2011](https://arxiv.org/html/2607.15907#bib.bib33)\)\. In a different line of study,Zhanget al\.\([2025](https://arxiv.org/html/2607.15907#bib.bib25)\)andZhanget al\.\([2025](https://arxiv.org/html/2607.15907#bib.bib25)\)described the platoon behavior using Markov chains and derived the macroscopic stochastic FD\. Moreover,Bramichet al\.\([2023](https://arxiv.org/html/2607.15907#bib.bib19)\)employed an advanced statistical regression framework, namely generalized additive models for location, scale, and shape \(GAMLSS\)\(Rigby and Stasinopoulos,[2005](https://arxiv.org/html/2607.15907#bib.bib32)\), to develop the SFD model\.

## 3Methodology

### 3\.1Problem Definition

This study treats traffic flow as a random variable conditioned on density\. Letρ∈\[0,ρjam\]\\rho\\in\[0,\\rho\_\{\\mathrm\{jam\}\}\]andq∈ℝ≥0q\\in\\mathbb\{R\}\_\{\\geq 0\}denote the traffic density and traffic flow, respectively, whereρjam\\rho\_\{\\mathrm\{jam\}\}is the traffic jam density\. The objective is to model the conditional probability distribution of traffic flow given density, denoted byp​\(q∣ρ\)p\(q\\mid\\rho\)\. We assume that, for each fixed densityρ\\rho, the conditional distributionp​\(q∣ρ\)p\(q\\mid\\rho\)admits an analytical probability density function \(PDF\) belonging to a parametric familyfPDFf\_\{\\mathrm\{PDF\}\}, parameterized by annn\-dimensional vector𝜼∈ℝn\\bm\{\\eta\}\\in\\mathbb\{R\}^\{n\}:

p​\(q∣ρ\)=fPDF​\(q∣ρ;𝜼\)\.p\(q\\mid\\rho\)=f\_\{\\mathrm\{PDF\}\}\\\!\\left\(q\\mid\\rho;\\bm\{\\eta\}\\right\)\.\(1\)
The admissible distributions are required to satisfy fundamental physical properties of traffic flow\. Under the zero\-density condition \(ρ=0\\rho=0\), the absence of vehicles implies that both traffic flow and its variability vanish\. Under the jam\-density condition \(ρ=ρjam\\rho=\\rho\_\{\\mathrm\{jam\}\}\), vehicles are fully congested, resulting in zero flow with negligible variability\. For intermediate density, the traffic flow value must be positive, and its variability must also be positive\.

These physical requirements impose the following moment\-based constraints on the conditional distributionp​\(q∣ρ\)p\(q\\mid\\rho\)\. Let𝔼​\[⋅\]\\mathbb\{E\}\[\\cdot\]and𝕊​\[⋅\]\\mathbb\{S\}\[\\cdot\]denote the expectation and standard deviation, respectively:

\{𝔼q∼p​\(q∣ρ\)​\[q\]=𝕊q∼p​\(q∣ρ\)​\[q\]=0,ρ∈\{0,ρjam\},0<𝔼q∼p​\(q∣ρ\)​\[q\]<∞,ρ∈\(0,ρjam\),0<𝕊q∼p​\(q∣ρ\)​\[q\]<∞,ρ∈\(0,ρjam\)\.\\begin\{cases\}\\mathbb\{E\}\_\{q\\sim p\(q\\mid\\rho\)\}\[q\]=\\mathbb\{S\}\_\{q\\sim p\(q\\mid\\rho\)\}\[q\]=0,&\\rho\\in\\\{0,\\rho\_\{\\mathrm\{jam\}\}\\\},\\\\ 0<\\mathbb\{E\}\_\{q\\sim p\(q\\mid\\rho\)\}\[q\]<\\infty,&\\rho\\in\(0,\\rho\_\{\\mathrm\{jam\}\}\),\\\\ 0<\\mathbb\{S\}\_\{q\\sim p\(q\\mid\\rho\)\}\[q\]<\\infty,&\\rho\\in\(0,\\rho\_\{\\mathrm\{jam\}\}\)\.\\end\{cases\}\(2\)The modeling problem addressed in this study is therefore to construct a flexible yet physically consistent parameterization ofp​\(q∣ρ\)p\(q\\mid\\rho\)that satisfies these constraints while accurately capturing the stochastic variability observed in empirical traffic data\.

### 3\.2Semiparametric Framework for SFD Models

A common data\-driven approach to modeling stochastic traffic flow relationships is to learn distributional parameters as functions of traffic density using nonparametric models such as neural networks\(Bishop,[1994](https://arxiv.org/html/2607.15907#bib.bib36); Theis and Bethge,[2015](https://arxiv.org/html/2607.15907#bib.bib37)\)\. While such approaches offer substantial modeling flexibility and expressive capability, they do not inherently guarantee physical consistency with respect to the fundamental properties of traffic flow\. In particular, without explicitly enforcing the boundary and non\-negativity conditions specified in Eq\. \([2](https://arxiv.org/html/2607.15907#S3.E2)\), unconstrained nonparametric models may produce physically implausible behavior, especially in regimes near zero density or jam density\.

To address this limitation, this study adopts a semiparametric modeling framework in which physical constraints are explicitly embedded into the parameterization of the conditional distributionp​\(q∣ρ\)p\(q\\mid\\rho\)\. The central idea begins with separating the distributional parameters into constrained and unconstrained components\.

###### Definition 1\(Partition of Distribution Parameters\)\.

Let𝛈∈ℝn\\bm\{\\eta\}\\in\\mathbb\{R\}^\{n\}\(n≥2n\\geq 2\) denote the parameter vector governing the conditional distributionp​\(q∣ρ\)p\(q\\mid\\rho\)\. The parameter vector𝛈\\bm\{\\eta\}is defined by a mutually exclusive partition:

𝜼𝖳=\[𝜼con𝖳𝜼free𝖳\],\\bm\{\\eta\}^\{\\mathsf\{T\}\}=\\begin\{bmatrix\}\\bm\{\\eta\}\_\{\\mathrm\{con\}\}^\{\\mathsf\{T\}\}&\\bm\{\\eta\}\_\{\\mathrm\{free\}\}^\{\\mathsf\{T\}\}\\end\{bmatrix\},\(3\)where:

- 1\.𝜼con∈ℝ2\\bm\{\\eta\}\_\{\\mathrm\{con\}\}\\in\\mathbb\{R\}^\{2\}represents theconstrainedcomponent determined by construction through physical constraints on the first two moments \(the conditional expectation and standard deviation\)\.
- 2\.𝜼free∈ℝn−2\\bm\{\\eta\}\_\{\\mathrm\{free\}\}\\in\\mathbb\{R\}^\{n\-2\}represents thefreecomponent that is unconstrained and capture higher\-order distributional characteristics\.

Under this framework, the distributional parameters are derived by integrating data\-driven learning with strict physical constraints\. The unconstrained parameters𝜼free\\bm\{\\eta\}\_\{\\mathrm\{free\}\}are learned directly from data using a neural network, leveraging its universal approximation capability\(Horniket al\.,[1989](https://arxiv.org/html/2607.15907#bib.bib29)\)\. In contrast, the constrained parameters𝜼con\\bm\{\\eta\}\_\{\\mathrm\{con\}\}are obtained analytically by enforcing consistency between the theoretical moments of the distribution and the macroscopic traffic characteristics\.

Letμ​\(𝜼\)\\mu\(\\bm\{\\eta\}\)andσ​\(𝜼\)\\sigma\(\\bm\{\\eta\}\)denote the theoretical expectation and standard deviation derived from the probability density functionfPDF​\(q∣ρ;𝜼\)f\_\{\\mathrm\{PDF\}\}\(q\\mid\\rho;\\ \\bm\{\\eta\}\)\. Meanwhile, we introduce twosemiparametric functionsm​\(ρ\)m\(\\rho\)ands​\(ρ\)s\(\\rho\), representing the conditional expectation and standard deviation of traffic flow as functions of traffic density\. These functions, whose detailed structure will be discussed in Section[3\.3](https://arxiv.org/html/2607.15907#S3.SS3), satisfy the physical constraints defined in Eq\. \([2](https://arxiv.org/html/2607.15907#S3.E2)\) by construction\. The consistency between the distributional moments and macroscopic traffic flow characteristics is described by a moment\-matching system as follows:

𝔼q∼fPDF​\[q\]\\displaystyle\\mathbb\{E\}\_\{q\\sim f\_\{\\mathrm\{PDF\}\}\}\[q\]=μ​\(\[𝜼con𝖳𝜼free𝖳\]𝖳\)=m​\(ρ;𝝋,𝒉\),\\displaystyle=\\mu\\left\(\\begin\{bmatrix\}\\bm\{\\eta\}\_\{\\mathrm\{con\}\}^\{\\mathsf\{T\}\}&\\bm\{\\eta\}\_\{\\mathrm\{free\}\}^\{\\mathsf\{T\}\}\\end\{bmatrix\}^\{\\mathsf\{T\}\}\\right\)=m\(\\rho;\\bm\{\\varphi\},\\bm\{h\}\),\(4\)𝕊q∼fPDF​\[q\]\\displaystyle\\mathbb\{S\}\_\{q\\sim f\_\{\\mathrm\{PDF\}\}\}\[q\]=σ​\(\[𝜼con𝖳𝜼free𝖳\]𝖳\)=s​\(ρ;𝝋,𝒉\),\\displaystyle=\\sigma\\left\(\\begin\{bmatrix\}\\bm\{\\eta\}\_\{\\mathrm\{con\}\}^\{\\mathsf\{T\}\}&\\bm\{\\eta\}\_\{\\mathrm\{free\}\}^\{\\mathsf\{T\}\}\\end\{bmatrix\}^\{\\mathsf\{T\}\}\\right\)=s\(\\rho;\\bm\{\\varphi\},\\bm\{h\}\),where the semiparametric functions consist of a parametric component governed by domain\-specific parameters𝝋\\bm\{\\varphi\}\(e\.g\., jam density\) and a nonparametric correction term𝒉\\bm\{h\}that captures the deviations from the underling traffic flow relationship\.

The correction term𝒉\\bm\{h\}and the unconstrained distribution parameters𝜼free\\bm\{\\eta\}\_\{\\mathrm\{free\}\}are generally unknown and must be inferred from data\. To provide a flexible representation of their dependence on traffic density, we introduce a neural network mapping that generates these values directly from the density input\. That is, a single neural network modelDNN​\(⋅\)\\mathrm\{DNN\}\(\\cdot\), parameterized by𝜽\\bm\{\\theta\}, is employed to jointly model the latent correction term𝒉\\bm\{h\}and the unconstrained parameters𝜼free\\bm\{\\eta\}\_\{\\mathrm\{free\}\}:

\[𝒉𝖳𝜼free𝖳\]𝖳=DNN​\(ρ;𝜽\)\.\\begin\{bmatrix\}\\bm\{h\}^\{\\mathsf\{T\}\}&\\bm\{\\eta\}\_\{\\mathrm\{free\}\}^\{\\mathsf\{T\}\}\\end\{bmatrix\}^\{\\mathsf\{T\}\}=\\mathrm\{DNN\}\(\\rho;\\bm\{\\theta\}\)\.\(5\)The constrained parameters𝜼con\\bm\{\\eta\}\_\{\\mathrm\{con\}\}are subsequently determined by solving the moment\-matching system in Eq\. \([4](https://arxiv.org/html/2607.15907#S3.E4)\)\. For the proposed framework to be well\-defined, it is necessary that this system admits a unique solution\. To guarantee this property, we establish the following proposition for location–scale families of distributions\.

###### Proposition 1\(Unique Solvability for Moment Matching\)\.

Suppose the conditional distributionfPDF​\(q∣ρ;𝛈\)f\_\{\\mathrm\{PDF\}\}\(q\\mid\\rho;\\bm\{\\eta\}\)belongs to a location\-scale family with well\-defined first and second moments\. Let𝛈con\\bm\{\\eta\}\_\{\\mathrm\{con\}\}represent the location and scale parameters\. Then, for any intermediate densityρ∈\(0,ρjam\)\\rho\\in\(0,\\rho\_\{\\mathrm\{jam\}\}\)and any admissible unconstrained parameter vector𝛈free\\bm\{\\eta\}\_\{\\mathrm\{free\}\}, the moment\-matching system defined in Eq\. \([4](https://arxiv.org/html/2607.15907#S3.E4)\) admits a unique solution for𝛈con\\bm\{\\eta\}\_\{\\mathrm\{con\}\}\.

###### Proof\.

Let𝜼con=\[η1,η2\]𝖳\\bm\{\\eta\}\_\{\\mathrm\{con\}\}=\[\\eta\_\{1\},\\eta\_\{2\}\]^\{\\mathsf\{T\}\}, whereη1\\eta\_\{1\}is the location parameter andη2\>0\\eta\_\{2\}\>0is the scale parameter\. By the definition of a location\-scale family, a random variableX∼fPDF​\(q∣ρ;𝜼\)X\\sim f\_\{\\mathrm\{PDF\}\}\(q\\mid\\rho;\\bm\{\\eta\}\)can be expressed as an affine transformation of a base random variableZZ:

X=η1\+η2​Z,X=\\eta\_\{1\}\+\\eta\_\{2\}Z,whereZZfollows the standardized distribution with locationη1=0\\eta\_\{1\}=0and scaleη2=1\\eta\_\{2\}=1\. Crucially, since the shape ofZZis governed solely by the unconstrained parameters𝜼free\\bm\{\\eta\}\_\{\\mathrm\{free\}\}, we can express its expectation and variance as explicit functions of𝜼free\\bm\{\\eta\}\_\{\\mathrm\{free\}\}, denoted byμ0​\(⋅\)\\mu\_\{0\}\(\\cdot\)andσ02​\(⋅\)\\sigma\_\{0\}^\{2\}\(\\cdot\)respectively:

𝔼​\[Z\]=μ0​\(𝜼free\),Var​\[Z\]=σ02​\(𝜼free\)\.\\mathbb\{E\}\[Z\]=\\mu\_\{0\}\(\\bm\{\\eta\}\_\{\\mathrm\{free\}\}\),\\quad\\mathrm\{Var\}\[Z\]=\\sigma\_\{0\}^\{2\}\(\\bm\{\\eta\}\_\{\\mathrm\{free\}\}\)\.By the linearity of expectation and properties of variance, the moments ofXXare given by:

𝔼​\[X\]=η1\+η2​μ0​\(𝜼free\),Var​\[X\]=η22​σ02​\(𝜼free\)\.\\mathbb\{E\}\[X\]=\\eta\_\{1\}\+\\eta\_\{2\}\\mu\_\{0\}\(\\bm\{\\eta\}\_\{\\mathrm\{free\}\}\),\\quad\\mathrm\{Var\}\[X\]=\\eta\_\{2\}^\{2\}\\sigma\_\{0\}^\{2\}\(\\bm\{\\eta\}\_\{\\mathrm\{free\}\}\)\.For any given densityρ∈\(0,ρjam\)\\rho\\in\(0,\\rho\_\{\\mathrm\{jam\}\}\), the corresponding flow expectationm​\(ρ\)m\(\\rho\)and standard deviations​\(ρ\)s\(\\rho\)are admissible provided they satisfy the boundary constraints in Eq\. \([2](https://arxiv.org/html/2607.15907#S3.E2)\), i\.e\.,

0<m​\(ρ\)<∞,0<s​\(ρ\)<∞\.0<m\(\\rho\)<\\infty,\\quad 0<s\(\\rho\)<\\infty\.To enforce consistency, the moment\-matching system is specified as follows:

η1\+η2​μ0​\(𝜼free\)=m​\(ρ\),η22​σ02​\(𝜼free\)=s2​\(ρ\)\.\\eta\_\{1\}\+\\eta\_\{2\}\\mu\_\{0\}\(\\bm\{\\eta\}\_\{\\mathrm\{free\}\}\)=m\(\\rho\),\\quad\\eta\_\{2\}^\{2\}\\sigma\_\{0\}^\{2\}\(\\bm\{\\eta\}\_\{\\mathrm\{free\}\}\)=s^\{2\}\(\\rho\)\.Sinceη2\>0\\eta\_\{2\}\>0ands​\(ρ\)\>0s\(\\rho\)\>0, taking the positive square root of the variance equation yields the unique solutions for the location and scale parameters:

η1=m​\(ρ\)−s​\(ρ\)​μ0​\(𝜼free\)σ0​\(𝜼free\),η2=s​\(ρ\)σ0​\(𝜼free\)\.\\eta\_\{1\}=m\(\\rho\)\-s\(\\rho\)\\frac\{\\mu\_\{0\}\(\\bm\{\\eta\}\_\{\\mathrm\{free\}\}\)\}\{\\sigma\_\{0\}\(\\bm\{\\eta\}\_\{\\mathrm\{free\}\}\)\},\\quad\\eta\_\{2\}=\\frac\{s\(\\rho\)\}\{\\sigma\_\{0\}\(\\bm\{\\eta\}\_\{\\mathrm\{free\}\}\)\}\.∎

The above proposition guarantees that the moment\-matching system admits a unique solution for location\-scale families of distributions\. Consequently, once the semiparametric functionsm​\(ρ;𝝋,𝒉\)m\(\\rho;\\bm\{\\varphi\},\\bm\{h\}\)ands​\(ρ;𝝋,𝒉\)s\(\\rho;\\bm\{\\varphi\},\\bm\{h\}\)and the unconstrained parameters𝜼free\\bm\{\\eta\}\_\{\\mathrm\{free\}\}are specified, the constrained parameters𝜼con\\bm\{\\eta\}\_\{\\mathrm\{con\}\}are uniquely determined\. Because the latent correction term𝒉\\bm\{h\}and the unconstrained parameters𝜼free\\bm\{\\eta\}\_\{\\mathrm\{free\}\}are generated by the neural network defined in Eq\. \([5](https://arxiv.org/html/2607.15907#S3.E5)\), the complete parameter vector𝜼\\bm\{\\eta\}becomes a deterministic function of densityρ\\rho, parameterized by the neural network weights𝜽\\bm\{\\theta\}and the physical parameters𝝋\\bm\{\\varphi\}\. We denote this mapping as

𝜼𝖳​\(ρ;𝜽,𝝋\)=\[𝜼con𝖳​\(ρ;𝜽,𝝋\)𝜼free𝖳​\(ρ;𝜽\)\]\\bm\{\\eta\}^\{\\mathsf\{T\}\}\\left\(\\rho;\\bm\{\\theta\},\\bm\{\\varphi\}\\right\)=\\begin\{bmatrix\}\\bm\{\\eta\}\_\{\\mathrm\{con\}\}^\{\\mathsf\{T\}\}\\left\(\\rho;\\bm\{\\theta\},\\bm\{\\varphi\}\\right\)&\\bm\{\\eta\}\_\{\\mathrm\{free\}\}^\{\\mathsf\{T\}\}\\left\(\\rho;\\bm\{\\theta\}\\right\)\\end\{bmatrix\}\(6\)The resulting modeling procedure can therefore be summarized as the following computational chain:

ρ→Eq\. \([5](https://arxiv.org/html/2607.15907#S3.E5)\)𝜽\(𝒉,𝜼free\)→Eq\. \([4](https://arxiv.org/html/2607.15907#S3.E4)\)𝝋\(𝜼con,𝜼free\)≡𝜼⟶fPDF​\(q∣ρ;𝜼\)\.\\rho\\xrightarrow\[\\text\{Eq\.~\\eqref\{eq:dnn\}\}\]\{\\bm\{\\theta\}\}\(\\bm\{h\},\\bm\{\\eta\}\_\{\\mathrm\{free\}\}\)\\xrightarrow\[\\text\{Eq\.~\\eqref\{eq:moment\_matching\}\}\]\{\\bm\{\\varphi\}\}\(\\bm\{\\eta\}\_\{\\mathrm\{con\}\},\\bm\{\\eta\}\_\{\\mathrm\{free\}\}\)\\equiv\\bm\{\\eta\}\\longrightarrow f\_\{\\mathrm\{PDF\}\}\(q\\mid\\rho;\\bm\{\\eta\}\)\.\(7\)Finally, model parameters are estimated via maximum likelihood using the observed dataset\{\(ρi,qi\)\}i=1N\\\{\(\\rho\_\{i\},q\_\{i\}\)\\\}\_\{i=1\}^\{N\}:

𝜽^,𝝋^=arg⁡max𝜽,𝝋​∑i=1Nlog⁡fPDF​\(qi∣ρi;𝜼​\(ρi;𝜽,𝝋\)\)\.\\hat\{\\bm\{\\theta\}\},\\hat\{\\bm\{\\varphi\}\}=\\arg\\max\_\{\\bm\{\\theta\},\\bm\{\\varphi\}\}\\sum\_\{i=1\}^\{N\}\\log f\_\{\\mathrm\{PDF\}\}\\\!\\left\(q\_\{i\}\\mid\\rho\_\{i\};\\bm\{\\eta\}\(\\rho\_\{i\};\\bm\{\\theta\},\\bm\{\\varphi\}\)\\right\)\.\(8\)The semiparametric framework is therefore well defined for location\-scale families of distributions, ensuring that the resulting conditional distribution satisfies the imposed physical constraints by model construction\. Other distribution families may also be incorporated into the framework, although additional modifications may be required\. The extension of the framework will be further discussed in Section[7](https://arxiv.org/html/2607.15907#S7)\. The overall modeling procedure is intuitively summarized in Fig\.[1](https://arxiv.org/html/2607.15907#S3.F1)\.

![Refer to caption](https://arxiv.org/html/2607.15907v1/figs/flow-diagram.png)Figure 1:Flowchart for Stochastic FD modeling under Semiparametric Framework
### 3\.3Structure Specification of Semiparametric Functions

This section specifies candidate functional forms for the expectationm​\(ρ\)m\(\\rho\)and standard deviations​\(ρ\)s\(\\rho\)of traffic flowqqgiven densityρ\\rho, associated with the conditional distributionp​\(q∣ρ\)p\(q\\mid\\rho\)\. Under the physical constraints in Eq\. \([2](https://arxiv.org/html/2607.15907#S3.E2)\), bothm​\(ρ\)m\(\\rho\)ands​\(ρ\)s\(\\rho\)belong to the same function class that satisfies specific boundary conditions\.

###### Definition 2\(Class of Physically Admissible Functions\)\.

Letℱphy\\mathcal\{F\}\_\{\\mathrm\{phy\}\}denote the class of functions

ℱphy=\{f∈C​\(\[0,ρjam\]\)∣f​\(ρ\)≥0,f​\(0\)=0,f​\(ρjam\)=0\},\\mathcal\{F\}\_\{\\mathrm\{phy\}\}=\\bigl\\\{f\\in C\(\[0,\\rho\_\{\\mathrm\{jam\}\}\]\)\\mid f\(\\rho\)\\geq 0,\\;f\(0\)=0,\\;f\(\\rho\_\{\\mathrm\{jam\}\}\)=0\\bigr\\\},\(9\)whereC​\(\[0,ρjam\]\)C\(\[0,\\rho\_\{\\mathrm\{jam\}\}\]\)denotes the space of continuous real\-valued functions defined on the interval\[0,ρjam\]\[0,\\rho\_\{\\mathrm\{jam\}\}\]\.

Among functions inℱphy\\mathcal\{F\}\_\{\\mathrm\{phy\}\}, the quadratic flow\-density relationship implied by the classical Greenshields model represents the simplest parametric example\. Building on this baseline, we introduce two semiparametric extensions that preserve physical admissibility while enabling data\-driven flexibility through NN\-based correction terms\.

#### 3\.3\.1Quadratic Function with Neural Correction \(QwNC\)

The classical Greenshields model assumes a linear relationship between traffic densityρ\\rhoand space\-mean speedvv\. Since traffic flow is defined asq=ρ​vq=\\rho v, this assumption implies a quadratic relationship between density and flow\. The Greenshields formulation is given by

vG​\(ρ\)=vfree​\(1−ρρjam\),qG​\(ρ\)=ρ​\(ρjam−ρ\)​vfreeρjam,v\_\{\\mathrm\{G\}\}\(\\rho\)=v\_\{\\mathrm\{free\}\}\\left\(1\-\\frac\{\\rho\}\{\\rho\_\{\\mathrm\{jam\}\}\}\\right\),\\qquad q\_\{\\mathrm\{G\}\}\(\\rho\)=\\rho\(\\rho\_\{\\mathrm\{jam\}\}\-\\rho\)\\frac\{v\_\{\\mathrm\{free\}\}\}\{\\rho\_\{\\mathrm\{jam\}\}\},\(10\)wherevfreev\_\{\\mathrm\{free\}\}denotes the free\-flow speed andρjam\\rho\_\{\\mathrm\{jam\}\}the jam density\. The resulting quadratic function attains zero flow at bothρ=0\\rho=0andρ=ρjam\\rho=\\rho\_\{\\mathrm\{jam\}\}, and therefore belongs toℱphy\\mathcal\{F\}\_\{\\mathrm\{phy\}\}\.

To enhance modeling flexibility while preserving physical consistency, the constant coefficientvfree/ρjamv\_\{\\mathrm\{free\}\}/\\rho\_\{\\mathrm\{jam\}\}is replaced by a density\-dependent correction termc​\(ρ\)c\(\\rho\)\. This term corresponds to one element of the latent vector𝒉\\bm\{h\}generated by the neural network defined in Eq\. \([5](https://arxiv.org/html/2607.15907#S3.E5)\)\. The resulting Quadratic Function with Neural Correction is expressed as

fQwNC​\(ρ\)=ρ​\(ρjam−ρ\)​c​\(ρ\)\.f\_\{\\mathrm\{QwNC\}\}\(\\rho\)=\\rho\(\\rho\_\{\\mathrm\{jam\}\}\-\\rho\)c\(\\rho\)\.\(11\)To ensure numerical stability and enforce physical non\-negativity, all positive quantities are mapped through the Softplus functionσ​\(x\)=ln⁡\(1\+ex\)\\sigma\(x\)=\\ln\(1\+e^\{x\}\)\. In particular, the jam density is parameterized asρjam=σ​\(ρj\)\\rho\_\{\\mathrm\{jam\}\}=\\sigma\(\\rho\_\{j\}\), whereρj\\rho\_\{j\}is an element of the physical parameter vector𝝋\\bm\{\\varphi\}\. The density\-dependent correction termc​\(ρ\)c\(\\rho\)is constrained analogously viaσ​\(c​\(ρ\)\)\\sigma\(c\(\\rho\)\), yielding the final formulation

fQwNC​\(ρ\)=ρ​\(σ​\(ρj\)−ρ\)​σ​\(c​\(ρ\)\)\.f\_\{\\mathrm\{QwNC\}\}\(\\rho\)=\\rho\\bigl\(\\sigma\(\\rho\_\{j\}\)\-\\rho\\bigr\)\\sigma\\bigl\(c\(\\rho\)\\bigr\)\.\(12\)
Although the QwNC formulation is motivated by extending the mean flow relationship implied by the Greenshields model, it satisfies the physical constraints in Eq\. \([2](https://arxiv.org/html/2607.15907#S3.E2)\) by construction\. Accordingly, the same functional form can be leveraged to model both the conditional expectation and the conditional standard deviation of traffic flow\.

#### 3\.3\.2Beta Function with Neural Correction \(BwNC\)

To further enhance structural flexibility beyond the quadratic form, we introduce a functional specification inspired by the shape characteristics of the Beta distribution\. Specifically, two shape parameters,aaandbb, are introduced to transform the quadratic termρ​\(σ​\(ρj\)−ρ\)\\rho\(\\sigma\(\\rho\_\{j\}\)\-\\rho\)into the more flexible formρa​\(σ​\(ρj\)−ρ\)b\\rho^\{a\}\\bigl\(\\sigma\(\\rho\_\{j\}\)\-\\rho\\bigr\)^\{b\}\. The parametersaaandbbregulate the curvature near low and high density regimes, respectively, enabling the model to capture skewed or asymmetric patterns observed in empirical data\.

To ensure physical admissibility and numerical stability, the Softplus function is applied to the exponents, enforcing positivity\. In addition, the density difference term,\(σ​\(ρj\)−ρ\)\(\\sigma\(\\rho\_\{j\}\)\-\\rho\), is rectified using amax⁡\(0,⋅\)\\max\(0,\\cdot\)operator to prevent negative values, thereby guaranteeing that the power function remains well\-defined across the entire density domain\. The resulting Beta Function with Neural Correction is given by

fBwNC\(ρ\)=ρσ​\(a\)max\(0,σ\(ρj\)−ρ\)σ​\(b\)σ\(c\(ρ\)\)\.f\_\{\\mathrm\{BwNC\}\}\(\\rho\)=\\rho^\{\\sigma\(a\)\}\\max\\bigl\(0,\\sigma\(\\rho\_\{j\}\)\-\\rho\\bigr\)^\{\\sigma\(b\)\}\\sigma\\bigl\(c\(\\rho\)\\bigr\)\.\(13\)This formulation preserves the physical boundary conditions and provides additional flexibility to capture asymmetric relationships\. The BwNC formulation can be used to model both the expectation and the standard deviation of conditional traffic flow distribution\.

## 4Specification of SFD Models

Given the semiparametric framework introduced in the preceding section, a SFD model is fully specified once the conditional distributionp​\(q∣ρ\)p\(q\\mid\\rho\)and the semiparametric forms of the conditional expectationm​\(ρ\)m\(\\rho\)and standard deviations​\(ρ\)s\(\\rho\)are defined\. In this section, the framework is instantiated by assuming that the conditional distribution follows a Skew\-Normal distribution\. For completeness, the Gaussian distribution is discussed as a special case obtained by setting skewness parameter as zero\.

### 4\.1Skew\-Normal Distribution

The stochastic fundamental diagram model is instantiated by assuming the conditional distributionp​\(q∣ρ\)p\(q\\mid\\rho\)to follow the Skew\-Normal distribution with parameter as𝜼=\{ξ,ω,α\}\\bm\{\\eta\}=\\\{\\xi,\\omega,\\alpha\\\}:

p​\(q∣ρ\)\\displaystyle p\(q\\mid\\rho\)=fSkewNormal​\(q∣ρ;ξ,ω,α\)\\displaystyle=f\_\{\\mathrm\{SkewNormal\}\}\\left\(q\\mid\\rho;\\ \\xi,\\omega,\\alpha\\right\)\(14\)=2π​1ω​exp⁡\(−12​\(q−ξω\)2\)​Φ​\(α⋅q−ξω\),\\displaystyle=\\sqrt\{\\frac\{2\}\{\\pi\}\}\\frac\{1\}\{\\omega\}\\exp\\left\(\-\\frac\{1\}\{2\}\\left\(\\frac\{q\-\\xi\}\{\\omega\}\\right\)^\{2\}\\right\)\\ \\Phi\\left\(\\alpha\\cdot\\frac\{q\-\\xi\}\{\\omega\}\\right\),whereΦ​\(⋅\)\\Phi\(\\cdot\)denote the standard normal cumulative distribution function\. Following the Proposition[1](https://arxiv.org/html/2607.15907#Thmproposition1), location parameterξ\\xiand scale parameterω\\omegaare regarded as constrained, i\.e\.,𝜼con=\[ξ,ω\]𝖳\\bm\{\\eta\}\_\{\\mathrm\{con\}\}=\[\\xi,\\ \\omega\]^\{\\mathsf\{T\}\}\. While the skewness parameterα\\alphaforms the unconstrained parameter𝜼free=\[α\]𝖳\\bm\{\\eta\}\_\{\\mathrm\{free\}\}=\[\\alpha\]^\{\\mathsf\{T\}\}, which is learned directly from data\.

Table 2:Summary of Functional Forms and Parameters for Stochastic FD Models under Skew\-Normal DistributionTwo model variants based on Skew\-Normal distribution, namely SN\-QwNC and SN\-BwNC, are established based on the functional forms defined in Section[3\.3](https://arxiv.org/html/2607.15907#S3.SS3)\. The SN\-QwNC model applies thefQwNCf\_\{\\mathrm\{QwNC\}\}for functions of the expectationm​\(ρ;𝝋,𝒉\)m\(\\rho;\\bm\{\\varphi\},\\bm\{h\}\)and standard deviations​\(ρ;𝝋,𝒉\)s\(\\rho;\\bm\{\\varphi\},\\bm\{h\}\), whereas the SN\-BwNC model utilizes thefBwNCf\_\{\\mathrm\{BwNC\}\}\. A summary of the functional forms and parameter sets is provided in Table[2](https://arxiv.org/html/2607.15907#S4.T2)\. For both the SN\-QwNC and SN\-BwNC specifications, the vector𝒉\\bm\{h\}contains two correction components, one associated with the conditional expectation and one with the conditional standard deviation\. The domain\-specific parameter vector𝝋\\bm\{\\varphi\}differs across the two variants: in the SN\-QwNC model,𝝋\\bm\{\\varphi\}contains only density\-related parameters \(e\.g\., jam density\), whereas in the SN\-BwNC model it additionally includes parameters controlling the curvature of the flow–density relationship\.

The correction terms𝒉\\bm\{h\}and the skewness parameterα\\alphaare learned jointly using a neural network\. Specifically, a three\-layer multilayer perceptron \(MLP\) with Sigmoid Linear Unit \(SiLU\) activation functions\(Hendrycks and Gimpel,[2016](https://arxiv.org/html/2607.15907#bib.bib23)\)is employed\.

The constrained parameters\{ξ,ω\}\\\{\\xi,\\omega\\\}are determined by matching the theoretical moments of the Skew\-Normal distribution to the semiparametric functionsm​\(ρ;𝝋,𝒉\)m\(\\rho;\\bm\{\\varphi\},\\bm\{h\}\)ands​\(ρ;𝝋,𝒉\)s\(\\rho;\\bm\{\\varphi\},\\bm\{h\}\)\. This yields the system of equations as follows:

ξ\+ω​\(α21\+α2​2π\)1/2\\displaystyle\\xi\+\\omega\\left\(\\frac\{\\alpha^\{2\}\}\{1\+\\alpha^\{2\}\}\\frac\{2\}\{\\pi\}\\right\)^\{1/2\}=m​\(ρ;𝝋,𝒉\),\\displaystyle=m\(\\rho;\\bm\{\\varphi\},\\bm\{h\}\),\(15\)ω2​\(1−α21\+α2​2π\)\\displaystyle\\omega^\{2\}\\left\(1\-\\frac\{\\alpha^\{2\}\}\{1\+\\alpha^\{2\}\}\\frac\{2\}\{\\pi\}\\right\)=s2​\(ρ;𝝋,𝒉\),\\displaystyle=s^\{2\}\(\\rho;\\bm\{\\varphi\},\\bm\{h\}\),from whichξ\\xiandω\\omegacan be expressed explicitly as functions ofα\\alpha,m​\(ρ;𝝋,𝒉\)m\(\\rho;\\bm\{\\varphi\},\\bm\{h\}\), ands​\(ρ;𝝋,𝒉\)s\(\\rho;\\bm\{\\varphi\},\\bm\{h\}\), i\.e\.,

ω\\displaystyle\\omega=s​\(ρ;𝝋,𝒉\)​\(1−α21\+α2​2π\)−1/2,\\displaystyle=s\(\\rho;\\bm\{\\varphi\},\\bm\{h\}\)\\left\(1\-\\frac\{\\alpha^\{2\}\}\{1\+\\alpha^\{2\}\}\\frac\{2\}\{\\pi\}\\right\)^\{\-1/2\},\(16\)ξ\\displaystyle\\xi=m​\(ρ;𝝋,𝒉\)−ω​\(α21\+α2​2π\)1/2\.\\displaystyle=m\(\\rho;\\bm\{\\varphi\},\\bm\{h\}\)\-\\omega\\left\(\\frac\{\\alpha^\{2\}\}\{1\+\\alpha^\{2\}\}\\frac\{2\}\{\\pi\}\\right\)^\{1/2\}\.
Consequently, all distributional parameters\{ξ,ω,α\}\\\{\\xi,\\omega,\\alpha\\\}become deterministic functions of the densityρ\\rho, parameterized by the neural network weights𝜽\\bm\{\\theta\}and the domain\-specific parameters𝝋\\bm\{\\varphi\}\. Model parameters are learned by minimizing a composite loss function consisting of the negative log\-likelihood and a jam\-density regularization term, i\.e\.,

ℓ=ℓNLL\+λ⋅ℓReg\.\\ell=\\ell\_\{\\mathrm\{NLL\}\}\+\\lambda\\cdot\\ell\_\{\\mathrm\{Reg\}\}\.\(17\)Given a dataset𝒟=\{\(ρi,qi\)\}i=1N\\mathcal\{D\}=\\\{\(\\rho\_\{i\},q\_\{i\}\)\\\}\_\{i=1\}^\{N\}of observed density\-flow pairs, the negative log\-likelihood under the Skew\-Normal model is

ℓNLL\\displaystyle\\ell\_\{\\mathrm\{NLL\}\}=−log⁡\(∏i=1Np​\(qi\|ρi\)\)\\displaystyle=\-\\log\\\!\\left\(\\prod\_\{i=1\}^\{N\}p\(q\_\{i\}\|\\rho\_\{i\}\)\\right\)\(18\)=N​log⁡ωi\+12​∑i=1N\(qi−ξiωi\)2−∑i=1Nlog⁡Φ​\(αi​qi−ξiωi\),\\displaystyle=N\\log\\omega\_\{i\}\+\\frac\{1\}\{2\}\\sum\_\{i=1\}^\{N\}\\left\(\\frac\{q\_\{i\}\-\\xi\_\{i\}\}\{\\omega\_\{i\}\}\\right\)^\{2\}\-\\sum\_\{i=1\}^\{N\}\\log\\Phi\\\!\\left\(\\alpha\_\{i\}\\frac\{q\_\{i\}\-\\xi\_\{i\}\}\{\\omega\_\{i\}\}\\right\),where

𝒉i,αi=DNN​\(ρi;𝜽\),\\bm\{h\}\_\{i\},\\,\\alpha\_\{i\}=\\mathrm\{DNN\}\(\\rho\_\{i\};\\bm\{\\theta\}\),ωi=s​\(ρi;𝝋,𝒉i\)​\(1−αi21\+αi2​2π\)−1/2,\\omega\_\{i\}=s\(\\rho\_\{i\};\\bm\{\\varphi\},\\bm\{h\}\_\{i\}\)\\left\(1\-\\frac\{\\alpha\_\{i\}^\{2\}\}\{1\+\\alpha\_\{i\}^\{2\}\}\\frac\{2\}\{\\pi\}\\right\)^\{\-1/2\},ξi=m​\(ρi;𝝋,𝒉i\)−ωi​\(αi21\+αi2​2π\)1/2\.\\xi\_\{i\}=m\(\\rho\_\{i\};\\bm\{\\varphi\},\\bm\{h\}\_\{i\}\)\-\\omega\_\{i\}\\left\(\\frac\{\\alpha\_\{i\}^\{2\}\}\{1\+\\alpha\_\{i\}^\{2\}\}\\frac\{2\}\{\\pi\}\\right\)^\{1/2\}\.A regularization term is introduced to stabilize the optimization process and to enforce physical consistency of the estimated jam density:

ℓReg=∑i=1Nmax⁡\(ρi−σ​\(ρj\),0\),\\ell\_\{\\mathrm\{Reg\}\}=\\sum\_\{i=1\}^\{N\}\\max\\bigl\(\\rho\_\{i\}\-\\sigma\(\\rho\_\{j\}\),\\,0\\bigr\),\(19\)whereσ​\(ρj\)\\sigma\(\\rho\_\{j\}\)denotes the jam density reparameterized by Softplus function\. The rationale of this regularization is to encourage the learned jam density to be no smaller than any observed traffic density in the dataset\. Whenever the estimated jam density falls below observed densities, a penalty is incurred to prevent physically implausible solutions\. No penalty is applied when the estimated jam density exceeds all observed densities\.

A relatively large weighting factorλ\\lambda, e\.g\.,λ=102\\lambda=10^\{2\}, is employed to ensure effective enforcement of the regularization during the training phase\. In practice, the regularization rapidly guides the jam density estimate toward a physically admissible range while maintaining stable convergence of the likelihood\-based optimization\.

### 4\.2Normal Distribution

The Normal distribution can be considered as a degenerate case of the Skew\-Normal distribution obtained by setting the skewness parameterα\\alphato zero\. Under this assumption, the conditional distribution of traffic flow given density becomes symmetric, and only the location and scale parameters,ξ\\xiandω\\omega, are required\. Consistent with the semiparametric framework, these parameters are treated as constrained parameters𝜼con\\bm\{\\eta\}\_\{\\mathrm\{con\}\}, while no distributional parameter is left as free, i\.e\.,𝜼free=∅\\bm\{\\eta\}\_\{\\mathrm\{free\}\}=\\emptyset\. The conditional probability density function is given by

p​\(q∣ρ\)=12​π​ω​exp⁡\(−12​\(q−ξω\)2\)\.p\(q\\mid\\rho\)=\\frac\{1\}\{\\sqrt\{2\\pi\}\\,\\omega\}\\exp\\\!\\left\(\-\\frac\{1\}\{2\}\\left\(\\frac\{q\-\\xi\}\{\\omega\}\\right\)^\{2\}\\right\)\.\(20\)Analogous to the Skew\-Normal specification, two Normal distribution based SFD models are defined: N\-QwNC and N\-BwNC\. These models employ the same semiparametric functional forms for the conditional expectation and standard deviation as their Skew\-Normal counterparts\. The only structural difference lies in the neural network output, which no longer includes a skewness parameter\. A summary of the functional forms and associated parameters is provided in Table[3](https://arxiv.org/html/2607.15907#S4.T3)\.

Table 3:Summary of Functional Forms and Parameters for Stochastic FD Models under Normal DistributionFor the Normal distribution, the constrained parameters are obtained as a special case of Eq \([16](https://arxiv.org/html/2607.15907#S4.E16)\) by setting the skewness parameterα=0\\alpha=0\. Under this condition, the location and scale parameters reduce to

ω=s​\(ρ;𝝋,𝒉\),ξ=m​\(ρ;𝝋,𝒉\)\.\\omega=s\(\\rho;\\bm\{\\varphi\},\\bm\{h\}\),\\qquad\\xi=m\(\\rho;\\bm\{\\varphi\},\\bm\{h\}\)\.\(21\)Since there is no skewness parameter, the neural network only outputs the correction terms associated with the semiparametric functions\. Given a dataset𝒟=\{\(ρi,qi\)\}i=1N\\mathcal\{D\}=\\\{\(\\rho\_\{i\},q\_\{i\}\)\\\}\_\{i=1\}^\{N\}, the loss function is defined as:

ℓ\\displaystyle\\ell=ℓNLL\+λ​ℓReg\\displaystyle=\\ell\_\{\\mathrm\{NLL\}\}\+\\lambda\\ell\_\{\\mathrm\{Reg\}\}\(22\)=∑i=1Nlog⁡s​\(ρi;𝝋,𝒉\)\+12​∑i=1N\(qi−m​\(ρi;𝝋,𝒉\)s​\(ρi;𝝋,𝒉\)\)2\+λ​∑i=1Nmax⁡\(ρi−σ​\(ρj\),0\)\.\\displaystyle=\\sum\_\{i=1\}^\{N\}\\log s\(\\rho\_\{i\};\\bm\{\\varphi\},\\bm\{h\}\)\+\\frac\{1\}\{2\}\\sum\_\{i=1\}^\{N\}\\left\(\\frac\{q\_\{i\}\-m\(\\rho\_\{i\};\\bm\{\\varphi\},\\bm\{h\}\)\}\{s\(\\rho\_\{i\};\\bm\{\\varphi\},\\bm\{h\}\)\}\\right\)^\{2\}\+\\lambda\\sum\_\{i=1\}^\{N\}\\max\\bigl\(\\rho\_\{i\}\-\\sigma\(\\rho\_\{j\}\),0\\bigr\)\.

## 5Computational Experiments

### 5\.1MAGIC Dataset

The proposed models are evaluated using the MAGIC \(Multiple Conditions UAV Group\-based High\-fidelity Comprehensive Vehicle Trajectory\) dataset\(Maet al\.,[2022](https://arxiv.org/html/2607.15907#bib.bib4)\), which provides high\-resolution vehicle trajectory data collected by unmanned aerial vehicles \(UAVs\)\. The data were recorded over a three\-hour period \(7:40–10:40 AM\) along a 4 km segment of the Shanghai Inner Ring Road\.

The dataset contains detailed information on vehicle type, position, velocity, and acceleration at a sampling frequency of 25 Hz\. From the full dataset, a 200 m two\-lane road segment was extracted to ensure coverage of a wide range of traffic conditions, spanning free\-flow to congested regimes\. Traffic variables were aggregated into11second intervals\. The aggregated speed is measured in kilometers per hour \(km/h\), while traffic density is expressed in vehicles per kilometer per lane \(veh/km/lane\)\. Traffic flow is computed as the product of density and speed, with units of vehicles per hour per lane \(veh/h/lane\)\.

### 5\.2Experiment Setup

A major challenge in our evaluation is the limited and imbalanced nature of the dataset\. Since only three hours of tabular data were collected from a single location, the limited data volume introduces high variance in the evaluation metrics\. Furthermore, the skewed distribution toward low\-density samples could overemphasize model performance in the free\-flow state while ignoring the performance in the congested state\.

Stratifiedkk\-fold cross\-validation combined with a sample\-weighting scheme \(as illustrated in Figure[2](https://arxiv.org/html/2607.15907#S5.F2)\) is used to address the aforementioned challenge\. The procedure is detailed as follows:

- 1\.Stratified Partitioning\. The entire density range is partitioned intob=10b=10equal bins, withNmN\_\{m\}data points inmm\-th bin\. The data within each bin is further split intok=5k=5equal folds\. A complete fold is constructed by taking one partition from each bin without replacement, resulting in a total ofkkfolds\.
- 2\.Sample Weighting\. For each experimental run, one fold𝒯j\\mathcal\{T\}\_\{j\}\(j∈\{1,…,k\}j\\in\\\{1,\\dots,k\\\}\) serves as the test set, while the remainingk−1k\-1folds are used for training\. Each test samplei∈𝒯ji\\in\\mathcal\{T\}\_\{j\}is weighted inversely proportional to its bin’s total population\. Specifically, the weight is wi=1Nb​\(i\),w\_\{i\}=\\frac\{1\}\{N\_\{b\(i\)\}\},\(23\)whereb​\(i\)b\(i\)is the bin index of sampleii\.

The weightwiw\_\{i\}is explicitly integrated into the calculation of each specific evaluation metric \(detailed in Section[5\.3](https://arxiv.org/html/2607.15907#S5.SS3)\) to ensure an unbiased assessment\. LetSjS\_\{j\}denote the aggregated score of a given metric evaluated on the test fold𝒯j\\mathcal\{T\}\_\{j\}\. After iterating through allkkfolds, the model’s final performance is summarized by the mean \(μS\\mu\_\{S\}\) and standard deviation \(σS\\sigma\_\{S\}\) of thekkaggregated scores:

μS=1k​∑j=1kSj,σS=1k​∑j=1k\(Sj−μS\)2\.\\mu\_\{S\}=\\frac\{1\}\{k\}\\sum\_\{j=1\}^\{k\}S\_\{j\},\\quad\\sigma\_\{S\}=\\sqrt\{\\frac\{1\}\{k\}\\sum\_\{j=1\}^\{k\}\\left\(S\_\{j\}\-\\mu\_\{S\}\\right\)^\{2\}\}\.\(24\)
![Refer to caption](https://arxiv.org/html/2607.15907v1/figs/experiment_setup.png)Figure 2:Stratified 5\-fold cross\-validation combined with a sample\-weighting scheme
### 5\.3Metrics

Five metrics are used for a comprehensive evaluation of the proposed models\. The first two are probabilistic metrics, which evaluate the quality of the predicted output distribution\. The remaining three are deterministic metrics, which assess the accuracy of the predicted conditional mean\. For notation clarity across all metrics, letqiq\_\{i\}denote the ground truth flow observation for sampleii\. LetfPDF\(⋅∣ρi;𝜼\(ρi;𝜽,𝝋\)\)f\_\{\\mathrm\{PDF\}\}\\left\(\\cdot\\mid\\rho\_\{i\};\\bm\{\\eta\}\(\\rho\_\{i\};\\bm\{\\theta\},\\bm\{\\varphi\}\)\\right\)denote the predicted conditional probability density function given densityρi\\rho\_\{i\}, where the distribution parameters𝜼\\bm\{\\eta\}are produced by a model parameterized by learnable weights𝜽\\bm\{\\theta\}and𝝋\\bm\{\\varphi\}\.

#### 5\.3\.1Probabilistic Metrics

##### Weighted Continuous Ranked Probability Score \(WCRPS\)

The CRPS is a proper scoring rule commonly used to evaluate the probabilistic forecasts\. For a predictive distributionffand an actual observationyy, it is defined as:

CRPS​\(f,y\)=∫−∞∞\(∫−∞xf​\(t\)​dt−𝟏x≥y\)2​dx\.\\text\{CRPS\}\(f,y\)=\\int\_\{\-\\infty\}^\{\\infty\}\\left\(\\int\_\{\-\\infty\}^\{x\}f\(t\)\\mathrm\{d\}t\-\\mathbf\{1\}\_\{x\\geq y\}\\right\)^\{2\}\\mathrm\{d\}x\.\(25\)Since integrating the cumulative distribution is often computationally intractable, CRPS can be approximated in a expectation form using Monte Carlo simulation\(Gneiting and Raftery,[2007](https://arxiv.org/html/2607.15907#bib.bib38)\):

CRPS\(f,y\)=𝔼X∼f\[\|X−y\|\]−12𝔼X,X′∼f\[\|X−X′\|\]\.\\text\{CRPS\}\(f,y\)=\\mathbb\{E\}\_\{X\\sim f\}\\bigl\[\|X\-y\|\\bigl\]\-\\frac\{1\}\{2\}\\mathbb\{E\}\_\{X,X^\{\\prime\}\\sim f\}\\bigl\[\|X\-X^\{\\prime\}\|\\bigl\]\.\(26\)When incorporating the sample weightswiw\_\{i\}, the aggregated Weighted CRPS for thejj\-th test fold𝒯j\\mathcal\{T\}\_\{j\}is calculated as:

Sj,WCRPS=∑i∈𝒯jwi⋅CRPS\(fPDF\(⋅∣ρi;𝜼\(ρi;𝜽,𝝋\)\),qi\)∑i∈𝒯jwi\.S\_\{j,\\mathrm\{WCRPS\}\}=\\frac\{\\sum\_\{i\\in\\mathcal\{T\}\_\{j\}\}w\_\{i\}\\cdot\\text\{CRPS\}\\left\(f\_\{\\mathrm\{PDF\}\}\\\!\\left\(\\cdot\\mid\\rho\_\{i\};\\bm\{\\eta\}\(\\rho\_\{i\};\\bm\{\\theta\},\\bm\{\\varphi\}\)\\right\),q\_\{i\}\\right\)\}\{\\sum\_\{i\\in\\mathcal\{T\}\_\{j\}\}w\_\{i\}\}\.\(27\)

##### Weighted Negative Log\-Likelihood \(WNLL\)

The Negative Log\-Likelihood measures the goodness\-of\-fit by penalizing low probabilities assigned to the true observations\. By incorporating the weight, the WNLL for the test fold is given by:

Sj,WNLL=−∑i∈𝒯jwi⋅log⁡fPDF​\(qi∣ρi;𝜼​\(ρi;𝜽,𝝋\)\)∑i∈𝒯jwi\.S\_\{j,\\mathrm\{WNLL\}\}=\-\\frac\{\\sum\_\{i\\in\\mathcal\{T\}\_\{j\}\}w\_\{i\}\\cdot\\log f\_\{\\mathrm\{PDF\}\}\\\!\\left\(q\_\{i\}\\mid\\rho\_\{i\};\\bm\{\\eta\}\(\\rho\_\{i\};\\bm\{\\theta\},\\bm\{\\varphi\}\)\\right\)\}\{\\sum\_\{i\\in\\mathcal\{T\}\_\{j\}\}w\_\{i\}\}\.\(28\)

#### 5\.3\.2Deterministic Metrics

For deterministic evaluation, we collapse the predictive distribution into a single point estimateq^i\\hat\{q\}\_\{i\}by taking its mathematical expectation:

q^i=𝔼X∼fPDF\(⋅∣ρi;𝜼\(ρi;𝜽,𝝋\)\)​\[X\]\.\\hat\{q\}\_\{i\}=\\mathbb\{E\}\_\{X\\sim f\_\{\\mathrm\{PDF\}\}\\\!\\left\(\\cdot\\mid\\rho\_\{i\};\\bm\{\\eta\}\(\\rho\_\{i\};\\bm\{\\theta\},\\bm\{\\varphi\}\)\\right\)\}\[X\]\.\(29\)
##### Weighted Mean Absolute Error \(WMAE\)

The WMAE measures the average magnitude of the absolute errors, weighted by the density bin frequency to prevent bias toward free\-flow states:

Sj,WMAE=∑i∈𝒯jwi​\|q^i−qi\|∑i∈𝒯jwi\.S\_\{j,\\mathrm\{WMAE\}\}=\\frac\{\\sum\_\{i\\in\\mathcal\{T\}\_\{j\}\}w\_\{i\}\|\\hat\{q\}\_\{i\}\-q\_\{i\}\|\}\{\\sum\_\{i\\in\\mathcal\{T\}\_\{j\}\}w\_\{i\}\}\.\(30\)

##### Root Weighted Mean Square Error \(RWMSE\)

To heavily penalize larger prediction errors while accounting for the skewed dataset, we compute the Root Weighted Mean Square Error as follows:

Sj,RWMSE=∑i∈𝒯jwi​\(q^i−qi\)2∑i∈𝒯jwi\.S\_\{j,\\mathrm\{RWMSE\}\}=\\sqrt\{\\frac\{\\sum\_\{i\\in\\mathcal\{T\}\_\{j\}\}w\_\{i\}\(\\hat\{q\}\_\{i\}\-q\_\{i\}\)^\{2\}\}\{\\sum\_\{i\\in\\mathcal\{T\}\_\{j\}\}w\_\{i\}\}\}\.\(31\)

##### Weighted Mean Absolute Percentage Error \(WMAPE\)

To provide a relative error perspective, we define the WMAPE as the ratio of the weighted absolute errors to the weighted absolute true values:

Sj,WMAPE=∑i∈𝒯jwi​\|q^i−qi\|∑i∈𝒯jwi​\|qi\|\.S\_\{j,\\mathrm\{WMAPE\}\}=\\frac\{\\sum\_\{i\\in\\mathcal\{T\}\_\{j\}\}w\_\{i\}\|\\hat\{q\}\_\{i\}\-q\_\{i\}\|\}\{\\sum\_\{i\\in\\mathcal\{T\}\_\{j\}\}w\_\{i\}\|q\_\{i\}\|\}\.\(32\)

### 5\.4Model Setup

During the training phase, the model \(configured with a hidden size of 16\) is updated in mini\-batches of 128 via the Adam optimizer\(Kingma and Ba,[2014](https://arxiv.org/html/2607.15907#bib.bib24)\)\. We employ a weight decay of10−510^\{\-5\}for regularization, alongsideβ\\betacoefficients of\(0\.9,0\.99\)\(0\.9,0\.99\)\. The learning rate is initially set to10−210^\{\-2\}and explicitly decayed to10−310^\{\-3\}halfway through the 200\-epoch training process\. To ensure numerical stability,log⁡Φ​\(⋅\)\\log\\Phi\(\\cdot\)is computed using the optimized functionlog\_ndtravailable in PyTorch\(Paszkeet al\.,[2019](https://arxiv.org/html/2607.15907#bib.bib26)\)and JAX\(Frostiget al\.,[2019](https://arxiv.org/html/2607.15907#bib.bib27)\)\.

### 5\.5Baselines

This study utilizes three recently proposed stochastic FD models as baselines, which are representative of two prevailing modeling strategies: 1\) random parameter modeling and 2\) stochastic error/variance modeling\. A common principle in both strategies is the modification of deterministic FD models\. To obtain accurate deterministic FD models, the weighted least squares method, adopted in\(Quet al\.,[2015](https://arxiv.org/html/2607.15907#bib.bib8)\), is applied to mitigate the influence of unbalanced data during the training phase\.

#### 5\.5\.1Random parameter modeling

The work ofChenget al\.\([2024](https://arxiv.org/html/2607.15907#bib.bib6)\)extends the deterministic S3 model\(Chenget al\.,[2021](https://arxiv.org/html/2607.15907#bib.bib9)\)by introducing stochasticity into its parameters\. Since the original formulation is not explicitly named in the original paper, we refer to it asS3\+LNfor clarity\. The model is defined as follows\. The S3 speed\-density and flow\-density relations are:

vS3​\(ρ\)=vfree\(1\+\(ρρcritical\)m\)2m,\\displaystyle v\_\{\\mathrm\{S3\}\}\(\\rho\)=\\frac\{v\_\{\\mathrm\{free\}\}\}\{\\left\(1\+\\left\(\\frac\{\\rho\}\{\\rho\_\{\\mathrm\{critical\}\}\}\\right\)^\{m\}\\right\)^\{\\frac\{2\}\{m\}\}\},qS3​\(ρ\)=ρ​vfree\(1\+\(ρρcritical\)m\)2m\\displaystyle\\quad q\_\{\\mathrm\{S3\}\}\(\\rho\)=\\frac\{\\rho\\,v\_\{\\mathrm\{free\}\}\}\{\\left\(1\+\\left\(\\frac\{\\rho\}\{\\rho\_\{\\mathrm\{critical\}\}\}\\right\)^\{m\}\\right\)^\{\\frac\{2\}\{m\}\}\}\(33\)In the stochastic formulation, the free\-flow speed and critical speed are treated as random variables following Log\-Normal distributions:

ln⁡\(vfree\)∼𝒩​\(μf,σf2\),\\displaystyle\\ln\(v\_\{\\mathrm\{free\}\}\)\\sim\\mathcal\{N\}\(\\mu\_\{f\},\\sigma^\{2\}\_\{f\}\),ln⁡\(vcritical\)∼𝒩​\(μc,σc2\),\\displaystyle\\quad\\ln\(v\_\{\\mathrm\{critical\}\}\)\\sim\\mathcal\{N\}\(\\mu\_\{c\},\\sigma^\{2\}\_\{c\}\),\(34\)and the parameter follows

1m∼𝒩​\(μf−μc2​ln⁡2,σf2\+σc2\(2​ln⁡2\)2\)\.\\frac\{1\}\{m\}\\sim\\mathcal\{N\}\\left\(\\frac\{\\mu\_\{f\}\-\\mu\_\{c\}\}\{2\\ln 2\},\\frac\{\\sigma^\{2\}\_\{f\}\+\\sigma^\{2\}\_\{c\}\}\{\(2\\ln 2\)^\{2\}\}\\right\)\.\(35\)This framework contains five trainable parameters:\{μf,σf,μc,σc,ρcritical\}\\\{\\mu\_\{f\},\\sigma\_\{f\},\\mu\_\{c\},\\sigma\_\{c\},\\rho\_\{\\mathrm\{critical\}\}\\\}\. The free\-flow speed and critical speed are parameterized by\{μf,σf\}\\\{\\mu\_\{f\},\\sigma\_\{f\}\\\}and\{μc,σc\}\\\{\\mu\_\{c\},\\sigma\_\{c\}\\\}, respectively\. The critical densityρcritical\\rho\_\{\\mathrm\{critical\}\}is estimated using the weighted least squares approach\.

#### 5\.5\.2Stochastic Error/Variance Modeling

The work ofLeiet al\.\([2024](https://arxiv.org/html/2607.15907#bib.bib7)\)extends deterministic fundamental diagram \(FD\) models by modeling the residual error using a Gaussian Process \(GP\)\. In this formulation, the overall speedv​\(ρ\)v\(\\rho\)is represented as the sum of a deterministic baseline modelfdet​\(ρ\)f\_\{\\mathrm\{det\}\}\(\\rho\)and a stochastic residual termϵ​\(ρ\)\\epsilon\(\\rho\):

v​\(ρ\)\\displaystyle v\(\\rho\)=fdet​\(ρ\)\+ϵ​\(ρ\),\\displaystyle=f\_\{\\mathrm\{det\}\}\(\\rho\)\+\\epsilon\(\\rho\),\(36\)ϵ​\(ρ\)\\displaystyle\\epsilon\(\\rho\)∼𝒢​𝒫​\(0,k​\(ρ,ρ′\)\),\\displaystyle\\sim\\mathcal\{GP\}\\bigl\(0,k\(\\rho,\\rho^\{\\prime\}\)\\bigr\),where the error termϵ​\(ρ\)\\epsilon\(\\rho\)follows a zero\-mean Gaussian Process governed by a Radial Basis Function \(RBF\) kernelk​\(ρ,ρ′\)k\(\\rho,\\rho^\{\\prime\}\):

k​\(ρ,ρ′\)=σ2​exp⁡\(−\(ρ−ρ′\)22​l2\)\.k\(\\rho,\\rho^\{\\prime\}\)=\\sigma^\{2\}\\exp\\left\(\-\\frac\{\(\\rho\-\\rho^\{\\prime\}\)^\{2\}\}\{2l^\{2\}\}\\right\)\.\(37\)This kernel introduces two additional trainable hyperparameters: the signal varianceσ2\\sigma^\{2\}and the length\-scalell, which control the magnitude and smoothness of the stochastic component, respectively\. The Greenshields and S3 deterministic models are used as baselines for the GP framework, resulted in two variants of models denoted asGS\+GPandS3\+GP\.

## 6Results

### 6\.1Metric Performance

Table 4:Performance of all evaluated stochastic FD models![Refer to caption](https://arxiv.org/html/2607.15907v1/figs/main_result_models/model_for_experiment_bwnc_wo_cdf/fold_0/fd.png)\(a\)SN\-BwNC
![Refer to caption](https://arxiv.org/html/2607.15907v1/figs/main_result_models/model_for_experiment_bwnc_g_wo_cdf/fold_0/fd.png)\(b\)N\-BwNC
![Refer to caption](https://arxiv.org/html/2607.15907v1/figs/main_result_models/model_for_experiment_qwnc_wo_cdf/fold_0/fd.png)\(c\)SN\-QwNC
![Refer to caption](https://arxiv.org/html/2607.15907v1/figs/main_result_models/model_for_experiment_qwnc_g_wo_cdf/fold_0/fd.png)\(d\)N\-QwNC

Figure 3:Modeling results from Stochastic FD Models under the Semiparametric Framework![Refer to caption](https://arxiv.org/html/2607.15907v1/figs/main_result_models/gp_gs/fold_0.png)\(a\)Greenshields Model with Gaussian Process
![Refer to caption](https://arxiv.org/html/2607.15907v1/figs/main_result_models/gp_s3/fold_0.png)\(b\)S3 Model with Gaussian Process
![Refer to caption](https://arxiv.org/html/2607.15907v1/figs/main_result_models/s3_ln/fold_0.png)\(c\)S3 Model with Log\-Normal Distribution

Figure 4:Modeling results from Baseline Stochastic FD ModelsPerformance metrics for probabilistic estimation across all evaluated models are summarized in Table[4](https://arxiv.org/html/2607.15907#S6.T4)\. Five evaluation metrics are reported for both speed\-density and flow\-density relations\. The reported values represent statistical mean and standard deviation using 5\-fold cross\-validation\. The NLL metric is omitted for theS3\+LNmodel because its probability density function does not admit a close\-form expression\.

Overall, the experimental results reveal that stochastic FD models developed under the proposed semiparametric framework consistently outperform the baseline approaches\. Among all baseline models, theS3\+LNmodel exhibits the poorest performance across most of the metrics\. In contrast, the Gaussian Process \(GP\) regression models achieve substantially better performance relative to theS3\+LNapproach\.

The table also reveals performance differences among the model variants\. For the baseline models, the GP regression based on the deterministic S3 model outperforms the formulation based on the Greenshields model\. This advantage is consistent across cross\-validation folds, as reflected by the relatively small standard deviations\.

Among the models proposed in this study, semiparametric models employing the Skew\-Normal \(SN\) distribution consistently outperform their Normal \(N\) distribution counterparts, indicating that modeling asymmetric uncertainty improves predictive performance\. Furthermore, the BwNC formulation demonstrates slight better mean performance than the QwNC formulation\. However, the considerable overlap in their standard deviations suggests that the performance difference between the two formulations is relatively modest\.

### 6\.2Modeling Analysis

Figure[3](https://arxiv.org/html/2607.15907#S6.F3)illustrates stochastic FD modeling results from semiparametric models proposed in this study, while Figure[4](https://arxiv.org/html/2607.15907#S6.F4)shows the corresponding results from baseline models\. In these visualizations, the red curve represents the conditional expected flow or speed for each density value, while the blue shading areas demonstrate the 90% and 99% confidence intervals\.

The semiparametric models produce physically consistent FD relationships\. Specifically, both flow and speed approach zero when traffic density approaches the estimated jam density\. The estimated jam density is approximately150150veh/km/lane, which is consistent with the typical observations in the traffic environment\. Furthermore, the predicted uncertainty covers the majority of the observed data points and gradually shrink as density approaches the jam density in the congestion regime\.

Furthermore, models based on Skew\-Normal distribution are able to capture the transition between free flow and congestion regimes, which occurs around a density of3535veh/km/lane\. In addition, the SN\-BwNC and N\-QwNC models generate smooth functional relationships without noticeable overfitting, suggesting that the semiparametric formulation provides a good balance between model flexibility and structural regularization\.

The results for theS3\+LNmodel reveal a pronounced bias toward higher values in heavy congestion regimes, as illustrated in Figure[4\(c\)](https://arxiv.org/html/2607.15907#S6.F4.sf3)\. Specifically, in both the speed\-density and flow\-density relationships, the empirical observations in the congested region consistently fall outside the 99% confidence interval of the model estimation\.

A comparison of model parameters betweenS3\+LNand the deterministic S3 model \(Table[5](https://arxiv.org/html/2607.15907#A1.T5)\) indicates that this performance degradation stems from the inherent inability of the deterministic model to characterize heavy congestion regimes\. Treating parameters as random variables primarily serves to introduce uncertainty without substantially improving the underlying fitting quality of the deterministic framework\. Consequently,S3\+LNinherits the structural drawbacks of theS3model, leading to persistent systematic bias in higher density regions\.

In contrast, Gaussian Process \(GP\) regression demonstrates a superior capacity for stochastic modeling, though it encounters challenges in generating consistent uncertainty across the entire density range\. As shown in Figure[4\(b\)](https://arxiv.org/html/2607.15907#S6.F4.sf2), the inherent bias of theS3model is successfully mitigated in theS3\+GPconfiguration\. Similarly, Figure[4\(a\)](https://arxiv.org/html/2607.15907#S6.F4.sf1)displays the results forS3\+GP, where the Gaussian Process effectively introduces necessary nonlinearity to the standard Greenshields speed–density model\. However, the performance in the heavy congestion region remains suboptimal\. The scarcity of training data in high\-density states leads to a significant localized increase in model uncertainty\. This variance is further magnified when transformed into the flow\-density relation, where the multiplicative effect of density results in a wide, undesirable predictive interval that limits the model’s practical utility for congestion forecasting\.

![Refer to caption](https://arxiv.org/html/2607.15907v1/figs/metrics_comparison_2x2.png)Figure 5:CRPS and MAE Comparison on Different Density Regions
### 6\.3Performance on Different Density Regions

The CRPS and MAE comparison of the stochastic FD models under different density regions is presented in Figure[5](https://arxiv.org/html/2607.15907#S6.F5)\. The density is categorized into four regimes: free flow \(0\-20 veh/km\), transition \(20\-60 veh/km\), light congestion \(60\-100 veh/km\), and heavy congestion \(100\-150 veh/km\)\.

Semiparametric models consistently outperform baseline counterparts across all density regimes\. This improvement is most pronounced in the heavy congestion region, where baseline models typically struggle with high variance\.

Regarding the flow\-density relationship, a divergence in performance trends is observed as density increases\. While the CRPS for baseline models degrades significantly in high\-density states, the semiparametric models exhibit enhanced predictive accuracy and stability under heavy congestion\.

For the speed\-density relationship, all models follow a non\-monotonic performance curve in the speed\-density relationship, with the highest error rates \(worst performance\) appeared in the transition region, likely due to the inherent instability of traffic flow during phase shifts\.

The deterministic performance of semiparametric SFD models is comparable to that of the Gaussian Process \(GP\) models\. This suggests that the potential driver for improved stochastic modeling is the integration of physical constraints, which effectively regularizes the model’s uncertainty in high\-congestion regions\.

### 6\.4Ablation on Regularization Loss

![Refer to caption](https://arxiv.org/html/2607.15907v1/figs/regulation_ablation.png)Figure 6:The convergence patterns during model training with v\.s\. without regularization loss in the objective function\.The influence of the regularization termℓReg\\ell\_\{\\mathrm\{Reg\}\}on training stability is evaluated in Figure[6](https://arxiv.org/html/2607.15907#S6.F6), which depicts the negative log\-likelihood \(NLL\) loss across five cross\-validation folds\. The blue and orange curves represent the training trajectories with and without the inclusion ofℓReg\\ell\_\{\\mathrm\{Reg\}\}, respectively\.

The experimental results indicate thatℓReg\\ell\_\{\\mathrm\{Reg\}\}is indispensable for numerical stability in both the SN\-QwNC and SN\-BwNC architectures\. In the absence of this regularization term, both models exhibit highly volatile loss curves and a significant reduction in convergence speed\. Notably, the SN\-BwNC model fails to achieve convergence in two out of the five folds whenℓReg\\ell\_\{\\mathrm\{Reg\}\}is removed\. These ablation results confirm that the proposed regularization loss is a critical component for ensuring both the stability and the overall effectiveness of the stochastic training process\.

## 7Discussion

The effectiveness of the proposed framework relies on its ability to propagate first and second\-moment constraints to the distribution parameters, specifically the constrained parameters𝜼con\\bm\{\\eta\}\_\{\\mathrm\{con\}\}\. Therefore, after determining constrained parameters, the primary theoretical challenge lies in the existence of a solution for the moment\-matching system in Eq\. \([4](https://arxiv.org/html/2607.15907#S3.E4)\)\. For a distributionffand values0<m,s<∞0<m,s<\\infty, the problem simplifies to:

\{𝔼X∼f​\[X\]=mVarX∼f​\[X\]=s2\\begin\{cases\}\\ \\mathbb\{E\}\_\{X\\sim f\}\[X\]&=m\\\\ \\ \\mathrm\{Var\}\_\{X\\sim f\}\[X\]&=s^\{2\}\\\\ \\end\{cases\}\(38\)
As established in Proposition[1](https://arxiv.org/html/2607.15907#Thmproposition1), any distribution belonging to a location\-scale family guarantees a unique solution to this system\. However, the applicability of our framework extends beyond this family: while some distributions naturally yield unique solutions, others become feasible once appropriate parameter constraints are imposed\.

### 7\.1Feasible Distribution from Non\-Location\-scale Family

#### 7\.1\.1Gamma Distribution

A primary example of a feasible non\-location\-scale distribution is the Gamma distribution, with probability density function as:

f​\(x;α,θ\)=xα−1​e−x/θθα​Γ​\(α\),for​x\>0,f\(x;\\alpha,\\theta\)=\\frac\{x^\{\\alpha\-1\}e^\{\-x/\\theta\}\}\{\\theta^\{\\alpha\}\\Gamma\(\\alpha\)\},\\quad\\text\{for \}x\>0,\(39\)whereΓ​\(⋅\)\\Gamma\(\\cdot\)is the gamma function, and\{α,θ\}\\\{\\alpha,\\theta\\\}are positive values\. The first two moments are defined as:

𝔼​\[X\]=α​θ,Var​\(X\)=α​θ2\.\\mathbb\{E\}\[X\]=\\alpha\\theta,\\qquad\\mathrm\{Var\}\(X\)=\\alpha\\theta^\{2\}\.\(40\)By defining the constrained parameters as𝜼con=\[α,θ\]𝖳\\bm\{\\eta\}\_\{\\mathrm\{con\}\}=\[\\alpha,\\ \\theta\]^\{\\mathsf\{T\}\}, we obtain a unique solution for the system of equations in Eq\. \([38](https://arxiv.org/html/2607.15907#S7.E38)\):

θ=s2m,α=m2s2\.\\theta=\\frac\{s^\{2\}\}\{m\},\\qquad\\alpha=\\frac\{m^\{2\}\}\{s^\{2\}\}\.\(41\)

#### 7\.1\.2Log\-Normal Distribution

The Log\-Normal distribution is another widely used non\-location\-scale family that fits our framework\. With its probability density function defined by

f​\(x;μ,σ\)=1x​σ​2​π​exp⁡\(−\(ln⁡x−μ\)22​σ2\),for​x\>0,f\(x;\\mu,\\sigma\)=\\frac\{1\}\{x\\sigma\\sqrt\{2\\pi\}\}\\exp\\left\(\-\\frac\{\(\\ln x\-\\mu\)^\{2\}\}\{2\\sigma^\{2\}\}\\right\),\\quad\\text\{for \}x\>0,\(42\)whereμ∈ℝ\\mu\\in\\mathbb\{R\}andσ\>0\\sigma\>0, its first two moments are:

𝔼​\[X\]=eμ\+σ2/2,Var​\(X\)=\(eσ2−1\)​e2​μ\+σ2\.\\mathbb\{E\}\[X\]=e^\{\\mu\+\\sigma^\{2\}/2\},\\quad\\mathrm\{Var\}\(X\)=\(e^\{\\sigma^\{2\}\}\-1\)e^\{2\\mu\+\\sigma^\{2\}\}\.\(43\)Here, the constrained parameters are defined as𝜼con=\[μ,σ\]𝖳\\bm\{\\eta\}\_\{\\mathrm\{con\}\}=\[\\mu,\\ \\sigma\]^\{\\mathsf\{T\}\}\. To adapt this to our framework, we substitute these theoretical moments into Eq\. \([38](https://arxiv.org/html/2607.15907#S7.E38)\)\. This yields a unique, closed\-form solution for the parameters:

σ=\(ln⁡\(1\+s2m2\)\)12,μ=ln⁡m−12​ln⁡\(1\+s2m2\)\.\\displaystyle\\sigma=\\left\(\\ln\\\!\\left\(1\+\\frac\{s^\{2\}\}\{m^\{2\}\}\\right\)\\right\)^\{\\frac\{1\}\{2\}\},\\qquad\\mu=\\ln m\-\\frac\{1\}\{2\}\\ln\\\!\\left\(1\+\\frac\{s^\{2\}\}\{m^\{2\}\}\\right\)\.\(44\)

### 7\.2Distributions Requiring Additional Constraints

#### 7\.2\.1Truncated Distributions

While many distributions inherently support the moment\-matching system, certain families, such as truncated distributions, require additional mathematical constraints to remain feasible\. As demonstrated byBhatia and Davis \([2000](https://arxiv.org/html/2607.15907#bib.bib39)\), the variance of any bounded probability distribution is strictly limited by its expectation, as well as its absolute lower boundLLand upper boundUU\. This relationship is expressed as:

Var​\[X\]≤\(𝔼​\[X\]−L\)​\(U−𝔼​\[X\]\)\.\\mathrm\{Var\}\[X\]\\leq\(\\mathbb\{E\}\[X\]\-L\)\(U\-\\mathbb\{E\}\[X\]\)\.\(45\)Consequently, the condition0<s<∞0<s<\\inftydoes not guarantee a solution for the system in Eq\. \([38](https://arxiv.org/html/2607.15907#S7.E38)\), particularly whenssis excessively large\. Translating the mathematical limitation into our physical constraints, an upper bound is needed for the standard deviation:

\{𝔼q∼p​\(q∣ρ\)​\[q\]=𝕊q∼p​\(q∣ρ\)​\[q\]=0,ρ∈\{0,ρjam\},0<𝔼q∼p​\(q∣ρ\)​\[q\]<∞,ρ∈\(0,ρjam\),0<𝕊q∼p​\(q∣ρ\)​\[q\]≤Smax​\(ρ\),ρ∈\(0,ρjam\)\.\\begin\{cases\}\\mathbb\{E\}\_\{q\\sim p\(q\\mid\\rho\)\}\[q\]=\\mathbb\{S\}\_\{q\\sim p\(q\\mid\\rho\)\}\[q\]=0,&\\rho\\in\\\{0,\\rho\_\{\\mathrm\{jam\}\}\\\},\\\\ 0<\\mathbb\{E\}\_\{q\\sim p\(q\\mid\\rho\)\}\[q\]<\\infty,&\\rho\\in\(0,\\rho\_\{\\mathrm\{jam\}\}\),\\\\ 0<\\mathbb\{S\}\_\{q\\sim p\(q\\mid\\rho\)\}\[q\]\\leq S\_\{\\max\}\(\\rho\),&\\rho\\in\(0,\\rho\_\{\\mathrm\{jam\}\}\)\.\\end\{cases\}\(46\)Smax​\(ρ\)S\_\{\\max\}\(\\rho\)represents its theoretical upper bound derived from Eq\. \([45](https://arxiv.org/html/2607.15907#S7.E45)\), given by:

Smax​\(ρ\)=\(𝔼q∼p​\(q∣ρ\)​\[q\]−L​\(ρ\)\)​\(U​\(ρ\)−𝔼q∼p​\(q∣ρ\)​\[q\]\),S\_\{\\max\}\(\\rho\)=\\sqrt\{\\left\(\\mathbb\{E\}\_\{q\\sim p\(q\\mid\\rho\)\}\[q\]\-L\(\\rho\)\\right\)\\left\(U\(\\rho\)\-\\mathbb\{E\}\_\{q\\sim p\(q\\mid\\rho\)\}\[q\]\\right\)\},\(47\)whereL​\(ρ\)L\(\\rho\)andU​\(ρ\)U\(\\rho\)denote the physical lower bound \(e\.g\., minimum flow\) and upper bound \(i\.e\., maximum capacity\) of flowqqat densityρ\\rho, respectively\. To guarantee the existence of a valid solution for the moment\-matching equations, the functional design of the standard deviations​\(ρ\)s\(\\rho\)must explicitly depend on the expectation functionm​\(ρ\)m\(\\rho\), as well as the local boundary functionsL​\(ρ\)L\(\\rho\)andU​\(ρ\)U\(\\rho\)\.

#### 7\.2\.2Gaussian Mixture Model

Another complex scenario arises when employing mixture distributions, such as a two\-component Gaussian Mixture Model \(GMM\)\. The probability density function is given by:

p​\(x\)=k​𝒩​\(x;μ1,σ12\)\+\(1−k\)​𝒩​\(x;μ2,σ22\),p\(x\)=k\\mathcal\{N\}\(x;\\mu\_\{1\},\\sigma\_\{1\}^\{2\}\)\+\(1\-k\)\\mathcal\{N\}\(x;\\mu\_\{2\},\\sigma\_\{2\}^\{2\}\),\(48\)wherek∈\(0,1\)k\\in\(0,1\)is the mixing weight\. The theoretical mean and variance can be analytically expressed as:

𝔼​\[X\]\\displaystyle\\mathbb\{E\}\[X\]=k​μ1\+\(1−k\)​μ2,\\displaystyle=k\\mu\_\{1\}\+\(1\-k\)\\mu\_\{2\},\(49\)Var​\(X\)\\displaystyle\\mathrm\{Var\}\(X\)=k​σ12\+\(1−k\)​σ22\+k​\(1−k\)​\(μ1−μ2\)2\.\\displaystyle=k\\sigma\_\{1\}^\{2\}\+\(1\-k\)\\sigma\_\{2\}^\{2\}\+k\(1\-k\)\(\\mu\_\{1\}\-\\mu\_\{2\}\)^\{2\}\.The parameters are partitioned into constrained parameters𝜼con=\[μ1,μ2\]𝖳\\bm\{\\eta\}\_\{\\mathrm\{con\}\}=\[\\mu\_\{1\},\\ \\mu\_\{2\}\]^\{\\mathsf\{T\}\}, and unconstrained parameters𝜼free=\[k,σ1,σ2\]𝖳\\bm\{\\eta\}\_\{\\mathrm\{free\}\}=\[k,\\ \\sigma\_\{1\},\\ \\sigma\_\{2\}\]^\{\\mathsf\{T\}\}\. Equating the moments tommands2s^\{2\}provides the following symmetric closed\-form solutions:

μ1=m±1−kk​Δ,μ2=m∓k1−k​Δ,\\mu\_\{1\}=m\\pm\\sqrt\{\\frac\{1\-k\}\{k\}\\Delta\},\\qquad\\mu\_\{2\}=m\\mp\\sqrt\{\\frac\{k\}\{1\-k\}\\Delta\}\\ ,\(50\)whereΔ=s2−\[k​σ12\+\(1−k\)​σ22\]\>0\\Delta=s^\{2\}\-\\left\[k\\sigma\_\{1\}^\{2\}\+\(1\-k\)\\sigma\_\{2\}^\{2\}\\right\]\>0\. A real solution necessitatesΔ\>0\\Delta\>0, thereby introducing a lower bound for the standard deviation\. Consequently, the physical constraints must be adjusted:

\{𝔼q∼p​\(q∣ρ\)​\[q\]=𝕊q∼p​\(q∣ρ\)​\[q\]=0,ρ∈\{0,ρjam\},0<𝔼q∼p​\(q∣ρ\)​\[q\]<∞,ρ∈\(0,ρjam\),Smin​\(ρ\)<𝕊q∼p​\(q∣ρ\)​\[q\]<∞,ρ∈\(0,ρjam\)\.\\begin\{cases\}\\mathbb\{E\}\_\{q\\sim p\(q\\mid\\rho\)\}\[q\]=\\mathbb\{S\}\_\{q\\sim p\(q\\mid\\rho\)\}\[q\]=0,&\\rho\\in\\\{0,\\rho\_\{\\mathrm\{jam\}\}\\\},\\\\ 0<\\mathbb\{E\}\_\{q\\sim p\(q\\mid\\rho\)\}\[q\]<\\infty,&\\rho\\in\(0,\\rho\_\{\\mathrm\{jam\}\}\),\\\\ S\_\{\\min\}\(\\rho\)<\\mathbb\{S\}\_\{q\\sim p\(q\\mid\\rho\)\}\[q\]<\\infty,&\\rho\\in\(0,\\rho\_\{\\mathrm\{jam\}\}\)\.\\end\{cases\}\(51\)whereSmin​\(ρ\)=k​\(ρ\)​σ12​\(ρ\)\+\(1−k​\(ρ\)\)​σ22​\(ρ\)S\_\{\\min\}\(\\rho\)=\\sqrt\{k\(\\rho\)\\sigma\_\{1\}^\{2\}\(\\rho\)\+\(1\-k\(\\rho\)\)\\sigma\_\{2\}^\{2\}\(\\rho\)\}\. Ultimately, the design of the standard deviation functions​\(ρ\)s\(\\rho\)must carefully accommodate the selected free parameters𝜼free\\bm\{\\eta\}\_\{\\mathrm\{free\}\}to ensure the conditionΔ\>0\\Delta\>0is continuously satisfied\.

## 8Conclusion

In this study, we proposed a novel semiparametric framework for modeling stochastic fundamental diagrams\. By partitioning distribution parameters into a constrained component and a data\-driven neural component, our framework reconciles physical consistency with empirical flexibility\. This semiparametric design ensures that essential traffic properties, such as boundary conditions at zero and jam density, are satisfied by construction while capturing the complex stochastic patterns inherent in traffic data\.

From a theoretical perspective, we established that the underlying moment\-matching system admits a unique solution for location–scale families, guaranteeing the well\-posedness of the parameterization\. We further demonstrated the framework’s extensibility by deriving feasibility conditions for non\-location–scale families, including truncated and mixture\-based distributions in the final discussion\.

Empirical evaluation using the MAGIC dataset demonstrates that our semiparametric approach consistently outperforms baseline models\. Specifically, the proposed models provide superior probabilistic accuracy and robust uncertainty quantification, particularly in congested regimes where the conventional stochastic extension of deterministic models often exhibit systematic bias\. Our results further highlight that incorporating asymmetric distributions significantly enhances the model’s ability to capture observed empirical variance\.

In summary, this framework provides a flexible, theoretically grounded foundation for stochastic FD modeling\. Future research will explore the integration of diverse distribution families, the extension of the model to multi\-regime and multi\-lane dynamics, and the inclusion of spatiotemporal dependencies to further improve the fidelity of stochastic traffic flow representations\.

## Acknowledgment

The authors thank the Swedish Transport Administration for supporting this work\.

## Appendix AParameters for the Log\-Normal\-Based S3 Model

Table[5](https://arxiv.org/html/2607.15907#A1.T5)shows marginal difference between theS3model parameters obtained using the original method\(Chenget al\.,[2021](https://arxiv.org/html/2607.15907#bib.bib9)\)\(i\.e\., the weighted least squares method\) and those derived from the method byChenget al\.\([2024](https://arxiv.org/html/2607.15907#bib.bib6)\)\.

Table 5:Parameter Value Comparison between Deterministic Model and Log\-Normal\-Based Stochastic Model
## References

- A\. Ahmed, D\. Ngoduy, M\. Adnan, and M\. A\. U\. Baig \(2021\)On the fundamental diagram and driving behavior modeling of heterogeneous traffic flow using UAV\-based data\.Transportation Research Part A: Policy and Practice148,pp\. 100–115\.Cited by:[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p1.1)\.
- L\. Bai, S\. C\. Wong, P\. Xu, A\. H\. F\. Chow, and W\. H\. K\. Lam \(2021\)Calibration of stochastic link\-based fundamental diagram with explicit consideration of speed heterogeneity\.Transportation Research Part B: Methodological150,pp\. 524–539\.Cited by:[§1](https://arxiv.org/html/2607.15907#S1.p3.1),[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p1.1)\.
- R\. Bhatia and C\. Davis \(2000\)A better bound on the variance\.The American Mathematical Monthly107\(4\),pp\. 353–357\.Cited by:[§7\.2\.1](https://arxiv.org/html/2607.15907#S7.SS2.SSS1.p1.2)\.
- C\. M\. Bishop \(1994\)Mixture density networks\.Technical reportAston University\.Cited by:[§3\.2](https://arxiv.org/html/2607.15907#S3.SS2.p1.1)\.
- D\. M\. Bramich, M\. Menendez, and L\. Ambühl \(2023\)FitFun: a modelling framework for successfully capturing the functional form and noise of observed traffic flow–density–speed relationships\.Transportation Research Part C: Emerging Technologies151,pp\. 104068\.Cited by:[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p3.1)\.
- Q\. Cheng, Y\. Lin, X\. S\. Zhou, and Z\. Liu \(2024\)Analytical formulation for explaining the variations in traffic states: a fundamental diagram modeling perspective with stochastic parameters\.European Journal of Operational Research312\(1\),pp\. 182–197\.Cited by:[Appendix A](https://arxiv.org/html/2607.15907#A1.p1.1),[§1](https://arxiv.org/html/2607.15907#S1.p3.1),[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p1.1),[§5\.5\.1](https://arxiv.org/html/2607.15907#S5.SS5.SSS1.p1.5),[Table 4](https://arxiv.org/html/2607.15907#S6.T4.48.48.5),[Table 4](https://arxiv.org/html/2607.15907#S6.T4.9.9.5)\.
- Q\. Cheng, Z\. Liu, Y\. Lin, and X\. S\. Zhou \(2021\)An S\-shaped three\-parameter \(S3\) traffic stream model with consistent car following relationship\.Transportation Research Part B: Methodological153,pp\. 246–271\.Cited by:[Appendix A](https://arxiv.org/html/2607.15907#A1.p1.1),[Table 1](https://arxiv.org/html/2607.15907#S2.T1.3.3.3),[§5\.5\.1](https://arxiv.org/html/2607.15907#S5.SS5.SSS1.p1.5)\.
- Z\. Cheng, X\. Wang, X\. Chen, M\. Trépanier, and L\. Sun \(2022\)Bayesian calibration of traffic flow fundamental diagrams using Gaussian processes\.IEEE Open Journal of Intelligent Transportation Systems3,pp\. 763–771\.Cited by:[§1](https://arxiv.org/html/2607.15907#S1.p3.1),[§1](https://arxiv.org/html/2607.15907#S1.p4.1),[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p2.1)\.
- J\. M\. Del Castillo and F\. G\. Benitez \(1995\)On the functional form of the speed\-density relationship—II: empirical investigation\.Transportation Research Part B: Methodological29\(5\),pp\. 391–406\.Cited by:[Table 1](https://arxiv.org/html/2607.15907#S2.T1.6.6.3)\.
- R\. Frostig, M\. J\. Johnson, and C\. Leary \(2019\)Compiling machine learning programs via high\-level tracing\.InSysML Conference 2018,Cited by:[§5\.4](https://arxiv.org/html/2607.15907#S5.SS4.p1.6)\.
- T\. Gneiting and A\. E\. Raftery \(2007\)Strictly proper scoring rules, prediction, and estimation\.Journal of the American Statistical Association102\(477\),pp\. 359–378\.External Links:[Document](https://dx.doi.org/10.1198/016214506000001437)Cited by:[§5\.3\.1](https://arxiv.org/html/2607.15907#S5.SS3.SSS1.Px1.p1.6)\.
- H\. Greenberg \(1959\)An analysis of traffic flow\.Operations Research7\(1\),pp\. 79–85\.Cited by:[§1](https://arxiv.org/html/2607.15907#S1.p2.1),[Table 1](https://arxiv.org/html/2607.15907#S2.T1.8.8.3)\.
- B\. D\. Greenshields, J\. R\. Bibbins, W\. S\. Channing, and H\. H\. Miller \(1935\)A study of traffic capacity\.InHighway Research Board Proceedings,Vol\.14,Washington, D\.C\.,pp\. 448–477\.Cited by:[§1](https://arxiv.org/html/2607.15907#S1.p2.1),[§2\.1](https://arxiv.org/html/2607.15907#S2.SS1.p1.1),[Table 1](https://arxiv.org/html/2607.15907#S2.T1.1.1.3)\.
- D\. Hendrycks and K\. Gimpel \(2016\)Gaussian error linear units \(GELUs\)\.arXiv preprint arXiv:1606\.08415\.Cited by:[§4\.1](https://arxiv.org/html/2607.15907#S4.SS1.p3.2)\.
- K\. Hornik, M\. Stinchcombe, and H\. White \(1989\)Multilayer feedforward networks are universal approximators\.Neural Networks2\(5\),pp\. 359–366\.Cited by:[§3\.2](https://arxiv.org/html/2607.15907#S3.SS2.p3.2)\.
- S\. E\. Jabari, J\. Zheng, and H\. X\. Liu \(2014\)A probabilistic stationary speed–density relation based on Newell’s simplified car\-following model\.Transportation Research Part B: Methodological68,pp\. 205–223\.Cited by:[§1](https://arxiv.org/html/2607.15907#S1.p3.1),[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p1.1)\.
- R\. Jayakrishnan, W\. K\. Tsai, and A\. Chen \(1995\)A dynamic traffic assignment model with traffic\-flow relationships\.Transportation Research Part C: Emerging Technologies3\(1\),pp\. 51–72\.Cited by:[Table 1](https://arxiv.org/html/2607.15907#S2.T1.2.2.3)\.
- D\. P\. Kingma and J\. Ba \(2014\)Adam: a method for stochastic optimization\.arXiv preprint arXiv:1412\.6980\.Cited by:[§5\.4](https://arxiv.org/html/2607.15907#S5.SS4.p1.6)\.
- Y\. Lei, Y\. Gong, and X\. T\. Yang \(2024\)Unraveling stochastic fundamental diagrams with empirical knowledge: modeling, limitations, and future directions\.Transportation Research Part C: Emerging Technologies169,pp\. 104851\.Cited by:[§1](https://arxiv.org/html/2607.15907#S1.p3.1),[§1](https://arxiv.org/html/2607.15907#S1.p4.1),[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p2.1),[§5\.5\.2](https://arxiv.org/html/2607.15907#S5.SS5.SSS2.p1.3),[Table 4](https://arxiv.org/html/2607.15907#S6.T4.14.14.6),[Table 4](https://arxiv.org/html/2607.15907#S6.T4.19.19.6),[Table 4](https://arxiv.org/html/2607.15907#S6.T4.53.53.6),[Table 4](https://arxiv.org/html/2607.15907#S6.T4.58.58.6)\.
- Z\. Liu, C\. Lyu, Z\. Wang, S\. Wang, P\. Liu, and Q\. Meng \(2023\)A Gaussian\-process\-based data\-driven traffic flow model and its application in road capacity analysis\.IEEE Transactions on Intelligent Transportation Systems24\(2\),pp\. 1544–1563\.Cited by:[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p2.1)\.
- W\. Ma, H\. Zhong, L\. Wang, L\. Jiang, and M\. Abdel\-Aty \(2022\)MAGIC dataset: multiple conditions unmanned aerial vehicle group\-based high\-fidelity comprehensive vehicle trajectory dataset\.Transportation Research Record2676\(5\),pp\. 793–805\.Cited by:[§5\.1](https://arxiv.org/html/2607.15907#S5.SS1.p1.1)\.
- G\. F\. Newell \(1961\)Nonlinear effects in the dynamics of car following\.Operations Research9\(2\),pp\. 209–229\.Cited by:[§1](https://arxiv.org/html/2607.15907#S1.p2.1),[Table 1](https://arxiv.org/html/2607.15907#S2.T1.4.4.3)\.
- D\. Ngoduy \(2011\)Multiclass first\-order traffic model using stochastic fundamental diagrams\.Transportmetrica7\(2\),pp\. 111–125\.Cited by:[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p3.1)\.
- D\. Ni, H\. K\. Hsieh, and T\. Jiang \(2018\)Modeling phase diagrams as stochastic processes with application in vehicular traffic flow\.Applied Mathematical Modelling53,pp\. 106–117\.Cited by:[§1](https://arxiv.org/html/2607.15907#S1.p3.1),[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p2.1)\.
- D\. Ni, J\. D\. Leonard, C\. Jia, and J\. Wang \(2016\)Vehicle longitudinal control and traffic stream modeling\.Transportation Science50\(3\),pp\. 1016–1031\.Cited by:[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p2.1)\.
- D\. Ni \(2015\)Traffic flow theory: characteristics, experimental methods, and numerical techniques\.Butterworth\-Heinemann\.Cited by:[§1](https://arxiv.org/html/2607.15907#S1.p4.1)\.
- M\. Papageorgiou, J\. Blosseville, and H\. Hadj\-Salem \(1989\)Macroscopic modelling of traffic flow on the Boulevard Périphérique in Paris\.Transportation Research Part B: Methodological23\(1\),pp\. 29–47\.Cited by:[Table 1](https://arxiv.org/html/2607.15907#S2.T1.5.5.3)\.
- A\. Paszke, S\. Gross, F\. Massa, A\. Lerer, J\. Bradbury, G\. Chanan, T\. Killeen, Z\. Lin, N\. Gimelshein, L\. Antiga,et al\.\(2019\)PyTorch: an imperative style, high\-performance deep learning library\.Advances in Neural Information Processing Systems32\.Cited by:[§5\.4](https://arxiv.org/html/2607.15907#S5.SS4.p1.6)\.
- X\. Qu, S\. Wang, and J\. Zhang \(2015\)On the fundamental diagram for freeway traffic: a novel calibration approach for single\-regime models\.Transportation Research Part B: Methodological73,pp\. 91–102\.Cited by:[§5\.5](https://arxiv.org/html/2607.15907#S5.SS5.p1.1)\.
- X\. Qu, J\. Zhang, and S\. Wang \(2017\)On the stochastic fundamental diagram for freeway traffic: model development, analytical properties, validation, and extensive applications\.Transportation Research Part B: Methodological104,pp\. 256–271\.Cited by:[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p1.1)\.
- R\. A\. Rigby and D\. M\. Stasinopoulos \(2005\)Generalized additive models for location, scale and shape\.Journal of the Royal Statistical Society Series C: Applied Statistics54\(3\),pp\. 507–554\.Cited by:[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p3.1)\.
- R\. Shi, Z\. Mo, K\. Huang, X\. Di, and Q\. Du \(2021\)A physics\-informed deep learning paradigm for traffic state and fundamental diagram estimation\.IEEE Transactions on Intelligent Transportation Systems23\(8\),pp\. 11688–11698\.Cited by:[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p3.1)\.
- P\. J\. Storm, M\. Mandjes, and B\. van Arem \(2022\)Efficient evaluation of stochastic traffic flow models using Gaussian process approximation\.Transportation Research Part B: Methodological164,pp\. 126–144\.Cited by:[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p3.1)\.
- L\. Theis and M\. Bethge \(2015\)Generative image modeling using spatial LSTMs\.Advances in Neural Information Processing Systems28\.Cited by:[§3\.2](https://arxiv.org/html/2607.15907#S3.SS2.p1.1)\.
- H\. Wang, J\. Li, Q\. Chen, and D\. Ni \(2011\)Logistic modeling of the equilibrium speed–density relationship\.Transportation Research Part A: Policy and Practice45\(6\),pp\. 554–566\.Cited by:[Table 1](https://arxiv.org/html/2607.15907#S2.T1.7.7.3)\.
- S\. Wang, X\. Chen, and X\. Qu \(2021\)Model on empirically calibrating stochastic traffic flow fundamental diagram\.Communications in Transportation Research1,pp\. 100015\.Cited by:[§1](https://arxiv.org/html/2607.15907#S1.p4.1),[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p1.1)\.
- X\. Zhang, J\. Sun, and J\. Sun \(2025\)On the stochastic fundamental diagram: a general micro\-macroscopic traffic flow modeling framework\.Communications in Transportation Research5,pp\. 100163\.Cited by:[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p3.1)\.
- J\. Zhou and F\. Zhu \(2020\)Modeling the fundamental diagram of mixed human\-driven and connected automated vehicles\.Transportation Research Part C: Emerging Technologies115,pp\. 102614\.Cited by:[§1](https://arxiv.org/html/2607.15907#S1.p3.1),[§2\.2](https://arxiv.org/html/2607.15907#S2.SS2.p1.1)\.

Similar Articles

Generative Diffusion Models of Stochastic Graph Signals

arXiv cs.LG

This paper proposes a unified denoising diffusion framework for conditional generation of graph signals, introducing a novel U-GNN architecture that extends U-Net to graph-structured data. The method is demonstrated on stock price forecasting and wireless resource allocation tasks.

A Graph-Based Control Interface for Traffic Signals on Heterogeneous Road Networks

arXiv cs.LG

This paper presents a graph-based traffic signal control interface using a shared graph neural network to assign scores to movements, with deterministic phase construction via incidence matrices. Experiments evaluate transfer across synthetic and city road networks, showing feasibility but sensitivity to distribution shifts.

Towards Continuous-time Causal Foundation Models

arXiv cs.LG

Proposes a continuity criterion for extending discrete-time causal prior-data fitted networks to continuous time using stochastic differential equations, introducing a taxonomy and fine-grid integration method that outperforms naive integration on irregular observation schedules.