SeT-Diff: Towards Semantic Foundation Models for HPC Telemetry and Time-Series
Summary
SeT-Diff proposes the first foundation model for HPC telemetry, using diffusion conditioned on semantic sensor descriptions to enable zero-shot generalization across tasks like imputation, forecasting, and virtual sensing, achieving an MAE of 0.0470 on reconstruction.
View Cached Full Text
Cached at: 07/28/26, 06:23 AM
# SeT-Diff: Towards Semantic Foundation Models for HPC Telemetry and Time-Series
Source: [https://arxiv.org/html/2607.22548](https://arxiv.org/html/2607.22548)
\(2026\)
###### Abstract\.
Data centers and their compute nodes require accurate and flexible digital twins capable of modeling the complex interplay of workloads, environmental parameters, and physical metrics\. Current machine learning approaches for HPC and its telemetry typically rely on a static subset of anonymous, fixed\-position sensor variables tailored to single tasks\. Consequently, these models become obsolete when target tasks change or sensor metrics vary\. We propose SeT\-Diff, the first foundational model for compute node telemetry and time\-series\. Unlike rigid architectures, our diffusion\-based approach conditions the generative process on each sensor’s semantic description, decoupling the system dynamics from the structure of the dataset\. Experiments on a real\-world supercomputer dataset demonstrate a Mean Absolute Error \(MAE\) of 0\.0470 on reconstruction tasks\. SeT\-Diff exhibits zero\-shot permutation stability, maintaining accuracy with negligible degradation even when sensors are shuffled\. A single pre\-trained model effectively performs data imputation, forecasting, and virtual sensing \- achieving a 0\.033 MAE in thermal inference \- making SeT\-Diff an effective data\-driven digital twin for HPC systems\.
HPC Telemetry, Diffusion Models, Semantic Conditioning, Foundation Models, Virtual Sensing, Zero\-Shot Adaptation
††journalyear:2026††copyright:cc††conference:Proceedings of the 23rd ACM International Conference on Computing Frontiers; May 19–21, 2026; Catania, Italy††booktitle:Proceedings of the 23rd ACM International Conference on Computing Frontiers \(CF ’26\), May 19–21, 2026, Catania, Italy††doi:10\.1145/3801487\.3806064††isbn:979\-8\-4007\-2568\-5/2026/05††ccs:Computing methodologies Model development and analysis††ccs:Computing methodologies Artificial intelligence††ccs:Computing methodologies Multi\-task learning## 1\.Introduction
High\-performance computing \(HPC\) systems generate continuous streams of multivariate telemetry data, instrumental for system management\(Borghesi and others,[2023](https://arxiv.org/html/2607.22548#bib.bib1); Antici and others,[2025](https://arxiv.org/html/2607.22548#bib.bib13)\), predictive maintenance\(Lima and others,[2021](https://arxiv.org/html/2607.22548#bib.bib17)\), anomaly detection\(Molan and others,[2023](https://arxiv.org/html/2607.22548#bib.bib19)\), and digital twins\(Maiterth and others,[2024](https://arxiv.org/html/2607.22548#bib.bib15),[2025](https://arxiv.org/html/2607.22548#bib.bib16)\)\. However, modeling this data is challenging due to high dimensionality, missing values\(Medaiyese and others,[2025](https://arxiv.org/html/2607.22548#bib.bib20)\), and dynamically changing sensor layouts\. Traditional time series models treat input signals as anonymous, fixed\-position numerical vectors, causing them to fail when configurations vary\.
To address this semantic gap, we propose SeT\-Diff, a context\-aware foundation model based on the denoising diffusion framework\. SeT\-Diff conditions the generation of multivariate time\-series not only on historical data but also on the semantic context of the sensors\. By integrating pre\-trained textual embeddings \(via Sentence\-BERT\(Reimers and Gurevych,[2019](https://arxiv.org/html/2607.22548#bib.bib14)\)\) describing each feature, the model associates a signal’s physical behavior with its meaning rather than its positional index\.
Our core contribution is a unified framework capable of zero\-shot generalization across multiple downstream tasks—including data imputation, system load forecasting, and unmeasured metric regression \(virtual sensing\)—without architectural changes\. Evaluated on the Marconi100 supercomputer dataset\(Borghesi and others,[2023](https://arxiv.org/html/2607.22548#bib.bib1)\), SeT\-Diff achieves a MAE of 0\.0470 in reconstruction tasks\. Furthermore, it maintains this accuracy indistinguishably between fixed and randomly shuffled sensor layouts \(OLS slope of 0\.99\) and achieves an MAE of 0\.033 in zero\-shot thermal regression\. This establishes SeT\-Diff as a versatile and robust digital twin for HPC systems\.
## 2\.Background and Related Work
As HPC systems scale to support generative Artificial Intelligence \(AI\) and gigawatt\-scale data centers, the complexity of their monitoring infrastructure has grown exponentially\. Modern Tier\-0 systems integrate diverse accelerators, multi\-tier memory hierarchies, and liquid cooling, generating gigabytes of high\-frequency telemetry daily\(Borghesi and others,[2023](https://arxiv.org/html/2607.22548#bib.bib1); Jadon and others,[2021](https://arxiv.org/html/2607.22548#bib.bib5)\)\. However, the usage of this data is currently limited to dashboarding and visualization\.
AI adoption in system monitoring has been fragmented into specialized models tailored to specific tasks, such as autoencoders for anomaly detection\(Molan and others,[2023](https://arxiv.org/html/2607.22548#bib.bib19)\)or graph networks for thermal regression\(Guindani and others,[2024](https://arxiv.org/html/2607.22548#bib.bib22)\)\. These supervised and unsupervised models suffer from a ”semantic gap”: they map anonymous numerical vectors based strictly on positional indices\. If the monitored metrics change order or composition, these rigid pipelines become obsolete\. Futhermore, these approaches are single task and trained on specific output features which cannot be changed at inference time\.
To move beyond single\-task models, recent works on Time Series Foundation Models, such as Time\-LLM\(Jin and others,[2024](https://arxiv.org/html/2607.22548#bib.bib6)\)and Chronos\(Ansari and others,[2024](https://arxiv.org/html/2607.22548#bib.bib7)\), adapt Transformer backbones for broad generalization\. However, they function as deterministic mappings lacking the unified generative capabilities required for complex multi\-task scenarios\. Conversely, Denoising Diffusion Probabilistic Models\(Ho and others,[2020](https://arxiv.org/html/2607.22548#bib.bib2)\)provide a powerful, non\-autoregressive framework for synthesizing high\-fidelity data and handling unobserved values as noise for simultaneous forecasting and imputation\. Yet, existing diffusion approaches for time\-series \(e\.g\., CSDI\(Tashiro and others,[2021](https://arxiv.org/html/2607.22548#bib.bib10)\)\) remain structurally rigid and semantically blind\. SeT\-Diff bridges this gap by fusing the generative versatility of diffusion models with the semantic awareness of language models, enabling true zero\-shot adaptation to evolving hardware topologies\.
## 3\.SeT\-Diff Foundational Model
We propose SeT\-Diff, the first multi\-task foundational model for compute node telemetry, acting as a flexible digital twin\. The core design philosophy decouples the modeling of physical system dynamics from the rigid structural constraints of sensor instrumentation\. Traditional models approximate the conditional distributionp\(xt\+1\|x0:t\)p\(x\_\{t\+1\}\|x\_\{0:t\}\)assuming a fixed indexiifor each feature\. In contrast, SeT\-Diff treats the multivariate time series not as a fixed matrix, but as a collection of interacting physical signals defined by semantic identity and statistical behavior\.
We leverage a DDPM framework for its ability to model complex distributions and its inherent robustness to noise\. The training objective reverses a forward diffusion process adding Gaussian noise to the input dataX∈ℝN×PX\\in\\mathbb\{R\}^\{N\\times P\}\. Unlike standard approaches, our denoising networkϵθ\\epsilon\_\{\\theta\}is explicitly conditioned on a rich context𝒞\\mathcal\{C\}describing what each sensor measures\. This allows learning permutation\-invariant representations: if ”CPU Temperature” moves channels, the model recognizes its context embedding rather than its position, ensuring consistent generation across reconfigured hardware layouts\.
### 3\.1\.Context Encoding
To transform raw metadata into a machine\-interpretable signal, a dual\-stream Context Encoder generates a context vector𝒞\\mathcal\{C\}for each of thePPfeatures, composed of two embedding types:
#### Semantic Embeddings \(EtextE\_\{text\}\)
The primary anchor for permutation stability is the textual description of each sensor \(e\.g\., ”Temperature in CPU Core 0, in celsius degrees”\)\. These define only the feature’s identity, explicitly excluding behavioral characteristics or inter\-feature relationships, which the model learns autonomously\. We utilize Sentence\-BERT to encode these into fixed\-size dense vectorsd∈ℝP×384d\\in\\mathbb\{R\}^\{P\\times 384\}\. These embeddings provide semantic meaning, freeing the model from positional reliance and enabling zero\-shot adaptation when sensors are reordered, removed, or newly introduced\.
#### Statistical Descriptors \(EstatE\_\{stat\}\)
Semantic embeddings lack the scale or stationary properties crucial for normalizing the generative process\. We compute a summary statistics vectors∈ℝP×9s\\in\\mathbb\{R\}^\{P\\times 9\}for each channel \(mean, quantiles, skewness, kurtosis\)\. These descriptors provide a structural prior about the expected distribution, stabilizing generation for signals with vastly different dynamic ranges \(e\.g\., RPM vs\. Volts\)\.
The final context𝒞\\mathcal\{C\}is obtained by projecting and concatenating these components, creating a conditioning tensor that accompanies the noisy input throughout denoising\.
### 3\.2\.Architecture
The backbone of SeT\-Diff is a specialized Transformer architecture handling both temporal dependencies \(dynamics\) and inter\-sensor correlations \(system topology\)\.
#### Input Projection
The noisy inputXtX\_\{t\}and context𝒞\\mathcal\{C\}are projected into a common hidden dimensionHH\. To seamlessly integrate the context, we prepend the embeddings to the raw time\-series\. Specifically, statistical and textual embeddings act as two additional ”timesteps” per sensor, resulting in an augmented tensorZt∈ℝ\(N\+2\)×P×HZ\_\{t\}\\in\\mathbb\{R\}^\{\(N\+2\)\\times P\\times H\}\. The attention mechanism thus attends to the sensor’s description and statistics exactly as it attends to its past values\.
#### Factorized Attention Mechanism
To efficiently model dependencies in high\-dimensional telemetry \(P≫100P\\gg 100\), we employ a factorized attention scheme alternating between two blocks:
- •Temporal Attention:Computes self\-attention across theNNtimesteps independently for each feature, capturing individual temporal evolution\.
- •Feature\-wise Attention:Computes self\-attention across thePPfeatures independently per timestep\. This captures instantaneous correlations \(e\.g\., power impact on temperature\), learning the system’s ”interaction graph\.”
This modular design is critical for handling missing data\. By masking within attention heads, SeT\-Diff ignores missing features or timesteps, ensuring generation is based solely on observed evidence\. Crucially, appropriately modifying these masking bits during generation programs different tasks in SeT\-Diff\.
### 3\.3\.Training Objective
Following the DDPM framework, our Transformer backbonefθ\(Zt,t\)f\_\{\\theta\}\(Z\_\{t\},t\)estimates the noise added to the original dataX0X\_\{0\}at timett\. The training lossℒθ\\mathcal\{L\}\_\{\\theta\}is the distance between the estimated noiseϵ^\\hat\{\\epsilon\}and actual noiseϵ\\epsilon:
ℒθ=𝔼\[‖ϵt−fθ\(Zt,t\)‖22\]\\mathcal\{L\}\_\{\\theta\}=\\mathbb\{E\}\\left\[\|\|\\epsilon\_\{t\}\-f\_\{\\theta\}\(Z\_\{t\},t\)\|\|\_\{2\}^\{2\}\\right\]
### 3\.4\.Generation and Multi\-Task Inference
The inference process is iterative and stochastic\. Starting from pure Gaussian noiseXT∼𝒩\(0,I\)X\_\{T\}\\sim\\mathcal\{N\}\(0,I\), the model progressively estimates the noise to remove to reachX0X\_\{0\}:
Xt−1=1αt\(Xt−1−αt1−α¯tϵ^θ\(Zt,t\)\)\+σtzX\_\{t\-1\}=\\frac\{1\}\{\\sqrt\{\\alpha\_\{t\}\}\}\\left\(X\_\{t\}\-\\frac\{1\-\\alpha\_\{t\}\}\{\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\}\\hat\{\\epsilon\}\_\{\\theta\}\(Z\_\{t\},t\)\\right\)\+\\sigma\_\{t\}z
## 4\.Multi\-task Capabilities and Inference Strategies
Unlike specialized architectures, SeT\-Diff serves as a unified framework requiring no retraining\. By leveraging the generative reverse diffusion process, we address distinct telemetry tasks solely by modifying a binary conditioning maskMMat inference time\. Formally, the inference process iteratively samples missing partsXmissX\_\{miss\}conditioned on observed dataXobsX\_\{obs\}and the semantic context𝒞\\mathcal\{C\}\.
### 4\.1\.Imputation and Data Recovery
Telemetry streams often suffer from missing data points due to transient sensor or network failures\. We frame imputation as a random in\-painting task\. The maskMimpM\_\{imp\}reflects the arbitrary pattern of missing values\. During reverse diffusion, the model harmonizes unobserved entries with observed ones via learned inter\-feature correlations \(e\.g\., reconstructing ”CPU Temperature” using ”Fan Speed” and ”Power Load”\), ensuring coherent log reconstruction under severe data loss\.
### 4\.2\.Virtual Sensing \(Regression\)
When comprehensive instrumentation is prohibitively expensive, operators infer ”hidden” physical quantities from available proxies\. We approach this as a feature\-wise regression\. To infer an unmonitored metricjj, maskMregM\_\{reg\}zeros out thejj\-th column for all timesteps \(M:,j=0M\_\{:,j\}=0\), while proxy columns remain visible\. Lacking numerical history for featurejj, the model relies entirely on its Semantic Context \(e\.g\., ”DRAM Power”\) to understand what physical quantity to generate, inferring values based on the dynamics of visible proxies\.
### 4\.3\.Probabilistic Forecasting
Deterministic forecasters often fail to capture the stochastic nature of workloads for proactive system management\. We treat forecasting as a temporal extension task\. Given a history context lengthLLand forecast horizonHH, maskMforeM\_\{fore\}observes timestampst∈\[1,L\]t\\in\[1,L\]and maskst∈\[L\+1,L\+H\]t\\in\[L\+1,L\+H\]\. By running stochastic denoising multiple times, SeT\-Diff generates a distribution of possible futures\. This ensemble of trajectories inherently quantifies aleatoric uncertainty, providing confidence intervals for critical metrics like peak temperature\.
Figure 1\.Unified Task Inference via Dynamic Masking Strategy\. The diagram illustrates how a single pre\-trained SeT\-Diff addresses distinct challenges by modifying the input mask during inference: \(a\) Imputation of random missing values; \(b\) Probabilistic Forecasting of future timesteps; \(c\) Virtual Sensing \(Regression\) of entire unobserved features
## 5\.Experimental Evaluation
We validate SeT\-Diff on real\-world HPC telemetry, focusing on multi\-task capability, zero\-shot permutation stability, and the efficacy of semantic conditioning\.
### 5\.1\.Experimental Setup
#### Dataset & Preprocessing
We use the M100 ExaData dataset\(Borghesi and others,[2023](https://arxiv.org/html/2607.22548#bib.bib1)\)\(20 months of Marconi100 operations\), aggregating out\-of\-band \(IPMI\) and in\-band \(Ganglia\) metrics\. Data is structured into rolling windows \(L=32L=32\), yielding∼\\sim40K windows withP=261P=261metrics\. We applied chronological splitting to prevent look\-ahead bias and standardized features using training\-set statistics Missing values in training use mean imputation; test\-set gaps remain intact to simulate real sparsit
#### Context & Implementation
Sensor descriptions were encoded via Sentence\-BERT into fixed embeddingsEtext∈ℝ261×384E\_\{text\}\\in\\mathbb\{R\}^\{261\\times 384\}\. For production feasibility, we employed a compact Transformer \(1 block,H=64H=64, 8 heads, 136\.3K parameters\) withT=1000T=1000diffusion steps\. Training on a single NVIDIA A100 GPU \(Adam, LR=10−310^\{\-3\}, batch 128\) took∼\\sim90 minutes for 50 epochs\. Generating a32×26132\\times 261window takes∼\\sim140 ms, proving lightweight efficiency\.
### 5\.2\.Results
Table[1](https://arxiv.org/html/2607.22548#S5.T1)reports the performance across operational scenarios under stochastic test\-time perturbations\. We compare SeT\-Diff against a structurally identical diffusion baseline that relies on standard positional indices instead of textual semantic embeddings\. While Mean Absolute Error \(MAE\) serves as our primary comparative metric, we additionally conduct an extended probabilistic evaluation on our proposed model using the Continuous Ranked Probability Score \(CRPS\) to quantify its generative uncertainty\.
Table 1\.Performance comparison\. SeT\-Diff outperforms the structurally identical positional baseline in point\-wise accuracy \(MAE\)\. Additionally, we report the CRPS for our proposed model to highlight its well\-calibrated probabilistic uncertainty quantification\.#### Impact of Semantic Conditioning
The integration of semantic context yields massive performance gains over the positional baseline\. In standard imputation, semantic embeddings significantly reduces the error from a 0\.154 MAE to a 0\.047 MAE\. This superiority is pronounced also under virtual sensing and forecasting tasks, showing how SeT\-Diff effectively leverages the remaining contextual descriptions to guide the generation, bounding the degradation to a MAE of∼0\.062\\sim 0\.062\.
#### Operational Performance and Zero\-Shot Stability
Beyond outperforming the baseline, SeT\-Diff demonstrates remarkable versatility as a unified modeling tool\. Under severe masking \(up to 50% feature loss\), it acts as a robust soft\-sensor \(0\.0622 MAE, 0\.0589 CRPS\) by capturing physical inter\-dependencies\. Similarly, it provides reliable temporal extrapolations for forecasting \(0\.0613 MAE, 0\.0583 CRPS\)\. Under random permutation of the input sensor array , SeT\-Diff yields performance virtually identical to its unperturbed state \(0\.0472 MAE vs\. 0\.0470 MAE\)\. This confirms that the model identifies signals strictly by their semantic meaning, exhibiting intrinsic zero\-shot robustness to hardware reconfigurations where traditional \(index\-based\) models naturally fail\. Furthermore, the CRPS values obtained for our model \(e\.g\., 0\.0299 for imputation\) are consistently lower than the corresponding MAE\. This confirms that the model’s generated predictive distributions are also sharp and well\-calibrated, providing operators with highly reliable confidence intervals
### 5\.3\.Operational Use\-Cases \(Virtual Sensing\)
We defined three Virtual Sensing scenarios \(Table[2](https://arxiv.org/html/2607.22548#S5.T2)\), masking the entire timeseries if selected target features to evaluate SeT\-Diff’s reconstructive capabilities using solely the context and the remaining sensors sequences\.
Table 2\.Virtual Sensing Case Studies: Input contexts vs\. Targets\. In brackets the \# of features\.- •A\. Power Estimation:SeT\-Diff can effectively map computational activity \(OS load\) to power consumption without requiring physical instrumentation, achieving a 0\.088 MAE\.
- •B\. Thermal Estimation:The model can identify the lagged correlation between power injection and temperature rise, to reconstruct internal thermal maps purely from power/ambient context, with a 0\.033 MAE
- •C\. GPU Estimation:Leveraging microarchitectural proxies \(clocks, utilization\), the model is able to impute GPU temperatures \(0\.0250 MAE\), though power recovery proves slightly more complex \(0\.1545 MAE\), an aspect reserved for future investigation\.
## 6\.Conclusion and Future Work
We presented SeT\-Diff, a context\-aware diffusion framework advancing the realization of Foundation Models for HPC telemetry\. By integrating semantic knowledge directly into the generative process, we address a pervasive challenge in data center operations: the rigidity of data\-driven models in the face of dynamic hardware configurations\.
Experiments on the M100 Exascale dataset demonstrate that SeT\-Diff is a robust system\-level tool\. Conditioning on textual embeddings grants it unique ”permutation stability,” maintaining consistent performance even when sensor layouts are completely reshuffled—a scenario that traditional models can’t handle\. Furthermore, a single pre\-trained model seamlessly handles data imputation, virtual sensing, and forecasting simply by altering the inference\-time masking strategy\. This ”train once, deploy everywhere” capability significantly reduces operational overhead by eliminating the need to train specialized models for every compute node variant\.
Future work will scale this paradigm along two axes\. First, we will expand the semantic vocabulary to include system topologies and executing processes\. Second, we will extend the training corpus across multiple architectures and heterogeneous sources\.
###### Acknowledgements\.
The activities of this work have been supported by EU \- NextGenerationEU with funds made available by the National Recovery and Resilience Plan \(PNRR\) Mission 4, Component 2, Investment 3\.3 \(D\.M\. 117/2023\), EuroHPC EUPEX \(g\.a\. 101033975\), EuroHPC JU SEANERGYS \(g\.a\. 101177590\), and CINECA
## References
- A\. F\. Ansariet al\.\(2024\)Chronos: learning the language of time series\.Transactions on Machine Learning Research\.Cited by:[§2](https://arxiv.org/html/2607.22548#S2.p3.1)\.
- F\. Anticiet al\.\(2025\)F\-data: a fugaku workload dataset for job\-centric predictive modelling in hpc systems\.Scientific Data12\(1\),pp\. 1321\.Cited by:[§1](https://arxiv.org/html/2607.22548#S1.p1.1)\.
- A\. Borghesiet al\.\(2023\)M100 exadata: a data collection campaign on the cineca’s marconi100 tier\-0 supercomputer\.Scientific Data10\(1\),pp\. 288\.Cited by:[§1](https://arxiv.org/html/2607.22548#S1.p1.1),[§1](https://arxiv.org/html/2607.22548#S1.p3.1),[§2](https://arxiv.org/html/2607.22548#S2.p1.1),[§5\.1](https://arxiv.org/html/2607.22548#S5.SS1.SSS0.Px1.p1.3)\.
- B\. Guindaniet al\.\(2024\)Exploring the utility of graph methods in hpc thermal modeling\.InICPE,pp\. 106–111\.Cited by:[§2](https://arxiv.org/html/2607.22548#S2.p2.1)\.
- J\. Hoet al\.\(2020\)Denoising diffusion probabilistic models\.Advances in neural information processing systems33,pp\. 6840–6851\.Cited by:[§2](https://arxiv.org/html/2607.22548#S2.p3.1)\.
- S\. Jadonet al\.\(2021\)Challenges and approaches to time\-series forecasting in data center telemetry: a survey\.arXiv preprint arXiv:2101\.04224\.Cited by:[§2](https://arxiv.org/html/2607.22548#S2.p1.1)\.
- M\. Jinet al\.\(2024\)Time\-llm: time series forecasting by reprogramming large language models\.InICLR,Cited by:[§2](https://arxiv.org/html/2607.22548#S2.p3.1)\.
- A\. L\. d\. C\. D\. Limaet al\.\(2021\)Smart predictive maintenance for high\-performance computing systems: a literature review\.The Journal of Supercomputing77\(11\),pp\. 13494–13513\.Cited by:[§1](https://arxiv.org/html/2607.22548#S1.p1.1)\.
- M\. Maiterthet al\.\(2024\)Visualizing an exascale data center digital twin: considerations, challenges and opportunities\.InVIS,Vol\.,pp\. 21–25\.External Links:[Document](https://dx.doi.org/10.1109/VIS55277.2024.00012)Cited by:[§1](https://arxiv.org/html/2607.22548#S1.p1.1)\.
- M\. Maiterthet al\.\(2025\)HPC digital twins for evaluating scheduling policies, incentive structures and their impact on power and cooling\.InSC,pp\. 1959–1969\.Cited by:[§1](https://arxiv.org/html/2607.22548#S1.p1.1)\.
- O\. Medaiyeseet al\.\(2025\)Hardware telemetry at scale: a case study on ssds endurance monitoring in datacenters\.InDSN\-S,pp\. 147–152\.Cited by:[§1](https://arxiv.org/html/2607.22548#S1.p1.1)\.
- M\. Molanet al\.\(2023\)RUAD: unsupervised anomaly detection in hpc systems\.Future Generation Computer Systems141,pp\. 542–554\.Cited by:[§1](https://arxiv.org/html/2607.22548#S1.p1.1),[§2](https://arxiv.org/html/2607.22548#S2.p2.1)\.
- N\. Reimers and I\. Gurevych \(2019\)Sentence\-bert: sentence embeddings using siamese bert\-networks\.arXiv preprint arXiv:1908\.10084\.Cited by:[§1](https://arxiv.org/html/2607.22548#S1.p2.1)\.
- Y\. Tashiroet al\.\(2021\)Csdi: conditional score\-based diffusion models for probabilistic time series imputation\.Advances in neural information processing systems34,pp\. 24804–24816\.Cited by:[§2](https://arxiv.org/html/2607.22548#S2.p3.1)\.Similar Articles
SevDiff: Severity-Conditioned Diffusion for Long-Tail Conflict Trajectory Generation
SevDiff is a severity-conditioned diffusion model for generating vehicle conflict trajectories with controlled time-to-collision values, achieving high hit-rate on a real-world dataset for ADAS evaluation.
Beyond Distribution Matching: Semantics-Consistent Tabular Diffusion with Weak Semantic Priors
SCTab-Diff is a semantics-consistent tabular diffusion framework that uses weak semantic priors to generate high-fidelity synthetic tabular data, improving distributional fidelity and semantic consistency over existing methods.
Quantum Generative Diffusion Model for Real-World Time Series
QDiffusion-TS is the first quantum generative diffusion model for real-world time series synthesis, replacing feed-forward components in a denoising transformer with quantum neural networks. It reduces trainable parameters by nearly three orders of magnitude and improves Wasserstein distance by 44% on financial data, with downstream forecasting gains up to 71% in RMSE.
MMDiff: Extending Diffusion Transformers for Multi-Modal Generation
MMDiff extends frozen diffusion transformers into multi-modal generative systems using lightweight decoders, achieving significant improvements in semantic segmentation and other perceptual tasks through multi-timestep feature fusion.
ReDiTT: Retrieval Augmented Conditional Diffusion Transformers for Asynchronous Time Series
This paper presents ReDiTT, a retrieval augmented conditional diffusion transformer for asynchronous time series prediction. The model retrieves structurally similar latent sequences as reference conditions to improve long-horizon forecasting and sample diversity, achieving state-of-the-art performance on seven real-world datasets.