Tianmu-TC: Physics-constraints Generative Artificial Intelligence for Global Tropical Cyclone Forecasting

arXiv cs.LG Papers

Summary

This paper introduces Tianmu-TC, a physics-constrained generative AI framework for global tropical cyclone forecasting that outperforms traditional systems in reliability and computational efficiency.

arXiv:2608.18500v1 Announce Type: new Abstract: Tropical cyclones (TCs) pose severe risks from strong winds and heavy rainfall. However, forecasting their track and intensity remains challenging due to chaotic atmosphere and the rapid amplification of initial condition errors, leading to growing forecast uncertainty. While numerical weather prediction (NWP) and deep learning models have made progress, they remain computationally demanding and often fail under complex meteorological scenarios. Here, we present Tianmu-TC, a physics-constraints generative framework for global TC forecasting. Trained on Western North Pacific data, Tianmu-TC leverages physics-constraints to generate controllable outputs with reduced uncertainty thus improving forecast reliability. Experiments show Tianmu-TC outperforms deterministic and ensemble meteorological artificial intelligence models and authoritative NWP systems such as ECMWF in global ocean basins, with significantly lower computational cost. We further show Tianmu-TC performs well in challenging scenarios such as data sparsity, anomaly tracks, rapid intensification and weakening. These findings suggest physics-constraints generative AI offers a promising approach for reliable, efficient global TC forecasting.
Original Article
View Cached Full Text

Cached at: 08/20/26, 10:27 AM

# Tianmu-TC: Physics-constraints Generative Artificial Intelligence for Global Tropical Cyclone Forecasting
Source: [https://arxiv.org/html/2608.18500](https://arxiv.org/html/2608.18500)
Pan MuCheng HuangAffiliation:College of Computer Science and Technology, Zhejiang University of Technology, Hangzhou & 310023, China\.Affiliation:School of Computer Science and Engineering, Tianjin University of Technology, Tianjin & 300384, China\.Hanting YanAffiliation:College of Computer Science and Technology, Zhejiang University of Technology, Hangzhou & 310023, China\.Affiliation:School of Earth Sciences, Zhejiang University, Hangzhou & 310058, China\.Yuchao ZhuAffiliation:College of Computer Science and Technology, Zhejiang University of Technology, Hangzhou & 310023, China\.Jinglin ZhangAffiliation:School of Control Science and Engineering, Shandong University, Jinan & 250061, China\.Shengyong ChenShoujuan ShuCong Bai

###### Abstract

Tropical cyclones \(TCs\) pose severe risks from strong winds and heavy rainfall\. However, forecasting their track and intensity remains challenging due to chaotic atmosphere and the rapid amplification of initial condition errors, leading to growing forecast uncertainty\. While numerical weather prediction \(NWP\) and deep learning models have made progress, they remain computationally demanding and often fail under complex meteorological scenarios\. Here, we present Tianmu\-TC, a physics\-constraints generative framework for global TC forecasting\. Trained on Western North Pacific data, Tianmu\-TC leverages physics\-constraints to generate controllable outputs with reduced uncertainty thus improving forecast reliability\. Experiments show Tianmu\-TC outperforms deterministic and ensemble meteorological artificial intelligence models and authoritative NWP systems such as ECMWF in global ocean basins, with significantly lower computational cost\. We further show Tianmu\-TC performs well in challenging scenarios such as data sparsity, anomaly tracks, rapid intensification and weakening\. These findings suggest physics\-constraints generative AI offers a promising approach for reliable, efficient global TC forecasting\.

## Introduction

Tropical cyclones \(TCs\) are complex weather systems that often bring strong winds and heavy rainfall, leading to natural disasters that cause significant damage to infrastructure and pose serious threats to human life and property\[[1](https://arxiv.org/html/2608.18500#bib.bib1)\]\. Therefore, accurate forecasting is crucial for effective prevention of disasters caused by TCs\. The track of a TC tells us where it is currently located, while its intensity represents its strength and the potential severity of the damage it may cause\. Forecasting these two attributes are crucial for determining the level of preparedness required\. Specifically, the TC track is a combination of locations, including longitude and latitude of the TC center\. TC intensity can represented by different attributes, with pressure and wind speed being the most representative\. Pressure refers to the lowest atmospheric pressure \(hPa\) at the TC center, and wind speed to the maximum sustained wind speed \(m/s\) over two minutes near the TC center\. These attributes are closely interconnected in physics\[[2](https://arxiv.org/html/2608.18500#bib.bib2),[3](https://arxiv.org/html/2608.18500#bib.bib3)\]\.

However, forecasting tropical cyclones inherently involves uncertainty, making accurate predictions a persistent challenge\. The atmosphere is a chaotic system, meaning that even small errors in the initial conditions can amplify over time—a phenomenon widely known as the butterfly effect—ultimately degrading forecast accuracy and increasing uncertainty\. This uncertainty becomes especially pronounced as the forecast horizon extends, with both track and intensity predictions becoming less reliable\. For instance, the predicted location of the TC center may deviate by several hundred kilometers, while intensity forecasts often fail to capture abrupt changes such asrapid intensificationorrapid weakening\. In practice, this leads to wide forecast spreads and rapidly growing uncertainty\.

To address these challenges, authoritative Numerical Weather Prediction \(NWP\) models\[[4](https://arxiv.org/html/2608.18500#bib.bib4),[5](https://arxiv.org/html/2608.18500#bib.bib5),[6](https://arxiv.org/html/2608.18500#bib.bib6)\]adopt ensemble forecasting strategies, which simulate multiple scenarios by perturbing the initial conditions\[[7](https://arxiv.org/html/2608.18500#bib.bib7)\]\. This approach has been widely adopted by official agencies such as the European Centre for Medium\-Range Weather Forecasts \(ECMWF\)\[[8](https://arxiv.org/html/2608.18500#bib.bib8)\], the National Hurricane Center \(OFCL\)\[[6](https://arxiv.org/html/2608.18500#bib.bib6)\], and the China Central Meteorological Observatory \(CMO\)\[[9](https://arxiv.org/html/2608.18500#bib.bib9)\]\. While ensemble methods help mitigate uncertainty, they require substantial computational resources to simulate global atmospheric processes\[[10](https://arxiv.org/html/2608.18500#bib.bib10)\], resulting in prohibitively long training and inference times\. This makes them less suitable for real\-time regional applications where rapid response is essential\. As the number of uncertainties being considered grows, the required computational resources increase significantly\. Thus, it is imperative to explore more efficient approaches for TC forecasting\.

Recently, the advancement of generative models has provided a novel perspective for addressing the challenges of uncertainty\[[11](https://arxiv.org/html/2608.18500#bib.bib11),[12](https://arxiv.org/html/2608.18500#bib.bib12),[13](https://arxiv.org/html/2608.18500#bib.bib13),[14](https://arxiv.org/html/2608.18500#bib.bib14),[15](https://arxiv.org/html/2608.18500#bib.bib15)\]with substantially lower computational cost\. By introducing controlled noise or perturbations, they technically circumvent the impact of initial errors and their iterative amplification, enabling the generation of future TC tracks and intensities\. Furthermore, their ability to generate a large number of diverse members, similar to ensemble forecasting, makes them highly suitable for predicting future tropical cyclone tracks and intensities\.

However, the diversity reflected in these ensemble members—while beneficial for representing uncertainty—can lead to overly dispersed outputs, including multiple plausible but divergent tracks, thereby compromising the precision required for disaster prevention and emergency response\. To address this, we introducephysics\-constraintsinto the generative process\. These constraints regulate the stochasticity of the model, enabling a more balanced trade\-off between diversity and accuracy\. The constraint conditions are formally defined and illustrated in Figure[1](https://arxiv.org/html/2608.18500#Sx1.F1)\.

![Refer to caption](https://arxiv.org/html/2608.18500v1/framework/fig1-0804.jpeg)Figure 1:Physics\-constraints for TC forecasting\.\(A\) Overall schematic: meteorological factors influencing the development of TCs\. \(B\) Proposed physics\-constraints govern the generation process and achieve a lower uncertainty score with the smallest prediction area\. \(C\) The data composition of atmospheric environment constraint\. The large\-scale wind field consists of the zonaluuand meridional windvvcomponents from 200 hPa to 850 hPa\. High\-pressure systems are composed of the South Asian High \(SAH,z​200z200\) and the Western Pacific Subtropical High \(WPSH,z​500z500\), wherezzdenotes geopotential height\. \(D\) TC tilt metrics: calculated from vorticity data at three pressure levels\. See details in Materials and Methods\.![Refer to caption](https://arxiv.org/html/2608.18500v1/framework/fig2-zeng.png)

Figure 2:Framework of Tianmu\-TC\.Tianmu\-TC consists of three key components:\(1\) Constraint Encoderfor preliminary constraint encoding,\(2\) Cross\-Modal Fusionmodule to establish correlations between constraints, and\(3\) Constraint guiding denoising moduleto perform the final denoising process\. TheyK∼y0y\_\{K\}\\sim y\_\{0\}represents the TC prediction results under different diffusion stepskk\. To facilitate clear visualization, only the predicted track is displayed here, whereas wind speed and pressure predictions are simultaneously generated\. FromyKy\_\{K\}toy0y\_\{0\}, the blue dashed line gradually converges from noise to near ground truth\. The single\-step denoising process fromyky\_\{k\}toyk−1y\_\{k\-1\}is implemented by thePhysics\-constraints Denoising Block, which operates under the guidance of physics\- constraints\.Firstly, one common constraint in prediction tasks is the historical values of the target attribute\. In TC forecasting, this typically includes the TC’s past track and intensity, which we refer to as thehistorical information constraint\.

Secondly, the generation and evolution of a TC are influenced by large\-scale atmospheric environment\. To account for this, we introducethe atmospheric environment constraint, which incorporate the major large\-scale circulation systems influencing TC activity include the South Asia High \(SAH\) in the upper troposphere, the WP Subtropical High \(WPSH\), and the monsoon trough\[[16](https://arxiv.org/html/2608.18500#bib.bib16),[17](https://arxiv.org/html/2608.18500#bib.bib17),[18](https://arxiv.org/html/2608.18500#bib.bib18)\]\.

Finally, beyond environmental forcing, TC attributes are also impacted by internal structural changes, which are often overlooked in prior work\. Therefore, we designed theinternal structure variation constraint, explicitly incorporating structural changes into the prediction process to control the diversity of generated outputs\.

The three constraints above are collectively referred to as thePhysics\-driven Meteorological Factor Constraints, namely, Physics\-constraints\. Based on these, we introduceTianmu\-TC, a physics\-constrained generative diffusion framework for global tropical cyclone forecasting, as shown in Figure[2](https://arxiv.org/html/2608.18500#Sx1.F2)\. By integrating Physics\-constraints, Tianmu\-TC transforms track and intensity prediction into a conditional generative process, where the initially chaotic generation results progressively become clearer and more accurate\. Our approach offers a new paradigm for efficient and robust global TC forecasting, highlighting the practicality and potential of generative modeling as a forecasting approach\.

Previous efforts to apply artificial intelligence to TC forecasting have largely focused on deterministic prediction by learning patterns from historical data, while often lacking explicit modeling of forecast uncertainty, as exemplified by classical spatiotemporal sequence models such as RNNs\[[19](https://arxiv.org/html/2608.18500#bib.bib19),[20](https://arxiv.org/html/2608.18500#bib.bib20)\], GRUs\[[21](https://arxiv.org/html/2608.18500#bib.bib21)\], Bi\-GRU\[[22](https://arxiv.org/html/2608.18500#bib.bib22)\], LSTMs\[[23](https://arxiv.org/html/2608.18500#bib.bib23)\], and ConvLSTMs\[[24](https://arxiv.org/html/2608.18500#bib.bib24)\]\. More recent work incorporates multi\-modal inputs—including historical TC attributes and gridded reanalysis variables—to better represent the surrounding atmospheric environment\[[25](https://arxiv.org/html/2608.18500#bib.bib25),[26](https://arxiv.org/html/2608.18500#bib.bib26),[27](https://arxiv.org/html/2608.18500#bib.bib27),[28](https://arxiv.org/html/2608.18500#bib.bib28),[29](https://arxiv.org/html/2608.18500#bib.bib29),[30](https://arxiv.org/html/2608.18500#bib.bib30),[31](https://arxiv.org/html/2608.18500#bib.bib31),[32](https://arxiv.org/html/2608.18500#bib.bib32)\]\. In parallel, large\-scale meteorological foundation models\[[33](https://arxiv.org/html/2608.18500#bib.bib33)\]such as Pangu\[[34](https://arxiv.org/html/2608.18500#bib.bib34)\]and Fengwu\[[35](https://arxiv.org/html/2608.18500#bib.bib35)\], together with recent probabilistic or ensemble AI weather models such as GenCast\[[11](https://arxiv.org/html/2608.18500#bib.bib11)\]and NeuralGCM\[[36](https://arxiv.org/html/2608.18500#bib.bib36)\], have achieved impressive skill by directly forecasting global ERA5 reanalysis fields\. Although these models can indirectly support TC forecasting through downstream tracking, they are not specifically optimized for TC\-focused tasks\. This often leads to underestimation of intensity, less compact ensemble distributions, and suboptimal performance in rare or extreme scenarios\. Additionally, their dependence on globally scaled datasets and large model sizes incurs substantial training and inference costs, making them less practical for real\-time or resource\-constrained forecasting settings\.

In summary, existing approaches struggle to balance forecasting accuracy, computational efficiency, and uncertainty reduction, and often underperform under complex or rare meteorological conditions, such as rapid changes in tropical cyclone intensity or track\. Our goal is to enhance prediction reliability by effectively reducing uncertainty, without incurring additional computational cost\.

![Refer to caption](https://arxiv.org/html/2608.18500v1/framework/c1c2c3-alltask-picture-xiufu.jpg)

Figure 3:Results under different constraints\.\(A\) Generation results of Typhoon Krosa at 00:00 UTC on August 16, 2019, under five constraint settings with decreasing diffusion stepskk\(from 100 to 0\)\. Shaded regions represent uncertainty\. Applying all constraints \(C1C\_\{1\},C2C\_\{2\}, andC3C\_\{3\}; rightmost column\) results in the most concentrated predictions with the lowest uncertainty\. \(B\) For the Typhoon Krosa at 00:00 UTC on August 16, 2019, corresponding to \(A\), the mean and variance of the predictions are visualized as spread distributions\. Smaller variance leads to a steeper curve, indicating lower prediction uncertainty\. \(C\) Average prediction errors across 146 TCs from 2017 to 2023\. The rows correspond to \(a\) track, \(b\) pressure, and \(c\) wind speed\.
## Results

Tianmu\-TC takes 48 hours of TC historical data as input and generates forecasts for the next 24 or 72 hours\. The input includes three core components: ① historical TC attributes from the International Best Track Archive for Climate Stewardship \(IBTrACS\[[37](https://arxiv.org/html/2608.18500#bib.bib37)\]\), ② ERA5\-derived environmental fields representing large\-scale atmospheric conditions, and ③ TC tilt metrics extracted from vorticity fields to capture internal structural dynamics\. In this study, the IBTrACS best track records are adopted as the ground truth in all evaluations, following standard practice in deep learning\-based forecasting\.

To rigorously assess generalizability, we evaluate Tianmu\-TC on a global test set spanning all ocean basins from 2017 to 2023\. This design ensures coverage of a wide range of TC scenarios, including varying intensities, tracks, and environmental conditions\. Notably, to avoid excessive training time and computational cost, the model is trained exclusively on Western North Pacific \(WP\) TCs from 1950 to 2016\. This highlights the model’s ability to generalize beyond its training domain, demonstrating robust performance across unseen basins despite being trained regionally\. The time resolution is 6 hours\. For fair comparison, following TCNM\[[31](https://arxiv.org/html/2608.18500#bib.bib31)\], the sampling number was set to 6, meaning that the model generates 6 possible tendencies\. Further details on the model architecture, training setup, and datasets are provided in the Materials and Methods section\.

Tianmu\-TC reduces forecast uncertainty through the incorporation of physics\-driven constraints, which consists of the historical information constraint \(C1C\_\{1\}\), the atmospheric environment constraint \(C2C\_\{2\}\), and the TC internal structure variation constraint \(C3C\_\{3\}\)\. We compare its performance against authoritative numerical weather prediction systems, deep learning–based methods, and deterministic and ensemble large\-scale meteorological foundation models across global test regions\. Additionally, we examine the model’s effectiveness under challenging scenarios such as data sparsity, anomaly tracks, landfall events,rapid intensificationandrapid weakening\.

The results demonstrate that Tianmu\-TC achieves a favorable balance between predictive accuracy and controlled output diversity, while offering strong robustness and fast inference\. Together with its compact and accurate multi\-trend forecasts, these findings suggest that generative AI, when properly constrained by physical principles, constitutes a reliable and effective approach for advancing global TC forecasting\.

### Physics\-constraints reduce prediction uncertainty

A central challenge in TC forecasting is how to effectively reduce predictive uncertainty\. Uncertainty arises naturally from the chaotic nature of the atmosphere and can be further amplified by model imperfections, making it essential to evaluate whether generated ensembles can provide not only accurate predictions but also reliable distributions\. To address this, we investigate the role of incorporating physics\-constraints into the generative process\. We evaluate their effectiveness by comparing generation results under five settings: unconstrained generation \(none of the constraints\), generation with partial constraints \(C1C\_\{1\},C1&C2C\_\{1\}\\&C\_\{2\},C1&C3C\_\{1\}\\&C\_\{3\}\), and generation with full physics\-constraints \(C1&C2&C3C\_\{1\}\\&C\_\{2\}\\&C\_\{3\}\)\. During reverse diffusion, the forecasts evolve from broad, highly uncertain distributions at diffusion stepk=100k=100to concentrated predictions near the ground truth atk=0k=0\.

These evaluations leads to three key observations:

1\. Full physics\-constraints yield the lowest uncertainty with the most concentrated prediction distributions\. When all three constraints are applied, Tianmu\-TC generates ensembles that converge tightly around the ground truth for both track and intensity\. As shown in Figure[3](https://arxiv.org/html/2608.18500#Sx1.F3)Aa–Ac, the final outcomes \(k=0k=0\) under full constraints form compact clusters, whereas the unconstrained case produces substantially larger deviations—track errors are about six times and intensity errors about three times larger\. This aspect is further illustrated by the probability density curves in Figure[3](https://arxiv.org/html/2608.18500#Sx1.F3)Ba–Bc, where the fully constrained results exhibit the steepest probability density curves, with data points highly concentrated, indicating concentrated forecasts and substantially reduced uncertainty\. The Figure[S4](https://arxiv.org/html/2608.18500#A1.F4)presents the uncertainty score corresponding to each predicted region, calculated as the MSE loss between the predictions and the ground truth\. This further supports the aforementioned conclusion\.

2\. Physics\-constraints enable the forecast distribution to better align with the ground\-truth distribution, thereby improving both local accuracy and global consistency\. As shown in Figure[3](https://arxiv.org/html/2608.18500#Sx1.F3)Ba–Bc, the probability density curves of Tianmu\-TC exhibit peak values and shapes closest to those of the ground truth, indicating that predictions are not only accurate at specific points but also capture the broader structural trends of tropical cyclones\. This distributional alignment highlights the model’s ability to reproduce realistic ensemble behavior, extending beyond reduced spread to better capturing the statistical consistency of forecasts with the ground truth\.

3\. Physics\-constraints reduce average prediction errors\. As shown in Figure[3](https://arxiv.org/html/2608.18500#Sx1.F3)Ca–Cc, incorporating all three constraints lowers the average errors across track, pressure, and intensity, confirming the higher accuracy of Tianmu\-TC compared with alternative settings\. Specifically for intensity prediction, theC3C\_\{3\}\(TC internal structure variation constraint\) enhances forecasts by encoding vertical structural changes within the TC, consistent with the well\-established relationship between cyclone structure and intensity evolution\[[38](https://arxiv.org/html/2608.18500#bib.bib38)\]\. This agreement reinforces the physical interpretability of the model\.

The effectiveness of physics\-constraints in reducing prediction uncertainty is consistently reflected in visualization patterns, distribution convergence, and average error metrics, thereby resulting in improved forecasting performance\.

![Refer to caption](https://arxiv.org/html/2608.18500v1/framework/global1-2.jpg)

Figure 4:Comparison of average forecast errors with large\-scale meteorological models and the authoritative NWP system during the test period from 2017 to 2023\.The testing regions cover six ocean basins: \(A\) Eastern North Pacific \(EP\), \(B\) North Atlantic \(NA\), \(C\) North Indian \(NI\), \(D\) South Indian \(SI\), \(E\) South Pacific \(SP\), and \(F\) Western North Pacific \(WP\)\. \(G\) Comparison of computational efficiency metrics with large\-scale meteorological models\.Figure 5:Visualization of track predictions for representative TCs across six ocean basins\.Each basin presents forecasts for a specific TC at three consecutive initialization times: \(A\) EP \(Jimena\), \(B\) NA \(Rene\), \(C\) NI \(Ockhi\), \(D\) SI \(Enawo\), \(E\) SP \(Gita\), and \(F\) WP \(Sonca\)\.
### Generalizability on global ocean areas

Tropical cyclones occur across multiple ocean basins, making cross\-regional generalization a critical requirement for forecasting models\. Although Tianmu\-TC is trained exclusively on Western North Pacific data, we evaluate its performance on global TC datasets spanning six basins to assess its transferability to real\-world applications\. As shown in Figure[4](https://arxiv.org/html/2608.18500#Sx2.F4), Tianmu\-TC outperforms large meteorological models such as Pangu\[[34](https://arxiv.org/html/2608.18500#bib.bib34)\]and Fengwu\[[35](https://arxiv.org/html/2608.18500#bib.bib35)\]in most regions, while achieving comparable overall performance\. In addition, Figure[4](https://arxiv.org/html/2608.18500#Sx2.F4)G shows that Tianmu\-TC achieves an eightfold reduction in training time, a tenfold reduction in inference time, and a threefold reduction in model size\. Two main conclusions emerge:

1. 1\.Efficiency and performance superiority\.Despite its lower training costs, smaller model size, and faster inference compared with large meteorological models, Tianmu\-TC achieves comparable track prediction and significantly better intensity prediction\. Notably, it significantly outperforms ECMWF \(the ECMWF global model\[[8](https://arxiv.org/html/2608.18500#bib.bib8)\]\) in the EP, NI, and SI basins \(Figure[4](https://arxiv.org/html/2608.18500#Sx2.F4)A,C,D\)\. In the NA, SP, and WP basins \(Figure[4](https://arxiv.org/html/2608.18500#Sx2.F4)B,E,F\), it delivers superior intensity prediction performance across all metrics and competitive track accuracy within 18 hours\. Moreover, Tianmu\-TC runs on a single NVIDIA A6000 GPU rather than supercomputing infrastructure, highlighting that deep learning can surpass operational NWP while requiring substantially fewer computational resources\.
2. 2\.Cross\-regional scalability\.Although trained only on WP data, Tianmu\-TC generalizes effectively to the other five basins \(Figure[4](https://arxiv.org/html/2608.18500#Sx2.F4)A–E\), capturing distinct TC dynamics in both hemispheres\. This demonstrates robust adaptability to diverse environments and strong potential for global TC forecasting\.

Table 1:Results on special TC scenarios on global area from 2017 to 2021\.Samples is the number of cases\. Lower values indicate better performance\. “Anomaly” consists of circular movement, turning on the ocean, and V\-shaped tracks\. “RI” denotesRapid Intensification, and “RW” denotesRapid Weakeningof TCs\.CategoryScenario \(Unit\)MethodsSamplesForecast Horizon\+6h\+12h\+18h\+24hTrackAnomaly \(km\)Pangu58244\.1355\.9072\.75104\.26Fengwu58244\.1949\.3365\.8995\.62ECMWF58240\.3250\.2661\.8278\.11Ours58216\.9017\.2834\.7470\.38Landfall \(km\)Pangu30447\.8662\.0987\.55125\.50Fengwu30443\.9652\.8065\.1983\.17ECMWF30435\.7147\.2457\.0876\.26Ours30416\.9820\.9036\.6166\.26IntensityRI \(m/s\)Pangu23122\.8931\.6841\.3049\.32Fengwu23121\.8929\.8538\.8846\.17ECMWF23117\.4118\.3924\.3228\.78Ours2311\.351\.294\.417\.59RW \(m/s\)Pangu32028\.6922\.0516\.0712\.01Fengwu32027\.8721\.2414\.9410\.49ECMWF32013\.7810\.738\.6426\.40Ours3202\.691\.633\.795\.74Figure 6:Visualization of complex prediction scenarios\.\(A\) Track prediction, \(B\) wind speed, and \(C\) pressure\. For \(A\), four challenging track cases are illustrated: \(Aa\) circular track, \(Ab\) post\-landfall prediction, \(Ac\) turning on the ocean, and \(Ad\) V\-shaped movement\. \(Ba\)–\(Bb\):rapid intensification; \(Bc\)–\(Bd\):rapid weakening\.Rapid intensificationis defined as a tropical cyclone’s wind speed increasing by more than 30 knots \(15\.43 m/s\) within 24 hours, whereasrapid weakeningrefers to the opposite process\. \(C\) displays the corresponding pressure prediction results for \(B\)\.For broader comparison, we also benchmarked Tianmu\-TC against authoritative forecasts provided by local meteorological agencies in Figure[4](https://arxiv.org/html/2608.18500#Sx2.F4)and Figure[S2](https://arxiv.org/html/2608.18500#A1.F2)\. The official forecast issued by the National Hurricane Center \(OFCL\)\[[6](https://arxiv.org/html/2608.18500#bib.bib6)\]was assessed in the EP and NA basins, while the NWP system of the China Central Meteorological Observatory \(CMO\)\[[9](https://arxiv.org/html/2608.18500#bib.bib9)\]was evaluated in the WP basin\. As shown in Figure[4](https://arxiv.org/html/2608.18500#Sx2.F4)A, Tianmu\-TC significantly outperforms OFCL in the EP basin\. Notably, Tianmu\-TC surpasses for the first time the authoritative NWP model operated by the CMO across all metrics in the WP basin \(Figure[4](https://arxiv.org/html/2608.18500#Sx2.F4)F\)\.

Most existing studies focus on the Western North Pacific, where tropical cyclones occur most frequently\. Therefore, we further compared Tianmu\-TC with deep learning–based approaches in the WP basin \(Figure[S3](https://arxiv.org/html/2608.18500#A1.F3)\)\. These include single\-task baselines using only historical track or intensity information \(DLM\[[27](https://arxiv.org/html/2608.18500#bib.bib27)\]and TCIF\-fusion\[[39](https://arxiv.org/html/2608.18500#bib.bib39)\]\) and multi\-task models \(SGAN\[[40](https://arxiv.org/html/2608.18500#bib.bib40)\], GBRNN\[[20](https://arxiv.org/html/2608.18500#bib.bib20)\], MMSTN\[[28](https://arxiv.org/html/2608.18500#bib.bib28)\], TCNM\[[31](https://arxiv.org/html/2608.18500#bib.bib31)\], TC\-Diffuser\[[32](https://arxiv.org/html/2608.18500#bib.bib32)\]and TAM\-CL\[[41](https://arxiv.org/html/2608.18500#bib.bib41)\]\)\. Tianmu\-TC consistently surpasses all of these methods across evaluation metrics, highlighting its advantages in both track and intensity prediction\.

In addition to aggregated performance metrics, Figure[5](https://arxiv.org/html/2608.18500#Sx2.F5)presents qualitative visualizations of track predictions across the six global ocean basins\. For each basin, we selected a representative tropical cyclone and visualized its predicted tracks at three consecutive forecast timestamps\. This enables the comparison of temporal continuity and accuracy of predictions among different models\. Tianmu\-TC shows consistent alignment with ground truth tracks over time, capturing not only general movement direction but also subtle tracks turns, demonstrating its reliability and temporal coherence in global\-scale forecasting\.

Overall, these experiments underscore the practical potential of Tianmu\-TC: it not only reduces training costs but also demonstrates exceptional competitiveness against both operational forecasts and state\-of\-the\-art deep learning approaches\.

### Comparison with ensemble meteorological AI model: GenCast & NeuralGCM

To compare with existing multi\-trend forecasting methods, GenCast and NeuralGCM were further selected as representative global machine\-learning weather forecasting models for supplementary evaluation\. These models were chosen because they provide both deterministic and ensemble forecasts, making them suitable references for assessing the effectiveness and competitiveness of Tianmu\-TC in multi\-trend tropical cyclone forecasting\. To ensure reproducibility, the publicly archived gridded forecast data for the year 2020 were used, and tropical cyclone centers and intensity diagnostics were derived directly from the original forecast fields\. Details of the data sources, selected forecast variables, tracking rules, and intensity diagnosis are provided in Supplementary Methods \(Sections[A\.2](https://arxiv.org/html/2608.18500#A1.SSx1.SSS2)and[A\.3](https://arxiv.org/html/2608.18500#A1.SSx1.SSS3)\)\.

This comparison is conducted only for the year 2020 because publicly released GenCast and NeuralGCM forecasts are only provided for this year, and this avoids potential temporal overlap with the training period of GenCast\. Since these forecasts are provided at 12 h intervals, the evaluation is performed at the shared forecast horizons of \+12 h and \+24 h\. In addition, because MSLP is not available in the original NeuralGCM configuration used here, only track prediction results are reported for NeuralGCM\.

Table 2:Comparison with deterministic NeuralGCM and GenCast forecasts\.Results are reported on test samples for the year 2020 at the overlapping forecast horizons of \+12 h and \+24 h\. Track error, pressure error, and wind speed error are measured in km, hPa, and m/s, respectively\. Lower values indicate better performance\. In “Inference time”, s denotes seconds\.ModelSamplesTrack Error \(km\)Pressure Error \(hPa\)Wind Speed Error \(m/s\)TrainingtimeInferencetime\+12h\+24hOverall\+12h\+24hOverall\+12h\+24hOverallNeuralGCM68467\.5283\.2875\.40//////3 weeks11\.9 sGenCast68443\.2562\.0052\.6310\.7110\.7510\.7310\.109\.9310\.025 days32 sTianmu\-TC68417\.5266\.3741\.950\.391\.460\.920\.722\.111\.412 days1\.23 sTable[2](https://arxiv.org/html/2608.18500#Sx2.T2)shows thedeterministic forecast comparison\. Tianmu\-TC clearly outperforms NeuralGCM in track prediction\. This gap may be partly associated with the relatively coarse\(0\.7∘\)\(0\.7^\{\\circ\}\)spatial resolution of the public deterministic NeuralGCM forecasts, which limits the precision of grid\-based TC center localization\. In contrast, Tianmu\-TC is trained with high\-precision best\-track coordinates and is specifically optimized for short\-term TC track prediction\. Compared with GenCast, Tianmu\-TC achieves lower errors for the \+12 h track, the overall track, and all intensity\-related metrics, while its \+24 h track error is only slightly higher\. These results indicate that Tianmu\-TC provides more accurate short\-term tropical cyclone intensity estimates and comparable track prediction skill relative to GenCast\. This comparison is also informative from the perspective of computational cost\. As a global medium\-range ensemble forecasting system, GenCast is built at a substantially larger scale: the original GenCast paper reports slightly more than 3\.5 days of pretraining and under 1\.5 days of 0\.25° fine\-tuning on 32 TPUv5 instances, while a single 15\-day global forecast takes approximately 8 minutes on a Cloud TPUv5 device\. Under a simple linear scaling by forecast length, this corresponds to approximately 32 s for a 24 h forecast\. In contrast, Tianmu\-TC is specifically designed for short\-term multi\-attribute tropical cyclone forecasting, requires only 1\.23 s for prediction, and can be trained on a single A6000 GPU in approximately 2 days\. These results suggest that, with substantially lower computational cost, a tropical\-cyclone\-specific model still has clear advantages for short\-term tropical cyclone track and intensity forecasting\.

Table 3:Comparison with ensemble NeuralGCM and GenCast forecasts\.Results are reported on test samples for the year 2020 at the overlapping forecast horizons of \+12 h and \+24 h\. The overall value is computed over the combined \+12 h and \+24 h samples\. NeuralGCM\-ENS, GenCast\-ENS and Tianmu\-TC\-ENS denote the ensemble versions of NeuralGCM, GenCast and Tianmu\-TC, respectively\. For a fair comparison, all ensemble\-based methods are evaluated using 50 members\.TaskModelSamplesMean ErrorStd\. ErrorSpreadBest\-of\-50\+12h\+24hOverall\+12h\+24hOverall\+12h\+24hOverall\+12h\+24hOverallTrack \(km\)NeuralGCM\-ENS521/50591\.75126\.87109\.0431\.8857\.1844\.3395\.02145\.47119\.8562\.3667\.1664\.72GenCast\-ENS521/50592\.28175\.74133\.3651\.91104\.0577\.5796\.40189\.52142\.2424\.7639\.8632\.19Tianmu\-TC\-ENS521/50528\.40109\.8668\.498\.7634\.1621\.2622\.3385\.2953\.3210\.4240\.9425\.44Pressure \(hPa\)NeuralGCM\-ENS/////////////GenCast\-ENS521/5059\.7210\.059\.881\.301\.921\.601\.342\.051\.696\.886\.276\.58Tianmu\-TC\-ENS521/5051\.664\.903\.260\.591\.721\.140\.671\.941\.290\.581\.721\.14Wind Speed \(m/s\)NeuralGCM\-ENS/////////////GenCast\-ENS521/5059\.349\.289\.311\.141\.521\.331\.181\.601\.396\.726\.156\.44Tianmu\-TC\-ENS521/5051\.013\.162\.070\.361\.060\.700\.421\.210\.810\.341\.220\.77

![Refer to caption](https://arxiv.org/html/2608.18500v1/framework/multitrend.jpeg)

Figure 7:Lifecycle visualization of multi\-trend track forecasts for Typhoon ATSANI \(WP, 2020\)\.The main panel shows the rolling 12\-h forecast path over the TC lifecycle, where each method is represented by the mean track of 50 ensemble members \(i\.e\., multiple possible tracks\), and the black curve denotes the best\-track positions\. The numbered panels show the spatial distributions of the 50 ensemble\-predicted TC positions for representative cases selected from the main panel, covering early, middle, and late lifecycle stages as well as unusual looping\-track predictions\. In each numbered panel, circles indicate the spatial dispersion of the forecast members, and the valuesNN,GG, andTTdenote the corresponding dispersion radii of NeuralGCM\-ENS, GenCast\-ENS, and Tianmu\-TC\-ENS, respectively\. The black star denotes the best\-track position\. The selected cases reveal different performance regimes, including accurate and biased situations\.Formulti\-trend prediction, a formulation analogous to ensemble forecasting in numerical weather prediction, Table[3](https://arxiv.org/html/2608.18500#Sx2.T3)evaluates 50 sampled trends from four complementary perspectives\. Mean Error and Std\. Error denote the mean and standard deviation of the errors across the 50 ensemble members or sampled trends, respectively\. Spread measures the dispersion of the 50 forecasts and is used as an uncertainty\-related indicator of the forecast distribution\. Specifically, for track prediction, Spread is computed as the 90th\-percentile radius of the distances between the 50 ensemble\-predicted TC positions and the ensemble center, while for pressure and wind speed it is computed from the dispersion of the 50 predicted scalar values\. Best\-of\-50 evaluates the error of the best member among the 50 forecasts, reflecting whether the forecast distribution contains a highly accurate candidate\.

First, Tianmu\-TC\-ENS improves the accuracy and stability of multi\-trend prediction\. For track prediction, Tianmu\-TC\-ENS reduces the overall Mean Error from 109\.04 km and 133\.36 km to 68\.49 km compared with NeuralGCM\-ENS and GenCast\-ENS, respectively\. The overall Std\. Error is also reduced from 44\.33 km and 77\.57 km to 21\.26 km, indicating that the sampled trends are not only closer to the ground truth on average, but also more stable across different members\. For intensity prediction, Tianmu\-TC\-ENS similarly reduces the overall Mean Error from 9\.88 hPa to 3\.26 hPa for pressure and from 9\.31 m/s to 2\.07 m/s for wind speed, further confirming its advantage in multi\-attribute TC prediction\.

Second, Tianmu\-TC\-ENS produces a more compact forecast distribution while preserving accurate candidates\. For track prediction, the overall Spread is reduced from 119\.85 km and 142\.24 km to 53\.32 km, showing that Tianmu\-TC\-ENS substantially suppresses unnecessary spatial dispersion among the 50 trends\. For pressure and wind speed, Tianmu\-TC\-ENS also reduces the overall Spread from 1\.69 hPa to 1\.29 hPa and from 1\.39 m/s to 0\.81 m/s, respectively\. Meanwhile, the Best\-of\-50 results show that this compactness does not come from a simple collapse of the forecast distribution: Tianmu\-TC\-ENS achieves the best overall Best\-of\-50 errors for track, pressure, and wind speed prediction, although GenCast\-ENS is slightly better for the \+24 h track Best\-of\-50 case\.

Overall, these results provide quantitative evidence that the physically constrained multi\-trend formulation of Tianmu\-TC\-ENS can improve the accuracy of sampled TC evolution trends while reducing unnecessary predictive dispersion\. This supports the central objective of Tianmu\-TC\-ENS, namely generating physically consistent and uncertainty\-compact multi\-trend forecasts for TC track and intensity prediction\.

Figure[7](https://arxiv.org/html/2608.18500#Sx2.F7)further evaluates the lifecycle behavior of multi\-trend track forecasts for Typhoon ATSANI\. Rather than focusing on a single initialization time, the figure examines forecasts throughout the entire TC lifecycle\. ATSANI is selected because it provides diverse and challenging forecast scenarios, including unusual looping\-track predictions, a pronounced track deflection near southern Taiwan, and subsequent evolution close to coastal and mountainous regions\. These scenarios enable forecast stability to be evaluated under varying environmental and topographic conditions\. Specifically, each method is represented by the mean track of 50 ensemble members, allowing a direct and fair comparison of lifecycle\-scale track evolution, the dispersion of 50 ensemble\-predicted TC positions, and the ability of the forecast distribution to cover the best\-track position\. As shown in the predicted positions panels of Figure[7](https://arxiv.org/html/2608.18500#Sx2.F7), the valuesNN,GG, andTTdenote the dispersion radii of NeuralGCM\-ENS, GenCast\-ENS, and Tianmu\-TC\-ENS, respectively, with smaller values indicating lower spatial spread among the 50 forecasts\. Tianmu\-TC\-ENS produces a more compact multi\-trend forecast distribution, indicating lower forecast uncertainty while still covering the best\-track position in most selected cases\. In cases①–⑤, the orange Tianmu\-TC\-ENS predicted positions are concentrated within a much smaller circle than the blue NeuralGCM\-ENS and green GenCast\-ENS predicted positions, and the best\-track position is generally located inside or close to the Tianmu\-TC\-ENS distribution\. For example, Tianmu\-TC\-ENS shows notably smaller dispersion in cases①, ②, and ④, where the orange predicted positions remain tightly clustered around the best\-track position\. Meanwhile, NeuralGCM\-ENS and GenCast\-ENS usually exhibit larger circles, indicating stronger ensemble spread\. This larger spread can also be beneficial, as it enables the global ensemble forecasts to cover the best\-track position in several cases, although at the cost of higher uncertainty\. These results suggest that Tianmu\-TC\-ENS provides a more concentrated multi\-trend forecast while maintaining useful best\-track coverage in most stages of the ATSANI lifecycle\.

The selected cases also reveal several limitations\. In case⑥, Tianmu\-TC\-ENS shows the largest positional bias among the three methods, whereas GenCast\-ENS places its members closer to the best\-track position\. This error may be related to the relatively fast storm displacement in the preceding stages, which increases the difficulty of short\-term track extrapolation\. In case⑦, all three methods exhibit large errors, while NeuralGCM\-ENS and GenCast\-ENS show particularly large dispersion\. This is likely associated with the late\-stage land interaction and the rapidly changing storm environment near landfall\. In addition, although GenCast\-ENS usually benefits from its larger spread, it fails to cover the best\-track position in case③\. These results indicate that Tianmu\-TC\-ENS can substantially reduce excessive ensemble dispersion and provide stable multi\-trend forecasts in most situations\. However, achieving both low dispersion and reliable best\-track coverage remains an important direction for future improvement\.

### Robust forecasting of abrupt track changes and RI/RW events

To evaluate the performance of Tianmu\-TC under challenging conditions, we examined its predictions for abrupt track changes and rapid intensity changes, which remain particularly difficult to forecast\[[42](https://arxiv.org/html/2608.18500#bib.bib42)\]\. As shown in Table[1](https://arxiv.org/html/2608.18500#Sx2.T1), across abrupt track changes \(denoted asanomalyin the table\), landfall,rapid intensification \(RI\), andrapid weakening \(RW\)cases, Tianmu\-TC consistently achieves the lowest forecast errors compared with the large meteorological models and operational NWP models \(Pangu, Fengwu, ECMWF\), demonstrating superior robustness under complex conditions\.Anomalyincludes circular movement, turning on the ocean, and V\-shaped tracks;landfallrefers to samples in which landfall occurs during at least half of the forecast period \(24 hours\)\. For track anomalies, the errors of Tianmu\-TC are reduced by up to a factor of six relative to Pangu and by a factor of three relative to Fengwu\. In landfall prediction, it substantially improves accuracy over ECMWF and achieves more than a twofold error reduction compared with Pangu\. For intensity variation, Tianmu\-TC reduces RI errors to less than one\-tenth of those from operational NWP models and also achieves the best performance under RW cases\. Since ECMWF dataset is limited, we additionally report comparisons on larger samples \(aligned with Tianmu\-TC, Pangu, Fengwu, see Table[S1](https://arxiv.org/html/2608.18500#A1.T1)\)\. The performance trends remain consistent, further validating the robustness of Tianmu\-TC\.

Besides, the corresponding visualization results are shown in Figure[6](https://arxiv.org/html/2608.18500#Sx2.F6)\. For tracks, four representative scenarios were analyzed: circular movement, post\-landfall prediction, turning on the ocean, and V\-shaped tracks\. For intensity, we focused on cases ofrapid intensificationandweakening\. It can be observed that Tianmu\-TC effectively captures the dynamics and transitions in these scenarios, providing reliable track forecasts even under irregular movements and maintaining high accuracy in predicting rapid intensity changes\.

These results confirm that the proposed physics\-constrained generative framework is capable of handling the most difficult TC scenarios, where both track and intensity variations remain highly uncertain in conventional approaches\.

### Accurate real\-time prediction under data sparsity

Given that the ERA5 reanalysis data used to encode the atmospheric environment constraint are not available in real time, we examined Tianmu\-TC’s performance in cases where ERA5 data were missing during testing \(see Table[4](https://arxiv.org/html/2608.18500#Sx2.T4)\)\. In such cases, critical environmental inputs—including wind fields \(uu,vv\) and geopotential height \(zz\) representing high\-pressure systems—were unavailable\.

By comparing Rows 9 and 12 in Table[4](https://arxiv.org/html/2608.18500#Sx2.T4), we find that the absence of ERA5 data during inference has negligible impact on prediction accuracy\. Notably, Tianmu\-TC maintains comparable performance using only the TC historical attributes from the previous 24 hours\.

This finding highlights the model’s robustness under data sparsity, where only partial inputs are available, and confirms its practical potential as a real\-time, end\-to\-end TC forecasting method, even in the absence of comprehensive environmental data\.

Table 4:Testing results under missing input conditions\.The last row corresponds to the complete model \(Tianmu\-TC\)\.Xh​i​sX\_\{his\},XE​R​A​5X\_\{ERA5\},Xt​i​l​tX\_\{tilt\}refer to TC historical attribute values, ERA5 data, and TC tilt values, respectively\.\+6​h\+6hto\+24​h\+24hindicate prediction horizons in hours\. Lower values represent better performance\.RowReal\-timeDataTrack Error \(km\)Pressure Error \(hPa\)Wind Speed Error \(m/s\)Xh​i​sX\_\{his\}XE​R​A​5X\_\{ERA5\}Xt​i​l​tX\_\{tilt\}\+6h\+12h\+18h\+24h\+6h\+12h\+18h\+24h\+6h\+12h\+18h\+24h1yes///114\.88228\.64331\.53428\.402\.784\.936\.046\.941\.572\.843\.724\.502yes6h//57\.30104\.16146\.70191\.522\.203\.083\.624\.151\.341\.932\.262\.583yes12h//34\.7157\.5283\.13116\.121\.311\.941\.912\.620\.621\.251\.171\.544yes18h//26\.4339\.9761\.8893\.991\.102\.031\.972\.670\.591\.131\.091\.475yes24h//20\.9730\.2146\.4577\.190\.961\.601\.612\.330\.500\.860\.861\.256yes30h//19\.1226\.4843\.1074\.860\.991\.611\.612\.280\.500\.830\.821\.237yes36h//18\.2123\.6039\.8171\.610\.961\.541\.552\.230\.490\.810\.801\.218yes42h//17\.9922\.7738\.7070\.160\.961\.531\.562\.240\.490\.810\.811\.209yes48h//17\.8021\.8937\.5169\.030\.961\.521\.552\.230\.490\.800\.801\.2010no48h✓/17\.8021\.8937\.5169\.030\.961\.521\.552\.230\.490\.800\.801\.2011no48h/✓17\.8021\.9137\.5569\.080\.961\.521\.552\.230\.490\.800\.801\.2012no48h✓✓17\.8021\.9137\.5569\.080\.961\.521\.552\.230\.490\.800\.801\.20![Refer to caption](https://arxiv.org/html/2608.18500v1/framework/saliency_mult_task_xiufu_small.jpg)

Figure 8:Analyzing the correlation between track and intensity prediction\.\(A\) Ablation results when predictions are made individually and jointly\. \(B\) Saliency maps of historical attribute values\. The columns correspond to \(a\) track, \(b\) pressure, and \(c\) wind speed\.
### Intensity prediction is dependent on track

A core challenge in tropical cyclone forecasting is the simultaneous prediction of both track and intensity\. Traditionally, these two tasks have been modeled independently, reflecting the fact that they are governed by different dominant physical mechanisms: track prediction is mainly influenced by the steering flow, such as the position and strength of the subtropical high\[[43](https://arxiv.org/html/2608.18500#bib.bib43)\], while intensity prediction depends more heavily on factors like sea surface temperature, vertical wind shear, and ocean\-atmosphere interactions\[[44](https://arxiv.org/html/2608.18500#bib.bib44),[45](https://arxiv.org/html/2608.18500#bib.bib45)\]\.

However, growing evidence suggests a strong coupling between track and intensity\. In our proposed approach, both attributes are predicted jointly in an end\-to\-end framework, allowing the model to learn potential cross\-task dependencies\. As shown in Figure[8](https://arxiv.org/html/2608.18500#Sx2.F8)A, this joint prediction leads to superior performance compared to separate forecasting\. Furthermore, the saliency analysis in Figure[8](https://arxiv.org/html/2608.18500#Sx2.F8)B highlights how intensity forecasts are sensitive to track\-related features, providing further evidence for the interdependence between these two TC attributes\.

A key conclusion is that while TC track and intensity are physically interconnected, their dependencies are asymmetric\. Track prediction appears to be largely independent of intensity, whereas intensity prediction is strongly dependent on the track\. As illustrated in Figure[8](https://arxiv.org/html/2608.18500#Sx2.F8), this is primarily because the environmental field is the dominant driver of TC intensity changes\. As a TC moves along its track, it encounters different environmental conditions \(e\.g\., ocean heat content, wind shear\), which in turn influence its structure and intensity\. Thus, accurate track forecasts provide crucial context for predicting intensity evolution, reinforcing the necessity of jointly modeling both attributes\.

## Discussion

As shown in Table[5](https://arxiv.org/html/2608.18500#Sx3.T5), Tianmu\-TC’s intensity forecasting maintains a clear advantage over the China Meteorological Observatory \(CMO\), and the European Centre for Medium\-Range Weather Forecasts global model, even at the \+72 h forecast horizon\. However, it must be acknowledged that this work still has some limitations\. Track forecasting performance becomes suboptimal beyond 24 h, mainly due to the lack of explicit atmospheric dynamic constraints, and limited long\-range temporal dependency modeling capability\. Addressing these issues will be the focus of our future work\.

Table 5:Compare the prediction errors at a 72\-hour forecast horizon with authoritative NWP methods in WP basin\.The testing years are from 2017 to 2023\.Prediction Horizon \(hours\)Error TypeMethod\+6​h\+6h\+12​h\+12h\+24​h\+24h\+48​h\+48h\+72​h\+72hTrack Error \(km\)CMO32\.452\.975\.5132\.1201\.1ECMWF42\.347\.065\.8116\.8193\.6Tianmu\-TC17\.821\.969\.1238\.7438\.9Pressure Error \(hPa\)CMO2\.64\.376\.259\.0111\.25ECMWF10\.911\.211\.812\.911\.7Tianmu\-TC1\.01\.52\.14\.86\.2Wind Speed Error \(m/s\)CMO1\.52\.33\.54\.96\.0ECMWF7\.27\.27\.27\.46\.9Tianmu\-TC0\.50\.81\.23\.04\.0
## Materials and Methods

### Experimental Setup

For fair comparison, following TCNM\[[31](https://arxiv.org/html/2608.18500#bib.bib31)\], the sampling number is 6, which means that the model generates 6 possible tendencies\. Similarly, Tianmu\-TC inputs 48 hours \(nn= 8, the time resolution is 6 hours\) of TC historical data and outputs 24 hours \(mm= 4\) of future TC attribute data\. Tianmu\-TC was deployed on the PyTorch framework\. Training was performed using Adam Optimizer with a learning rate of 0\.0001, a batch size of 256, and a duration of 48 hours\.

All experiments, including other deep learning methods used for comparison, were conducted on an NVIDIA RTX A6000 GPU\. Random seeds were fixed during both training and testing to ensure reproducibility\. The number of training epochs was set to 350 based on model convergence criteria\. Absolute errors were computed between predicted results and ground truth for track \(distance in km\), pressure \(hPa\), and wind speed \(m/s\)\.

### Data preparation and availability

We organize a new dataset, which mainly consists of three parts: historical TC attributes, ERA5 data representing external environment, and TC tilt metrics representing TC internal structure\. This dataset encompassesallthe 1786 TCs data from 1950 to 2023 over the Western North Pacific\. So the datasets contain sufficiently diverse conditions of TCs\. 80% of the TC data from 1950 to 2016 was allocated for training, 20% for validation, and the data from 2017 to 2023 were reserved for testing\. Environmental data from MGTCF—including future move direction, intensity class and future intensity change direction—were also utilized\. Further details are described below\.

Historical TC attributesare derived from International Best Track Archive for Climate Stewardship \(IBTrACS\[[37](https://arxiv.org/html/2608.18500#bib.bib37)\]\), including essential TC center information such as longitude \(∘\), latitude \(∘\), pressure \(hPa\) and wind speed \(m/s\)\. Time resolution is 6 hours\.

ERA5 dataare derived from the fifth\-generation atmospheric reanalysis of the global climate \(ERA5\)\[[46](https://arxiv.org/html/2608.18500#bib.bib46)\]by the European Centre for Medium\-Range Weather Forecasts \(ECMWF\)\. We selected the geopotential height data at 200 hPa and 500 hPa, represented asz​200z200andz​500z500, which indicate high\-pressure systems, as well as the large\-scale wind field data, includingu​200u200,u​500u500,u​700u700,u​850u850,v​200v200,v​500v500,v​700v700, andv​850v850\. Here,uurepresents the zonal wind component \(east\-west direction\), andvvrepresents the meridional wind component \(north\-south direction\)\. The 200 hPa data reflect upper\-level information of tropical cyclones, the 500 hPa data represent mid\-level information, while the 700 hPa and 850 hPa data capture lower\-level \(near\-surface\) information of tropical cyclones\. The image size is 20∘×\\times20∘, centered on the TC eye, covering a radius of 10∘around the TC’s eye\. The spatial resolution is 0\.25∘, resulting in a data size of 81×\\times81\.

TC tilt metricsare derived from vorticity data, which are calculated from wind field data\. As shown in Figure[1](https://arxiv.org/html/2608.18500#Sx1.F1)D, vorticity at 200 hPa, 500 hPa, and 850 hPa are first computed using the correspondinguuandvvcomponents for each pressure level, as described in Equation[1](https://arxiv.org/html/2608.18500#Sx4.E1)\. The resulting vorticity data for these three pressure levels are referred to asv​o​200vo200,v​o​500vo500, andv​o​850vo850, respectively\. Next, takingv​o​200vo200as an example, as shown in Figure[1](https://arxiv.org/html/2608.18500#Sx1.F1)D, the image size is 81×\\times81\. The TC eye is marked by a black diamond at the center of the image, located at position \(41, 41\), denoted as \(♢x\\diamondsuit\_\{x\},♢y\\diamondsuit\_\{y\}\)\. The three local minima closest to the TC eye are marked as red circles, and their average coordinates are calculated to obtain the center of the three local minima, which is marked as a blue triangle \(△xp​l\\bigtriangleup\_\{x\}^\{pl\},△yp​l\\bigtriangleup\_\{y\}^\{pl\}\)\.p​l∈\{200,500,850\}pl\\in\\\{200,500,850\\\}refers to pressure level\.

v​o=∂v∂x−∂u∂yvo=\\frac\{\\partial v\}\{\\partial x\}\-\\frac\{\\partial u\}\{\\partial y\}\(1\)∂x=∂y×cos⁡\(π180×avg\_latitude\)\\partial x=\\partial y\\times\\cos\\left\(\\frac\{\\pi\}\{180\}\\times\\text\{avg\\\_latitude\}\\right\)\(2\)where∂y=111\.32×0\.25\\partial y=111\.32\\times 0\.25, 111\.32 represents the distance in kilometers corresponding to 1 degree of longitude\. The value 0\.25 refers to the spatial resolution\. Thus,∂y\\partial yrepresents the distance between data points in the longitudinal direction, with units in kilometers\. Similarly,∂x\\partial xrepresents the distance between data points in the latitudinal direction, which varies depending on the latitude\. Here, it is simplified to the average latitude \(avg\_latitude\)\. Since the training data is based on tropical cyclones in the WP, the average latitude is set to 15∘\.

Dcen200=\(△x200−♢x\)2\+\(△y200−♢y\)2D\_\{\\text\{cen200\}\}=\\sqrt\{\\left\(\\bigtriangleup\_\{x\}^\{200\}\-\\diamondsuit\_\{x\}\\right\)^\{2\}\+\\left\(\\bigtriangleup\_\{y\}^\{200\}\-\\diamondsuit\_\{y\}\\right\)^\{2\}\}\(3\)Dcen850=\(△x850−♢x\)2\+\(△y850−♢y\)2D\_\{\\text\{cen850\}\}=\\sqrt\{\\left\(\\bigtriangleup\_\{x\}^\{850\}\-\\diamondsuit\_\{x\}\\right\)^\{2\}\+\\left\(\\bigtriangleup\_\{y\}^\{850\}\-\\diamondsuit\_\{y\}\\right\)^\{2\}\}\(4\)D28=11\+\(△x200−△x850\)2\+\(△y200−△y850\)2D\_\{\\text\{28\}\}=\\frac\{1\}\{1\+\\sqrt\{\\left\(\\bigtriangleup\_\{x\}^\{200\}\-\\bigtriangleup\_\{x\}^\{850\}\\right\)^\{2\}\+\\left\(\\bigtriangleup\_\{y\}^\{200\}\-\\bigtriangleup\_\{y\}^\{850\}\\right\)^\{2\}\}\}\(5\)D58=11\+\(△x500−△x850\)2\+\(△y500−△y850\)2D\_\{\\text\{58\}\}=\\frac\{1\}\{1\+\\sqrt\{\\left\(\\bigtriangleup\_\{x\}^\{500\}\-\\bigtriangleup\_\{x\}^\{850\}\\right\)^\{2\}\+\\left\(\\bigtriangleup\_\{y\}^\{500\}\-\\bigtriangleup\_\{y\}^\{850\}\\right\)^\{2\}\}\}\(6\)whereD28D\_\{28\}represents the alignment between 200 hPa and 850 hPa;D58D\_\{58\}represents the alignment between 500 hPa and 850 hPa\.Dcen200D\_\{\\text\{cen200\}\}andDcen850D\_\{\\text\{cen850\}\}measure the distance between the TC eye and local minima center at 200 hPa and 850 hPa, respectively\. The same calculations were performed for each historical time step, resulting in a sequential series of TC tilt metrics\.

### Tianmu\-TC network design

Overall, Tianmu\-TC can be divided into three main modules: constraint encoder, cross\-modal fusion, and constraint guiding denoising\. The constraint guiding denoising consists ofKKPhysics\-constraints Denoising Blocks\.

Constraint Encoder: How to design the 3 constraints?

LSTM\-based Encoder for encoding TC track and intensity, which is designed based on the encoder from Trajectron\+\+\[[47](https://arxiv.org/html/2608.18500#bib.bib47)\]\. The input data has a shape of\(T,4\)\(T,4\), whereT=8T=8, representing the historical attributes of the tropical cyclone over8×68\\times 6hours=48=48hours\. The four values correspond to longitude, latitude, pressure, and wind speed\. After encoding, the outputF1F\_\{1\}has a shape of\(256\)\(256\)\.

CNN\-based Encoderencodes selected ERA5 data\. The data includesu​200u200,u​500u500,u​700u700,u​850u850,v​200v200,v​500v500,v​700v700,v​850v850,z​200z200, andz​500z500, totaling 10 variables\. Therefore, the data shape is\(T,10,81,81\)\(T,10,81,81\), whereT=8T=8\. The predicted attributes are all the attributes in and around the TC center, which we define as a highly task\-relevant area\. To enhance task\-relevant information, the Central Information Enhancement \(CenIE\) block enhances the center point of the original ERA5 data \(see in the Figure[S1](https://arxiv.org/html/2608.18500#A1.F1)of supplementary materials\)\. As a result, the input data size becomes\(T,20,81,81\)\(T,20,81,81\)\. We implement channel fusion through 6 layers of CNN with ReLU activation functions and 3 layers of spatial attention\. This process outputs data with a shape of\(T,1,16,16\)\(T,1,16,16\), where the channel dimensionCCis reduced from 20 to 1, and the height \(HH\) and width \(WW\) are downscaled\.

Another LSTM\-based encoder encodes TC tilt metrics\. The input data has a shape of\(T,4\)\(T,4\), whereT=8T=8\. The 4 values represent four metrics:Dcen200,Dcen850,D28,D58D\_\{\\text\{cen200\}\},D\_\{\\text\{cen850\}\},D\_\{28\},D\_\{58\}\. A single\-layer LSTM is used to encode the data, resulting inF3F\_\{3\}with a size of256256\.

Cross\-modal Fusion: Establishing relationships among the three constraints\.

The outputs from the constraint encoder have shapes ofF1F\_\{1\}\(256\),\(T,1,16,16\)\(T,1,16,16\), andF3F\_\{3\}\(256\), respectively\. From the input data shapes, it can also be observed that TC track and intensity, as well as TC tilt values, are one\-dimensional attribute values, while ERA5 data is two\-dimensional image data\. Therefore, it is necessary to match and fuse these two modalities\.

As shown in Figure[2](https://arxiv.org/html/2608.18500#Sx1.F2), we select the encoding result \(256\) of TC track and intensity, which is semantically most relevant to the prediction target, and first reshape it into a shape of\(1,16,16\)\(1,16,16\)\. We initializeHHto 0, with a shape of\(1,16,16\)\(1,16,16\)\.

The iteration is performed over each historical time step, withTTranging from 1 to 8\. AtT=1T=1, the corresponding slice from\(T,1,16,16\)\(T,1,16,16\)with shape\(1,16,16\)\(1,16,16\)is selected and concatenated with the reshaped encodings of TC track, intensity, andHH, yielding a final shape of\(3,16,16\)\(3,16,16\)\. This is then sent into theSA\(self\-attention module\) to capture semantic associations among the three, producing a newHHwith a shape of\(1,16,16\)\(1,16,16\)\. This process is repeated to updateHHuntilT=8T=8\. After reshape the finalHH, the feature size is \(256\), which is the result of temporal fusion of ERA5 data, denoted asF2F\_\{2\}\.

Finally, a 4\-layer Transformer encoder is used to fuseF1F\_\{1\},F2F\_\{2\}, andF3F\_\{3\}to obtainC1C\_\{1\},C2C\_\{2\}, andC3C\_\{3\}\. In fact, these three constraints are fused into a complete feature of size \(256\), denoted asCC\.

Constraint guiding denoising

A diffusion sequence is\(y0,y1,…,yK\)\(y\_\{0\},y\_\{1\},\.\.\.,y\_\{K\}\), whereKKis the maximum diffusion step\. The diffusion process gradually adds noise into the ground truth regiony0y\_\{0\}of TC forecasting\. Correspondingly, the reverse diffusion sequence is\(yK,yK−1,…,y0\)\(y\_\{K\},y\_\{K\-1\},\.\.\.,y\_\{0\}\)\. This process utilizes historical data as conditions, and gradually denoises the standard GaussianyKy\_\{K\}to obtain the 1D future TC attributesy0y\_\{0\}\. This process reduces prediction noise\.

The diffusion processis defined asq⁡\(yk\|y0\):=𝒩⁡\(yk,αk¯​y0,\(1−αk¯\)​I\)q\(y\_\{k\}\|y\_\{0\}\):=\\mathcal\{N\}\(y\_\{k\};\\sqrt\{\\bar\{\\alpha\_\{k\}\}\}y\_\{0\},\(1\-\\bar\{\\alpha\_\{k\}\}\)\\mathrm\{I\}\), thus

yk=αk¯​y0\+\(1−αk¯\)​εy\_\{k\}=\\sqrt\{\\bar\{\\alpha\_\{k\}\}\}y\_\{0\}\+\\sqrt\{\(1\-\\bar\{\\alpha\_\{k\}\}\)\}\\varepsilon\(7\)αk¯=∏s=1kαs\\bar\{\\alpha\_\{k\}\}=\\prod\_\{s=1\}^\{k\}\\alpha\_\{s\},α1,α2,…​αk\\alpha\_\{1\},\\alpha\_\{2\},\.\.\.\\alpha\_\{k\}are fixed variance schedulers, the noise variableε∼𝒩⁡\(0,I\)\\varepsilon\\sim\\mathcal\{N\}\(0,\\mathrm\{I\}\)\. As the diffusion stepkkincreases,yk∼𝒩⁡\(0,I\)y\_\{k\}\\sim\\mathcal\{N\}\(0,\\mathrm\{I\}\)\. This meansz0z\_\{0\}is gradually destroyed into a Gaussian noise distribution\. The training loss is:

L⁡\(θ,ψ\)=𝔼ε,z0,k​‖ε−ε\(θ,ψ\)​\(yk,k,X\)‖L\(\\theta,\\psi\)=\\mathbb\{E\}\_\{\\varepsilon,z\_\{0\},k\}\|\|\\varepsilon\-\\varepsilon\_\{\(\\theta,\\psi\)\}\(y\_\{k\},k,X\)\|\|\(8\)θ\\thetaare parameters of Physics\-constraints Denoising model,ψ\\psiare parameters of constraint encoder and cross\-modal fusion,XXare the input of the model\.

The generation of increasingly reliable TC predictions is formulated as areverse diffusion process, guided by three specifically designed constraints that progressively reduce uncertainty\.

The reverse diffusion process with Bi\-condition is defined as:

pθ​\(yk−1\|yk,C\):=𝒩⁡\(yk−1,μθ​\(yk,k,C\),Σθ​\(yk,k\)\)\\textstyle p\_\{\\theta\}\(y\_\{k\-1\}\|y\_\{k\},C\):=\\mathcal\{N\}\(y\_\{k\-1\};\\mu\_\{\\theta\}\(y\_\{k\},k,C\);\\Sigma\_\{\\theta\}\(y\_\{k\},k\)\)\(9\)μθ​\(⋅\)=1αk​\(yk−βk1−αk¯​εθ​\(yk,k,C\)\)\\mu\_\{\\theta\}\(\\cdot\)=\\frac\{1\}\{\\sqrt\{\\alpha\_\{k\}\}\}\(y\_\{k\}\-\\frac\{\\beta\_\{k\}\}\{\\sqrt\{1\-\\bar\{\\alpha\_\{k\}\}\}\}\\varepsilon\_\{\\theta\}\(y\_\{k\},k,C\)\)\(10\)Σθ​\(⋅\)=σk2​I=βk​I\\Sigma\_\{\\theta\}\(\\cdot\)=\\sigma\_\{k\}^\{2\}\\mathrm\{I\}=\\beta\_\{k\}\\mathrm\{I\}\(11\)thus:

yk−1=1αk​\(yk−βk1−αk¯​εθ​\(yk,k,C\)\)\+βk​ey\_\{k\-1\}=\\frac\{1\}\{\\sqrt\{\\alpha\_\{k\}\}\}\\left\(y\_\{k\}\-\\frac\{\\beta\_\{k\}\}\{\\sqrt\{1\-\\bar\{\\alpha\_\{k\}\}\}\}\\varepsilon\_\{\\theta\}\(y\_\{k\},k,C\)\\right\)\+\\sqrt\{\\beta\_\{k\}\}e\(12\)In Equation[12](https://arxiv.org/html/2608.18500#Sx4.E12),eeis a random variable in standard Gaussian Distribution,βk=1−αk\\beta\_\{k\}=1\-\\alpha\_\{k\}\.εθ\\varepsilon\_\{\\theta\}is implemented by the Physics\-constraints denoising block in Figure[2](https://arxiv.org/html/2608.18500#Sx1.F2)\. The output of the block is:

yk′=εθ​\(yk,k,C\)y\_\{k\}^\{\\prime\}=\\varepsilon\_\{\\theta\}\(y\_\{k\},k,C\)\(13\)inspired by MID\[[48](https://arxiv.org/html/2608.18500#bib.bib48)\],εθ\\varepsilon\_\{\\theta\}is implemented mainly by a Transformer encoder and 4\-layer gating units, which use constraintsCCto generate Gate and bias, and perform a gating transform for latent noiseyky\_\{k\}\.

## References and Notes

- \[1\]T\. R\. Knutson, J\. L\. McBride, J\. Chan, K\. Emanuel, G\. Holland, C\. Landsea, I\. Held, J\. P\. Kossin, A\. K\. Srivastava, M\. Sugi, Tropical cyclones and climate change\.*Nat\. Geosci\.*3, 157–163 \(2010\)\.
- \[2\]J\. Callaghan, R\. Smith, The relationship between maximum surface wind speeds and central pressure in tropical cyclones\.Aust\. Meteorol\. Mag\.47\(3\), 191–202 \(1998\)\.
- \[3\]J\. A\. Knaff, R\. M\. Zehr, Reexamination of tropical cyclone wind–pressure relationships\.Weather Forecast\.22\(1\), 71–88 \(2007\)\.
- \[4\]P\. Bauer, A\. Thorpe, G\. Brunet, The quiet revolution of numerical weather prediction\.Nature525\(7567\), 47–55 \(2015\)\.
- \[5\]R\. Kimura, Numerical weather prediction\.J\. Wind Eng\. Ind\. Aerodyn\.90\(12\-15\), 1403–1414 \(2002\)\.
- \[6\]National Hurricane Center, Forecast Verification\.National Hurricane Center,https://www\.nhc\.noaa\.gov/verification/verify3\.shtml\. Accessed August 5, 2025\.
- \[7\]Leutbecher, Martin and Palmer, Tim N, Ensemble forecasting, InJournal of computational physics, 227, 7, 3515–3539 \(2008\)\.
- \[8\]ECMWF, IFS documentation\-Cy37r2, operational implementation\. \(European Centre for Medium\-Range Weather Forecasts, 2011\)\.
- \[9\]Central Meteorological Observatory CMO, Typhoon network of Central Meteorological Observatory, 2019,http://typhoon\.nmc\.cn/web\.html\(Accessed: 2003\-03\-10\)\.
- \[10\]F\. Sanders, A\. C\. Pike, J\. P\. Gaertner, A barotropic model for operational prediction of tracks of tropical storms\.J\. Appl\. Meteorol\. Climatol\.14, 265–280 \(1975\)\.
- \[11\]I\. Price, A\. Sanchez\-Gonzalez, F\. Alet, T\. R\. Andersson, A\. El\-Kadi, D\. Masters, T\. Ewalds, J\. Stott, S\. Mohamed, P\. Battaglia, et al\., Probabilistic weather forecasting with machine learning\.Nature637, 84–90 \(2025\)\.
- \[12\]Z\. Gao, X\. Shi, B\. Han, H\. Wang, X\. Jin, D\. Maddix, C\. Li, Y\. Wang, et al\., PreDiff: Precipitation nowcasting with latent diffusion models\.Advances in Neural Information Processing Systems36, 78621–78656 \(2023\)\.
- \[13\]J\. Liu, C\. Xu, S\. Han, L\. Song, P\. Wang, T\. Zhang, Uncertainty\-aware precipitation nowcasting with diffusion model simulating precipitation evolution processes\.Eng\. Appl\. Artif\. Intell\.178, 115140 \(2026\)\.
- \[14\]Z\. Chen, B\. Peng, T\. Zhai, D\. Adu\-Ampratwum, X\. Ning, Generating 3D small binding molecules using shape\-conditioned diffusion models with guidance\.Nat\. Mach\. Intell\.1–13 \(2025\)\.
- \[15\]M\. Mardani, N\. Brenowitz, Y\. Cohen, J\. Pathak, C\. Y\. Chen, C\. C\. Liu, M\. Pritchard, et al\., Residual corrective diffusion modeling for km\-scale atmospheric downscaling\.Commun\. Earth Environ\.6, 124 \(2025\)\.
- \[16\]C\. Wang, B\. Wang, Impacts of the South Asian high on tropical cyclone genesis in the South China Sea\.Clim\. Dyn\.56, 2279–2288 \(2021\)\.
- \[17\]L\. Wu, Z\. Wen, R\. Huang, R\. Wu, Possible linkage between the monsoon trough variability and the tropical cyclone activity over the western North Pacific\.Mon\. Weather Rev\.140, 140–150 \(2012\)\.
- \[18\]Y\. Sun, Z\. Zhong, L\. Yi, T\. Li, M\. Chen, H\. Wan, Y\. Wang, K\. Zhong, Dependence of the relationship between the tropical cyclone track and western Pacific subtropical high intensity on initial storm size: A numerical investigation\.J\. Geophys\. Res\. Atmos\.120, 11–451 \(2015\)\.
- \[19\]M\. Moradi Kordmahalleh, M\. Gorji Sefidmazgi, A\. Homaifar, A sparse recurrent neural network for trajectory prediction of Atlantic hurricanes\. InProc\. Genetic and Evolutionary Comput\. Conf\. \(GECCO\), 957–964 \(2016\)\.
- \[20\]S\. Alemany, J\. Beltran, A\. Perez, S\. Ganzfried, Predicting hurricane trajectories using a recurrent neural network\. InProc\. AAAI Conf\. Artif\. Intell\.33, 468–475 \(2019\)\.
- \[21\]K\. Cho, B\. van Merriënboer, C\. Gulcehre, D\. Bahdanau, F\. Bougares, H\. Schwenk, Y\. Bengio, Learning phrase representations using RNN encoder–decoder for statistical machine translation\. InProc\. 2014 Conf\. Empir\. Methods Nat\. Lang\. Process\. \(EMNLP\), 1724–1734 \(2014\)\.
- \[22\]T\. Song, Y\. Li, F\. Meng, P\. Xie, D\. Xu, A novel deep learning model by BiGRU with attention mechanism for tropical cyclone track prediction in the Northwest Pacific\.J\. Appl\. Meteorol\. Climatol\.61, 3–12 \(2022\)\.
- \[23\]S\. Gao, P\. Zhao, B\. Pan, Y\. Li, M\. Zhou, J\. Xu, S\. Zhong, Z\. Shi, A nowcasting model for the prediction of typhoon tracks based on a long short term memory neural network\.Acta Oceanol\. Sin\.37, 8–12 \(2018\)\.
- \[24\]S\. Kim, H\. Kim, J\. Lee, S\. Yoon, S\. E\. Kahou, K\. Kashinath, Mr\. Prabhat, Deep\-Hurricane\-Tracker: Tracking and forecasting extreme climate events\. InProc\. IEEE Winter Conf\. Appl\. Comput\. Vis\. \(WACV\), 1761–1769 \(2019\)\.
- \[25\]X\. Geng, Z\. Liu, Z\. Shi, Spatio\-Temporal Alignment and Track\-To\-Velocity Module for Tropical Cyclone Forecast\.Remote Sens\.15, 4938 \(2023\)\.
- \[26\]Y\. Wu, X\. Geng, Z\. Liu, Z\. Shi, Tropical cyclone forecast using multitask deep learning framework\.IEEE Geosci\. Remote Sens\. Lett\.19, 1–5 \(2021\)\.
- \[27\]B\. Pan, X\. Xu, Z\. Shi, Tropical cyclone intensity prediction based on recurrent neural networks\.Electron\. Lett\.55, 413–415 \(2019\)\.
- \[28\]C\. Huang, C\. Bai, S\. Chan, J\. Zhang, MMSTN: A multi\-modal spatial\-temporal network for tropical cyclone short\-term prediction\.Geophys\. Res\. Lett\.49, e2021GL096898 \(2022\)\.
- \[29\]C\. Huang, C\. Bai, S\. Chan, J\. Zhang, Y\. Wu, MGTCF: Multi\-generator tropical cyclone forecasting with heterogeneous meteorological data\. InProc\. AAAI Conf\. Artif\. Intell\.37, 5096–5104 \(2023\)\.
- \[30\]X\. Wang, K\. Chen, L\. Liu, T\. Han, B\. Li, L\. Bai, Global tropical cyclone intensity forecasting with multi\-modal multi\-scale causal autoregressive model\.ICASSP 2025\-2025 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\), 1–5 \(2025\)\.
- \[31\]C\. Huang and P\. Mu and J\. Zhang and S\. Chan and S\. Zhang and H\. Yan and S\. Chen and C\. Bai, Benchmark dataset and deep learning method for global tropical cyclone forecasting\.Nature Commun\.16\(1\), 5923 \(2025\)\.
- \[32\]S\. Zhang and P\. Mu and C\. Huang and J\. Zhang and C\. Bai, TC\-Diffuser: Bi\-Condition Multi\-Modal Diffusion for Tropical Cyclone Forecasting\.Proc\. AAAI Conf\. Artif\. Intell\.39, 1120\-1128 \(2025\)\.
- \[33\]C\. Bodnar, W\. P\. Bruinsma, A\. Lucic, M\. Stanley, A\. Allen, J\. Brandstetter, et al\., A foundation model for the Earth system\.Nature641, 1180–1187 \(2025\)\.
- \[34\]K\. Bi, L\. Xie, H\. Zhang, X\. Chen, X\. Gu, Q\. Tian, Accurate medium\-range global weather forecasting with 3D neural networks\.Nature619, 533–538 \(2023\)\.
- \[35\]K\. Chen, T\. Han, F\. Ling, J\. Gong, L\. Bai, X\. Wang, J\.\-J\. Luo, B\. Fei, W\. Zhang, X\. Chen, The operational medium\-range deterministic weather forecasting can be extended beyond a 10\-day lead time\.Commun\. Earth Environ\.6\(1\), 518 \(2025\)\.
- \[36\]D\. Kochkov, J\. Yuval, I\. Langmore, P\. Norgaard, J\. Smith, G\. Mooers, M\. Klöwer, J\. Lottes, S\. Rasp, P\. Düben, et al\., Neural general circulation models for weather and climate\.Nature632, 1060–1066 \(2024\)\.
- \[37\]Knapp, K\. R\., Kruk, M\. C\., Levinson, D\. H\., Diamond, H\. J\., Neumann, C\. J\. The international best track archive for climate stewardship \(ibtracs\): Unifying tropical cyclone data\.Bull\. Am\. Meteorol\. Soc\.91, 363 – 376 \(2010\)\.
- \[38\]Y\. Wang, C\.\-C\. Wu, Current understanding of tropical cyclone structure and intensity changes—a review\.Meteorol\. Atmos\. Phys\.87\(4\), 257–278 \(2004\)\.
- \[39\]C\. Wang, X\. Li, G\. Zheng, Tropical cyclone intensity forecasting using model knowledge guided deep learning model\.Environ\. Res\. Lett\.19, 024006 \(2024\)\.
- \[40\]A\. Gupta, J\. Johnson, L\. Fei\-Fei, S\. Savarese, A\. Alahi, Social GAN: Socially acceptable trajectories with generative adversarial networks\. InProc\. IEEE Conf\. Comput\. Vis\. Pattern Recognit\., 2255–2264 \(2018\)\.
- \[41\]T\. Li, M\. Lai, S\. Nie, H\. Liu, Z\. Liang, W\. Lv, Tropical cyclone trajectory based on satellite remote sensing prediction and time attention mechanism ConvLSTM model\.Big Data Res\.36, 100439 \(2024\)\.
- \[42\]C\.\-Y\. Lee, M\. K\. Tippett, A\. H\. Sobel, S\. J\. Camargo, Rapid intensification and the bimodal distribution of tropical cyclone intensity\.Nat\. Commun\.7\(1\), 10625 \(2016\)\.
- \[43\]Q\. Wu, X\. Wang, L\. Tao, Interannual and interdecadal impact of Western North Pacific Subtropical High on tropical cyclone activity\.Clim\. Dyn\.54\(3\), 2237–2248 \(2020\)\.
- \[44\]M\. L\. M\. Wong, J\. C\. L\. Chan, Tropical cyclone intensity in vertical wind shear\.J\. Atmos\. Sci\.61\(15\), 1859–1876 \(2004\)\.
- \[45\]L\. R\. Schade, Tropical cyclone intensity and sea surface temperature\.J\. Atmos\. Sci\.57\(18\), 3122–3130 \(2000\)\.
- \[46\]H\. Hersbach, B\. Bell, P\. Berrisford, G\. Biavati, A\. Horányi, J\. Muñoz Sabater, J\. Nicolas, C\. Peubey, R\. Radu, I\. Rozum, ERA5 hourly data on single levels from 1979 to present\.Copernicus Climate Change Serv\. \(C3S\) Climate Data Store \(CDS\)10, 10\.24381 \(2018\)\.
- \[47\]T\. Salzmann, B\. Ivanovic, P\. Chakravarty, M\. Pavone, Trajectron\+\+: Dynamically\-feasible trajectory forecasting with heterogeneous data\. InProc\. Eur\. Conf\. Comput\. Vis\. \(ECCV\), 683–700 \(2020\)\.
- \[48\]T\. Gu, G\. Chen, J\. Li, C\. Lin, Y\. Rao, J\. Zhou, J\. Lu, Stochastic trajectory prediction via motion indeterminacy diffusion\. InProc\. IEEE/CVF Conf\. Comput\. Vis\. Pattern Recognit\., 17113–17122 \(2022\)\.

## Acknowledgments

##### Funding:

This work is partially supported by Zhejiang Provincial Natural Science Foundation of China under Grant No\. RG25F020001 and LR21F020002, as well as the Natural Science Foundation of China under Grant No\. 62202429\.

##### Author contributions:

Conceptualization: S\.Z\., P\.M\., C\.H\., C\.B\., and S\.S\. Methodology: S\.Z\., P\.M\., S\.S\., C\.H\., H\.Y\., and Y\.Z\. Software: S\.Z\., P\.M\., Y\.Z\., C\.H\. Validation: S\.Z\., P\.M\., H\.Y\., Y\.Z\., C\.B\. and S\.S\. Formal analysis: S\.Z\., P\.M\., C\.B\., S\.S\., H\.Y\., C\.H\., J\.Zh\., and S\.C\. Investigation: S\.Z\., P\.M\., C\.H\., S\.S\., and C\.B\. Resources: C\.B\., S\.S\., S\.C\., and J\.Zh\. Data curation: S\.Z\., C\.H\., Y\.Z\., and H\.Y\. Writing—original draft: S\.Z\., P\.M\. Writing—review and editing: C\.B\., S\.S\., P\.M\., C\.H\., J\.Zh\., and S\.C\. Visualization: S\.Z\., P\.M\., and S\.S\. Supervision: C\.B\. and S\.S\. Project administration: C\.B\. and S\.S\. Funding acquisition: C\.B\.

##### Competing interests:

There are no competing interests to declare\.

##### Data and materials availability:

Our dataset consists of three main components:

1. 1\.Historical TC attributes, sourced from our previously organized dataset, available at https://zenodo\.org/records/15009527\.
2. 2\.ERA5 data, provided by ECMWF, which can be accessed from the https://cds\.climate\.copernicus\.eu/\.
3. 3\.TC tilt metrics, derived from vorticity data, which are computed using wind field data\. The processing code is available at https://github\.com/Zjut\-MultimediaPlus/Tianmu\-TC\.

In our model, these datasets are preprocessed intoPKLfiles, which can be obtained from Google Drive: https://drive\.google\.com/file/d/1XpfByEZkZHAybXgB5p2YsR5KZhHrtVei/view?usp=drive\_link and https://drive\.google\.com/file/d/1aiJaUH035YOIbsS9Q1Y9GGmyKW1HiJU1/view?usp=drive\_link\.

All the code and processed data are public on Github \( https://github\.com/Zjut\-MultimediaPlus/Tianmu\-TC\)\. This code will be updated and changed over time\.

### Supplementary materials

Supplementary Text Table S1 Figures S1 to S4 Movies S1 to S3

## Supplementary Materials for Tianmu\-TC: Physics\-constraints Generative Artificial Intelligence for Global Tropical Cyclone Forecasting

Shiqi Zhang†, Pan Mu†, Cheng Huang, Hanting Yan, Yuchao Zhu, Jinglin Zhang, Shengyong Chen∗, Shoujuan Shu∗, Cong Bai∗ ∗Corresponding author\. Email: csy@tjut\.edu\.cn, sjshu@zju\.edu\.cn, congbai@zjut\.edu\.cn †These authors contributed equally to this work\.

#### This PDF file includes:

Supplementary Text Table S1 Figures S1 to S4 Captions for Movies S1 to S3

#### Other Supplementary Materials for this manuscript:

Movies S1 to S3

### Supplementary Text

#### A\.1Diffusion Models for Multi\-Trend Generation

Diffusion models are inherently well\-suited for multi\-trend generation due to their ability to capture complex, multi\-modal data distributions\. Their probabilistic framework allows for flexible sampling, enabling the generation of diverse trends by conditioning on specific features\. Additionally, the iterative refinement mechanism in the reverse diffusion process ensures scalability and accurate representation of overlapping patterns\. Empirical successes in generative tasks further validate their compatibility with multi\-trend dynamics, making diffusion models a natural choice for such challenges\.

For fair comparison, following TCNM\[[31](https://arxiv.org/html/2608.18500#bib.bib31)\], the sampling number is 6, which means that the model generates 6 possible tendencies\. Similarly, Tianmu\-TC inputs 48 hours \(nn= 8, the time resolution is 6 hours\) of TC historical data and outputs 24 hours \(mm= 4\) of 1D future TC attribute data\.

#### A\.2NeuralGCM Tropical Cyclone Tracking

NeuralGCM provides publicly available forecast data for 2020 \(https://neuralgcm\.readthedocs\.io/en/latest/neuralgcm\_datasets\.html\)\. In this study, tropical cyclone tracking is performed according to the available forecast variables and the tracking parameters reported in Table I7 of the original NeuralGCM paper\. The original NeuralGCM setting uses sea level pressure rather than MSLP, while the released forecast variables used in this study do not include SLP\. Therefore, NeuralGCM cannot be used for intensity prediction and is only evaluated for track prediction\.

For the deterministic forecast, we use the finest\-resolution 0\.7° NeuralGCM forecast data provided officially \(gs://weatherbench2/datasets/neuralgcm\_deterministic\)\. According to Table I7 of the original NeuralGCM paper, vorticity tracking is used for tropical cyclone tracking\. The criteria are: vorticity at 850 hPa change of±0\.00006​s−1\\pm 0\.00006~\\mathrm\{s\}^\{\-1\}over 5\.5 GCD, together with a difference between geopotential surfaces at 300 and 500 hPa that decreases by at least−25\.8​m2​s−2\-25\.8~\\mathrm\{m\}^\{2\}\\mathrm\{s\}^\{\-2\}over 6\.5 GCD\. Here, GCD denotes great\-circle distance\. For the ensemble forecast \(gs://weatherbench2/datasets/neuralgcm\_ens\), due to the high computational cost, the original NeuralGCM paper reduced the highest resolution to 1\.4° for ensemble forecasts\. The tracking method is consistent with that used for the deterministic forecast, and the ensemble size is 50 members\.

#### A\.3GenCast Tropical Cyclone Tracking and Intensity Extraction

This section describes how tropical cyclone centers and intensity variables were extracted from the public GenCast forecast archive \(gs://weatherbench2/datasets/gencast\)\. We first summarize the original TempestExtremes tracker used in GenCast, then describe the modifications required by the limited pressure\-level variables available in WeatherBench2, followed by the candidate review procedure and the final intensity extraction strategy\.

The original GenCast used TempestExtremes v2\.1 to track tropical cyclones\. The tracker consists of two main stages\. In the first stage, DetectNodes identifies candidate cyclone centers at each time step\. Candidate locations are first determined from local minima in mean sea level pressure \(MSL\), and each low\-pressure center is required to be associated with a sufficiently compact closed low\-pressure structure\. In addition, the original tracker requires the candidate storm to be colocated with an upper\-level warm\-core structure in the 500–300 hPa geopotential thickness field\. In the second stage, StitchNodes links candidate nodes across time into physically plausible cyclone trajectories, and further filters non\-tropical\-cyclone trajectories using criteria related to the maximum displacement between adjacent time steps, the maximum temporal gap, the minimum track duration, wind speed, surface elevation, and latitude\. Because GenCast forecasts are evaluated at 12 h intervals, the original study adjusted the StitchNodes connection range from the default 8° to 12° in great\-circle distance\. This larger range allows candidate centers at adjacent forecast times to be linked even when a tropical cyclone travels farther within a 12 h interval\.

In this study, the public WeatherBench2 GenCast archive provides pressure\-level variables only at 500, 700, and 850 hPa\. The 300 hPa geopotential height required to reproduce the original 300–500 hPa warm\-core criterion is therefore unavailable\. As a result, the original TempestExtremes tracking protocol used in GenCast cannot be fully reproduced\. To remain as consistent as possible with the original tracking rationale under the constraints of the public archive, we retained the core DetectNodes criteria based on local MSLP minima and closed low\-pressure structures\. Following the 12° displacement scale used for 12 h tracking in GenCast, we implemented a 12° local search window for each lead time\. The search is anchored at the best\-track center for the first lead time and, in the stepwise procedure, at the previously predicted center for the subsequent lead time\. It should be emphasized that the goal of this study is not to redetect the genesis, termination, or complete life cycle of all global cyclones\. Instead, for each given best\-track sample, we diagnose the \+12 h and \+24 h tropical cyclone centers and intensities from the original GenCast gridded forecast fields\. Therefore, the full global StitchNodes trajectory construction, minimum\-duration filtering, and elevation and latitude filtering are not applied\.

Specifically, for each initialization time and forecast lead time, MSLP, 850 hPa horizontal winds, and 10 m winds are first extracted within a 12° latitude–longitude window around the tracking anchor\. The tracking anchor is initialized as the best\-track center at the forecast initialization time\. In the stepwise tracking procedure, the search anchor for the later lead time is updated to the predicted center from the previous lead time\. Candidate centers are generated from local minima in the MSLP field, sorted by their MSLP values in ascending order, and at most 500 candidates are examined\. Each candidate is then tested for the presence of a sufficiently compact closed low\-pressure structure\.

Because the 300 hPa warm\-core criterion cannot be reproduced, 850 hPa relative vorticity is further used as an auxiliary constraint on the low\-level cyclonic structure\. The 850 hPa relative vorticity is computed from the 850 hPa horizontal wind field\. For candidates in the Northern Hemisphere, cyclonic vorticity corresponds to a positive vorticity core, so the maximum relative vorticity within a 1° radius of the candidate center is used\. For candidates in the Southern Hemisphere, cyclonic vorticity corresponds to a negative vorticity core, so the minimum relative vorticity within a 1° radius of the candidate center is used\. This criterion does not require the vorticity center and the MSLP minimum to fall on exactly the same grid point; rather, it requires a consistent low\-level cyclonic structure near the candidate low\-pressure center\.

The final center is selected using a hierarchical priority strategy\. Among the four lowest local MSLP minima, candidates that satisfy both the SLP closed\-contour criterion and the 850 hPa vorticity\-structure criterion, and that are close to the tracking anchor, are given the highest priority\. All candidates are retained, and whether each candidate passes the SLP and SLP\+VORT criteria is recorded for subsequent quality diagnosis\.

Because this tracker cannot fully reproduce the TempestExtremes stitching procedure used in the original GenCast, we further performed candidate\-level quality review of the automatic tracking results\. For every 2020 test sample and lead time, we saved the candidate table, MSLP field, 850 hPa vorticity field, 10 m wind field, and the corresponding visualizations\. The review was based on three criteria\. First, the selected center should be a local MSLP minimum located within the innermost closed low\-pressure structure, rather than a grid point on the outer edge of the closed low or along a low\-pressure trough\. Second, the candidate should be associated with a consistent cyclonic vorticity core in the 850 hPa relative vorticity field, namely a local positive\-vorticity maximum in the Northern Hemisphere and a local negative\-vorticity minimum in the Southern Hemisphere\. Third, the predicted centers at \+12 h and \+24 h should be temporally continuous, avoiding jumps to unrelated low\-pressure systems between adjacent lead times\. For the small number of samples with clear jumps in the automatic tracking results, corrections were restricted to selecting, from the candidate table saved during automatic tracking, a candidate center satisfying the above physical\-structure and temporal\-continuity criteria\. No new center coordinates were manually specified outside the saved candidate set\. To ensure traceability, we saved and provided the predicted tracks, track errors, candidate tables, and visualizations of MSLP, 850 hPa relative vorticity, and 10 m wind speed for all 2020 test samples\. Overall, the reviewed centers are consistent with the closed MSLP structure, the low\-level cyclonic vorticity structure, and the temporal continuity between adjacent lead times\.

After the diagnostic center is determined, GenCast intensity diagnostics are further extracted\. The intensity task in this study includes two variables: pressure and wind speed\. Pressure refers to the lowest atmospheric pressure at the TC center, in hPa, and wind speed refers to the two\-minute maximum sustained wind speed near the TC center, in m/s\. For the GenCast gridded forecast fields, we extract the minimum MSLP within a 1° radius of the diagnostic center as the minimum central pressure diagnostic, and the maximum 10 m wind speed within a 3° radius as the near\-center wind speed diagnostic\. It should be noted that the public GenCast archive provides instantaneous gridded wind fields at discrete forecast times, rather than two\-minute maximum sustained wind observations\. Therefore, the GenCast wind speed error reported in this study should be interpreted as the diagnostic error of the near\-center maximum wind speed derived from the instantaneous 10 m wind field, rather than a strict error of the two\-minute maximum sustained wind speed\.

Table S1:Extended\-Sample Comparison on special TC scenarios on global area from 2017 to 2021\.Samples is the number of cases\. Lower values indicate better performance\. “Anomaly” consists of circular movement, turning on the ocean, and V\-shaped tracks\. “RI” denotesRapid Intensification, and “RW” denotesRapid Weakening\.CategoryScenario \(Unit\)MethodsSamplesForecast Horizon\+6h\+12h\+18h\+24hTrackAnomaly \(km\)Pangu110356\.2562\.9677\.58110\.94Fengwu110344\.1949\.3365\.8995\.62Ours110318\.0817\.7236\.9373\.57Landfall \(km\)Pangu58259\.8476\.14101\.26131\.16Fengwu58252\.3362\.0476\.1493\.80Ours58217\.5919\.0638\.8671\.78IntensityRI \(m/s\)Pangu33020\.7929\.4139\.1747\.02Fengwu33019\.9227\.7136\.8243\.99Ours3301\.371\.314\.638\.04RW \(m/s\)Pangu46428\.1721\.5315\.5011\.37Fengwu46427\.4220\.7614\.4810\.08Ours4642\.701\.703\.895\.78![Refer to caption](https://arxiv.org/html/2608.18500v1/framework/fig3-2.png)

Figure S1:Central Information Enhancement \(CenIE\) block\.Considering that the center of selected ERA5 data corresponds to the location of the TC center, specifically, the longitude and latitude coordinates in the track task\. Besides, pressure represents the lowest pressure near the TC center, and wind speed denotes the two\-minute maximum sustained wind speed near the TC center\. In other words, the predicted attributes are all the attributes in and around the TC center\. Therefore, we define TC center as a highly task\-relevant area\. To enhance task\-relevant information in the selected ERA5 data, CenIE block enhances the center pointEa,bt\+i′E\_\{a,b\}^\{t\+i^\{\\prime\}\}of original 2D dataEt\+i′E\_\{t\+i\}^\{\\prime\}andEt\+i=c​o​n​c​a​t​\(Et\+i′,Et\+i′−Ea,bt\+i′\)E\_\{t\+i\}=concat\(E\_\{t\+i\}^\{\\prime\},E\_\{t\+i\}^\{\\prime\}\-E\_\{a,b\}^\{t\+i^\{\\prime\}\}\), whereCCin the figure denotes channel concatenation,i∈\[1,n\]i\\in\[1,n\], theaaandbbdenote the coordinates in the geopotential height map,HHandWWdenote the width and height of the map, andnndenotes the length of the historical data\. That is, the value of each pixel point inEt\+i′E\_\{t\+i\}^\{\\prime\}is subtracted from that of the current central pointEa,bt\+i′E\_\{a,b\}^\{t\+i^\{\\prime\}\}, and concatenated with the originalEt\+i′E\_\{t\+i\}^\{\\prime\}\. Finally, CenIE block outputs the enhancement of central informationEt\+iE\_\{t\+i\}\. This data preprocessing method is very simple and does not need GPU resources\.![Refer to caption](https://arxiv.org/html/2608.18500v1/framework/supple-compare.jpg)

Figure S2:Comparison with additional operational models in the EP and NA basins\.We include authoritative forecasting systems issued by the National Hurricane Center \(NHC\): the official forecast \(OFCL\), and the statistical\-dynamical model combination OCD5, which merges CLP5 \(track\) and DSHP \(intensity\)\. Other models include EMXI \(ECMWF global model from the previous cycle\), HCCA \(a weighted consensus of multiple dynamical and statistical models\), as well as recent large models Pangu and Fengwu\. Testing period: from 2017 to 2023\.![Refer to caption](https://arxiv.org/html/2608.18500v1/framework/global_wp.jpg)

Figure S3:More comparison methods on WP\.![Refer to caption](https://arxiv.org/html/2608.18500v1/framework/c1c2c3_uqscore.jpg)

Figure S4:Quantitative results of uncertainty for\(a\) track, \(b\) pressure and \(c\) wind speed\.##### Caption for Movie S1\.

The reduction of prediction uncertainty in track forecasting\.Animated version of Figure[3](https://arxiv.org/html/2608.18500#Sx1.F3)\(a\)\. The red solid line represents the ground truth, the cyan dashed lines represent the predictions of 6 tendencies\. Under the guidance of proposed physics\-constrained generative AI, six tendencies gradually converge into a region with low uncertainty\.

##### Caption for Movie S2\.

The reduction of prediction uncertainty in pressure forecasting\.Animated version of Figure[3](https://arxiv.org/html/2608.18500#Sx1.F3)\(b\)\. The blue solid line represents the ground truth, the cyan dashed lines represent the predictions of 6 tendencies\.

##### Caption for Movie S3\.

The reduction of prediction uncertainty in wind speed forecasting\.Animated version of Figure[3](https://arxiv.org/html/2608.18500#Sx1.F3)\(c\)\. The yellow solid line represents the ground truth, the cyan dashed lines represent the predictions of 6 tendencies\.

Similar Articles

TC-Next: Zero-Shot Multimodal Cyclone Forecasting

arXiv cs.LG

TC-Next is a multimodal deep learning model that leverages foundation model forecasts and satellite imagery for zero-shot tropical cyclone track and intensity forecasting, showing significant error reduction over conventional trackers.

How we're supporting better tropical cyclone prediction with AI

Google DeepMind Blog

Google DeepMind and Google Research launched Weather Lab, an interactive platform featuring experimental AI-based tropical cyclone prediction models that can forecast cyclone formation, track, intensity, and shape up to 15 days ahead. The models are being validated in partnership with the U.S. National Hurricane Center to improve forecast accuracy and support real-time warnings.

Real-time probabilistic tsunami forecasting via generative AI

arXiv cs.LG

This paper introduces a probabilistic ensemble model based on a conditional diffusion model for real-time tsunami inundation forecasting, offering uncertainty quantification in contrast to deterministic warnings. Validated with 2011 Tohoku-oki data, it demonstrates that generative AI can shift tsunami forecasting from deterministic to probabilistic approaches.