WildfireSpreadBench:指标在野火蔓延预测中决定模型
摘要
本文介绍了WildfireSpreadBench,用于基准测试野火蔓延预测模型,揭示了评估指标如AP与F1可能导致不同的模型排名,并影响操作适用性。
arXiv:2609.22191v1 Announce Type: new
Abstract: Machine learning is being increasingly used to predict where active wildfires will burn the following day, helping inform evacuation boundaries and containment lines. Most models are evaluated using Average Precision (AP), which summarizes performance across all decision thresholds, although acting on a forecast requires choosing one. We benchmarked five discriminative architectures and one generative model on WildfireSpreadTS using a shared evaluation pipeline and two input configurations. We found that model rankings varied depending on whether performance was measured by AP or by threshold-dependent metrics like F1 and IoU. The highest-AP model flagged 4 to 5 times the area that burned and ranked fifth of six on F1 and IoU, and the most recall-heavy model flagged 16 to 23 times. Models with more usable predictions had AP scores 24 to 37 lower. Across architectures, we identified three distinct prediction profiles: over-predicting, balanced, and under-predicting, which AP alone could not distinguish. Expanding the input from 7 to 23 channels changed AP by 0.03 on average, against a 0.21 to 0.24 spread across architectures. These results show AP alone can favor models whose predictions are poorly suited for operational wildfire forecasting.
查看缓存全文
缓存时间: 2026/09/22 09:17
# The Metric Decides the Modelin Wildfire Spread Prediction
Source: [https://arxiv.org/html/2609.22191](https://arxiv.org/html/2609.22191)
## WildfireSpreadBench: The Metric Decides the Model in Wildfire Spread Prediction
###### Abstract
Machine learning is being increasingly used to predict where active wildfires will burn the following day, helping inform evacuation boundaries and containment lines\. Most models are evaluated using Average Precision \(AP\), which summarizes performance across all decision thresholds, although acting on a forecast requires choosing one\. We benchmarked five discriminative architectures and one generative model on WildfireSpreadTS using a shared evaluation pipeline and two input configurations\. We found that model rankings varied depending on whether performance was measured by AP or by threshold\-dependent metrics like F1 and IoU\. The highest\-AP model flagged 4 to 5 times the area that burned and ranked fifth of six on F1 and IoU, and the most recall\-heavy model flagged 16 to 23 times\. Models with more usable predictions had AP scores 24 to 37% lower\. Across architectures, we identified three distinct prediction profiles: over\-predicting, balanced, and under\-predicting, which AP alone could not distinguish\. Expanding the input from 7 to 23 channels changed AP by 0\.03 on average, against a 0\.21 to 0\.24 spread across architectures\. These results show AP alone can favor models whose predictions are poorly suited for operational wildfire forecasting\.
## 1Introduction
Fire seasons in the western United States have grown longer and more severe, largely due to anthropogenic warming\([Abatzoglou and Williams, 2016](https://arxiv.org/html/2609.22191#bib.bib4)\), while smoke exposure has increased alongside them\([Burke et al\., 2021](https://arxiv.org/html/2609.22191#bib.bib5)\)\. Where and how a fire burns today shapes where it burns tomorrow, and incident teams have to anticipate that to advise on evacuation and containment lines\. Progress in technology has made this job easier; physics\-based rate\-of\-spread models\([Sullivan, 2009](https://arxiv.org/html/2609.22191#bib.bib3)\)now sit alongside models trained on satellite data\([Huot et al\., 2022](https://arxiv.org/html/2609.22191#bib.bib2)\), and WildfireSpreadTS\([Gerard et al\., 2023b](https://arxiv.org/html/2609.22191#bib.bib1)\)gave the field a shared benchmark\. Convolutional, recurrent and attention\-based architectures have been benchmarked on it\([Lahrichi et al\., 2026](https://arxiv.org/html/2609.22191#bib.bib17)\), and generative ones on simulated fire data\([Yu et al\., 2026](https://arxiv.org/html/2609.22191#bib.bib6)\)\. The majority are scored with AP, which is defensible since it needs no threshold, and precision–recall beats ROC\-AUC when the positive class is rare\([Davis and Goadrich, 2006](https://arxiv.org/html/2609.22191#bib.bib13);[Saito and Rehmsmeier, 2015](https://arxiv.org/html/2609.22191#bib.bib14)\)\. However, deployment is different since fire teams draw outlines on a map, which means picking a threshold\. If AP and threshold metrics disagree on model performance, evaluating solely on AP may favor the wrong one for field use\. We tested six architectures to find:
- •A unified benchmarkof all six architectures under one metric implementation\.
- •Evidence that metrics outrank architecture, with AP and F1 picking different winners\.
- •A reporting standardpairing AP with F1, IoU, and precision–recall at a unified threshold\.
## 2Related Work
Datasets and multi\-temporal learning\.The advancement and rigorous evaluation of spatio\-temporal architectures depend entirely on the availability of standardized, continuous data\. Early foundational benchmarks, such as the Next Day Wildfire Spread dataset, catalyzed initial deep learning research by providing high\-resolution static snapshots of historical fires alongside environmental variables\([Huot et al\., 2022](https://arxiv.org/html/2609.22191#bib.bib2)\)\. While useful, such static transitions miss the continuous, sequential evolution of a spreading fire\. The WildfireSpreadTS dataset addresses this critical limitation by providing contiguous, 24\-hour multi\-modal time\-series observations\([Gerard et al\., 2023b](https://arxiv.org/html/2609.22191#bib.bib1)\)\. Rather than a cross\-sectional snapshot of the fire, WildfireSpreadTS captures 23 input data channels, including meteorological factors, land coverage, and ground truth active fire area across multi\-day windows\. This lets researchers evaluate architectures on their capacity to learn complex temporal dynamics and sequential physical interactions rather than pattern matching alone\([Gerard et al\., 2023b](https://arxiv.org/html/2609.22191#bib.bib1)\)\.
Standardized evaluation metrics for imbalanced domains\.A major hurdle in wildfire spread prediction training and evaluation is extreme class imbalance\. Because wildfire datasets consist of very few “active fire” pixels in comparison to “no fire” pixels, designing a loss function and evaluation metric that captures the accuracy of predicting fire spread can be challenging\([Andrianarivony and Akhloufi, 2024](https://arxiv.org/html/2609.22191#bib.bib15)\)\. A model could achieve over 90% accuracy simply by predicting that no fire will occur anywhere\. Consequently, robust benchmarking requires specialized metrics designed for imbalanced semantic segmentation\. Intersection over Union \(IoU\) and the F1\-score measure spatial overlap directly, heavily penalizing models that over\-predict the fire class\. To further analyze a model’s operational performance, Precision \(the proportion of predicted fire pixels that actually burned\) and Recall \(the proportion of actual fire pixels successfully predicted by the model\) reported independently are extremely useful\. For instance, in the context of emergency management, maximizing Recall is often prioritized to prevent the catastrophic under\-prediction of a fire’s leading edge\([Rösch et al\., 2024](https://arxiv.org/html/2609.22191#bib.bib19)\)\. Finally, Average Precision summarizes the precision–recall curve across all operational thresholds\([Gerard et al\., 2023b](https://arxiv.org/html/2609.22191#bib.bib1)\)\.
The gap\.While WildfireSpreadTS provides the temporal data, and Flow Matching enables efficient probabilistic forecasting, the literature lacks a direct comparison between optimized deterministic and generative architectures\. This study bridges that gap, evaluating established baselines alongside a custom BCE U\-Net and an adapted Flow Matching model under one set of metrics\.
## 3Methods
Data\.WildfireSpreadTS\([Gerard et al\., 2023b](https://arxiv.org/html/2609.22191#bib.bib1)\)comprises 13,607 daily images across 607 U\.S\. fire events from January 2018 to October 2021, at 375 m resolution over the 23 channels described in Section[2](https://arxiv.org/html/2609.22191#S2)\. We write dayttfor the most recent observed day of a fire and dayt\+1t\{\+\}1for the day after, which is the day whose active fire mask every model predicts\. ConvLSTM and UTAE receive dayst−4t\{\-\}4throughtt, whereas the other four receive dayttalone\. All models are trained on 2018 and 2019 and tested on the held\-out 2021 season; the four dataset baselines also validate on 2020\. To separate the effects of every model architecture and input data, we evaluate two configurations\. Vegetation uses seven channels of reflectance, vegetation indices and active\-fire data, while All uses all 23 channels\.
Models\.Five are discriminative, meaning they focus on predicting labels\. Logistic Regression, a pixel\-wise linear baseline; ResNet18 U\-Net, the encoder–decoder released with the dataset\([Gerard et al\., 2023b](https://arxiv.org/html/2609.22191#bib.bib1);[Ronneberger et al\., 2015](https://arxiv.org/html/2609.22191#bib.bib11);[He et al\., 2016](https://arxiv.org/html/2609.22191#bib.bib12)\); ConvLSTM\([Shi et al\., 2015](https://arxiv.org/html/2609.22191#bib.bib9)\), which models day\-to\-day dependence recurrently; UTAE\([Garnot and Landrieu, 2021](https://arxiv.org/html/2609.22191#bib.bib10)\), a U\-Net with a temporal attention encoder for satellite time series; and our own BCE U\-Net, a segmentation U\-Net trained with a positive\-weighted binary cross\-entropy objective\. Flow Matching\([Lipman et al\., 2023](https://arxiv.org/html/2609.22191#bib.bib7);[Liu et al\., 2023](https://arxiv.org/html/2609.22191#bib.bib8)\)learns a continuous\-time velocity field carrying noise to the data distribution and integrates it as an ODE at inference, so its prediction is sampled rather than thresholded from one forward pass\. We include this modeling approach because fire spread is stochastic, and the authors of the dataset expected label\-noise\-tolerant methods to aid in this regard\([Gerard et al\., 2023b](https://arxiv.org/html/2609.22191#bib.bib1)\)\.
## 4Results
Table[1](https://arxiv.org/html/2609.22191#S4.T1)reports all twelve runs, Figure[1](https://arxiv.org/html/2609.22191#S4.F1)plots them two ways, and Figure[2](https://arxiv.org/html/2609.22191#S4.F2)shows one 2021 test fire\.
Table 1:The models predict the wildfire spread 24h ahead based on two input feature sets:*Vegetation*and*All*\. The performance displayed is the test set AP, F1, IoU, Precision, and Recall for the held\-out 2021 year, with threshold metrics computed at 0\.5\. Rows are grouped by operating\-point profile, and bold marks the best\.Figure 1:\(a\)Model rank under AP against rank under F1, for both feature sets; bolded lines mark ranks that move three places or more\.\(b\)The same twelve runs as precision–recall operating points, with iso\-F1 contours\.Figure 2:Active fire on daytt, the observed fire on dayt\+1t\{\+\}1, and the ConvLSTM and BCE U\-Net predictions at a 0\.9 threshold, with the observed footprint outlined on both prediction panels\. BCE U\-Net covers 2\.1×\\timesthe observed area and merges the observed patches into roughly a dozen blobs; ConvLSTM covers 0\.9×\\times\.
## 5Discussion and Climate Impact
Metric analysis\.AP is the standard metric in machine learning for class\-imbalanced datasets\([Davis and Goadrich, 2006](https://arxiv.org/html/2609.22191#bib.bib13);[Saito and Rehmsmeier, 2015](https://arxiv.org/html/2609.22191#bib.bib14)\)\. However, it may be over\-relied upon in the context of wildfire spread\. This is demonstrated most strongly in this study by the BCE U\-Net model, which earned the highest AP score of all models across both feature sets\. Despite its high AP, it may be a poor tool for practical fire prediction because it over\-predicts, as Figure[2](https://arxiv.org/html/2609.22191#S4.F2)shows, and because AP does not describe behavior at any one operating point\([Maier\-Hein et al\., 2024](https://arxiv.org/html/2609.22191#bib.bib16)\)\.[Rösch et al\. \(2024\)](https://arxiv.org/html/2609.22191#bib.bib19)argue over\-prediction is the more tolerable failure, since false positives can be revised\. At 16 to 23 times the observed area, that tolerance runs out\. We argue for a holistic approach to model benchmarking in similar domains that require standards for prediction outcomes\.
Dataset and architecture\.Wildfire training datasets have been a focus of the field for the past half\-decade, from Next Day Wildfire Spread\([Huot et al\., 2022](https://arxiv.org/html/2609.22191#bib.bib2)\)to WildfireSpreadTS and its expansion, WildfireSpreadTS\+\. A key finding affirmed by this study is that higher\-complexity datasets do not appear to be the immediate answer to improving machine learning applications in wildfire spread prediction\. Where[Gerard et al\. \(2023b\)](https://arxiv.org/html/2609.22191#bib.bib1)found that the full channel set reduced AP for their temporal models, we found the seven vegetation channels gave metrics very similar to the full set, while remaining more computationally efficient\. This, combined with the WildfireSpreadTS\+ finding that four further years of training data did not improve accuracy\([Lahrichi et al\., 2026](https://arxiv.org/html/2609.22191#bib.bib17)\), suggests that raw data volume is not the primary bottleneck in wildfire machine learning\. In stark contrast, altering the model architecture leads to massive changes in output characteristics\. Specifically, a model’s mathematical framework for navigating class imbalance ultimately dictates its viability in this field, with different solutions offering drastically different results\([Lin et al\., 2017](https://arxiv.org/html/2609.22191#bib.bib18)\)\. For instance, this study’s data demonstrates that an emergency response team utilizing a balanced architecture like Convolutional LSTM rather than a recall\-maximizing UTAE architecture would see an increase in precision by a factor of nine or more at standard thresholds \(0\.5\)\. This degree of variability is not present when tailoring feature sets or gathering more input variables, indicating that the immediate future direction should prioritize model architecture over the aggregation of more data\.
Climate impact\.As anthropogenic warming intensifies fire seasons\([Abatzoglou and Williams, 2016](https://arxiv.org/html/2609.22191#bib.bib4);[Burke et al\., 2021](https://arxiv.org/html/2609.22191#bib.bib5)\), effective climate adaptation requires deployable machine learning\. However, the highest\-AP model here flags four to five times the area that burned\. Unnecessary evacuation is costly and disruptive, so a metric that does not surface it is a poor guide to field use\. That said, to build genuine climate resilience, the machine learning community must align evaluation with operational reality\. We urge future researchers to prioritize a more holistic and humanistic approach to model evaluation, focusing on the practical impacts of their optimizations rather than traditional standardized machine learning metrics\.
Limitations\.WildfireSpreadTS covers United States wildfires from 2018 to 2021 with a single held\-out test year, so the results here describe relative model behavior on this benchmark rather than validated performance in other geographies or fire regimes\. Secondly, active fire labels derived from satellite detection carry noise from false positives and missed detections\. That noise affects every model evaluated here, and it is part of why the comparison includes a generative approach at all\.
## 6Conclusion
We present WildfireSpreadBench, a unified comparison of generative and discriminative architectures for next\-day wildfire spread prediction under a shared evaluation protocol and a metric set broader than Average Precision alone\. Discriminative models lead on every metric in both feature configurations, with BCE U\-Net taking AP, while ConvLSTM and ResNet18 U\-Net take F1, IoU, and precision\. The generative model never leads, and ResNet18 U\-Net beats it on all five metrics in both configurations, so the label\-noise robustness that motivated its inclusion does not appear in these results\. Flow Matching is still the only architecture that behaves conservatively, covering 0\.8 times the observed burned area, where BCE U\-Net covers 4 to 5 times\. This contrast highlights an important difference between overall predictive performance and the practical behavior of the resulting fire maps\. Which metric is reported, therefore, matters more than the architecture family, and far more than the feature set\. These results show that model selection can change substantially depending on whether evaluation emphasizes ranking quality or thresholded spatial predictions\. Future work should widen the evaluation rather than the input\. This calls for multiple held\-out years on varied fire data, while newer generative formulations should receive the same treatment, and the prioritized metrics should be validated against firefighters’ judgment in real\-world working settings\. Code and configurations are available at[https://anonymous\.4open\.science/r/OfficialWildfireSpreadBench](https://anonymous.4open.science/r/OfficialWildfireSpreadBench), covering all six models, the shared data pipeline, and the evaluation routine behind every number in Table[1](https://arxiv.org/html/2609.22191#S4.T1)\.
## References
- Abatzoglou and Williams \(2016\)J\. T\. Abatzoglou and A\. P\. WilliamsImpact of anthropogenic climate change on wildfire across western US forests\.Proceedings of the National Academy of Sciences113\(42\),pp\. 11770–11775\.Cited by:[§1](https://arxiv.org/html/2609.22191#S1.p1.1),[§5](https://arxiv.org/html/2609.22191#S5.p3.1)\.
- Andrianarivony and Akhloufi \(2024\)H\. S\. Andrianarivony and M\. A\. AkhloufiMachine learning and deep learning for wildfire spread prediction: a review\.Fire7\(12\),pp\. 482\.External Links:[Document](https://dx.doi.org/10.3390/fire7120482)Cited by:[§2](https://arxiv.org/html/2609.22191#S2.p2.1)\.
- Bogenspergeret al\.\(2025\)L\. Bogensperger, D\. Narnhofer, A\. Falk, K\. Schindler, and T\. PockFlowSDF: flow matching for medical image segmentation using distance transforms\.International Journal of Computer Vision133\(7\),pp\. 4864–4876\.External Links:[Document](https://dx.doi.org/10.1007/s11263-025-02373-y)Cited by:[Appendix B](https://arxiv.org/html/2609.22191#A2.p7.1)\.
- Burkeet al\.\(2021\)M\. Burke, A\. Driscoll, S\. Heft\-Neal, J\. Xue, J\. Burney, and M\. WaraThe changing risk and burden of wildfire in the United States\.Proceedings of the National Academy of Sciences118\(2\),pp\. e2011048118\.Cited by:[§1](https://arxiv.org/html/2609.22191#S1.p1.1),[§5](https://arxiv.org/html/2609.22191#S5.p3.1)\.
- Davis and Goadrich \(2006\)J\. Davis and M\. GoadrichThe relationship between Precision\-Recall and ROC curves\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§1](https://arxiv.org/html/2609.22191#S1.p1.1),[§5](https://arxiv.org/html/2609.22191#S5.p1.1)\.
- Garnot and Landrieu \(2021\)V\. S\. F\. Garnot and L\. LandrieuPanoptic segmentation of satellite image time series with convolutional temporal attention networks\.InIEEE/CVF International Conference on Computer Vision \(ICCV\),Cited by:[Appendix B](https://arxiv.org/html/2609.22191#A2.p3.1),[§3](https://arxiv.org/html/2609.22191#S3.p2.1)\.
- Gerardet al\.\(2023a\)S\. Gerard, Y\. Zhao, and J\. SullivanWildfireSpreadTS reference implementation\.Note:Repository updates on the corrected dataset class and the angular\-feature transformExternal Links:[Link](https://github.com/SebastianGer/WildfireSpreadTS)Cited by:[Appendix A](https://arxiv.org/html/2609.22191#A1.p12.1),[Appendix B](https://arxiv.org/html/2609.22191#A2.p4.1),[Appendix C](https://arxiv.org/html/2609.22191#A3.p8.1)\.
- Gerardet al\.\(2023b\)S\. Gerard, Y\. Zhao, and J\. SullivanWildfireSpreadTS: a dataset of multi\-modal time series for wildfire spread prediction\.InAdvances in Neural Information Processing Systems \(NeurIPS\) Datasets and Benchmarks Track,External Links:[Link](https://openreview.net/forum?id=RgdGkPRQ03)Cited by:[Table 2](https://arxiv.org/html/2609.22191#A1.T2),[Appendix A](https://arxiv.org/html/2609.22191#A1.p1.1),[Appendix A](https://arxiv.org/html/2609.22191#A1.p11.1),[Appendix A](https://arxiv.org/html/2609.22191#A1.p3.1),[Appendix A](https://arxiv.org/html/2609.22191#A1.p4.1),[Appendix A](https://arxiv.org/html/2609.22191#A1.p6.1),[Appendix A](https://arxiv.org/html/2609.22191#A1.p7.1),[Appendix B](https://arxiv.org/html/2609.22191#A2.p11.1),[Appendix B](https://arxiv.org/html/2609.22191#A2.p2.1),[Appendix B](https://arxiv.org/html/2609.22191#A2.p3.1),[Table 5](https://arxiv.org/html/2609.22191#A3.T5),[Appendix C](https://arxiv.org/html/2609.22191#A3.p7.1),[Appendix C](https://arxiv.org/html/2609.22191#A3.p8.1),[§1](https://arxiv.org/html/2609.22191#S1.p1.1),[§2](https://arxiv.org/html/2609.22191#S2.p1.1),[§2](https://arxiv.org/html/2609.22191#S2.p2.1),[§3](https://arxiv.org/html/2609.22191#S3.p1.1),[§3](https://arxiv.org/html/2609.22191#S3.p2.1),[§5](https://arxiv.org/html/2609.22191#S5.p2.1)\.
- Heet al\.\(2016\)K\. He, X\. Zhang, S\. Ren, and J\. SunDeep residual learning for image recognition\.InIEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),Cited by:[Appendix B](https://arxiv.org/html/2609.22191#A2.p3.1),[§3](https://arxiv.org/html/2609.22191#S3.p2.1)\.
- Huotet al\.\(2022\)F\. Huot, R\. L\. Hu, N\. Goyal, T\. Sankar, M\. Ihme, and Y\. ChenNext day wildfire spread: a machine learning dataset to predict wildfire spreading from remote\-sensing data\.IEEE Transactions on Geoscience and Remote Sensing60,pp\. 1–13\.Cited by:[§1](https://arxiv.org/html/2609.22191#S1.p1.1),[§2](https://arxiv.org/html/2609.22191#S2.p1.1),[§5](https://arxiv.org/html/2609.22191#S5.p2.1)\.
- Lahrichiet al\.\(2026\)S\. Lahrichi, J\. Bova, J\. Johnson, and J\. MalofImproved wildfire spread prediction with time\-series data and the WSTS\+ benchmark\.InIEEE/CVF Winter Conference on Applications of Computer Vision \(WACV\),Cited by:[§1](https://arxiv.org/html/2609.22191#S1.p1.1),[§5](https://arxiv.org/html/2609.22191#S5.p2.1)\.
- Linet al\.\(2017\)T\. Lin, P\. Goyal, R\. Girshick, K\. He, and P\. DollárFocal Loss for Dense Object Detection\.InIEEE International Conference on Computer Vision \(ICCV\),pp\. 2980–2988\.Cited by:[§5](https://arxiv.org/html/2609.22191#S5.p2.1)\.
- Lipmanet al\.\(2023\)Y\. Lipman, R\. T\. Q\. Chen, H\. Ben\-Hamu, M\. Nickel, and M\. LeFlow matching for generative modeling\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[Appendix B](https://arxiv.org/html/2609.22191#A2.p6.1),[§3](https://arxiv.org/html/2609.22191#S3.p2.1)\.
- Liuet al\.\(2023\)X\. Liu, C\. Gong, and Q\. LiuFlow straight and fast: learning to generate and transfer data with rectified flow\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[Appendix B](https://arxiv.org/html/2609.22191#A2.p6.1),[§3](https://arxiv.org/html/2609.22191#S3.p2.1)\.
- Maier\-Heinet al\.\(2024\)L\. Maier\-Hein, A\. Reinke, P\. Godau, M\. D\. Tizabi,et al\.Metrics reloaded: recommendations for image analysis validation\.Nature Methods21\(2\),pp\. 195–212\.External Links:[Document](https://dx.doi.org/10.1038/s41592-023-02151-z)Cited by:[§5](https://arxiv.org/html/2609.22191#S5.p1.1)\.
- Ronnebergeret al\.\(2015\)O\. Ronneberger, P\. Fischer, and T\. BroxU\-Net: convolutional networks for biomedical image segmentation\.InMedical Image Computing and Computer\-Assisted Intervention \(MICCAI\),Cited by:[Appendix B](https://arxiv.org/html/2609.22191#A2.p3.1),[§3](https://arxiv.org/html/2609.22191#S3.p2.1)\.
- Röschet al\.\(2024\)M\. Rösch, M\. Nolde, T\. Ullmann, and T\. RiedlingerData\-driven wildfire spread modeling of european wildfires using a spatiotemporal graph neural network\.Fire7\(6\),pp\. 207\.External Links:[Document](https://dx.doi.org/10.3390/fire7060207)Cited by:[§2](https://arxiv.org/html/2609.22191#S2.p2.1),[§5](https://arxiv.org/html/2609.22191#S5.p1.1)\.
- Saito and Rehmsmeier \(2015\)T\. Saito and M\. RehmsmeierThe precision\-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets\.PLoS ONE10\(3\),pp\. e0118432\.Cited by:[§1](https://arxiv.org/html/2609.22191#S1.p1.1),[§5](https://arxiv.org/html/2609.22191#S5.p1.1)\.
- Shiet al\.\(2015\)X\. Shi, Z\. Chen, H\. Wang, D\. Yeung, W\. Wong, and W\. WooConvolutional LSTM network: a machine learning approach for precipitation nowcasting\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[Appendix B](https://arxiv.org/html/2609.22191#A2.p3.1),[§3](https://arxiv.org/html/2609.22191#S3.p2.1)\.
- Sullivan \(2009\)A\. L\. SullivanWildland surface fire spread modelling, 1990–2007\. 1: physical and quasi\-physical models\.International Journal of Wildland Fire18\(4\),pp\. 349–368\.Cited by:[§1](https://arxiv.org/html/2609.22191#S1.p1.1)\.
- Yuet al\.\(2026\)W\. Yu, A\. Ghosh, T\. S\. Finn, R\. Arcucci, M\. Bocquet, and S\. ChengA probabilistic approach to wildfire spread prediction using a denoising diffusion surrogate model\.Geoscientific Model Development19\(2\),pp\. 1027–1054\.External Links:[Document](https://dx.doi.org/10.5194/gmd-19-1027-2026)Cited by:[§1](https://arxiv.org/html/2609.22191#S1.p1.1)\.
## Appendix ADataset
Composition\.WildfireSpreadTS\[[Gerard et al\., 2023b](https://arxiv.org/html/2609.22191#bib.bib1)\]contains 13,607 daily images across 607 wildfire events in the contiguous United States, spanning January 2018 to October 2021 at 375 m ground resolution\. Events were selected from GlobFire as fires larger than 1000 hectares, with four buffer days added before and after each event, and each is converted to one HDF5 bundle of shape\[T,23,H,W\]\[T,23,H,W\]\. Image sizes vary between304×207304\\times 207and356×308356\\times 308because of how Google Earth Engine partitions its output, soHHandWWdiffer between fires\. The four years are unevenly represented with 176 events being from 2018, 74 from 2019, 201 from 2020 and 156 from 2021\.
Channel groups\.The 23 channels described in Section[2](https://arxiv.org/html/2609.22191#S2)come from seven source products and fall into the eight groups in Table[2](https://arxiv.org/html/2609.22191#A1.T2)\. Active fire is derived from the VIIRS 375 m active fire product, surface reflectance from VNP09GA bands I1, I2 and M11, the vegetation indices from VNP13A1, observed weather and drought from GRIDMET, forecast weather from the Global Forecast System, land cover from the MODIS yearly product under the IGBP classification, and topography from NASA SRTM\. Every other source is resampled bilinearly to the 375 m resolution of the active fire maps and aggregated into 24\-hour windows beginning at midnight\.
Table 2:The 23 released channels by group, following the feature groupings used in the ablation studies of[Gerard et al\. \[2023b\]](https://arxiv.org/html/2609.22191#bib.bib1), and which groups the*Vegetation*configuration draws on\.The weather group holds minimum and maximum surface temperature, total precipitation, wind speed, wind direction and specific humidity\. The energy release component and the Palmer Drought Severity Index come from the same product but are grouped separately, following the ablation studies in[Gerard et al\. \[2023b\]](https://arxiv.org/html/2609.22191#bib.bib1)\. The forecast group mirrors the weather group with two exceptions: the Global Forecast System supplies mean temperature in place of a minimum and a maximum, and it reports wind as speed and direction after conversion from its native u and v components\. Topography holds elevation together with the slope and aspect derived from it\.
From 23 channels to 40 model inputs\.The models do not read the released channels directly\. Preprocessing expands the categorical land cover channel into a one\-hot encoding over the 17 IGBP classes and appends the binarized active fire mask after standardization, so the tensor reaching the model carries 40 channels rather than 23\. The published parameter count for the logistic regression baseline confirms this width independently, since that model is a single3×33\\times 3convolution and its 361 parameters are exactly40×9\+140\\times 9\+1\[[Gerard et al\., 2023b](https://arxiv.org/html/2609.22191#bib.bib1)\]\. Under our*Vegetation*configuration the same model has 64 parameters\.
Continuous channels are standardized with means and standard deviations estimated on the training years alone\. Two groups are held out of that standardization\. The angular channels, meaning wind direction, aspect and forecast wind direction, are mapped through a sine transform instead, since a linear rescaling of a quantity that wraps at 360 degrees carries no meaning\. The land cover class is left categorical for the one\-hot expansion\. The statistics are fixed constants selected by the declared training years, so no test\-year statistic can reach training\.
Feature configurations\.In the released channel order the reflectance and vegetation index channels occupy indices 0 to 4 and the active fire channel occupies index 22, which after the one\-hot expansion of land cover lands at index 38, with the appended binary mask at 39\.*Vegetation*therefore keeps processed indices\[0,1,2,3,4,38,39\]\[0,1,2,3,4,38,39\], corresponding to five reflectance and vegetation\-index channels and two active\-fire channels\.*All*keeps every processed channel\. Both settings follow[Gerard et al\. \[2023b\]](https://arxiv.org/html/2609.22191#bib.bib1), who report a*Vegetation*and an*All*feature set and always include the fire masks rather than counting them as a feature group\. The two numeric labels used in the body count different quantities\. The 7 of*Vegetation*is a model input width, since the active fire channel enters twice, once standardized and once binarized, and it draws on 6 released channels\. The 23 of*All*is a released channel count, which the same preprocessing widens to 40 model inputs\.
Split protocol\.We use fold 0 of the cross\-validation protocol defined by the dataset authors\. Training uses the 2018 and 2019 fire seasons, validation uses 2020 for the models that use a validation set, and testing uses the held\-out 2021 season\. Standardization statistics come from 2018 and 2019 only, matching the training years\. The full protocol in[Gerard et al\. \[2023b\]](https://arxiv.org/html/2609.22191#bib.bib1)runs all twelve permutations of years across train, validation and test, which they describe as necessary given how much the yearly distributions differ\. Fold 0 also holds out the easiest of the four years by their own persistence measure\. Every number in Table[1](https://arxiv.org/html/2609.22191#S4.T1)comes from fold 0, so the comparisons are paired across architectures, not averaged over folds\.
Crops and augmentation\.Training and validation draw random 128\-by\-128\-pixel crops\. The crop window favors regions containing fire pixels instead of being placed uniformly at random, and horizontal flips, vertical flips and 90\-degree rotations are applied as augmentation, with the angular channels rotated to match so that wind direction and aspect stay physically consistent with the transformed image\.
Test\-time cropping\.Test evaluation does not use 128\-by\-128 crops\. Each scene is center\-cropped to the nearest multiple of 32 in each dimension, which is what the U\-Net encoders require, and scored at batch size 1 because scene sizes differ between fires\. ConvLSTM, the only model that needs a fixed input size, runs tiled inference across the scene\. On the final row and column the crop window is aligned to the bottom and right edges and overlapping predictions are overwritten, and the tiles are aggregated into a full\-resolution map before any metric is computed\.
Test\-set alignment\.The number of samples a fire contributes depends on how many leading days a model consumes, so the test set would otherwise differ between the one\-day and five\-day models\. The dataset class exposes a test adjustment that skips the firstnadj−nleadingn\_\{\\text\{adj\}\}\-n\_\{\\text\{leading\}\}samples of each fire\. Setting it to five throughout gives a skip of four for the one\-day models and zero for the five\-day models, so all six are scored on an identical set of targets beginning on the sixth day of each fire\.
Class imbalance\.About 0\.1% of pixels in the dataset carry an active fire detection, and the rest do not\[[Gerard et al\., 2023b](https://arxiv.org/html/2609.22191#bib.bib1)\]\. At that prevalence a model predicting no fire anywhere reaches roughly 99\.9% accuracy, which is why Section[2](https://arxiv.org/html/2609.22191#S2)treats accuracy as uninformative here and why Table[1](https://arxiv.org/html/2609.22191#S4.T1)reports precision and recall separately\. It also bears on how the AP column should be read, since AP is prevalence\-dependent and the base rate, not anything about the models, sets its absolute scale\.
Known issues in the released pipeline\.Two are worth recording, because both affect any work built on this benchmark\. The dataset class contained a bug that the authors found after publication and corrected in the repository, and they report that the corrected version gives slightly higher performance while leaving the trends unchanged\[[Gerard et al\., 2023a](https://arxiv.org/html/2609.22191#bib.bib21)\]\. Separately, in February 2026 the authors noted that the angular channels are transformed through sine only, where sine and cosine together would be needed to preserve direction, so wind direction and aspect reach the model with some information lost\[[Gerard et al\., 2023a](https://arxiv.org/html/2609.22191#bib.bib21)\]\. Both issues apply to every model we evaluate, so they do not favor any architecture, but the second one falls entirely inside the*All*configuration and gives one concrete reason why the additional channels may deliver less than they should\.
## Appendix BModels and Training
One harness evaluates all six architectures\. Every model exposes a prediction function returning a per\-pixel score in\[0,1\]\[0,1\], and that function is handed to a single evaluation routine that computes AP, F1, IoU, precision and recall identically for all of them, so differences between rows in Table[1](https://arxiv.org/html/2609.22191#S4.T1)reflect the models and not differences in metric implementation\. Each model is trained separately for each of the two feature configurations, giving the twelve runs reported in the body\. That routine issrc/evaluation/unified\_eval\.pyin the released code,[https://anonymous\.4open\.science/r/OfficialWildfireSpreadBench](https://anonymous.4open.science/r/OfficialWildfireSpreadBench)\.
ConvLSTM and UTAE receive a five\-day window, dayst−4t\{\-\}4throughtt, which is the setting both were designed for and the reason they are included\. The remaining four receive dayttalone\. UTAE also receives the day\-of\-year of each observation, which its temporal attention encoder uses as a positional signal\. The four baselines released with the dataset are optimized with AdamW at its default parameters,β1=0\.9\\beta\_\{1\}=0\.9,β2=0\.999\\beta\_\{2\}=0\.999andλ=0\.01\\lambda=0\.01\[[Gerard et al\., 2023b](https://arxiv.org/html/2609.22191#bib.bib1)\], and we keep those settings\. The released sweep configurations train for 10k steps, except logistic regression at 520\. Checkpoints are written on best validation loss, but the metrics in Table[1](https://arxiv.org/html/2609.22191#S4.T1)come from the unified evaluation pass at the end of each run, so they reflect the final weights rather than the selected checkpoint\.
Table 3:Model and training configuration\. Dayttis the most recent observed day of a fire, and every model predicts the active fire mask on dayt\+1t\{\+\}1\. Parameter counts are for the*All*configuration\. The four dataset baselines use the configurations released with the benchmark\.Discriminative baselines\.Logistic Regression is a single3×33\\times 3convolution applied to the input channels, following the released implementation, and it is the floor for what the channel set alone supports\. ResNet18 U\-Net is the segmentation\-models\-pytorch encoder–decoder released with the dataset\[[Gerard et al\., 2023b](https://arxiv.org/html/2609.22191#bib.bib1),[Ronneberger et al\., 2015](https://arxiv.org/html/2609.22191#bib.bib11),[He et al\., 2016](https://arxiv.org/html/2609.22191#bib.bib12)\]\. ConvLSTM\[[Shi et al\., 2015](https://arxiv.org/html/2609.22191#bib.bib9)\]replaces the fully connected gates of an LSTM with convolutions, so day\-to\-day dependence is modeled recurrently while spatial structure is preserved, and it is used here as a single block followed by a convolution producing the segmentation map\. UTAE\[[Garnot and Landrieu, 2021](https://arxiv.org/html/2609.22191#bib.bib10)\]is a U\-Net whose encoder applies simplified multi\-head self\-attention across the temporal dimension at the bottleneck, with the resulting attention map upscaled and applied to every skip connection\.
Positive\-class weighting\.Two of the six models optimize a binary cross\-entropy with an explicit positive\-class weight, and they are the two recall\-maximizers\. The released training script sets the UTAE fire\-class weight to the inverse relative frequency of that class in the training years, overriding the value written in the configuration file\[[Gerard et al\., 2023a](https://arxiv.org/html/2609.22191#bib.bib21)\]\. At the 0\.104% base rate of 2018 and 2019 that gives 964\. Our BCE U\-Net uses a fixed positive weight of 50, more than an order of magnitude smaller, and the two models sit in that same order in Figure[1](https://arxiv.org/html/2609.22191#S4.F1)\(b\), with UTAE further from the diagonal than BCE U\-Net at precision 0\.042 to 0\.061 against 0\.186 to 0\.203\. The three remaining discriminative baselines use unweighted overlap losses, Dice for ResNet18 U\-Net and Logistic Regression and Jaccard for ConvLSTM, and all three sit near the diagonal\. The size of the weight orders the operating points, not the architecture family\.
BCE U\-Net\.A segmentation U\-Net with 96 base channels, three downsampling levels and multi\-scale injection of the conditioning input at every level, trained with binary cross\-entropy under a positive weight of 50\. The model is optimized with AdamW at learning rate2×10−42\\times 10^\{\-4\}, weight decay10−410^\{\-4\}, batch size 16, gradient clipping at 1\.0 and a cosine annealing schedule\. A positive weight shifts the decision boundary away from 0\.5 by construction, so recall of 0\.86 to 0\.90 at precision of 0\.19 to 0\.20 is the expected direction of effect\. Section[5](https://arxiv.org/html/2609.22191#S5)does not claim the behavior is surprising\. Its claim is that AP does not reveal it, since exact AP depends only on the ranking of pixel scores\. Our estimate is computed on a fixed threshold grid, described in Appendix[C](https://arxiv.org/html/2609.22191#A3), and is therefore also mildly sensitive to the scale on which those scores are expressed\.
Flow Matching\.Flow matching\[[Lipman et al\., 2023](https://arxiv.org/html/2609.22191#bib.bib7),[Liu et al\., 2023](https://arxiv.org/html/2609.22191#bib.bib8)\]learns a continuous\-time velocity field carrying a simple prior to the data distribution and integrates it as an ODE at inference\. Applied directly to a binary next\-day fire mask it fails at this sparsity, for two reasons\. Integrating from the current fire mask makes the regression target the displacement between consecutive masks, which is overwhelmingly positive wherever new fire appears\. The learned field is therefore biased toward growth by construction, and integration amplifies that bias into large connected over\-predictions\. A binary mask is also a poorly posed regression target for a continuous field, since it is almost entirely zeros with sparse unit spikes\.
We therefore adopt the FlowSDF formulation\[[Bogensperger et al\., 2025](https://arxiv.org/html/2609.22191#bib.bib20)\]\. Integration starts from Gaussian noise instead of from the current mask, and the conditioning enters through the network instead of through the integration start, which removes the positive bias structurally\. The regression target is the truncated signed distance function of the dayt\+1t\{\+\}1mask, negative inside the fire, positive outside and zero at the boundary, truncated at 3 pixels\. The distance field is dense and smooth, so the loss carries signal everywhere and the model has to use the conditioning\. The loss is a plain mean squared error on the velocity field with no asymmetric weighting, which the balanced target makes unnecessary\. It is the only model in Table[1](https://arxiv.org/html/2609.22191#S4.T1)that under\-predicts\.
Signed distance values are standardized once using statistics estimated over the training masks and applied identically at training and test time, with mean 2\.9699 and standard deviation 0\.3237\. The zero level set of the raw field, which is the mask boundary, therefore maps to the fixed normalized value−mean/std=−9\.175\-\\text\{mean\}/\\text\{std\}=\-9\.175, so mask recovery is deterministic\. Because the distance distribution is right\-skewed at this sparsity, a small positive residual skew of about 0\.1 survives standardization, which the network learns around\. The recovered field is mapped to a bounded per\-pixel score in\[0,1\]\[0,1\]by a monotone transform that sends a raw signed distance of zero to exactly 0\.5, so the threshold metrics are computed at the model’s own mask boundary and AP is computed on a continuous score for all six models\. That score occupies a narrower band than the sigmoid outputs of the other five, which the fixed\-grid AP estimate of Appendix[C](https://arxiv.org/html/2609.22191#A3)is not fully invariant to\.
The velocity field is a conditional network with 128 base channels and a 256\-dimensional time embedding, trained for 100 epochs at batch size 16 with learning rate5×10−55\\times 10^\{\-5\}, 500 warmup steps, weight decay10−410^\{\-4\}, and gradient clipping at 0\.5\. Inference integrates 50 Euler steps from noise, and because the starting point is sampled, a single forward pass gives one draw from the model and not a deterministic map\. The row in Table[1](https://arxiv.org/html/2609.22191#S4.T1)is one such draw, taken without a fixed random seed\. Training also skips the optimizer step on any batch whose loss is non\-finite or more than five times a running mean, which guards against a single pathological batch corrupting the weights\.
Protocol for the two custom models\.BCE U\-Net and Flow Matching use the same year split and the same preprocessing as the baselines, but they are trained without a validation set, so no model selection was performed for them at all\. The periodic evaluation used to monitor them during development ran on the 2021 test set, so their formulation and hyperparameters were chosen with test metrics visible\. Their rows are best read as an upper bound on what each approach reaches on this split rather than as a clean held\-out estimate\.
Persistence baseline\.Carrying the dayttactive fire mask forward unchanged as the dayt\+1t\{\+\}1prediction is the persistence forecast, a standard reference baseline in short\-range weather forecasting\. It takes no parameters and reads no channels beyond the fire mask, so it is identical under both feature configurations\. On the 2021 test year it reaches AP 0\.287 and F1 0\.535, against a four\-year mean of AP 0\.193 and F1 0\.432\[[Gerard et al\., 2023b](https://arxiv.org/html/2609.22191#bib.bib1)\]\. Every architecture in Table[1](https://arxiv.org/html/2609.22191#S4.T1)exceeds it on AP, but only ResNet18 U\-Net, ConvLSTM and Logistic Regression exceed it on F1\. BCE U\-Net, UTAE and Flow Matching all fall below a parameter\-free baseline once a threshold is applied, in both feature configurations, which is a second reading of the same split the body draws between ranking quality and thresholded output\.
## Appendix CEvaluation Protocol
Metrics\.Precision is the fraction of predicted fire pixels that burned, recall the fraction of burned pixels that were predicted, F1 their harmonic mean, and IoU the intersection over union of predicted and observed fire pixels\. All four are computed at threshold 0\.5 on the per\-pixel score\. Average Precision summarizes the precision–recall curve, estimated on a fixed grid of 200 uniformly spaced thresholds so that memory stays bounded over the full test season\. That estimate is close to exact AP for scores spread across\[0,1\]\[0,1\]and drifts low for scores concentrated in a narrow band, so it should be read as an approximation rather than as the exact statistic\. Scores and targets are pooled across the whole 2021 test set before the metrics are computed, not averaged per fire, so large fires contribute in proportion to their pixel count\.
Redundancy between F1 and IoU\.Two of the reported metrics are algebraically determined by the others\. For a binary task,IoU=F1/\(2−F1\)\\text\{IoU\}=\\text\{F1\}/\(2\-\\text\{F1\}\)exactly\. This holds across all twelve runs in Table[1](https://arxiv.org/html/2609.22191#S4.T1)to within 0\.001, which is rounding\. The IoU columns therefore corroborate nothing the F1 columns do not already establish, and the two are best read as one measurement and not as two that agree\. We report both because both are community standards in this domain, not because they are independent evidence\.
Predicted\-to\-observed burned area\.The ratio of predicted burned area to observed burned area isR/P\\text\{R\}/\\text\{P\}, since predicted positives areTP/P\\text\{TP\}/\\text\{P\}and actual positives areTP/R\\text\{TP\}/\\text\{R\}\. It is recoverable from Table[1](https://arxiv.org/html/2609.22191#S4.T1)without further measurement, and it is the quantity behind the area figures quoted in the abstract and the conclusion\. Table[4](https://arxiv.org/html/2609.22191#A3.T4)gives it for all twelve runs\.
Table 4:Predicted burned area divided by observed burned area, equal to R/P, at threshold 0\.5 over the 2021 test year\. A value of 1\.00 matches the burned extent, above 1\.00 over\-predicts and below 1\.00 under\-predicts\.The three operating\-point profiles separate cleanly here\. The recall\-maximizers sit between 4\.24 and 22\.95, the balanced models within 0\.96 to 1\.21, and Flow Matching below 1\.00 in both configurations\. Flow Matching is the only architecture under 1\.00 in both, which is the basis for calling it conservative\. ResNet18 U\-Net falls marginally below on*All*at 0\.96 while sitting above on*Vegetation*, so it does not hold that profile across configurations and is grouped with the balanced models throughout\.
AP separates ConvLSTM and UTAE on*Vegetation*by 0\.403 against 0\.383, about five percent, while the area the two flag differs by a factor of roughly fifteen\. Exact AP is invariant to monotone rescaling of the scores, so a difference of that size lies outside what it can express by construction, and no threshold sweep recovers it either, because the quantity is a property of the operating point and not of the ranking\.
Thresholds in Figure[2](https://arxiv.org/html/2609.22191#S4.F2)\.Figure[2](https://arxiv.org/html/2609.22191#S4.F2)is rendered at threshold 0\.9 rather than the 0\.5 used in Table[1](https://arxiv.org/html/2609.22191#S4.T1), so that the BCE U\-Net panel stays legible, and it shows a single 2021 test fire rather than the whole test year\. Its 2\.1×\\timesand the 4\.82×\\timesin Table[4](https://arxiv.org/html/2609.22191#A3.T4)are therefore not the same measurement\. The figure is the more conservative of the two, since raising the threshold shrinks the predicted region, and the over\-prediction it shows is the amount that survives at a threshold chosen to flatter the model\.
Comparison to the published baselines\.Our harness reproduces the persistence baseline for 2021 exactly, at AP 0\.287 against the published 0\.287\[[Gerard et al\., 2023b](https://arxiv.org/html/2609.22191#bib.bib1)\]\. Because persistence is parameter\-free, that match tests the data loading and the target construction without any training in the way\. It does not exercise the continuous\-score path of the AP estimator, since persistence emits a binary score\. Table[5](https://arxiv.org/html/2609.22191#A3.T5)sets the four shared architectures against their published values\.
Table 5:AP for the four architectures shared with[Gerard et al\. \[2023b\]](https://arxiv.org/html/2609.22191#bib.bib1)\. Published values are means over the full twelve\-fold cross\-validation with the standard deviation across folds\. Ours are fold 0, with 2021 held out\.All eight differences are positive and none exceeds 1\.2 standard deviations of the published spread across folds\. Two documented effects explain the direction\. The first is that 2021 is the easiest of the four years by the dataset’s own measure, since persistence scores AP 0\.287 there against a four\-year mean of 0\.193, so any model tested on 2021 alone should sit above a twelve\-fold average\[[Gerard et al\., 2023b](https://arxiv.org/html/2609.22191#bib.bib1)\]\. The second is the corrected dataset class, which the authors report gives slightly higher performance than the version the published numbers were produced with\[[Gerard et al\., 2023a](https://arxiv.org/html/2609.22191#bib.bib21)\]\. We report the comparison as a check on the harness rather than as a like\-for\-like result, since one fold and twelve folds are not the same experiment\.相似文章
评估用于野火后泥石流预测的机器学习模型
本文使用美国地质调查局(USGS)流域尺度数据,系统评估了包括TabPFN基础模型在内的15种机器学习模型在野火后泥石流预测中的表现,发现TabPFN取得了最佳性能(威胁评分为0.637),并且合成数据增强能改善大多数模型的性能。
PyroAdapt: 在空间异质性和时间偏移下适应野火预测
PyroAdapt 提出了一个预训练-检索-排序框架,用于适应空间异质性和时间分布偏移的野火预测模型,提高检测精度并在预算约束下优先处理火灾易发地点。
FireWorldBench:通过耦合场火灾动力学评估复杂物理世界智能
FireWorldBench是一个基准,通过耦合场火灾动力学评估多模态大语言模型中的复杂物理世界智能,专注于预测、感知和因果推理等任务。
评估基础模型在极端环境事件中的泛化能力:以加州野火PM2.5为例
本文系统评估了时间序列基础模型(TSFMs)在预测野火烟雾导致的极端PM2.5浓度方面的表现,使用了加州12年的数据集。结果表明,像BiLSTM这样完全训练的循环基线模型优于TSFMs,挑战了更大预训练模型主导环境预测的假设。
数据驱动的火区划分用于改进短期野火预测
本文提出了一种无监督的火区划分方法,将分水岭检测与K-means聚类相结合,以改进短期野火预测,并在法国多个省和多种预测模型上表现出相较于基于网格的方法的一致性能提升。