Systematic Evaluation of TabPFN-TS for Zero-Shot Probabilistic Heat Load Forecasting in District Heating Networks

arXiv cs.LG Papers

Summary

This study systematically evaluates TabPFN-TS for zero-shot probabilistic heat load forecasting in district heating networks, comparing it with state-of-the-art time-series foundation models like Chronos-2 and machine-learning baselines.

arXiv:2608.20024v1 Announce Type: new Abstract: District heating energy hubs require reliable heat load forecasts for efficient operational scheduling. Conventional forecasting workflows train system-specific models on historical data, which can become burdensome when networks change through new consumers, retrofits, or changing operating regimes. Zero-shot time-series foundation models and in-context forecasting offer a promising alternative: they can adapt at inference time from recent observations rather than by repeated retraining. This study systematically evaluates TabPFN-TS against time-series foundation models and trained machine-learning baselines for probabilistic heat load forecasting in district heating networks. Unlike foundation models pretrained on large collections of real time series, TabPFN-TS relies on synthetic pretraining data, which avoids direct pretraining-test overlap but raises the question of whether the learned prior captures district heating dynamics. We analyze covariate choice, context length, temporal resolution, and prediction horizon on representative operating weeks, validate the selected configuration over a full year, and test transferability on a second network. The results identify hourly 24-hour forecasting with a 12-week rolling context and ambient temperature as a parsimonious high-performing configuration; longer context windows do not improve accuracy. TabPFN-TS remains close to Chronos-2 in deterministic accuracy, reaching CVRMSE values of 13.06% versus 12.48% on the main dataset, and lies within the critical-difference threshold in the daily-rank comparison. Although Chronos-2 achieves the lowest aggregate full-year error, TabPFN-TS shows better empirical calibration. Finally, the diagnostic findings motivate a Multi-Resolution Residual-Correction Forecaster that combines a low-frequency Base Forecaster with a short-horizon Residual Forecaster to improve longer-horizon planning accuracy.
Original Article
View Cached Full Text

Cached at: 08/21/26, 10:31 AM

# Systematic Evaluation of TabPFN-TS for Zero-Shot Probabilistic Heat Load Forecasting in District Heating Networks
Source: [https://arxiv.org/html/2608.20024](https://arxiv.org/html/2608.20024)
Karim K\. Ben HichamAffiliation:Process Systems Engineering, RWTH Aachen University,\[\-0\.2em\] Forckenbeckstr\. 51, 52074 Aachen, GermanyKai DerzsiPhilipp AlthausAlexander MitsosAffiliation:Process Systems Engineering, RWTH Aachen University,\[\-0\.2em\] Forckenbeckstr\. 51, 52074 Aachen, GermanyDirk Müller\[0\.6em\]Chair of Energy Efficient BuildingsIndoor ClimateRWTH Aachen University\[\-0\.2em\]Mathieustr\. 10, 52074 Aachen, Germany\[0\.4em\]Corresponding author:[ben\.spoek@eonerc\.rwth\-aachen\.de](mailto:[email protected])

###### Abstract

District heating energy hubs require reliable heat load forecasts for efficient operational scheduling\. Conventional forecasting workflows usually train system\-specific models on historical data, which can become burdensome when networks change through new consumers, retrofits, or changing operating regimes\. Zero\-shot time\-series foundation models and in\-context forecasting therefore offer a promising alternative: they can adapt at inference time from recent observations rather than by repeated retraining\. This study systematically evaluates TabPFN\-TS against state\-of\-the\-art time\-series foundation models and trained machine\-learning baselines for probabilistic heat load forecasting in district heating networks\. Unlike foundation models pretrained on large collections of real time series, TabPFN\-TS relies on synthetic pretraining data, which avoids direct pretraining\-test overlap but raises the question of whether the learned prior captures district heating dynamics\. We analyze covariate choice, context length, temporal resolution, and prediction horizon on representative operating weeks, validate the selected configuration over a full year, and test transferability on a second network\. The results identify hourly 24\-hour forecasting with a 12\-week rolling context and ambient temperature as a parsimonious high\-performing configuration; longer context windows do not improve accuracy\. TabPFN\-TS remains close to Chronos\-2 in deterministic accuracy, reaching CVRMSE values of 13\.06% versus 12\.48% on the main dataset, and lies within the critical\-difference threshold in the daily\-rank comparison\. Although Chronos\-2 achieves the lowest aggregate full\-year error, TabPFN\-TS shows better empirical calibration\. Finally, the diagnostic findings motivate a Multi\-Resolution Residual\-Correction Forecaster that combines a low\-frequency Base Forecaster with a short\-horizon Residual Forecaster to improve longer\-horizon planning accuracy\.

Keywords:District heating; heat load forecasting; TabPFN\-TS; prior\-data fitted networks; time\-series foundation models; zero\-shot forecasting; probabilistic forecasting

## 1Introduction

To meet the climate targets of the European Union[11](https://arxiv.org/html/2608.20024#bib.bib11), the heat supply of buildings must be decarbonized\. District heating networks provide one promising pathway toward this goal[31](https://arxiv.org/html/2608.20024#bib.bib32)because they enable the integration of waste heat and renewable heat sources[25](https://arxiv.org/html/2608.20024#bib.bib26)\. Realizing this potential requires coordinating heat sources and storage under fluctuating availability, making accurate heat\-demand forecasts particularly important for cost\- and emission\-optimal operation[47](https://arxiv.org/html/2608.20024#bib.bib47)\. Heat demand in district heating networks is shaped by multiple superimposed drivers[29](https://arxiv.org/html/2608.20024#bib.bib30), including weather\-dependent space\-heating demand[49](https://arxiv.org/html/2608.20024#bib.bib49), comparatively stochastic domestic hot water consumption[23](https://arxiv.org/html/2608.20024#bib.bib24), and additional operational, daily, and seasonal variations[7](https://arxiv.org/html/2608.20024#bib.bib7)\.

Further, district heating networks are systems that evolve continuously over time\. Buildings may be connected to or disconnected from the network, renovated through improved insulation, equipped with additional heating or storage technologies, or subject to changes in occupancy and usage patterns\. Such changes may not be fully known to the operator of the network\. Consequently, heat load forecasting in district heating networks is challenging because it must account simultaneously for weather\-driven seasonal trends, stochastic consumption components, operational effects, and gradual changes in the connected building stock, whose evolution can change the underlying demand dynamics\.[8](https://arxiv.org/html/2608.20024#bib.bib8)These characteristics make heat load forecasting with conventional, manually specified models difficult, particularly when detailed and current information on the connected buildings and their operation is unavailable\. Data\-driven forecasting methods are therefore attractive for this task because they can directly leverage operational data, but many conventional supervised machine\-learning models typically require explicit training and retraining in response to such changes\.[5](https://arxiv.org/html/2608.20024#bib.bib5)

Recent advances in tabular foundation models, such as TabPFN[21](https://arxiv.org/html/2608.20024#bib.bib22)and TabICL[41](https://arxiv.org/html/2608.20024#bib.bib43), are highly promising in this regard\. These models achieve state\-of\-the\-art performance on tabular prediction benchmarks[21](https://arxiv.org/html/2608.20024#bib.bib22);[41](https://arxiv.org/html/2608.20024#bib.bib43)entirely through in\-context learning, i\.e\., without task\-specific training and generally without expensive hyperparameter tuning\. Moreover, time\-series extensions such as TabPFN\-TS[22](https://arxiv.org/html/2608.20024#bib.bib23), which apply simple time\-series featurization, are competitive with leading forecasting models[37](https://arxiv.org/html/2608.20024#bib.bib38)despite not being pretrained on time\-series data\. However, because TabPFN\-TS instead relies on a synthetic prior[22](https://arxiv.org/html/2608.20024#bib.bib23), it remains unclear whether this prior adequately captures the complex heat\-demand dynamics of district heating networks\. We therefore investigate TabPFN\-TS’s suitability for heat load forecasting and assess whether it can produce accurate point predictions and calibrated predictive distributions directly from recent heat\-demand observations and known future weather covariates\. Chronos\-2[3](https://arxiv.org/html/2608.20024#bib.bib3)is used as a strong benchmark from the class of time\-series foundation models, while trained models serve as classical statistical, machine\-learning, and deep\-learning benchmarks\.

The contributions of this work are as follows:

- •We provide a detailed evaluation of TabPFN\-TS for probabilistic heat load forecasting in district heating networks\.
- •We investigate how covariate choice and context length affect TabPFN\-TS forecasts to identify configurations that achieve strong performance with limited input\-data requirements\.
- •We compare hourly and 15\-minute resolutions to assess whether TabPFN\-TS can capture higher\-frequency heat load dynamics that are relevant for district heating operation\.
- •We evaluate forecast horizons from 4 hours to one week to characterize horizon\-dependent TabPFN\-TS performance and its interaction with suitable context length\.
- •We compare TabPFN\-TS with classical machine\-learning methods and Chronos\-2 as a strong reference from the class of time\-series foundation models\.
- •Based on the empirical results, we derive a Multi\-Resolution Residual\-Correction Forecaster and validate it on heat load data from a district heating network\.

## 2Related Work

This section reviews the forecasting methods and foundation\-model concepts relevant to the present study\. It first reviews established forecasting approaches for district heating networks and their operational challenges\. It then introduces tabular Prior\-Data Fitted Networks, TabPFN, and the TabPFN\-TS extension\. Finally, it discusses recent foundation\-model benchmarks in energy forecasting\.

### 2\.1Heat Load Forecasting in District Heating Networks

Heat load forecasting methods range from physically informed models to statistical, machine\-learning, and deep\-learning approaches\.[36](https://arxiv.org/html/2608.20024#bib.bib37)Physically motivated and gray\-box methods incorporate knowledge of building and district\-heating dynamics[10](https://arxiv.org/html/2608.20024#bib.bib10);[35](https://arxiv.org/html/2608.20024#bib.bib36), while engineered load\-shape indicators provide a more empirical representation of operational demand patterns[16](https://arxiv.org/html/2608.20024#bib.bib17)\. Classical data\-driven approaches formulate forecasting as supervised regression, including linear regression and autoregressive models with exogenous inputs[12](https://arxiv.org/html/2608.20024#bib.bib12);[6](https://arxiv.org/html/2608.20024#bib.bib6)\.

Machine\-learning methods extend these formulations through nonlinear relationships\. Common examples include support vector regression[38](https://arxiv.org/html/2608.20024#bib.bib39), decision\-tree models[13](https://arxiv.org/html/2608.20024#bib.bib13), and gradient\-boosting methods such as XGBoost[43](https://arxiv.org/html/2608.20024#bib.bib45)\. Neural approaches include feed\-forward networks[24](https://arxiv.org/html/2608.20024#bib.bib25), nonlinear ARX models[40](https://arxiv.org/html/2608.20024#bib.bib42), convolutional\-recurrent architectures[45](https://arxiv.org/html/2608.20024#bib.bib46), recurrent neural networks[26](https://arxiv.org/html/2608.20024#bib.bib27), attention\-based LSTMs[51](https://arxiv.org/html/2608.20024#bib.bib50), and Transformer\-based models such as Informer[18](https://arxiv.org/html/2608.20024#bib.bib19)\. Rather than replacing earlier methods, these model families provide increasingly flexible alternatives for representing heat\-demand dynamics\.

Most DHN load\-forecasting models are fitted to a specific network and may require retraining as system dynamics change\.[17](https://arxiv.org/html/2608.20024#bib.bib18)Adapting to concept drift remains challenging[30](https://arxiv.org/html/2608.20024#bib.bib31), as changes must be detected, suitable adaptation windows selected, and recurring concepts accounted for[48](https://arxiv.org/html/2608.20024#bib.bib48)\. Such recurrence is particularly relevant in district\-heating systems, where seasonal shifts in the relative dominance of space\-heating and domestic\-hot\-water demand create recurring regimes[29](https://arxiv.org/html/2608.20024#bib.bib30)\. These limitations motivate forecasting methods that can adapt through context rather than repeated parameter fitting\. The following subsection introduces in\-context learning for tabular prediction and its extension to time\-series forecasting\.

### 2\.2Tabular Prediction with TabPFN and Its Extension to Time\-Series Data

TabPFN[21](https://arxiv.org/html/2608.20024#bib.bib22)is a foundation model for tabular prediction based on in\-context learning\. It is pretrained on a large collection of fully synthetic supervised\-learning tasks and, at inference time, predicts query labels from labeled context examples\. In contrast to deep neural networks or gradient\-boosted decision trees, this happens without requiring a conventional gradient\-based training loop for each target data set\.[20](https://arxiv.org/html/2608.20024#bib.bib21);[21](https://arxiv.org/html/2608.20024#bib.bib22);[41](https://arxiv.org/html/2608.20024#bib.bib43)

More formally, the model receives both labeled context examplesDc=\(Xc,yc\)D\_\{c\}=\(X\_\{c\},y\_\{c\}\)and unlabeled query examplesx∗x\_\{\*\}and produces predictionsy∗y\_\{\*\}in a single forward pass\. In this sense, the training data of the target task become context data for the pretrained model\. The Bayesian interpretation is central to Prior\-Data Fitted Networks \(PFNs\)\.[34](https://arxiv.org/html/2608.20024#bib.bib35)Letϕ\\phidenote a latent data\-generating task or process\. The posterior predictive distribution can be written as

p⁡\(y∗∣x∗,Dc\)=∫p⁡\(y∗∣x∗,ϕ\)​p​\(ϕ∣Dc\)​𝑑ϕ\.p\(y\_\{\*\}\\mid x\_\{\*\},D\_\{c\}\)=\\int p\(y\_\{\*\}\\mid x\_\{\*\},\\phi\)\\,p\(\\phi\\mid D\_\{c\}\)\\,d\\phi\.\(1\)Using Bayes’ rule,

p⁡\(ϕ∣Dc\)=p⁡\(Dc∣ϕ\)​p​\(ϕ\)p⁡\(Dc\),p\(\\phi\\mid D\_\{c\}\)=\\frac\{p\(D\_\{c\}\\mid\\phi\)p\(\\phi\)\}\{p\(D\_\{c\}\)\},\(2\)so that

p⁡\(y∗∣x∗,Dc\)\\displaystyle p\(y\_\{\*\}\\mid x\_\{\*\},D\_\{c\}\)=1p⁡\(Dc\)​∫p⁡\(y∗∣x∗,ϕ\)​p​\(Dc∣ϕ\)​p​\(ϕ\)​𝑑ϕ\\displaystyle=\\dfrac\{1\}\{p\(D\_\{c\}\)\}\\int p\(y\_\{\*\}\\mid x\_\{\*\},\\phi\)\\,p\(D\_\{c\}\\mid\\phi\)\\,p\(\\phi\)\\,d\\phi\(3\)∝∫p⁡\(y∗∣x∗,ϕ\)​p​\(Dc∣ϕ\)​p​\(ϕ\)​dϕ\.\\displaystyle\\propto\\int p\(y\_\{\*\}\\mid x\_\{\*\},\\phi\)\\,p\(D\_\{c\}\\mid\\phi\)\\,p\(\\phi\)\\,d\\phi\.Computing this integral explicitly would generally be intractable\. A PFN instead amortizes the computation: it is trained to map\(Dc,x∗\)\(D\_\{c\},x\_\{\*\}\)directly to an approximation of the posterior predictive distribution in a forward pass\. This approximation is Bayesian with respect to the synthetic pretraining prior over data\-generating tasks, rather than an explicit posterior over the real\-world target system\. The practical consequence is that TabPFN removes the standard per\-data\-set gradient\-based training loop in its default use\.[20](https://arxiv.org/html/2608.20024#bib.bib21);[21](https://arxiv.org/html/2608.20024#bib.bib22)No classical hyperparameter search is required for basic application, although inference\-time ensembling, preprocessing variants, fine\-tuning, or hyperparameter tuning can still be used to improve performance\.

TabPFN\-TS, as proposed by Hoo et al\.[22](https://arxiv.org/html/2608.20024#bib.bib23), can be directly used for time series forecasting by reformulating a time series as a tabular regression problem\. Each time step becomes a row in a table\. The feature set includes temporal features such as a running time index, cyclic calendar features, a non\-cyclic year feature, automatically extracted seasonal features, and optionally time\-varying covariates whose future values are known at prediction time\. The historical observations form the labeled contextDcD\_\{c\}, while the future time points form query rows with known featuresx∗x\_\{\*\}and unknown targetsy∗y\_\{\*\}\. TabPFN then outputs a probabilistic forecast for the full prediction horizon in a non\-autoregressive forward pass\. This is notable because the same study reports strong GIFT\-Eval[1](https://arxiv.org/html/2608.20024#bib.bib1)performance for TabPFN\-TS, although the underlying TabPFN model was pretrained only on synthetic non\-temporal tabular data\. The synthetic pretraining paradigm is important for the present application: since TabPFN is pretrained on synthetic supervised prediction tasks rather than on collections of real district\-heating time series, the risk of direct or indirect pretraining\-test overlap with historical district\-heating backtesting data is substantially reduced\. In addition, the probabilistic forecast can be interpreted as an approximation to the posterior predictive distribution under the learned synthetic prior\. This makes TabPFN\-TS particularly interesting for heat load forecasting, where reliable uncertainty estimates are operationally relevant and where point forecasts alone do not fully describe the stochastic demand process\. Hoo et al\.[22](https://arxiv.org/html/2608.20024#bib.bib23)report that TabPFN\-TS can distinguish signal from noise, adapt its probabilistic predictions accordingly, incorporate covariates, and handle seasonal patterns, while trend recognition, especially for linear and exponential trends, remains a reported weakness\.

### 2\.3Foundation\-Model Benchmarks in Energy Forecasting

Recent work indicates that time\-series foundation models can be useful in energy\-forecasting applications\. Meyer et al\.[33](https://arxiv.org/html/2608.20024#bib.bib34)show that foundation models can provide performance comparable to approaches trained from scratch for household electricity\-demand forecasts\. Koch et al\.[28](https://arxiv.org/html/2608.20024#bib.bib29)study thermal\-energy forecasting and compare multisource transfer\-learning strategies with time\-series foundation models, including TimesFM, Toto, and Chronos\-2\. Chronos\-2 achieves the strongest performance among the considered time\-series foundation models\. They further find that multisource transfer learning requires at least 16 building data sets before it outperforms time\-series foundation models\.

Direct TabPFN\-TS benchmarks in energy forecasting provide the closest context for this study: Hertel et al\.[19](https://arxiv.org/html/2608.20024#bib.bib20)compare TabPFN\-TS and Chronos\-2 for load forecasting in electricity grids at transmission\-system\-operator level and use Shapley\-based explanations to increase interpretability\. In their study, Chronos\-2 outperforms TabPFN\-TS in both the univariate and covariate settings\. Obermeier et al\.[37](https://arxiv.org/html/2608.20024#bib.bib38)benchmark several time\-series foundation models on energy\-forecasting tasks, including Chronos\-2 and TabPFN\-TS in a covariate setting\. The benchmark includes a data set from Flensburg’s district\-heating network as part of a heat load prediction subset\. For this heat load subset and overall, covariate models outperform univariate models\. Chronos\-2 again outperforms TabPFN\-TS overall and on the heat load subset, although TabPFN\-TS obtains some wins over Chronos\-2 in the covariate setting on other subcategories and several wins overall\. Chronos\-2[3](https://arxiv.org/html/2608.20024#bib.bib3)is therefore used as a strong foundation\-model benchmark\.

Alkhulaifi et al\.[2](https://arxiv.org/html/2608.20024#bib.bib2)show the importance of feature engineering for time\-series forecasting with TabPFN\. TabPFN\-TS largely automates this step by embedding generic temporal and seasonal feature construction into the forecasting pipeline\.

Ramachandran et al\.[42](https://arxiv.org/html/2608.20024#bib.bib44)compare several time\-series foundation models and classical machine\-learning baselines with a method that combines a continuous wavelet transform with LSTMs for heat\-demand forecasting\. Their proposed method outperforms the benchmark approaches\. However, it is specifically engineered for the considered forecasting problem and is therefore less directly transferable than general\-purpose time\-series foundation model approaches\.

Overall, the literature motivates the evaluation of TabPFN\-TS as a zero\-shot probabilistic covariate model for heat load forecasting and Chronos\-2 as a strong time\-series foundation\-model benchmark\. Existing work does not yet establish whether TabPFN\-TS is suitable for zero\-shot probabilistic heat load forecasting in DHNs, which application\-specific configuration is appropriate with respect to covariate choice, context length, temporal resolution, and forecast horizon, how it performs in terms of deterministic accuracy and probabilistic calibration, how it compares with Chronos\-2 and trained baselines, and whether a selected setup transfers to another DHN\.

## 3Methodology

This section defines the forecasting setup and evaluation protocol\. It first specifies how rolling forecasts are generated with TabPFN\-TS and its benchmarks, and then describes the deterministic, probabilistic, ranking, and computational metrics used for model comparison\.

### 3\.1Forecasting Models and Prediction Setup

The forecasting workflow uses rolling historical context windows comprising past heat load and observed weather variables\. Known future covariates comprise weather variables over the prediction horizon and calendar features; unless stated otherwise, ambient temperature is the only weather variable\. The prediction target is the future heat load\. TabPFN\-TS is evaluated without task\-specific gradient\-based training or manual data preprocessing\. It produces quantile forecasts, with the 50th quantile, denotedq50q\_\{50\}, used as the deterministic prediction and the remaining quantiles used for probabilistic evaluation\.

To contextualize the performance of TabPFN\-TS, Chronos\-2 is included as a strong pretrained time\-series foundation\-model benchmark for the same operational forecasting task where corresponding runs are available\. The benchmark comparison uses the same input\-output task, while contrasting two foundation\-model approaches with different pretraining paradigms and inference architectures\. TabPFN\-TS forecasts through the tabularized time\-series representation described above, whereas Chronos\-2 is used directly as a pretrained time\-series model\. Both models receive the same future covariate values over the prediction horizon and, where permitted by model constraints, the same rolling context windows\. The maximum number of context points that can be passed to Chronos\-2 is 8192 time steps\. If a selected context window exceeds this limit, only the most recent 8192 time steps are passed as context to Chronos\-2, and results are reported at the effective context length\.

Since Chronos\-2 was pretrained on real\-world time\-series data, potential train\-test contamination between the pretraining corpus and the downstream evaluation data must be considered\. Based on the disclosed Chronos\-2 pretraining corpus, leakage of time\-series training data into the 2024 heat load forecasting data sets for the district heating networks in Munich or Flensburg appears unlikely\. The real pretraining data sets listed in Appendix A of the Chronos\-2 documentation are univariate time series from broad domains such as electricity, solar and wind generation, weather, transport, web, cloud operations, and simulated U\.S\. building energy, but no district\-heating heat load data from Munich or Flensburg are reported[3](https://arxiv.org/html/2608.20024#bib.bib3)\. Moreover, Chronos\-2’s multivariate and covariate\-informed training tasks are stated to rely entirely on synthetic data\.

AutoGluon TimeSeries[5](https://arxiv.org/html/2608.20024#bib.bib5);[44](https://arxiv.org/html/2608.20024#bib.bib41)is used as a benchmark suite for the selected hourly day\-ahead forecasting task\. The benchmark contains several model families\. Persistence baselines are represented by Naive and SeasonalNaive\. Naive repeats the last observed heat load value over the prediction horizon\. SeasonalNaive uses a seasonal period of 24 hours and therefore predicts each future hour from the corresponding observed value on the previous day\. Classical statistical baselines include exponential smoothing \(ETS\), automatic ETS \(AutoETS\), and automatic autoregressive integrated moving average \(AutoARIMA\)\. Tabular machine\-learning baselines include direct LightGBM \(LGBM\) and XGBoost \(XGB\) models as well as a recursive LGBM model\. Neural baselines include DeepAR, TemporalFusionTransformer \(TFT\), SimpleFeedForward, and DLinear\. AutoGluon’s weighted ensemble is reported as AG Ensemble\.

The predictors are trained on 2023 target values and evaluated on rolling 24\-hour forecasts in 2024\. For a transfer validation on a new data set, the AutoGluon models are trained from scratch on the corresponding 2023 data before being evaluated on 2024 data\. The predictors are configured with ambient temperature as a known covariate, so future ambient\-temperature values are supplied over each prediction horizon\. Models with native known\-covariate support can use this information directly, including the direct and recursive tabular models, DeepAR, and TemporalFusionTransformer\. Other models in the benchmark, such as the persistence, statistical, SimpleFeedForward, and DLinear baselines, do not use the future covariate in the same way\.

The AutoGluon runs are not constrained to the same rolling context used for the selected TabPFN\-TS setup\. Instead, each prediction call receives the observed target history from the start of 2023 up to the respective forecast start\. Individual AutoGluon model classes then use this history according to their own defaults: statistical models may truncate long series internally, neural models use model\-specific context lengths, and tabular models construct lag, time, and covariate features\.

In the main experiments, future ambient temperature is provided as the realized temperature over the prediction horizon\. This setting can be interpreted as a perfect\-weather\-forecast assumption: the heat load forecasting problem is isolated from errors in numerical weather prediction, and model behavior is evaluated conditional on known future weather\.

To assess the sensitivity to weather\-covariate uncertainty, an additional analysis was performed using archived 24\-hour\-ahead temperature forecasts, i\.e\., weather information that would have been available before the respective valid time\. The analysis uses the Munich heat load data\. The archived forecasts are obtained from the Open\-Meteo Previous Runs API[52](https://arxiv.org/html/2608.20024#bib.bib52)at the Munich district coordinates using the DWD ICON\-D2/ICON Seamless output\. The queried forecast variable represents the hourly 2 m air temperature that was predicted 24 hours before each valid timestamp\. It therefore follows a rolling fixed\-lead\-time reference frame and is not equivalent to using the most recent complete weather forecast run available at the heat load forecast start, because the 24\-hour heat load horizon is assembled from fixed\-lead hourly forecast values rather than from one coherent forecast run\.

Archived weather forecasts covered the period from 20 January 2024 to 30 December 2024\. On this common evaluation interval, two otherwise identical settings are compared: a perfect\-weather variant using realized future temperature and a forecast\-weather variant using the corresponding archived temperature forecast\. The heat load history, forecast starts, target values, and historical context weather are identical in both variants, so the effect of replacing realized future weather with forecasted future weather is isolated\. Forecast starts with an incomplete 24\-hour forecast\-temperature horizon are excluded from both variants\.

All experiments are run on the RWTH Aachen University HPC cluster\. Per job one NVIDIA H100 GPU with 95,830 MiB memory, 8 CPU cores, and 64 GB RAM was used\. Nodes are equipped with Intel Xeon Platinum 8468 CPUs, with two sockets and 48 physical cores per socket\. The software environment used Python 3\.12\.3 and CUDA 12\.6\.3, with tabpfn\-time\-series 1\.1\.0, tabpfn 8\.0\.3, chronos\-forecasting 2\.2\.2, and autogluon\.timeseries 1\.5\.0\.

### 3\.2Evaluation Metrics and Approach for Model Comparison

Point\-forecast quality is assessed using the coefficient of variation of the root mean squared error \(CVRMSE\), the coefficient of determination \(R2R^\{2\}\), and mean absolute error \(MAE\)\. Fornnevaluation samples, letyty\_\{t\}denote the target value,y^t\\hat\{y\}\_\{t\}the corresponding point forecast, andy¯=1n​∑t=1nyt\\bar\{y\}=\\frac\{1\}\{n\}\\sum\_\{t=1\}^\{n\}y\_\{t\}the mean target value of the evaluation set\. In this study,yty\_\{t\}corresponds to the measured heat load\.

The three metrics are defined as

CVRMSE\\displaystyle\\mathrm\{CVRMSE\}=100⋅1n​∑t=1n\(y^t−yt\)2y¯,\\displaystyle=100\\cdot\\frac\{\\sqrt\{\\frac\{1\}\{n\}\\sum\_\{t=1\}^\{n\}\\left\(\\hat\{y\}\_\{t\}\-y\_\{t\}\\right\)^\{2\}\}\}\{\\bar\{y\}\},\(4\)R2\\displaystyle R^\{2\}=1−∑t=1n\(yt−y^t\)2∑t=1n\(yt−y¯\)2,\\displaystyle=1\-\\frac\{\\sum\_\{t=1\}^\{n\}\\left\(y\_\{t\}\-\\hat\{y\}\_\{t\}\\right\)^\{2\}\}\{\\sum\_\{t=1\}^\{n\}\\left\(y\_\{t\}\-\\bar\{y\}\\right\)^\{2\}\},\(5\)MAE\\displaystyle\\mathrm\{MAE\}=1n​∑t=1n\|y^t−yt\|\.\\displaystyle=\\frac\{1\}\{n\}\\sum\_\{t=1\}^\{n\}\\left\|\\hat\{y\}\_\{t\}\-y\_\{t\}\\right\|\.\(6\)
CVRMSE normalizes the root mean squared error by the mean target value, making the error comparable across operating regimes and between the two district heating networks\.

Probabilistic forecast quality is assessed from the stored quantile forecasts\. Central prediction intervals are evaluated by their empirical coverage, absolute interval width, and relative interval width\. Let𝟏​\{⋅\}\\mathbf\{1\}\\\{\\cdot\\\}denote the indicator function, which equals one if the condition in braces is true and zero otherwise\. ForKKcentral prediction intervals with nominal coveragesckc\_\{k\}, lower boundslk,tl\_\{k,t\}, and upper boundsuk,tu\_\{k,t\}, the empirical coverage of intervalkkis

c^k=1n∑t=1n𝟏\{lk,t≤yt≤uk,t\}\.\\hat\{c\}\_\{k\}=\\frac\{1\}\{n\}\\sum\_\{t=1\}^\{n\}\\mathbf\{1\}\\left\\\{l\_\{k,t\}\\leq y\_\{t\}\\leq u\_\{k,t\}\\right\\\}\.\(7\)
Calibration over multiple central intervals is summarized by the mean absolute coverage error \(MACE\), which averages the absolute differences between nominal and empirical coverage:

MACE=100K​∑k=1K\|c^k−ck\|\.\\mathrm\{MACE\}=\\frac\{100\}\{K\}\\sum\_\{k=1\}^\{K\}\\left\|\\hat\{c\}\_\{k\}\-c\_\{k\}\\right\|\.\(8\)
A MACE of zero would indicate perfect empirical calibration for the evaluated intervals; higher values indicate larger coverage deviations in percentage points\. Distributional forecast quality is evaluated with the continuous ranked probability score \(CRPS\)\. LetFt​\(z\)F\_\{t\}\(z\)denote the predictive cumulative distribution function for time steptt, whereyty\_\{t\}is the observed heat load\. Using the same indicator\-function notation, the CRPS is defined as

CRPS=1n∑t=1n∫−∞∞\(Ft\(z\)−𝟏\{z≥yt\}\)2dz\.\\mathrm\{CRPS\}=\\frac\{1\}\{n\}\\sum\_\{t=1\}^\{n\}\\int\_\{\-\\infty\}^\{\\infty\}\\left\(F\_\{t\}\(z\)\-\\mathbf\{1\}\\\{z\\geq y\_\{t\}\\\}\\right\)^\{2\}\\,dz\.\(9\)
Since the probabilistic forecasts are stored as discrete quantiles, CRPS is approximated by trapezoidal integration of the pinball loss over the stored quantile levels\. Lower MACE and CRPS values indicate better probabilistic forecasts with respect to their respective criteria\.

For the full\-year model comparison, daily forecast blocks are also used to compare the consistency of the algorithms over time\. Each issued 24\-hour forecast is treated as one block, and CVRMSE,R2R^\{2\}, and MAE are computed over the 24 hourly prediction points for each model and forecast start\. Models are ranked within each block, using ascending ranks for CVRMSE and MAE and descending ranks forR2R^\{2\}; ties are assigned average ranks\. The reported mean ranks therefore describe how consistently a model performs well across individual forecast days, not only its aggregate full\-year error\. Pairwise win\-rate matrices and critical\-difference diagrams are based on the daily CVRMSE ranks\.

To capture how forecast errors vary over time, MAE entries include a±\\pmterm\. The value preceding±\\pmdenotes the aggregate MAE computed over all evaluated predictions\. The value following±\\pmdenotes the standard deviation of block\-level errors: one MAE value is first computed for each issued 24\-hour forecast, and the standard deviation is then taken over these forecast\-level MAE values\. This quantity therefore reflects changes in forecast difficulty and model consistency across operating conditions, not uncertainty from repeated model training or stochastic initialization\.

For experiments with overlapping forecast windows, metrics are computed after collapsing predictions to a single operational trajectory: for each target timestamp, the forecast with the latest forecast start strictly before or equal to the target timestamp is retained\. Thus each realization contributes once to the error metrics\.

The long\-term metrics evaluate the forecast information that would be available for planning decisions over a longer look\-ahead window, where the total heat energy over the next hours can be more relevant than the exact phase of short\-term oscillations\. For the long\-term evaluation, a rolling 12\-hour energy metric is used\. At every hourly anchor, the forecast available in operation is integrated over the following 12 hours and compared with the realized heat energy over the same interval\.

A final evaluation criterion is computational cost\. To quantify it, the mean forecast\-loop runtime and the operational real\-time factor \(RTF\) are reported\. The operational real\-time factor is defined as the ratio between the total wall\-clock computation time required to generate forecasts and the total elapsed operating time over the evaluation periodTevalT\_\{\\mathrm\{eval\}\},

RTFop=∑itcomp,iTeval,\\mathrm\{RTF\}\_\{\\mathrm\{op\}\}=\\frac\{\\sum\_\{i\}t\_\{\\mathrm\{comp\},i\}\}\{T\_\{\\mathrm\{eval\}\}\},\(10\)wheretcomp,it\_\{\\mathrm\{comp\},i\}denotes the computation time of forecastii\. This metric captures the computational burden of a forecasting scheme in an operational setting, accounting for both the cost of individual forecasts and their update frequency\. Values below one indicate that the forecast is computed faster than real time\.

## 4Case Study and Experimental Design

This section introduces two district\-heating case studies with complementary forecasting challenges\. The Munich system is a comparatively small and expanding network\. The system dynamics are subject to ongoing change due to the expansion\. Because its operational data are proprietary, we additionally consider the Flensburg system, for which openly available data enable reproducible evaluation and which represents a substantially larger network with comparatively stable system dynamics\. The section then describes both systems and their operational and weather data before presenting the experimental design and a derived forecasting architecture\.

### 4\.1District Heating Networks

The data used for model development and primary evaluation originate from a district\-heating network serving a mixed\-use urban district in Munich, Germany\. The district comprises a heterogeneous building stock, including offices, retail, hotels, restaurants, cultural and educational facilities, and industrial uses\. As a result, the aggregated heat demand exhibits both weather\-dependent space\-heating demand and a year\-round domestic hot water consumption\.

Heat is supplied by a centralized energy hub connected to the district heating network\. The generation portfolio includes two combined heat and power units \(1111 kW thermal each\), a gas boiler \(1870 kW thermal\), a high\-temperature heat pump \(1284 kW thermal\), and a power\-to\-heat unit \(250 kW\), supported by two 12m3\\mathrm\{m\}^\{\\mathrm\{3\}\}thermal buffer storages\. The system is operated using a rule\-based control strategy that prioritizes combined heat and power generation, while the heat pump, gas boiler, and power\-to\-heat unit provide supplementary heat depending on operating conditions\.

For this study, the most relevant characteristic of the district is its diverse and aggregated heat\-demand profile, which combines strong seasonal weather dependence with a persistent non\-weather\-dependent base load\. This makes the network a representative case for district heat load forecasting\.

The Munich district is still expanding, and its energy system is therefore continuously changing\. This expansion is accompanied by a substantial increase in annual heat demand over the complete calendar years available in the data set\. The annual development is shown in Supplementary Figure S1\.

The case study uses operational heat demand data from the district heating network at 15\-minute resolution\. Weather covariates are obtained exclusively from German Weather Service \(DWD\) stations and aligned with the 15\-minute heat load data\. Temperature, relative humidity, wind speed, and precipitation are taken from station 03379 München\-Stadt\.[9](https://arxiv.org/html/2608.20024#bib.bib9)As this station has no solar irradiation sensor, irradiation data are obtained from station 05404 Weihenstephan\-Dürnast\. In 2024, these data were 99\.47% complete; remaining gaps were filled using surrounding DWD stations\. All weather variables are treated as known future inputs during forecasting\.

Although the native temporal resolution of the heat load data is 15 minutes, large thermal buffer storages partially smooth the network demand\. Consequently, this high temporal resolution is not strictly required for heat\-generation scheduling\. Selected experiments therefore also use hourly aggregated data to assess forecasting performance at a coarser, operationally relevant resolution\.

The measured heat load signal contains occasional anomalous periods, including zero heat demand, potentially indicating maintenance or faults, and repeated identical values over several time steps, which indicate signal outages\. For model development and analysis, we therefore select clean representative weeks\.

Candidate weeks are restricted to complete Monday–Sunday weeks without detected heat load anomalies\. Weeks containing signal outages are excluded, where a signal outage is defined as three or more consecutive identical heat load values\. From the remaining weeks, three operating regimes are selected based on ambient temperature: the coldest week represents winter operation, the hottest week represents summer operation, and the week with the largest within\-week temperature variability represents transitional operation\. The selected winter, transitional, and summer weeks in 2024 are January 15–21, April 22–28, and August 12–18, respectively\. Their detailed ambient\-temperature statistics are provided in Supplementary Table S1\.

These representative weeks are used to study seasonal effects and to compare forecast horizons, context lengths, and temporal resolutions under characteristic operating conditions\. In contrast, the final full\-year simulations are performed without removing potentially anomalous data points\. This evaluates the selected configuration under continuous operational conditions and follows a data\-driven setting in which no manual preprocessing of the target signal is applied for the final assessment\.

To characterize the dominant time scales in the measured heat demand, a Fourier spectrum of the full available 15\-minute heat load series is shown in Figure[1](https://arxiv.org/html/2608.20024#S4.F1)\. The very\-long\-period components above one year should not be interpreted as stationary cycles; they reflect the expansion of the Munich DHN and the associated increase in annual heat demand\. Further details of this annual development are provided in Supplementary Figure S1\. The spectrum also shows a dominant seasonal component close to one year, mainly attributable to space\-heating demand, and a clear diurnal component around 24 h\. Additional peaks at harmonics of the daily cycle, including 12 h, 8 h, 6 h, and 4 h, indicate that the diurnal load pattern contains substantial harmonic content rather than being well described by a single sinusoidal component\. The pronounced 12 h component is consistent with two dominant demand periods per day, for example, morning and evening domestic\-hot\-water peaks\. In contrast, no pronounced weekly peak is visible, suggesting that the aggregated DHN signal contains no strong strictly periodic seven\-day component\.

Figure 1:Fourier magnitude spectra of the full available 15\-minute heat load time series: \(a\) full period range and \(b\) periods up to one day\. The x\-axis is shown as period rather than frequency; marked periods indicate seasonal, weekly, daily, and sub\-daily components\.The data used in the validation originate from the district heating network of Flensburg, Germany\. The network has developed historically from 1969 into a city\-scale district heating system with more than 700 km of district heating pipelines and a connection rate of over 90% among private households\. This is the same network analyzed by Obermeier et al\.[37](https://arxiv.org/html/2608.20024#bib.bib38), discussed in Section[2\.3](https://arxiv.org/html/2608.20024#S2.SS3)\.

Heat is supplied from a central heating plant connected to the district heating network\. As of 2024, the heat supply infrastructure comprises six generation units and two large thermal storage tanks\.[46](https://arxiv.org/html/2608.20024#bib.bib14)The measured heat load signal reflects the aggregated demand supplied by this city\-scale network\. Heat loads are provided at hourly resolution from January 2014 to December 2024\.[14](https://arxiv.org/html/2608.20024#bib.bib15)Weather data for Flensburg are provided by DWD, using hourly air\-temperature observations from station 01666 Glücksburg\-Meierwik\.

### 4\.2Experimental Design

The experimental design proceeds in three stages\. First, configuration diagnostics on selected representative weeks are used to study how TabPFN\-TS responds to context length, temporal resolution, forecast horizon, weather\-covariate choice, and context\-data selection\. These experiments define the configuration used for the main full\-year evaluation\. Second, the selected configuration is evaluated over the complete year 2024 and compared with Chronos\-2 and trained AutoGluon baselines; the same setup is then applied to the Flensburg data set to assess transferability\. This stage also includes probabilistic forecast\-quality evaluation and a sensitivity analysis with archived weather forecasts\. Third, findings from the configuration diagnostics are used to derive an adapted forecasting architecture for high\-resolution operation\. This Multi\-Resolution Residual\-Correction Forecaster is introduced in the following subsection and evaluated in a separate full\-year simulation\.

### 4\.3Multi\-Resolution Residual\-Correction Forecaster

We propose a two\-stage architecture called the Multi\-Resolution Residual\-Correction Forecaster \(MRRC Forecaster\)\. It combines a long\-horizon base forecast with a short\-horizon residual correction at higher temporal resolution\.

In the MRRC Forecaster, the first component is referred to as the Base Forecaster\. It provides the general daily load trajectory at hourly resolution needed for scheduling and storage planning\. The Base Forecaster’s point forecast is linearly interpolated to the 15\-minute evaluation grid\. The second component is referred to as the Residual Forecaster\. It forecasts the remaining high\-frequency error of this baseline\. The architecture aims to improve the representation of short\-term dynamics while retaining an operational 24\-hour forecast\. The final prediction is obtained by adding the predicted residual back to the interpolated baseline\.

From a boosting perspective, the setup can be interpreted as a single residual\-correction step\. An initial modelF⁡\(x\)F\(x\)produces the base prediction, while a second modelh⁡\(x\)h\(x\)predicts the residualsr=y−F⁡\(x\)r=y\-F\(x\)\. The final prediction is then given byy^=F⁡\(x\)\+h⁡\(x\)\\hat\{y\}=F\(x\)\+h\(x\)\. In the MRRC Forecaster, the Base Forecaster corresponds toF⁡\(x\)F\(x\), and the Residual Forecaster corresponds toh⁡\(x\)h\(x\)\.

For the Base Forecaster, the context window is chosen based on the 24\-hour forecast results described in Section[5\.1](https://arxiv.org/html/2608.20024#S5.SS1)\. The forecast horizon of the Base Forecaster is set to 24 hours and a new forecast is issued every 12 hours\. Thus, at every point in time a forecast for at least the 12 hours ahead is available\.

Because residual prediction differs from direct heat load forecasting, the context\-window results from the main configuration experiments are not directly transferable to the Residual Forecaster\. Its context window is therefore tuned separately between 2 hours and 14 days\. The forecast horizon of the Residual Forecaster is set to 2 hours and a new forecast is issued every hour\. Hence, at every point in time, a high\-frequency forecast for at least the next hour is available\. The structure of the MRRC Forecaster is illustrated in Figure[2](https://arxiv.org/html/2608.20024#S4.F2)\.

A related multi\-scale residual\-aware forecasting setup for TSFMs has recently been proposed by Biswas et al\.[4](https://arxiv.org/html/2608.20024#bib.bib4)\.

Figure 2:Schematic structure of the Multi\-Resolution Residual\-Correction Forecaster\.We compare the MRRC Forecaster with two baseline setups\. The first is the Base Forecaster linearly interpolated to 15\-minute resolution without the residual correction\. The second baseline is a direct high\-frequency, high\-resolution predictor: it uses 15\-minute resolution, a 13\-hour prediction horizon, and hourly re\-issuing\. This setup ensures that at least 12 hours of forecast are available at every time step while matching the short\-term resolution of the MRRC Forecaster\. It also has access to the same context window as the MRRC Forecaster\. In the following, this setup is referred to as the High Frequency & High Resolution setup\. This comparison is deliberately strict: it has the same resolution, update frequency, and context window as the MRRC Forecaster, so it tests whether the two\-stage architecture adds value beyond simply forecasting directly at high resolution with frequent updates\.

Two sets of forecast\-quality metrics are reported to reflect the different operational requirements\. The short\-term metrics assess the latest available 15\-minute operational trajectory\. For the hourly updated high\-resolution and MRRC setups, this corresponds to the first hour after each forecast update and therefore quantifies how well the most recent forecast tracks the realized high\-resolution load profile\. In contrast, the long\-term metrics assess the forecast information available for planning decisions\. This long\-term evaluation uses the rolling 12\-hour energy metric described in Section[3\.2](https://arxiv.org/html/2608.20024#S3.SS2)\.

## 5Results

This section presents the empirical results\. It first reports the configuration experiments on selected operating weeks, then evaluates the selected setup over the full year and on the transfer network\. It further analyzes probabilistic forecast quality and weather sensitivity, and finally evaluates the derived MRRC Forecaster\.

### 5\.1Configuration Experiments

First, the required length of historical context data was investigated for the 24\-hour forecast task\. Context windows of 1 day, 1 week, 4 weeks, 12 weeks, and 1 year were evaluated for both 15\-minute and hourly forecast resolutions\. For TabPFN\-TS, a context window of 12 weeks provided the strongest CVRMSE in both time resolutions, as shown in Figure[3](https://arxiv.org/html/2608.20024#S5.F3)\. TheR2R^\{2\}and MAE results follow the same qualitative trend\. At hourly resolution, the 12\-week context achieved a CVRMSE of 11\.42%, anR2R^\{2\}of 0\.970, and an MAE of 102\.6 kW; at 15\-minute resolution, it achieved a CVRMSE of 15\.71%, anR2R^\{2\}of 0\.945, and an MAE of 138\.7 kW\. Extending the recent context to one year did not further improve the TabPFN\-TS point forecast in either resolution\. This suggests that very long recent context windows can include operating regimes that are less representative of the forecast period\. In this experiment Chronos\-2 slightly outperforms TabPFN\-TS on every metric for all resolutions and context window lengths\. For Chronos\-2, the best hourly result was also obtained with a 12\-week context window, while the 15\-minute result was best with a 4\-week context window and then remained close for the longer 12\-week context\.

Figure 3:Overall CVRMSE, CRPS, and mean forecast\-loop runtime per forecast start for 24\-hour heat load forecasts as a function of context length for the selected summer, winter, and transitional operating weeks\.The probabilistic score follows the same main pattern as the deterministic metrics\. For TabPFN\-TS, CRPS improved with increasing context length up to 12 weeks, reaching 70\.7 kW at hourly resolution and 97\.3 kW at 15\-minute resolution, before worsening again for the one\-year context\. Chronos\-2 showed the same optimum at hourly resolution, while the 15\-minute CRPS reached its minimum at 4 weeks and changed only slightly for longer contexts\. Thus, the CRPS results support the choice of a moderate recent context window and do not indicate a probabilistic benefit from using the full one\-year context\.

The corresponding forecast\-loop runtime comparison is shown in the lower panel of Figure[3](https://arxiv.org/html/2608.20024#S5.F3)\. For the 12\-week context, the mean forecast\-loop time per forecast start was 1\.41 s and 2\.07 s for TabPFN\-TS at hourly and 15\-minute resolution, respectively, compared to 0\.06 s and 0\.07 s for Chronos\-2\. The setup times of TabPFN\-TS and Chronos\-2 are similar and range from 24\.6 s to 55\.4 s across the timed runs\.

The influence of the forecast horizon was then evaluated separately\. First, a shorter 4\-hour horizon was evaluated at 15\-minute resolution, with forecasts issued every 4 hours\. Figure[4](https://arxiv.org/html/2608.20024#S5.F4)shows how the context\-window length affects this short\-horizon task\. When only very short context windows are available, namely 4 hours or 12 hours, TabPFN\-TS outperforms Chronos\-2\. For medium context windows of 1 day, 2 days, and 1 week, Chronos\-2 achieves lower CVRMSE values\. For longer context windows of 4 weeks and 12 weeks, both models perform very similarly\.

The shorter horizon also changes the context\-length optimum for TabPFN\-TS\. Whereas the 24\-hour forecast at 15\-minute resolution performed best with a 12\-week context window, the 4\-hour forecast performed best with a 4\-week context window\. Compared with the best TabPFN\-TS setup for 24\-hour forecasts at 15\-minute resolution, the best 4\-hour setup reduced CVRMSE from 15\.71% to 14\.04%, corresponding to a decrease of 1\.67 percentage points, or 10\.6% relative to the 24\-hour setup\. This comparison reflects the combined effect of the shorter forecast horizon and the higher forecast\-update frequency\.

Figure 4:Overall CVRMSE for 4\-hour heat load forecasts at 15\-minute resolution as a function of context length\.To separate the effect of the forecast horizon from the effect of the update frequency, an additional TabPFN\-TS run issued 24\-hour forecasts every 4 hours and was evaluated with the non\-overlapping latest\-forecast\-wins strategy defined in Section[3\.2](https://arxiv.org/html/2608.20024#S3.SS2)\. With a 12\-week context window, this setting achieved a CVRMSE of 14\.17% and anR2R^\{2\}of 0\.955\. This is only 0\.13 percentage points higher than the best 4\-hour TabPFN\-TS forecast, which used a 4\-week context window\. Thus, most of the improvement over the daily 24\-hour 15\-minute\-resolution forecast is explained by the higher forecast\-update frequency\. Once forecasts are reissued every 4 hours, extending the prediction horizon from 4 hours to 24 hours has only a small additional effect on the non\-overlapping evaluation metric\.

Finally, the selected hourly setup was evaluated for a longer 168\-hour forecast horizon to test whether the configuration remains suitable for week\-ahead prediction\. For weekly forecasts at hourly resolution, the optimal context\-window length for TabPFN\-TS was again 12 weeks\. The TabPFN\-TS weekly forecast achieved a CVRMSE of 14\.57%, anR2R^\{2\}of 0\.952, and an MAE of 129\.0 kW, compared with 11\.42%, 0\.970, and 102\.6 kW for 24\-hour forecasts issued daily with the same 12\-week context\. The weekly prediction therefore resulted in a 27\.6% higher CVRMSE, 1\.9% lowerR2R^\{2\}, and 25\.7% higher MAE\. Figure[5](https://arxiv.org/html/2608.20024#S5.F5)reports the corresponding metrics by operating regime\. On the selected operating weeks, TabPFN\-TS slightly outperformed Chronos\-2 for the weekly horizon in terms of aggregate CVRMSE\. This result did not hold in the full\-year weekly simulation: Chronos\-2 achieved a CVRMSE of 15\.64% compared with 17\.06% for TabPFN\-TS\. CompleteR2R^\{2\}and MAE results are provided in Supplementary Table S2\.

Figure 5:CVRMSE andR2R^\{2\}for daily 24\-hour and weekly 168\-hour forecasts at hourly resolution with a 12\-week context window for TabPFN\-TS and Chronos\-2\.These experiments identify a 12\-week recent context as a robust default for TabPFN\-TS, while showing that the attainable error level is governed not only by context length but also by temporal aggregation, forecast\-update frequency, and forecast horizon\.

After the context\-length, resolution, and horizon diagnostics, the sensitivity to weather\-covariate choice was evaluated for the hourly 24\-hour forecast setup with a 12\-week context window\. Table[1](https://arxiv.org/html/2608.20024#S5.T1)reports the aggregate performance over all selected operating weeks\.

Table 1:Weather feature selection for 24\-hour hourly forecasts with a 12\-week context window for TabPFN\-TS and Chronos\-2\.TabPFN\-TSChronos\-2Weather covariatesCVRMSE \(%\)R2R^\{2\}MAE \(kW\)CVRMSE \(%\)R2R^\{2\}MAE \(kW\)Amb\. temp\.11\.420\.970102\.6±\\pm58\.710\.600\.97496\.0±\\pm51\.6Amb\. temp\., relative humidity11\.430\.970102\.2±\\pm59\.410\.620\.97496\.7±\\pm52\.9Amb\. temp\., wind speed11\.090\.97298\.8±\\pm55\.510\.560\.97596\.0±\\pm49\.8Amb\. temp\., precipitation11\.160\.972100\.1±\\pm54\.810\.590\.97596\.0±\\pm51\.4Amb\. temp\., precip\., wind sp\.11\.250\.97199\.9±\\pm56\.510\.380\.97694\.2±\\pm48\.5Amb\. temp\., solar irradiation11\.330\.971101\.5±\\pm57\.510\.470\.97595\.3±\\pm50\.4All weather covariates11\.090\.97298\.8±\\pm55\.910\.320\.97695\.6±\\pm47\.3Including weather covariates beyond ambient temperature does not lead to a practically significant improvement for either TabPFN\-TS or Chronos\-2\. The variation in MAE across forecast starts is also not materially reduced by the larger feature sets\. While the configuration with all weather covariates achieves the best or tied\-best aggregate scores for both foundation models, the ordering of the intermediate feature sets is not consistent between TabPFN\-TS and Chronos\-2\. For example, including wind speed gives one of the strongest TabPFN\-TS results, whereas the combination of precipitation and wind speed is more favorable for Chronos\-2\. This indicates that the small changes in accuracy should not be overinterpreted as a robust ranking of individual weather variables\. At the same time, no additional weather covariate causes a clear deterioration in prediction quality\. The mean prediction\-call time increases by 2\.4% for TabPFN\-TS and 3\.3% for Chronos\-2 when moving from ambient temperature only to all weather covariates\. For robustness, transferability, and operational simplicity, ambient temperature alone is used in the subsequent experiments, because it is measured at many sites and is commonly available as a forecast variable\.

Next, we examine whether the composition of the context window itself can improve TabPFN\-TS forecasts\. This experiment tests whether augmenting the recent context with seasonally similar historical observations improves TabPFN\-TS forecasts when the recent context alone may not contain comparable operating conditions\. The spectral analysis of the Munich heat load data discussed in Section[4\.1](https://arxiv.org/html/2608.20024#S4.SS1)revealed a strong seasonal component, with the largest peak at a period of approximately one year\. We therefore investigate whether adding relevant calendar\-aligned observations from the previous year improves prediction accuracy\.

Here, relevant observations denote historical data from a window around the current calendar date, shifted by one year\. The recent\-plus\-relevant setup combines 6 weeks of recent context with a 6\-week block from the same calendar period one year earlier, so that the total amount of context data matches the observed optimal 12\-week recent\-data context\. According to the TabPFN\-TS documentation, automatic temporal feature engineering must be disabled for this non\-contiguous context configuration\. Therefore, the recent\-plus\-relevant experiment used only calendar features and disabled the running\-index and automatic seasonal features\. This experiment was carried out only for TabPFN\-TS, because Chronos\-2 expects strictly regular time steps without gaps in the context\.

No performance benefit was obtained from adding relevant data\. Table[2](https://arxiv.org/html/2608.20024#S5.T2)also shows that the default automatic temporal feature engineering improved prediction performance by 2\.3% on CVRMSE compared with the recent\-only setup\.

Table 2:Effect of context\-data selection on the daily hourly baseline\. "Recent 12 \(auto feat\.\)" is the default automatic\-feature setup; "Recent 12" disables automatic temporal features; "Recent 6 \+ Relevant 6" combines 6 recent weeks with 6 weeks from the same calendar period one year earlier\.SetupCVRMSE \(%\)R2R^\{2\}MAE \(kW\)Recent 12 \(auto feat\.\)11\.420\.970102\.6±\\pm58\.7Recent 1211\.690\.969106\.9±\\pm63\.1Recent 6 \+ Relevant 612\.690\.963113\.0±\\pm68\.5Taken together, the context\-length, temporal\-resolution, forecast\-horizon, weather\-covariate, and context\-data diagnostics motivate using 24\-hour\-ahead forecasts at hourly resolution with a 12\-week rolling context window and ambient temperature as the only weather covariate in the subsequent full\-year and transfer analyses\. Figure[6](https://arxiv.org/html/2608.20024#S5.F6)shows the resulting TabPFN\-TS and Chronos\-2 forecasts for the selected winter, transitional, and summer weeks\.

Figure 6:Munich predictions for the selected 24\-hour forecast setup at hourly resolution with a 12\-week context window\. Rows show the winter, transitional, and summer representative weeks; columns compare TabPFN\-TS and Chronos\-2\. The shaded areas show the q10–q90 prediction intervals\.
### 5\.2Seasonal Effects on Forecast Accuracy and Uncertainty

Figure[5](https://arxiv.org/html/2608.20024#S5.F5)shows that forecast performance depends strongly on the operating period\. For daily 24\-hour forecasts, the CVRMSE is lowest in winter, followed by the transitional period and summer\. TheR2R^\{2\}score follows a different pattern: winter and summer are comparable, while the transitional week achieves the highestR2R^\{2\}\. During the transitional week, the loss in accuracy when moving from daily to weekly forecasts is particularly visible\.

The seasonal pattern is also reflected in the probabilistic forecasts\. As an example, Table[3](https://arxiv.org/html/2608.20024#S5.T3)reports the empirical coverage and width of the 80% prediction interval bounded byq10q\_\{10\}andq90q\_\{90\}\. The relative interval width is lowest in winter and highest in summer, with the transitional period in between\. The empirical coverage is highest in winter for both TabPFN\-TS and Chronos\-2\. The prediction intervals tend to be conservative in this period\. In the transitional period, the coverage is lowest for both models, indicating overconfident probabilistic forecasts\. Overall, TabPFN\-TS is closer to the nominal 80% coverage, while Chronos\-2 is more overconfident, especially in the transitional and summer periods\.

Table 3:Probabilistic interval quality evaluated onq10q\_\{10\}toq90q\_\{90\}for daily issued forecasts at hourly resolution with a 12\-week context on the selected representative weeks\.PeriodModelCov\. \(%\)Width \(kW\)Rel\. w\. \(%\)WinterTabPFN\-TS88\.1498\.922\.1Chronos\-289\.3488\.421\.6TransitionalTabPFN\-TS75\.0370\.328\.5Chronos\-264\.3317\.724\.5SummerTabPFN\-TS79\.2114\.836\.7Chronos\-269\.093\.930\.0OverallTabPFN\-TS80\.8328\.025\.4Chronos\-274\.2300\.023\.2
### 5\.3Full\-Year Benchmark and Transfer Validation

The results of the full\-year simulations are reported in Table[4](https://arxiv.org/html/2608.20024#S5.T4)\. The entries are sorted by aggregate CVRMSE within each network\. On the Munich data set, Chronos\-2 achieved the lowest aggregate error, while TabPFN\-TS outperformed the trained AutoGluon models across the reported accuracy metrics\. To assess whether models perform consistently when transferred to a different data set, the rank\-shift column reports the change in CVRMSE position in Flensburg relative to the Munich ordering\. The relative error metrics were generally lower for Flensburg than for Munich, especially for the strongest models\. Chronos\-2, TabPFN\-TS, and TFT remained the three strongest models in both networks, whereas the naive baselines stayed at the lower end of the ranking\. The largest rank changes occurred for AutoETS and ETS, which lost five CVRMSE positions in Flensburg, and for Feed\-forward, Recursive LGBM, and Seasonal naive, which gained three positions each\. Chronos\-2 again achieved the lowest aggregate error, while TabPFN\-TS remained ahead of the trained AutoGluon models\.

Table 4:Full\-year benchmark and transfer\-validation results for hourly 24\-hour forecasts in 2024\.Aggregate metricsMean daily rankModelRank shiftCVRMSE \(%\)R2R^\{2\}MAE \(kW\)CVRMSER2R^\{2\}MAEMunichChronos\-212\.480\.96285\.2±\\pm49\.33\.053\.013\.11TabPFN\-TS13\.060\.95990\.0±\\pm51\.53\.893\.863\.83TFT14\.570\.948101\.6±\\pm60\.05\.865\.845\.92Direct LGBM15\.850\.939113\.3±\\pm69\.26\.496\.416\.49DeepAR16\.430\.935116\.4±\\pm78\.06\.666\.626\.72AG Ensemble16\.430\.935116\.4±\\pm78\.06\.666\.626\.72Direct XGB17\.360\.927126\.3±\\pm77\.28\.458\.408\.62AutoETS18\.450\.917132\.5±\\pm85\.07\.757\.737\.84ETS18\.450\.917132\.5±\\pm85\.07\.757\.737\.84DLinear18\.650\.916134\.6±\\pm85\.58\.668\.688\.74Feed\-forward19\.400\.909141\.4±\\pm88\.79\.589\.629\.66AutoARIMA20\.510\.898148\.1±\\pm92\.610\.1110\.139\.90Recursive LGBM21\.120\.892151\.8±\\pm95\.410\.6510\.7110\.54Seasonal naive22\.550\.877161\.6±\\pm102\.511\.6411\.7211\.52Naive23\.720\.864172\.0±\\pm108\.812\.7912\.9212\.55FlensburgChronos\-29\.820\.9707913\.6±\\pm5187\.23\.983\.984\.14TabPFN\-TS10\.370\.9668437\.3±\\pm5411\.14\.774\.774\.60TFT10\.690\.9648669\.2±\\pm6199\.95\.075\.075\.06DeepAR↑1\\uparrow 110\.720\.9648677\.9±\\pm6003\.65\.015\.015\.12AG Ensemble↑1\\uparrow 110\.720\.9648677\.9±\\pm6003\.65\.015\.015\.12Direct LGBM↓2\\downarrow 211\.120\.9619380\.2±\\pm5728\.96\.206\.206\.50Direct XGB11\.430\.9599716\.2±\\pm5923\.96\.896\.897\.07Feed\-forward↑3\\uparrow 313\.050\.94610665\.8±\\pm7458\.07\.447\.447\.53DLinear↑1\\uparrow 113\.620\.94111200\.1±\\pm7702\.88\.648\.648\.58Recursive LGBM↑3\\uparrow 313\.750\.94011456\.8±\\pm7405\.59\.239\.239\.22Seasonal naive↑3\\uparrow 315\.530\.92412775\.2±\\pm8593\.810\.6810\.6810\.40AutoARIMA15\.800\.92112820\.7±\\pm9293\.59\.669\.669\.55AutoETS↓5\\downarrow 521\.210\.85818255\.5±\\pm12529\.312\.4312\.4312\.38ETS↓5\\downarrow 521\.210\.85818255\.5±\\pm12529\.312\.4312\.4312\.38Naive22\.390\.84218766\.4±\\pm14250\.612\.5612\.5612\.34The critical\-difference diagrams are shown in Fig\.[7](https://arxiv.org/html/2608.20024#S5.F7)\. For both data sets, TabPFN\-TS lies within the critical difference of Chronos\-2\. On the Munich data set, the two time\-series foundation models form the leading group and outperform all trained AutoGluon baselines in the daily\-rank comparison\. In Flensburg, the leading group is broader: TFT, DeepAR, and AG Ensemble also lie within the critical difference of the foundation\-model results\. Detailed pairwise daily CVRMSE win rates are provided in Supplementary Figure S2\.

Figure 7:Critical\-difference diagrams for the full\-year based on daily CVRMSE ranks\.
### 5\.4Probabilistic Forecast Quality

Probabilistic forecast quality is compared between TabPFN\-TS and Chronos\-2\. Table[5](https://arxiv.org/html/2608.20024#S5.T5)compares representative 50% and 80% central prediction intervals with their empirical coverage for both models, while Figure[8](https://arxiv.org/html/2608.20024#S5.F8)reports their aggregate probabilistic scores for the full\-year hourly forecasts\. Coverage results for all evaluated intervals are provided in Supplementary Table S3\. TabPFN\-TS is better calibrated, as indicated by its lower MACE\. Chronos\-2 achieves a lower CRPS, indicating better distributional accuracy under this score, but tends to be slightly overconfident and is less well calibrated in terms of empirical coverage\.

Figure 8:Probabilistic full\-year forecast scores for TabPFN\-TS and Chronos\-2\.Table 5:Empirical coverage of representative central prediction intervals for the full\-year hourly forecasts\. Interval width is reported in kW for both networks\.TabPFN\-TSChronos\-2Interval \(%\)Emp\. cov\. \(%\)Width \(kW\)Rel\. width \(%\)Emp\. cov\. \(%\)Width \(kW\)Rel\. width \(%\)Munich50\.0049\.92145\.113\.9847\.77132\.312\.7580\.0080\.50288\.627\.8077\.73260\.325\.07Flensburg50\.0050\.731370211\.3548\.371269810\.5280\.0079\.822712622\.4878\.192472420\.49
### 5\.5Weather\-Forecast Sensitivity

Table[6](https://arxiv.org/html/2608.20024#S5.T6)summarizes the resulting change in point and probabilistic forecast quality\. Both models are affected, with Chronos\-2 showing slightly greater sensitivity in this experiment: its CVRMSE increases by 9\.5%, compared with 8\.3% for TabPFN\-TS\. The degradation is visible for both models in point\-forecast accuracy, distributional scores, and calibration\.

Table 6:Forecast\-quality change when replacing realized future temperature with archived 24\-hour\-ahead temperature forecasts for the Munich sensitivity subset\.Point forecastProbabilistic forecastModelWeather cov\.CVRMSE \(%\)R2R^\{2\}MAE \(kW\)CRPS \(kW\)MACE \(%\)TabPFN\-TSRealized13\.640\.95487\.461\.90\.49Forecasted14\.77 \(\+8\.3%\)0\.946 \(\-0\.8%\)97\.7 \(\+11\.7%\)68\.8 \(\+11\.1%\)2\.71Chronos\-2Realized12\.890\.95982\.258\.21\.72Forecasted14\.12 \(\+9\.5%\)0\.950 \(\-0\.9%\)92\.6 \(\+12\.6%\)65\.3 \(\+12\.1%\)4\.91
### 5\.6Multi\-Resolution Residual\-Correction Forecaster Performance

The MRRC Forecaster is motivated by four findings from the preceding experiments:

1. 1\.The 24\-hour forecast at hourly resolution with a 12\-week context window provides a robust forecast of the slowly varying daily load trend, see Figure[3](https://arxiv.org/html/2608.20024#S5.F3)\.
2. 2\.The high\-resolution experiments depicted in Figure[4](https://arxiv.org/html/2608.20024#S5.F4)show that a higher forecast\-update frequency improves the prediction of short\-term dynamics\.
3. 3\.For the more frequently updated high\-resolution tasks, shorter context windows are sufficient and can even be optimal, suggesting that recent data is more relevant for high\-frequency behavior, see Figure[4](https://arxiv.org/html/2608.20024#S5.F4)\.
4. 4\.As reported in Section[5\.1](https://arxiv.org/html/2608.20024#S5.SS1), reducing the forecast horizon improves short\-term accuracy further compared with issuing longer high\-resolution forecasts\.

Together, these observations motivate a forecaster that combines long\-term forecasting with a longer context window and short\-term forecasting with a shorter forecast horizon\. Following the configuration experiments in Section[5\.1](https://arxiv.org/html/2608.20024#S5.SS1), the Base Forecaster uses the selected hourly 24\-hour forecast setup with a 12\-week rolling context window and is reissued every 12 hours\. The context window of the Residual Forecaster was tuned separately for the residual\-prediction task and set to 7 days\. Table[7](https://arxiv.org/html/2608.20024#S5.T7)reports the resulting full\-year comparison of the three TabPFN\-TS setups\.

Table 7:Deterministic full\-year comparison of the TabPFN\-TS MRRC Forecaster with the Base Forecaster and High Frequency & High Resolution setups\.Short\-term forecastLong\-term forecastComputationSetupCVRMSE \(%\)R2R^\{2\}MAE \(kW\)12 h E\-CVRMSE \(%\)12 h E\-bias \(%\)RTFBase Forecaster18\.640\.919124\.9±\\pm57\.87\.56\-1\.231\.72×10−51\.72\{\\times\}10^\{\-5\}High Frequency & High Resolution17\.960\.925114\.2±\\pm54\.47\.41\-1\.234\.21×10−44\.21\{\\times\}10^\{\-4\}MRRC Forecaster17\.950\.925114\.0±\\pm54\.17\.21\-1\.182\.16×10−42\.16\{\\times\}10^\{\-4\}

The MRRC Forecaster achieves similar short\-term results as the High Frequency & High Resolution setup\. For the long\-term 12\-hour energy metric, it reduces the integrated error by 2\.7%, and reduces the bias magnitude by 4\.1%\. Across the reported accuracy metrics, the MRRC Forecaster is the leading setup and has the lowest variance based on short\-term MAE\. In addition, it reduces the RTF by 48\.7% compared with the High Frequency & High Resolution setup, as the different forecasting schemes differ in how often computationally expensive predictions are issued\. Compared with the Base Forecaster, the MRRC Forecaster reduces the short\-term CVRMSE by 3\.7%, the 12 h E\-CVRMSE by 4\.6%, and the bias magnitude by 4\.1%\. The Base Forecaster requires less computation and is 12\.6 times faster than the MRRC Forecaster\.

For Chronos\-2, the MRRC Forecaster reduced the 12\-hour E\-CVRMSE by 6\.2% and the RTF by 21\.7% relative to the High Frequency & High Resolution setup, while increasing short\-term CVRMSE by 2\.6%\. Complete results are provided in Supplementary Table S4\.

## 6Discussion

Overall, Chronos\-2 outperformed TabPFN\-TS in deterministic forecasting, consistent with Obermeier et al\.[37](https://arxiv.org/html/2608.20024#bib.bib38)\. This advantage may reflect Chronos\-2’s native time\-series architecture and pretraining objective, which are more directly aligned with temporal structure than TabPFN\-TS’s tabular reformulation under a synthetic prior\. Nevertheless, TabPFN\-TS’s competitive performance supports the potential of PFN\-based in\-context heat load forecasting, particularly if future variants use priors or pretraining tasks tailored to district\-heating dynamics\.

The context\-length diagnostics in Section[5\.1](https://arxiv.org/html/2608.20024#S5.SS1)show that forecast accuracy does not improve monotonically with the amount of historical context, as can be seen in Figure[3](https://arxiv.org/html/2608.20024#S5.F3)and Figure[4](https://arxiv.org/html/2608.20024#S5.F4), contradicting the assumption that more data necessarily improve forecasts\. Optimal context length therefore remains application\-specific even for foundation models\. This may reflect superimposed demand processes: space heating is strongly seasonal and weather\-driven, whereas domestic hot water is more stochastic and exhibits different, weaker temporal patterns\. Without explicit separation, long contexts may include regimes irrelevant to the current forecast and obscure rather than improve in\-context prediction\.

The configuration experiments showed that hourly forecasts achieved substantially lower error metrics than forecasts at 15\-minute resolution\. This is likely due to temporal aggregation: hourly averaging smooths high\-frequency dynamics that are difficult to predict at 15\-minute resolution\. One potential driver of these dynamics is domestic\-hot\-water draw\-off behavior\. The exact timing of individual draw\-off events contains a stochastic component and is therefore inherently difficult to forecast[15](https://arxiv.org/html/2608.20024#bib.bib16), whereas broader temporal patterns in domestic\-hot\-water demand are easier to capture[32](https://arxiv.org/html/2608.20024#bib.bib33)\. Additional sources of high\-frequency variability may include hydraulic interactions between buildings[50](https://arxiv.org/html/2608.20024#bib.bib51), varying transport delays between the energy hub and consumers, delay\-dependent heat losses, hysteresis\-controlled pressure pumps, and other operational effects\. The short\-horizon experiments also revealed a setting in which TabPFN\-TS outperformed Chronos\-2 in deterministic accuracy: for the 4\-hour forecast task at 15\-minute resolution, TabPFN\-TS performed better than Chronos\-2 when only very short context windows of 4 or 12 hours were available, as shown in Figure[4](https://arxiv.org/html/2608.20024#S5.F4)\. This is consistent with energy\-demand forecasting results by Kim et al\.[27](https://arxiv.org/html/2608.20024#bib.bib28), who report strong TabPFN performance under short\-context conditions, and with the broader finding that TabPFN can deliver accurate predictions under low data availability[21](https://arxiv.org/html/2608.20024#bib.bib22)\.

The seasonal effects in Section[5\.2](https://arxiv.org/html/2608.20024#S5.SS2)provide complementary evidence on how demand composition affects forecast accuracy and uncertainty\. The lower winter daily CVRMSE in Figure[5](https://arxiv.org/html/2608.20024#S5.F5)is consistent with a larger share of weather\-driven space\-heating demand, creating a stronger and more predictable baseload\. In summer, the smaller space\-heating share and greater relative contribution of stochastic domestic hot water demand are expected to increase the relative forecast error\. The higherR2R^\{2\}during the transitional period results from greater variation in measured demand, allowing more variance to be explained despite the relative error not being lowest\. The greater accuracy loss from daily to weekly forecasts in this period may arise because the longer horizon spans rapidly changing load and weather conditions, making multi\-day extrapolation harder than in the more stable winter and summer periods\. The relative interval widths in Table[3](https://arxiv.org/html/2608.20024#S5.T3)similarly suggest that relative uncertainty rises as space\-heating demand becomes less dominant and stochastic components gain importance\. The low empirical coverage during the transitional period may be related to the 12\-week context window containing operating conditions less representative of this rapidly changing regime\.

Neither TabPFN\-TS nor Chronos\-2 materially benefited from weather covariates beyond ambient temperature\. This is consistent with Petković et al\.[39](https://arxiv.org/html/2608.20024#bib.bib40), who favor compact weather\-input sets over large sets of meteorological covariates for heat load prediction\. However, additional covariates also caused no clear deterioration in forecast accuracy\. Because their computational cost was negligible in the tested setting, including all reliably available weather covariates remains defensible when sufficiently accurate future values are available\.

The transfer results in Section[5\.3](https://arxiv.org/html/2608.20024#S5.SS3)and Table[4](https://arxiv.org/html/2608.20024#S5.T4)should be interpreted in light of network scale\. The lower relative errors in the Flensburg validation case may partly reflect aggregation: domestic hot water demand is comparatively stochastic at the individual\-building level but becomes more deterministic as more buildings are connected\. Because Flensburg is a city\-scale district\-heating system whereas Munich is a smaller mixed\-use district, stronger aggregation in Flensburg may smooth idiosyncratic consumption events and contribute to its lower relative forecasting errors\.

Our literature review in Section[2\.1](https://arxiv.org/html/2608.20024#S2.SS1)identified retraining under concept drift as a central challenge in heat load forecasting\. We therefore hypothesized that rolling\-context in\-context learning could adapt implicitly to gradual changes in system behavior without explicit retraining\. Results from the dynamically evolving Munich network support this hypothesis: Chronos\-2 and TabPFN\-TS performed strongly and consistently outperformed task\-specific models trained on historical data from 2023\. Between training and evaluation, the network expanded, while changes in consumer composition and control strategies may also have altered its time\-series dynamics\. By conditioning on recent observations, the in\-context learners could account for these changes directly, whereas models fitted to fixed historical data remained tied to past system conditions\. Their smaller advantage in the more stable Flensburg case further supports this interpretation\. Although the results do not establish causality, they indicate that rolling\-context inference may be particularly beneficial for district\-heating networks undergoing structural or operational change\. In such settings, time\-series foundation models with in\-context learning may offer a conceptual advantage over conventional approaches requiring explicit retraining to adapt to concept drift\.

The experiment with seasonally relevant context should be interpreted cautiously\. As described in Section[5\.1](https://arxiv.org/html/2608.20024#S5.SS1), adding calendar\-aligned observations from the previous year did not improve forecasts\. However, the 2023 and 2024 data span the ongoing expansion of the Munich district and associated demand growth\. The added block may therefore represent a different system configuration\. Calendar alignment also only approximates physical similarity; selecting context by ambient temperature or clustered operating conditions may be more meaningful, although nonstationarity would remain challenging\. Seasonally relevant context may thus be more beneficial in stationary district\-heating systems\.

The full\-year experiments in Section[5\.3](https://arxiv.org/html/2608.20024#S5.SS3)used realized future temperature and thus assume perfect weather forecasts\. As expected, replacing realized temperatures with archived 24\-hour\-ahead forecasts reduced deterministic and probabilistic forecast quality, but the forecasts remained usable at the tested 24\-hour horizon\. Because the sensitivity analysis covered only this horizon and weather\-forecast uncertainty generally increases with lead time, larger effects may occur for week\-ahead forecasting or other longer\-horizon operational planning tasks\.

The MRRC Forecaster discussed in Section[4\.3](https://arxiv.org/html/2608.20024#S4.SS3)should be viewed as an architecture derived from the diagnostics, not as a universally superior model\. It decomposes demand into a slowly varying daily component and a short\-term high\-resolution residual, potentially reflecting the slower, weather\-driven dynamics of space heating and the faster, more stochastic dynamics of domestic hot water\. Although it does not explicitly separate these end uses, it represents their characteristic time scales separately\. For TabPFN\-TS, the MRRC Forecaster improved the 12\-hour energy metric and reduced computational cost relative to the direct High Frequency & High Resolution setup while maintaining similar short\-term trajectory accuracy\. The High Frequency & High Resolution setup issues a forecast using a 12\-week context window and a 13\-hour horizon every hour\. The MRRC Forecaster instead issues a 24\-hour forecast with a 12\-week context window only every 12 hours and supplements it with hourly updates using a 7\-day context window and a 2\-hour horizon\. By reducing expensive long\-context forecasts, it lowers operational prediction time, which may matter on less powerful hardware\. Its benefit therefore depends on whether operational priorities favor short\-term trajectory accuracy, longer\-horizon energy planning, or lower computational cost\.

## 7Conclusion & Future Work

This work investigated whether TabPFN\-TS is suitable for zero\-shot probabilistic heat load forecasting in district heating networks\. The results indicate that TabPFN\-TS is a suitable and competitive forecasting approach, although Chronos\-2 achieved the strongest overall accuracy in the reported benchmarks\. A robust configuration for TabPFN\-TS was obtained with hourly 24\-hour forecasts, a 12\-week rolling context window, and ambient temperature as the weather covariate\.

The experiments show that rolling\-context foundation\-model forecasting can perform well in an evolving district heating network without repeated task\-specific retraining\. TabPFN\-TS transferred to the Flensburg validation case, provided empirically well\-calibrated prediction intervals, and outperformed the trained AutoGluon baselines in the main full\-year comparison\. At the same time, the diagnostic experiments showed that configuration choices remain important: more context data, higher temporal resolution, and larger weather\-covariate sets did not automatically improve performance\.

The derived Multi\-Resolution Residual\-Correction Forecaster further illustrates how the diagnostic findings can be translated into an operational forecasting architecture\. By combining a longer\-horizon base forecast with a short\-term residual correction, it improved the balance between high\-resolution forecast accuracy, longer\-horizon energy prediction, and computational cost\. Overall, the results suggest that time\-series foundation models, including TabPFN\-TS, are a promising low\-engineering\-effort option for probabilistic heat load forecasting in district heating networks, especially when recent operating conditions are more informative than fixed historical training data\.

The MRRC Forecaster warrants further investigation beyond the Munich network and across different operating conditions\. Its separation of slow load trends and high\-frequency deviations may also transfer to other energy\-system forecasting tasks\. Future work should tune the Residual Forecaster horizon, test multiple weak learners, and compare combinations of time\-series foundation models or classical machine\-learning algorithms as Base and Residual Forecasters\.

A second direction is to adapt foundation models more closely to district\-heating heat load forecasting by preconditioning them on real\-world time series, simulation data, or synthetic data containing the relevant dynamics\. Task\-specific prior data or further weight adaptation may improve forecast accuracy\. However, if target\-network data are included in the prior or used for adaptation, robustness under concept drift must be evaluated separately\.

Future work should extend forecasts beyond heat demand to mass flow rates as well as supply and return temperatures\. The delivered heat rate is proportional to mass flow rate and temperature difference, while the efficiency and usability of sustainable sources such as heat pumps and waste\-heat streams depend on the required temperature level\. Spatially distributed multivariate forecasts within the network could further support decentralized heat\-source integration by providing more detailed information for operational planning and control\.

## CRediT Authorship Contribution Statement

Ben Spoek:Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Data curation, Visualization, Writing – original draft, Project administration\.

Karim K\. Ben Hicham:Conceptualization, Methodology, Software, Validation, Writing – review and editing\.

Kai Derzsi:Data curation, Methodology, Validation, Formal analysis, Writing – review and editing\.

Philipp Althaus:Conceptualization, Supervision, Writing – review and editing\.

Alexander Mitsos:Supervision, Project administration, Funding acquisition, Writing – review and editing\.

Dirk Müller:Supervision, Project administration, Funding acquisition, Writing – review and editing\.

## Funding

We gratefully acknowledge the financial support provided by the Federal Ministry for Economic Affairs and Energy \(BMWE\) under funding code 03EN3118\. We acknowledge support of the Werner Siemens Foundation in the frame of the WSS Research Center “catalaix”\. The funders had no role in the study design, data collection, analysis, interpretation, manuscript preparation, or decision to submit the article\.

## Declaration of Competing Interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper\.

## Code & Data Availability

The code used for the experiments is available on GitHub at[https://github\.com/benspoek/tsfm\-heatload](https://github.com/benspoek/tsfm-heatload)and on Zenodo at[https://zenodo\.org/records/21511738](https://zenodo.org/records/21511738)\. The Munich district heating dataset used in this study is not publicly available due to third\-party restrictions on the underlying operational data\. The Flensburg district heating dataset is a third\-party dataset and is publicly available\.[14](https://arxiv.org/html/2608.20024#bib.bib15)

## Acknowledgments and Computational Resources

Computations were performed with computing resources granted by RWTH Aachen University under project rwth2083\.

## Declaration of Generative AI and AI\-Assisted Technologies in the Manuscript Preparation Process

During the preparation of this work, the authors used OpenAI Codex to assist with code review and editing, manuscript organization, language refinement, and consistency checks\. The authors reviewed and edited the output as needed and take full responsibility for the content of the published article\.

## References

- Aksuet al\.\(2024\)T\. Aksu, G\. Woo, J\. Liu, X\. Liu, C\. Liu, S\. Savarese, C\. Xiong, and D\. SahooGIFT\-Eval: A benchmark for general time series forecasting model evaluation\.arXiv\.Note:arXiv preprint arXiv:2410\.10393External Links:[Document](https://dx.doi.org/10.48550/arXiv.2410.10393),[Link](https://arxiv.org/abs/2410.10393)Cited by:[§2\.2](https://arxiv.org/html/2608.20024#S2.SS2.p3.1.2)\.
- Alkhulaifiet al\.\(2025\)N\. Alkhulaifi, A\. L\. Bowler, D\. Pekaslan, N\. J\. Watson, and I\. TrigueroAutoEnergy: An automated feature engineering algorithm for energy consumption forecasting with AutoML\.Knowl\.\-Based Syst\.329,pp\. 114300\.External Links:ISSN 09507051,[Document](https://dx.doi.org/10.1016/j.knosys.2025.114300)Cited by:[§2\.3](https://arxiv.org/html/2608.20024#S2.SS3.p3.1.1)\.
- Ansariet al\.\(2025\)A\. F\. Ansari, O\. Shchur, J\. Küken, A\. Auer, B\. Han, P\. Mercado, S\. S\. Rangapuram, H\. Shen, L\. Stella, X\. Zhang, M\. Goswami, S\. Kapoor, D\. C\. Maddix, P\. Guerron, T\. Hu, J\. Yin, N\. Erickson, P\. M\. Desai, H\. Wang, H\. Rangwala, G\. Karypis, Y\. Wang, and M\. Bohlke\-SchneiderChronos\-2: From Univariate to Universal Forecasting\.arXiv\.Note:arXiv preprint arXiv:2510\.15821External Links:[Document](https://dx.doi.org/10.48550/arXiv.2510.15821),[Link](https://arxiv.org/abs/2510.15821)Cited by:[§1](https://arxiv.org/html/2608.20024#S1.p3.1.7),[§2\.3](https://arxiv.org/html/2608.20024#S2.SS3.p2.1.3),[§3\.1](https://arxiv.org/html/2608.20024#S3.SS1.p3.1.1)\.
- Biswaset al\.\(2026\)A\. Biswas, M\. Kamal, R\. Krambroeckers, M\. M\. L\. Elahi, S\. Momen, N\. Mohammed, and S\. RahmanOne Step Closer to Ground Truth: A Multi\-Scale Residual\-Aware Representation Learning Pipeline for Predicting Time Series Data\.arXiv\.Note:arXiv preprint arXiv:2606\.10678External Links:[Document](https://dx.doi.org/10.48550/arXiv.2606.10678),[Link](https://arxiv.org/abs/2606.10678)Cited by:[§4\.3](https://arxiv.org/html/2608.20024#S4.SS3.p6.1.1)\.
- Celik and Vanschoren \(2021\)B\. Celik and J\. VanschorenAdaptation Strategies for Automated Machine Learning on Evolving Data\.IEEE Trans\. Pattern Anal\. Mach\. Intell\.43\(9\),pp\. 3067–3078\.External Links:ISSN 1939\-3539,[Document](https://dx.doi.org/10.1109/TPAMI.2021.3062900)Cited by:[§1](https://arxiv.org/html/2608.20024#S1.p2.1.2),[§3\.1](https://arxiv.org/html/2608.20024#S3.SS1.p4.1.1)\.
- Dahlet al\.\(2017\)M\. Dahl, A\. Brun, and G\. B\. AndresenUsing ensemble weather predictions in district heating operation and load forecasting\.Appl\. Energy193,pp\. 455–465\.External Links:[Document](https://dx.doi.org/10.1016/j.apenergy.2017.02.066)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p1.1.4)\.
- Danget al\.\(2022\)L\. M\. Dang, S\. Lee, Y\. Li, C\. Oh, T\. N\. Nguyen, H\. Song, and H\. MoonDaily and seasonal heat usage patterns analysis in heat networks\.Sci\. Rep\.12\(1\),pp\. 9165\.External Links:ISSN 2045\-2322,[Document](https://dx.doi.org/10.1038/s41598-022-13030-6)Cited by:[§1](https://arxiv.org/html/2608.20024#S1.p1.1.8)\.
- Derzsiet al\.\(2026\)K\. Derzsi, J\. Klingebiel, and D\. MüllerConcept drift detection in district heating: A characterization framework and detector\-centric benchmark\.Energy AI25,pp\. 100788\.External Links:ISSN 26665468,[Document](https://dx.doi.org/10.1016/j.egyai.2026.100788)Cited by:[§1](https://arxiv.org/html/2608.20024#S1.p2.1.1)\.
- Deutscher Wetterdienst \(DWD\) \(2026\)Deutscher Wetterdienst \(DWD\)10\-minute station observation data for Munich \(ID 03379\) for the period 2023–2024\.DWD Climate Data Center\.Note:\[dataset\]External Links:[Link](https://opendata.dwd.de/climate_environment/CDC/observations_germany/climate/10_minutes/air_temperature/historical/)Cited by:[§4\.1](https://arxiv.org/html/2608.20024#S4.SS1.p5.1.1)\.
- Dotzauer \(2002\)E\. DotzauerSimple model for prediction of loads in district\-heating systems\.Appl\. Energy73\(3–4\),pp\. 277–284\.External Links:[Document](https://dx.doi.org/10.1016/S0306-2619%2802%2900078-8)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p1.1.2)\.
- European Parliament and Council of the European Union \(2021\)European Parliament and Council of the European UnionRegulation \(EU\) 2021/1119 of the European Parliament and of the Council of 30 June 2021 establishing the framework for achieving climate neutrality and amending Regulations \(EC\) No 401/2009 and \(EU\) 2018/1999 \(‘European Climate Law’\)\.Official Journal of the European Union\.Note:L 243External Links:[Link](https://eur-lex.europa.eu/eli/reg/2021/1119/oj)Cited by:[§1](https://arxiv.org/html/2608.20024#S1.p1.1.1)\.
- Fang and Lahdelma \(2016\)T\. Fang and R\. LahdelmaEvaluation of a multiple linear regression model and SARIMA model in forecasting heat demand for district heating system\.Appl\. Energy179,pp\. 544–552\.External Links:ISSN 0306\-2619,[Document](https://dx.doi.org/10.1016/j.apenergy.2016.06.133)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p1.1.4)\.
- Finkenrathet al\.\(2022\)M\. Finkenrath, T\. Faber, F\. Behrens, and S\. LeiprechtHolistic modelling and optimisation of thermal load forecasting, heat generation and plant dispatch for a district heating network\.Energy250,pp\. 123666\.External Links:[Document](https://dx.doi.org/10.1016/j.energy.2022.123666)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p2.1.2)\.
- Freißmannet al\.\(2025\)J\. Freißmann, M\. Fritz, I\. Tuschy, and Stadtwerke Flensburg GmbHNetwork Data of the District Heating System for the city of Flensburg from 2020–2024\.Zenodo\.Note:\[dataset\]External Links:[Document](https://dx.doi.org/10.5281/zenodo.17177421),[Link](https://doi.org/10.5281/zenodo.17177421)Cited by:[§4\.1](https://arxiv.org/html/2608.20024#S4.SS1.p12.1.2),[Code & Data Availability](https://arxiv.org/html/2608.20024#Sx4.p1.1.1)\.
- Fuenteset al\.\(2018\)E\. Fuentes, L\. Arce, and J\. SalomA review of domestic hot water consumption profiles for application in systems and buildings energy performance analysis\.Renew\. Sustain\. Energy Rev\.81,pp\. 1530–1547\.External Links:ISSN 1364\-0321,[Document](https://dx.doi.org/10.1016/j.rser.2017.05.229)Cited by:[§6](https://arxiv.org/html/2608.20024#S6.p3.1.1)\.
- Gadd and Werner \(2013\)H\. Gadd and S\. WernerDaily heat load variations in Swedish district heating systems\.Appl\. Energy106,pp\. 47–55\.External Links:ISSN 03062619,[Document](https://dx.doi.org/10.1016/j.apenergy.2013.01.030)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p1.1.3)\.
- Gamaet al\.\(2014\)J\. Gama, I\. Žliobaitė, A\. Bifet, M\. Pechenizkiy, and A\. BouchachiaA survey on concept drift adaptation\.ACM Comput\. Surv\.46\(4\),pp\. 1–37\.External Links:ISSN 0360\-0300, 1557\-7341,[Document](https://dx.doi.org/10.1145/2523813)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p3.1.1)\.
- Gonget al\.\(2022\)M\. Gong, Y\. Zhao, J\. Sun, C\. Han, G\. Sun, and B\. YanLoad forecasting of district heating system based on Informer\.Energy253,pp\. 124179\.External Links:[Document](https://dx.doi.org/10.1016/j.energy.2022.124179)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p2.1.9)\.
- Hertelet al\.\(2026\)M\. Hertel, A\. Nikoltchovska, S\. Pütz, R\. Mikut, B\. Schäfer, and V\. HagenmeyerExplainable load forecasting with covariate\-informed time series foundation models\.InProceedings of the 17th ACM International Conference on Future and Sustainable Energy Systems \(e\-Energy ’26\),pp\. 612–626\.External Links:[Document](https://dx.doi.org/10.1145/3744255.3811724)Cited by:[§2\.3](https://arxiv.org/html/2608.20024#S2.SS3.p2.1.1)\.
- Hollmannet al\.\(2023\)N\. Hollmann, S\. Müller, K\. Eggensperger, and F\. HutterTabPFN: a transformer that solves small tabular classification problems in a second\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=cp5PvcI6w8_)Cited by:[§2\.2](https://arxiv.org/html/2608.20024#S2.SS2.p1.1.2),[§2\.2](https://arxiv.org/html/2608.20024#S2.SS2.p2.4.1)\.
- Hollmannet al\.\(2025\)N\. Hollmann, S\. Müller, L\. Purucker, A\. Krishnakumar, M\. Körfer, S\. B\. Hoo, R\. T\. Schirrmeister, and F\. HutterAccurate predictions on small data with a tabular foundation model\.Nature637\(8045\),pp\. 319–326\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-08328-6)Cited by:[§1](https://arxiv.org/html/2608.20024#S1.p3.1.1),[§1](https://arxiv.org/html/2608.20024#S1.p3.1.3),[§2\.2](https://arxiv.org/html/2608.20024#S2.SS2.p1.1.1),[§2\.2](https://arxiv.org/html/2608.20024#S2.SS2.p1.1.2),[§2\.2](https://arxiv.org/html/2608.20024#S2.SS2.p2.4.1),[§6](https://arxiv.org/html/2608.20024#S6.p3.1.5)\.
- Hooet al\.\(2025\)S\. B\. Hoo, S\. Müller, D\. Salinas, and F\. HutterFrom tables to time: Extending TabPFN\-v2 to time series forecasting\.arXiv\.Note:arXiv preprint arXiv:2501\.02945External Links:[Document](https://dx.doi.org/10.48550/arXiv.2501.02945),[Link](https://arxiv.org/abs/2501.02945)Cited by:[§1](https://arxiv.org/html/2608.20024#S1.p3.1.4),[§1](https://arxiv.org/html/2608.20024#S1.p3.1.6),[§2\.2](https://arxiv.org/html/2608.20024#S2.SS2.p3.1.1),[§2\.2](https://arxiv.org/html/2608.20024#S2.SS2.p3.1.3)\.
- Ivankoet al\.\(2021\)D\. Ivanko, Å\. L\. Sørensen, and N\. NordSplitting measurements of the total heat demand in a hotel into domestic hot water and space heating heat use\.Energy219,pp\. 119685\.External Links:ISSN 03605442,[Document](https://dx.doi.org/10.1016/j.energy.2020.119685)Cited by:[§1](https://arxiv.org/html/2608.20024#S1.p1.1.7)\.
- Izadyaret al\.\(2015\)N\. Izadyar, H\. C\. Ong, S\. Shamshirband, H\. Ghadamian, and C\. W\. TongIntelligent forecasting of residential heating demand for the district heating system based on the monthly overall natural gas consumption\.Energy Build\.104,pp\. 208–214\.External Links:[Document](https://dx.doi.org/10.1016/j.enbuild.2015.07.006)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p2.1.4)\.
- Jodeiriet al\.\(2022\)A\.M\. Jodeiri, M\.J\. Goldsworthy, S\. Buffa, and M\. CozziniRole of sustainable heat sources in transition towards fourth generation district heating – A review\.Renew\. Sustain\. Energy Rev\.158,pp\. 112156\.External Links:ISSN 13640321,[Document](https://dx.doi.org/10.1016/j.rser.2022.112156)Cited by:[§1](https://arxiv.org/html/2608.20024#S1.p1.1.3)\.
- Katoet al\.\(2008\)K\. Kato, M\. Sakawa, K\. Ishimaru, S\. Ushiro, and T\. ShibanoHeat load prediction through recurrent neural network in district heating and cooling systems\.In2008 IEEE International Conference on Systems, Man and Cybernetics,pp\. 1401–1406\.External Links:[Document](https://dx.doi.org/10.1109/ICSMC.2008.4811482)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p2.1.7)\.
- Kimet al\.\(2025\)S\. Kim, J\. Lim, J\. Byun, J\. Kim, and H\. KimData\-efficient electricity consumption forecasting with a tabular foundation model\.InProceedings of the 2025 International Conference on Information and Communication Technology Convergence \(ICTC\),pp\. 489–493\.External Links:[Document](https://dx.doi.org/10.1109/ICTC66702.2025.11389094),ISBN 979\-8\-3315\-5678\-5Cited by:[§6](https://arxiv.org/html/2608.20024#S6.p3.1.4)\.
- Kochet al\.\(2026\)F\. Koch, F\. Raisch, and B\. TischlerThermal\-GEMs: Generalized models for building thermal dynamics\.InThe 13th ACM International Conference on Systems for Energy\-Efficient Buildings, Cities, and Transportation \(BuildSys ’26\),Banff, Canada,pp\. 115–125\.External Links:[Document](https://dx.doi.org/10.1145/3744256.3812565)Cited by:[§2\.3](https://arxiv.org/html/2608.20024#S2.SS3.p1.1.2)\.
- Leiriaet al\.\(2023\)D\. Leiria, H\. Johra, A\. Marszal\-Pomianowska, and M\. Z\. PomianowskiA methodology to estimate space heating and domestic hot water energy demand profile in residential buildings from low\-resolution heat meter data\.Energy263,pp\. 125705\.External Links:ISSN 03605442,[Document](https://dx.doi.org/10.1016/j.energy.2022.125705)Cited by:[§1](https://arxiv.org/html/2608.20024#S1.p1.1.5),[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p3.1.4)\.
- Luet al\.\(2019\)J\. Lu, A\. Liu, F\. Dong, F\. Gu, J\. Gama, and G\. ZhangLearning under Concept Drift: A Review\.IEEE Trans\. Knowl\. Data Eng\.31\(12\),pp\. 2346–2363\.External Links:ISSN 1041\-4347, 1558\-2191, 2326\-3865,[Document](https://dx.doi.org/10.1109/TKDE.2018.2876857)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p3.1.2)\.
- Lundet al\.\(2014\)H\. Lund, S\. Werner, R\. Wiltshire, S\. Svendsen, J\. E\. Thorsen, F\. Hvelplund, and B\. V\. Mathiesen4th Generation District Heating \(4GDH\)\.Energy68,pp\. 1–11\.External Links:ISSN 03605442,[Document](https://dx.doi.org/10.1016/j.energy.2014.02.089)Cited by:[§1](https://arxiv.org/html/2608.20024#S1.p1.1.2)\.
- Maltais and Gosselin \(2021\)L\. Maltais and L\. GosselinPredictability analysis of domestic hot water consumption with neural networks: From single units to large residential buildings\.Energy229,pp\. 120658\.External Links:ISSN 0360\-5442,[Document](https://dx.doi.org/10.1016/j.energy.2021.120658)Cited by:[§6](https://arxiv.org/html/2608.20024#S6.p3.1.2)\.
- Meyeret al\.\(2025\)M\. Meyer, D\. R\. Zapata Gonzalez, S\. B\. Kaltenpoth, and O\. MüllerBenchmarking time series foundation models for short\-term household electricity load forecasting\.IEEE Access13,pp\. 218141–218153\.External Links:2410\.09487,[Document](https://dx.doi.org/10.1109/ACCESS.2025.3648056)Cited by:[§2\.3](https://arxiv.org/html/2608.20024#S2.SS3.p1.1.1)\.
- Mülleret al\.\(2022\)S\. Müller, N\. Hollmann, S\. Pineda Arango, J\. Grabocka, and F\. HutterTransformers can do bayesian inference\.InThe Tenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=KSugKcbNf9)Cited by:[§2\.2](https://arxiv.org/html/2608.20024#S2.SS2.p2.1.1)\.
- Nielsen and Madsen \(2006\)H\. A\. Nielsen and H\. MadsenModelling the heat consumption in district heating systems using a grey\-box approach\.Energy Build\.38\(1\),pp\. 63–71\.External Links:[Document](https://dx.doi.org/10.1016/j.enbuild.2005.05.002)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p1.1.2)\.
- Ntakoliaet al\.\(2022\)C\. Ntakolia, A\. Anagnostis, S\. Moustakidis, and N\. KarcaniasMachine learning applied on the district heating and cooling sector: a review\.Energy Syst\.13\(1\),pp\. 1–30\.External Links:ISSN 1868\-3975,[Document](https://dx.doi.org/10.1007/s12667-020-00405-9)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p1.1.1)\.
- Obermeieret al\.\(2026\)M\. Obermeier, M\. Pruckner, F\. Haselbeck, and A\. ZeiselmairFETS benchmark: Foundation models outperform dataset\-specific machine learning in energy time series forecasting\.arXiv\.Note:arXiv preprint arXiv:2604\.22328External Links:[Document](https://dx.doi.org/10.48550/arXiv.2604.22328),[Link](https://arxiv.org/abs/2604.22328)Cited by:[§1](https://arxiv.org/html/2608.20024#S1.p3.1.5),[§2\.3](https://arxiv.org/html/2608.20024#S2.SS3.p2.1.2),[§4\.1](https://arxiv.org/html/2608.20024#S4.SS1.p11.1.1),[§6](https://arxiv.org/html/2608.20024#S6.p1.1.1)\.
- Parket al\.\(2010\)T\. C\. Park, U\. S\. Kim, L\. Kim, B\. W\. Jo, and Y\. K\. YeoHeat consumption forecasting using partial least squares, artificial neural network and support vector regression techniques in district heating systems\.Korean J\. Chem\. Eng\.27\(4\),pp\. 1063–1071\.External Links:[Document](https://dx.doi.org/10.1007/s11814-010-0220-9)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p2.1.1)\.
- Petkovićet al\.\(2015\)D\. Petković, M\. Protić, S\. Shamshirband, S\. Akib, M\. Raos, and D\. MarkovićEvaluation of the most influential parameters of heat load in district heating systems\.Energy Build\.104,pp\. 264–274\.External Links:[Document](https://dx.doi.org/10.1016/j.enbuild.2015.06.074)Cited by:[§6](https://arxiv.org/html/2608.20024#S6.p5.1.1)\.
- Powellet al\.\(2014\)K\. M\. Powell, A\. Sriprasad, W\. J\. Cole, and T\. F\. EdgarHeating, cooling, and electrical load forecasting for a large\-scale district energy system\.Energy74,pp\. 877–885\.External Links:[Document](https://dx.doi.org/10.1016/j.energy.2014.07.064)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p2.1.5)\.
- Quet al\.\(2025\)J\. Qu, D\. Holzmüller, G\. Varoquaux, and M\. Le MorvanTabICL: a tabular foundation model for in\-context learning on large data\.InProceedings of the 42nd International Conference on Machine Learning,A\. Singh, M\. Fazel, D\. Hsu, S\. Lacoste\-Julien, F\. Berkenkamp, T\. Maharaj, K\. Wagstaff, and J\. Zhu \(Eds\.\),Proceedings of Machine Learning Research, Vol\.267,pp\. 50817–50847\.Cited by:[§1](https://arxiv.org/html/2608.20024#S1.p3.1.2),[§1](https://arxiv.org/html/2608.20024#S1.p3.1.3),[§2\.2](https://arxiv.org/html/2608.20024#S2.SS2.p1.1.2)\.
- Ramachandranet al\.\(2026\)A\. Ramachandran, S\. Chatterjee, T\. F\. B\. Neergaard, M\. Oberndoerfer, A\. Maier, and S\. BayerA deep learning framework for heat demand forecasting using time\-frequency representations of decomposed features\.Energy AI24,pp\. 100704\.External Links:2603\.01137,[Document](https://dx.doi.org/10.1016/j.egyai.2026.100704)Cited by:[§2\.3](https://arxiv.org/html/2608.20024#S2.SS3.p4.1.1)\.
- Runge and Saloux \(2023\)J\. Runge and E\. SalouxA comparison of prediction and forecasting artificial intelligence models to estimate the future energy demand in a district heating system\.Energy269,pp\. 126661\.External Links:[Document](https://dx.doi.org/10.1016/j.energy.2023.126661)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p2.1.3)\.
- Shchuret al\.\(2023\)O\. Shchur, A\. C\. Turkmen, N\. Erickson, H\. Shen, A\. Shirkov, T\. Hu, and B\. WangAutoGluon\-TimeSeries: AutoML for probabilistic time series forecasting\.InProceedings of the Second International Conference on Automated Machine Learning,A\. Faust, R\. Garnett, C\. White, F\. Hutter, and J\. R\. Gardner \(Eds\.\),Proceedings of Machine Learning Research, Vol\.224,pp\. 9/1–21\.Cited by:[§3\.1](https://arxiv.org/html/2608.20024#S3.SS1.p4.1.1)\.
- Songet al\.\(2021\)J\. Song, L\. Zhang, G\. Xue, Y\. Ma, S\. Gao, and Q\. JiangPredicting hourly heating load in a district heating system based on a hybrid CNN\-LSTM model\.Energy Build\.243,pp\. 110998\.External Links:[Document](https://dx.doi.org/10.1016/j.enbuild.2021.110998)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p2.1.6)\.
- Stadtwerke Flensburg \(2024\)Stadtwerke FlensburgEinzigartiges Fernwärmenetz: Anschlussquote von über 90 %\.Stadtwerke Flensburg\.External Links:[Link](https://www.stadtwerke-flensburg.de/unternehmen/nachhaltigkeit/transformationsplan/flensburger-fernwaerme)Cited by:[§4\.1](https://arxiv.org/html/2608.20024#S4.SS1.p12.1.1)\.
- Vandermeulenet al\.\(2018\)A\. Vandermeulen, B\. Van Der Heijde, and L\. HelsenControlling district heating and cooling networks to unlock flexibility: A review\.Energy151,pp\. 103–115\.External Links:ISSN 03605442,[Document](https://dx.doi.org/10.1016/j.energy.2018.03.034)Cited by:[§1](https://arxiv.org/html/2608.20024#S1.p1.1.4)\.
- Webbet al\.\(2016\)G\. I\. Webb, R\. Hyde, H\. Cao, H\. L\. Nguyen, and F\. PetitjeanCharacterizing concept drift\.Data Min\. Knowl\. Discov\.30\(4\),pp\. 964–994\.External Links:ISSN 1384\-5810, 1573\-756X,[Document](https://dx.doi.org/10.1007/s10618-015-0448-4)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p3.1.3)\.
- Wojdyga \(2008\)K\. WojdygaAn influence of weather conditions on heat demand in district heating systems\.Energy Build\.40\(11\),pp\. 2009–2014\.External Links:[Document](https://dx.doi.org/10.1016/j.enbuild.2008.05.008)Cited by:[§1](https://arxiv.org/html/2608.20024#S1.p1.1.6)\.
- Xuet al\.\(2009\)B\. Xu, L\. Fu, and H\. DiField investigation on consumer behavior and hydraulic performance of a district heating system in Tianjin, China\.Build\. Environ\.44\(2\),pp\. 249–259\.External Links:ISSN 0360\-1323,[Document](https://dx.doi.org/10.1016/j.buildenv.2008.03.002)Cited by:[§6](https://arxiv.org/html/2608.20024#S6.p3.1.3)\.
- Xueet al\.\(2020\)P\. Xue, C\. Qi, J\. Li, L\. Kong, and X\. SongHeating load prediction based on attention long short term memory: A case study of Xingtai\.Energy203,pp\. 117846\.External Links:[Document](https://dx.doi.org/10.1016/j.energy.2020.117846)Cited by:[§2\.1](https://arxiv.org/html/2608.20024#S2.SS1.p2.1.8)\.
- Zippenfenig \(2023\)P\. ZippenfenigOpen\-meteo\.com weather API\.Zenodo\.Note:\[software\]External Links:[Document](https://dx.doi.org/10.5281/zenodo.7970649),[Link](https://doi.org/10.5281/zenodo.7970649)Cited by:[§3\.1](https://arxiv.org/html/2608.20024#S3.SS1.p8.1.1)\.

## Supplementary Information

This Supplementary Information provides additional details on the development of the Munich district heating network, the selection of representative operating weeks, the daily model\-comparison results, the probabilistic coverage evaluation, the full\-year week\-ahead forecasts, and the Chronos\-2 Multi\-Resolution Residual\-Correction Forecaster comparison\.

## Appendix S1Munich Network Development and Representative Weeks

The Munich district heating network expanded over the available observation period, and its annual heat demand increased accordingly\. Figure[S1](https://arxiv.org/html/2608.20024#A1.F1)shows the annual heat supplied during each complete calendar year in the data set\.

Figure S1:Annual heat provided by the Munich district heating network for complete calendar years\. Values are computed by integrating the 15\-minute heat load signal in kW and converting the result to GWh\.Configuration diagnostics were conducted on complete Monday–Sunday weeks without detected heat load anomalies\. Weeks containing signal outages, defined as three or more consecutive identical heat load values, were excluded\. From the remaining weeks, the coldest week represents winter operation, the hottest week represents summer operation, and the week with the largest within\-week ambient\-temperature variability represents transitional operation\. Table[S1](https://arxiv.org/html/2608.20024#A1.T1)reports the selected weeks and their temperature statistics\.

Table S1:Selected representative weeks from 2024 and their ambient\-temperature statistics\.RegimeMin\. \(∘C\)Mean \(∘C\)Max\. \(∘C\)DatesWinter\-8\.60\-0\.5111\.20Jan\. 15–21Transitional0\.407\.9724\.60Apr\. 22–28Summer16\.8023\.1632\.20Aug\. 12–18
## Appendix S2Extended Daily\-Rank Model Comparison

Figure[S2](https://arxiv.org/html/2608.20024#A2.F2)shows pairwise daily win rates based on CVRMSE for the full\-year Munich and Flensburg evaluations\. Each cell gives the percentage of daily forecasts for which the row model achieved a lower CVRMSE than the column model, with exact ties counted as half a win\.

![Refer to caption](https://arxiv.org/html/2608.20024v1/Supplementary_Figure_S2.png)Figure S2:Pairwise daily CVRMSE win rates for the full\-year Munich and Flensburg evaluations\.
## Appendix S3Full\-Year Week\-Ahead Forecasting

The selected\-week analysis suggested a slight CVRMSE advantage of TabPFN\-TS over Chronos\-2 for weekly 168\-hour forecasts\. To test whether this behavior was specific to the selected weeks, an additional full\-year simulation was conducted for the Munich data set\. The full\-year result in Table[S2](https://arxiv.org/html/2608.20024#A3.T2)follows the main benchmark pattern: Chronos\-2 outperforms TabPFN\-TS on the week\-ahead task, achieving a lower CVRMSE and MAE and a higherR2R^\{2\}\.

Table S2:Full\-year Munich results for non\-overlapping weekly 168\-hour forecasts at hourly resolution\.ModelCVRMSE \(%\)R2R^\{2\}MAE \(kW\)Chronos\-215\.640\.941109\.7TabPFN\-TS17\.060\.929120\.2
## Appendix S4Extended Probabilistic Evaluation

Table[S3](https://arxiv.org/html/2608.20024#A4.T3)reports empirical coverage and interval width for all evaluated central prediction intervals\. The results compare TabPFN\-TS and Chronos\-2 for the full\-year hourly forecasts in both district heating networks\.

Table S3:Empirical coverage of all evaluated central prediction intervals for the full\-year hourly forecasts\. Interval width is reported in kW for both networks\.TabPFN\-TSChronos\-2Interval \(%\)Emp\. cov\.\(%\)Width\(kW\)Rel\. width\(%\)Emp\. cov\.\(%\)Width\(kW\)Rel\. width\(%\)Munich50\.0049\.92145\.113\.9847\.77132\.312\.7560\.0060\.13182\.717\.6057\.76166\.616\.0580\.0080\.50288\.627\.8077\.73260\.325\.0790\.0090\.90389\.737\.5488\.20348\.433\.5695\.0095\.62496\.147\.7995\.63477\.445\.9998\.0098\.21661\.463\.7297\.41554\.953\.45Flensburg50\.0050\.731370211\.3548\.371269810\.5260\.0060\.221724214\.2958\.881595413\.2280\.0079\.822712622\.4878\.192472420\.4990\.0090\.453634030\.1188\.183269127\.0995\.0095\.754577237\.9394\.844379036\.2998\.0098\.286007649\.7896\.705045041\.81
## Appendix S5Chronos\-2 MRRC Forecaster Comparison

The Chronos\-2 comparison uses the same three forecasting setups as the TabPFN\-TS evaluation\. The Base Forecaster issues hourly 24\-hour forecasts with a 12\-week context every 12 hours\. The High Frequency & High Resolution setup issues 13\-hour forecasts at 15\-minute resolution every hour\. The Multi\-Resolution Residual\-Correction \(MRRC\) Forecaster combines the Base Forecaster with hourly residual updates using a 7\-day context and a 2\-hour horizon\.

Table S4:Deterministic full\-year comparison of the Chronos\-2 MRRC Forecaster with the Base Forecaster and High Frequency & High Resolution setups\.Short\-term forecastSetupCVRMSE \(%\)R2R^\{2\}MAE \(kW\)Base Forecaster18\.210\.923121\.2±\\pm54\.7High Frequency & High Resolution17\.720\.927107\.7±\\pm52\.8MRRC Forecaster18\.180\.923109\.0±\\pm54\.1
Long\-term forecast and computationSetup12 hE\-CVRMSE \(%\)12 hE\-bias \(%\)RTFBase Forecaster6\.89\-0\.515\.80×10−75\.80\{\\times\}10^\{\-7\}High Frequency & High Resolution7\.120\.558\.01×10−68\.01\{\\times\}10^\{\-6\}MRRC Forecaster6\.68\-0\.476\.27×10−66\.27\{\\times\}10^\{\-6\}

The Chronos\-2 results point in a similar direction to the TabPFN\-TS results, although the short\-term trajectory does not improve with the MRRC Forecaster\. The High Frequency & High Resolution setup achieves the lowest short\-term CVRMSE and MAE\. Relative to this direct setup, the MRRC Forecaster increases short\-term CVRMSE by 2\.6%, while reducing the 12\-hour E\-CVRMSE by 6\.2%, the bias magnitude by 14\.5%, and the RTF by 21\.7%\. The MRRC Forecaster remains 10\.8 times more computationally expensive than the Base Forecaster alone\.

Similar Articles

A Benchmark for Electrical Load Forecasting Across Grid Levels: Time-Series Transformers Outperform Established Methods

arXiv cs.LG

This paper presents a comprehensive benchmark for electrical load forecasting across grid levels, evaluating ten methods and finding that Transformer-based approaches consistently outperform established methods, reducing forecast error by 6.6–10.7%. The standard Transformer achieves superior performance over a novel flexible architecture, and the foundation model Chronos-2 shows competitive zero-shot performance on some datasets.

Unified Zero-Shot Time Series Forecasting: A Darts Foundation

arXiv cs.LG

Darts, a popular open-source Python library for time series analysis, introduces a unified FoundationModel class collection that integrates multiple time series foundation models (Chronos-2, TimesFM 2.5, TiRex, PatchTST-FM) for zero-shot and fine-tuned forecasting with standardized interfaces and minimal dependencies.

TabPFN-3: Technical Report

arXiv cs.LG

TabPFN-3 is a new foundation model for tabular data, pretrained on synthetic data, that scales to 1M training rows while reducing training and inference time, achieving state-of-the-art performance on tabular prediction, time series, and relational data.