EXAONE Demand 1.0:一个用于需求预测的时间序列基础模型
摘要
EXAONE Demand 1.0 是一个专为需求预测设计的时间序列基础模型,使用需求特定的语料库和适配器,在22个数据集上超越现有模型。
查看缓存全文
缓存时间: 2026/09/28 09:51
# EXAONE Demand 1.0:A Time Series Foundation Model for Demand Forecasting
Source: [https://arxiv.org/html/2609.30880](https://arxiv.org/html/2609.30880)
LG AI Research\*Affiliation:Seunghan Lee, Sangjun Han, Jun Seo, Junhyeok Kang, Jaehoon Lee, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Minjae Kim, Sungdong Yoo, Soonyoung Lee, Wonbin Ahn†
Time series foundation models \(TSFMs\) are pretrained on series from diverse domains, where demand series make up only a small fraction\. Demand data has properties that such corpora rarely contain: Short histories, frequent zeros, censoring by stock\-outs, and exogenous events that the series does not record\. To this end, we proposeEXAONE Demand, built on 1\) ademand\-specific corpusand 2\) ademand\-aware adapter\. For the corpus, we assemble 11\.3M series and 48\.4B observations from 73 sources, and asynthetic generatorsupplies the behaviour that open demand data under\-represents\. For the adapter, we attach low\-rank branches to a frozen general\-domain backbone, one for each of the four demand classes \(smooth, intermittent, erratic, and lumpy\), and a router that reads eight scale\-free statistics of the input series decides how much each branch contributes\. We build EXAONE Demand in two versions, one trained on real\-world and synthetic demand together and one trained on the synthetic corpus alone\. On 22 held\-out datasets, both versions outperform 36 TSFMs, and real\-world demand adds a gain over synthetic data alone\.Model[https://huggingface\.co/LG\-AI\-Research/EXAONE\-Demand\-1\.0](https://huggingface.co/LG-AI-Research/EXAONE-Demand-1.0) Code[https://github\.com/LGAI\-Research/EXAONE\-Forecast](https://github.com/LGAI-Research/EXAONE-Forecast)
†††Corresponding author\.## 1Introduction
Demand forecasting predicts how much of a product, a service, or a resource will be requested over a future period, and that prediction decides what an organisation stocks, staffs, and ships\[[51](https://arxiv.org/html/2609.30880#bib.bib51)\]\. Its two directions of error carry different costs, whereunder\-forecastingleaves demand unmet andover\-forecastingcommits resources that are never used\. Demand behaves unlike thesmoothandlongseries that dominate general forecasting benchmarks, and it is therefore studied as a problem of its own, with its own error measures and its own taxonomy of series\[[66](https://arxiv.org/html/2609.30880#bib.bib66)\]\.
Time series foundation models \(TSFMs\) are pretrained at scale to forecast across domains\[[4](https://arxiv.org/html/2609.30880#bib.bib4),[72](https://arxiv.org/html/2609.30880#bib.bib72),[17](https://arxiv.org/html/2609.30880#bib.bib17),[26](https://arxiv.org/html/2609.30880#bib.bib26)\], and they reach strong zero\-shot accuracy on general benchmarks\[[25](https://arxiv.org/html/2609.30880#bib.bib25),[2](https://arxiv.org/html/2609.30880#bib.bib2)\]\. Their corpora are assembled forbreadthrather than fordemand, and the properties that make demand hard are diluted by the far more numerous other series\. We argue thatdemand forecasting needs a corpus of its own, one that reflects these properties instead of averaging them out\.
Three properties of demand data\.Three properties recur across our demand series and are rare in general corpora, and Figure[2](https://arxiv.org/html/2609.30880#S1.F2)shows all three in eight series, one drawn from each of eight open sources of the corpus\.
- •Intermittency\.A large share of series arezeromost of the time, broken by irregular non\-zero periods, which theaverage demand interval\(ADI\) and thesquared coefficient of variation\(CV2\) characterise\[[66](https://arxiv.org/html/2609.30880#bib.bib66)\]\.
- •Shortness\.Demand is recorded at the granularity at which it isordered, which is weekly or monthly per SKU and per store, leaving individual series far shorter than the ones general benchmarks are built on\.
- •Exogenous and censored structure\.Promotions, holidays, and life\-cycles drive level shifts whose cause isabsent from the series, and observed sales are acensoredview of latent demand whenever inventory runs out\.
Figure 1:Representative real\-world demand series\.One series is drawn per source from the corpus\. Retail and SKU\-level series are dominated by zeros and irregular spikes, while energy and transport series are long with strong daily and weekly cycles\.Figure 2:Performance on demand datasets\.
A demand corpus does not hold one kind of series but four, wheresmooth,intermittent,erratic, andlumpydemand arrive mixed together\. They ask different questions of a forecaster, whereintermittencyturns the question intowhendemand arrives, whileerraticandlumpybehaviour turns it intohow mucharrives at once\. A corpus that holds all four is not enough by itself, since one set of weights fitted to the mixture has to serve both questions, and moving it toward one behaviour costs accuracy on the other, a trade\-off that a single adapter has no means to resolve\.
Figure 3:Overview of EXAONE Demand\.To this end, we proposeEXAONE Demand, a TSFM for demand forecasting thatspecialises by demand behaviour\. As Figure[3](https://arxiv.org/html/2609.30880#S1.F3)lays out, it rests on two components: 1\) Ademand\-specific corpusbuilt from real\-world demand sources and a synthetic generator, and 2\) ademand\-aware adapterthat mixes low\-rank branches by the behaviour of the input series\. Specifically, we adapt a general\-domain TSFM pretrained on KernelSynth\[[4](https://arxiv.org/html/2609.30880#bib.bib4)\]alone, namely synthetic series drawn from Gaussian process kernels, keeping its weights frozen and training only five low\-rank branches, one shared by every series and four tied to the demand classes, where arouterdecides how much each of the four contributes\. We build EXAONE Demand in two versions, one trained on real\-world and synthetic demand together and one trained on the synthetic corpus alone\. As shown in Figure[2](https://arxiv.org/html/2609.30880#S1.F2), both versions outperform the state\-of\-the\-art \(SoTA\) TSFMs on 22 held\-out demand datasets, with real\-world demand adding a gain over synthetic data\.
## 2Construction of Demand Dataset
We construct the demand dataset in the following three steps, as illustrated in Figure[4](https://arxiv.org/html/2609.30880#S2.F4):
- •Collection\(§[2\.1](https://arxiv.org/html/2609.30880#S2.SS1)\): We gather various sources and keep only the series that measure demand\.
- •Preprocessing\(§[2\.2](https://arxiv.org/html/2609.30880#S2.SS2)\): We convert every source into a common schema and index every series we keep\.
- •Train\-test split\(§[2\.3](https://arxiv.org/html/2609.30880#S2.SS3)\): We assign sources by their published role and remove the training series that leak\.
A synthetic corpus is built in parallel, and the real and synthetic corpora stay apart until the last step, where every synthetic series is used for training\. Section[3](https://arxiv.org/html/2609.30880#S3)describes the design and the generating process behind it\.
Figure 4:Pipeline for demand dataset construction\.1\) Collection:Real sources are grouped by their published role, and the synthetic corpus forms a group of its own below\.2\) Preprocessing:The two corpora run on separate tracks\. The real track filters, converts, and verifies each source before removing leaked series, while the synthetic track generates, validates, and labels its own\.3\) Train\-test split:Of the five roles, only the evaluation suites reach the test side, and the other four are used for training, together with every one of the synthetic series\.Table 1:Collected sources by domain\.The count is over the 93 collected sources, and the last two rows are source types rather than domains, since each of them holds many datasets under a single published name\.DomainSourcesRepresentative contentRetail and e\-commerce24Store and SKU sales, grocery baskets, fashion and online retailEnergy23Household and industrial electricity load, national demandTransport14Taxi and ride\-hailing pickups, bike share, transit ridership, port callsHealthcare6Prescription dispensing at practice, pharmacy, and hospital levelService and hospitality4Restaurant covers, hotel and booking demandTelecommunications3Cell\-grid trafficEconomics3Retail sales indices, industrial production, trade volumesWeb and compute2Page requests, virtual\-machine resource usageGeneral\-purpose corpora7Time\-300B, GIFT\-Eval, LOTSA, Chronos, fev\-bench, MonashCompetition benchmarks7M4, M5, VN1, Rossmann, Favorita, and related### 2\.1Collection
The first decision is notwhereto find data butwhatto accept as demand\. This cannot be taken from the data, as nothing in a published file marks a series as demand and no metadata field records it\. We therefore state a definition of our own, taking a demand series to be a measurement of*consumption, sales, or utilisation per unit of time by an identifiable entity*\. What the definition rejects matters as much as what it keeps, since the rejected series are often the ones that look most like demand\. Three rejections came up in almost every source, and we state them as rules:
- •1\) Supply is not demand\.We keep a series only when it measures what was taken, not what was produced\. Solar and wind generation are therefore excluded, and so is theexportreading of a residential smart meter, while theimportreading of the same meter is kept, since it records what the household actually drew\.
- •2\) Physical quantities are not demand\.Transformer oil temperature, air quality, river flow, and weather measure the state of a system rather than anyone’s use of it\. Several of them are standard forecasting benchmarks, and they therefore reach us inside the broad collections described below rather than as sources of their own\.
- •3\) Inventory is not demand\.The number of bikes available at a docking station is supply, and the number of trips starting there is demand\. The same file often holds both, and we take only the second\.
Two cases are not decided by the rules above, since the question there is not whether a series is demand but what a single number in it stands for\. \(1\)Sources with no series:Event logs such as taxi trips, bike rentals, and municipal service requests store one row per event\. They become demand only after the events are counted per entity and per time step, which makes those two settings the definition of demand for that source\. \(2\)Sources with several candidate values:In pharmaceutical data we take prescription*items*rather than*quantity*, since quantity is measured in tablets for one drug and in millilitres for another, and the two cannot be summed on a common scale\.
Sources\.Demand data is published in three types of sources, and relying on a single type yields a biased corpus, as benchmark suites and public data archives provide almost disjoint material\. We therefore collect from all three:
- •Type 1\) Forecasting benchmarks and competitions\.M4, M5, VN1, Rossmann, Favorita, and related retail competitions, which define the conventional evaluation of demand forecasting in the literature\.
- •Type 2\) General\-purpose corpora\.Broad collections assembled by other groups to cover time series in general, from which we keep only the subsets that are relevant to demand and discard the rest\.
- •Type 3\) Domain\-specific open data\.National health services, transport authorities, energy regulators, and statistical agencies, which publish demand directly but are absent from forecasting benchmarks\.
We collect 93 sources in total, of which 73 are converted to a common schema\. The remaining 20 are excluded because 1\) their contents are not tabular, 2\) they duplicate sources already collected, or 3\) they require credentials we do not hold\. Table[1](https://arxiv.org/html/2609.30880#S2.T1)summarises the collection by domain, and Figure[6](https://arxiv.org/html/2609.30880#S2.F6)shows the distribution of the converted corpus\.
Published roles\.Each source is released with a role determined by its publisher, which we record rather than assign\. We identify five roles: 1\)Pretrain only, corpora with no held\-out part, 2\)No split, releases published as a single undivided file, 3\)Train only, releases that publish a training file while keeping the held\-out part private, 4\)Train and test, releases that ship a fixed train and test pair, and 5\)Evaluation, suites published to be scored against\. Table[2](https://arxiv.org/html/2609.30880#S2.T2)lists the sources in each role, and §[2\.3](https://arxiv.org/html/2609.30880#S2.SS3)describes how each role is used in training and in evaluation\.
Table 2:Collected sources by published role\.The role is a property of the release rather than a decision of ours, and what a source finally contributes is settled by the leakage check rather than by the role alone\.Published roleSourcesSeriesSources in this groupPretrain only88,542,433Time\-300B\[[63](https://arxiv.org/html/2609.30880#bib.bib63)\], GIFT\-Eval pretrain\[[2](https://arxiv.org/html/2609.30880#bib.bib2)\], Chronos\[[4](https://arxiv.org/html/2609.30880#bib.bib4)\], LOTSA\[[72](https://arxiv.org/html/2609.30880#bib.bib72)\], EIA\-930\[[68](https://arxiv.org/html/2609.30880#bib.bib68)\], IMF PortWatch\[[32](https://arxiv.org/html/2609.30880#bib.bib32)\], NYC TLC\[[53](https://arxiv.org/html/2609.30880#bib.bib53)\], NYC 311\[[56](https://arxiv.org/html/2609.30880#bib.bib56)\]No split67999,102Goiener smart meters\[[28](https://arxiv.org/html/2609.30880#bib.bib28)\], UCI electricity load\[[67](https://arxiv.org/html/2609.30880#bib.bib67)\], Citi Bike\[[12](https://arxiv.org/html/2609.30880#bib.bib12)\], Divvy\[[19](https://arxiv.org/html/2609.30880#bib.bib19)\], NHS dispensing\[[54](https://arxiv.org/html/2609.30880#bib.bib54)\], Online Retail II\[[9](https://arxiv.org/html/2609.30880#bib.bib9)\], OPSD\[[57](https://arxiv.org/html/2609.30880#bib.bib57)\], AEMO\[[6](https://arxiv.org/html/2609.30880#bib.bib6)\], and 59 othersTrain only\(Held\-out private\)3815,098Web Traffic\[[35](https://arxiv.org/html/2609.30880#bib.bib35)\], Instacart\[[64](https://arxiv.org/html/2609.30880#bib.bib64)\], Kaggle sales forecasting\[[59](https://arxiv.org/html/2609.30880#bib.bib59)\]Train and test1150,489M4\[[50](https://arxiv.org/html/2609.30880#bib.bib50)\], M5\[[51](https://arxiv.org/html/2609.30880#bib.bib51)\], Favorita\[[36](https://arxiv.org/html/2609.30880#bib.bib36)\], VN1\[[69](https://arxiv.org/html/2609.30880#bib.bib69)\], Rossmann\[[34](https://arxiv.org/html/2609.30880#bib.bib34)\], Walmart sales forecast\[[1](https://arxiv.org/html/2609.30880#bib.bib1)\], Walmart recruiting\[[33](https://arxiv.org/html/2609.30880#bib.bib33)\], Store Item Demand\[[37](https://arxiv.org/html/2609.30880#bib.bib37)\], Panama load \(two releases\)\[[49](https://arxiv.org/html/2609.30880#bib.bib49)\], Food demand\[[38](https://arxiv.org/html/2609.30880#bib.bib38)\]Evaluation3899,618GIFT\-Eval\[[2](https://arxiv.org/html/2609.30880#bib.bib2)\], fev\-bench\[[62](https://arxiv.org/html/2609.30880#bib.bib62)\], BOOM\[[18](https://arxiv.org/html/2609.30880#bib.bib18)\]
### 2\.2Preprocessing
Preprocessing brings sources of all three types into a single schema\. The demand filter of §[2\.1](https://arxiv.org/html/2609.30880#S2.SS1)applies to all of them, but at a different level in each\. For competitions and domain\-specific open data \(types 1 and 3\) the source as a whole is demand, and the filter only decides which columns to take\. A general\-purpose corpus \(type 2\) holds demand and non\-demand datasets side by side, and there the filter removes whole datasets\. Of the seven general\-purpose corpora in Table[1](https://arxiv.org/html/2609.30880#S2.T1), six are converted, since Monash is already contained in the Chronos and LOTSA collections, and the filter keeps 269 of the 343 datasets they contain and drops the remaining 74\. This large drop rate is expected, as these corpora are assembled to covertime seriesrather thandemand, and the dropped datasets are mostly weather, generation, and air quality, which rules 1\) and 2\) of the demand filter in §[2\.1](https://arxiv.org/html/2609.30880#S2.SS1)already exclude\.
General\-purpose corpora\.These six are not single datasets but collections, each one containing many datasets under a single name, and the corpus itself therefore has no single domain\. Four were published for pretraining \(Time\-300B\[[63](https://arxiv.org/html/2609.30880#bib.bib63)\], the GIFT\-Eval pretraining split\[[2](https://arxiv.org/html/2609.30880#bib.bib2)\], the Chronos training collection\[[4](https://arxiv.org/html/2609.30880#bib.bib4)\], LOTSA\[[72](https://arxiv.org/html/2609.30880#bib.bib72)\]\) and two as evaluation suites \(fev\-bench\[[62](https://arxiv.org/html/2609.30880#bib.bib62)\], GIFT\-Eval\[[2](https://arxiv.org/html/2609.30880#bib.bib2)\]\)\. We therefore assign a domain to each dataset inside them rather than to the corpus as a whole\. These six hold 83% of all series, and this one step therefore fixes the domain of most of the corpus\.
Conversion and verification\.All sources are converted to a single schema of\(series\_id, timestamp, target\), together with a per\-series index recording the shape, demand character, and provenance of each series\. The converted counts are then verified against numbers establishedoutsidethe pipeline, such as competition documentation or the published number of entities, rather than against the conversion log itself\. Appendix[C](https://arxiv.org/html/2609.30880#A3)records the schema in full and the four quantities that every source is required to agree with before it is admitted to the corpus\.
\[1\] QuantityValue\[2\] Series lengthPointsSeries11,306,740P25128Observations48,355,437,247P501,373Converted sources73P758,761Collected sources93P9921,383\[3\] ClassSeries\[4\] Zero ratioShareSmooth6,340,416P500\.00Erratic4,141,168P900\.06Intermittent727,598P990\.98Lumpy97,558P1001\.00Table 3:Statistics of demand datasets\.Counts and percentiles are over the whole converted corpus, real\-world and synthetic together\.
Figure 5:Zeros in the corpus\.Zeros are absent from most series, and intermittency is therefore concentrated in a minority\.
Figure 6:Corpus composition by series count\.Every series carries one domain of the converted corpus\.Corpus statistics\.The converted corpus contains 11\.3M series and 48\.4B observations from 73 sources\. Aggregate size alone does not indicate whether a corpus is suited to demand, as a small number of long high\-frequency series can dominate any total\. Table[3](https://arxiv.org/html/2609.30880#S2.T3)therefore reports theshapeof the corpus instead, namely how long its series are, how they split across the four demand classes, and what share of their values is zero at each percentile\.
Two characteristics of the corpus affect all downstream design decisions\. First, the length distribution is heavily skewed, where the median series has 1,373 points while the lower quartile has only 128\. Second, intermittency is concentrated, where zeros are absent from the median series but dominate the upper percentiles, as shown in Figure[5](https://arxiv.org/html/2609.30880#S2.F5)\. The intermittent and lumpy classes, the hard cases that motivate the corpus, are therefore a minority by count, and the generator of Section[3](https://arxiv.org/html/2609.30880#S3)is built to correct this imbalance by generating more series of the two rare classes\.
### 2\.3Train\-Test Split
Roles and their use\.The five roles in Table[2](https://arxiv.org/html/2609.30880#S2.T2)reduce to a single rule, namely that only evaluation suites are held out, while all other roles are training candidates until the leakage check below removes whatever overlaps the test side\.
- •Evaluationis the only role held out, as these three sources exist to be scored against, and using any part of them for training would invalidate the very measurement that they were published to provide\.
- •Pretrain onlyandNo splitare used for training in full, as neither withholds anything we are obliged to respect\. What is removed from them is determined by the leakage check below rather than by the role\.
- •Train only \(held\-out private\)is also used for training in full, since the hidden part was never published by its authors, and there is therefore no test set of our own that any part of it could contaminate\.
- •Train and testcontributesbothparts to training, as its test part is the competition’s own rather than the benchmark we evaluate on, and any coincidence between the two is caught by the leakage check below\.
This rule concerns how a source was published rather than what it contains, and it is therefore not sufficient on its own, since two sources can publish the same underlying series under different roles without any record of the overlap\.
Train and test separation\.Evaluation uses the held\-out evaluation suites, and all remaining series are available for training\. Separating them by dataset name is not sufficient, as the same underlying series is redistributed under different names across sources\. We therefore separate the two sides at the level of the values themselves, where a redistributed copy cannot hide behind the name it was republished under\. Any training candidate whose values match a test series isremoved rather than down\-weighted\. Table[4](https://arxiv.org/html/2609.30880#S2.T4)reports what the check leaves on each side, counted in series and in time points, together with what it removes from the corpus on the training side\.
Table 4:Train and test separation\.Removed is what the leakage check takes out of training, and the shares are over the whole corpus\.SeriesTime pointsValueShareValueShareTrain10,049,11488\.88%47,419,572,86298\.06%Test899,6187\.96%810,403,2801\.68%Removed358,0083\.17%125,461,1050\.26%We apply the same measurement to the pretraining data of the models we compare against, as discussed in §[5\.1](https://arxiv.org/html/2609.30880#S5.SS1)\. That data was fixed before the benchmarks and was never filtered against them, and the comparison is thereforenot symmetric, which we state rather than leave implicit\. The synthetic corpus is untouched by this check, since it holds no real observation that the leakage check could match, and every one of its series is available to the training mixture in full\.
## 3Synthetic Demand Dataset
Table 5:Design of the synthetic generator by demand properties\.Each property has a component of its own in the sampling process, and each component is turned on or off for every series with its strength drawn at random, so that the properties combine freely rather than in the fixed groups that a fixed recipe would give\.FamilyPropertyGenerating mechanismCountsDiscretenessNegative\-binomial sampling from a multiplicative intensityOverdispersionDispersion parameter sampled per seriesIntermittencyBernoulli mask over the intensity, targeted to a demand classCalendarDay\-of\-week profileMultiplicative weekday factorsSeasonalityFourier components at annual and weekly periodsMoving holidaysLunar New Year, Chuseok, Easter, and Thanksgiving as shifting datesPromotionUpliftThree\-phase response of build\-up, peak, and post\-promotion dipCannibalisationNegative cross\-effect on related seriesPrice elasticityLog\-linear response of intensity to a sampled price pathLife\-cycleProduct launchBass diffusion for adoptionDiscontinuationDecaying tail with an absorbing zero stateObservationStock\-out censoring\(s,S\)\(s,S\)inventory policy clipping observed sales below latent demandReporting gapsContiguous missing spansThe real\-world corpus of Section[2](https://arxiv.org/html/2609.30880#S2)is broad but uneven, since several of the behaviours that decide whether a demand forecast is useful are thin in it, and one of them cannot appear in observed data at all\. We therefore generate a synthetic corpus alongside the real one, designed to cover exactly those regions of demand behaviour\.
### 3\.1Motivation
Why synthetic data\.Open demand data under\-represents several behaviours that matter operationally, and in some cases cannot represent them at all\. Promotions and holidays are visible in the series only as unexplained level shifts\. The calendar and the promotional plan that caused them are not published with the data\. Product life\-cycles are truncated, since a dataset released at one point in time rarely covers both the launch and the discontinuation of the same item\. Most importantly, observed sales are a censored view of latent demand\. When inventory runs out, a zero records an empty shelf rather than an absence of customers, and no real dataset distinguishes the two\.
We therefore generate a synthetic corpus of 47,500 series and 59\.7M points alongside the real one\. Its purpose is not to substitute for real demand but tocover the regions of demand behaviour that open data leaves thin\. The generating process is known, and it therefore records what observation alone cannot, as described in §[3\.2](https://arxiv.org/html/2609.30880#S3.SS2)\.
### 3\.2Generating Process
Generating process\.A series is built in three stages, where a multiplicative intensity sets the expected level over time, a count process turns that intensity into integer demand, and an observation model decides what is actually recorded\. Table[5](https://arxiv.org/html/2609.30880#S3.T5)lists the properties the generator reproduces and the mechanism behind each\. Every property is switched on independently with a sampled strength, which lets the corpus cover smooth seasonal demand and series dominated by long zero runs and abrupt life\-cycle transitions, together with the mixtures that lie between the two extremes\. The corpus is organised into eight generators, each configured to imitate a distinct operational setting rather than a distinct statistical family\. This keeps the corpus interpretable, since a mixture can be described in terms of the business situations it emphasises\. Figure[7](https://arxiv.org/html/2609.30880#S3.F7)shows one series from each of the eight generators below\.
The corpus records the pre\-censoring latent demand, the underlying intensity, and a stock\-out flag next to every observed value\. This is a signalno real dataset can provide, since in observed sales data a zero caused by no customer and a zero caused by an empty shelf are indistinguishable in the file as the publisher releases it\.
Figure 7:Representative synthetic demand series\.One series per generator\. Spare parts and B2B orders realise the intermittent and lumpy regimes that open data supplies only sparsely, new product realises a full launch\-to\-discontinuation life\-cycle, and e\-commerce carries explicit promotional responses of its own\.
### 3\.3Validation and Use
Validation\.Generation is controlled rather than left to an unconstrained sampler, as the Syntetos–Boylan class is an*input*to the generator, which produces series to hit a target class, and the realised classification is then checked against that target, where the two agree in 95\.5% of cases\. Seven further validation checks, covering class distribution, zero ratios, seasonality recovery, promotion response, censoring rate, length distribution, and determinism under a fixed seed, all pass on the final corpus, which confirms that the properties that the generator was written to produce are the properties the corpus actually carries rather than ones merely assumed from the generating process\.
Use in training\.The synthetic corpus contains no real observations by construction, where no value of it can match an evaluation series, and the leakage check therefore does not apply to it\. It is therefore used in its entirety for training, with no held\-out portion\. All 47,500 series are available to every mixture, and the train and test separation reported in §[2\.3](https://arxiv.org/html/2609.30880#S2.SS3)concerns the real\-world corpus alone, where a leaked series would otherwise be scored against itself\.
## 4EXAONE\-Demand
Figure 8:EXAONE Demand on the four demand classes\.\(1\)Raw context: The window that the model receives, with no rescaling of its values\. \(2\)8 Statistics: The scale\-free summary that the window is reduced to\. \(3\)Router: One weight per routed branch, with the always\-on shared branch added at a fixed weight\.Motivation\.Demand series do not share one behaviour, since the corpus of Section[2](https://arxiv.org/html/2609.30880#S2)spanssmoothseasonal sales,intermittentspare\-part orders,erraticpromotional spikes, andlumpywholesale batches\. The hard cases are a minority by count while being the reason the corpus exists, where one adapter fitted to all four spends most of its capacity on the easy majority, and we therefore let it specialise by demand behaviour rather than average across them\.
Overview of EXAONE Demand\.EXAONE Demand keeps a pretrained TSFM frozen and adapts it with a mixture\-of\-experts \(MoE\) adapter\[[61](https://arxiv.org/html/2609.30880#bib.bib61),[74](https://arxiv.org/html/2609.30880#bib.bib74)\], whose five experts are low\-rank branches and whose mixing weights follow the demand behaviour of the input series\. Four of them are routed, each tied to one of the four demand classes, and the fifth is always on and shared by every series as a shared expert\[[15](https://arxiv.org/html/2609.30880#bib.bib15)\], so that it carries the correction that every demand series needs and leaves the routed four to carry only the differences between the classes\.
The pipeline has five stages, and we describe each in turn:
- •Demand statistics\(§[4\.1](https://arxiv.org/html/2609.30880#S4.SS1)\): We summarise the raw context window into eight scale\-free statistics\.
- •Continuous demand membership\(§[4\.2](https://arxiv.org/html/2609.30880#S4.SS2)\): We turn ADI and CV2into a membership over the four demand classes\.
- •Routing\(§[4\.3](https://arxiv.org/html/2609.30880#S4.SS3)\): We route on the eight statistics and supervise the router with that membership\.
- •Mixture of low\-rank branches\(§[4\.4](https://arxiv.org/html/2609.30880#S4.SS4)\): We mix the shared and the four routed branches in every frozen projection\.
- •Training\(§[4\.5](https://arxiv.org/html/2609.30880#S4.SS5)\): We train only the low\-rank branches and the router, keeping the backbone frozen\.
Figure[8](https://arxiv.org/html/2609.30880#S4.F8)draws the whole pipeline, with four real series, one from each of the four demand classes, and the statistics and the membership that each of them implies at the input of the adapter\.
### 4\.1Stage 1: Demand Statistics
The router reads eight statistics of the raw series, and each of them answers one question about the demand behaviour that separates the four classes\. We list them in the order they enter the router input:
- \[1\]Zero share\(Zero\): The fraction of the window with no demand, which is the plainest sign of intermittency\.
- \[2\]Average inter\-demand interval\(ADI\): The mean spacing between two non\-zero periods, which is the first axis of the Syntetos–Boylan classification, separating smooth from intermittent demand\.
- \[3\]Squared coefficient of variation\(CV2\): The dispersion of the non\-zero values around their own mean, which is the second axis of the same scheme, separating steady from erratic order sizes\.
- \[4\]Trend correlation\(Trend\): The correlation between the series and time, which distinguishes a product in its launch or in its decline phase from one that has held a stable level\.
- \[5\]First\-order autocorrelation\(AC\(1\)\): How much one period predicts the next, which tells a seasonal or otherwise persistent series apart from one whose movements are close to noise\.
- \[6\]Fraction of upward steps\(Up\): The share of consecutive pairs that increase, which describes the shape of the path the demand takes rather than its level or its spread\.
- \[7\]Coefficient of variation\(CV\): The dispersion of the whole window, zeros included, which reacts to the spikes that the dispersion of the non\-zero values alone would miss\.
- \[8\]Length\(Length\): The number of observed periods, which says how much evidence the other seven rest on\.
The first three are ratios over counts, the next three describe the shape of the path, and the last two describe its spread and its size, and the eight together therefore cover both what the two classical statistics of §[4\.2](https://arxiv.org/html/2609.30880#S4.SS2)rest on and what they leave out\. We now define each of them, from the raw window to the vector the router receives\.
Letx∈ℝLx\\in\\mathbb\{R\}^\{L\}be the raw context window of one series and letm∈\{0,1\}Lm\\in\\\{0,1\\\}^\{L\}mark its observed entries, withn=∑tmtn=\\sum\_\{t\}m\_\{t\}\. We first decide which of those entries count as a demand event, and because units differ by three orders of magnitude across the corpus, the test is relative to the scale of the series rather than absolute,
ϵ=10−3⋅1n∑tmt\|xt\|,zt=mt⋅\[\|xt\|\>ϵ\],nz=∑tzt\.\\epsilon=10^\{\-3\}\\cdot\\frac\{1\}\{n\}\\sum\_\{t\}m\_\{t\}\\,\|x\_\{t\}\|,\\qquad z\_\{t\}=m\_\{t\}\\cdot\\mathbb\{1\}\\\!\\left\[\\,\|x\_\{t\}\|\>\\epsilon\\,\\right\],\\qquad n\_\{z\}=\\sum\_\{t\}z\_\{t\}\.\(1\)
From these we form the two axes of the Syntetos–Boylan classification\[[66](https://arxiv.org/html/2609.30880#bib.bib66)\], namely the average inter\-demand interval and the squared coefficient of variation of the non\-zero values,
ADI=nnz,μz=1nz∑tzt\|xt\|,CV2=1nzμz2∑tzt\(\|xt\|−μz\)2\.\\mathrm\{ADI\}=\\frac\{n\}\{n\_\{z\}\},\\qquad\\mu\_\{z\}=\\frac\{1\}\{n\_\{z\}\}\\sum\_\{t\}z\_\{t\}\|x\_\{t\}\|,\\qquad\\mathrm\{CV\}^\{2\}=\\frac\{1\}\{n\_\{z\}\\mu\_\{z\}^\{2\}\}\\sum\_\{t\}z\_\{t\}\\left\(\|x\_\{t\}\|\-\\mu\_\{z\}\\right\)^\{2\}\.\(2\)
Four further quantities describe the shape of the series\. Withx¯\\bar\{x\}the mean over observed entries,x~t=mt\(xt−x¯\)\\tilde\{x\}\_\{t\}=m\_\{t\}\(x\_\{t\}\-\\bar\{x\}\)the centred series, andt~\\tilde\{t\}the centred time index, these are the trend correlation, the first\-order autocorrelation, the fraction of upward steps, and the overall coefficient of variation,
ρtrend=⟨t~,x~⟩‖t~‖‖x~‖,ρ1=⟨x~1:L−1,x~2:L⟩∥x~1:L−1∥∥x~2:L∥,u=1n−1∑t\[xt\+1\>xt\],CV=σ\|x¯\|\.\\rho\_\{\\text\{trend\}\}=\\frac\{\\langle\\tilde\{t\},\\tilde\{x\}\\rangle\}\{\\\|\\tilde\{t\}\\\|\\,\\\|\\tilde\{x\}\\\|\},\\quad\\rho\_\{1\}=\\frac\{\\langle\\tilde\{x\}\_\{1:L\-1\},\\tilde\{x\}\_\{2:L\}\\rangle\}\{\\\|\\tilde\{x\}\_\{1:L\-1\}\\\|\\,\\\|\\tilde\{x\}\_\{2:L\}\\\|\},\\quad u=\\frac\{1\}\{n\-1\}\\sum\_\{t\}\\mathbb\{1\}\\\!\\left\[x\_\{t\+1\}\>x\_\{t\}\\right\],\\quad\\mathrm\{CV\}=\\frac\{\\sigma\}\{\|\\bar\{x\}\|\}\.\(3\)
The router input is the concatenation of eight numbers, where the three ratio\-valued ones, namely entries 2, 3, and 7, and the length in entry 8 are compressed on a logarithmic axis so that no single statistic dominates the first layer,
s=\[1−nzn,13log\(1\+ADI\),13log\(1\+CV2\),ρtrend,ρ1,2u−1,13log\(1\+CV\),18logn\]∈ℝ8\.s=\\Big\[\\;1\-\\tfrac\{n\_\{z\}\}\{n\},\\;\\;\\tfrac\{1\}\{3\}\\log\(1\{\+\}\\mathrm\{ADI\}\),\\;\\;\\tfrac\{1\}\{3\}\\log\(1\{\+\}\\mathrm\{CV\}^\{2\}\),\\;\\;\\rho\_\{\\text\{trend\}\},\\;\\;\\rho\_\{1\},\\;\\;2u\-1,\\;\\;\\tfrac\{1\}\{3\}\\log\(1\{\+\}\\mathrm\{CV\}\),\\;\\;\\tfrac\{1\}\{8\}\\log n\\;\\Big\]\\in\\mathbb\{R\}^\{8\}\.\(4\)
Table 6:Example router input for four demand series\.Each row is the vectorssof Equation[4](https://arxiv.org/html/2609.30880#S4.E4)as the router receives it, and the four series are drawn from four different real\-world sources and are of four different lengths\.ClassSource8 StatisticsZeroADICV2TrendAC\(1\)UpCVLengthSmoothCiti Bike0\.000\.230\.04−\-0\.160\.36−\-0\.870\.100\.56IntermittentPort calls0\.981\.260\.00−\-0\.09−\-0\.02−\-0\.950\.670\.70ErraticSKU orders0\.030\.240\.32−\-0\.040\.10−\-0\.660\.280\.51LumpyRetail sales0\.860\.690\.14−\-0\.010\.00−\-0\.790\.470\.63Table[6](https://arxiv.org/html/2609.30880#S4.T6)showsssfor four series taken from the corpus, one from each of the four demand classes\. The classes separate on different entries, which is why eight numbers are kept rather than the two that define them\. The smooth and erratic series agree on the zero share and on the interval, and are separated by the dispersion of their values, while the intermittent and lumpy series are told apart by the interval and by that same dispersion together\.
Every entry ofssis invariant to the unit of the series\. A router that can see the magnitude of the values learns to recognise thesourcerather than thebehaviour, and a source it has not seen then falls outside everything it has learned\. The statistics are computed once per series, before the encoder, and the same eight numbers are handed to every adapted layer, so that no layer re\-derives them from the hidden state that the backbone passes to it\.
### 4\.2Stage 2: Continuous Demand Membership
Table 7:Classical demand classes\.A low ADI means that demand arrives often, and a high CV2means that its size varies widely\.CV2<0\.49\\mathrm\{CV\}^\{2\}<0\.49CV2≥0\.49\\mathrm\{CV\}^\{2\}\\geq 0\.49ADI<1\.32\\mathrm\{ADI\}<1\.32SmoothFrequency↑\\uparrow, Variability↓\\downarrowErraticFrequency↑\\uparrow, Variability↑\\uparrowADI≥1\.32\\mathrm\{ADI\}\\geq 1\.32IntermittentFrequency↓\\downarrow, Variability↓\\downarrowLumpyFrequency↓\\downarrow, Variability↑\\uparrowThe classical scheme of the intermittent demand literature\[[14](https://arxiv.org/html/2609.30880#bib.bib14),[65](https://arxiv.org/html/2609.30880#bib.bib65),[66](https://arxiv.org/html/2609.30880#bib.bib66)\]assigns a series to one of four classes by cutting Equation[2](https://arxiv.org/html/2609.30880#S4.E2)atADI=1\.32\\mathrm\{ADI\}=1\.32andCV2=0\.49\\mathrm\{CV\}^\{2\}=0\.49, as shown in Table[7](https://arxiv.org/html/2609.30880#S4.T7)\. However, demand series do not respect those cuts\. They spreadcontinuouslyacross the two axes, and a series just past a boundary is not a different kind of object from the one that lies just before it\. Series at ADI=1\.30=1\.30and1\.341\.34are nearly alike, yet the cuts split them\.
We therefore keep the two axes but soften the cuts, and since both quantities are ratios, we measure the distance to each cut on a logarithmic axis, where the gap betweenADI=1\.2\\mathrm\{ADI\}=1\.2and1\.41\.4carries the same weight as the gap between8\.08\.0and9\.39\.3,
a=σ\(logADI−log1\.32τa\),c=σ\(logCV2−log0\.49τc\),a=\\sigma\\\!\\left\(\\frac\{\\log\\mathrm\{ADI\}\-\\log 1\.32\}\{\\tau\_\{a\}\}\\right\),\\qquad c=\\sigma\\\!\\left\(\\frac\{\\log\\mathrm\{CV\}^\{2\}\-\\log 0\.49\}\{\\tau\_\{c\}\}\\right\),\(5\)withσ\\sigmathe logistic function andτa,τc\\tau\_\{a\},\\tau\_\{c\}the widths of the cuts\. Hereaareads as howintermittentthe series is andccas howerratic, and treating the two axes as independent gives a membership whose entries sum to one,
π=\[\(1−a\)\(1−c\),a\(1−c\),\(1−a\)c,ac\]=\[πsmooth,πinter,πerratic,πlumpy\]\.\\pi=\\big\[\\,\(1\-a\)\(1\-c\),\\;\\;a\(1\-c\),\\;\\;\(1\-a\)c,\\;\\;a\\,c\\,\\big\]\\;=\\;\[\\,\\pi\_\{\\text\{smooth\}\},\\;\\pi\_\{\\text\{inter\}\},\\;\\pi\_\{\\text\{erratic\}\},\\;\\pi\_\{\\text\{lumpy\}\}\\,\]\.\(6\)
A series that is clearly one thing gets a membership close to a corner, while a series between two behaviours keepsbothof them\. The lumpy example of Table[6](https://arxiv.org/html/2609.30880#S4.T6)hasπ=\(0\.00,0\.46,0\.00,0\.53\)\\pi=\(0\.00,0\.46,0\.00,0\.53\), which says it is lumpy by a narrow margin over intermittent, a split between the two classes that a hard label would have thrown away\.
Figure 9:One adapted projection\.The four routed branches are weighed by the gate, and the shared branch is always on\.Figure 10:Auxiliary supervision on one erratic series\.The bars are the router output, and the chain below turns the membership into the target that the cross\-entropy pulls that output towards during training\.
### 4\.3Stage 3: Routing
Each of the four routed branches is meant to carry one demand behaviour, and a rule has to decide how much of each branch a given series receives, which makes the membership of Equation[6](https://arxiv.org/html/2609.30880#S4.E6)the obvious candidate for that rule\. It rests on two of the eight statistics, however, and has to be recomputed from ADI and CV2for every series that arrives\. Therefore, we learn that rule instead, as a small network that reads the whole vectorssand that we call therouter, while the membership is kept as a supervision signal rather than as the mixture itself\. The auxiliary term of §[4\.5](https://arxiv.org/html/2609.30880#S4.SS5)makes that arrangement precise, and it is that term which ties each routed branch to one demand class\.
That router is a two\-layer network applied toss, followed by a softmax,
q=softmax\(W2ϕ\(W1s\+b1\)τ\)∈Δ3,W1∈ℝh×8,W2∈ℝ4×h,q=\\mathrm\{softmax\}\\\!\\left\(\\frac\{W\_\{2\}\\,\\phi\(W\_\{1\}s\+b\_\{1\}\)\}\{\\tau\}\\right\)\\in\\Delta^\{3\},\\qquad W\_\{1\}\\in\\mathbb\{R\}^\{h\\times 8\},\\;\\;W\_\{2\}\\in\\mathbb\{R\}^\{4\\times h\},\(7\)whereϕ\\phiis a GELU nonlinearity,τ\\tauis a temperature, andhhis the hidden width\. The router produces four weights, one per demand class of Equation[6](https://arxiv.org/html/2609.30880#S4.E6)\. The shared branch is not routed at all, as it takes a fixed weightwwwhile the routed branches divide whatever remains,
g=\[w,\(1−w\)q1,\(1−w\)q2,\(1−w\)q3,\(1−w\)q4\]∈Δ4\.g=\\big\[\\,w,\\;\\;\(1\-w\)\\,q\_\{1\},\\;\\;\(1\-w\)\\,q\_\{2\},\\;\\;\(1\-w\)\\,q\_\{3\},\\;\\;\(1\-w\)\\,q\_\{4\}\\,\\big\]\\in\\Delta^\{4\}\.\(8\)One gate is produced per series rather than per token, and every token of a series is adapted by the same mixture\.
Figure[10](https://arxiv.org/html/2609.30880#S4.F10)draws what one adapted projection then holds, namely the frozen weight, the shared branch, and the four routed ones that the gate weighs\. Each adapted layer carries its own router, and every one of them reads the same vectorssand decides its own mixture from it\. The router is small on purpose and cannot learn anything beyond a partition of the eight\-dimensional space of statistics, which is all the design asks of it and all the supervision teaches\.
Figure 11:What the router produces on real demand series\.Each panel holds a context window and the weight that the gate gives to each of the four routed branches, against the target that supervision asks for, where the four bars carry the smooth, intermittent, erratic, and lumpy branches of the mixture in that same order\.Nothing in Equation[7](https://arxiv.org/html/2609.30880#S4.E7)ties routed branchccto any particular behaviour\. Left alone the routed branches areinterchangeable, and whichever partition the gate settles on is anaccident of initialisation\. We therefore supervise the gate with the membership of Equation[6](https://arxiv.org/html/2609.30880#S4.E6), using a cross\-entropy between the two distributions,
ℒ=ℒforecast\+λ⋅ℒaux,ℒaux=−∑c=14π~clogqc\.\\mathcal\{L\}=\\mathcal\{L\}\_\{\\text\{forecast\}\}\\;\+\\;\\lambda\\cdot\\mathcal\{L\}\_\{\\text\{aux\}\},\\qquad\\mathcal\{L\}\_\{\\text\{aux\}\}=\-\\sum\_\{c=1\}^\{4\}\\tilde\{\\pi\}\_\{c\}\\log q\_\{c\}\.\(9\)
The loss is written onqqrather than ongg, which leaves the shared branch out of it and supervises only the four routed branches\. That is what gives routed branchcca meaning of its own, in that branch 1 is pulled toward smooth series and branch 2 toward intermittent ones, while a series that is0\.60\.6lumpy and0\.40\.4intermittent asks for both in that proportion\.
The targetπ~\\tilde\{\\pi\}is a sharpened form of the membership of Equation[6](https://arxiv.org/html/2609.30880#S4.E6),
π~c=πc1/T∑c′πc′1/T,\\tilde\{\\pi\}\_\{c\}\\;=\\;\\frac\{\\pi\_\{c\}^\{1/T\}\}\{\\sum\_\{c^\{\\prime\}\}\\pi\_\{c^\{\\prime\}\}^\{1/T\}\},\(10\)whereT<1T<1is the sharpening temperature\. We sharpen because series near the cuts of Equation[2](https://arxiv.org/html/2609.30880#S4.E2)carry a membership that is close to a tie, and a target such as\[0,0\.40,0,0\.60\]\[0,0\.40,0,0\.60\]is too weak to settle which branch should lead\. The alternative is a hard label and a negative log\-likelihood on the assigned class, which pullsqqtoward a one\-hot vector\. We do not use it, because it contradicts the premise of Equation[6](https://arxiv.org/html/2609.30880#S4.E6), namely that a demand series is amixture of behavioursrather than amember of one class\.
Figure[10](https://arxiv.org/html/2609.30880#S4.F10)follows one erratic series through that chain, from the membership we measure to the target we sharpen and the weights the router finally produces\. The auxiliary term is atraining signal only, and at inference the router reads the eight statistics of the raw window and no membership is computed, as only the loss needs it\.
Figure[11](https://arxiv.org/html/2609.30880#S4.F11)puts the trained router of EXAONE Demand on six real series to show what those weights look like once training is done\. The four panels named by a class are the series on which the router is the most decisive, and on them the weight it produces sits almost exactly on the target that the supervision asks for\. The two panels namedMixedare the series on which it is the least decisive, and the target there is spread across three branches while the router spreads its weight in the same way, which is what a series that mixes two behaviours is meant to look like at the gate\.
### 4\.4Stage 4: Mixture of Low\-Rank Branches
We replace six projections in each encoder block, namely the four attention projections and the two feed\-forward projections, and leave the rest of the block untouched\. Each adapted layer keeps its frozen weightW0W\_\{0\}and addsE=5E=5low\-rank branches mixed bygg,
y=W0x\+∑e=04ge⋅αre⋅BeAex,Ae∈ℝre×din,Be∈ℝdout×re\.y\\;=\\;W\_\{0\}\\,x\\;\+\\;\\sum\_\{e=0\}^\{4\}g\_\{e\}\\cdot\\frac\{\\alpha\}\{r\_\{e\}\}\\cdot B\_\{e\}A\_\{e\}\\,x,\\qquad A\_\{e\}\\in\\mathbb\{R\}^\{r\_\{e\}\\times d\_\{\\text\{in\}\}\},\\;\\;B\_\{e\}\\in\\mathbb\{R\}^\{d\_\{\\text\{out\}\}\\times r\_\{e\}\}\.\(11\)Branch 0 is the shared one and branches 1 to 4 are the routed ones\. The shared branch is given the larger rank, so that it alone holds as much capacity as a plain low\-rank adapter, while each routed branch adds a small correction on top of it\. A shared branch that is too small would leave part of the common correction to the routed branches, and that residue would blur the routing that the supervision of Stage 3 asks for\. The scaleα/re\\alpha/r\_\{e\}follows the low\-rank convention and is applied per branch, so that a change of rank does not change the scale at which a branch enters the sum\.
Written naively, Equation[11](https://arxiv.org/html/2609.30880#S4.E11)forms one output tensor per branch\. We instead stack the branches along the rank axis and evaluate the sum with two matrix products,
y=W0x\+B¯\(\(A¯x\)⊙\(g⊗𝟙r\)\),A¯=\[A0A4\],B¯=\[B0⋯B4\],y=W\_\{0\}x\+\\bar\{B\}\\,\\big\(\(\\bar\{A\}x\)\\odot\(g\\otimes\\mathbb\{1\}\_\{r\}\)\\big\),\\qquad\\bar\{A\}=\\begin\{bmatrix\}A\_\{0\}\\\\ \\vdots\\\\ A\_\{4\}\\end\{bmatrix\},\\quad\\bar\{B\}=\\begin\{bmatrix\}B\_\{0\}&\\cdots&B\_\{4\}\\end\{bmatrix\},\(12\)where the only intermediate tensor has width∑ere\\sum\_\{e\}r\_\{e\}rather thanEdoutE\\,d\_\{\\text\{out\}\}\.
### 4\.5Stage 5: Training
OnlyAeA\_\{e\},BeB\_\{e\}, and the router of Equation[7](https://arxiv.org/html/2609.30880#S4.E7)receive gradients\. Everything else stays frozen, including the embedding and the normalisation layers\. The loss is Equation[9](https://arxiv.org/html/2609.30880#S4.E9), and the weightλ\\lambdais held fixed for every run that we report\.
The three choices below are what keep the four routed branches apart from each other, and each of them does so in a different way, namely how the router starts, what it reads, and how much weight the shared branch is given\.
Router starts at random\.WithW2=0W\_\{2\}=0the softmax of Equation[7](https://arxiv.org/html/2609.30880#S4.E7)is exactly uniform for every input, every routed branch receives the same gradient, and the symmetry between them never breaks\. A fixed uniform mixture ofEEbranches of rankrrequals a single adapter of rankErErby Equation[11](https://arxiv.org/html/2609.30880#S4.E11), and the architecture then reduces to the thing it was meant to improve on\. DrawingW2W\_\{2\}at random costs nothing, since the adapter output is multiplied byBe=0B\_\{e\}=0at step zero and the adapted layer therefore returnsW0xW\_\{0\}xexactly, no matter which weights the router assigns\.
Router reads the raw series rather than the hidden state\.A gate applied to hidden states learns to recognise something the backbone encodes for itself, and its softmax flattens for the same reason as above\. The raw series gives the router information that the encodernever sees, since the encoder is fed a normalised window\.
Shared weight balances the two parts\.The weightwwof Equation[8](https://arxiv.org/html/2609.30880#S4.E8)decides how much of the adapter is common and how much is routed\. Near zero the routed branches carry almost everything and near one they carry almost nothing, and we therefore setwwbetween the two, where both parts contribute, atw=0\.5w=0\.5in every run that we report\.
## 5Experiments
### 5\.1Datasets and Metrics
Evaluation suite\.We evaluate on 22 held\-out datasets, none of which takes any part in training, as verified by the leakage check of Section[2](https://arxiv.org/html/2609.30880#S2)at the level ofraw valuesrather thandataset names\. Every model is scored on the same window, the same horizon, and the same quantile levels, with no per\-model tuning of any kind\.
Metrics\.We report the six metrics of Table[8](https://arxiv.org/html/2609.30880#S5.T8), namely the mean absolute scaled error \(MASE\), the normalised deviation \(ND\), the weighted quantile loss \(WQL\), the mean absolute percentage error \(MAPE\), the mean absolute error \(MAE\), and the mean scaled interval score \(MSIS\), and we aggregate every one of them across the 22 datasets by geometric mean, which gives every dataset the same share of the aggregate whatever level its errors sit at\.
Letyty\_\{t\}be the observed value at stepttof a horizon of lengthhh, lety^t\\hat\{y\}\_\{t\}be the median forecast, lety^t\(q\)\\hat\{y\}^\{\(q\)\}\_\{t\}be the forecast at quantile levelqq, and letx1,…,xnx\_\{1\},\\dots,x\_\{n\}be the training portion of the same series at seasonal periodmm, which the first and the last of the six use alone\. The six of them are written as follows,
MASE=1h∑t=1h\|yt−y^t\|1n−m∑t=m\+1n\|xt−xt−m\|,ND=∑t=1h\|yt−y^t\|∑t=1h\|yt\|,\\mathrm\{MASE\}=\\frac\{\\frac\{1\}\{h\}\\sum\_\{t=1\}^\{h\}\\lvert y\_\{t\}\-\\hat\{y\}\_\{t\}\\rvert\}\{\\frac\{1\}\{n\-m\}\\sum\_\{t=m\+1\}^\{n\}\\lvert x\_\{t\}\-x\_\{t\-m\}\\rvert\},\\qquad\\mathrm\{ND\}=\\frac\{\\sum\_\{t=1\}^\{h\}\\lvert y\_\{t\}\-\\hat\{y\}\_\{t\}\\rvert\}\{\\sum\_\{t=1\}^\{h\}\\lvert y\_\{t\}\\rvert\},\(13\)
WQL=∑q∈Q∑t=1h2\(q\[yt−y^t\(q\)\]\+\+\(1−q\)\[y^t\(q\)−yt\]\+\)\|Q\|∑t=1h\|yt\|,MAPE=1h∑t=1h\|yt−y^t\|\|yt\|,\\mathrm\{WQL\}=\\frac\{\\sum\_\{q\\in Q\}\\sum\_\{t=1\}^\{h\}2\\,\\big\(q\\,\[\\,y\_\{t\}\-\\hat\{y\}^\{\(q\)\}\_\{t\}\\,\]\_\{\+\}\+\(1\-q\)\\,\[\\,\\hat\{y\}^\{\(q\)\}\_\{t\}\-y\_\{t\}\\,\]\_\{\+\}\\big\)\}\{\\lvert Q\\rvert\\sum\_\{t=1\}^\{h\}\\lvert y\_\{t\}\\rvert\},\\qquad\\mathrm\{MAPE\}=\\frac\{1\}\{h\}\\sum\_\{t=1\}^\{h\}\\frac\{\\lvert y\_\{t\}\-\\hat\{y\}\_\{t\}\\rvert\}\{\\lvert y\_\{t\}\\rvert\},\(14\)
MAE=1h∑t=1h\|yt−y^t\|,MSIS=1h∑t=1h\(ut−lt\+2α\[lt−yt\]\+\+2α\[yt−ut\]\+\)1n−m∑t=m\+1n\|xt−xt−m\|,\\mathrm\{MAE\}=\\frac\{1\}\{h\}\\sum\_\{t=1\}^\{h\}\\lvert y\_\{t\}\-\\hat\{y\}\_\{t\}\\rvert,\\qquad\\mathrm\{MSIS\}=\\frac\{\\frac\{1\}\{h\}\\sum\_\{t=1\}^\{h\}\\Big\(u\_\{t\}\-l\_\{t\}\+\\tfrac\{2\}\{\\alpha\}\\big\[\\,l\_\{t\}\-y\_\{t\}\\,\\big\]\_\{\+\}\+\\tfrac\{2\}\{\\alpha\}\\big\[\\,y\_\{t\}\-u\_\{t\}\\,\\big\]\_\{\+\}\\Big\)\}\{\\frac\{1\}\{n\-m\}\\sum\_\{t=m\+1\}^\{n\}\\lvert x\_\{t\}\-x\_\{t\-m\}\\rvert\},\(15\)
whereQQis the set of nine quantile levels that both of our models emit,\[⋅\]\+\[\\,\\cdot\\,\]\_\{\+\}keeps the positive part of its argument, andltl\_\{t\}andutu\_\{t\}are the bounds of the central1−α1\-\\alphainterval atα=0\.2\\alpha=0\.2\. Each of the six divides the error by a different quantity, and that divisor is what the metric is about\. All six are non\-negative and unbounded above, and lower is better on each:
- •MASE: How much better than doing nothing a model is, since the divisor is the error that a seasonal naive forecaster makes on the history of the same series\. A value above one says that repeating last season would have been the better choice of the two, and every row of Table[8](https://arxiv.org/html/2609.30880#S5.T8)sits above one on the present evaluation suite\.
- •ND: The share of the demand that the forecast misplaces, since the divisor is the total volume of the horizon\. A value of0\.140\.14reads as fourteen units misplaced for every hundred demanded, the currency a planner stocks in\.
- •WQL: How good the whole predictive distribution is rather than its middle alone, since the error at each quantile level is charged asymmetrically\. Missing above the0\.90\.9quantile costs nine times as much as missing below it\.
- •MAPE: The error as a fraction of the demand at each step taken separately, which is the most familiar of the six and the least suited to this setting\. A value of0\.260\.26reads as a quarter of the demand missed on the average period\.
- •MAE: The average absolute error, and the only one of the six that keeps the unit of the series\. Its level says nothing across datasets, while its ordering across the models of one suite still says what the others say\.
- •MSIS: How good the interval is rather than the point or the distribution, charging the width of the central80%80\\%band plus a penalty for every observation outside it\. It is the one metric that widening the band alone cannot win\.
Average rank and win rate\.To separatehow mucha model is ahead fromhow oftenit is ahead, we report two further quantities beside the six metrics, both of which read the per\-dataset MASE rather than the aggregate of it\. An aggregate is a single number that a handful of datasets can carry, and a model that is far ahead on three datasets and behind on the other nineteen can lead on it, which is the reading that these two are there to rule out\.
Theaverage rankof a model is its mean position across the 22 datasets, where a model that is second on every one of them scores 2 however wide or narrow the gaps around it are\. Thewin rateis the share of the pairwise comparisons that a model wins, and with 38 other rows and 22 datasets there are 836 comparisons for each of them\. A model wins one of those when its MASE on that dataset is the lower of the two, and a tie counts as a win for neither model\.
### 5\.2Experimental Setup
Baselines\.We compare against the 36 TSFMs of Table[8](https://arxiv.org/html/2609.30880#S5.T8)and the frozen backbone that we adapt, allrun by usrather thanquoted from their papers\. The comparison covers Chronos\-Bolt \(Base, Small\) and Chronos\-T5 \(Base, Small\)\[[4](https://arxiv.org/html/2609.30880#bib.bib4)\], Chronos\-2 and Chronos\-2 \(Synthetic\)\[[3](https://arxiv.org/html/2609.30880#bib.bib3)\], TiRex\-1\.1\[[5](https://arxiv.org/html/2609.30880#bib.bib5)\], TimesFM\-1\.0, TimesFM\-2\.0 and TimesFM\-2\.5\[[17](https://arxiv.org/html/2609.30880#bib.bib17)\], Moirai\-1\.1\-R \(Small, Base, Large\) and Moirai\-2\.0\-R\[[72](https://arxiv.org/html/2609.30880#bib.bib72)\], Timer\-S1\[[47](https://arxiv.org/html/2609.30880#bib.bib47)\], Sundial\[[46](https://arxiv.org/html/2609.30880#bib.bib46)\], Toto \(Open Base\)\[[13](https://arxiv.org/html/2609.30880#bib.bib13)\], TabPFN\-TS\[[31](https://arxiv.org/html/2609.30880#bib.bib31)\], FlowState and Granite FlowState\-R1\[[27](https://arxiv.org/html/2609.30880#bib.bib27)\], TTM\-R1 and TTM\-R2\[[21](https://arxiv.org/html/2609.30880#bib.bib21)\], PatchTST\-FM\-R1 and Granite PatchTST\-FM\-R1\[[55](https://arxiv.org/html/2609.30880#bib.bib55)\], Kairos \(10M, 23M, 50M\)\[[22](https://arxiv.org/html/2609.30880#bib.bib22)\], TempoPFN\[[52](https://arxiv.org/html/2609.30880#bib.bib52)\], Reverso\[[23](https://arxiv.org/html/2609.30880#bib.bib23)\], VisionTS\[[10](https://arxiv.org/html/2609.30880#bib.bib10)\], Lag\-Llama\[[58](https://arxiv.org/html/2609.30880#bib.bib58)\], CleanTS\[[20](https://arxiv.org/html/2609.30880#bib.bib20)\], and YingLong \(6M, 50M, 110M, 300M\)\[[70](https://arxiv.org/html/2609.30880#bib.bib70)\]\. No baseline is selected by how well it does on the 22 datasets, and every model, ours included, is scored under the same protocol, with no chance to tune on it at any stage of the study that this section reports and the appendix records in full\.
Implementation\.The adapter of Section[4](https://arxiv.org/html/2609.30880#S4)is attached to the four attention projections and the two feed\-forward projections of every encoder block, and the backbone stays frozen throughout\. The shared branch is given the larger rank and half of the mixture weight, and the router is trained with the auxiliary term of Equation[9](https://arxiv.org/html/2609.30880#S4.E9)alongside the forecast loss\. Appendix[A](https://arxiv.org/html/2609.30880#A1)records every hyperparameter behind the tables and figures of this section\.
### 5\.3Two Versions of EXAONE Demand
Table 8:Comparison with TSFMs\.Both of our models are shown against every baseline we ran and against the frozen backbone that we adapt, on the 22 evaluation datasets, with the rows ordered and numbered \(\#\) by MASE\.\#ModelRankingError metricsAvg\. rankWin rateMASENDWQLMAPEMAEMSIS1EXAONE Demand4\.0991\.9%1\.06670\.14090\.11380\.256560\.2410\.702EXAONE Demand \(Synthetic\)5\.5588\.0%1\.07420\.14200\.11470\.258360\.7310\.633TiRex\-1\.1\[[5](https://arxiv.org/html/2609.30880#bib.bib5)\]7\.6482\.5%1\.08180\.14510\.11610\.291162\.0610\.754Chronos\-2\[[3](https://arxiv.org/html/2609.30880#bib.bib3)\]8\.5980\.0%1\.08850\.14970\.14060\.301064\.0214\.485EXAONE Backbone \(Zero\-shot\)7\.9581\.7%1\.10500\.14450\.11690\.260961\.8211\.416TimesFM\-2\.5\[[17](https://arxiv.org/html/2609.30880#bib.bib17)\]10\.0576\.2%1\.10730\.14860\.12150\.303263\.5812\.317Chronos\-Bolt \(Base\)\[[4](https://arxiv.org/html/2609.30880#bib.bib4)\]11\.0573\.6%1\.11220\.14930\.11910\.294163\.8411\.148Timer\-S1\[[47](https://arxiv.org/html/2609.30880#bib.bib47)\]12\.2770\.3%1\.11970\.14680\.11710\.299762\.7712\.089Reverso\[[23](https://arxiv.org/html/2609.30880#bib.bib23)\]11\.2773\.0%1\.12910\.15100\.15100\.277364\.6045\.1710Chronos\-Bolt \(Small\)\[[4](https://arxiv.org/html/2609.30880#bib.bib4)\]14\.0065\.8%1\.13310\.14850\.11930\.299863\.5311\.5711Chronos\-T5 \(Base\)\[[4](https://arxiv.org/html/2609.30880#bib.bib4)\]8\.6879\.8%1\.13360\.14940\.12190\.271263\.9013\.3312Toto \(Open Base\)\[[13](https://arxiv.org/html/2609.30880#bib.bib13)\]11\.7371\.8%1\.13990\.14950\.12140\.267763\.9311\.6413Chronos\-2 \(Synthetic\)\[[3](https://arxiv.org/html/2609.30880#bib.bib3)\]15\.9560\.6%1\.15150\.15940\.14520\.308868\.1716\.1114Moirai\-2\.0\-R\[[72](https://arxiv.org/html/2609.30880#bib.bib72)\]12\.0570\.9%1\.15860\.15310\.12600\.280265\.4813\.5615Chronos\-T5 \(Small\)\[[4](https://arxiv.org/html/2609.30880#bib.bib4)\]12\.2770\.3%1\.16230\.15240\.12490\.275165\.1813\.8716Sundial\[[46](https://arxiv.org/html/2609.30880#bib.bib46)\]15\.6861\.4%1\.16900\.15430\.13250\.300666\.0019\.2617TabPFN\-TS\[[31](https://arxiv.org/html/2609.30880#bib.bib31)\]20\.2349\.4%1\.21720\.17190\.13530\.352073\.5212\.2118FlowState\[[27](https://arxiv.org/html/2609.30880#bib.bib27)\]19\.2352\.0%1\.23630\.17050\.13570\.342272\.9412\.1419Granite FlowState\-R1\[[27](https://arxiv.org/html/2609.30880#bib.bib27)\]18\.6853\.5%1\.24190\.17100\.13630\.346973\.1212\.6720Moirai\-1\.1\-R \(Large\)\[[72](https://arxiv.org/html/2609.30880#bib.bib72)\]18\.2354\.7%1\.25720\.16400\.13220\.302770\.1612\.5121Moirai\-1\.1\-R \(Base\)\[[72](https://arxiv.org/html/2609.30880#bib.bib72)\]19\.0952\.4%1\.27240\.16380\.13310\.310870\.0612\.6722TimesFM\-1\.0\[[17](https://arxiv.org/html/2609.30880#bib.bib17)\]20\.0549\.9%1\.27940\.17270\.14270\.344873\.8716\.2723PatchTST\-FM\-R1\[[55](https://arxiv.org/html/2609.30880#bib.bib55)\]21\.2346\.8%1\.29840\.16930\.13840\.331272\.4112\.9124Moirai\-1\.1\-R \(Small\)\[[72](https://arxiv.org/html/2609.30880#bib.bib72)\]22\.5043\.4%1\.35700\.17530\.14230\.323074\.9813\.5825TempoPFN\[[52](https://arxiv.org/html/2609.30880#bib.bib52)\]21\.6845\.6%1\.39680\.17940\.14750\.371676\.7319\.6126Granite PatchTST\-FM\-R1\[[55](https://arxiv.org/html/2609.30880#bib.bib55)\]23\.4141\.0%1\.40610\.17900\.14930\.348176\.5416\.2527Kairos \(50M\)\[[22](https://arxiv.org/html/2609.30880#bib.bib22)\]22\.2344\.1%1\.46410\.19050\.15720\.364881\.4618\.9728Kairos \(23M\)\[[22](https://arxiv.org/html/2609.30880#bib.bib22)\]26\.0034\.2%1\.53910\.20810\.17370\.373089\.0220\.8729TTM\-R1\[[21](https://arxiv.org/html/2609.30880#bib.bib21)\]27\.9129\.2%1\.57140\.20140\.20140\.404786\.1362\.8630TTM\-R2\[[21](https://arxiv.org/html/2609.30880#bib.bib21)\]29\.7324\.4%1\.63990\.21250\.21250\.438790\.8965\.6031Kairos \(10M\)\[[22](https://arxiv.org/html/2609.30880#bib.bib22)\]26\.9131\.8%1\.64630\.20670\.17050\.390988\.4221\.0532VisionTS\[[10](https://arxiv.org/html/2609.30880#bib.bib10)\]30\.2723\.0%1\.90690\.24440\.24440\.5455104\.5576\.2833Lag\-Llama\[[58](https://arxiv.org/html/2609.30880#bib.bib58)\]29\.7324\.4%2\.03280\.24370\.20780\.4339104\.2424\.9734TimesFM\-2\.0\[[17](https://arxiv.org/html/2609.30880#bib.bib17)\]33\.1415\.4%2\.46840\.25480\.19890\.4632108\.9820\.0935CleanTS\[[20](https://arxiv.org/html/2609.30880#bib.bib20)\]33\.8613\.5%3\.15660\.37050\.31120\.7209158\.4845\.1636YingLong \(6M\)\[[70](https://arxiv.org/html/2609.30880#bib.bib70)\]36\.007\.9%3\.73290\.46720\.40510\.8661199\.8373\.9637YingLong \(110M\)\[[70](https://arxiv.org/html/2609.30880#bib.bib70)\]36\.825\.7%3\.84430\.47730\.42890\.8865204\.1488\.7938YingLong \(300M\)\[[70](https://arxiv.org/html/2609.30880#bib.bib70)\]36\.865\.6%3\.84910\.48790\.43190\.8870208\.6985\.4339YingLong \(50M\)\[[70](https://arxiv.org/html/2609.30880#bib.bib70)\]37\.414\.2%3\.88530\.48610\.43480\.8876207\.9387\.12We build two models that differ only in their training data\.EXAONE Demand, the model of Section[4](https://arxiv.org/html/2609.30880#S4), is trained on real\-world and synthetic demand together, andEXAONE Demand \(Synthetic\)on the synthetic corpus alone, with no real\-world series taking part at any point of training\. The second model exists because open demand data carries licences that a model trained on it inherits, while one trained on series we generatedourselvesinherits none of them\. The question is whether the synthetic corpus canstand on its ownrather than onlyfill gaps\.
### 5\.4Main Results
Table[8](https://arxiv.org/html/2609.30880#S5.T8)reports both of our models against every baseline we ran, on the 22 evaluation datasets and under the protocol of §[5\.1](https://arxiv.org/html/2609.30880#S5.SS1)\. Both of them lead the table on every one of the six metrics as well as on the average rank and the win rate, and neither was tuned against the suite that scores them\. The two take the first and the second place on five of the six metrics, and they divide those two places between themselves on MSIS, where the synthetic\-only model is the better of the pair by a margin of only0\.070\.07, and we therefore read MSIS as a tie between our two models rather than as a win\.
EXAONE Demand wins91\.9%91\.9\\%of its836836pairwise comparisons against the other3838rows, where the strongest baseline wins82\.5%82\.5\\%of the comparisons among its own peers\. The synthetic\-only model wins88\.0%88\.0\\%and stays ahead of every baseline in the table, which shows that the synthetic corpus can carry a model on its own\.
### 5\.5Results by Dataset
An aggregate can hide a group of datasets on which a model loses while it still wins on the geometric mean, and a demand model is exactly the case where that matters, since the hard classes are a minority of the suite\. Table[9](https://arxiv.org/html/2609.30880#S5.T9)therefore reports every one of the datasets separately rather than folded into a single number\.
The spread across the rows of that table is wider than the spread across its columns, since some datasets are several times harder than the rest for every baseline we ran, and a few rows of that kind move the aggregate further than any architecture choice of Section[4](https://arxiv.org/html/2609.30880#S4)moves it, as Bitbrains \(random\), M4 yearly, and M4 daily show\.
Neither of our models is last on any of the 22 datasets, each of them is ahead of all 36 baselines on 7 of them, and the real\-world model is ahead of the synthetic\-only one on 16 of the 22\. The gap between the two is therefore a lead held across the suite rather than the work of a handful of rows on which one of them happens to do unusually well, and the same holds of the gap between either of them and the strongest of the 36 baselines in Table[8](https://arxiv.org/html/2609.30880#S5.T8)\.
Table 9:Per\-dataset MASE on the 22 evaluation datasets\.Lower is better, the letter in brackets is the sampling interval, and the five baselines are the strongest of the 36 by their aggregate over the suite\. Real\+Synth is the model trained on real\-world and synthetic demand, while Synth is the one trained on the synthetic corpus alone\.DatasetEXAONE DemandTSFM BaselineReal\+SynthSynthTiRex\-1\.1\[[5](https://arxiv.org/html/2609.30880#bib.bib5)\]Chronos\-2\[[3](https://arxiv.org/html/2609.30880#bib.bib3)\]TimesFM\-2\.5\[[17](https://arxiv.org/html/2609.30880#bib.bib17)\]Chronos\-Bolt \(Base\)\[[4](https://arxiv.org/html/2609.30880#bib.bib4)\]Timer\-S1\[[47](https://arxiv.org/html/2609.30880#bib.bib47)\]Bitbrains \(fast storage\) \(H\)1\.0201\.0081\.0160\.9951\.1201\.0821\.126Bitbrains \(random\) \(H\)5\.8295\.8325\.8595\.8535\.9515\.9265\.955BizITObs L2C \(H\)0\.4410\.4450\.4870\.4390\.5200\.4280\.407Car parts \(M\)0\.8280\.8470\.9160\.9000\.9420\.9030\.893Electricity \(D\)1\.3751\.3831\.4661\.4151\.4361\.5131\.456Electricity \(H\)1\.0101\.0090\.9120\.8880\.9420\.8860\.897Electricity \(W\)1\.4941\.4681\.4781\.4561\.4911\.5171\.614Hierarchical sales \(D\)0\.7460\.7610\.7680\.7680\.7860\.7650\.776Hierarchical sales \(W\)0\.7190\.7220\.7470\.7450\.7480\.7640\.755Hospital \(M\)0\.7590\.7650\.7850\.8070\.7630\.8280\.839Loop Seattle \(D\)0\.8890\.8910\.9000\.9280\.8810\.9090\.896Loop Seattle \(H\)0\.8170\.8190\.8200\.8160\.8800\.8980\.808M4 daily \(D\)3\.0793\.0463\.0633\.1643\.2023\.0703\.336M4 hourly \(H\)0\.7630\.8050\.7080\.8000\.7240\.8740\.797M4 monthly \(M\)0\.9090\.9250\.9380\.9430\.9630\.9781\.038M4 quarterly \(Q\)1\.1491\.1501\.1181\.1771\.1851\.2101\.259M4 weekly \(W\)2\.0662\.1501\.9732\.0701\.9802\.1542\.442M4 yearly \(Y\)3\.1653\.2173\.2383\.2563\.6353\.3283\.692M\-dense \(D\)0\.6450\.6400\.6780\.6810\.6520\.6760\.631M\-dense \(H\)0\.7860\.7830\.8020\.8070\.8300\.7990\.776Restaurant \(D\)0\.6760\.6770\.7040\.7130\.6960\.7300\.717SZ taxi \(H\)0\.5660\.5670\.5820\.5910\.5730\.5800\.5981st1^\{\\text\{st\}\}Count101321142nd2^\{\\text\{nd\}\}Count21312310
### 5\.6Forecast Visualization
Figure[12](https://arxiv.org/html/2609.30880#S5.F12)places EXAONE Demand beside the two strongest baselines of Table[8](https://arxiv.org/html/2609.30880#S5.T8), TiRex\-1\.1 and Chronos\-2, on six series of the evaluation suite, and our model reaches the lowest MASE of the three on every one of them\. On the spare\-part and hospital series, the two baselines forecast a return of demand that never arrives, while our model stays at zero on the first and holds the level on the second\. On the two electricity series, demand collapses to zero within the horizon and TiRex\-1\.1 keeps repeating the cycle it has read, while our model follows the collapse\. On the two hourly series, M\-dense and M4 hourly, our model tracks the height of the sharp peaks most closely\. The panels are chosen to show what the adaptation does when it works, and a selected curve cannot carry the claim that the aggregate carries\.
Figure 12:Forecasts on six demand series\.Each line is the median forecast of one model, and EXAONE Demand reaches the lowest MASE of the three on all six series\. On Car parts \(M\) and Electricity \(D\), where demand falls to zero, TiRex\-1\.1 and Chronos\-2 keep forecasting a positive level, while EXAONE Demand follows the zeros\.
### 5\.7Effect of Training Data
To isolate what real\-world demand contributes, we compare runs that share the architecture, learning rate, target set, and seed, and differ only in whetherreal\-world series take part in training\. The left panel of Figure[13](https://arxiv.org/html/2609.30880#S5.F13)sweeps the share of synthetic series in the mixture from0\.050\.05to1\.01\.0, where the last setting is the synthetic corpus alone\. The curve has an interior minimum at the share we adopt, and it rises sharply once no real\-world series is left in the mixture, yet both of our models stay well below TiRex\-1\.1, the strongest baseline of Table[8](https://arxiv.org/html/2609.30880#S5.T8)\. The right panel places our two models side by side on each of the 22 datasets, where real\-world demand lowers the error on 16 of them and raises it on 6, and the datasets it helps most are helped several times as much as the ones it costs anything on are hurt\. What the real\-world corpus buys is therefore held across the suite rather than produced by a handful of rows\.
Figure 13:What the training data contributes\.\[Left\] The share of synthetic series in the mixture, where each point is the median of three seeds except the synthetic\-only point at1\.01\.0, which is the median of its own runs\. \[Right\] The change in MASE from adding real\-world demand to the synthetic corpus, on the 22 datasets in the evaluation suite\.
## 6Conclusion
We took a general TSFM as given and asked what a corpus and an adapter built fordemandare worth\. The corpus is assembled from real\-world sources, verified against outside counts, separated from the evaluation data at the level ofraw values, and completed by a synthetic generator\. On this corpus, EXAONE Demand trains an MoE adapter whose low\-rank branches are mixed by the demand behaviour of each series, and both versions, one trained on real\-world and synthetic demand and one on synthetic data alone, outperform every compared TSFM on 22 held\-out datasets\.
Limitation and future works\.The router reads the context window alone, and the events that drive demand, namely promotions, holidays, and price changes, reach the model only through the trace they leave in the series, which makes a covariate\-aware router the natural next step\. The four classes and the number of routed branches are fixed rather than learned, and stock\-outs are modelled only by the synthetic generator, which leaves the recovery of latent demand from censored sales open for the next version of the model, together with a router that is supervised by forecast error\.
## References
- \[1\]Aslan Ahmedov\.Walmart sales forecast\.Kaggle dataset,[https://www\.kaggle\.com/datasets/aslanahmedov/walmart\-sales\-forecast](https://www.kaggle.com/datasets/aslanahmedov/walmart-sales-forecast), 2022\.
- \[2\]Taha Aksu, Gerald Woo, Juncheng Liu, Xu Liu, Chenghao Liu, Silvio Savarese, Caiming Xiong, and Doyen Sahoo\.GIFT\-Eval: A Benchmark For General Time Series Forecasting Model Evaluation\.InNeurIPS 2024 TSALM Workshop, 2024\.
- \[3\]Abdul Fatir Ansari, Oleksandr Shchur, Jaris Küken, Andreas Auer, Boran Han, Pedro Mercado, Syama Sundar Rangapuram, Huibin Shen, Lorenzo Stella, Xiyuan Zhang, Mononito Goswami, Shubham Kapoor, Danielle C\. Maddix, Pablo Guerron, Tony Hu, Junming Yin, Nick Erickson, Prateek Mutalik Desai, Hao Wang, Huzefa Rangwala, George Karypis, Yuyang Wang, and Michael Bohlke\-Schneider\.Chronos\-2: From Univariate to Universal Forecasting\.[https://arxiv\.org/abs/2510\.15821](https://arxiv.org/abs/2510.15821), 2025\.
- \[4\]Abdul Fatir Ansari, Lorenzo Stella, Caner Turkmen, Xiyuan Zhang, Pedro Mercado, Huibin Shen, Oleksandr Shchur, Syama Sundar Rangapuram, Sebastian Pineda Arango, Shubham Kapoor, Jasper Zschiegner, Danielle C\. Maddix, Hao Wang, Michael W\. Mahoney, Kari Torkkola, Andrew Gordon Wilson, Michael Bohlke\-Schneider, and Yuyang Wang\.Chronos: Learning the Language of Time Series\.TMLR, 2024\.
- \[5\]Andreas Auer, Patrick Podest, Daniel Klotz, Sebastian Böck, Günter Klambauer, and Sepp Hochreiter\.TiRex: Zero\-Shot Forecasting Across Long and Short Horizons with Enhanced In\-Context Learning\.InNeurIPS, 2025\.
- \[6\]Australian Energy Market Operator\.Aggregated price and demand data\.[https://aemo\.com\.au/energy\-systems/electricity/national\-electricity\-market\-nem/data\-nem/aggregated\-data](https://aemo.com.au/energy-systems/electricity/national-electricity-market-nem/data-nem/aggregated-data), 2026\.
- \[7\]Shaojie Bai, J\. Zico Kolter, and Vladlen Koltun\.An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling\.[https://arxiv\.org/abs/1803\.01271](https://arxiv.org/abs/1803.01271), 2018\.
- \[8\]George E\. P\. Box and Gwilym M\. Jenkins\.Time Series Analysis: Forecasting and Control\.Holden\-Day, 1970\.
- \[9\]Daqing Chen\.Online Retail II\.UCI Machine Learning Repository,[https://doi\.org/10\.24432/C5CG6D](https://doi.org/10.24432/C5CG6D), 2012\.
- \[10\]Mouxiang Chen, Lefei Shen, Zhuo Li, Xiaoyun Joy Wang, Jianling Sun, and Chenghao Liu\.VisionTS: Visual Masked Autoencoders Are Free\-Lunch Zero\-Shot Time Series Forecasters\.InICML, 2025\.
- \[11\]Si\-An Chen, Chun\-Liang Li, Nate Yoder, Sercan O\. Arik, and Tomas Pfister\.TSMixer: An All\-MLP Architecture for Time Series Forecasting\.TMLR, 2023\.
- \[12\]Citi Bike system data\.[https://citibikenyc\.com/system\-data](https://citibikenyc.com/system-data), 2026\.
- \[13\]Ben Cohen, Emaad Khwaja, Kan Wang, Charles Masson, Elise Ramé, Youssef Doubli, and Othmane Abou\-Amal\.Toto: Time Series Optimized Transformer for Observability\.[https://arxiv\.org/abs/2407\.07874](https://arxiv.org/abs/2407.07874), 2024\.
- \[14\]J\. D\. Croston\.Forecasting and Stock Control for Intermittent Demands\.Operational Research Quarterly, 23\(3\):289–303, 1972\.
- \[15\]Damai Dai, Chengqi Deng, Chenggang Zhao, R\. X\. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y\. Wu, Zhenda Xie, Y\. K\. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang\.DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture\-of\-Experts Language Models\.InACL, pages 1280–1297, 2024\.
- \[16\]Abhimanyu Das, Weihao Kong, Andrew Leach, Shaan Mathur, Rajat Sen, and Rose Yu\.Long\-term Forecasting with TiDE: Time\-series Dense Encoder\.TMLR, 2023\.
- \[17\]Abhimanyu Das, Weihao Kong, Rajat Sen, and Yichen Zhou\.A decoder\-only foundation model for time\-series forecasting\.InICML, 2024\.
- \[18\]Datadog\.BOOM: Benchmark of observability metrics\.[https://huggingface\.co/datasets/Datadog/BOOM](https://huggingface.co/datasets/Datadog/BOOM), 2025\.
- \[19\]Divvy system data\.[https://divvybikes\.com/system\-data](https://divvybikes.com/system-data), 2026\.
- \[20\]EINK\.CleanTS\-65M\.Hugging Face model,[https://huggingface\.co/EINK/CleanTS\-65M](https://huggingface.co/EINK/CleanTS-65M), 2026\.
- \[21\]Vijay Ekambaram, Arindam Jati, Pankaj Dayama, Sumanta Mukherjee, Nam H\. Nguyen, Wesley M\. Gifford, Chandra Reddy, and Jayant Kalagnanam\.Tiny Time Mixers \(TTMs\): Fast Pre\-trained Models for Enhanced Zero/Few\-Shot Forecasting of Multivariate Time Series\.InNeurIPS, 2024\.
- \[22\]Kun Feng, Shaocheng Lan, Yuchen Fang, Wenchao He, Sihan Lu, Shuqi Gu, Lintao Ma, Xingyu Lu, and Kan Ren\.Kairos: Toward Adaptive and Parameter\-Efficient Time Series Foundation Models\.[https://arxiv\.org/abs/2509\.25826](https://arxiv.org/abs/2509.25826), 2025\.
- \[23\]Xinghong Fu, Yanhong Li, Georgios Papaioannou, and Yoon Kim\.Reverso: Efficient Time Series Foundation Models for Zero\-shot Forecasting\.[https://arxiv\.org/abs/2602\.17634](https://arxiv.org/abs/2602.17634), 2026\.
- \[24\]Azul Garza, Cristian Challu, and Max Mergenthaler\-Canseco\.TimeGPT\-1\.[https://arxiv\.org/abs/2310\.03589](https://arxiv.org/abs/2310.03589), 2023\.
- \[25\]Rakshitha Godahewa, Christoph Bergmeir, Geoffrey I\. Webb, Rob J\. Hyndman, and Pablo Montero\-Manso\.Monash Time Series Forecasting Archive\.InNeurIPS Datasets and Benchmarks Track, 2021\.
- \[26\]Mononito Goswami, Konrad Szafer, Arjun Choudhry, Yifu Cai, Shuo Li, and Artur Dubrawski\.MOMENT: A Family of Open Time\-series Foundation Models\.InICML, 2024\.
- \[27\]Lars Graf, Thomas Ortner, Stanisław Woźniak, and Angeliki Pantazi\.FlowState: Sampling\-Rate\-Equivariant Time\-Series Forecasting\.InICML, 2026\.
- \[28\]Carlos Quesada Granja, Cruz Enrique Borges Hernández, Leire Astigarraga, and Chris Merveille\.GoiEner smart meters data\.Zenodo,[https://doi\.org/10\.5281/zenodo\.7362094](https://doi.org/10.5281/zenodo.7362094), 2022\.
- \[29\]Albert Gu and Tri Dao\.Mamba: Linear\-Time Sequence Modeling with Selective State Spaces\.InCOLM, 2024\.
- \[30\]Tao Hong and Shu Fan\.Probabilistic Electric Load Forecasting: A Tutorial Review\.International Journal of Forecasting, 32\(3\):914–938, 2016\.
- \[31\]Shi Bin Hoo, Samuel Müller, David Salinas, and Frank Hutter\.From Tables to Time: Extending TabPFN\-v2 to Time Series Forecasting\.[https://arxiv\.org/abs/2501\.02945](https://arxiv.org/abs/2501.02945), 2025\.
- \[32\]International Monetary Fund\.IMF PortWatch\.[https://portwatch\.imf\.org/](https://portwatch.imf.org/), 2026\.
- \[33\]Kaggle\.Walmart recruiting – store sales forecasting\.[https://www\.kaggle\.com/competitions/walmart\-recruiting\-store\-sales\-forecasting](https://www.kaggle.com/competitions/walmart-recruiting-store-sales-forecasting), 2014\.
- \[34\]Kaggle\.Rossmann store sales\.[https://www\.kaggle\.com/competitions/rossmann\-store\-sales](https://www.kaggle.com/competitions/rossmann-store-sales), 2015\.
- \[35\]Kaggle\.Web traffic time series forecasting\.[https://www\.kaggle\.com/competitions/web\-traffic\-time\-series\-forecasting](https://www.kaggle.com/competitions/web-traffic-time-series-forecasting), 2017\.
- \[36\]Kaggle\.Corporación favorita grocery sales forecasting\.[https://www\.kaggle\.com/competitions/favorita\-grocery\-sales\-forecasting](https://www.kaggle.com/competitions/favorita-grocery-sales-forecasting), 2018\.
- \[37\]Kaggle\.Store item demand forecasting challenge\.[https://www\.kaggle\.com/competitions/demand\-forecasting\-kernels\-only](https://www.kaggle.com/competitions/demand-forecasting-kernels-only), 2018\.
- \[38\]Kaggle\.Food demand forecasting\.Kaggle dataset,[https://www\.kaggle\.com/datasets/kannanaikkal/food\-demand\-forecasting](https://www.kaggle.com/datasets/kannanaikkal/food-demand-forecasting), 2026\.
- \[39\]Seunghan Lee, Juri Hong, Kibok Lee, and Taeyoung Park\.Sequential Order\-Robust Mamba for Time Series Forecasting\.InNeurIPS 2024 TSALM Workshop, 2024\.
- \[40\]Seunghan Lee, Jaehoon Lee, Jun Seo, Junhyeok Kang, Sangjun Han, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Soonyoung Lee, and Wonbin Ahn\.Beyond Magnitude and Shape: A Direction\-Aware Loss for Time Series Forecasting\.[https://arxiv\.org/abs/2608\.01857](https://arxiv.org/abs/2608.01857), 2026\.
- \[41\]Seunghan Lee, Jaehoon Lee, Jun Seo, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Minjae Kim, Sungdong Yoo, Junhyeok Kang, Sangjun Han, Soonyoung Lee, and Wonbin Ahn\.EXAONE Finance 1\.0: An Attention\-free Time Series Foundation Model for Financial Time Series\.[https://arxiv\.org/abs/2609\.04239](https://arxiv.org/abs/2609.04239), 2026\.
- \[42\]Seunghan Lee, Kibok Lee, and Taeyoung Park\.ANT: Adaptive Noise Schedule for Time Series Diffusion Models\.InNeurIPS, 2024\.
- \[43\]Seunghan Lee, Taeyoung Park, and Kibok Lee\.Learning to Embed Time Series Patches Independently\.InICLR, 2024\.
- \[44\]Seunghan Lee, Taeyoung Park, and Kibok Lee\.Dataset\-Driven Channel Masks in Transformers for Multivariate Time Series\.InICASSP, 2026\.
- \[45\]Xu Liu, Juncheng Liu, Gerald Woo, Taha Aksu, Yuxuan Liang, Roger Zimmermann, Chenghao Liu, Junnan Li, Silvio Savarese, Caiming Xiong, and Doyen Sahoo\.Moirai\-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts\.InICML, 2025\.
- \[46\]Yong Liu, Guo Qin, Zhiyuan Shi, Zhi Chen, Caiyin Yang, Xiangdong Huang, Jianmin Wang, and Mingsheng Long\.Sundial: A Family of Highly Capable Time Series Foundation Models\.InICML, 2025\.
- \[47\]Yong Liu, Haoran Zhang, Chenyu Li, Xiangdong Huang, Jianmin Wang, and Mingsheng Long\.Timer: Generative Pre\-trained Transformers Are Large Time Series Models\.InICML, 2024\.
- \[48\]Donghao Luo and Xue Wang\.ModernTCN: A Modern Pure Convolution Structure for General Time Series Analysis\.InICLR, 2024\.
- \[49\]Ernesto Aguilar Madrid\.Short\-term electricity load forecasting \(Panama case study\)\.Mendeley Data,[https://doi\.org/10\.17632/byx7sztj59\.1](https://doi.org/10.17632/byx7sztj59.1), 2021\.
- \[50\]Spyros Makridakis, Evangelos Spiliotis, and Vassilios Assimakopoulos\.The M4 competition: 100,000 time series and 61 forecasting methods\.International Journal of Forecasting, 36\(1\):54–74, 2020\.
- \[51\]Spyros Makridakis, Evangelos Spiliotis, and Vassilios Assimakopoulos\.The M5 competition: Background, organization, and implementation\.International Journal of Forecasting, 38\(4\):1325–1336, 2022\.
- \[52\]Vladyslav Moroshan, Julien Siems, Arber Zela, Timur Carstensen, and Frank Hutter\.TempoPFN: Synthetic Pre\-training of Linear RNNs for Zero\-shot Time Series Forecasting\.[https://arxiv\.org/abs/2510\.25502](https://arxiv.org/abs/2510.25502), 2025\.
- \[53\]New York City Taxi and Limousine Commission\.TLC trip record data\.[https://www\.nyc\.gov/site/tlc/about/tlc\-trip\-record\-data\.page](https://www.nyc.gov/site/tlc/about/tlc-trip-record-data.page), 2026\.
- \[54\]NHS Business Services Authority\.Prescribing and dispensing open data\.[https://opendata\.nhsbsa\.net/](https://opendata.nhsbsa.net/), 2026\.
- \[55\]Yuqi Nie, Nam H\. Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam\.A Time Series is Worth 64 Words: Long\-term Forecasting with Transformers\.InICLR, 2023\.
- \[56\]NYC Open Data\.311 service requests from 2020 to present\.[https://data\.cityofnewyork\.us/Social\-Services/311\-Service\-Requests\-from\-2020\-to\-Present/erm2\-nwe9](https://data.cityofnewyork.us/Social-Services/311-Service-Requests-from-2020-to-Present/erm2-nwe9), 2026\.
- \[57\]Open Power System Data\.[https://open\-power\-system\-data\.org/](https://open-power-system-data.org/), 2020\.
- \[58\]Kashif Rasul, Arjun Ashok, Andrew Robert Williams, Hena Ghonia, Rishika Bhagwatkar, Arian Khorasani, Mohammad Javad Darvishi Bayazi, George Adamopoulos, Roland Riachi, Nadhir Hassen, Marin Biloš, Sahil Garg, Anderson Schneider, Nicolas Chapados, Alexandre Drouin, Valentina Zantedeschi, Yuriy Nevmyvaka, and Irina Rish\.Lag\-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting\.[https://arxiv\.org/abs/2310\.08278](https://arxiv.org/abs/2310.08278), 2023\.
- \[59\]Rohit Sahoo\.Superstore sales dataset\.Kaggle dataset,[https://www\.kaggle\.com/datasets/rohitsahoo/sales\-forecasting](https://www.kaggle.com/datasets/rohitsahoo/sales-forecasting), 2026\.
- \[60\]David Salinas, Valentin Flunkert, Jan Gasthaus, and Tim Januschowski\.DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks\.International Journal of Forecasting, 36\(3\):1181–1191, 2020\.
- \[61\]Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean\.Outrageously Large Neural Networks: The Sparsely\-Gated Mixture\-of\-Experts Layer\.InICLR, 2017\.
- \[62\]Oleksandr Shchur, Abdul Fatir Ansari, Caner Turkmen, Lorenzo Stella, Nick Erickson, Pablo Guerron, Michael Bohlke\-Schneider, and Yuyang Wang\.fev\-bench: A Realistic Benchmark for Time Series Forecasting\.[https://arxiv\.org/abs/2509\.26468](https://arxiv.org/abs/2509.26468), 2025\.
- \[63\]Xiaoming Shi, Shiyu Wang, Yuqi Nie, Dianqi Li, Zhou Ye, Qingsong Wen, and Ming Jin\.Time\-MoE: Billion\-Scale Time Series Foundation Models with Mixture of Experts\.InICLR, 2025\.
- \[64\]Jeremy Stanley\.The Instacart online grocery shopping dataset 2017\.[https://tech\.instacart\.com/3\-million\-instacart\-orders\-open\-sourced\-d40d29ead6f2](https://tech.instacart.com/3-million-instacart-orders-open-sourced-d40d29ead6f2), 2017\.
- \[65\]Aris A\. Syntetos and John E\. Boylan\.On the Bias of Intermittent Demand Estimates\.International Journal of Production Economics, 71\(1\-3\):457–466, 2001\.
- \[66\]Aris A\. Syntetos, John E\. Boylan, and J\. D\. Croston\.On the categorization of demand patterns\.Journal of the Operational Research Society, 56\(5\):495–503, 2005\.
- \[67\]Artur Trindade\.ElectricityLoadDiagrams20112014\.UCI Machine Learning Repository,[https://doi\.org/10\.24432/C58C86](https://doi.org/10.24432/C58C86), 2015\.
- \[68\]U\.S\. Energy Information Administration\.Hourly electric grid monitor \(form EIA\-930\)\.[https://www\.eia\.gov/electricity/gridmonitor/](https://www.eia.gov/electricity/gridmonitor/), 2026\.
- \[69\]VN1 forecasting accuracy challenge\.[https://www\.datasource\.ai/en/home/data\-science\-competitions\-for\-startups/phase\-2\-vn1\-forecasting\-accuracy\-challenge/description](https://www.datasource.ai/en/home/data-science-competitions-for-startups/phase-2-vn1-forecasting-accuracy-challenge/description), 2024\.
- \[70\]Xue Wang, Tian Zhou, Jinyang Gao, Bolin Ding, and Jingren Zhou\.Output Scaling: YingLong\-Delayed Chain of Thought in a Large Pretrained Time Series Forecasting Model\.[https://arxiv\.org/abs/2506\.11029](https://arxiv.org/abs/2506.11029), 2025\.
- \[71\]Zihan Wang, Fanheng Kong, Shi Feng, Ming Wang, Xiaocui Yang, Han Zhao, Daling Wang, and Yifei Zhang\.Is Mamba Effective for Time Series Forecasting?Neurocomputing, 619:129178, 2025\.
- \[72\]Gerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong, Silvio Savarese, and Doyen Sahoo\.Unified Training of Universal Time Series Forecasting Transformers\.InICML, 2024\.
- \[73\]Shifeng Xie, Vasilii Feofanov, Jianfeng Zhang, Themis Palpanas, and Ievgen Redko\.CauKer: Classification Time Series Foundation Models Can Be Pretrained on Synthetic Data\.InICLR, 2026\.
- \[74\]Ted Zadouri, Ahmet Üstün, Arash Ahmadian, Beyza Ermiş, Acyr Locatelli, and Sara Hooker\.Pushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction Tuning\.InICLR, 2024\.
- \[75\]Ailing Zeng, Muxi Chen, Lei Zhang, and Qiang Xu\.Are Transformers Effective for Time Series Forecasting?InAAAI, 2023\.
- \[76\]Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang\.Informer: Beyond Efficient Transformer for Long Sequence Time\-Series Forecasting\.InAAAI, volume 35, pages 11106–11115, 2021\.
- \[77\]Zhuohang Zhu, Haodong Chen, Qiang Qu, and Vera Chung\.FinCast: A Foundation Model for Financial Time\-Series Forecasting\.InCIKM, pages 4539–4549, 2025\.
## Appendix AHyperparameters
This appendix records every setting behind the runs of §[5](https://arxiv.org/html/2609.30880#S5), so that the model of §[4](https://arxiv.org/html/2609.30880#S4)can be reproduced from the description alone\. Table[10](https://arxiv.org/html/2609.30880#A1.T10)lists them in four groups, namely the frozen backbone, the adapter of Equation[11](https://arxiv.org/html/2609.30880#S4.E11), the router and its supervision, and the optimisation\. Every value is held fixed across runs, and the data mixture of §[5\.7](https://arxiv.org/html/2609.30880#S5.SS7)is the one axis that we sweep across the whole study, at four shares of synthetic series and three seeds for each\.
Table 10:Hyperparameters of EXAONE Demand\.The backbone group describes the frozen model that we adapt, and the remaining three groups describe what we add to it and how it is trained on the demand corpus\.GroupSettingValueBackbone\(frozen\)Blocks24Model width1,024Embedding width2,048Attention heads16Context length8,192Input and output patch16Output patches per step4Quantile levels21, from0\.010\.01to0\.990\.99AdapterAdapted projectionsWq,Wk,Wv,WoW\_\{q\},W\_\{k\},W\_\{v\},W\_\{o\}, and the two feed\-forward projectionsAdapted layers24×6=14424\\times 6=144BranchesEE5, namely 1 shared and 4 routedShared rankr0r\_\{0\}16Routed rankr1…r4r\_\{1\}\\dots r\_\{4\}4Scaleα\\alpha, dropout8,0\.050\.05Router andsupervisionRouter widthhh32Router temperatureτ\\tau1\.01\.0Router learning\-rate multiplier100Shared weightww0\.50\.5Soft cutsτa,τc\\tau\_\{a\},\\tau\_\{c\}0\.350\.35,0\.600\.60SharpeningTT0\.50\.5Auxiliary weightλ\\lambda0\.10\.1OptimisationSteps9,000Learning rate1\.5×10−51\.5\\times 10^\{\-5\}Weight decay0\.10\.1Warm\-up steps300Batch size per device12, accumulated over 4 stepsDevices and precision8 GPUs with distributed data parallel \(DDP\),bf16mixedValidation intervalEvery 1,500 stepsThe backbone is loaded from its zero\-shot checkpoint, and no weight of it receives a gradient at any point\. The router learning\-rate multiplier deserves a note of its own, because the router is the one part of the model that has to move far from its initialisation within the same budget as the branches\. At the base learning rate it barely leaves the uniform mixture, which produces the collapse described in §[4\.5](https://arxiv.org/html/2609.30880#S4.SS5)without any other symptom in training\.
## Appendix BRelated Work
Demand forecasting\.Demand forecasting has a literature of its own, in which intermittent demand is handled by estimators that forecast the size of a non\-zero demand and the interval between two of them separately\[[14](https://arxiv.org/html/2609.30880#bib.bib14),[65](https://arxiv.org/html/2609.30880#bib.bib65)\], and a series is placed in one of four classes by its average inter\-demand interval and by the dispersion of its non\-zero values\[[66](https://arxiv.org/html/2609.30880#bib.bib66)\]\. Classical statistical models fit one estimator per series\[[8](https://arxiv.org/html/2609.30880#bib.bib8)\], while forecasting competitions and deep global models fit one model to a collection of related series\[[51](https://arxiv.org/html/2609.30880#bib.bib51),[60](https://arxiv.org/html/2609.30880#bib.bib60)\], and energy load forecasting has followed the same path\[[30](https://arxiv.org/html/2609.30880#bib.bib30)\]\. Our router is supervised by the classical taxonomy, and it carries the separation between classes into a foundation model without giving up the single set of weights that makes one useful in practice\.
General time series forecasting models\.Task\-specific forecasting models are trained on one dataset at a time, where patch\-based transformers made long contexts affordable\[[76](https://arxiv.org/html/2609.30880#bib.bib76),[55](https://arxiv.org/html/2609.30880#bib.bib55)\], and linear, convolutional, and state space models showed that much of the gain survives without attention\[[75](https://arxiv.org/html/2609.30880#bib.bib75),[11](https://arxiv.org/html/2609.30880#bib.bib11),[16](https://arxiv.org/html/2609.30880#bib.bib16),[43](https://arxiv.org/html/2609.30880#bib.bib43),[7](https://arxiv.org/html/2609.30880#bib.bib7),[48](https://arxiv.org/html/2609.30880#bib.bib48),[29](https://arxiv.org/html/2609.30880#bib.bib29),[71](https://arxiv.org/html/2609.30880#bib.bib71),[39](https://arxiv.org/html/2609.30880#bib.bib39)\]\. Around these models, prior work has studied what the loss should reward\[[40](https://arxiv.org/html/2609.30880#bib.bib40)\]and how to generate series\[[42](https://arxiv.org/html/2609.30880#bib.bib42)\], each on one dataset at a time\.
Time series foundation models\.TSFMs are pretrained on a corpus drawn from many domains and transfer to series that they have never seen\[[24](https://arxiv.org/html/2609.30880#bib.bib24),[4](https://arxiv.org/html/2609.30880#bib.bib4),[72](https://arxiv.org/html/2609.30880#bib.bib72),[17](https://arxiv.org/html/2609.30880#bib.bib17),[58](https://arxiv.org/html/2609.30880#bib.bib58),[47](https://arxiv.org/html/2609.30880#bib.bib47),[26](https://arxiv.org/html/2609.30880#bib.bib26),[63](https://arxiv.org/html/2609.30880#bib.bib63),[45](https://arxiv.org/html/2609.30880#bib.bib45),[21](https://arxiv.org/html/2609.30880#bib.bib21),[46](https://arxiv.org/html/2609.30880#bib.bib46),[13](https://arxiv.org/html/2609.30880#bib.bib13),[5](https://arxiv.org/html/2609.30880#bib.bib5),[3](https://arxiv.org/html/2609.30880#bib.bib3),[44](https://arxiv.org/html/2609.30880#bib.bib44)\], and they are ranked by benchmarks assembled in the same way\[[25](https://arxiv.org/html/2609.30880#bib.bib25),[2](https://arxiv.org/html/2609.30880#bib.bib2)\]\. Mixture\-of\-experts layers\[[61](https://arxiv.org/html/2609.30880#bib.bib61)\]have entered TSFMs as a way to scale capacity\[[63](https://arxiv.org/html/2609.30880#bib.bib63),[45](https://arxiv.org/html/2609.30880#bib.bib45)\], whereas we use a mixture of low\-rank experts\[[74](https://arxiv.org/html/2609.30880#bib.bib74)\]with a shared expert\[[15](https://arxiv.org/html/2609.30880#bib.bib15)\]to specialise one frozen backbone by demand behaviour\. Synthetic data is the standard answer to a corpus that cannot be collected at the required scale\[[4](https://arxiv.org/html/2609.30880#bib.bib4),[73](https://arxiv.org/html/2609.30880#bib.bib73)\], and specialising a general backbone to one domain has been carried furthest in finance\[[77](https://arxiv.org/html/2609.30880#bib.bib77),[41](https://arxiv.org/html/2609.30880#bib.bib41)\], where a companion effort to ours adapts the same backbone\. The present report is its demand\-side counterpart, where the properties that a general corpus under\-represents are intermittency, short histories, censoring, and exogenous events\.
## Appendix CConversion Schema and Verification
This appendix records the schema that every source is converted into, the per\-series index that is built alongside it, and the checks that a converted source is required to pass before it is admitted to the corpus, as referred to in §[2\.2](https://arxiv.org/html/2609.30880#S2.SS2)\.
All sources are converted to a single schema of\(series\_id, timestamp, target\), and we build a per\-series index next to the values themselves, since the checks below compare counts and need every series described first\. The index records three kinds of information for every series:
- •Shape\.Length, observation count, zero ratio, and missing ratio\.
- •Demand character\.Average inter\-demand interval, squared coefficient of variation, and the resulting Syntetos–Boylan class, all computed on the converted series rather than on the raw file as it was published\.
- •Provenance\.Source, domain, licence code, commercial\-use flag, and split role\.
Verification compares the converted counts against numbers establishedoutsidethe pipeline, such as competition documentation, API totals, or the published number of entities, rather than against the conversion log\. Four quantities are compared for every source, namely the number of series, the number of observations, the range of timestamps, and the number of distinct entities where the publisher states one\. A source is admitted to the corpus once all four agree with the published figures, and the per\-series index above is what the first two of them are counted from\.
Timestamps are taken from the columns that the source itself carries, and no index is synthesised from an assumed start date\. Where a source renames its columns between releases, as NHS prescribing data does in turningSTP\_CODEintoICB\_CODE, the names are unified before conversion so that every release reaches the same schema\.
## Appendix DDomain Breakdown of the Corpus
Figure[6](https://arxiv.org/html/2609.30880#S2.F6)shows the corpus by series count, and Table[11](https://arxiv.org/html/2609.30880#A4.T11)gives the same breakdown numerically while adding the observation count, which tells a different story\. The two views disagree because the domains differ in how their series are shaped, and a domain that is large by one of the two measures can be small by the other\.
Energy and transport holdfew but very longseries sampled at high frequency, such as meter readings and road sensors\. Retail and health holdmany but shortones, since demand there is recorded per SKU or per practice over a limited window\. Energy therefore takes 67\.5% of all observations from 39\.9% of the series, whereas retail takes 17\.2% of the series and only 1\.3% of the observations, a gap of fifty times on observations against two on series\.
How a domain is assigned\.Most sources carry one domain, but the six general\-purpose corpora of §[2\.2](https://arxiv.org/html/2609.30880#S2.SS2)hold many datasets under a single name, and a domain assigned to the corpus would say nothing about the series inside it\. The domain is recoverable from the series identifier, which keeps the name of its source dataset, and we map those 153 dataset names onto the same domain vocabulary that the rest of the corpus uses, so thatno series is left without a domain\. The mapping rests on three kinds of evidence: 1\) The top\-level folders of Time\-300B, which are already domain names such as energy, transport, and sales, 2\) the per\-dataset domain fields recorded when subsets of fev\-bench were selected, and 3\) the published description of each remaining dataset, read one at a time, which mapsdominickandrossmannto retail,electricityto energy, andazure\_vm\_tracesto compute\. The mapping changes only how series are reported in this appendix, and never which of them are used for training the model\.
Table 11:Corpus composition by domain\.Shares are over the whole converted corpus, including the synthetic portion\. Series that arrived inside the general\-purpose corpora are counted under the domain of the dataset they came from, not under the corpus that delivered them, so that the shares reflect the data rather than its packaging\.DomainSeriesShareObservationsShareEnergy4,507,84239\.87%32,646,256,62067\.51%Retail1,942,68217\.18%641,490,8111\.33%Transport1,578,37313\.96%9,742,740,72120\.15%Compute and cloud ops1,133,43110\.02%3,686,572,8727\.62%Web traffic880,7557\.79%1,348,375,7292\.79%Health817,9017\.23%32,301,7830\.07%Economics372,3393\.29%100,580,8990\.21%Synthetic47,5000\.42%59,638,4510\.12%Telecom20,0000\.18%90,925,3120\.19%Tourism3,5760\.03%460,3400\.00%Logistics2,0650\.02%5,790,2600\.01%Public services2760\.00%303,4490\.00%Total11,306,740100\.00%48,355,437,247100\.00%相似文章
APEX:一种面向无线边缘运营的网络原生时间序列基础模型,用于预测与异常检测
APEX是一个网络原生的解码器专用Transformer,针对无线边缘遥测数据的预测与异常检测而设计,预训练数据来自约4500个生产网络。在DHCP退化基准测试中,其MAE比最佳通用时间序列基础模型低18%,并能在边缘硬件上实现亚秒级推理。
K-EXAONE 2.0 技术报告
K-EXAONE 2.0 是 LG AI Research 推出的开源权重多语言 MoE 基础模型,总参数达 750B,激活参数为 37B,支持 10 种语言和 256K 上下文,在智能体编码、长上下文理解和安全性方面均有显著提升。
K-EXAONE 2.0 Technical Report
LG AI Research presents K-EXAONE 2.0, a 750B-parameter MoE foundation model upcycled from K-EXAONE, supporting 256K context and six languages, with self-speculative decoding for efficient inference.
Cadence:基于时间序列基础模型的需求时间序列有界误差有损压缩
Cadence 提出了一种使用基础模型和自适应算术编码器的有界误差有损压缩方法,用于时间序列,相比经典预测器实现了显著性能提升。
统一零样本时间序列预测:Darts基础
Darts,一个广受欢迎的开源Python时间序列分析库,引入了一个统一的FoundationModel类集合,该集合整合了多种时间序列基础模型(Chronos-2、TimesFM 2.5、TiRex、PatchTST-FM),通过标准化接口和最小依赖实现零样本和微调预测。