When Do Corrective Features Help? An Agent for Corrective Feature Discovery on Black-Box Forecasters
Summary
The paper introduces Crafter, an agent for corrective feature discovery that mines the residual of frozen black-box forecasters using compositional search and LLM-generated features, achieving up to 27% error reduction across six datasets and backbones.
View Cached Full Text
Cached at: 08/07/26, 07:48 AM
# When Do Corrective Features Help? An Agent for Corrective Feature Discovery on Black-Box Forecasters
Source: [https://arxiv.org/html/2608.05207](https://arxiv.org/html/2608.05207)
Fangxin Wang1,Ziyi Zhang2,Diyi Zhuang3,Langzhou He1,Shiyu Wang, Baichuan Mo4,Philip S\. Yu1 1University of Illinois Chicago,2Texas A&M University, 3Massachusetts Institute of Technology,4Tsinghua University
###### Abstract
Frozen pretrained forecasters often fail in structured, recurring ways that are costly to repair through fine\-tuning\. We study*corrective feature discovery*: mining interpretable features of the frozen forecaster’s residual to drive a lightweight post\-hoc corrector\. Prior automated feature engineering models the data\-generating process; corrective features instead model the*model\-failure*process\. We presentCrafter\(Corrective Residual Agent with Feature\-based Temporal Exploration and Reasoning\), which keeps the backbone frozen and mines its residual with two generators of different character: a compositional search over the raw input channels, and a large language model \(LLM\) that proposes named feature combinations, binary flags, and short executable code\. A single validation\-grounded gate accepts or rejects every candidate blind to its origin, and a validation\-selected corrector applies the survivors or leaves the forecast unchanged\. The same source\-blindness lets prior feature\-engineering systems run through the identical pipeline, soCrafterdoubles as an instrument that attributes forecast changes to the feature source alone\. Across six public datasets and six frozen backbones,Craftersurpasses every dedicated feature\-engineering system at every feature budget, roughly doubling the corrector\-only lift and cutting the weakest backbones’ error by up to27%27\\%\. The gains are robust to the LLM backend and persist on top of fine\-tuned backbones\.
When Do Corrective Features Help? An Agent for Corrective Feature Discovery on Black\-Box Forecasters
## 1Introduction
Large\-scale forecasting is now served by pretrained backbones deployed*frozen*\(Liuet al\.,[2024](https://arxiv.org/html/2608.05207#bib.bib9); Ansariet al\.,[2024](https://arxiv.org/html/2608.05207#bib.bib10); Liuet al\.,[2025a](https://arxiv.org/html/2608.05207#bib.bib12)\)\. An e\-commerce platform or grocery chain predicts sales for hundreds of thousands of store–product series from a single model, because fine\-tuning one per item or regime is infeasible\. The frozen model still fails in structured, recurring ways: it misses promotions, lags regime shifts, and underprices event spikes\. These errors are business\-critical and must be interpretable to act on, so the question is not how to train a better forecaster but how to*correct*a deployed one cheaply, without modifying it\.
A natural way to correct a fixed forecaster is to add external, interpretable features around it\. This connects the problem to automated feature engineering, but with a different target: once the backbone is frozen, the object of discovery changes from the target process itself to the residual of a particular deployed model\. The relevant question is not which regularities help predict the future, but which remain unmodeled by this backbone\. We therefore study*corrective feature discovery*: automatically finding interpretable features of the residual, the signal of where and how a particular frozen forecaster fails\. The residual is not leftover noise a priori—it carries structure induced by the backbone’s blind spots, concentrated around promotions, rare events, and local regimes—so a corrective feature is model\-dependent and doubles as a diagnosis: each accepted feature names a mechanism the deployed model misses\.
Two properties make the discovery hard\. First, the most valuable corrective features are*named*from domain semantics rather than enumerated from the raw schema; where they matter, a statistical feature bank and a syntactic search improve the forecast only modestly, while named features close most of the remaining gap \(§[5\.1](https://arxiv.org/html/2608.05207#S5.SS1)\)\. Second, a feature’s value is*regime\-dependent*: it concentrates where the backbone errs and vanishes where the backbone is already accurate, so the same source helps on one backbone–dataset cell and is inert or harmful on another\. Whether a feature source or corrector configuration helps therefore depends on the regime, and this dependence has not been characterized systematically\.
Crafteraddresses the method and the characterization together\. It keeps the backbone frozen and mines the residual with two generators of different character: a compositional search over the raw channels, and an LLM that proposes named combinations, binary flags, and short executable code\. A single validation\-grounded gate admits candidates blind to their origin, and a validation\-selected corrector applies the survivors or leaves the forecast unchanged when none help\. This source\-blind design merges two heterogeneous generators into one method, and also turns the pipeline into an instrument for fairly evaluating external feature\-engineering system\.
We evaluateCrafteron six public datasets and six frozen backbone families using a rolling\-origin, multi\-seed protocol\. All methods are run in the same evaluation harness and, where applicable, with the same LLM backend\. Across this controlled comparison,Crafterconsistently outperforms three dedicated feature\-engineering systems\. This advantage is robust to feature budget, remains under a second LLM backend, and also holds when the corrected backbone is fine\-tuned rather than used zero\-shot\. On the two datasets where we test temporal reuse, the discovered features transfer across backtest windows\. Finally, we identify the limits of residual correction: in saturated cells, no feature source yields a reliable improvement\.
#### Contributions\.
\(C1\) A source\-blind framework for corrective feature discovery\.We formulate feature discovery for a frozen forecaster as residual correction, and instantiate it with a single validation\-grounded gate that admits heterogeneous sources blind to their origin\. The framework is both a deployable correction layer and an evaluation instrument: any external feature\-engineering system runs through the identical gate, corrector, and budget, so a difference in the forecast is attributable to the feature source alone\.\(C2\) A characterization of when corrective features help\.Across six datasets and six backbone families we map where residual correction is beneficial, inert, or harmful\. Correction pays in proportion to the residual’s exploitable structure: large on weak backbones with informative forecast covariates, where LLM\-named features are the differentiator, and neutral once a strong backbone saturates the series\. Budget amplifies that margin rather than creating it\.
## 2Related Work
#### Frozen foundation forecasters\.
Pretrained time\-series models are often deployed zero\-shot, with a single backbone served frozen across many series\(Ansariet al\.,[2024](https://arxiv.org/html/2608.05207#bib.bib10); Liuet al\.,[2024](https://arxiv.org/html/2608.05207#bib.bib9),[2025a](https://arxiv.org/html/2608.05207#bib.bib12)\)\. This avoids per\-series fine\-tuning, but the frozen model leaves a backbone\-specific residual it cannot remove on its own\. We take that residual as the object to model and leave the backbone untouched, consistent with broader lightweight post\-training approaches that add capabilities to frozen foundation models without updating their pretrained weights\(Houlsbyet al\.,[2019](https://arxiv.org/html/2608.05207#bib.bib38); Huet al\.,[2021](https://arxiv.org/html/2608.05207#bib.bib37); Heet al\.,[2025](https://arxiv.org/html/2608.05207#bib.bib36)\)\.
#### Automated and LLM\-based feature engineering\.
A long line of work builds features to improve a trainable predictor\. Statistical banks extract many candidate features and select among them\(Christet al\.,[2018](https://arxiv.org/html/2608.05207#bib.bib17); Cerqueiraet al\.,[2020](https://arxiv.org/html/2608.05207#bib.bib1); Costa,[2021](https://arxiv.org/html/2608.05207#bib.bib2)\), and search\-based engines compose operators over the raw schema with reinforcement learning\(Khuranaet al\.,[2018](https://arxiv.org/html/2608.05207#bib.bib18)\), Monte\-Carlo tree search\(Huanget al\.,[2022](https://arxiv.org/html/2608.05207#bib.bib19)\), or expert\-level generation\(Zhanget al\.,[2022](https://arxiv.org/html/2608.05207#bib.bib16)\)\. A recent line prompts an LLM to write feature code from dataset context\(Hollmannet al\.,[2023](https://arxiv.org/html/2608.05207#bib.bib3); Namet al\.,[2024](https://arxiv.org/html/2608.05207#bib.bib4); Abhyankaret al\.,[2025](https://arxiv.org/html/2608.05207#bib.bib5); Murrayet al\.,[2025](https://arxiv.org/html/2608.05207#bib.bib8)\), and AutoCT pairs an LLM with tree search for tabular features\(Liuet al\.,[2025b](https://arxiv.org/html/2608.05207#bib.bib20)\)\. Two things separate this work from ours\. Its objective is*predictive*: the features train a downstream model rather than explain a fixed one\. And its strongest LLM methods target*tabular*data\. We run three of these systems as feature generators inside our own pipeline \(§[4](https://arxiv.org/html/2608.05207#S4)\), which puts every source on one footing\.
#### Correcting a frozen forecaster\.
Modeling the error of a fixed forecaster with a second learner is a classic idea\(Zhang,[2003](https://arxiv.org/html/2608.05207#bib.bib21)\)\. Recent methods correct a deployed forecaster without retraining it\. Post\-Training Corrections picks a sequence of corrections from a fixed pool\(Cherkaouiet al\.,[2025](https://arxiv.org/html/2608.05207#bib.bib23)\)\. Theδ\\delta\-Adapter learns a sparse input mask and a bounded residual head, and uses no LLM\(Lianget al\.,[2026](https://arxiv.org/html/2608.05207#bib.bib24)\)\. Other methods adapt a frozen forecaster online or through a memory module\(Lyuet al\.,[2026](https://arxiv.org/html/2608.05207#bib.bib25); Daiet al\.,[2026a](https://arxiv.org/html/2608.05207#bib.bib26)\), and AutoGluon\-TimeSeries regresses on raw covariates inside its pipeline\(Shchuret al\.,[2023](https://arxiv.org/html/2608.05207#bib.bib22)\)\. These methods correct the forecast, but they act on inputs that are already given\. None*discovers*new interpretable features of the residual\.
#### Discovering residual features to correct\.
Closest to our setting is NSR\-Boost, which keeps a model frozen and uses an LLM to write symbolic residual experts for an aggregator to combine\(Daiet al\.,[2026b](https://arxiv.org/html/2608.05207#bib.bib27)\)\. It is tabular, it uses the LLM as its only generator, and it runs no search\. Work on forecasting explanation and interpretable correction analyzes or adjusts model behavior, but does not discover corrective features\(Wanget al\.,[2023](https://arxiv.org/html/2608.05207#bib.bib6); Lopez and Sobieczky,[2024](https://arxiv.org/html/2608.05207#bib.bib7)\)\.Crafterdiffers in three ways\. It mines the residual of a frozen*forecaster*\. It draws candidates from two different generators, a compositional search and an LLM, and admits them through one validation gate that is blind to their source\. Because the gate is source\-blind, the same pipeline also measures*which*source and configuration help in*which*regime, which the work above does not study\.
Figure 1:TheCrafterframework\.The forecaster stays frozen\. Two generators mine its residual for structured feature specifications: a compositional Monte\-Carlo tree search and an LLM proposer, optionally coupled \(Crafter†\)\. A source\-blind gate admits a candidate only when it explains validation error the corrector cannot\. A validation\-selected corrector based on correlationρ\\rhoand budget then applies the survivors, or leaves the forecast unchanged\. The loop runsRRrounds\.
## 3Method
We presentCrafter\(CorrectiveResidualAgent withFeature\-basedTemporalExploration andReasoning\), an agent for*corrective feature discovery*\. It leaves the forecasting backbone frozen and mines interpretable features of its residual, the signal of*where*and*how*the backbone fails, to drive a lightweight post\-hoc corrector\. This section describes the framework and the single source\-blind gate at its core \(§[3\.1](https://arxiv.org/html/2608.05207#S3.SS1)\), its two feature generators,*Exploration*\(a compositional search\) and*Reasoning*\(a language model\), with their optional coupling \(§[3\.2](https://arxiv.org/html/2608.05207#S3.SS2)\), the gate \(§[3\.3](https://arxiv.org/html/2608.05207#S3.SS3)\), and the validation\-selected corrector \(§[3\.4](https://arxiv.org/html/2608.05207#S3.SS4)\)\.
### 3\.1Overview
Illustrated in Fig\.[1](https://arxiv.org/html/2608.05207#S2.F1),Craftercorrects a frozen forecaster without modifying it\. Given the backbone forecasty^\\hat\{y\}over a horizon, it mines interpretable features of the residual and fits a lightweight gradient\-boosted\-tree corrector on them, which then adjusts the forecast or leaves it unchanged\. Quality is measured by the weighted mean absolute percentage error \(wMAPE\) of the horizon total \(the corrector target and hyperparameters are in App\. A and D\)\. Covariates enter only as corrector features, never as backbone inputs; §[5\.4](https://arxiv.org/html/2608.05207#S5.SS4)shows this routing is what makes the gain possible\.
Every candidate feature is a*structured specification*: a typed record drawn from a fixed grammar \(App\. C\) that applies transforms and an operator to a few input channels and compiles to one numeric column\. Specifications are searchable, de\-duplicable across rounds, and interpretable by name, so a run can be audited feature by feature\.Craftermines the residual with two generators of different character—an atomic search over the raw channels and an LLM that proposes named features \(§[3\.2](https://arxiv.org/html/2608.05207#S3.SS2.SSS0.Px2)\)—and routes both into a single*source\-blind*acceptance gate \(§[3\.3](https://arxiv.org/html/2608.05207#S3.SS3)\) that keeps a candidate only when it explains validation error the corrector cannot already account for, whichever generator produced it\.
Crafterruns forRRrounds \(Alg\. S1, App\. B\)\. Each round fits the corrector, diagnoses where it still errs on a held\-out validation block, mines new candidates from both generators, and keeps those that pass the gate\. A validation\-selected corrector then applies the survivors or leaves the backbone unchanged when none help \(§[3\.4](https://arxiv.org/html/2608.05207#S3.SS4)\)\.
### 3\.2Feature generators
Craftermines the residual with two generators of different character\. A compositional search composes the raw channels into new features, and an LLM names mechanisms the raw schema does not contain\. Both feed candidates into the single gate of §[3\.3](https://arxiv.org/html/2608.05207#S3.SS3), which decides what survives\.
#### Compositional search\.
The first generator builds*compositional*features in the tradition of operator\-based automated feature engineering\(Cerqueiraet al\.,[2020](https://arxiv.org/html/2608.05207#bib.bib1); Costa,[2021](https://arxiv.org/html/2608.05207#bib.bib2); Zhanget al\.,[2022](https://arxiv.org/html/2608.05207#bib.bib16)\), namedatomicin this paper\. A feature applies a transform to one channel, or an operator to two transformed channels, drawn from a fixed vocabulary ofT=15T\{=\}15transforms andB=8B\{=\}8operators \(App\. C\)\. The depth\-one and depth\-two expressions number far too many to score exhaustively, about6×1056\{\\times\}10^\{5\}forC≈25C\{\\approx\}25channels, soCraftersearches\. It uses a Monte\-Carlo tree search\(Kocsis and Szepesvári,[2006](https://arxiv.org/html/2608.05207#bib.bib14)\)whose nodes are UCB1 bandits\(Aueret al\.,[2002](https://arxiv.org/html/2608.05207#bib.bib13)\)\. Channels are grouped into mechanism*families*such asvolatilityortrend; a family pools its channels, so one reward updates the whole mechanism and the search finds productive families from few samples before refining to a single channel\. A root\-to\-leaf path picks a family and arity, then the channels, then the operator, each node conditioned on the choices above it \(App\. C\)\. An arm is credited by two signals: a cheap proxy at scoring time, the rank correlationρ\\rhobetween the candidate and the still\-unexplained validation residual, and the realized gain in validation wMAPE after the corrector is refit\. The bandit selects arms by UCB1 and updates each arm’s value by the incremental sample mean of these rewards; the selection, update, and reward rules are Eq\. \(S3\)–\(S5\) in App\. D\. The realized gain ties the search to the deployed model: a feature is worth what keeping it actually buys\.
#### LLM\-proposed features\.
The search only recombines existing channels; the LLM*invents*features that name mechanisms the schema lacks\. Organized into roles and queried a few times per round, it reasons over the run’s own evidence, not generic knowledge alone: which features the gate accepted, which it rejected, and which mechanism families are still uncovered\. APlannerreads this state and proposes a few exploration*directions*\. ASpecifierturns each direction into candidate specs built from existing channels, surfacing semantic compositions a syntactic search would not prioritize: on theepfelectricity\-price series it composes arenewable\_penetration, renewable generation over total load, which captures the price suppression a covariate\-blind backbone misses\. ABase Proposergoes further and invents new named base columns for mechanisms the schema lacks, which join the channel pool and become searchable next round\. AReflectioncall revises the next round’s directions when few candidates survive the gate\. The LLM proposes three kinds of features: 1\)LLM\-combo: combinations of existing channels, 2\)LLM\-flag: binary flags, and 3\)LLM\-code: short executable code \(accepted examples of every kind in App\. K\)\. LLM only*proposes*: which candidates survive is decided by the gate of §[3\.3](https://arxiv.org/html/2608.05207#S3.SS3), not by the model\. All role prompts are in App\. H\.
#### Coupling the two generators\.
The two can also run as a coupled loop\. A new base column invented by the LLM, once it passes the gate, joins the channel pool, so the search can compose it with other channels in later rounds\. Conversely, each round’s merged accepted set, family coverage, and rejected candidates condition the next LLM proposal, steering it toward thin families and away from past rejects\. Decided by whether allowing coupling, we have bothCrafterandCrafter†\(coupled\) evaluated in the experiments\.
### 3\.3The source\-blind acceptance gate
Both generators feed one gate, and that gate is what makesCraftera single method rather than two pipelines\. A candidate is accepted only if it explains validation error the corrector cannot*already*account for: the Spearman correlation between the candidate and the unexplained validation residual must exceed a thresholdτ\\tau, the candidate must not be collinear with an already\-accepted feature, and per\-family and per\-round budgets must hold\. The same test applies to every candidate, compositional or semantic\. Deciding acceptance on one validation\-grounded scale makes the two generators commensurable without either reading the other’s internal scores: the merged pool self\-selects against the residual, not against any generator’s private confidence\.
The gate is agnostic to a feature’s origin, with a consequence beyond merging the two generators\. Any external feature\-engineering system runs through the identical pipeline by swapping only the generator, with the gate, corrector, and selection unchanged\.Crafterthen doubles as an*instrument*: a difference in the final forecast is attributable to the feature source, because nothing else moves \(§[5\.1](https://arxiv.org/html/2608.05207#S5.SS1)\)\. The same source\-blindness that unifies the two generators keeps that comparison fair\.
### 3\.4Validation\-selected corrector
The features that pass the gate feed a corrector\.Crafterfits a small*corrector set*—an identity \(None\) map plus an additive and a multiplicative gradient\-boosted corrector—and keeps the one with the best validation wMAPE\. Trees are chosen because the residual is nonlinear and regime\-dependent and because per\-feature importances keep the correction auditable \(App\. D\)\. TheNoneoption floors the risk*on validation*\(a round never ships a corrector worse there than leaving the backbone alone\), making correction an asymmetric bet\. This is a validation\-time safeguard, not a test guarantee: under rolling\-origin shift a low\-headroom cell can still regress out\-of\-sample \(e\.g\.,bizitobs/Chronos, §[5\.1](https://arxiv.org/html/2608.05207#S5.SS1)\)\.
\(a\)Correction before and after fine\-tuning\.
\(b\)Same\-budget fairness\.
Figure 2:Correction holds across training and feature budgets\.Mean lift in test wMAPE \(%, higher is better\) over the66datasets\. Colours are shared across panels; the twoCraftervariants share a hue\.\(a\)Correcting the zero\-shot backbone \(lift overRaw\) and the fine\-tuned one \(lift overFT\)\.Crafterhelps in both, with the largest gains on weak backbones\.\(b\)Lift overRawagainst budgetKK; “All” is uncapped\. BothCraftervariants beat the best external at everyKK; the no\-LLM corrector stays below\. Per\-cell results in App\. Table S12\.
## 4Experimental Setup
#### Datasets\.
We use six public forecasting datasets spanning the regimes where covariate information varies in kind and strength \(listed in the order used throughout\):epf\-de\(day\-ahead electricity prices; a single long series with*forecast\-type*renewable covariates\)\(Lagoet al\.,[2021](https://arxiv.org/html/2608.05207#bib.bib30)\), two promotion\-driven retail panels \(rossmann,rohlik\)\(Kaggle,[2015](https://arxiv.org/html/2608.05207#bib.bib33),[2024](https://arxiv.org/html/2608.05207#bib.bib35)\), an hourly business\-operations panel \(bizitobs\-l2c\)\(Palaskaret al\.,[2024](https://arxiv.org/html/2608.05207#bib.bib32)\), a store\-sales panel \(favorita\)\(Kaggle,[2017](https://arxiv.org/html/2608.05207#bib.bib34)\), and the M5 weekly competition data \(m5\)\(Makridakiset al\.,[2022](https://arxiv.org/html/2608.05207#bib.bib31)\)\. Together they cover single\- vs\. many\-series, covariate\-rich vs\. covariate\-poor, and weak\- vs\. saturated\-backbone settings\.
#### Frozen backbones\.
Our grid corrects six pretrained backbones used zero\-shot, listed in order of release: Timer\(Liuet al\.,[2024](https://arxiv.org/html/2608.05207#bib.bib9)\), Chronos\(Ansariet al\.,[2024](https://arxiv.org/html/2608.05207#bib.bib10)\), Moirai\(Wooet al\.,[2024](https://arxiv.org/html/2608.05207#bib.bib28)\), and Toto\(Cohenet al\.,[2024](https://arxiv.org/html/2608.05207#bib.bib29)\), followed by Chronos\-2\(Ansariet al\.,[2025](https://arxiv.org/html/2608.05207#bib.bib11)\)and Moirai\-2\.0\(Liuet al\.,[2025a](https://arxiv.org/html/2608.05207#bib.bib12)\), spanning the first foundation\-model wave to current leaders\. Chronos and Timer are primarily univariate or channel\-independent forecasters; Moirai and Toto can handle multivariate time\-series inputs; and Chronos\-2 explicitly exposes a covariate\-aware forecasting interface\. In every grid we nonetheless run*all*backbones covariate\-blind, so the backbone forecast never sees the covariates and every method corrects the*same*residual—any covariate value is then attributable to the corrector, not the backbone \(routing them into the backbone instead is at best neutral; §[5\.4](https://arxiv.org/html/2608.05207#S5.SS4)\)\. The grid is6×6=366\\times 6=36cells\.
#### Protocol\.
We use a rolling\-origin backtest: on cached forecasts we re\-cut three expanding train:val:test splits \(6:1:16\{:\}1\{:\}1,7:1:17\{:\}1\{:\}1,8:1:18\{:\}1\{:\}1\) with no test leakage, validate on three blocked CV windows, runR=3R\{=\}3rounds, and report last\-round test wMAPE averaged over the three splits and three seeds \(nine runs per cell\)\. We assess significance with a one\-sided Wilcoxon signed\-rank test over the paired cells and decompose run\-to\-run variance into seed and split components \(App\. D\); split variance dominates\. The public datasets may appear in either LLM’s pretraining corpus, but the leakage surface is small: the LLM emits only feature*specifications*and code reading past history and the backbone forecast, never test\-period values\. Because seeds do not change the cached forecasts, the grid is GPU\-free and cheap; compute and timing details are in App\. D\.
#### Baselines and backend\.
Every feature\-engineering method—bothCraftervariants and the external systems—uses the same GPT\-5\.2 backend and an identical harness \(same cached residual, same corrector set, same validation top\-KKselection\), so only the feature*generator*differs and the comparison is fair by construction\. We report twoCraftervariants, the two rightmost \(highlighted\) columns of App\. Table S12:Crafter, the full system, andCrafter†, which additionally enables the two source\-coupling channels of §[3\.2](https://arxiv.org/html/2608.05207#S3.SS2.SSS0.Px3)\. The baselines are: the frozen backbone \(Raw\); a backbone fine\-tune \(FT; a uniform head\-only fine\-tune of every backbone, with a stronger full\-parameter Moirai variant in App\. F\); a covariate\-only corrector \(Cov\) and the no\-LLM corrector \(Crafter\-noLLM\), which ablate the LLM and the search; and three external feature\-engineering systems adapted to our setting—TSFresh\(a statistical feature bank\)\(Christet al\.,[2018](https://arxiv.org/html/2608.05207#bib.bib17)\), CAAFE \(LLM\-written feature code\)\(Hollmannet al\.,[2023](https://arxiv.org/html/2608.05207#bib.bib3)\), and LLM\-FE \(LLM\-guided evolutionary search\)\(Abhyankaret al\.,[2025](https://arxiv.org/html/2608.05207#bib.bib5)\)\. App\. E states what each adaptation keeps and changes\.
## 5Results
Corrective feature discovery improves a frozen forecaster without retraining it, and the size of that improvement is governed by residual headroom rather than by backbone identity, feature budget, or LLM backend\. We evaluate on six datasets and six pretrained backbones under one rolling\-origin, multi\-seed protocol in which every feature\-engineering method passes the same validation gate and downstream corrector, so measured differences trace to the discovered features rather than to the corrector \(App\. D; full per\-cell grid in Table S12\)\. This section establishes five points:Crafteris state of the art against prior feature engineering and helps weak backbones most \(§[5\.1](https://arxiv.org/html/2608.05207#S5.SS1)\); the wrapper generalizes across backbone families \(§[5\.3](https://arxiv.org/html/2608.05207#S5.SS3)\) and survives non\-standard input paths \(§[5\.4](https://arxiv.org/html/2608.05207#S5.SS4)\); the decisive gain comes from the LLM \(§[5\.2](https://arxiv.org/html/2608.05207#S5.SS2)\); and the learned features transfer across time \(§[5\.5](https://arxiv.org/html/2608.05207#S5.SS5)\)\.
Table 1:Cross\-LLM robustness\.Per\-backbone summary under both LLM backends\. Both variants improve every backbone, most on the weak ones, so the backbone\-strength picture is LLM\-agnostic; dataset\-level increments still depend on the backend \(§[5\.2](https://arxiv.org/html/2608.05207#S5.SS2)\)\.Rawis the backbone’s mean zero\-shot wMAPE×100\\times 100over the66datasets\.Lift%is the mean reduction vs\.Raw;winscounts datasets \(of66\) where the variant matches or beats the best external\.

Figure 3:Feature provenance and transfer\.\(a\)Features proposed vs\. selected by source, with survival rate\. LLM kinds supply about three\-quarters of the selected set, and LLM\-code is almost always pruned\.\(b\)Features learned on one split and replayed on a later one track the target split’s own features along the diagonal\. The off\-diagonal points are the Moirai shift corner, mostly single\-seriesepf\(§[5\.5](https://arxiv.org/html/2608.05207#S5.SS5); App\. L\)\.
### 5\.1Corrective discovery improves frozen and fine\-tuned backbones
Figure[2\(a\)](https://arxiv.org/html/2608.05207#S3.F2.sf1)summarizes the main comparison across six datasets and six backbone families; the full per\-cell grid with external baselines is in App\. Table S12\. Under the shared harness of §[4](https://arxiv.org/html/2608.05207#S4), bothCraftervariants outperform each of the three dedicated feature\-engineering systems, with per\-baseline paired Wilcoxon tests significant atp<0\.01p\{<\}0\.01\. The result holds both when aggregating by dataset and when aggregating by backbone \(Figure[2\(a\)](https://arxiv.org/html/2608.05207#S3.F2.sf1)and Table[1](https://arxiv.org/html/2608.05207#S5.T1)\), showing that corrective discovery is not a single\-dataset artifact\.
The gain is not a feature\-budget artifact\.The same\-budget sweep rules the explanation thatCrafterwins simply because it is allowed to propose more candidates\. In Fig\.[2\(b\)](https://arxiv.org/html/2608.05207#S3.F2.sf2)\(per\-budget win rates and mean differences in App\. Table S11\),Crafterbeats the best external feature\-engineering system at every feature budgetKK\. Even atK=5K\{=\}5,Crafterexceeds the best external system at its own uncapped budget, and the margin saturates by moderate budgets rather than appearing only in the “All” setting\. The external systems’ curve is comparatively flat, indicating that their additional candidates do not translate into comparable residual\-correction gains under the shared gate and corrector\. Thus the main result reflects feature quality, not merely feature count\.
Fine\-tuning is not a substitute for residual correction\.We also compare against a uniform head\-only fine\-tune,FT, applied identically across backbones\. Fine\-tuning improves the frozen forecasts, but it does not remove the structured residuals thatCrafterexploits\. Correcting the zero\-shot backbone beats fine\-tuning alone on most cells, and applyingCrafteron top of the fine\-tuned backbone still yields additional lift overFT\(Fig\.[2\(a\)](https://arxiv.org/html/2608.05207#S3.F2.sf1), top vs\. bottom; the full per\-cell zero\-shot\-versus\-fine\-tuned grid is App\. Table S1\)\. The same conclusion holds for the stronger full\-parameter in\-domain fine\-tune of the covariate\-capable Moirai models, reported in App\. Table S2\. These results placeCrafteras a deployment\-time correction layer rather than a replacement for backbone adaptation: it improves a frozen model, and it remains useful after fine\-tuning when exploitable residual structure remains\.
### 5\.2Search matters through coupling, not only standalone lift
We next discuss where the gain in §[5\.1](https://arxiv.org/html/2608.05207#S5.SS1)comes from\. The source\-blind gate in the pipeline serve as an attribution device: external feature\-engineering systems, compositional search, and LLM\-generated candidates are all filtered by the same validation rule and passed to the same corrector\.
Search defines the enumerable feature space\.The compositional search adds little standalone lift on these public benchmarks \(the noLLM column of Table S12; a component\-level with/without\-LLM comparison is App\. Table S5\), but this null result is informative rather than incidental\. It shows what can be recovered by syntactic enumeration over the observed schema\. When this search and the external feature banks fail to close the residual gap, but named LLM features do, the missing signal is not merely another lag, rolling aggregate, or low\-order interaction\. It is a semantic feature that must be named in a form aligned with the backbone’s residual\. The search arm therefore acts as a control for the LLM comparison: it measures the enumerable floor, so the remaining gain is identified as a semantic feature gap rather than assumed to be one\.
The clearest case isepf\. External banks and syntactic search add little beyond the covariate\-only corrector, whileCrafterremoves much more of the frozen error using named forecast\-covariate features such asrenewable\_penetration\(theepfrows of Table S12; representative accepted features in App\. K\)\. This is the regime corrective discovery is designed for: the relevant information is present in the covariates, but it does not become useful until it is expressed as a feature that matches the backbone’s failure mode\. Similar, though smaller, effects appear on multi\-series retail panels, where domain\-named interactions generalize across related store–product series\.
Coupling turns search from a control into a mechanism\.The coupled variant,Crafter†, enables the two cross\-feeding generators in §[3\.2](https://arxiv.org/html/2608.05207#S3.SS2.SSS0.Px3): search results summarize which enumerable candidates survived the gate, and that feedback is exposed to the LLM in later proposal rounds\. Search alone adds little net lift, yet feeding its accepted and rejected candidates back to the LLM makesCrafter†stronger on balance: it improves onCrafteron more cells, gives per\-backbone lift at least as large on all but one backbone, and has its clearest gains onepf\(per\-backbone lift in Table[1](https://arxiv.org/html/2608.05207#S5.T1); per\-cell in Table S12\)\. Since the search contributes little alone, that gain is the cross\-feeding steering the LLM toward features it would not otherwise propose\. Its cost is variance, especially on volatile single\-series cells such asepf, where a small change in the feature set can move a cell in either direction \(seed versus split variance in App\. D\)\. We therefore reportCrafterandCrafter†as two endpoints of an explore–exploit trade\-off\.
Feature\-level evidence supports the same mechanism\.The provenance analysis in Fig\.[3](https://arxiv.org/html/2608.05207#S5.F3)a shows that most selected features come from LLM\-generated semantic sources, while atomic search features survive at a comparable rate but are limited by the smaller enumerable pool \(per\-source and per\-dataset counts in App\. J\)\. This matches the performance attribution: the LLM supplies most of the selected feature mass, while search marks the reliable enumerable candidates and provides feedback for coupling\. Code\-generated features are rarely selected, probably constrained by limited program synthesis\. Instead,Crafterworks by combining three roles: a corrector that tests whether a feature actually reduces residual error, a search process that establishes and feeds back the enumerable space, and an LLM that names semantic features outside that space\.
### 5\.3When do corrective features help?
The central condition is residual headroom, and App\. Fig\. S1 is its most direct support: plotting every backbone–dataset cell’s corrected gain against the backbone’s ownRawerror,Crafter’s lift concentrates where the raw error is large and collapses to zero where the backbone has already saturated the series\. Corrective features help when the frozen backbone leaves structured errors that external signals can explain; they are neutral or harmful otherwise\. This condition is stronger than backbone identity, release order, or feature budget\.
Headroom, not backbone identity, determines the gain\.The clearest evidence comes from comparing the same dataset across backbone families\. Onbizitobs, correction gives large lift when the raw backbone is weak, as with Timer and Moirai, but becomes neutral or negative when the raw forecast already saturates the series, as with Chronos and Moirai\-2\.0\. The gain is governed by how much residual structure the backbone leaves behind, not by which foundation model produced the forecast\. This also explains why newer or stronger backbones can still benefit on some datasets, such asrossmannandepf, while older or weaker backbones do not automatically improve everywhere\.
The useful feature source depends on the kind of residual structure\.When the remaining error is tied to semantic forecast covariates, named features matter most\. This is the case onepf, where external feature banks and syntactic search add little beyond the covariate\-only corrector, but LLM\-named features close much more of the frozen error \(theepfrows of Table S12, whose selected set is LLM\-dominated in App\. Table S7\)\. When the available covariates are already directly informative, as on the retail panels, the covariate\-only corrector already captures a large share of the gain and additional searched features are often redundant \(therossmannandfavoritarows of Table S12\)\. When the residual contains no stable structure, no feature source reliably helps and every column collapses ontoRaw\(the saturatedm5rows, and the flat right\-hand region of App\. Fig\. S1\)\. The question is therefore not whether one source is globally best, but which source matches the residual left by a particular backbone on a particular dataset\.
Budget amplifies headroom but does not create it\.The same\-budget sweep gives the feature\-budget axis of the characterization \(Fig\.[2\(b\)](https://arxiv.org/html/2608.05207#S3.F2.sf2); App\. Table S11\)\. IncreasingKKwidensCrafter’s margin in cells where residual headroom exists, and the gains saturate by moderate budgets\. But a larger budget does not turn saturated cells into wins: if the backbone has already absorbed the available structure, more candidate features mostly add variance and selection risk\. Budget is therefore an amplifier, not a switch\. It helps exploit residual structure when that structure is present, but cannot manufacture structure where the residual is already close to noise\. The practical guidance is therefore simple: deploy corrective feature discovery when validation residuals remain structured, especially in weak\-backbone or semantic\-covariate regimes; treat feature budget as a way to scale gains in those regimes; and avoid forcing correction in saturated cells where no feature source passes the validation gate reliably\.
### 5\.4Robustness to backend and input path
The preceding sections use a fixed LLM backend and route all auxiliary information through the corrector\. We check that the method and the headroom characterization survive two deployment choices: which LLM proposes features, and where covariates or multivariate inputs are injected\.
Changing the LLM backend preserves the characterization\.Repeating the full grid with DeepSeek\-3\.4\-pro yields the same qualitative pattern as the main backend: bothCraftervariants improve every backbone on average, and the gains remain largest where the raw backbone leaves more residual headroom \(Table[1](https://arxiv.org/html/2608.05207#S5.T1); per\-cell results in App\. Table S13\)\. Thus the conclusion of §[5\.3](https://arxiv.org/html/2608.05207#S5.SS3)is not an artifact of one LLM\. The backend does affect individual cells, especially volatile single\-series cases such asepf, so we do not claim backend\-invariant feature sets or identical per\-dataset increments\. The robust claim is weaker and more useful: changing the LLM changes which named features are found, but it does not change the deployment rule that correction helps when exploitable residual structure remains\.
Covariate\-capable backbones do not replace the corrector\.A second deployment choice is whether to feed covariates directly to a backbone that accepts them, or to keep the backbone forecast\-only and route covariates to the residual corrector\. The direct route does not dominate\. Toto ingests covariates harmlessly, with little change inRaw, while Moirai\-2\.0 can degrade substantially when given the same inputs directly \(App\. Table S3\)\. In contrast, the correction layer is stable across this choice: on the retail panels,Crafterreaches similar corrected accuracy whether or not the covariates were first exposed to the backbone, and it can recover even when covariate ingestion degrades the raw forecast\. This supports the design choice in §[3\.1](https://arxiv.org/html/2608.05207#S3.SS1): keep the backbone fixed and route auxiliary signals to a learnable residual corrector\.
### 5\.5Learned features transfer across time
A corrective feature is useful only if it captures persistent structure rather than a single\-window artifact\. We test this by*replaying*the features accepted on one split onto a later split, with the LLM and search turned off, so only the feature set’s origin changes; the target split’s own features serve as an upper bound, and a covariate\-only corrector as a lower bound\. Across two datasets, five backbones, and every ordered split pair, transferred features match the target’s own features at the median and stay within a few percent on most cells \(Fig\.[3](https://arxiv.org/html/2608.05207#S5.F3)b; App\. Table S9\)\. Performance decays only mildly with split distance, while catastrophic failures concentrate in the Moirai shift corner, especially the volatile single\-seriesepf\. Despite low overlap in accepted feature names across splits, performance transfers because different runs often discover functionally equivalent features\. Corrective feature discovery therefore recovers reusable residual structure, not per\-window noise\.
## 6Conclusion
We studied corrective feature discovery: correcting a frozen black\-box forecaster by mining interpretable features of its residual instead of retraining it\.Craftermines the residual with a compositional search and a language model, and admits their candidates through one validation\-grounded gate that is blind to a feature’s origin\. The shared gate turns two generators into a single method and lets any external feature\-engineering system run through the identical pipeline, soCrafteralso serves as an instrument that attributes a change in the forecast to the feature source alone\. Correction pays in proportion to the residual’s exploitable structure: large on weak backbones with informative forecast covariates, where LLM\-named features are the differentiator, and neutral once a strong backbone already fits the series—we report the saturated cells alongside the wins\. The finding holds across six backbones and two LLM backends, is not bought with a larger feature budget, and survives fine\-tuning\. Because features learned on one window transfer to the next on the datasets we test \(§[5\.5](https://arxiv.org/html/2608.05207#S5.SS5)\), a natural next step is a lightweight feature\-memory that carries mined features forward to amortize the search and warm\-start the corrector, adapting as a backbone’s errors evolve across windows\.
## References
- LLM\-fe: automated feature engineering for tabular data with llms as evolutionary optimizers\.arXiv preprint arXiv:2503\.14434\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2608.05207#S4.SS0.SSS0.Px4.p1.2)\.
- A\. F\. Ansari, O\. Shchur, J\. Küken, A\. Auer, B\. Han, P\. Mercado, S\. S\. Rangapuram, H\. Shen, L\. Stella, X\. Zhang, M\. Goswami, S\. Kapoor, D\. C\. Maddix, P\. Guerron, T\. Hu, J\. Yin, N\. Erickson, P\. M\. Desai, H\. Wang, H\. Rangwala, G\. Karypis, Y\. Wang, and M\. Bohlke\-Schneider \(2025\)Chronos\-2: from univariate to universal forecasting\.External Links:2510\.15821,[Document](https://dx.doi.org/10.48550/arXiv.2510.15821),[Link](https://arxiv.org/abs/2510.15821)Cited by:[§4](https://arxiv.org/html/2608.05207#S4.SS0.SSS0.Px2.p1.1)\.
- A\. F\. Ansari, L\. Stella, A\. C\. Turkmen, X\. Zhang, P\. Mercado, H\. Shen, O\. Shchur, S\. S\. Rangapuram, S\. P\. Arango, S\. Kapoor, J\. Zschiegner, D\. C\. Maddix, H\. Wang, M\. W\. Mahoney, K\. Torkkola, A\. G\. Wilson, M\. Bohlke\-Schneider, and B\. Wang \(2024\)Chronos: learning the language of time series\.Transactions on Machine Learning Research\.Note:Expert CertificationExternal Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=gerNCVqqtR)Cited by:[§1](https://arxiv.org/html/2608.05207#S1.p1.1),[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2608.05207#S4.SS0.SSS0.Px2.p1.1)\.
- P\. Auer, N\. Cesa\-Bianchi, and P\. Fischer \(2002\)Finite\-time analysis of the multiarmed bandit problem\.Machine Learning47\(2–3\),pp\. 235–256\.Cited by:[§3\.2](https://arxiv.org/html/2608.05207#S3.SS2.SSS0.Px1.p1.5)\.
- V\. Cerqueira, N\. Moniz, and C\. Soares \(2020\)VEST: automatic feature engineering for forecasting\.arXiv preprint arXiv:2010\.07137\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2608.05207#S3.SS2.SSS0.Px1.p1.5)\.
- H\. Cherkaoui, M\. Tiomoko, G\. Paolo, Y\. Zhang, M\. Yu, K\. Zhang, and H\. Tiomoko Ali \(2025\)Post\-training corrections for improved time\-series forecasting\.arXiv preprint arXiv:2505\.15354\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px3.p1.1)\.
- M\. Christ, N\. Braun, J\. Neuffer, and A\. W\. Kempa\-Liehr \(2018\)Time series FeatuRe extraction on basis of scalable hypothesis tests \(tsfresh – a Python package\)\.Neurocomputing307,pp\. 72–77\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2608.05207#S4.SS0.SSS0.Px4.p1.2)\.
- B\. Cohen, E\. Khwaja, K\. Wang, C\. Masson, E\. Ramé, Y\. Doubli, and O\. Abou\-Amal \(2024\)Toto: time series optimized transformer for observability\.arXiv preprint arXiv:2407\.07874\.Cited by:[§4](https://arxiv.org/html/2608.05207#S4.SS0.SSS0.Px2.p1.1)\.
- P\. M\. P\. Costa \(2021\)Autofits: automated feature engineering for irregular time\-series\.Master’s Thesis,Universidade do Porto\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2608.05207#S3.SS2.SSS0.Px1.p1.5)\.
- X\. Dai, Y\. Liu, H\. Xia, Y\. Hu, Z\. Dong, J\. Yang, and Q\. Xu \(2026a\)Learning the context of errors: black\-box online adaptation of time series foundation models\.arXiv preprint arXiv:2606\.14222\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px3.p1.1)\.
- Z\. Dai, D\. Ma, J\. Tong, M\. Han, J\. Yang, H\. Liu, H\. Fei, and Q\. Yang \(2026b\)NSR\-Boost: a neuro\-symbolic residual boosting framework for industrial legacy models\.arXiv preprint arXiv:2601\.10457\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px4.p1.1)\.
- L\. He, J\. Zhu, F\. Wang, J\. Liu, H\. Xu, Y\. Zhao, P\. S\. Yu, and Q\. Wu \(2025\)Can molecular foundation models know what they don’t know? a simple remedy with preference optimization\.External Links:2509\.25509,[Link](https://arxiv.org/abs/2509.25509)Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px1.p1.1)\.
- N\. Hollmann, S\. Müller, and F\. Hutter \(2023\)Large language models for automated data science: introducing caafe for context\-aware automated feature engineering\.arXiv preprint arXiv:2305\.03403\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2608.05207#S4.SS0.SSS0.Px4.p1.2)\.
- N\. Houlsby, A\. Giurgiu, S\. Jastrzebski, B\. Morrone, Q\. de Laroussilhe, A\. Gesmundo, M\. Attariyan, and S\. Gelly \(2019\)Parameter\-efficient transfer learning for nlp\.External Links:1902\.00751,[Link](https://arxiv.org/abs/1902.00751)Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px1.p1.1)\.
- E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen \(2021\)LoRA: low\-rank adaptation of large language models\.External Links:2106\.09685,[Link](https://arxiv.org/abs/2106.09685)Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px1.p1.1)\.
- Y\. Huang, Y\. Zhou, M\. Hefenbrock, T\. Riedel, L\. Fang, and M\. Beigl \(2022\)Automatic feature engineering through Monte Carlo tree search\.InMachine Learning and Knowledge Discovery in Databases \(ECML PKDD\),Lecture Notes in Computer Science, Vol\.13715\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px2.p1.1)\.
- Kaggle \(2015\)Rossmann store sales\.Note:Kaggle competition,[https://www\.kaggle\.com/c/rossmann\-store\-sales](https://www.kaggle.com/c/rossmann-store-sales)Cited by:[§4](https://arxiv.org/html/2608.05207#S4.SS0.SSS0.Px1.p1.1)\.
- Kaggle \(2017\)Corporación favorita grocery sales forecasting\.Note:Kaggle competition,[https://www\.kaggle\.com/c/favorita\-grocery\-sales\-forecasting](https://www.kaggle.com/c/favorita-grocery-sales-forecasting)Cited by:[§4](https://arxiv.org/html/2608.05207#S4.SS0.SSS0.Px1.p1.1)\.
- Kaggle \(2024\)Rohlik sales forecasting challenge\.Note:Kaggle competition,[https://www\.kaggle\.com/competitions/rohlik\-sales\-forecasting\-challenge](https://www.kaggle.com/competitions/rohlik-sales-forecasting-challenge)Cited by:[§4](https://arxiv.org/html/2608.05207#S4.SS0.SSS0.Px1.p1.1)\.
- U\. Khurana, H\. Samulowitz, and D\. Turaga \(2018\)Feature engineering for predictive modeling using reinforcement learning\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.32\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px2.p1.1)\.
- L\. Kocsis and C\. Szepesvári \(2006\)Bandit based monte\-carlo planning\.InMachine Learning: ECML 2006,pp\. 282–293\.Cited by:[§3\.2](https://arxiv.org/html/2608.05207#S3.SS2.SSS0.Px1.p1.5)\.
- J\. Lago, G\. Marcjasz, B\. De Schutter, and R\. Weron \(2021\)Forecasting day\-ahead electricity prices: a review of state\-of\-the\-art algorithms, best practices and an open\-access benchmark\.Applied Energy293,pp\. 116983\.External Links:[Document](https://dx.doi.org/10.1016/j.apenergy.2021.116983)Cited by:[§4](https://arxiv.org/html/2608.05207#S4.SS0.SSS0.Px1.p1.1)\.
- D\. Liang, Q\. Li, Y\. Wang, J\. Chen, H\. Zhang, X\. Cui, Q\. Wang, and S\. Li \(2026\)The forecast after the forecast: a post\-processing shift in time series\.arXiv preprint arXiv:2601\.20280\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px3.p1.1)\.
- C\. Liu, T\. Aksu, J\. Liu, X\. Liu, H\. Yan, Q\. Pham, S\. Savarese, D\. Sahoo, C\. Xiong, and J\. Li \(2025a\)Moirai 2\.0: when less is more for time series forecasting\.arXiv preprint arXiv:2511\.11698\.Cited by:[§1](https://arxiv.org/html/2608.05207#S1.p1.1),[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2608.05207#S4.SS0.SSS0.Px2.p1.1)\.
- F\. Liu, H\. Wang, J\. Cho, D\. Roth, and A\. W\. Lo \(2025b\)AutoCT: automating interpretable clinical trial prediction with LLM agents\.arXiv preprint arXiv:2506\.04293\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px2.p1.1)\.
- Y\. Liu, H\. Zhang, C\. Li, X\. Huang, J\. Wang, and M\. Long \(2024\)Timer: generative pre\-trained transformers are large time series models\.InProceedings of the 41st International Conference on Machine Learning,pp\. 32369–32399\.Cited by:[§1](https://arxiv.org/html/2608.05207#S1.p1.1),[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2608.05207#S4.SS0.SSS0.Px2.p1.1)\.
- A\. Lopez and F\. Sobieczky \(2024\)Surrogate modeling for explainable predictive time series corrections\.arXiv preprint arXiv:2412\.19897\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px4.p1.1)\.
- S\. Lyu, S\. Zhong, T\. Chen, W\. Ruan, Q\. Liu, T\. Lv, Q\. Wen, R\. C\. Wong, and Y\. Liang \(2026\)TS\-Memory: plug\-and\-play memory for time series foundation models\.arXiv preprint arXiv:2602\.11550\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px3.p1.1)\.
- S\. Makridakis, E\. Spiliotis, and V\. Assimakopoulos \(2022\)M5 accuracy competition: results, findings, and conclusions\.International Journal of Forecasting38\(4\),pp\. 1346–1364\.External Links:[Document](https://dx.doi.org/10.1016/j.ijforecast.2021.11.013)Cited by:[§4](https://arxiv.org/html/2608.05207#S4.SS0.SSS0.Px1.p1.1)\.
- A\. Murray, D\. Dervovic, and M\. Cashmore \(2025\)ELATE: evolutionary language model for automated time\-series engineering\.arXiv preprint arXiv:2508\.14667\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Nam, K\. Kim, S\. Oh, J\. Tack, J\. Kim, and J\. Shin \(2024\)Optimized feature generation for tabular data via llms with decision tree reasoning\.arXiv preprint arXiv:2406\.08527\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px2.p1.1)\.
- S\. Palaskar, V\. Ekambaram, A\. Jati, N\. Gantayat, A\. Saha, S\. Nagar, N\. H\. Nguyen, P\. Dayama, R\. Sindhgatta, P\. Mohapatra, H\. Kumar, J\. Kalagnanam, N\. Hemachandra, and N\. Rangaraj \(2024\)AutoMixer for improved multivariate time\-series forecasting on business and it observability data\.Proceedings of the AAAI Conference on Artificial Intelligence38\(21\),pp\. 22962–22968\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v38i21.30336)Cited by:[§4](https://arxiv.org/html/2608.05207#S4.SS0.SSS0.Px1.p1.1)\.
- O\. Shchur, A\. C\. Turkmen, N\. Erickson, H\. Shen, A\. Shirkov, T\. Hu, and B\. Wang \(2023\)AutoGluon\-TimeSeries: AutoML for probabilistic time series forecasting\.InProceedings of the Second International Conference on Automated Machine Learning \(AutoML\),Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px3.p1.1)\.
- Z\. Wang, I\. Miliou, I\. Samsten, and P\. Papapetrou \(2023\)Counterfactual explanations for time series forecasting\.arXiv preprint arXiv:2310\.08137\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px4.p1.1)\.
- G\. Woo, C\. Liu, A\. Kumar, C\. Xiong, S\. Savarese, and D\. Sahoo \(2024\)Unified training of universal time series forecasting transformers\.InProceedings of the 41st International Conference on Machine Learning \(ICML\),Cited by:[§4](https://arxiv.org/html/2608.05207#S4.SS0.SSS0.Px2.p1.1)\.
- G\. P\. Zhang \(2003\)Time series forecasting using a hybrid ARIMA and neural network model\.Neurocomputing50,pp\. 159–175\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px3.p1.1)\.
- T\. Zhang, Z\. Zhang, Z\. Fan, H\. Luo, F\. Liu, Q\. Liu, W\. Cao, and J\. Li \(2022\)OpenFE: automated feature generation with expert\-level performance\.arXiv preprint arXiv:2211\.12507\.Cited by:[§2](https://arxiv.org/html/2608.05207#S2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2608.05207#S3.SS2.SSS0.Px1.p1.5)\.
## Appendix
Tables, figures, algorithms and equations below carry an S prefix \(Table S1, Fig\. S1, Alg\. S1, Eq\. \(S1\)\); plain references \(§3\.2, Fig\. 2a, Table 1\) point to the main paper above\.
## Appendix ACorrection Target and Metric
For each series and forecast origin we observe a historyx1:Tx\_\{1:T\}, a frozen backbone forecasty^1:H\\hat\{y\}\_\{1:H\}over horizonHH, and \(at training time\) the truthy1:Hy\_\{1:H\}\.Crafterpredicts a per\-series multiplicative correction factor
y~1:H=f⋅y^1:H,f=clip\(∑hyh∑hy^h,ℓ,u\),\[ℓ,u\]=\[0\.3,3\.0\],\\begin\{split\}\\tilde\{y\}\_\{1:H\}&=f\\cdot\\hat\{y\}\_\{1:H\},\\\\ f&=\\mathrm\{clip\}\\\!\\Big\(\\tfrac\{\\sum\_\{h\}y\_\{h\}\}\{\\sum\_\{h\}\\hat\{y\}\_\{h\}\},\\,\\ell,\\,u\\Big\),\\quad\[\\ell,u\]=\[0\.3,3\.0\],\\end\{split\}\(S1\)with the clip admitting genuine multiplicative structure \(e\.g\., promotions, price changes\) while rejecting heavy\-tailed ratios from near\-zero denominators; an additive\-offset targetf\+=∑hyh−∑hy^hf^\{\+\}=\\sum\_\{h\}y\_\{h\}\-\\sum\_\{h\}\\hat\{y\}\_\{h\}is also supported and selected against within the corrector set \(§3\.4\)\. We report the weighted mean absolute percentage error of the horizon total,
wMAPE=∑i\|∑hy~i,h−∑hyi,h\|∑i\|∑hyi,h\|\.wMAPE=\\frac\{\\sum\_\{i\}\\big\|\\sum\_\{h\}\\tilde\{y\}\_\{i,h\}\-\\sum\_\{h\}y\_\{i,h\}\\big\|\}\{\\sum\_\{i\}\\big\|\\sum\_\{h\}y\_\{i,h\}\\big\|\}\.\(S2\)We score the horizon*total*because it is the quantity deployment consumes—replenishment, procurement, and capacity decisions are driven by aggregate demand over the lead time—and the quantity the correction factor of Eq\. \([S1](https://arxiv.org/html/2608.05207#A1.E1)\) acts on\. The comparison stays fair: every method, the frozen backbone included, is scored on the same total\.
## Appendix BTheCrafterRound
Alg\.[S1](https://arxiv.org/html/2608.05207#alg1)gives one round of the loop sketched in §3\.1\. A run repeats it forRRrounds, threading the accepted\-feature setℱ\\mathcal\{F\}, the search arm statisticsΘ\\Theta, and the channel pool𝒞\\mathcal\{C\}across rounds\. The two generators of §3\.2 appear as*Exploration*\(compositional UCB search, §3\.2\) and*Reasoning*\(the LLM roles, §3\.2\); both feed the single source\-blind gate of §3\.3 \(line[6](https://arxiv.org/html/2608.05207#alg1.l6)\), which keeps only candidates whose Spearman correlation with the*unexplained*validation residual exceedsτ\\tau, that are not collinear with an accepted feature, and that fit the per\-family and per\-round budgets\. After refitting, a per\-tier permutation step prunes the lowest\-value features byΔ\\Deltavalidation\-wMAPE \(line[8](https://arxiv.org/html/2608.05207#alg1.l8)\), and the kept/trimmed outcome plus the realized validation gain is credited back to the search arms as UCB feedback\. The cross\-round coupling shown—LLM\-named base columns entering𝒞\\mathcal\{C\}so the search can compose them \(line[7](https://arxiv.org/html/2608.05207#alg1.l7)\)—is one of the two channelsCrafter†enables \(§3\.2\); it is off in the decoupledCrafter\.
Algorithm S1OneCrafterround0:accepted features
ℱ\\mathcal\{F\}, arm statistics
Θ\\Theta, channel pool
𝒞\\mathcal\{C\}, frozen backbone forecast
y^\\hat\{y\}
1:fit the corrector set on
ℱ\\mathcal\{F\}; keep the head \(None/additive/multiplicative\) with lowest validation wMAPE \{Nonefloors risk, §3\.4\}
2:
r←r\\leftarrowunexplained validation residual under the kept corrector; compute family coverage
3:
D←Planner\(coverage,ℱ,blacklist\)D\\leftarrow\\textsc\{Planner\}\(\\text\{coverage\},\\,\\mathcal\{F\},\\,\\text\{blacklist\}\)\{Reasoning: exploration directions, §3\.2\}
4:
PR←\(⋃d∈DSpecifier\(d\)\)∪BaseProposer\(\)P\_\{\\mathrm\{R\}\}\\leftarrow\\big\(\\textstyle\\bigcup\_\{d\\in D\}\\textsc\{Specifier\}\(d\)\\big\)\\cup\\textsc\{BaseProposer\}\(\)\{Reasoning: LLM\-named features\}
5:
PE←UCB\-Sample\(Θ,𝒞\)P\_\{\\mathrm\{E\}\}\\leftarrow\\textsc\{UCB\-Sample\}\(\\Theta,\\,\\mathcal\{C\}\)\{Exploration: compositional search, §3\.2\}
6:
𝒜←\{c∈dedup\(PE∪PR\):\|ρ\(c,r\)\|\>τ,non\-collinear,within budgets\}\\mathcal\{A\}\\leftarrow\\big\\\{\\,c\\in\\mathrm\{dedup\}\(P\_\{\\mathrm\{E\}\}\\cup P\_\{\\mathrm\{R\}\}\):\|\\rho\(c,r\)\|\>\\tau,\\ \\text\{non\-collinear\},\\ \\text\{within budgets\}\\,\\big\\\}\{source\-blind gate, §3\.3\}
7:
𝒞←𝒞∪\{new base columns in𝒜\}\\mathcal\{C\}\\leftarrow\\mathcal\{C\}\\cup\\\{\\text\{new base columns in \}\\mathcal\{A\}\\\}\{Reasoning
→\\toExploration coupling \(off by default\)\}
8:refit the corrector on
ℱ∪𝒜\\mathcal\{F\}\\cup\\mathcal\{A\};
𝒯←\\mathcal\{T\}\\leftarrowfeatures below the per\-tier top\-
KKby permutation
Δ\\Deltaval\-wMAPE \{prune\}
9:update
Θ\\Thetawith UCB credit from kept vs\. trimmed importances and the validation\-wMAPE gain
10:return
\(ℱ∪𝒜\)∖𝒯,Θ,𝒞\(\\mathcal\{F\}\\cup\\mathcal\{A\}\)\\setminus\\mathcal\{T\},\\ \\Theta,\\ \\mathcal\{C\}
## Appendix CFeature Grammar
A corrective feature is a structured specification: a typed record that compiles deterministically to one numeric column over items\. Specifications take one of three types, and every field is drawn from a fixed vocabulary, so the space is bounded, searchable, and canonicalizable\. The atomic search emitscomborecords; the LLM emitscombo,flag, andcoderecords, which the rest of the paper namesLLM\-combo,LLM\-flag, andLLM\-code\.
- •combo— a pairwise expressionpost\(τa\(ca\)⊕τb\(cb\)\)\\textsc\{post\}\\big\(\\tau\_\{a\}\(c\_\{a\}\)\\,\\oplus\\,\\tau\_\{b\}\(c\_\{b\}\)\\big\), whereca,cbc\_\{a\},c\_\{b\}are input channels \(base features or covariates\),τ∙\\tau\_\{\\bullet\}are unary transforms,⊕\\oplusis a binary operator, andpostis an optional output transform with an optional clip \(a single transformed channel whencb,⊕c\_\{b\},\\oplusare omitted\)\.
- •flag— a0/10/1indicator formed by combining a small set of thresholded conditions\{\(c,op,v\)\}\\\{\(c,\\,\\mathrm\{op\},\\,v\)\\\}withall/any\.
- •code— a short Python function over the target history and the backbone forecast, restricted by an abstract\-syntax\-tree allow\-list tonumpy,math, andscipy\.\{signal,stats,fft,special\}with no I/O,exec, oreval\.
Vocabulary\.The search ranges overB=8B\{=\}8binary operators⊕∈\{add,sub,mul,div,abs\_diff,min,max,safe\_div\}\\oplus\\in\\\{\\texttt\{add\},\\texttt\{sub\},\\texttt\{mul\},\\texttt\{div\},\\texttt\{abs\\\_diff\},\\texttt\{min\},\\texttt\{max\},\\texttt\{safe\\\_div\}\\\}andT=15T\{=\}15unary transformsτ∈\\tau\\in\{null,neg,abs,log,log1p,signlog,sqrt,square,exp,sigmoid,tanh,softplus,zscore,winsorize,ema\}; flag conditions draw from the operators\{<,≤,\>,≥,=,≠,between,in,not\_in\}\\\{<,\\leq,\>,\\geq,=,\\neq,\\texttt\{between\},\\texttt\{in\},\\texttt\{not\\\_in\}\\\}\. The executor additionally admits a few rarer transforms and operators \(e\.g\.,geo\_mean, quantile summaries\); these are reachable by the LLM but outside the atomic search’s UCB vocabulary\.
Feature families\.Each channel belongs to one mechanism*family*, the unit L1 routes over \(§3\.2\)\. Representative families and example channels:level\(mean\_sales,zero\_ratio\);trend\(trend\_ratio,trend\_consistency\);seasonality\(autocorr7,seasonal\_strength\);volatility\(cv,outlier\_rate\);shift\(regime\_shift,shift\_magnitude\);anomaly\(spike\_ratio,consecutive\_zeros\)\.
Example\.The JSON record\{type:combo, feat\_a:cv, feat\_b:mean\_sales, op:div, transform\_a:log, transform\_b:null\}denoteslog\(cv\)/mean\_sales\\log\(\\mathrm\{cv\}\)/\\mathrm\{mean\\\_sales\}and is named accordingly\. A*canonical key*maps equivalent records to one identifier: the display name and all bookkeeping fields \(source, family, scores\) are dropped, the empty/null/identity transforms collapse to a single sentinel, numeric thresholds and clips are rounded, andflagconditions are sorted, so the key is an order\- and rename\-invariant hash of the remaining grammar fields\. It drives cross\-round de\-duplication and a rejection blacklist that suppresses any candidate already tried and refused\.
## Appendix DCorrector and Hyperparameters
The corrector is a histogram gradient\-boosted regressor \(200200trees, max depth44, learning rate0\.10\.1, min\-samples\-leaf2020\)\. The corrector set comprises\{None,additive,multiplicative\-factor\}\\\{\\textsc\{None\},\\text\{additive\},\\text\{multiplicative\-factor\}\\\}correctors, the best selected by validation wMAPE; the multiplicative\-factor head is the strongest on weak\-backbone×\\timesforecast\-covariate cells, where a per\-row learned factor is more flexible than per\-horizon offsets\. We fix its clip to\[0\.3,3\.0\]\[0\.3,3\.0\]rather than a dataset\-adaptive bound: an adaptive bound is occasionally tighter on clean retail but blows up on weak\-backbone×\\timesLLM cells, so we trade a little adaptivity to remove that failure mode\. Acceptance gate: a candidate is kept iff\|ρSpearman\(feature,unexplained residual\)\|≥τ\|\\rho\_\{\\mathrm\{Spearman\}\}\(\\text\{feature\},\\,\\text\{unexplained residual\}\)\|\\geq\\tau, pairwise\|ρ\|≤0\.75\|\\rho\|\\leq 0\.75, with at most two features per family and ten per round\. The threshold is a single global constantτ=0\.05\\tau\{=\}0\.05, fixed across all datasets rather than tuned per series, so there is no per\-dataset grid\. Search defaults: a four\-layer conditioned UCB1 tree with exploration constantκ=1\.4\\kappa\{=\}1\.4,R=3R\{=\}3rounds, andK=4K\{=\}4planner directions; the selection, update, and reward rules follow\.
#### Bandit selection and update\.
The compositional search is a Monte\-Carlo tree of UCB1 bandits with one arm per choice at each layer \(family, then arity, then channels, then operator\)\. Every armaastores a visit countnan\_\{a\}and a running mean rewardμ^a\\hat\{\\mu\}\_\{a\}; each L1 family arm keeps one\(na,μ^a\)\(n\_\{a\},\\hat\{\\mu\}\_\{a\}\)pair per arity, so univariate and bivariate evidence never mix\. At a node whose children are𝒜\\mathcal\{A\}with total visitsN=∑a′∈𝒜na′N=\\sum\_\{a^\{\\prime\}\\in\\mathcal\{A\}\}n\_\{a^\{\\prime\}\}, the descent selects the arm that maximizes the UCB1 score,
a⋆=argmaxa∈𝒜μ^a\+κlnNna,a^\{\\star\}=\\arg\\max\_\{a\\in\\mathcal\{A\}\}\\ \\hat\{\\mu\}\_\{a\}\+\\kappa\\sqrt\{\\frac\{\\ln N\}\{n\_\{a\}\}\},\(S3\)visiting any never\-tried arm \(na=0n\_\{a\}\{=\}0\) first\. Once an arm’s candidate has been scored, or a round has finished, the arm is credited with a rewardrrand updated by the incremental sample mean
na←na\+1,μ^a←μ^a\+1na\(r−μ^a\),n\_\{a\}\\leftarrow n\_\{a\}\+1,\\qquad\\hat\{\\mu\}\_\{a\}\\leftarrow\\hat\{\\mu\}\_\{a\}\+\\frac\{1\}\{n\_\{a\}\}\\bigl\(r\-\\hat\{\\mu\}\_\{a\}\\bigr\),\(S4\)applied to every arm on the selected root\-to\-leaf path; a decayed copy is also propagated to sibling contexts that share the family or base channel\. The reward combines two signals\. A fast per\-candidate proxy is available already at scoring time,
rproxy=α\|ρ\|\+β𝕀\[accepted\],α=0\.4,β=0\.3,r\_\{\\mathrm\{proxy\}\}=\\alpha\\,\\lvert\\rho\\rvert\+\\beta\\,\\mathbb\{I\}\[\\text\{accepted\}\],\\qquad\\alpha\{=\}0\.4,\\ \\beta\{=\}0\.3,\(S5\)whereρ\\rhois the Spearman correlation between the candidate and the still\-unexplained validation residual and𝕀\[accepted\]\\mathbb\{I\}\[\\text\{accepted\}\]marks passing the acceptance gate above\. After the corrector is refit at each round’s end, a realized signal supplements the proxy: every surviving feature is credited in proportion to its gradient\-boosting importance, plus aγ=0\.3\\gamma\{=\}0\.3round\-level term on the round’s validation\-wMAPE improvement\. The proxy steers exploration cheaply within a round; the realized signal ties the search to the deployed corrector, so a feature is ultimately worth what keeping it actually buys\.
#### Why gradient\-boosted trees\.
The residual is nonlinear and regime\-dependent, concentrating in promotions, regime shifts, and rare events, so a linear corrector cannot fit it\. Gradient\-boosted trees have the capacity and keep the correction auditable: a per\-feature importance names*which*mechanism drives a backbone’s residual\. Read across backbones, the corrector*diagnoses*each model’s blind spots, naming what a given backbone systematically misses \(§5\.2\)\.
#### Protocol and compute\.
All backbone forecasts are computed once on NVIDIA A100/A40 GPUs and cached; because random seeds do not change a frozen backbone’s forecast, the entire multi\-seed grid \(seeds4242–4444×\\timesthree rolling splits per cell\) reuses the warm cache with*no*GPU recomputation, and the residual\-correction loop itself runs on CPU\. The stack is Python 3\.12, with PyTorch 2\.x for the backbone runners and LightGBM 4\.x for the corrector\. A full single\-seed sweep of the grid completes in∼30\{\\sim\}30minutes on warm caches—orders of magnitude cheaper than per\-dataset backbone fine\-tuning, which requires a GPU training run for every \(dataset, backbone\) pair\. We report last\-round \(R=3R\{=\}3\) test wMAPE, averaged over the three splits and three seeds, and decompose variance into seed \(optimizer/LLM\-sampling\) and split \(temporal\-drift\) components; split variance dominates\. Quantified per cell on the DeepSeek grid \(wMAPE×100\\times 100\): the median seed\-std ofCrafteris0\.490\.49against a median split\-std of∼3\.2\{\\sim\}3\.2; the frozen backbone is seed\-deterministic, the no\-LLM andTSFreshcolumns sit near0\.020\.02, and the LLM methods’ seed noise concentrates onepf\.
#### Per\-arity statistics and the univariate transform\.
Each L1 family arm stores two independent \(count, mean\) pairs, one per arity, so univariate and bivariate evidence never mix\. We do not search the univariate transform: with onlyT=15T\{=\}15options and that branch serving mainly to surface new channels for later composition, a bandit there would chiefly drain samples from the bivariate branch, where theBT2B\\,T^\{2\}operator\-triple choice carries the reward\. Because the bivariate branch is sampled far more often, its visit count grows much faster; we periodically rebalance the two arities’ counts so the exploration bonusκlnN/n\\kappa\\sqrt\{\\ln N/n\}does not lock onto the under\-sampled univariate arity\.
## Appendix EExternal\-System Adaptation
The three external systems run*inside*theCrafterharness: the same cached backbone residual, the same base and covariate columns, the same gradient\-boosted corrector and corrector\-set selection, and the same validation top\-KKfeature selector\. Only the candidate generator is swapped, so the comparison isolates feature generation \(§4\)\. Each generator is kept as close to its published form as the forecasting setting allows\.
- •TSFreshextracts its standard statistical bank over each series history through the library’s own extractor \(∼30\{\\sim\}30statistics per series\), from which the shared selector keeps the topKK\.
- •CAAFEtargets classification, so its LLM iteration loop’s validation\-accuracy acceptance test runs on the residual ratio discretized into three bins \(over\-, well\-, and under\-predicted\); the accepted features themselves remain continuous columns for the shared corrector\.
- •LLM\-FEruns its evolutionary loop, with LLM\-proposed mutations and crossovers over a population of feature expressions and fitness the validation correlation against the same residual target\.
The two LLM\-based systems use the same backend, per\-run seed, and feature budget asCrafter\. We do not run the systems’ native end\-to\-end pipelines, because they target tabular prediction with their own downstream model: a native run would change the corrector, the target, and the metric at once, and a difference in the final forecast could then no longer be attributed to the feature source\. This inside\-the\-harness design is the point of the instrument \(§3\.3\); its cost is that we measure each system as a feature generator for residual correction, not its native end\-to\-end performance\.
## Appendix FFine\-tuning baseline
TheFTbaseline in Table[S12](https://arxiv.org/html/2608.05207#A14.T12)applies a*uniform head\-only*fine\-tune to every backbone \(training the output head with the backbone frozen\), early\-stopping on validation wMAPE with the test window untouched, on splits identical toCrafter; a constant recipe keeps the comparison fair across backbones\. As a stronger check, the two Moirai backbones, which ship an official recipe, additionally receive a full\-parameter in\-domain fine\-tune \(Table[S2](https://arxiv.org/html/2608.05207#A6.T2)\): it reuses the officialuni2tsMoiraiFinetunemodule and loss but masks the lastHHpatches \(eval\-aligned\) rather than at random, selects checkpoints on validation wMAPE rather than NLL, and trains in fp32 with gradient clipping and warmup\. The comprehensive grid \(Table[S1](https://arxiv.org/html/2608.05207#A6.T1); all six backbones×\\timessix datasets, both backbone branches×\\timesfour correction gradients\) reports this head\-onlyFTnext to each correction level\. The grids are single\-seed \(42\), split\-averaged, so absolute values differ slightly from the33\-seed main table, most visibly on the volatile single\-seriesepf; the before/after comparison of Fig\. 2a \(top vs\. bottom\) uses both branches within the same seed and is unaffected\. Fine\-tuning requires a GPU training run for every \(dataset, backbone\) pair, orders of magnitude more costly than the GPU\-free correction loop \(App\.[D](https://arxiv.org/html/2608.05207#A4.SS0.SSS0.Px3)\)\.
Table S1:Fine\-tuning does not remove the gain \(full grid\)\.Test wMAPE×100\\times 100\(lower is better\) for the four correction gradients—no correction \(Rawzero\-shot /FTfine\-tuned\), covariate\-only corrector \(Cov\),Crafter, andCrafter†—on the zero\-shot and the fine\-tuned backbone, over all66backbones×\\times66datasets \(chr2is covariate\-aware Chronos\-2, single\-seed; its covariate\-only corrector was not re\-run on the fine\-tuned branch, so the twoCovcolumns carry the same value\)\. Fine\-tuning is the uniform head\-only probe \(full\-parameter Moirai in Table[S2](https://arxiv.org/html/2608.05207#A6.T2)\)\. On the non\-saturated datasets \(epf,rossmann,rohlik\) the orderingCrafter<Cov<\\textsc\{Crafter\}\{\}<\\textsc\{Cov\}<no\-correction holds on*both*branches;m5is saturated and all gradients coincide\. TheCrafterandCrafter†columns \(gray\) are highlighted in both branches\.Boldmarks the best gradient within each branch \(aCraftervariant when it ties\)\. Single seed \(42\), split\-averaged\.Table S2:Full\-parameter Moirai stress test\.The four correction gradients on a*fully*fine\-tuned Moirai backbone—no correction \(FT\), covariate\-only corrector \(Cov\),Crafter, andCrafter†—alongside the zero\-shot branch, for the two covariate\-capable Moirai models \(test wMAPE×100\\times 100, lower is better\)\.Crafterstill helps on66of88fine\-tuned cells; fine\-tuning rarely beats correcting the frozen backbone\. TheCrafter/Crafter†columns \(gray\) are highlighted;boldmarks a corrector that beats both no\-correction baselines in its branch \(Raw/Covon the left,FT/Covon the right\), the Table[S12](https://arxiv.org/html/2608.05207#A14.T12)rule\. Exceptions:favorita,bizitobsunder Moirai\-2\.0\. Single seed \(42\), split\-averaged\.
## Appendix GCovariate routing and native multivariate forecasting
These two ablations support §5\.4: the corrector’s gain is decoupled from how a covariate\-capable or multivariate\-capable backbone uses its richer input path\.
#### Covariates in the backbone vs\. the corrector\.
Table[S3](https://arxiv.org/html/2608.05207#A7.T3)feeds known covariates to the frozen backbone \(Raw\+\+cov\) instead of routing them to the corrector\. Toto ingests them harmlessly \(Δ≤0\\Delta\\leq 0on every dataset\) while Moirai\-2\.0 misuses them, degradingRawby up to45\.8%45\.8\\%\(rossmann\); neither backbone turns covariate access into a real gain\. Routing the same covariates to the corrector instead yields large gains, andCrafterreaches comparable accuracy whether or not the backbone was given the covariates—recovering even where the cov\-fed backbone degraded its ownRaw\(Moirai\-2\.0rossmann,31\.1→13\.831\.1\{\\to\}13\.8vs\.11\.811\.8on the univariate backbone\)\. The corrector’s gain is therefore largely independent of how, or whether, the backbone ingests covariates; the lone volatileepfseries is the one exception\. Timer and Chronos are univariate and cannot ingest covariates\.
Table S3:Covariates belong in the corrector, not the backbone\(DeepSeek\-3\.4\-pro backend, test wMAPE×100\\times 100, three\-split mean, seed 42\)\.Rawis the univariate backbone;Raw\+\+cov feeds covariates to the frozen backbone, withΔ\\Deltathe resulting change inRaw\(positive==worse\)\.Crafter\(uni\) andCrafter\(cov\) place the corrector on the univariate and the cov\-fed backbone respectively \(the twograycolumns;bold==beats bothRawandRaw\+\+cov in that row, the Table[S12](https://arxiv.org/html/2608.05207#A14.T12)rule\)\. Feeding covariates to the backbone is at best neutral \(Toto,Δ≤0\\Delta\\leq 0throughout\) and often harmful \(Moirai\-2\.0\), whereas routing them to the corrector yields large gains; the twoCraftercolumns match across the retail panels, so the corrector’s gain is decoupled from how the backbone routes covariates \(epfexcepted\)\. Timer and Chronos are univariate and cannot ingest covariates; Moirai covariate injection was run single\-split only;bizitobshas no covariates\.m5is evaluated on this grid’s own series basis, so itsRawdiffers slightly from Table[S12](https://arxiv.org/html/2608.05207#A14.T12); compare within the table\.
#### Native multivariate forecasting\.
bizitobsis the one panel with several jointly forecastable channels\. Table[S4](https://arxiv.org/html/2608.05207#A7.T4)compares per\-channel \(uni\) with joint \(mv\) backbone forecasting on the two natively multivariate backbones\. Joint forecasting improves Toto’s any\-variateRawand badly degrades the small Moirai\-2\.0, but the corrector behaves identically on either forecast \(the same small lift in both the uni and mv columns\), so the uni/mv gap is set by the backbone, not the wrapper\. Timer and Chronos are univariate; Moirai is multivariate\-capable but not run here\. Importantly, native multivariate forecasting changes the backbone, not the wrapper\. Onbizitobs, the only panel where several channels can be forecast jointly, native multivariate forecasting is strongly backbone\-dependent: it improves Toto’s raw forecast but degrades the smaller Moirai\-2\.0 model \(App\. Table[S4](https://arxiv.org/html/2608.05207#A7.T4)\)\. The correction layer behaves similarly on either raw forecast\. Thus the univariate–multivariate gap is mainly a property of the backbone’s input path, whereasCrafterremains a black\-box residual correction layer that applies after either choice\.
Table S4:Native multivariate forecasting onbizitobs\(DeepSeek\-3\.4\-pro backend, test wMAPE×100\\times 100, three\-split mean, seed 42\)\. uni==per\-channel, mv==joint forecast\.Craftershows a similar lift overRawin both columns; the uni/mv difference comes from the backbone’s joint\-forecast quality\.
## Appendix HLLM Roles and Prompts
Craftercalls a single LLM provider in four roles: a fast tier runs thePlanner, a reasoning tier runs theSpecifierandReflection, and aBase Proposerinvents new base columns\. The legacyRerankeris disabled—the merged candidate pool is gated deterministically by the source\-blind rule of §3\.3, not by an LLM\. Across a run \(three rounds\) the LLM is queried about a dozen times \(∼4\{\\sim\}4per round\): each exploration round runs aPlannerthat proposes the round’s directions, aSpecifierper direction, aBase Proposer, and aReflectioncall only when fewer than three candidates pass the gate\. Every system prompt is parameterized by a per\-dataset domain preamble \(the dataset name, a one\-line domain hint, and a role\-specific note\)\. We show each role’s system and user templates below; run\-time inputs appear as\{placeholders\}and the model must reply in strict JSON\.
#### Planner \(fast tier\)\.
*Inputs:*round index, validation\-wMAPE trajectory, per\-family coverage, under\-served families, chronic failure subgroups, and rejection fingerprints by family\.*Output:*exactlyKKexploration directions\.
> *System\.*“You are a feature\-exploration planner for\{domain\}time\-series forecasting\. Each round you see current feature coverage, chronic failure subgroups, recent rejection fingerprints grouped by family, and the wMAPE trend over recent rounds\. OutputKKstructured exploration*directions*; each tells the downstream Specifier which family/subgroup to design for, what type of signal \(combo,flag, or both\), and which base feats to use\. Return strict JSON\{directions:\[\{family, target\_subgroup, base\_feat\_pool, spec\_type\_hint, mechanism\_hint, rationale\}\]\}\. Constraints: return exactlyKKdirections;familyfrom the allowed names;base\_feat\_poola subset of the supplied feats; do not repeat a \(family, subgroup\) pair; prioritize under\-served families and chronic subgroups\.” *User\.*“Round\{r\},K=\{k\}K\{=\}\\texttt\{\\\{k\\\}\}\. wMAPE trajectory \(last 5\):\{\.\.\.\}\. Family names, coverage, under\-served families, chronic subgroups \(top 8\), and rejections by family over the last 3 rounds:\{\.\.\.\}\. Existing dynamic features \(top 30, withρ\\rho\):\{\.\.\.\}\. Available base feats:\{\.\.\.\}\. Output exactlyKKdirections as JSON\.”
#### Specifier \(reasoning tier\)\.
*Inputs:*one Planner direction \(family, optional subgroup, base\-feat pool, signal\-type hint, mechanism hint\), up to eight chronic failure cases, and in\-scope rejections to avoid\.*Output:*5–8 named specs \(combo/flag/code\)\.
> *System\.*“You are a feature engineer for\{domain\}time\-series forecasting\. You receive one exploration direction \(family \+ optional target subgroup \+ allowed base\-feat pool \+ mechanism hint\)\. Strictly within that direction, produce 5–8 structured feature specs \(combo,flag, orcode\)\. Hard constraints:combo/flagmay use only feats inbase\_feat\_pool;codespecs may read the target history and the backbone forecast but may import onlynumpy,math, andscipy\.\{signal,stats,fft,special\}, with no I/O,exec,eval,open, loops, ortry; do not resubmit a spec identical to a rejected one on \(op, feat set, transforms\); every spec must passvalidate\_spec\. Output JSON undermissing\_features, each entry acombo\(feature\_name, feat\_a, feat\_b, op, transform\_a, transform\_b, post\_transform\), aflag\(conditions, combine, then, else\), or acodespec \(compute\_func\_code\)\. Guidance: the atomic search already covers plaincombo/unary features, so prioritizeflag\(multi\-condition thresholds the search cannot express\) andcode\(FFT, autocorrelation, Hurst, windowed statistics\); of the 5–8 specs keepcombo≤2\\leq 2andflag\+code≥60%\\geq 60\\%\.” *User\.*“family, target\_subgroup, base\_feat\_pool, spec\_type\_hint, mechanism\_hint:\{\.\.\.\}\. Chronic failure cases \(for inspiration\):\{\.\.\.\}\. Rejected fingerprints in this direction \(avoid the same mechanism\):\{\.\.\.\}\. Produce 5–8 specs as strict JSON undermissing\_features\.”
#### Reflection \(reasoning tier; fires only when<3\{<\}3candidates pass the gate in a round\)\.
*Inputs:*one rejected candidate spec, its rejection reason \(pair\-rho\-highormarginal\-rho\-low\), and the direction’s mechanism hint\.*Output:*exactly one rewritten spec with a different mechanism\.
> *System\.*“You are a feature\-improvement expert for\{domain\}time\-series forecasting\. You receive a candidate rejected by the novelty gate, its rejection reason, and the direction’s mechanism hint\. Rewrite one new spec with a*different*mechanism—you must change at least one oftransform,feat\_b, orop; renaming alone does not count\. Output JSON undermissing\_featurescontaining exactly one spec\.” *User\.*“Rejected candidate:\{spec\}\. reject\_reason:\{\.\.\.\}\. mechanism\_hint:\{\.\.\.\}\. Rewrite one spec with a different mechanism\.”
#### Base Proposer \(invents new base columns that join the channel pool and become searchable next round\)\.
*Inputs:*round index, existing base\-feature names, recent residual statistics, and chronic worst subgroups\.*Output:*new univariate base\-feature functions\.
> *System\.*“You design new base\-level features for\{domain\}time\-series forecasting that feed the residual\-correction GBDT\. Each feature is a Python functioncompute\_NAME\(history, chronos\_pred\)returning one scalar, wherechronos\_predis the backbone forecast\. Allowed:numpy\(asnp\), arithmetic, and simple control flow\. Forbidden: anyimport, I/O,exec/eval/open, and referencing bothhistoryandchronos\_predin one function \(each base feat must summarize a single input; cross\-source interactions are left to the atomic search\)\. Output strict JSON underproposals, each\{feature\_name, family, rationale, code\}withfamilyfrom the allowed set\. Target patterns the existing base set under\-represents and keep proposals mutually orthogonal\.” *User\.*“Round\{r\}\. Generate\{n\}new base features\. Existing base names \(avoid duplicates\):\{\.\.\.\}\. Recent residual stats:\{\.\.\.\}\. Chronic worst subgroups:\{\.\.\.\}\. Output strict JSON with aproposalslist; each function takes \(history,chronos\_pred\) and returns a float\.”
## Appendix IFramework Comparison: With vs\. Without the LLM
Table[S5](https://arxiv.org/html/2608.05207#A9.T5)contrasts theCrafterpipeline with and without the LLM agent at the component level\. The two columns are theCrafter\-noLLM andCraftercolumns of Table[S12](https://arxiv.org/html/2608.05207#A14.T12); the counts are means per cell over the GPT\-5\.2 grid \(selected==admitted to the corrector\)\. Adding the compositional search on top of the covariate\-only corrector moves accuracy by≈0\{\\approx\}0\(better\-or\-equal on10/3010/30cells, mean\+0\.006\+0\.006wMAPE\); adding the LLM on top of the search is the decisive step \(mean−0\.037\-0\.037wMAPE, better on28/3028/30cells\), and the gain concentrates on the weak\-backbone cells where a semantic signal exists\. At the level of accepted features, the LLM also dominates the final corrector: about three\-quarters of admitted features are LLM\-proposed, with LLM\-combo features the most efficient and LLM\-code features almost always pruned \(per\-source counts in Table[S6](https://arxiv.org/html/2608.05207#A10.T6), broken down by dataset and backbone in §[J](https://arxiv.org/html/2608.05207#A10)\)\.
Table S5:Component\-level comparison ofCrafterwith vs\. without the LLM agent\. Counts are per\-cell means over the GPT\-5\.2 grid\.
## Appendix JFeature\-source anatomy by dataset and backbone
Table[S6](https://arxiv.org/html/2608.05207#A10.T6)aggregates the per\-source*proposed*and*selected*feature counts and survival rates over the grid \(the source for Fig\. 3a\); Tables[S7](https://arxiv.org/html/2608.05207#A10.T7)and[S8](https://arxiv.org/html/2608.05207#A10.T8)break them down by dataset and by backbone \(GPT\-5\.2 fine\-tuning grid, pooled over the twoCraftervariants and three splits\)\. Three patterns hold\. The selected total rises with residual headroom: the weak\-backbone, signal\-richfavoritacells admit the most features, while the small or saturatedepfandbizitobscells sit at the floor\. LLM\-flag features are admitted mainly where the data has explicit calendar structure \(favorita\); onepfandbizitobsmost are trimmed\. LLM\-code features are pruned almost everywhere\. The mix is insensitive to the backbone \(Moirai≈\{\\approx\}Moirai\-2\.0\)\.
Table S6:Feature\-source anatomy \(aggregate\)\.Per\-cell mean*proposed*\(P\) and*selected*\(S, admitted to the corrector\) feature counts by source, with survival rateS/PS/\\text\{P\}, over the fine\-tuning grid \(both Moirai backbones, three splits\)\. LLM features supply about three\-quarters of the selected set; LLM\-code is almost always pruned\.Table S7:Feature\-source anatomy by dataset \(per\-cell mean, proposed/selected\)\.PP,SSare the totals; LLM%\(S\) is the LLM share of selected features\.Table S8:Feature\-source anatomy by backbone \(per\-cell mean, proposed/selected\)\.
## Appendix KRepresentative Accepted Features
The examples below illustrate what each source contributes\. They are drawn from the per\-feature audit logs of the DeepSeekCrafter†runs \(rolling split88, seed4242\), which record every candidate’s acceptance correlationρ\\rhoand its split gain in the final corrector\. “Share” is the feature’s fraction of its cell’s total corrector gain; gains are target\-scale dependent, so shares compare only within a cell\.
#### Atomic\.
Compositions the search assembled from raw channels; the recorded name encodes the expression, e\.g\.,binary\_safe\_div\_log\-cov\_Promo\_softplus\-dow\. That feature,safe\_div\(log\(Promo\),softplus\(dow\)\)\\mathrm\{safe\\\_div\}\(\\log\(\\texttt\{Promo\}\),\\,\\mathrm\{softplus\}\(\\texttt\{dow\}\)\)\(rossmann, Moirai\-2;ρ=0\.61\\rho\{=\}0\.61, share50%50\\%\), is a promotion\-by\-weekday interaction, and the identical spec is also accepted in the decoupledCrafterrun of the same cell\.min\(softplus\(hour\),is\_weekend\)\\min\(\\mathrm\{softplus\}\(\\texttt\{hour\}\),\\,\\texttt\{is\\\_weekend\}\)\(bizitobs, Timer;ρ=−0\.78\\rho\{=\}\{\-\}0\.78, share48%48\\%\) gates the hourly IT\-traffic residual by weekend and time of day\.min\(log\(holiday\),tanh\(is\_weekend\)\)\\min\(\\log\(\\texttt\{holiday\}\),\\,\\tanh\(\\texttt\{is\\\_weekend\}\)\)\(rohlik\) is selected under four of the five backbones; for a tree corrector it reduces to a re\-derived holiday indicator, typical of how the search re\-expresses information the schema already carries rather than adding semantics\.
#### LLM\-combo\.
Named combinations with a stated mechanism\.renewable\_penetration, the renewable\-generation forecast over the load forecast \(epf, Moirai;ρ=−0\.81\\rho\{=\}\{\-\}0\.81, share81%81\\%\), was proposed because “values above 1\.0 directly signal surplus conditions …which modulates price volatility”; a renewables\-over\-load variant is selected on all fiveepfbackbones\.baseline\_mean\_x\_open, the context mean×\\timesOpen\(rossmann, Moirai; share88%88\\%\), is the expected sales level when the store is open\.event\_ratio, the last value over the context mean \(favorita, Chronos; share37%37\\%\), was proposed to flag “sudden drops typical of earthquake shocks”\.
#### LLM\-flag\.
The LLM proposes the regime and its channels; the numeric thresholds are pool quantiles filled in by the enumerator\. Onrossmann\(Moirai\-2; share14%14\\%\) an anomaly direction targeting abnormal promotion days yields𝟏\[ctx\_std\>q75∧Promo\>0∧Open≥1\]\\mathbf\{1\}\[\\texttt\{ctx\\\_std\}\{\>\}q\_\{75\}\\wedge\\texttt\{Promo\}\{\>\}0\\wedge\\texttt\{Open\}\{\\geq\}1\], an open, high\-volatility store running a promotion\. Onbizitobs\(Toto;ρ=−0\.61\\rho\{=\}\{\-\}0\.61, share18%18\\%\),𝟏\[is\_weekend∨h<6\]\\mathbf\{1\}\[\\texttt\{is\\\_weekend\}\\vee h\{<\}6\]marks the low\-load regime of the API\-traffic series\. Onfavorita\(Moirai\),𝟏\[onpromotion≤0∧holiday≤0\]\\mathbf\{1\}\[\\texttt\{onpromotion\}\{\\leq\}0\\wedge\\texttt\{holiday\}\{\\leq\}0\]separates quiet baseline days from event days; holiday\-conditioned flags are selected on four of the fivefavoritabackbones\.
#### LLM\-code\.
Selected code features are rare \(Table[S6](https://arxiv.org/html/2608.05207#A10.T6)\), and the survivors compute statistics the grammar cannot express\.trend\_r\_squared\(favorita, Toto; share55%55\\%\) fits a line to the last2828days and returns itsR2R^\{2\}, proposed as “how strongly a linear model explains recent sales, complementary to slope magnitude”\.mk\_trend\_tau\(m5, Timer; share28%28\\%\) is Kendall’sτ\\tauof demand against time, a monotone\-trend strength robust to outliers; its accepted source is
```
def compute_mk_trend_tau(history, chronos_pred):
import numpy as np
from scipy import stats
if len(history) < 5:
return 0.0
tau, p = stats.kendalltau(
np.arange(len(history)), history)
return float(tau) if not np.isnan(tau) else 0.0
```
## Appendix LFeature transferability \(faithful replay\)
We test whether features learned on one backtest window stay useful on a later one\. For an ordered split pairS→TS\\\!\\to\\\!Twe load the accepted feature specifications \(combinations and synthesized code\) from the source cell, force\-carry them into the target split’s accepted set, re\-materialize the code features on the target data, and refit and evaluate the corrector with no new proposals \(LLM and search off, validation gating on\)\. Two references on the same target bound the effect:*self*, the target’s own accepted features replayed back \(a faithfulness upper bound that should reproduce nativeCrafter\), and*cov*, a covariate\-only corrector \(the no\-transfer lower bound\)\. The grid is two datasets \(epf,bizitobs\)×\\timesfive backbones×\\timesthree split pairs \(6→76\\\!\\to\\\!7,7→87\\\!\\to\\\!8,6→86\\\!\\to\\\!8\)×\\times\{cov, self, xfer\}=90=90cells, all completing\. The self column at splitSSreproduces nativeCrafterclosely, so the replay is faithful\. Table[S9](https://arxiv.org/html/2608.05207#A12.T9)summarizes the3030source\-cell triples and Table[S10](https://arxiv.org/html/2608.05207#A12.T10)lists every cell\.
Three findings hold\. Transferred features match the target’s own at the median \(rel\.\+0\.2%\+0\.2\\%,17/3017/30within±5%\\pm 5\\%\); the mean is pulled up only by a heavy tail\. Distance decay is mild: the median for the two\-split jump \(6→86\\\!\\to\\\!8\) is only a couple of percent above the adjacent\-split median\. The7/307/30catastrophic transfers \(worse than self by\>15%\>15\\%\) concentrate on the single\-seriesepfand on Moirai backbones—the shift\-sensitive corner of §5\.2—while the multivariatebizitobspanel transfers with median\+0\.0%\+0\.0\\%\. Low feature\-name overlap across splits therefore reflects equivalent\-feature redundancy, not non\-transferability\.
Table S9:Transferability summary over the3030source\-cell triples\. “rel\.” is the relative test wMAPE gap to the reference \(negative favors transfer\)\.Table S10:Per\-cell faithful replay \(test wMAPE×100\\times 100, lower is better\)\.*cov*: covariate\-only target corrector \(no\-transfer lower bound\);*self*: target’s own accepted features replayed \(faithfulness upper bound\);*xfer*: source\-split features replayed on the target\.Δs\\Delta\_\{\\text\{s\}\},Δc\\Delta\_\{\\text\{c\}\}are the relative gaps of xfer to self and to cov in %\. Pairs7→87\\\!\\to\\\!8and6→86\\\!\\to\\\!8share the same target \(split88\), so their cov/self match\.
## Appendix MFeature\-budget sweep
Table[S11](https://arxiv.org/html/2608.05207#A13.T11)sweeps the feature capKKover the3030GPT\-5\.2 cells \(each a three\-split mean\), backing Fig\. 2b and the budget claim of §5\.1\. The externals here are the honest budget\-5050sweep \(grid\_allbb\_uncap\), so theirK≥10K\{\\geq\}10columns are genuinely re\-generated rather than repeats of the small cap\. The win rate overRawis constant at80%80\\%across everyKK: the budget never changes which cellsCrafterwins, only by how much\. The margin over the best external system instead widens withKKand then saturates atK=20K\{=\}20\(K=20/50/K\{=\}20/50/all are nearly identical\), soK=20K\{=\}20is the practical operating point\. Split by residual headroom, the budget amplifies an existing gap rather than creating one: on raw\-weak cells \(external movesRawby≥2%\{\\geq\}2\\%\) the mean lift overRawreaches−0\.131\-0\.131atK=20K\{=\}20\(20/2020/20cells\), whereas on raw\-strong cells it sits near\+0\.008\+0\.008and a larger cap, if anything, fits more noise\.
Table S11:Feature\-budget sweep \(GPT\-5\.2,3030cells, three\-split mean\)\. Win rates and mean differences in test wMAPE forCrafter\(decoupled\) vs\.Rawand vs\. the best external system \(negative mean favorsCrafter\)\.
## Appendix NFull per\-cell results
Table[S12](https://arxiv.org/html/2608.05207#A14.T12)gives every cell of the GPT\-5\.2 main grid \(the source for Fig\. 2a\); Table[S13](https://arxiv.org/html/2608.05207#A14.T13)is the DeepSeek\-3\.4\-pro reproduction\. Figure[S1](https://arxiv.org/html/2608.05207#A14.F1)plots the per\-cell gain against backbone competence: the lift overRawtrends positive with the backbone’s own error but is noisy, which is why the body reports the cleaner per\-backbone aggregate \(Fig\. 2a\) rather than the raw scatter\. A single quantity—how much exploitable residual the frozen forecast leaves—orders the gain: the wrapper is net\-positive on the weak, high\-error cells to the right and neutral\-to\-negative on the saturated, low\-error cells to the left, with no separate dependence on dataset or on which foundation model produced the point\. The per\-cell scatter is the headroom condition of §5\.2 before averaging\.
Table S12:Main results\.Crafter\(two rightmost,graycolumns\) vs\. the frozen backbone \(Raw\), a head\-only backbone fine\-tune \(FT\), a covariate\-only corrector \(Cov\), the no\-LLM corrector \(noLLM\), and three external feature\-engineering systems, on the6×6=366\{\\times\}6\{=\}36\-cell grid \(test wMAPE×100\\times 100, lower is better\)\. Every feature\-engineering method uses thesame GPT\-5\.2 backendand harness, so only the feature*generator*differs\. The noLLM, external, andCraftercolumns average three rolling splits×\\timesthree seeds;Raw,Cov, andFTinvolve no LLM or search randomness and are single\-run on the same splits\. Per\-cell seed variability is quantified on the DeepSeek companion grid \(Table[S13](https://arxiv.org/html/2608.05207#A14.T13)caption; App\.[D](https://arxiv.org/html/2608.05207#A4.SS0.SSS0.Px3)\)\.chris Chronos \(univariate\);chr2is the covariate\-aware Chronos\-2, reported single\-seed \(42\) on the same splits\.Crafteris the full system;Crafter†adds source coupling \(§5\.2 of the main paper\)\.Bold==aCraftervariant beats*every*non\-fine\-tuned baseline in that row \(both may bold\); if neither does, the row’s best method is bold\. A star \(∗\) marks a boldCraftervariant that*also*beatsFT\.Raw,CovandFTare LLM\-independent and shared with Table[S13](https://arxiv.org/html/2608.05207#A14.T13); the stronger full\-parameter Moirai fine\-tune is in Table[S2](https://arxiv.org/html/2608.05207#A6.T2)\.Table S13:DeepSeek\-3\.4\-pro full grid\(cross\-LLM reproduction; test wMAPE×100\\times 100,6×6=366\{\\times\}6\{=\}36cells, lower is better\)\. Companion to the GPT\-5\.2 Table[S12](https://arxiv.org/html/2608.05207#A14.T12)with the same columns; the per\-backbone summary is in Table 1 of the main paper\.Raw,CovandFTare LLM\-independent and shared with Table[S12](https://arxiv.org/html/2608.05207#A14.T12)\. noLLM,Crafterand the externals are the DeepSeek runs;Crafter†is the DeepSeek source\-coupling variant \(3 splits, single\-seed\)\.Boldand star \(∗\) follow Table[S12](https://arxiv.org/html/2608.05207#A14.T12)\. Across the three seeds the median per\-cell std of the split\-averaged value is0\.020\.02\(noLLM\),0\.010\.01\(TSFresh\),0\.050\.05\(CAAFE\),0\.080\.08\(LLM\-FE\), and0\.490\.49\(Crafter\);Rawis seed\-deterministic, and the largestCrafterstds all sit onepfandbizitobs\. Thechr2\(Chronos\-2\) rows are single\-seed \(42\) on the same three splits, matching how Chronos\-2 is reported in Table[S12](https://arxiv.org/html/2608.05207#A14.T12)\.Figure S1:Gain vs\. backbone competence \(per cell, GPT\-5\.2\)\.Each point is one \(dataset, backbone\) cell:xxis the frozen backbone’s Raw wMAPE×100\\times 100,yyisCrafter’s lift over Raw\. The gain trends positive with backbone error \(more exploitable residual\) but is noisy across datasets; the per\-backbone means \(Fig\. 2a\) summarize the trend\. Residual headroom, not dataset or backbone identity, orders the gain: points on the high\-error right are consistently corrected, while saturated low\-error points on the left cluster near zero or below, the per\-cell form of the headroom condition \(§5\.2\)\.Similar Articles
Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction
This paper introduces DARC, a diagnosis-guided recovery harness that makes agent self-correction selective by profiling failure modes and pruning mismatched interventions before test-time correction, improving performance on ALFWorld, AppWorld, and XBRL Finance.
Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents
Forecast-Dojo is a replayable environment for benchmarking and training LLM forecasting agents, combining resolved prediction-market questions with dated news to enable repeated evaluation and learning from outcomes.
Downside-Controlled Online Forecast Combination under Delayed and Revised Outcomes
This paper proposes a method for improving frozen forecasters by combining static and online correctors to control downside risk, demonstrating gains up to 11.5% in benchmarks and reduced errors in electricity load forecasting.
Explanations-Driven Active Feature Acquisition for Algorithmic Recourse
The paper proposes an Explanation-Driven Feature Acquisition (EDFA) method that unifies counterfactual, semifactual, and alterfactual explanations to jointly address algorithmic recourse and feature acquisition, aiming for lower-cost and more actionable recourse with validity guarantees.
Search Discipline for Long-Horizon Research Agents
This paper identifies a failure mode in long-horizon research agents where optimizing an aggregate metric can select candidates that improve the headline number but break critical subgroups (inversion). It proposes a search-discipline protocol with an external control loop that audits candidates based on disaggregated behavior rather than the score.