ReLOBGen: Replayable Limit Order Book Message Generation
Summary
The paper proposes ReLOBGen, a limit order book message generator that ensures replayability by construction by selecting referenced orders from the current resting book and masking invalid tokens, achieving 100% replayability in 500-message rollouts and a 2.7-3.6x speedup per replayed message over the LOBS5 baseline.
View Cached Full Text
Cached at: 09/30/26, 09:39 AM
# Replayable Limit Order Book Message Generation
Source: [https://arxiv.org/html/2609.35867](https://arxiv.org/html/2609.35867)
Kiseop LeeAffiliation:Department of Statistics, Purdue UniversityEmail:[kiseop@purdue\.edu](mailto:)Bohyung HanAffiliation:ECE &Affiliation:IPAI, Seoul National UniversityEmail:[bhhan@snu\.ac\.kr](mailto:)
###### Abstract
We propose ReLOBGen, a method for generating limit order book \(LOB\) messages that are replayable by construction\. Replayability is required for closed\-loop market simulation, yet existing LOB message generators may produce non\-replayable raw messages, i\.e\., messages inconsistent with the current market state\. These generators therefore rely on post\-hoc correction or rejection followed by resampling, which may alter the replayed message distribution or increase inference cost\. ReLOBGen instead ensures replayability during generation: it selects the referenced order from the resting orders in the current LOB and then generates the remaining message fields to be consistent with that order and the market state\. For realistic reference selection, ReLOBGen samples from a learned distribution over eligible resting orders, efficiently computed from cached order representations and a context\-dependent query\. It then enforces the consistency of the remaining fields by masking out invalid tokens\. Together, these components enable efficient generation of realistic messages without post\-hoc correction or resampling\. In 500\-message rollouts, ReLOBGen achieves 100% replayability, improves market realism, particularly for top\-of\-book statistics and the relative prices of LOB messages, and provides a2\.7–3\.6×2\.7\\text\{\-\-\}3\.6\\timesspeedup per replayed message over the LOBS5 baseline\.
## 1Introduction
A limit order book \(LOB\) contains buy and sell orders waiting to be traded, called resting orders\. Each resting order is characterized by its price, size, and timestamp\. LOB messages specify changes to the book: add events introduce new resting orders, while cancellation and execution events reduce or remove a referenced resting order\. Recent work\([Nagy et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib1);[Wheeler and Varner, 2024](https://arxiv.org/html/2609.35867#bib.bib2)\)models sequences of LOB messages autoregressively, learning to generate the next message from preceding messages and LOB states\. These models have a range of potential applications in financial markets, such as training and evaluating trading strategies and analyzing market impact\([Nagy et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib1);[Li et al\., 2025](https://arxiv.org/html/2609.35867#bib.bib4)\)\. To simulate a market for these applications, the model and a rule\-based LOB simulator\([Byrd et al\., 2020](https://arxiv.org/html/2609.35867#bib.bib9);[Frey et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib10)\)operate in a closed loop, as illustrated in[Figure1](https://arxiv.org/html/2609.35867#S1.F1)\. At each step, the model generates a message, and the simulator updates the LOB according to that message; we refer to this update as*replay*\. The updated LOB then conditions the next generation step\.
However, an LOB message generator may produce raw messages that cannot be replayed against the current market state, such as a cancellation message that references a nonexistent resting order or whose size exceeds the referenced order’s remaining size\. For example, over 25% of raw generation attempts in our reproduction of the LOBS5 baseline are not replayable as generated\. Existing approaches therefore rely on method\-specific post\-processing to obtain a replayable message from a raw outputM^t\\widehat\{M\}\_\{t\}\. This includes correction or rejection followed by resampling if correction fails, as illustrated in the post\-processing stage of[Figure1](https://arxiv.org/html/2609.35867#S1.F1)\. These post\-processing procedures are often heuristic: LOBS5\([Nagy et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib1)\)and MarketGPT\([Wheeler and Varner, 2024](https://arxiv.org/html/2609.35867#bib.bib2)\)apply different fallback matching rules to identify the order referenced by a cancellation message, whereas MarS\([Li et al\., 2025](https://arxiv.org/html/2609.35867#bib.bib4)\)projects the generated cancellation price to the nearest available resting\-order price\. LOBS5 and MarketGPT also interpret the generated cancellation quantity differently: LOBS5 treats it as an absolute amount, whereas MarketGPT preserves the generated cancellation ratio\. Such corrections may alter message semantics and shift the distribution of replayed messages, while resampling increases inference cost\. This problem is structural: replayability depends on the exact current market state, which a finite message context and aggregated LOB states may not fully specify \(see[AppendixD](https://arxiv.org/html/2609.35867#A4)\)\.
We introduce ReLOBGen, which generates replayable LOB messages by construction and thus requires neither post\-hoc correction nor resampling\. Its design satisfies two jointly sufficient conditions for replayability: reference eligibility and event\-order compatibility—a non\-add event must reference an eligible resting order, and the event order must be consistent with the event, the market state, and any reference order it is applied to\. ReLOBGen’s message representation encodes the reference order using the target resting order’s current attributes, such as price and remaining quantity, so that these attributes can constrain the event\-order fields\. During generation, ReLOBGen first generates the event and then selects the reference order from the resting orders eligible for that event, ensuring reference eligibility\. Finally, it generates the event\-order fields while masking out invalid token values, ensuring event\-order compatibility\. For realistic reference selection, ReLOBGen samples from a learned distribution over eligible resting orders, efficiently computed by dot\-product scoring of cached order representations against a context\-dependent query\.
Empirically, ReLOBGen achieves 100% replayability, improves rollout realism, and generates messages2\.7–3\.6×2\.7\\text\{\-\-\}3\.6\\timesfaster than the LOBS5 baseline\([Nagy et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib1)\)in 500\-message rollouts on GOOG and INTC\. Using the same S5\([Smith et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib11)\)backbone as the baseline, ReLOBGen achieves this speedup through fewer autoregressive decoding steps and the elimination of resampling\. In terms of rollout realism, ReLOBGen more closely matches the marginal event\-type distribution and shows broad improvements across unconditional, conditional, and market\-impact evaluations with LOB\-Bench\([Nagy et al\., 2025](https://arxiv.org/html/2609.35867#bib.bib5)\)\. These improvements are particularly pronounced in metrics capturing top\-of\-book statistics and the prices of submitted and canceled orders relative to the current book\. An ablation shows that learned reference selection improves rollout realism over uniform selection from the same eligible resting orders\.
Overall, ReLOBGen makes every generated message replayable by construction—selecting reference orders from the eligible resting orders and using the selected orders to constrain event\-order generation\. Our results show that this guarantee can be achieved alongside improved rollout realism and computational efficiency\. Generated messages are replayed without modification, preserving their semantics and avoiding distribution changes introduced by method\-specific correction rules\. Together, these results establish ReLOBGen as a practical method for efficiently generating replayable and realistic LOB messages\.
Figure 1:Market simulation with an LOB message generator\.At each step, the generator produces a raw messageM^t\\widehat\{M\}\_\{t\}\. WhenM^t\\widehat\{M\}\_\{t\}is not replayable, existing approaches correct it when possible or otherwise reject it and resample\. ReLOBGen, however, generates replayable messages by construction and requires no post\-processing\. The LOB simulator replays the resulting messageMtM\_\{t\}to update the market state, which conditions the next generation step\.
## 2Related work
### 2\.1Modeling financial market microstructure
Data\-driven modeling of financial markets spans multiple targets, from future price movements to the evolution of the limit order book\. DeepLOB\([Zhang et al\., 2019](https://arxiv.org/html/2609.35867#bib.bib6)\)predicts future mid\-price movements from recent LOB observations\. Because the mid\-price provides only a compressed representation of market state, subsequent work directly forecasts or generates the joint evolution of prices and queued volumes across multiple book levels\([Jung and Lee, 2025](https://arxiv.org/html/2609.35867#bib.bib7);[Backhouse et al\., 2025](https://arxiv.org/html/2609.35867#bib.bib8)\)\.
Another line of work models order flow—the sequence of events that update the book—using stochastic processes\.[Smith et al\. \(2003\)](https://arxiv.org/html/2609.35867#bib.bib15)treat order flow as independent random events and analyze continuous double\-auction statistics, while[Cont et al\. \(2010\)](https://arxiv.org/html/2609.35867#bib.bib16)model limit, market, and cancellation events as Poisson flows calibrated to high\-frequency data\.[Huang et al\. \(2015\)](https://arxiv.org/html/2609.35867#bib.bib17)make event intensities depend on the current book state in a queue\-reactive model, whereas[Bacry et al\. \(2015\)](https://arxiv.org/html/2609.35867#bib.bib18)describe multivariate Hawkes processes that capture temporal clustering and self\- and cross\-excitation through history\-dependent intensities\.
More recent work uses neural autoregressive models to generate order flow\. LOBS5\([Nagy et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib1)\)uses an S5\-based state\-space model\([Smith et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib11)\)to generate tokenized LOB messages from LOBSTER data\([Huang and Polak, 2011](https://arxiv.org/html/2609.35867#bib.bib19)\)for two Nasdaq equities, conditioned on past messages and LOB states\. MarketGPT\([Wheeler and Varner, 2024](https://arxiv.org/html/2609.35867#bib.bib2)\)uses a decoder\-only Transformer\([Vaswani et al\., 2017](https://arxiv.org/html/2609.35867#bib.bib12)\)to generate tokenized ITCH messages from message histories, with full\-depth Nasdaq pretraining across 20 equities and asset\-specific fine\-tuning\. MarS\([Li et al\., 2025](https://arxiv.org/html/2609.35867#bib.bib4)\)combines order\-level and order\-batch models to generate orders conditioned on past orders and LOB states, training on 500 liquid Chinese equities\. TradeFM\([Kawawa\-Beaudan et al\., 2026](https://arxiv.org/html/2609.35867#bib.bib3)\)scales Transformer\-based order\-flow generation to over 9,000 US equities, conditioning on LOB message streams\.
### 2\.2Evaluating message\-generation models
Message\-generation models are commonly evaluated using next\-token perplexity and market realism in autoregressive rollouts\([Nagy et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib1);[Kawawa\-Beaudan et al\., 2026](https://arxiv.org/html/2609.35867#bib.bib3)\)\. LOB\-Bench\([Nagy et al\., 2025](https://arxiv.org/html/2609.35867#bib.bib5)\)assesses realism by comparing distributions of market microstructure statistics, such as spreads and message inter\-arrival times, extracted from real and generated messages and LOB trajectories\. The 21 metrics span six groups: top\-of\-book conditions \(*State*\), event timing \(*Times*\), resting liquidity \(*Volumes*\), the locations of order submissions and cancellations, measured by price distance and book\-level rank \(*Depths*and*Levels*, respectively\), and trading activity and order\-flow imbalance \(*Trades*\)\. These comparisons cover unconditional and conditional distributions, as well as error accumulation measured by changes in distributional divergence over the rollout horizon\. It also evaluates market impact through price responses to order events and temporal correlations between order events\. Additionally, we evaluate raw\-message replayability by measuring how often generated messages require correction or rejection and whether they violate reference eligibility or event\-order compatibility\.
## 3Problem formulation
### 3\.1LOB message generation model
We first define the main objects used in LOB message generation\. Let𝒪t−1\\mathcal\{O\}\_\{t\-1\}denote the set of resting orders at stept−1t\-1\. An LOB messageMt=\(Et,Rt,Xt\)M\_\{t\}=\(E\_\{t\},R\_\{t\},X\_\{t\}\)describes a change to this set:EtE\_\{t\}specifies the event type and side \(bid or ask\),RtR\_\{t\}denotes the reference order, andXtX\_\{t\}denotes the event order\. We consider four event types: add inserts a new resting order, cancel partially reduces a referenced order’s remaining quantity, delete removes the referenced order entirely, and execute reduces a referenced order’s remaining quantity through a trade\. The orders involved in these events are represented byXtX\_\{t\}andRtR\_\{t\}, each containing price, size, and timestamp fields\. For add events,XtX\_\{t\}describes the newly submitted order, and no reference order is required\. For non\-add events,RtR\_\{t\}identifies the referenced resting order, andXtX\_\{t\}specifies the price and quantity of the operation on that order\. ReplayingMtM\_\{t\}updates the resting\-order set to𝒪t\\mathcal\{O\}\_\{t\}\. Aggregating the remaining quantities in𝒪t\\mathcal\{O\}\_\{t\}by side and price and retaining the bestNNoccupied price levels on each side yields theNN\-level LOB stateBt\(N\)B\_\{t\}^\{\(N\)\}, which does not retain individual resting orders\.
An LOB message generation modelπθ\\pi\_\{\\theta\}learns the conditional distribution of the next messageMtM\_\{t\}given the precedingLLLOB messagesMt−L:t−1M\_\{t\-L:t\-1\}andNN\-level LOB statesBt−L:t−1\(N\)B\_\{t\-L:t\-1\}^\{\(N\)\}\. Each message is tokenized and modeled autoregressively using architectures such as Transformers\([Vaswani et al\., 2017](https://arxiv.org/html/2609.35867#bib.bib12)\)or state\-space models\([Smith et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib11)\)\. The generation model used in this work follows LOBS5\([Nagy et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib1)\), using an S5 architecture and representing each messageMtM\_\{t\}as a sequence of 22 field\-wise tokenszt,1:22z\_\{t,1:22\}\(see[Table8](https://arxiv.org/html/2609.35867#A2.T8)\)\. The model is trained to minimize the standard autoregressive negative log\-likelihood:
ℒNLL\(θ\)=−𝔼t,k\[logπθ\(zt,k∣Mt−L:t−1,Bt−L:t−1\(N\),zt,<k\)\]\.\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{NLL\}\}\(\\theta\)=\-\\mathbb\{E\}\_\{t,k\}\\left\[\\log\\pi\_\{\\theta\}\\left\(z\_\{t,k\}\\mid M\_\{t\-L:t\-1\},B\_\{t\-L:t\-1\}^\{\(N\)\},z\_\{t,<k\}\\right\)\\right\]\.\(1\)
Table 1:Event\-specific replayability conditions\.Here,side=Et\[side\]\\mathrm\{side\}=E\_\{t\}\[\\mathrm\{side\}\],𝒪t−1\(side\)\\mathcal\{O\}\_\{t\-1\}^\{\(\\mathrm\{side\}\)\}contains the resting orders on that side, andRtR\_\{t\}carries the referenced order’s current attributes\.BestAsk\\operatorname\{BestAsk\}andBestBid\\operatorname\{BestBid\}denote the lowest ask and highest bid prices in the current LOB state\.FrontOfQueue\\operatorname\{FrontOfQueue\}returns the oldest order resting at the best price on that side\. Under the LOB simulator’s price–time priority, an execute event allows only this order as its reference\.
### 3\.2Market simulation and replayability
To simulate a market trajectory autoregressively, the message generation modelπθ\\pi\_\{\\theta\}and the LOB simulator alternate between message generation and replay, as illustrated in[Figure1](https://arxiv.org/html/2609.35867#S1.F1)\. At each step, the model generates a message conditioned on the preceding message and LOB state histories\. The simulator replays the messageMtM\_\{t\}against the current resting\-order set:
𝒪t\\displaystyle\\mathcal\{O\}\_\{t\}=Replay\(𝒪t−1,Mt\),\\displaystyle=\\operatorname\{Replay\}\(\\mathcal\{O\}\_\{t\-1\},M\_\{t\}\),\(2\)whereReplay\\operatorname\{Replay\}denotes the event\-specific deterministic state\-transition operator used by the LOB simulator\. The messageMtM\_\{t\}and the LOB stateBt\(N\)B\_\{t\}^\{\(N\)\}aggregated from𝒪t\\mathcal\{O\}\_\{t\}are appended to their respective histories to condition the next generation step\.
However, a messageM^t\\widehat\{M\}\_\{t\}generated byπθ\\pi\_\{\\theta\}may not be replayable against the current resting\-order set𝒪t−1\\mathcal\{O\}\_\{t\-1\}\. For example, a non\-add event may reference an order that is absent from𝒪t−1\\mathcal\{O\}\_\{t\-1\}or ineligible for that event\. Even with an eligible reference, the event order may violate event\-specific constraints, such as exceeding the reference order’s remaining quantity; an add event may instead violate the price constraint imposed by the current market state\. In our reproduction of LOBS5 \(ref\.\-last\), over 25% of raw generation attempts require correction or rejection on both stocks \([Table2](https://arxiv.org/html/2609.35867#S5.T2)\)\. The larger MarketGPT model also reports that approximately 7% of generated messages cannot be corrected and require resampling\([Wheeler and Varner, 2024](https://arxiv.org/html/2609.35867#bib.bib2)\)\. Based on the LOB simulator’s replay rules, we define replayability as follows\.
###### Definition 1\(Replayability\)\.
A messageMt=\(Et,Rt,Xt\)M\_\{t\}=\(E\_\{t\},R\_\{t\},X\_\{t\}\)is replayable with respect to the current resting\-order set𝒪t−1\\mathcal\{O\}\_\{t\-1\}if it satisfies both conditions:
- •Reference eligibility\.For a non\-add event, the reference orderRtR\_\{t\}must belong to the eligible resting\-order set𝒞t\(Et,𝒪t−1\)\\mathcal\{C\}\_\{t\}\(E\_\{t\},\\mathcal\{O\}\_\{t\-1\}\)defined in[Table1](https://arxiv.org/html/2609.35867#S3.T1)\.
- •Event\-order compatibility\.The event orderXtX\_\{t\}satisfies the constraints in[Table1](https://arxiv.org/html/2609.35867#S3.T1)\.
When a generated messageM^t\\widehat\{M\}\_\{t\}is not replayable, existing approaches use method\-specific post\-processing to obtain a replayable messageMtM\_\{t\}\([Nagy et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib1);[Wheeler and Varner, 2024](https://arxiv.org/html/2609.35867#bib.bib2);[Li et al\., 2025](https://arxiv.org/html/2609.35867#bib.bib4)\)\. This includes correction or rejection followed by resampling, with event\-specific procedures detailed in[AppendixA](https://arxiv.org/html/2609.35867#A1)\. Although this post\-processing makesMtM\_\{t\}replayable, it introduces two limitations\. First, method\-specific corrections may modify generated fields or reinterpret their semantics, shifting the distribution of replayed messages away from that of raw model outputs\. Second, resampling rejected messages increases inference cost\. These limitations motivate generating messages that are replayable by construction while preserving rollout realism\.
## 4ReLOBGen: Replayable LOB message generation
ReLOBGen generates each LOB message by first generating the eventEtE\_\{t\}, then resolving the reference orderRtR\_\{t\}, and finally generating the event orderXtX\_\{t\}\. For non\-add events, resolving the referenced resting order first allows the event order’s price and quantity to be generated within the constraints imposed by that order, which makes reference\-first generation a natural choice\. We adapt the message representation to support this reference\-first generation process \([Section4\.1](https://arxiv.org/html/2609.35867#S4.SS1)\)\. During generation, ReLOBGen first generates the eventEtE\_\{t\}\. For non\-add events, ReLOBGen then selects the reference order from the eligible resting\-order set \([Section4\.2](https://arxiv.org/html/2609.35867#S4.SS2)\)\. Finally, it generates the event order to be consistent with the event, any selected reference order, and the current market state by masking out invalid token values \([Section4\.3](https://arxiv.org/html/2609.35867#S4.SS3)\)\. Together, these mechanisms produce replayable messages by construction without post\-hoc correction or resampling\.
\(a\) Reference\-first message representation\(b\) Resting\-order selection\(c\) Inference\-time maskingFigure 2:ReLOBGen: mechanisms for replayability by construction\.\(a\) The message representation placesRtR\_\{t\}beforeXtX\_\{t\}and encodes the referenced order’s remaining size rather than its initial submitted size\. \(b\) ReLOBGen samplesRtR\_\{t\}from a learned distribution over the eligible resting\-order set𝒞t\\mathcal\{C\}\_\{t\}\. This distribution is efficiently computed by applying a softmax to the scaled dot products between a query derived from the generation context \(the hidden state after processingEtE\_\{t\}\) and the cached order keys\. \(c\) ReLOBGen masks invalid token values and samples from the remaining valid support duringXtX\_\{t\}generation\.### 4\.1Reference\-first message representation
For ReLOBGen, we make two modifications to the LOBS5\([Nagy et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib1)\)message representation, as illustrated in[Figure2](https://arxiv.org/html/2609.35867#S4.F2)\(a\), while retaining its field\-wise tokenization and vocabulary \(see[Table8](https://arxiv.org/html/2609.35867#A2.T8)in[SectionB\.1](https://arxiv.org/html/2609.35867#A2.SS1)\)\. First, we placeRtR\_\{t\}beforeXtX\_\{t\}in the token sequence, yielding the orderEt→Rt→XtE\_\{t\}\\rightarrow R\_\{t\}\\rightarrow X\_\{t\}and making the reference available to determine the constraints onXtX\_\{t\}before its generation\. Second, we representRtR\_\{t\}using the referenced resting order’s current attributes so that these constraints reflect the current market state\. Specifically, we setRt\[size\]R\_\{t\}\[\\mathrm\{size\}\]to the order’s remaining size in𝒪t−1\\mathcal\{O\}\_\{t\-1\}rather than the initial submitted size used by LOBS5\. We train the message generation modelπθ\\pi\_\{\\theta\}on this representation using the negative log\-likelihood objective in[Section3\.1](https://arxiv.org/html/2609.35867#S3.SS1)\.
### 4\.2Resting\-order selection for reference eligibility
#### Resting\-order selection\.
ReLOBGen guarantees the reference eligibility condition in[Definition1](https://arxiv.org/html/2609.35867#Thmdefinition1)by selectingRtR\_\{t\}only from the eligible resting\-order set𝒞t\(Et,𝒪t−1\)\\mathcal\{C\}\_\{t\}\(E\_\{t\},\\mathcal\{O\}\_\{t\-1\}\)specified in[Table1](https://arxiv.org/html/2609.35867#S3.T1)\. Note that the number of model forward passes per message is reduced by selectingRtR\_\{t\}rather than autoregressively generating its eight tokens, and non\-add events with no eligible resting orders are masked during event generation\.
While selecting any eligible resting order guarantees reference eligibility, the particular choice still affects the realism of the generated LOB messages\. ReLOBGen therefore samplesRtR\_\{t\}from a learned categorical distributionpϕ\(⋅∣𝐡t,𝒞t\)p\_\{\\phi\}\(\\cdot\\mid\\mathbf\{h\}\_\{t\},\\mathcal\{C\}\_\{t\}\)over the eligible resting orders, where𝐡t\\mathbf\{h\}\_\{t\}denotes the generation context up to and includingEtE\_\{t\}\. With the base model frozen, we train this distribution by minimizing the cross\-entropy for the observed reference orderRt⋆R\_\{t\}^\{\\star\}:
ℒsel=−𝔼t\[logpϕ\(Rt⋆∣𝐡t,𝒞t\)\]\.\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{sel\}\}=\-\\mathbb\{E\}\_\{t\}\\left\[\\log p\_\{\\phi\}\(R\_\{t\}^\{\\star\}\\mid\\mathbf\{h\}\_\{t\},\\mathcal\{C\}\_\{t\}\)\\right\]\.\(3\)
#### Efficient resting\-order selection with cached keys\.
ReLOBGen implements resting\-order selection efficiently by exploiting a structural property of LOB messages: each message affects only one resting\-order entry\. We therefore cache a key𝐤o\\mathbf\{k\}\_\{o\}for each resting order and reuse the keys of unchanged orders\. Using query–key scoring similar to dense retrieval\([Karpukhin et al\., 2020](https://arxiv.org/html/2609.35867#bib.bib13)\), each eligible resting order𝐨∈𝒞t\(Et,𝒪t−1\)\\mathbf\{o\}\\in\\mathcal\{C\}\_\{t\}\(E\_\{t\},\\mathcal\{O\}\_\{t\-1\}\)is scored by the scaled dot product between its cached key𝐤o\\mathbf\{k\}\_\{o\}and a query𝐪t\\mathbf\{q\}\_\{t\}representing the current generation context, as illustrated in[Figure2](https://arxiv.org/html/2609.35867#S4.F2)\(b\):
sϕ\(𝐨,𝐡t\)=𝐪t⊤𝐤o/d,\\displaystyle s\_\{\\phi\}\(\\mathbf\{o\},\\mathbf\{h\}\_\{t\}\)=\\mathbf\{q\}\_\{t\}^\{\\top\}\\mathbf\{k\}\_\{o\}/\\sqrt\{d\},\(4\)whereddis the key and query dimension\. Applying a softmax over these scores within the eligible resting\-order set yieldspϕ\(⋅∣𝐡t,𝒞t\)p\_\{\\phi\}\(\\cdot\\mid\\mathbf\{h\}\_\{t\},\\mathcal\{C\}\_\{t\}\)\.
To keep cached keys reusable as the market moves, we encode resting orders relative to a fixed mid\-price anchor rather than the current mid\-price, while the query incorporates the current generation context and the displacement of the current mid\-price from this anchor\. We represent generation context𝐡t\\mathbf\{h\}\_\{t\}using the base model’s hidden state and compute the order key and shared query using lightweight MLPs,forderf\_\{\\mathrm\{order\}\}andfqueryf\_\{\\mathrm\{query\}\}:
𝐤o\\displaystyle\\mathbf\{k\}\_\{o\}=forder\(𝐨,midanchor\),\\displaystyle=f\_\{\\mathrm\{order\}\}\(\\mathbf\{o\};\\mathrm\{mid\}\_\{\\mathrm\{anchor\}\}\),𝐪t\\displaystyle\\mathbf\{q\}\_\{t\}=fquery\(𝐡t,midt−1,midanchor\)\.\\displaystyle=f\_\{\\mathrm\{query\}\}\(\\mathbf\{h\}\_\{t\},\\mathrm\{mid\}\_\{t\-1\};\\mathrm\{mid\}\_\{\\mathrm\{anchor\}\}\)\.\(5\)Here,midanchor\\mathrm\{mid\}\_\{\\mathrm\{anchor\}\}is the fixed mid\-price anchor, andmidt−1\\mathrm\{mid\}\_\{t\-1\}is the mid\-price at stept−1t\-1\.
At the start of each rollout, we setmidanchor\\mathrm\{mid\}\_\{\\mathrm\{anchor\}\}to the initial mid\-price\. When an order is added or updated, we compute its key from the order’s current attributes; when it is removed, we discard its key; we maintain cached keys for unaffected orders\. When trainingforderf\_\{\\mathrm\{order\}\}andfqueryf\_\{\\mathrm\{query\}\}, we samplemidanchor\\mathrm\{mid\}\_\{\\mathrm\{anchor\}\}from earlier mid\-prices in the context to expose them to price drift relative to the anchor\.
### 4\.3Inference\-time masking for event\-order compatibility
ReLOBGen guarantees event\-order compatibility by restricting each field inXtX\_\{t\}to the valid support determined byEtE\_\{t\},RtR\_\{t\}, and the current market state, as specified in[Table1](https://arxiv.org/html/2609.35867#S3.T1)\.[Figure2](https://arxiv.org/html/2609.35867#S4.F2)\(c\) illustrates this process: the model masks invalid tokens and samples from the remaining support, as in constrained decoding for structured generation\([Scholak et al\., 2021](https://arxiv.org/html/2609.35867#bib.bib14)\)\. When the valid support is a singleton, ReLOBGen directly assigns the unique valid token, further reducing the number of autoregressive model forward passes\. The pretrained model assigns negligible probability mass to invalid tokens on teacher\-forced test windows—less than0\.14%0\.14\\%on GOOG and0\.29%0\.29\\%on INTC, suggesting a limited impact of forcing and masking\.
## 5Experiments
### 5\.1Experimental setup
#### Data and model training\.
We use market\-by\-order \(MBO\) data for GOOG and INTC from Databento111[https://databento\.com/portal/catalog/us\-equities\#XNAS\.ITCH](https://databento.com/portal/catalog/us-equities#XNAS.ITCH), with 2025 data for training and separate periods in January 2026 for validation and testing\. We follow the preprocessing and field\-wise tokenization of LOBS5\([Nagy et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib1)\), except for the reference\-field semantics described in[Section4\.1](https://arxiv.org/html/2609.35867#S4.SS1)\(see[SectionB\.1](https://arxiv.org/html/2609.35867#A2.SS1)for details\)\. We train 35M\-parameter models using the same S5 backbone as the scaled\-up LOBS5 in[Nagy et al\. \(2025\)](https://arxiv.org/html/2609.35867#bib.bib5), with reference\-last and reference\-first token orders\. We additionally train two lightweight MLP projection heads with 2\.3M parameters in total for learned reference selection \(see[SectionB\.2](https://arxiv.org/html/2609.35867#A2.SS2)for training details\)\.
#### Rollout evaluation\.
We compare the following configurations on the same 1,000 rollout windows per stock, with 500 successfully replayed messages per rollout, matching the rollout length reported in LOB\-Bench\([Nagy et al\., 2025](https://arxiv.org/html/2609.35867#bib.bib5)\)\(see[SectionB\.3](https://arxiv.org/html/2609.35867#A2.SS3)for details\)\. ReLOBGen combines the reference\-first model with learned selection from the eligible resting\-order set and inference\-time masking\. Our baselines autoregressively generate messages with either model and apply the LOBS5 post\-processing procedure \(see[SectionA\.1](https://arxiv.org/html/2609.35867#A1.SS1)for details\); we denote them by LOBS5 \(ref\.\-last\) and LOBS5 \(ref\.\-first\)\. Ref\.\-first \+ uniform uses the reference\-first model and uniform reference selection with inference\-time masking\. All configurations start from the same exact resting\-order set, which is maintained by the JAX\-LOB simulator\([Frey et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib10)\)and available to both ReLOBGen and baseline post\-processing\. We assess replayability on raw generation attempts before post\-processing and evaluate rollout realism from replayed messages and resulting LOB trajectories using LOB\-Bench and the marginal event\-type distribution\.
Table 2:Post\-processing rates, replayability violations, and runtime\.Runtime is measured on an NVIDIA RTX A6000 GPU and reported as mean±\\pmstandard deviation\. Aborted rollouts—59 and 17 for LOBS5 \(ref\.\-last\) and LOBS5 \(ref\.\-first\) on GOOG, respectively, and two for each baseline on INTC—are excluded from evaluation\. Neither ReLOBGen nor Ref\.\-first \+ uniform requires any restarts on either stock\. Lower is better, and bold indicates the lowest value\.
### 5\.2Replayability
[Table2](https://arxiv.org/html/2609.35867#S5.T2)summarizes post\-processing and replayability violation rates, with an event\-wise breakdown in[SectionC\.1](https://arxiv.org/html/2609.35867#A3.SS1)\. For non\-add events, event\-order compatibility is assessed against the generated reference fields, regardless of whether the reference is eligible\. Both LOBS5 baselines frequently require correction or rejection, with reference eligibility violations far more frequent than event\-order compatibility violations\. This gap may partly reflect the difficulty of identifying eligible resting orders, many of which are not individually specified by the finite message context or the aggregated LOB state, as analyzed in[AppendixD](https://arxiv.org/html/2609.35867#A4)\. Both baselines also require rollout restarts after 100 consecutive failed message\-generation attempts\. On the other hand, ReLOBGen satisfies both replayability conditions by construction, recording zero violations and requiring no post\-processing throughout the evaluated rollouts\. Note that LOBS5 \(ref\.\-last\) still requires post\-processing when its resting\-order set is initialized with INIT orders, as shown in[SectionC\.2](https://arxiv.org/html/2609.35867#A3.SS2)\.
Relative to LOBS5 \(ref\.\-last\), ReLOBGen achieves2\.7×2\.7\\timesand3\.6×3\.6\\timesspeedups in runtime per replayed message on GOOG and INTC, respectively\. Runtime covers the full closed\-loop simulation, including message generation, post\-processing when required, and simulator state updates\. These gains come from reduced autoregressive decoding and the elimination of post\-processing\. Selecting the reference order rather than generating its tokens autoregressively and forcing singleton\-valid field tokens reduces non\-add decoding from 17 to 7–8 model forward passes\. Eliminating rejection also removes the1/\(1−r\)1/\(1\-r\)factor in attempts per replayed message, whererrdenotes the rejection rate\. The larger speedup on INTC might reflect higher reference\-matching costs in the baseline due to its larger resting\-order population\.
Table 3:LOB\-Bench evaluation: metric\-group summary of unconditionalL1L\_\{1\}distances\.Overall averages all metrics, while each metric\-group column averages the metrics in its group\. Lower is better, and bold indicates the lowest value\.Table 4:LOB\-Bench evaluation: conditionalL1L\_\{1\}distances and market\-impact discrepancies\.Market\-impact discrepancies cover market orders \(MO\), limit orders \(LO\), and cancellations \(CA\), with subscripts 0 and 1 for events without and with an immediate mid\-price change\. Conditional 99% percentile bootstrap confidence intervals have half\-widths below 0\.004\. Lower is better, and bold indicates the lowest value\.
### 5\.3Rollout realism
Beyond generating fully replayable messages, ReLOBGen also improves rollout realism over the LOBS5 baseline, with broadly lower unconditionalL1L\_\{1\}distances across metric groups, as shown in[Table3](https://arxiv.org/html/2609.35867#S5.T3)\. These improvements are particularly pronounced in the State, Depths, and Levels metric groups\. State captures top\-of\-book statistics, while Depths and Levels characterize the locations of order submissions and cancellations, measured by price distance and book\-level rank, respectively\. At the individual\-metric level, ReLOBGen achieves lower point estimates than LOBS5 on 18 of 21 metrics for GOOG, with statistically significant improvements on 16 of those 18\. On INTC, it achieves lower point estimates on 20 of 21 metrics, with statistically significant improvements on 18 of those 20\. Statistical significance is assessed using non\-overlapping 99% bootstrap confidence intervals\. DetailedL1L\_\{1\}and Wasserstein results are shown in[Figures5](https://arxiv.org/html/2609.35867#A3.F5)and[6](https://arxiv.org/html/2609.35867#A3.F6), respectively, with metric\-group averages for Wasserstein distances reported in[Table12](https://arxiv.org/html/2609.35867#A3.T12), all in[SectionC\.3](https://arxiv.org/html/2609.35867#A3.SS3)\.
[Table4](https://arxiv.org/html/2609.35867#S5.T4)shows that ReLOBGen achieves lower average conditional and market\-impact discrepancies than LOBS5 on both stocks, reducing conditional averages \(0\.41 to 0\.32 on GOOG; 0\.20 to 0\.10 on INTC\) and market\-impact averages \(10\.6 to 5\.2 on GOOG; 5\.6 to 4\.5 on INTC\)\.[Figure3](https://arxiv.org/html/2609.35867#S5.F3)\(a\) further illustrates the difference in lagged price\-response curves: LOBS5 tends to produce largely flat responses, whereas ReLOBGen better captures those observed in the real data\. Regarding error accumulation,[Figure3](https://arxiv.org/html/2609.35867#S5.F3)\(b\) shows lower errors for ReLOBGen on the two illustrated metrics, suggesting potential for more stable generation over longer rollout horizons\. ReLOBGen also better preserves the marginal event\-type distribution, reducing the total variation distance from 9\.6 to 6\.2 percentage points on GOOG, as detailed in[Table14](https://arxiv.org/html/2609.35867#A3.T14)in[SectionC\.4](https://arxiv.org/html/2609.35867#A3.SS4)\. Additional LOB\-Bench evaluation results, including conditional Wasserstein results \([Table13](https://arxiv.org/html/2609.35867#A3.T13)\), the full market\-impact response curves \([Figure7](https://arxiv.org/html/2609.35867#A3.F7)\), and error accumulation for all metrics \([Figures8](https://arxiv.org/html/2609.35867#A3.F8)and[9](https://arxiv.org/html/2609.35867#A3.F9)\), are provided in[SectionC\.3](https://arxiv.org/html/2609.35867#A3.SS3)\.
\(a\) Market impact\(b\) Error accumulationFigure 3:LOB\-Bench evaluation: market\-impact response curves and error accumulation on GOOG\.\(a\) Responses after market orders without \(MO0MO\_\{0\}\) and with \(MO1MO\_\{1\}\) an immediate mid\-price change; shading shows 99% bootstrap confidence intervals\. \(b\)L1L\_\{1\}distance over successive 100\-message intervals\.Table 5:Teacher\-forced NLLs under reference\-last and reference\-first token orders\.For fields represented by multiple tokens, NLLs are summed across tokens\. Shaded columns highlight the price and size fields of the reference and event orders\. Lower is better\.
### 5\.4Ablation studies
#### Effect of learned reference selection\.
Although uniform selection generates replayable messages, it does not necessarily improve rollout realism\. As shown in[Table3](https://arxiv.org/html/2609.35867#S5.T3), it only slightly reduces the overallL1L\_\{1\}distance relative to LOBS5 \(ref\.\-last\) on GOOG and even increases it on INTC, from 0\.22 to 0\.30\. Thus, which eligible order is selected matters for rollout realism\.
[Tables3](https://arxiv.org/html/2609.35867#S5.T3)and[4](https://arxiv.org/html/2609.35867#S5.T4)show that learned selection improves overall rollout realism over uniform selection on both stocks\. Learned selection improves reference\-order choices for cancel and delete events, with lowerL1L\_\{1\}distances on all five related metrics for both stocks, as shown in[Figure5](https://arxiv.org/html/2609.35867#A3.F5)in[SectionC\.3](https://arxiv.org/html/2609.35867#A3.SS3): log time\-to\-cancel and the bid\- and ask\-side cancellation depth and level distributions\. The gains are especially pronounced on INTC, where the top ten price levels of the LOB contain 2\.7 times as many resting orders as in GOOG, consistent with the potential benefit of learned over uniform selection in larger eligible sets\.
#### Effect of reference\-first factorization\.
We evaluate field\-wise teacher\-forced NLLs under the two token orders on 10,000 windows from the test split\.[Table5](https://arxiv.org/html/2609.35867#S5.T5)shows similar overall NLLs and similar field\-wise NLLs for fields other than price and size\. The price and size NLLs suggest that reference\-first ordering redistributes predictive uncertainty from the event\-order fields to the reference\-order fields that now precede them: reference\-order NLLs increase, while the corresponding event\-order NLLs decrease\. This redistribution is consistent with the price and size dependencies in[Table1](https://arxiv.org/html/2609.35867#S3.T1), which allow ground\-truth reference\-order fields to inform subsequent event\-order prediction under teacher forcing\. During inference, ReLOBGen selects an eligible resting order from a learned distribution and provides its current attributes before generating the event order, allowing subsequent generation to be conditioned and constrained by the selected reference\. This use of the reference during generation may contribute to ReLOBGen’s improved rollout realism\.
## 6Conclusion
We introduced ReLOBGen, a method for generating LOB messages that are replayable by construction against the current market state\. We organized generation around two replayability conditions—reference eligibility and event\-order compatibility—and satisfied them through a reference\-first factorization that selects an eligible resting order before constraining event\-order generation through inference\-time masking\. We further proposed a learned, efficiently computable distribution over eligible resting orders for realistic reference selection\. Empirically, ReLOBGen achieved 100% replayability in 500\-message rollouts on GOOG and INTC, improved most evaluated realism metrics over the LOBS5 baseline, and achieved a2\.7–3\.6×2\.7\\text\{\-\-\}3\.6\\timesspeedup per replayed message\. These results show that replayability by construction can be achieved alongside improved rollout realism and computational efficiency\. Future work can build on ReLOBGen to train and evaluate trading strategies and analyze market impact in closed\-loop market simulation\.
## References
- Backhouseet al\.\(2025\)A\. Backhouse, K\. Li, J\. Foerster, A\. Calinescu, and S\. ZohrenPainting the market: generative diffusion models for financial limit order book simulation and forecasting\.arXiv preprint arXiv:2509\.05107\.Cited by:[§2\.1](https://arxiv.org/html/2609.35867#S2.SS1.p1.1)\.
- Bacryet al\.\(2015\)E\. Bacry, I\. Mastromatteo, and J\. MuzyHawkes processes in finance\.Market Microstructure and Liquidity1\(01\),pp\. 1550005\.Cited by:[§2\.1](https://arxiv.org/html/2609.35867#S2.SS1.p2.1)\.
- Byrdet al\.\(2020\)D\. Byrd, M\. Hybinette, and T\. H\. BalchABIDES: towards high\-fidelity multi\-agent market simulation\.InProceedings of the 2020 ACM SIGSIM Conference on Principles of Advanced Discrete Simulation,Cited by:[§1](https://arxiv.org/html/2609.35867#S1.p1.1)\.
- Contet al\.\(2010\)R\. Cont, S\. Stoikov, and R\. TalrejaA stochastic model for order book dynamics\.Operations research58\(3\),pp\. 549–563\.Cited by:[§2\.1](https://arxiv.org/html/2609.35867#S2.SS1.p2.1)\.
- Freyet al\.\(2023\)S\. Y\. Frey, K\. Li, P\. Nagy, S\. Sapora, C\. Lu, S\. Zohren, J\. Foerster, and A\. CalinescuJAX\-LOB: a gpu\-accelerated limit order book simulator to unlock large scale reinforcement learning for trading\.InICAIF,Cited by:[§1](https://arxiv.org/html/2609.35867#S1.p1.1),[§5\.1](https://arxiv.org/html/2609.35867#S5.SS1.SSS0.Px2.p1.1)\.
- Huang and Polak \(2011\)R\. Huang and T\. PolakLOBSTER: limit order book reconstruction system\.Available at SSRN 1977207\.Cited by:[§B\.1](https://arxiv.org/html/2609.35867#A2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2609.35867#S2.SS1.p3.1)\.
- Huanget al\.\(2015\)W\. Huang, C\. Lehalle, and M\. RosenbaumSimulating and analyzing order book data: the queue\-reactive model\.Journal of the American Statistical Association110\(509\),pp\. 107–122\.Cited by:[§2\.1](https://arxiv.org/html/2609.35867#S2.SS1.p2.1)\.
- Jung and Lee \(2025\)J\. Jung and K\. LeeAttention\-based reading, highlighting, and forecasting of the limit order book\.Quantitative Finance\.Cited by:[§2\.1](https://arxiv.org/html/2609.35867#S2.SS1.p1.1)\.
- Karpukhinet al\.\(2020\)V\. Karpukhin, B\. Oguz, S\. Min, P\. Lewis, L\. Wu, S\. Edunov, D\. Chen, and W\. YihDense passage retrieval for open\-domain question answering\.InEMNLP,Cited by:[§4\.2](https://arxiv.org/html/2609.35867#S4.SS2.SSS0.Px2.p1.1)\.
- Kawawa\-Beaudanet al\.\(2026\)M\. Kawawa\-Beaudan, S\. Sood, K\. Papasotiriou, D\. Borrajo, and M\. VelosoTradeFM: a generative foundation model for trade\-flow and market microstructure\.arXiv preprint arXiv:2602\.23784\.Cited by:[§2\.1](https://arxiv.org/html/2609.35867#S2.SS1.p3.1),[§2\.2](https://arxiv.org/html/2609.35867#S2.SS2.p1.1)\.
- Liet al\.\(2025\)J\. Li, Y\. Liu, W\. Liu, S\. Fang, L\. Wang, C\. Xu, and J\. BianMarS: a financial market simulation engine powered by generative foundation model\.InICLR,Cited by:[§A\.3](https://arxiv.org/html/2609.35867#A1.SS3),[§1](https://arxiv.org/html/2609.35867#S1.p1.1),[§1](https://arxiv.org/html/2609.35867#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.35867#S2.SS1.p3.1),[§3\.2](https://arxiv.org/html/2609.35867#S3.SS2.p3.1)\.
- Nagyet al\.\(2025\)P\. Nagy, S\. Frey, K\. Li, B\. Sarkar, S\. Vyetrenko, S\. Zohren, A\. Calinescu, and J\. FoersterLOB\-Bench: benchmarking generative ai for finance–an application to limit order book data\.InICML,Cited by:[§B\.2](https://arxiv.org/html/2609.35867#A2.SS2.p1.1),[§1](https://arxiv.org/html/2609.35867#S1.p4.1),[§2\.2](https://arxiv.org/html/2609.35867#S2.SS2.p1.1),[§5\.1](https://arxiv.org/html/2609.35867#S5.SS1.SSS0.Px1.p1.1),[§5\.1](https://arxiv.org/html/2609.35867#S5.SS1.SSS0.Px2.p1.1)\.
- Nagyet al\.\(2023\)P\. Nagy, S\. Frey, S\. Sapora, K\. Li, A\. Calinescu, S\. Zohren, and J\. FoersterGenerative ai for end\-to\-end limit order book modelling: a token\-level autoregressive generative model of message flow using a deep state space network\.InICAIF,Cited by:[§A\.1](https://arxiv.org/html/2609.35867#A1.SS1),[§B\.1](https://arxiv.org/html/2609.35867#A2.SS1.SSS0.Px2.p2.1),[§B\.1](https://arxiv.org/html/2609.35867#A2.SS1.SSS0.Px3.p1.1),[§B\.1](https://arxiv.org/html/2609.35867#A2.SS1.p1.1),[§B\.2](https://arxiv.org/html/2609.35867#A2.SS2.p1.1),[§1](https://arxiv.org/html/2609.35867#S1.p1.1),[§1](https://arxiv.org/html/2609.35867#S1.p2.1),[§1](https://arxiv.org/html/2609.35867#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.35867#S2.SS1.p3.1),[§2\.2](https://arxiv.org/html/2609.35867#S2.SS2.p1.1),[§3\.1](https://arxiv.org/html/2609.35867#S3.SS1.p2.1),[§3\.2](https://arxiv.org/html/2609.35867#S3.SS2.p3.1),[§4\.1](https://arxiv.org/html/2609.35867#S4.SS1.p1.1),[§5\.1](https://arxiv.org/html/2609.35867#S5.SS1.SSS0.Px1.p1.1)\.
- Scholaket al\.\(2021\)T\. Scholak, N\. Schucher, and D\. BahdanauPICARD: parsing incrementally for constrained auto\-regressive decoding from language models\.InEMNLP,Cited by:[§4\.3](https://arxiv.org/html/2609.35867#S4.SS3.p1.1)\.
- Smithet al\.\(2003\)E\. Smith, J\. D\. Farmer, L\. Gillemot, and S\. KrishnamurthyStatistical theory of the continuous double auction\.Quantitative Finance3\(6\),pp\. 481–514\.Cited by:[§2\.1](https://arxiv.org/html/2609.35867#S2.SS1.p2.1)\.
- Smithet al\.\(2023\)J\. T\. Smith, A\. Warrington, and S\. W\. LindermanSimplified state space layers for sequence modeling\.InICLR,Cited by:[§1](https://arxiv.org/html/2609.35867#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.35867#S2.SS1.p3.1),[§3\.1](https://arxiv.org/html/2609.35867#S3.SS1.p2.1)\.
- Vaswaniet al\.\(2017\)A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. PolosukhinAttention is all you need\.InNeurIPS,Cited by:[§2\.1](https://arxiv.org/html/2609.35867#S2.SS1.p3.1),[§3\.1](https://arxiv.org/html/2609.35867#S3.SS1.p2.1)\.
- Wheeler and Varner \(2024\)A\. Wheeler and J\. D\. VarnerMarketGPT: developing a pre\-trained transformer \(gpt\) for modeling financial time series\.arXiv preprint arXiv:2411\.16585\.Cited by:[§A\.2](https://arxiv.org/html/2609.35867#A1.SS2),[§1](https://arxiv.org/html/2609.35867#S1.p1.1),[§1](https://arxiv.org/html/2609.35867#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.35867#S2.SS1.p3.1),[§3\.2](https://arxiv.org/html/2609.35867#S3.SS2.p2.1),[§3\.2](https://arxiv.org/html/2609.35867#S3.SS2.p3.1)\.
- Zhanget al\.\(2019\)Z\. Zhang, S\. Zohren, and S\. RobertsDeepLOB: deep convolutional neural networks for limit order books\.IEEE Transactions on Signal Processing\.Cited by:[§2\.1](https://arxiv.org/html/2609.35867#S2.SS1.p1.1)\.
## Appendix APost\-processing in prior methods
Prior LOB message generators do not guarantee that their raw outputs can be replayed on the current LOB\. Hence, their market\-simulation pipelines rely on model\-specific, often heuristic, post\-processing to convert raw outputs into simulator actions\.
The design of this post\-processing depends on whether a generated message is treated as an exchange\-side record of an event that has already occurred or as a trader\-side instruction submitted to the matching engine\. LOBS5 corrects generated messages into exchange\-side records, whereas MarS generates trader\-side instructions\. MarketGPT incorporates aspects of both designs\.
### A\.1LOBS5\([Nagy et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib1)\)
LOBS5 uses the generated\-message field layoutM^t=\(E^t,X^t,R^t\)\\widehat\{M\}\_\{t\}=\(\\widehat\{E\}\_\{t\},\\widehat\{X\}\_\{t\},\\widehat\{R\}\_\{t\}\), with the reference fields appearing last\. Here,E^t\\widehat\{E\}\_\{t\}denotes event type and side,X^t\\widehat\{X\}\_\{t\}denotes the event order, andR^t\\widehat\{R\}\_\{t\}denotes the referenced order\.
#### E^t\[type\]=Add\\widehat\{E\}\_\{t\}\[\\mathrm\{type\}\]=\\mathrm\{Add\}\.
- •Rejection:LOBS5 rejects the message ifX^t\[price\]\\widehat\{X\}\_\{t\}\[\\mathrm\{price\}\]is marketable against the opposite best price\.
- •Correction:Otherwise, LOBS5 replaces any generated non\-N/A reference field inR^t\\widehat\{R\}\_\{t\}with N/A\. No other correction is applied\.
#### E^t\[type\]∈\{Cancel,Delete\}\\widehat\{E\}\_\{t\}\[\\mathrm\{type\}\]\\in\\\{\\mathrm\{Cancel\},\\mathrm\{Delete\}\\\}\.
- •Rejection:LOBS5 matches a resting ordero∈𝒪t−1o\\in\\mathcal\{O\}\_\{t\-1\}toR^t\\widehat\{R\}\_\{t\}under\(price,size,time\)→\(price,size\)→\(price\)\(\\text\{price\},\\text\{size\},\\text\{time\}\)\\rightarrow\(\\text\{price\},\\text\{size\}\)\\rightarrow\(\\text\{price\}\)\. The final price\-only fallback applies only to INIT liquidity, which represents the initial book state at each side–price level as a single order\. The message is rejected if no suchoois found\.
- •Correction:If such an orderoois found, LOBS5 sets Rt=o,Xt\[price\]=Rt\[price\],Xt\[size\]=min\{X^t\[size\],Rt\[size\]\}\.R\_\{t\}=o,\\quad X\_\{t\}\[\\mathrm\{price\}\]=R\_\{t\}\[\\mathrm\{price\}\],\\quad X\_\{t\}\[\\mathrm\{size\}\]=\\min\\\{\\widehat\{X\}\_\{t\}\[\\mathrm\{size\}\],R\_\{t\}\[\\mathrm\{size\}\]\\\}\.It setsEt\[type\]E\_\{t\}\[\\mathrm\{type\}\]to Cancel ifXt\[size\]<Rt\[size\]X\_\{t\}\[\\mathrm\{size\}\]<R\_\{t\}\[\\mathrm\{size\}\], and to Delete otherwise\.
- •Our reproduction:We first attempt an exact match against the reconstructed resting\-order set, then fall back to full and relaxed matching within the message history, following the original code\. We use exact resting\-order initialization and compare it with INIT\-order initialization in[SectionC\.2](https://arxiv.org/html/2609.35867#A3.SS2)\.
#### E^t\[type\]=Execute\\widehat\{E\}\_\{t\}\[\\mathrm\{type\}\]=\\mathrm\{Execute\}\.
- •Rejection:The message is rejected if no executable order is found\.
- •Correction:Otherwise, LOBS5 ignoresR^t\\widehat\{R\}\_\{t\}and setsRt=FrontOfQueue\(𝒪t−1\(E^t\[side\]\)\)R\_\{t\}=\\operatorname\{FrontOfQueue\}\(\\mathcal\{O\}\_\{t\-1\}^\{\(\\widehat\{E\}\_\{t\}\[\\mathrm\{side\}\]\)\}\)\. It then setsXt\[price\]=Rt\[price\]X\_\{t\}\[\\mathrm\{price\}\]=R\_\{t\}\[\\mathrm\{price\}\]andXt\[size\]=min\{X^t\[size\],Rt\[size\]\}X\_\{t\}\[\\mathrm\{size\}\]=\\min\\\{\\widehat\{X\}\_\{t\}\[\\mathrm\{size\}\],R\_\{t\}\[\\mathrm\{size\}\]\\\}\.
### A\.2MarketGPT\([Wheeler and Varner, 2024](https://arxiv.org/html/2609.35867#bib.bib2)\)
MarketGPT uses the generated\-message field layoutM^t=\(E^t,X^t,R^t\)\\widehat\{M\}\_\{t\}=\(\\widehat\{E\}\_\{t\},\\widehat\{X\}\_\{t\},\\widehat\{R\}\_\{t\}\), with an additionalX^t\[remainingsize\]\\widehat\{X\}\_\{t\}\[\\mathrm\{remaining\\ size\}\]field and five event types: Add, Execute, Execute\-at\-different\-price, Cancel/Delete, and Replace\. Execute\-at\-different\-price is unsupported in the released generated rollouts, so we omit it below\. For Cancel/Delete,X^t\[size\]\\widehat\{X\}\_\{t\}\[\\mathrm\{size\}\]andX^t\[remainingsize\]\\widehat\{X\}\_\{t\}\[\\mathrm\{remaining\\ size\}\]denote the canceled and remaining quantities, respectively\.
#### E^t\[type\]=Add\\widehat\{E\}\_\{t\}\[\\mathrm\{type\}\]=\\mathrm\{Add\}\.
- •Simulator handling:The generated message is submitted directly as a limit order without correction or rejection\. If it is marketable, the matching engine executes it against the opposite book in price–time priority, and any unfilled quantity rests at the generated limit price\.
#### E^t\[type\]=Execute\\widehat\{E\}\_\{t\}\[\\mathrm\{type\}\]=\\mathrm\{Execute\}\.
- •Simulator handling:MarketGPT ignoresR^t\\widehat\{R\}\_\{t\}andX^t\[price\]\\widehat\{X\}\_\{t\}\[\\mathrm\{price\}\]and treats the generated side and size as a market order\. It repeatedly matches opposite\-side resting orders in price–time priority until the generated size is filled or the opposite book is exhausted, discarding any unfilled quantity\.
#### E^t\[type\]∈\{Cancel,Delete\}\\widehat\{E\}\_\{t\}\[\\mathrm\{type\}\]\\in\\\{\\mathrm\{Cancel\},\\mathrm\{Delete\}\\\}\.
- •Rejection:MarketGPT matches a resting ordero∈𝒪t−1o\\in\\mathcal\{O\}\_\{t\-1\}to\(X^t\[price\],X^t\[size\]\+X^t\[remainingsize\],R^t\[time\]\)\(\\widehat\{X\}\_\{t\}\[\\mathrm\{price\}\],\\widehat\{X\}\_\{t\}\[\\mathrm\{size\}\]\+\\widehat\{X\}\_\{t\}\[\\mathrm\{remaining\\ size\}\],\\widehat\{R\}\_\{t\}\[\\mathrm\{time\}\]\)under\(price,size,time\)→\(price,size\)→\(price,time\)→\(price\)\(\\text\{price\},\\text\{size\},\\text\{time\}\)\\rightarrow\(\\text\{price\},\\text\{size\}\)\\rightarrow\(\\text\{price\},\\text\{time\}\)\\rightarrow\(\\text\{price\}\)\. The final price\-only fallback selects the front\-of\-queue resting order atX^t\[price\]\\widehat\{X\}\_\{t\}\[\\mathrm\{price\}\]\. The message is rejected if no suchoois found\.
- •Correction:If such an orderoois found, letρ^t=X^t\[size\]/\(X^t\[size\]\+X^t\[remainingsize\]\)\\widehat\{\\rho\}\_\{t\}=\\widehat\{X\}\_\{t\}\[\\mathrm\{size\}\]/\(\\widehat\{X\}\_\{t\}\[\\mathrm\{size\}\]\+\\widehat\{X\}\_\{t\}\[\\mathrm\{remaining\\ size\}\]\)denote the generated cancellation ratio\. MarketGPT sets Rt=o,Xt\[size\]=⌊Rt\[size\]ρ^t⌋,Xt\[remainingsize\]=Rt\[size\]−Xt\[size\],R\_\{t\}=o,\\quad X\_\{t\}\[\\mathrm\{size\}\]=\\left\\lfloor R\_\{t\}\[\\mathrm\{size\}\]\\widehat\{\\rho\}\_\{t\}\\right\\rfloor,\\quad X\_\{t\}\[\\mathrm\{remaining\\ size\}\]=R\_\{t\}\[\\mathrm\{size\}\]\-X\_\{t\}\[\\mathrm\{size\}\],preserving the generated cancellation ratio for the matched resting order up to integer rounding\.
#### E^t\[type\]=Replace\\widehat\{E\}\_\{t\}\[\\mathrm\{type\}\]=\\mathrm\{Replace\}\.
- •Rejection:MarketGPT matches a resting ordero∈𝒪t−1o\\in\\mathcal\{O\}\_\{t\-1\}toR^t\\widehat\{R\}\_\{t\}under\(price,size,time\)→\(price,size\)→\(price,time\)→\(price\)\(\\text\{price\},\\text\{size\},\\text\{time\}\)\\rightarrow\(\\text\{price\},\\text\{size\}\)\\rightarrow\(\\text\{price\},\\text\{time\}\)\\rightarrow\(\\text\{price\}\)\. The final price\-only fallback selects the front\-of\-queue resting order atR^t\[price\]\\widehat\{R\}\_\{t\}\[\\mathrm\{price\}\]\. The message is rejected if no suchoois found\.
- •Correction:If such an orderoois found, MarketGPT setsRt=oR\_\{t\}=oand retainsXt=X^tX\_\{t\}=\\widehat\{X\}\_\{t\}\.
### A\.3MarS\([Li et al\., 2025](https://arxiv.org/html/2609.35867#bib.bib4)\)
In our notation, MarS represents each generated output asM^t=\(E^t,X^t\)\\widehat\{M\}\_\{t\}=\(\\widehat\{E\}\_\{t\},\\widehat\{X\}\_\{t\}\)without reference\-order fields, whereE^t∈\{Bid,Ask,Cancel\}\\widehat\{E\}\_\{t\}\\in\\\{\\mathrm\{Bid\},\\mathrm\{Ask\},\\mathrm\{Cancel\}\\\}\. It interprets each output as a trader\-side instruction to the simulated clearing house rather than as a complete exchange message\.
#### E^t∈\{Bid,Ask\}\\widehat\{E\}\_\{t\}\\in\\\{\\mathrm\{Bid\},\\mathrm\{Ask\}\\\}\.
- •Simulator handling:MarS submits the generated Bid or Ask directly as a limit order without correction or rejection\. If it is marketable, the simulated clearing house executes it against opposite\-side resting orders, and any unfilled quantity rests at the generated price\.
#### E^t=Cancel\\widehat\{E\}\_\{t\}=\\mathrm\{Cancel\}\.
- •Rejection:If no resting order is available for cancellation, MarS discards the generated cancellation by returning an empty simulator action\. The discarded output is not added to the model’s order history, and the agent immediately samples again without rolling back the elapsed simulation time\.
- •Correction:MarS projects the generated cancellation onto the current resting\-order set𝒪t−1\\mathcal\{O\}\_\{t\-1\}by selecting an available resting ordero∈𝒪t−1o\\in\\mathcal\{O\}\_\{t\-1\}whose priceo\[price\]o\[\\mathrm\{price\}\]is closest toX^t\[price\]\\widehat\{X\}\_\{t\}\[\\mathrm\{price\}\]\. It assigns actual order IDs and sides and distributesX^t\[size\]\\widehat\{X\}\_\{t\}\[\\mathrm\{size\}\]across the resting orders ato\[price\]o\[\\mathrm\{price\}\], capped by their available size\.
## Appendix BExperimental details
### B\.1Preprocessing
We convert XNAS\.ITCH MBO data for GOOG and INTC from Databento222[https://databento\.com/portal/catalog/us\-equities\#XNAS\.ITCH](https://databento.com/portal/catalog/us-equities#XNAS.ITCH)into LOBSTER\-compatible messages\([Huang and Polak, 2011](https://arxiv.org/html/2609.35867#bib.bib19)\)and follow the preprocessing pipeline of LOBS5\([Nagy et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib1)\), except for the reference\-order representation\. See below for details\.
#### Databento MBO to LOBSTER\-compatible message\.
Within each stream identified by\(publisher\_id, instrument\_id, channel\_id\), consecutive records in feed order with the samesequencevalue form an event group, which is converted according to[Table6](https://arxiv.org/html/2609.35867#A2.T6)\. Under these conversion rules, individual fields are mapped as shown in[Table7](https://arxiv.org/html/2609.35867#A2.T7), with format conversions applied as needed\.
Table 6:Conversion of Databento MBO action patterns to LOBSTER\-compatible messages\.Reset clears the order state once at session start without emitting a message\. Sole trades \(single\-Tgroups withside=N\) may correspond to Type 5 but are omitted because their side is unknown\.Table 7:Field correspondence between Databento MBO and LOBSTER\-compatible messages\.
#### Message representation\.
[Table8](https://arxiv.org/html/2609.35867#A2.T8)summarizes normalization and encoding in ReLOBGen\. All fields share a single vocabulary of 12,011 tokens, including three special tokens for masking, hiding, and not\-applicable values\. The event order’s absolute timestamp is determined by adding its inter\-arrival time to the previous event’s timestamp, rather than generated by the model\.
LOBS5\([Nagy et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib1)\)copies the encoded field values from the referenced order’s original Add message into the reference\-order fields, retaining the price encoding relative to the mid\-price at submission rather than the current mid\-price\.
Table 8:Normalization and encoding in ReLOBGen\.Prices are encoded as tick offsets from the tick\-aligned mid\-price, clipped to\[−999,999\]\[\-999,999\]:Δp=clip\(\(porder−pmid\)/τ,−999,999\)\\Delta p=\\operatorname\{clip\}\(\(p\_\{\\mathrm\{order\}\}\-p\_\{\\mathrm\{mid\}\}\)/\\tau,\-999,999\)\. Here,pmid=τ⌊\(pask,1\+pbid,1\)/\(2τ\)⌋p\_\{\\mathrm\{mid\}\}=\\tau\\lfloor\(p\_\{\\mathrm\{ask\},1\}\+p\_\{\\mathrm\{bid\},1\}\)/\(2\\tau\)\\rfloor,pask,1p\_\{\\mathrm\{ask\},1\}andpbid,1p\_\{\\mathrm\{bid\},1\}denote the best ask and bid prices, respectively, andτ=100\\tau=100\. Each time component is encoded in base 1000, with one token per three\-digit group\.
#### Book representation\.
We reconstruct book states by replaying the message sequence\. Then, following LOBS5\([Nagy et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib1)\), we represent each book state as a 501\-dimensional continuous vector comprising the mid\-price change in ticks and signed volumes on a 500\-slot price grid centered on the tick\-aligned mid\-price:
𝐛t\\displaystyle\\mathbf\{b\}\_\{t\}=\(Δpmid,t/τ,xt,0,…,xt,499\)∈ℝ501,\\displaystyle=\\left\(\\Delta p\_\{\\mathrm\{mid\},t\}/\\tau,x\_\{t,0\},\\ldots,x\_\{t,499\}\\right\)\\in\\mathbb\{R\}^\{501\},\(6\)whereΔpmid,t=pmid,t−pmid,t−1\\Delta p\_\{\\mathrm\{mid\},t\}=p\_\{\\mathrm\{mid\},t\}\-p\_\{\\mathrm\{mid\},t\-1\}andτ=100\\tau=100\. We map the top 10 ask and bid levels onto this grid by their tick offsets from the mid\-price\. Each slot stores the total quantity divided by 1000, negative for asks and positive for bids\. Levels outside the grid are omitted, and empty slots are zero\.
#### Filtering\.
For training, we retain only Type 1 \(Add\), Type 2 \(Cancel\), Type 3 \(Delete\), and Type 4 \(Execute\) messages whose prices lie between the 10th\-best bid and 10th\-best ask prices at the time of the event\. We use only 500\-message sequences fully contained within regular trading hours \(09:30–16:00 ET\)\.
### B\.2Training configuration
We follow the scaled\-up LOBS5 architecture used in LOB\-Bench\([Nagy et al\., 2023](https://arxiv.org/html/2609.35867#bib.bib1);[Nagy et al\., 2025](https://arxiv.org/html/2609.35867#bib.bib5)\)and likewise use one year of training data for both stocks, covering Jan\. 1–Dec\. 31, 2025, with Jan\. 1–15, 2026 for validation\. We train the message generation modelπθ\\pi\_\{\\theta\}first, followed by the query and order projection headsfqueryf\_\{\\mathrm\{query\}\}andforderf\_\{\\mathrm\{order\}\}on the same data periods\.[Table9](https://arxiv.org/html/2609.35867#A2.T9)summarizes the training settings, and[Table10](https://arxiv.org/html/2609.35867#A2.T10)details the architecture of each component\. Training budgets count message windows forπθ\\pi\_\{\\theta\}and window–mid\-price pairs for the projection heads, with 256 mid\-prices per window\.
Table 9:Training configuration\.In both training stages, learning rates use linear warmup over the first 10% of training steps, followed by cosine decay to 10% of their peak values\.Table 10:Model architecture\.πθ\\pi\_\{\\theta\}is the message generation model;fqueryf\_\{\\mathrm\{query\}\}andforderf\_\{\\mathrm\{order\}\}are the query and order projection heads, respectively\.
### B\.3Rollout protocol
For each stock, we use 100 rollout windows from each of 10 test days, January 16–30, 2026\. Each rollout starts from the exact resting\-order set within the top 20 price levels on each side, providing a buffer of deeper resting orders that may enter the top 10 price levels during the rollout\. Consistent with training, model LOB\-state inputs and eligible reference sets are restricted to the current top 10 price levels on each side\. We do not apply temperature scaling or top\-kk/top\-ppfiltering during either resting\-order selection or token sampling\. To avoid spending excessive computation on stalled rollouts, we restart a rollout with a different random seed after 100 consecutive failed attempts\.
## Appendix CAdditional experimental results
### C\.1Event\-wise replayability
[Table11](https://arxiv.org/html/2609.35867#A3.T11)reports post\-processing and replayability violation rates by event type on GOOG and INTC\. Both LOBS5 baselines rarely require post\-processing for add events, but frequently reject cancel and delete events and correct all execute events\. Across non\-add event types, reference violations are more frequent than event\-order violations on both stocks\. ReLOBGen records no violations and requires no post\-processing for any event type\.
Table 11:Event\-wise post\-processing rates and replayability violations\.Positive rates below0\.1%0\.1\\%are reported as <0\.1%\. Lower is better, and bold indicates the lowest value\.
### C\.2Effect of resting\-order initialization on replayability
We compare two ways of initializing the resting\-order set at the start of a rollout\. INIT\-order initialization, as used in the original LOBS5 implementation, represents the total volume at each side–price level of the initial LOB state as a single INIT order, whereas exact resting\-order initialization preserves the individual orders reconstructed from market\-by\-order data\. INIT orders do not correspond to individual orders in the model’s training messages and can only be matched to generated references through price\-only matching during post\-processing\. Our evaluation uses exact resting\-order initialization because the exact resting\-order set is available from our data\.
[Figure4](https://arxiv.org/html/2609.35867#A3.F4)shows that LOBS5 \(ref\.\-last\) still requires correction and rejection under INIT\-order initialization on GOOG\. With INIT orders, rejection rates start lower but approach those under exact initialization by the end of the rollout\. Together with the higher correction rates, this pattern suggests that the initial reduction in rejection reflects the looser matching available for INIT orders, rather than improved replayability of the raw generated messages\.
Figure 4:Effect of resting\-order initialization on replayability\.We report rejection and correction rates for LOBS5 \(ref\.\-last\) on GOOG under INIT\-order and exact resting\-order initialization\. Rates are computed over generation attempts grouped into 20\-message bins by accepted\-message position\.
### C\.3Additional LOB\-Bench results
#### Unconditional distributions\.
ReLOBGen achieves broadly lower unconditionalL1L\_\{1\}and Wasserstein distances than LOBS5 across metric groups \([Tables3](https://arxiv.org/html/2609.35867#S5.T3)and[12](https://arxiv.org/html/2609.35867#A3.T12)\)\. At the individual\-metric level, ReLOBGen achieves lower Wasserstein point estimates on 16 of 21 metrics for GOOG, with statistically significant improvements on 14 of those 16\. On INTC, it achieves lower point estimates on 20 of 21 metrics, with statistically significant improvements on 18 of those 20, using the same CI non\-overlap criterion as in[Section5\.3](https://arxiv.org/html/2609.35867#S5.SS3)\.[Figures5](https://arxiv.org/html/2609.35867#A3.F5)and[6](https://arxiv.org/html/2609.35867#A3.F6)show the individual metric distances and their 99% bootstrap confidence intervals forL1L\_\{1\}and Wasserstein distances, respectively\.
Table 12:LOB\-Bench evaluation: metric\-group summary of unconditional Wasserstein distances\.Overall averages all metrics, while each metric\-group column averages the metrics in its group\. Lower is better, and bold indicates the lowest value\.Figure 5:LOB\-Bench evaluation: unconditionalL1L\_\{1\}distances\.Log time\-to\-cancel\* measures time\-to\-cancel for all orders canceled during the rollout using exact resting\-order information, including cancellations of orders present in the initial LOB that were excluded from the original LOB\-Bench implementation\. Error bars indicate 99% percentile bootstrap confidence intervals\. Lower is better\.Figure 6:LOB\-Bench evaluation: unconditional Wasserstein distances\.Log time\-to\-cancel\* measures time\-to\-cancel for all orders canceled during the rollout using exact resting\-order information, including cancellations of orders present in the initial LOB that were excluded from the original LOB\-Bench implementation\. Error bars indicate 99% percentile bootstrap confidence intervals\. Lower is better\.
#### Conditional distributions\.
ReLOBGen achieves lower conditional Wasserstein point estimates than LOBS5 on 5 of 6 metrics across the two stocks \([Table13](https://arxiv.org/html/2609.35867#A3.T13)\)\.
Table 13:LOB\-Bench evaluation: conditional Wasserstein distances\.Conditional 99% percentile bootstrap confidence intervals have half\-widths below 0\.012\. Lower is better, and bold indicates the lowest value\.
#### Market\-impact responses\.
[Figure7](https://arxiv.org/html/2609.35867#A3.F7)shows price\-response curves following market orders, limit orders, and cancellations on GOOG and INTC\. LOBS5 produces largely flat response curves across all six event types on GOOG and across the three event types without an immediate mid\-price change on INTC\. ReLOBGen more closely matches the real responses, with lower market\-impact discrepancies than LOBS5 on 9 of 12 metrics across the two stocks \([Table4](https://arxiv.org/html/2609.35867#S5.T4)\)\.
Figure 7:LOB\-Bench evaluation: market\-impact response curves\.Responses are plotted against event lag for market orders \(MO\), limit orders \(LO\), and cancellations \(CA\), with subscripts 0 and 1 for events without and with an immediate mid\-price change\. Shaded regions indicate 99% bootstrap confidence intervals\.
#### Error accumulation\.
[Figures8](https://arxiv.org/html/2609.35867#A3.F8)and[9](https://arxiv.org/html/2609.35867#A3.F9)showL1L\_\{1\}and Wasserstein divergences between generated and real distributions on GOOG and INTC at 100\-message intervals\. ReLOBGen generally exhibits lower divergence values and flatter slopes than LOBS5 over the rollout horizon\.
GOOGINTC
Figure 8:LOB\-Bench evaluation: error accumulation inL1L\_\{1\}distances\.Log time\-to\-cancel\* measures time\-to\-cancel for all orders canceled during the rollout using exact resting\-order information, including cancellations of orders present in the initial LOB that were excluded from the original LOB\-Bench implementation\. Lower is better\.GOOGINTC
Figure 9:LOB\-Bench evaluation: error accumulation in Wasserstein distances\.Log time\-to\-cancel\* measures time\-to\-cancel for all orders canceled during the rollout using exact resting\-order information, including cancellations of orders present in the initial LOB that were excluded from the original LOB\-Bench implementation\. Lower is better\.
### C\.4Marginal event\-type distributions
[Table14](https://arxiv.org/html/2609.35867#A3.T14)compares marginal event\-type distributions in actual and generated messages on GOOG and INTC\. ReLOBGen achieves the lowest total variation distance among the evaluated methods on both stocks, indicating a closer match to the actual event\-type distribution\.
Table 14:Marginal event\-type frequencies in actual and generated messages\.TV denotes the total variation distance from the actual distribution, reported in percentage points\. Lower TV is better, and bold indicates the lowest value\.
## Appendix DResting\-order coverage by history length
[Figure10](https://arxiv.org/html/2609.35867#A4.F10)shows resting\-order coverage as a function of model\-history length for GOOG and INTC\. Let𝒪t−1\(10\)\\mathcal\{O\}\_\{t\-1\}^\{\(10\)\}denote the subset of𝒪t−1\\mathcal\{O\}\_\{t\-1\}within the current top 10 price levels on each side\. We define coverage as
Ct\(L\)=\|\{o∈𝒪t−1\(10\):oappears inMt−L:t−1\}\|\|𝒪t−1\(10\)\|,\\displaystyle C\_\{t\}\(L\)=\\frac\{\\left\|\\left\\\{o\\in\\mathcal\{O\}\_\{t\-1\}^\{\(10\)\}:o\\text\{ appears in \}M\_\{t\-L:t\-1\}\\right\\\}\\right\|\}\{\|\\mathcal\{O\}\_\{t\-1\}^\{\(10\)\}\|\},\(7\)and report the mean ofCt\(L\)C\_\{t\}\(L\)across windows\. An order appears in the history if its raw order ID occurs in the precedingLLfiltered messages\.
At the 500\-message context length used by our models, coverage is only 60\.4% for GOOG and 28\.6% for INTC\. Coverage increases to 86\.2% and 65\.6%, respectively, with 5,000 preceding messages, with diminishing gains as history length grows\. Even at this length, a substantial fraction of the current resting orders remains absent from the message history\.
Figure 10:Mean resting\-order coverage by history length on GOOG and INTC\.Coverage is evaluated during regular trading hours on the test split\.Similar Articles
Do LLMs Understand Limit Order Book Dynamics?
This paper investigates whether large language models trained on synthetic limit order book data develop an accurate world model, finding that while they generate valid sequences, they have systematic errors leading to biased and spurious forecasts.
ReLoRA: Knowledge-Reusing Adaptation for Fast Rollout of Evolving LLM Services
ReLoRA is a knowledge-reusing adaptation framework that efficiently restores service-ready LoRA adapters for evolving LLM services, reducing time-to-readiness by up to 8.9× and improving accuracy by up to 4.6% through adaptive initialization and scheduled regularization.
Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments
This arXiv paper presents a unified LLMOps architecture for real-time, enterprise-ready LLM deployments, integrating data ingestion, continual learning, RAG, and feedback loops. It introduces components like AIPO, STAR+FAR, and SAGE to address knowledge staleness, hallucination, and latency-cost trade-offs in regulated sectors.
Code2LoRA: Hypernetwork-Generated Adapters for Code Language Models under Software Evolution
Code2LoRA introduces a hypernetwork that generates LoRA adapters from a repository in a single forward pass, allowing frozen code LLMs to adapt to repository context without extra tokens, and supporting evolving codebases efficiently. It also delivers RepoPeftBench, a benchmark for repo-conditioned code modeling.
I built LOLM: a lower-cost LLM agent with live control decisions and sealed run receipts
The builder of LOLM announces a hybrid Transformer-SSM language model and agent system from Qira, featuring a controller for live decisions, run receipts, a CLI, coding sandbox, MCP support, and lower-cost hosted access.