MINT: A Universal Zero-Shot Predictor for Transaction Data
Summary
MINT is a framework that connects pretrained transaction sequence encoders to decoder-only LLMs for zero-shot predictive tasks on financial transaction data, achieving state-of-the-art performance with reduced resources.
View Cached Full Text
Cached at: 08/17/26, 10:22 AM
# A Universal Zero-Shot Predictor for Transaction Data Source: [https://arxiv.org/html/2608.14198](https://arxiv.org/html/2608.14198) ## MINT: A Universal Zero\-Shot Predictor for Transaction DataDOI:[XXXXXXX\.XXXXXXX](https://doi.org/XXXXXXX.XXXXXXX)Conference:Make sure to enter the correct conference title from your rights confirmation email; June 03–05, 2018; Woodstock, NYISBN:978\-1\-4503\-XXXX\-X/2018/06123\-A56\-BU3CCS:Do Not Use This Code Generate the Correct Terms for Your PaperCCS:Do Not Use This Code Generate the Correct Terms for Your PaperCCS:Do Not Use This Code Generate the Correct Terms for Your PaperCCS:Do Not Use This Code Generate the Correct Terms for Your Paper Parameswaran KamalarubanNote:Both authors contributed equally to this research\.Affiliation:Risk and Security AI Lab, Visa Inc\.,United Kingdomemail:[kaparame@visa\.com](mailto:[email protected])Viktor DrobnyiAffiliation:Risk and Security AI Lab, Visa Inc\.,United Kingdomemail:[vdrobnyi@visa\.com](mailto:[email protected]),Maeve MadiganAffiliation:Risk and Security AI Lab, Visa Inc\.,United Kingdomemail:[mmadigan@visa\.com](mailto:[email protected]),Julia RozanovaAffiliation:Risk and Security AI Lab, Visa Inc\.,United Kingdomemail:[yrozanov@visa\.com](mailto:[email protected]),David SuttonAffiliation:Risk and Security AI Lab, Visa Inc\.,United Kingdomemail:[dsutton@visa\.com](mailto:[email protected])andStuart BurrellAffiliation:Risk and Security AI Lab, Visa Inc\.,United Kingdomemail:[sburrell@visa\.com](mailto:[email protected]) 2026© , 2026; ###### Abstract\. Banks analyse sequential financial transaction data to perform many tasks, including fraud prevention, credit risk assessment and offer personalization\. To improve the predictive accuracy of these tasks, Payments Foundation Models encode transaction sequence data as rich contextual embeddings, which can then be provided to task\-specific models as features\. However, these Foundation Models are not designed for flexible zero\-shot reasoning across novel downstream prediction tasks, limiting their adaptability and utility\. Existing LLM\-based approaches to zero\-shot prediction often fail to fully exploit the predictive signal within transaction data, while relying on costly text serialization or task\-specific architectures that scale poorly\. To address these limitations, we present the Multimodal Instruction Network for Transactions \(MINT\), a framework that connects a pretrained transaction sequence encoder to a decoder\-only LLM through lightweight embedding injection, transaction\-language alignment, and instruction tuning\. We find thatMINTachieves state\-of\-the\-art predictive question\-answering performance in both in\-distribution and out\-of\-distribution questions, while substantially reducing input tokens, latency, and memory consumption compared to text\-serialization baselines\. Through comprehensive analyses of representations, alignment strategies, training data, and history length, we establish that compact transaction embeddings are a superior approach to transaction representation than text serialization for multimodal reasoning and zero\-shot prediction tasks\. ## 1\.Introduction Figure 1\.Overview of MINT\. Transaction histories are encoded into compact transaction embeddings, enabling predictive question answering and financial reasoning through a natural\-language interface\. Example applications include forecasting future customer behavior, such as transaction volume and average transaction amount over the next 30 days, alongside summarization, fraud\-risk assessment, and financial analysis\.Financial behavior is naturally represented as a sequence of discrete events\. Each customer generates a temporally ordered history of transactions, where every transaction consists of heterogeneous categorical and numerical attributes \(e\.g\., merchant, amount, channel, and timestamp\)\. Reasoning over such data requires modeling both relationships among fields within a transaction and temporal dependencies across transactions, often over long horizons\. These challenges have made transaction understanding a central problem in domains such as customer analytics, risk assessment, financial planning, and fraud detection\. Large language models \(LLMs\) provide a compelling interface for transaction understanding because they support natural\-language interaction, instruction following, and explanation generation\. A straightforward approach is to serialize transaction histories into text and append them to prompts\. While effective for some retrieval\-style tasks, this strategy quickly becomes inefficient as histories grow longer and often struggles to exploit the structured numerical and temporal signals present in transaction data\([11](https://arxiv.org/html/2608.14198#bib.bib2);[14](https://arxiv.org/html/2608.14198#bib.bib10);[28](https://arxiv.org/html/2608.14198#bib.bib11);[31](https://arxiv.org/html/2608.14198#bib.bib12)\)\. Recent work has therefore explored multimodal transaction reasoning systems that combine transaction sequence encoders with LLMs, allowing models to condition on compact transaction representations rather than long textual serializations\([7](https://arxiv.org/html/2608.14198#bib.bib3);[25](https://arxiv.org/html/2608.14198#bib.bib9);[15](https://arxiv.org/html/2608.14198#bib.bib1);[27](https://arxiv.org/html/2608.14198#bib.bib8);[1](https://arxiv.org/html/2608.14198#bib.bib5);[22](https://arxiv.org/html/2608.14198#bib.bib7)\)\. At the same time, self\-supervised pretraining on large\-scale transaction corpora has produced increasingly powerful transaction foundation models\([21](https://arxiv.org/html/2608.14198#bib.bib16);[23](https://arxiv.org/html/2608.14198#bib.bib15);[20](https://arxiv.org/html/2608.14198#bib.bib17)\)\. These encoders learn transferable representations from behavioral sequences and achieve strong performance across a variety of downstream financial tasks\. Together, these developments suggest a natural modular design: use a pretrained transaction sequence encoder to extract behavioral signals and reserve LLM capacity for language understanding, reasoning, and generation\. Following this principle, we present the Multimodal Instruction Network for Transactions \(MINT\), a multimodal transaction reasoning framework that couples a pretrained transaction sequence encoder with a decoder\-only LLM through embedding injection\. Transaction histories are first encoded into dense behavioral representations, which are projected into the LLM embedding space using a lightweight connector\. The transaction sequence encoder remains frozen throughout multimodal training, while LoRA adapters efficiently adapt the LLM for downstream reasoning\. Inspired by modern vision\-language model training pipelines,MINTcombines transaction\-language alignment on caption data with instruction tuning on question\-answering and reasoning tasks, enabling efficient conditioning on long transaction histories without costly textual serialization\. Figure 2\.Accuracy\-efficiency trade\-off on predictive QA\. The x\-axis shows input token count and the y\-axis shows accuracy; filled and unfilled markers correspond to ID and OOD evaluation, respectively\. MINT occupies a more favorable Pareto region than competing methods, with MINT \(h=1h=1\) providing the strongest predictive QA performance while balancing the performance and inference efficiency\. MINT also shows the least degradation from ID to OOD among top performing methods\.Our main contributions are: - •IntroduceMINT, a multimodal transaction reasoning framework that integrates a pretrained transaction sequence encoder with a decoder\-only LLM through embedding injection, modality alignment, and parameter\-efficient adaptation \(see Figure[1](https://arxiv.org/html/2608.14198#S1.F1)\)\. - •Validatethat transaction embeddings provide a stronger representation than textual serialization for predictive reasoning, achieving state\-of\-the\-art predictive QA performance in both in\-distribution \(ID\) and out\-of\-distribution \(OOD\) settings while reducing input tokens, latency, and GPU memory usage \(see Figure[2](https://arxiv.org/html/2608.14198#S1.F2)\)\. - •Presenta comprehensive study of multimodal transaction reasoning, analyzing transaction representations, model capacity, training\-data mixtures, history length, robustness, and efficiency, yielding practical insights for the design of future transaction\-language systems\. ## 2\.Related Work Event sequence foundation models\.Transaction histories can be represented as structured event sequences with heterogeneous attributes and timestamps\. Prior work on tabular and event\-sequence modeling has shown the importance of hierarchical architectures that separately encode field\-level information within events and temporal dependencies across events\([21](https://arxiv.org/html/2608.14198#bib.bib16);[29](https://arxiv.org/html/2608.14198#bib.bib29);[4](https://arxiv.org/html/2608.14198#bib.bib30)\)\. More recently, large\-scale self\-supervised pretraining on transaction corpora has established pretrained transaction sequence encoders as strong foundation models for behavioral data\([23](https://arxiv.org/html/2608.14198#bib.bib15);[16](https://arxiv.org/html/2608.14198#bib.bib14);[6](https://arxiv.org/html/2608.14198#bib.bib27);[3](https://arxiv.org/html/2608.14198#bib.bib26);[9](https://arxiv.org/html/2608.14198#bib.bib24);[26](https://arxiv.org/html/2608.14198#bib.bib23);[20](https://arxiv.org/html/2608.14198#bib.bib17)\)\. In particular,[23](https://arxiv.org/html/2608.14198#bib.bib15)showed that generative autoregressive pretraining yields highly effective transaction representations, outperforming both traditional feature\-engineering approaches and alternative self\-supervised objectives for transaction modeling\([13](https://arxiv.org/html/2608.14198#bib.bib20);[8](https://arxiv.org/html/2608.14198#bib.bib21);[5](https://arxiv.org/html/2608.14198#bib.bib22)\)\. Their results demonstrated that autoregressively pretrained transaction encoders transfer well across a diverse range of downstream tasks\. This motivates our use of a pretrained transaction sequence encoder and makes the corresponding task\-specific classifier a strong baseline throughout our evaluation\. Multimodal event sequence reasoning\.Recent work has explored adapting LLMs to time\-series and event\-sequence reasoning, highlighting challenges in long\-context processing, numerical reasoning, and faithful generation\([11](https://arxiv.org/html/2608.14198#bib.bib2);[14](https://arxiv.org/html/2608.14198#bib.bib10);[28](https://arxiv.org/html/2608.14198#bib.bib11);[31](https://arxiv.org/html/2608.14198#bib.bib12)\)\. To address these limitations, multimodal approaches combine specialized sequence encoders with LLMs through projection layers, cross\-attention modules, or other fusion mechanisms\([7](https://arxiv.org/html/2608.14198#bib.bib3);[25](https://arxiv.org/html/2608.14198#bib.bib9);[15](https://arxiv.org/html/2608.14198#bib.bib1);[27](https://arxiv.org/html/2608.14198#bib.bib8)\)\. Transaction\-focused reasoning systems extend this paradigm to financial event sequences\([1](https://arxiv.org/html/2608.14198#bib.bib5);[22](https://arxiv.org/html/2608.14198#bib.bib7)\)\. However, existing approaches typically omit one or more ingredients that have become standard in modern multimodal reasoning systems, including explicit modality\-alignment training, chain\-of\-thought supervision, and comprehensive evaluation across both in\-distribution and out\-of\-distribution settings\. A key distinction of our approach is the decoupling of transaction representation learning from reasoning\-model training\. Prior work either jointly optimizes the transaction sequence encoder, connector, and LLM\([1](https://arxiv.org/html/2608.14198#bib.bib5)\)or employs more complex fusion architectures\([22](https://arxiv.org/html/2608.14198#bib.bib7)\)\. In contrast, following contemporary vision\-language model training practices\([18](https://arxiv.org/html/2608.14198#bib.bib13);[19](https://arxiv.org/html/2608.14198#bib.bib19)\), we first pretrain a transaction sequence encoder on large\-scale transaction data, then freeze it during modality alignment and instruction tuning\. This separation substantially simplifies training, allows transaction representations to be learned efficiently without repeatedly processing massive datasets through the LLM, and enables the use of a lightweight MLP projector rather than more complex architectures such as Q\-Former\([17](https://arxiv.org/html/2608.14198#bib.bib18)\)\. Furthermore, unlike approaches that rely on task identifiers, task embeddings, or task\-specific control tokens\([1](https://arxiv.org/html/2608.14198#bib.bib5);[22](https://arxiv.org/html/2608.14198#bib.bib7)\), our formulation treats transaction understanding as a unified instruction\-following problem\. As a result, the model is not tied to a fixed set of task definitions and can naturally generalize to previously unseen question types and instructions\. Beyond predictive accuracy, we also provide a comprehensive evaluation spanning robustness, efficiency, latency, memory consumption, and detailed ablations of the major design choices\. Alignment of event\-sequence and text representations\.A complementary line of work aligns event\-sequence representations with semantic text embeddings using contrastive objectives\([10](https://arxiv.org/html/2608.14198#bib.bib6);[30](https://arxiv.org/html/2608.14198#bib.bib4)\)\. These methods primarily aim to learn transferable representations for downstream predictive tasks\. In contrast, we use transaction\-language alignment as an intermediate stage in a multimodal reasoning training pipeline, ultimately targeting generative tasks such as transaction captioning and question answering\. ## 3\.Model Architecture Figure 3\.MINTarchitecture: frozen transaction sequence encoder produces per\-transaction embeddings, a trainable connector projects them into the LLM hidden space, and a decoder\-only LLM adapted via LoRA generates text outputs conditioned on injected embeddings\.Our model comprises: \(i\) afrozen transaction sequence encoderproducing per\-transaction sequence embeddings, \(ii\) atrainable connectormapping transaction embeddings into the LLM hidden space, and \(iii\) adecoder\-only LLMadapted via parameter\-efficient fine\-tuning \(LoRA\)\. Frozen Transaction Sequence Encoder\.Let𝐓=\{\(xi,ti\)\}i=1N\\mathbf\{T\}=\\\{\(x\_\{i\},t\_\{i\}\)\\\}\_\{i=1\}^\{N\}denote a customer’s transaction history, wherexix\_\{i\}is a transaction record andtit\_\{i\}its timestamp\. Following prior work on tabular time\-series modeling\([21](https://arxiv.org/html/2608.14198#bib.bib16);[23](https://arxiv.org/html/2608.14198#bib.bib15)\), we use a Transformer\-based transaction sequence encoder\. Numerical attributes are log\-transformed, categorical attributes are embedded via field\-specific lookup tables, and a field\-level Transformer models interactions within each transaction\. A sequence\-level Transformer then captures temporal dependencies across transactions, producing contextualized embeddingsei=E\(xi∣\{\(xj,tj\)\}j=1i\)∈ℝdee\_\{i\}=E\(x\_\{i\}\\mid\\\{\(x\_\{j\},t\_\{j\}\)\\\}\_\{j=1\}^\{i\}\)\\in\\mathbb\{R\}^\{d\_\{e\}\}\. The encoder is pretrained with an autoregressive next\-event prediction objective, ℒAR=−∑i=1Nlogp\(xi∣x<i\),\\mathcal\{L\}\_\{\\text\{AR\}\}=\-\\sum\_\{i=1\}^\{N\}\\log p\(x\_\{i\}\\mid x\_\{<i\}\),to learn customer behavioral patterns, wherexi=D\(ei−1\)x\_\{i\}=D\(e\_\{i\-1\}\)for a learnable decoderDD\. During multimodal training, the encoder remains frozen and serves solely as a transaction feature extractor\. Connector \(Modality Projector\)\.The connectorgϕg\_\{\\phi\}\(parameterized byϕ\\phi\) maps the transaction embeddingei∈ℝdee\_\{i\}\\in\\mathbb\{R\}^\{d\_\{e\}\}into the LLM hidden spaceℝd\\mathbb\{R\}^\{d\}:zi=gϕ\(ei\)∈ℝdz\_\{i\}=g\_\{\\phi\}\(e\_\{i\}\)\\in\\mathbb\{R\}^\{d\}\. We use a multi\-layer perceptron \(MLP\) with normalization: gϕ\(e\)=LN\(W2σ\(W1LN\(e\)\)\),g\_\{\\phi\}\(e\)~=~\\mathrm\{LN\}\\\!\\left\(W\_\{2\}\\,\\sigma\\\!\\left\(W\_\{1\}\\,\\mathrm\{LN\}\(e\)\\right\)\\right\),whereσ\\sigmais GELU, LN denotes layer norm,W1∈ℝh×deW\_\{1\}\\in\\mathbb\{R\}^\{h\\times d\_\{e\}\},W2∈ℝd×hW\_\{2\}\\in\\mathbb\{R\}^\{d\\times h\}, andhhis the connector hidden size\. We study the effect of connector capacity in the ablations presented in Section[6\.4](https://arxiv.org/html/2608.14198#S6.SS4)\. LLM Backbone and Low Rank Adaptation\.We use instruction\-tuned decoder\-only LLMs and adapt them via LoRA\([12](https://arxiv.org/html/2608.14198#bib.bib28)\)\. Base LLM weights are frozen; gradients flow only through LoRA parameters and the connector parametersϕ\\phi\. This yields parameter\-efficient specialization while preserving general linguistic competence\. Embedding Injection\.Let𝐇∈ℝL×d\\mathbf\{H\}\\in\\mathbb\{R\}^\{L\\times d\}denote the LLM input embeddings for a prompt containingNN<emb\>placeholder tokens\. Each placeholder is replaced with a projected transaction embeddingziz\_\{i\}: 𝐇\[π\(i\)\]←zi,i=1,…,N,\\mathbf\{H\}\[\\pi\(i\)\]\\leftarrow z\_\{i\},\\quad i=1,\\dots,N,whereπ\(i\)\\pi\(i\)denotes the position of theii\-th<emb\>token\. The resulting embedding sequence is processed by the LLM for autoregressive generation\. This approach enables long transaction histories to be injected as compact embeddings rather than serialized into text\. ## 4\.Datasets We formulate transaction understanding as a unified instruction\-following problem that combines transaction history embeddings with natural\-language instructions\. The model is trained on two task families:transaction captioning, which generates summaries of customer behavior \(e\.g\., spending patterns, merchant preferences, and temporal trends\), andtransaction question answering \(QA\), which answers questions grounded in transaction history\. We focus on*multiple\-choice*QA, including both*extractive*questions, where the correct answer is directly supported by observed transactions, and*predictive*questions, which require inference or forecasting from historical behavior\. Depending on the task, the model may additionally generate natural\-language rationales alongside the selected answer\. Prompt Templates\.We represent a transaction history using special tokens interleaved with injected transaction embeddings\. We add<txn\>,</txn\>, and<emb\>to the tokenizer vocabulary\. The<emb\>token is a placeholder whose embedding is replaced at runtime by projected transaction vectors\. We use a fixed\-length history ofNNtransactions per example\. Input: Transaction history of a customer: Transaction at \{time\}:<txn\><emb\></txn\> Transaction at \{time\}:<txn\><emb\></txn\> \.\.\. All tasks are expressed in a unified instruction\-following format with structured output tags\. Captioning: Instruction: Write a structured report \(or semantic description\) of the transaction history\. Provide your summary in this format:<summary\>\.\.\.</summary\> Response: QA with rationale: Instruction: Answer the following question based on the transaction history\. Provide your answer in this format:<answer\>\.\.\.</answer\> Provide your reasoning in this format:<reason\>\.\.\.</reason\> Question: \{question\} Response: Transaction data\.We use a large\-scale proprietary dataset of anonymized card transactions collected from a real\-world payment network\. Each transaction record contains a customer identifier, timestamp, and a set of structured attributes comprising numerical features \(e\.g\., transaction amount\) and high\-cardinality categorical features \(e\.g\., merchant identity, merchant category, and location\)\. Due to the sensitive nature of the data, specific attribute names cannot be disclosed\. Transactions are grouped by customer and ordered chronologically to form behavioral histories that capture both short\-term dynamics and long\-term spending patterns\. The transaction encoder is pre\-trained on a substantially larger corpus spanning billions of transactions\. For multimodal instruction tuning and evaluation, we construct customer\-level train/validation/test splits \(80%/10%/10%\), corresponding to approximately 134k/14k/19k transactions\. This demonstrates that, when initialized from large\-scale pre\-training, MINT can achieve strong multimodal reasoning performance with relatively limited supervision\. To prevent information leakage, all splits are constructed at the customer level\. Caption data\.We construct captioning data via programmatic generation and LLM\-based synthesis\. Structured reports are generated from templates defined over transaction attributes and temporal aggregates, while semantic captions are synthesized by prompting a teacher LLM \(Mistral\-7B\-Instruct\-v0\.3\([2](https://arxiv.org/html/2608.14198#bib.bib31)\)\) with textualized transaction histories\. The resulting dataset contains approximately 134k/14k/19k train/validation/test examples\. QA data\.We generate QA datasets using programmatic templates defined over transaction attributes and temporal aggregates\. Ground\-truth answers are computed via deterministic query execution \(e\.g\., filtering, aggregation, and ranking operations\), ensuring verifiability\. The corpus comprises two task families:*extractive QA*, containing 5\.74M/605k/813k train/validation/test examples spanning 43 question types, and*predictive QA*, containing 3\.47M/366k/492k examples spanning 26 question types\. To evaluate generalization, we additionally construct out of distribution \(OOD\) test sets with previously unseen question types, consisting of 170k extractive QA examples across 9 question types and 189k predictive QA examples across 10 question types\. Chain\-of\-thought data\.To provide reasoning supervision, we synthesize chain\-of\-thought \(CoT\) rationales using a teacher LLM \(Mistral\-7B\-Instruct\-v0\.3\([2](https://arxiv.org/html/2608.14198#bib.bib31)\)\) conditioned on the transaction history, question, and ground\-truth answer\. Specifically, the model receives the full transaction context together with the QA pair and generates a step\-by\-step explanation of the reasoning process\. Generated rationales are filtered using answer\-consistency checks, length constraints, and manual auditing of a subset of examples\. The final CoT datasets comprise 94k/16k/26k extractive QA examples spanning 8 question types and 362k/50k/80k predictive QA examples covering all 26 predictive question types\. ## 5\.Training We train the model in two stages: \(i\) modality alignment, which trains the connector which maps transaction representations into the LLM token space, and \(ii\) supervised fine\-tuning \(SFT\), which jointly optimizes the connector and LoRA adapters on a multi\-task mixture\. Letx=\(𝐓,𝐪,𝐲\)x=\(\\mathbf\{T\},\\mathbf\{q\},\\mathbf\{y\}\)denote a training example, where𝐓\\mathbf\{T\}is a transaction history,𝐪\\mathbf\{q\}is an instruction or question, and𝐲=\(y1,…,y\|𝐲\|\)\\mathbf\{y\}=\(y\_\{1\},\\dots,y\_\{\|\\mathbf\{y\}\|\}\)is the target response\. The frozen encoder producese1:N=E\(𝐓\)e\_\{1:N\}=E\(\\mathbf\{T\}\), which is mapped by the connectorgϕg\_\{\\phi\}to injected embeddingsz1:N=gϕ\(e1:N\)z\_\{1:N\}=g\_\{\\phi\}\(e\_\{1:N\}\)\. Conditioning on𝐪\\mathbf\{q\}and the injected embeddings, the decoder\-only LLM models \(1\)pθ,Δ\(𝐲∣𝐓,𝐪\)=∏t=1\|𝐲\|pθ,Δ\(yt∣𝐲<t,z1:N,𝐪\),p\_\{\\theta,\\Delta\}\(\\mathbf\{y\}\\mid\\mathbf\{T\},\\mathbf\{q\}\)~=~\\prod\_\{t=1\}^\{\|\\mathbf\{y\}\|\}p\_\{\\theta,\\Delta\}\(y\_\{t\}\\mid\\mathbf\{y\}\_\{<t\},z\_\{1:N\},\\mathbf\{q\}\),whereθ\\thetaare frozen base parameters andΔ\\Deltaare LoRA parameters \(when enabled\)\. Modality Alignment\.The goal is to learngϕg\_\{\\phi\}such that injected transaction embeddings are compatible with the LLM token representations\. We train the connector on captioning data𝒟cap\\mathcal\{D\}\_\{\\text\{cap\}\}using a completion\-only autoregressive loss: \(2\)ℒalign\(ϕ\)=𝔼\(𝐓,𝐪,𝐲\)∼𝒟cap\[−∑t∈ℳlogpθ,Δ=0\(yt∣𝐲<t,z1:N,𝐪\)\],\\mathcal\{L\}\_\{\\text\{align\}\}\(\\phi\)~=~\\mathbb\{E\}\_\{\(\\mathbf\{T\},\\mathbf\{q\},\\mathbf\{y\}\)\\sim\\mathcal\{D\}\_\{\\text\{cap\}\}\}\\\!\\left\[\-\\sum\_\{t\\in\\mathcal\{M\}\}\\log p\_\{\\theta,\\Delta=0\}\(y\_\{t\}\\mid\\mathbf\{y\}\_\{<t\},z\_\{1:N\},\\mathbf\{q\}\)\\right\],whereℳ\\mathcal\{M\}denotes the index set corresponding to response tokens \(excluding prompt tokens\)\. During alignment,θ\\thetais frozen and LoRA is disabled \(Δ=0\\Delta=0\); onlyϕ\\phiis updated\. Supervised Fine\-Tuning \(SFT\)\.We then jointly optimize connectorϕ\\phiand LoRAΔ\\Deltaon a mixture of captioning and QA datasets,𝒟=𝒟cap∪𝒟qa\\mathcal\{D\}=\\mathcal\{D\}\_\{\\text\{cap\}\}\\cup\\mathcal\{D\}\_\{\\text\{qa\}\}, using a completion\-only autoregressive loss: \(3\)ℒsft\(ϕ,Δ\)=𝔼\(𝐓,𝐪,𝐲\)∼𝒟\[−∑t∈ℳlogpθ,Δ\(yt∣𝐲<t,z1:N,𝐪\)\]\.\\mathcal\{L\}\_\{\\text\{sft\}\}\(\\phi,\\Delta\)~=~\\mathbb\{E\}\_\{\(\\mathbf\{T\},\\mathbf\{q\},\\mathbf\{y\}\)\\sim\\mathcal\{D\}\}\\\!\\left\[\-\\sum\_\{t\\in\\mathcal\{M\}\}\\log p\_\{\\theta,\\Delta\}\(y\_\{t\}\\mid\\mathbf\{y\}\_\{<t\},z\_\{1:N\},\\mathbf\{q\}\)\\right\]\.We interleave captioning and QA to preserve generation quality while improving grounded reasoning\. For QA examples with rationales,𝐲\\mathbf\{y\}concatenates the reasoning trace and final answer; for answer\-only examples,𝐲\\mathbf\{y\}contains only the short answer\. Decoupled Representation and Reasoning Learning\.MINT separates transaction representation learning from language\-model reasoning\. The transaction sequence encoder is pretrained on large\-scale unlabeled transaction sequences using a next\-transaction prediction objective and remains frozen throughout multimodal training\. Modality alignment and SFT then learn only the connector and LoRA adaptation parameters\. This design treats transaction embeddings as reusable representations that can be injected directly into the LLM, avoiding repeated processing of large transaction corpora during multimodal training\. As a result, pretrained transaction representations can be aligned with the LLM using only a few hundred thousand supervised transaction captioning and QA examples, enabling efficient adaptation to domain\-specific reasoning tasks\. ## 6\.Experiments Table 1\.Overall accuracy on extractive and predictive QA tasks, together with inference efficiency metrics \(TTFT, decode throughput, peak VRAM, and input token count\) measured on Predictive QA\. Results marked with†are trained and evaluated on the same question set and are included for reference only\.Table 2\.Ablation study of MINT\. We analyze the effects of input representation, training data composition, decoding strategy, transaction embedding design, and model capacity on extractive and predictive QA performance \(overall accuracy\) in both ID and OOD settings\.Table 3\.Impact of LLM\-LTM pairing on MINT performance \(overall accuracy\)\. Results are reported for extractive and predictive QA under both ID and OOD evaluation settings across different Qwen3 backbone sizes and pretrained LTMs\.We evaluate on two task families: extractive QA and predictive QA\. Each task family includes both in\-distribution \(ID\) and out\-of\-distribution \(OOD\) question splits\. Unless otherwise stated, we report results for a default MINT configuration using a Qwen3\-1\.7B instruction\-tuned decoder\-only LLM\. To study scaling behavior, we additionally instantiate MINT with Qwen3 models at two other sizes \(0\.6B and 4B\), enabling a controlled analysis of scaling across both the language backbone and the transaction sequence encoder\. ### 6\.1\.Hyperparameters Unless otherwise stated, all experiments use a Qwen3\-1\.7B instruction tuned LLM, a 20M\-parameter transaction sequence encoder, and a two\-layer modality projector with hidden dimension 1024\. The transaction sequence encoder produces 256\-dimensional embeddings, from which the final five transaction embeddings are projected into the 2048\-dimensional LLM embedding space\. For modality alignment, we use an equal mixture of structured report and semantic caption examples\. During SFT, we adapt the LLM using LoRA \(rankr=64r=64, scalingα=2r\\alpha=2r, dropout=0\.05=0\.05\) applied to attention and MLP projections\. The default SFT mixture consists of structured report \(12\.5%\), semantic caption \(12\.5%\), extractive QA \(25%\), predictive QA \(25%\), extractive QA with CoT rationales \(12\.5%\), and predictive QA with CoT rationales \(12\.5%\)\. In thew/o CoTablation, CoT examples are removed and the remaining tasks are reweighted proportionally\. Both modality alignment and SFT are optimized with AdamW using a learning rate of10−410^\{\-4\},\(β1,β2\)=\(0\.9,0\.95\)\(\\beta\_\{1\},\\beta\_\{2\}\)=\(0\.9,0\.95\), weight decay0\.010\.01, warmup ratio0\.030\.03, gradient clipping of1\.01\.0, and mixed\-precision bf16 training\. We train on four A100 GPUs with an effective batch size of 128 samples per optimization step using gradient checkpointing throughout\. Modality alignment is performed for two epochs on 160k samples per epoch, whereas SFT is run for a maximum of five epochs on 320k samples per epoch, with early stopping based on validation performance\. Unless otherwise specified in ablation studies, transaction embeddings are extracted from the encoder’s final layer, timestamps are included in transaction prompts, and decoding uses deterministic greedy generation\. ### 6\.2\.Baselines Text\-serialized LLM baselines\.We serialize each transaction as a structured text record and prepend the full history to the prompt: Input: Transaction history of a customer: Transaction at \{time\}: amount=\{…\}, merchant=\{…\}, city=\{…\}, … Transaction at \{time\}: amount=\{…\}, merchant=\{…\}, city=\{…\}, … \.\.\. Our primary text\-only baseline isLLM SFT, which uses the same training data and LoRA rank as MINT but operates directly on serialized transaction records\. We also evaluated the backbone LLM in a zero\-shot setting; however, it frequently failed to produce outputs conforming to the required answer format, complicating reliable automatic evaluation, and performed substantially worse than the LLM SFT baseline\. For all LLM\-based models, we use a maximum training sequence length of 2048 tokens per input example\. Due to context\-length constraints and computational cost, the LLM SFT baseline cannot accommodate transaction histories of lengthh≥10h\\geq 10, wherehhdenotes the number of historical transactions provided as context\. In contrast, MINT encodes transaction histories into a fixed\-size representation, enabling efficient support for longer histories\. To ensure a fair comparison, MINT and the LLM SFT baseline share the same architecture, training procedure, prompt template, and hyperparameter settings wherever applicable; the only differences are those inherent to multimodal modeling, such as the connector module and its associated parameters\. In particular, the LLM SFT baseline operates directly on raw transaction features, whereas MINT consumes transaction embeddings produced by the transaction sequence encoder\. Additional baselines\.To further contextualize performance, we include: \(i\)Prior sampling, which samples answers according to their empirical frequencies in the training data; \(ii\)task\-specific classifiers, lightweight supervised heads trained on frozen transaction embeddings; and \(iii\) acontrastive alignment model, which aligns transaction embeddings with LLM\-generated transaction summaries using frozen transaction and text encoders and modality\-specific projection heads\. Given transaction and text embeddings\(zitxn,zitext\)\(z\_\{i\}^\{\\text\{txn\}\},z\_\{i\}^\{\\text\{text\}\}\), the model is trained using a symmetric InfoNCE loss, ℒalign=\\displaystyle\\mathcal\{L\}\_\{\\text\{align\}\}~=~12\(ℒtxn→text\+ℒtext→txn\),\\displaystyle\\frac\{1\}\{2\}\\left\(\\mathcal\{L\}\_\{\\text\{txn\}\\rightarrow\\text\{text\}\}\+\\mathcal\{L\}\_\{\\text\{text\}\\rightarrow\\text\{txn\}\}\\right\),ℒtxn→text=\\displaystyle\\mathcal\{L\}\_\{\\text\{txn\}\\rightarrow\\text\{text\}\}~=~−1B∑i=1Blogexp\(⟨zitxn,zitext⟩/τ\)∑j=1Bexp\(⟨zitxn,zjtext⟩/τ\),\\displaystyle\-\\frac\{1\}\{B\}\\sum\_\{i=1\}^\{B\}\\log\\frac\{\\exp\(\\langle z^\{\\text\{txn\}\}\_\{i\},z^\{\\text\{text\}\}\_\{i\}\\rangle/\\tau\)\}\{\\sum\_\{j=1\}^\{B\}\\exp\(\\langle z^\{\\text\{txn\}\}\_\{i\},z^\{\\text\{text\}\}\_\{j\}\\rangle/\\tau\)\},where positives correspond to matched transaction\-summary pairs and all other examples in the minibatch serve as negatives\. ### 6\.3\.Main Results Table[1](https://arxiv.org/html/2608.14198#S6.T1)reports performance across extractive QA and predictive QA \(ID and OOD\), as well as inference metrics \(latency and memory\) to contextualize quality\-efficiency trade\-offs\. Predictive QA\.Under both ID and OOD evaluation, MINT \(h=1h=1\) clearly outperforms LLM SFT \(h=1h=1,h=5h=5\)\. This is the central empirical result of the paper: a single injected transaction embedding is sufficient to outperform text\-serialized histories on the harder, more consequential task of forecasting future behavior\. Extractive QA\.Under ID evaluation, LLM SFT \(h=1h=1,h=5h=5\) is relatively stronger than MINT \(h=1h=1\), likely because the richer surface\-level information in serialized text directly benefits extraction\-style questions\. Both MINT and LLM SFT exhibit monotonic performance gains with increasing history lengthhh\. Under OOD evaluation, MINT \(h=1h=1\) surpasses LLM SFT \(h=1h=1,h=5h=5\)\. Effect of history length\.Across all tasks, LLM SFT performance increases monotonically withhh, at the cost of growing inference overhead\. MINT does not exhibit this monotonic behavior except in extractive QA \(ID\); for predictive QA and OOD tasks, additional history embeddings provide no consistent benefit, and can even hurt performance, indicating that MINT extracts the relevant signal from very limited context\. Additional baselines\.Both LLM SFT and MINT outperform prior sampling and task\-specific classifiers on ID QA tasks\. Prior sampling and task\-specific classifiers are not directly applicable in the OOD setting, as they require training on the target questions; the OOD results reported for these methods are therefore included for reference only\. The substantially lower accuracy of the contrastive alignment baseline indicates that alignment alone is insufficient for effective QA, highlighting the importance of task\-specific supervision beyond caption training\. Inference efficiency\.We measure time\-to\-first\-token \(TTFT\), decode throughput \(tokens/sec\), and peak VRAM at a fixed batch size of 8 while varying history lengthhh, comparing embedding injection \(MINT\) against tokenized serialization \(LLM SFT\)\. Although increasinghhincreases computational cost for both approaches, MINT scales substantially more efficiently\. Notably, MINT \(h=5h=5\) remains more efficient than LLM SFT \(h=1h=1\), achieving 18% lower TTFT, 2\.4×\\timeshigher decode throughput, 11% lower peak VRAM, and 26% fewer input tokens\. These results demonstrate that embedding injection effectively mitigates the token\-length bottleneck inherent to text serialization\. ### 6\.4\.Ablations and Analysis We conduct ablations to understand the impact of architectural design, training strategy, and data composition \(Tables[2](https://arxiv.org/html/2608.14198#S6.T2)and[3](https://arxiv.org/html/2608.14198#S6.T3)\)\. Inclusion of transaction timestamps in the prompt\.Including the timestamp alongside each transaction embedding in the context prompt considerably improves extractive QA \(ID\) performance, but considerably hurts predictive QA \(OOD\) performance\. Extractive QA \(OOD\) and predictive QA \(ID\) are largely insensitive to this choice\. Decoding strategy\.Deterministic \(temperature00\) and stochastic \(e\.g\., temperature0\.70\.7, top\-p=0\.9p=0\.9\) decoding yield roughly similar performance overall, with deterministic decoding slightly favored on predictive QA \(ID\) task\. Rationale supervision\.Incorporating rationale\-augmented QA data \(𝒟qa\-cot\\mathcal\{D\}\_\{\\text\{qa\-cot\}\}\) alongside captioning \(𝒟cap\\mathcal\{D\}\_\{\\text\{cap\}\}\) and standard QA \(𝒟qa\\mathcal\{D\}\_\{\\text\{qa\}\}\) considerably improves performance across tasks, underscoring the value of CoT supervision even when it is not used explicitly at inference time\. LoRA rank\.Varying LoRA rankr∈\{32,64,128\}r\\in\\\{32,64,128\\\}has little effect on performance overall, except for predictive QA \(OOD\), wherer=32r=32performs considerably better\. This suggests moderate ranks offer the best capacity\-efficiency trade\-off\. Connector hidden size\.Varying the connector’s hidden sizeℓ∈\{256,512,1024\}\\ell\\in\\\{256,512,1024\\\}has minimal impact on ID performance, butℓ=512\\ell=512considerably improves both extractive and predictive QA \(OOD\), indicating a capacity sweet spot for out\-of\-distribution generalization\. Number of injected transaction embeddings \(from the final encoder layer\)\.Increasinghh\(i\) monotonically improves extractive QA \(ID\), \(ii\) leaves predictive QA \(ID\) largely unchanged, and \(iii\) is best ath=1h=1for both extractive and predictive OOD, with no clear benefit from largerhh\. This suggests that history length should be adapted per task type rather than fixed globally\. Number of injected transaction embeddings \(from the second\-to\-last encoder layer\)\.Motivated by prior work showing that intermediate transformer layers can yield stronger representations than the final layer for downstream tasks, we additionally evaluate embeddings from the second\-to\-last layer of the transaction sequence encoder\([24](https://arxiv.org/html/2608.14198#bib.bib25)\)\. Second\-to\-last\-layer embeddings show the same qualitative trends as the final layer for ID tasks \(monotonic gains for extractive QA, robustness for predictive QA\), but no clear pattern for OOD tasks\. Comparing layers directly: for a givenhh, second\-to\-last\-layer embeddings are consistently better on ID tasks but generally weaker on OOD tasks than final\-layer embeddings, revealing a layer\-depth trade\-off between in\-distribution precision and out\-of\-distribution robustness\. Encoder/LLM capacity scaling \(Table[3](https://arxiv.org/html/2608.14198#S6.T3)\)\.We jointly vary transaction sequence encoder size \(20M / 120M / 720M\) and LLM scale \(0\.6B\-4B\)\. On extractive QA \(ID\), the 4B LLM yields slightly better performance; on predictive QA \(ID\), all three LLM sizes perform similarly\. For both ID tasks, performance is largely robust to transaction sequence encoder size at a given LLM size\. On OOD tasks, the 0\.6B LLM is noticeably weaker, while the 1\.7B model performs best overall; for the 1\.7B model specifically, the 120M transaction sequence encoder yields the best OOD performance\. Taken together, Qwen3\-1\.7B offers the most favorable balance of quality and inference cost among the scales tested\. As an additional analysis, we evaluated information loss introduced by modality projection using linear classifiers trained on projected embeddings\. The original transaction embeddings achieved accuracies of 0\.756 on extractive QA and 0\.752 on predictive QA\. After projection, the contrastive alignment model retained similar performance \(0\.766 and 0\.732\), while the MINT projector achieved 0\.724 and 0\.74 accuracy on extractive and predictive QA, respectively\. These results suggest that both projectors preserve most of the task\-relevant signal in the transaction representations, with only modest degradation relative to the original embedding space\. To assess retention of general\-purpose reasoning capabilities, we evaluate the LoRA\-adapted models on standard language\-model benchmarks\. Both MINT and LLM\-SFT exhibit substantial performance degradation relative to the base Qwen3\-1\.7B model\. For example, GSM8K accuracy drops from 0\.412 to 0\.026 \(MINT\) and 0\.033 \(LLM\-SFT\), while ARC\-Challenge declines from 0\.425 to 0\.306/0\.332 and BoolQ from 0\.776 to 0\.493/0\.492\. These results reflect the well\-known trade\-off between domain specialization and general\-purpose reasoning\. Importantly, the degradation is observed only with the LoRA adapters enabled\. Since adaptation is performed via PEFT and leaves the underlying foundation model unchanged, users can simply disable the adapters to recover the original capabilities of the base LLM, enabling seamless switching between transaction\-specialized and general\-purpose modes\. ## 7\.Conclusion We presentedMINT, a multimodal transaction reasoning framework that integrates a pretrained transaction sequence encoder with a decoder\-only LLM through embedding injection\. Across extractive QA, predictive QA, and transaction captioning tasks,MINTdemonstrates that compact transaction representations can support effective reasoning while maintaining favorable inference efficiency\. Beyond task performance, we conducted a systematic study of the multimodal transaction reasoning design space, examining the effects of transaction\-history representation, modality projection, training\-data composition, prompt design, and inference\-time trade\-offs\. Compared to prior work,MINTincorporates both captioning and chain\-of\-thought supervision during training, achieves state\-of\-the\-art zero\-shot predictive QA performance on a large\-scale real\-world transaction dataset, and provides one of the most comprehensive evaluations of multimodal transaction reasoning to date, spanning predictive performance, robustness, latency, memory consumption, and detailed ablation analyses\. Our findings reveal complementary strengths between transaction embeddings and text serialization: embeddings excel at predictive reasoning, whereas serialized histories remain highly effective for extractive retrieval\. We hypothesize that the predictive advantage arises from the encoder’s autoregressive self\-supervised pretraining objective, which is naturally aligned with forecasting future behavior\. Overall, our results highlight pretrained transaction foundation models as a powerful substrate for efficient multimodal transaction reasoning\. Limitations\.Our study uses proprietary transaction data that cannot be publicly released due to privacy and regulatory constraints\. To support reproducibility, we provide detailed descriptions of the datasets, prompts, model architectures, training procedures, and hyperparameters\. More broadly, transaction reasoning systems raise important concerns regarding privacy, profiling, fairness, and potential misuse\. Any real\-world deployment should incorporate robust privacy protections, access controls, auditing, bias assessment, and appropriate human oversight\. Future work\.Several promising directions remain\. First, we plan to extend beyond structured QA toward open\-ended transaction reasoning and discovery tasks, including comparative analysis of customer behavior, context\-aware behavioral assessment, and fraud\-pattern discovery\. Second, improving reliability through confidence\-aware reasoning, e\.g\., generating\(answer, rationale, confidence\)tuples, may lead to better calibrated decision\-support systems\. Third, reinforcement\-based post\-training offers a promising avenue for improving reasoning quality while maintaining compact multimodal representations\. Finally, we intend to explore hybrid approaches that combine transaction embeddings with selectively retrieved textual context and prompt\-compression techniques, potentially improving both reasoning performance and scalability\. Key FindingsF1\. Predictive reasoning benefits from transaction embeddings\.MINT consistently outperforms LLM SFT on predictive QA, suggesting that transaction embeddings capture predictive signals more effectively than textual serialization\.F2\. Text serialization is strong for extractive retrieval\.Serialized transaction histories provide rich contextual detail that benefits extractive QA, especially as the number of transactions included in the prompt increases\.F3\. Embedding injection scales more efficiently than text serialization\.MINT achieves comparable or stronger performance while requiring substantially fewer input tokens, lower latency, and less GPU memory\.F4\. Additional transaction history is task dependent\.Increasing history length consistently improves extractive QA but offers limited gains for predictive QA, suggesting that optimal context size for MINT should depend on the reasoning task\.F5\. ID and OOD settings prefer different representations\.Earlier transaction sequence encoder layers \(second\-last layer\) consistently improve ID performance, whereas final\-layer embeddings generally provide stronger OOD generalization\.F6\. Chain\-of\-thought supervision improves performance\.Adding CoT data to the training mixture improves performance broadly across tasks, particularly for extractive QA\. ## References - Abdullaevaet al\.\(2026\)I\. Abdullaevaet al\.ESQA: Event Sequences Question Answering\.IEEE Access\.Cited by:[§1](https://arxiv.org/html/2608.14198#S1.p2.1),[§2](https://arxiv.org/html/2608.14198#S2.p2.1),[§2](https://arxiv.org/html/2608.14198#S2.p3.1)\. - AI \(2024\)M\. AIMistral\-7B\-Instruct\-v0\.3\.Note:[https://huggingface\.co/mistralai/Mistral\-7B\-Instruct\-v0\.3](https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.3)Cited by:[§4](https://arxiv.org/html/2608.14198#S4.p10.1),[§4](https://arxiv.org/html/2608.14198#S4.p12.1)\. - Aminianet al\.\(2025\)G\. Aminianet al\.FraudTransformer: Time\-Aware GPT for Transaction Fraud Detection\.arXiv preprint arXiv:2509\.23712\.Cited by:[§2](https://arxiv.org/html/2608.14198#S2.p1.1)\. - Azorinet al\.\(2024\)R\. Azorin, Z\. Ben\-Houidi, M\. Gallo, A\. Finamore, and P\. MichiardiFine\-grained Attention in Hierarchical Transformers for Tabular Time\-series\.arXiv preprint arXiv:2406\.15327\.Cited by:[§2](https://arxiv.org/html/2608.14198#S2.p1.1)\. - Babaevet al\.\(2022\)D\. Babaevet al\.CoLES: Contrastive Learning for Event Sequences with Self\-Supervision\.InSIGMOD,Cited by:[§2](https://arxiv.org/html/2608.14198#S2.p1.1)\. - Braithwaiteet al\.\(2025\)D\. Braithwaiteet al\.Your Spending Needs Attention: Modeling Financial Habits with Transformers\.arXiv preprint arXiv:2507\.23267\.Cited by:[§2](https://arxiv.org/html/2608.14198#S2.p1.1)\. - Chowet al\.\(2024\)W\. Chow, L\. Gardiner, H\. T\. Hallgrímsson, M\. A\. Xu, and S\. Y\. RenTowards Time\-Series Reasoning with LLMs\.arXiv preprint arXiv:2409\.11376\.Cited by:[§1](https://arxiv.org/html/2608.14198#S1.p2.1),[§2](https://arxiv.org/html/2608.14198#S2.p2.1)\. - Clarket al\.\(2020\)K\. Clark, M\. Luong, Q\. V\. Le, and C\. D\. ManningELECTRA: Pre\-training Text Encoders as Discriminators Rather Than Generators\.InICLR,Cited by:[§2](https://arxiv.org/html/2608.14198#S2.p1.1)\. - Douet al\.\(2025\)Y\. Douet al\.TransactionGPT\.arXiv preprint arXiv:2511\.08939\.Cited by:[§2](https://arxiv.org/html/2608.14198#S2.p1.1)\. - Fadeevet al\.\(2025\)E\. Fadeevet al\.LATTE: Learning Aligned Transactions and Textual Embeddings for Bank Clients\.arXiv preprint arXiv:2508\.10021\.Cited by:[§2](https://arxiv.org/html/2608.14198#S2.p4.1)\. - Gruveret al\.\(2023\)N\. Gruver, M\. Finzi, S\. Qiu, and A\. G\. WilsonLarge Language Models Are Zero\-Shot Time Series Forecasters\.NeurIPS\.Cited by:[§1](https://arxiv.org/html/2608.14198#S1.p2.1),[§2](https://arxiv.org/html/2608.14198#S2.p2.1)\. - Huet al\.\(2022\)E\. J\. Huet al\.LoRA: Low\-Rank Adaptation of Large Language Models\.ICLR\.Cited by:[§3](https://arxiv.org/html/2608.14198#S3.p4.1)\. - Jhaet al\.\(2012\)S\. Jha, M\. Guillen, and J\. C\. WestlandEmploying transaction aggregation strategy to detect credit card fraud\.Expert Systems with Applications\.Cited by:[§2](https://arxiv.org/html/2608.14198#S2.p1.1)\. - Konget al\.\(2025\)Y\. Konget al\.Time\-MQA: Time Series Multi\-Task Question Answering with Context Enhancement\.arXiv preprint arXiv:2503\.01875\.Cited by:[§1](https://arxiv.org/html/2608.14198#S1.p2.1),[§2](https://arxiv.org/html/2608.14198#S2.p2.1)\. - Langeret al\.\(2025\)P\. Langeret al\.OpenTSLM: Time\-Series Language Models for Reasoning over Multivariate Medical Text\- and Time\-Series Data\.arXiv preprint arXiv:2510\.02410\.Cited by:[§1](https://arxiv.org/html/2608.14198#S1.p2.1),[§2](https://arxiv.org/html/2608.14198#S2.p2.1)\. - Liet al\.\(2025\)G\. Liet al\.PANTHER: Generative Pretraining Beyond Language for Sequential User Behavior Modeling\.NeurIPS\.Cited by:[§2](https://arxiv.org/html/2608.14198#S2.p1.1)\. - Liet al\.\(2023\)J\. Li, D\. Li, S\. Savarese, and S\. HoiBLIP\-2: Bootstrapping Language\-Image Pre\-training with Frozen Image Encoders and Large Language Models\.InICML,Cited by:[§2](https://arxiv.org/html/2608.14198#S2.p3.1)\. - Liuet al\.\(2023\)H\. Liu, C\. Li, Q\. Wu, and Y\. J\. LeeVisual Instruction Tuning\.NeurIPS\.Cited by:[§2](https://arxiv.org/html/2608.14198#S2.p3.1)\. - Marafiotiet al\.\(2025\)A\. Marafiotiet al\.SmolVLM: Redefining small and efficient multimodal models\.InCOLM,Cited by:[§2](https://arxiv.org/html/2608.14198#S2.p3.1)\. - Ostroukhovet al\.\(2026\)M\. Ostroukhovet al\.PRAGMA: Revolut Foundation Model\.arXiv preprint arXiv:2604\.08649\.Cited by:[§1](https://arxiv.org/html/2608.14198#S1.p3.1),[§2](https://arxiv.org/html/2608.14198#S2.p1.1)\. - Padhiet al\.\(2021\)I\. Padhiet al\.Tabular Transformers for Modeling Multivariate Time Series\.InICASSP,Cited by:[§1](https://arxiv.org/html/2608.14198#S1.p3.1),[§2](https://arxiv.org/html/2608.14198#S2.p1.1),[§3](https://arxiv.org/html/2608.14198#S3.p2.1)\. - Ramanet al\.\(2024\)N\. Raman, S\. Ganesh, and M\. VelosoScalable Representation Learning for Multimodal Tabular Transactions\.arXiv preprint arXiv:2410\.07851\.Cited by:[§1](https://arxiv.org/html/2608.14198#S1.p2.1),[§2](https://arxiv.org/html/2608.14198#S2.p2.1),[§2](https://arxiv.org/html/2608.14198#S2.p3.1)\. - Skalskiet al\.\(2023\)P\. Skalski, D\. Sutton, S\. Burrell, I\. Perez, and J\. WongTowards a Foundation Purchasing Model: Pretrained Generative Autoregression on Transaction Sequences\.InICAIF,Cited by:[§1](https://arxiv.org/html/2608.14198#S1.p3.1),[§2](https://arxiv.org/html/2608.14198#S2.p1.1),[§3](https://arxiv.org/html/2608.14198#S3.p2.1)\. - Skeanet al\.\(2025\)O\. Skeanet al\.Layer by Layer: Uncovering Hidden Representations in Language Models\.InICML,Cited by:[§6\.4](https://arxiv.org/html/2608.14198#S6.SS4.p8.1)\. - Xieet al\.\(2024\)Z\. Xieet al\.ChatTS: Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and Reasoning\.arXiv preprint arXiv:2412\.03104\.Cited by:[§1](https://arxiv.org/html/2608.14198#S1.p2.1),[§2](https://arxiv.org/html/2608.14198#S2.p2.1)\. - Yehet al\.\(2026\)C\. M\. Yehet al\.TREASURE: The Visa Payment Foundation Model for High\-Volume Transaction Understanding\.InKDD,Cited by:[§2](https://arxiv.org/html/2608.14198#S2.p1.1)\. - Yuet al\.\(2025\)F\. Yu, H\. Zhao, and T\. ZhouTS\-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning\.arXiv preprint arXiv:2510\.03519\.Cited by:[§1](https://arxiv.org/html/2608.14198#S1.p2.1),[§2](https://arxiv.org/html/2608.14198#S2.p2.1)\. - Zhanget al\.\(2025a\)J\. Zhanget al\.TimeMaster: Training Time\-Series Multimodal LLMs to Reason via Reinforcement Learning\.arXiv preprint arXiv:2506\.13705\.Cited by:[§1](https://arxiv.org/html/2608.14198#S1.p2.1),[§2](https://arxiv.org/html/2608.14198#S2.p2.1)\. - Zhang and Yan \(2023\)Y\. Zhang and J\. YanCrossformer: Transformer Utilizing Cross\-Dimension Dependency for Multivariate Time Series Forecasting\.InICLR,Cited by:[§2](https://arxiv.org/html/2608.14198#S2.p1.1)\. - Zhanget al\.\(2025b\)Y\. Zhanget al\.SensorLM: Learning the Language of Wearable Sensors\.arXiv preprint arXiv:2506\.09108\.Cited by:[§2](https://arxiv.org/html/2608.14198#S2.p4.1)\. - Zhouet al\.\(2025\)Y\. Zhouet al\.Time Series Forecasting as Reasoning: A Slow\-Thinking Approach with Reinforced LLMs\.arXiv preprint arXiv:2506\.10630\.Cited by:[§1](https://arxiv.org/html/2608.14198#S1.p2.1),[§2](https://arxiv.org/html/2608.14198#S2.p2.1)\.
Similar Articles
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction
This paper introduces AdaMTP, an adaptive training paradigm for multi-token prediction that dynamically aligns prediction horizons with sequence predictability using entropy-based segmentation, consistently outperforming standard MTP on math, code, and general benchmarks across three LLM backbones.
Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs
This paper introduces Zero-Shot Embedding Drift Detection (ZEDD), a lightweight framework that detects prompt injection attacks in LLMs by measuring semantic shifts in embedding space, achieving over 93% accuracy with less than 3% false positive rate across multiple architectures.
MINT: Tensor Decomposition on Stacked Recurrence Matrices for Time Series Data Mining
This paper introduces MINT, a method that stacks recurrence (self-similarity) matrices of multiple time series into a tensor and applies tensor decomposition to mine co-clustered cross-sensor patterns. Experiments on transit, electricity, wind turbine, and traffic data show effective co-clustering of motifs in regular time series.
A Foundation Model for Multimodal Event Sequences in Financial Applications
This paper presents a foundation transformer model pretrained on multimodal event sequences for financial applications, unifying heterogeneous data sources and using next-event prediction to learn general-purpose representations. The approach outperforms traditional task-specific models and was deployed at a major Eastern European bank, yielding measurable business improvements.
MindZero: Learning Online Mental Reasoning With Zero Annotations
MindZero introduces a self-supervised reinforcement learning framework that trains multimodal large language models for efficient and robust online mental reasoning without requiring mental state annotations, outperforming model-based methods in accuracy and efficiency.