SuperThoughts: Reasoning Tokens in Superposition

arXiv cs.LG Papers

Summary

SuperThoughts compresses consecutive chain-of-thought tokens into latent representations and decodes two tokens per step, achieving ~20–30% CoT length reduction with minimal accuracy loss on math reasoning benchmarks, while doubling inference throughput.

arXiv:2606.13862v1 Announce Type: new Abstract: Long Chain-of-Thought (CoT) reasoning improves LLM problem-solving but is computationally expensive due to sequential token generation. While recent works explore reasoning in continuous latent spaces to bypass discrete token generation, they often struggle with training stability and fail to scale to complex, long-horizon tasks due to lack of supervision signal. We propose SuperThoughts, which compresses pairs of consecutive CoT tokens into single latent representations and decodes two tokens per step via a lightweight Multi-Token Prediction (MTP) module. This preserves discrete token supervision at training time while doubling throughput at inference time. We finetune Qwen2.5-Math-1.5B-Instruct, Qwen2.5-Math-7B-Instruct, Qwen2.5-Math-14B-Instruct, and evaluate on MATH500, AMC, OlympiadBench, and GPQA-Diamond. With a confidence-based adaptive mechanism that falls back to standard decoding when uncertain, SuperThoughts achieves $\sim$20--30\% CoT length reduction while maintaining accuracy with minimal degradation (1-2 points accuracy drop on most tasks).
Original Article
View Cached Full Text

Cached at: 06/15/26, 09:08 AM

# SuperThoughts: Reasoning Tokens in Superposition
Source: [https://arxiv.org/html/2606.13862](https://arxiv.org/html/2606.13862)
Zheyang Xiongw,m, Shivam Garg∗m, Max Yu∗i, Vaishnavi Shrivastavam, Haoyu Zhaop,m Anastasios Kyrillidisr,Dimitris Papailiopoulosw,m wUniversity of Wisconsin\-Madison,mMicrosoft Research,iIndependent pPrinceton University,rRice University

###### Abstract

Long Chain\-of\-Thought \(CoT\) reasoning improves LLM problem\-solving but is computationally expensive due to sequential token generation\. While recent works explore reasoning in continuous latent spaces to bypass discrete token generation, they often struggle with training stability and fail to scale to complex, long\-horizon tasks due to lack of supervision signal\. We propose SuperThoughts, which compresses pairs of consecutive CoT tokens into single latent representations and decodes two tokens per step via a lightweight Multi\-Token Prediction \(MTP\) module\. This preserves discrete token supervision at training time while doubling throughput at inference time\. We finetune Qwen2\.5\-Math\-1\.5B\-Instruct, Qwen2\.5\-Math\-7B\-Instruct, Qwen2\.5\-Math\-14B\-Instruct, and evaluate on MATH500, AMC, OlympiadBench, and GPQA\-Diamond\. With a confidence\-based adaptive mechanism that falls back to standard decoding when uncertain, SuperThoughts achieves∼\\sim20–30% CoT length reduction while maintaining accuracy with minimal degradation \(1\-2 points accuracy drop on most tasks\)\.![Refer to caption](https://arxiv.org/html/2606.13862v1/x1.png)Figure 1:Comparison between SuperThoughts and HAMburger\(Liu and Zhang,[2025](https://arxiv.org/html/2606.13862#bib.bib39)\)on trained Qwen2\.5\-1\.5B\-Math\-Instruct\.

00footnotetext:∗Equal contribution\. Email:<zheyang@cs\.wisc\.edu\>\. Correspondence:<dimitris@papail\.io\>\.## 1Introduction

Large language models \(LLMs\) solve complex problems by generating explicit Chain\-of\-Thought \(CoT\) sequences before arriving at a final answer\(Weiet al\.,[2022](https://arxiv.org/html/2606.13862#bib.bib9)\)\. We can view each CoT token as a unit of compute \(one forward pass\), and longer chains mean more computation spent before reaching the answer\. Recent successes such as OpenAI o1\(Jaechet al\.,[2024](https://arxiv.org/html/2606.13862#bib.bib11)\)and DeepSeek\-R1\(Guoet al\.,[2025](https://arxiv.org/html/2606.13862#bib.bib12)\)demonstrate that this additional test\-time compute substantially improves performance\(Snellet al\.,[2024](https://arxiv.org/html/2606.13862#bib.bib13)\)\.

This raises a question:*why must the model reason in discrete token space?*The vocabulary of a language model is a finite, human\-interpretable set of symbols, yet the model’s internal representations live in a continuous, high\-dimensional vector space\. If reasoning could occur directly in this richer latent space, the model might express more intermediate computations per step, achieving the same quality with fewer steps, or better quality with the same compute\.

Recent work explores*latent reasoning*, which aims to bypass discrete token generation\.Haoet al\.\([2024](https://arxiv.org/html/2606.13862#bib.bib16)\)propose COCONUT that trains models to reason with continuous latent thoughts that are never decoded into language\.Cheng and Van Durme \([2024](https://arxiv.org/html/2606.13862#bib.bib19)\)compress chain\-of\-thought into dense representations via knowledge distillation\. Other approaches explore hybrid schemes that interleave latent and discrete tokens\(Suet al\.,[2025](https://arxiv.org/html/2606.13862#bib.bib25); Shenet al\.,[2025b](https://arxiv.org/html/2606.13862#bib.bib21); Zhanget al\.,[2025a](https://arxiv.org/html/2606.13862#bib.bib24)\)\.

However, these methods face a key challenge:*the lack of intermediate supervision*\. Standard CoT training benefits from token\-level cross\-entropy loss at every reasoning step, providing dense gradient signal throughout the reasoning chain\. When reasoning occurs in an unconstrained latent space, this supervision vanishes and the model must learn to produce useful intermediate representations without any direct feedback on what those representations should encode\. This makes training unstable and prone to representational drift, particularly for long\-horizon tasks where errors compound across many latent steps\. As a result, prior latent reasoning methods have been demonstrated primarily on simple settings and often struggle to match the performance of explicit CoT on challenging benchmarks\.

*Can we train the model to reason in a richer, superposed space while keeping intermediate supervision?*

![Refer to caption](https://arxiv.org/html/2606.13862v1/x2.png)Figure 2:Comparison of three generation strategies for producing tokens “b” through “e”\.\(a\) Standard:Each forward pass consumes one token and predicts one token, requiring 4 steps\.\(b\) Standard \+ MTP:A Multi\-Token Prediction head predicts an additional token per step, but inputs remain single tokens, still requiring 4 steps\.\(c\) SuperThoughts:Token pairs are fused into superposed embeddings as input, and two tokens are decoded per step via MTP, halving the required forward passes to 2 steps\. Green denotes main model predictions; blue denotes MTP predictions\.In this work, we explore a natural first step toward this goal\. We proposeSuperThoughts, a framework that compresses*pairs*of consecutive CoT tokens into single latent representations during reasoning\. At each step, the model consumes a superposed embedding of two tokens and predicts two discrete tokens: one from the main model backbone and one from a lightweight Multi\-Token Prediction \(MTP\) module\(Gloeckleet al\.,[2024](https://arxiv.org/html/2606.13862#bib.bib5); Liuet al\.,[2024](https://arxiv.org/html/2606.13862#bib.bib7)\)\. This halves the number of forward passes required while maintaining token\-level cross\-entropy supervision throughout training\.

Our main contributions are:

1. 1\.We propose SuperThoughts, an architecture that compresses token pairs into single representations via a Compressor and decodes two tokens per step using a Main Module and an MTP Module\.
2. 2\.We develop a two\-stage training protocol that first aligns the compressed latent space via distillation\(Bertonet al\.,[2025](https://arxiv.org/html/2606.13862#bib.bib38)\), then jointly trains all components end\-to\-end with discrete token supervision\.
3. 3\.We introduce a confidence\-based adaptive inference mechanism that falls back to standard decoding when the MTP module is uncertain, trading throughput for accuracy on difficult reasoning steps\.
4. 4\.We evaluate on MATH500\(Hendryckset al\.,[2021](https://arxiv.org/html/2606.13862#bib.bib55)\), AMC23\(MAA,[2023](https://arxiv.org/html/2606.13862#bib.bib58)\), OlympiadBench\(Heet al\.,[2024](https://arxiv.org/html/2606.13862#bib.bib56)\)and GPQA\-Diamond\(Reinet al\.,[2024](https://arxiv.org/html/2606.13862#bib.bib57)\), achieving2020–35%35\\%CoT length reduction while maintaining accuracy within11–22points of the baseline\.

## 2Related Works

#### Latent Reasoning in LLMs\.

When prompted with a question, LLMs can generate intermediate reasoning via discrete tokens before answering the question, and such reasoning process is termed chain\-of\-thought \(CoT\)\(Weiet al\.,[2022](https://arxiv.org/html/2606.13862#bib.bib9)\)\. Recently, several works focus on using CoT states beyond discrete tokens\.Haoet al\.\([2024](https://arxiv.org/html/2606.13862#bib.bib16)\); Yueet al\.\([2025](https://arxiv.org/html/2606.13862#bib.bib20)\); Shenet al\.\([2025b](https://arxiv.org/html/2606.13862#bib.bib21)\)introduce methods that directly feed the last continuous hidden state as input embedding for the next step\. However, these methods either require complicated training curriculum or only consider simple settings\.Giannouet al\.\([2025](https://arxiv.org/html/2606.13862#bib.bib22)\); Zhanget al\.\([2025a](https://arxiv.org/html/2606.13862#bib.bib24)\); Denget al\.\([2025](https://arxiv.org/html/2606.13862#bib.bib28)\); Shenet al\.\([2025a](https://arxiv.org/html/2606.13862#bib.bib27)\)generate first and then compress the newly generated tokens, but only save context length and involve attention mask manipulations that are not compatible with modern inference engines\(Kwonet al\.,[2023](https://arxiv.org/html/2606.13862#bib.bib49); Zhenget al\.,[2024](https://arxiv.org/html/2606.13862#bib.bib50)\)\.Cheng and Van Durme \([2024](https://arxiv.org/html/2606.13862#bib.bib19)\); Suet al\.\([2025](https://arxiv.org/html/2606.13862#bib.bib25)\); Tanet al\.\([2025](https://arxiv.org/html/2606.13862#bib.bib29)\)trains the model to compress discrete CoT into latent tokens and during inference generate latent tokens directly\. Several works explore composing multiple next token choices into a latent input token\(Zhanget al\.,[2025b](https://arxiv.org/html/2606.13862#bib.bib17); Zhuanget al\.,[2025](https://arxiv.org/html/2606.13862#bib.bib18); Zhuet al\.,[2025](https://arxiv.org/html/2606.13862#bib.bib23); Jain and Rappazzo,[2025](https://arxiv.org/html/2606.13862#bib.bib26); Wuet al\.,[2025](https://arxiv.org/html/2606.13862#bib.bib30); Tanget al\.,[2026](https://arxiv.org/html/2606.13862#bib.bib36); Gozetenet al\.,[2026](https://arxiv.org/html/2606.13862#bib.bib31)\)\.Penget al\.\([2026](https://arxiv.org/html/2606.13862#bib.bib60)\)pretrains LLMs with token superposition and yields pretraining time speedup\.

#### Compressed Input Context\.

In addition to latent reasoning, there have been many works that compress more information into input embeddings\. Prefix Tuning\(Li and Liang,[2021](https://arxiv.org/html/2606.13862#bib.bib32)\)uses a learned soft embedding prefix to condition the LLM\. Many works compress input context tokens to save context length\(Jianget al\.,[2023](https://arxiv.org/html/2606.13862#bib.bib33); Liet al\.,[2023](https://arxiv.org/html/2606.13862#bib.bib34); Muet al\.,[2023](https://arxiv.org/html/2606.13862#bib.bib35); Bertonet al\.,[2025](https://arxiv.org/html/2606.13862#bib.bib38); Feldman and Artzi,[2025](https://arxiv.org/html/2606.13862#bib.bib37)\)\.

#### Reducing Discrete CoT Tokens\.

Many methods have also produced shorter discrete CoT sequences through Reinforcement Learning\(Aggarwal and Welleck,[2025](https://arxiv.org/html/2606.13862#bib.bib40); Shrivastavaet al\.,[2025](https://arxiv.org/html/2606.13862#bib.bib41)\)and fine\-tuning\(Xiaet al\.,[2025](https://arxiv.org/html/2606.13862#bib.bib42)\)\. Notably, these discrete CoT tokens length reduction methods are orthogonal to SuperThoughts\.

#### Variable Compute Per Token\.

Recent work has explored adaptive compute allocation in language models by moving beyond uniform token\-level processing, such as BLT\(Pagnoniet al\.,[2025](https://arxiv.org/html/2606.13862#bib.bib44)\)and H\-Net\(Hwanget al\.,[2025](https://arxiv.org/html/2606.13862#bib.bib43)\)segmenting bytes into dynamically\-sized patches, and DLCM\(Quet al\.,[2026](https://arxiv.org/html/2606.13862#bib.bib45)\)learning variable\-length semantic concepts on top of tokens\.Liu and Zhang \([2025](https://arxiv.org/html/2606.13862#bib.bib39)\)propose HAMBURGER, which similarly fuses multiple tokens into a single input embedding via a compositional embedder and decodes several tokens per forward through a micro\-step decoder\.

#### Multi\-token Prediction\.

Traditionally, LLMs are trained with next\-token prediction loss where the model is provided with a prefix and asked to predict the next token that follows the prefix\(Radfordet al\.,[2019](https://arxiv.org/html/2606.13862#bib.bib3)\)\.Bachmann and Nagarajan \([2024](https://arxiv.org/html/2606.13862#bib.bib4)\)argue that teacher\-forcing in next\-token prediction results in inaccurate next\-token predictor and proposes a solution that learns to predict multiple tokens\.Gloeckleet al\.\([2024](https://arxiv.org/html/2606.13862#bib.bib5)\)pre\-train LLMs from scratch that predicts multiple future tokens at once using multiple output heads and show that multi\-token prediction \(MTP\) is better than next\-token prediction \(NTP\) on larger models\. DeepSeek\-V3\(Liuet al\.,[2024](https://arxiv.org/html/2606.13862#bib.bib7)\)also train the model with MTP objective but use a lightweight MTP module instead of an independent output head\.Ahnet al\.\([2025](https://arxiv.org/html/2606.13862#bib.bib8)\)propose joint multi\-token prediction \(JTP\) by employing a representation bottleneck that encourages the model to encode richer information in the output hidden state\.

Despite MTP predicting multiple tokens at once, at inference time current MTP architectures can only utilize the extra tokens for self\-speculative decoding\(Liuet al\.,[2024](https://arxiv.org/html/2606.13862#bib.bib7); Gloeckleet al\.,[2024](https://arxiv.org/html/2606.13862#bib.bib5); Caiet al\.,[2024](https://arxiv.org/html/2606.13862#bib.bib6)\), since the main model still needs to populate KV entries for tokens the MTP modules generate\. Critically, this does not reduce the total FLOPs at inference time\. The main model must still perform a full forward pass over every accepted token, meaning self\-speculative decoding with MTP targets latency reduction under low GPU utilization rather than computational efficiency\.

#### Scaling Test\-Time Compute\.

Recent scaling laws suggest that optimizing test\-time compute can outperform simply increasing parameter counts\(Snellet al\.,[2024](https://arxiv.org/html/2606.13862#bib.bib13)\)\. Leading reasoning models, such as OpenAI o1\(Jaechet al\.,[2024](https://arxiv.org/html/2606.13862#bib.bib11)\)and DeepSeek\-R1\(Guoet al\.,[2025](https://arxiv.org/html/2606.13862#bib.bib12)\), utilize Reinforcement Learning extended CoT sequences\.

## 3Methods

![Refer to caption](https://arxiv.org/html/2606.13862v1/x3.png)Figure 3:Overview of SuperThoughts architecture\.At each stepii, the Compressor encodes a CoT token pair\(c2​i−1,c2​i\)\(c\_\{2i\-1\},c\_\{2i\}\)into a single latent vector𝒙i\{\\bm\{x\}\}\_\{i\}via a learned2​H→H2H\\to Hcompressor, whereHHis the dimension of a single token embedding\. The Main module processes𝒙i\{\\bm\{x\}\}\_\{i\}to produce hidden state𝒉i\{\\bm\{h\}\}\_\{i\}and predicts the next odd\-indexed tokenc2​i\+1c\_\{2i\+1\}\. The MTP module then receives a projection𝒙i′\{\\bm\{x\}\}^\{\\prime\}\_\{i\}that combines three inputs – the previous even token, the just\-predicted odd token from Main, and the Main hidden state – via a learned3​H→H3H\\to Hprojection, and predicts the corresponding even\-indexed tokenc2​i\+2c\_\{2i\+2\}\. Both modules share the same output LM head\. This design enables the model to consume two tokens and generate two tokens per step\.Standard Chain\-of\-Thought \(CoT\) reasoning generates a sequence of discrete tokensc1:L=\(c1,…,cL\)c\_\{1:L\}=\(c\_\{1\},\\dots,c\_\{L\}\)autoregressively, requiringLLforward passes to produce a reasoning chain of lengthLL\. We introduceSuperThoughts, a framework that halves this computational cost by processing and generating tokens in pairs\. At each reasoning step, the model consumes two tokens and predicts two tokens, reducing the number of forward passes fromLLtoL/2L/2while preserving discrete token supervision\.

In this section, we detail:\(1\) Architecture \([Section3\.1](https://arxiv.org/html/2606.13862#S3.SS1)\):The three components of our model that includes a Compressor, a Main module, and a lightweight Multi\-Token Prediction \(MTP\) Module;\(2\) Training \([Section3\.2](https://arxiv.org/html/2606.13862#S3.SS2)\):A two\-stage protocol that first aligns the compressed latent space via distillation, then jointly trains all components; and\(3\) Adaptive Inference \([Section3\.3](https://arxiv.org/html/2606.13862#S3.SS3)\):A decoding algorithm that dynamically falls back to standard single\-token generation when model confidence is low\.

### 3\.1SuperThoughts Architecture

Our model processes reasoning chains by operating on superposed token pairs during the thinking process rather than individual CoT tokens\. We structure each example as a sequence

q1:Lq​<think\>⏟prompt tokens​c1:Lc​</think\>​a1:La⏟response tokens,\\underbrace\{q\_\{1:L\_\{q\}\}\\ \\texttt\{<think\>\}\}\_\{\\text\{prompt tokens\}\}\\ \\underbrace\{c\_\{1:L\_\{c\}\}\\ \\texttt\{</think\>\}\\ a\_\{1:L\_\{a\}\}\}\_\{\\text\{response tokens\}\},whereq1:Lqq\_\{1:L\_\{q\}\}denotes question tokens,c1:Lcc\_\{1:L\_\{c\}\}denotes CoT tokens, anda1:Laa\_\{1:L\_\{a\}\}denotes answer tokens\. The sequence includes special delimiter tokens<think\>and</think\>, with<think\>appended to the prompt to initiate reasoning\. We reorganize the CoT sequencec1:Lcc\_\{1:L\_\{c\}\}into a sequence of pairs, reducing the effective reasoning length fromLcL\_\{c\}toS=Lc/2S=L\_\{c\}/2steps, where we assumeLcL\_\{c\}is even andSSrepresents the number of superposition steps; ifLcL\_\{c\}is odd, we pad with a special token to maintain the pair structure\.

Our model consists of three components: \(1\) Compressor, \(2\) Main module and \(3\) Multi\-Token Prediction \(MTP\) Module\. At CoT phase, for each stepii, the Compressor encodes a token pair\(c2​i−1,c2​i\)\(c\_\{2i\-1\},c\_\{2i\}\)into a single latent vector, from which the Main module predicts the next tokenc2​i\+1c\_\{2i\+1\}and the MTP module predicts the next\-next tokenc2​i\+2c\_\{2i\+2\}\.

#### Compressor\.

LetEmb​\(⋅\)∈ℝH\\texttt\{Emb\}\(\\cdot\)\\in\\mathbb\{R\}^\{H\}denote the token embedding function\. For each stepi=1,…,Si=1,\\ldots,S, the compressorComp​\(⋅\)\\texttt\{Comp\}\(\\cdot\)maps the pair\(Emb​\(c2​i−1\),Emb​\(c2​i\)\)\(\\texttt\{Emb\}\(c\_\{2i\-1\}\),\\texttt\{Emb\}\(c\_\{2i\}\)\)into a single compressed vector𝒙i∈ℝH\{\\bm\{x\}\}\_\{i\}\\in\\mathbb\{R\}^\{H\}\. We explore two implementations for the compressor: either aLinear Projection, where we concatenate the token embeddings and project them using a learnable matrixP∈ℝH×2​HP\\in\\mathbb\{R\}^\{H\\times 2H\}:

𝒙i=P​\[Emb​\(c2​i−1\)Emb​\(c2​i\)\],\{\\bm\{x\}\}\_\{i\}=P\\begin\{bmatrix\}\\texttt\{Emb\}\(c\_\{2i\-1\}\)\\\\ \\texttt\{Emb\}\(c\_\{2i\}\)\\end\{bmatrix\},or aTransformer Block, where we process the pair using a small Transformer layer and extract the hidden state corresponding to the second token \(denoted by the subscript22\):

𝒙i=TF​\(Emb​\(c2​i−1\),Emb​\(c2​i\)\)2,\{\\bm\{x\}\}\_\{i\}=\\texttt\{TF\}\\big\(\\texttt\{Emb\}\(c\_\{2i\-1\}\),\\texttt\{Emb\}\(c\_\{2i\}\)\\big\)\_\{2\},where\[⋅,⋅\]\[\\cdot,\\cdot\]denotes sequence concatenation\. These compressed vectors𝒙1:S\{\\bm\{x\}\}\_\{1:S\}serve as the inputs to the Main module\.

#### Main module\.

The Main module \(base LLM\) is the primary reasoning backbone, responsible for evolving the latent reasoning state and predictingodd\-indexedtokens\. At each time step, it takes as input the compressed representation of two tokens, and outputs a latent reasoning state, which is fed to a language modeling head to predict the next token\. In more detail, stepi=0i=0, the Main module takes in<think\>, produces𝒉0\{\\bm\{h\}\}\_\{0\}and predictsc1c\_\{1\}; at stepsi≥1i\\geq 1, it takes in𝒙i\{\\bm\{x\}\}\_\{i\}, produces the hidden state𝒉i\{\\bm\{h\}\}\_\{i\}and predictsc2​i\+1c\_\{2i\+1\}\. At the final stepi=Si=S, it predicts the closing delimiter</think\>from𝒉S\{\\bm\{h\}\}\_\{S\}\. The Main and MTP modules share the same output language modeling head\.

#### MTP module\.

The MTP module is responsible for predictingeven\-indexedtokens\. At each step, it takes as input the hidden representation output from Main module, some of the past tokens, and predicts next to next token\. In more detail, at stepii, it takes in the previous even tokenc2​ic\_\{2i\}\(withc0=<PAD\>c\_\{0\}=\\texttt\{<PAD\>\}\), the current odd tokenc2​i\+1c\_\{2i\+1\}just predicted by Main and the hidden state𝒉i\{\\bm\{h\}\}\_\{i\}from Main\. These are compressed into𝒙i′∈ℝH\{\\bm\{x\}\}^\{\\prime\}\_\{i\}\\in\\mathbb\{R\}^\{H\}via a learnable projectionP′∈ℝH×3​HP^\{\\prime\}\\in\\mathbb\{R\}^\{H\\times 3H\}:

𝒙i′=P′​\[RMSNorm​\(Emb​\(c2​i\)\)RMSNorm​\(Emb​\(c2​i\+1\)\)RMSNorm​\(𝒉i\)\]\.\{\\bm\{x\}\}^\{\\prime\}\_\{i\}=P^\{\\prime\}\\begin\{bmatrix\}\\texttt\{RMSNorm\}\(\\texttt\{Emb\}\(c\_\{2i\}\)\)\\\\ \\texttt\{RMSNorm\}\(\\texttt\{Emb\}\(c\_\{2i\+1\}\)\)\\\\ \\texttt\{RMSNorm\}\(\{\\bm\{h\}\}\_\{i\}\)\\end\{bmatrix\}\.A one\-layer transformer then produces𝒉i′\{\\bm\{h\}\}^\{\\prime\}\_\{i\}and predictsc2​i\+2c\_\{2i\+2\}\.

### 3\.2Training Strategy

Training a model to reason in latent space can be unstable if the compressed representations drift significantly from the pre\-trained language manifold\. To mitigate this, we employ a two\-stage training protocol: first, we warm\-start the compressor via knowledge distillation to align the latent space; second, we jointly train the entire model using standard cross\-entropy loss\.

#### Stage 1: Training Compressor module via latent distillation\.

![Refer to caption](https://arxiv.org/html/2606.13862v1/x4.png)Figure 4:Compressor training via latent distillation\.Top \(Teacher\):The frozen Base LLM processes the full discrete token sequence\.Bottom \(Student\):The same frozen Base LLM receives compressed representations𝒙i\{\\bm\{x\}\}\_\{i\}from the Compressor, which fuses each CoT token pair \(e\.g\., “It’s” \+ “Currently”→\\to𝒙1\{\\bm\{x\}\}\_\{1\}\)\. The Compressor is trained to minimize the Smooth\-L1L\_\{1\}distance between teacher and student hidden states across all layers at corresponding positions\. Positions used to compute the distillation loss are marked in red\.Before end\-to\-end training, we train the Compressor module via distillation, followingBertonet al\.\([2025](https://arxiv.org/html/2606.13862#bib.bib38)\)\. Consider the Main module processing the discrete sequence\[c1,c2,…,c2​S\]\[c\_\{1\},c\_\{2\},\\ldots,c\_\{2S\}\]one token at a time \(theteacher\), versus the Main module processing compressed pairs\[𝒙1,…,𝒙S\]\[\{\\bm\{x\}\}\_\{1\},\\ldots,\{\\bm\{x\}\}\_\{S\}\]where𝒙i=Comp​\(c2​i−1,c2​i\)\{\\bm\{x\}\}\_\{i\}=\\texttt\{Comp\}\(c\_\{2i\-1\},c\_\{2i\}\)\(thestudent\)\. We train the Compressor so that the student’s hidden state after𝒙i\{\\bm\{x\}\}\_\{i\}matches the teacher’s hidden state \(layer\-wise\) afterc2​ic\_\{2i\}\. In effect, the model should produce the same hidden states whether processing tokens discretely or in compressed form\.

We define the set of distillation targetsDDas pairs of corresponding positions in the teacher \(uncompressed\) and student \(compressed\) sequences\. FollowingBertonet al\.\([2025](https://arxiv.org/html/2606.13862#bib.bib38)\), we include all answer token positions\. Crucially, we extend their approach by also including allevenCoT token positions, enforcing alignment within the reasoning chain itself and not just the final answer\.

The loss is computed as the Smooth\-L1L\_\{1\}distance between the teacher’s hidden statesH\(ℓ\)H^\{\(\\ell\)\}and the student’s hidden statesH~\(ℓ\)\\tilde\{H\}^\{\(\\ell\)\}across all layersℓ\\elland target pairs\(t,t′\)∈D\(t,t^\{\\prime\}\)\\in D:

ℒdistill=∑ℓ1σ\(ℓ\)​\|D\|​∑\(t,t′\)∈DSmoothL1β⁡\(Ht\(ℓ\),H~t′\(ℓ\)\),\\mathcal\{L\}\_\{\\text\{distill\}\}=\\sum\_\{\\ell\}\\frac\{1\}\{\\sigma^\{\(\\ell\)\}\|D\|\}\\sum\_\{\(t,t^\{\\prime\}\)\\in D\}\\operatorname\{SmoothL1\}\_\{\\beta\}\\bigl\(\{H\}^\{\(\\ell\)\}\_\{t\},\\,\\tilde\{H\}^\{\(\\ell\)\}\_\{t^\{\\prime\}\}\\bigr\),whereσ\(ℓ\)=Std⁡\(HD\(ℓ\)\)\\sigma^\{\(\\ell\)\}=\\operatorname\{Std\}\(H^\{\(\\ell\)\}\_\{D\}\)is layer\-wise normalization and the Smooth\-L1L\_\{1\}distanceSmoothL1β⁡\(u,v\)\\operatorname\{SmoothL1\}\_\{\\beta\}\(u,v\)is defined as1d∑i=0dSmoothL1β\(u,v\)i\\frac\{1\}\{d\}\\sum^\{d\}\_\{i=0\}\\operatorname\{SmoothL1\}\_\{\\beta\}\(u,v\)\_\{i\}with

SmoothL1β\(u,v\)i=\{12​\(ui−vi\)2β,\|ui−vi\|<β,\|ui−vi\|−β2,otherwise\.\\operatorname\{SmoothL1\}\_\{\\beta\}\(u,v\)\_\{i\}=\\begin\{cases\}\\dfrac\{1\}\{2\}\\dfrac\{\(u\_\{i\}\-v\_\{i\}\)^\{2\}\}\{\\beta\},&\|u\_\{i\}\-v\_\{i\}\|<\\beta,\\\\\[6\.0pt\] \|u\_\{i\}\-v\_\{i\}\|\-\\dfrac\{\\beta\}\{2\},&\\text\{otherwise\}\.\\end\{cases\}

#### Stage 2: Joint training with cross\-entropy loss\.

Once the compressor is aligned, we train the full model \(Main \+ MTP \+ Compressor\) end\-to\-end\. This essentially minimizes the cross entropy loss for all tokens predicted \(including those coming from the Main module and the MTP module\)\. Letℓ​\(y∣𝒉\):=CE​\(y,head​\(𝒉\)\)\\ell\(y\\mid\{\\bm\{h\}\}\):=\\mathrm\{CE\}\\\!\\bigl\(y,\\mathrm\{head\}\(\{\\bm\{h\}\}\)\\bigr\)denote token\-level cross\-entropy \(CE\) loss\. Define the Main targetsyimain=c2​i\+1y^\{\\mathrm\{main\}\}\_\{i\}=c\_\{2i\+1\}fori=0,1,…,S−1i=0,1,\\ldots,S\-1, andySmain=</think\>y^\{\\mathrm\{main\}\}\_\{S\}=\\texttt\{</think\>\}\. The CoT losses are

ℒNTPCoT\\displaystyle\\mathcal\{L\}^\{\\mathrm\{CoT\}\}\_\{\\mathrm\{NTP\}\}=1S\+1​∑i=0Sℓ​\(yimain∣𝒉i\),\\displaystyle=\\frac\{1\}\{S\+1\}\\sum\_\{i=0\}^\{S\}\\ell\\\!\\left\(y^\{\\mathrm\{main\}\}\_\{i\}\\mid\{\\bm\{h\}\}\_\{i\}\\right\),ℒMTPCoT\\displaystyle\\mathcal\{L\}^\{\\mathrm\{CoT\}\}\_\{\\mathrm\{MTP\}\}=1S​∑i=0S−1ℓ​\(c2​i\+2∣𝒉i′\)\.\\displaystyle=\\frac\{1\}\{S\}\\sum\_\{i=0\}^\{S\-1\}\\ell\\\!\\left\(c\_\{2i\+2\}\\mid\{\\bm\{h\}\}^\{\\prime\}\_\{i\}\\right\)\.Letℒanswer\\mathcal\{L\}\_\{\\mathrm\{answer\}\}denote the standard next\-token CE loss on the answer tokens\. We define

ℒNTP:=ℒanswer\+ℒNTPCoT,ℒTraining=ℒNTP\+λ​ℒMTPCoT\.\\mathcal\{L\}\_\{\\mathrm\{NTP\}\}:=\\mathcal\{L\}\_\{\\mathrm\{answer\}\}\+\\mathcal\{L\}^\{\\mathrm\{CoT\}\}\_\{\\mathrm\{NTP\}\},\\mathcal\{L\}\_\{\\mathrm\{Training\}\}=\\mathcal\{L\}\_\{\\mathrm\{NTP\}\}\+\\lambda\\,\\mathcal\{L\}^\{\\mathrm\{CoT\}\}\_\{\\mathrm\{MTP\}\}\.
Since we train the MTP module from scratch, we first freeze the Main and Compressor module to train the MTP module, after which we unfreeze all modules and jointly train them\.

### 3\.3Confidence\-based Adaptive Inference

While our model can process two CoT tokens within one step with superposed tokens, this sometimes can degrade the performance\. For example, if the model needs to decode two “hard” tokens, a single superposition step may lack sufficient computational capacity to predict both correctly\. Ideally, the model should allocate more compute to “hard tokens” via discrete reasoning, while processing “easy tokens” efficiently with superposition\.

We implement this by looking at the confidence of the MTP module\. At each stepii, after the Main module predicts the odd tokenc2​i\+1c\_\{2i\+1\}, the MTP module produces𝒉i′\{\\bm\{h\}\}^\{\\prime\}\_\{i\}that will be used to predictc2​i\+2c\_\{2i\+2\}\. Let

piMTP=maxj∈V⁡softmax​\(head​\(𝒉i′\)\)j\\displaystyle p\_\{i\}^\{\\text\{MTP\}\}=\\max\_\{j\\in V\}\\text\{softmax\}\(\\text\{head\}\(\{\\bm\{h\}\}^\{\\prime\}\_\{i\}\)\)\_\{j\}be the maximum probability of the MTP prediction at stepiiandτ\\taube a threshold\. IfpiMTP<τp\_\{i\}^\{\\text\{MTP\}\}<\\tau, this means the MTP module is not confident about the prediction\. In this case, we reject the MTP prediction\. On the next stepi\+1i\+1, instead of feeding a compressed pair to the Main module, we inputEmb​\(c2​i\+1\)\\texttt\{Emb\}\(c\_\{2i\+1\}\)directly and re\-predictc2​i\+2c\_\{2i\+2\}using the more powerful Main module, after which MTP tries to predictc2​i\+3c\_\{2i\+3\}followed by the same acceptance check\. This fallback mechanism allows the model to self\-regulate its speed, processing two tokens for easy text while slowing down for difficult reasoning steps\. We analyze the inference cost in Appendix[B](https://arxiv.org/html/2606.13862#A2)\.

## 4Experiments

Table 1:Accuracy and average correct CoT length of Qwen\-2\.5\-Math\-Instruct\-1\.5B/7B models on three benchmarks\. We trained two variants of SuperThoughts model, one with a projection matrix as the Compressor and another with a 1\-layer Transformer as the Compressor\. The baseline CoT is trained on the same dataset as we train SuperThoughts\.### 4\.1Experimental Setup

In this section, we introduce our experiment to train a model that reasons in superposition\. We use Qwen2\.5\-Math\-1\.5B\-Instruct and Qwen2\.5\-Math\-7B\-Instruct as the model we start from and post\-train it to reason in superposition\. The Main module is initialized the same as Qwen2\.5 and the MTP module is initialized using the weights of the last layer\. The projection matrices are initialized as a map that averages the input embeddings\.

To train the model, we curate a synthetic reasoning dataset\. We collect questions fromAlbalaket al\.\([2025](https://arxiv.org/html/2606.13862#bib.bib14)\)andMoshkovet al\.\([2025](https://arxiv.org/html/2606.13862#bib.bib15)\)\. For each question, we let Qwen2\.5 to generate 10 responses with CoT\. We filter out incorrect responses and further filter down to∼\\sim1\.5M responses\. For each sample, we consider the whole response as the CoT part and the “\\boxed\{…\}” part as the answer part\. The dataset preparation process is detailed in Appendix[C](https://arxiv.org/html/2606.13862#A3)\.

The Main module is initialized the same as the pretrained Qwen2\.5 model\. In stage 1 training, we choose a projection matrix or a 1\-layer transformer as two variants of the Compressor module and uses learning rate1×10−41\\times 10^\{\-4\}\.

In stage 2 training, we first freeze the Compressor and the Main module and only train MTP with learning rate5×10−45\\times 10^\{\-4\}\. Then for the joint training, we use learning rate1×10−51\\times 10^\{\-5\}and train for 2 epochs\. For model trained without using adaptive inference, we useλ=1\.0\\lambda=1\.0; for model trained to perform adaptive inference, we useλ=0\.02\\lambda=0\.02\. This is because since uncertain MTP predictions can be rejected and re\-predicted by the Main module in the subsequent step, we prioritize the Main module’s accuracy and thus do not need to optimize the MTP loss as aggressively\. We train Qwen2\.5 on the same dataset as a baseline and discuss our baseline choice in Appendix[A](https://arxiv.org/html/2606.13862#A1)\.

### 4\.2Results

SuperThoughts saves CoT length by 20\-30% with confidence\-based adaptive inference while only having a small accuracy drops \(1\-2 points on most tasks\)\.

Table[1](https://arxiv.org/html/2606.13862#S4.T1)presents our main results on three mathematical reasoning benchmarks: MATH500, OlympiadBench, and AMC23\. We evaluate both compressor variants \(Projection and Transformer\) across different inference settings\.

#### Uniform superposition reduces length but hurts accuracy\.

Without adaptive inference, SuperThoughts reduces CoT length by approximately half but incurs substantial accuracy drops\. On1\.5B\(Projection\), CoT length decreases by4747–54%54\\%across benchmarks, but accuracy drops by14\.714\.7–21\.321\.3points\. On7B\(Projection\), compression rates are similar \(4848–53%53\\%\), but accuracy drops are notably smaller at5\.65\.6–12\.112\.1points\. This*scale effect*suggests that larger models better tolerate aggressive token compression\. Both compressor variants exhibit similar behavior, indicating the bottleneck is per\-step compute capacity rather than the compression mechanism\.

#### Adaptive inference recovers accuracy\.

Confidence\-based adaptive inference substantially closes the accuracy gap while retaining meaningful CoT length reductions\. On1\.5B\(Projection\) withτ=0\.999\\tau=0\.999, MATH500 accuracy matches the baseline \(73\.0%73\.0\\%vs\.72\.4%72\.4\\%\) with a36%36\\%CoT reduction\. OlympiadBench and AMC23 remain within0\.90\.9–1\.61\.6points of baseline while achieving2929–30%30\\%reductions\. On7B\(Projection\),τ=0\.9999\\tau=0\.9999offers a balanced trade\-off:3030–34%34\\%CoT reduction with accuracy within0\.90\.9–2\.22\.2points across all benchmarks\.

Notably, while largerτ\\tausometimes achieves best accuracy, increasingτ\\tau*does not always improve accuracy*\. On1\.5B\(Projection\) MATH500, accuracy is the highest atτ=0\.999\\tau=0\.999\. Similarly, on7B\(Projection\) AMC23, the best adaptive accuracy occurs atτ=0\.9999\\tau=0\.9999\. We hypothesize this reflects noise or thatτ=0\.999\\tau=0\.999already provides a sufficiently high confidence threshold\.

#### Compressor comparison\.

The Projection and Transformer compressors perform similarly\. Without adaptive inference, Projection shows a slight edge \(e\.g\.,30\.7%30\.7\\%vs\.27\.9%27\.9\\%on 7B OlympiadBench\), but this gap disappears with adaptive decoding\. Given its simplicity and lower computational cost, the linear projection is the preferred choice\.

#### Comparison with HAMburger\.

We train HAMburger\(Liu and Zhang,[2025](https://arxiv.org/html/2606.13862#bib.bib39)\)using the same data on Qwen2\.5\-1\.5B and compare against SuperThoughts\. We choose confidence∈\{0\.93,0\.95,0\.99,0\.995,0\.999,0\.9995,0\.9999,1\.0\}\\in\\\{0\.93,0\.95,0\.99,0\.995,0\.999,0\.9995,0\.9999,1\.0\\\}and for SuperThoughts we chooseτ∈\{0\.99,0\.993,0\.995,0\.999,0\.9999,0\.99999\}\\tau\\in\\\{0\.99,0\.993,0\.995,0\.999,0\.9999,0\.99999\\\}\. For each configuration we record the CoT length and the accuracy in FigureLABEL:fig:hamburger\_comp\. SuperThoughts achieves higher accuracy at every compression level, and shorter CoT at every accuracy level, than HAMburger\.

#### Beyond mathematical reasoning\.

Table 2:Accuracy and average correct CoT length of Qwen\-2\.5\-Instruct\-14B trained models222Note that this is a non\-Math model so the accuracies on Math benchmarks are lower than the 7B\-Math model in Table[1](https://arxiv.org/html/2606.13862#S4.T1); for the Projection compressor we add an additional RMSNorm after the compressor\.on four benchmarks\.To test whetherSuperThoughtsapplies to reasoning capabilities beyond Math, we add additional science questions fromGuhaet al\.\([2025](https://arxiv.org/html/2606.13862#bib.bib54)\)and train Qwen2\.5\-14B\-Instruct \(a non\-Math Instruct model\) following the same training paradigm\. Table[2](https://arxiv.org/html/2606.13862#footnote2)shows thatSuperThoughtsworks on domains other than Math\.

The gap between theoretical speedup \(CoT length reduction\) and the actual speedup \(wall\-clock time reduction\) is smaller as the model gets larger\.

#### Inference wall\-clock time analysis\.

We run additional experiment to measure the theoretical speedup \(generation length reduction\) vs actual speedup \(generation time reduction\)\. We implement SuperThoughts using nano\-vLLM\(GeeeekExplorer,[2025](https://arxiv.org/html/2606.13862#bib.bib59)\)for fast inference\. We run MATH500 \(500 questions\) on 1\.5B, 7B and 14B model, and for each model we chooseτ∈\{0\.999,0\.9995,0\.9999,0\.99995,0\.99999\}\\tau\\in\\\{0\.999,0\.9995,0\.9999,0\.99995,0\.99999\\\}\. For each generation configuration, we record generation length reductionRlenR\_\{\\text\{len\}\}and wallclock generation time reductionRtimeR\_\{\\text\{time\}\}:

Rlen=Lbaseline−LSuperThoughtsLbaseline,R\_\{\\text\{len\}\}=\\frac\{L\_\{\\text\{baseline\}\}\-L\_\{\\text\{SuperThoughts\}\}\}\{L\_\{\\text\{baseline\}\}\},Rtime=Tbaseline−TSuperThoughtsTbaseline,R\_\{\\text\{time\}\}=\\frac\{T\_\{\\text\{baseline\}\}\-T\_\{\\text\{SuperThoughts\}\}\}\{T\_\{\\text\{baseline\}\}\},whereLbaselineL\_\{\\text\{baseline\}\}is the baseline average generation length,LSuperThoughtsL\_\{\\text\{SuperThoughts\}\}is the SuperThoughts average generation length,TbaselineT\_\{\\text\{baseline\}\}is the baseline generation time andTSuperThoughtsT\_\{\\text\{SuperThoughts\}\}is the SuperThoughts generation time\. We plotRlenR\_\{\\text\{len\}\}versusRtimeR\_\{\\text\{time\}\}in Figure[5](https://arxiv.org/html/2606.13862#S4.F5)\. We can see that the additional overhead from SuperThoughts \(compressor, MTP and adaptive fallback\) has less influence as the model gets larger\. For example, with projection compressor, on 1\.5B model, 30\.25% reduction in generation length gives 21\.25% reduction in generation time; on 7B model, 33\.17% reduction in generation length gives 25\.72% reduction in generation time; on 14B model, 32\.75% reduction in generation length gives 28\.3% reduction in generation time\. This is because while the compressor, MTP, and adaptive fallback adds additional overhead \(e\.g\., additional kernel launches and other CPU activities\), such overhead takes a smaller percentage of the total generation time as the model gets larger\.

![Refer to caption](https://arxiv.org/html/2606.13862v1/figures/genlen_vs_gentime.png)Figure 5:Generation length reduction vs\. wall\-clock time inference speedup on 1\.5B, 7B and 14B models\.

## 5Discussion

Our results show that SuperThoughts successfully compresses CoT reasoning while preserving accuracy: with adaptive inference, we achieve20−35%20\-35\\%length reduction within points below baseline across all benchmarks\.

The adaptive mechanism also speaks to a broader question:*how should models allocate compute across reasoning steps?*Standard decoding spends the same compute on every token, yet intuitively some steps can be harder than other steps\. Our confidence\-based adaptive mechanism can be viewed as a simple compute scheduler: superpose when the MTP module is confident, fall back to discrete tokens when it is not fully confident\.

We also observe that larger models tolerate aggressive compression better\. Under uniform superposition \(no adaptive fallback\), the 7B model drops5−125\-12points while the 1\.5B model drops14−2114\-21points\. This suggests larger models have enough spare capacity per step to fit two tokens more reliably\. If the trend holds at greater scale, even larger models might tolerate superposing three or more tokens, or need fewer fallbacks to discrete decoding\.

Finally, we train the MTP module from scratch because Qwen2\.5 lacks a native MTP module\. Recent models like Qwen3\-Next\(Team,[2025](https://arxiv.org/html/2606.13862#bib.bib51)\)and and MiMo\(Xiaomiet al\.,[2025](https://arxiv.org/html/2606.13862#bib.bib52)\)are pre\-trained with native MTP modules\. Starting from a native MTP module would simplify training and likely improve results, since the MTP module is already aligned with the backbone\. As native MTP becomes standard, SuperThoughts becomes easier to apply\.

## References

- L1: controlling how long a reasoning model thinks with reinforcement learning\.InSecond Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=4jdIxXBNve)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px3.p1.1)\.
- K\. Ahn, A\. Lamb, and J\. Langford \(2025\)Efficient joint prediction of multiple future tokens\.arXiv preprint arXiv:2503\.21801\.Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px5.p1.1)\.
- A\. Albalak, D\. Phung, N\. Lile, R\. Rafailov, K\. Gandhi, L\. Castricato, A\. Singh, C\. Blagden, V\. Xiang, D\. Mahan, and N\. Haber \(2025\)Big\-math: a large\-scale, high\-quality math dataset for reinforcement learning in language models\.External Links:2502\.17387,[Link](https://arxiv.org/abs/2502.17387)Cited by:[§4\.1](https://arxiv.org/html/2606.13862#S4.SS1.p2.1)\.
- G\. Bachmann and V\. Nagarajan \(2024\)The pitfalls of next\-token prediction\.InProceedings of the 41st International Conference on Machine Learning,R\. Salakhutdinov, Z\. Kolter, K\. Heller, A\. Weller, N\. Oliver, J\. Scarlett, and F\. Berkenkamp \(Eds\.\),Proceedings of Machine Learning Research, Vol\.235,pp\. 2296–2318\.External Links:[Link](https://proceedings.mlr.press/v235/bachmann24a.html)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px5.p1.1)\.
- G\. Berton, J\. Unnikrishnan, S\. Tran, and M\. Shah \(2025\)CompLLM: compression for long context q&a\.External Links:2509\.19228,[Link](https://arxiv.org/abs/2509.19228)Cited by:[item 2](https://arxiv.org/html/2606.13862#S1.I1.i2.p1.1),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2606.13862#S3.SS2.SSS0.Px1.p1.5),[§3\.2](https://arxiv.org/html/2606.13862#S3.SS2.SSS0.Px1.p2.1)\.
- T\. Cai, Y\. Li, Z\. Geng, H\. Peng, J\. D\. Lee, D\. Chen, and T\. Dao \(2024\)MEDUSA: simple llm inference acceleration framework with multiple decoding heads\.InProceedings of the 41st International Conference on Machine Learning,ICML’24\.Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px5.p2.1)\.
- J\. Cheng and B\. Van Durme \(2024\)Compressed chain of thought: efficient reasoning through dense representations\.arXiv preprint arXiv:2412\.13171\.Cited by:[§1](https://arxiv.org/html/2606.13862#S1.p3.1),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Deng, L\. Pang, Z\. Wei, S\. Xu, Z\. Duan, K\. Xu, Y\. Song, H\. Shen, and X\. Cheng \(2025\)Latent reasoning in llms as a vocabulary\-space superposition\.External Links:2510\.15522,[Link](https://arxiv.org/abs/2510.15522)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- Y\. Feldman and Y\. Artzi \(2025\)Simple context compression: mean\-pooling and multi\-ratio training\.arXiv preprint arXiv:2510\.20797\.Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px2.p1.1)\.
- GeeeekExplorer \(2025\)Nano\-vllm\.GitHub\.Note:[https://github\.com/GeeeekExplorer/nano\-vllm](https://github.com/GeeeekExplorer/nano-vllm)Cited by:[§4\.2](https://arxiv.org/html/2606.13862#S4.SS2.SSS0.Px6.p1.3)\.
- A\. Giannou, L\. Yang, K\. Lee, R\. D\. Nowak, and D\. Papailiopoulos \(2025\)Stoic reasoner: dual\-mode transformers that compress to think and decompress to speak\.InThe 5th Workshop on Mathematical Reasoning and AI at NeurIPS 2025,External Links:[Link](https://openreview.net/forum?id=2RTdyYfa0v)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- F\. Gloeckle, B\. Y\. Idrissi, B\. Rozière, D\. Lopez\-Paz, and G\. Synnaeve \(2024\)Better & faster large language models via multi\-token prediction\.InProceedings of the 41st International Conference on Machine Learning,ICML’24\.Cited by:[§1](https://arxiv.org/html/2606.13862#S1.p6.1),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px5.p1.1),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px5.p2.1)\.
- H\. A\. Gozeten, M\. E\. Ildiz, X\. Zhang, H\. Harutyunyan, A\. S\. Rawat, and S\. Oymak \(2026\)Continuous chain of thought enables parallel exploration and reasoning\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=sTPKDKn5ig)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- E\. Guha, R\. Marten, S\. Keh, N\. Raoof, G\. Smyrnis, H\. Bansal, M\. Nezhurina, J\. Mercat, T\. Vu, Z\. Sprague, A\. Suvarna, B\. Feuer, L\. Chen, Z\. Khan, E\. Frankel, S\. Grover, C\. Choi, N\. Muennighoff, S\. Su, W\. Zhao, J\. Yang, S\. Pimpalgaonkar, K\. Sharma, C\. C\. Ji, Y\. Deng, S\. Pratt, V\. Ramanujan, J\. Saad\-Falcon, J\. Li, A\. Dave, A\. Albalak, K\. Arora, B\. Wulfe, C\. Hegde, G\. Durrett, S\. Oh, M\. Bansal, S\. Gabriel, A\. Grover, K\. Chang, V\. Shankar, A\. Gokaslan, M\. A\. Merrill, T\. Hashimoto, Y\. Choi, J\. Jitsev, R\. Heckel, M\. Sathiamoorthy, A\. G\. Dimakis, and L\. Schmidt \(2025\)OpenThoughts: data recipes for reasoning models\.External Links:2506\.04178,[Link](https://arxiv.org/abs/2506.04178)Cited by:[§4\.2](https://arxiv.org/html/2606.13862#S4.SS2.SSS0.Px5.p1.1)\.
- D\. Guo, D\. Yang, H\. Zhang, J\. Song, R\. Zhang, R\. Xu, Q\. Zhu, S\. Ma, P\. Wang, X\. Bi,et al\.\(2025\)Deepseek\-r1: incentivizing reasoning capability in llms via reinforcement learning\.arXiv preprint arXiv:2501\.12948\.Cited by:[§1](https://arxiv.org/html/2606.13862#S1.p1.1),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px6.p1.1)\.
- S\. Hao, S\. Sukhbaatar, D\. Su, X\. Li, Z\. Hu, J\. Weston, and Y\. Tian \(2024\)Training large language models to reason in a continuous latent space\.arXiv preprint arXiv:2412\.06769\.Cited by:[Appendix A](https://arxiv.org/html/2606.13862#A1.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2606.13862#S1.p3.1),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- C\. He, R\. Luo, Y\. Bai, S\. Hu, Z\. L\. Thai, J\. Shen, J\. Hu, X\. Han, Y\. Huang, Y\. Zhang, J\. Liu, L\. Qi, Z\. Liu, and M\. Sun \(2024\)OlympiadBench: a challenging benchmark for promoting agi with olympiad\-level bilingual multimodal scientific problems\.External Links:2402\.14008Cited by:[item 4](https://arxiv.org/html/2606.13862#S1.I1.i4.p1.4)\.
- D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. Steinhardt \(2021\)Measuring mathematical problem solving with the math dataset\.NeurIPS\.Cited by:[item 4](https://arxiv.org/html/2606.13862#S1.I1.i4.p1.4)\.
- S\. Hwang, B\. Wang, and A\. Gu \(2025\)Dynamic chunking for end\-to\-end hierarchical sequence modeling\.External Links:2507\.07955,[Link](https://arxiv.org/abs/2507.07955)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px4.p1.1)\.
- A\. Jaech, A\. Kalai, A\. Lerer, A\. Richardson, A\. El\-Kishky, A\. Low, A\. Helyar, A\. Madry, A\. Beutel, A\. Carney,et al\.\(2024\)Openai o1 system card\.arXiv preprint arXiv:2412\.16720\.Cited by:[§1](https://arxiv.org/html/2606.13862#S1.p1.1),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px6.p1.1)\.
- A\. Jain and B\. Rappazzo \(2025\)Learning to reason with mixture of tokens\.External Links:2509\.21482,[Link](https://arxiv.org/abs/2509.21482)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- H\. Jiang, Q\. Wu, C\. Lin, Y\. Yang, and L\. Qiu \(2023\)LLMLingua: compressing prompts for accelerated inference of large language models\.External Links:2310\.05736,[Link](https://arxiv.org/abs/2310.05736)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px2.p1.1)\.
- W\. Kwon, Z\. Li, S\. Zhuang, Y\. Sheng, L\. Zheng, C\. H\. Yu, J\. Gonzalez, H\. Zhang, and I\. Stoica \(2023\)Efficient memory management for large language model serving with pagedattention\.InProceedings of the 29th Symposium on Operating Systems Principles,SOSP ’23,New York, NY, USA,pp\. 611–626\.External Links:ISBN 9798400702297,[Link](https://doi.org/10.1145/3600006.3613165),[Document](https://dx.doi.org/10.1145/3600006.3613165)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- X\. L\. Li and P\. Liang \(2021\)Prefix\-tuning: optimizing continuous prompts for generation\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),C\. Zong, F\. Xia, W\. Li, and R\. Navigli \(Eds\.\),Online,pp\. 4582–4597\.External Links:[Link](https://aclanthology.org/2021.acl-long.353/),[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.353)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px2.p1.1)\.
- Y\. Li, B\. Dong, F\. Guerin, and C\. Lin \(2023\)Compressing context to enhance inference efficiency of large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 6342–6353\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.391/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.391)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Liu, B\. Feng, B\. Xue, B\. Wang, B\. Wu, C\. Lu, C\. Zhao, C\. Deng, C\. Zhang, C\. Ruan,et al\.\(2024\)Deepseek\-v3 technical report\.arXiv preprint arXiv:2412\.19437\.Cited by:[§1](https://arxiv.org/html/2606.13862#S1.p6.1),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px5.p1.1),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px5.p2.1)\.
- J\. Liu and C\. Zhang \(2025\)HAMburger: accelerating llm inference via token smashing\.External Links:2505\.20438,[Link](https://arxiv.org/abs/2505.20438)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px4.p1.1),[§4\.2](https://arxiv.org/html/2606.13862#S4.SS2.SSS0.Px4.p1.2)\.
- MAA \(2023\)AMC 2023 problems\.External Links:[Link](https://artofproblemsolving.com/wiki/index.php/2023_AMC_12A_Problems)Cited by:[item 4](https://arxiv.org/html/2606.13862#S1.I1.i4.p1.4)\.
- I\. Moshkov, D\. Hanley, I\. Sorokin, S\. Toshniwal, C\. Henkel, B\. Schifferer, W\. Du, and I\. Gitman \(2025\)AIMO\-2 winning solution: building state\-of\-the\-art mathematical reasoning models with openmathreasoning dataset\.arXiv preprint arXiv:2504\.16891\.Cited by:[§4\.1](https://arxiv.org/html/2606.13862#S4.SS1.p2.1)\.
- J\. Mu, X\. L\. Li, and N\. Goodman \(2023\)Learning to compress prompts with gist tokens\.InThirty\-seventh Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=2DtxPCL3T5)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Pagnoni, R\. Pasunuru, P\. Rodriguez, J\. Nguyen, B\. Muller, M\. Li, C\. Zhou, L\. Yu, J\. E\. Weston, L\. Zettlemoyer, G\. Ghosh, M\. Lewis, A\. Holtzman, and S\. Iyer \(2025\)Byte latent transformer: patches scale better than tokens\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 9238–9258\.External Links:[Link](https://aclanthology.org/2025.acl-long.453/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.453),ISBN 979\-8\-89176\-251\-0Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px4.p1.1)\.
- B\. Peng, T\. Gigant, and J\. Quesnelle \(2026\)Efficient pre\-training with token superposition\.arXiv preprint arXiv:2605\.06546\.Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- X\. Qu, S\. Wang, Z\. Huang, K\. Hua, F\. Yin, R\. Zhu, J\. Zhou, Q\. Min, Z\. Wang, Y\. Li, T\. Zhang, H\. Xing, Z\. Zhang, Y\. Song, T\. Zheng, Z\. Zeng, C\. Lin, G\. Zhang, and W\. Huang \(2026\)Dynamic large concept models: latent reasoning in an adaptive semantic space\.External Links:2512\.24617,[Link](https://arxiv.org/abs/2512.24617)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px4.p1.1)\.
- A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, I\. Sutskever,et al\.\(2019\)Language models are unsupervised multitask learners\.OpenAI blog1\(8\),pp\. 9\.Cited by:[Appendix B](https://arxiv.org/html/2606.13862#A2.p2.8),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px5.p1.1)\.
- D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. Bowman \(2024\)GPQA: a graduate\-level google\-proof q&a benchmark\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=Ti67584b98)Cited by:[item 4](https://arxiv.org/html/2606.13862#S1.I1.i4.p1.4)\.
- S\. Z\. Shen, R\. Shao, C\. Wang, S\. Yang, V\. Berges, G\. Ghosh, P\. W\. Koh, L\. Zettlemoyer, Y\. Kim, J\. E\. Weston, D\. Sontag, and W\. Yih \(2025a\)HybridCoT: interleaving latent and text chain\-of\-thought for efficient reasoning\.InNeurIPS 2025 Workshop on Efficient Reasoning,External Links:[Link](https://openreview.net/forum?id=NRGRrHmq1H)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- Z\. Shen, H\. Yan, L\. Zhang, Z\. Hu, Y\. Du, and Y\. He \(2025b\)CODI: compressing chain\-of\-thought into continuous space via self\-distillation\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 677–693\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.36/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.36),ISBN 979\-8\-89176\-332\-6Cited by:[Appendix A](https://arxiv.org/html/2606.13862#A1.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2606.13862#S1.p3.1),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- V\. Shrivastava, A\. Awadallah, V\. Balachandran, S\. Garg, H\. Behl, and D\. Papailiopoulos \(2025\)Sample more to think less: group filtered policy optimization for concise reasoning\.External Links:2508\.09726,[Link](https://arxiv.org/abs/2508.09726)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px3.p1.1)\.
- C\. Snell, J\. Lee, K\. Xu, and A\. Kumar \(2024\)Scaling llm test\-time compute optimally can be more effective than scaling model parameters\.arXiv preprint arXiv:2408\.03314\.Cited by:[§1](https://arxiv.org/html/2606.13862#S1.p1.1),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px6.p1.1)\.
- D\. Su, H\. Zhu, Y\. Xu, J\. Jiao, Y\. Tian, and Q\. Zheng \(2025\)Token assorted: mixing latent and text tokens for improved language model reasoning\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=hYfOPXrbUr)Cited by:[§1](https://arxiv.org/html/2606.13862#S1.p3.1),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- W\. Tan, J\. Li, J\. Ju, Z\. Luo, J\. Luan, and R\. Song \(2025\)Think silently, think fast: dynamic latent compression of llm reasoning chains\.External Links:2505\.16552,[Link](https://arxiv.org/abs/2505.16552)Cited by:[Appendix A](https://arxiv.org/html/2606.13862#A1.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- Y\. Tang, L\. Dong, Y\. Hao, Q\. Dong, F\. Wei, and J\. Gu \(2026\)Multiplex thinking: reasoning via token\-wise branch\-and\-merge\.arXiv preprint arXiv:2601\.08808\.Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- Q\. Team \(2025\)Qwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§5](https://arxiv.org/html/2606.13862#S5.p4.1)\.
- J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. V\. Le, D\. Zhou,et al\.\(2022\)Chain\-of\-thought prompting elicits reasoning in large language models\.Advances in neural information processing systems35,pp\. 24824–24837\.Cited by:[§1](https://arxiv.org/html/2606.13862#S1.p1.1),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Wu, J\. Lu, Z\. Ren, G\. Hu, Z\. Wu, D\. Dai, and H\. Wu \(2025\)LLMs are single\-threaded reasoners: demystifying the working mechanism of soft thinking\.External Links:2508\.03440,[Link](https://arxiv.org/abs/2508.03440)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- H\. Xia, C\. T\. Leong, W\. Wang, Y\. Li, and W\. Li \(2025\)TokenSkip: controllable chain\-of\-thought compression in LLMs\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 3351–3363\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.165/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.165),ISBN 979\-8\-89176\-332\-6Cited by:[Appendix A](https://arxiv.org/html/2606.13862#A1.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px3.p1.1)\.
- L\. Xiaomi, B\. Xia, B\. Shen, D\. Zhu, D\. Zhang, G\. Wang, H\. Zhang, H\. Liu, J\. Xiao, J\. Dong,et al\.\(2025\)MiMo: unlocking the reasoning potential of language model–from pretraining to posttraining\.arXiv preprint arXiv:2505\.07608\.Cited by:[§5](https://arxiv.org/html/2606.13862#S5.p4.1)\.
- Z\. Yue, B\. Jin, H\. Zeng, H\. Zhuang, Z\. Qin, J\. Yoon, L\. Shang, J\. Han, and D\. Wang \(2025\)Hybrid latent reasoning via reinforcement learning\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=LjtgTpWH71)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Zhang, Y\. Zhu, M\. Sun, Y\. Luo, S\. Qiao, L\. Du, D\. Zheng, H\. Chen, and N\. Zhang \(2025a\)LightThinker: thinking step\-by\-step compression\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 13307–13328\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.673/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.673),ISBN 979\-8\-89176\-332\-6Cited by:[§1](https://arxiv.org/html/2606.13862#S1.p3.1),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- Z\. Zhang, X\. He, W\. Yan, A\. Shen, C\. Zhao, S\. Wang, Y\. Shen, and X\. E\. Wang \(2025b\)Soft thinking: unlocking the reasoning potential of llms in continuous concept space\.arXiv preprint arXiv:2505\.15778\.Cited by:[Appendix A](https://arxiv.org/html/2606.13862#A1.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- L\. Zheng, L\. Yin, Z\. Xie, C\. Sun, J\. Huang, C\. H\. Yu, S\. Cao, C\. Kozyrakis, I\. Stoica, J\. E\. Gonzalez, C\. Barrett, and Y\. Sheng \(2024\)SGLang: efficient execution of structured language model programs\.InProceedings of the 38th International Conference on Neural Information Processing Systems,NIPS ’24,Red Hook, NY, USA\.External Links:ISBN 9798331314385Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- H\. Zhu, S\. Hao, Z\. Hu, J\. Jiao, S\. Russell, and Y\. Tian \(2025\)Reasoning by superposition: a theoretical perspective on chain of continuous thought\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=UdOEZgWJLc)Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.
- Y\. Zhuang, L\. Liu, C\. Singh, J\. Shang, and J\. Gao \(2025\)Text generation beyond discrete token sampling\.arXiv preprint arXiv:2505\.14827\.Cited by:[§2](https://arxiv.org/html/2606.13862#S2.SS0.SSS0.Px1.p1.1)\.

## Appendix ADiscussion on Baseline

We report the standard discrete CoT as our baseline choice\. In this section we discuss other related methods and why we didn’t choose them as baselines\.

#### Methods that reduce the discrete tokens\.

Methods like TokenSkip\(Xiaet al\.,[2025](https://arxiv.org/html/2606.13862#bib.bib42)\)cuts the number of discrete CoT tokens by finetuning the model on a more concise reasoning trace\. TokenSkip operates entirely within the discrete token space – it reduces the number of tokens generated but still decodes one token per forward pass\.SuperThoughts, by contrast, reduces the number of forward passes themselves by decoding multiple tokens per step, targeting a different axis of efficiency\. The two approaches are complementary and could in principle be combined\.

#### MTP for self\-speculative decoding\.

MTP module can be used as a draft model to speculatively predict future tokens and verify with the target model\. However, speculative decoding does not reduce total FLOPs – the main model must still perform full forward passes to populate KV cache entries for every accepted token\. It can speed up inference in low\-utilization regimes with small batch sizes by increasing GPU utilization, but when serving large batches at high throughput, the GPU is already compute\-bound and no speed\-up is achieved\. On the other hand,SuperThoughtsreduces total FLOPs by adaptively generating multiple tokens per forward pass, in contrast to standard autoregressive decoding which produces a single token per step\. Therefore, speculative decoding is not a suitable baseline forSuperThoughts, as it targets a fundamentally different bottleneck – latency under low utilization – rather than reducing total FLOPs, which is the focus of our work\.

#### Latent reasoning\.

Latent reasoning methods like COCONUT\(Haoet al\.,[2024](https://arxiv.org/html/2606.13862#bib.bib16)\), CODI\(Shenet al\.,[2025b](https://arxiv.org/html/2606.13862#bib.bib21)\)and CoLaR\(Tanet al\.,[2025](https://arxiv.org/html/2606.13862#bib.bib29)\)directly use the latent from the last layer as the input to the next reasoning step\. However, these methods are designed for tasks and training regimes that involve short reasoning sequences \(e\.g\., 20–60 tokens\), and we find that they do not translate effectively to our settin\. For example we tried to train CODI in our setting and the trained model can hardly produce a correct answer \(accuracy 0\.1%\)\. We conduct experiment with Soft Thinking\(Zhanget al\.,[2025b](https://arxiv.org/html/2606.13862#bib.bib17)\)that aggregates token embeddings from multiple candidates within a single step\. However, Table[3](https://arxiv.org/html/2606.13862#A1.T3)shows that Soft Thinking is not significantly better than Standard CoT \(except 1\.5B on AMC23\) and the CoT length is longer\.

Table 3:Accuracy and average correct CoT length of Qwen\-2\.5\-Math\-Instruct\-1\.5B/7B models on three benchmarks, comparing the Standard CoT baseline against Soft Thinking\.

## Appendix BInference Cost Analysis

In this subsection we analyze how much compute canSuperThoughtssave compared to standard discrete CoT decoding\. For simplicity we assume batch sizeB=1B=1; the analysis extends directly toB\>1B\>1sinceBBcancels in the ratioCST/CbaseC\_\{\\text\{ST\}\}/C\_\{\\text\{base\}\}, andSuperThoughtssupports batched generation as token superposition operates independently within each sequence\. LetLqL\_\{q\}be the question \(prompt\) length,LcL\_\{c\}be the CoT length andLaL\_\{a\}be the response length\. The inference cost can be split into prefilling cost and decoding cost, while at prefilling stage the model process the question part one time and at the decoding stage the model generates CoT and answer autoregressively\. For reasoning intensive tasks like Math, usuallyLc\+La≫LqL\_\{c\}\+L\_\{a\}\\gg L\_\{q\}, meaning that the decoding cost dominates\. Also sinceLc≫LaL\_\{c\}\\gg L\_\{a\}, for simplicity, ignore prompt and answer: letLq=0,La=0L\_\{q\}=0,L\_\{a\}=0andL=LcL=L\_\{c\}be the context length\.

We decompose compute into \(1\)*core attention*\(theQ​K⊤QK^\{\\top\}andAttn⋅V\\text\{Attn\}\\cdot Vmatmuls\), which scales asO​\(L2​H\)O\(L^\{2\}H\), and \(2\)*linear layers*\(FFN \+ QKV/output projections\), which scale asO​\(L​H2\)O\(LH^\{2\}\)\. LetHHbe the embedding dimension andTTbe the number of layers\. For GPT\-2333Here we use GPT\-2 architecture for simplicity\. For other architecture like Qwen, only the constant term changes and the analysis still holds\.\(Radfordet al\.,[2019](https://arxiv.org/html/2606.13862#bib.bib3)\)Transformer block, the core attention cost \(number of multiplications\) is2​T​L2​H2TL^\{2\}Hand the linear layers cost is12​T​L​H212TLH^\{2\}\.

For SuperThoughts, letSSbe the number of superposition steps and letL′L^\{\\prime\}be the CoT length, which includes both \(1\) steps when we superpose tokens and \(2\) steps when we use discrete tokens during adaptive inference\. Note that without adaptive inference,L′=S=L/2L^\{\\prime\}=S=L/2because we superpose at each CoT step\. we break down the inference cost for each module:

- •\(Main module\)Core attention cost:2​T​L′⁣2​H2TL^\{\\prime 2\}H\. Linear layers cost:12​T​L′​H212TL^\{\\prime\}H^\{2\}\.
- •\(MLP module\)Core attention cost:2​L′⁣2​H2L^\{\\prime 2\}H\. Linear layers cost:12​L′​H212L^\{\\prime\}H^\{2\}\. Additional projection cost:3​L′​H23L^\{\\prime\}H^\{2\}\.
- •\(Projection Compressor\)Projection cost:2​S​H22SH^\{2\}\.
- •\(Transformer Compressor\)Transformer cost:8​S​H\+24​S​H28SH\+24SH^\{2\}\.

Note that for the Projection Compressor, the cost is2​S​H22SH^\{2\}not2​L′​H22L^\{\\prime\}H^\{2\}because if we use adaptive inference and when we reject the MTP token, we do not compress tokens at the next CoT step\. Ignoring the additional minor projection cost, a simple mental model could be that SuperThoughts hasL′L\\frac\{L^\{\\prime\}\}\{L\}CoT steps as the standard discrete model but each step costs\(T\+1\)/\(T\)\(T\+1\)/\(T\)compute\. Note that since the core attention cost is quadratic inLLandL′L^\{\\prime\}, this can underestimate the compute saved depending on how large isLL\.

For SuperThoughts with Projection Compressor, ignoring these small projection\-style overheads, a simple mental model is that SuperThoughts runsL′L^\{\\prime\}decoding steps instead ofLL, but each step executes aTT\-layer backbone plus an extra 1\-layer MTP, i\.e\., a per\-step multiplier of\(T\+1\)/T\(T\+1\)/T\. This multiplier applies to both the dense\-matmul term \(∝L​H2\\propto LH^\{2\}\) and the core attention\-mixing term \(∝L2​H\\propto L^\{2\}H\)\. Approximating with the linear term gives

CSTCbase≈L′L⋅T\+1T,\\frac\{C\_\{\\text\{ST\}\}\}\{C\_\{\\text\{base\}\}\}\\approx\\frac\{L^\{\\prime\}\}\{L\}\\cdot\\frac\{T\+1\}\{T\},which typically overestimates this ratio \(and thus underestimates compute savings\), since the core attention term scales asT\+1T​\(L′L\)2\\frac\{T\+1\}\{T\}\\left\(\\frac\{L^\{\\prime\}\}\{L\}\\right\)^\{2\}\.

## Appendix CAlgorithms for Dataset Preparation

Input:Question

qq, answer

aa, chain token ids

𝐜=\(c1,…,cN\)\\mathbf\{c\}=\(c\_\{1\},\\dots,c\_\{N\}\), token probabilities

𝐩=\(p1,…,pN\)\\mathbf\{p\}=\(p\_\{1\},\\dots,p\_\{N\}\), window size

kk, special treatment mode

ℳ∈\{prob,none\}\\mathcal\{M\}\\in\\\{\\texttt\{prob\},\\texttt\{none\}\\\}
Output:NTP tensors

\(𝐱,𝐆,𝐕,𝐲\)\(\\mathbf\{x\},\\mathbf\{G\},\\mathbf\{V\},\\mathbf\{y\}\)and MTP tensors

\(𝐱mtp,𝐆mtp,𝐕mtp,𝐲mtp\)\(\\mathbf\{x\}\_\{\\text\{mtp\}\},\\mathbf\{G\}\_\{\\text\{mtp\}\},\\mathbf\{V\}\_\{\\text\{mtp\}\},\\mathbf\{y\}\_\{\\text\{mtp\}\}\)
1

21ex

//Step 1: Adaptive token selection \(mode\-specific, see Algorithms[2](https://arxiv.org/html/2606.13862#algorithm2)\-\-[3](https://arxiv.org/html/2606.13862#algorithm3)\)

3

𝐬←AdaptiveSelectℳ​\(𝐜,𝐩\)\\mathbf\{s\}\\leftarrow\\textnormal\{\{AdaptiveSelect\}\}\_\{\\mathcal\{M\}\}\(\\mathbf\{c\},\\mathbf\{p\}\);

4

51ex

//Step 2: Window\-aligned padding

6

𝐜~,𝐦←WindowAlignedPad​\(𝐜,𝐬,k,ℳ\)\\tilde\{\\mathbf\{c\}\},\\,\\mathbf\{m\}\\leftarrow\\textnormal\{\{WindowAlignedPad\}\}\(\\mathbf\{c\},\\mathbf\{s\},k,\\mathcal\{M\}\);

7

81ex

//Step 3: Build NTP sequence

9

𝐱,𝐆,𝐕,𝐲←BuildSequence​\(q,a,𝐜~,𝐦,k,false\)\\mathbf\{x\},\\,\\mathbf\{G\},\\,\\mathbf\{V\},\\,\\mathbf\{y\}\\leftarrow\\textnormal\{\{BuildSequence\}\}\(q,a,\\tilde\{\\mathbf\{c\}\},\\mathbf\{m\},k,\\texttt\{false\}\);

10

111ex

//Step 4: Build MTP sequence \(shift chain right byk−1k\-1\)

12

𝐜~mtp←\(cot\_pad,…,cot\_pad⏟k−1,c~1,…,c~\|𝐜~\|\)\\tilde\{\\mathbf\{c\}\}\_\{\\text\{mtp\}\}\\leftarrow\(\\underbrace\{\\textsc\{cot\\\_pad\},\\dots,\\textsc\{cot\\\_pad\}\}\_\{k\-1\},\\,\\tilde\{c\}\_\{1\},\\dots,\\tilde\{c\}\_\{\|\\tilde\{\\mathbf\{c\}\}\|\}\);

13

𝐦mtp←\(true,…,true⏟k−1,m1,…,m\|𝐦\|\)\\mathbf\{m\}\_\{\\text\{mtp\}\}\\leftarrow\(\\underbrace\{\\texttt\{true\},\\dots,\\texttt\{true\}\}\_\{k\-1\},\\,m\_\{1\},\\dots,m\_\{\|\\mathbf\{m\}\|\}\);

14

𝐱mtp,𝐆mtp,𝐕mtp,𝐲mtp←BuildSequence​\(q,a,𝐜~mtp,𝐦mtp,k,true\)\\mathbf\{x\}\_\{\\text\{mtp\}\},\\,\\mathbf\{G\}\_\{\\text\{mtp\}\},\\,\\mathbf\{V\}\_\{\\text\{mtp\}\},\\,\\mathbf\{y\}\_\{\\text\{mtp\}\}\\leftarrow\\textnormal\{\{BuildSequence\}\}\(q,a,\\tilde\{\\mathbf\{c\}\}\_\{\\text\{mtp\}\},\\mathbf\{m\}\_\{\\text\{mtp\}\},k,\\texttt\{true\}\);

15

161exreturn

\(𝐱,𝐆,𝐕,𝐲\),\(𝐱mtp,𝐆mtp,𝐕mtp,𝐲mtp\)\(\\mathbf\{x\},\\mathbf\{G\},\\mathbf\{V\},\\mathbf\{y\}\),\\;\(\\mathbf\{x\}\_\{\\text\{mtp\}\},\\mathbf\{G\}\_\{\\text\{mtp\}\},\\mathbf\{V\}\_\{\\text\{mtp\}\},\\mathbf\{y\}\_\{\\text\{mtp\}\}\);

Algorithm 1SuperThoughts Dataset PreparationInput:Token ids

𝐜=\(c1,…,cN\)\\mathbf\{c\}=\(c\_\{1\},\\dots,c\_\{N\}\), token probabilities

𝐩=\(p1,…,pN\)\\mathbf\{p\}=\(p\_\{1\},\\dots,p\_\{N\}\)
Output:Boolean mask

𝐬=\(s1,…,sN\)\\mathbf\{s\}=\(s\_\{1\},\\dots,s\_\{N\}\)
1

1exSample

α∼𝒰​\[αmin,αmax\]\\alpha\\sim\\mathcal\{U\}\[\\alpha\_\{\\min\},\\,\\alpha\_\{\\max\}\];

//Here0≤αmin<αmax≤10\\leq\\alpha\_\{\\text\{min\}\}<\\alpha\_\{\\text\{max\}\}\\leq 1is a fraction

2

m←⌊α⋅N⌋m\\leftarrow\\lfloor\\alpha\\cdot N\\rfloor;

3

ℐ←\\mathcal\{I\}\\leftarrowindices of the

mmsmallest values in

𝐩\\mathbf\{p\};

4

si←𝟙​\[i∈ℐ\]∀i∈\{1,…,N\}s\_\{i\}\\leftarrow\\mathbb\{1\}\[i\\in\\mathcal\{I\}\]\\quad\\forall\\,i\\in\\\{1,\\dots,N\\\};

5return

𝐬\\mathbf\{s\};

Algorithm 2AdaptiveSelectprob\\textsc\{AdaptiveSelect\}\_\{\\texttt\{prob\}\}: Fraction\-Based Probability Selection1

si≔false∀is\_\{i\}\\coloneqq\\texttt\{false\}\\quad\\forall\\,i;

Algorithm 3AdaptiveSelectnone\\textsc\{AdaptiveSelect\}\_\{\\texttt\{none\}\}: No Selection \(Baseline\)Input:Token ids

𝐜\\mathbf\{c\}, special mask

𝐬\\mathbf\{s\}, window size

kk, mode

ℳ\\mathcal\{M\}
Output:Padded ids

𝐜~\\tilde\{\\mathbf\{c\}\}, padding mask

𝐦\\mathbf\{m\}\(

mi=true⇒m\_\{i\}=\\texttt\{true\}\\Rightarrowinserted pad\)

1

21ex

pad\_right←\(ℳ≠prob\-frac\)\\texttt\{pad\\\_right\}\\leftarrow\(\\mathcal\{M\}\\neq\\texttt\{prob\-frac\}\);

//isolate specials in own window?

𝐜~←\[\],𝐦←\[\],j←0\\tilde\{\\mathbf\{c\}\}\\leftarrow\[\\,\],\\quad\\mathbf\{m\}\\leftarrow\[\\,\],\\quad j\\leftarrow 0;

//

jj: position within current window

3

4for*i←1i\\leftarrow 1to\|𝐜\|\|\\mathbf\{c\}\|*do

5if*sis\_\{i\}*then

6if*j≠0j\\neq 0*then//pad to finish current window

7

𝐜~\+⁣=\[cot\_pad\]k−j,𝐦\+⁣=\[true\]k−j\\tilde\{\\mathbf\{c\}\}\\mathrel\{\{\+\}\{=\}\}\[\\textsc\{cot\\\_pad\}\]^\{k\-j\},\\quad\\mathbf\{m\}\\mathrel\{\{\+\}\{=\}\}\[\\texttt\{true\}\]^\{k\-j\};

8

9

𝐜~\+⁣=ci,𝐦\+⁣=false\\tilde\{\\mathbf\{c\}\}\\mathrel\{\{\+\}\{=\}\}c\_\{i\},\\quad\\mathbf\{m\}\\mathrel\{\{\+\}\{=\}\}\\texttt\{false\};

10

j←1j\\leftarrow 1;

11if*pad\_right*then//fill rest of window

12

𝐜~\+⁣=\[cot\_pad\]k−1,𝐦\+⁣=\[true\]k−1\\tilde\{\\mathbf\{c\}\}\\mathrel\{\{\+\}\{=\}\}\[\\textsc\{cot\\\_pad\}\]^\{k\-1\},\\quad\\mathbf\{m\}\\mathrel\{\{\+\}\{=\}\}\[\\texttt\{true\}\]^\{k\-1\};

13

j←0j\\leftarrow 0;

14

15

16else

17

𝐜~\+⁣=ci,𝐦\+⁣=false\\tilde\{\\mathbf\{c\}\}\\mathrel\{\{\+\}\{=\}\}c\_\{i\},\\quad\\mathbf\{m\}\\mathrel\{\{\+\}\{=\}\}\\texttt\{false\};

18

j←\(j\+1\)modkj\\leftarrow\(j\+1\)\\bmod k;

19

20

21return

𝐜~,𝐦\\tilde\{\\mathbf\{c\}\},\\,\\mathbf\{m\};

Algorithm 4WindowAlignedPad: Window\-Aligned PaddingInput:Question

qq, answer

aa, padded chain

𝐜~\\tilde\{\\mathbf\{c\}\}, padding mask

𝐦\\mathbf\{m\}, window size

kk, flagis\_mtp

Output:Input ids

𝐱\\mathbf\{x\}, grouped CoT inputs

𝐆∈ℤL×k\\mathbf\{G\}\\in\\mathbb\{Z\}^\{L\\times k\}, valid masks

𝐕∈\{0,1\}L×k\\mathbf\{V\}\\in\\\{0,1\\\}^\{L\\times k\}, targets

𝐲\\mathbf\{y\}
1

21ex

//Tokenize with chat template

𝐭←ChatTemplate​\(q,a\)\\mathbf\{t\}\\leftarrow\\textsc\{ChatTemplate\}\(q,a\);

//contains <think\>…\\dots</think\>

3

b←index of<think\>in​𝐭b\\leftarrow\\text\{index of \}\\texttt\{<think\>\}\\text\{ in \}\\mathbf\{t\};

4

e←index of</think\>in​𝐭e\\leftarrow\\text\{index of \}\\texttt\{</think\>\}\\text\{ in \}\\mathbf\{t\};

5

61ex

//Group padded chain intokk\-windows

7

Gl←\(c~\(l−1\)​k\+1,…,c~l​k\)G\_\{l\}\\leftarrow\(\\tilde\{c\}\_\{\(l\-1\)k\+1\},\\dots,\\tilde\{c\}\_\{lk\}\)for

l=1,…,Lcotl=1,\\dots,L\_\{\\text\{cot\}\};

8

Vl←\(¬m\(l−1\)​k\+1,…,¬ml​k\)V\_\{l\}\\leftarrow\(\\lnot\\,m\_\{\(l\-1\)k\+1\},\\dots,\\lnot\\,m\_\{lk\}\)for

l=1,…,Lcotl=1,\\dots,L\_\{\\text\{cot\}\};

9

101ex

//Construct input ids: replace reasoning span with placeholders

11

𝐱←\[𝐭1:b,pad,…,pad⏟Lcot,𝐭e:\|t\|\]\\mathbf\{x\}\\leftarrow\[\\,\\mathbf\{t\}\_\{1:b\},\\;\\underbrace\{\\textsc\{pad\},\\dots,\\textsc\{pad\}\}\_\{L\_\{\\text\{cot\}\}\},\\;\\mathbf\{t\}\_\{e:\|t\|\}\\,\];

12

rstart←b\+1,rend←b\+Lcotr\_\{\\text\{start\}\}\\leftarrow b\+1,\\quad r\_\{\\text\{end\}\}\\leftarrow b\+L\_\{\\text\{cot\}\};

𝐆\[rstart:rend\]←\(G1,…,GLcot\)\\mathbf\{G\}\[r\_\{\\text\{start\}\}:r\_\{\\text\{end\}\}\]\\leftarrow\(G\_\{1\},\\dots,G\_\{L\_\{\\text\{cot\}\}\}\);

//elsewhere filled withpad

𝐕\[rstart:rend\]←\(V1,…,VLcot\)\\mathbf\{V\}\[r\_\{\\text\{start\}\}:r\_\{\\text\{end\}\}\]\\leftarrow\(V\_\{1\},\\dots,V\_\{L\_\{\\text\{cot\}\}\}\);

//elsewhere false

cot\_maski←⋁j=1kVi,j\\texttt\{cot\\\_mask\}\_\{i\}\\leftarrow\\bigvee\_\{j=1\}^\{k\}V\_\{i,j\};

//true if positioniihas any valid CoT token

13

141ex

//Construct targets

15for*i←1i\\leftarrow 1toLL*do

16if**cot\_mask*i\\texttt\{cot\\\_mask\}\_\{i\}*then

𝐲i←\\mathbf\{y\}\_\{i\}\\leftarrowfirst

Gl​\(i\)\+1​\[j\]G\_\{l\(i\)\+1\}\[j\]such that

Gl​\(i\)\+1​\[j\]≠cot\_padG\_\{l\(i\)\+1\}\[j\]\\neq\\textsc\{cot\\\_pad\};

//next window’s first valid token

17

18else if*is\_mtpori≤r*end*i\\leq r\_\{\\text\{end\}\}*then

𝐲i←ignore\\mathbf\{y\}\_\{i\}\\leftarrow\\textsc\{ignore\};

//mask out question region; MTP masks answer too

19

20else

𝐲i←𝐱i\+1\\mathbf\{y\}\_\{i\}\\leftarrow\\mathbf\{x\}\_\{i\+1\};

//standard next\-token prediction for answer

21

22

23if*¬*is\_mtp*\\lnot\\,\\texttt\{is\\\_mtp\}*then

𝐲b←G1​\[1\]\\mathbf\{y\}\_\{b\}\\leftarrow G\_\{1\}\[1\];

//at <think\>, predict first CoT token

24

25

261exreturn

𝐱,𝐆,𝐕,𝐲\\mathbf\{x\},\\,\\mathbf\{G\},\\,\\mathbf\{V\},\\,\\mathbf\{y\};

Algorithm 5BuildSequence: Assemble Input and Target Tensors

Similar Articles

MUX: Continuous Reasoning via Multiplexed Tokens

arXiv cs.AI

MUX proposes a method for lossless continuous reasoning by distilling discrete reasoning steps into multiplexed latent tokens that encode a superposition of subwords, achieving higher bandwidth and enabling parallel exploration in language model reasoning tasks.