Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion
Summary
Orthrus is a dual-architecture framework that combines autoregressive LLMs with diffusion models for fast parallel token generation while maintaining exact inference fidelity via shared KV caches and consensus mechanisms, achieving up to 7.8x speedup.
View Cached Full Text
Cached at: 05/14/26, 04:16 AM
Paper page - Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion
Source: https://huggingface.co/papers/2605.12825
Abstract
Orthrus is a dual-architecture framework that combines autoregressive LLMs with diffusion models to achieve fast parallel token generation while maintaining exact inference fidelity through shared KV caches and consensus mechanisms.
We introduce Orthrus, a simple and efficientdual-architecture frameworkthat unifies the exact generation fidelity ofautoregressive Large Language Models(LLMs) with the high-speedparallel token generationofdiffusion models. The sequential nature of standard autoregressive decoding represents a fundamental bottleneck for high-throughput inference. While diffusion language models attempt to break this barrier via parallel generation, they suffer from significant performance degradation, high training costs, and a lack of rigorous convergence guarantees. Orthrus resolves this dichotomy natively. Designed to seamlessly integrate into existingTransformers, the framework augments a frozen LLM with a lightweight, trainable module to create a parallel diffusion view alongside the standard autoregressive view. In this unified system, both views attend to the exact same high-fidelity Key-Value (KV) cache; the autoregressive head executes context pre-filling to construct accurate KV representations, while the diffusion head executes parallel generation. By employing an exactconsensus mechanismbetween the two views, Orthrus guaranteeslossless inference, delivering up to a 7.8x speedup with only an O(1) memory cache overhead and minimal parameter additions.
View arXiv pageView PDFGitHub1Add to collection
Get this paper in your agent:
hf papers read 2605\.12825
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper3
#### chiennv/Orthrus-Qwen3-8B Text Generation• 10B• Updatedabout 2 hours ago
#### chiennv/Orthrus-Qwen3-4B Text Generation• 5B• Updatedabout 2 hours ago • 20
#### chiennv/Orthrus-Qwen3-1.7B Text Generation• 2B• Updatedabout 2 hours ago
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.12825 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.12825 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
Orthrus-Qwen3: up to 7.8×tokens/forward on Qwen3, identical output distribution
Orthrus is a dual-architecture framework that combines autoregressive LLM fidelity with diffusion model speed, delivering up to 7.8x speedup on Qwen3 models while guaranteeing identical output distribution.
Orthrus-Qwen3-8B : up to 7.8×tokens/forward on Qwen3-8B, frozen backbone, provably identical output distribution
Introduces Orthrus, a method that injects a trainable diffusion attention module into a frozen autoregressive transformer to achieve up to 7.8× tokens per forward pass and ~6× wall-clock speedup on MATH-500, with provably identical output distribution to the base Qwen3-8B model. The approach requires minimal additional parameters and training, and avoids the TTFT penalty of external drafters.
@alec_helbling: Diffusion LMs generate multiple tokens in parallel. However, iterative unmasking repeatedly updates token states, limit…
The post describes how diffusion language models face KV-cache reuse limitations due to iterative unmasking, and introduces Block Diffusion as a solution that decodes blocks left-to-right for efficient caching while generating in parallel.
Set Diffusion: Interpolating Token Orderings Between Autoregression and Diffusion for Fast and Flexible Decoding
Set Diffusion introduces a new class of language models that interpolates between autoregressive and diffusion models by factorizing token generation over flexible-position, flexible-length token sets. This enables faster decoding and flexible token ordering, achieving better speed-quality tradeoffs on reasoning, summarization, and unconditional generation tasks.
DiffRetriever: Parallel Representative Tokens for Retrieval with Diffusion Language Models
This paper introduces DiffRetriever, a method that uses diffusion language models to generate multiple representative tokens in parallel for efficient information retrieval, outperforming autoregressive baselines in speed and accuracy.