parallel-decoding

Tag

Cards List
#parallel-decoding

FastGuide: Accelerating Reward Guidance for Diffusion Large Language Models

arXiv cs.CL ↗ · 19h ago Cached

FastGuide 是一种自适应的并行与自回归混合解码方法,通过复用奖励模型反向传播计算、利用 KV 缓存和稀疏注意力重计算来加速扩散大语言模型的梯度奖励引导,在三个奖励基准上实现最高 4.4× 加速同时保持相近生成质量。

0 favorites 0 likes
#parallel-decoding

Less Uniform Discrete Diffusion is More Powerful and Scalable

arXiv cs.CL ↗ · 19h ago Cached

This paper proposes LUDI, a less uniform diffusion language modeling framework that fixes over-uniform training objectives and condition-target confusion in uniform diffusion LMs, enabling a 7B-scale UDLM with 3x-token-per-step speedup over autoregressive decoding and competitive complex reasoning performance.

0 favorites 0 likes
#parallel-decoding

WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models

Hugging Face Daily Papers ↗ · 4d ago Cached

WaveFront Decoding introduces a training-free self-speculative decoding framework for looped language models that reduces latency by concurrently batching drafting and verification, achieving up to 4.81x speedup on Huginn-3.5B.

0 favorites 0 likes
#parallel-decoding

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

Hugging Face Daily Papers ↗ · 2026-09-22 Cached

Flash-dLLM is a training-free inference acceleration framework for diffusion LLMs that uses IO-aware KV caching and parallel decoding to achieve significant speedups and memory efficiency improvements.

0 favorites 0 likes
#parallel-decoding

Scaling E-Commerce Attribute Extraction with Parallel Decoding

arXiv cs.CL ↗ · 2026-09-10 Cached

The paper introduces a two-stage LLM pipeline using fine-tuned Qwen3-4B with Hyper-Parallel Decoding to extract purchase-discriminative attributes in e-commerce, achieving 85% accuracy with 92% cost reduction.

0 favorites 0 likes
#parallel-decoding

alibaba-pai/MiniMax-H3-Acc-LoRAs

Hugging Face Models Trending ↗ · 2026-08-26 Cached

Alibaba PAI releases LoRA checkpoints for accelerating MiniMax-H3 video generation using Parallel Decoding Distillation, enabling efficient inference in 8 steps.

0 favorites 0 likes
#parallel-decoding

CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing

arXiv cs.CL ↗ · 2026-08-17 Cached

The paper introduces Consistency Forcing (CForce), a distillation technique for diffusion large language models that improves parallel decoding by aligning early-stage predictions with later stages, enhancing speed-quality trade-offs.

0 favorites 0 likes
#parallel-decoding

Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models

arXiv cs.CL ↗ · 2026-08-13 Cached

Proposes Ripple-Pivot Search, a training-free decoding method for diffusion large language models that proactively commits mid-entropy pivot positions to reduce uncertainty and accelerate parallel decoding, achieving 4-10x speedup.

0 favorites 0 likes
#parallel-decoding

PaDoc: Layout-Grounded Parallel Decoding for Document Parsing

Hugging Face Daily Papers ↗ · 2026-08-06 Cached

PaDoc introduces a layout-grounded parallel decoding method for end-to-end document parsers, decoupling layout and content decoding to reduce decoding depth and improve throughput. It achieves state-of-the-art results on OmniDocBenchFull and significantly speeds up inference compared to sequential baselines.

0 favorites 0 likes
#parallel-decoding

Parallel Decoding for Video Generation (10 minute read)

TLDR AI ↗ · 2026-07-30 Cached

NVIDIA introduces Parallel Decoding Distillation (PDD) for accelerating image and video generation, enabling high-quality outputs with fewer neural function evaluations on models like LTX-2.3 and Wan2.1-14B.

0 favorites 0 likes
#parallel-decoding

Parallel Decoding Distillation for Fast Image and Video Generation

Hugging Face Daily Papers ↗ · 2026-07-28 Cached

Parallel Decoding Distillation (PDD) is a trajectory-based distillation method that accelerates image and video generation by predicting multiple denoising steps per network evaluation, achieving state-of-the-art performance with 4-8 NFEs on models like LTX-2.3, Wan14B, and Qwen-Image.

0 favorites 0 likes
#parallel-decoding

DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding

arXiv cs.AI ↗ · 2026-07-24 Cached

Proposes DC-Leap, a training-free framework that accelerates diffusion large language models by introducing dynamic contiguous verification and draft-guided decoding, achieving up to 105× speedup with comparable generation quality.

0 favorites 0 likes
#parallel-decoding

HPD-Parsing: Hierarchical Parallel Document Parsing

Hugging Face Daily Papers ↗ · 2026-07-21 Cached

HPD-Parsing introduces a hierarchical parallel decoding paradigm for VLM-based document parsing, replacing full-page autoregressive generation to achieve 4,752 tokens per second throughput (2.62x faster than prior models) while maintaining competitive accuracy.

0 favorites 0 likes
#parallel-decoding

Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models

arXiv cs.CL ↗ · 2026-07-20 Cached

Proposes AdaLook, an adaptive multi-step lookahead decoding framework for masked diffusion language models that dynamically determines rollout depth and branch expansion based on candidate-score variance, achieving better accuracy-decoding steps trade-off compared to existing one-step lookahead decoding methods.

0 favorites 0 likes
#parallel-decoding

@rohanpaul_ai: Spatially Speculative Decoding (SSD) sped up autoregressive image models up to 13.28X by predicting image rows in paral…

X AI KOLs Timeline ↗ · 2026-07-14 Cached

Spatially Speculative Decoding (SSD) accelerates autoregressive image models by predicting entire rows in parallel using small helper networks, achieving up to 13.28x speedup while maintaining benchmark performance.

0 favorites 0 likes
#parallel-decoding

DeLS-Spec: Decoupled Long-Short Contexts for Parallel Speculative Drafting

arXiv cs.CL ↗ · 2026-07-09 Cached

DeLS-Spec decouples long- and short-context modeling in speculative decoding by adding a lightweight local head to DFlash, achieving consistent speedups without full retraining. It requires only standard next-token prediction training for the local head and improves acceptance length on Qwen3 benchmarks.

0 favorites 0 likes
#parallel-decoding

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning

Hugging Face Daily Papers ↗ · 2026-07-03 Cached

This paper introduces PadCaptioner, a 3B parameter model for omni-modal dense video captioning that uses parallelized autoregressive decoding to achieve high efficiency and quality, outperforming 7B counterparts. A latent planning mechanism enables lossless parallel generation by exploiting weak local dependencies among events.

0 favorites 0 likes
#parallel-decoding

Understanding Evaluation Illusion in Diffusion Large Language Models

arXiv cs.CL ↗ · 2026-06-30 Cached

This paper identifies evaluation inconsistencies in diffusion LLM decoding methods, showing that prompt template choice significantly impacts rankings, and proposes guidelines for reliable evaluation.

0 favorites 0 likes
#parallel-decoding

Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLM

arXiv cs.CL ↗ · 2026-06-26 Cached

This paper proposes Dynamic-dLLM, a training-free framework that accelerates diffusion large language models by dynamically allocating cache-update budgets and calibrating decoding thresholds, achieving over 3x speedup on models like LLaDA and Dream while maintaining performance.

0 favorites 0 likes
#parallel-decoding

What is Speculative Decoding? (trending on paperswithco.de) [R]

Reddit r/MachineLearning ↗ · 2026-06-17

Speculative decoding is an inference optimization technique that uses a fast draft model to propose future tokens verified in parallel by a larger model, improving LLM generation speed. The article highlights its trending status on Papers with Code and a recent SGLang blog post about state-of-the-art latencies using DFlash models.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback