parallel-decoding

Tag

Cards List
#parallel-decoding

WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models

Hugging Face Daily Papers ↗ · 4d ago Cached

WaveFront Decoding introduces a training-free self-speculative decoding framework for looped language models that reduces latency by concurrently batching drafting and verification, achieving up to 4.81x speedup on Huginn-3.5B.

0 favorites 0 likes
#parallel-decoding

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

Hugging Face Daily Papers ↗ · 2026-09-22 Cached

Flash-dLLM is a training-free inference acceleration framework for diffusion LLMs that uses IO-aware KV caching and parallel decoding to achieve significant speedups and memory efficiency improvements.

0 favorites 0 likes
#parallel-decoding

Scaling E-Commerce Attribute Extraction with Parallel Decoding

arXiv cs.CL ↗ · 2026-09-10 Cached

The paper introduces a two-stage LLM pipeline using fine-tuned Qwen3-4B with Hyper-Parallel Decoding to extract purchase-discriminative attributes in e-commerce, achieving 85% accuracy with 92% cost reduction.

0 favorites 0 likes
#parallel-decoding

alibaba-pai/MiniMax-H3-Acc-LoRAs

Hugging Face Models Trending ↗ · 2026-08-26 Cached

Alibaba PAI releases LoRA checkpoints for accelerating MiniMax-H3 video generation using Parallel Decoding Distillation, enabling efficient inference in 8 steps.

0 favorites 0 likes
#parallel-decoding

CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing

arXiv cs.CL ↗ · 2026-08-17 Cached

The paper introduces Consistency Forcing (CForce), a distillation technique for diffusion large language models that improves parallel decoding by aligning early-stage predictions with later stages, enhancing speed-quality trade-offs.

0 favorites 0 likes
#parallel-decoding

Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models

arXiv cs.CL ↗ · 2026-08-13 Cached

Proposes Ripple-Pivot Search, a training-free decoding method for diffusion large language models that proactively commits mid-entropy pivot positions to reduce uncertainty and accelerate parallel decoding, achieving 4-10x speedup.

0 favorites 0 likes
#parallel-decoding

PaDoc: Layout-Grounded Parallel Decoding for Document Parsing

Hugging Face Daily Papers ↗ · 2026-08-06 Cached

PaDoc introduces a layout-grounded parallel decoding method for end-to-end document parsers, decoupling layout and content decoding to reduce decoding depth and improve throughput. It achieves state-of-the-art results on OmniDocBenchFull and significantly speeds up inference compared to sequential baselines.

0 favorites 0 likes
#parallel-decoding

Parallel Decoding for Video Generation (10 minute read)

TLDR AI ↗ · 2026-07-30 Cached

NVIDIA introduces Parallel Decoding Distillation (PDD) for accelerating image and video generation, enabling high-quality outputs with fewer neural function evaluations on models like LTX-2.3 and Wan2.1-14B.

0 favorites 0 likes
#parallel-decoding

Parallel Decoding Distillation for Fast Image and Video Generation

Hugging Face Daily Papers ↗ · 2026-07-28 Cached

Parallel Decoding Distillation (PDD) is a trajectory-based distillation method that accelerates image and video generation by predicting multiple denoising steps per network evaluation, achieving state-of-the-art performance with 4-8 NFEs on models like LTX-2.3, Wan14B, and Qwen-Image.

0 favorites 0 likes
#parallel-decoding

DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding

arXiv cs.AI ↗ · 2026-07-24 Cached

Proposes DC-Leap, a training-free framework that accelerates diffusion large language models by introducing dynamic contiguous verification and draft-guided decoding, achieving up to 105× speedup with comparable generation quality.

0 favorites 0 likes
#parallel-decoding

HPD-Parsing: Hierarchical Parallel Document Parsing

Hugging Face Daily Papers ↗ · 2026-07-21 Cached

HPD-Parsing introduces a hierarchical parallel decoding paradigm for VLM-based document parsing, replacing full-page autoregressive generation to achieve 4,752 tokens per second throughput (2.62x faster than prior models) while maintaining competitive accuracy.

0 favorites 0 likes
#parallel-decoding

Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models

arXiv cs.CL ↗ · 2026-07-20 Cached

Proposes AdaLook, an adaptive multi-step lookahead decoding framework for masked diffusion language models that dynamically determines rollout depth and branch expansion based on candidate-score variance, achieving better accuracy-decoding steps trade-off compared to existing one-step lookahead decoding methods.

0 favorites 0 likes
#parallel-decoding

@rohanpaul_ai: Spatially Speculative Decoding (SSD) sped up autoregressive image models up to 13.28X by predicting image rows in paral…

X AI KOLs Timeline ↗ · 2026-07-14 Cached

Spatially Speculative Decoding (SSD) accelerates autoregressive image models by predicting entire rows in parallel using small helper networks, achieving up to 13.28x speedup while maintaining benchmark performance.

0 favorites 0 likes
#parallel-decoding

DeLS-Spec: Decoupled Long-Short Contexts for Parallel Speculative Drafting

arXiv cs.CL ↗ · 2026-07-09 Cached

DeLS-Spec decouples long- and short-context modeling in speculative decoding by adding a lightweight local head to DFlash, achieving consistent speedups without full retraining. It requires only standard next-token prediction training for the local head and improves acceptance length on Qwen3 benchmarks.

0 favorites 0 likes
#parallel-decoding

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning

Hugging Face Daily Papers ↗ · 2026-07-03 Cached

This paper introduces PadCaptioner, a 3B parameter model for omni-modal dense video captioning that uses parallelized autoregressive decoding to achieve high efficiency and quality, outperforming 7B counterparts. A latent planning mechanism enables lossless parallel generation by exploiting weak local dependencies among events.

0 favorites 0 likes
#parallel-decoding

Understanding Evaluation Illusion in Diffusion Large Language Models

arXiv cs.CL ↗ · 2026-06-30 Cached

This paper identifies evaluation inconsistencies in diffusion LLM decoding methods, showing that prompt template choice significantly impacts rankings, and proposes guidelines for reliable evaluation.

0 favorites 0 likes
#parallel-decoding

Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLM

arXiv cs.CL ↗ · 2026-06-26 Cached

This paper proposes Dynamic-dLLM, a training-free framework that accelerates diffusion large language models by dynamically allocating cache-update budgets and calibrating decoding thresholds, achieving over 3x speedup on models like LLaDA and Dream while maintaining performance.

0 favorites 0 likes
#parallel-decoding

What is Speculative Decoding? (trending on paperswithco.de) [R]

Reddit r/MachineLearning ↗ · 2026-06-17

Speculative decoding is an inference optimization technique that uses a fast draft model to propose future tokens verified in parallel by a larger model, improving LLM generation speed. The article highlights its trending status on Papers with Code and a recent SGLang blog post about state-of-the-art latencies using DFlash models.

0 favorites 0 likes
#parallel-decoding

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models

Hugging Face Daily Papers ↗ · 2026-06-17 Cached

PerceptionDLM introduces a multimodal diffusion language model that enables parallel region perception via structured attention masking and efficient prompting, achieving faster inference without sacrificing caption quality. Experiments show competitive performance with substantial speed improvements for multi-region perception tasks.

0 favorites 0 likes
#parallel-decoding

Why might DiffusionGemma be better at tool calls than its benchmark quality suggests

Reddit r/LocalLLaMA ↗ · 2026-06-16

Analyzes how DiffusionGemma's bidirectional attention and parallel block generation could potentially yield higher valid tool call rates due to its ability to revise tokens, even though its base quality is lower than Gemma 4.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback