diffusion-llm

Tag

Cards List
#diffusion-llm

Mercury 2.5

Hacker News Top ↗ · 2026-09-08 Cached

Inception Labs releases Mercury 2.5, a diffusion-based language model with improved intelligence, speed, and cost-efficiency for production use in search, voice, and coding.

0 favorites 0 likes
#diffusion-llm

@StefanoErmon: Today we're excited to announce Mercury 2.5 It’s the most capable diffusion LLM on the market. It is a 40% jump in inte…

X AI KOLs Timeline ↗ · 2026-09-08 Cached

Mercury 2.5 is announced as the most capable diffusion language model, with a 40% increase in intelligence over Mercury 2, operating at over 1,100 tokens/sec on NVIDIA GPUs, and optimized for production with low latency and cost.

0 favorites 0 likes
#diffusion-llm

Answer First, Reason Later: Commitment Order in Diffusion LLMs

arXiv cs.CL ↗ · 2026-08-07 Cached

This paper investigates why diffusion LLMs fail at reasoning tasks: unconstrained token commitment freezes answers early and collapses to answer-only outputs. The authors identify commitment order as the root cause and propose a training-free, frontier-gated decoding intervention that recovers performance while preserving parallel decoding.

0 favorites 0 likes
#diffusion-llm

DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding

arXiv cs.AI ↗ · 2026-07-24 Cached

Proposes DC-Leap, a training-free framework that accelerates diffusion large language models by introducing dynamic contiguous verification and draft-guided decoding, achieving up to 105× speedup with comparable generation quality.

0 favorites 0 likes
#diffusion-llm

@mr_r0b0t: Big week for open-weights! inclusionAI just dropped LLaDA2.2-flash100B MoE diffusion LLM built for agents Levenshtein E…

X AI KOLs Following ↗ · 2026-07-23 Cached

inclusionAI released LLaDA2.2-flash, a 100B MoE diffusion LLM built for agents with Levenshtein Editing, 128K context, and up to 2.3× higher throughput, achieving strong scores on agentic benchmarks like τ²-Bench and PinchBench under Apache 2.0.

0 favorites 0 likes
#diffusion-llm

Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLM

arXiv cs.CL ↗ · 2026-06-26 Cached

This paper proposes Dynamic-dLLM, a training-free framework that accelerates diffusion large language models by dynamically allocating cache-update budgets and calibrating decoding thresholds, achieving over 3x speedup on models like LLaDA and Dream while maintaining performance.

0 favorites 0 likes
#diffusion-llm

Learning from the Self-future: On-policy Self-distillation for dLLMs

arXiv cs.CL ↗ · 2026-06-17 Cached

Introduces d-OPSD, the first on-policy self-distillation framework for diffusion large language models, using suffix conditioning and step-level supervision to outperform RLVR and SFT baselines on reasoning benchmarks.

0 favorites 0 likes
#diffusion-llm

Efficient On-Device Diffusion LLM Inference with Mobile NPU

arXiv cs.LG ↗ · 2026-06-15 Cached

This paper presents llada.cpp, an NPU-aware inference framework for accelerating diffusion large language models (dLLMs) on smartphones. It introduces three techniques—Multi-Block Speculative Decoding, Dual-Path Progressive Revision, and Swap-Optimized Memory Runtime—to align dLLM inference with mobile NPU characteristics, achieving 17-42x latency reduction over CPU baseline.

0 favorites 0 likes
#diffusion-llm

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload

Hugging Face Daily Papers ↗ · 2026-05-19 Cached

TIDE is a lossless inference system for diffusion large language models that leverages temporal stability of expert activations to reduce I/O overhead and computation, achieving up to 1.4-1.5x throughput improvements on single GPU-CPU systems.

0 favorites 0 likes
#diffusion-llm

PSD: Pushing the Pareto Frontier of Diffusion LLMs via Parallel Speculative Decoding

arXiv cs.CL ↗ · 2026-05-18 Cached

This paper introduces Parallel Speculative Decoding (PSD), a training-free framework that accelerates diffusion LLM inference by jointly improving spatial and temporal efficiency, achieving up to 5.5× tokens per forward pass with comparable quality to greedy decoding.

0 favorites 0 likes
#diffusion-llm

@DivyanshT91162: Autoregressive LLMs might already be getting replaced Someone built dLLM — an open-source library that can turn ANY aut…

X AI KOLs Timeline ↗ · 2026-05-16 Cached

dLLM is an open-source library that converts any autoregressive LLM into a diffusion LLM, enabling parallel decoding and faster text generation.

0 favorites 0 likes
#diffusion-llm

Why there isn't any top LLM providers investing on diffusion LLM?

Reddit r/singularity ↗ · 2026-05-11

This article questions why major LLM providers are not investing in Diffusion LLMs despite recent advancements like Mercury 2. It explores potential fundamental issues or hardware bottlenecks hindering broader adoption.

0 favorites 0 likes
#diffusion-llm

$R^2$-dLLM: Accelerating Diffusion Large Language Models via Spatio-Temporal Redundancy Reduction

arXiv cs.CL ↗ · 2026-04-22 Cached

R²-dLLM introduces spatio-temporal redundancy reduction techniques that cut diffusion LLM decoding steps by up to 75% while preserving generation quality, addressing a key deployment bottleneck.

0 favorites 0 likes
← Back to home

Submit Feedback