diffusion-large-language-models

Tag

Cards List
#diffusion-large-language-models

LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models

arXiv cs.AI · 2026-07-21 Cached

LaCache introduces a training-free acceleration framework for diffusion-based LLMs that uses lossless caching and precision-adaptive inference to eliminate redundant computation during denoising steps, achieving up to 40.2× end-to-end speedup while maintaining task accuracy.

0 favorites 0 likes
#diffusion-large-language-models

Understanding Evaluation Illusion in Diffusion Large Language Models

arXiv cs.CL · 2026-06-30 Cached

This paper identifies evaluation inconsistencies in diffusion LLM decoding methods, showing that prompt template choice significantly impacts rankings, and proposes guidelines for reliable evaluation.

0 favorites 0 likes
#diffusion-large-language-models

FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models

arXiv cs.LG · 2026-06-08 Cached

This paper proposes FAIR-Calib, a two-stage post-training quantization framework for diffusion large language models that addresses the instability of token commitments during iterative refinement. It achieves state-of-the-art results on LLaDA and Dream models under low-bit quantization.

0 favorites 0 likes
#diffusion-large-language-models

dMoE: dLLMs with Learnable Block Experts

Hugging Face Daily Papers · 2026-05-29 Cached

This paper proposes dMoE, a block-level mixture-of-experts framework for diffusion large language models that aggregates token-level expert distributions into block-level routing, reducing activated experts and memory usage while maintaining performance.

0 favorites 0 likes
← Back to home

Submit Feedback