dMoE: dLLMs with Learnable Block Experts

Hugging Face Daily Papers Papers

Summary

This paper proposes dMoE, a block-level mixture-of-experts framework for diffusion large language models that aggregates token-level expert distributions into block-level routing, reducing activated experts and memory usage while maintaining performance.

Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive models, offering competitive performance while naturally supporting parallel decoding. However, as dLLMs are increasingly integrated with Mixture-of-Experts (MoE) architectures to scale model capacity, a fundamental mismatch arises between block parallel decoding and token-level expert selection. Specifically, each dLLM forward pass processes multiple tokens with bidirectional dependencies, whereas conventional MoE layers route each token independently. This mismatch substantially increases the number of uniquely activated experts, making inference increasingly memory-bound. To address this, we propose dMoE, a simple yet effective block-level MoE framework. The central idea of dMoE is to aggregate token-level expert distributions within each block into a unified block-level expert distribution, which is then used to guide expert routing in a more coherent manner. In this way, dMoE substantially reduces the number of uniquely activated experts during inference without sacrificing performance, thereby mitigating the memory-bound bottleneck. Extensive experiments across a variety of benchmarks demonstrate the effectiveness of dMoE. On average, dMoE reduces the number of uniquely activated experts from 69.5 to 14.6 while retaining 99.11% of the original performance. Meanwhile, it reduces memory usage by 76.64% to 79.84% and achieves 1.14times to 1.66times end-to-end latency speedup. Code is available at: https://github.com/fscdc/dMoE
Original Article
View Cached Full Text

Cached at: 06/01/26, 03:17 AM

Paper page - dMoE: dLLMs with Learnable Block Experts

Source: https://huggingface.co/papers/2605.30876

Abstract

Diffusion large language models combined with mixture-of-experts architectures face a mismatch between block parallel decoding and token-level expert selection, which dMoE addresses by aggregating token-level distributions into block-level routing to reduce activated experts and improve efficiency.

Diffusion Large Language Models(dLLMs) have recently emerged as a promising alternative toautoregressive models, offering competitive performance while naturally supportingparallel decoding. However, as dLLMs are increasingly integrated withMixture-of-Experts(MoE) architectures to scale model capacity, a fundamental mismatch arises betweenblock parallel decodingandtoken-level expert selection. Specifically, each dLLM forward pass processes multiple tokens with bidirectional dependencies, whereas conventional MoE layers route each token independently. This mismatch substantially increases the number of uniquely activated experts, making inference increasingly memory-bound. To address this, we propose dMoE, a simple yet effective block-level MoE framework. The central idea of dMoE is to aggregate token-level expert distributions within each block into a unifiedblock-level expert distribution, which is then used to guideexpert routingin a more coherent manner. In this way, dMoE substantially reduces the number of uniquely activated experts during inference without sacrificing performance, thereby mitigating thememory-bound bottleneck. Extensive experiments across a variety of benchmarks demonstrate the effectiveness of dMoE. On average, dMoE reduces the number of uniquely activated experts from 69.5 to 14.6 while retaining 99.11% of the original performance. Meanwhile, it reduces memory usage by 76.64% to 79.84% and achieves 1.14times to 1.66timesend-to-end latencyspeedup. Code is available at: https://github.com/fscdc/dMoE

View arXiv pageView PDFProject pageGitHub16Add to collection

Get this paper in your agent:

hf papers read 2605\.30876

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2605.30876 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2605.30876 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2605.30876 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Xiaomi-TabLDM: A Tabular Foundation Model Technical Report

arXiv cs.AI

Introduces Xiaomi-TabLDM, a tabular foundation model that leverages synthetic data and in-context learning for superior prediction accuracy without task-specific fine-tuning, achieving top rankings on multiple benchmarks.

SGD-KV: Summarization Guided KV Cache Compression

arXiv cs.CL

SGD-KV is a framework that uses summarization to guide KV cache compression in large language models, reducing memory usage by up to 75% for contexts up to 1M tokens while achieving state-of-the-art performance on long-context benchmarks.