cMoLLM at Scale: Horizontal Scaling Laws for Mixture-of-LLMs
Summary
Introduces cMoLLM, a convolutionally gated mixture-of-LLMs that routes over end-to-end streams through fully differentiable dynamic convolution, improving perplexity and downstream accuracy under matched compute.
View Cached Full Text
Cached at: 07/28/26, 06:25 AM
# Horizontal Scaling Laws for Convolutionally-Gated Mixture-of-LLMs
Source: [https://arxiv.org/html/2607.22577](https://arxiv.org/html/2607.22577)
Yemin WangMingda LiuLetian LiShuaishuai CaoZhengxiao HeRyan Dong
###### Abstract
Scaling large language models \(LLMs\) has driven their success, yet dense Transformers couple capacity and computation: every parameter is activated for every token, making training and inference costs grow linearly with model size—a critical bottleneck as models approach trillion\-parameter regimes\. We aim to scale capacity through MoE\-style mixture throughout the LLM pipeline rather than only the FFN\. Prior pipeline\-level approaches include ParaScale\(Chenet al\.,[2025](https://arxiv.org/html/2607.22577#bib.bib10)\), which introduces virtual tokens and parallel streams but incurs substantial overhead and suffers from homogenized routing and gradient collapse, and AltUp\(Baykalet al\.,[2023](https://arxiv.org/html/2607.22577#bib.bib11)\), which uses an auxiliary prediction branch but offers limited adaptivity and slow convergence\. We establish that MoE\-style mixture layers can be reformulated as variable\-kernel dynamic convolutions, where each expert corresponds to a1×11\{\\times\}1convolutional kernel and routing implements input\-conditioned kernel aggregation\. Building on this equivalence, we introduce cMoLLM: a convolutionally gated mixture\-of\-LLMs that routes over end\-to\-end streams through fully differentiable dynamic convolution\. In GPT\-2\-style models trained on FineWeb, cMoLLM improves language modeling perplexity and downstream GLUE and SQuAD accuracy under matched compute, with better stream utilization, more stable optimization, and favorable scaling compared to ParaScale\- and AltUp\-style baselines\.
Mixture\-of\-LLMs, Pipeline\-Level Scaling, Dynamic Convolution, Convolutionally\-Gated Routing, Large Language Models
Figure 1:Teaser of cMoLLM\.cMoLLM scales capacity at the pipeline level via convolutionally\-gated mixture over end\-to\-end streams, yielding better perplexity and downstream accuracy than dense baselines under matched compute, without virtual tokens or auxiliary heads\.Table 1:Systematic qualitative comparison of pipeline\-level capacity scaling methods: ParaScale, AltUp, and cMoLLM\.## 1Introduction
The remarkable success of large language models \(LLMs\)\(OpenAI,[2025](https://arxiv.org/html/2607.22577#bib.bib42); Anthropic,[2025](https://arxiv.org/html/2607.22577#bib.bib41); Yanget al\.,[2025](https://arxiv.org/html/2607.22577#bib.bib43)\)is closely tied to scale: increasing model capacity consistently improves performance across diverse tasks\(Radfordet al\.,[2019](https://arxiv.org/html/2607.22577#bib.bib26); Brownet al\.,[2020](https://arxiv.org/html/2607.22577#bib.bib27); Kaplanet al\.,[2020](https://arxiv.org/html/2607.22577#bib.bib1); Hoffmannet al\.,[2022](https://arxiv.org/html/2607.22577#bib.bib2); Yanget al\.,[2026](https://arxiv.org/html/2607.22577#bib.bib40); Chenget al\.,[2025](https://arxiv.org/html/2607.22577#bib.bib39)\)\. Training and inference cost grow roughly linearly with model size, as every parameter is activated for every token\. As models approach trillion\-parameter regimes, this linear compute–capacity coupling has become a dominant bottleneck, motivating the search for more parameter\-efficient ways to expand capacity\.
A natural direction is to use a Mixture\-of\-Experts \(MoE\)–style approach to scale capacity: route tokens across multiple sub\-models and combine their outputs\. Most prior MoE work applies this only to thefeed\-forward network \(FFN\): experts are FFN blocks, and routing is typically discrete \(e\.g\., Top\-KK\), which leads to expert collapse, skewed utilization, and brittle training\(Shazeeret al\.,[2017](https://arxiv.org/html/2607.22577#bib.bib5); Feduset al\.,[2022](https://arxiv.org/html/2607.22577#bib.bib6)\)\. We instead pursue pipeline\-level mixture: routing across entire LLM pipelines \(or stream\-wise sub\-models\) rather than FFN\-only experts\.
Two works are directly relevant\.AltUp\(Baykalet al\.,[2023](https://arxiv.org/html/2607.22577#bib.bib11)\)widens token representations and uses a virtual prediction branch to update inactive blocks; it increases effective capacity with small parameter overhead but relies on a fixed, hand\-designed structure that is not fully adaptive and often converges slowly\. This increases sequence length and thus compute; moreover, routing can homogenize across streams, leading to gradient collapse and training instability\. We seek a new pipeline\-level scaling mechanism that keeps the MoE\-style routing across sub\-models intuition, avoids virtual tokens, auxiliary prediction heads, and the pitfalls above\.
Our starting point is a theoretical insight: MoE\-style mixture layers can be exactly rewritten as dynamic convolutions with input\-dependent \(variable\) kernels\. Formally, each expert corresponds to a1×11\{\\times\}1convolutional kernel; the router performs input\-dependent mixing of these kernels \([Theorem4\.1](https://arxiv.org/html/2607.22577#S4.Thmtheorem1)\)\. This establishes MoE–dynamic convolution equivalence and gives a unified lens for analyzing sparse, conditional computation\. We then approximate this ideal: each “stream” uses a distinct convolutional kernel, and we mix streams via soft, fully differentiable gating—no Top\-KK, no low\-rank factorization\.
Building on this, we proposecMoLLM: a convolutionally\-gated mixture\-of\-LLMs that applies conditional computation to the entire Transformer pipeline \([Figure1](https://arxiv.org/html/2607.22577#S0.F1)\)\. We maintain a small set of end\-to\-end “streams,” each associated with its own1×11\{\\times\}1kernel; a lightweight gating network produces input\-dependent mixture weights, and the mixed kernel is applied via standard \(grouped\) pointwise convolution\. All streams are combined through a stable, fully differentiable dynamic convolution—no virtual tokens, no auxiliary heads, no Top\-KKor low\-rank—yielding parameter\-efficient pipeline\-level capacity scaling with hardware\-friendly convolutional primitives\.
Our contributions are as follows:
- •Theoretical:We establish the formal equivalence between MoE\-style mixture layers and dynamic convolutions with variable kernels \([Theorem4\.1](https://arxiv.org/html/2607.22577#S4.Thmtheorem1)\)\. This provides a unified theoretical framework for analyzing and designing sparse, conditional computation models, beyond FFN\-level MoE\.
- •Methodological:We introducecMoLLM, a pipeline\-level convolutionally\-gated mixture whose kernel\-sharing and stream structure follow directly from the convolution view\. It achieves parameter\-efficient scaling and training stability without virtual tokens, auxiliary prediction branches, Top\-KKrouting, or low\-rank factorization\.
- •Empirical:On GPT\-2–style models trained on FineWeb, cMoLLM matches or improves perplexity, GLUE, and SQuAD over dense baselines under comparable cost, with better stream utilization, more stable training dynamics, and favorable scaling \(see[Section5\.3](https://arxiv.org/html/2607.22577#S5.SS3)\) compared to ParaScale\- and AltUp\-style pipeline mixtures\.
Figure 2:Architecture \(Fig\. 2\)\.cMoLLM block: input𝐗\\mathbf\{X\}passes through a gating network to produce soft mixture weights\{gk\}\\\{g\_\{k\}\\\}; each stream has a1×11\{\\times\}1kernel𝐊k\\mathbf\{K\}\_\{k\}; the effective kernel𝐊~\(𝐱\)=∑kgk\(𝐱\)𝐊k\\widetilde\{\\mathbf\{K\}\}\(\\mathbf\{x\}\)=\\sum\_\{k\}g\_\{k\}\(\\mathbf\{x\}\)\\mathbf\{K\}\_\{k\}is applied via grouped1×11\{\\times\}1convolution\. No virtual tokens, no Top\-KK, no auxiliary heads\.
## 2Related Work
Mixture\-of\-Experts and Pipeline\-Level Scaling\.Classical MoE approaches\(Shazeeret al\.,[2017](https://arxiv.org/html/2607.22577#bib.bib5); Feduset al\.,[2022](https://arxiv.org/html/2607.22577#bib.bib6); Lepikhinet al\.,[2021](https://arxiv.org/html/2607.22577#bib.bib7)\)sparsify thefeed\-forward network \(FFN\)layer using Top\-KKrouting, where each token is routed to a sparse subset of expert FFN blocks\. Large\-scale MoE systems such as GLaM\(Duet al\.,[2022](https://arxiv.org/html/2607.22577#bib.bib14)\), BASE Layers\(Lewiset al\.,[2021](https://arxiv.org/html/2607.22577#bib.bib13)\), DeepSpeed\-MoE\(Rajbhandariet al\.,[2022](https://arxiv.org/html/2607.22577#bib.bib15)\), and Vision MoE models\(Riquelmeet al\.,[2021](https://arxiv.org/html/2607.22577#bib.bib33)\)focus primarily on FFN\-level sparsity and system\-level optimizations for training and inference\. In contrast, we target pipeline\-level mixture over entire LLM streams, leveraging the MoE–dynamic convolution equivalence \([Section4\.1](https://arxiv.org/html/2607.22577#S4.SS1)\) to design convolutionally\-gated routing without Top\-KKtruncation\. The most directly relevant prior work includesParaScale\(Chenet al\.,[2025](https://arxiv.org/html/2607.22577#bib.bib10)\), which scales capacity via parallel virtual token streams but increases compute and suffers from gradient collapse, andAltUp\(Baykalet al\.,[2023](https://arxiv.org/html/2607.22577#bib.bib11)\), which uses a hand\-designed auxiliary prediction branch but converges slowly\.[Table1](https://arxiv.org/html/2607.22577#S0.T1)provides a systematic comparison; cMoLLM avoids virtual tokens and auxiliary heads, using fully differentiable dynamic convolution for stable, parameter\-efficient scaling\.
Dynamic Convolution and Conditional Computation\.Dynamic convolution\(Jiaet al\.,[2016](https://arxiv.org/html/2607.22577#bib.bib19); Yanget al\.,[2019](https://arxiv.org/html/2607.22577#bib.bib16); Chenet al\.,[2020](https://arxiv.org/html/2607.22577#bib.bib20); Wuet al\.,[2019](https://arxiv.org/html/2607.22577#bib.bib17); Maet al\.,[2020](https://arxiv.org/html/2607.22577#bib.bib18); Zhanget al\.,[2020](https://arxiv.org/html/2607.22577#bib.bib38)\)adapts convolution kernels based on input, typically within CNN backbones or sequence models\. More broadly, conditional computation and routing networks\(Bengioet al\.,[2013](https://arxiv.org/html/2607.22577#bib.bib21); Rosenbaumet al\.,[2018](https://arxiv.org/html/2607.22577#bib.bib22); McGill and Perona,[2017](https://arxiv.org/html/2607.22577#bib.bib37)\)learn to select computation paths depending on inputs\. Our key theoretical contribution is establishing the formal equivalence between MoE\-style mixture layers and dynamic convolutions with variable kernels \([Theorem4\.1](https://arxiv.org/html/2607.22577#S4.Thmtheorem1)\), which enables us to design a pipeline\-level, convolutionally\-gated mixture over LLM streams with explicit probabilistic routing and load\-balancing objectives\.
Parameter\-Efficient Methods and Surveys\.Parameter\-efficient fine\-tuning methods such as adapters\(Houlsbyet al\.,[2019](https://arxiv.org/html/2607.22577#bib.bib24)\)and LoRA\(Huet al\.,[2022](https://arxiv.org/html/2607.22577#bib.bib23)\)add or reparameterize a small set of trainable weights on top of frozen backbones, primarily targeting efficient fine\-tuning\. cMoLLM instead modifies the pretraining architecture to improve the capacity–compute trade\-off; these approaches are complementary and could in principle be combined\. Comprehensive surveys of MoE models and their applications in LLMs and beyond are provided byCaiet al\.\([2025](https://arxiv.org/html/2607.22577#bib.bib34)\); Mu and Lin \([2025](https://arxiv.org/html/2607.22577#bib.bib36)\); Liuet al\.\([2026](https://arxiv.org/html/2607.22577#bib.bib35)\), which situate our contribution within the broader MoE landscape\.
## 3Preliminaries
We first introduce the notation used throughout the paper and briefly review the core concepts that underpin our formulation\.
### 3\.1Notation and Assumptions
We use bold lowercase letters \(e\.g\.,𝐱\\mathbf\{x\}\) for vectors, bold uppercase letters \(e\.g\.,𝐖\\mathbf\{W\}\) for matrices, and calligraphic letters \(e\.g\.,𝒮\\mathcal\{S\}\) for sets\. The hidden dimension of the Transformer isdd, the intermediate FFN dimension isdffd\_\{\\mathrm\{ff\}\}, the sequence length isLLand the number of experts/streams isNN\. For𝐗∈ℝL×d\\mathbf\{X\}\\in\\mathbb\{R\}^\{L\\times d\}, theii\-th token is denoted by𝐱i\\mathbf\{x\}\_\{i\}\.
Throughout, we make the following mild assumptions\.
###### Assumption 3\.1\(Bounded Inputs\)\.
There existsR\>0R\>0such that‖𝐱‖2≤R\\\|\\mathbf\{x\}\\\|\_\{2\}\\leq Rfor all token representations𝐱\\mathbf\{x\}encountered during training and evaluation\.
###### Assumption 3\.2\(Lipschitz Nonlinearity\)\.
The activation functionσ\\sigmaisLσL\_\{\\sigma\}\-Lipschitz and satisfiesσ\(0\)=0\\sigma\(0\)=0\(e\.g\., ReLU or GELU\)\.
These conditions are standard in theoretical analyses of deep networks and MoE architectures, and they suffice for establishing the complexity and stability results in this paper\.
### 3\.2Transformer Feed\-Forward Network
A standard Transformer FFN applies two linear projections with a nonlinearity\(Vaswaniet al\.,[2017](https://arxiv.org/html/2607.22577#bib.bib25)\):
FFN\(𝐱\)=𝐖2σ\(𝐖1𝐱\+𝐛1\)\+𝐛2,\\mathrm\{FFN\}\(\\mathbf\{x\}\)=\\mathbf\{W\}\_\{2\}\\,\\sigma\(\\mathbf\{W\}\_\{1\}\\mathbf\{x\}\+\\mathbf\{b\}\_\{1\}\)\+\\mathbf\{b\}\_\{2\},\(1\)where𝐱∈ℝd\\mathbf\{x\}\\in\\mathbb\{R\}^\{d\}is the input token representation,𝐖1∈ℝdff×d\\mathbf\{W\}\_\{1\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{ff\}\}\\times d\},𝐖2∈ℝd×dff\\mathbf\{W\}\_\{2\}\\in\\mathbb\{R\}^\{d\\times d\_\{\\mathrm\{ff\}\}\}, andσ\\sigmais typically GELU or ReLU\. The FFN is applied independently to each token\.
### 3\.3Mixture\-of\-Experts Layer
An MoE layer replaces a single feed\-forward network with a collection ofNNexperts\{Ek\}k=1N\\\{E\_\{k\}\\\}\_\{k=1\}^\{N\}and a gating functionGGthat routes each input to a subset of experts\(Shazeeret al\.,[2017](https://arxiv.org/html/2607.22577#bib.bib5)\):
MoE\(𝐱\)=∑k=1Ngk\(𝐱\)Ek\(𝐱\),\\mathrm\{MoE\}\(\\mathbf\{x\}\)=\\sum\_\{k=1\}^\{N\}g\_\{k\}\(\\mathbf\{x\}\)\\,E\_\{k\}\(\\mathbf\{x\}\),\(2\)whereg\(𝐱\)∈ℝNg\(\\mathbf\{x\}\)\\in\\mathbb\{R\}^\{N\}denotes the routing weights produced by the gate\. In practice, only the top\-KKentries ofg\(𝐱\)g\(\\mathbf\{x\}\)are nonzero, yielding sparse computation\. Each expertEkE\_\{k\}is typically a two\-layer FFN with its own parameters\.
### 3\.4Pointwise Convolution as a Linear Transform
A pointwise \(1×11\{\\times\}1\) convolution applies the same linear map across token positions\. Given an input sequence𝐗∈ℝL×din\\mathbf\{X\}\\in\\mathbb\{R\}^\{L\\times d\_\{\\mathrm\{in\}\}\}and a kernel𝐊∈ℝdout×din\\mathbf\{K\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times d\_\{\\mathrm\{in\}\}\}, the operation is given by
Conv1×1\(𝐗;𝐊\)=𝐗𝐊⊤,\\mathrm\{Conv\}\_\{1\{\\times\}1\}\(\\mathbf\{X\};\\mathbf\{K\}\)=\\mathbf\{X\}\\mathbf\{K\}^\{\\top\},\(3\)which corresponds to applying an identical token\-wise linear projection fromℝdin\\mathbb\{R\}^\{d\_\{\\mathrm\{in\}\}\}toℝdout\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\}\. This observation serves as a key building block for our subsequent analysis\.
## 4Method: cMoLLM
We now present our main theoretical result and the cMoLLM architecture\.
### 4\.1MoE as Dynamic Convolution: A Formal Equivalence
We show that MoE layers admit an equivalent interpretation as dynamic pointwise convolutions\. In this subsection we focus on linear experts and the structure of the router; extensions to nonlinear experts are discussed in[AppendixA](https://arxiv.org/html/2607.22577#A1)and summarized in[Corollary4\.2](https://arxiv.org/html/2607.22577#S4.Thmtheorem2)\.
###### Theorem 4\.1\(MoE–Dynamic Convolution Equivalence\)\.
Let𝐱∈ℝd\\mathbf\{x\}\\in\\mathbb\{R\}^\{d\}denote a token representation\. Consider an MoE layer withNNlinear expertsEk\(𝐱\)=𝐱𝐖k⊤E\_\{k\}\(\\mathbf\{x\}\)=\\mathbf\{x\}\\mathbf\{W\}\_\{k\}^\{\\top\}, where𝐖k∈ℝdout×d\\mathbf\{W\}\_\{k\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times d\}, and routing weights\{gk\(𝐱\)\}k=1N\\\{g\_\{k\}\(\\mathbf\{x\}\)\\\}\_\{k=1\}^\{N\}satisfying∑k=1Ngk\(𝐱\)=1\\sum\_\{k=1\}^\{N\}g\_\{k\}\(\\mathbf\{x\}\)=1for all𝐱\\mathbf\{x\}\. Then the MoE output can be written as a dynamic1×11\{\\times\}1convolution:
MoE\(𝐱\)=Conv1×1\(𝐱;𝐊~\(𝐱\)\)=𝐱𝐊~\(𝐱\)⊤,\\mathrm\{MoE\}\(\\mathbf\{x\}\)=\\mathrm\{Conv\}\_\{1\{\\times\}1\}\\\!\\Bigl\(\\mathbf\{x\};\\,\\tilde\{\\mathbf\{K\}\}\(\\mathbf\{x\}\)\\Bigr\)=\\mathbf\{x\}\\,\\tilde\{\\mathbf\{K\}\}\(\\mathbf\{x\}\)^\{\\top\},\(4\)where the effective kernel is given by
𝐊~\(𝐱\):=∑k=1Ngk\(𝐱\)𝐖k\.\\tilde\{\\mathbf\{K\}\}\(\\mathbf\{x\}\):=\\sum\_\{k=1\}^\{N\}g\_\{k\}\(\\mathbf\{x\}\)\\,\\mathbf\{W\}\_\{k\}\.\(5\)Since𝐊~\(𝐱\)\\tilde\{\\mathbf\{K\}\}\(\\mathbf\{x\}\)depends on the input𝐱\\mathbf\{x\}, the MoE layer is precisely a form of dynamic convolution\.
###### Proof\.
By linearity of matrix multiplication:
∑k=1Ngk\(𝐱\)⋅\(𝐱𝐖k⊤\)\\displaystyle\\sum\_\{k=1\}^\{N\}g\_\{k\}\(\\mathbf\{x\}\)\\cdot\(\\mathbf\{x\}\\mathbf\{W\}\_\{k\}^\{\\top\}\)=𝐱\(∑k=1Ngk\(𝐱\)𝐖k\)⊤\\displaystyle=\\mathbf\{x\}\\Bigl\(\\textstyle\\sum\_\{k=1\}^\{N\}g\_\{k\}\(\\mathbf\{x\}\)\\mathbf\{W\}\_\{k\}\\Bigr\)^\{\\\!\\top\}=Conv1×1\(𝐱;∑k=1Ngk\(𝐱\)𝐖k\)\.\\displaystyle=\\mathrm\{Conv\}\_\{1\{\\times\}1\}\\\!\\Bigl\(\\mathbf\{x\};\\,\\textstyle\\sum\_\{k=1\}^\{N\}g\_\{k\}\(\\mathbf\{x\}\)\\mathbf\{W\}\_\{k\}\\Bigr\)\.\(6\)The weighted sum of expert weight matrices forms a dynamic kernel that varies with𝐱\\mathbf\{x\}\. \(Classical MoE uses Top\-KKrouting so onlyKKterms are nonzero; we use soft mixing over all streams\.\) See[AppendixA](https://arxiv.org/html/2607.22577#A1)for extension to nonlinear experts\. ∎
###### Corollary 4\.2\(Nonlinear Two\-Layer Experts\)\.
Under[3\.2](https://arxiv.org/html/2607.22577#S3.Thmtheorem2), consider two\-layer experts of the formEk\(𝐱\)=σ\(𝐱𝐖k,1⊤\)𝐖k,2⊤E\_\{k\}\(\\mathbf\{x\}\)=\\sigma\(\\mathbf\{x\}\\mathbf\{W\}\_\{k,1\}^\{\\top\}\)\\mathbf\{W\}\_\{k,2\}^\{\\top\}with routing weights\{gk\(𝐱\)\}k=1N\\\{g\_\{k\}\(\\mathbf\{x\}\)\\\}\_\{k=1\}^\{N\}satisfying∑kgk\(𝐱\)=1\\sum\_\{k\}g\_\{k\}\(\\mathbf\{x\}\)=1\. Then the MoE layer can still be written as a dynamic1×11\{\\times\}1convolution with an input\-dependent effective kernel that incorporates a data\-dependent mask induced byσ\\sigma; see[AppendixA](https://arxiv.org/html/2607.22577#A1)for details\.
Assumptions[3\.1](https://arxiv.org/html/2607.22577#S3.Thmtheorem1)and[3\.2](https://arxiv.org/html/2607.22577#S3.Thmtheorem2)are not needed for[Theorem4\.1](https://arxiv.org/html/2607.22577#S4.Thmtheorem1)itself, but they will be used in our later complexity, stability, and toy\-model analyses that build on this equivalence\.
This equivalence motivates our approach: we implement pipeline\-level mixture as an explicit dynamic convolution, leveraging efficient convolutional primitives\.
### 4\.2cMoLLM Block Design
Based on[Theorem4\.1](https://arxiv.org/html/2607.22577#S4.Thmtheorem1), we designcMoLLMas a pipeline\-level mixture\.[Figure2](https://arxiv.org/html/2607.22577#S1.F2)illustrates the architecture: gating, per\-stream kernels, and dynamic convolution\.
Figure 3:Experimental results \(Fig\. 3\)\.Left: per\-stream gating weight distribution across layers\. Right: validation loss and perplexity vs\. stream countnnor training steps\. cMoLLM maintains stable utilization and scaling gains\.Gating Network\.Given input𝐗∈ℝL×d\\mathbf\{X\}\\in\\mathbb\{R\}^\{L\\times d\}, a lightweight MLP produces routing logits:
𝐇flat=Flatten\(𝐘streams\)∈ℝL×\(N⋅d\),\\displaystyle\\mathbf\{H\}\_\{\\text\{flat\}\}=\\text\{Flatten\}\(\\mathbf\{Y\}\_\{\\text\{streams\}\}\)\\in\\mathbb\{R\}^\{L\\times\(N\\cdot d\)\},\(7\)𝐇mid=σ\(𝐇flat𝐖g\(1\)\+𝐛g\(1\)\)∈ℝL×d,\\displaystyle\\mathbf\{H\}\_\{\\text\{mid\}\}=\\sigma\\bigl\(\\mathbf\{H\}\_\{\\text\{flat\}\}\\mathbf\{W\}\_\{g\}^\{\(1\)\}\+\\mathbf\{b\}\_\{g\}^\{\(1\)\}\\bigr\)\\in\\mathbb\{R\}^\{L\\times d\},𝐇logits=𝐇mid𝐖g\(2\)\+𝐛g\(2\)∈ℝL×N,\\displaystyle\\mathbf\{H\}\_\{\\text\{logits\}\}=\\mathbf\{H\}\_\{\\text\{mid\}\}\\mathbf\{W\}\_\{g\}^\{\(2\)\}\+\\mathbf\{b\}\_\{g\}^\{\(2\)\}\\in\\mathbb\{R\}^\{L\\times N\},where𝐖g∈ℝd×N\\mathbf\{W\}\_\{g\}\\in\\mathbb\{R\}^\{d\\times N\}\. We apply a softmax over streams to obtain mixture weights:
gk\(𝐱\)=softmax\(𝐡\)k,k∈\[N\]\.g\_\{k\}\(\\mathbf\{x\}\)=\\mathrm\{softmax\}\(\\mathbf\{h\}\)\_\{k\},\\quad k\\in\[N\]\.\(8\)This fully differentiable gating avoids discrete Top\-KKtruncation; all streams participate via soft weights\. We support multiple gating variants \(simple,context\_aware,multi\_head,adaptive\), evaluated in[Section5](https://arxiv.org/html/2607.22577#S5)\.
Expert Kernel Generator\.Each streamkkhas its own1×11\{\\times\}1convolutional kernel𝐊k∈ℝdff×d\\mathbf\{K\}\_\{k\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{ff\}\}\\times d\}\. The effective kernel is their mixture weighted bygk\(𝐱\)g\_\{k\}\(\\mathbf\{x\}\); we use no low\-rank factorization or shared base in our main setup\.
Dynamic Grouped Convolution\.Let a set of mixture weights\{gk\(𝐱\)\}k=1N\\\{g\_\{k\}\(\\mathbf\{x\}\)\\\}\_\{k=1\}^\{N\}be given\. We form a dynamic mixture of convolutional kernels by
𝐊~\(𝐱\)=∑k=1Ngk\(𝐱\)𝐊k\.\\tilde\{\\mathbf\{K\}\}\(\\mathbf\{x\}\)=\\sum\_\{k=1\}^\{N\}g\_\{k\}\(\\mathbf\{x\}\)\\,\\mathbf\{K\}\_\{k\}\.\(9\)Practically, we compute branch\-wise outputs in parallel and fuse them:
𝐗k=Convk\(𝐗;𝐊k\),k∈\[N\],\\mathbf\{X\}\_\{k\}=\\mathrm\{Conv\}\_\{k\}\(\\mathbf\{X\};\\mathbf\{K\}\_\{k\}\),\\quad k\\in\[N\],\(10\)and concatenate the per\-branch outputs along the representation dimension:
𝐗input=\[𝐗1,𝐗2,…,𝐗N\]∈ℝL×N×d,\\mathbf\{X\}\_\{\\text\{input\}\}=\[\\mathbf\{X\}\_\{1\},\\mathbf\{X\}\_\{2\},\\ldots,\\mathbf\{X\}\_\{N\}\]\\in\\mathbb\{R\}^\{L\\times N\\times d\},\(11\)whereLLis the sequence length andddis the feature dimension \(hidden size\)\.
To model interactions among the branches, we apply a transformer block to the stream of parallel outputs:
𝐘stream=TRANSFORMERS\(𝐗input;𝐊~\(𝐗\)\)\.\\mathbf\{Y\}\_\{\\mathrm\{stream\}\}=\\mathrm\{TRANSFORMERS\}\\bigl\(\\mathbf\{X\}\_\{\\text\{input\}\};\\;\\tilde\{\\mathbf\{K\}\}\(\\mathbf\{X\}\)\\bigr\)\.\(12\)HereTRANSFORMERS\(⋅\)\\mathrm\{TRANSFORMERS\}\(\\cdot\)denotes one or more Transformer layers, and the attention \(or feed\-forward\) computations may be conditioned on the dynamic kernels𝐊~\(𝐗\)\\tilde\{\\mathbf\{K\}\}\(\\mathbf\{X\}\)\.
ParaScale Output Aggregation\.Let the hidden states have shape𝐇∈ℝS×B×H\\mathbf\{H\}\\in\\mathbb\{R\}^\{S\\times B\\times H\}\. If the ParaScale module usesNps\>1N\_\{\\mathrm\{ps\}\}\>1parallel scales, we first reorganize the hidden states along the parallel dimension:
𝐇→𝐇^=reshape\(𝐇,S,B/Nps,Nps,H\),\\mathbf\{H\}\\rightarrow\\hat\{\\mathbf\{H\}\}=\\operatorname\{reshape\}\\bigl\(\\mathbf\{H\},\\,S,\\,B/N\_\{\\mathrm\{ps\}\},\\,N\_\{\\mathrm\{ps\}\},\\,H\\bigr\),\(13\)and flatten the parallel axis:
𝐇flat=reshape\(𝐇^,S,B/Nps,NpsH\)\.\\mathbf\{H\}\_\{\\text\{flat\}\}=\\operatorname\{reshape\}\\bigl\(\\hat\{\\mathbf\{H\}\},\\,S,\\,B/N\_\{\\mathrm\{ps\}\},\\,N\_\{\\mathrm\{ps\}\}\\,H\\bigr\)\.\(14\)We compute an aggregation weight via a learned layer:
𝜶s,b=softmax\(𝐖agg𝐇flats,b,∗\)∈ℝNps,‖𝜶s,b‖1=1\.\\bm\{\\alpha\}\_\{s,b\}=\\operatorname\{softmax\}\\bigl\(\\mathbf\{W\}\_\{\\mathrm\{agg\}\}\\,\\mathbf\{H\}\_\{\\text\{flat\}\}^\{s,b,\*\}\\bigr\)\\in\\mathbb\{R\}^\{N\_\{\\mathrm\{ps\}\}\},\\quad\\\|\\bm\{\\alpha\}\_\{s,b\}\\\|\_\{1\}=1\.\(15\)
Smoothness and Gradient\-Aware Measures\.1\) Per\-step attention smoothing\. Let𝜶s,b\\bm\{\\alpha\}\_\{s,b\}be the per\-step, per\-sample attention over theNpsN\_\{\\mathrm\{ps\}\}parallel scales\. We define the smoothed attention as
𝜶s,bsmooth=\(1−β\)𝜶s,b\+β1Nps1,β∈\[0,1\]\.\\bm\{\\alpha\}\_\{s,b\}^\{\\,\\text\{smooth\}\}=\(1\-\\beta\)\\,\\bm\{\\alpha\}\_\{s,b\}\+\\beta\\,\\frac\{1\}\{N\_\{\\mathrm\{ps\}\}\}\\,\\mathbf\{1\},\\qquad\\beta\\in\[0,1\]\.\(16\)
2\) Gradient\-aware aggregation\. The weighted sum uses the smoothed weights:
𝐇weighteds,b=∑i=1Npsαs,b\(i\)smooth𝐇flats,b,i,\\mathbf\{H\}\_\{\\mathrm\{weighted\}\}^\{s,b\}=\\sum\_\{i=1\}^\{N\_\{\\mathrm\{ps\}\}\}\\alpha\_\{s,b\}^\{\(i\)\\,\\text\{smooth\}\}\\,\\mathbf\{H\}\_\{\\text\{flat\}\}^\{s,b,i\},\(17\)followed by a linear projection and a mean residual term:
𝐇aggs,b=𝐅𝐇weighteds,b\+𝐇means,b,𝐇outs,b=Proj\(𝐇aggs,b\)\.\\mathbf\{H\}\_\{\\mathrm\{agg\}\}^\{s,b\}=\\mathbf\{F\}\\,\\mathbf\{H\}\_\{\\mathrm\{weighted\}\}^\{s,b\}\+\\mathbf\{H\}\_\{\\text\{mean\}\}^\{s,b\},\\quad\\mathbf\{H\}\_\{\\mathrm\{out\}\}^\{s,b\}=\\mathrm\{Proj\}\\bigl\(\\mathbf\{H\}\_\{\\mathrm\{agg\}\}^\{s,b\}\\bigr\)\.\(18\)
### 4\.3cMoLLM Forward Pass: Pseudocode
[Algorithm1](https://arxiv.org/html/2607.22577#alg1)gives pseudocode for a single Transformer layer with a cMoLLM block \(no Top\-KK, no low\-rank\)\.
Algorithm 1Forward pass of a Transformer layer with cMoLLM1:Input:
𝐗∈ℝL×d\\mathbf\{X\}\\in\\mathbb\{R\}^\{L\\times d\}, number of streams
NN, kernels
\{𝐊k\}k=1N\\\{\\mathbf\{K\}\_\{k\}\\\}\_\{k=1\}^\{N\},
𝐖down\\mathbf\{W\}\_\{\\mathrm\{down\}\}
2:Self\-Attention:
𝐇←SelfAttention\(𝐗\)\\mathbf\{H\}\\leftarrow\\mathrm\{SelfAttention\}\(\\mathbf\{X\}\)
3:Gating logits:
𝐆←𝐇𝐖g\+𝐛g∈ℝL×N\\mathbf\{G\}\\leftarrow\\mathbf\{H\}\\mathbf\{W\}\_\{g\}\+\\mathbf\{b\}\_\{g\}\\in\\mathbb\{R\}^\{L\\times N\}
4:Soft mixture weights:
gi,k←softmax\(𝐆i,:\)kg\_\{i,k\}\\leftarrow\\mathrm\{softmax\}\(\\mathbf\{G\}\_\{i,:\}\)\_\{k\}for each token
ii, stream
kk
5:Mixed kernel:
𝐊~i←∑k=1Ngi,k𝐊k\\widetilde\{\\mathbf\{K\}\}\_\{i\}\\leftarrow\\sum\_\{k=1\}^\{N\}g\_\{i,k\}\\,\\mathbf\{K\}\_\{k\}
6:Dynamic1×11\{\\times\}1conv:
𝐔i←𝐇i𝐊~i⊤\\mathbf\{U\}\_\{i\}\\leftarrow\\mathbf\{H\}\_\{i\}\\,\\widetilde\{\\mathbf\{K\}\}\_\{i\}^\{\\top\}
7:Nonlinearity & down\-proj:
𝐘i←𝐖downσ\(𝐔i\)\\mathbf\{Y\}\_\{i\}\\leftarrow\\mathbf\{W\}\_\{\\mathrm\{down\}\}\\,\\sigma\(\\mathbf\{U\}\_\{i\}\)
8:Output:
𝐘\\mathbf\{Y\}\(residual \+ norm as usual\)
### 4\.4CNN as a Constrained Implicit MoE
The equivalence in[Theorem4\.1](https://arxiv.org/html/2607.22577#S4.Thmtheorem1)suggests that standard convolutional layers can be interpreted as special cases of MoE under strong constraints\. We formalize this intuition for a single linear or convolutional layer\.
###### Proposition 4\.4\(CNN as Implicit Constrained MoE\)\.
Consider a linear layer𝐲=𝐖𝐱\\mathbf\{y\}=\\mathbf\{W\}\\mathbf\{x\}with𝐖∈ℝdout×d\\mathbf\{W\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times d\}and activationσ\\sigmasatisfying[3\.2](https://arxiv.org/html/2607.22577#S3.Thmtheorem2)\. Let𝐰j⊤\\mathbf\{w\}\_\{j\}^\{\\top\}denote thejj\-th row of𝐖\\mathbf\{W\}and define per\-output “experts”Ej\(𝐱\)=σ\(𝐰j⊤𝐱\)E\_\{j\}\(\\mathbf\{x\}\)=\\sigma\(\\mathbf\{w\}\_\{j\}^\{\\top\}\\mathbf\{x\}\)forj∈\[dout\]j\\in\[d\_\{\\mathrm\{out\}\}\]\. Then the layer can be written as an MoE ofdoutd\_\{\\mathrm\{out\}\}experts with:
1. 1\.static, parameter\-free gating determined solely by the sign pattern of pre\-activations;
2. 2\.unnormalized gating weights given byσ\(𝐰j⊤𝐱\)\\sigma\(\\mathbf\{w\}\_\{j\}^\{\\top\}\\mathbf\{x\}\);
3. 3\.dense computation, as all experts are evaluated for every input\.
In particular, a standard CNN layer implements an implicit, highly constrained MoE over its output channels\.
###### Proof\.
Write the layer output as𝐲=σ\(𝐖𝐱\)\\mathbf\{y\}=\\sigma\(\\mathbf\{W\}\\mathbf\{x\}\), whereσ\\sigmais applied element\-wise\. For each output dimensionjj, we haveyj=σ\(𝐰j⊤𝐱\)y\_\{j\}=\\sigma\(\\mathbf\{w\}\_\{j\}^\{\\top\}\\mathbf\{x\}\), which we interpret as the output of expertEjE\_\{j\}\. Define the gating weightgj\(𝐱\)=1g\_\{j\}\(\\mathbf\{x\}\)=1for alljjand inputs\. Then the overall output can be written as
𝐲=∑j=1doutgj\(𝐱\)Ej\(𝐱\)=∑j=1doutEj\(𝐱\),\\mathbf\{y\}=\\sum\_\{j=1\}^\{d\_\{\\mathrm\{out\}\}\}g\_\{j\}\(\\mathbf\{x\}\)\\,E\_\{j\}\(\\mathbf\{x\}\)=\\sum\_\{j=1\}^\{d\_\{\\mathrm\{out\}\}\}E\_\{j\}\(\\mathbf\{x\}\),\(19\)which matches the standard layer\. The nonlinearityσ\\sigmainduces a data\-dependent binary mask𝕀\(𝐰j⊤𝐱\>0\)\\mathbb\{I\}\(\\mathbf\{w\}\_\{j\}^\{\\top\}\\mathbf\{x\}\>0\)on each expert, yielding a form of fixed, unnormalized gating as discussed in[Remark4\.3](https://arxiv.org/html/2607.22577#S4.Thmtheorem3)\. Unlike explicit MoE, all experts are always evaluated, so computation is dense\. ∎
This proposition connects classical CNNs to MoE: they sit at one end of a spectrum with static, dense, and unnormalized routing, whereas cMoLLM uses learned, normalized routing via dynamic convolutions \(soft mixture over all streams, no Top\-KK\)\. Viewed through this lens, standard CNNs are implicit, constrained MoE models, andcMoLLMcan be seen as relaxing these constraints—moving along a continuum from fixed, dense routing to learned, normalized, and capacity\-controllable routing over pipeline\-level streams\.
### 4\.5Gating Variants \(Convolution Types\)
We implement four gating variants corresponding to thesimple,context\_aware,multi\_head, andadaptiveoptions evaluated in[Section5](https://arxiv.org/html/2607.22577#S5)\.
#### 4\.5\.1Simple Gated Convolution
Two parallel convolutions are applied: one for feature extraction and one for gate generation\. The feature output is element\-wise multiplied by the sigmoid\-activated gate output\. Gating is local and token\-wise, with no global context\.
#### 4\.5\.2Context\-Aware Gated Convolution
A context analyzer \(global average pooling plus an MLP with 4:1 compression\) produces channel\-wise modulation weights from the full sequence; these are combined with the local gate convolution output\. Gating thus depends on both local patterns and sequence\-level statistics\.
#### 4\.5\.3Multi\-Head Gated Convolution
Multiple independent gate convolutions operate on the same input; each head produces gates for a disjoint subset of output channels\. Outputs are concatenated and fused by a learned layer\. Heads can specialize in different temporal scales or aspects of the input\.
#### 4\.5\.4Adaptive Gated Convolution
An adaptive network \(global pooling plus a two\-layer MLP\) outputs three bounded parameters \(sensitivity, bias, magnitude\) from global input statistics\. The gate is sigmoid\(scaled gate\-conv output \+ bias\) times magnitude\. Gating behavior is adjusted per sequence \(e\.g\., dynamic range or complexity\)\.
### 4\.6Training Strategies
Load Balancing Loss\.To encourage balanced use of streams, we add an auxiliary loss\(Feduset al\.,[2022](https://arxiv.org/html/2607.22577#bib.bib6)\):
ℒbal=α⋅N⋅∑k=1Npk2,\\mathcal\{L\}\_\{\\mathrm\{bal\}\}=\\alpha\\cdot N\\cdot\\sum\_\{k=1\}^\{N\}p\_\{k\}^\{2\},\(20\)wherepk=1L∑i=1Lgk\(𝐱i\)p\_\{k\}=\\frac\{1\}\{L\}\\sum\_\{i=1\}^\{L\}g\_\{k\}\(\\mathbf\{x\}\_\{i\}\)is the average gating probability for streamkkover the sequence\. Minimizing∑kpk2\\sum\_\{k\}p\_\{k\}^\{2\}discourages collapse onto a few streams\. We useα=0\.01\\alpha=0\.01in experiments\.
Total Loss\.The total training objective isℒ=ℒLM\+ℒbal\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{LM\}\}\+\\mathcal\{L\}\_\{\\mathrm\{bal\}\}; we use no kernel regularization, progressive sparsification, or low\-rank terms\.
### 4\.7Computational Complexity Analysis
We compare per\-token cost of dense FFN, Top\-KKMoE, and cMoLLM\. LetFLOPs\(⋅\)\\mathrm\{FLOPs\}\(\\cdot\)denote leading\-order FLOPs per token\.
Dense FFN:FLOPsdense≈2ddff\\mathrm\{FLOPs\}\_\{\\mathrm\{dense\}\}\\approx 2dd\_\{\\mathrm\{ff\}\}\(up\- and down\-projections\)\. Top\-KKMoE withNNexperts:FLOPsmoe≈2Kddff\+FLOPsgate\\mathrm\{FLOPs\}\_\{\\mathrm\{moe\}\}\\approx 2Kdd\_\{\\mathrm\{ff\}\}\+\\mathrm\{FLOPs\}\_\{\\mathrm\{gate\}\}\.
In cMoLLM we compute𝐔i=∑k=1Ngi,k\(𝐇i𝐊k⊤\)\\mathbf\{U\}\_\{i\}=\\sum\_\{k=1\}^\{N\}g\_\{i,k\}\(\\mathbf\{H\}\_\{i\}\\mathbf\{K\}\_\{k\}^\{\\top\}\)then𝐖downσ\(𝐔i\)\\mathbf\{W\}\_\{\\mathrm\{down\}\}\\sigma\(\\mathbf\{U\}\_\{i\}\)\. TheNNstream products costN⋅d⋅dffN\\cdot d\\cdot d\_\{\\mathrm\{ff\}\}, and the down\-projectiondff⋅dd\_\{\\mathrm\{ff\}\}\\cdot d\. Thus
FLOPscMoLLM≈\(N\+1\)ddff\+FLOPsgate\.\\mathrm\{FLOPs\}\_\{\\mathrm\{cMoLLM\}\}\\approx\(N\+1\)dd\_\{\\mathrm\{ff\}\}\+\\mathrm\{FLOPs\}\_\{\\mathrm\{gate\}\}\.\(21\)For smallNN\(e\.g\.,N=4N\{=\}4\), this is on the same order as dense, while affordingNNdistinct kernels and pipeline\-level mixture\. Convolution primitives often yield favorable throughput in practice\.
Combining the above, we can summarize the horizontal scaling behavior as follows\. Fix a compute budgetFLOPs0\\mathrm\{FLOPs\}\_\{0\}and FFN dimensiondffd\_\{\\mathrm\{ff\}\}\. Then there exists a constantCC\(absorbing gating overhead\) such that
FLOPscMoLLM≤C⋅FLOPsdensewheneverN≤C−1,\\mathrm\{FLOPs\}\_\{\\mathrm\{cMoLLM\}\}\\leq C\\cdot\\mathrm\{FLOPs\}\_\{\\mathrm\{dense\}\}\\quad\\text\{whenever \}N\\leq C\-1,\(22\)while the number of distinct kernels \(and thus the effective capacity of the mixture\) grows linearly withNN\. Top\-KKMoE trades compute for sparsity: increasingNNat fixedKKprimarily increases parameter count but leaves per\-token compute approximatelyO\(Kddff\)O\(Kdd\_\{\\mathrm\{ff\}\}\)\. cMoLLM therefore realizes a horizontal scaling law: for boundedNN, we scale capacity roughly linearly inNNwhile keeping per\-token compute within a constant factor of the dense baseline\.
Table 2:Full experimental results \(3 seeds; mean±\\pmstd\): validation loss, perplexity \(PPL\), GLUE \(%\), SQuAD v2 \(%\)\. “conv” = gating variant;nn= streams\.Bold= best \(SOTA\) in column\.Table 3:Scaling experiments \(3 seeds; mean±\\pmstd\) across model sizes \([Table4](https://arxiv.org/html/2607.22577#S4.T4)\)\.mh= multi\-head cMoLLM;base= dense baseline\.Bold= SOTA in column\.
### 4\.8Intuitive Summary
We summarize cMoLLM intuitively\. MoE\-style mixture linearly combines expert outputs;[Theorem4\.1](https://arxiv.org/html/2607.22577#S4.Thmtheorem1)shows this is equivalent to mixing1×11\{\\times\}1convolution kernels—i\.e\., dynamic convolution with variable kernels\. We approximate this with per\-stream kernels𝐊k\\mathbf\{K\}\_\{k\}and soft mixture weightsgk\(𝐱\)g\_\{k\}\(\\mathbf\{x\}\); no Top\-KK, no low\-rank\.
cMoLLM is implemented with standard1×11\{\\times\}1convolutions and a lightweight gating network\. Recipe: \(i\) keep the Transformer backbone; \(ii\) replace FFN blocks with cMoLLM blocks \(streams \+ gating\); \(iii\) tuneNNto trade capacity vs\. compute\. Experiments \([Section5](https://arxiv.org/html/2607.22577#S5)\) show better perplexity and GLUE at similar compute, with more stable stream utilization than ParaScale\- and AltUp\-style pipeline mixtures\. In addition, a cluster\-structured toy model in[AppendixB](https://arxiv.org/html/2607.22577#A2)formalizes when routing \(and by equivalence, dynamic convolution\) can achieve Bayes\-optimal performance while any single linear classifier suffers a nontrivial error floor, illustrating the potential benefits of cMoLLM\-style conditional computation\.
Table 4:Model scales used in scaling experiments\. Config 1 = small \(85M\), Config 2 = medium \(350M\), Config 3 = large \(760M\)\.Table 5:Shared training hyperparameters across all variants\.
## 5Experiments
We evaluate cMoLLM on language modeling, comparing to a dense GPT\-2 baseline under matched training setups\.
### 5\.1Experimental Setup
Model Configuration\.We adopt a GPT\-2\-style architecture\(Radfordet al\.,[2019](https://arxiv.org/html/2607.22577#bib.bib26)\)as our base model\. We evaluate at three scales \([Table4](https://arxiv.org/html/2607.22577#S4.T4)\): small \( 85M\), medium \( 350M\), and large \( 760M\) parameters, varying layers, hidden size, and attention heads\.[Table5](https://arxiv.org/html/2607.22577#S4.T5)gives shared hyperparameters\.
Training Data\.We train on FineWeb\(Penedoet al\.,[2024](https://arxiv.org/html/2607.22577#bib.bib28)\), a large\-scale curated web corpus\.
Baselines\.Our main baseline is the dense GPT\-2 model in which all layers are standard Transformer blocks\. For cMoLLM we keep the backbone identical and replace the FFN stack with convolutionally\-gated mixture streams\.
Scaling Factors\.We use scaling factorsn∈\{1,2,4,8\}n\\in\\\{1,2,4,8\\\}, wherennis the number of parallel cMoLLM streams\.
Gating Variants\.We evaluate gating types:simple,context\_aware,multi\_head, andadaptive\.
Reporting\.All experiments use3 random seeds\. We reportmean±\\pmstandard deviation \(std\)for all metrics in tables\.
### 5\.2Main Results
We report language modeling \(loss, perplexity\), GLUE, and SQuAD v2 accuracy for each cMoLLM variant \([Table2](https://arxiv.org/html/2607.22577#S4.T2)\)\.
Experiment figures\.[Figure3](https://arxiv.org/html/2607.22577#S4.F3)shows stream utilization and horizontal scaling: per\-stream gating weight distribution across layers \(left\) and validation loss / perplexity vs\. stream countnnor training steps \(right\)\. Together with[Tables2](https://arxiv.org/html/2607.22577#S4.T2)and[3](https://arxiv.org/html/2607.22577#S4.T3), the figure confirms balanced stream usage and gains asnnincreases over an optimal range, without collapse\.
### 5\.3Scaling Laws
We consolidate scaling behavior along two axes: \(i\)*horizontal*scaling in the number of streamsNNat fixed model size \(as in[Section4\.7](https://arxiv.org/html/2607.22577#S4.SS7)\), and \(ii\)*model\-size*scaling from 85M to 760M parameters\.
Horizontal scaling \(stream count\)\.As shown in[Section4\.7](https://arxiv.org/html/2607.22577#S4.SS7), cMoLLM satisfies a horizontal scaling law: for boundedNN, effective capacity grows roughly linearly inNNwhile per\-token FLOPs remain within a constant factor of the dense baseline \([Equation21](https://arxiv.org/html/2607.22577#S4.E21)\)\.[Figure3](https://arxiv.org/html/2607.22577#S4.F3)and[Table2](https://arxiv.org/html/2607.22577#S4.T2)confirm that validation loss and perplexity improve asnnincreases from 1 to 4–8, with best loss/PPL atmulti\_headn=8n\{=\}8; beyond that, some gating variants show mild over\-streaming \(e\.g\.,simpleatn=8n\{=\}8\), whilemulti\_headandadaptiveremain stable\.
Model\-size scaling\.[Table3](https://arxiv.org/html/2607.22577#S4.T3)reports results across the three scales in[Table4](https://arxiv.org/html/2607.22577#S4.T4)\(85M, 350M, 760M\)\. At every scale, multi\-head cMoLLM \(mh\) outperforms the dense baseline \(base\) on loss, PPL, GLUE, and SQuAD v2\. The gains are consistent with standard scaling: loss and perplexity decrease as model size increases, and the relative advantage of cMoLLM over the dense baseline is preserved \(e\.g\.,mh\-760Machieves SOTA across all four metrics\)\. This supports that the convolutionally\-gated pipeline mixture scales favorably with both stream count and parameter count under matched training setups\.
## 6Conclusion
We have presentedcMoLLM, a convolutionally\-gated mixture\-of\-LLMs that scales capacity at the pipeline level\. Our main theoretical contribution is the formal equivalence between MoE\-style mixture layers and dynamic convolutions with variable kernels, providing a unified framework for analyzing and designing sparse, conditional computation beyond FFN\-level MoE\. cMoLLM instantiates this via per\-stream kernels and soft gating—no Top\-KK, no low\-rank, no virtual tokens or auxiliary heads—yielding parameter\-efficient, stable pipeline\-level scaling\.
Experiments on GPT\-2–style models trained on FineWeb show that cMoLLM improves perplexity, GLUE, and SQuAD under matched compute, with better stream utilization and training stability than ParaScale\- and AltUp\-style pipeline mixtures\. Scaling laws \(horizontal scaling in stream count and model\-size scaling\) are analyzed in[Section5\.3](https://arxiv.org/html/2607.22577#S5.SS3)\. The convolution\-based design enables efficient implementation and favorable scaling in practice\. Extended results on downstream tasks and scaling are provided inLABEL:app:extended\.
Limitations and Future Work\.Experiments are at GPT\-2 scale; validation at 7B\+ parameters is needed\. Future work: \(i\) scale cMoLLM to larger models and distributed training; \(ii\) extend the MoE–convolution equivalence to attention; \(iii\) combine with retrieval or other conditional compute mechanisms\.
## Impact Statement
This work introduces cMoLLM, a convolutionally\-gated mixture\-of\-LLMs that scales capacity efficiently without auxiliary routing mechanisms\.Positive impacts:Parameter\-efficient scaling reduces training/inference costs and energy consumption; the convolution\-based design is hardware\-friendly\.Risks:Misuse concerns \(e\.g\., misleading content\) remain, though we introduce no new failure modes beyond existing LLMs\. Experiments are limited to GPT\-2 scale \( 760M parameters\); validation at 7B\+ is needed\.Summary:We believe the net effect is beneficial by improving LLM scaling efficiency\.
## References
- Anthropic \(2025\)Introducing Claude 4\.Note:[https://www\.anthropic\.com/news/claude\-4](https://www.anthropic.com/news/claude-4)Cited by:[§1](https://arxiv.org/html/2607.22577#S1.p1.1)\.
- C\. Baykal, D\. J\. Cutler, N\. Dikkala, N\. Ghosh, R\. Panigrahy, and X\. Wang \(2023\)Alternating updates for efficient transformers\.Cited by:[Appendix F](https://arxiv.org/html/2607.22577#A6.p6.1),[§1](https://arxiv.org/html/2607.22577#S1.p3.1),[§2](https://arxiv.org/html/2607.22577#S2.p1.2)\.
- Y\. Bengio, N\. Léonard, and A\. C\. Courville \(2013\)Estimating or propagating gradients through stochastic neurons for conditional computation\.CoRRabs/1308\.3432\.External Links:[Link](http://arxiv.org/abs/1308.3432),1308\.3432Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p2.1)\.
- T\. B\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell, S\. Agarwal, A\. Herbert\-Voss, G\. Krueger, T\. Henighan, R\. Child, A\. Ramesh, D\. M\. Ziegler, J\. Wu, C\. Winter, C\. Hesse, M\. Chen, E\. Sigler, M\. Litwin, S\. Gray, B\. Chess, J\. Clark, C\. Berner, S\. McCandlish, A\. Radford, I\. Sutskever, and D\. Amodei \(2020\)Language models are few\-shot learners\.Cited by:[§1](https://arxiv.org/html/2607.22577#S1.p1.1)\.
- W\. Cai, J\. Jiang, F\. Wang, J\. Tang, S\. Kim, and J\. Huang \(2025\)A survey on mixture of experts in large language models\.IEEE Trans\. Knowl\. Data Eng\.37\(7\),pp\. 3896–3915\.External Links:[Link](https://doi.org/10.1109/TKDE.2025.3554028),[Document](https://dx.doi.org/10.1109/TKDE.2025.3554028)Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p3.1)\.
- M\. Chen, B\. Hui, Z\. Cui, J\. Yang, D\. Liu, J\. Sun, J\. Lin, and Z\. Liu \(2025\)Parallel scaling law for language models\.CoRRabs/2505\.10475\.External Links:[Link](https://doi.org/10.48550/arXiv.2505.10475),[Document](https://dx.doi.org/10.48550/ARXIV.2505.10475),2505\.10475Cited by:[Appendix F](https://arxiv.org/html/2607.22577#A6.p6.1),[§2](https://arxiv.org/html/2607.22577#S2.p1.2)\.
- Y\. Chen, X\. Dai, M\. Liu, D\. Chen, L\. Yuan, and Z\. Liu \(2020\)Dynamic convolution: attention over convolution kernels\.In2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13\-19, 2020,pp\. 11027–11036\.External Links:[Link](https://openaccess.thecvf.com/content%5C_CVPR%5C_2020/html/Chen%5C_Dynamic%5C_Convolution%5C_Attention%5C_Over%5C_Convolution%5C_Kernels%5C_CVPR%5C_2020%5C_paper.html),[Document](https://dx.doi.org/10.1109/CVPR42600.2020.01104)Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p2.1)\.
- Z\. Chen, Y\. Deng, Y\. Wu, Q\. Gu, and Y\. Li \(2022\)Towards understanding the mixture\-of\-experts layer in deep learning\.Cited by:[§B\.2](https://arxiv.org/html/2607.22577#A2.SS2.p3.1),[§B\.3](https://arxiv.org/html/2607.22577#A2.SS3.p2.1),[Appendix B](https://arxiv.org/html/2607.22577#A2.p1.1)\.
- X\. Cheng, K\. Zeng, Z\. Cao, L\. Dai, W\. Gao, F\. Han, A\. Jian, F\. Hong, W\. Hu, Z\. Huang, D\. Kong, J\. Leng, Z\. Liao, P\. Liu, J\. Lin, X\. Ma, J\. Ruan, J\. Song, X\. Tan, R\. Xiao, W\. Yu, W\. Zhan, H\. Zhang, C\. Zhou, H\. Zhou, S\. Zheng, R\. Chen, S\. Chen, Z\. Chen, Y\. Dong, Y\. Fan, Y\. Fang, Y\. Gan, S\. Guo, Q\. He, C\. Hu, B\. Li, D\. Li, X\. Li, Y\. Li, C\. Liu, X\. Liu, J\. Lv, Q\. Ma, J\. Pan, C\. Qin, C\. Sun, W\. Sun, Z\. Wang, A\. Wuerkaixi, X\. Yang, F\. Yuan, Y\. Zhu, T\. Zhai, J\. Zhang, R\. Zhang, Y\. Xu, Y\. Zhao, Y\. Wang, X\. Cai, Y\. Hu, C\. Liu, L\. Pan, X\. Wang, B\. Xiao, W\. Yao, Q\. Zhou, and B\. Zhu \(2025\)Higher satisfaction, lower cost: A technical report on how llms revolutionize meituan’s intelligent interaction systems\.CoRRabs/2510\.13291\.External Links:[Link](https://doi.org/10.48550/arXiv.2510.13291),[Document](https://dx.doi.org/10.48550/ARXIV.2510.13291),2510\.13291Cited by:[§1](https://arxiv.org/html/2607.22577#S1.p1.1)\.
- N\. Du, Y\. Huang, A\. M\. Dai, S\. Tong, D\. Lepikhin, Y\. Xu, M\. Krikun, Y\. Zhou, A\. W\. Yu, O\. Firat, B\. Zoph, L\. Fedus, M\. P\. Bosma, Z\. Zhou, T\. Wang, Y\. E\. Wang, K\. Webster, M\. Pellat, K\. Robinson, K\. S\. Meier\-Hellstern, T\. Duke, L\. Dixon, K\. Zhang, Q\. V\. Le, Y\. Wu, Z\. Chen, and C\. Cui \(2022\)GLaM: efficient scaling of language models with mixture\-of\-experts\.InInternational Conference on Machine Learning, ICML 2022, 17\-23 July 2022, Baltimore, Maryland, USA,K\. Chaudhuri, S\. Jegelka, L\. Song, C\. Szepesvári, G\. Niu, and S\. Sabato \(Eds\.\),Proceedings of Machine Learning Research,pp\. 5547–5569\.External Links:[Link](https://proceedings.mlr.press/v162/du22c.html)Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p1.2)\.
- W\. Fedus, B\. Zoph, and N\. Shazeer \(2022\)Switch transformers: scaling to trillion parameter models with simple and efficient sparsity\.J\. Mach\. Learn\. Res\.23,pp\. 120:1–120:39\.External Links:[Link](https://jmlr.org/papers/v23/21-0998.html)Cited by:[§1](https://arxiv.org/html/2607.22577#S1.p2.1),[§2](https://arxiv.org/html/2607.22577#S2.p1.2),[§4\.6](https://arxiv.org/html/2607.22577#S4.SS6.p1.5)\.
- J\. Hoffmann, S\. Borgeaud, A\. Mensch, E\. Buchatskaya, T\. Cai, E\. Rutherford, D\. de Las Casas, L\. A\. Hendricks, J\. Welbl, A\. Clark, T\. Hennigan, E\. Noland, K\. Millican, G\. van den Driessche, B\. Damoc, A\. Guy, S\. Osindero, K\. Simonyan, E\. Elsen, J\. W\. Rae, O\. Vinyals, and L\. Sifre \(2022\)Training compute\-optimal large language models\.CoRRabs/2203\.15556\.External Links:[Link](https://doi.org/10.48550/arXiv.2203.15556),[Document](https://dx.doi.org/10.48550/ARXIV.2203.15556),2203\.15556Cited by:[§1](https://arxiv.org/html/2607.22577#S1.p1.1)\.
- N\. Houlsby, A\. Giurgiu, S\. Jastrzebski, B\. Morrone, Q\. de Laroussilhe, A\. Gesmundo, M\. Attariyan, and S\. Gelly \(2019\)Parameter\-efficient transfer learning for NLP\.InProceedings of the 36th International Conference on Machine Learning, ICML 2019, 9\-15 June 2019, Long Beach, California, USA,K\. Chaudhuri and R\. Salakhutdinov \(Eds\.\),Proceedings of Machine Learning Research,pp\. 2790–2799\.External Links:[Link](http://proceedings.mlr.press/v97/houlsby19a.html)Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p3.1)\.
- E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen \(2022\)LoRA: low\-rank adaptation of large language models\.External Links:[Link](https://openreview.net/forum?id=nZeVKeeFYf9)Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p3.1)\.
- X\. Jia, B\. D\. Brabandere, T\. Tuytelaars, and L\. V\. Gool \(2016\)Dynamic filter networks\.pp\. 667–675\.Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p2.1)\.
- J\. Kaplan, S\. McCandlish, T\. Henighan, T\. B\. Brown, B\. Chess, R\. Child, S\. Gray, A\. Radford, J\. Wu, and D\. Amodei \(2020\)Scaling laws for neural language models\.CoRRabs/2001\.08361\.External Links:[Link](https://arxiv.org/abs/2001.08361),2001\.08361Cited by:[§1](https://arxiv.org/html/2607.22577#S1.p1.1)\.
- D\. Lepikhin, H\. Lee, Y\. Xu, D\. Chen, O\. Firat, Y\. Huang, M\. Krikun, N\. Shazeer, and Z\. Chen \(2021\)GShard: scaling giant models with conditional computation and automatic sharding\.External Links:[Link](https://openreview.net/forum?id=qrwe7XHTmYb)Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p1.2)\.
- M\. Lewis, S\. Bhosale, T\. Dettmers, N\. Goyal, and L\. Zettlemoyer \(2021\)BASE layers: simplifying training of large, sparse models\.InProceedings of the 38th International Conference on Machine Learning, ICML 2021, 18\-24 July 2021, Virtual Event,M\. Meila and T\. Zhang \(Eds\.\),Proceedings of Machine Learning Research,pp\. 6265–6274\.External Links:[Link](http://proceedings.mlr.press/v139/lewis21a.html)Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p1.2)\.
- J\. Liu, P\. Tang, W\. Wang, Y\. Ren, X\. Hou, P\. Heng, M\. Guo, and C\. Li \(2026\)A survey on inference optimization techniques for mixture of experts models\.ACM Comput\. Surv\.58\(10\),pp\. 247:1–247:37\.External Links:[Link](https://doi.org/10.1145/3794845),[Document](https://dx.doi.org/10.1145/3794845)Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p3.1)\.
- N\. Ma, X\. Zhang, J\. Huang, and J\. Sun \(2020\)WeightNet: revisiting the design space of weight networks\.InComputer Vision \- ECCV 2020 \- 16th European Conference, Glasgow, UK, August 23\-28, 2020, Proceedings, Part XV,A\. Vedaldi, H\. Bischof, T\. Brox, and J\. Frahm \(Eds\.\),Lecture Notes in Computer Science,pp\. 776–792\.External Links:[Link](https://doi.org/10.1007/978-3-030-58555-6%5C_46),[Document](https://dx.doi.org/10.1007/978-3-030-58555-6%5F46)Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p2.1)\.
- M\. McGill and P\. Perona \(2017\)Deciding how to decide: dynamic routing in artificial neural networks\.InProceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6\-11 August 2017,D\. Precup and Y\. W\. Teh \(Eds\.\),Proceedings of Machine Learning Research,pp\. 2363–2372\.External Links:[Link](http://proceedings.mlr.press/v70/mcgill17a.html)Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p2.1)\.
- S\. Mu and S\. Lin \(2025\)A comprehensive survey of mixture\-of\-experts: algorithms, theory, and applications\.CoRRabs/2503\.07137\.External Links:[Link](https://doi.org/10.48550/arXiv.2503.07137),[Document](https://dx.doi.org/10.48550/ARXIV.2503.07137),2503\.07137Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p3.1)\.
- OpenAI \(2025\)Introducing GPT\-4\.1 in the API\.Note:[https://openai\.com/index/gpt\-4\-1/](https://openai.com/index/gpt-4-1/)Cited by:[§1](https://arxiv.org/html/2607.22577#S1.p1.1)\.
- G\. Penedo, H\. Kydlícek, L\. B\. Allal, A\. Lozhkov, M\. Mitchell, C\. A\. Raffel, L\. von Werra, and T\. Wolf \(2024\)The fineweb datasets: decanting the web for the finest text data at scale\.Cited by:[Appendix F](https://arxiv.org/html/2607.22577#A6.p2.1),[§5\.1](https://arxiv.org/html/2607.22577#S5.SS1.p2.1)\.
- A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, and I\. Sutskever \(2019\)Language models are unsupervised multitask learners\.External Links:[Link](https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf)Cited by:[Appendix F](https://arxiv.org/html/2607.22577#A6.p3.8),[§1](https://arxiv.org/html/2607.22577#S1.p1.1),[§5\.1](https://arxiv.org/html/2607.22577#S5.SS1.p1.1)\.
- S\. Rajbhandari, C\. Li, Z\. Yao, M\. Zhang, R\. Y\. Aminabadi, A\. A\. Awan, J\. Rasley, and Y\. He \(2022\)DeepSpeed\-moe: advancing mixture\-of\-experts inference and training to power next\-generation AI scale\.InInternational Conference on Machine Learning, ICML 2022, 17\-23 July 2022, Baltimore, Maryland, USA,K\. Chaudhuri, S\. Jegelka, L\. Song, C\. Szepesvári, G\. Niu, and S\. Sabato \(Eds\.\),Proceedings of Machine Learning Research,pp\. 18332–18346\.External Links:[Link](https://proceedings.mlr.press/v162/rajbhandari22a.html)Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p1.2)\.
- P\. Rajpurkar, J\. Zhang, K\. Lopyrev, and P\. Liang \(2016\)SQuAD: 100, 000\+ questions for machine comprehension of text\.pp\. 2383–2392\.External Links:[Link](https://doi.org/10.18653/v1/d16-1264),[Document](https://dx.doi.org/10.18653/V1/D16-1264)Cited by:[Appendix F](https://arxiv.org/html/2607.22577#A6.p2.1)\.
- C\. Riquelme, J\. Puigcerver, B\. Mustafa, M\. Neumann, R\. Jenatton, A\. S\. Pinto, D\. Keysers, and N\. Houlsby \(2021\)Scaling vision with sparse mixture of experts\.pp\. 8583–8595\.Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p1.2)\.
- C\. Rosenbaum, T\. Klinger, and M\. Riemer \(2018\)Routing networks: adaptive selection of non\-linear functions for multi\-task learning\.External Links:[Link](https://openreview.net/forum?id=ry8dvM-R-)Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p2.1)\.
- N\. Shazeer, A\. Mirhoseini, K\. Maziarz, A\. Davis, Q\. V\. Le, G\. E\. Hinton, and J\. Dean \(2017\)Outrageously large neural networks: the sparsely\-gated mixture\-of\-experts layer\.External Links:[Link](https://openreview.net/forum?id=B1ckMDqlg)Cited by:[§1](https://arxiv.org/html/2607.22577#S1.p2.1),[§2](https://arxiv.org/html/2607.22577#S2.p1.2),[§3\.3](https://arxiv.org/html/2607.22577#S3.SS3.p1.3)\.
- A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, L\. Kaiser, and I\. Polosukhin \(2017\)Attention is all you need\.pp\. 5998–6008\.Cited by:[§3\.2](https://arxiv.org/html/2607.22577#S3.SS2.p1.5)\.
- A\. Wang, A\. Singh, J\. Michael, F\. Hill, O\. Levy, and S\. R\. Bowman \(2019\)GLUE: A multi\-task benchmark and analysis platform for natural language understanding\.In7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6\-9, 2019,External Links:[Link](https://openreview.net/forum?id=rJ4km2R5t7)Cited by:[Appendix F](https://arxiv.org/html/2607.22577#A6.p2.1)\.
- F\. Wu, A\. Fan, A\. Baevski, Y\. N\. Dauphin, and M\. Auli \(2019\)Pay less attention with lightweight and dynamic convolutions\.External Links:[Link](https://openreview.net/forum?id=SkVhlh09tX)Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p2.1)\.
- B\. Yang, G\. Bender, Q\. V\. Le, and J\. Ngiam \(2019\)CondConv: conditionally parameterized convolutions for efficient inference\.pp\. 1305–1316\.Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p2.1)\.
- X\. Yang, L\. Li, A\. Wuerkaixi, X\. Cheng, C\. Liu, K\. Zeng, X\. Cai, and W\. Jiang \(2026\)Towards self\-robust llms: intrinsic prompt noise resistance via coipo\.CoRRabs/2603\.03314\.External Links:[Link](https://doi.org/10.48550/arXiv.2603.03314),[Document](https://dx.doi.org/10.48550/ARXIV.2603.03314),2603\.03314Cited by:[§1](https://arxiv.org/html/2607.22577#S1.p1.1)\.
- X\. Yang, B\. Tang, Y\. Wang, Z\. Ji, and W\. Jiang \(2025\)Can llms write fast system\-aware numerical computation code?\.InIEEE International Conference on Systems, Man, and Cybernetics, SMC 2025, Vienna, Austria, October 5\-8, 2025,pp\. 672–675\.External Links:[Link](https://doi.org/10.1109/SMC58881.2025.11343280),[Document](https://dx.doi.org/10.1109/SMC58881.2025.11343280)Cited by:[§1](https://arxiv.org/html/2607.22577#S1.p1.1)\.
- Y\. Zhang, J\. Zhang, Q\. Wang, and Z\. Zhong \(2020\)DyNet: dynamic convolution for accelerating convolutional neural networks\.CoRRabs/2004\.10694\.External Links:[Link](https://arxiv.org/abs/2004.10694),2004\.10694Cited by:[§2](https://arxiv.org/html/2607.22577#S2.p2.1)\.
\*
## Appendix AProof of[Theorem4\.1](https://arxiv.org/html/2607.22577#S4.Thmtheorem1): Extension to Nonlinear Experts
We extend the equivalence result to two\-layer experts with nonlinear activations\.
Setup\.LetEk\(𝐱\)=σ\(𝐱𝐖k,1⊤\)𝐖k,2⊤E\_\{k\}\(\\mathbf\{x\}\)=\\sigma\(\\mathbf\{x\}\\mathbf\{W\}\_\{k,1\}^\{\\top\}\)\\mathbf\{W\}\_\{k,2\}^\{\\top\}be a two\-layer expert with activationσ\\sigma\(e\.g\., ReLU, GELU\)\. The MoE output is:
MoE\(𝐱\)=∑k=1Ngk\(𝐱\)σ\(𝐱𝐖k,1⊤\)𝐖k,2⊤\.\\mathrm\{MoE\}\(\\mathbf\{x\}\)=\\sum\_\{k=1\}^\{N\}g\_\{k\}\(\\mathbf\{x\}\)\\,\\sigma\(\\mathbf\{x\}\\mathbf\{W\}\_\{k,1\}^\{\\top\}\)\\mathbf\{W\}\_\{k,2\}^\{\\top\}\.\(23\)
Linear Case\.Whenσ=id\\sigma=\\mathrm\{id\}\(identity\), we recover the result in[Theorem4\.1](https://arxiv.org/html/2607.22577#S4.Thmtheorem1):
∑kgk\(𝐱\)𝐱𝐖k,1⊤𝐖k,2⊤\\displaystyle\\sum\_\{k\}g\_\{k\}\(\\mathbf\{x\}\)\\,\\mathbf\{x\}\\mathbf\{W\}\_\{k,1\}^\{\\top\}\\mathbf\{W\}\_\{k,2\}^\{\\top\}=𝐱\(∑kgk\(𝐱\)𝐖k,2𝐖k,1\)⊤\\displaystyle=\\mathbf\{x\}\\Bigl\(\\textstyle\\sum\_\{k\}g\_\{k\}\(\\mathbf\{x\}\)\\,\\mathbf\{W\}\_\{k,2\}\\mathbf\{W\}\_\{k,1\}\\Bigr\)^\{\\\!\\top\}=Conv1×1\(𝐱;∑kgk\(𝐱\)𝐖k,2𝐖k,1\)\.\\displaystyle=\\mathrm\{Conv\}\_\{1\{\\times\}1\}\\\!\\Bigl\(\\mathbf\{x\};\\,\\textstyle\\sum\_\{k\}g\_\{k\}\(\\mathbf\{x\}\)\\,\\mathbf\{W\}\_\{k,2\}\\mathbf\{W\}\_\{k,1\}\\Bigr\)\.\(24\)
Nonlinear Case \(ReLU\)\.Forσ=ReLU\\sigma=\\mathrm\{ReLU\}, we can rewrite element\-wise:
ReLU\(z\)=𝕀\(z\>0\)⋅z,\\mathrm\{ReLU\}\(z\)=\\mathbb\{I\}\(z\>0\)\\cdot z,\(25\)where𝕀\(⋅\)\\mathbb\{I\}\(\\cdot\)is the indicator function\. The ReLU acts as a data\-dependent binary gate\.
Define the diagonal masking matrix:
𝐌k\(𝐱\)=diag\(𝕀\(𝐱𝐖k,1⊤\>0\)\)∈\{0,1\}dff×dff\.\\mathbf\{M\}\_\{k\}\(\\mathbf\{x\}\)=\\mathrm\{diag\}\\bigl\(\\mathbb\{I\}\(\\mathbf\{x\}\\mathbf\{W\}\_\{k,1\}^\{\\top\}\>0\)\\bigr\)\\in\\\{0,1\\\}^\{d\_\{\\mathrm\{ff\}\}\\times d\_\{\\mathrm\{ff\}\}\}\.\(26\)
Then:
σ\(𝐱𝐖k,1⊤\)=𝐌k\(𝐱\)\(𝐱𝐖k,1⊤\)=\(𝐱𝐖k,1⊤\)𝐌k\(𝐱\)\.\\sigma\(\\mathbf\{x\}\\mathbf\{W\}\_\{k,1\}^\{\\top\}\)=\\mathbf\{M\}\_\{k\}\(\\mathbf\{x\}\)\\,\(\\mathbf\{x\}\\mathbf\{W\}\_\{k,1\}^\{\\top\}\)=\(\\mathbf\{x\}\\mathbf\{W\}\_\{k,1\}^\{\\top\}\)\\,\\mathbf\{M\}\_\{k\}\(\\mathbf\{x\}\)\.\(27\)
The expert output becomes:
Ek\(𝐱\)=𝐱𝐖k,1⊤𝐌k\(𝐱\)𝐖k,2⊤=𝐱\(𝐖k,2𝐌k\(𝐱\)𝐖k,1\)⊤\.E\_\{k\}\(\\mathbf\{x\}\)=\\mathbf\{x\}\\,\\mathbf\{W\}\_\{k,1\}^\{\\top\}\\,\\mathbf\{M\}\_\{k\}\(\\mathbf\{x\}\)\\,\\mathbf\{W\}\_\{k,2\}^\{\\top\}=\\mathbf\{x\}\\,\\bigl\(\\mathbf\{W\}\_\{k,2\}\\,\\mathbf\{M\}\_\{k\}\(\\mathbf\{x\}\)\\,\\mathbf\{W\}\_\{k,1\}\\bigr\)^\{\\top\}\.\(28\)
This is a convolution with an input\-dependent effective kernel:
𝐊keff\(𝐱\)=𝐖k,2𝐌k\(𝐱\)𝐖k,1\.\\mathbf\{K\}\_\{k\}^\{\\mathrm\{eff\}\}\(\\mathbf\{x\}\)=\\mathbf\{W\}\_\{k,2\}\\,\\mathbf\{M\}\_\{k\}\(\\mathbf\{x\}\)\\,\\mathbf\{W\}\_\{k,1\}\.\(29\)
The full MoE output:
MoE\(𝐱\)=∑k=1Ngk\(𝐱\)𝐱\(𝐊keff\(𝐱\)\)⊤=𝐱\(∑k=1Ngk\(𝐱\)𝐊keff\(𝐱\)\)⊤\.\\mathrm\{MoE\}\(\\mathbf\{x\}\)=\\sum\_\{k=1\}^\{N\}g\_\{k\}\(\\mathbf\{x\}\)\\,\\mathbf\{x\}\\,\\bigl\(\\mathbf\{K\}\_\{k\}^\{\\mathrm\{eff\}\}\(\\mathbf\{x\}\)\\bigr\)^\{\\top\}=\\mathbf\{x\}\\Bigl\(\\textstyle\\sum\_\{k=1\}^\{N\}g\_\{k\}\(\\mathbf\{x\}\)\\,\\mathbf\{K\}\_\{k\}^\{\\mathrm\{eff\}\}\(\\mathbf\{x\}\)\\Bigr\)^\{\\\!\\top\}\.\(30\)
This can be viewed as a dynamic convolution withNNinput\-dependent kernels, where both the routing weightsgk\(𝐱\)g\_\{k\}\(\\mathbf\{x\}\)and the effective kernels𝐊keff\(𝐱\)\\mathbf\{K\}\_\{k\}^\{\\mathrm\{eff\}\}\(\\mathbf\{x\}\)depend on the input\.
Interpretation\.The nonlinear case reveals that:
- •Each expert implements a gated linear unit \(GLU\)\-like computation with data\-dependent masking\.
- •The MoE output is a weighted combination of masked convolutions\.
- •Both outer gating \(gkg\_\{k\}\) and inner gating \(𝐌k\\mathbf\{M\}\_\{k\}\) are input\-dependent, creating a two\-level conditional computation structure\.
This analysis motivates designing cMoLLM with explicit control over both routing and activation patterns\.
□\\square
## Appendix BA Cluster\-Structured Toy Model
We provide a simple toy model illustrating when MoE\-style routing \(and hence cMoLLM\) is provably beneficial for cluster\-structured data, in the spirit ofChenet al\.\([2022](https://arxiv.org/html/2607.22577#bib.bib12)\)\.
### B\.1Problem Setup
Consider a binary classification problem with input spaceℝd\\mathbb\{R\}^\{d\}and two clusters per class\. Letμ1,\+,μ1,−,μ2,\+,μ2,−∈ℝd\\mu\_\{1,\+\},\\mu\_\{1,\-\},\\mu\_\{2,\+\},\\mu\_\{2,\-\}\\in\\mathbb\{R\}^\{d\}be four mean vectors, and letσ2I\\sigma^\{2\}Ibe a shared covariance\. We define the data distribution as
𝐱∣y=\+1∼12𝒩\(μ1,\+,σ2I\)\+12𝒩\(μ2,\+,σ2I\),\\displaystyle\\mathbf\{x\}\\mid y=\+1\\sim\\tfrac\{1\}\{2\}\\mathcal\{N\}\(\\mu\_\{1,\+\},\\sigma^\{2\}I\)\+\\tfrac\{1\}\{2\}\\mathcal\{N\}\(\\mu\_\{2,\+\},\\sigma^\{2\}I\),\(31\)𝐱∣y=−1∼12𝒩\(μ1,−,σ2I\)\+12𝒩\(μ2,−,σ2I\),\\displaystyle\\mathbf\{x\}\\mid y=\-1\\sim\\tfrac\{1\}\{2\}\\mathcal\{N\}\(\\mu\_\{1,\-\},\\sigma^\{2\}I\)\+\\tfrac\{1\}\{2\}\\mathcal\{N\}\(\\mu\_\{2,\-\},\\sigma^\{2\}I\),\(32\)with priorℙ\(y=\+1\)=ℙ\(y=−1\)=12\\mathbb\{P\}\(y=\+1\)=\\mathbb\{P\}\(y=\-1\)=\\tfrac\{1\}\{2\}\. We assume that clusters\(μ1,\+,μ1,−\)\(\\mu\_\{1,\+\},\\mu\_\{1,\-\}\)and\(μ2,\+,μ2,−\)\(\\mu\_\{2,\+\},\\mu\_\{2,\-\}\)are well separated and lie in different “regions” of the input space\.
Formally, suppose there exists a unit vector𝐮∈ℝd\\mathbf\{u\}\\in\\mathbb\{R\}^\{d\}and scalarsa<ba<bsuch that
𝐮⊤μ1,\+,𝐮⊤μ1,−\\displaystyle\\mathbf\{u\}^\{\\top\}\\mu\_\{1,\+\},\\mathbf\{u\}^\{\\top\}\\mu\_\{1,\-\}<a−γ,\\displaystyle<a\-\\gamma,\(33\)𝐮⊤μ2,\+,𝐮⊤μ2,−\\displaystyle\\mathbf\{u\}^\{\\top\}\\mu\_\{2,\+\},\\mathbf\{u\}^\{\\top\}\\mu\_\{2,\-\}\>b\+γ,\\displaystyle\>b\+\\gamma,\(34\)for some marginγ\>0\\gamma\>0, and that the Bayes\-optimal decision boundary within each cluster\-pair is approximately linear in a \(potentially different\) direction\.
### B\.2Expressivity of a Single Linear Classifier
Letflin\(𝐱\)=sign\(𝐰⊤𝐱\)f\_\{\\mathrm\{lin\}\}\(\\mathbf\{x\}\)=\\mathrm\{sign\}\(\\mathbf\{w\}^\{\\top\}\\mathbf\{x\}\)be a linear classifier\. Because the two class\-conditional mixtures overlap across clusters, a single hyperplane must simultaneously separate both\(μ1,\+,μ1,−\)\(\\mu\_\{1,\+\},\\mu\_\{1,\-\}\)and\(μ2,\+,μ2,−\)\(\\mu\_\{2,\+\},\\mu\_\{2,\-\}\)\. When the optimal separating directions within the two regions are sufficiently misaligned, any single hyperplane incurs a non\-negligible error\.
The following statement summarizes this limitation at a high level\.
###### Proposition B\.1\(Limitation of Single Linear Classifier\)\.
Under the cluster separation conditions above, suppose that the optimal separating directions for the first and second cluster\-pairs differ by an angle of at leastθ0\>0\\theta\_\{0\}\>0\. Then there exists a constantε0=ε0\(θ0,γ,σ\)\>0\\varepsilon\_\{0\}=\\varepsilon\_\{0\}\(\\theta\_\{0\},\\gamma,\\sigma\)\>0such that any linear classifierflin\(𝐱\)=sign\(𝐰⊤𝐱\)f\_\{\\mathrm\{lin\}\}\(\\mathbf\{x\}\)=\\mathrm\{sign\}\(\\mathbf\{w\}^\{\\top\}\\mathbf\{x\}\)has misclassification error at leastε0\\varepsilon\_\{0\}\.
The proof follows standard arguments for mixtures of Gaussians with incompatible linear separators and is omitted for brevity; seeChenet al\.\([2022](https://arxiv.org/html/2607.22577#bib.bib12)\)\.
### B\.3Two\-Expert MoE/cMoLLM Construction
Now consider a two\-expert MoE \(or cMoLLM\) model with a simple router that partitions space along direction𝐮\\mathbf\{u\}:
g1\(𝐱\)=𝕀\(𝐮⊤𝐱≤τ\),g2\(𝐱\)=𝕀\(𝐮⊤𝐱\>τ\),g\_\{1\}\(\\mathbf\{x\}\)=\\mathbb\{I\}\(\\mathbf\{u\}^\{\\top\}\\mathbf\{x\}\\leq\\tau\),\\quad g\_\{2\}\(\\mathbf\{x\}\)=\\mathbb\{I\}\(\\mathbf\{u\}^\{\\top\}\\mathbf\{x\}\>\\tau\),\(35\)for some thresholdτ∈\(a,b\)\\tau\\in\(a,b\)\. Let each expert be a linear classifier specialized to one region:
E1\(𝐱\)\\displaystyle E\_\{1\}\(\\mathbf\{x\}\)=sign\(𝐰1⊤𝐱\),\\displaystyle=\\mathrm\{sign\}\(\\mathbf\{w\}\_\{1\}^\{\\top\}\\mathbf\{x\}\),\(36\)E2\(𝐱\)\\displaystyle E\_\{2\}\(\\mathbf\{x\}\)=sign\(𝐰2⊤𝐱\),\\displaystyle=\\mathrm\{sign\}\(\\mathbf\{w\}\_\{2\}^\{\\top\}\\mathbf\{x\}\),\(37\)with𝐰1\\mathbf\{w\}\_\{1\}optimized for the first cluster\-pair and𝐰2\\mathbf\{w\}\_\{2\}for the second\. The overall prediction is
fmoe\(𝐱\)=g1\(𝐱\)E1\(𝐱\)\+g2\(𝐱\)E2\(𝐱\)\.f\_\{\\mathrm\{moe\}\}\(\\mathbf\{x\}\)=g\_\{1\}\(\\mathbf\{x\}\)E\_\{1\}\(\\mathbf\{x\}\)\+g\_\{2\}\(\\mathbf\{x\}\)E\_\{2\}\(\\mathbf\{x\}\)\.\(38\)
###### Theorem B\.2\(Toy Cluster Model: Benefit of Routing\)\.
In the cluster\-structured setting above, there exist parameters\(𝐮,τ,𝐰1,𝐰2\)\(\\mathbf\{u\},\\tau,\\mathbf\{w\}\_\{1\},\\mathbf\{w\}\_\{2\}\)such that the two\-expert MoE \(or cMoLLM\) classifierfmoef\_\{\\mathrm\{moe\}\}attains misclassification error arbitrarily close to the Bayes\-optimal error as the marginγ\\gammaincreases and the covarianceσ2\\sigma^\{2\}decreases, while any single linear classifierflinf\_\{\\mathrm\{lin\}\}suffers error at leastε0\>0\\varepsilon\_\{0\}\>0as in[PropositionB\.1](https://arxiv.org/html/2607.22577#A2.Thmtheorem1)\.
###### Proof Sketch\.
Because𝐮\\mathbf\{u\}separates the two cluster regions with marginγ\\gamma, choosingτ∈\(a,b\)\\tau\\in\(a,b\)ensures that, with high probability \(increasing asγ/σ\\gamma/\\sigmagrows\), samples from\(μ1,\+,μ1,−\)\(\\mu\_\{1,\+\},\\mu\_\{1,\-\}\)fall into the first region and samples from\(μ2,\+,μ2,−\)\(\\mu\_\{2,\+\},\\mu\_\{2,\-\}\)into the second\. Within each region, the problem reduces to a two\-component Gaussian mixture that is linearly separable by an appropriately chosen𝐰i\\mathbf\{w\}\_\{i\}\. Thus, the routed classifierfmoef\_\{\\mathrm\{moe\}\}can implement the Bayes\-optimal decision rule up to an exponentially small error inγ2/σ2\\gamma^\{2\}/\\sigma^\{2\}\. On the other hand,[PropositionB\.1](https://arxiv.org/html/2607.22577#A2.Thmtheorem1)implies that any single hyperplane must compromise between the two misaligned regions, incurring a constant error floorε0\\varepsilon\_\{0\}even asγ\\gammagrows\. Hence, for sufficiently well\-separated clusters, the routed MoE/cMoLLM strictly outperforms any single linear classifier\. ∎
This toy example provides a simple, concrete setting where conditional computation \(and by equivalence, dynamic convolution\) is provably beneficial, aligning with the broader conclusions ofChenet al\.\([2022](https://arxiv.org/html/2607.22577#bib.bib12)\)\.
## Appendix CImplementation Details
Kernel Generator\.Each stream kernel𝐊k\\mathbf\{K\}\_\{k\}is initialized with Xavier initialization\. We use no shared base or low\-rank factorization in the main experiments\.
Gating Network\.The gating network is a single linear layer𝐖g∈ℝd×N\\mathbf\{W\}\_\{g\}\\in\\mathbb\{R\}^\{d\\times N\}\. Forcontext\_awaregating, we add layer normalization before the projection\. Formulti\_headgating, we useH=4H=4heads with dimensiond/Hd/Heach\.
Optimization\.We use AdamW withβ1=0\.9\\beta\_\{1\}=0\.9,β2=0\.95\\beta\_\{2\}=0\.95, weight decay0\.10\.1, and a cosine learning\-rate schedule with linear warmup over the first 2000 steps\.
## Appendix DHyperparameter Sensitivity
Table 6:Sensitivity to key hyperparameters \(3 seeds; best value over mean±\\pmstd\)\.[Table6](https://arxiv.org/html/2607.22577#A4.T6)summarizes the tested ranges and best values\. cMoLLM is relatively robust to hyperparameter choices\. The load\-balancing coefficientα\\alphahas the largest impact; values too small encourage stream collapse, while values too large hurt performance\.
## Appendix EPlain Language Summary
This paper introducescMoLLM, a new way to scale large language models \(LLMs\) more efficiently\. Traditional LLMs activate all parameters for every token, making training and inference expensive as models grow larger\. Mixture\-of\-Experts \(MoE\) approaches try to address this by routing tokens to only a subset of “expert” networks, but most prior work applies this only to feed\-forward layers and uses discrete Top\-KKrouting, which can be unstable\. Our key insight is that MoE\-style mixture layers can be exactly rewritten as dynamic convolutions: each expert corresponds to a convolution kernel, and the router mixes these kernels based on the input\. We use this equivalence to design cMoLLM, which applies mixture routing to the entire LLM pipeline \(not just feed\-forward layers\) using soft, fully differentiable gating—no Top\-KK, no virtual tokens, no auxiliary prediction branches\. cMoLLM maintains a small set of parallel “streams,” each with its own convolution kernel; a lightweight gating network produces input\-dependent mixture weights, and the mixed kernel is applied via standard pointwise convolution, yielding parameter\-efficient capacity scaling with stable training dynamics\. On GPT\-2–style models trained on FineWeb, cMoLLM improves perplexity and GLUE accuracy under matched compute, with better stream utilization and training stability than prior pipeline\-level scaling methods like ParaScale and AltUp\. The convolution\-based design enables efficient implementation and favorable scaling in practice\.
## Appendix FReproducibility Statement
Our implementation is based on PyTorch and follows standard Transformer architectures\. The codebase will be made publicly available upon acceptance, including full model implementation \(cMoLLM blocks, gating networks, training loop\), training scripts with hyperparameter configurations, evaluation scripts for language modeling and downstream tasks, and preprocessing scripts for the FineWeb dataset\.
We use FineWeb\(Penedoet al\.,[2024](https://arxiv.org/html/2607.22577#bib.bib28)\)for pretraining, which is publicly available\. For downstream evaluation, we use standard benchmarks GLUE\(Wanget al\.,[2019](https://arxiv.org/html/2607.22577#bib.bib29)\)and SQuAD\(Rajpurkaret al\.,[2016](https://arxiv.org/html/2607.22577#bib.bib31)\), all of which are publicly available\.
All hyperparameters are reported in[Sections5\.1](https://arxiv.org/html/2607.22577#S5.SS1),[C](https://arxiv.org/html/2607.22577#A3)and[D](https://arxiv.org/html/2607.22577#A4)\. Implementation details and hyperparameter sensitivity analysis are provided in[AppendicesC](https://arxiv.org/html/2607.22577#A3)and[D](https://arxiv.org/html/2607.22577#A4)\. We use a GPT\-2–style architecture\(Radfordet al\.,[2019](https://arxiv.org/html/2607.22577#bib.bib26)\)with 12 layers, hidden dimension 768, FFN dimension 3072, sequence length 4096, and learning rate6×10−56\\times 10^\{\-5\}with cosine schedule and 2000\-step warmup\. We use AdamW optimizer withβ1=0\.9\\beta\_\{1\}=0\.9,β2=0\.95\\beta\_\{2\}=0\.95, and weight decay0\.10\.1\. For cMoLLM, we evaluate stream countsN∈\{1,2,4,8\}N\\in\\\{1,2,4,8\\\}; best validation loss in[Table2](https://arxiv.org/html/2607.22577#S4.T2)is achieved atn=8n\{=\}8formulti\_head\. We useN=8N\{=\}8as a representative configuration in sensitivity and scaling tables\. Load\-balancing coefficientα=0\.01\\alpha=0\.01\.
Experiments were conducted on NVIDIA A100 GPUs\. Training a single cMoLLM model \(N=4N=4\) for the reported experiments requires approximately 8 A100 GPU\-days\. Exact hardware specifications and software versions \(PyTorch, CUDA, etc\.\) will be documented in the code repository\.
Language modeling metrics \(loss, perplexity\) are computed on validation splits\. Downstream task evaluation follows standard protocols: GLUE use development set accuracy; SQuAD uses F1 score\. All experimental results use3 random seedsand are reported asmean±\\pmstandard deviation \(std\)throughout the paper \(main tables, scaling tables, and appendix\)\. A plain language summary is provided in[AppendixE](https://arxiv.org/html/2607.22577#A5)\.
We compare against a dense GPT\-2 baseline \(standard Transformer\), ParaScale\(Chenet al\.,[2025](https://arxiv.org/html/2607.22577#bib.bib10)\)\(reimplemented with virtual token streams\), and AltUp\(Baykalet al\.,[2023](https://arxiv.org/html/2607.22577#bib.bib11)\)\(reimplemented with auxiliary prediction branch\)\. All baselines use identical training data, hyperparameters where applicable, and evaluation protocols to ensure fair comparison\. Reproducibility details are provided in[AppendixF](https://arxiv.org/html/2607.22577#A6)\.
## Appendix GUse of LLM
The authors used generative AI tools \(Grammarly, ChatGPT\) only for grammar checking and language polishing\. All technical content, experimental design, data analysis, and conclusions were generated and verified by the human authors\. The use of AI tools does not affect the originality or authorship of this work\.Similar Articles
Scaling LLMs horizontally: hidden-state coupling without weight modification [R]
Residual Coupling (RC) connects frozen language models in parallel using lightweight learned linear bridges, enabling horizontal scaling without weight modification. It reduces perplexity by up to 80.7% compared to MoE and improves accuracy on TruthfulQA by 9.1 percentage points.
Accelerating Dense LLMs via L0-regularized Mixture-of-Experts
This paper proposes L0-MoE, a lightweight Mixture-of-Experts approach using L0-regularization to accelerate dense Large Language Models with up to 2.5x speedup while maintaining competitive performance.
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
SMELT is a method that loops middle layers in Mixture-of-Experts Transformers to improve training efficiency and downstream performance while matching compute, parameter, and cache budgets, leading to faster loss reduction and practical gains.
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
A research paper presenting scaling laws and design principles for mixture-of-experts diffusion language models, trained as LLaDA MoE v2 (30B-A3B) on 23.5T tokens, approaching Qwen3 on several benchmarks with fewer pretraining tokens.
Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025)
A deep dive into vLLM's architecture and components for high-throughput LLM inference, covering scheduling, paged attention, continuous batching, advanced features, scaling, serving, and benchmarking.