expert-routing

Tag

Cards List
#expert-routing

A Declarative-Procedural Perspective on Expert Routing in Bilingual Mixture-of-Experts Language Models

arXiv cs.CL · 2026-08-18 Cached

This paper investigates whether bilingual Mixture-of-Experts (MoE) language models develop linguistically structured expert routing. It finds that interpretable linguistic organization emerges within MoE routing patterns, and that curriculum training influences specialization in language balance.

0 favorites 0 likes
#expert-routing

TEXAS: Task-Expert-Aware Supervision for Downstream Mixture-of-Experts LLM Adaptation

arXiv cs.CL · 2026-08-10 Cached

Presents TEXAS, a method for downstream adaptation of Mixture-of-Experts LLMs that discovers task-relevant experts via correctness-conditioned activations and applies token-level supervision allocation, improving performance across multiple benchmarks.

0 favorites 0 likes
#expert-routing

TIER-MoE: Trust-Informed Expert Routing via Conditional Modality Risk for Multimodal Fusion in Biomedical Classification

arXiv cs.LG · 2026-07-31 Cached

This paper introduces TIER-MoE, a risk-guided subspace mixture-of-experts model for multimodal biomedical classification that estimates sample-specific modality reliability from out-of-fold predictions and routes modalities to experts, improving performance and calibration on four public datasets.

0 favorites 0 likes
#expert-routing

OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis

Hugging Face Daily Papers · 2026-07-27 Cached

OPERA proposes a multi-agent ensemble framework that treats expert weight assignment as an offline policy learning problem for universal biomedical image analysis, enabling test-time adaptation without retraining and consistently improving performance across 9 datasets and 30+ baselines.

0 favorites 0 likes
#expert-routing

PADD: Path-Aligned Decompression Distillation for Non-Router Teacher to Guide MoE Student Learning

arXiv cs.CL · 2026-06-10 Cached

Proposes PADD, a framework for distilling knowledge from dense teachers into mixture-of-experts (MoE) students, addressing the challenge of learning routing policies without a router in the teacher. The method involves four stages and shows improvements on mathematical reasoning benchmarks.

0 favorites 0 likes
#expert-routing

dMoE: dLLMs with Learnable Block Experts

Hugging Face Daily Papers · 2026-05-29 Cached

This paper proposes dMoE, a block-level mixture-of-experts framework for diffusion large language models that aggregates token-level expert distributions into block-level routing, reducing activated experts and memory usage while maintaining performance.

0 favorites 0 likes
#expert-routing

Notes on pretraining parallelisms and failed training runs (12 minute read)

TLDR AI · 2026-05-18 Cached

A technical deep-dive into common causes of failed pretraining runs in large language models, including causality-breaking issues in expert routing and numerical precision bugs, with examples from Llama 4, Gemini 2 Pro, and GPT-4.

0 favorites 0 likes
#expert-routing

Expert Routing for Communication-Efficient MoE via Finite Expert Banks

arXiv cs.LG · 2026-05-08 Cached

The paper introduces an information-theoretic framework for communication-efficient expert routing in sparse mixture-of-experts models, treating the gate as a stochastic channel and deriving practical mutual information estimators to analyze accuracy-rate tradeoffs over finite expert banks.

0 favorites 0 likes
← Back to home

Submit Feedback