mixture-of-experts

Tag

Cards List
#mixture-of-experts

Motif 3 (314B A13B, NVFP4 available) seems good!? What are your experiences with it so far?

Reddit r/LocalLLaMA · 3h ago

A newly released large MoE model, Motif 3 (314B with 13B active parameters, NVFP4 available), seems promising; the author asks for community experiences.

0 favorites 0 likes
#mixture-of-experts

Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts

arXiv cs.LG · 10h ago Cached

This paper introduces MOSAIC, a framework that jointly optimizes sparse Mixture-of-Experts model architecture and hardware systems for large-scale pretraining, showing that compute-optimal sparsity is not necessarily cluster-optimal when MFU, communication costs, and parallel layouts are considered.

0 favorites 0 likes
#mixture-of-experts

Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation

arXiv cs.LG · 10h ago Cached

This paper proposes UniF-MoE, a unified framework for token-adaptive Mixture-of-Experts computation that first shares reusable computation across experts and then routes the remaining residual demand, improving performance while reducing activated computation, latency, and memory on DomainBed and GLUE benchmarks.

0 favorites 0 likes
#mixture-of-experts

@heyshrutimishra: NVIDIA just dropped Nemotron 3.5 Lightning 30 billion parameters. Only 3 billion active. Built for the execution layer …

X AI KOLs Following · yesterday Cached

NVIDIA released Nemotron 3.5 Lightning, a 30B-parameter MoE model with only 3B active parameters, optimized for agent execution tasks. It claims faster, cheaper tool calls and agent execution while staying fully open-source under OpenMDW-1.1.

0 favorites 0 likes
#mixture-of-experts

NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI

NVIDIA Blog · yesterday Cached

NVIDIA announced Nemotron 3.5 Lightning, a 30B mixture-of-experts open model optimized for high-volume agentic AI workloads, alongside NeMo Switchyard, an open-source library for intelligent model routing across heterogeneous model ecosystems.

0 favorites 0 likes
#mixture-of-experts

DoGMA: A Central-Dogma-Guided Foundation Model for Multi-Omics Alignment and Multi-Task Learning in Oncology

arXiv cs.LG · yesterday Cached

DoGMA is a central-dogma-guided foundation model for pan-cancer multi-omics analysis, using a Transformer-MoE architecture with directed attention to align DNA-RNA-protein flows and pretraining via masked hierarchical omics reconstruction. It shows strong performance across cancer representation learning, survival prediction, and metastasis prediction tasks.

0 favorites 0 likes
#mixture-of-experts

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism

arXiv cs.CL · yesterday Cached

A technical survey of Mixture-of-Experts architectures in LLMs, organizing evolution along expert granularity, topology, routing, load balancing, and execution, and proposing complementary views of architectural milestones and control planes.

0 favorites 0 likes
#mixture-of-experts

EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference

arXiv cs.LG · yesterday Cached

This paper presents EasyBalance, a cross-layer load balancing strategy for distributed Mixture-of-Experts (MoE) inference that schedules and jointly executes workloads from different layers to mitigate GPU idling without modifying expert-device mappings, reducing idle time by over 40% in experiments.

0 favorites 0 likes
#mixture-of-experts

When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes

arXiv cs.LG · yesterday Cached

This paper investigates how trace-driven evaluation can mislead assessments of MoE expert caching, identifying replay semantics, workload contamination, and operating regimes as confounding axes that can reverse policy rankings. After correcting these issues, it shows that a large offline-optimal gap overstates the gains actually recovered by lightweight causal caching mechanisms.

0 favorites 0 likes
#mixture-of-experts

Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models

arXiv cs.LG · yesterday Cached

This paper proposes using router weight sensitivity under lightweight fine-tuning (e.g., LoRA) to identify and prune experts in Mixture-of-Experts models, enabling significant memory and latency reductions with minimal accuracy loss.

0 favorites 0 likes
#mixture-of-experts

Shape Mutating Expert Compression:LorExperts and BTExperts

arXiv cs.LG · yesterday Cached

This paper introduces LorExperts and BTExperts, router-preserving compression methods for Mixture-of-Experts LLMs that cluster experts and represent non-dominant members as low-rank corrections, improving compression quality over prior methods like D2-MoE.

0 favorites 0 likes
#mixture-of-experts

TEXAS: Task-Expert-Aware Supervision for Downstream Mixture-of-Experts LLM Adaptation

arXiv cs.CL · 2d ago Cached

Presents TEXAS, a method for downstream adaptation of Mixture-of-Experts LLMs that discovers task-relevant experts via correctness-conditioned activations and applies token-level supervision allocation, improving performance across multiple benchmarks.

0 favorites 0 likes
#mixture-of-experts

Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrast

arXiv cs.AI · 2d ago Cached

The paper presents Contribution-Contrast (CoCo), a novel response-level interpretation method for Mixture-of-Experts reward models, which captures routing and preference behavior more faithfully than routing-weight-based approaches.

0 favorites 0 likes
#mixture-of-experts

EntropyMoE: Entropy-Aware Sparse Expert Routing for Tokenizer-Free LLMs

arXiv cs.AI · 2d ago Cached

EntropyMoE introduces an entropy-aware Mixture-of-Experts architecture for tokenizer-free LLMs, using dynamic byte patches as routing units to enable sparse conditional computation. Experiments show it achieves the lowest held-out bits-per-byte among baselines while maintaining downstream accuracy.

0 favorites 0 likes
#mixture-of-experts

Motif 3: Technical Report

Hugging Face Daily Papers · 2d ago Cached

Motif 3 is a 314B-parameter Mixture-of-Experts language model with 13.2B active parameters per token, featuring Grouped Differential Latent Attention and trained on 12.5T tokens, demonstrating competitive performance across reasoning, coding, and long-context tasks.

0 favorites 0 likes
#mixture-of-experts

UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models

Hugging Face Daily Papers · 3d ago Cached

UniMoMo compresses MoE-based recommendation models by merging experts based on functional behavior and routing traffic, preserving quality while speeding up inference.

0 favorites 0 likes
#mixture-of-experts

@rohanpaul_ai: Alibaba released Qwen3.8-Max, a 2.4 trillion parameter model that activates only about 95 bn parameters per token. A th…

X AI KOLs Timeline · 4d ago Cached

Alibaba released Qwen3.8-Max, a 2.4 trillion-parameter sparse MoE model with 95B active parameters per token, 1M token context, and strong agentic and benchmark results, including autonomously coding for days, circuit design, and outperforming rivals on Terminal Bench and PaperBench.

0 favorites 0 likes
#mixture-of-experts

An Emerging Retail Portfolio Management Application: Personalized, Tax-Aware Reinforcement Learning with Natural Language Goals

arXiv cs.LG · 5d ago Cached

Presents an emerging retail portfolio management application that uses personalized, tax-aware reinforcement learning with natural language goal input, featuring a three-phase pipeline and integration with live brokerage APIs.

0 favorites 0 likes
#mixture-of-experts

How Modalities Learn Together (49 minute read)

TLDR AI · 5d ago Cached

A systematic study from Meta FAIR, Reality Labs, and Oxford on multimodal pretraining, revealing asymmetric knowledge flow between modalities, synergy vs. competition dynamics, the benefits of early unification, and efficient training recipes validated with 13.5B MoE models.

0 favorites 0 likes
#mixture-of-experts

MESH: Memory-Efficient Sinkhorn Optimization for Mixture-of-Experts Training

arXiv cs.LG · 6d ago Cached

This paper introduces MESH, a memory-efficient Sinkhorn-based optimizer for Mixture-of-Experts (MoE) training that restores temporal momentum without storing full optimizer state, reducing memory by 62.5% while maintaining competitive evaluation loss compared to AdamW.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback