sparse-mixture-of-experts

Tag

Cards List
#sparse-mixture-of-experts

MedMix: Specialization-Consistent Federated Sparse MoEs under Modality Heterogeneity

arXiv cs.LG · 2d ago Cached

MedMix is a semantic-alignment framework for federated multimodal sparse Mixture-of-Experts that addresses modality heterogeneity by coordinating routing and expert specialization, achieving improved performance in medical AI datasets.

0 favorites 0 likes
#sparse-mixture-of-experts

Sparse By Design (5 minute read)

TLDR AI · 2026-07-21 Cached

Moonshot's Kimi K3, a 2.8 trillion parameter open weights model with 896 experts (16 active per token), exemplifies the trend of scaling total parameters while holding active compute constant, and uses attention compression to reduce KV cache size, making frontier inference more accessible but with high storage costs.

0 favorites 0 likes
#sparse-mixture-of-experts

silx-ai/Quasar-Preview

Hugging Face Models Trending · 2026-06-08 Cached

SILX AI releases Quasar-Preview, an 18B parameter MoE foundation model with 2B active parameters and experimental 5M-token context, built on a hybrid recurrent/attention architecture and designed for decentralized training via Bittensor SN24.

0 favorites 0 likes
#sparse-mixture-of-experts

HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts

arXiv cs.LG · 2026-05-15 Cached

HodgeCover uses higher-order topological coverage to compress sparse Mixture-of-Experts layers by addressing irreducible mergeability barriers that pairwise signals miss, matching state-of-the-art baselines on expert reduction and leading on aggressive compression.

0 favorites 0 likes
← Back to home

Submit Feedback