moE

Tag

Cards List
#moE

@AdinaYakup: Ling 3.0 tiny a 7.9B/1.3B hybrid reasoning MoE https://huggingface.co/inclusionAI/Ling-3.0-tiny…

X AI KOLs Timeline · 2026-08-11 Cached

InclusionAI introduces Ling-3.0-tiny, a 7.9B-parameter hybrid reasoning MoE model with only 1.3B active parameters per token, optimized for efficient local and edge deployment.

0 favorites 0 likes
#moE

inclusionAI/Ling-3.0-flash · Hugging Face

Reddit r/LocalLLaMA · 2026-08-04 Cached

inclusionAI released Ling-3.0-flash, a native hybrid reasoning model with 124B total/5.1B active parameters using a hybrid linear attention architecture (KDA+MLA) and sparse MoE. It matches or outperforms its 1T-class predecessor Ring-2.6-1T while being far more compute-efficient, with built-in agentic and long-context optimizations.

0 favorites 0 likes
#moE

@kimmonismus: Holy, China strikes again: Qwen3.8-Max reportedly worked autonomously for 16 days while costing 80% less than GPT-5.6 S…

X AI KOLs Timeline · 2026-08-03 Cached

Alibaba announces Qwen3.8-Max, a 2.4T-parameter MoE frontier model with open weights coming next week, claiming autonomous operation for 16 days and significantly lower cost than GPT-5.6 Sol and Claude Fable 5.

0 favorites 0 likes
#moE

Kwaipilot/KAT-Coder-V2.5-Dev

Hugging Face Models Trending · 2026-07-23 Cached

KAT-Coder-V2.5-Dev is an open-weight MoE coding model with 35B total parameters (3B active), achieving state-of-the-art results on agentic coding benchmarks through SFT and RL training.

0 favorites 0 likes
#moE

@Tech2Wild: Running GLM-5.2 at home the FULL 744B, all 256 experts, UNPRUNED across 4× NVIDIA DGX Spark (GB10). 200K context · MTP …

X AI KOLs Following · 2026-07-05 Cached

A detailed recipe for running the unpruned GLM-5.2 model (744B parameters, 256 experts) across 4 NVIDIA DGX Spark nodes with 200K context, achieving up to 60.5 tok/s aggregate. Includes performance benchmarks, credits, and patches.

0 favorites 0 likes
#moE

README_EN.md · openpangu/openPangu-2.0-Flash at main

Reddit r/LocalLLaMA · 2026-07-01 Cached

openPangu-2.0-Flash is a 92B-parameter MoE model with 6B activated parameters, trained on Ascend, featuring 512k context length and fast thinking capabilities. It achieves strong performance on reasoning and coding benchmarks, using architectural innovations like MLA attention and multi-token prediction.

0 favorites 0 likes
#moE

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training

Hugging Face Daily Papers · 2026-05-09 Cached

This paper explores structured pruning and knowledge distillation techniques for compressing large Mixture-of-Experts (MoE) models during pre-training. It demonstrates that progressive pruning and combined distillation strategies, such as multi-token prediction distillation, improve downstream performance, exemplified by compressing Qwen3-Next-80A3B to a more efficient 23A2B model.

0 favorites 0 likes
← Back to home

Submit Feedback