moe

Tag

Cards List
#moe

Qwen3.8 Max (0902) scores 45 on the Artificial Analysis Intelligence Index, up 5 points in a month and back on top of China's leaderboard, nosing out GLM-5.3 (44.9) and Kimi K3 (43.8)

Reddit r/LocalLLaMA ↗ · 2026-09-16

Qwen3.8 Max has been upgraded and now leads China's AI leaderboard with a score of 45 on the Artificial Analysis Intelligence Index, surpassing GLM-5.3 and Kimi K3 after a 30-day improvement to the 2.4T MoE model.

0 favorites 0 likes
#moe

10%+ performance improvement on MoE ssd-streaming with expert-lookahead

Reddit r/LocalLLaMA ↗ · 2026-09-15

An implementation of expert lookahead achieves over 10% performance improvement for MoE models running on low-memory devices using slotstream, with additional gains from a correction model.

0 favorites 0 likes
#moe

dealignai/DeepSeek-V4.1-Flash-UNCENSORED-FP8

Hugging Face Models Trending ↗ · 2026-09-10 Cached

A permanently modified version of DeepSeek-V4.1-Flash with surgically removed safety guardrails, retaining full capabilities including vision, reasoning, and tools, and demonstrating 100% compliance on HarmBench evaluations.

0 favorites 0 likes
#moe

Proposed architecture for inferencing sparse MOE models increasing Active parameters using layered + linear decay. Succinct reasoning without any model training or fine tune. [p]

Reddit r/MachineLearning ↗ · 2026-09-06

A llama.cpp implementation that expands MoE expert routing beyond native top-K during inference using adaptive thresholds and layer-specific linear decay, without requiring model retraining or fine-tuning.

0 favorites 0 likes
#moe

Expert expansion with llama.cpp

Reddit r/LocalLLaMA ↗ · 2026-09-06

A developer built a custom branch of llama.cpp that implements expert expansion for Mixture-of-Experts models, tested it on Metal, and is seeking cross-platform feedback.

0 favorites 0 likes
#moe

@ssahoo_: As per Artificial Analysis' updated Evals, our K2-Horizon-400B-Moe model still beats Thinking Machine's Inkling model. …

X AI KOLs Timeline ↗ · 2026-09-05 Cached

IFM AI's K2-Horizon-400B-MoE continues to outperform Thinking Machine's Inkling on Artificial Analysis benchmarks, according to a team member highlighting the year-old lab's progress.

0 favorites 0 likes
#moe

Routing Is Not Enough: Diagnosing Intra-Adapter Subspace Contention in MoE+LoRA Fine-Tuning

arXiv cs.LG ↗ · 2026-09-04 Cached

This paper diagnoses intra-adapter contention in MoE+LoRA fine-tuning and introduces SpawnLoRA to dynamically add sub-adapters, reducing negative transfer across domains.

0 favorites 0 likes
#moe

@percyliang: A year ago, David was the only FTE on Marin. Today, thanks to Open Athena, Marin has 10 FTE. Read his post to better un…

X AI KOLs Following ↗ · 2026-09-03 Cached

Marin has grown from one to ten full-time employees since joining Open Athena a year ago. The team is currently in the middle of a large-scale 535B/A23B Mixture of Experts model training run.

0 favorites 0 likes
#moe

Ant launches Ling-3.0-flash-Fin for finance workflows; OpenRouter access is free for one month

Reddit r/ArtificialInteligence ↗ · 2026-08-28

Ant's Ling team has released Ling-3.0-flash-Fin, a finance-enhanced AI model designed for financial workflows, with free access on OpenRouter for one month and plans to open-source the weights.

0 favorites 0 likes
#moe

ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration

arXiv cs.LG ↗ · 2026-08-27 Cached

ExFold is a unified training-free framework that accelerates MoE model inference by folding excluded expert contributions into retained experts, achieving up to 1.41× speedup while maintaining high quality.

0 favorites 0 likes
#moe

@threerouter_com: Tonight at 23:00, Alibaba is going to open-source Qwen3.8-Flash-Next on ModelScope, and the news is accurate. But honestly, releasing it without benchmarks, I'm skeptical. The official reason is to adapt for the complete Qwen4 family, meaning 'you try it first, and we'll add the report card later.' The architecture looks impressive—Qwen…

X AI KOLs Timeline ↗ · 2026-08-26 Cached

Alibaba will open-source the Qwen3.8-Flash-Next model on ModelScope at 23:00 tonight, featuring Qwen4's GDN hybrid layers and sparse attention architecture, but lacking benchmark data.

0 favorites 0 likes
#moe

Quants impact for agentic use and local LLMs?

Reddit r/AI_Agents ↗ · 2026-08-23

The author shares findings from testing quantization impacts on local LLMs for agentic use, revealing that many quants are statistically indistinguishable, MoEs are less affected than dense models, and significant degradation occurs below Q4 quantization.

0 favorites 0 likes
#moe

@PyTorch: Leverage a PyTorch-native fine-tuning library with Day-0 Hugging Face checkpoint support for the recently released Alib…

X AI KOLs Following ↗ · 2026-08-20 Cached

PyTorch provides a native fine-tuning library with Day-0 Hugging Face checkpoint support for Alibaba's open-weight Qwen3.8-2.4T-A95B model, enabling efficient training and deployment on NVIDIA systems with configurable reasoning.

0 favorites 0 likes
#moe

@with_gene2626: wtf? 397b nvfp4 is a 3 or 4 spark setup... showing better benchmarks then all of the current open weight models? anyone…

X AI KOLs Following ↗ · 2026-08-19 Cached

Ornith-1.5, a family of open-source LLMs, is introduced with variants up to 397B MoE, achieving state-of-the-art performance among comparable models and rivaling Claude Opus in benchmarks.

0 favorites 0 likes
#moe

@MinLiBuilds: https://x.com/MinLiBuilds/status/2089338660386992295

X AI KOLs Timeline ↗ · 2026-08-17 Cached

This article compares the performance of NVIDIA DGX Spark and a modified RTX 4090 in locally deploying the Qwen3.8-27B and Ling-3.0-flash models, providing benchmark data and purchase recommendations.

0 favorites 0 likes
#moe

@natolambert: Happy to see @NVIDIAAI released their expert models for MOPD. Starting to make research there much more accessible (tho…

X AI KOLs Following ↗ · 2026-08-14 Cached

NVIDIA released their expert models for MOPD, making AI research more accessible with the NVIDIA-Nemotron-Labs-Teacher-Competition-Coding model designed for competitive programming and code reasoning.

0 favorites 0 likes
#moe

Are we getting Qwen 3.8 35-A3B?

Reddit r/LocalLLaMA ↗ · 2026-08-14

Speculation about the upcoming Qwen 3.8 release, questioning whether it will be a dense 27B model or a MoE variant like the previous 35B-A3B, with discussion of performance implications for local hardware.

0 favorites 0 likes
#moe

Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation

arXiv cs.AI ↗ · 2026-08-14 Cached

This paper introduces Dual-Flow Transformers, an architecture that decouples prefill and decode computation by adding an auxiliary flow for continuation prediction while sharing weights and the primary KV cache, enabling phase-specific compute allocation and improved efficiency.

0 favorites 0 likes
#moe

12GB VRAM gang, what's our plan?

Reddit r/LocalLLaMA ↗ · 2026-08-11

Discussion about running LLMs on 12GB VRAM, noting current focus on dense models like Muse Glimmer 30B and Qwen 3.8 27B, and questioning whether upgrading to 24GB VRAM is needed.

0 favorites 0 likes
#moe

I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examples

Reddit r/LocalLLaMA ↗ · 2026-08-11

The author trained a 40M-parameter connector on 100K examples to give DeepSeek V4 Flash basic vision, freezing both the language model and MoonViT image encoder, demonstrating a low-cost approach to turning a text-only MoE into a basic VLM.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback