moe-model

Tag

Cards List
#moe-model

XingChen-AGI/Xing4.0-29B-A4B MoE

Reddit r/LocalLLaMA · 7h ago

Xing4.0-29B-A4B is a next-generation MoE large language model developed by China Telecom, featuring 29B total parameters with 4B active per token, native support for 256K context length, and optimization for Ascend NPU with agent-oriented architecture for complex engineering tasks.

0 favorites 0 likes
#moe-model

Running Qwen 3.8 next on 16vram+32ram - A useful/fun post for the gpu poors

Reddit r/LocalLLaMA · 3d ago

A Reddit user shares how they successfully ran the Qwen 3.8 Next MoE model on a system with 16GB VRAM and 32GB RAM using aggressive quantization and specific llama.cpp settings, achieving usable performance for large models on limited hardware.

0 favorites 0 likes
#moe-model

DeepSeek V4.1 Flash is getting surprisingly close to GPT-5.6 Sol territory, while being absurdly cheap

Reddit r/singularity · 2026-09-10

DeepSeek has released V4.1 Flash, a 552B MoE model with efficient active parameters, achieving performance close to GPT-5.6 Sol at a much lower cost.

0 favorites 0 likes
#moe-model

DeepSeek-V4-Flash-Vision-Exp (285B MoE) on 10-12x RTX 3090 — spec decoding, vision

Reddit r/LocalLLaMA · 2026-09-09

Running the DeepSeek-V4-Flash-Vision-Exp 285B MoE model on 10-12x RTX 3090 GPUs achieves over 60-120 tok/s decode speeds with vision and tool support, fully documented for reproducibility.

0 favorites 0 likes
#moe-model

@AlmustyFX: This is the kind of local AI test that actually matters. A 7.9B MoE running at 152 tok/s on an M4 Pro with 64GB unified…

X AI KOLs Timeline · 2026-08-29 Cached

A 7.9B MoE model runs at 152 tokens per second on an M4 Pro with 64GB unified memory, enabling offline processing of sensitive contract data and demonstrating the practical use of local AI.

0 favorites 0 likes
#moe-model

Today I hit 181 toks/s (aggregate) on Qwen3.8-Flash-Next on 2x DGX Sparks

Reddit r/LocalLLaMA · 2026-08-28

Achieved 181 tokens per second aggregate throughput on the Qwen3.8-Flash-Next model using a 2x DGX Spark cluster with optimizations like NVMe mapping and speculative decoding.

0 favorites 0 likes
#moe-model

llama.cpp slower on P-Cores than on E-Cores with MoE Model and GPU+CPU offloading?

Reddit r/LocalLLaMA · 2026-07-24

Observation that llama.cpp runs slower on P-cores than E-cores when running Mixture of Experts models with GPU+CPU offloading.

0 favorites 0 likes
#moe-model

Qwen3.5 122B is the best?

Reddit r/LocalLLaMA · 2026-07-09

A user shares their experience comparing several large language models (Qwen, Gemma) on complex tool-calling tasks, finding Qwen3.5 122B the most reliable, while criticizing smaller MoE models for instability.

0 favorites 0 likes
#moe-model

@sudoingX: anyone running a 16gb card, stop scrolling. @pupposandro and @davideciffa got qwen 35b-a3b down to 13.3gb, measured on …

X AI KOLs Timeline · 2026-06-10 Cached

A technique called luce spark allows Qwen 35B-a3B MoE model to run on a 16GB GPU (like RTX 3090) by learning which experts are frequently used and streaming the rest from RAM, achieving ~100 tok/s without VRAM bottleneck.

0 favorites 0 likes
#moe-model

Llama.cpp B9406 MTP mmproj fix

Reddit r/LocalLLaMA · 2026-05-29

Llama.cpp release B9406 fixes a crash (GGML_ASSERT) when using MTP with MoE vision models like Qwen3.6-35B-A3B.

0 favorites 0 likes
#moe-model

Re. what ever happened to Cohere’s Command-A series of models?

Reddit r/LocalLLaMA · 2026-05-20

Cohere launches Command A+, its first Mixture-of-Experts model, released under Apache 2.0 with efficient quantization for 1-2 GPU deployment, prioritizing practicality and open access for developers.

0 favorites 0 likes
← Back to home

Submit Feedback