@AdinaYakup: Step-3.7-Flash New VL model from @StepFun_ai 198B / 11B active - MoE 256K context 3 reasoning level Up to 400 tokens/sec
Summary
StepFun releases Step-3.7-Flash, a new large vision-language MoE model with 198B parameters (11B active), 256K context, and up to 400 tokens/sec inference speed.
View Cached Full Text
Cached at: 05/29/26, 02:11 PM
Step-3.7-Flash 🔥 New VL model from @StepFun_ai
✨ 198B / 11B active - MoE ✨ 256K context ✨ 3 reasoning level ✨ Up to 400 tokens/sec 🤯 https://t.co/zNgbnsMEsl
Similar Articles
stepfun-ai/Step-3.7-Flash
Step 3.7 Flash is a 198B-parameter sparse MoE vision-language model with 11B active parameters per token, supporting 256k context and three reasoning levels, designed for high-throughput agentic workflows.
@nathanhabib1011: Step-3.7-Flash from @StepFun_ai is a silent winner. Super impressive results, the best model under 500B params on HF le…
Step-3.7-Flash from StepFun_ai is highlighted as the best model under 500B parameters on Hugging Face leaderboards, with strong multimodal performance.
stepfun-ai/Step-3.7-Flash-GGUF
StepFun releases GGUF quantizations of their 198B-parameter sparse MoE vision-language model Step-3.7-Flash, enabling local deployment with up to 256K context and selectable reasoning levels.
@modal: Day 0 support for Step 3.7 Flash on Modal. - 198B parameter MoE with 11B active - 256K context - 3 reasoning levels - N…
Modal announces day 0 support for the Step 3.7 Flash AI model, a 198B parameter MoE with 11B active parameters, 256K context, three reasoning levels, and native image and video understanding.
@modal: DeepSeek-V4-Flash has 284B total parameters with 13B active per token. Combined with a hybrid compressed attention mech…
DeepSeek-V4-Flash is a 284B-parameter MoE model with 13B active parameters per token, featuring a hybrid compressed attention mechanism that reduces KV cache needs for 1M-token contexts. It can be served with SGLang on Modal for fast decoding on a single B300.