Naive-N0.5-Flash - 309B-A15.5B
Summary
Naive-N0.5-Flash is an open-weight 309B Mixture-of-Experts AI model with 15.5B active parameters, optimized for coding and AI research, featuring a native 1M-token context window and high inference speeds.
View Cached Full Text
Cached at: 09/27/26, 07:48 PM
NaiveAI/Naive-N0.5-Flash · Hugging Face
Source: https://huggingface.co/NaiveAI/Naive-N0.5-Flash
Building Frontier AI with AI
https://huggingface.co/NaiveAI/Naive-N0.5-Flash#introductionIntroduction
Naive-N0.5-Flash is an open-weight309B MoE model with 15.5B active parameters, built forcoding and AI R&D. It supports anative 1M-token context windowthrough a hybrid of Sliding-Window Attention (SWA) and lightweight DeepSeek Sparse Attention (DSA), with no full-attention layers.
https://huggingface.co/NaiveAI/Naive-N0.5-Flash#key-featuresKey Features
- **Native 1M context, without full attention.**Naive-N0.5-Flash combines Sliding-Window Attention (SWA) and lightweight DeepSeek Sparse Attention (DSA) with GQA4 at a predominantly 5:1 SWA–DSA layout. The entire network remains local or sparse, with no full-attention layers.
- **AI-optimized inference up to 2,000 tokens/s.**NaiveRT, our inference system for Naive-N0.5-Flash, was built and optimized through AI-centered R&D. It combines mega-kernel fusion, Programmatic Dependent Launch (PDL), and speculative decoding, delivering 50 tokens/s per user in Standard mode and up to 2,000 tokens/s in Ultrafast mode. See the NaiveRT case study in thetechnical blogfor the implementation and optimization process.
- **Open weights and API.**Model weights and inference code are released under the MIT license. API access will also be provided, with pricing set at $0.10 / $0.40 / $0.01 per million tokens for input, output, and cache reads, respectively.
https://huggingface.co/NaiveAI/Naive-N0.5-Flash#model-architectureModel Architecture
PropertySpecificationArchitectureMixture-of-Experts (MoE)Total parameters309BActive parameters15.5BContext lengthNative 1M tokensTransformer layers48Attention-layer composition39 SWA layers + 9 DSA layersAttention mechanismHybrid SWA–DSASWA window128 tokensDSA token selectionTop 2,048 tokens for backbone attentionDSA KV groups4 (GQA4)Indexer query heads16
https://huggingface.co/NaiveAI/Naive-N0.5-Flash#hybrid-swadsa-attentionHybrid SWA–DSA Attention
Naive-N0.5-Flash builds on the open-weight MiMo-V2.5 base model, which has a simple architecture with strong foundational capabilities in world knowledge and deep research. Most layers use Sliding-Window Attention (SWA), whose per-token decoding cost does not grow with context length, while a small number of global-attention layers preserve long-range information. At million-token context lengths, however, these global-attention layers account for much of the decoding overhead.
Naive-N0.5-Flash replaces the global-attention layers with DeepSeek Sparse Attention (DSA). A lightweight indexer scores the full history, while the backbone computes attention only over a selected subset of tokens. Although the indexer still scans the full history and the full KV cache is retained, sparse attention substantially reduces attention computation and memory access. Adapting the model to this new attention structure was one objective of continued pretraining.
Figure 1. The hybrid attention stack and DSA module.
The network consists ofeight six-layer modules. A standard module contains five SWA layers followed by one DSA layer, with the first layer of the first module also replaced by DSA. SWA uses a128-token window, while DSA selects thetop 2,048 tokensfor backbone attention. Both attention types incorporate sink bias.
Unlike the original MLA-based DSA implementation, Naive-N0.5-Flash replaces MLA with grouped-query attention (GQA) usingfour KV groups. For the architecture design process and indexer efficiency comparison, see model architecture in thetechnical blog.
https://huggingface.co/NaiveAI/Naive-N0.5-Flash#training-overviewTraining Overview
Following the architectural changes, Naive-N0.5-Flash completed 3.25T tokens of multi-stage training with a native 1M-token context window: 50B tokens of Indexer Warmup, 3T tokens of Sparse Attention Training, and 200B tokens of Learning Rate Decay. This process adapted the model to its new sparse attention architecture while substantially improving its AI R&D and coding capabilities. See thetechnical blogfor training details.
https://huggingface.co/NaiveAI/Naive-N0.5-Flash#evaluation-resultsEvaluation Results
Figure 2. Coding and agentic task results. Naive-N0.5-Flash is highlighted in yellow.
Figure 3. AI research and systems optimization results. Metric directions are indicated in the figure.
Evaluation setup and metric notes**Evaluation setup.**Unless otherwise noted, our evaluations of Naive-N0.5-Flash use Claude Code 2.1.207 with a 1M-token context window, temperature 1.0, and top-p 0.95. The harness exposes only basic file I/O and Bash tools.
Sources for reported benchmark scores are as follows:
- GLM-5.3 and GLM-5.3-Flash:GLM-5.3 blogandGLM-5.3-Flash blog, respectively.
- Kimi-K3:Kimi-K3 model page.
- Qwen-3.8-Max:Qwen-3.8-Max blog.
- Hy4-preview:Hy4-preview model page.
- DeepSeek-V4.1-Flash:DeepSeek-V4.1-Flash model page.
- Step-5-preview:Step-5-Preview-BF16 model page.
- Fable-5 (w/ fallback):GLM-5.3 blog.
- **SWE-Bench Pro:**GPT-5.6-Sol, Opus-5, and Opus-5.5 scores are drawn from theGPT-5.6 blogand theClaude Opus 5.5 System Card.
- **DeepSWE v1.1:**The Muse-Spark-1.3 score comes from itsMuse-Spark-1.3 blog. GPT-5.6-Sol and Opus-5 scores come from theDeepSWE v1.1 leaderboard. The Opus-5.5 score comes from theClaude Opus 5.5 System Card.
- **Terminal-Bench 2.1:**Muse-Spark-1.3, GPT-5.6-Sol, and Opus-5 scores come from theMuse-Spark-1.3 blog. The GPT-6-Astra score comes from theTerminal-Bench 2.1 leaderboard.
- **ALE-CLI:**GPT-5.6-Sol, GPT-6-Astra, Muse-Spark-1.3, Opus-5, and Opus-5.5 scores come from theALE-CLI leaderboard.
- **FrontierSWE v1:**We calculate the Dominance score using the competing systems’ results as of August 23, 2026.
- **ProgramBench:**We report the Almost@1 score. GPT-5.6-Sol and Opus-5 scores come from theProgramBench leaderboard.
- **MLE-bench-30:**Gemini-3.5-Flash, Gemini-3.6-Flash, Grok-4.5, and GPT-5.6-Luna scores come from theGemini 3.6 Flash model card. Following its evaluation protocol, we report the average position score of Naive-N0.5-Flash.
- **PaperBench:**MiniMax M3, Opus-4.7, GPT-5.5, and Gemini-3.1-Pro scores come from theMiniMax M3 model page.
- **SOL-ExecBench, NanoChat AutoResearch, and NanoGPT SpeedRun:**Naive-N0.5-Flash scores were obtained using our in-house AutoResearch harness. Recursive Superintelligence Inc. scores come from itsresearch article. Following the SOL-ExecBench update, we use its updated score from theleaderboard.
https://huggingface.co/NaiveAI/Naive-N0.5-Flash#deploymentDeployment
Naive-N0.5-Flash supports FP8 mixed-precision inference. For general use, we recommend setting the sampling parameters totemperature=1\.0andtop\_p=0\.95.
https://huggingface.co/NaiveAI/Naive-N0.5-Flash#quick-start-with-transformersQuick Start with Transformers
Naive-N0.5-Flash requires FP8-capable NVIDIA GPUs. The model weights occupy approximately 315 GB; allow additional GPU memory for inference.
pip install "transformers[torch,kernels]>=5.17.0"
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "NaiveAI/Naive-N0.5-Flash-FP8"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
dtype="auto",
device_map="auto",
)
inputs = tokenizer.apply_chat_template(
[{"role": "user", "content": "Hello!"}],
add_generation_prompt=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
output = model.generate(**inputs, max_new_tokens=2048)
response = tokenizer.decode(
output[0, inputs["input_ids"].shape[1]:],
skip_special_tokens=True,
)
print(response)
https://huggingface.co/NaiveAI/Naive-N0.5-Flash#licenseLicense
Naive-N0.5-Flash is released under the MIT License.
https://huggingface.co/NaiveAI/Naive-N0.5-Flash#citationCitation
If you find Naive-N0.5-Flash useful in your research or work, please cite:
@misc{naiveai2026naiven05flash,
title = {Naive-N0.5-Flash: Building Frontier AI with AI},
author = {{NaiveAI Team}},
year = {2026},
url = {https://naive.ai/en/research/}
}
https://huggingface.co/NaiveAI/Naive-N0.5-Flash#acknowledgmentsAcknowledgments
Naive-N0.5-Flash builds on the work of the open-source community and gives back to it. We thank the Xiaomi MiMo team for making their MiMo-V2.5 base model publicly available, the DeepSeek team for their work on DeepSeek Sparse Attention (DSA), and the SGLang team and community for their open-source inference infrastructure.
https://huggingface.co/NaiveAI/Naive-N0.5-Flash#contactContact
For questions, feedback, or collaboration, please contact us at[email protected]or follow us on X at @naiveailab. You can also find our open-source projects and model releases onGitHubandHugging Face.
Similar Articles
Mach-Mind-4-Flash Technical Report
This technical report introduces Mach-Mind-4-Flash, a 35B-parameter Mixture-of-Experts agentic model with 3B activated parameters that matches or surpasses 100B-class models through post-training optimization alone. It presents a novel training infrastructure with dynamic multi-teacher scheduling, multi-teacher on-policy distillation, and hybrid median-length policy optimization for token efficiency.
@AntLingAGI: Introducing Ling-2.6-flash, an instruct model with 104B total parameters and 7.4B active parameters. Ling-2.6-flash is …
Ling-2.6-flash is a 104B-total/7.4B-active sparse instruct model optimized for token efficiency, aiming to cut costs and boost throughput on agent tasks.
@AdinaYakup: Step-3.7-Flash New VL model from @StepFun_ai 198B / 11B active - MoE 256K context 3 reasoning level Up to 400 tokens/sec
StepFun releases Step-3.7-Flash, a new large vision-language MoE model with 198B parameters (11B active), 256K context, and up to 400 tokens/sec inference speed.
README_EN.md · openpangu/openPangu-2.0-Flash at main
openPangu-2.0-Flash is a 92B-parameter MoE model with 6B activated parameters, trained on Ascend, featuring 512k context length and fast thinking capabilities. It achieves strong performance on reasoning and coding benchmarks, using architectural innovations like MLA attention and multi-token prediction.
@VukRosic99: Li Auto's Mach-Mind-4-Flash: a 35B MoE with only 3B active params that matches 100B-class models through post-training …
Li Auto's Mach-Mind-4-Flash is a 35B MoE model with only 3B active parameters that, through post-training optimization using a unified RL/on-policy distillation loss, achieves performance matching or surpassing 100B-parameter models across multiple benchmarks.

