{"version":1,"type":"story","url":"https://digestai.news/story/naiveai-releases-naive-n0-5-flash-a-309b-open-weight-moe-model","json":"https://digestai.news/story/naiveai-releases-naive-n0-5-flash-a-309b-open-weight-moe-model.json","markdown":"https://digestai.news/story/naiveai-releases-naive-n0-5-flash-a-309b-open-weight-moe-model.md","slug":"naiveai-releases-naive-n0-5-flash-a-309b-open-weight-moe-model","headline":"NaiveAI releases Naive-N0.5-Flash, a 309B open-weight MoE model","summary":"NaiveAI has released Naive-N0.5-Flash, an open-weight Mixture-of-Experts (MoE) model with 309B total parameters and 15.5B active parameters. Designed for coding and AI research and development, the model features a native 1M-token context window achieved through a hybrid of Sliding-Window Attention (SWA) and DeepSeek Sparse Attention (DSA), eliminating the need for full-attention layers. The weights are available under the MIT license, and an API is offered with pricing of $0.10 per million input tokens, $0.40 per million output tokens, and $0.01 per million cache read tokens.\n\nThe model builds on the MiMo-V2.5 base and underwent 3.25T tokens of multi-stage training to adapt to its sparse attention architecture. NaiveAI’s inference system, NaiveRT, supports FP8 mixed-precision inference on NVIDIA GPUs, delivering up to 2,000 tokens per second in Ultrafast mode. The release includes technical details on the hybrid attention stack, which uses a 128-token SWA window and DSA selecting the top 2,048 tokens for backbone attention.\n\nEvaluations were conducted using Claude Code 2.1.207 with a 1M-token context window. The model is benchmarked against systems like GPT-5.6-Sol, Opus-5.5, and GLM-5.3 on tasks including SWE-Bench Pro, Terminal-Bench 2.1, and MLE-bench-30. NaiveAI acknowledges the contributions of the Xiaomi MiMo, DeepSeek, and SGLang teams to the development of the model and its infrastructure.","keyPoints":["NaiveAI released Naive-N0.5-Flash, a 309B MoE model with 15.5B active parameters under MIT license.","The model supports a native 1M-token context window using hybrid Sliding-Window and Sparse Attention.","API pricing is set at $0.10 per million input tokens and $0.40 per million output tokens."],"whyItMatters":"This release provides a high-performance, open-weight option for developers needing long-context coding capabilities. Its MIT license and competitive API pricing lower barriers for integrating advanced AI R&D tools into local or cloud environments.","category":{"slug":"models","name":"Generative AI & Models","url":"https://digestai.news/category/models"},"entities":{"companies":["NaiveAI","Xiaomi","DeepSeek","SGLang"],"models":["Naive-N0.5-Flash","MiMo-V2.5","GPT-5.6-Sol","Opus-5.5","GLM-5.3"],"people":[]},"firstPublishedAt":"2026-09-27T18:48:03Z","updatedAt":"2026-09-27T18:48:03Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"huggingface.co","title":"Naive-N0.5-Flash - 309B-A15.5B","url":"https://huggingface.co/NaiveAI/Naive-N0.5-Flash","publishedAt":"2026-09-27T18:48:03Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[{"site":"Reddit","url":"https://www.reddit.com/r/LocalLLaMA/comments/1wrs58t/naiven05flash_309ba155b/","points":null}],"thread":null,"cite":{"text":"Digest AI, \"NaiveAI releases Naive-N0.5-Flash, a 309B open-weight MoE model\", 27 September 2026, https://digestai.news/story/naiveai-releases-naive-n0-5-flash-a-309b-open-weight-moe-model","publisher":"Digest AI","title":"NaiveAI releases Naive-N0.5-Flash, a 309B open-weight MoE model","datePublished":"2026-09-27T18:48:03Z","url":"https://digestai.news/story/naiveai-releases-naive-n0-5-flash-a-309b-open-weight-moe-model"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}