StepFun launches Step 5 Preview with 1M-token context and $1 input pricing
StepFun announced Step 5 Preview, its flagship mixture‑of‑experts model aimed at agentic software engineering, knowledge work, and finance tasks. The model totals about 600 billion parameters but activates roughly 27 billion per token, giving a 1 million‑token context window and supporting text, image, and video input. StepFun says the model is available now via a hosted API and on its platform,…
Key points
- Step 5 Preview is a 600B‑parameter MoE model activating 27B weights per token, with a 1 M‑token context window.
- API pricing is $1.00 per 1 M input tokens and $2.70 per 1 M output tokens.
- Open weights are planned for release on October 15 2026; until then the model is available via hosted API only.
The company lists API pricing at $1.00 per 1 M input tokens and $2.70 per 1 M output tokens. Internal benchmarks show a 3× end‑to‑end speedup for long‑horizon reinforcement‑learning workloads and a throughput of 99.8 tokens per second. On StepFun’s own coding benchmarks the model scored 67.7 on DeepSWE v1.1, 49.0 on StepCodeBench, and 80.5 on ProgramBench, while independent analysis from Artificial Analysis gave it an Intelligence Index score of 44, well above the 24 median for similarly priced models. The model also demonstrated hardware tuning gains, reaching 508 TFLOPS on an H100 kernel compared with 493 TFLOPS for Claude Opus 5.
Model pages: Step 5 Preview → · GPT-6 Astra →
The story so far
2 episodes →- StepFun launches Step 5 Preview with 1M-token context and $1 input pricingthis story
StepFun Launches Step 5 Preview: A 600B-Total, 27B-Active MoE Model With 1M Context for Long-Horizon Agentic Work
MarkTechPost · 21 September 2026
StepFun has released Step 5 Preview, its new flagship model for agentic work. The target workloads are software engineering, professional knowledge work, and finance. The main pitch is cost. StepFun team states the model delivers comparable intelligence at a substantially lower task cost. That is the ‘Pareto frontier’ framing in the launch title.
Is it deployable? Yes, as a hosted API and on the StepFun platform. Self-hosting waits for open weights. StepFun says open weights land on October 15, 2026. By simple arithmetic, 600B parameters need about 1.2 TB in BF16, before KV cache. Plan for multi-GPU server hardware once weights ship.
What StepFun Shipped
Step 5 Preview is a sparse Mixture-of-Experts (MoE) model. It holds about 600B total parameters and activates about 27B per token. That is roughly 4.5% of the weights per token.
The official model documentation lists these specs:
- Model ID:
step-5-preview - Context window: 1M tokens
- Input: text, images, and video
- Output: text
- Reasoning effort:
low,medium, andhigh - Streaming, tool calling, JSON Mode, JSON Schema, and prompt caching
On research tasks, StepFun team states the model coordinated 950 web fetches in a single agent action. StepFun team also documents a Claude Code integration through its Step Plan.
Architecture: Narrow and Deep
StepFun did not widen the network. It stacked 92 Transformer layers in a narrow-deep layout, according to Pandaily. The research team argues deeper stacks give longer paths for implicit multi-hop reasoning. This matters during long prefill, when agents search, run code, and read tool returns.
Training leans on on-policy, long-horizon reinforcement learning. StepFun cites bit-wise train and inference alignment across MoE routing. Other listed techniques include MTP-3 speculative decoding, FP8 MoE, and KV-cache offload. StepFun reports more than 3x end-to-end speedup for long-horizon RL.
Interactive Explainer
Benchmarks: Company-Reported vs Independent
Step 5 Preview ran at High effort, while rivals ran at Max.
StepFun reports these results, via RuntimeWire:
On coding, StepFun reports 67.7 on DeepSWE v1.1, 49.0 on StepCodeBench, and 80.5 on ProgramBench. GPT-6 Astra and Claude Opus 5 stay ahead on all 3. StepCodeBench is StepFun’s own benchmark.
StepFun also ran 2 agent experiments lasting 24 hours each. In the first, the model tuned an H100 kernel to 508 TFLOPS, against 493 for Claude Opus 5. In the second, it raised Qwen3-30B-A3B on AIME24 from 53.3% to 60% through automated post-training.
The independent check comes from Artificial Analysis. It scores Step 5 Preview at 44 on its Intelligence Index. The median for reasoning models in a similar price tier is 24. It measured output at 99.8 tokens per second on StepFun’s API.
Pricing
StepFun’s API list prices per 1M tokens:
Artificial Analysis puts the medians for comparable models at $1.88 input and $10.00 output. There is 1 catch. The model generated 160M output tokens on the index run, against a 92M median. Verbose reasoning eats part of the per-token savings.
Key Takeaways
- StepFun’s Step 5 Preview is a 600B-total, 27B-active MoE model.
- It offers a 1M-token context with text, image, and video input.
- API pricing is $1.00 input and $2.70 output per 1M tokens.
- Artificial Analysis scores it 44 on its Intelligence Index.
- Open weights are scheduled for October 15, 2026.
FAQ
- What is Step 5 Preview? It is StepFun’s flagship MoE model for agentic coding, knowledge work, and finance.
- Is Step 5 Preview open weight? Not yet. StepFun schedules open weights for October 15, 2026.
- How large is the context window? 1M tokens, per StepFun’s documentation.
- How much does it cost? $1.00 per 1M input tokens and $2.70 per 1M output tokens.
Check out the Technical Details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.
This text was published by MarkTechPost and written by Michal Sutter. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Generative AI & Models
All →- mini-AGI is a continual learning byte-level model that trains on a single 8 GB GPU · 1 src
- Alibaba's Qwen-Image-2.1 claims to beat closed models in image generation · 1 src
- Tencent unveils Gander model that talks while handling background tasks · 1 src
- Runway says it wants real-time AI video generation as a live stream · 1 src
- Alibaba releases Qwen3.8-Omni-Flash, a 1M-token omni-modal model with agentic audio‑video · 3 src
Comments
via GitHub Discussions