/v1/chat/completionsStep 3.7 Flash
step-3.7-flashStep 3.7 Flash is a high-efficiency multimodal Mixture-of-Experts (MoE) model from StepFun, designed for production-scale agentic and developer workloads. With native text and image understanding, a 256K context window, and flexible reasoning capabilities, it is well suited for coding, visual and document analysis, tool calling, search-augmented tasks, and complex multi-step agent workflows. Its sparse 198B architecture activates only about 11B parameters per token, balancing strong reasoning and multimodal performance with efficient inference.
Total Context
256Ktokens
Max Output
256Ktokens
Released
May 29, 2026
Modalities
Step 3.7 Flash Price
| Input Price | Output Price | Cache Read |
|---|---|---|
| $0.2/M | $1.15/M | $0.04/M |
How do I use Step 3.7 Flash through the API?
/v1/responses/v1/messagesStep 3.7 Flash Benchmark
Step 3.7 Flash
30.9
/100
Artificial Analysis Intelligence Index
Artificial Analysis broad capability aggregate
Index score
39.6
/100
Artificial Analysis Coding Index
Artificial Analysis software task aggregate
Index score
Metrics sourced from Artificial Analysis
Step 3.7 Flash FAQs
Common questions about using step-3.7-flash through TokenHub.
What is Step 3.7 Flash?+
Step 3.7 Flash is StepFun’s open-weight sparse mixture-of-experts vision-language model for agentic workflows, coding, search, tool use, and image understanding.
What is Step 3.7 Flash best used for?+
It fits tool-using agents, repository-scale coding, visual UI and document analysis, research workflows, and other high-throughput tasks that combine perception and reasoning.
What are the main strengths of Step 3.7 Flash?+
Native vision, long-context processing, three reasoning levels, a focus on stable tool use, and open deployment options make it flexible for production agent systems.
What limitations or tradeoffs should I consider?+
It is a very large model to self-host, so local deployment can require substantial memory and compute. Its outputs can still be wrong; validate important results and compare quality, latency, and cost on your own workload.
How do I use Step 3.7 Flash through the TokenHub API?+
Choose Step 3.7 Flash as the model in your TokenHub API request, then use the endpoint and credentials shown in your TokenHub account. Follow TokenHub’s current guidance for supported text and image payloads.
How is Step 3.7 Flash priced and made available?+
TokenHub pricing, regions, and quotas can change, so check the live model listing and your account before deployment. The model also has open weights and upstream hosted options, while TokenHub availability is governed by TokenHub.
When should I choose Step 3.7 Flash over another model?+
Choose it when you need a fast multimodal agent model with tool use and open deployment choices. For simple text-only tasks, compare a smaller model; for maximum accuracy, benchmark stronger alternatives on your own data.
Step 3.7 Flash Media and Demos
Selected public announcements, tests, and community discussions about step-3.7-flash.
X (Twitter)
Reddit
YouTube