StepFun
Step 5 Preview
Total Context
1M
Max Output
1M
Released
N/A
step-3.7-flashStepFun Step 3.7 Flash is a high-efficiency multimodal Mixture-of-Experts (MoE) model from StepFun, built for production agent workflows and developer workloads. It supports native text and image understanding, a 256K context window, and flexible reasoning for coding, visual and document analysis, tool calling, search-augmented tasks, and multi-step agents. Its sparse 198B architecture activates about 11B parameters per token, balancing reasoning and multimodal performance with efficient inference. Use the StepFun API through TokenHub to review the current Step 3.7 Flash price, model ID, supported endpoints, and limits before integrating it into production.
Context Window
256K tokens
Maximum Output
256K tokens
Release Date
May 29, 2026
Modalities
| Input Price | Output Price | Cache Read |
|---|---|---|
| $0.2/M | $1.15/M | $0.04/M |
Reasoning
Tool calling
Temperature parameter
Attachments
Knowledge Base
Endpoint Protocols
Step 3.7 Flash combines native image understanding, selectable reasoning depth, and high-throughput agentic coding for frequent production workloads.
A dedicated vision encoder lets the model interpret screenshots, interface layouts, diagrams, charts, and other image inputs.
Low, medium, and high reasoning levels let developers balance response speed with the depth required by each task.
The sparse architecture is designed for rapid production inference while supporting multi-file coding and long-running agent workflows.
Step 3.7 Flash is suited to screenshot-to-code work, repository debugging, financial document analysis, and multi-source research agents.
Interpret a UI screenshot or wireframe and generate structured front-end code that reflects its layout and components.
Trace behavior across multiple files, isolate a defect from an issue report, and generate a patch that can be checked with tests.
Run iterative searches, compare evidence across sources, and synthesize verified findings from text, images, and dense reports.
import OpenAI from "openai"
const client = new OpenAI({
apiKey: process.env.TOKENHUB_API_KEY,
baseURL: "https://us-api.tokenhub.com/v1",
})
const result = await client.responses.create({})
console.log(result.output_text)Step 3.7 Flash
| Index score | ||
|---|---|---|
| Artificial Analysis Intelligence Index | Artificial Analysis broad capability aggregate | 30.9 |
| Artificial Analysis Coding Index | Artificial Analysis software task aggregate | 39.6 |
Metrics sourced from Artificial Analysis
Answers about the StepFun API for Step 3.7 Flash, including multimodal capabilities, agent workflows, pricing, context limits, and TokenHub integration.
Step 3.7 Flash is a high-efficiency multimodal Mixture-of-Experts model from StepFun. It is designed for production agent and developer workloads, combining text and image understanding, a 256K context window, and flexible reasoning for coding, document analysis, tool calling, and multi-step tasks.
Choose Step 3.7 Flash for responsive coding assistants, visual or document analysis, search-augmented research, structured extraction, and tool-using agent workflows. It is a practical fit when one task needs both multimodal input and multi-step reasoning rather than a text-only response.
The model is designed for native text and image understanding, flexible reasoning, and tool-oriented agent tasks. Before implementation, confirm the current TokenHub model page for the exact input formats, endpoint support, request schema, and tool-calling behavior available to your account.
A 256K context window can make Step 3.7 Flash useful for large documents, visual materials, retrieved evidence, and multi-step task state. Its sparse architecture is intended to balance capability with inference efficiency, but actual latency and cost still depend on input size, output length, reasoning settings, and the current TokenHub pricing.
Create a TokenHub API key, select the current Step 3.7 Flash model ID shown in the catalog, and use the endpoint and request example published on its model page. Check the live documentation before sending production traffic, especially if your application includes image input, streaming, or tool calls.
Use the current TokenHub price table rather than a fixed number from an older article. Estimate cost from your typical input tokens, output tokens, cache use where available, request volume, and the amount of image or long-context content in each workflow. Test representative prompts before committing a production budget.
Prioritize Step 3.7 Flash when your workload combines image or document understanding, coding or research, tool calling, and multi-step agent execution—and you value inference efficiency. For simple text-only requests, compare smaller and cheaper models; for a high-stakes task, evaluate Step 3.7 Flash and alternatives on the same prompts, latency target, and budget.
Use one API key to access Step 3.7 Flash and more AI models through TokenHub.
Step 3.7 Flash Media and Demos
Selected public announcements, tests, and community discussions about step-3.7-flash.
X (Twitter)
Reddit
YouTube