Step 3.7 Flash

step-3.7-flash

StepFun Step 3.7 Flash is a high-efficiency multimodal Mixture-of-Experts (MoE) model from StepFun, built for production agent workflows and developer workloads. It supports native text and image understanding, a 256K context window, and flexible reasoning for coding, visual and document analysis, tool calling, search-augmented tasks, and multi-step agents. Its sparse 198B architecture activates about 11B parameters per token, balancing reasoning and multimodal performance with efficient inference. Use the StepFun API through TokenHub to review the current Step 3.7 Flash price, model ID, supported endpoints, and limits before integrating it into production.

Total Context

256Ktokens

Max Output

256Ktokens

Released

May 29, 2026

Modalities

Step 3.7 Flash Price

Input PriceOutput PriceCache Read
$0.2/M$1.15/M$0.04/M

How do I use Step 3.7 Flash through the API?

POSTopenai/v1/chat/completions
POSTopenai-response/v1/responses
POSTanthropic/v1/messages

Step 3.7 Flash Benchmark

Step 3.7 Flash

30.9

/100

Artificial Analysis Intelligence Index

Artificial Analysis broad capability aggregate

Index score

39.6

/100

Artificial Analysis Coding Index

Artificial Analysis software task aggregate

Index score

Metrics sourced from Artificial Analysis

Step 3.7 Flash Media and Demos

Selected public announcements, tests, and community discussions about step-3.7-flash.

X (Twitter)

Open social media source
View post on X
Open social media source
View post on X
Open social media source
View post on X

Reddit

YouTube

Open social media source
Watch on YouTube
Open social media source
Watch on YouTube
Open social media source
Watch on YouTube

Step 3.7 Flash API FAQs

Answers about the StepFun API for Step 3.7 Flash, including multimodal capabilities, agent workflows, pricing, context limits, and TokenHub integration.

What is Step 3.7 Flash?+

Step 3.7 Flash is a high-efficiency multimodal Mixture-of-Experts model from StepFun. It is designed for production agent and developer workloads, combining text and image understanding, a 256K context window, and flexible reasoning for coding, document analysis, tool calling, and multi-step tasks.

What is Step 3.7 Flash best used for?+

Choose Step 3.7 Flash for responsive coding assistants, visual or document analysis, search-augmented research, structured extraction, and tool-using agent workflows. It is a practical fit when one task needs both multimodal input and multi-step reasoning rather than a text-only response.

Does Step 3.7 Flash support multimodal input, reasoning, and tool calling?+

The model is designed for native text and image understanding, flexible reasoning, and tool-oriented agent tasks. Before implementation, confirm the current TokenHub model page for the exact input formats, endpoint support, request schema, and tool-calling behavior available to your account.

How does the 256K context window and sparse MoE architecture affect use?+

A 256K context window can make Step 3.7 Flash useful for large documents, visual materials, retrieved evidence, and multi-step task state. Its sparse architecture is intended to balance capability with inference efficiency, but actual latency and cost still depend on input size, output length, reasoning settings, and the current TokenHub pricing.

How do I call Step 3.7 Flash through the StepFun API on TokenHub?+

Create a TokenHub API key, select the current Step 3.7 Flash model ID shown in the catalog, and use the endpoint and request example published on its model page. Check the live documentation before sending production traffic, especially if your application includes image input, streaming, or tool calls.

How should I evaluate Step 3.7 Flash API pricing?+

Use the current TokenHub price table rather than a fixed number from an older article. Estimate cost from your typical input tokens, output tokens, cache use where available, request volume, and the amount of image or long-context content in each workflow. Test representative prompts before committing a production budget.

When should I choose Step 3.7 Flash over another model?+

Prioritize Step 3.7 Flash when your workload combines image or document understanding, coding or research, tool calling, and multi-step agent execution—and you value inference efficiency. For simple text-only requests, compare smaller and cheaper models; for a high-stakes task, evaluate Step 3.7 Flash and alternatives on the same prompts, latency target, and budget.