Step 3.7 Flash

step-3.7-flash

StepFun Step 3.7 Flash is a high-efficiency multimodal Mixture-of-Experts (MoE) model from StepFun, built for production agent workflows and developer workloads. It supports native text and image understanding, a 256K context window, and flexible reasoning for coding, visual and document analysis, tool calling, search-augmented tasks, and multi-step agents. Its sparse 198B architecture activates about 11B parameters per token, balancing reasoning and multimodal performance with efficient inference. Use the StepFun API through TokenHub to review the current Step 3.7 Flash price, model ID, supported endpoints, and limits before integrating it into production.

Context Window

256K tokens

Maximum Output

256K tokens

Release Date

May 29, 2026

Modalities

Step 3.7 Flash Pricing

Input PriceOutput PriceCache Read
$0.2/M$1.15/M$0.04/M

Step 3.7 Flash API Capabilities

Reasoning

Supported

Tool calling

Supported

Temperature parameter

Supported

Attachments

Supported

Knowledge Base

—

Endpoint Protocols

Responses APIMessages APICompletions API

Step 3.7 Flash Model Highlights

Step 3.7 Flash combines native image understanding, selectable reasoning depth, and high-throughput agentic coding for frequent production workloads.

Native Visual Understanding

A dedicated vision encoder lets the model interpret screenshots, interface layouts, diagrams, charts, and other image inputs.

Selectable Reasoning Depth

Low, medium, and high reasoning levels let developers balance response speed with the depth required by each task.

High-Throughput Agentic Coding

The sparse architecture is designed for rapid production inference while supporting multi-file coding and long-running agent workflows.

Step 3.7 Flash Use Cases

Step 3.7 Flash is suited to screenshot-to-code work, repository debugging, financial document analysis, and multi-source research agents.

Screenshot to Code

Interpret a UI screenshot or wireframe and generate structured front-end code that reflects its layout and components.

Repository Debugging

Trace behavior across multiple files, isolate a defect from an issue report, and generate a patch that can be checked with tests.

Multi-Source Research

Run iterative searches, compare evidence across sources, and synthesize verified findings from text, images, and dense reports.

How to Use Step 3.7 Flash via the TokenHub API

Create API key
import OpenAI from "openai"

const client = new OpenAI({
  apiKey: process.env.TOKENHUB_API_KEY,
  baseURL: "https://us-api.tokenhub.com/v1",
})

const result = await client.responses.create({})
console.log(result.output_text)

Step 3.7 Flash Benchmarks

Step 3.7 Flash

Index score
Artificial Analysis Intelligence IndexArtificial Analysis broad capability aggregate30.9
Artificial Analysis Coding IndexArtificial Analysis software task aggregate39.6

Metrics sourced from Artificial Analysis

Step 3.7 Flash Media and Demos

Selected public announcements, tests, and community discussions about step-3.7-flash.

X (Twitter)

View post on X
View post on X
View post on X

Reddit

YouTube

Watch on YouTube
Watch on YouTube
Watch on YouTube

Step 3.7 Flash API FAQs

Answers about the StepFun API for Step 3.7 Flash, including multimodal capabilities, agent workflows, pricing, context limits, and TokenHub integration.

What is Step 3.7 Flash?+

Step 3.7 Flash is a high-efficiency multimodal Mixture-of-Experts model from StepFun. It is designed for production agent and developer workloads, combining text and image understanding, a 256K context window, and flexible reasoning for coding, document analysis, tool calling, and multi-step tasks.

What is Step 3.7 Flash best used for?+

Choose Step 3.7 Flash for responsive coding assistants, visual or document analysis, search-augmented research, structured extraction, and tool-using agent workflows. It is a practical fit when one task needs both multimodal input and multi-step reasoning rather than a text-only response.

Does Step 3.7 Flash support multimodal input, reasoning, and tool calling?+

The model is designed for native text and image understanding, flexible reasoning, and tool-oriented agent tasks. Before implementation, confirm the current TokenHub model page for the exact input formats, endpoint support, request schema, and tool-calling behavior available to your account.

How does the 256K context window and sparse MoE architecture affect use?+

A 256K context window can make Step 3.7 Flash useful for large documents, visual materials, retrieved evidence, and multi-step task state. Its sparse architecture is intended to balance capability with inference efficiency, but actual latency and cost still depend on input size, output length, reasoning settings, and the current TokenHub pricing.

How do I call Step 3.7 Flash through the StepFun API on TokenHub?+

Create a TokenHub API key, select the current Step 3.7 Flash model ID shown in the catalog, and use the endpoint and request example published on its model page. Check the live documentation before sending production traffic, especially if your application includes image input, streaming, or tool calls.

How should I evaluate Step 3.7 Flash API pricing?+

Use the current TokenHub price table rather than a fixed number from an older article. Estimate cost from your typical input tokens, output tokens, cache use where available, request volume, and the amount of image or long-context content in each workflow. Test representative prompts before committing a production budget.

When should I choose Step 3.7 Flash over another model?+

Prioritize Step 3.7 Flash when your workload combines image or document understanding, coding or research, tool calling, and multi-step agent execution—and you value inference efficiency. For simple text-only requests, compare smaller and cheaper models; for a high-stakes task, evaluate Step 3.7 Flash and alternatives on the same prompts, latency target, and budget.

Ready to use Step 3.7 Flash?

Use one API key to access Step 3.7 Flash and more AI models through TokenHub.

Create API key