DeepSeek
DeepSeek V4.1 Flash
Total Context
1M
Max Output
384K
Released
N/A
deepseek-v4-flashDeepSeek V4 Flash keeps the V4 family’s 1M-token context window but uses a lighter MoE configuration, commonly described as 284B total parameters with 13B activated parameters. The emphasis is throughput: fast inference, lower cost per call, and production workloads that still need long-context handling. It is the better fit when the task volume is high and the workload benefits from V4-style long-context architecture without always requiring the deepest reasoning tier.
Context Window
1M tokens
Maximum Output
384K tokens
Release Date
Apr 24, 2026
Modalities
| Input Price | Output Price | Cache Read |
|---|---|---|
| $0.2857/M | $1.1429/M | $0.0057/M |
Reasoning
Tool calling
Temperature parameter
Attachments
Knowledge Base
Endpoint Protocols
DeepSeek V4 Flash combines agentic coding, adjustable reasoning effort and efficient long-context processing for demanding development work.
Post-training strengthens performance on terminal operation, repository modification and multi-step software-engineering tasks.
Low, high and max reasoning levels let developers balance deliberation depth against response time for different workloads.
The V4 architecture supports million-token context processing, helping the model work across large codebases and extensive document collections.
DeepSeek V4 Flash is suited to repository engineering, long-document analysis and multi-step automation that combine reasoning with code execution.
Inspect dependencies across a codebase, identify affected files and implement coordinated changes for software maintenance tasks.
Review extensive technical specifications or document collections, connect information across sections and produce structured findings.
Plan and execute workflows involving terminal commands, code changes and result checks to complete bounded operational tasks.
import OpenAI from "openai"
const client = new OpenAI({
apiKey: process.env.TOKENHUB_API_KEY,
baseURL: "https://us-api.tokenhub.com/v1",
})
const result = await client.chat.completions.create({})
console.log(result.choices[0]?.message?.content)| Index score | ||
|---|---|---|
| Artificial Analysis Intelligence Index | Artificial Analysis broad capability aggregate | 40.3 |
| Artificial Analysis Coding Index | Artificial Analysis software task aggregate | 38.7 |
| Knowledge & Reasoning | ||
| GPQA | Advanced science problem solving | 89.4% |
| HLE | Broad expert-level exam set | 32.1% |
| Coding & Engineering | ||
| SciCode | Scientific coding challenges | 44.9% |
| Terminal-Bench Hard | Hard terminal task execution | 35.6% |
| Instruction Following & Agent Tasks | ||
| IFBench | Prompt constraint adherence | 79.2% |
| AA-LCR | Long-context reasoning | 63% |
| τ²-Bench | Agent workflow tasks | 95.0% |
Metrics sourced from Artificial Analysis
DeepSeek V4 Flash: capabilities, use cases, limits, and TokenHub guidance.
DeepSeek V4 Flash is a DeepSeek model for fast, efficient reasoning and agent work.
Best for latency-sensitive applications, agent workflows and high-volume requests, especially when speed and cost efficiency is the priority.
Key strength: a smaller design with faster, more economical inference and switchable thinking and non-thinking modes.
It has less headroom on the hardest reasoning and engineering tasks. For maximum answer quality, consider DeepSeek V4 Pro.
Use the exact ID shown by TokenHub; follow your account docs and verify current features.
Use one API key to access DeepSeek V4 Flash and more AI models through TokenHub.
Media and Discussions
Selected public videos and posts related to this model.
X (Twitter)
Reddit
YouTube