Gemini 3.6 Flash
Total Context
1M
Max Output
65.5K
Released
N/A
gemini-2.5-flashGemini 2.5 Flash is the balanced price-performance model in the Gemini 2.5 family, combining thinking capability with lower latency and cost. Google materials position it between Pro-level depth and Flash-Lite efficiency. It is suitable for production tasks that need reasoning, multimodal input, and practical throughput.
Context Window
1M tokens
Maximum Output
65.5K tokens
Release Date
Jun 17, 2025
Modalities
| Input Price | Output Price | Cache Read |
|---|---|---|
| $0.3/M | $2.5/M | $0.03/M |
Reasoning
Tool calling
Temperature parameter
Attachments
Knowledge Base
Endpoint Protocols
Gemini 2.5 Flash combines adaptive reasoning, multimodal understanding and a large context window for responsive, high-volume workloads.
The model can apply thinking when a task requires additional analysis, balancing reasoning depth with responsiveness.
It accepts text, images, video and audio, enabling a single model to analyze information across several media types.
A 1,048,576-token input window supports analysis of extensive documents, media collections and consolidated application context.
Gemini 2.5 Flash is suited to large-scale content processing, multimodal analysis and interactive applications that still require reasoning.
Analyze images, recordings and videos alongside text to produce searchable summaries, classifications and extracted findings.
Classify, summarize and extract information from large streams of content while retaining reasoning for less routine cases.
Build responsive assistants that interpret user content, reason through requests and return useful text responses with low perceived delay.
import OpenAI from "openai"
const client = new OpenAI({
apiKey: process.env.TOKENHUB_API_KEY,
baseURL: "https://us-api.tokenhub.com/v1",
})
const result = await client.chat.completions.create({})
console.log(result.choices[0]?.message?.content)| Index score | ||
|---|---|---|
| Artificial Analysis Intelligence Index | Artificial Analysis broad capability aggregate | 14.1 |
| Artificial Analysis Coding Index | Artificial Analysis software task aggregate | 17.8 |
| Artificial Analysis Math Index | Artificial Analysis math reasoning aggregate | 60.3 |
| Knowledge & Reasoning | ||
| MMLU-Pro | Advanced multi-task knowledge | 80.9% |
| GPQA | Advanced science problem solving | 68.3% |
| HLE | Broad expert-level exam set | 5.1% |
| Coding & Engineering | ||
| LiveCodeBench | Live coding problems | 49.5% |
| SciCode | Scientific coding challenges | 29.1% |
| Terminal-Bench Hard | Hard terminal task execution | 12.1% |
| Math | ||
| MATH-500 | Advanced math problem solving | 93.2% |
| AIME | Competition math problems | 50% |
| AIME 2025 | Competition math problems | 60.3% |
| Instruction Following & Agent Tasks | ||
| IFBench | Prompt constraint adherence | 39.0% |
| AA-LCR | Long-context reasoning | 45.9% |
| τ²-Bench | Agent workflow tasks | 14.9% |
Metrics sourced from Artificial Analysis
Understand what Gemini 2.5 Flash is, its best uses, distinguishing strengths, practical tradeoffs, and safe TokenHub integration guidance.
Gemini 2.5 Flash is Google’s balanced Gemini 2.5 Flash model for high-volume, low-latency tasks that still benefit from thinking. It remains a defined model generation, but newer models in the same family may be preferable for new evaluations.
Best-fit scenarios include high-volume application requests, reliable execution of multi-step agent workflows, and analysis of text and visual inputs. Test representative inputs and define measurable acceptance criteria before production.
Key strengths include a strong balance of quality, speed, and cost, fast response times, and strong reasoning on difficult problems. This combination is especially useful for reliable execution of multi-step agent workflows.
Consider another model when the task needs the strongest Pro-tier reasoning, the project can adopt a newer Gemini generation, or the workflow cannot include human review for important decisions. Verify important factual, legal, financial, medical, or operational outputs with qualified human review.
In TokenHub, select the exact model identifier displayed for Gemini 2.5 Flash, use the endpoint documented for your account, and authenticate with your TokenHub credentials. Confirm the TokenHub-exposed input types, tools, grounding options, and model lifecycle rather than assuming full Gemini API parity.
Use one API key to access Gemini 2.5 Flash and more AI models through TokenHub.
Media and Discussions
Selected public videos and posts related to this model.
X (Twitter)
Reddit
YouTube