Gemini 3.6 Flash
Total Context
1M
Max Output
65.5K
Released
N/A
gemini-2.5-flash-liteGemini 2.5 Flash-Lite is Google’s fastest and most budget-friendly Gemini 2.5 option. Official docs highlight low latency, low cost, multimodal support, thinking budgets, and tool integrations such as grounding and code execution. It is best described for classification, translation, routing, extraction, and high-scale workloads.
Context Window
1M tokens
Maximum Output
65.5K tokens
Release Date
Jun 17, 2025
Modalities
| Input Price | Output Price | Cache Read |
|---|---|---|
| $0.1/M | $0.4/M | $0.01/M |
Reasoning
Tool calling
Temperature parameter
Attachments
Knowledge Base
Endpoint Protocols
Gemini 2.5 Flash-Lite prioritizes low latency and efficient multimodal processing for frequent, lightweight tasks at scale.
The model is optimized for fast execution of lightweight requests, supporting responsive applications and high-frequency processing.
It can interpret text, images, video, audio and PDF files while producing text responses for downstream applications.
Its 1,048,576-token input limit allows lightweight processing to cover long documents and sizeable collections in fewer requests.
Gemini 2.5 Flash-Lite fits high-volume classification, straightforward extraction and latency-sensitive content processing.
Assign categories, priorities or moderation labels to large streams of text and multimodal content with fast turnaround.
Extract routine fields and facts from messages, forms, images or PDF files for indexing and downstream automation.
Condense lengthy documents and PDFs into short overviews, key points or routing metadata for content pipelines.
Replace these path values before running: {model}
curl 'https://us-api.tokenhub.com/v1beta/models/{model}:generateContent' \
-X 'POST' \
-H "Authorization: Bearer $TOKENHUB_API_KEY"| Index score | ||
|---|---|---|
| Artificial Analysis Intelligence Index | Artificial Analysis broad capability aggregate | 11.4 |
| Artificial Analysis Coding Index | Artificial Analysis software task aggregate | 9.5 |
| Artificial Analysis Math Index | Artificial Analysis math reasoning aggregate | 53.3 |
| Knowledge & Reasoning | ||
| MMLU-Pro | Advanced multi-task knowledge | 75.9% |
| GPQA | Advanced science problem solving | 62.5% |
| HLE | Broad expert-level exam set | 6.4% |
| Coding & Engineering | ||
| LiveCodeBench | Live coding problems | 59.3% |
| SciCode | Scientific coding challenges | 19.3% |
| Terminal-Bench Hard | Hard terminal task execution | 4.5% |
| Math | ||
| MATH-500 | Advanced math problem solving | 96.9% |
| AIME | Competition math problems | 70.3% |
| AIME 2025 | Competition math problems | 53.3% |
| Instruction Following & Agent Tasks | ||
| IFBench | Prompt constraint adherence | 49.9% |
| AA-LCR | Long-context reasoning | 51.3% |
| τ²-Bench | Agent workflow tasks | 18.4% |
Metrics sourced from Artificial Analysis
Understand what Gemini 2.5 Flash-Lite is, its best uses, distinguishing strengths, practical tradeoffs, and safe TokenHub integration guidance.
Gemini 2.5 Flash-Lite is Google’s most economical Gemini 2.5 model for simple, high-frequency multimodal processing. It remains a defined model generation, but newer models in the same family may be preferable for new evaluations.
Best-fit scenarios include large-scale classification and routing, simple structured data extraction, and high-volume translation. Test representative inputs and define measurable acceptance criteria before production.
Key strengths include cost-efficient scaling, fast response times, and support for varied multimodal inputs. This combination is especially useful for simple structured data extraction.
Consider another model when the workload involves difficult multi-step reasoning, the project can adopt a newer Gemini generation, or the workflow cannot include human review for important decisions. Verify important factual, legal, financial, medical, or operational outputs with qualified human review.
In TokenHub, select the exact model identifier displayed for Gemini 2.5 Flash-Lite, use the endpoint documented for your account, and authenticate with your TokenHub credentials. Confirm the TokenHub-exposed input types, tools, grounding options, and model lifecycle rather than assuming full Gemini API parity.
Use one API key to access Gemini 2.5 Flash-Lite and more AI models through TokenHub.
Media and Discussions
Selected public videos and posts related to this model.
X (Twitter)
Reddit
YouTube