Gemini 2.5 Flash-Lite

gemini-2.5-flash-lite

Gemini 2.5 Flash-Lite is Google’s fastest and most budget-friendly Gemini 2.5 option. Official docs highlight low latency, low cost, multimodal support, thinking budgets, and tool integrations such as grounding and code execution. It is best described for classification, translation, routing, extraction, and high-scale workloads.

Context Window

1M tokens

Maximum Output

65.5K tokens

Release Date

Jun 17, 2025

Modalities

Gemini 2.5 Flash-Lite Pricing

Input PriceOutput PriceCache Read
$0.1/M$0.4/M$0.01/M

Gemini 2.5 Flash-Lite API Capabilities

Reasoning

Supported

Tool calling

Supported

Temperature parameter

Supported

Attachments

Supported

Knowledge Base

2025-01-01

Endpoint Protocols

geminiCompletions APIMessages API

Gemini 2.5 Flash-Lite Model Highlights

Gemini 2.5 Flash-Lite prioritizes low latency and efficient multimodal processing for frequent, lightweight tasks at scale.

Low-Latency Processing

The model is optimized for fast execution of lightweight requests, supporting responsive applications and high-frequency processing.

Broad Multimodal Input

It can interpret text, images, video, audio and PDF files while producing text responses for downstream applications.

Large Input Window

Its 1,048,576-token input limit allows lightweight processing to cover long documents and sizeable collections in fewer requests.

Gemini 2.5 Flash-Lite Use Cases

Gemini 2.5 Flash-Lite fits high-volume classification, straightforward extraction and latency-sensitive content processing.

High-Volume Classification

Assign categories, priorities or moderation labels to large streams of text and multimodal content with fast turnaround.

Simple Data Extraction

Extract routine fields and facts from messages, forms, images or PDF files for indexing and downstream automation.

Rapid Document Summaries

Condense lengthy documents and PDFs into short overviews, key points or routing metadata for content pipelines.

How to Use Gemini 2.5 Flash-Lite via the TokenHub API

Create API key

Replace these path values before running: {model}

curl 'https://us-api.tokenhub.com/v1beta/models/{model}:generateContent' \
  -X 'POST' \
  -H "Authorization: Bearer $TOKENHUB_API_KEY"

Gemini 2.5 Flash-Lite Benchmarks

Index score
Artificial Analysis Intelligence IndexArtificial Analysis broad capability aggregate11.4
Artificial Analysis Coding IndexArtificial Analysis software task aggregate9.5
Artificial Analysis Math IndexArtificial Analysis math reasoning aggregate53.3
Knowledge & Reasoning
MMLU-ProAdvanced multi-task knowledge75.9%
GPQAAdvanced science problem solving62.5%
HLEBroad expert-level exam set6.4%
Coding & Engineering
LiveCodeBenchLive coding problems59.3%
SciCodeScientific coding challenges19.3%
Terminal-Bench HardHard terminal task execution4.5%
Math
MATH-500Advanced math problem solving96.9%
AIMECompetition math problems70.3%
AIME 2025Competition math problems53.3%
Instruction Following & Agent Tasks
IFBenchPrompt constraint adherence49.9%
AA-LCRLong-context reasoning51.3%
τ²-BenchAgent workflow tasks18.4%

Metrics sourced from Artificial Analysis

Media and Discussions

Selected public videos and posts related to this model.

X (Twitter)

View post on X
View post on X
View post on X

Reddit

YouTube

Watch on YouTube
Watch on YouTube
Watch on YouTube

Frequently asked questions about Gemini 2.5 Flash-Lite

Understand what Gemini 2.5 Flash-Lite is, its best uses, distinguishing strengths, practical tradeoffs, and safe TokenHub integration guidance.

What kind of model is Gemini 2.5 Flash-Lite?+

Gemini 2.5 Flash-Lite is Google’s most economical Gemini 2.5 model for simple, high-frequency multimodal processing. It remains a defined model generation, but newer models in the same family may be preferable for new evaluations.

What should teams use Gemini 2.5 Flash-Lite for?+

Best-fit scenarios include large-scale classification and routing, simple structured data extraction, and high-volume translation. Test representative inputs and define measurable acceptance criteria before production.

Where does Gemini 2.5 Flash-Lite have a clear technical advantage?+

Key strengths include cost-efficient scaling, fast response times, and support for varied multimodal inputs. This combination is especially useful for simple structured data extraction.

When should a team choose another model instead of Gemini 2.5 Flash-Lite?+

Consider another model when the workload involves difficult multi-step reasoning, the project can adopt a newer Gemini generation, or the workflow cannot include human review for important decisions. Verify important factual, legal, financial, medical, or operational outputs with qualified human review.

What should be checked before integrating Gemini 2.5 Flash-Lite with TokenHub?+

In TokenHub, select the exact model identifier displayed for Gemini 2.5 Flash-Lite, use the endpoint documented for your account, and authenticate with your TokenHub credentials. Confirm the TokenHub-exposed input types, tools, grounding options, and model lifecycle rather than assuming full Gemini API parity.

Ready to use Gemini 2.5 Flash-Lite?

Use one API key to access Gemini 2.5 Flash-Lite and more AI models through TokenHub.

Create API key