Gemini 3.6 Flash
Total Context
1M
Max Output
65.5K
Released
N/A
gemini-3.1-flash-liteGemini 3.1 Flash Lite is the high-efficiency multimodal model in the Gemini 3.1 family. Model cards describe low latency, high-volume use, support for text, image, video, audio, and PDFs, and lightweight agent tasks. It is best framed for extraction, classification, routing, and production-scale multimodal workloads.
Context Window
1M tokens
Maximum Output
65.5K tokens
Release Date
May 7, 2026
Modalities
| Input Price | Output Price | Cache Read | Cache Create 5m |
|---|---|---|---|
| $0.25/M | $1.5/M | $0.025/M | $0.0833/M |
Reasoning
Tool calling
Temperature parameter
Attachments
Knowledge Base
Endpoint Protocols
Gemini 3.1 Flash-Lite combines low-latency multimodal processing, long context and lightweight reasoning for frequent, high-volume workloads.
Responds efficiently to straightforward, high-frequency requests where turnaround time and operating cost are primary constraints.
Accepts text, images, video, audio and PDFs for scalable classification, extraction and transformation pipelines.
Can allocate additional reasoning when a simple task needs more accuracy without switching to a larger model.
Gemini 3.1 Flash-Lite is suited to bulk translation, multimedia transcription and structured extraction or routing pipelines.
Translate large streams of chat messages, reviews or support tickets while preserving concise output formatting.
Convert recordings, voice notes or video speech into searchable text without a separate speech-recognition pipeline.
Classify requests, extract entities into a defined schema and route each item to the appropriate downstream process.
import OpenAI from "openai"
const client = new OpenAI({
apiKey: process.env.TOKENHUB_API_KEY,
baseURL: "https://us-api.tokenhub.com/v1",
})
const result = await client.chat.completions.create({})
console.log(result.choices[0]?.message?.content)Gemini 3.1 Flash-Lite
| Index score | ||
|---|---|---|
| Artificial Analysis Intelligence Index | Artificial Analysis broad capability aggregate | 25 |
| Artificial Analysis Coding Index | Artificial Analysis software task aggregate | 30.1 |
| Knowledge & Reasoning | ||
| GPQA | Advanced science problem solving | 82.2% |
| HLE | Broad expert-level exam set | 16.2% |
| Coding & Engineering | ||
| SciCode | Scientific coding challenges | 41.9% |
| Terminal-Bench Hard | Hard terminal task execution | 24.2% |
| Instruction Following & Agent Tasks | ||
| IFBench | Prompt constraint adherence | 77.2% |
| AA-LCR | Long-context reasoning | 65.3% |
| τ²-Bench | Agent workflow tasks | 31.3% |
Metrics sourced from Artificial Analysis
Understand what Gemini 3.1 Flash-Lite is, its best uses, distinguishing strengths, practical tradeoffs, and safe TokenHub integration guidance.
Gemini 3.1 Flash-Lite is Google’s low-latency, cost-efficient Gemini 3-series model for frequent lightweight multimodal tasks. It is a current public model in its provider’s documentation, though availability can vary by platform.
Best-fit scenarios include large-scale classification and routing, simple structured data extraction, and high-volume translation. Test representative inputs and define measurable acceptance criteria before production.
Key strengths include fast response times, cost-efficient scaling, and support for varied multimodal inputs. This combination is especially useful for simple structured data extraction.
Consider another model when the task needs the strongest Pro-tier reasoning, the workload requires nuanced long-form generation or difficult reasoning, or the workflow cannot include human review for important decisions. Verify important factual, legal, financial, medical, or operational outputs with qualified human review.
In TokenHub, select the exact model identifier displayed for Gemini 3.1 Flash-Lite, use the endpoint documented for your account, and authenticate with your TokenHub credentials. Confirm the TokenHub-exposed input types, tools, grounding options, and model lifecycle rather than assuming full Gemini API parity.
Use one API key to access Gemini 3.1 Flash-Lite and more AI models through TokenHub.
Media and Discussions
Selected public videos and posts related to this model.
X (Twitter)
Reddit
YouTube