Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic…
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic…
| Plan | Price | Description |
|---|---|---|
| API Usage-Based Pricing | Input: $0.25 / Output: $1.50 per 1M tokens | Pay based on actual API usage. Supports a 1M-token context window, with output pricing including reasoning tokens. Audio input is priced at $0.50 per 1M tokens. Ideal for cost-efficient AI applications and high-volume workloads. |
| API Batch / Flex Mode | Input: $0.125 / Output: $0.75 per 1M tokens | Batch and Flex processing modes offer approximately 50% lower pricing than standard API usage. Designed for large-scale processing and non-real-time workloads. |
| API Free Tier | Free (Limited) | Google AI Studio provides free API access with usage limits and rate restrictions. Suitable for testing, prototyping, and lightweight AI projects. |
| Context Caching | $0.025 per 1M tokens | Supports cached context reads at $0.025 per 1M tokens, with storage priced at $1 per 1M tokens per hour. Helps reduce costs for repeated inputs and long-context applications. |
Multimodal Input & Long Context: Supports diverse input formats including text, image, audio, video, and PDF. Features a 1M token context window, enabling processing of entire long documents or large datasets in a single pass.
Adjustable "Thinking Levels": Offers four reasoning depth levels — minimal, low, medium, and high. Use lower levels for simple tasks (e.g., translation) to save cost and time, and higher levels for complex tasks (e.g., UI generation) to boost quality, flexibly adapting to mixed workloads.
Output & Integration Capabilities: Supports text output up to 64K tokens (including executable code), along with structured outputs and tool calling, facilitating the development of agentic applications.
Gemini 3.1 Flash Lite primarily targets cost-sensitive developers and enterprises seeking large-scale, low-latency solutions:
High-volume task developers: Process frequent repetitive tasks like translation, content moderation, and data extraction at extremely low pricing ($0.25/million input tokens), effectively controlling large-scale deployment costs.
Real-time application builders: With 2.5x faster first-token response and 45% faster output, ideal for low-latency scenarios including real-time translation, mobile interactions, and data dashboards.
Agent application developers: Flexibly switch among 4 reasoning levels — use lower levels for simple tasks to cut costs, higher levels for complex tasks (e.g., UI generation) to maintain quality. Serves as an ideal foundation model for versatile agents.
Enterprises across industries: E-commerce, fashion, and others can leverage its low cost to integrate AI into previously budget-constrained business processes, enabling scalable deployment.
| Model | Context | Pricing | API | Released | Global Heat |
|---|---|---|---|---|---|
![]()
Gemini 3.1 Flash Lite
|
1M | Paid | YES | 2026-05 |
77/100
|
| 1M | — | YES | 2026-08 |
95/100
|
|
| 1M | — | YES | 2026-07 |
96/100
|
|
| 1M | — | YES | 2026-07 |
74/100
|
|
| 131K | Paid | YES | 2026-06 |
80/100
|
Comments (0)