Gemini 3.5 Flash is Google’s high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution…
Gemini 3.5 Flash is Google’s high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution…
| Plan | Price | Description |
|---|---|---|
| Standard API Usage | Input $1.50 / Output $9.00 per 1M Tokens | Supports up to a 1M-token context window. Output pricing includes reasoning tokens. Cached input is approximately $0.15 per 1M Tokens. Suitable for everyday AI application development. |
| Batch / Flex API Mode | Input $0.75 / Output $4.50 per 1M Tokens | Designed for batch processing and non-real-time workloads. Pricing is approximately 50% lower than standard API usage, making it suitable for large-scale content processing and automation tasks. |
| Priority API Mode | Input $2.70 / Output $16.20 per 1M Tokens | Provides higher-priority resource access with faster response performance. Priced at approximately 1.8× the standard API rate, ideal for applications requiring higher reliability and lower latency. |
| Free API Tier | Free (Limited Usage) | Google AI Studio provides free access for testing, learning, and lightweight development with usage limits. |
1.Agentic Execution:Autonomously executes multi-step workflows, scoring 83.6% on MCP Atlas and 78.4% on OSWorld-Verified, supporting hours-long unattended operation.
2.Code Generation & Programming:Scores 76.2% on Terminal-Bench 2.1 and 55.1% on SWE-Bench Pro. Delivers a full application in ~10 minutes, significantly faster than competing models.
3.Multimodal Understanding:Processes text, image, audio, and video holistically. Achieves 84.2% on CharXiv Reasoning and 83.6% on MMMU-Pro.
4.Long-Context Processing:Natively supports 128k+ tokens, scoring 77.3% on MRCR v2, ideal for large-scale document analysis and long conversation histories.
Developers & Programmers:Excels at code generation, debugging, and refactoring. Delivers a full application in ~10 minutes, achieving significant efficiency gains over competing models. Ideal for daily coding assistance and review support.
Enterprises Building Agent Workflows:Specializes in multi-step automation tasks (83.6% on MCP Atlas). Applicable to software pipelines, financial processing, customer onboarding, OCR data extraction, and tax workflows.
Large-Scale Deployments Prioritizing Cost & Speed:Priced at $1.50/1M tokens input and $9.00/1M tokens output—roughly one-third the cost of competitors. Output speed of ~289 tokens/sec is 4x faster, making it ideal for production use cases demanding high throughput and low latency.
Users Handling Multimodal Content:Processes text, images, audio, video, and PDFs holistically. Excels in academic literature reasoning (84.2% on CharXiv Reasoning), suitable for mixed-information processing involving charts, screenshots, and documents.
General Users:Available for free in the Gemini app and Google Search AI Mode. Ideal for daily Q&A, document organization, and content summarization tasks.
| Model | Context | Pricing | API | Released | Global Heat |
|---|---|---|---|---|---|
![]()
Gemini 3.5 Flash
|
1M | Paid | YES | 2026-05 |
93/100
|
| 1M | — | YES | 2026-08 |
95/100
|
|
| 1M | — | YES | 2026-07 |
96/100
|
|
| 1M | — | YES | 2026-07 |
74/100
|
|
| 131K | Paid | YES | 2026-06 |
80/100
|
Comments (0)