Gemini 3.8 Flash

4.90
Gemini 3.8 Flash is a more interesting upgrade than its version number suggests. Google launched Gemini 3.7 Flash just three weeks ago, and now 3.8 Flash is already here. That pace is hard to ignore.
Advertisement 728 × 90
CompanyGoogle
Context1M
Released2026-09
Updated2026-09-07

Gemini 3.8 Flash Overview

Gemini 3.8 Flash was released on September 2, 2026, as part of the Gemini 3 family.

It can work with text, images, audio, and video, and generates text as its output. Google calls it its “most intelligent workhorse model.”

That description actually fits the direction of this release pretty well.

3.8 Flash is built for coding, AI agents, and problems that take more than one or two steps to solve. When a task gets complicated, it can spend more time reasoning, use tools, check what it gets back, and keep working from there.

In other words, Flash is starting to spend more of its time thinking before it answers. And it’s still fast.

Google also released Gemini 3.8 Flash Cyber, a version aimed at cybersecurity work such as finding vulnerabilities and automatically fixing them. It is being made available to trusted security teams through the Fairwind program.

Gemini 3.8 Flash Pricing

PlanPriceDescription
Free $0/month Basic Gemini features, but no access to Gemini 3.8 Flash
AI Plus $7.99/month For everyday AI use and lighter tasks; no Gemini 3.8 Flash
AI Pro $19.99/month Includes Gemini 3.8 Flash; a good fit for coding and heavier AI use
AI Ultra $99.99/month or $199.99/month Includes Gemini 3.8 Flash and is aimed at heavy and professional users
API Pay-as-you-go $0.75/M input tokens and $3.75/M output tokens through Dec. 31, 2026; regular pricing starts at $1.50/M and $7.50/M

Cached input is priced at 10% of the standard input rate.

One date is worth remembering if you're building with the API: the promotional pricing ends on December 31, 2026. After that, both input and output prices double.

And because 3.8 Flash is more willing to spend tokens on reasoning, your real-world bill can be higher than the headline price suggests.

Gemini 3.8 Flash Key Features

1. Coding is where 3.8 Flash really starts to stand out

This is probably the biggest reason to pay attention to the new model.

On DeepSWE v1.1, Gemini 3.8 Flash scored 73.7%, compared with 65.3% for 3.7 Flash and 74.0% for Claude Opus 5.

On Terminal-Bench 2.1, it scored 89.4%, slightly ahead of Opus 5 at 89.1%.

That puts Flash in a very different conversation from the old “fast model for simple tasks” idea. Complex coding, terminal work, and agent-style development are clearly part of the target now.

2. It’s getting better at specialized work

Finance and legal tasks are two good examples.

On Vals Finance Agent V2 and Harvey's Legal Agent Benchmark, 3.8 Flash scored higher than both 3.7 Flash and Claude Opus 5.

It also scored 54.9% on HLE-Verified, which covers questions across STEM, humanities, and professional fields.

That doesn't mean it can suddenly solve every expert-level problem. The more useful takeaway is that it has become better suited to tasks where the answer takes several rounds of reasoning to get right.

3. You can control how hard it thinks

Gemini 3.8 Flash gives developers three reasoning levels: low, medium, and high thinking effort.

In a third-party intelligence test, the model scored 59 at high effort, 57 at medium, and 52 at low.

That gives developers a useful trade-off. Simple requests don't need maximum reasoning. For complicated code or longer agent workflows, you can turn it up and spend more tokens when the extra work is actually worth it.

4. It thinks more without completely giving up its speed

3.8 Flash can generate around 300 tokens per second, roughly 4–6 times faster than many mainstream models.

Its p50 latency is also about 52% faster than 3.6 Flash.

So Google hasn't simply traded speed for reasoning. Flash is still fast. It just has more room to think when the task calls for it.

⚡ How It Has Changed From Gemini 3.7 Flash

This isn't a ground-up rebuild of 3.7 Flash.

Google is still building on the same general technical foundation, including the model architecture, training data, hardware, and software stack.

The bigger change is how the model handles a difficult job.

Instead of rushing toward an answer, 3.8 Flash can spend more time reasoning, call tools, check the results, and keep iterating. Across 14 benchmark groups, it took first place in 8.

There is a price for that, though: a single task costs about 40% more than with 3.7 Flash.

And that makes the upgrade a lot easier to judge.

For a quick function or simple question, you may not notice much. For a large codebase, a long-running agent, or a complicated analysis task, those extra reasoning steps can be much more useful.

Summary

The good stuff

  • Stronger performance in software engineering, agents, and professional tasks
  • 73.7% on DeepSWE v1.1 and 89.4% on Terminal-Bench 2.1
  • Around 300 tokens/sec output speed
  • 1M-token context window, enough for roughly 1,500 pages of A4 text
  • Three thinking-effort levels
  • API promotional pricing matches 3.7 Flash through the end of 2026

The trade-offs

The biggest thing that gives me pause is the price.

Harder tasks already tend to use more tokens, and 3.8 Flash costs about 40% more per task than 3.7 Flash. If you're running agents all day, that difference can show up pretty quickly on the bill.

There's also the access issue. In the Gemini app, 3.8 Flash is limited to Pro and Ultra. API promotional pricing ends on December 31, 2026, after which both input and output prices double.

And then there's the gap between benchmarks and real work.

Benchmarks can look great on paper. Real projects are another story.

One evaluation found that 3.8 Flash could generate runnable 3D scene code in two or three minutes, but the final result could still be a case of “it runs, but it’s not right.”

That's a useful reminder: a high benchmark score doesn't guarantee a clean result when you drop the model into a real project.

Who should use it?

If you spend a lot of time on complex coding, AI agents, or tasks that require several rounds of reasoning, 3.8 Flash is worth a look.

It's also interesting for finance, legal, and other professional workflows where the model needs to work through information step by step.

For businesses, the math is pretty simple: if the extra reasoning saves enough human time, the higher token cost may be worth it.

Who probably doesn't need it?

If you're mainly using AI for chat, summaries, or basic rewriting, 3.8 Flash is probably more power than you need.

Developers watching API costs closely don't need to rush into an upgrade either. A lot of straightforward tasks can still be handled perfectly well by 3.7 Flash.

And if you just want to try it without paying for Pro, waiting is a perfectly reasonable option.

Google's direction here is pretty clear to me: Flash is no longer trying to be just the fast and cheap model.

It's starting to go after the harder jobs, too.

That's a good thing for users. Just keep one thing in mind: the better the model gets at thinking, the smarter your bill may get at finding your wallet.

Comments (0)

Leave a comment

Advertisement 728 × 90

Google Model Comparison

Model Context Pricing API Released Global Heat
Gemini 3.8 Flash
1M YES 2026-09
98/100
1M YES 2026-08
95/100
1M YES 2026-07
97/100
1M YES 2026-07
72/100
131K Paid YES 2026-06
80/100

Similar Models

Gemini 3.6 Flash
97
Google
Gemini 3.7 Flash
95
Just three weeks after the last update, Google has released another new model. Gemini 3.7 Flash is more than a minor refresh. Its coding, Agent, and automation capabilities have all improved noticeably, while API pricing is cut in half through the end of 2026. It may not be the most powerful model available, but for developers, the value proposition is hard to ignore.
Google
Gemini 3.1 Pro
94
Google DeepMind’s latest Gemini Pro model delivers advanced reasoning, multimodal understanding, coding support, and enterprise AI capabilities for professional applications.
google
Gemini 3.5 Flash
92
Gemini 3.5 Flash delivers near-Pro intelligence at Flash-tier cost and speed: Pro-level coding proficiency, parallel agentic execution, all at the same price point as a Flash model.
Google
Nano Banana 2
80
Google's latest image model delivers pro-grade quality at blazing speed, excels at complex generation and iterative editing, and combines deep understanding with high cost-effectiveness.
Google
Nano Banana Pro
80
Google's most powerful image model, featuring precise text rendering, multi-image fusion, and 4K output, built for professional design.
Google
Gemini 3.1 Flash Lite
78
Google's lightweight multimodal model, supporting text, image, audio, video, and PDF. Delivers low latency and high throughput for large-scale agentic workloads.
Google
Gemini 3.5 Flash Lite
72
Google’s lightweight multimodal AI model optimized for high-volume tasks, AI Agent workflows, and cost-efficient AI applications.
Google

Related Tools

Related News