Gemini 3.8 Flash: Speed and Cost Tests of Google's Lightweight Model

Price is the most notable thing about Gemini 3.8 Flash. Until the end of 2026, it costs a small fraction of what flagship models charge. Even so, it scores 90.8

Price is the most notable thing about Gemini 3.8 Flash. Until the end of 2026, it costs a small fraction of what flagship models charge. Even so, it scores 90.8% on Terminal-Bench 2.1, higher than both GPT-5.6 Terra and Claude Sonnet 5. However, it isn't cheap and fast in every scenario. At the high thinking level, the first token takes 16 seconds on average, and prices double starting January 1, 2027. The right choice depends on whether your workload is "batch and long-running tasks" or "real-time conversation." What Is Gemini 3.8 Flash: Positioning and Specs Gemini 3.8 Flash is the lightweight workhorse model Google released on September 2, 2026. It is also the third Flash version Google has shipped in six weeks. Google calls it "our most intelligent workhorse model" and positions it for software engineering, agentic tasks, and multi-step professional reasoning. The goal is near-frontier capability at Flash-tier pricing. On the same day, Google also released Gemini 3.8 Flash Cyber, a security-focused version available only to trusted cyber defenders through the Fairwind Program. See the full announcement on the Official Google Blog: Introducing Gemini 3.8 Flash and 3.8 Flash Cyber . On the technical side, 3.8 Flash's context window can hold roughly an entire mid-sized codebase. According to the official Gemini API model documentation , the key specs are: Model ID : gemini-3.8-flash Input limit : 1,048,576 tokens; Output limit : 65,536 tokens Input types : text, images, video, audio, PDF; output is text only Thinking levels : supports low, medium, and high; does not support minimal Supported features : context caching, code execution, structured output, function calling, Google Search and Google Maps grounding, Batch API, Flex/Priority inference, Computer use (Preview) Not supported : audio generation, image generation, Live API Because it doesn't support the Live API, 3.8 Flash isn't suited to real-time voice conversation. Teams building voice customer service wil

FAQ

What Is Gemini 3.8 Flash: Positioning and Specs

Gemini 3.8 Flash is the lightweight workhorse model Google released on September 2, 2026. It is also the third Flash version Google has shipped in six weeks. Google calls it "our most intelligent workhorse model" and positions it for software engineering, agentic tasks, and multi-step professional reasoning. The goal is near-frontier capability at Flash-tier pricing. On the same day, Google also released Gemini 3.8 Flash Cyber, a security-focused version available only to trusted cyber defenders

Capability Tests: Where It Improved and Where It Stood Still

Gemini 3.8 Flash improved most on terminal operations and agentic tasks, while general knowledge reasoning barely changed. According to "Gemini 3.8 Flash scores 90.8% on Terminal-Bench 2.1, compared with 81.6% for 3.7 Flash" (Source: DataCamp) . On the same test, GPT-5.6 Terra scores 87.4% and Claude Sonnet 5 scores 80.4%. Changes on other benchmarks vary widely: τ³-bench Banking (agentic tasks) : 30.9% up to 38.1%, a gain of 7.2 percentage points SWE-Atlas : 48.0% up to 51.9%, a gain of 3.9 per

When to Choose 3.8 Flash and When Not To

Whether to use Gemini 3.8 Flash depends mainly on how much latency the task can tolerate and whether it is an agentic or engineering task. Use the criteria below: Good Fits Background code changes, test fixes, and terminal automation Analysis of long documents, long videos, and entire PDFs, making full use of the 1-million-token context Multi-step agentic workflows in professional fields such as finance and law (Google cites leading results on Vals Finance Agent V2 and the Harvey Legal Agent Ben

Related Articles

Related Guidebooks

Reviewed and verified by FeiYueh · Last verified 2026-10-03. Independently maintained — not AI-generated boilerplate.

← Back to Blog