On this page

The short answer

Gemini 3.8 Flash is a current candidate for multimodal and agent workflows. Compare its quality and complete task cost at the reasoning level you intend to use. Budget for the announced end of introductory pricing, and treat published high-effort benchmarks as evidence about that configuration rather than every request.

What to take away

  • Introductory pricing ends December 31, 2026.
  • High-effort benchmark results do not describe default latency.
  • Test evidence support for each input modality.

Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.

A dated model snapshot. Published metrics describe the named model, settings, and test. They are not Rubrex experiments or a universal ranking. Prices and availability can change after the research date.

Compare models and check the data date →

Which Gemini release is covered

Google’s current model guide lists Gemini 3.8 Flash as generally available for text-output application work. Its specification supports text, image, video, audio, and PDF input, with an input limit of 1,048,576 tokens and output limit of 65,536. This article focuses on that general-purpose model. Specialized Gemini Live and text-to-speech releases have separate capabilities and evaluation requirements; a shared version number should not be taken to mean they are interchangeable endpoints.

Evidence: Gemini 3.8 Flash model specification [1]

Budget beyond the introductory price

Google lists introductory base rates of $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. It states that standard rates of $1.50 and $7.50 take effect January 1, 2027. Include both periods in a product forecast. Our recommendation is to hold observed token usage constant for a simple sensitivity calculation, then separately test whether changes to reasoning effort alter quality, output volume, and task duration.

Evidence: What is new in Gemini 3.8 Flash [2]

Google specifications checked September 25, 2026. USD per million tokens; these base rates exclude separate tool charges and other billing options.
Model or propertyPublished valueHow to interpret it
Through December 31, 2026$0.75 input / $3.75 outputGoogle introductory base rates
From January 1, 2027$1.50 input / $7.50 outputGoogle announced standard base rates
Input / output limits1,048,576 / 65,536 tokensSeparate documented limits

What the independent metrics support

The saved Epoch AI results show Gemini 3.8 Flash at high reasoning effort on GPQA Diamond and FrontierMath Tiers 1-3 v2. The first tests difficult science questions; the second tests mathematics with Python access. Their scores are not directly comparable and should not be collapsed into a single quality percentage. These records do not measure the default medium setting or the latency of your application. Use them as task-specific evidence and keep missing measurements visible rather than filling them with estimates.

Evidence: Epoch AI: GPQA Diamond [3]Epoch AI: FrontierMath Tiers 1–3 v2 [4]

Epoch AI internal evaluations, CC BY 4.0. Snapshot checked September 25, 2026. Standard error is not a confidence interval; scores across different benchmarks are not interchangeable.
BenchmarkExact configurationScoreStandard errorRun date
GPQA DiamondGemini 3.8 Flash (high)95.39%1.40 percentage points2026-09-02
FrontierMath Tiers 1–3 v2Gemini 3.8 Flash (high)68.42%2.76 percentage points2026-09-02

Evaluate migration and multimodal behavior

Before switching an application, review Google’s current migration requirements and run compatibility checks on the exact endpoint and SDK. The latest guide warns that minimal thinking is unsupported for this model. Test the actual low, medium, and high configurations your product may use. For document or video work, verify that the answer identifies the supporting passage or frame, not merely a plausible conclusion. Record accepted outcomes, evidence support, p95 task duration, and billed cost, including any grounding or tool charges.

Evidence: What is new in Gemini 3.8 Flash [2]

Limits of the evidence

This briefing combines Google specifications and Epoch AI records, not original Rubrex experiments. Benchmark results do not establish multimodal reliability or p95 latency. Pricing and migration guidance can change; the article records what was checked on September 25, 2026.

Common questions

Is the listed introductory price permanent?

No. Google’s checked guide lists it through December 31, 2026 and announces higher standard base rates from January 1, 2027.

Which reasoning level should we use?

Compare supported settings on your tasks. Measure accepted quality and complete latency instead of choosing from a benchmark run at a different setting.

Sources & further reading

Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.

  1. Gemini 3.8 Flash model specification Google · Living reference · Maintained documentationReviewed: model ID, input/output limits, modalities, and reasoning settings. Accessed September 25, 2026.
  2. What is new in Gemini 3.8 Flash Google · Living reference · Maintained documentationReviewed: introductory pricing dates and migration requirements. Accessed September 25, 2026.
  3. Epoch AI: GPQA Diamond Epoch AI · 2026 · Independent benchmarkReviewed: benchmark definition and saved internal evaluation records; CC BY 4.0. Accessed September 25, 2026.
  4. Epoch AI: FrontierMath Tiers 1–3 v2 Epoch AI · 2026 · Independent benchmarkReviewed: benchmark definition and saved internal evaluation records; CC BY 4.0. Accessed September 25, 2026.

Questions or corrections? Write to Rubrex. Read our editorial approach.

From reading to a repeatable evaluation

Put a quality bar around your use case.

The AI Reliability Sprint covers one use case, a baseline, failure analysis, and one improvement experiment in 10 business days.

Explore the Sprint