Google Gemini 3.8 Flash Has a Mix of Improved Benchmarks

Gemini 3.8 Flash is Google’s newest Flash-tier model released September 2, 2026. It is the most capable Flash model yet for mid-difficulty long-horizon software engineering, autonomous agents, and enterprise workflows, while keeping Flash-level speed and the same introductory pricing as 3.7 Flash. It is a value model.

Pricing is the same as 3.7 Flash

Introductory (through Dec 31, 2026): $0.75 / 1M input, $3.75 / 1M output
Regular (from Jan 1, 2027): $1.50 / $7.50

Where it is strong
Vals Finance Agent v2 — 61.4% (best)
Harvey’s Legal Agent Benchmark — 10.0% (best. all models score poorly here)
Terminal-bench 2.1 — 89.4% (narrowly best)
CharXiv Reasoning (charts, no tools) — 86.2% (best)
LVBench long video — 87.8% agentic / 87.1% static (best)
HLE-Verified — 54.9% (best)
BioMysteryBench (Human Difficult) — 56.5% (best)
LABBench2 — 86.2% (best)

These wins cluster in multimodal understanding (video + charts), domain-specific professional agents (finance, legal, biology), and mid-difficulty terminal/agentic coding.

Where it is behind

DeepSWE v1 (hard long-horizon SWE) — 71.0% vs Opus 5’s 74.0% and GPT-5.6 Sol’s 72.7%
GDPVal-AA v2 knowledge-work Elo — 1545 vs Opus 5’s 1824
Terminal-bench 4.0 (harder general agents) — 19.1% vs Opus 5’s 51.8%
OSWorld-2.0 computer use — 59.0% vs Opus 5’s 75.4%
GDP.PDF — 35.0% vs GPT-5.6 Sol’s 40.0%

1 thought on “Google Gemini 3.8 Flash Has a Mix of Improved Benchmarks”

  1. 3.7 Flash is decent. Compared to Grok 4.6 it does some things better, but for some things Grok is the way better choice so in practical sense my opinion is mixed. I dont believe the marketing and hype.

    For some reason Alphabet is not releasing Pro versions. I mean old Grok versions were not so good really, but they released them. Google decided the other way.

    Reply

Leave a Comment