News · High impactBack to News

Google Releases Gemini 3.8 Flash: Breaking 90% on Terminal-Bench with Robust Long-Horizon Agentic Coding

Google launched Gemini 3.8 Flash and the specialized 3.8 Flash Cyber, targeting end-to-end terminal execution and complex code repair. Retaining a 1M context window and low inference costs, the model leverages iterative tool loops and deeper reasoning to score 90.8% on Terminal-Bench 2.1 and 61.6% on SWE-bench Pro.

Crossing 90% on Terminal-Bench: Flash Architecture Pivots to Autonomous Recovery

Google officially released Gemini 3.8 Flash on September 2, 2026. The headline milestone is not another conversational benchmark, but an operational breakthrough where Terminal-Bench 2.1 crossed 90.8%, marking a 9.2 percentage point jump over Gemini 3.7 Flash.

Core Engineering Benchmarks

  • 90.8%

    Terminal-Bench 2.1 terminal operations and system maintenance success rate, up 9.2% over 3.7 Flash.

  • 61.6%

    SWE-bench Pro score resolving real-world GitHub issues with automated patch generation.

  • 54.9%

    HLE-Verified score on rigorous multidisciplinary expert-level STEM reasoning.

Arriving just three weeks after 3.7 Flash, this marks Google's third Flash iteration within six weeks. Such an aggressive cadence underscores a decisive research pivot: frontier competition has shifted from one-shot code generation to long-horizon autonomous error recovery in live runtime environments. Starting today, developers can call the model via Google AI Studio, Vertex AI, and the Gemini API, with native support across environments like Antigravity.

Resilient Agent Loops: Patching Code Directly Against stderr

Gemini 3.8 Flash features a standard 1M-token context window and supports up to 64,000 output tokens per request, with a knowledge cutoff of March 2026. In long-horizon development scenarios, the primary value of a large context window is not passive retrieval, but powering iterative tool calling and autonomous error recovery.

When confronting multi-file refactoring or ambiguous compilation failures, the model avoids fragile one-shot guesses. Instead, it activates chain-of-thought exploration, searching repositories, running speculative builds, and refining patches directly against runtime stderr output. Developers can freely adjust reasoning effort budgets based on latency requirements, keeping compute minimal for routine tasks while allocating deeper exploration for complex bugs.

Cyber Variant and Pricing: Restricted Defensive Rollout with Extended Half-Price Inference

Addressing the classic defender's dilemma, Google introduced the specialized Gemini 3.8 Flash Cyber model. The variant achieves an exploit discovery rate exceeding 70% and sits on the CWE-Bench automated patching frontier; however, Google has restricted access, offering this cyber edition exclusively via the Fairwind Program to trusted partners rather than the general public.

On the operational economics side, long-horizon agents repeatedly inspect repository state and ingest terminal feedback, making them extraordinarily price-sensitive. Google is extending promotional rates through December 31, 2026, pricing input tokens at just $0.75 per million ($3.75 output; standard rates of $1.50/$7.50 take effect in 2027), lowering exploration overhead across tens of thousands of steps as it integrates across Google Workspace automation tools and enterprise agent frameworks.

Next step

Keep tracking Gemini 3.8 Flash

Continue along the same topic.

Open entity record