Google

Tier B · Specialist or second-choice models.

Gemini 3.7 Flash

Fast, cheap, and surprisingly good for smaller UI builds.

API pricing

USD per million tokens · Gemini API Standard paid tier
Input
$0.75
Output
$3.75
Cached input
$0.075
Context window
1.05M
Pricing details and specifications

Try the visual runs

Open the generated scene to test it, or compare that prompt with another model.

16 runs

City Scroll Journey

Cinematic scroll journey testing AI-generated scene continuity, scroll-scrubbed camera motion, and art-directed landing craft.

Mechanical Watch Simulator

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Stormwind Trebuchet Simulator

Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Editorial verdict

Where this model fits

Back to all benchmarks →

Gemini 3.7 Flash sits at the top of B tier. Speed and cost still make it useful for one-shot UI and mobile-friendly layouts, even though the team no longer sees Google as leading the overall model race.

Best for

  • Small UI generation
  • Mobile-friendly layouts
  • Cheap fast iterations

Watch out

  • Quota and regional availability can be inconsistent
  • Do not overload it with intense projects

Why it is ranked here

  1. Google calls Gemini 3.7 Flash its most capable Flash model for agentic workflows and multimodal reasoning.
  2. The team settles on high B: useful and efficient, but no longer close to the overall lead.
  3. The main operational risks remain quota and regional availability.

Evidence

2026-09-07

Superbash editorial model ranking

Takeaway: Gemini 3.7 Flash is currently placed in Tier B.

The September 2026 editorial roster places Gemini 3.7 Flash at rank 10.

Open source →
2026-08-14

Gemini 3.7 Flash model page

Takeaway: Use Gemini 3.7 Flash for fast UI generation, not oversized builds.

Google documents a stable one-million-token model with multimodal input, a 65,536-token output limit, and temporary 2026 pricing. Superbash still treats quota availability as an operational caveat.

Open source →

How Gemini 3.7 Flash compares

Three relevant peers, side by side.

  • Gemini 3.7 Flash This model
  • Gemini 3.6 Flash
  • Claude Sonnet 5
  • GPT-5.6 Terra

6 benchmarks · 4 models

Artificial Analysis Intelligence Index

Composite intelligence · Higher is better

Source
Gemini 3.7 Flash56
Gemini 3.6 Flash52
Claude Sonnet 555
GPT-5.6 Terra57

FrontierCode 1.1 Main

Production coding · Higher is better

Source
Gemini 3.7 Flash43.6%
Gemini 3.6 Flash34.4%
Claude Sonnet 542.7%
GPT-5.6 Terra41.3%

DeepSWE v1.1

Long-horizon coding · Higher is better

Source
Gemini 3.7 Flash65.3%
Gemini 3.6 Flash48.6%
Claude Sonnet 553.8%
GPT-5.6 Terra69.6%

Terminal-Bench 2.1

Agentic coding · Higher is better

Source
Gemini 3.7 Flash85.8%
Gemini 3.6 Flash78.0%
Claude Sonnet 580.4%
GPT-5.6 Terra87.4%

Terminal-Bench 3.0

General agent capabilities · Higher is better

Source
Gemini 3.7 Flash14.9%
Gemini 3.6 Flash5.4%
Claude Sonnet 514.6%
GPT-5.6 Terra20.8%

AutomationBench

Enterprise automation · Higher is better

Source
Gemini 3.7 Flash30.4%
Gemini 3.6 Flash17.0%
Claude Sonnet 510.7%
GPT-5.6 Terra23.6%

Official benchmark profile

Google’s published Gemini 3.7 Flash coding, agent, multimodal, and research results.

Google DeepMind sourceAugust 13, 2026Google DeepMind model card →
Composite intelligence

Artificial Analysis Intelligence Index

56
Production coding

FrontierCode 1.1 Main

43.6%
Long-horizon coding

DeepSWE v1.1

65.3%
Full official benchmark table20 scores and peer comparisons
BenchmarkAreaScoreComparison
Artificial Analysis Intelligence IndexComposite intelligence56
FrontierCode 1.1 MainProduction coding43.6%
Gemini 3.7 Flash43.6%
Claude Sonnet 542.7%
GPT-5.6 Terra41.3%
Gemini 3.6 Flash34.4%
DeepSWE v1.1Long-horizon coding65.3%
Gemini 3.7 Flash65.3%
GPT-5.6 Terra69.6%
Muse Spark 1.254.9%
Claude Sonnet 553.8%
Code ArenaWeb development1588 Elo
Terminal-Bench 2.1Agentic coding85.8%
Gemini 3.7 Flash85.8%
GPT-5.6 Terra87.4%
Muse Spark 1.282.9%
Claude Sonnet 580.4%
Terminal-Bench 3.0General agent capabilities14.9%
Gemini 3.7 Flash14.9%
Claude Sonnet 514.6%
GPT-5.6 Terra20.8%
Gemini 3.6 Flash5.4%
AutomationBenchEnterprise automation30.4%
Gemini 3.7 Flash30.4%
GPT-5.6 Terra23.6%
Gemini 3.6 Flash17.0%
Claude Sonnet 510.7%
GDPVal-AA v2Knowledge work1525 Elo
Harvey LAB-AALegal workflows90.7%
Gemini 3.7 Flash90.7%
Claude Sonnet 590.1%
GPT-5.6 Terra85.2%
Gemini 3.6 Flash85.1%
GDP.pdfDocument reasoning34.0%
Gemini 3.7 Flash34.0%
Claude Sonnet 528.0%
GPT-5.6 Terra24.7%
Gemini 3.6 Flash22.0%
CharXiv Reasoning — no toolsChart reasoning84.5%
Gemini 3.7 Flash84.5%
Gemini 3.6 Flash85.2%
GPT-5.6 Terra85.9%
Claude Sonnet 577.0%
CharXiv Reasoning — with toolsTool-assisted chart reasoning88.7%
Gemini 3.7 Flash88.7%
Claude Sonnet 588.3%
Gemini 3.6 Flash89.4%
LVBenchLong-video understanding85.4%
Gemini 3.7 Flash85.4%
Gemini 3.6 Flash84.2%
GPT-5.6 Terra78.9%
Claude Sonnet 568.5%
GDM-MRCR v2 (8-needle)Long context97.0%
Gemini 3.7 Flash97.0%
GPT-5.6 Terra93.5%
Gemini 3.6 Flash91.8%
Claude Sonnet 581.5%
OSWorld 2.0Computer use47.9%
Gemini 3.7 Flash47.9%
GPT-5.6 Terra50.2%
Gemini 3.6 Flash33.8%
Agent’s Last ExamDesktop agents26.3%
Gemini 3.7 Flash26.3%
GPT-5.6 Terra28.0%
Gemini 3.6 Flash24.2%
Claude Sonnet 533.3%
HLE-VerifiedExpert reasoning53.6%
Gemini 3.7 Flash53.6%
Gemini 3.6 Flash51.2%
GPT-5.6 Terra51.1%
Claude Sonnet 531.0%
BioMysteryBench — human-solvableBioinformatics reasoning87.1%
Gemini 3.7 Flash87.1%
Claude Sonnet 587.5%
GPT-5.6 Terra83.8%
Gemini 3.6 Flash80.6%
BioMysteryBench — human-difficultBioinformatics reasoning43.5%
Gemini 3.7 Flash43.5%
Gemini 3.6 Flash41.2%
GPT-5.6 Terra49.4%
Claude Sonnet 534.1%
LABBench2Biology research82.1%
Gemini 3.7 Flash82.1%
GPT-5.6 Terra81.2%
Claude Sonnet 580.1%
Gemini 3.6 Flash76.1%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-08 ↗
Status
Current
API model ID
gemini-3.7-flash
Context
1.05M
Max output
66K
License
Not publicly verified
Access route
Gemini app, Gemini Enterprise, Gemini Enterprise Agent Platform, Google AI Studio, Gemini API, and Google Antigravity.
Modalities
text input, image input, audio input, video input, PDF input, text output
Hardware
Not publicly verified
Gemini API Standard paid tier · USD / 1M tokens
$0.75 input · $0.075 cached input · $3.75 output

Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached input, and $7.50 output per 1M tokens. Output includes thinking tokens. Cache storage costs $0.50 per 1M tokens per hour during 2026, then $1.00; grounding is billed separately.

Benchmark sources

Google’s August 2026 model card publishes exact-model results across coding, terminal agents, enterprise automation, document and video understanding, long context, computer use, and scientific reasoning. Each row below retains the score type and the material harness details from Google’s accompanying methodology.

Gemini 3.7 Flash Model Card

Google says Gemini scores are pass@1 unless noted, use the exact gemini-3.7-flash API model ID with default sampling unless specified, and average multiple trials on smaller benchmarks. Competitor values are generally provider-reported; the methodology identifies exceptions.

Evaluation settings

  • Composite model-intelligence score in Google’s August 2026 comparison table.
  • Score from the official public FrontierCode leaderboard.
  • Google-computed result using a mini-SWE-agent harness, LiteLLM 1.96, and high thinking.
  • Elo rating from the official public Web Dev leaderboard.
  • Google-computed result using the default Terminus 2 agent harness.
  • Google-computed result using a mini-SWE-agent harness and LiteLLM 1.96; two tasks were modified for Google’s container environment.
  • Private-set result sourced from the official public AutomationBench leaderboard.
  • Elo rating sourced from the Artificial Analysis public leaderboard.
  • Complex legal-workflow result sourced from the Artificial Analysis public leaderboard.
  • Google-computed expert PDF document-comprehension result.
  • Google-computed information-synthesis result without tools.
  • Google-computed result with Google Search and code execution enabled.
  • Google-computed result using 1,024 frames and no tools.
  • Google-computed cumulative score at an average 128K-token context.
  • Partial score, maxed over three single-attempt runs at 1080p with a 500-step cap, screenshot-only observation, and the official evaluator.
  • Pass rate using the default ALE-Claw harness.
  • Google-computed accuracy on the full 1,811-item verified set; 689 uncertain original HLE items were excluded.
  • Google-computed result with a Linux terminal, bioinformatics tools, Python, R, and restricted-domain internet access.
  • Harder split with the same terminal, bioinformatics, Python, R, and restricted-domain internet tools.
  • Google-computed real-world biology result with a Linux terminal, bioinformatics tools, Python, R, and internet access.