Google AI Just Released Gemini 3.7 Flash: A Coding and Agent Model at $0.75/1M Input Tokens

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane August 13, 2026 2 min read
Google AI Just Released Gemini 3.7 Flash: A Coding and Agent Model at $0.75/1M Input Tokens

Google has released Gemini 3.7 Flash, a refinement of the 3.6 Flash model rather than a new base system. It arrived three weeks after its predecessor and accepts text, images, audio, and video within a 1M-token context window. The knowledge cutoff remains March 2026. Gains concentrate in software engineering, document-heavy knowledge work, and web development.

Pricing is the real argument

The sharper argument is price. The model ships at $0.75 per 1M input tokens and $3.75 per 1M output tokens. This rate is introductory and expires December 31, 2026; from January 1, 2027, costs double to $1.50 and $7.50. The current rate is roughly a third the blended cost of Claude Sonnet 5 or GPT-5.6 Terra.

Is it deployable

Access runs through hosted surfaces only. There are no open weights. Consumers reach it through Gemini Spark on Google AI Pro and Ultra plans. Businesses use the Gemini API, Google AI Studio, Google Antigravity, Android Studio, the Gemini Enterprise Agent Platform, or the Gemini Enterprise app. Teams with data-residency or air-gap requirements are excluded because there is nothing to self-host.

Startups and mid-market teams gain the most, as the introductory price makes always-on agents affordable without a Pro-tier budget. Regulated enterprises get a governed path through Gemini Enterprise. Google’s own eval set points at legal, financial services, biosciences, and enterprise operations. The Harvey LAB-AA, GDP.pdf, and AutomationBench results indicate the target sectors.

Applications include long-running coding agents, document-heavy back-office automation, UI generation from screenshots or design systems, and PDF-to-structured-data pipelines.

The benchmark picture

On FrontierCode 1.1 Main, which measures production code quality, Gemini 3.7 Flash scores 43.6% against 34.4% for 3.6 Flash. On DeepSWE v1.1, a long-horizon software engineering eval, it reaches 65.3%. On WebDev Arena it posts an Elo of 1588 versus 1538, the top score in Google’s comparison table.

Document and workflow results move further. GDP.pdf, an expert PDF comprehension eval, goes from 22.0% to 34.0%. AutomationBench, a private enterprise workflow set, goes from 17.0% to 30.4% — ahead of both Claude Sonnet 5 at 10.7% and GPT-5.6 Terra at 23.6%. Long-context retrieval on GDM-MRCR v2 at 128k reaches 97.0%.

GPT-5.6 Terra is ahead on DeepSWE (69.6%), Terminal-bench 2.1 (87.4%), Terminal-bench 3.0 (20.8%), and OSWorld-2.0 (50.2%). On GDPval-AA v2 knowledge work, 3.7 Flash scores 1525 Elo against 1598 for Sonnet 5 and 1628 for Muse Spark 1.2. CharXiv Reasoning is a regression: 84.5% without tools, down from 85.2% for 3.6 Flash. On the Artificial Analysis Intelligence Index, 3.7 Flash scores 56, against 57 for both GPT-5.6 Terra and Muse Spark 1.2.

What it means

At an 80/20 input-output mix, that is a blended $1.35 per 1M tokens today against $3.60 for Sonnet 5 and $4.00 for GPT-5.6 Terra. For teams running agents at volume, the intelligence-per-dollar gap is the reason to evaluate, not the individual eval wins. Developers using this model will see improved performance in generating code and handling complex documents without needing to change their existing workflow significantly.

Scroll to Top