Google Cuts Gemini 3.7 Flash API Prices in Half for AI Agents

Google Cuts Gemini 3.7 Flash API Prices in Half for AI Agents

August 16, 2026

MOUNTAIN VIEW, California, August 16, 2026, 14:37 PDT

  • Gemini 3.7 Flash costs $0.75 per million input tokens through year-end.
  • Google cut both introductory input and output rates by 50%.

Google parent Alphabet (NASDAQ:GOOGL) has launched Gemini 3.7 Flash with introductory API prices half those of Gemini 3.6 Flash. The model targets coding and automated business workflows that require repeated tool calls.

The price cut matters most for long-running agents. Those systems plan tasks, query software and revise results across many steps. Token charges can accumulate quickly.

Google set the introductory rate at $0.75 per million input tokens. Output costs $3.75 per million tokens. Both offers run through the end of 2026.

API price per million tokensGemini 3.7 Flash introductory rateGemini 3.6 Flash original rateReduction
Input$0.75$1.5050%
Output$3.75$7.5050%
Offer periodThrough end-2026Original launch rateTemporary for 3.7

A simple mixed workload now costs $4.50 for one million input and output tokens each. The equivalent 3.6 Flash bill was $9.00. Larger agent deployments preserve the same percentage saving.

Illustrative workloadGemini 3.7 FlashGemini 3.6 Flash original rateCalculated saving
1M input + 1M output tokens$4.50$9.00$4.50
10M input + 2M output tokens$15.00$30.00$15.00
100M input + 10M output tokens$112.50$225.00$112.50
Calculated from Google’s published per-token rates; excludes other cloud and tool costs.

Gemini 3.7 Flash arrived only three weeks after version 3.6. Google says it improves debugging, issue resolution and production-ready code generation. Independent results across real workloads remain limited.

The model is available to developers through Google’s AI services. It also powers Gemini Spark, the company’s subscription agent for AI Pro and Ultra users in more than 160 countries.

Google AI productStatus on August 16Primary roleAccess detail
Gemini 3.7 FlashAvailableCoding and multi-step agent workflowsDeveloper services and Gemini Spark
Gemini 3.6 FlashAvailable predecessorGeneral Flash workloadsOriginal API pricing was twice the 3.7 offer
Gemini 3.5 ProStill in partner testingPremium flagship modelNo launch date disclosed

Google describes 3.7 Flash as its “most intelligent workhorse model yet.” The phrase signals a commercial priority: broad deployment at lower cost, rather than a new flagship benchmark leader. Tom’s Guide hands-on test

That distinction is important. Gemini 3.5 Pro remains in partner testing after earlier timing slipped. Google has not provided a release date for the premium model.

One early Spark test used Gmail and Drive to assemble deadlines, bills and appointments. It produced linked evidence and draft actions. It also missed files with unclear names, showing that workflow quality still depends on organization and prompts.

External rankings also temper the launch claims. An Artificial Analysis composite cited by MarketWatch placed Gemini 3.7 Flash ninth among current models. Such rankings vary by test mix and can change quickly.

The pricing therefore carries the clearest measurable change. A 50% cut can make retries, tool calls and verification steps cheaper. It does not prove that agents will finish tasks correctly.

Businesses should compare total workflow cost, not token rates alone. Search, storage, databases, external APIs and human review can exceed the model bill.

Risks: The discounted prices are introductory and may change after 2026. Google has not published enough independent production evidence to verify reliability, security or savings across complex enterprise agents.

BEZ KABLI • EXTENDED COVERAGE

Further analysis

What is Gemini 3.7 Flash?
Gemini 3.7 Flash is Google's latest lower-cost AI model for coding and multi-step agent workflows. It is designed to plan tasks, use software tools and complete automated business processes.
How much does Gemini 3.7 Flash cost?
The introductory API price is $0.75 per million input tokens and $3.75 per million output tokens. The rates run through the end of 2026. Google has not disclosed the later standard price.
How much cheaper is it than Gemini 3.6 Flash?
Both introductory rates are 50% lower than Gemini 3.6 Flash's original prices. A workload using one million input and one million output tokens would cost $4.50, down from $9.00.
Where is Gemini 3.7 Flash available?
Developers can access the model through Google's AI services. It also powers Gemini Spark for eligible Google AI Pro and Ultra subscribers in more than 160 countries. Product and regional availability can vary.
Does the lower price guarantee cheaper AI agents?
No. Token charges are only one cost. Tool calls, cloud infrastructure, databases, retries and human review also matter. Weak task completion can erase the model-price saving.
Is Gemini 3.7 Flash Google's new flagship model?
No. Google positions Flash as a fast, cost-focused workhorse. Gemini 3.5 Pro remains the planned premium model, but Google has not announced its release date.

Artur Ślesik

Artur Ślesik is a technology and financial markets journalist at Bez-kabli.pl, covering artificial intelligence, semiconductors, technology stocks and emerging innovations. A graduate of Warsaw University of Technology, he combines a technical background with market analysis to explain how new technologies are shaping industries, businesses and investment trends worldwide.