Google's Gemini 3.7 Flash cuts token prices to $0.75 per million input, half of 3.6 Flash, as the flagship 3.5 Pro remains delayed.
Google's Gemini 3.7 Flash cuts token prices to $0.75 per million input, half of 3.6 Flash, as the flagship 3.5 Pro remains delayed.

Google's Gemini 3.7 Flash cuts token prices to $0.75 per million input, half of 3.6 Flash, as the flagship 3.5 Pro remains delayed.
Google's Gemini 3.7 Flash, out three weeks after 3.6 Flash, halves token prices to $0.75 per million input and $3.75 per million output as the flagship 3.5 Pro stays unreleased.
"Gemini 3.7 Flash is a direct result of developer feedback and algorithmic innovations," Tulsee Doshi, senior director at Google, said. The model delivers "substantial improvements" in coding and agentic performance, she added.
Benchmarks show gains across the board. DeepSWE v1.1, which measures software engineering issue resolution, rose to 65.3 percent from 49 percent with 3.6 Flash. FrontierCode 1.1 Main climbed to 43.6 percent from 34.4 percent, and WebDev Arena Elo improved to 1588 from 1538. For knowledge work, GDP.pdf document processing reached 34 percent versus 22 percent, while AutomationBench business workflow execution jumped to 30.4 percent from 17 percent.
The release comes as Alphabet faces mounting pressure on AI competitiveness. Google promised Gemini 3.5 Pro would launch in June at I/O, but the model remains in testing with partners. Reports suggest Gemini's coding capabilities have not kept pace with recent advances from OpenAI and Anthropic, and Google DeepMind recently underwent a leadership overhaul that saw Demis Hassabis step aside for deputy Koray Kavukcuoglu.
The introductory rate of $0.75 per million input tokens and $3.75 per million output tokens, available through the end of 2026, is half of what 3.6 Flash charged at launch. OpenAI's GPT 5.6 Luna, its Flash-class model, is priced at $0.20 per million input and $1.20 per million output tokens, still undercutting Google's offering. The pricing pressure reflects a broader industry shift toward cheaper inference as open-source models gain traction.
For developers building autonomous AI systems that plan tasks, use software tools, and complete multi-step workflows, the cost difference is material. A developer processing 100 million input tokens monthly would pay $75 with Gemini 3.7 Flash versus $20 with OpenAI's Luna. Google is betting that the model's improved coding and agentic performance justifies the premium, even as Anthropic has also permanently reduced pricing on its Sonnet 5 model to compete with open-source alternatives.
Gemini 3.7 Flash is live in the Gemini API, AI Studio, and Gemini Enterprise. For individual users, availability is narrower: the model powers the Spark agent in the Gemini app but only for AI Pro and Ultra subscribers, while the standard chatbot interface continues to run on 3.6 Flash. Google also highlighted updated safeguards against misuse in chemical, biological, radiological, and nuclear domains and cyber offense, while enabling beneficial use cases.
The company says the model "thinks more diligently," putting more effort into multi-step planning and tool calls, with more disciplined execution requiring less manual oversight and fewer retries across engineering workflows. For UI generation, the model shows high design adherence and parity based on reference inputs, whether screenshots, images, or full design systems.
The stakes extend beyond model benchmarks. Co-founder Sergey Brin has urged key AI staff to focus on Gemini, and the two original technical co-leads of the model departed to co-found a startup. CEO Sundar Pichai defended Google's AI strategy on the July earnings call, pushing back on concerns that the company has fallen behind in AI coding. For Alphabet, competitive AI models are essential to Google Cloud's growth trajectory, and the half-price introductory rate may compress near-term margins but could help retain developers considering OpenAI's cheaper Luna models. The accelerated Flash cadence — multiple releases edging toward 4.0 — suggests Google is prioritizing incremental wins over a risky flagship launch that could invite unfavorable comparisons to OpenAI and Anthropic's latest models.
This article is for informational purposes only and does not constitute investment advice.