What happened
Google released Gemini 3.7 Flash on August 13, calling it its most capable workhorse model for coding and agent-style workloads, and shipping it just three weeks after the previous model in the family. Google's published benchmarks show sizable jumps: the DeepSWE v1.1 software-engineering score rose from 49.0 percent to 65.3 percent versus the prior Flash model, with similar gains on web-development and automation tests. The pricing is the headline for buyers: introductory API rates of 75 cents per million input tokens and 3.75 dollars per million output tokens, about half what the previous Flash model cost, but only through December 31, 2026. On the consumer side, the model is initially available through Spark, Google's AI agent integrated into Chrome, and reporting indicates it is also being used in AI Mode in Google Search.
Why it matters for your business
Cheaper mid-tier models are what make AI features affordable in the software small businesses actually buy, from website chat to invoice processing. If a vendor quotes you for a custom AI feature, the raw model cost just dropped again, and quotes should reflect that. If your team builds anything on AI APIs directly, two cautions apply. First, that half-price rate is labeled introductory and expires December 31, so budget for the possibility that your per-request cost changes in January. Second, Google is now shipping models every few weeks. That pace rewards keeping your integration flexible, testing a new model on your own tasks before switching, and avoiding designs that hard-wire you to one provider's model of the month. The practical move is boring but useful: note what you pay per thousand requests today, so you can tell whether next quarter's model changes actually save you money.
