Google introduced Gemini 3.7 Flash on Thursday, three weeks after Gemini 3.6 Flash, and the release note calls it “our most intelligent workhorse model yet for coding and agents.” The headline numbers are all aimed at the middle of the market: the introductory price is half of what 3.6 Flash cost at launch ($0.75 per million input tokens, $3.75 per million output), and the benchmark jumps cluster exactly where developers feel them - FrontierCode 1.1 Main goes from 34.4% to 43.6%, DeepSWE v1.1 from 49.0% to 65.3%, and the AutomationBench business-workflow eval from 17.0% to 30.4%. The model also starts powering Spark, Google’s 24/7 personal agent for AI Pro and Ultra subscribers in more than 160 countries.
The cadence is the story before the benchmarks. Three weeks between Flash versions is release discipline that used to belong to open-source labs shipping weekly, and Google says the speed is “a direct result of developer feedback and algorithmic innovations.” The price footnote is doing heavy lifting too: the introductory rate expires December 31, 2026, after which input tokens double to $1.50 per million. Hacker News gave the launch 676 points and 380 comments, and the thread read less like a product announcement and more like a pricing war tribunal.
🎩 Cask’s Take
The thread’s opening shot came from a skeptic who compared 3.7 Flash with Grok 4.6: “Hard to understand why anyone would choose 3.7 Flash under these conditions… is Deepmind still a frontier lab?” The counter came with receipts - one commenter pulled Artificial Analysis and found Grok 4.6 takes $1,068 to run their benchmark suite while Gemini 3.7 Flash takes $485. Less than half the real-world cost, and that was the whole argument in miniature: nobody could agree on which model was the value king, because the answer depends on which bill you pay.
Then the DeepSeek shadow fell over the room. Another commenter asked why anyone would pick Flash over DeepSeek V4 Flash or Pro at 13-26x cheaper with comparable intelligence, except for multimodal input. That is the uncomfortable sentence Google cannot answer in a blog post: the price floor for workhorse models is now set in Shenzhen, not Mountain View, and the footnote that prices double in January reads like a hedge against a market Google does not control. When even the vendor signals the price will not hold, the intro rate is a trial coupon, not a price.
The real debate underneath was about what a workhorse model is for. Sonnet 5 took the night’s beating - “arguably the most cost ineffective model to ever be released” - while the quiet counterpoint came from people who run workloads at scale: “The only thing that works at scale is gemini flash,” said one commenter who ingests a million documents an hour on GCP, and latency is the differentiator when your app uses AI invisibly. Three weeks between flashes is the new normal for the tier that has to be cheap enough to be invisible. The frontier gets the essays; the workhorse tier gets the billing department.