[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$fFyHFlYjuZiY3lqR9z7y4hZS8lwbsd2Cl_hv83nbT540":3},{"article":4,"related":19},{"id":5,"slug":6,"title":7,"seo_title":7,"description":8,"keywords":9,"content":10,"category":11,"image_url":12,"source_guid":13,"published_at":14,"created_at":15,"updated_at":16,"source_url":17,"source_name":18},1330,"google-ships-gemini-38-flash-its-third-budget-model-in-6-weeks","Gemini 3.8 Flash: Google's Third Budget AI Model in 6 Weeks","Google launches Gemini 3.8 Flash and a cybersecurity variant, its third low-cost model since July, as Gemini 3.5 Pro and 4 remain unannounced.","[\"Gemini 3.8 Flash\",\"Google DeepMind\",\"Gemini Flash Cyber\",\"Claude Opus 5\",\"GPT-5.6 Sol\",\"AI benchmarks\",\"Artificial Analysis\"]","\u003Cp>Google released Gemini 3.8 Flash on September 2, 2026, just three weeks after Gemini 3.7 Flash and six weeks after the start of that release cycle, making it the company's third budget-tier model launch in that span, as \u003Ca href=\"https:\u002F\u002Fthe-decoder.com\u002Fgemini-3-8-flash-is-googles-third-budget-model-in-six-weeks-while-frontier-models-remain-mia\u002F\" rel=\"noopener noreferrer\">first reported by the-decoder.com\u003C\u002Fa>. The release comes in two forms: a general-purpose reasoning and coding model, and a specialized version called 3.8 Flash Cyber built for defensive cybersecurity work.\u003C\u002Fp>\n\n\u003Cp>On Google's own DeepSWE v1.1 benchmark for long-horizon software engineering tasks, Gemini 3.8 Flash scores 73.7 percent, just under Claude Opus 5's 74.0 percent and ahead of Claude Sonnet 5 (53.8%), GPT-5.6 Sol (72.7%), and its own predecessor 3.7 Flash (65.3%). Independent testing from Artificial Analysis puts the model's Intelligence Index at 59, three points above 3.7 Flash, tying it with GPT-5.6 Sol at xhigh reasoning and Grok 4.6 at medium reasoning. Google is keeping introductory pricing at $0.75 per million input tokens and $3.75 per million output tokens, matching 3.7 Flash, before raising to $1.50 and $7.50 in January 2027. For comparison, Claude Opus 5 costs $5.00 and $25.00 per million tokens, and GPT-5.6 Sol runs $4.00 and $20.00. \u003Ca href=\"\u002Fnews\u002Fgoogles-ai-chip-push-what-it-means-for-efficiency-and-competition\" rel=\"noopener noreferrer\">Gemini 3.8 Flash\u003C\u002Fa> offers additional context on this topic.\u003C\u002Fp>\n\n\u003Ch2>The catch: it works harder, and that costs more\u003C\u002Fh2>\n\n\u003Cp>Google attributes the performance jump partly to the model running more reasoning steps and calling tools iteratively on complex tasks, which the company describes as the model \"working harder.\" That drives up token consumption even though the per-token price hasn't moved. Artificial Analysis found the cost per task rose about 40 percent versus 3.7 Flash, from $0.40 to $0.58, even as 3.8 Flash still sits on the Pareto frontier as the cheapest model at its intelligence level. Google itself recommends developers stick with 3.7 Flash, or dial down reasoning levels on 3.8 Flash, for workloads where token efficiency matters more than raw capability. \u003Ca href=\"\u002Fnews\u002Fspacexai-grok-46-surpasses-kimi-k3-matches-gpt-56-sol\" rel=\"noopener noreferrer\">GPT-5.6 Sol\u003C\u002Fa> offers additional context on this topic.\u003C\u002Fp>\n\n\u003Cp>The Cyber variant isn't publicly available. Google distributes it only through its Fairwind Program to government agencies, critical infrastructure operators, and software maintainers, with looser safety restrictions than the consumer model since it's meant for defensive security work. On CyberGym, a benchmark for finding vulnerabilities in C\u002FC++, it scores 86.2 percent, ahead of the prior 3.5 Flash Cyber (77.5%), GPT-5.6 Sol (83.6%), and GPT-5.5-Cyber (85.6%). On the Gray Swan prompt injection benchmark, it posts a 5.5 percent attack success rate, second only to Claude Opus 5's 4.8 percent and far below DeepSeek V4 Pro (60.1%), Kimi K3 (52.7%), and Grok 4.6 (51.8%).\u003C\u002Fp>\n\n\u003Ch2>What this cadence signals\u003C\u002Fh2>\n\n\u003Cp>Three Flash releases in six weeks is an unusually fast cycle for a company that used to ship major model updates a few times a year. That pace reads two ways. One is that Google has found a genuine edge in cheap, fast, iteratively-improving models and is pressing it while it can, especially with 3.8 Flash landing near or above Opus 5 on coding benchmarks at a fraction of the price. The other is that the frontier models everyone actually wants to see, Gemini 3.5 Pro and Gemini 4, are still nowhere, and stacking budget releases is a way to stay in the headlines while those remain unfinished.\u003C\u002Fp>\n\n\u003Cp>DeepMind's new head, Koray Kavukcuoglu, said publicly that Google isn't just optimizing for price-performance and still intends to compete on raw capability. That statement only matters if a frontier release actually follows it. Until Gemini 3.5 Pro or Gemini 4 ships, the company's story is being told entirely by its cheapest tier, which is a strange position for a lab that wants to be seen as leading the field rather than undercutting it.\u003C\u002Fp>\n\n\u003Cp>The pricing math is the more concrete story for anyone actually building on these models. A per-token price that stays flat while real task cost rises 40 percent is a pattern worth watching as reasoning models increasingly decide for themselves how much compute to spend on a given problem. It means the sticker price on a model card is becoming less useful as a budgeting tool, and teams comparing Gemini, Claude, and GPT costs will need to benchmark against their own workloads rather than trust the per-million-token rate card. Google's own advice to use the older, cheaper 3.7 Flash for cost-sensitive work is a tacit admission of that.\u003C\u002Fp>\n\n\u003Cp>The Cyber variant's restricted distribution is also worth noting on its own terms. Keeping a more capable, less restricted model behind a vetting program for governments and infrastructure operators is a sensible way to hand out extra capability without widening the pool of people who could misuse it. Its strong showing against prompt injection, second only to Opus 5, suggests defensive security is becoming a real competitive category between the major labs rather than an afterthought bolted onto general models.\u003C\u002Fp>","AI & Machine Learning","https:\u002F\u002Fseedwire.co\u002Fapi\u002Fimages\u002Farticles\u002F1788394374372-39oiiq5e356.png","f7e114e24f3c4a5f929257cd81291b5741627be857aface8224c7482e78ffd5d","2026-09-02T16:59:29.000Z","2026-09-03T00:12:54.811Z","2026-09-03 18:20:31","https:\u002F\u002Fthe-decoder.com\u002Fgemini-3-8-flash-is-googles-third-budget-model-in-six-weeks-while-frontier-models-remain-mia\u002F","the-decoder.com",[20,27,34,41],{"id":21,"slug":22,"title":23,"description":24,"category":11,"image_url":25,"published_at":26},1338,"ai-generated-intel-nearly-triggered-us-boarding-of-chinese-ship","False AI-Assisted Intelligence Nearly Led to Ship Boarding","A false AI-assisted intelligence report nearly prompted US forces to board a Chinese ship, raising questions about evidence checks before military action.","https:\u002F\u002Fseedwire.co\u002Fapi\u002Fimages\u002Farticles\u002F1789834667387-suur25h1pi8.webp","2026-09-19T00:13:24.748Z",{"id":28,"slug":29,"title":30,"description":31,"category":11,"image_url":32,"published_at":33},1336,"openai-will-now-disclose-misaligned-model-behavior-faster","OpenAI Will Now Disclose Misaligned Model Behavior Faster","OpenAI unveiled a framework for publicly reporting AI misalignment and disclosed incidents including models uploading files unprompted. Why timing matters.",null,"2026-09-16T22:07:24.000Z",{"id":35,"slug":36,"title":37,"description":38,"category":11,"image_url":39,"published_at":40},1331,"gpt-6-astra-lands-and-openai-finally-says-agi","GPT-6 Astra Lands, and OpenAI Finally Says \"AGI\"","OpenAI shipped GPT-6 Astra with benchmark gains over Sol and Anthropic's Fable models; Brockman floats the AGI label, but pricing and access details temper t...","https:\u002F\u002Fseedwire.co\u002Fapi\u002Fimages\u002Farticles\u002F1788480788914-3di1idw6ekn.png","2026-09-03T19:25:40.000Z",{"id":42,"slug":43,"title":44,"description":45,"category":11,"image_url":32,"published_at":46},1337,"muse-spark-13-ships-without-the-mode-behind-its-best-scores","Muse Spark 1.3 Ships Without the Mode Behind Its Best Scores","Meta's Muse Spark 1.3 joins the frontier cluster on independent rankings, but its top scores come from a max mode not yet shipping, and prices did not move.","2026-09-03T16:19:49.000Z"]