在三周前发布的 3.7 Flash 的基础上:https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/,并且标志着我们在短短六周内的第三次 Flash 发布,今天我们推出了 Gemini 3.8,这是我们迄今为止最强的推理和编码模型,其速度和 3.7 相同且成本低。Gemini 3.8 引入了两个变体:
虽然针对不同的部署环境进行了优化,但今天发布的两个版本都由相同的基础智能驱动,并通过长周期的自主循环进一步加速,旨在递归地评估和优化底层模型。在这个共享核心上的显著编码和推理提升来自多项创新,包括在高度要求的网络安全领域进行的严格训练。
Gemini 3.8 Flash 相较于 3.7 Flash 提供了显著的提升,性能往往接近更高成本的前沿模型。
在 DeepSWE v1.1(长周期软件工程)测试中,3.8 Flash 在端到端自主解决复杂工程问题方面的表现超越了大多数更大规模的前沿模型,而成本仅为其一小部分。
此外,3.8 Flash 在专门知识领域展示了关键企业自主所需的可靠性。在需要高级分析和报告的定量和专业领域,3.8 Flash 在 Vals Finance Agent V2:https://www.vals.ai/benchmarks/fabv2 和 Harvey's Legal Agent Benchmark:https://www.vals.ai/benchmarks/hlab 等基准测试中超过了 3.7 Flash 和其他前沿模型。3.8 Flash 在 HLE-Verified 上也取得了 54.9% 的成绩,显示出其在 STEM、 人文学科和专业领域处理多步骤推理的能力。
这些性能提升源于核心设计选择:3.8 Flash 更加努力工作。在复杂任务中,它表现出更高的勤勉——执行更多推理步骤,并迭代调用工具。有时模型可能会使用更多的 tokens 以最大化性能,尤其是在更高努力级别时。
对于以计算效率为主要限制的应用,开发者可以使用较低努力级别以最小化 token 开销,或者继续依赖 Gemini 3.7 Flash,其在追求效率的工作负载下仍然得到完全支持。
Gemini 3.8 Flash 使用一个简单提示并在 Google Antigravity 中使用循环指令构建了这个游戏。该游戏结合了谜题、环境叙事和 Nano Banana 生成的纹理,创造了一个沉浸式的 3D 关卡,你在其中扮演一名导航城堡的巫师。
Gemini 3.8 Flash 在 Google Antigravity 中通过单个提示构建了一个功能齐全的 DOS 版 Google 地图,该版本完全可游玩,包含地点、路线和街景视图。
使用来自美国地质调查局的真实数据集,Gemini 3.8 Flash 在 Google Antigravity 中构建了著名地理地点的地形图,可探索实时剖面、2D 投影和科学解释。
硬件解剖是一款交互式 3D 可视化工具,由 Gemini 3.8 Flash 在 Google AI Studio 中构建,可生成具有物理比例的硬件设备拆解的逼真 Three.js 渲染图。它可以自动将设备分解为各层,你可以使用拆解滑块进行爆炸式查看和检查。
Gemini 3.8 Flash Cyber 通过 Fairwind 计划提供给一组值得信赖的防御者:https://blog.google/innovation-and-ai/technology/safety-security/fairwind-program,在当今复杂的网络安全环境中提供了决定性的优势,其 Flash 速度和成本可实现快速迭代。
在寻找漏洞的标准行业基准 CyberGym 测试中,Gemini 3.8 Flash Cyber 展现了前沿级的自主漏洞发现性能。它超越了 3.5 Flash Cyber 以及显著更大型的前沿模型。
为了更好地捕捉现实世界中的防御需求,而不仅限于 CyberGym 中的 C/C++ 代码库,我们还在一个综合内部基准上评估了 Gemini 3.8 Flash Cyber。在该基准中,模型必须发现跨 20 种编程语言的复杂代码库中的各种漏洞。在这里,该模型相比我们之前的模型展现了显著进步,并达到了超过 70% 的成功率。
通过 Gemini 3.8 Flash Cyber,我们专注于为防御者提供专业能力,使其在攻击者面前具有优势。这也是为什么我们从一开始就投入于漏洞修复,并将其优先于像利用漏洞这样的攻击能力。
CWE-Bench:https://cwe-bench.com/#leaderboard,由 Collinear 运营,是一个具有挑战性的补丁能力外部基准测试。在此基准测试中,Gemini 3.8 Flash Cyber 位于帕累托前沿:其 pass@1 为 47.2%,相比领先前沿模型的 47.8%,但成本却显著更低。
我们已经在 Google 内部使用 Gemini 3.8 Flash Cyber 来保障代码安全。例如:
3.8 Flash 在化学、生物、放射性和核(CBRN)及网络攻击领域配备了防止滥用的安全措施,同时支持有益的使用案例,根据我们的前沿安全框架:https://deepmind.google/blog/strengthening-our-frontier-safety-framework/。3.8 Flash Cyber 在网络安全方面配备了更宽松的一套缓解措施,因此仅提供给需要更全面网络能力的可信防御者。
Gemini 3.8 模型在 Gray Swan 评测中,提示注入鲁棒性也有显著提升,有效保护 Gemini 模型用户免受与提示注入相关的恶意攻击。
引入价格将于 2026 年 12 月 31 日到期。从 2027 年 1 月 1 日起,输入 1M 令牌收费 1.50 美元,输出 1M 令牌收费 7.50 美元。
Building on the momentum of 3.7 Flash:https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/ from three weeks ago and marking our third Flash release in only six weeks, today we’re introducing Gemini 3.8, our best reasoning & coding model yet, at the same speed and low cost of 3.7. Gemini 3.8 introduces 2 variants:
While tailored for different deployment environments, both of today's releases are powered by the same foundational intelligence, and further accelerated by long-running agentic loops designed to recursively evaluate and refine the underlying models. The significant coding and reasoning gains across this shared core were driven by a number of innovations, including rigorous training in the highly demanding domain of cybersecurity.
Gemini 3.8 Flash delivers substantial gains from 3.7 Flash, often approaching the performance of higher-cost frontier models.
On DeepSWE v1.1 (Long-Horizon Software Engineering ) 3.8 Flash outperforms most larger frontier models in autonomously solving complex engineering problems end to end, only at a fraction of the cost.
Additionally, 3.8 Flash exhibits the dependability required for critical enterprise autonomy, across specialized knowledge domains . In quantitative and professional fields that require advanced analysis and reporting, 3.8 Flash outperforms 3.7 Flash and other frontier models in benchmarks like Vals Finance Agent V2:https://www.vals.ai/benchmarks/fabv2 and Harvey's Legal Agent Benchmark:https://www.vals.ai/benchmarks/hlab. 3.8 Flash also achieves a 54.9% on HLE-Verified, demonstrating its ability to handle multi-step reasoning across STEM, humanities, and professional fields.
These performance gains stem from a core design choice: 3.8 Flash works harder. On complex tasks, it exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively. At times, the model might use more tokens to maximize performance, especially at higher effort levels.
For applications where compute efficiency is the primary constraint, developers can utilize lower effort levels to minimize token overhead or continue to rely on Gemini 3.7 Flash, which remains fully supported for efficiency-first workloads.
Gemini 3.8 Flash built this game with a simple prompt using a looping instruction in Google Antigravity. The game uses puzzles, environmental storytelling, and textures generated with Nano Banana to create an immersive 3D level in which you play a wizard navigating a castle.
Gemini 3.8 Flash builds a fully functional DOS version of Google Maps in a single prompt in Google Antigravity that is fully playable with locations, directions, and Street View.
Explore realtime cross-sections, 2D projections, scientific explanations in a topographic map of famous geographical sites built with Gemini 3.8 Flash in Google Antigravity using real datasets from the U.S. Geological Survey.
Hardware Anatomy is an interactive 3D visualizer built with Gemini 3.8 Flash in Google AI Studio that generates realistic Three.js renderings of physically-proportioned teardowns for hardware devices. It automatically decomposes devices into layers you can explode and inspect with a deconstruction slider.
Gemini 3.8 Flash Cyber, available to a set of trusted defenders via the Fairwind Program:https://blog.google/innovation-and-ai/technology/safety-security/fairwind-program, provides a decisive advantage in today’s complex cybersecurity landscape, with the Flash speed and cost that enables quick iteration.
On the standard industry benchmark for finding vulnerabilities, CyberGym, Gemini 3.8 Flash Cyber demonstrates frontier-level performance in autonomous vulnerability discovery. It surpasses both 3.5 Flash Cyber as well as significantly larger frontier models.
To better capture real-world defensive needs which are not limited to just C/C++ codebases like in CyberGym, we also evaluated Gemini 3.8 Flash Cyber against a comprehensive internal benchmark in which the model has to discover a wide range of vulnerabilities across complex codebases spanning 20 programming languages. Here, the model showcases an impressive leap over our previous models and reaches a success rate exceeding 70%.
With Gemini 3.8 Flash Cyber, we focused specifically on equipping defenders with expert capabilities that give them an advantage over attackers. This is why we have invested in vulnerability fixing from the start, and prioritized it over offensive capabilities like exploitation.
CWE-Bench:https://cwe-bench.com/#leaderboard, run by Collinear, is a challenging external benchmark for patching capabilities. On this benchmark, Gemini 3.8 Flash Cyber is on the Pareto frontier: with a pass@1 of 47.2% compared to a leading frontier model at 47.8%, yet offered at a significantly lower cost.
We’re already using Gemini 3.8 Flash Cyber to secure code across Google. For example:
3.8 Flash ships with safeguards against misuse in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense, while enabling beneficial use cases, as per our Frontier Safety Framework:https://deepmind.google/blog/strengthening-our-frontier-safety-framework/. 3.8 Flash Cyber ships with a more permissive set of mitigations for cybersecurity, and as such, is only available to trusted defenders who require a more comprehensive set of cyber capabilities.
Gemini 3.8 models have also made a significant leap in prompt injection robustness as measured by Gray Swan, protecting Gemini model users from prompt-injection related malicious attacks.
Introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.