内部评估结果非常引人注目。在 Google 的 Big Sleep:https://cloud.google.com/blog/products/identity-security/cloud-ciso-perspectives-our-big-sleep-agent-makes-big-leap 评估中,Flash Cyber 显著超越了主线 3.5 Flash 和 3.6 Flash。在 V8 JavaScript 引擎上,它在固定调用次数下发现了 55 个独特的已确认问题。相比之下,主线 3.5 Flash 为 47 个,Claude Opus 4.6 为 36 个。它捕获了另外两种模型未发现的 10 个问题。在一次现实测试中,Google 的云漏洞研究团队使用它在两小时内在公共 API 中发现了远程代码执行漏洞。
反应分裂在意料之中。开发者欢迎其价格和效率。延迟发布的旗舰产品受到了最强烈的批评。在 Hacker News:https://news.ycombinator.com/item?id=48993130 上,有人认为 Google 过度宣传其无法可靠提供的容量,并引用了令人沮丧的实际编码体验。受限制的 Flash Cyber 发布引发了关于谁应持有自动化漏洞利用工具的双用途讨论。
下面的仪表板按平台汇总了该讨论。它是定性的编辑综合,而非抓取的数据集,方法说明已嵌入其中。
Gemini 3.6 Flash 和 3.5 Flash-Lite 从今天起可用。开发者可以通过 Google AI Studio:https://aistudio.google.com/prompts/new_chat?model=gemini-3.6-flash 和 Android Studio:https://developer.android.com/studio 访问它们。Gemini 3.6 Flash 也在 Google Antigravity:https://antigravity.google/ 并正在 GitHub Copilot:https://github.blog/changelog/2026-07-21-gemini-3-6-flash-is-now-available-in-github-copilot/ 推出。企业可以在 Gemini 企业代理平台中获得两种模型,3.6 Flash 在 Gemini 企业应用中提供。所有人都可以通过 Gemini 应用使用它们,3.5 Flash-Lite 正在 Google 搜索中推出。请从开发者指南开始:https://ai.google.dev/gemini-api/docs/latest-model。
Asif Razzaq 是 Marktechpost Media Inc. 的首席执行官。作为一位有远见的企业家和工程师,Asif 致力于利用人工智能的潜力来造福社会。他最近的努力是推出一个人工智能媒体平台 Marktechpost,该平台以其对机器学习和深度学习新闻的深入报道而脱颖而出,既技术可靠,又易于广大观众理解。该平台每月浏览量超过 200 万次,显示出其在观众中的受欢迎程度。
构建一个智能事件场地运营商 [完整代码]:https://pxllnk.co/twdn5
谢谢!我们的团队会尽快与您联系 🙌
Developers building production agents need higher token efficiency, lower latency, and more reliable performance. Today, Google has released three new Gemini models:https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/. The lineup is Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. All three sit in the Flash tier, which Google tunes for speed, cost, and high-volume agentic work rather than maximum reasoning depth.
Gemini 3.6 Flash is the new default workhorse. It builds on 3.5 Flash and targets coding, knowledge work, and multimodal tasks. The main point is efficiency. On the Artificial Analysis Index:https://artificialanalysis.ai/models/gemini-3-6-flash, 3.6 Flash uses 17% fewer output tokens than 3.5 Flash. On the DeepSWE benchmark by Datacurve:https://deepswe.datacurve.ai/, Google reports up to a 65% reduction. The model also takes fewer reasoning steps and tool calls per multi-step workflow.
Pricing moves down alongside efficiency. Gemini 3.6 Flash is priced at $1.50 per 1M input tokens and $7.50 per 1M output tokens. The output rate drops from the previous $9.00 on 3.5 Flash. Lower verbosity and a lower output price reduces the total cost per agentic task.
Quality gains accompany the efficiency gains. On DeepSWE, 3.6 Flash scores 49% versus 37% for 3.5 Flash. On MLE Bench, it reaches 63.9% versus 49.7%. On OSWorld-Verified, it hits 83.0% versus 78.4%. On GDPval-AA v2, a knowledge-work benchmark, it scores 1421 versus 1349. Computer use is now a built-in client-side tool through the Gemini API and Gemini Enterprise. Early customers including Hebbia and Harvey cite gains in document parsing, chart and data analysis, and report drafting.
Google is shipping 3.6 Flash with enhanced Frontier Safety:https://deepmind.google/blog/strengthening-our-frontier-safety-framework/ safeguards. These cover Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber-offense misuse. Full details are in the 3.6 Flash model card:https://deepmind.google/models/model-cards/gemini-3-6-flash/.
The interactive explainer below lets you compare each model against its predecessor and estimate token cost at your own volume.
Gemini 3.5 Flash-Lite is highlighted for low-latency and high-throughput jobs. Target use cases include agentic search and document processing. As measured by Artificial Analysis:https://artificialanalysis.ai/models/gemini-3-5-flash-lite, it runs at 350 output tokens per second. Pricing is $0.30 per 1M input tokens and $2.50 per 1M output tokens.
The model clears the prior 3.1 Flash-Lite by wide margins. On Terminal-Bench 2.1, it scores 54% versus 31%. On GDM-MRCR v2, a long-context benchmark, it reaches 72.2% versus 60.1%. On GDPval-AA v2, it scores 1140 versus 642. Notably, Flash-Lite also beats the older 3 Flash on some evals. It leads on SWE-Bench Pro at 54.2% versus 49.6% and on OSWorld-Verified at 74.0% versus 65.1%.
Flash-Lite exposes configurable thinking levels: minimal, low, and higher. Developers can prioritize low-cost, low-latency execution for high-volume tasks. They can also engage higher thinking levels for multi-step subagent workloads. Computer use is a built-in tool here too.
Gemini 3.5 Flash Cyber:https://deepmind.google/blog/introducing-gemini-3-5-flash-cyber/ is the most specialized release. It is built on 3.5 Flash and fine-tuned to find, validate, and patch software vulnerabilities. The design premise is the search-space problem. Finding deep flaws means exploring an immense execution search space. A single call to one massive model becomes a bottleneck.
The answer is a cheap model called many times. Inside CodeMender:https://deepmind.google/blog/introducing-codemender-an-ai-agent-for-code-security/, Google’s code-security agent, multiple 3.5 Flash Cyber agents run in parallel. CodeMender invokes the model up to five times, then merges the sub-agent findings into one report. On the CyberGym benchmark, this setup reaches competitive performance against much larger models.
The internal evaluations are striking. On Google’s Big Sleep:https://cloud.google.com/blog/products/identity-security/cloud-ciso-perspectives-our-big-sleep-agent-makes-big-leap evaluation, Flash Cyber significantly surpassed mainline 3.5 Flash and 3.6 Flash. On the V8 JavaScript engine, it found 55 unique confirmed issues at a fixed number of invocations. That compares to 47 for mainline 3.5 Flash and 36 for Claude Opus 4.6. It caught 10 issues the other two models missed. In one real-world test, Google’s Cloud Vulnerability Research team used it to find remote-code-execution flaws in public APIs within two hours.
Reaction split along predictable lines. Builders welcomed the price and efficiency. The delayed flagship drew the loudest criticism. On Hacker News:https://news.ycombinator.com/item?id=48993130, some argued Google is over-selling capacity it cannot reliably provision, citing frustrating hands-on coding sessions. The gated Flash Cyber release opened a dual-use debate about who should hold automated exploit-finding tools.
The dashboard below aggregates that discussion by platform. It is a qualitative editorial synthesis, not a scraped dataset, and the method note is embedded.
Gemini 3.6 Flash and 3.5 Flash-Lite are available starting today. Developers can access them through the Gemini API via Google AI Studio:https://aistudio.google.com/prompts/new_chat?model=gemini-3.6-flash and Android Studio:https://developer.android.com/studio. Gemini 3.6 Flash is also in Google Antigravity:https://antigravity.google/ and rolling out in GitHub Copilot:https://github.blog/changelog/2026-07-21-gemini-3-6-flash-is-now-available-in-github-copilot/. Enterprises get both models in the Gemini Enterprise Agent Platform, with 3.6 Flash in the Gemini Enterprise app. Everyone can use them via the Gemini app, and 3.5 Flash-Lite is rolling out in Google Search. Start with the Developer Guide:https://ai.google.dev/gemini-api/docs/latest-model.
Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.
Build an Agentic Event Venue Operator [Full Codes]:https://pxllnk.co/twdn5