Gemini 3.5 Pro,谷歌的下一代前沿模型,仍未公开发布。OpenAI 已经提供了 GPT-5.6 Sol 类模型,Anthropic 有 Fable 和 Mythos,中国的实验室如 Moonshot 提供 Kimi K3,智谱推出 GLM-5.2,都在逼近前沿性能。即使 Meta 最近也发布了一个在编程能力上超过谷歌现有产品线的模型。
谷歌没有选择在顶端竞争,而是推出了另一款专注于效率和降低价格的 Flash 更新。根据最近一篇彭博社报道:https://finance.yahoo.com/technology/ai/articles/google-gemini-launch-delayed-tech-180715298.html,旗舰模型因公司致力于提升编码性能而延迟数月。只要 3.5 Pro 仍处于私人测试阶段,谷歌就没有在市场顶端竞争的公开模型。
保持对人工智能的了解。内容清晰、有用,无废话。
关注 The Decoder 获取人工智能新闻、背景故事和专家分析。
The Decoder:https://the-decoder.com/
Google is expanding the Gemini lineup with two Flash models and a specialized Cyber version. But the anticipated frontier model, Gemini 3.5 Pro, is still missing.
Google has announced three new models in the Gemini Flash family:https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/: 3.6 Flash, 3.5 Flash-Lite, and the cybersecurity model 3.5 Flash Cyber. Buried in the announcement is the fact that Google's anticipated flagship, Gemini 3.5 Pro, is still being tested exclusively with partners and will ship "as soon as it is ready."
Google says pretraining for Gemini 4 is already underway. The company calls it its "most ambitious training run" yet and says it is "excited by the progress." That reads like damage control, and it makes clear that Google knows what the market expects but can't deliver yet. Ad
According to the benchmark aggregator Artificial Analysis Index:https://artificialanalysis.ai/models/gemini-3-6-flash, Gemini 3.6 Flash is expected to use about 17 percent fewer output tokens than 3.5 Flash. Google says the savings reach 65 percent on specific benchmarks such as DeepSWE. Google has cut the price to $1.50 per million input tokens and $7.50 per million output tokens, making it much cheaper than the earlier 3.1 Pro model, which 3.6 Flash consistently beats in benchmarks. Ad DEC_D_Incontent-1
Google also reports gains over 3.5 Flash. DeepSWE rises from 37 to 49 percent, MLE Bench from 49.7 to 63.9 percent, and OSWorld-Verified from 78.4 to 83 percent. The GDPval-AA v2 knowledge work benchmark improves from 1,349 to 1,421 points.
Computer Use is now a built-in client-side tool in the Gemini API and Gemini Enterprise. Google has also added stronger Frontier Safety safeguards against CBRN misuse and cyberattacks. CBRN refers to chemical, biological, radiological, and nuclear threats. Ad
Despite gains on multimodal tasks and a one million token context window, Google still trails the best models from competitors in the US and China. Logan Kilpatrick, a member of the technical staff, responded to criticism on X:https://x.com/OfficialLoganK/status/2079590903090663838, saying the explicit goal was efficiency, usability, and lower cost, and that performance still improved in the process.
The smaller Gemini 3.5 Flash-Lite is tuned for low latency and high throughput. According to Artificial Analysis:https://artificialanalysis.ai/models/gemini-3-5-flash-lite, it produces 350 output tokens per second. It costs $0.30 per million input tokens and $2.50 per million output tokens. Ad DEC_D_Incontent-2
Google says Flash-Lite beats the older 3 Flash on several agentic and coding benchmarks, including SWE-Bench Pro and OSWorld-Verified. Compared with its direct predecessor, 3.1 Flash-Lite, its Terminal-Bench 2.1 score rises from 31 to 54 percent. Ad
Google introduced Gemini 3.5 Flash at its last I/O conference as the centerpiece of its agent strategy. The company later added native computer use, allowing the model to operate browsers, desktops, and mobile devices on its own.
Gemini 3.5 Flash Cyber is based on 3.5 Flash and tuned for cybersecurity work. Google has built it into CodeMender, Google DeepMind's code security agent. Several Flash Cyber subagents work in parallel and combine their results into one report. On the CyberGym benchmark, the model scores 83.2 percent, within two points of OpenAI's GPT-5.5-Cyber at 85.6 percent, despite being a much smaller model.
Google's Big Sleep team:https://blog.google/innovation-and-ai/technology/safety-security/cybersecurity-updates-summer-2025/ put the model to work hunting critical flaws in Chrome and Safari, and it outperformed both the standard Flash models and Anthropic's Claude Opus 4.6. When scanning commits in the V8 JavaScript engine, Flash Cyber turned up 55 confirmed unique findings. The standard 3.5 Flash found 47, Opus 4.6 found 36, and Google says ten of Flash Cyber's findings didn't show up in any other model's results.
In another test, Google's Cloud Vulnerability Research Team used the model to scan public APIs. It found remote code execution flaws within two hours and produced a working exploit that bypassed security protections.
Google says the model is just as useful for offense as it is for defense, so the company is keeping access tight. Only governments and trusted partners can use 3.5 Flash Cyber through CodeMender as part of a pilot program. Google is making the CodeMender agent:https://the-decoder.com/googles-codemender-is-designed-to-automatically-find-and-fix-security-flaws-in-software/ itself broadly available in preview through the Gemini Enterprise Agent Platform, but that version runs on the standard Gemini models. It supports C/C++, Go, Java, Python, Ruby, Rust, and TypeScript.
Gemini 3.5 Pro, Google's next frontier model, still isn't publicly available. OpenAI is already serving that class with GPT-5.6 Sol, Anthropic has Fable and Mythos, and Chinese labs like Moonshot with Kimi K3 and Zhipu with GLM-5.2 are closing in on frontier performance. Even Meta recently released a model that outperforms Google's current lineup at writing code.
Instead of competing at the top, Google shipped another Flash update focused on efficiency and lower prices. According to a recent Bloomberg report:https://finance.yahoo.com/technology/ai/articles/google-gemini-launch-delayed-tech-180715298.html, the flagship model is months behind schedule as the company works on improving its coding performance. As long as 3.5 Pro stays in private testing, Google doesn't have a public model that competes at the top of the market.
Stay in the loop on AI. Clear, useful, no fluff.
Follow The Decoder for AI news, background stories and expert analyses.