{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-23T08:01:28.298Z","headline":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","description":"Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","url":"https://www.aioga.com/news/cmrnb8azf00t8bic685er9bs3/","mainEntityOfPage":"https://www.aioga.com/news/cmrnb8azf00t8bic685er9bs3/","datePublished":"2026-07-16T09:07:58.000Z","dateModified":"2026-07-16T09:07:58.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://the-decoder.com/gemma-4-gets-a-stealth-update-that-fixes-tool-calling-bugs-and-truncated-responses-under-the-same-name","https://aihot.virxact.com/items/cmrnb8azf00t8bic685er9bs3"],"canonicalUrl":"https://www.aioga.com/news/cmrnb8azf00t8bic685er9bs3/","directAnswer":{"@type":"Answer","text":"Aioga 编辑摘要：Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。 Aioga 将其归入「模型更新」方向，重点关注它对真实使用和行业竞争的影响。","url":"https://www.aioga.com/news/cmrnb8azf00t8bic685er9bs3/","dateCreated":"2026-07-16T09:07:58.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"the-decoder.com source article","url":"https://the-decoder.com/gemma-4-gets-a-stealth-update-that-fixes-tool-calling-bugs-and-truncated-responses-under-the-same-name","datePublished":"2026-07-16T09:07:58.000Z","provider":{"@type":"Organization","name":"the-decoder.com","url":"https://the-decoder.com/gemma-4-gets-a-stealth-update-that-fixes-tool-calling-bugs-and-truncated-responses-under-the-same-name"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmrnb8azf00t8bic685er9bs3","datePublished":"2026-07-16T09:07:58.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmrnb8azf00t8bic685er9bs3"}}],"aggregationSource":"The Decoder：AI News（RSS）","originalPublisher":{"name":"the-decoder.com","url":"https://the-decoder.com/gemma-4-gets-a-stealth-update-that-fixes-tool-calling-bugs-and-truncated-responses-under-the-same-name"},"article":{"id":"cmrnb8azf00t8bic685er9bs3","slug":"cmrnb8azf00t8bic685er9bs3","url":"https://www.aioga.com/news/cmrnb8azf00t8bic685er9bs3/","title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","title_en":"Gemma 4 gets a stealth update that fixes tool calling bugs and truncated responses under the same name","summary":"Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","source":"The Decoder：AI News（RSS）","sourceUrl":"https://the-decoder.com/gemma-4-gets-a-stealth-update-that-fixes-tool-calling-bugs-and-truncated-responses-under-the-same-name","aiHotUrl":"https://aihot.virxact.com/items/cmrnb8azf00t8bic685er9bs3","publishedAt":"2026-07-16T09:07:58.000Z","category":"模型更新","score":48,"selected":false,"articleBody":["Google shipped an update to its open AI model Gemma 4 that speeds up performance on Nvidia Hopper GPUs, fixes tool calling bugs, and addresses problems with truncated responses. Turning on Flash Attention 4 boosts the speed at which the model processes incoming prompts by 25 to 70 percent, according to Google. Time to first token drops by up to 31 percent. Google also fixed bugs in tool calling, the feature that lets the model trigger external tools on its own.","Google says it also cut down on cases where the model would cut answers short or return incomplete responses. For image processing, users can manually raise the \"max_soft_tokens\" parameter from 280 to 1,120 to get sharper OCR results and support resolutions up to 2.51 megapixels. Google put up an interactive configurator on Hugging Face：https://huggingface.co/spaces/google/gemma4_vision_token_budget for that. The published benchmarks only compare the 31B and E4B variants against their predecessors, but the Hugging Face repository：https://huggingface.co/collections/google/gemma-4 shows that all parameter sizes in this model generation got updated, including the newest 12B release：https://the-decoder.com/google-deepminds-gemma-4-12b-squeezes-multimodal-ai-onto-a-laptop-with-just-16-gb-of-ram/. The community has pushed back on Google shipping the update under the same \"Gemma 4\" name：https://the-decoder.com/googles-gemma-4-is-now-available-with-apache-2-0-licensing-for-the-first-time/ instead of tagging it as a separate version like \"Gemma 4.1.\"","Stay in the loop on AI. Clear, useful, no fluff.","Follow The Decoder for AI news, background stories and expert analyses.","The Decoder：https://the-decoder.com/"],"articleImages":[{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/07/Google-Gemma-4-31B-net-improvements-scaled-1.jpeg","alt":"Gemma 4 31B: BFCL 74,2 %(+0,4), TB2 25,8 %(+4,5), Tau2 Retail 77,6 %(+3,1), Airline 84 %(+2,0), Telecom 62,7 %(+10,1).","afterParagraph":0,"url":"/media/articles/cmrnb8azf00t8bic685er9bs3/2137a614a06bab7e.jpg"}],"mediaStatus":"ok","articleBodyZh":["谷歌发布了其开源 AI 模型 Gemma 4 的更新，该更新加快了在 Nvidia Hopper GPU 上的性能，修复了工具调用错误，并解决了回答被截断的问题。谷歌表示，开启 Flash Attention 4 可以将模型处理输入提示的速度提升 25% 到 70%。首次生成令牌的时间最多可减少 31%。谷歌还修复了工具调用功能中的错误，该功能允许模型自行触发外部工具。","谷歌表示，还减少了模型回答被截断或返回不完整响应的情况。对于图像处理，用户可以手动将 \"max_soft_tokens\" 参数从 280 提高到 1,120，以获得更清晰的 OCR 结果，并支持最高 2.51 兆像素的分辨率。谷歌在 Hugging Face 上提供了一个互动配置器：https://huggingface.co/spaces/google/gemma4_vision_token_budget。已发布的基准测试仅比较了 31B 和 E4B 版本与其前代模型的性能，但 Hugging Face 仓库：https://huggingface.co/collections/google/gemma-4 显示该代模型所有参数大小都已更新，包括最新的 12B 版本：https://the-decoder.com/google-deepminds-gemma-4-12b-squeezes-multimodal-ai-on-to-a-laptop-with-just-16-gb-of-ram/。社区对谷歌以相同的 \"Gemma 4\" 名称发布更新表示不满，而不是将其标记为单独的版本，如 \"Gemma 4.1\"：https://the-decoder.com/googles-gemma-4-is-now-available-with-apache-2-0-licensing-for-the-first-time/。","保持 AI 动态。内容清晰、有用，无废话。","关注 The Decoder 获取 AI 新闻、背景故事和专家分析。","The Decoder：https://the-decoder.com/"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Aioga 编辑摘要：Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。 Aioga 将其归入「模型更新」方向，重点关注它对真实使用和行业竞争的影响。","background":"背景分析：模型与研究类动态需要结合能力边界、开放方式、成本、可用性和真实任务表现判断，单项指标领先不等于已经形成稳定采用。","viewpoint":"Aioga 判断：这条动态更适合作为行业观察信号，当前信息足以建立线索，但不足以推导长期结论。","implications":"影响分析：对相关团队而言，短期应先核对来源、可用范围和实际成本，再判断是否值得接入或跟进。","nextStep":"后续观察：继续观察官方文档、实际可用性、价格变化、开发者反馈和竞品回应。","evidenceRefs":["title","summary","articleBody"],"confidence":"medium","status":"published","aiGenerated":false,"autoApproved":true,"generatedBy":"rule-safe-fallback","generatedAt":"2026-07-23T08:10:16.875Z","sourceHash":"708f4b6f511052a9","validation":{"passed":true,"mode":"rule-safe-fallback","checks":["schema","length","source-attribution","no-html"]}},"tags":["模型更新","The Decoder：AI News（RSS）"],"translations":{"zh-CN":{"title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","summary":"Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能 - Aioga AI资讯","description":"Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","url":"https://www.aioga.com/news/cmrnb8azf00t8bic685er9bs3/"},"en":{"title":"Gemma 4 gets a stealth update that fixes tool calling bugs and truncated responses under the same name","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under Models. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"Models","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 gets a stealth update that fixes tool calling bugs and truncated responses under the same name - Aioga AI News","description":"Aioga tracks this update from The Decoder：AI News（RSS） under Models. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工","url":"https://www.aioga.com/en/news/cmrnb8azf00t8bic685er9bs3/"},"ja":{"title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","summary":"Aiogaは「モデル更新」の動きとして、The Decoder：AI News（RSS） からの更新を追跡しています。Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"モデル更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能 - Aioga AIニュース","description":"Aiogaは「モデル更新」の動きとして、The Decoder：AI News（RSS） からの更新を追跡しています。Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断","url":"https://www.aioga.com/ja/news/cmrnb8azf00t8bic685er9bs3/"},"ko":{"title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","summary":"Aioga는 The Decoder：AI News（RSS）의 업데이트를 모델 업데이트 흐름으로 추적합니다. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"모델 업데이트","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能 - Aioga AI 뉴스","description":"Aioga는 The Decoder：AI News（RSS）의 업데이트를 모델 업데이트 흐름으로 추적합니다. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断","url":"https://www.aioga.com/ko/news/cmrnb8azf00t8bic685er9bs3/"},"es":{"title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","summary":"Aioga sigue esta actualización de The Decoder：AI News（RSS） dentro de Modelos. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"Modelos","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能 - Aioga Noticias de IA","description":"Aioga sigue esta actualización de The Decoder：AI News（RSS） dentro de Modelos. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31","url":"https://www.aioga.com/es/news/cmrnb8azf00t8bic685er9bs3/"},"fr":{"title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","summary":"Aioga suit cette mise à jour de The Decoder：AI News（RSS） dans la catégorie Modèles. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"Modèles","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能 - Aioga Actualités IA","description":"Aioga suit cette mise à jour de The Decoder：AI News（RSS） dans la catégorie Modèles. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延","url":"https://www.aioga.com/fr/news/cmrnb8azf00t8bic685er9bs3/"},"de":{"title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能 - Aioga KI-News","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调","url":"https://www.aioga.com/de/news/cmrnb8azf00t8bic685er9bs3/"},"pt-BR":{"title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能 - Aioga Notícias de IA","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调","url":"https://www.aioga.com/pt-BR/news/cmrnb8azf00t8bic685er9bs3/"},"ru":{"title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能 - Aioga Новости ИИ","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调","url":"https://www.aioga.com/ru/news/cmrnb8azf00t8bic685er9bs3/"},"ar":{"title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能 - Aioga أخبار الذكاء الاصطناعي","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调","url":"https://www.aioga.com/ar/news/cmrnb8azf00t8bic685er9bs3/"},"hi":{"title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能 - Aioga AI समाचार","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调","url":"https://www.aioga.com/hi/news/cmrnb8azf00t8bic685er9bs3/"},"it":{"title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能 - Aioga Notizie IA","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调","url":"https://www.aioga.com/it/news/cmrnb8azf00t8bic685er9bs3/"},"nl":{"title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能 - Aioga AI-nieuws","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调","url":"https://www.aioga.com/nl/news/cmrnb8azf00t8bic685er9bs3/"},"tr":{"title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能 - Aioga AI Haberleri","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调","url":"https://www.aioga.com/tr/news/cmrnb8azf00t8bic685er9bs3/"},"vi":{"title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能 - Tin tức AI Aioga","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调","url":"https://www.aioga.com/vi/news/cmrnb8azf00t8bic685er9bs3/"},"id":{"title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能 - Berita AI Aioga","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调","url":"https://www.aioga.com/id/news/cmrnb8azf00t8bic685er9bs3/"},"th":{"title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能 - ข่าว AI Aioga","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调","url":"https://www.aioga.com/th/news/cmrnb8azf00t8bic685er9bs3/"},"pl":{"title":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调用错误和响应截断问题，Gemma 4 31B 在电信场景的智能体推理性能提升 10.1%。","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Gemma 4 静默更新：修复工具调用错误并提升 Hopper GPU 性能 - Aioga Wiadomości AI","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 模型更新. Google 为开源模型 Gemma 4 推送更新，在 Nvidia Hopper GPU 上启用 Flash Attention 4 后，提示词处理速度提升 25-70%，首 token 延迟降低 31%。更新还修复了工具调","url":"https://www.aioga.com/pl/news/cmrnb8azf00t8bic685er9bs3/"}}}}