{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-09-21T03:00:43.626Z","headline":"Google 员工分享如何设计可信的 AI 评测的五条规则","description":"Google Developer Relations 团队的 Jan-Felix Schmakeit 介绍为 Google 产品的 Agent Skills 设计评测的方法，主张像写单元测试一样建立结构化、自动化的评测流程，而非在终端做 vibe testing。","url":"https://www.aioga.com/news/cmtiwj88v070uro9y27rlys4c/","mainEntityOfPage":"https://www.aioga.com/news/cmtiwj88v070uro9y27rlys4c/","datePublished":"2026-09-01T16:35:00.000Z","dateModified":"2026-09-01T16:35:00.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://dev.to/googleai/how-to-design-ai-evaluations-you-can-actually-trust-41c3","https://aihot.virxact.com/items/cmtiwj88v070uro9y27rlys4c"],"canonicalUrl":"https://www.aioga.com/news/cmtiwj88v070uro9y27rlys4c/","directAnswer":{"@type":"Answer","text":"Google Developer Relations 团队成员 Jan-Felix Schmakeit 分享了为 Google 产品 Agent Skills 设计 AI 评测的方法，主张建立类似单元测试的结构化、自动化流程，减少在终端进行“凭感觉测试”。","url":"https://www.aioga.com/news/cmtiwj88v070uro9y27rlys4c/","dateCreated":"2026-09-01T16:35:00.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"dev.to source article","url":"https://dev.to/googleai/how-to-design-ai-evaluations-you-can-actually-trust-41c3","datePublished":"2026-09-01T16:35:00.000Z","provider":{"@type":"Organization","name":"dev.to","url":"https://dev.to/googleai/how-to-design-ai-evaluations-you-can-actually-trust-41c3"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmtiwj88v070uro9y27rlys4c","datePublished":"2026-09-01T16:35:00.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmtiwj88v070uro9y27rlys4c"}}],"aggregationSource":"Google AI：DEV 作者专属（RSS）","originalPublisher":{"name":"dev.to","url":"https://dev.to/googleai/how-to-design-ai-evaluations-you-can-actually-trust-41c3"},"geoDeepAnswer":null,"article":{"id":"cmtiwj88v070uro9y27rlys4c","slug":"cmtiwj88v070uro9y27rlys4c","url":"https://www.aioga.com/news/cmtiwj88v070uro9y27rlys4c/","title":"Google 员工分享如何设计可信的 AI 评测的五条规则","title_en":"","summary":"Google Developer Relations 团队的 Jan-Felix Schmakeit 介绍为 Google 产品的 Agent Skills 设计评测的方法，主张像写单元测试一样建立结构化、自动化的评测流程，而非在终端做 vibe testing。","source":"Google AI：DEV 作者专属（RSS）","sourceUrl":"https://dev.to/googleai/how-to-design-ai-evaluations-you-can-actually-trust-41c3","aiHotUrl":"https://aihot.virxact.com/items/cmtiwj88v070uro9y27rlys4c","publishedAt":"2026-09-01T16:35:00.000Z","category":"行业动态","score":58,"selected":false,"articleBody":["As part of my work at Google, we are publishing a suite of Agent Skills for Google products and technologies on GitHub：https://github.com/google/skills. These agent skills：https://agentskills.io/home are designed to help AI agents interact with our technologies. But how do you test that these skills are useful and work as expected? My team in Developer Relations has been focused on this question, because having reliable signals on their performance is critical to help us improve them over time.","Just as you wouldn't deploy a production API without writing unit tests, you should apply the same standard to your AI agents. As Joe Spiro showed in the Designing AI Evals ：https://dev.to/googleai/designing-ai-evals-clarity-now-and-visualization-next-4eii post series, scaling AI tools means moving beyond \"vibe testing\" in a terminal. Instead, you should set up a structured, automated evaluation pipeline to benchmark your integration. The evaluations (evals) are the actions you asked the agent to perform, which are graded using scorers (for example rubrics) that assert whether the agent succeeded. We'll focus on evaluations in this post and tackle tips for scoring rubrics in the next post.","However, AI evaluations cost real tokens. You need to make sure that you use these tokens as efficiently as possible. They need to provide real value that helps you build better tools. Writing good evaluations is critical. Poor evaluations provide false signals, waste your token budget, and create noise in your metrics.","Here are five rules we learned to design better evaluations you can trust. Follow them to ensure that every token you spend produces a useful metric.","Before writing evaluations, you need to understand the setup and limitations of your chosen framework. This includes systems like Harbor, Inspect AI, or integrations in development tools like in the Agent Development Kit：https://docs.cloud.google.com/gemini-enterprise-agent-platform/optimize/evaluation/agent-evaluation. Does it use an ephemeral sandbox? What tools are available? How is the output captured","If your evaluations show a high baseline accuracy (i.e., without your agent tool), it might not prove its value, or the evaluation prompts are too easy.","You cannot grade an agent on something you did not explicitly ask it to do. Your evaluation prompts and graders should be complementary. This means that they should only test for things included in the prompt.","Agents possess inherent model knowledge and might skip your custom tools entirely to arrive at the correct answer. (That on its own is some useful feedback!)","A strong evaluation suite tests real, diverse use cases. But testing the same capability repeatedly causes overfitting and creates noisy metrics.","You cannot improve AI tools if you can't measure them accurately. If you treat AI evaluations with the same focus as traditional unit tests, you improve the quality of your metrics and get more robust signals.","By applying these five rules, you eliminate false signals that waste your token budget. Instead of generating noise, your test suite gives you actionable feedback you can use to guide your engineering decisions and improve your tools.","Figuring out what to test is only the first step. A well-designed evaluation is only useful if the scorer grading answers is reliable and returns meaningful results. In my next post, we will look at how to test. You will learn how to write lean, atomic rubrics that minimize ambiguity for an LLM grader and make every token count.","Photo by William Warby：https://unsplash.com/@wwarby on Unsplash：https://unsplash.com/photos/gray-and-yellow-measures-WahfNoqbYnM","Templates let you quickly answer FAQs or store snippets for re-use.","Are you sure you want to hide this comment? It will become hidden in your post, but will still be visible via the comment's permalink：#.","For further actions, you may consider blocking this person and/or reporting abuse：/report-abuse","Google AI Studio is the fastest way to start building with Gemini. Ready to build?","DEV Community：/ — A space to discuss and keep up software development and manage your software career","Built on Forem：https://www.forem.com — the open source：https://dev.to/t/opensource software that powers DEV：https://dev.to and other inclusive communities.","Made with love and Ruby on Rails：https://dev.to/t/rails. DEV Community &copy; 2016 - 2026.","We're a place where coders share, stay up-to-date and grow their careers."],"articleImages":[{"sourceUrl":"https://media2.dev.to/dynamic/image/width=256,height=,fit=scale-down,gravity=auto,format=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F8j7kvp660rqzt99zui8e.png","alt":"pic","afterParagraph":12,"url":"/media/articles/cmtiwj88v070uro9y27rlys4c/f75d1e7bc8b434f4.webp"},{"sourceUrl":"https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg","alt":"","afterParagraph":20,"url":"/media/articles/cmtiwj88v070uro9y27rlys4c/ae2a1867ed3e47a7.jpg"},{"sourceUrl":"https://assets.dev.to/assets/exploding-head-daceb38d627e6ae9b730f36a1e390fca556a4289d5a41abb2c35068ad3e2c4b5.svg","alt":"","afterParagraph":20,"url":"/media/articles/cmtiwj88v070uro9y27rlys4c/c9631abcb96b6799.jpg"}],"mediaStatus":"ok","articleBodyZh":["作为我在谷歌工作的一部分，我们在 GitHub 上发布了一套针对谷歌产品和技术的代理技能：https://github.com/google/skills。这些代理技能：https://agentskills.io/home 旨在帮助 AI 代理与我们的技术进行交互。但你如何测试这些技能是否有用并按预期工作呢？我在开发者关系团队一直专注于这个问题，因为拥有关于它们性能的可靠信号对于帮助我们随着时间改进这些技能至关重要。","正如你不会在没有编写单元测试的情况下部署生产 API 一样，你也应该对 AI 代理应用同样的标准。正如 Joe Spiro 在“设计 AI 评估”系列文章中展示的：https://dev.to/googleai/designing-ai-evals-clarity-now-and-visualization-next-4eii，扩展 AI 工具意味着要超越在终端中“凭感觉测试”。相反，你应该建立一个结构化的、自动化的评估管道来基准测试你的集成。评估（evals）是你要求代理执行的动作，这些动作使用评分器（例如评分标准）评分，以判断代理是否成功。本文将重点介绍评估，并在下一篇文章中讨论评分标准的技巧。","然而，AI 评估需要消耗真实的代币。你需要确保尽可能高效地使用这些代币。它们需要提供真正的价值，帮助你构建更好的工具。编写好的评估至关重要。糟糕的评估会提供错误信号，浪费你的代币预算，并在指标中产生噪音。","以下是我们学到的五条设计更可靠评估规则。遵循它们可以确保你花费的每一个代币都产生有用的指标。","在编写评估之前，你需要了解所选框架的设置和限制。这包括 Harbor、Inspect AI 等系统，或在开发工具中的集成，如代理开发工具包中的集成：https://docs.cloud.google.com/gemini-enterprise-agent-platform/optimize/evaluation/agent-evaluation。它是否使用临时沙箱？有哪些可用工具？如何捕获输出？","如果你的评估显示出较高的基线准确率（即没有你的代理工具时），可能无法证明它的价值，或者评估提示过于简单。","你不能根据你没有明确要求代理完成的事情来对其打分。你的评估提示和评分者应该是互补的。这意味着它们只应该测试提示中包含的内容。","代理具有固有的模型知识，可能会完全跳过你的自定义工具来得出正确答案。（这本身就是一些有用的反馈！）","强大的评估套件测试真实、多样化的用例。但重复测试相同的能力会导致过拟合并产生噪音指标。","如果你不能准确衡量 AI 工具，就无法改进它们。如果你像对待传统单元测试一样关注 AI 评估，你将提高指标的质量并获得更稳健的信号。","通过应用这五条规则，你可以消除浪费令牌预算的虚假信号。你的测试套件将不再产生噪音，而是提供可行的反馈，用于指导你的工程决策并改进工具。","确定测试内容只是第一步。只有当评分答案的评分者可靠并返回有意义的结果时，精心设计的评估才有用。在我的下一篇文章中，我们将探讨如何测试。你将学习如何编写精简、原子化的评分标准，最大限度地减少 LLM 评分者的歧义，让每个令牌都发挥作用。","照片由 William Warby 拍摄：https://unsplash.com/@wwarby，来自 Unsplash：https://unsplash.com/photos/gray-and-yellow-measures-WahfNoqbYnM","模板让你可以快速回答常见问题或存储片段以便重复使用。","你确定要隐藏这条评论吗？它将在你的帖子中被隐藏，但仍可通过评论的永久链接查看：#。","如需进一步操作，你可以考虑屏蔽此人和/或举报滥用：/report-abuse","Google AI Studio 是使用 Gemini 开始构建的最快方式。准备好开始了吗？","DEV Community：/ — 一个讨论、跟进软件开发并管理软件职业的空间","基于 Forem 构建：https://www.forem.com — 开源软件：https://dev.to/t/opensource，支持 DEV：https://dev.to 及其他包容性社区。","用爱和 Ruby on Rails 制作：https://dev.to/t/rails。DEV 社区 &copy; 2016 - 2026。","我们是一个程序员分享、保持最新信息并发展职业的地方。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Google Developer Relations 团队成员 Jan-Felix Schmakeit 分享了为 Google 产品 Agent Skills 设计 AI 评测的方法，主张建立类似单元测试的结构化、自动化流程，减少在终端进行“凭感觉测试”。","background":"Google 正在发布面向其产品和技术的 Agent Skills。文章将评测定义为要求代理完成的行动，并通过评分器判断是否成功，同时提醒评测会消耗真实 token，设计不当可能产生错误信号和指标噪声。","viewpoint":"Aioga 判断：文章的核心价值在于把代理评测从一次性体验判断转向可重复的测试流程，但评测结果仍取决于框架限制、任务难度、提示与评分标准是否匹配。","implications":"可能影响：采用代理评测的团队需要先确认沙盒、可用工具和输出捕获方式，并检查基础准确率是否过高、任务是否过易；高分不代表自定义工具一定被使用，也不足以单独证明代理工具的价值。","nextStep":"后续观察：应关注文章后续关于评分量规的讨论，以及 Google Agent Skills 评测流程如何处理工具调用、输出捕获和基线结果等问题。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-09-01T17:43:30.631Z","sourceHash":"2eef1d519e8ada25","review":{"approved":true,"groundedness":95,"clarity":92,"duplicationRisk":18,"blockingIssues":[],"notes":["“核心价值”属于观点判断，已明确标注为 Aioga 判断，不构成事实冒充。","“高分不代表自定义工具一定被使用”是基于原文关于代理可能跳过自定义工具的合理概括。","“后续观察”中的关注建议属于编辑性延伸，来源未明确说明具体评测流程将如何处理这些问题，但未被表述为既成事实。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","editorial-labels","inference-boundary","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","Google AI：DEV 作者专属（RSS）"],"translations":{"zh-CN":{"title":"Google 员工分享如何设计可信的 AI 评测的五条规则","summary":"Google Developer Relations 团队的 Jan-Felix Schmakeit 介绍为 Google 产品的 Agent Skills 设计评测的方法，主张像写单元测试一样建立结构化、自动化的评测流程，而非在终端做 vibe testing。","category":"行业动态","source":"dev.to","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Google 员工分享如何设计可信的 AI 评测的五条规则 - Aioga AI资讯","description":"Google Developer Relations 团队的 Jan-Felix Schmakeit 介绍为 Google 产品的 Agent Skills 设计评测的方法，主张像写单元测试一样建立结构化、自动化的评测流程，而非在终端做 vibe testing。","url":"https://www.aioga.com/news/cmtiwj88v070uro9y27rlys4c/","articleBody":["作为我在谷歌工作的一部分，我们在 GitHub 上发布了一套针对谷歌产品和技术的代理技能：https://github.com/google/skills。这些代理技能：https://agentskills.io/home 旨在帮助 AI 代理与我们的技术进行交互。但你如何测试这些技能是否有用并按预期工作呢？我在开发者关系团队一直专注于这个问题，因为拥有关于它们性能的可靠信号对于帮助我们随着时间改进这些技能至关重要。","正如你不会在没有编写单元测试的情况下部署生产 API 一样，你也应该对 AI 代理应用同样的标准。正如 Joe Spiro 在“设计 AI 评估”系列文章中展示的：https://dev.to/googleai/designing-ai-evals-clarity-now-and-visualization-next-4eii，扩展 AI 工具意味着要超越在终端中“凭感觉测试”。相反，你应该建立一个结构化的、自动化的评估管道来基准测试你的集成。评估（evals）是你要求代理执行的动作，这些动作使用评分器（例如评分标准）评分，以判断代理是否成功。本文将重点介绍评估，并在下一篇文章中讨论评分标准的技巧。","然而，AI 评估需要消耗真实的代币。你需要确保尽可能高效地使用这些代币。它们需要提供真正的价值，帮助你构建更好的工具。编写好的评估至关重要。糟糕的评估会提供错误信号，浪费你的代币预算，并在指标中产生噪音。","以下是我们学到的五条设计更可靠评估规则。遵循它们可以确保你花费的每一个代币都产生有用的指标。","在编写评估之前，你需要了解所选框架的设置和限制。这包括 Harbor、Inspect AI 等系统，或在开发工具中的集成，如代理开发工具包中的集成：https://docs.cloud.google.com/gemini-enterprise-agent-platform/optimize/evaluation/agent-evaluation。它是否使用临时沙箱？有哪些可用工具？如何捕获输出？","如果你的评估显示出较高的基线准确率（即没有你的代理工具时），可能无法证明它的价值，或者评估提示过于简单。","你不能根据你没有明确要求代理完成的事情来对其打分。你的评估提示和评分者应该是互补的。这意味着它们只应该测试提示中包含的内容。","代理具有固有的模型知识，可能会完全跳过你的自定义工具来得出正确答案。（这本身就是一些有用的反馈！）","强大的评估套件测试真实、多样化的用例。但重复测试相同的能力会导致过拟合并产生噪音指标。","如果你不能准确衡量 AI 工具，就无法改进它们。如果你像对待传统单元测试一样关注 AI 评估，你将提高指标的质量并获得更稳健的信号。","通过应用这五条规则，你可以消除浪费令牌预算的虚假信号。你的测试套件将不再产生噪音，而是提供可行的反馈，用于指导你的工程决策并改进工具。","确定测试内容只是第一步。只有当评分答案的评分者可靠并返回有意义的结果时，精心设计的评估才有用。在我的下一篇文章中，我们将探讨如何测试。你将学习如何编写精简、原子化的评分标准，最大限度地减少 LLM 评分者的歧义，让每个令牌都发挥作用。","照片由 William Warby 拍摄：https://unsplash.com/@wwarby，来自 Unsplash：https://unsplash.com/photos/gray-and-yellow-measures-WahfNoqbYnM","模板让你可以快速回答常见问题或存储片段以便重复使用。","你确定要隐藏这条评论吗？它将在你的帖子中被隐藏，但仍可通过评论的永久链接查看：#。","如需进一步操作，你可以考虑屏蔽此人和/或举报滥用：/report-abuse","Google AI Studio 是使用 Gemini 开始构建的最快方式。准备好开始了吗？","DEV Community：/ — 一个讨论、跟进软件开发并管理软件职业的空间","基于 Forem 构建：https://www.forem.com — 开源软件：https://dev.to/t/opensource，支持 DEV：https://dev.to 及其他包容性社区。","用爱和 Ruby on Rails 制作：https://dev.to/t/rails。DEV 社区 &copy; 2016 - 2026。","我们是一个程序员分享、保持最新信息并发展职业的地方。"]},"en":{"title":"Google employees share five rules for designing trustworthy AI reviews","summary":"Jan-Felix Schmakeit from Google's Developer Relations team introduces methods for designing evaluation for Google's Agent Skills, advocating for a structured, automated evaluation process similar to writing unit tests, rather than vibe testing on the terminal.","category":"Industry","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Google employees share five rules for designing trustworthy AI reviews - Aioga AI News","description":"Jan-Felix Schmakeit from Google's Developer Relations team introduces methods for designing evaluation for Google's Agent Skills, advocating for a structured, automated evaluation...","url":"https://www.aioga.com/en/news/cmtiwj88v070uro9y27rlys4c/","contentTranslated":true,"sourceHash":"78121c4bfa5172c3","translatedAt":"2026-09-01T17:22:35.019Z"},"ja":{"title":"Googleの従業員が信頼できるAIレビューをデザインするための5つのルールを共有しています","summary":"Googleの開発者関係チームのJan-Felix Schmakeitは、Googleのエージェントスキルの評価設計手法を紹介し、端末でのバイブテストではなく、ユニットテストを書くような構造化された自動評価プロセスを提唱しています。","category":"業界動向","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Googleの従業員が信頼できるAIレビューをデザインするための5つのルールを共有しています - Aioga AIニュース","description":"Googleの開発者関係チームのJan-Felix Schmakeitは、Googleのエージェントスキルの評価設計手法を紹介し、端末でのバイブテストではなく、ユニットテストを書くような構造化された自動評価プロセスを提唱しています。","url":"https://www.aioga.com/ja/news/cmtiwj88v070uro9y27rlys4c/","contentTranslated":true,"sourceHash":"78121c4bfa5172c3","translatedAt":"2026-09-01T17:22:35.468Z"},"ko":{"title":"구글 직원들이 신뢰할 수 있는 AI 리뷰를 설계하기 위한 다섯 가지 규칙을 공유합니다","summary":"구글 개발자 관계팀의 얀-펠릭스 슈메이키트는 구글 에이전트 스킬 평가 설계 방법을 소개하며, 터미널에서의 분위기 테스트 대신 단위 테스트 작성과 유사한 구조적이고 자동화된 평가 프로세스를 권장합니다.","category":"업계 동향","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"구글 직원들이 신뢰할 수 있는 AI 리뷰를 설계하기 위한 다섯 가지 규칙을 공유합니다 - Aioga AI 뉴스","description":"구글 개발자 관계팀의 얀-펠릭스 슈메이키트는 구글 에이전트 스킬 평가 설계 방법을 소개하며, 터미널에서의 분위기 테스트 대신 단위 테스트 작성과 유사한 구조적이고 자동화된 평가 프로세스를 권장합니다.","url":"https://www.aioga.com/ko/news/cmtiwj88v070uro9y27rlys4c/","contentTranslated":true,"sourceHash":"78121c4bfa5172c3","translatedAt":"2026-09-01T17:22:44.422Z"},"es":{"title":"Los empleados de Google comparten cinco reglas para diseñar reseñas de IA fiables","summary":"Jan-Felix Schmakeit, del equipo de Relaciones con Desarrolladores de Google, presenta métodos para diseñar la evaluación de las Habilidades de Agente de Google, defendiendo un proceso de evaluación estructurado y automatizado similar a escribir pruebas unitarias, en lugar de pruebas de vibración en el terminal.","category":"Industria","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Los empleados de Google comparten cinco reglas para diseñar reseñas de IA fiables - Aioga Noticias de IA","description":"Jan-Felix Schmakeit, del equipo de Relaciones con Desarrolladores de Google, presenta métodos para diseñar la evaluación de las Habilidades de Agente de Google, defendiendo un proc...","url":"https://www.aioga.com/es/news/cmtiwj88v070uro9y27rlys4c/","contentTranslated":true,"sourceHash":"78121c4bfa5172c3","translatedAt":"2026-09-01T17:22:44.537Z"},"fr":{"title":"Les employés de Google partagent cinq règles pour concevoir des avis fiables sur l’IA","summary":"Jan-Felix Schmakeit, de l’équipe Relations Développeurs de Google, présente des méthodes pour concevoir l’évaluation des compétences d’agent de Google, en préconisant un processus d’évaluation structuré et automatisé, similaire à l’écriture de tests unitaires, plutôt que des tests d’ambiance sur le terminal.","category":"Industrie","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Les employés de Google partagent cinq règles pour concevoir des avis fiables sur l’IA - Aioga Actualités IA","description":"Jan-Felix Schmakeit, de l’équipe Relations Développeurs de Google, présente des méthodes pour concevoir l’évaluation des compétences d’agent de Google, en préconisant un processus...","url":"https://www.aioga.com/fr/news/cmtiwj88v070uro9y27rlys4c/","contentTranslated":true,"sourceHash":"78121c4bfa5172c3","translatedAt":"2026-09-01T17:22:52.409Z"},"de":{"title":"Google-Mitarbeiter teilen fünf Regeln für die Erstellung vertrauenswürdiger KI-Bewertungen","summary":"Jan-Felix Schmakeit vom Developer Relations-Team von Google stellt Methoden zur Gestaltung der Bewertung für Googles Agent Skills ein und plädiert für einen strukturierten, automatisierten Bewertungsprozess, ähnlich dem Schreiben von Unit-Tests, anstatt Vibe-Tests am Terminal.","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Google-Mitarbeiter teilen fünf Regeln für die Erstellung vertrauenswürdiger KI-Bewertungen - Aioga KI-News","description":"Jan-Felix Schmakeit vom Developer Relations-Team von Google stellt Methoden zur Gestaltung der Bewertung für Googles Agent Skills ein und plädiert für einen strukturierten, automat...","url":"https://www.aioga.com/de/news/cmtiwj88v070uro9y27rlys4c/","contentTranslated":true,"sourceHash":"78121c4bfa5172c3","translatedAt":"2026-09-01T17:22:53.445Z"},"pt-BR":{"title":"Funcionários do Google compartilham cinco regras para criar avaliações confiáveis de IA","summary":"Jan-Felix Schmakeit, da equipe de Relações com Desenvolvedores do Google, apresenta métodos para projetar a avaliação das Habilidades de Agente do Google, defendendo um processo estruturado e automatizado de avaliação, semelhante à escrita de testes unitários, em vez de testes de vibração no terminal.","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Funcionários do Google compartilham cinco regras para criar avaliações confiáveis de IA - Aioga Notícias de IA","description":"Jan-Felix Schmakeit, da equipe de Relações com Desenvolvedores do Google, apresenta métodos para projetar a avaliação das Habilidades de Agente do Google, defendendo um processo es...","url":"https://www.aioga.com/pt-BR/news/cmtiwj88v070uro9y27rlys4c/","contentTranslated":true,"sourceHash":"78121c4bfa5172c3","translatedAt":"2026-09-01T17:23:02.545Z"},"ru":{"title":"Сотрудники Google делятся пятью правилами для создания надёжных отзывов об ИИ","summary":"Ян-Феликс Шмакейт из команды по работе с разработчиками Google представляет методы разработки оценки для навыков агентов Google, выступая за структурированный, автоматизированный процесс оценки, похожий на написание модульных тестов, вместо тестирования вибрации на терминале.","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Сотрудники Google делятся пятью правилами для создания надёжных отзывов об ИИ - Aioga Новости ИИ","description":"Ян-Феликс Шмакейт из команды по работе с разработчиками Google представляет методы разработки оценки для навыков агентов Google, выступая за структурированный, автоматизированный п...","url":"https://www.aioga.com/ru/news/cmtiwj88v070uro9y27rlys4c/","contentTranslated":true,"sourceHash":"78121c4bfa5172c3","translatedAt":"2026-09-01T17:23:02.590Z"},"ar":{"title":"موظفو جوجل يشاركون خمس قواعد لتصميم مراجعات ذكاء اصطناعي موثوقة","summary":"يقدم يان-فيليكس شماكيت من فريق علاقات المطورين في جوجل طرقا لتصميم التقييم لمهارات الوكلاء في جوجل، ويدعو إلى عملية تقييم منظمة وآلية مشابهة لكتابة اختبارات الوحدة، بدلا من اختبار الاهتزاز على الطرفية.","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"موظفو جوجل يشاركون خمس قواعد لتصميم مراجعات ذكاء اصطناعي موثوقة - Aioga أخبار الذكاء الاصطناعي","description":"يقدم يان-فيليكس شماكيت من فريق علاقات المطورين في جوجل طرقا لتصميم التقييم لمهارات الوكلاء في جوجل، ويدعو إلى عملية تقييم منظمة وآلية مشابهة لكتابة اختبارات الوحدة، بدلا من اختبار...","url":"https://www.aioga.com/ar/news/cmtiwj88v070uro9y27rlys4c/","contentTranslated":true,"sourceHash":"78121c4bfa5172c3","translatedAt":"2026-09-01T17:23:11.100Z"},"hi":{"title":"Google के कर्मचारी भरोसेमंद AI समीक्षाएँ डिज़ाइन करने के लिए पाँच नियम साझा करते हैं","summary":"Google की डेवलपर रिलेशंस टीम के जान-फेलिक्स श्मेकिट ने Google के एजेंट कौशल के लिए मूल्यांकन डिजाइन करने के तरीकों का परिचय दिया, जो टर्मिनल पर वाइब परीक्षण के बजाय यूनिट परीक्षण लिखने के समान एक संरचित, स्वचालित मूल्यांकन प्रक्रिया की वकालत करता है।","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Google के कर्मचारी भरोसेमंद AI समीक्षाएँ डिज़ाइन करने के लिए पाँच नियम साझा करते हैं - Aioga AI समाचार","description":"Google की डेवलपर रिलेशंस टीम के जान-फेलिक्स श्मेकिट ने Google के एजेंट कौशल के लिए मूल्यांकन डिजाइन करने के तरीकों का परिचय दिया, जो टर्मिनल पर वाइब परीक्षण के बजाय यूनिट परीक्षण ल...","url":"https://www.aioga.com/hi/news/cmtiwj88v070uro9y27rlys4c/","contentTranslated":true,"sourceHash":"78121c4bfa5172c3","translatedAt":"2026-09-01T17:23:11.740Z"},"it":{"title":"I dipendenti di Google condividono cinque regole per progettare recensioni affidabili sull'IA","summary":"Jan-Felix Schmakeit del team Developer Relations di Google introduce metodi per progettare la valutazione delle Competenze Agenti di Google, sostenendo un processo di valutazione strutturato e automatizzato simile alla scrittura di test unitari, piuttosto che ai vibe testing sul terminale.","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"I dipendenti di Google condividono cinque regole per progettare recensioni affidabili sull'IA - Aioga Notizie IA","description":"Jan-Felix Schmakeit del team Developer Relations di Google introduce metodi per progettare la valutazione delle Competenze Agenti di Google, sostenendo un processo di valutazione s...","url":"https://www.aioga.com/it/news/cmtiwj88v070uro9y27rlys4c/","contentTranslated":true,"sourceHash":"78121c4bfa5172c3","translatedAt":"2026-09-01T17:23:20.858Z"},"nl":{"title":"Google-medewerkers delen vijf regels voor het ontwerpen van betrouwbare AI-reviews","summary":"Jan-Felix Schmakeit van het Developer Relations-team van Google introduceert methoden voor het ontwerpen van evaluatie voor Google's Agent Skills, en pleit voor een gestructureerd, geautomatiseerd evaluatieproces vergelijkbaar met het schrijven van unittests, in plaats van vibe-testen op de terminal.","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Google-medewerkers delen vijf regels voor het ontwerpen van betrouwbare AI-reviews - Aioga AI-nieuws","description":"Jan-Felix Schmakeit van het Developer Relations-team van Google introduceert methoden voor het ontwerpen van evaluatie voor Google's Agent Skills, en pleit voor een gestructureerd,...","url":"https://www.aioga.com/nl/news/cmtiwj88v070uro9y27rlys4c/","contentTranslated":true,"sourceHash":"78121c4bfa5172c3","translatedAt":"2026-09-01T17:23:20.681Z"},"tr":{"title":"Google çalışanları, güvenilir yapay zeka incelemeleri tasarlamak için beş kural paylaşıyor","summary":"Google'ın Geliştirici İlişkileri ekibinden Jan-Felix Schmakeit, Google'ın Ajan Becerileri için değerlendirme tasarım yöntemlerini tanıtıyor; terminalde vibe testi yerine birim testleri yazmaya benzer yapılandırılmış, otomatik bir değerlendirme sürecini savunuyor.","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Google çalışanları, güvenilir yapay zeka incelemeleri tasarlamak için beş kural paylaşıyor - Aioga AI Haberleri","description":"Google'ın Geliştirici İlişkileri ekibinden Jan-Felix Schmakeit, Google'ın Ajan Becerileri için değerlendirme tasarım yöntemlerini tanıtıyor; terminalde vibe testi yerine birim test...","url":"https://www.aioga.com/tr/news/cmtiwj88v070uro9y27rlys4c/","contentTranslated":true,"sourceHash":"78121c4bfa5172c3","translatedAt":"2026-09-01T17:23:28.804Z"},"vi":{"title":"Nhân viên Google chia sẻ năm quy tắc để thiết kế các đánh giá AI đáng tin cậy","summary":"Jan-Felix Schmakeit từ nhóm Quan hệ Nhà phát triển của Google giới thiệu các phương pháp thiết kế đánh giá cho Kỹ năng Đại lý của Google, ủng hộ một quy trình đánh giá có cấu trúc, tự động tương tự như viết bài kiểm thử đơn vị, thay vì kiểm thử vibe trên thiết bị đầu cuối.","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Nhân viên Google chia sẻ năm quy tắc để thiết kế các đánh giá AI đáng tin cậy - Tin tức AI Aioga","description":"Jan-Felix Schmakeit từ nhóm Quan hệ Nhà phát triển của Google giới thiệu các phương pháp thiết kế đánh giá cho Kỹ năng Đại lý của Google, ủng hộ một quy trình đánh giá có cấu trúc,...","url":"https://www.aioga.com/vi/news/cmtiwj88v070uro9y27rlys4c/","contentTranslated":true,"sourceHash":"78121c4bfa5172c3","translatedAt":"2026-09-01T17:23:29.588Z"},"id":{"title":"Karyawan Google membagikan lima aturan untuk merancang ulasan AI yang dapat dipercaya","summary":"Jan-Felix Schmakeit dari tim Hubungan Pengembang Google memperkenalkan metode untuk merancang evaluasi untuk Agent Skills Google, menganjurkan proses evaluasi yang terstruktur dan otomatis mirip dengan menulis unit test, bukan vibe testing di terminal.","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Karyawan Google membagikan lima aturan untuk merancang ulasan AI yang dapat dipercaya - Berita AI Aioga","description":"Jan-Felix Schmakeit dari tim Hubungan Pengembang Google memperkenalkan metode untuk merancang evaluasi untuk Agent Skills Google, menganjurkan proses evaluasi yang terstruktur dan...","url":"https://www.aioga.com/id/news/cmtiwj88v070uro9y27rlys4c/","contentTranslated":true,"sourceHash":"78121c4bfa5172c3","translatedAt":"2026-09-01T17:23:38.601Z"},"th":{"title":"พนักงาน Google แชร์กฎ 5 ข้อสําหรับการออกแบบรีวิว AI ที่น่าเชื่อถือ","summary":"Jan-Felix Schmakeit จากทีม Developer Relations ของ Google แนะนําวิธีการออกแบบการประเมินผลสําหรับ Agent Skills ของ Google โดยสนับสนุนกระบวนการประเมินผลแบบอัตโนมัติที่มีโครงสร้างคล้ายกับการเขียน unit test แทนที่จะทดสอบ vibe บนเทอร์มินัล","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"พนักงาน Google แชร์กฎ 5 ข้อสําหรับการออกแบบรีวิว AI ที่น่าเชื่อถือ - ข่าว AI Aioga","description":"Jan-Felix Schmakeit จากทีม Developer Relations ของ Google แนะนําวิธีการออกแบบการประเมินผลสําหรับ Agent Skills ของ Google โดยสนับสนุนกระบวนการประเมินผลแบบอัตโนมัติที่มีโครงสร้างคล้า...","url":"https://www.aioga.com/th/news/cmtiwj88v070uro9y27rlys4c/","contentTranslated":true,"sourceHash":"78121c4bfa5172c3","translatedAt":"2026-09-01T17:23:37.741Z"},"pl":{"title":"Pracownicy Google dzielą się pięcioma zasadami dotyczącymi projektowania wiarygodnych recenzji AI","summary":"Jan-Felix Schmakeit z zespołu ds. relacji z deweloperami Google przedstawia metody projektowania ewaluacji dla umiejętności agentów Google, promując uporządkowany, zautomatyzowany proces oceny podobny do pisania testów jednostkowych, zamiast testów atmosfery na terminalu.","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Pracownicy Google dzielą się pięcioma zasadami dotyczącymi projektowania wiarygodnych recenzji AI - Aioga Wiadomości AI","description":"Jan-Felix Schmakeit z zespołu ds. relacji z deweloperami Google przedstawia metody projektowania ewaluacji dla umiejętności agentów Google, promując uporządkowany, zautomatyzowany...","url":"https://www.aioga.com/pl/news/cmtiwj88v070uro9y27rlys4c/","contentTranslated":true,"sourceHash":"78121c4bfa5172c3","translatedAt":"2026-09-01T17:23:47.208Z"}},"evidenceTier":"verified-news","reviewStatus":"automated-ingest","indexable":true,"editorialCover":""}}