宣布 Kotlin 版 ADK 1.0:在 Kotlin、Android 及更多平台构建生产就绪的 AI 代理
如何在 ADK 中评估实时和语音代理
推动开发者卓越:深入了解项目冲刺
When developers first work on harness engineering for agentic coding systems, they often fall into the same trap: they run common end-to-end benchmarks like Terminal-Bench and DeepSWE, watch a composite score move by a few percentage points, and have no idea why it changed.
End-to-end benchmarks are the de facto for evaluating model performance and determining what needs deeper investigation, but the challenge is that those investigations come at a high cost.
Behavioral evaluations are often a better measure of confidence on whether the behaviors you expect actually do happen and whether you’re moving in the right direction instead of backsliding when it comes to regressions or new model changes. They can serve as your iteration partner and help give insight into why certain changes move the needle in one way or another.
Here’s our take on behavioral evaluation, including approaches that have helped us keep agent systems reliable as models evolve.
Most teams evaluate AI agents like they would evaluate a student taking an exam. They hand the agent a large codebase, give it a time limit, and measure its success based on how many tests pass or fail.
When that score drops, what went wrong?
End-to-end benchmarks don’t typically directly answer these questions.
Behavioral evaluations function like integration tests for improving agent harness operation. When you have a rich enough behavioral eval set, you have a baseline for the behavior you're targeting from your agent, and you're able to iteratively improve the prompt to get there.
Instead of measuring whether the agent solved an entire multi-file refactor, a behavioral eval measures discrete, observable actions:
Instead of setting up a complex evaluation harness on day one, use this time to follow your hunches and run experiments.
When bootstrapping an agent from scratch, you start with developer instinct and dogfooding . Until you have built an agent capable of dogfooding its own codebase, handling boilerplate, writing its own markdown renderer, and executing routine developer tasks, it doesn’t make sense to run evaluations.
Evals belong to the second phase of development: ensuring forward progress and guarding against regressions .
The primary purpose of an evaluation suite is not to celebrate when you make the agent 2% better; it is to give you unshakeable confidence that a new prompt tweak, tool schema change, or model upgrade did not make the agent holistically worse .
A robust harness evaluation framework separates behavioral assertions into fast, deterministic, unit-style checks that run locally.
Shifting your focus to these smaller, observable actions creates a reliable safety net. You can confidently iterate on your system prompts or switch to a different model, because you’ll know immediately if you've accidentally broken a core behavior.
Behavioral evals assert on intermediate execution steps, like specific tool calls or file modifications, instead of final string equality:
With a rich suite of behavioral evals, you can automate your prompt engineering. For example, you can set up a loop where an LLM tweaks its own system prompt, iterating until a failing test finally passes, all while the rest of your test suite acts similar to how a CI/CD-style guardrail operates. This helps you ensure that the changes don’t break any existing features.
There are a few things you can do from the start to make this process repeatable. I suggest you start small with a three-step behavioral testing loop:
Your agent doesn't need a higher benchmark score to get started. It needs an evaluation harness that keeps it honest.
To build a stable, resilient harness, you have to stop treating your model like a black box passing a final exam, and start treating your harness like standard software that requires unit and integration testing.
While behavioral evaluations are a core pillar of harness engineering, they aren’t a replacement for larger, end-to-end evaluation suites. They’re actually complementary. Macro benchmarks verify the final destination and micro behavioral evals serve as a partner that enables safe, rapid iteration. When you adopt both, you'll have higher confidence levels when iterating, like when making prompt changes, building out new features, or even deploying brand-new models.
Announcing ADK for Kotlin 1.0: Building Production-Ready AI Agents in Kotlin, Android, and Beyond
How to Evaluate Live & Voice Agents in ADK
Driving Developer Excellence: Inside the Program Sprints
情报判断
Aioga 编辑摘要
Google 开发者博客介绍用于 AI 编码智能体的行为评估方法,建议在端到端基准之外,对工具调用、是否先运行验证器等中间执行步骤编写断言,以判断目标行为是否发生,并识别回归。
可能影响:评估体系可能需要同时覆盖端到端结果与关键行为断言;单一综合分数不足以解释性能变化,也不代表智能体的执行过程可靠。团队应根据预期行为建立可重复的基线。 后续观察:需要关注 Google 是否披露更完整的行为评估案例、断言设计方式及其与现有基准的结合效果,并观察该方法能否稳定识别模型或提示词变化带来的回归。