{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-10-06T08:00:48.640Z","headline":"OpenRouter 教程：提示词或模型变更后如何对 AI Agent 做回归测试","description":"OpenRouter 发布 AI Agent 回归测试教程：每次提示词、模型、工具定义或检索设置变更后，重跑锁定的用例集并对照书面行为契约检查。","url":"https://www.aioga.com/news/lq71il0lehssgkular7zs5ybl/","mainEntityOfPage":"https://www.aioga.com/news/lq71il0lehssgkular7zs5ybl/","datePublished":"2026-09-30T00:00:00.000Z","dateModified":"2026-09-30T00:00:00.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://openrouter.ai/blog/tutorials/ai-agent-regression-testing-after-a-prompt-or-model-change","https://aihot.news/items/lq71il0lehssgkular7zs5ybl"],"canonicalUrl":"https://www.aioga.com/news/lq71il0lehssgkular7zs5ybl/","directAnswer":{"@type":"Answer","text":"OpenRouter 发布 AI Agent 回归测试教程，建议在提示词、模型、工具定义或检索设置变更后，重新运行锁定用例集，并依据书面行为契约检查工具调用、参数、策略遵守情况及信息补充请求。","url":"https://www.aioga.com/news/lq71il0lehssgkular7zs5ybl/","dateCreated":"2026-09-30T00:00:00.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"OpenRouter：Announcements source article","url":"https://openrouter.ai/blog/tutorials/ai-agent-regression-testing-after-a-prompt-or-model-change","datePublished":"2026-09-30T00:00:00.000Z","provider":{"@type":"Organization","name":"OpenRouter：Announcements","url":"https://openrouter.ai/blog/tutorials/ai-agent-regression-testing-after-a-prompt-or-model-change"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.news/items/lq71il0lehssgkular7zs5ybl","datePublished":"2026-09-30T00:00:00.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.news/items/lq71il0lehssgkular7zs5ybl"}}],"aggregationSource":"OpenRouter：Announcements","originalPublisher":{"name":"OpenRouter：Announcements","url":"https://openrouter.ai/blog/tutorials/ai-agent-regression-testing-after-a-prompt-or-model-change"},"geoDeepAnswer":null,"article":{"id":"lq71il0lehssgkular7zs5ybl","slug":"lq71il0lehssgkular7zs5ybl","url":"https://www.aioga.com/news/lq71il0lehssgkular7zs5ybl/","title":"OpenRouter 教程：提示词或模型变更后如何对 AI Agent 做回归测试","title_en":"","summary":"OpenRouter 发布 AI Agent 回归测试教程：每次提示词、模型、工具定义或检索设置变更后，重跑锁定的用例集并对照书面行为契约检查。","source":"OpenRouter：Announcements","sourceUrl":"https://openrouter.ai/blog/tutorials/ai-agent-regression-testing-after-a-prompt-or-model-change","aiHotUrl":"https://aihot.news/items/lq71il0lehssgkular7zs5ybl","publishedAt":"2026-09-30T00:00:00.000Z","category":"行业动态","score":72,"selected":true,"articleBody":["A ~author/family-latest alias always resolves to the newest concrete model in a family. That’s convenient in production and a problem in a regression test, because the model can change between runs without any change in your repository. Our latest model resolution：https://openrouter.ai/docs/guides/routing/routers/latest-resolution docs describe the mechanism and recommend a concrete model slug when you need a fixed version for reproducibility. This guide covers the locked case set and the behavioral contract, then the model-swap case in detail.","Code regression testing rests on a known input, a known correct output, and a diff that tells you when the output changed. Three properties of an agent break that.","Two correct answers rarely look alike. A text diff against a golden answer fails on behavior that was never wrong. What holds still is structural. You check whether the agent called the right tool with the right arguments, respected the policy, and asked for the piece of information it was missing.","The model is a moving part. A model selected through a ~author/family-latest alias can change without a commit in your repository, and the part that changed is the one doing most of the reasoning. The model field in every OpenRouter response reports the concrete model that served the request. Reading it back is the cheapest way to notice that the model answering your calls is no longer the model you tested.","A pass expires when the baseline moves. Comparing against the same fixed set of cases every time is what turns “it seems fine” into a claim you can defend.","Agents drift on changes that a traditional test suite has no reason to look at. We group them into three kinds.","The third row is the one that is easiest to miss. A new chunking strategy for retrieved documents, an added field in a tool response, or a longer history can push content the agent relied on out of what it sees, and none of it touches the prompt. What you see is rarely an error. A support agent that used to quote the refund policy accurately starts paraphrasing it from memory, because the paragraph it relied on now falls outside the retrieved chunk, and the transcript reads just as fluently either way. A prompt edit has the same property. Tightening one sentence to fix one complaint can change which tool fires on an unrelated case.","Everything downstream depends on the case set, so build it before you think about automation.","Include representative cases that cover the requests your agent handles most often, a few edge cases such as ambiguous input or a request that sits on a policy boundary, and at least one case built to test a rule you never want broken. For a support agent that means a routine refund, a request with no order ID, and a refund above whatever limit your policy sets.","Once the set exists, stop editing it casually. Adding, removing, or rewording a case breaks comparability with every past run, and you lose the ability to tell a real regression from a different test. Every edit turns the set into a new experiment, so treat changes with the care you would give a schema migration.","For each case, write two things. The structural assertion says what the agent should do, such as calling lookup_order before acting and leaving escalate_to_human alone on a routine refund. The hard invariant says what the agent must never do, such as approving a refund above $500 without a human. That $500 is an example application policy rather than anything OpenRouter sets. The number in your own contract comes from your business rules. Most cases only need the structural assertion. The hard invariant is the one you want as an automatic ship-blocker, with no threshold and no judgment call attached.","Here is one case expressed as a plain API call, pinned to a concrete model, printing back both the model that served it and the tools it chose. The request sets no max_tokens , because a truncated response can cut off the tool call’s JSON and report a failure that has nothing to do with the agent’s decision.","The same case in Python with the OpenAI SDK pointed at our base URL.","The same case in TypeScript with fetch .","The mechanics are simple once the cases and contracts exist. A few details decide whether the run catches anything.","Trigger the run on the change. Re-run the full case set whenever a prompt, model, tool definition, or retrieval setting changes. A suite that runs only when someone remembers to run it will eventually miss the change that mattered.","Score the delta as well as the pass. A case that a judge scored well last month and scores lower today hasn’t failed, and it’s still a regression worth opening. Treat a meaningful score drop the way you would treat a failing test. Check that your rubric can move before you rely on it, because a rubric that scores every answer alike reports a clean pass while measuring nothing.","Match the check to the case. Deterministic cases, where you can name the exact tool and argument you expect, get exact or structural checks. Open-ended cases, such as whether an explanation is accurate and correctly scoped, need a judge model, because no single correct string exists to match against.","Ori Eval covers both shapes in one file. Assertions such as run.tool('lookup_order').toBeCalled() , run.toComplete() , run.toCostAtMost(0.01) , and run.toFinishWithin(30_000) handle the structural side. setupJudge({ minScore: 0.8 }) scores open-ended cases from 0 to 1 against criteria you write. Ori also resolves one harness and one model per run and holds them for every test in that run, so two runs of the same eval files use the same configuration.","Switching models on OpenRouter is a configuration change rather than a rewrite. That only helps if you can show that behavior stayed put when you made the switch.","Price is usually what starts the conversation. Two models we serve today sit at opposite ends of the price range, and both list tools in their supported parameters.","Checked September 18, 2026, against the live Claude Fable 5.1：https://openrouter.ai/anthropic/claude-fable-5.1 and Gemini 3.8 Flash：https://openrouter.ai/google/gemini-3.8-flash endpoint data. Gemini 3.8 Flash prices are for the standard tier. Both Google providers also serve flex and priority tiers at different prices. Prices change, so recheck before you plan around a ratio.","A thirteenfold difference in input price is reason enough to try the swap. The run is what earns the right to ship it. The mechanics are the locked case set and the contract you already have, with one variable moved. Pin the prompt, the tool definitions, the tool results, the case set, the judge, and the inference parameters, then run the suite against the candidate before any real traffic reaches it. When a result moves, you know the model moved it.","Pinning includes the slug itself. An alias like ~anthropic/claude-fable-latest routes to the newest concrete model in that family and updates whenever the author publishes a new version. Name the exact version on both sides of the comparison, and read the response’s model field to confirm what served each call.","Pinning also includes the inference parameters, and the two models don’t accept the same ones. Each entry in the models endpoint：https://openrouter.ai/docs/api/api-reference/models/list-all-models-and-their-properties has a supported_parameters array. Gemini 3.8 Flash lists temperature . Claude Fable 5.1 doesn’t, so with default routing a temperature value sent to it is ignored by the provider rather than applied, and setting it on one side of the comparison doesn’t hold the other side still. With require_parameters set, a parameter that no endpoint of the model supports means the request isn’t routed at all, so leave temperature out of the Claude side. Both models list reasoning and accept low , medium , and high as efforts, while their default efforts differ. The harness below sets reasoning.effort to medium for both and sets provider.require_parameters to true , so we only route each request to a provider endpoint that supports every parameter in it. See provider routing：https://openrouter.ai/docs/guides/routing/provider-selection#requiring-providers-to-support-all-parameters for the field. If every model in your comparison lists temperature , set it explicitly as well.","Here is the diff in its smallest useful form. It runs the same three cases against both slugs and separates a structural miss from a broken policy. The agent’s first move on a refund is a lookup, so the harness runs a short tool loop with fixed order records rather than reading a single response. The order data is pinned along with everything else, so a tool result can’t vary between runs.","The loop passes each assistant turn back with its reasoning_details unchanged, which our reasoning tokens：https://openrouter.ai/docs/guides/best-practices/reasoning-tokens#preserving-reasoning docs describe for tool calling with reasoning models. It stops after the first decision tool or after four turns, whichever comes first, and records every tool the model called in order.","The same comparison in Ori Eval loops over the two slugs in a single *.eval.ts file. Ori runs the tool loop for you and exposes the calls through run.tool() . The judge is pinned to a concrete model too. setupJudge() without an agent option grades on a default model, and we pass an explicit agent so the grading model is part of the pinned configuration.","Four ori eval flags map onto the pinning this section describes.","--baseline chooses what the run’s report is compared against. It takes last , best , or model: . The last form is the model swap as a single argument, comparing this run against a stored run of another model. The comparison is reporting only and doesn’t change the exit code. It reads run history from .ori/eval/history.jsonl , so it needs an Ori workspace, and a comparison is only possible between runs that included exactly the same eval files. --no-history keeps a run out of that file.","--hermetic gives the agent a fresh temporary workspace instead of your project directory, which keeps your ori.md , AGENTS.md , CLAUDE.md , and skill directories out of the run. Those files are context the agent reads, which makes them a variable you can change without noticing that you changed the agent. A CLAUDE.md or AGENTS.md at the repository root is checked in, edited often, and read on every run. Before it runs anything, ori eval lists the agent’s working directory and the instruction files and skill directories it discovered on stderr, so the run tells you what it had in front of it.","--dry-run loads every discovered eval and runs no tests, so a parse error or an unresolved import fails before any model call rather than after one. It needs no credentials, and it confirms nothing about whether an eval passes.","--pilot runs a strided sample of n cases from each eval file that wraps its cases in pilotCases() and reports the measured and estimated cost. It’s a cost measurement, not a comparison.","We ran this suite against both slugs three times on September 7, 2026, with openai/gpt-6-astra grading the open-ended case, and the swap held. Every structural case passed on both sides, the tool sequences came back identical, and no hard invariant moved.","Three runs per model on September 7, 2026, with openai/gpt-6-astra grading the open-ended case. The judge in that run reported on a 0 to 10 scale and returned 10.0 every time. Ori Eval’s setupJudge() reports from 0 to 1, and minScore is set on that scale. Eighteen structural case runs, all passing, with the same tool sequence every time. Those runs set no reasoning effort, so each model ran at its default. Model behavior changes, so read this as one dated measurement rather than a standing claim about either model.","That is the outcome you want from a swap. The run before it taught us more than these three did.","The first run failed on both sides, and our harness caused it. Both models failed the policy case with the same note. Neither had called issue_refund , so no invariant tripped.","The first version of the harness sent one request and read one response. An agent’s first move on that case is a lookup, so it never reached the decision the contract was written about, and the contract failed a decision the agent never had a chance to make. The tool loop in the harness above is the fix. A case that fails on both sides is a broken test, and the candidate column means nothing until that is fixed.","The judge returned the maximum score for every answer, from both models, across all three runs. A rubric that nothing can fail has no resolving power, and it will sit in your suite looking like coverage while detecting nothing. Ours asked whether the answer states the $500 limit and avoids inventing policy, which both models clear. A judge case earns its place only when a worse answer would score lower, so calibrate it by feeding it an answer you know is bad and confirming that the score moves.","Model routing and fallbacks change which provider or model serves a request. They don’t test anything. The eval run is what makes an easy switch safe to act on.","Treating every dip as a release blocker trains a team to ignore the gate, so the last piece of the pattern is deciding what deserves attention.","Judges drift, and they carry biases, including a preference for longer answers over shorter ones that are better. We haven’t measured judge agreement rates ourselves, so treat a single judge score as one signal. One failing case out of twenty isn’t automatically a reason to block a release, because it may be judge noise on a borderline case. Set a threshold for how many cases can fail before you treat it as a real signal, and re-run anything flaky before you trust a single failure.","Hard invariants are the exception, and they shouldn’t have a threshold. A refund above the limit or a skipped escalation goes to human review every time, no matter how many other cases passed. Stylistic drift is a judgment call. A broken policy boundary is a defect.","AI agent regression testing means re-running a locked, versioned set of test cases every time a prompt, model, tool definition, or retrieval setting changes, then checking that the agent’s behavior still matches a written contract. Because two correct answers from an agent rarely use the same words, the check is structural. It covers which tools the agent called with which arguments and whether policy boundaries held, rather than a text diff against a golden output.","Re-run your existing locked case set against the edited prompt with everything else held still, then compare each result to that case’s contract. Keep the case set unchanged so the comparison is valid. Pin the model to a concrete slug and set inference parameters explicitly so the prompt is the only variable. Apply structural assertions for deterministic cases and a judge score for open-ended ones. Treat a broken hard invariant as a release blocker and a score drop as something to investigate. Wiring the suite to a job scoped to your prompt files means the run happens on the change rather than when someone remembers.","Any test runner that can call your agent, assert on the tools it called, and score open-ended answers will work. The framework matters less than the locked case set and the written contract behind it. You can drive the agent from a general-purpose test runner such as pytest, Jest, or bun test with your own assertions, use an evaluation product that stores runs and diffs them for you, or use Ori Eval：https://openrouter.ai/docs/guides/ori/eval, which provides tool-call assertions, an LLM judge, and a pinned harness in *.eval.ts files.","A unit test diffs against one known-correct output. Agent regression testing checks structural assertions and hard invariants, because the output text varies between runs. The scope also differs. A unit test assumes the runtime under your code is stable, while an agent’s model can change when a provider ships a new version behind an alias or when you swap models yourself.","A behavioral contract is a per-case definition of what correct means. It has a structural assertion about which tool the agent called and what arguments it received, plus, where a policy exists, a hard invariant the agent must never violate, such as a refund limit or an escalation rule. Writing both down separates the cases where a judgment call is acceptable from the rule that should stop a release.","Run them on every change that touches a prompt, model, tool definition, or retrieval setting, triggered by the paths those files live in. Keep them out of the unit-test job that fires on every commit, because eval runs send requests to real models and cost money. A scheduled run catches provider-side changes that arrive without a commit of your own.","A model can grade its own output, but an independent judge on a different model stops one model’s blind spots from shaping both the answer and the score. Ori Eval’s setupJudge() creates a separate agent on its own grading model for that reason, and the measured run in this guide used a third model to grade both candidates. Treat the judge score as one signal, watch for known biases such as favoring longer answers, and keep hard invariants in deterministic structural checks where no judge is involved.","If you adopt one thing from this guide, adopt the rule that no model swap ships without a documented pass. That means pinning a concrete slug in the tests rather than an alias, giving the suite its own job keyed to the paths that hold your prompts, models, tools, and retrieval settings, and scheduling a separate run to catch the provider-side changes that arrive without a commit of your own. Together, those three turn a swap from a judgment call into a decision with a record behind it.","The rest follows from the case set. Ours found nothing wrong with the candidate model and two things wrong with our own harness, which is a good reason to start before you think you need to. Browse the model catalog：https://openrouter.ai/models to pick a candidate, and read the Ori Eval guide：https://openrouter.ai/docs/guides/ori/eval for the harness that runs your cases against it.","Model usage data, product updates, and research reports. One email each week."],"articleImages":[{"sourceUrl":"https://openrouter.ai/blog/images/ai-agent-regression-testing-after-a-prompt-or-model-change.png","alt":"AI Agent Regression Testing After a Prompt or Model Change","afterParagraph":0,"url":"/media/articles/lq71il0lehssgkular7zs5ybl/d262acf3670a408c.png"},{"sourceUrl":"https://openrouter.ai/blog/images/agent-regression-testing-pattern.png","alt":"Diagram of the regression-testing pattern: a locked case set, one change to a prompt line, model slug, or tool or retrieval setting, a pinned re-run with the same cases and harness, and two outcomes, behavior held or behavior broke on an un-escalated refund","afterParagraph":0,"url":"/media/articles/lq71il0lehssgkular7zs5ybl/26d9b36132a43c3f.png"}],"mediaStatus":"ok","articleBodyZh":["A ~author/family-latest 别名总是解析为该系列中最新的具体模型。在生产环境中这很方便，但在回归测试中可能成为问题，因为模型可能在运行之间发生变化，而您的仓库没有任何更改。我们的最新模型解析：https://openrouter.ai/docs/guides/routing/routers/latest-resolution 文档描述了该机制，并建议在需要可重复性的固定版本时使用具体的模型标识。本指南涵盖了锁定情况集和行为契约，然后详细介绍了模型交换的情况。","代码回归测试依赖于已知输入、已知正确输出以及在输出发生变化时可以告诉你的差异。而代理的三个特性会破坏这一点。","两个正确答案很少完全相同。针对黄金答案的文本差异检查会在从未出错的行为上失败。保持不变的是结构性部分。你需要检查代理是否调用了正确的工具并传入正确参数，是否遵守了策略，以及是否请求了它缺失的信息。","模型是一个动态部分。通过 ~author/family-latest 别名选择的模型可以在仓库没有提交的情况下发生变化，而发生变化的部分正是承担大部分推理的部分。每个 OpenRouter 响应中的 model 字段报告了提供请求服务的具体模型。读取该字段是发现回答你请求的模型不再是你测试过的模型的最经济的方法。","当基线移动时，测试通过就会过期。每次都与同一固定案例集进行比较，才能将“看起来没问题”转化为你可以维护的声明。","代理会在传统测试套件无理由关注的变动上产生漂移。我们将它们分为三类。","第三行是最容易被忽略的。一种用于检索文档的新分块策略、工具响应中新添加的字段，或者更长的历史记录，都可能将代理依赖的内容推到其无法看到的范围之外，而这些都不会触及提示内容。你所看到的通常并不是错误。一个曾经能够准确引用退款政策的客服代理，可能会开始根据记忆进行解读，因为它依赖的那段文字现在不在检索到的分块中，无论怎样，成绩记录看起来仍然很流畅。提示编辑也有同样的特性。为了修复一个投诉而紧缩某一句话，可能会改变对于不相关案例触发的工具。","所有下游操作都依赖于案例集合，所以在考虑自动化之前先建立它。","案例集合应包括能代表你代理处理最频繁请求的常规案例，一些边缘情况，例如输入模糊或请求位于政策边界上的案例，以及至少一个用于测试你绝不希望被破坏的规则的案例。对于客服代理来说，这意味着一个常规退款案例、一个没有订单号的请求，以及一个超过政策设定限额的退款请求。","一旦集合存在，就不要随意编辑它。添加、删除或改写一个案例都会破坏与以往每次运行的可比性，并且你将丧失识别实际回归与不同测试的能力。每一次编辑都会把集合变成一个新的实验，因此对待这些变动的态度应像对待数据库模式迁移一样谨慎。","对于每个案例，写两件事。结构性断言说明代理应该做什么，例如在处理前调用查找订单函数，对于常规退款则保持不升级人为处理。硬性不变条件则说明代理绝不能做什么，例如未经人工批准就批准超过500美元的退款。这个500美元只是应用策略示例，而非OpenRouter设定。你自己合同中的数字来源于你的业务规则。大多数案例只需要结构性断言。硬性不变条件是你希望自动阻止发货的条件，不附带阈值，也不需判断。","这里是一个以简单 API 调用表示的案例，固定到具体模型，打印返回提供服务的模型以及它选择的工具。请求没有设置 max_tokens，因为截断的响应可能会切断工具调用的 JSON 并报告一个与代理决策无关的失败。","使用 OpenAI SDK 指向我们的基础 URL 的 Python 版本相同的案例。","在 TypeScript 中使用 fetch 的相同情况。","一旦案例和合同存在，机制就很简单。一些细节决定运行时是否捕捉到任何内容。","在变更上触发运行。每当提示、模型、工具定义或检索设置发生变化时，重新运行完整案例集。仅在有人记得运行它时才运行的测试套件最终会错过重要的变化。","对增量和通过情况进行评分。一个上个月评审判定得分很高的案例，如果今天评分较低，并不意味着它失败了，但仍然值得打开作为回归问题。像对待失败测试一样处理有意义的分数下降。在依赖评分标准之前，检查你的评分标准是否会变化，因为对所有回答评分相同的标准会报告通过，但实际上没有测量到任何内容。","将检查与案例匹配。确定性案例，您可以命名期望的具体工具和参数，使用精确或结构化检查。开放性案例，如解释是否准确和范围是否正确，需要使用评审模型，因为不存在可匹配的单一正确字符串。","Ori Eval 在一个文件中覆盖这两种情况。断言如 run.tool('lookup_order').toBeCalled()、run.toComplete()、run.toCostAtMost(0.01) 和 run.toFinishWithin(30_000) 处理结构化部分。setupJudge({ minScore: 0.8 }) 根据你编写的标准从 0 到 1 对开放性案例进行评分。Ori 还在每次运行中解析一个测试工具和一个模型，并在该运行的每个测试中保持它们，因此同一评测文件的两次运行使用相同配置。","在 OpenRouter 上切换模型是配置更改而不是重写。只有当你能够证明行为在切换时保持不变，这才有帮助。","价格通常是开始讨论的起点。我们今天提供的两个模型价格处于价格范围的两端，并且都在其支持的参数中列出了工具。","截至2026年9月18日，根据实时Claude Fable 5.1（https://openrouter.ai/anthropic/claude-fable-5.1）和Gemini 3.8 Flash（https://openrouter.ai/google/gemini-3.8-flash）端点数据进行了检查。Gemini 3.8 Flash的价格适用于标准层。两家Google提供商也提供灵活层和优先层，价格不同。价格会变动，因此在以比例计划之前请重新确认。","输入价格相差十三倍，这已经足够作为尝试交换的理由。运行才是获得发布权的条件。机制是锁定的案例集和你已有的合同，只有一个变量被移动。固定提示、工具定义、工具结果、案例集、评判者和推理参数，然后在任何真实流量到达之前，将测试套件运行在候选模型上。当结果发生变化时，你就知道模型发生了变化。","固定还包括模型自身。像 ~anthropic/claude-fable-latest 这样的别名会路由到该系列中最新的具体模型，并在作者发布新版本时更新。在比较的两边指定确切版本，并读取响应的 model 字段以确认每次调用所使用的模型。","固定还包括推理参数，并且两个模型不接受相同的参数。Models端点的每个条目（https://openrouter.ai/docs/api/api-reference/models/list-all-models-and-their-properties）都有一个 supported_parameters 数组。Gemini 3.8 Flash列出 temperature，而Claude Fable 5.1没有，因此默认路由时发送给Claude的 temperature 值会被提供商忽略而不是应用，而且在比较的一边设置它并不能让另一边保持不变。当设置 require_parameters 为 true 时，如果模型的任何端点不支持某个参数，请求将根本不会被路由，因此Claude方请不要使用temperature。两个模型都列出 reasoning，并接受 low、medium 和 high 作为努力等级，而它们的默认努力等级不同。下述测试架构将 reasoning.effort 对两者都设置为 medium，并将 provider.require_parameters 设置为 true，因此我们只会将每个请求路由到支持其所有参数的提供商端点。有关提供商路由，请参阅：https://openrouter.ai/docs/guides/routing/provider-selection#requiring-providers-to-support-all-parameters。如果比较中的每个模型都列出 temperature，则也要显式设置它。","这是差异的最小可用形式。它对同三个案例同时运行在两个 slugs 上，并且将结构性错误与策略破损分开。代理在退款上的第一步是查找，因此测试工具运行一个短的工具循环，使用固定的订单记录，而不是读取单个响应。订单数据与其他所有内容一起被固定，因此工具的结果在不同运行之间不会变化。","循环将每次助手的回合返回，并且 reasoning_details 保持不变，我们的 reasoning tokens：https://openrouter.ai/docs/guides/best-practices/reasoning-tokens#preserving-reasoning 文档描述了用于带有推理模型的工具调用。循环在第一次决策工具调用后或者运行四次之后停止，以先发生者为准，并按顺序记录模型调用的每个工具。","在 Ori Eval 中的相同比较在单个 *.eval.ts 文件中对两个 slugs 循环执行。Ori 会为你运行工具循环，并通过 run.tool() 暴露调用。评判器也固定到具体模型。setupJudge() 如果没有代理选项，就在默认模型上评分，而我们传递一个明确代理，那么评分模型就是固定配置的一部分。","四个 ori eval 标志映射到这一节描述的固定设置。","--baseline 选择该运行的报告与哪个进行比较。它可以取 last、best 或 model:。最后一种形式是模型交换作为单一参数，将此次运行与另一个模型的已存运行进行比较。比较仅用于报告，不改变退出代码。它从 .ori/eval/history.jsonl 读取运行历史，因此需要 Ori 工作区，并且比较仅可以在包含完全相同 eval 文件的运行之间进行。--no-history 会让一次运行不进入该文件。","--hermetic 给代理一个新的临时工作区，而不是你的项目目录，这样就可以将 ori.md、AGENTS.md、CLAUDE.md 以及技能目录排除在运行之外。这些文件是代理读取的上下文，会成为一个变量，你可以在不注意的情况下修改代理。存储库根目录的 CLAUDE.md 或 AGENTS.md 会被检入，经常编辑，并在每次运行时读取。运行任何操作之前，ori eval 会在 stderr 列出代理的工作目录以及它发现的指令文件和技能目录，因此运行时会显示它面前的内容。","--dry-run 会加载每个发现的评估并不运行测试，因此解析错误或未解决的导入会在任何模型调用之前失败，而不是之后。它不需要凭证，也不会确认评估是否通过。","--pilot 会从每个评估文件中抽取 n 个案例的间隔样本，将其案例包装在 pilotCases() 中，并报告测量和估算的成本。这是成本测量，不是比较。","我们在 2026 年 9 月 7 日对两个模型各运行了三次该测试套件，由 openai/gpt-6-astra 对开放式案例进行评分，交换结果保持不变。每个结构案例在双方都通过，工具序列完全相同，没有任何硬性不变量变化。","每个模型在 2026 年 9 月 7 日运行三次，由 openai/gpt-6-astra 对开放式案例评分。该次运行中的评判以 0 到 10 的量表报告，每次返回 10.0。Ori Eval 的 setupJudge() 报告范围为 0 到 1，并在该量表上设置 minScore。十八次结构案例运行，全部通过，每次工具序列相同。这些运行没有设置推理努力，因此每个模型按默认运行。模型行为会变化，因此应将其视为一次特定日期的测量，而非对任一模型的长期声明。","这是你希望从交换中得到的结果。交换前的一次运行给我们的信息比这三次更多。","第一次运行双方都失败，这是由我们的测试工具引起的。两个模型在策略案例上都失败，并显示相同的备注。两个模型都没有调用 issue_refund，因此没有触发不变量。","测试工具的第一个版本发送了一次请求并读取一次响应。代理在该案例的第一次行动是一次查找，因此从未到达合同所规定的决策，合同在代理根本没有机会做出决定时失败了。上述测试工具中的工具循环是修复方法。双方都失败的案例是一个错误的测试，候选列在问题修复前没有意义。","法官对所有答案在三次运行中均给出了最高分，无论来源于哪种模型。一套没有任何失败的评分标准是没有判别力的，它会像覆盖率一样存在于你的评估套件中，但实际上侦测不到问题。我们的评分会问答案是否说明了500美元的上限并且避免虚构政策，这两点两种模型都满足。一个评分案例只有在较差的答案得分会更低时才有意义，因此要通过提供你知道不好的答案并确认分数变化来调整它。","模型路由和回退会改变哪家提供商或模型处理请求。它们并不测试任何东西。评估运行才是使轻松切换安全可行的关键。","将每一次下降都视为发布阻滞会训练团队忽视关卡，因此模式的最后一环是决定什么值得注意。","评分者会漂移，并且带有偏见，包括对较长答案的偏好，而这些答案可能并不比短答案更好。我们自己没有测量评分者的一致率，所以将单个评分视为一种信号。二十个案例中出现一次失败并不自动意味着必须阻止发布，因为这可能是评分者在边界案例上的噪声。在你将单个失败视为真实信号之前，设定一个阈值来确定多少案例失败后才算，并对任何不稳定的结果重新运行。","硬性不变项是例外，它们不应有阈值。每次超过限制的退款或跳过的升级都需要人工审核，无论其他案例通过多少。风格偏移属于判断性问题。破损的政策边界属于缺陷。","AI代理回归测试意味着每次提示、模型、工具定义或检索设置变化时，都要重新运行一组锁定、版本化的测试案例，并检查代理的行为是否仍然符合书面契约。因为代理的两个正确答案很少使用相同的词语，所以检查是结构性的。它涵盖了代理调用了哪些工具以及使用了哪些参数，以及政策边界是否保持，而不是针对标准输出的文本差异对比。","在其他条件保持静止的情况下，重新运行你已有的锁定案例集，对照已编辑的提示词，然后将每个结果与该案例的契约进行比较。保持案例设置不变，以确保比较有效。将模型钉在具体的条形上，并明确设置推理参数，使提示词成为唯一变量。对于确定性案例应用结构性断言，对开放式案例应用法官评分。将破损的硬不变量视为释放阻碍，将分数下降视为需要调查的因素。将套件连接到指向提示文件的作业，意味着运行发生在变更时，而不是在有人记起时。","任何能调用你的代理、对调用的工具进行断言并获得开放式答案的测试运行者都可以工作。框架本身不如锁定的案例集和背后的书面合同重要。你可以用自己的断言从通用测试运行工具如pytest、jest或bun test驱动代理，使用存储运行并进行差分的评估产品，或者使用Ori Eval：https：//openrouter.ai/docs/guides/ori/eval，它提供工具调用断言、LLM裁判和钉顶的线束，存在*.eval.ts文件中。","单元测试会与一个已知正确的输出进行差异。代理回归测试检查结构断言和硬不变量，因为输出文本在不同运行中会变化。范围也不同。单元测试假设代码下的运行时间稳定，而代理模型可能会在供应商通过别名发布新版本或你自己交换模型时发生变化。","行为合同是对正确含义的逐案定义。它有关于代理人调用了哪个工具及其收到的论据的结构性断言，此外，如果存在政策，则存在一个代理人绝不能违反的硬不变量，如退款限额或升级规则。将两者写出来区分了判断可接受的情况和应当停止释放的规则。","在每次涉及提示、模型、工具定义或检索设置的更改时运行这些文件，这些更改由这些文件所在的路径触发。不要让它们出现在每次提交时触发的单元测试作业中，因为评估运行会向真实模型发送请求，且需要花费成本。定时运行则捕捉提供者端的变更，这些更改没有提交你自己提交。","一个模型可以评价自己的输出，但一个在不同模型上的独立评审可以阻止单个模型的盲点影响答案和评分。因此，Ori Eval 的 setupJudge() 会为其自身的评分模型创建一个独立代理，本指南中测量运行使用了第三个模型来对两个候选模型进行评分。将评分视为一个参考信号，注意已知偏差，例如偏好较长的答案，并在没有评审参与的确定性结构检查中保持严格不变性。","如果你从本指南中只采纳一条原则，那就采纳这条：没有文档化通过，就不要进行模型替换。这意味着在测试中固定具体的版本，而不是别名，为测试套件分配自己的任务，任务键为保存你的提示、模型、工具和检索设置的路径，并安排一次单独运行以捕捉供应商端未提交的更改。结合这三点，可以将模型替换从一种判断变为有记录支撑的决策。","其余内容取决于案例集。我们的案例发现候选模型没有问题，而我们自己的测试平台存在两个问题，这也说明了在觉得自己需要之前开始测试是明智的。浏览模型目录：https://openrouter.ai/models 以选择候选模型，并阅读 Ori Eval 指南：https://openrouter.ai/docs/guides/ori/eval 来了解运行测试案例所用的测试平台。","模型使用数据、产品更新和研究报告。每周发送一封邮件。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"OpenRouter 发布 AI Agent 回归测试教程，建议在提示词、模型、工具定义或检索设置变更后，重新运行锁定用例集，并依据书面行为契约检查工具调用、参数、策略遵守情况及信息补充请求。","background":"教程指出，author/family-latest 别名会解析到某个模型系列中最新的具体模型，模型可能在代码库不变时发生变化。OpenRouter 响应中的 model 字段会报告实际服务请求的具体模型，固定版本复现时建议使用具体模型标识。","viewpoint":"Aioga 判断：对于输出存在多种正确表达的 Agent，单纯比较固定文本可能误报。以工具调用、参数、策略和信息请求等结构化行为作为检查对象，可能更适合识别实际行为变化，但仍依赖用例集和行为契约的明确性。","implications":"可能影响：模型别名、提示词、工具定义或检索设置的变化，都可能改变 Agent 行为；这不代表每次变化都会造成错误。团队需要维护固定用例和书面契约，并记录响应中的具体模型，以提高回归结果的可解释性。","nextStep":"后续观察：应关注锁定用例集是否覆盖工具调用、策略遵守和信息缺口处理，并在每次相关配置变化后对照行为契约复测。同时需要检查实际响应中的 model 字段，确认测试所用模型是否发生变化。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-10-01T17:26:22.788Z","sourceHash":"e2df7f0d0a4bcdef","review":{"approved":true,"groundedness":94,"clarity":93,"duplicationRisk":12,"blockingIssues":[],"notes":["“提高回归结果的可解释性”属于基于来源内容的合理推论，来源明确支持记录具体模型以发现模型变化，但未直接使用“可解释性”表述。","可将“某个模型系列中最新的具体模型”进一步表述为“该系列当前最新的具体模型”，以更贴近来源原意。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","editorial-labels","inference-boundary","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","OpenRouter：Announcements"],"translations":{"zh-CN":{"title":"OpenRouter 教程：提示词或模型变更后如何对 AI Agent 做回归测试","summary":"OpenRouter 发布 AI Agent 回归测试教程：每次提示词、模型、工具定义或检索设置变更后，重跑锁定的用例集并对照书面行为契约检查。","category":"行业动态","source":"openrouter.ai","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"OpenRouter 教程：提示词或模型变更后如何对 AI Agent 做回归测试 - Aioga AI资讯","description":"OpenRouter 发布 AI Agent 回归测试教程：每次提示词、模型、工具定义或检索设置变更后，重跑锁定的用例集并对照书面行为契约检查。","url":"https://www.aioga.com/news/lq71il0lehssgkular7zs5ybl/","articleBody":["A ~author/family-latest 别名总是解析为该系列中最新的具体模型。在生产环境中这很方便，但在回归测试中可能成为问题，因为模型可能在运行之间发生变化，而您的仓库没有任何更改。我们的最新模型解析：https://openrouter.ai/docs/guides/routing/routers/latest-resolution 文档描述了该机制，并建议在需要可重复性的固定版本时使用具体的模型标识。本指南涵盖了锁定情况集和行为契约，然后详细介绍了模型交换的情况。","代码回归测试依赖于已知输入、已知正确输出以及在输出发生变化时可以告诉你的差异。而代理的三个特性会破坏这一点。","两个正确答案很少完全相同。针对黄金答案的文本差异检查会在从未出错的行为上失败。保持不变的是结构性部分。你需要检查代理是否调用了正确的工具并传入正确参数，是否遵守了策略，以及是否请求了它缺失的信息。","模型是一个动态部分。通过 ~author/family-latest 别名选择的模型可以在仓库没有提交的情况下发生变化，而发生变化的部分正是承担大部分推理的部分。每个 OpenRouter 响应中的 model 字段报告了提供请求服务的具体模型。读取该字段是发现回答你请求的模型不再是你测试过的模型的最经济的方法。","当基线移动时，测试通过就会过期。每次都与同一固定案例集进行比较，才能将“看起来没问题”转化为你可以维护的声明。","代理会在传统测试套件无理由关注的变动上产生漂移。我们将它们分为三类。","第三行是最容易被忽略的。一种用于检索文档的新分块策略、工具响应中新添加的字段，或者更长的历史记录，都可能将代理依赖的内容推到其无法看到的范围之外，而这些都不会触及提示内容。你所看到的通常并不是错误。一个曾经能够准确引用退款政策的客服代理，可能会开始根据记忆进行解读，因为它依赖的那段文字现在不在检索到的分块中，无论怎样，成绩记录看起来仍然很流畅。提示编辑也有同样的特性。为了修复一个投诉而紧缩某一句话，可能会改变对于不相关案例触发的工具。","所有下游操作都依赖于案例集合，所以在考虑自动化之前先建立它。","案例集合应包括能代表你代理处理最频繁请求的常规案例，一些边缘情况，例如输入模糊或请求位于政策边界上的案例，以及至少一个用于测试你绝不希望被破坏的规则的案例。对于客服代理来说，这意味着一个常规退款案例、一个没有订单号的请求，以及一个超过政策设定限额的退款请求。","一旦集合存在，就不要随意编辑它。添加、删除或改写一个案例都会破坏与以往每次运行的可比性，并且你将丧失识别实际回归与不同测试的能力。每一次编辑都会把集合变成一个新的实验，因此对待这些变动的态度应像对待数据库模式迁移一样谨慎。","对于每个案例，写两件事。结构性断言说明代理应该做什么，例如在处理前调用查找订单函数，对于常规退款则保持不升级人为处理。硬性不变条件则说明代理绝不能做什么，例如未经人工批准就批准超过500美元的退款。这个500美元只是应用策略示例，而非OpenRouter设定。你自己合同中的数字来源于你的业务规则。大多数案例只需要结构性断言。硬性不变条件是你希望自动阻止发货的条件，不附带阈值，也不需判断。","这里是一个以简单 API 调用表示的案例，固定到具体模型，打印返回提供服务的模型以及它选择的工具。请求没有设置 max_tokens，因为截断的响应可能会切断工具调用的 JSON 并报告一个与代理决策无关的失败。","使用 OpenAI SDK 指向我们的基础 URL 的 Python 版本相同的案例。","在 TypeScript 中使用 fetch 的相同情况。","一旦案例和合同存在，机制就很简单。一些细节决定运行时是否捕捉到任何内容。","在变更上触发运行。每当提示、模型、工具定义或检索设置发生变化时，重新运行完整案例集。仅在有人记得运行它时才运行的测试套件最终会错过重要的变化。","对增量和通过情况进行评分。一个上个月评审判定得分很高的案例，如果今天评分较低，并不意味着它失败了，但仍然值得打开作为回归问题。像对待失败测试一样处理有意义的分数下降。在依赖评分标准之前，检查你的评分标准是否会变化，因为对所有回答评分相同的标准会报告通过，但实际上没有测量到任何内容。","将检查与案例匹配。确定性案例，您可以命名期望的具体工具和参数，使用精确或结构化检查。开放性案例，如解释是否准确和范围是否正确，需要使用评审模型，因为不存在可匹配的单一正确字符串。","Ori Eval 在一个文件中覆盖这两种情况。断言如 run.tool('lookup_order').toBeCalled()、run.toComplete()、run.toCostAtMost(0.01) 和 run.toFinishWithin(30_000) 处理结构化部分。setupJudge({ minScore: 0.8 }) 根据你编写的标准从 0 到 1 对开放性案例进行评分。Ori 还在每次运行中解析一个测试工具和一个模型，并在该运行的每个测试中保持它们，因此同一评测文件的两次运行使用相同配置。","在 OpenRouter 上切换模型是配置更改而不是重写。只有当你能够证明行为在切换时保持不变，这才有帮助。","价格通常是开始讨论的起点。我们今天提供的两个模型价格处于价格范围的两端，并且都在其支持的参数中列出了工具。","截至2026年9月18日，根据实时Claude Fable 5.1（https://openrouter.ai/anthropic/claude-fable-5.1）和Gemini 3.8 Flash（https://openrouter.ai/google/gemini-3.8-flash）端点数据进行了检查。Gemini 3.8 Flash的价格适用于标准层。两家Google提供商也提供灵活层和优先层，价格不同。价格会变动，因此在以比例计划之前请重新确认。","输入价格相差十三倍，这已经足够作为尝试交换的理由。运行才是获得发布权的条件。机制是锁定的案例集和你已有的合同，只有一个变量被移动。固定提示、工具定义、工具结果、案例集、评判者和推理参数，然后在任何真实流量到达之前，将测试套件运行在候选模型上。当结果发生变化时，你就知道模型发生了变化。","固定还包括模型自身。像 ~anthropic/claude-fable-latest 这样的别名会路由到该系列中最新的具体模型，并在作者发布新版本时更新。在比较的两边指定确切版本，并读取响应的 model 字段以确认每次调用所使用的模型。","固定还包括推理参数，并且两个模型不接受相同的参数。Models端点的每个条目（https://openrouter.ai/docs/api/api-reference/models/list-all-models-and-their-properties）都有一个 supported_parameters 数组。Gemini 3.8 Flash列出 temperature，而Claude Fable 5.1没有，因此默认路由时发送给Claude的 temperature 值会被提供商忽略而不是应用，而且在比较的一边设置它并不能让另一边保持不变。当设置 require_parameters 为 true 时，如果模型的任何端点不支持某个参数，请求将根本不会被路由，因此Claude方请不要使用temperature。两个模型都列出 reasoning，并接受 low、medium 和 high 作为努力等级，而它们的默认努力等级不同。下述测试架构将 reasoning.effort 对两者都设置为 medium，并将 provider.require_parameters 设置为 true，因此我们只会将每个请求路由到支持其所有参数的提供商端点。有关提供商路由，请参阅：https://openrouter.ai/docs/guides/routing/provider-selection#requiring-providers-to-support-all-parameters。如果比较中的每个模型都列出 temperature，则也要显式设置它。","这是差异的最小可用形式。它对同三个案例同时运行在两个 slugs 上，并且将结构性错误与策略破损分开。代理在退款上的第一步是查找，因此测试工具运行一个短的工具循环，使用固定的订单记录，而不是读取单个响应。订单数据与其他所有内容一起被固定，因此工具的结果在不同运行之间不会变化。","循环将每次助手的回合返回，并且 reasoning_details 保持不变，我们的 reasoning tokens：https://openrouter.ai/docs/guides/best-practices/reasoning-tokens#preserving-reasoning 文档描述了用于带有推理模型的工具调用。循环在第一次决策工具调用后或者运行四次之后停止，以先发生者为准，并按顺序记录模型调用的每个工具。","在 Ori Eval 中的相同比较在单个 *.eval.ts 文件中对两个 slugs 循环执行。Ori 会为你运行工具循环，并通过 run.tool() 暴露调用。评判器也固定到具体模型。setupJudge() 如果没有代理选项，就在默认模型上评分，而我们传递一个明确代理，那么评分模型就是固定配置的一部分。","四个 ori eval 标志映射到这一节描述的固定设置。","--baseline 选择该运行的报告与哪个进行比较。它可以取 last、best 或 model:。最后一种形式是模型交换作为单一参数，将此次运行与另一个模型的已存运行进行比较。比较仅用于报告，不改变退出代码。它从 .ori/eval/history.jsonl 读取运行历史，因此需要 Ori 工作区，并且比较仅可以在包含完全相同 eval 文件的运行之间进行。--no-history 会让一次运行不进入该文件。","--hermetic 给代理一个新的临时工作区，而不是你的项目目录，这样就可以将 ori.md、AGENTS.md、CLAUDE.md 以及技能目录排除在运行之外。这些文件是代理读取的上下文，会成为一个变量，你可以在不注意的情况下修改代理。存储库根目录的 CLAUDE.md 或 AGENTS.md 会被检入，经常编辑，并在每次运行时读取。运行任何操作之前，ori eval 会在 stderr 列出代理的工作目录以及它发现的指令文件和技能目录，因此运行时会显示它面前的内容。","--dry-run 会加载每个发现的评估并不运行测试，因此解析错误或未解决的导入会在任何模型调用之前失败，而不是之后。它不需要凭证，也不会确认评估是否通过。","--pilot 会从每个评估文件中抽取 n 个案例的间隔样本，将其案例包装在 pilotCases() 中，并报告测量和估算的成本。这是成本测量，不是比较。","我们在 2026 年 9 月 7 日对两个模型各运行了三次该测试套件，由 openai/gpt-6-astra 对开放式案例进行评分，交换结果保持不变。每个结构案例在双方都通过，工具序列完全相同，没有任何硬性不变量变化。","每个模型在 2026 年 9 月 7 日运行三次，由 openai/gpt-6-astra 对开放式案例评分。该次运行中的评判以 0 到 10 的量表报告，每次返回 10.0。Ori Eval 的 setupJudge() 报告范围为 0 到 1，并在该量表上设置 minScore。十八次结构案例运行，全部通过，每次工具序列相同。这些运行没有设置推理努力，因此每个模型按默认运行。模型行为会变化，因此应将其视为一次特定日期的测量，而非对任一模型的长期声明。","这是你希望从交换中得到的结果。交换前的一次运行给我们的信息比这三次更多。","第一次运行双方都失败，这是由我们的测试工具引起的。两个模型在策略案例上都失败，并显示相同的备注。两个模型都没有调用 issue_refund，因此没有触发不变量。","测试工具的第一个版本发送了一次请求并读取一次响应。代理在该案例的第一次行动是一次查找，因此从未到达合同所规定的决策，合同在代理根本没有机会做出决定时失败了。上述测试工具中的工具循环是修复方法。双方都失败的案例是一个错误的测试，候选列在问题修复前没有意义。","法官对所有答案在三次运行中均给出了最高分，无论来源于哪种模型。一套没有任何失败的评分标准是没有判别力的，它会像覆盖率一样存在于你的评估套件中，但实际上侦测不到问题。我们的评分会问答案是否说明了500美元的上限并且避免虚构政策，这两点两种模型都满足。一个评分案例只有在较差的答案得分会更低时才有意义，因此要通过提供你知道不好的答案并确认分数变化来调整它。","模型路由和回退会改变哪家提供商或模型处理请求。它们并不测试任何东西。评估运行才是使轻松切换安全可行的关键。","将每一次下降都视为发布阻滞会训练团队忽视关卡，因此模式的最后一环是决定什么值得注意。","评分者会漂移，并且带有偏见，包括对较长答案的偏好，而这些答案可能并不比短答案更好。我们自己没有测量评分者的一致率，所以将单个评分视为一种信号。二十个案例中出现一次失败并不自动意味着必须阻止发布，因为这可能是评分者在边界案例上的噪声。在你将单个失败视为真实信号之前，设定一个阈值来确定多少案例失败后才算，并对任何不稳定的结果重新运行。","硬性不变项是例外，它们不应有阈值。每次超过限制的退款或跳过的升级都需要人工审核，无论其他案例通过多少。风格偏移属于判断性问题。破损的政策边界属于缺陷。","AI代理回归测试意味着每次提示、模型、工具定义或检索设置变化时，都要重新运行一组锁定、版本化的测试案例，并检查代理的行为是否仍然符合书面契约。因为代理的两个正确答案很少使用相同的词语，所以检查是结构性的。它涵盖了代理调用了哪些工具以及使用了哪些参数，以及政策边界是否保持，而不是针对标准输出的文本差异对比。","在其他条件保持静止的情况下，重新运行你已有的锁定案例集，对照已编辑的提示词，然后将每个结果与该案例的契约进行比较。保持案例设置不变，以确保比较有效。将模型钉在具体的条形上，并明确设置推理参数，使提示词成为唯一变量。对于确定性案例应用结构性断言，对开放式案例应用法官评分。将破损的硬不变量视为释放阻碍，将分数下降视为需要调查的因素。将套件连接到指向提示文件的作业，意味着运行发生在变更时，而不是在有人记起时。","任何能调用你的代理、对调用的工具进行断言并获得开放式答案的测试运行者都可以工作。框架本身不如锁定的案例集和背后的书面合同重要。你可以用自己的断言从通用测试运行工具如pytest、jest或bun test驱动代理，使用存储运行并进行差分的评估产品，或者使用Ori Eval：https：//openrouter.ai/docs/guides/ori/eval，它提供工具调用断言、LLM裁判和钉顶的线束，存在*.eval.ts文件中。","单元测试会与一个已知正确的输出进行差异。代理回归测试检查结构断言和硬不变量，因为输出文本在不同运行中会变化。范围也不同。单元测试假设代码下的运行时间稳定，而代理模型可能会在供应商通过别名发布新版本或你自己交换模型时发生变化。","行为合同是对正确含义的逐案定义。它有关于代理人调用了哪个工具及其收到的论据的结构性断言，此外，如果存在政策，则存在一个代理人绝不能违反的硬不变量，如退款限额或升级规则。将两者写出来区分了判断可接受的情况和应当停止释放的规则。","在每次涉及提示、模型、工具定义或检索设置的更改时运行这些文件，这些更改由这些文件所在的路径触发。不要让它们出现在每次提交时触发的单元测试作业中，因为评估运行会向真实模型发送请求，且需要花费成本。定时运行则捕捉提供者端的变更，这些更改没有提交你自己提交。","一个模型可以评价自己的输出，但一个在不同模型上的独立评审可以阻止单个模型的盲点影响答案和评分。因此，Ori Eval 的 setupJudge() 会为其自身的评分模型创建一个独立代理，本指南中测量运行使用了第三个模型来对两个候选模型进行评分。将评分视为一个参考信号，注意已知偏差，例如偏好较长的答案，并在没有评审参与的确定性结构检查中保持严格不变性。","如果你从本指南中只采纳一条原则，那就采纳这条：没有文档化通过，就不要进行模型替换。这意味着在测试中固定具体的版本，而不是别名，为测试套件分配自己的任务，任务键为保存你的提示、模型、工具和检索设置的路径，并安排一次单独运行以捕捉供应商端未提交的更改。结合这三点，可以将模型替换从一种判断变为有记录支撑的决策。","其余内容取决于案例集。我们的案例发现候选模型没有问题，而我们自己的测试平台存在两个问题，这也说明了在觉得自己需要之前开始测试是明智的。浏览模型目录：https://openrouter.ai/models 以选择候选模型，并阅读 Ori Eval 指南：https://openrouter.ai/docs/guides/ori/eval 来了解运行测试案例所用的测试平台。","模型使用数据、产品更新和研究报告。每周发送一封邮件。"]},"en":{"title":"OpenRouter Tutorial: How to Perform Regression Testing on an AI Agent After Prompt or Model Changes","summary":"OpenRouter has published a tutorial on regression testing for AI Agents: after every change to a prompt, model, tool definition, or retrieval setting, rerun a fixed set of test cases and check the results against a written behavioral contract.","category":"Industry","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"OpenRouter Tutorial: How to Perform Regression Testing on an AI Agent After Prompt or Model Changes - Aioga AI News","description":"OpenRouter has published a tutorial on regression testing for AI Agents: after every change to a prompt, model, tool definition, or retrieval setting, rerun a fixed set of test cas...","url":"https://www.aioga.com/en/news/lq71il0lehssgkular7zs5ybl/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"f7505f384ad7b9f4","translatedAt":"2026-09-30T03:12:19.910Z"},"ja":{"title":"OpenRouter チュートリアル：プロンプトやモデルを変更した後に AI エージェントの回帰テストを行う方法","summary":"OpenRouterはAIエージェントの回帰テストチュートリアルを公開しました：プロンプト、モデル、ツールの定義、または検索設定が変更されるたびに、ロックされたテストケースセットを再実行し、書面による行動契約と照らし合わせて確認します。","category":"業界動向","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"OpenRouter チュートリアル：プロンプトやモデルを変更した後に AI エージェントの回帰テストを行う方法 - Aioga AIニュース","description":"OpenRouterはAIエージェントの回帰テストチュートリアルを公開しました：プロンプト、モデル、ツールの定義、または検索設定が変更されるたびに、ロックされたテストケースセットを再実行し、書面による行動契約と照らし合わせて確認します。","url":"https://www.aioga.com/ja/news/lq71il0lehssgkular7zs5ybl/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"f7505f384ad7b9f4","translatedAt":"2026-09-30T03:12:36.328Z"},"ko":{"title":"OpenRouter 튜토리얼: 프롬프트나 모델 변경 후 AI 에이전트에 대한 회귀 테스트 수행 방법","summary":"OpenRouter, AI 에이전트 회귀 테스트 튜토리얼 공개: 프롬프트, 모델, 도구 정의 또는 검색 설정이 변경될 때마다 고정된 테스트 케이스 모음을 다시 실행하고 문서화된 동작 계약과 대조해 확인하세요.","category":"업계 동향","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"OpenRouter 튜토리얼: 프롬프트나 모델 변경 후 AI 에이전트에 대한 회귀 테스트 수행 방법 - Aioga AI 뉴스","description":"OpenRouter, AI 에이전트 회귀 테스트 튜토리얼 공개: 프롬프트, 모델, 도구 정의 또는 검색 설정이 변경될 때마다 고정된 테스트 케이스 모음을 다시 실행하고 문서화된 동작 계약과 대조해 확인하세요.","url":"https://www.aioga.com/ko/news/lq71il0lehssgkular7zs5ybl/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"f7505f384ad7b9f4","translatedAt":"2026-09-30T03:12:23.125Z"},"es":{"title":"Tutorial de OpenRouter: Cómo realizar pruebas de regresión en el Agente de IA después de cambiar el prompt o el modelo","summary":"OpenRouter publica un tutorial de prueba de regresión de agentes de IA: cada vez que se cambien los prompts, el modelo, la definición de herramientas o la configuración de recuperación, vuelva a ejecutar el conjunto de casos bloqueados y verifique contra el contrato de comportamiento escrito.","category":"Industria","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"Tutorial de OpenRouter: Cómo realizar pruebas de regresión en el Agente de IA después de cambiar el prompt o el modelo - Aioga Noticias de IA","description":"OpenRouter publica un tutorial de prueba de regresión de agentes de IA: cada vez que se cambien los prompts, el modelo, la definición de herramientas o la configuración de recupera...","url":"https://www.aioga.com/es/news/lq71il0lehssgkular7zs5ybl/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"f7505f384ad7b9f4","translatedAt":"2026-09-30T03:18:25.326Z"},"fr":{"title":"Tutoriel OpenRouter : Comment effectuer un test de régression sur un agent IA après un changement de prompt ou de modèle","summary":"OpenRouter publie le tutoriel de test de régression pour les agents IA : à chaque modification des invites, du modèle, de la définition des outils ou des paramètres de récupération, relancez l'ensemble des cas verrouillés et vérifiez-les par rapport au contrat de comportement écrit.","category":"Industrie","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"Tutoriel OpenRouter : Comment effectuer un test de régression sur un agent IA après un changement de prompt ou de modèle - Aioga Actualités IA","description":"OpenRouter publie le tutoriel de test de régression pour les agents IA : à chaque modification des invites, du modèle, de la définition des outils ou des paramètres de récupération...","url":"https://www.aioga.com/fr/news/lq71il0lehssgkular7zs5ybl/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"f7505f384ad7b9f4","translatedAt":"2026-09-30T03:18:19.462Z"},"de":{"title":"OpenRouter Tutorial: Wie man nach Änderungen von Prompts oder Modellen einen Regressionstest für einen AI-Agent durchführt","summary":"OpenRouter veröffentlicht Tutorial zum Regressions-Test von AI Agents: Nach jeder Änderung von Prompt, Modell, Tool-Definition oder Abruf-Einstellungen die gesperrte Testfallgruppe erneut ausführen und mit dem schriftlich festgelegten Verhaltensvertrag abgleichen.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"OpenRouter Tutorial: Wie man nach Änderungen von Prompts oder Modellen einen Regressionstest für einen AI-Agent durchführt - Aioga KI-News","description":"OpenRouter veröffentlicht Tutorial zum Regressions-Test von AI Agents: Nach jeder Änderung von Prompt, Modell, Tool-Definition oder Abruf-Einstellungen die gesperrte Testfallgruppe...","url":"https://www.aioga.com/de/news/lq71il0lehssgkular7zs5ybl/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"f7505f384ad7b9f4","translatedAt":"2026-09-30T03:18:19.472Z"},"pt-BR":{"title":"Tutorial do OpenRouter: Como realizar testes de regressão no Agente de IA após alterações no prompt ou no modelo","summary":"OpenRouter lançou tutorial de testes de regressão de agentes de IA: sempre que houver alterações nos prompts, modelos, definições de ferramentas ou configurações de recuperação, execute novamente o conjunto de casos bloqueados e verifique em comparação com o contrato de comportamento por escrito.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"Tutorial do OpenRouter: Como realizar testes de regressão no Agente de IA após alterações no prompt ou no modelo - Aioga Notícias de IA","description":"OpenRouter lançou tutorial de testes de regressão de agentes de IA: sempre que houver alterações nos prompts, modelos, definições de ferramentas ou configurações de recuperação, ex...","url":"https://www.aioga.com/pt-BR/news/lq71il0lehssgkular7zs5ybl/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"f7505f384ad7b9f4","translatedAt":"2026-09-30T03:24:07.283Z"},"ru":{"title":"Учебник OpenRouter: как проводить регрессионное тестирование AI-агента после изменения подсказок или моделей","summary":"OpenRouter выпустил руководство по регрессионному тестированию AI-агентов: каждый раз после изменения подсказок, моделей, определений инструментов или настроек поиска повторно запускайте закрепленный набор случаев и проверяйте их в соответствии с письменным контрактом поведения.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"Учебник OpenRouter: как проводить регрессионное тестирование AI-агента после изменения подсказок или моделей - Aioga Новости ИИ","description":"OpenRouter выпустил руководство по регрессионному тестированию AI-агентов: каждый раз после изменения подсказок, моделей, определений инструментов или настроек поиска повторно запу...","url":"https://www.aioga.com/ru/news/lq71il0lehssgkular7zs5ybl/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"f7505f384ad7b9f4","translatedAt":"2026-09-30T03:24:08.830Z"},"ar":{"title":"دليل OpenRouter: كيفية إجراء اختبارات الانحدار على وكيل الذكاء الاصطناعي بعد تغيير المطالبات أو النماذج","summary":"أصدرت OpenRouter دليل اختبار عودة وكيل الذكاء الاصطناعي: بعد كل تغيير في الكلمات المفتاحية، أو النموذج، أو تعريف الأدوات، أو إعدادات الاسترجاع، أعد تشغيل مجموعة الحالات المحجوزة وراجعها مقابل عقد السلوك المكتوب.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"دليل OpenRouter: كيفية إجراء اختبارات الانحدار على وكيل الذكاء الاصطناعي بعد تغيير المطالبات أو النماذج - Aioga أخبار الذكاء الاصطناعي","description":"أصدرت OpenRouter دليل اختبار عودة وكيل الذكاء الاصطناعي: بعد كل تغيير في الكلمات المفتاحية، أو النموذج، أو تعريف الأدوات، أو إعدادات الاسترجاع، أعد تشغيل مجموعة الحالات المحجوزة ور...","url":"https://www.aioga.com/ar/news/lq71il0lehssgkular7zs5ybl/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"f7505f384ad7b9f4","translatedAt":"2026-09-30T03:24:02.775Z"},"hi":{"title":"OpenRouter ट्यूटोरियल: संकेत शब्द या मॉडल परिवर्तन के बाद AI एजेंट पर रिग्रेशन टेस्ट कैसे करें","summary":"OpenRouter ने AI एजेंट रिग्रेशन टेस्टिंग ट्यूटोरियल जारी किया: हर बार प्रॉम्प्ट, मॉडल, टूल्स की परिभाषा या रिक्वेस्ट सेटिंग्स बदलने के बाद, लॉक किए गए टेस्ट केस सेट को फिर से चलाएँ और लिखित व्यवहार अनुबंध के अनुसार जांच करें।","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"OpenRouter ट्यूटोरियल: संकेत शब्द या मॉडल परिवर्तन के बाद AI एजेंट पर रिग्रेशन टेस्ट कैसे करें - Aioga AI समाचार","description":"OpenRouter ने AI एजेंट रिग्रेशन टेस्टिंग ट्यूटोरियल जारी किया: हर बार प्रॉम्प्ट, मॉडल, टूल्स की परिभाषा या रिक्वेस्ट सेटिंग्स बदलने के बाद, लॉक किए गए टेस्ट केस सेट को फिर से चलाएँ...","url":"https://www.aioga.com/hi/news/lq71il0lehssgkular7zs5ybl/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"f7505f384ad7b9f4","translatedAt":"2026-09-30T03:29:40.705Z"},"it":{"title":"Tutorial OpenRouter: come eseguire test di regressione su un agente AI dopo modifiche al prompt o al modello","summary":"OpenRouter ha pubblicato il tutorial di test di regressione per agenti AI: ogni volta che vengono modificati i prompt, il modello, la definizione degli strumenti o le impostazioni di ricerca, rieseguire il set di casi bloccati e confrontarli con il contratto comportamentale scritto.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"Tutorial OpenRouter: come eseguire test di regressione su un agente AI dopo modifiche al prompt o al modello - Aioga Notizie IA","description":"OpenRouter ha pubblicato il tutorial di test di regressione per agenti AI: ogni volta che vengono modificati i prompt, il modello, la definizione degli strumenti o le impostazioni...","url":"https://www.aioga.com/it/news/lq71il0lehssgkular7zs5ybl/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"f7505f384ad7b9f4","translatedAt":"2026-09-30T03:29:56.826Z"},"nl":{"title":"OpenRouter-tutorial: hoe je AI-agents regressietest na wijzigingen in prompts of modellen","summary":"OpenRouter heeft een tutorial over regressietests voor AI-agents gepubliceerd: voer na elke wijziging van een prompt, model, tooldefinitie of ophaalinstelling de vastgelegde testcases opnieuw uit en controleer de resultaten aan de hand van het schriftelijke gedragscontract.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"OpenRouter-tutorial: hoe je AI-agents regressietest na wijzigingen in prompts of modellen - Aioga AI-nieuws","description":"OpenRouter heeft een tutorial over regressietests voor AI-agents gepubliceerd: voer na elke wijziging van een prompt, model, tooldefinitie of ophaalinstelling de vastgelegde testca...","url":"https://www.aioga.com/nl/news/lq71il0lehssgkular7zs5ybl/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"f7505f384ad7b9f4","translatedAt":"2026-09-30T03:29:52.442Z"},"tr":{"title":"OpenRouter Kılavuzu: İpuçları veya model değiştikten sonra AI Ajan üzerinde regresyon testi nasıl yapılır","summary":"OpenRouter, AI Agent regresyon testi için bir eğitim yayımladı: Her istem, model, araç tanımı veya retrieval ayarı değişikliğinden sonra, sabitlenmiş test senaryosu kümesini yeniden çalıştırın ve yazılı davranış sözleşmesine göre kontrol edin.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"OpenRouter Kılavuzu: İpuçları veya model değiştikten sonra AI Ajan üzerinde regresyon testi nasıl yapılır - Aioga AI Haberleri","description":"OpenRouter, AI Agent regresyon testi için bir eğitim yayımladı: Her istem, model, araç tanımı veya retrieval ayarı değişikliğinden sonra, sabitlenmiş test senaryosu kümesini yenide...","url":"https://www.aioga.com/tr/news/lq71il0lehssgkular7zs5ybl/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"f7505f384ad7b9f4","translatedAt":"2026-09-30T03:35:14.392Z"},"vi":{"title":"Hướng dẫn OpenRouter: Cách kiểm thử hồi quy AI Agent sau khi thay đổi prompt hoặc mô hình","summary":"OpenRouter phát hành hướng dẫn kiểm thử hồi quy cho AI Agent: sau mỗi lần thay đổi lời nhắc, mô hình, định nghĩa công cụ hoặc cài đặt truy xuất, hãy chạy lại bộ ca kiểm thử đã được cố định và đối chiếu với hợp đồng hành vi bằng văn bản.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"Hướng dẫn OpenRouter: Cách kiểm thử hồi quy AI Agent sau khi thay đổi prompt hoặc mô hình - Tin tức AI Aioga","description":"OpenRouter phát hành hướng dẫn kiểm thử hồi quy cho AI Agent: sau mỗi lần thay đổi lời nhắc, mô hình, định nghĩa công cụ hoặc cài đặt truy xuất, hãy chạy lại bộ ca kiểm thử đã được...","url":"https://www.aioga.com/vi/news/lq71il0lehssgkular7zs5ybl/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"f7505f384ad7b9f4","translatedAt":"2026-09-30T03:35:20.240Z"},"id":{"title":"Tutorial OpenRouter: Cara melakukan pengujian regresi pada AI Agent setelah perubahan prompt atau model","summary":"OpenRouter menerbitkan tutorial pengujian regresi AI Agent: setiap kali terjadi perubahan pada prompt, model, definisi alat, atau pengaturan retrieval, jalankan kembali kumpulan kasus uji yang telah dikunci dan periksa hasilnya berdasarkan kontrak perilaku tertulis.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"Tutorial OpenRouter: Cara melakukan pengujian regresi pada AI Agent setelah perubahan prompt atau model - Berita AI Aioga","description":"OpenRouter menerbitkan tutorial pengujian regresi AI Agent: setiap kali terjadi perubahan pada prompt, model, definisi alat, atau pengaturan retrieval, jalankan kembali kumpulan ka...","url":"https://www.aioga.com/id/news/lq71il0lehssgkular7zs5ybl/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"f7505f384ad7b9f4","translatedAt":"2026-09-30T03:35:07.072Z"},"th":{"title":"บทช่วยสอน OpenRouter: วิธีทดสอบการถดถอยสำหรับ AI Agent หลังจากเปลี่ยนพรอมต์หรือโมเดล","summary":"OpenRouter เผยแพร่คู่มือการทดสอบการกลับมาของ AI Agent: ทุกครั้งที่มีการเปลี่ยนแปลงคำสั่ง, โมเดล, การกำหนดเครื่องมือหรือการตั้งค่าการค้นหา ให้รันชุดเคสที่ถูกล็อกใหม่และตรวจสอบกับสัญญาพฤติกรรมเป็นลายลักษณ์อักษร","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"บทช่วยสอน OpenRouter: วิธีทดสอบการถดถอยสำหรับ AI Agent หลังจากเปลี่ยนพรอมต์หรือโมเดล - ข่าว AI Aioga","description":"OpenRouter เผยแพร่คู่มือการทดสอบการกลับมาของ AI Agent: ทุกครั้งที่มีการเปลี่ยนแปลงคำสั่ง, โมเดล, การกำหนดเครื่องมือหรือการตั้งค่าการค้นหา ให้รันชุดเคสที่ถูกล็อกใหม่และตรวจสอบกับสัญ...","url":"https://www.aioga.com/th/news/lq71il0lehssgkular7zs5ybl/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"f7505f384ad7b9f4","translatedAt":"2026-09-30T03:40:59.376Z"},"pl":{"title":"Poradnik OpenRouter: jak przeprowadzać testy regresyjne agentów AI po zmianie promptu lub modelu","summary":"OpenRouter opublikował poradnik dotyczący testów regresji agentów AI: po każdej zmianie promptów, modeli, definicji narzędzi lub ustawień wyszukiwania należy ponownie uruchomić zablokowany zestaw przypadków testowych i sprawdzić wyniki względem pisemnego kontraktu zachowania.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"Poradnik OpenRouter: jak przeprowadzać testy regresyjne agentów AI po zmianie promptu lub modelu - Aioga Wiadomości AI","description":"OpenRouter opublikował poradnik dotyczący testów regresji agentów AI: po każdej zmianie promptów, modeli, definicji narzędzi lub ustawień wyszukiwania należy ponownie uruchomić zab...","url":"https://www.aioga.com/pl/news/lq71il0lehssgkular7zs5ybl/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"f7505f384ad7b9f4","translatedAt":"2026-09-30T03:40:50.225Z"}},"evidenceTier":"verified-news","reviewStatus":"editorial-selected","indexable":true,"editorialCover":""}}