OpenRouter 发布 AI Agent 回归测试教程:每次提示词、模型、工具定义或检索设置变更后,重跑锁定的用例集并对照书面行为契约检查。
行业动态OpenRouter:Announcements
今日 AI 情报摘要
OpenRouter 发布 AI Agent 回归测试教程:每次提示词、模型、工具定义或检索设置变更后,重跑锁定的用例集并对照书面行为契约检查。
中文正文 · AI 翻译
A ~author/family-latest 别名总是解析为该系列中最新的具体模型。在生产环境中这很方便,但在回归测试中可能成为问题,因为模型可能在运行之间发生变化,而您的仓库没有任何更改。我们的最新模型解析:https://openrouter.ai/docs/guides/routing/routers/latest-resolution 文档描述了该机制,并建议在需要可重复性的固定版本时使用具体的模型标识。本指南涵盖了锁定情况集和行为契约,然后详细介绍了模型交换的情况。
其余内容取决于案例集。我们的案例发现候选模型没有问题,而我们自己的测试平台存在两个问题,这也说明了在觉得自己需要之前开始测试是明智的。浏览模型目录:https://openrouter.ai/models 以选择候选模型,并阅读 Ori Eval 指南:https://openrouter.ai/docs/guides/ori/eval 来了解运行测试案例所用的测试平台。
模型使用数据、产品更新和研究报告。每周发送一封邮件。
A ~author/family-latest alias always resolves to the newest concrete model in a family. That’s convenient in production and a problem in a regression test, because the model can change between runs without any change in your repository. Our latest model resolution:https://openrouter.ai/docs/guides/routing/routers/latest-resolution docs describe the mechanism and recommend a concrete model slug when you need a fixed version for reproducibility. This guide covers the locked case set and the behavioral contract, then the model-swap case in detail.
Code regression testing rests on a known input, a known correct output, and a diff that tells you when the output changed. Three properties of an agent break that.
Two correct answers rarely look alike. A text diff against a golden answer fails on behavior that was never wrong. What holds still is structural. You check whether the agent called the right tool with the right arguments, respected the policy, and asked for the piece of information it was missing.
The model is a moving part. A model selected through a ~author/family-latest alias can change without a commit in your repository, and the part that changed is the one doing most of the reasoning. The model field in every OpenRouter response reports the concrete model that served the request. Reading it back is the cheapest way to notice that the model answering your calls is no longer the model you tested.
A pass expires when the baseline moves. Comparing against the same fixed set of cases every time is what turns “it seems fine” into a claim you can defend.
Agents drift on changes that a traditional test suite has no reason to look at. We group them into three kinds.
The third row is the one that is easiest to miss. A new chunking strategy for retrieved documents, an added field in a tool response, or a longer history can push content the agent relied on out of what it sees, and none of it touches the prompt. What you see is rarely an error. A support agent that used to quote the refund policy accurately starts paraphrasing it from memory, because the paragraph it relied on now falls outside the retrieved chunk, and the transcript reads just as fluently either way. A prompt edit has the same property. Tightening one sentence to fix one complaint can change which tool fires on an unrelated case.
Everything downstream depends on the case set, so build it before you think about automation.
Include representative cases that cover the requests your agent handles most often, a few edge cases such as ambiguous input or a request that sits on a policy boundary, and at least one case built to test a rule you never want broken. For a support agent that means a routine refund, a request with no order ID, and a refund above whatever limit your policy sets.
Once the set exists, stop editing it casually. Adding, removing, or rewording a case breaks comparability with every past run, and you lose the ability to tell a real regression from a different test. Every edit turns the set into a new experiment, so treat changes with the care you would give a schema migration.
For each case, write two things. The structural assertion says what the agent should do, such as calling lookup_order before acting and leaving escalate_to_human alone on a routine refund. The hard invariant says what the agent must never do, such as approving a refund above $500 without a human. That $500 is an example application policy rather than anything OpenRouter sets. The number in your own contract comes from your business rules. Most cases only need the structural assertion. The hard invariant is the one you want as an automatic ship-blocker, with no threshold and no judgment call attached.
Here is one case expressed as a plain API call, pinned to a concrete model, printing back both the model that served it and the tools it chose. The request sets no max_tokens , because a truncated response can cut off the tool call’s JSON and report a failure that has nothing to do with the agent’s decision.
The same case in Python with the OpenAI SDK pointed at our base URL.
The same case in TypeScript with fetch .
The mechanics are simple once the cases and contracts exist. A few details decide whether the run catches anything.
Trigger the run on the change. Re-run the full case set whenever a prompt, model, tool definition, or retrieval setting changes. A suite that runs only when someone remembers to run it will eventually miss the change that mattered.
Score the delta as well as the pass. A case that a judge scored well last month and scores lower today hasn’t failed, and it’s still a regression worth opening. Treat a meaningful score drop the way you would treat a failing test. Check that your rubric can move before you rely on it, because a rubric that scores every answer alike reports a clean pass while measuring nothing.
Match the check to the case. Deterministic cases, where you can name the exact tool and argument you expect, get exact or structural checks. Open-ended cases, such as whether an explanation is accurate and correctly scoped, need a judge model, because no single correct string exists to match against.
Ori Eval covers both shapes in one file. Assertions such as run.tool('lookup_order').toBeCalled() , run.toComplete() , run.toCostAtMost(0.01) , and run.toFinishWithin(30_000) handle the structural side. setupJudge({ minScore: 0.8 }) scores open-ended cases from 0 to 1 against criteria you write. Ori also resolves one harness and one model per run and holds them for every test in that run, so two runs of the same eval files use the same configuration.
Switching models on OpenRouter is a configuration change rather than a rewrite. That only helps if you can show that behavior stayed put when you made the switch.
Price is usually what starts the conversation. Two models we serve today sit at opposite ends of the price range, and both list tools in their supported parameters.
Checked September 18, 2026, against the live Claude Fable 5.1:https://openrouter.ai/anthropic/claude-fable-5.1 and Gemini 3.8 Flash:https://openrouter.ai/google/gemini-3.8-flash endpoint data. Gemini 3.8 Flash prices are for the standard tier. Both Google providers also serve flex and priority tiers at different prices. Prices change, so recheck before you plan around a ratio.
A thirteenfold difference in input price is reason enough to try the swap. The run is what earns the right to ship it. The mechanics are the locked case set and the contract you already have, with one variable moved. Pin the prompt, the tool definitions, the tool results, the case set, the judge, and the inference parameters, then run the suite against the candidate before any real traffic reaches it. When a result moves, you know the model moved it.
Pinning includes the slug itself. An alias like ~anthropic/claude-fable-latest routes to the newest concrete model in that family and updates whenever the author publishes a new version. Name the exact version on both sides of the comparison, and read the response’s model field to confirm what served each call.
Pinning also includes the inference parameters, and the two models don’t accept the same ones. Each entry in the models endpoint:https://openrouter.ai/docs/api/api-reference/models/list-all-models-and-their-properties has a supported_parameters array. Gemini 3.8 Flash lists temperature . Claude Fable 5.1 doesn’t, so with default routing a temperature value sent to it is ignored by the provider rather than applied, and setting it on one side of the comparison doesn’t hold the other side still. With require_parameters set, a parameter that no endpoint of the model supports means the request isn’t routed at all, so leave temperature out of the Claude side. Both models list reasoning and accept low , medium , and high as efforts, while their default efforts differ. The harness below sets reasoning.effort to medium for both and sets provider.require_parameters to true , so we only route each request to a provider endpoint that supports every parameter in it. See provider routing:https://openrouter.ai/docs/guides/routing/provider-selection#requiring-providers-to-support-all-parameters for the field. If every model in your comparison lists temperature , set it explicitly as well.
Here is the diff in its smallest useful form. It runs the same three cases against both slugs and separates a structural miss from a broken policy. The agent’s first move on a refund is a lookup, so the harness runs a short tool loop with fixed order records rather than reading a single response. The order data is pinned along with everything else, so a tool result can’t vary between runs.
The loop passes each assistant turn back with its reasoning_details unchanged, which our reasoning tokens:https://openrouter.ai/docs/guides/best-practices/reasoning-tokens#preserving-reasoning docs describe for tool calling with reasoning models. It stops after the first decision tool or after four turns, whichever comes first, and records every tool the model called in order.
The same comparison in Ori Eval loops over the two slugs in a single *.eval.ts file. Ori runs the tool loop for you and exposes the calls through run.tool() . The judge is pinned to a concrete model too. setupJudge() without an agent option grades on a default model, and we pass an explicit agent so the grading model is part of the pinned configuration.
Four ori eval flags map onto the pinning this section describes.
--baseline chooses what the run’s report is compared against. It takes last , best , or model: . The last form is the model swap as a single argument, comparing this run against a stored run of another model. The comparison is reporting only and doesn’t change the exit code. It reads run history from .ori/eval/history.jsonl , so it needs an Ori workspace, and a comparison is only possible between runs that included exactly the same eval files. --no-history keeps a run out of that file.
--hermetic gives the agent a fresh temporary workspace instead of your project directory, which keeps your ori.md , AGENTS.md , CLAUDE.md , and skill directories out of the run. Those files are context the agent reads, which makes them a variable you can change without noticing that you changed the agent. A CLAUDE.md or AGENTS.md at the repository root is checked in, edited often, and read on every run. Before it runs anything, ori eval lists the agent’s working directory and the instruction files and skill directories it discovered on stderr, so the run tells you what it had in front of it.
--dry-run loads every discovered eval and runs no tests, so a parse error or an unresolved import fails before any model call rather than after one. It needs no credentials, and it confirms nothing about whether an eval passes.
--pilot runs a strided sample of n cases from each eval file that wraps its cases in pilotCases() and reports the measured and estimated cost. It’s a cost measurement, not a comparison.
We ran this suite against both slugs three times on September 7, 2026, with openai/gpt-6-astra grading the open-ended case, and the swap held. Every structural case passed on both sides, the tool sequences came back identical, and no hard invariant moved.
Three runs per model on September 7, 2026, with openai/gpt-6-astra grading the open-ended case. The judge in that run reported on a 0 to 10 scale and returned 10.0 every time. Ori Eval’s setupJudge() reports from 0 to 1, and minScore is set on that scale. Eighteen structural case runs, all passing, with the same tool sequence every time. Those runs set no reasoning effort, so each model ran at its default. Model behavior changes, so read this as one dated measurement rather than a standing claim about either model.
That is the outcome you want from a swap. The run before it taught us more than these three did.
The first run failed on both sides, and our harness caused it. Both models failed the policy case with the same note. Neither had called issue_refund , so no invariant tripped.
The first version of the harness sent one request and read one response. An agent’s first move on that case is a lookup, so it never reached the decision the contract was written about, and the contract failed a decision the agent never had a chance to make. The tool loop in the harness above is the fix. A case that fails on both sides is a broken test, and the candidate column means nothing until that is fixed.
The judge returned the maximum score for every answer, from both models, across all three runs. A rubric that nothing can fail has no resolving power, and it will sit in your suite looking like coverage while detecting nothing. Ours asked whether the answer states the $500 limit and avoids inventing policy, which both models clear. A judge case earns its place only when a worse answer would score lower, so calibrate it by feeding it an answer you know is bad and confirming that the score moves.
Model routing and fallbacks change which provider or model serves a request. They don’t test anything. The eval run is what makes an easy switch safe to act on.
Treating every dip as a release blocker trains a team to ignore the gate, so the last piece of the pattern is deciding what deserves attention.
Judges drift, and they carry biases, including a preference for longer answers over shorter ones that are better. We haven’t measured judge agreement rates ourselves, so treat a single judge score as one signal. One failing case out of twenty isn’t automatically a reason to block a release, because it may be judge noise on a borderline case. Set a threshold for how many cases can fail before you treat it as a real signal, and re-run anything flaky before you trust a single failure.
Hard invariants are the exception, and they shouldn’t have a threshold. A refund above the limit or a skipped escalation goes to human review every time, no matter how many other cases passed. Stylistic drift is a judgment call. A broken policy boundary is a defect.
AI agent regression testing means re-running a locked, versioned set of test cases every time a prompt, model, tool definition, or retrieval setting changes, then checking that the agent’s behavior still matches a written contract. Because two correct answers from an agent rarely use the same words, the check is structural. It covers which tools the agent called with which arguments and whether policy boundaries held, rather than a text diff against a golden output.
Re-run your existing locked case set against the edited prompt with everything else held still, then compare each result to that case’s contract. Keep the case set unchanged so the comparison is valid. Pin the model to a concrete slug and set inference parameters explicitly so the prompt is the only variable. Apply structural assertions for deterministic cases and a judge score for open-ended ones. Treat a broken hard invariant as a release blocker and a score drop as something to investigate. Wiring the suite to a job scoped to your prompt files means the run happens on the change rather than when someone remembers.
Any test runner that can call your agent, assert on the tools it called, and score open-ended answers will work. The framework matters less than the locked case set and the written contract behind it. You can drive the agent from a general-purpose test runner such as pytest, Jest, or bun test with your own assertions, use an evaluation product that stores runs and diffs them for you, or use Ori Eval:https://openrouter.ai/docs/guides/ori/eval, which provides tool-call assertions, an LLM judge, and a pinned harness in *.eval.ts files.
A unit test diffs against one known-correct output. Agent regression testing checks structural assertions and hard invariants, because the output text varies between runs. The scope also differs. A unit test assumes the runtime under your code is stable, while an agent’s model can change when a provider ships a new version behind an alias or when you swap models yourself.
A behavioral contract is a per-case definition of what correct means. It has a structural assertion about which tool the agent called and what arguments it received, plus, where a policy exists, a hard invariant the agent must never violate, such as a refund limit or an escalation rule. Writing both down separates the cases where a judgment call is acceptable from the rule that should stop a release.
Run them on every change that touches a prompt, model, tool definition, or retrieval setting, triggered by the paths those files live in. Keep them out of the unit-test job that fires on every commit, because eval runs send requests to real models and cost money. A scheduled run catches provider-side changes that arrive without a commit of your own.
A model can grade its own output, but an independent judge on a different model stops one model’s blind spots from shaping both the answer and the score. Ori Eval’s setupJudge() creates a separate agent on its own grading model for that reason, and the measured run in this guide used a third model to grade both candidates. Treat the judge score as one signal, watch for known biases such as favoring longer answers, and keep hard invariants in deterministic structural checks where no judge is involved.
If you adopt one thing from this guide, adopt the rule that no model swap ships without a documented pass. That means pinning a concrete slug in the tests rather than an alias, giving the suite its own job keyed to the paths that hold your prompts, models, tools, and retrieval settings, and scheduling a separate run to catch the provider-side changes that arrive without a commit of your own. Together, those three turn a swap from a judgment call into a decision with a record behind it.
The rest follows from the case set. Ours found nothing wrong with the candidate model and two things wrong with our own harness, which is a good reason to start before you think you need to. Browse the model catalog:https://openrouter.ai/models to pick a candidate, and read the Ori Eval guide:https://openrouter.ai/docs/guides/ori/eval for the harness that runs your cases against it.
Model usage data, product updates, and research reports. One email each week.
情报判断
Aioga 编辑摘要
OpenRouter 发布 AI Agent 回归测试教程,建议在提示词、模型、工具定义或检索设置变更后,重新运行锁定用例集,并依据书面行为契约检查工具调用、参数、策略遵守情况及信息补充请求。
背景分析
教程指出,author/family-latest 别名会解析到某个模型系列中最新的具体模型,模型可能在代码库不变时发生变化。OpenRouter 响应中的 model 字段会报告实际服务请求的具体模型,固定版本复现时建议使用具体模型标识。