{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-10-06T08:00:48.640Z","headline":"OpenRouter 教程：如何测试 AI Agent 的工具调用准确性","description":"OpenRouter 发布教程，讲解如何测试 AI Agent 的工具调用准确性，将失败拆分为工具选择错误和参数错误两类分别测试。","url":"https://www.aioga.com/news/qu0zwrg9img0cfv03vtgmopbs/","mainEntityOfPage":"https://www.aioga.com/news/qu0zwrg9img0cfv03vtgmopbs/","datePublished":"2026-09-30T00:00:00.000Z","dateModified":"2026-09-30T00:00:00.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://openrouter.ai/blog/tutorials/how-to-test-tool-calling-accuracy-in-ai-agents","https://aihot.news/items/qu0zwrg9img0cfv03vtgmopbs"],"canonicalUrl":"https://www.aioga.com/news/qu0zwrg9img0cfv03vtgmopbs/","directAnswer":{"@type":"Answer","text":"OpenRouter 发布教程，介绍如何测试 AI Agent 的工具调用准确性，并将失败拆分为工具选择错误与参数错误两类，分别检查模型是否选对工具及是否传入正确参数。","url":"https://www.aioga.com/news/qu0zwrg9img0cfv03vtgmopbs/","dateCreated":"2026-09-30T00:00:00.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"OpenRouter：Announcements source article","url":"https://openrouter.ai/blog/tutorials/how-to-test-tool-calling-accuracy-in-ai-agents","datePublished":"2026-09-30T00:00:00.000Z","provider":{"@type":"Organization","name":"OpenRouter：Announcements","url":"https://openrouter.ai/blog/tutorials/how-to-test-tool-calling-accuracy-in-ai-agents"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.news/items/qu0zwrg9img0cfv03vtgmopbs","datePublished":"2026-09-30T00:00:00.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.news/items/qu0zwrg9img0cfv03vtgmopbs"}}],"aggregationSource":"OpenRouter：Announcements","originalPublisher":{"name":"OpenRouter：Announcements","url":"https://openrouter.ai/blog/tutorials/how-to-test-tool-calling-accuracy-in-ai-agents"},"geoDeepAnswer":null,"article":{"id":"qu0zwrg9img0cfv03vtgmopbs","slug":"qu0zwrg9img0cfv03vtgmopbs","url":"https://www.aioga.com/news/qu0zwrg9img0cfv03vtgmopbs/","title":"OpenRouter 教程：如何测试 AI Agent 的工具调用准确性","title_en":"","summary":"OpenRouter 发布教程，讲解如何测试 AI Agent 的工具调用准确性，将失败拆分为工具选择错误和参数错误两类分别测试。","source":"OpenRouter：Announcements","sourceUrl":"https://openrouter.ai/blog/tutorials/how-to-test-tool-calling-accuracy-in-ai-agents","aiHotUrl":"https://aihot.news/items/qu0zwrg9img0cfv03vtgmopbs","publishedAt":"2026-09-30T00:00:00.000Z","category":"行业动态","score":72,"selected":true,"articleBody":["An agent can fail in two places when it uses a tool. It can choose the wrong tool, or it can choose the right one and send the wrong arguments.","Those failures tell you different things. If an agent calls lookup_order instead of refund_order , the problem is tool selection. If it calls refund_order with the wrong order_id , it chose the right tool and passed the wrong arguments.","This guide covers three ways to test tool-calling behavior. The first is a reference-free large language model (LLM) judge. The second is deterministic argument checks. The third is trajectory comparison. It then shows how to run the same test cases against several tool-capable models through OpenRouter.","Start with the tool decision itself, then inspect the call the model produced.","Suppose an internal support agent can use lookup_order , issue_refund , and search_docs .","If the user asks what the refund policy says, search_docs is the appropriate tool. If they ask to refund order ord_7281 , the agent may need to look up the order before issuing the refund. Your test set should also include cases where the model already has enough information to answer and should not call a tool.","The no-tool case matters because checking only whether a response contains tool_calls isn’t enough. A model that calls an unnecessary or incorrect function still produces a tool call.","When one tool is clearly expected, compare the returned tool name with the expected one in code. If several tools could reasonably solve the request, exact matching can reject a valid choice. That’s where an LLM judge is more useful.","After the model chooses a tool, inspect the arguments it generated. This check has two parts, structure and values.","Structural validation catches malformed JSON, missing required fields, incorrect types, invalid enum values, and parameters the tool doesn’t accept.","A structurally valid call can still contain the wrong value.","If order_id is defined as a string, that payload satisfies the schema. It’s still wrong if the user asked about ord_7281 .","Existing eval frameworks make the same split. DeepEval has separate Tool Correctness：https://deepeval.com/docs/metrics-tool-correctness and Argument Correctness：https://deepeval.com/docs/metrics-argument-correctness metrics, and Phoenix has a separate evaluator for tool selection：https://arize.com/docs/phoenix/evaluation/pre-built-metrics/tool-selection.","A reference-free judge grades a tool call without a fixed expected answer. This approach is useful when correctness depends on context or several choices could be valid, so there is no single value you can compare against in code.","For tool selection, give the judge the user’s request, the tools available to the agent, and the model’s output. Then ask it to decide whether the selected tool was appropriate, including whether the model should have avoided using a tool at all.","Consider a research agent with both web_search and search_internal_docs . There may not be one correct choice. The better tool can depend on what the user asked and what information is already available in the conversation.","The same issue appears in arguments. A search query, description, or date range may fit the schema and still fail to represent what the user meant. If there is no fixed value you can compare against, a judge can evaluate the meaning instead.","A reference-free judge still needs clear instructions for what counts as correct. Keep those instructions and the judge model fixed when comparing candidate models, and check the judge’s decisions against a small set of cases you have reviewed yourself before you use it across the full dataset.","If an equality check, schema validator, or business rule can answer the same question reliably, use that instead.","Not every argument error needs another model call. If the tool schema can prove the failure, validate it in code.","The same schema you send to the model can validate the arguments it returns. It catches a missing order_id , a string where include_items expects a boolean, or an undeclared field such as customer_email .","Tool selection and argument checks cover individual calls. Multi-step agents can also fail in the sequence of calls they make. When that sequence is part of the requirement, trajectory comparison tests it directly.","A refund workflow might require three calls in order.","Strict matching only makes sense when order is required. LangSmith’s trajectory evaluators：https://docs.langchain.com/langsmith/trajectory-evals support strict, unordered, subset, and superset matching for this reason. A strict check enforces one sequence. The other modes accept different orderings or require only a particular set of calls.","Agent benchmarks handle this the same way. In τ²-bench：https://github.com/sierra-research/tau2-bench/blob/main/docs/evaluation.md, the recorded action list is one reference trajectory that is replayed to derive a target database end state. Any sequence of tool calls that produces an equivalent end state passes the database check.","If lookup_customer and lookup_subscription can happen in either order, don’t fail one sequence because your reference used the other. Grade the required calls or the resulting state instead.","A test case can use more than one check. For example, you can compare the tool name, validate its arguments against JSON Schema, and then compare known argument values with the expected payload.","Once you define the test cases and graders, you can run the same harness against each candidate model.","We expose one tool-calling interface：https://openrouter.ai/docs/guides/features/tool-calling across supported models, so you don’t need a separate provider integration for each model you want to compare.","The example sets tool_choice to \"auto\" . That’s the default when you supply tools, and setting it explicitly makes the no-tool test easier to follow.","This example uses the OpenAI Python SDK with our OpenAI-compatible endpoint. Install the dependencies first.","Set OPENROUTER_API_KEY in your environment, then run the same test cases against each candidate model.","A correct no-tool case counts toward tool selection and the overall result, and there is no schema or argument payload to grade. For cases that return tools, the harness validates every call before comparing the returned values with the expected payload.","The harness sets reasoning.effort to low for every candidate so that a difference in default reasoning effort doesn’t show up as a difference in tool-calling accuracy. Each of the three models above lists low in the supported_efforts array of its reasoning object in the models endpoint. It also sets provider.require_parameters to true . A model’s supported_parameters list can include a parameter that only some of that model’s provider endpoints accept, and with default routing a provider that doesn’t support a parameter still receives the request and ignores it. With require_parameters set, we only route the request to providers that support every parameter in it, so each graded response ran at the effort the harness asked for. See provider routing：https://openrouter.ai/docs/guides/routing/provider-selection#requiring-providers-to-support-all-parameters for the field. The harness doesn’t set temperature , because openai/gpt-5.6-sol doesn’t list temperature in supported_parameters . If every candidate in your list accepts temperature , set it explicitly as well.","The model IDs above are examples. Each entry in the models endpoint：https://openrouter.ai/docs/api/api-reference/models/list-all-models-and-their-properties has a supported_parameters array. A model supports this harness when that array includes tools and tool_choice . Check the current tool-calling models collection：https://openrouter.ai/collections/tool-calling-models before you fix a candidate list in a long-lived eval suite.","For a real comparison, run each test case more than once so a model’s score isn’t based on a single response. A larger suite should also include the harder cases your application sees, such as missing arguments, similar tool descriptions, multiple calls, and requests where no tool should be used.","This harness evaluates one tool-calling turn. For a multi-step agent, collect calls across the full trace and compare the sequence or the resulting state, depending on what the workflow requires.","Provider routing also affects what your comparison measures. Auto Exacto：https://openrouter.ai/docs/guides/routing/auto-exacto runs by default on every request that includes tools and reorders providers for your chosen model, so it can change which provider endpoint serves a tool-calling request. One of its inputs is the Tool Call Error Rate. For each request that includes tools, we inspect every tool call the model returned and classify structural failures as InvalidJson , UnknownName , or SchemaMismatch , validating arguments against the parameters schema you supplied under JSON Schema Draft 7. That metric measures provider behavior. It doesn’t replace the local schema validation in your own harness, which is why the example above validates arguments itself.","If you want to test the routing setup your application will use in production, leave Auto Exacto enabled for every candidate. If you want an endpoint-level comparison, use our provider routing controls：https://openrouter.ai/docs/guides/routing/provider-selection to pin the provider. Set the order field in the provider object to that provider’s slug and set allow_fallbacks to false , so every request goes to one endpoint.","Keep the prompts, tools, test cases, judge model, and evaluation criteria the same between runs. Set sampling and reasoning parameters such as temperature and reasoning explicitly rather than relying on defaults, and check that each candidate supports the parameters you set. Don’t set max_tokens in the harness. A truncated response can cut off the tool call’s JSON and show up as an InvalidJson failure that has nothing to do with the model’s tool selection.","If you’d rather not maintain the cross-model runner yourself, Ori Eval：https://openrouter.ai/blog/announcements/ori-eval/ runs your agent against candidate models, asserts on the tools it called and the tools it avoided, and grades open-ended answers with an LLM judge.","A tool-calling eval can give you misleading results if the test cases or grading rules are too narrow.","Model usage data, product updates, and research reports. One email each week."],"articleImages":[{"sourceUrl":"https://openrouter.ai/blog/images/how-to-test-tool-calling-accuracy-in-ai-agents.png","alt":"How to Test Tool-Calling Accuracy in AI Agents","afterParagraph":0,"url":"/media/articles/qu0zwrg9img0cfv03vtgmopbs/1f7f68a0ad668a6d.png"},{"sourceUrl":"https://openrouter.ai/blog/images/ai-agent-regression-testing-after-a-prompt-or-model-change_2f6zbV.webp","alt":"AI Agent Regression Testing After a Prompt or Model Change","afterParagraph":41,"url":"/media/articles/qu0zwrg9img0cfv03vtgmopbs/d69409a6ab76363d.webp"}],"mediaStatus":"ok","articleBodyZh":["当代理使用工具时，可能在两个环节失败。它可能选择错误的工具，或者选择了正确的工具却传递了错误的参数。","这些失败告诉你不同的信息。如果代理调用了 lookup_order 而不是 refund_order，问题出在工具选择上。如果它调用了 refund_order 但传递了错误的 order_id，那么它选择了正确的工具但传递了错误的参数。","本指南涵盖了测试工具调用行为的三种方法。第一种是无参考的大语言模型（LLM）评判。第二种是确定性的参数检查。第三种是轨迹比较。然后展示了如何通过 OpenRouter 对几个支持工具的模型运行相同的测试用例。","首先从工具决策本身开始，然后检查模型产生的调用。","假设内部支持代理可以使用 lookup_order、issue_refund 和 search_docs。","如果用户询问退款政策的内容，search_docs 是适当的工具。如果他们要求退款订单 ord_7281，代理可能需要先查找订单再进行退款。你的测试集还应包括模型已经有足够信息回答、无需调用工具的情况。","无工具的情况很重要，因为仅检查响应中是否包含 tool_calls 并不足够。调用了不必要或错误函数的模型仍然会产生工具调用。","当明显期待某个工具时，在代码中将返回的工具名称与预期名称进行比较。如果几个工具都能合理解决请求，精确匹配可能拒绝一个有效选择。这时 LLM 评判更有用。","在模型选择工具之后，检查它生成的参数。这个检查有两个部分：结构和数值。","结构验证可以发现格式错误的 JSON、缺少必填字段、类型错误、无效的枚举值，以及工具不接受的参数。","结构有效的调用仍然可能包含错误的值。","如果 order_id 定义为字符串，该负载符合模式。如果用户询问的是 ord_7281，那么它仍然是错误的。","现有的评估框架采用相同的划分。DeepEval 有独立的工具正确性量化指标：https://deepeval.com/docs/metrics-tool-correctness 和论证正确性量化指标：https://deepeval.com/docs/metrics-argument-correctness，而 Phoenix 则有一个独立的工具选择评估器：https://arize.com/docs/phoenix/evaluation/pre-built-metrics/tool-selection。","无参考判官会在没有固定预期答案的情况下对工具调用进行评分。当正确性依赖于上下文或多种选择都可能有效时，这种方法非常有用，因此在代码中无法对比单一值。","对于工具选择，向判官提供用户请求、代理可用的工具以及模型输出。然后让它决定所选工具是否合适，包括模型是否完全不应该使用工具。","考虑一个同时具备 web_search 和 search_internal_docs 功能的研究代理。可能不存在唯一正确的选择。更优的工具取决于用户的提问以及对话中已有的信息。","同样的问题也出现在参数上。搜索查询、描述或日期范围可能符合模式要求，但仍未能准确表示用户意图。如果没有固定值可对比，判官可以评估其含义。","无参考判官仍然需要明确说明什么算作正确。在比较候选模型时，请保持这些说明和判官模型不变，并在将其应用于完整数据集之前，用你自己审核过的一小组案例检查判官的决定。","如果等值检查、模式验证器或业务规则能够可靠地回答同一个问题，请使用它们。","并非每个参数错误都需要再次调用模型。如果工具模式可以证明失败，请在代码中进行验证。","你发送给模型的同样模式可以验证它返回的参数。它可以捕捉缺失的 order_id、预期布尔值却提供了字符串的 include_items、或未声明字段如 customer_email。","工具选择和参数检查覆盖单次调用。多步骤代理的失败也可能出现在它们调用的序列中。当该序列是要求的一部分时，轨迹比较可以直接测试它。","退款工作流可能需要按顺序进行三次调用。","只有在顺序要求时，严格匹配才有意义。LangSmith 的轨迹评估器：https://docs.langchain.com/langsmith/trajectory-evals 因此支持严格匹配、无序匹配、子集匹配和超集匹配。严格检查会强制执行一个序列。其他模式接受不同的顺序或只要求特定的一组调用。","代理基准测试也是以相同的方式处理的。在 τ²-bench：https://github.com/sierra-research/tau2-bench/blob/main/docs/evaluation.md 中，记录的操作列表是一个参考轨迹，它会被重放以得出目标数据库的最终状态。任何产生等效最终状态的工具调用序列都能通过数据库检查。","如果 lookup_customer 和 lookup_subscription 可以按任意顺序发生，不要因为你的参考使用了另一种顺序而让一个序列失败。应对所需的调用或结果状态进行评分。","一个测试用例可以使用多个检查。例如，你可以比较工具名称，用 JSON Schema 验证其参数，然后将已知参数值与期望的负载进行比较。","一旦定义了测试用例和评分器，你就可以对每个候选模型运行相同的测试框架。","我们提供了一个工具调用接口：https://openrouter.ai/docs/guides/features/tool-calling，适用于支持的模型，因此你不需要为每个想要比较的模型单独进行提供者集成。","示例将 tool_choice 设置为 \"auto\"。这是提供工具时的默认值，显式设置它可以使无工具测试更易于理解。","此示例使用 OpenAI Python SDK 与我们兼容 OpenAI 的端点。请先安装依赖项。","在环境中设置 OPENROUTER_API_KEY，然后对每个候选模型运行相同的测试用例。","一个正确的无工具用例计入工具选择和整体结果评分，并且没有模式或参数负载需要评分。对于返回工具的用例，测试框架会验证每次调用，然后将返回值与期望的负载进行比较。","该测试工具将 reasoning.effort 为每个候选项设置为低，以便默认推理努力的差异不会显示为工具调用准确性的差异。上述三种模型在 models 接口的 reasoning 对象的 supported_efforts 数组中都列出了 low。它还将 provider.require_parameters 设置为 true。一个模型的 supported_parameters 列表可以包含只有该模型部分提供者端点接受的参数，而在默认路由下，不支持某个参数的提供者仍然会收到请求并忽略它。设置 require_parameters 后，我们只将请求路由到支持请求中每个参数的提供者，因此每个评分响应都按测试工具要求的努力级别运行。关于提供者路由，请参见：https://openrouter.ai/docs/guides/routing/provider-selection#requiring-providers-to-support-all-parameters。该测试工具不设置 temperature，因为 openai/gpt-5.6-sol 在 supported_parameters 中未列出 temperature。如果你的候选列表中的每个模型都接受 temperature，请也显式设置它。","上述模型 ID 仅为示例。models 接口的每个条目：https://openrouter.ai/docs/api/api-reference/models/list-all-models-and-their-properties 都有一个 supported_parameters 数组。当该数组包含 tools 和 tool_choice 时，模型支持此测试工具。在你在长期评估套件中固定候选列表之前，请检查当前的工具调用模型集合：https://openrouter.ai/collections/tool-calling-models。","为了进行真实比较，每个测试用例应运行多次，以免模型的分数基于单次响应。较大的测试套件还应包括你的应用可能遇到的更复杂情况，例如缺失参数、工具描述相似、多次调用以及不应使用任何工具的请求。","该测试工具评估一次工具调用。对于多步骤代理，请收集整个跟踪过程中的调用，并根据工作流需要比较调用序列或最终状态。","提供者路由也会影响你的比较衡量指标。Auto Exacto：https://openrouter.ai/docs/guides/routing/auto-exacto 默认在每个包含工具的请求上运行，并为你选择的模型重新排序提供者，因此它可能会更改哪个提供者端点处理工具调用请求。其输入之一是工具调用错误率。对于每个包含工具的请求，我们会检查模型返回的每个工具调用，并将结构性失败分类为 InvalidJson、UnknownName 或 SchemaMismatch，按照你在 JSON Schema Draft 7 下提供的参数架构验证参数。该指标衡量提供者的行为。它不能替代你自己运行环境中的本地架构验证，这也是上面示例自行验证参数的原因。","如果你想测试应用在生产中使用的路由设置，请对每个候选模型保持 Auto Exacto 启用。如果你想进行端点级比较，请使用我们的提供者路由控制：https://openrouter.ai/docs/guides/routing/provider-selection 固定提供者。在提供者对象中将 order 字段设置为该提供者的 slug，并将 allow_fallbacks 设置为 false，这样每个请求都会发送到同一个端点。","在运行之间保持提示、工具、测试用例、评判模型和评估标准相同。明确设置采样和推理参数，如 temperature 和 reasoning，而不要依赖默认值，并检查每个候选模型是否支持你设置的参数。在运行环境中不要设置 max_tokens。响应被截断可能导致工具调用的 JSON 被截断，并显示为 InvalidJson 错误，这与模型的工具选择无关。","如果你不想自己维护跨模型运行器，Ori Eval：https://openrouter.ai/blog/announcements/ori-eval/ 可以对你的代理进行候选模型测试，断言其调用的工具和未调用的工具，并使用 LLM 评审对开放式答案进行评分。","如果测试用例或评分规则过于狭窄，工具调用评估可能会给出误导性结果。","模型使用数据、产品更新和研究报告。每周发送一封邮件。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"OpenRouter 发布教程，介绍如何测试 AI Agent 的工具调用准确性，并将失败拆分为工具选择错误与参数错误两类，分别检查模型是否选对工具及是否传入正确参数。","background":"教程提出三种测试方式：无参考答案的大语言模型评审、确定性的参数检查和轨迹比较。材料还强调，测试集应覆盖应调用工具、无需调用工具，以及多个工具均可能适用的场景。","viewpoint":"Aioga 判断：将工具选择与参数传递分开评估，有助于更清晰地定位工具调用失败，但不同测试方法的适用条件并不相同，不能仅凭是否出现工具调用来判断准确性。","implications":"可能影响：工具调用评测需要同时关注工具名称、参数内容和调用必要性。单一的精确匹配可能不足以覆盖多个合理工具选择的情况，测试设计需要结合具体任务判断。","nextStep":"后续观察：可关注该教程所述方法在不同工具集和模型测试中的实际表现，尤其是无工具调用、工具选择存在多种合理答案，以及参数检查结果之间的差异。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-10-01T17:26:34.514Z","sourceHash":"9636eb1c4c5b0bbf","review":{"approved":true,"groundedness":96,"clarity":93,"duplicationRisk":12,"blockingIssues":[],"notes":["“Aioga 判断”明确标注为观点，未将其冒充为来源事实。","“不能仅凭是否出现工具调用来判断准确性”与原文关于无工具场景及错误工具调用的说明一致。","“实际表现”属于合理的后续观察建议，不是对来源材料的事实扩展。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","editorial-labels","inference-boundary","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","OpenRouter：Announcements"],"translations":{"zh-CN":{"title":"OpenRouter 教程：如何测试 AI Agent 的工具调用准确性","summary":"OpenRouter 发布教程，讲解如何测试 AI Agent 的工具调用准确性，将失败拆分为工具选择错误和参数错误两类分别测试。","category":"行业动态","source":"openrouter.ai","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"OpenRouter 教程：如何测试 AI Agent 的工具调用准确性 - Aioga AI资讯","description":"OpenRouter 发布教程，讲解如何测试 AI Agent 的工具调用准确性，将失败拆分为工具选择错误和参数错误两类分别测试。","url":"https://www.aioga.com/news/qu0zwrg9img0cfv03vtgmopbs/","articleBody":["当代理使用工具时，可能在两个环节失败。它可能选择错误的工具，或者选择了正确的工具却传递了错误的参数。","这些失败告诉你不同的信息。如果代理调用了 lookup_order 而不是 refund_order，问题出在工具选择上。如果它调用了 refund_order 但传递了错误的 order_id，那么它选择了正确的工具但传递了错误的参数。","本指南涵盖了测试工具调用行为的三种方法。第一种是无参考的大语言模型（LLM）评判。第二种是确定性的参数检查。第三种是轨迹比较。然后展示了如何通过 OpenRouter 对几个支持工具的模型运行相同的测试用例。","首先从工具决策本身开始，然后检查模型产生的调用。","假设内部支持代理可以使用 lookup_order、issue_refund 和 search_docs。","如果用户询问退款政策的内容，search_docs 是适当的工具。如果他们要求退款订单 ord_7281，代理可能需要先查找订单再进行退款。你的测试集还应包括模型已经有足够信息回答、无需调用工具的情况。","无工具的情况很重要，因为仅检查响应中是否包含 tool_calls 并不足够。调用了不必要或错误函数的模型仍然会产生工具调用。","当明显期待某个工具时，在代码中将返回的工具名称与预期名称进行比较。如果几个工具都能合理解决请求，精确匹配可能拒绝一个有效选择。这时 LLM 评判更有用。","在模型选择工具之后，检查它生成的参数。这个检查有两个部分：结构和数值。","结构验证可以发现格式错误的 JSON、缺少必填字段、类型错误、无效的枚举值，以及工具不接受的参数。","结构有效的调用仍然可能包含错误的值。","如果 order_id 定义为字符串，该负载符合模式。如果用户询问的是 ord_7281，那么它仍然是错误的。","现有的评估框架采用相同的划分。DeepEval 有独立的工具正确性量化指标：https://deepeval.com/docs/metrics-tool-correctness 和论证正确性量化指标：https://deepeval.com/docs/metrics-argument-correctness，而 Phoenix 则有一个独立的工具选择评估器：https://arize.com/docs/phoenix/evaluation/pre-built-metrics/tool-selection。","无参考判官会在没有固定预期答案的情况下对工具调用进行评分。当正确性依赖于上下文或多种选择都可能有效时，这种方法非常有用，因此在代码中无法对比单一值。","对于工具选择，向判官提供用户请求、代理可用的工具以及模型输出。然后让它决定所选工具是否合适，包括模型是否完全不应该使用工具。","考虑一个同时具备 web_search 和 search_internal_docs 功能的研究代理。可能不存在唯一正确的选择。更优的工具取决于用户的提问以及对话中已有的信息。","同样的问题也出现在参数上。搜索查询、描述或日期范围可能符合模式要求，但仍未能准确表示用户意图。如果没有固定值可对比，判官可以评估其含义。","无参考判官仍然需要明确说明什么算作正确。在比较候选模型时，请保持这些说明和判官模型不变，并在将其应用于完整数据集之前，用你自己审核过的一小组案例检查判官的决定。","如果等值检查、模式验证器或业务规则能够可靠地回答同一个问题，请使用它们。","并非每个参数错误都需要再次调用模型。如果工具模式可以证明失败，请在代码中进行验证。","你发送给模型的同样模式可以验证它返回的参数。它可以捕捉缺失的 order_id、预期布尔值却提供了字符串的 include_items、或未声明字段如 customer_email。","工具选择和参数检查覆盖单次调用。多步骤代理的失败也可能出现在它们调用的序列中。当该序列是要求的一部分时，轨迹比较可以直接测试它。","退款工作流可能需要按顺序进行三次调用。","只有在顺序要求时，严格匹配才有意义。LangSmith 的轨迹评估器：https://docs.langchain.com/langsmith/trajectory-evals 因此支持严格匹配、无序匹配、子集匹配和超集匹配。严格检查会强制执行一个序列。其他模式接受不同的顺序或只要求特定的一组调用。","代理基准测试也是以相同的方式处理的。在 τ²-bench：https://github.com/sierra-research/tau2-bench/blob/main/docs/evaluation.md 中，记录的操作列表是一个参考轨迹，它会被重放以得出目标数据库的最终状态。任何产生等效最终状态的工具调用序列都能通过数据库检查。","如果 lookup_customer 和 lookup_subscription 可以按任意顺序发生，不要因为你的参考使用了另一种顺序而让一个序列失败。应对所需的调用或结果状态进行评分。","一个测试用例可以使用多个检查。例如，你可以比较工具名称，用 JSON Schema 验证其参数，然后将已知参数值与期望的负载进行比较。","一旦定义了测试用例和评分器，你就可以对每个候选模型运行相同的测试框架。","我们提供了一个工具调用接口：https://openrouter.ai/docs/guides/features/tool-calling，适用于支持的模型，因此你不需要为每个想要比较的模型单独进行提供者集成。","示例将 tool_choice 设置为 \"auto\"。这是提供工具时的默认值，显式设置它可以使无工具测试更易于理解。","此示例使用 OpenAI Python SDK 与我们兼容 OpenAI 的端点。请先安装依赖项。","在环境中设置 OPENROUTER_API_KEY，然后对每个候选模型运行相同的测试用例。","一个正确的无工具用例计入工具选择和整体结果评分，并且没有模式或参数负载需要评分。对于返回工具的用例，测试框架会验证每次调用，然后将返回值与期望的负载进行比较。","该测试工具将 reasoning.effort 为每个候选项设置为低，以便默认推理努力的差异不会显示为工具调用准确性的差异。上述三种模型在 models 接口的 reasoning 对象的 supported_efforts 数组中都列出了 low。它还将 provider.require_parameters 设置为 true。一个模型的 supported_parameters 列表可以包含只有该模型部分提供者端点接受的参数，而在默认路由下，不支持某个参数的提供者仍然会收到请求并忽略它。设置 require_parameters 后，我们只将请求路由到支持请求中每个参数的提供者，因此每个评分响应都按测试工具要求的努力级别运行。关于提供者路由，请参见：https://openrouter.ai/docs/guides/routing/provider-selection#requiring-providers-to-support-all-parameters。该测试工具不设置 temperature，因为 openai/gpt-5.6-sol 在 supported_parameters 中未列出 temperature。如果你的候选列表中的每个模型都接受 temperature，请也显式设置它。","上述模型 ID 仅为示例。models 接口的每个条目：https://openrouter.ai/docs/api/api-reference/models/list-all-models-and-their-properties 都有一个 supported_parameters 数组。当该数组包含 tools 和 tool_choice 时，模型支持此测试工具。在你在长期评估套件中固定候选列表之前，请检查当前的工具调用模型集合：https://openrouter.ai/collections/tool-calling-models。","为了进行真实比较，每个测试用例应运行多次，以免模型的分数基于单次响应。较大的测试套件还应包括你的应用可能遇到的更复杂情况，例如缺失参数、工具描述相似、多次调用以及不应使用任何工具的请求。","该测试工具评估一次工具调用。对于多步骤代理，请收集整个跟踪过程中的调用，并根据工作流需要比较调用序列或最终状态。","提供者路由也会影响你的比较衡量指标。Auto Exacto：https://openrouter.ai/docs/guides/routing/auto-exacto 默认在每个包含工具的请求上运行，并为你选择的模型重新排序提供者，因此它可能会更改哪个提供者端点处理工具调用请求。其输入之一是工具调用错误率。对于每个包含工具的请求，我们会检查模型返回的每个工具调用，并将结构性失败分类为 InvalidJson、UnknownName 或 SchemaMismatch，按照你在 JSON Schema Draft 7 下提供的参数架构验证参数。该指标衡量提供者的行为。它不能替代你自己运行环境中的本地架构验证，这也是上面示例自行验证参数的原因。","如果你想测试应用在生产中使用的路由设置，请对每个候选模型保持 Auto Exacto 启用。如果你想进行端点级比较，请使用我们的提供者路由控制：https://openrouter.ai/docs/guides/routing/provider-selection 固定提供者。在提供者对象中将 order 字段设置为该提供者的 slug，并将 allow_fallbacks 设置为 false，这样每个请求都会发送到同一个端点。","在运行之间保持提示、工具、测试用例、评判模型和评估标准相同。明确设置采样和推理参数，如 temperature 和 reasoning，而不要依赖默认值，并检查每个候选模型是否支持你设置的参数。在运行环境中不要设置 max_tokens。响应被截断可能导致工具调用的 JSON 被截断，并显示为 InvalidJson 错误，这与模型的工具选择无关。","如果你不想自己维护跨模型运行器，Ori Eval：https://openrouter.ai/blog/announcements/ori-eval/ 可以对你的代理进行候选模型测试，断言其调用的工具和未调用的工具，并使用 LLM 评审对开放式答案进行评分。","如果测试用例或评分规则过于狭窄，工具调用评估可能会给出误导性结果。","模型使用数据、产品更新和研究报告。每周发送一封邮件。"]},"en":{"title":"OpenRouter Tutorial: How to Test the Tool Invocation Accuracy of an AI Agent","summary":"OpenRouter release tutorial, explaining how to test the accuracy of AI Agent tool calls, splitting failures into two categories: wrong tool selection and parameter errors for separate testing.","category":"Industry","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"OpenRouter Tutorial: How to Test the Tool Invocation Accuracy of an AI Agent - Aioga AI News","description":"OpenRouter release tutorial, explaining how to test the accuracy of AI Agent tool calls, splitting failures into two categories: wrong tool selection and parameter errors for separ...","url":"https://www.aioga.com/en/news/qu0zwrg9img0cfv03vtgmopbs/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"7569195cd17ab66e","translatedAt":"2026-09-30T03:12:19.910Z"},"ja":{"title":"OpenRouterチュートリアル：AI Agentのツール呼び出し精度をテストする方法","summary":"OpenRouter の公開チュートリアルで、AIエージェントのツール呼び出し精度をテストする方法を解説し、失敗をツール選択ミスとパラメータミスの2つに分けてそれぞれテストします。","category":"業界動向","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"OpenRouterチュートリアル：AI Agentのツール呼び出し精度をテストする方法 - Aioga AIニュース","description":"OpenRouter の公開チュートリアルで、AIエージェントのツール呼び出し精度をテストする方法を解説し、失敗をツール選択ミスとパラメータミスの2つに分けてそれぞれテストします。","url":"https://www.aioga.com/ja/news/qu0zwrg9img0cfv03vtgmopbs/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"7569195cd17ab66e","translatedAt":"2026-09-30T03:12:36.328Z"},"ko":{"title":"OpenRouter 튜토리얼: AI 에이전트의 도구 호출 정확성을 테스트하는 방법","summary":"OpenRouter가 AI Agent의 도구 호출 정확도를 테스트하는 방법을 설명하는 튜토리얼을 공개했습니다. 실패를 도구 선택 오류와 매개변수 오류로 나누어 각각 테스트합니다.","category":"업계 동향","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"OpenRouter 튜토리얼: AI 에이전트의 도구 호출 정확성을 테스트하는 방법 - Aioga AI 뉴스","description":"OpenRouter가 AI Agent의 도구 호출 정확도를 테스트하는 방법을 설명하는 튜토리얼을 공개했습니다. 실패를 도구 선택 오류와 매개변수 오류로 나누어 각각 테스트합니다.","url":"https://www.aioga.com/ko/news/qu0zwrg9img0cfv03vtgmopbs/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"7569195cd17ab66e","translatedAt":"2026-09-30T03:12:23.125Z"},"es":{"title":"Tutorial de OpenRouter: cómo probar la precisión de las llamadas a herramientas de los agentes de IA","summary":"OpenRouter publicó un tutorial que explica cómo probar la precisión de las llamadas a herramientas de un AI Agent, dividiendo los fallos en dos categorías: errores de selección de herramientas y errores de parámetros, para probarlas por separado.","category":"Industria","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"Tutorial de OpenRouter: cómo probar la precisión de las llamadas a herramientas de los agentes de IA - Aioga Noticias de IA","description":"OpenRouter publicó un tutorial que explica cómo probar la precisión de las llamadas a herramientas de un AI Agent, dividiendo los fallos en dos categorías: errores de selección de...","url":"https://www.aioga.com/es/news/qu0zwrg9img0cfv03vtgmopbs/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"7569195cd17ab66e","translatedAt":"2026-09-30T03:18:25.326Z"},"fr":{"title":"Tutoriel OpenRouter : Comment tester la précision des appels d'outils d'un agent IA","summary":"Tutoriel de publication d'OpenRouter, expliquant comment tester la précision des appels d'outils d'un agent IA, en divisant les échecs en deux catégories : erreur de sélection d'outil et erreur de paramètre, et en les testant séparément.","category":"Industrie","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"Tutoriel OpenRouter : Comment tester la précision des appels d'outils d'un agent IA - Aioga Actualités IA","description":"Tutoriel de publication d'OpenRouter, expliquant comment tester la précision des appels d'outils d'un agent IA, en divisant les échecs en deux catégories : erreur de sélection d'ou...","url":"https://www.aioga.com/fr/news/qu0zwrg9img0cfv03vtgmopbs/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"7569195cd17ab66e","translatedAt":"2026-09-30T03:18:19.462Z"},"de":{"title":"OpenRouter Anleitung: Wie man die Genauigkeit der Toolaufrufe eines KI-Agenten testet","summary":"OpenRouter Veröffentlichung Tutorial, erklärt, wie man die Genauigkeit der Werkzeugaufrufe von AI-Agenten testet und Misserfolge in zwei Kategorien unterteilt: falsche Werkzeugwahl und fehlerhafte Parameter, die jeweils separat getestet werden.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"OpenRouter Anleitung: Wie man die Genauigkeit der Toolaufrufe eines KI-Agenten testet - Aioga KI-News","description":"OpenRouter Veröffentlichung Tutorial, erklärt, wie man die Genauigkeit der Werkzeugaufrufe von AI-Agenten testet und Misserfolge in zwei Kategorien unterteilt: falsche Werkzeugwahl...","url":"https://www.aioga.com/de/news/qu0zwrg9img0cfv03vtgmopbs/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"7569195cd17ab66e","translatedAt":"2026-09-30T03:18:19.472Z"},"pt-BR":{"title":"Tutorial do OpenRouter: como testar a precisão das chamadas de ferramentas de um agente de IA","summary":"Tutorial de lançamento do OpenRouter, explicando como testar a precisão da chamada de ferramentas do Agente de IA, dividindo as falhas em duas categorias para teste: escolha incorreta da ferramenta e erro de parâmetro.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"Tutorial do OpenRouter: como testar a precisão das chamadas de ferramentas de um agente de IA - Aioga Notícias de IA","description":"Tutorial de lançamento do OpenRouter, explicando como testar a precisão da chamada de ferramentas do Agente de IA, dividindo as falhas em duas categorias para teste: escolha incorr...","url":"https://www.aioga.com/pt-BR/news/qu0zwrg9img0cfv03vtgmopbs/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"7569195cd17ab66e","translatedAt":"2026-09-30T03:24:07.283Z"},"ru":{"title":"Руководство по OpenRouter: как проверить точность вызова инструментов AI-агентом","summary":"OpenRouter опубликовал руководство, в котором объясняется, как тестировать точность вызова инструментов AI Agent, разделяя сбои на две категории — ошибки выбора инструмента и ошибки параметров — и тестируя их отдельно.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"Руководство по OpenRouter: как проверить точность вызова инструментов AI-агентом - Aioga Новости ИИ","description":"OpenRouter опубликовал руководство, в котором объясняется, как тестировать точность вызова инструментов AI Agent, разделяя сбои на две категории — ошибки выбора инструмента и ошибк...","url":"https://www.aioga.com/ru/news/qu0zwrg9img0cfv03vtgmopbs/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"7569195cd17ab66e","translatedAt":"2026-09-30T03:24:08.830Z"},"ar":{"title":"دليل OpenRouter: كيفية اختبار دقة استدعاء الأدوات بواسطة وكيل الذكاء الاصطناعي","summary":"نشرت OpenRouter教程ًا يشرح كيفية اختبار دقة استدعاء الأدوات لدى AI Agent، مع تقسيم حالات الفشل إلى فئتين واختبار كل منهما على حدة: أخطاء اختيار الأداة وأخطاء المعلمات.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"دليل OpenRouter: كيفية اختبار دقة استدعاء الأدوات بواسطة وكيل الذكاء الاصطناعي - Aioga أخبار الذكاء الاصطناعي","description":"نشرت OpenRouter教程ًا يشرح كيفية اختبار دقة استدعاء الأدوات لدى AI Agent، مع تقسيم حالات الفشل إلى فئتين واختبار كل منهما على حدة: أخطاء اختيار الأداة وأخطاء المعلمات.","url":"https://www.aioga.com/ar/news/qu0zwrg9img0cfv03vtgmopbs/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"7569195cd17ab66e","translatedAt":"2026-09-30T03:24:02.775Z"},"hi":{"title":"OpenRouter ट्यूटोरियल: AI Agent के टूल कॉल करने की सटीकता का परीक्षण कैसे करें","summary":"OpenRouter ने एक ट्यूटोरियल जारी किया, जिसमें बताया गया कि AI एजेंट के टूल कॉल की सटीकता को कैसे टेस्ट करें, और विफलताओं को टूल चयन में गलती और पैरामीटर में गलती के दो प्रकारों में विभाजित करके अलग-अलग टेस्ट किया जाए।","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"OpenRouter ट्यूटोरियल: AI Agent के टूल कॉल करने की सटीकता का परीक्षण कैसे करें - Aioga AI समाचार","description":"OpenRouter ने एक ट्यूटोरियल जारी किया, जिसमें बताया गया कि AI एजेंट के टूल कॉल की सटीकता को कैसे टेस्ट करें, और विफलताओं को टूल चयन में गलती और पैरामीटर में गलती के दो प्रकारों में...","url":"https://www.aioga.com/hi/news/qu0zwrg9img0cfv03vtgmopbs/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"7569195cd17ab66e","translatedAt":"2026-09-30T03:29:40.705Z"},"it":{"title":"Tutorial su OpenRouter: come testare l'accuratezza delle chiamate agli strumenti degli agenti AI","summary":"Tutorial di OpenRouter, che spiega come testare la precisione delle chiamate agli strumenti degli AI Agent, suddividendo i fallimenti in due categorie: errore nella scelta dello strumento ed errore nei parametri, e testandole separatamente.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"Tutorial su OpenRouter: come testare l'accuratezza delle chiamate agli strumenti degli agenti AI - Aioga Notizie IA","description":"Tutorial di OpenRouter, che spiega come testare la precisione delle chiamate agli strumenti degli AI Agent, suddividendo i fallimenti in due categorie: errore nella scelta dello st...","url":"https://www.aioga.com/it/news/qu0zwrg9img0cfv03vtgmopbs/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"7569195cd17ab66e","translatedAt":"2026-09-30T03:29:56.826Z"},"nl":{"title":"OpenRouter-handleiding: hoe de nauwkeurigheid van AI-agent toolopdrachten te testen","summary":"OpenRouter publiceert een handleiding waarin wordt uitgelegd hoe de nauwkeurigheid van AI Agent-toolaanroepen kan worden getest, waarbij mislukkingen worden opgesplitst in twee categorieën: verkeerde toolkeuze en verkeerde parameters, die afzonderlijk worden getest.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"OpenRouter-handleiding: hoe de nauwkeurigheid van AI-agent toolopdrachten te testen - Aioga AI-nieuws","description":"OpenRouter publiceert een handleiding waarin wordt uitgelegd hoe de nauwkeurigheid van AI Agent-toolaanroepen kan worden getest, waarbij mislukkingen worden opgesplitst in twee cat...","url":"https://www.aioga.com/nl/news/qu0zwrg9img0cfv03vtgmopbs/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"7569195cd17ab66e","translatedAt":"2026-09-30T03:29:52.442Z"},"tr":{"title":"OpenRouter Eğitimi: AI Agent'ın araç çağrılarının doğruluğu nasıl test edilir","summary":"OpenRouter, yapay zekâ ajanlarının araç çağırma doğruluğunun nasıl test edileceğini anlatan bir eğitim yayımladı. Eğitimde başarısızlıklar araç seçimi hataları ve parametre hataları olarak ikiye ayrılarak ayrı ayrı test ediliyor.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"OpenRouter Eğitimi: AI Agent'ın araç çağrılarının doğruluğu nasıl test edilir - Aioga AI Haberleri","description":"OpenRouter, yapay zekâ ajanlarının araç çağırma doğruluğunun nasıl test edileceğini anlatan bir eğitim yayımladı. Eğitimde başarısızlıklar araç seçimi hataları ve parametre hatalar...","url":"https://www.aioga.com/tr/news/qu0zwrg9img0cfv03vtgmopbs/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"7569195cd17ab66e","translatedAt":"2026-09-30T03:35:14.392Z"},"vi":{"title":"Hướng dẫn OpenRouter: Cách kiểm tra độ chính xác khi gọi công cụ của AI Agent","summary":"OpenRouter đã phát hành hướng dẫn giải thích cách kiểm tra độ chính xác khi AI Agent gọi công cụ, chia lỗi thành hai loại để kiểm tra riêng: lỗi chọn công cụ và lỗi tham số.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"Hướng dẫn OpenRouter: Cách kiểm tra độ chính xác khi gọi công cụ của AI Agent - Tin tức AI Aioga","description":"OpenRouter đã phát hành hướng dẫn giải thích cách kiểm tra độ chính xác khi AI Agent gọi công cụ, chia lỗi thành hai loại để kiểm tra riêng: lỗi chọn công cụ và lỗi tham số.","url":"https://www.aioga.com/vi/news/qu0zwrg9img0cfv03vtgmopbs/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"7569195cd17ab66e","translatedAt":"2026-09-30T03:35:20.240Z"},"id":{"title":"Tutorial OpenRouter: Cara Menguji Akurasi Pemanggilan Alat oleh AI Agent","summary":"OpenRouter menerbitkan tutorial yang menjelaskan cara menguji keakuratan pemanggilan alat oleh AI Agent, dengan membagi kegagalan menjadi dua kategori, yaitu kesalahan pemilihan alat dan kesalahan parameter, untuk diuji secara terpisah.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"Tutorial OpenRouter: Cara Menguji Akurasi Pemanggilan Alat oleh AI Agent - Berita AI Aioga","description":"OpenRouter menerbitkan tutorial yang menjelaskan cara menguji keakuratan pemanggilan alat oleh AI Agent, dengan membagi kegagalan menjadi dua kategori, yaitu kesalahan pemilihan al...","url":"https://www.aioga.com/id/news/qu0zwrg9img0cfv03vtgmopbs/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"7569195cd17ab66e","translatedAt":"2026-09-30T03:35:07.072Z"},"th":{"title":"บทช่วยสอน OpenRouter: วิธีทดสอบความแม่นยำในการเรียกใช้เครื่องมือของ AI Agent","summary":"OpenRouter เผยแพร่บทเรียน อธิบายวิธีการทดสอบความถูกต้องในการเรียกใช้งานเครื่องมือของ AI Agent โดยจะแยกความล้มเหลวออกเป็นสองประเภท คือ การเลือกเครื่องมือผิดและพารามิเตอร์ผิด เพื่อทดสอบแยกกัน","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"บทช่วยสอน OpenRouter: วิธีทดสอบความแม่นยำในการเรียกใช้เครื่องมือของ AI Agent - ข่าว AI Aioga","description":"OpenRouter เผยแพร่บทเรียน อธิบายวิธีการทดสอบความถูกต้องในการเรียกใช้งานเครื่องมือของ AI Agent โดยจะแยกความล้มเหลวออกเป็นสองประเภท คือ การเลือกเครื่องมือผิดและพารามิเตอร์ผิด เพื่อทด...","url":"https://www.aioga.com/th/news/qu0zwrg9img0cfv03vtgmopbs/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"7569195cd17ab66e","translatedAt":"2026-09-30T03:40:59.376Z"},"pl":{"title":"Samouczek OpenRouter: jak testować dokładność wywoływania narzędzi przez agentów AI","summary":"OpenRouter opublikował poradnik wyjaśniający, jak testować dokładność wywoływania narzędzi przez agentów AI, dzieląc błędy na dwa typy: błędy wyboru narzędzia i błędy parametrów, które należy testować osobno.","category":"行业动态","source":"OpenRouter：Announcements（RSS）","aggregationSource":"OpenRouter：Announcements（RSS）","pageTitle":"Samouczek OpenRouter: jak testować dokładność wywoływania narzędzi przez agentów AI - Aioga Wiadomości AI","description":"OpenRouter opublikował poradnik wyjaśniający, jak testować dokładność wywoływania narzędzi przez agentów AI, dzieląc błędy na dwa typy: błędy wyboru narzędzia i błędy parametrów, k...","url":"https://www.aioga.com/pl/news/qu0zwrg9img0cfv03vtgmopbs/","contentTranslated":true,"translationStatus":"translated","translationRetryAt":"","translationError":"","sourceHash":"7569195cd17ab66e","translatedAt":"2026-09-30T03:40:50.225Z"}},"evidenceTier":"verified-news","reviewStatus":"editorial-selected","indexable":true,"editorialCover":""}}