{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-23T07:21:26.498Z","headline":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","description":"Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","url":"https://www.aioga.com/news/cmrrhq17q004xbihkjmdnpwtt/","mainEntityOfPage":"https://www.aioga.com/news/cmrrhq17q004xbihkjmdnpwtt/","datePublished":"2026-07-19T07:19:14.000Z","dateModified":"2026-07-19T07:19:14.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.marktechpost.com/2026/07/19/perplexity-ai-releases-wandr-an-open-benchmark-evaluating-research-agents-that-must-search-wide-and-deep","https://aihot.virxact.com/items/cmrrhq17q004xbihkjmdnpwtt"],"canonicalUrl":"https://www.aioga.com/news/cmrrhq17q004xbihkjmdnpwtt/","directAnswer":{"@type":"Answer","text":"Aioga 编辑摘要：Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。 Aioga 将其归入「论文研究」方向，重点关注它对真实使用和行业竞争的影响。","url":"https://www.aioga.com/news/cmrrhq17q004xbihkjmdnpwtt/","dateCreated":"2026-07-19T07:19:14.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"marktechpost.com source article","url":"https://www.marktechpost.com/2026/07/19/perplexity-ai-releases-wandr-an-open-benchmark-evaluating-research-agents-that-must-search-wide-and-deep","datePublished":"2026-07-19T07:19:14.000Z","provider":{"@type":"Organization","name":"marktechpost.com","url":"https://www.marktechpost.com/2026/07/19/perplexity-ai-releases-wandr-an-open-benchmark-evaluating-research-agents-that-must-search-wide-and-deep"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmrrhq17q004xbihkjmdnpwtt","datePublished":"2026-07-19T07:19:14.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmrrhq17q004xbihkjmdnpwtt"}}],"aggregationSource":"MarkTechPost（RSS）","originalPublisher":{"name":"marktechpost.com","url":"https://www.marktechpost.com/2026/07/19/perplexity-ai-releases-wandr-an-open-benchmark-evaluating-research-agents-that-must-search-wide-and-deep"},"article":{"id":"cmrrhq17q004xbihkjmdnpwtt","slug":"cmrrhq17q004xbihkjmdnpwtt","url":"https://www.aioga.com/news/cmrrhq17q004xbihkjmdnpwtt/","title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","title_en":"Perplexity AI Releases WANDR： An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep","summary":"Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","source":"MarkTechPost（RSS）","sourceUrl":"https://www.marktechpost.com/2026/07/19/perplexity-ai-releases-wandr-an-open-benchmark-evaluating-research-agents-that-must-search-wide-and-deep","aiHotUrl":"https://aihot.virxact.com/items/cmrrhq17q004xbihkjmdnpwtt","publishedAt":"2026-07-19T07:19:14.000Z","category":"论文研究","score":37,"selected":false,"articleBody":["Research agents already handle real knowledge work today. Teams delegate competitive mapping, due diligence, and literature review to them. However, most benchmarks test a single answer, not large evidence-backed collections. Perplexity targets that gap with a new open benchmark.","Perplexity released WANDR (Wide ANd Deep Research)：https://github.com/perplexityai/wandr. It is an open benchmark and evaluation harness. It is built around 500 realistic, challenging data-collection tasks for knowledge work. WANDR is the wide sibling of Perplexity’s DRACO benchmark for deep research. DRACO asks whether an agent produces an accurate, complete, objective long-form report. WANDR instead asks whether it can build a large collection with evidence.","At its core, WANDR tests two demands together. Wide means discovering a large, often open-ended set of qualifying entities. Deep means investigating every entity enough to support each claim with evidence. Combining both changes the problem for agents. A few compelling examples are not enough here. A polished narrative built on incomplete research also falls short.","To capture this, WANDR uses a composable qualification key hierarchy . One task might request company(n) -> employee(m) -> url(k) . This means n qualifying companies, m employees each, and k supporting pages each. Every complete path through the tree gets validated independently. The same structure can represent a flat list, nested search, or matrix.","To ground that hierarchy, consider the released ceo_cfo_appointments task. It asks for at least 70 US-based companies. Each must have a CEO or CFO appointment first announced between March 1 and April 30, 2026. For each, the agent supplies one authoritative appointment page. A subtask adds a listing-authority page per company. Together, the task requires 140 source-backed records.","Concretely, the two hierarchies and one submitted record look like this:","Beyond single examples, WANDR builds its tasks from real usage. It starts from de-identified patterns seen in production, not synthetic prompts. A semi-automated pipeline then turns those patterns into tasks. The pipeline runs four stages: seeding, authoring, admission, and curation. It uses an interleaved author-critic loop with mechanical linting.","As a result, the median task asks for 50 members and 245 records overall. Across all 500 tasks, WANDR calls for 170,495 source-backed records. Tasks split into 167 lower, 166 middle, and 167 higher difficulty. Difficulty depends on per-record work, not scale alone.","Unlike fixed answer keys, WANDR grades each claim against cited evidence. Every record contains an item, URL, selected excerpts, and answer. The grader re-fetches the page during evaluation. It checks whether the page is usable and in scope. It then verifies the excerpts truly appear and support every requirement.","These binary record verdicts then roll up through the hierarchy. Precision measures the quality of what a system submitted. Recall measures quality-adjusted completion, filling any shortfall with zeros. Soft scores give partial credit to incomplete members. Hard scores count only members whose full subtree is correct.","Using that method, Perplexity ran six production systems on all 500 tasks. Its own Search as Code (SaC) system leads. Still, no system comes close to solving the benchmark.","With more effort, Perplexity reaches 0.447 soft F1 at the xhigh setting. Cost across settings spans more than four orders of magnitude. It ranges from $0.03 per task up to $324.83 per task.","Beyond the leaderboard, four findings stand out. First, partial progress is common, but complete coverage is not. Every system shows soft recall below soft precision. Second, scale compounds the problem sharply. Deeper hierarchies hurt most, since each branch adds a failure point. Third, discovery is the first structural bottleneck. Top-level discovery completion ranges from 0.611 to 0.951 across systems. Under-delivery, not duplicate merging, explains most missing volume. Fourth, finding a usable page is usually easy. Turning it into complete evidence is the hard part. For Perplexity, 41.4% of pages miss a substantive requirement. Also, 57.5% of excerpts fail to support the full claim. Its soft F1 falls from 0.531 under a retrieval-only check to 0.363 under the full verdict.","Notably, Search as Code fits this task shape well. An agent can express retrieval, filtering, fan-out, joins, deduplication, and stopping logic as a program. Deterministic compute then handles repeated operations outside the model context.","Practically, WANDR maps to jobs teams already automate. A market analyst needs every qualifying competitor, with matching evidence for each. A due-diligence team needs dozens of companies, then ownership, executives, and financing. Talent sourcing needs many candidates, each with supporting profile pages. WANDR tests exactly these wide-and-deep collection patterns at professional scale.","Because grading is per-record, teams can localize failures precisely. The score tree isolates loss to discovery, enrichment, or evidence extraction. This diagnosis helps engineers improve one weak stage at a time.","Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us ：https://forms.gle/wbash1wF6efRj8G58","Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.","Build an Agentic Event Venue Operator [Full Codes]：https://pxllnk.co/twdn5","Thanks! Our team will contact you soon"],"articleImages":[{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2019/06/Screen-Shot-2021-09-14-at-9.02.24-AM-300x300.png","alt":"","afterParagraph":16,"url":"/media/articles/cmrrhq17q004xbihkjmdnpwtt/787a6d54564e8e19.webp"},{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2026/07/blog19132-35-100x70.png","alt":"10 Open-Source No-Code Platforms for Building LLM Apps, RAG Systems, and AI Agents","afterParagraph":17,"url":"/media/articles/cmrrhq17q004xbihkjmdnpwtt/330c02f0ca218ba6.webp"},{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2026/07/blog19132-34-100x70.png","alt":"Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2","afterParagraph":17,"url":"/media/articles/cmrrhq17q004xbihkjmdnpwtt/a80f941097ee6a50.webp"},{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2026/07/blog19132-33-100x70.png","alt":"Fine-Tuning Qwen3 with LoRA Using NVIDIA NeMo AutoModel","afterParagraph":17,"url":"/media/articles/cmrrhq17q004xbihkjmdnpwtt/6fe6f7b6366a31dc.webp"},{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2026/07/blog19132-32-100x70.png","alt":"NVIDIA Released DeepStream 9.1","afterParagraph":17,"url":"/media/articles/cmrrhq17q004xbihkjmdnpwtt/1e950c77753bb01b.webp"}],"mediaStatus":"ok","articleBodyZh":["研究代理如今已经处理真正的知识工作。各队将竞争性地图绘制、尽职调查和文献综述委托给他们。然而，大多数基准测试的是单一答案，而非大规模的证据支持的收集。Perplexity通过新的开放基准针对这一差距。","Perplexity 发布了 WANDR（广泛 ANd 深度研究）：https：//github.com/perplexityai/wandr。它是一个开放的基准和评估工具。它围绕500个现实且具有挑战性的数据收集任务构建，用于知识工作。WANDR是Perplexity的DRACO深度研究基准的“广义兄弟”。DRACO询问代理人是否能产出准确、完整、客观的长篇报告。WANDR则询问是否能建立一个有证据的大型收藏。","从核心上讲，WANDR共同测试了两个需求。广义意味着发现一大类且通常开放的合格实体。深度意味着对每个实体进行足够的调查，以提供证据支持每个说法。两者结合后，代理的问题就不同了。仅凭几个有说服力的例子是不够的。建立在不完整研究基础上的精致叙述同样不足。","为了捕捉这一点，WANDR使用可组合的资格密钥层级结构。一个任务可能会请求公司（n） -> employee（m） -> url（k）。这意味着n家合格公司，每家有m名员工，每家有k页支持。树中每一条完整路径都会被独立验证。相同的结构可以表示平面列表、嵌套搜索或矩阵。","为了建立这种层级结构，考虑已释放的ceo_cfo_appointments任务。它要求至少有70家美国公司。每位职位必须在2026年3月1日至4月30日期间首次公布首席执行官或首席财务官任命。每个人，代理会提供一个权威的预约页面。子任务为每家公司添加一个列表权威页面。该任务合计需要140条源代码支持的记录。","具体来说，这两个层级和一份提交的记录如下：","除了单个示例，WANDR的任务还基于真实使用。它从制作中出现的去识别模式开始，而非合成提示。半自动化的管道将这些模式转化为任务。流程分为四个阶段：播种、创作、录取和策展。它采用了交错的作者-评论循环和机械式的线条。","因此，中位任务需要 50 名成员和总共 245 条记录。在所有 500 个任务中，WANDR 需要 170,495 条有来源支持的记录。任务分为 167 个低难度、166 个中等难度和 167 个高难度。难度取决于每条记录的工作量，而不仅仅是规模。","与固定答案键不同，WANDR 根据引用的证据对每个声明进行评分。每条记录包含一个条目、URL、选择的摘录和答案。评分者在评估过程中会重新获取页面。它会检查页面是否可用且在范围内，然后验证摘录是否真实出现并支持每个要求。","这些二进制记录判定随后在层级中汇总。精确度衡量系统提交内容的质量。召回率衡量质量调整后的完成度，任何不足部分以零计入。软评分为不完整的成员给予部分积分。硬评分仅计算其完整子树正确的成员。","使用这种方法，Perplexity 在所有 500 个任务上运行了六个生产系统。其自有的“代码搜索”（Search as Code，SaC）系统领先。然而，没有任何系统接近解决该基准。","经过更多努力，Perplexity 在 xhigh 设置下达到 0.447 的软 F1 分数。各设置的成本跨度超过四个数量级，从每个任务 0.03 美元到 324.83 美元不等。","在排行榜之外，有四个发现尤为突出。首先，部分进展很常见，但完全覆盖很少。每个系统的软召回均低于软精度。其次，规模会大幅加剧问题。更深的层级最为不利，因为每个分支都会增加失败点。第三，发现是第一个结构性瓶颈。各系统顶层发现完成率为 0.611 到 0.951。遗漏的数量主要由交付不足而非重复合并导致。第四，找到可用页面通常很容易，将其转化为完整证据才是难点。对于 Perplexity，41.4% 的页面缺少实质性要求。此外，57.5% 的摘录未能支持完整声明。在仅检索检查下，其软 F1 为 0.531，而在完整判定下降至 0.363。","值得注意的是，“代码搜索”非常适合该任务结构。代理可以将检索、筛选、分支、连接、去重和停止逻辑表现为程序，然后由确定性计算处理模型上下文之外的重复操作。","实际上，WANDR 对应于团队已经自动化的工作。市场分析师需要每一个符合条件的竞争对手，并为每一个提供匹配的证据。尽职调查团队需要数十家公司，然后是所有权、管理层和融资信息。人才招聘需要许多候选人，每个候选人都有支持的资料页面。WANDR 在专业规模上正是测试这些广而深的收集模式。","因为评分是按记录进行的，团队可以精确定位失败的地方。评分树将损失隔离到发现、补充或证据提取环节。这种诊断帮助工程师一次改进一个薄弱环节。","需要与我们合作推广您的 GitHub 仓库或 Hugging Face 页面或产品发布或网络研讨会等吗？请与我们联系：https://forms.gle/wbash1wF6efRj8G58","Asif Razzaq 是 Marktechpost Media Inc. 的首席执行官。作为一位具有远见的企业家和工程师，Asif 致力于利用人工智能的潜力造福社会。他最近的努力是推出人工智能媒体平台 Marktechpost，该平台以深入覆盖机器学习和深度学习新闻而脱颖而出，这些新闻既有技术深度，又易于广大受众理解。该平台拥有超过每月两百万的浏览量，显示其在受众中的受欢迎程度。","构建一个有代理功能的活动场地运营商 [完整代码]：https://pxllnk.co/twdn5","谢谢！我们的团队将很快与您联系"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Aioga 编辑摘要：Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。 Aioga 将其归入「论文研究」方向，重点关注它对真实使用和行业竞争的影响。","background":"背景分析：模型与研究类动态需要结合能力边界、开放方式、成本、可用性和真实任务表现判断，单项指标领先不等于已经形成稳定采用。","viewpoint":"Aioga 判断：这条动态更适合作为行业观察信号，当前信息足以建立线索，但不足以推导长期结论。","implications":"影响分析：对相关团队而言，短期应先核对来源、可用范围和实际成本，再判断是否值得接入或跟进。","nextStep":"后续观察：继续观察官方文档、实际可用性、价格变化、开发者反馈和竞品回应。","evidenceRefs":["title","summary","articleBody"],"confidence":"medium","status":"published","aiGenerated":false,"autoApproved":true,"generatedBy":"rule-safe-fallback","generatedAt":"2026-07-23T07:30:40.558Z","sourceHash":"c85a2522e0d8df26","validation":{"passed":true,"mode":"rule-safe-fallback","checks":["schema","length","source-attribution","no-html"]}},"tags":["论文研究","MarkTechPost（RSS）"],"translations":{"zh-CN":{"title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","summary":"Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"论文研究","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力 - Aioga AI资讯","description":"Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","url":"https://www.aioga.com/news/cmrrhq17q004xbihkjmdnpwtt/"},"en":{"title":"Perplexity AI Releases WANDR： An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep","summary":"Aioga tracks this update from MarkTechPost（RSS） under Research. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"Research","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity AI Releases WANDR： An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep - Aioga AI News","description":"Aioga tracks this update from MarkTechPost（RSS） under Research. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准","url":"https://www.aioga.com/en/news/cmrrhq17q004xbihkjmdnpwtt/"},"ja":{"title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","summary":"Aiogaは「論文研究」の動きとして、MarkTechPost（RSS） からの更新を追跡しています。Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"論文研究","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力 - Aioga AIニュース","description":"Aiogaは「論文研究」の動きとして、MarkTechPost（RSS） からの更新を追跡しています。Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","url":"https://www.aioga.com/ja/news/cmrrhq17q004xbihkjmdnpwtt/"},"ko":{"title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","summary":"Aioga는 MarkTechPost（RSS）의 업데이트를 연구 흐름으로 추적합니다. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"연구","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力 - Aioga AI 뉴스","description":"Aioga는 MarkTechPost（RSS）의 업데이트를 연구 흐름으로 추적합니다. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","url":"https://www.aioga.com/ko/news/cmrrhq17q004xbihkjmdnpwtt/"},"es":{"title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","summary":"Aioga sigue esta actualización de MarkTechPost（RSS） dentro de Investigación. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"Investigación","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力 - Aioga Noticias de IA","description":"Aioga sigue esta actualización de MarkTechPost（RSS） dentro de Investigación. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，","url":"https://www.aioga.com/es/news/cmrrhq17q004xbihkjmdnpwtt/"},"fr":{"title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","summary":"Aioga suit cette mise à jour de MarkTechPost（RSS） dans la catégorie Recherche. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"Recherche","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力 - Aioga Actualités IA","description":"Aioga suit cette mise à jour de MarkTechPost（RSS） dans la catégorie Recherche. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领","url":"https://www.aioga.com/fr/news/cmrrhq17q004xbihkjmdnpwtt/"},"de":{"title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","summary":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"论文研究","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力 - Aioga KI-News","description":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分","url":"https://www.aioga.com/de/news/cmrrhq17q004xbihkjmdnpwtt/"},"pt-BR":{"title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","summary":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"论文研究","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力 - Aioga Notícias de IA","description":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分","url":"https://www.aioga.com/pt-BR/news/cmrrhq17q004xbihkjmdnpwtt/"},"ru":{"title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","summary":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"论文研究","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力 - Aioga Новости ИИ","description":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分","url":"https://www.aioga.com/ru/news/cmrrhq17q004xbihkjmdnpwtt/"},"ar":{"title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","summary":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"论文研究","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力 - Aioga أخبار الذكاء الاصطناعي","description":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分","url":"https://www.aioga.com/ar/news/cmrrhq17q004xbihkjmdnpwtt/"},"hi":{"title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","summary":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"论文研究","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力 - Aioga AI समाचार","description":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分","url":"https://www.aioga.com/hi/news/cmrrhq17q004xbihkjmdnpwtt/"},"it":{"title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","summary":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"论文研究","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力 - Aioga Notizie IA","description":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分","url":"https://www.aioga.com/it/news/cmrrhq17q004xbihkjmdnpwtt/"},"nl":{"title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","summary":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"论文研究","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力 - Aioga AI-nieuws","description":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分","url":"https://www.aioga.com/nl/news/cmrrhq17q004xbihkjmdnpwtt/"},"tr":{"title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","summary":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"论文研究","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力 - Aioga AI Haberleri","description":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分","url":"https://www.aioga.com/tr/news/cmrrhq17q004xbihkjmdnpwtt/"},"vi":{"title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","summary":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"论文研究","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力 - Tin tức AI Aioga","description":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分","url":"https://www.aioga.com/vi/news/cmrrhq17q004xbihkjmdnpwtt/"},"id":{"title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","summary":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"论文研究","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力 - Berita AI Aioga","description":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分","url":"https://www.aioga.com/id/news/cmrrhq17q004xbihkjmdnpwtt/"},"th":{"title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","summary":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"论文研究","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力 - ข่าว AI Aioga","description":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分","url":"https://www.aioga.com/th/news/cmrrhq17q004xbihkjmdnpwtt/"},"pl":{"title":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力","summary":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分仅 0.447。","category":"论文研究","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perplexity 发布开放基准 WANDR，评估智能体的广泛搜索与深度验证能力 - Aioga Wiadomości AI","description":"Aioga tracks this update from MarkTechPost（RSS） under 论文研究. Perplexity 发布开放基准 WANDR，包含 500 个真实数据收集任务，要求智能体同时完成广泛发现与深度验证。其自有的 Search as Code 系统以 0.363 的 Soft F1 领先，但所有系统均远未解决该基准，最高分","url":"https://www.aioga.com/pl/news/cmrrhq17q004xbihkjmdnpwtt/"}}}}