美团LongCat推出LoHoSearch,一个基于762万实体维基百科知识图谱自动生成问题的搜索智能体基准,旨在解决BrowseComp等现有基准趋于饱和的问题。
在11个前沿模型测试中,最佳得分仅34.74%,远低于当前模型在BrowseComp上约90%的成绩;
上下文策略仅带来+6.8个百分点的提升。
该基准包含544道问题、11个领域,采用树与图结构,已开源。
Meituan LongCat has launched LoHoSearch, a search agent benchmark that automatically generates questions based on a 7.62 million-entity Wikipedia knowledge...
Meituan LongCat has launched LoHoSearch, a search agent benchmark that automatically generates
questions based on a 7.62 million-entity Wikipedia knowledge graph, aiming to address the saturation problem of existing benchmarks like BrowseComp. In tests with 11 cutting-edge models, the highest score was only 34.74%, far below the current models’ approximately 90% performance on BrowseComp; contextual strategies only brought an improvement of 6.8 percentage points. This benchmark includes 544 questions across 11 domains, uses tree and graph structures, and has been open-sourced.
美团LongCat推出LoHoSearch,一个基于762万实体维基百科知识图谱自动生成问题的搜索智能体基准,旨在解决BrowseComp等现有基准趋于饱和的问题。
在11个前沿模型测试中,最佳得分仅34.74%,远低于当前模型在BrowseComp上约90%的成绩;
上下文策略仅带来+6.8个百分点的提升。
该基准包含544道问题、11个领域,采用树与图结构,已开源。
美团LongCat发布搜索智能体基准LoHoSearch。该基准基于含762万实体的维基百科知识图谱自动生成问题,共收录544道题,覆盖11个领域,并采用树与图结构,目前已开源。
发布方称,LoHoSearch旨在应对BrowseComp等现有基准趋于饱和的问题。材料显示,当前模型在BrowseComp上的成绩约为90%,因此需要更具挑战性的搜索智能体评测。
Aioga判断,LoHoSearch的主要价值在于提供一组更难的搜索任务,并通过知识图谱及树、图结构组织问题。其能否成为通用评测标准,仍需更多公开测试与使用反馈验证。
在对11个前沿模型的测试中,最佳得分仅为34.74%;材料同时称,上下文策略只带来6.8个百分点提升。值得关注的是,该结果显示参测模型在这套题目上仍有较大提升空间。 后续可重点关注开源基准的题目生成与评测设置、不同模型在544道题和11个领域中的表现,以及更多测试能否复现最佳得分和上下文策略提升幅度。
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.
Ingestion channel: Summary aggregation · Source domain: x.com
Source: X:美团 LongCat (@Meituan_LongCat)
Original link: Open original source
Aioga archive: Open intelligence page
Content record: social-summary · Updated: 2026-07-17T14:08:29.000Z

统一接入主流 AI 模型 API,为开发、测试与生产环境提供稳定调用入口。
立即访问 api.w173.comAioga aggregates global AI updates and preserves source information for verification and citation.