{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-23T07:21:26.498Z","headline":"美团LongCat发布LoHoSearch：更难搜索智能体基准","description":"美团LongCat推出LoHoSearch，一个基于762万实体维基百科知识图谱自动生成问题的搜索智能体基准，旨在解决BrowseComp等现有基准趋于饱和的问题。在11个前沿模型测试中，最佳得分仅34.74%，远低于当前模型在BrowseComp上约90%的成绩；上下文策略仅带来+6.8个百分点的提升。该基准包含544道问题、11个领域，采用树与图结构，已开源。","url":"https://www.aioga.com/news/cmrp0qw3s090ybitorwj6diwi/","mainEntityOfPage":"https://www.aioga.com/news/cmrp0qw3s090ybitorwj6diwi/","datePublished":"2026-07-17T14:08:29.000Z","dateModified":"2026-07-17T14:08:29.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://x.com/Meituan_LongCat/status/2078119654632124547","https://aihot.virxact.com/items/cmrp0qw3s090ybitorwj6diwi"],"canonicalUrl":"https://www.aioga.com/news/cmrp0qw3s090ybitorwj6diwi/","directAnswer":{"@type":"Answer","text":"美团LongCat发布搜索智能体基准LoHoSearch。该基准基于含762万实体的维基百科知识图谱自动生成问题，共收录544道题，覆盖11个领域，并采用树与图结构，目前已开源。","url":"https://www.aioga.com/news/cmrp0qw3s090ybitorwj6diwi/","dateCreated":"2026-07-17T14:08:29.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"X：美团 LongCat (@Meituan_LongCat) source article","url":"https://x.com/Meituan_LongCat/status/2078119654632124547","datePublished":"2026-07-17T14:08:29.000Z","provider":{"@type":"Organization","name":"X：美团 LongCat (@Meituan_LongCat)","url":"https://x.com/Meituan_LongCat/status/2078119654632124547"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmrp0qw3s090ybitorwj6diwi","datePublished":"2026-07-17T14:08:29.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmrp0qw3s090ybitorwj6diwi"}}],"aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","originalPublisher":{"name":"X：美团 LongCat (@Meituan_LongCat)","url":"https://x.com/Meituan_LongCat/status/2078119654632124547"},"article":{"id":"cmrp0qw3s090ybitorwj6diwi","slug":"cmrp0qw3s090ybitorwj6diwi","url":"https://www.aioga.com/news/cmrp0qw3s090ybitorwj6diwi/","title":"美团LongCat发布LoHoSearch：更难搜索智能体基准","title_en":"BrowseComp went from 30% → 90% in 10 months. Search-agent benchmarks are saturating. So we built Lo…","summary":"美团LongCat推出LoHoSearch，一个基于762万实体维基百科知识图谱自动生成问题的搜索智能体基准，旨在解决BrowseComp等现有基准趋于饱和的问题。在11个前沿模型测试中，最佳得分仅34.74%，远低于当前模型在BrowseComp上约90%的成绩；上下文策略仅带来+6.8个百分点的提升。该基准包含544道问题、11个领域，采用树与图结构，已开源。","source":"X：美团 LongCat (@Meituan_LongCat)","sourceUrl":"https://x.com/Meituan_LongCat/status/2078119654632124547","aiHotUrl":"https://aihot.virxact.com/items/cmrp0qw3s090ybitorwj6diwi","publishedAt":"2026-07-17T14:08:29.000Z","category":"论文研究","score":75,"selected":true,"articleBody":["美团LongCat推出LoHoSearch，一个基于762万实体维基百科知识图谱自动生成问题的搜索智能体基准，旨在解决BrowseComp等现有基准趋于饱和的问题。","在11个前沿模型测试中，最佳得分仅34.74%，远低于当前模型在BrowseComp上约90%的成绩；","上下文策略仅带来+6.8个百分点的提升。","该基准包含544道问题、11个领域，采用树与图结构，已开源。"],"articleImages":[],"mediaStatus":"none","articleBodyZh":[],"translationStatus":"","bodyOrigin":"social-summary","editorial":{"summary":"美团LongCat发布搜索智能体基准LoHoSearch。该基准基于含762万实体的维基百科知识图谱自动生成问题，共收录544道题，覆盖11个领域，并采用树与图结构，目前已开源。","background":"发布方称，LoHoSearch旨在应对BrowseComp等现有基准趋于饱和的问题。材料显示，当前模型在BrowseComp上的成绩约为90%，因此需要更具挑战性的搜索智能体评测。","viewpoint":"Aioga判断，LoHoSearch的主要价值在于提供一组更难的搜索任务，并通过知识图谱及树、图结构组织问题。其能否成为通用评测标准，仍需更多公开测试与使用反馈验证。","implications":"在对11个前沿模型的测试中，最佳得分仅为34.74%；材料同时称，上下文策略只带来6.8个百分点提升。值得关注的是，该结果显示参测模型在这套题目上仍有较大提升空间。","nextStep":"后续可重点关注开源基准的题目生成与评测设置、不同模型在544道题和11个领域中的表现，以及更多测试能否复现最佳得分和上下文策略提升幅度。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-07-22T21:31:43.821Z","sourceHash":"bcbcafb253c86408","review":{"approved":true,"groundedness":95,"clarity":92,"duplicationRisk":18,"blockingIssues":[],"notes":["background中“因此需要更具挑战性的搜索智能体评测”属于基于现有基准趋于饱和所作的概括性判断，而非材料直接给出的独立事实；其与来源所述研发目的基本一致，不构成阻断问题。","viewpoint已明确标注为“Aioga判断”，且对成为通用评测标准持保留态度，没有将预测冒充事实。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","low-source-overlap","no-html","independent-ai-review"]}},"tags":["论文研究","X：美团 LongCat (@Meituan_LongCat)"],"translations":{"zh-CN":{"title":"美团LongCat发布LoHoSearch：更难搜索智能体基准","summary":"美团LongCat推出LoHoSearch，一个基于762万实体维基百科知识图谱自动生成问题的搜索智能体基准，旨在解决BrowseComp等现有基准趋于饱和的问题。在11个前沿模型测试中，最佳得分仅34.74%，远低于当前模型在BrowseComp上约90%的成绩；上下文策略仅带来+6.8个百分点的提升。该基准包含544道问题、11个领域，采用树与图结构，已开源。","category":"论文研究","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"美团LongCat发布LoHoSearch：更难搜索智能体基准 - Aioga AI资讯","description":"美团LongCat推出LoHoSearch，一个基于762万实体维基百科知识图谱自动生成问题的搜索智能体基准，旨在解决BrowseComp等现有基准趋于饱和的问题。在11个前沿模型测试中，最佳得分仅34.74%，远低于当前模型在BrowseComp上约90%的成绩；上下文策略仅带来+6.8个百分点的提升。该基准包含544道问题、11个领域，采用树与图结构，已","url":"https://www.aioga.com/news/cmrp0qw3s090ybitorwj6diwi/"},"en":{"title":"Meituan LongCat Releases LoHoSearch: A Harder Benchmark for Search Agents","summary":"Meituan LongCat has launched LoHoSearch, a search agent benchmark that automatically generates questions based on a 7.62 million-entity Wikipedia knowledge graph, aiming to address the saturation problem of existing benchmarks like BrowseComp. In tests with 11 cutting-edge models, the highest score was only 34.74%, far below the current models’ approximately 90% performance on BrowseComp; contextual strategies only brought an improvement of 6.8 percentage points. This benchmark includes 544 questions across 11 domains, uses tree and graph structures, and has been open-sourced.","category":"Research","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"Meituan LongCat Releases LoHoSearch: A Harder Benchmark for Search Agents - Aioga AI News","description":"Meituan LongCat has launched LoHoSearch, a search agent benchmark that automatically generates questions based on a 7.62 million-entity Wikipedia knowledge graph, aiming to address","url":"https://www.aioga.com/en/news/cmrp0qw3s090ybitorwj6diwi/","contentTranslated":true,"sourceHash":"b31004ee5d9a3259","translatedAt":"2026-07-18T03:55:38.495Z"},"ja":{"title":"美団LongCatはLoHoSearchを発表：より難しい検索エージェントベンチマーク","summary":"美団LongCatはLoHoSearchを立ち上げました。これは762万のエンティティを持つウィキペディア知識グラフに基づき自動生成された質問で構成された検索エージェントベンチマークであり、BrowseCompなど既存のベンチマークが飽和傾向にある問題の解決を目指しています。11の最先端モデルでのテストでは、最高スコアはわずか34.74％で、現在のモデルがBrowseCompで約90％を達成しているのに比べて大幅に低く、文脈戦略によっても6.8ポイントしか向上しませんでした。このベンチマークは544問、11分野を含み、ツリーとグラフ構造を採用しており、既にオープンソース化されています。","category":"論文研究","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"美団LongCatはLoHoSearchを発表：より難しい検索エージェントベンチマーク - Aioga AIニュース","description":"美団LongCatはLoHoSearchを立ち上げました。これは762万のエンティティを持つウィキペディア知識グラフに基づき自動生成された質問で構成された検索エージェントベンチマークであり、BrowseCompなど既存のベンチマークが飽和傾向にある問題の解決を目指しています。11の最先端モデルでのテストでは、最高スコアはわずか34.74％で、現在のモデルがB","url":"https://www.aioga.com/ja/news/cmrp0qw3s090ybitorwj6diwi/","contentTranslated":true,"sourceHash":"b31004ee5d9a3259","translatedAt":"2026-07-18T03:55:41.725Z"},"ko":{"title":"메이투안 LongCat, LoHoSearch 발표: 더 어려운 검색 지능 에이전트 벤치마크","summary":"메이투안 LongCat은 LoHoSearch를 출시했습니다. 이는 762만 개 엔티티 위키백과 지식 그래프를 기반으로 자동 생성된 질문을 사용하는 검색 지능 에이전트 벤치마크로, BrowseComp 등 기존 벤치마크가 포화 상태에 이르는 문제를 해결하는 것을 목표로 합니다. 11개의 최첨단 모델 테스트에서 최고 점수는 단 34.74%로, 현재 모델이 BrowseComp에서 기록하는 약 90%의 성적에 비해 훨씬 낮습니다. 문맥 전략은 단 6.8 포인트 향상에 그쳤습니다. 이 벤치마크는 544개의 질문과 11개의 분야를 포함하며, 트리 및 그래프 구조를 사용하고 있으며, 오픈소스로 공개되어 있습니다.","category":"연구","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"메이투안 LongCat, LoHoSearch 발표: 더 어려운 검색 지능 에이전트 벤치마크 - Aioga AI 뉴스","description":"메이투안 LongCat은 LoHoSearch를 출시했습니다. 이는 762만 개 엔티티 위키백과 지식 그래프를 기반으로 자동 생성된 질문을 사용하는 검색 지능 에이전트 벤치마크로, BrowseComp 등 기존 벤치마크가 포화 상태에 이르는 문제를 해결하는 것을 목표로 합니다. 11개의 최첨단 모델 테스트에서 최고 점수는 단","url":"https://www.aioga.com/ko/news/cmrp0qw3s090ybitorwj6diwi/","contentTranslated":true,"sourceHash":"b31004ee5d9a3259","translatedAt":"2026-07-18T03:55:44.355Z"},"es":{"title":"Meituan LongCat lanza LoHoSearch: un referente de búsqueda de agentes inteligentes más desafiante","summary":"Meituan LongCat ha lanzado LoHoSearch, un referente de agentes de búsqueda basado en preguntas generadas automáticamente a partir de un gráfico de conocimiento de Wikipedia con 7,62 millones de entidades, destinado a abordar el problema de la saturación de referentes existentes como BrowseComp. En pruebas con 11 modelos de vanguardia, la puntuación más alta fue solo del 34,74%, muy por debajo del aproximadamente 90% que los modelos actuales alcanzan en BrowseComp; las estrategias de contexto solo aportaron una mejora de 6,8 puntos porcentuales. Este referente incluye 544 preguntas y 11 áreas, utiliza estructuras de árbol y grafo, y ya es de código abierto.","category":"Investigación","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"Meituan LongCat lanza LoHoSearch: un referente de búsqueda de agentes inteligentes más desafiante - Aioga Noticias de IA","description":"Meituan LongCat ha lanzado LoHoSearch, un referente de agentes de búsqueda basado en preguntas generadas automáticamente a partir de un gráfico de conocimiento de Wikipedia con 7,6","url":"https://www.aioga.com/es/news/cmrp0qw3s090ybitorwj6diwi/","contentTranslated":true,"sourceHash":"b31004ee5d9a3259","translatedAt":"2026-07-18T03:55:46.242Z"},"fr":{"title":"Meituan LongCat lance LoHoSearch : un nouveau benchmark pour les agents de recherche plus difficiles","summary":"Meituan LongCat a lancé LoHoSearch, un benchmark pour agents de recherche basé sur 7,62 millions d’entités du graphe de connaissances Wikipédia, avec des questions générées automatiquement, visant à résoudre le problème de saturation des benchmarks existants tels que BrowseComp. Lors du test de 11 modèles de pointe, le meilleur score n’a atteint que 34,74 %, bien inférieur aux environ 90 % des modèles actuels sur BrowseComp ; la stratégie contextuelle n’a apporté qu’une amélioration de 6,8 points. Ce benchmark comprend 544 questions, 11 domaines et utilise des structures arborescentes et graphiques, et il est open source.","category":"Recherche","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"Meituan LongCat lance LoHoSearch : un nouveau benchmark pour les agents de recherche plus difficiles - Aioga Actualités IA","description":"Meituan LongCat a lancé LoHoSearch, un benchmark pour agents de recherche basé sur 7,62 millions d’entités du graphe de connaissances Wikipédia, avec des questions générées automat","url":"https://www.aioga.com/fr/news/cmrp0qw3s090ybitorwj6diwi/","contentTranslated":true,"sourceHash":"b31004ee5d9a3259","translatedAt":"2026-07-18T03:55:49.218Z"},"de":{"title":"Meituan LongCat veröffentlicht LoHoSearch: einen schwierigeren Agenten-Benchmark, nach dem man suchen kann","summary":"Meituan LongCat startete LoHoSearch, einen Suchagenten-Benchmark, der auf 7,62 Millionen physischen Wikipedia-Wissensgraphen basiert und automatisch Probleme generiert, mit dem Ziel, die Überlastung bestehender Benchmarks wie BrowseComp zu adressieren. Unter den 11 Frontier-Modell-Tests lag die beste Punktzahl nur bei 34,74 %, weit unter dem aktuellen Modell-Wert von etwa 90 % auf BrowseComp; Die kontextuelle Strategie brachte nur eine Verbesserung um +6,8 Prozentpunkte. Dieser Benchmark enthält 544 Fragen und 11 Domains, verwendet einen Baum und eine Graphenstruktur und ist Open Source.","category":"论文研究","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"Meituan LongCat veröffentlicht LoHoSearch: einen schwierigeren Agenten-Benchmark, nach dem man suchen kann - Aioga KI-News","description":"Meituan LongCat startete LoHoSearch, einen Suchagenten-Benchmark, der auf 7,62 Millionen physischen Wikipedia-Wissensgraphen basiert und automatisch Probleme generiert, mit dem Zie","url":"https://www.aioga.com/de/news/cmrp0qw3s090ybitorwj6diwi/","contentTranslated":true,"sourceHash":"b31004ee5d9a3259","translatedAt":"2026-07-19T12:15:11.141Z"},"pt-BR":{"title":"Meituan LongCat lança LoHoSearch: um benchmark de agente mais difícil de buscar","summary":"A Meituan LongCat lançou o LoHoSearch, um benchmark de agentes de busca baseado em 7,62 milhões de grafos físicos de conhecimento da Wikipédia que geram automaticamente problemas, visando lidar com a saturação de benchmarks existentes como o BrowseComp. Entre os 11 testes do modelo Frontier, a melhor pontuação foi de apenas 34,74%, muito abaixo da pontuação de cerca de 90% do modelo atual no BrowseComp; A estratégia contextual trouxe apenas uma melhora de +6,8 pontos percentuais. Este benchmark contém 544 perguntas e 11 domínios, utiliza estrutura de árvore e grafo, e é de código aberto.","category":"论文研究","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"Meituan LongCat lança LoHoSearch: um benchmark de agente mais difícil de buscar - Aioga Notícias de IA","description":"A Meituan LongCat lançou o LoHoSearch, um benchmark de agentes de busca baseado em 7,62 milhões de grafos físicos de conhecimento da Wikipédia que geram automaticamente problemas, ","url":"https://www.aioga.com/pt-BR/news/cmrp0qw3s090ybitorwj6diwi/","contentTranslated":true,"sourceHash":"b31004ee5d9a3259","translatedAt":"2026-07-19T12:15:09.314Z"},"ru":{"title":"Meituan LongCat выпускает LoHoSearch: более сложный для поиска бенчмарк агентов","summary":"Meituan LongCat запустил LoHoSearch — бенчмарк поискового агента, основанный на 7,62 миллионах физических графов знаний Википедии, который автоматически генерирует задачи, с целью компенсировать перенасыщение существующих бенчмарков, таких как BrowseComp. Среди 11 тестов Frontier Model лучший балл составил всего 34,74%, что значительно ниже примерно 90% текущей модели на BrowseComp; Контекстуальная стратегия принесла только улучшение на +6,8 процентных пункта. Этот бенчмарк содержит 544 вопроса и 11 доменов, использует структуру дерева и графов и является открытым исходным кодом.","category":"论文研究","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"Meituan LongCat выпускает LoHoSearch: более сложный для поиска бенчмарк агентов - Aioga Новости ИИ","description":"Meituan LongCat запустил LoHoSearch — бенчмарк поискового агента, основанный на 7,62 миллионах физических графов знаний Википедии, который автоматически генерирует задачи, с целью ","url":"https://www.aioga.com/ru/news/cmrp0qw3s090ybitorwj6diwi/","contentTranslated":true,"sourceHash":"b31004ee5d9a3259","translatedAt":"2026-07-19T12:15:10.005Z"},"ar":{"title":"ميتوان لونغكات تصدر LoHoSearch: معيار أصعب للبحث عن وكلاء","summary":"أطلقت Meituan LongCat LoHoSearch، وهو معيار لوكيل بحث يعتمد على 7.62 مليون رسم بياني للمعرفة في ويكيبيديا ينتج مشاكل تلقائيا، بهدف معالجة تشبع المعايير الحالية مثل BrowseComp. من بين 11 اختبارا للنماذج الرائدة، كانت أفضل نتيجة فقط 34.74٪، وهي أقل بكثير من الدرجة التي بلغت حوالي 90٪ للنموذج الحالي على BrowseComp؛ الاستراتيجية السياقية جلبت تحسنا بنسبة +6.8 نقطة مئوية فقط. يحتوي هذا المعيار على 544 سؤالا و11 مجالا، ويستخدم هيكل شجرة ورسم بياني، وهو مفتوح المصدر.","category":"论文研究","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"ميتوان لونغكات تصدر LoHoSearch: معيار أصعب للبحث عن وكلاء - Aioga أخبار الذكاء الاصطناعي","description":"أطلقت Meituan LongCat LoHoSearch، وهو معيار لوكيل بحث يعتمد على 7.62 مليون رسم بياني للمعرفة في ويكيبيديا ينتج مشاكل تلقائيا، بهدف معالجة تشبع المعايير الحالية مثل BrowseComp. من ب","url":"https://www.aioga.com/ar/news/cmrp0qw3s090ybitorwj6diwi/","contentTranslated":true,"sourceHash":"b31004ee5d9a3259","translatedAt":"2026-07-19T12:15:10.899Z"},"hi":{"title":"Meituan LongCat ने LoHoSearch जारी किया: खोजने के लिए एक कठिन एजेंट बेंचमार्क","summary":"Meituan LongCat ने LoHoSearch लॉन्च किया, जो 7.62 मिलियन भौतिक विकिपीडिया ज्ञान ग्राफ़ पर आधारित एक खोज एजेंट बेंचमार्क है जो स्वचालित रूप से समस्याएं उत्पन्न करता है, जिसका उद्देश्य BrowseComp जैसे मौजूदा बेंचमार्क की संतृप्ति को संबोधित करना है। 11 फ्रंटियर मॉडल परीक्षणों में, सर्वश्रेष्ठ स्कोर केवल 34.74% था, जो ब्राउज़कॉम्प पर वर्तमान मॉडल के लगभग 90% स्कोर से बहुत कम था; प्रासंगिक रणनीति ने केवल +6.8 प्रतिशत अंक का सुधार किया। इस बेंचमार्क में 544 प्रश्न और 11 डोमेन हैं, एक ट्री और ग्राफ़ संरचना का उपयोग करता है, और यह खुला स्रोत है।","category":"论文研究","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"Meituan LongCat ने LoHoSearch जारी किया: खोजने के लिए एक कठिन एजेंट बेंचमार्क - Aioga AI समाचार","description":"Meituan LongCat ने LoHoSearch लॉन्च किया, जो 7.62 मिलियन भौतिक विकिपीडिया ज्ञान ग्राफ़ पर आधारित एक खोज एजेंट बेंचमार्क है जो स्वचालित रूप से समस्याएं उत्पन्न करता है, जिसका उद्देश","url":"https://www.aioga.com/hi/news/cmrp0qw3s090ybitorwj6diwi/","contentTranslated":true,"sourceHash":"b31004ee5d9a3259","translatedAt":"2026-07-19T12:15:12.241Z"},"it":{"title":"Meituan LongCat rilascia LoHoSearch: un benchmark per agenti più difficile da cercare","summary":"Meituan LongCat ha lanciato LoHoSearch, un benchmark per agenti di ricerca basato su 7,62 milioni di grafici fisici di conoscenza di Wikipedia che generano automaticamente problemi, con l'obiettivo di affrontare la saturazione di benchmark esistenti come BrowseComp. Tra gli 11 test dei modelli frontier, il punteggio migliore è stato solo del 34,74%, molto al di sotto del punteggio di circa il 90% del modello attuale su BrowseComp; La strategia contestuale ha portato solo un miglioramento del +6,8 punti percentuali. Questo benchmark contiene 544 domande e 11 domini, utilizza una struttura ad albero e grafico, ed è open source.","category":"论文研究","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"Meituan LongCat rilascia LoHoSearch: un benchmark per agenti più difficile da cercare - Aioga Notizie IA","description":"Meituan LongCat ha lanciato LoHoSearch, un benchmark per agenti di ricerca basato su 7,62 milioni di grafici fisici di conoscenza di Wikipedia che generano automaticamente problemi","url":"https://www.aioga.com/it/news/cmrp0qw3s090ybitorwj6diwi/","contentTranslated":true,"sourceHash":"b31004ee5d9a3259","translatedAt":"2026-07-19T12:15:09.986Z"},"nl":{"title":"Meituan LongCat brengt LoHoSearch uit: een moeilijkere agent benchmark om naar te zoeken","summary":"Meituan LongCat lanceerde LoHoSearch, een zoekagent-benchmark gebaseerd op 7,62 miljoen fysieke Wikipedia-kennisgrafieken die automatisch problemen genereren, met als doel de verzadiging van bestaande benchmarks zoals BrowseComp aan te pakken. Van de 11 frontiermodeltests was de beste score slechts 34,74%, ver onder de huidige model die ongeveer 90% op BrowseComp heeft; Contextuele strategie bracht slechts een verbetering van +6,8 procentpunt. Deze benchmark bevat 544 vragen en 11 domeinen, gebruikt een boom- en grafiekstructuur en is open source.","category":"论文研究","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"Meituan LongCat brengt LoHoSearch uit: een moeilijkere agent benchmark om naar te zoeken - Aioga AI-nieuws","description":"Meituan LongCat lanceerde LoHoSearch, een zoekagent-benchmark gebaseerd op 7,62 miljoen fysieke Wikipedia-kennisgrafieken die automatisch problemen genereren, met als doel de verza","url":"https://www.aioga.com/nl/news/cmrp0qw3s090ybitorwj6diwi/","contentTranslated":true,"sourceHash":"b31004ee5d9a3259","translatedAt":"2026-07-19T12:15:09.519Z"},"tr":{"title":"Meituan LongCat, araması daha zor bir ajan kıyaslası olan LoHoSearch'i yayınladı","summary":"Meituan LongCat, 7,62 milyon fiziksel Wikipedia bilgi grafiklerine dayanan ve otomatik olarak problem üreten LoHoSearch adlı bir arama ajanı kıyaslaması başlattı; bu kıyaslama, BrowseComp gibi mevcut kıyaslamaların doygunluğunu ele almayı amaçladı. 11 sınır model testi arasında en iyi puan sadece %34,74 idi; bu, mevcut modelin BrowseComp'taki yaklaşık %90 puanının çok altındaydı; Bağlamsal strateji sadece +6,8 puanlık bir iyileşme getirdi. Bu kıyaslama 544 soru ve 11 alan içerir, ağaç ve grafik yapısı kullanır ve açık kaynaklıdır.","category":"论文研究","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"Meituan LongCat, araması daha zor bir ajan kıyaslası olan LoHoSearch'i yayınladı - Aioga AI Haberleri","description":"Meituan LongCat, 7,62 milyon fiziksel Wikipedia bilgi grafiklerine dayanan ve otomatik olarak problem üreten LoHoSearch adlı bir arama ajanı kıyaslaması başlattı; bu kıyaslama, Bro","url":"https://www.aioga.com/tr/news/cmrp0qw3s090ybitorwj6diwi/","contentTranslated":true,"sourceHash":"b31004ee5d9a3259","translatedAt":"2026-07-19T12:15:11.879Z"},"vi":{"title":"Meituan LongCat phát hành LoHoSearch: một điểm chuẩn đại lý khó tìm kiếm hơn","summary":"Meituan LongCat đã ra mắt LoHoSearch, một điểm chuẩn của tác nhân tìm kiếm dựa trên 7,62 triệu biểu đồ tri thức vật lý của Wikipedia tự động tạo ra các vấn đề, nhằm giải quyết sự bão hòa của các điểm chuẩn hiện có như BrowseComp. Trong số 11 bài kiểm tra mô hình biên giới, điểm số tốt nhất chỉ là 34,74%, thấp hơn nhiều so với điểm số khoảng 90% của mô hình hiện tại trên BrowseComp; Chiến lược theo ngữ cảnh chỉ mang lại sự cải thiện +6,8 điểm phần trăm. Điểm chuẩn này chứa 544 câu hỏi và 11 miền, sử dụng cấu trúc cây và đồ thị, đồng thời là mã nguồn mở.","category":"论文研究","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"Meituan LongCat phát hành LoHoSearch: một điểm chuẩn đại lý khó tìm kiếm hơn - Tin tức AI Aioga","description":"Meituan LongCat đã ra mắt LoHoSearch, một điểm chuẩn của tác nhân tìm kiếm dựa trên 7,62 triệu biểu đồ tri thức vật lý của Wikipedia tự động tạo ra các vấn đề, nhằm giải quyết sự b","url":"https://www.aioga.com/vi/news/cmrp0qw3s090ybitorwj6diwi/","contentTranslated":true,"sourceHash":"b31004ee5d9a3259","translatedAt":"2026-07-19T12:15:11.438Z"},"id":{"title":"Meituan LongCat merilis LoHoSearch: tolok ukur agen yang lebih sulit untuk dicari","summary":"Meituan LongCat meluncurkan LoHoSearch, tolok ukur agen pencarian berdasarkan 7,62 juta grafik pengetahuan Wikipedia fisik yang secara otomatis menghasilkan masalah, yang bertujuan untuk mengatasi kejenuhan tolok ukur yang ada seperti BrowseComp. Di antara 11 tes model perbatasan, skor terbaik hanya 34,74%, jauh di bawah skor sekitar 90% model saat ini di BrowseComp; Strategi kontekstual hanya membawa peningkatan +6,8 poin persentase. Tolok ukur ini berisi 544 pertanyaan dan 11 domain, menggunakan struktur pohon dan grafik, dan bersifat open source.","category":"论文研究","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"Meituan LongCat merilis LoHoSearch: tolok ukur agen yang lebih sulit untuk dicari - Berita AI Aioga","description":"Meituan LongCat meluncurkan LoHoSearch, tolok ukur agen pencarian berdasarkan 7,62 juta grafik pengetahuan Wikipedia fisik yang secara otomatis menghasilkan masalah, yang bertujuan","url":"https://www.aioga.com/id/news/cmrp0qw3s090ybitorwj6diwi/","contentTranslated":true,"sourceHash":"b31004ee5d9a3259","translatedAt":"2026-07-19T12:15:09.285Z"},"th":{"title":"Meituan LongCat เปิดตัว LoHoSearch: เกณฑ์มาตรฐานตัวแทนที่ยากกว่าในการค้นหา","summary":"Meituan LongCat เปิดตัว LoHoSearch ซึ่งเป็นเกณฑ์มาตรฐานของตัวแทนการค้นหาโดยอิงจากกราฟความรู้ทางกายภาพของ Wikipedia 7.62 ล้านรายการที่สร้างปัญหาโดยอัตโนมัติโดยมีเป้าหมายเพื่อจัดการกับความอิ่มตัวของเกณฑ์มาตรฐานที่มีอยู่เช่น BrowseComp ในบรรดาการทดสอบโมเดลพรมแดน 11 รายการ คะแนนที่ดีที่สุดเพียง 34.74% ซึ่งต่ํากว่าคะแนนประมาณ 90% ของโมเดลปัจจุบันใน BrowseComp มาก กลยุทธ์ตามบริบทนํามาซึ่งการปรับปรุงจุดเปอร์เซ็นต์ +6.8 เท่านั้น เกณฑ์มาตรฐานนี้ประกอบด้วยคําถาม 544 ข้อและ 11 โดเมน ใช้โครงสร้างต้นไม้และกราฟ และเป็นโอเพ่นซอร์ส","category":"论文研究","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"Meituan LongCat เปิดตัว LoHoSearch: เกณฑ์มาตรฐานตัวแทนที่ยากกว่าในการค้นหา - ข่าว AI Aioga","description":"Meituan LongCat เปิดตัว LoHoSearch ซึ่งเป็นเกณฑ์มาตรฐานของตัวแทนการค้นหาโดยอิงจากกราฟความรู้ทางกายภาพของ Wikipedia 7.62 ล้านรายการที่สร้างปัญหาโดยอัตโนมัติโดยมีเป้าหมายเพื่อจัดการก","url":"https://www.aioga.com/th/news/cmrp0qw3s090ybitorwj6diwi/","contentTranslated":true,"sourceHash":"b31004ee5d9a3259","translatedAt":"2026-07-19T12:15:12.287Z"},"pl":{"title":"Meituan LongCat wprowadza LoHoSearch: trudniejszy benchmark do wyszukiwania agentów","summary":"Meituan LongCat uruchomił LoHoSearch, benchmark agentów wyszukiwania oparty na 7,62 miliona fizycznych wykresów wiedzy Wikipedii, które automatycznie generują problemy, mające na celu rozwiązanie przesycenia istniejących benchmarków, takich jak BrowseComp. Spośród 11 testów modeli Frontier najlepszy wynik wyniósł tylko 34,74%, znacznie poniżej około 90% obecnego modelu na BrowseComp; Strategia kontekstowa przyniosła poprawę o +6,8 punktu procentowego. Ten benchmark zawiera 544 pytania i 11 dziedzin, wykorzystuje drzewo i strukturę grafu oraz jest open source.","category":"论文研究","source":"X：美团 LongCat (@Meituan_LongCat)","aggregationSource":"X：美团 LongCat (@Meituan_LongCat)","pageTitle":"Meituan LongCat wprowadza LoHoSearch: trudniejszy benchmark do wyszukiwania agentów - Aioga Wiadomości AI","description":"Meituan LongCat uruchomił LoHoSearch, benchmark agentów wyszukiwania oparty na 7,62 miliona fizycznych wykresów wiedzy Wikipedii, które automatycznie generują problemy, mające na c","url":"https://www.aioga.com/pl/news/cmrp0qw3s090ybitorwj6diwi/","contentTranslated":true,"sourceHash":"b31004ee5d9a3259","translatedAt":"2026-07-19T12:15:11.730Z"}}}}