{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-28T06:20:51.496Z","headline":"LangChain 发布 Deep Agents 新基准评测方案","description":"LangChain 重新设计了 Deep Agents 的基准评测方式，并在 Harbor 平台上运行覆盖编码、对话和检索三类任务的评估流程。该评测方案用于指导 Deep Agents 的版本迭代与变更发布。","url":"https://www.aioga.com/news/cms3dpwfe0apero3flnym5kgh/","mainEntityOfPage":"https://www.aioga.com/news/cms3dpwfe0apero3flnym5kgh/","datePublished":"2026-07-24T02:31:15.000Z","dateModified":"2026-07-24T02:31:15.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.langchain.com/blog/how-we-benchmark-deep-agents","https://aihot.virxact.com/items/cms3dpwfe0apero3flnym5kgh"],"canonicalUrl":"https://www.aioga.com/news/cms3dpwfe0apero3flnym5kgh/","directAnswer":{"@type":"Answer","text":"Aioga 编辑摘要：LangChain 重新设计了 Deep Agents 的基准评测方式，并在 Harbor 平台上运行覆盖编码、对话和检索三类任务的评估流程。 Aioga 将其归入「技巧观点」方向，重点关注它对真实使用和行业竞争的影响。","url":"https://www.aioga.com/news/cms3dpwfe0apero3flnym5kgh/","dateCreated":"2026-07-24T02:31:15.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"langchain.com source article","url":"https://www.langchain.com/blog/how-we-benchmark-deep-agents","datePublished":"2026-07-24T02:31:15.000Z","provider":{"@type":"Organization","name":"langchain.com","url":"https://www.langchain.com/blog/how-we-benchmark-deep-agents"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cms3dpwfe0apero3flnym5kgh","datePublished":"2026-07-24T02:31:15.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cms3dpwfe0apero3flnym5kgh"}}],"aggregationSource":"LangChain：Blog（RSS）","originalPublisher":{"name":"langchain.com","url":"https://www.langchain.com/blog/how-we-benchmark-deep-agents"},"article":{"id":"cms3dpwfe0apero3flnym5kgh","slug":"cms3dpwfe0apero3flnym5kgh","url":"https://www.aioga.com/news/cms3dpwfe0apero3flnym5kgh/","title":"LangChain 发布 Deep Agents 新基准评测方案","title_en":"How We Benchmark Deep Agents","summary":"LangChain 重新设计了 Deep Agents 的基准评测方式，并在 Harbor 平台上运行覆盖编码、对话和检索三类任务的评估流程。该评测方案用于指导 Deep Agents 的版本迭代与变更发布。","source":"LangChain：Blog（RSS）","sourceUrl":"https://www.langchain.com/blog/how-we-benchmark-deep-agents","aiHotUrl":"https://aihot.virxact.com/items/cms3dpwfe0apero3flnym5kgh","publishedAt":"2026-07-24T02:31:15.000Z","category":"技巧观点","score":47,"selected":false,"articleBody":["Agent design is hard, in no small part because evaluation of agents is hard. As we develop Deep Agents - our open source, model agnostic agent harness - we are constantly faced with many decisions: how to prompt, which tools to include, which middleware to include, etc. As we iterate, we want to have a robust evaluation set to benchmark our decisions against. We recently revamped our evaluation framework, and in this blog we will share how we think about evaluating Deep Agents.","Our recent evaluation work has centered around identifying a set of end-to-end evals. Previously, we had smaller, “unit”-style tests. We still have those, but as we’ve seen agent tasks get longer and longer running, we’ve moved to more end-to-end evals.","In order to accomplish this, we used Harbor：https://www.langchain.com/blog/unified-stack-for-evaluating-agents as an eval runner. Harbor is a popular open source framework for running agent evals, most well known for powering Terminal Bench (a leading coding benchmark).","To use Harbor, you provide three things:","Each dataset：https://www.harborframework.com/docs/core-concepts#dataset has tasks：https://www.harborframework.com/docs/core-concepts#task, which consist of:","Compared to simpler LLM evaluation, there are two main differences:","Today we run three benchmarks, each covering a distinct kind of agent work. Deep Agents is a general purpose harness, so we need to benchmark its capability across multiple different domains.","This is just the start - we will grow these sets of tasks over time.","We use these benchmarks to iterate with confidence. When we are making decisions, benchmarking changes allows to have confidence we are moving in the right direction.","A concrete example of this is how we used it to prepare for a 0.7 release of Deep Agents. As part of this release, we are looking to slim down the harness and remove prompting that may have once been necessary but no longer is. This pays off for someone running Deep Agents: fewer tokens per run and more of the model's attention on the instructions they wrote.","Two concrete changes we are considering: removing the todo-list middleware, and significantly slimming down the system prompt. We are using this benchmark to help decide whether those inclusions are still necessary for the agent harness or not.","LangSmith, our agent engineering platform, helps developers debug every agent decision, eval changes, and deploy in one click."],"articleImages":[],"mediaStatus":"none","articleBodyZh":["代理设计很难，很大程度上是因为代理的评估也很难。当我们开发 Deep Agents——我们的开源、模型无关的代理框架——时，我们不断面临许多决策：如何提示、包含哪些工具、包含哪些中间件等。在迭代过程中，我们希望拥有一套稳健的评估集合来衡量我们的决策。我们最近重新设计了评估框架，在这篇博客中，我们将分享我们如何看待 Deep Agents 的评估。","我们最近的评估工作集中在确定一组端到端评估。之前，我们有较小的“单元”式测试。我们仍然保留这些测试，但随着代理任务运行时间越来越长，我们转向了更多的端到端评估。","为此，我们使用了 Harbor：https://www.langchain.com/blog/unified-stack-for-evaluating-agents 作为评估运行器。Harbor 是一个流行的开源框架，用于运行代理评估，以支持 Terminal Bench（一个领先的编码基准）而闻名。","使用 Harbor 时，你需要提供三个内容：","每个数据集（https://www.harborframework.com/docs/core-concepts#dataset）都有任务（https://www.harborframework.com/docs/core-concepts#task），其组成部分包括：","与简单的 LLM 评估相比，有两个主要区别：","目前我们运行三个基准测试，每个测试涵盖不同类型的代理工作。Deep Agents 是一个通用框架，因此我们需要在多个不同领域中基准其能力。","这只是一个开始——我们会随着时间增长这些任务集合。","我们使用这些基准测试来自信地进行迭代。当我们做出决策时，基准测试的变化让我们有信心朝着正确的方向前进。","一个具体的例子是我们如何使用它为 Deep Agents 0.7 发布做准备。作为该版本的一部分，我们希望精简框架，并移除曾经必要但现在不再需要的提示。这对使用 Deep Agents 的人来说是有好处的：每次运行的 token 数更少，模型可以将更多注意力集中在他们编写的指令上。","我们正在考虑的两个具体变化：移除待办事项中间件，以及显著精简系统提示。我们正在使用这个基准来帮助决定这些包含项对于代理工具是否仍然必要。","LangSmith，我们的代理工程平台，可以帮助开发者调试每个代理决策、评估变更，并一键部署。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Aioga 编辑摘要：LangChain 重新设计了 Deep Agents 的基准评测方式，并在 Harbor 平台上运行覆盖编码、对话和检索三类任务的评估流程。 Aioga 将其归入「技巧观点」方向，重点关注它对真实使用和行业竞争的影响。","background":"背景分析：模型与研究类动态需要结合能力边界、开放方式、成本、可用性和真实任务表现判断，单项指标领先不等于已经形成稳定采用。","viewpoint":"Aioga 判断：这条动态更适合作为行业观察信号，当前信息足以建立线索，但不足以推导长期结论。","implications":"影响分析：对相关团队而言，短期应先核对来源、可用范围和实际成本，再判断是否值得接入或跟进。","nextStep":"后续观察：继续观察官方文档、实际可用性、价格变化、开发者反馈和竞品回应。","evidenceRefs":["title","summary","articleBody"],"confidence":"medium","status":"published","aiGenerated":false,"autoApproved":true,"generatedBy":"rule-safe-fallback","generatedAt":"2026-07-28T06:29:10.560Z","sourceHash":"95c8d73c01e0e85b","validation":{"passed":true,"mode":"rule-safe-fallback","checks":["schema","length","source-attribution","no-html"]}},"tags":["技巧观点","LangChain：Blog（RSS）"],"translations":{"zh-CN":{"title":"LangChain 发布 Deep Agents 新基准评测方案","summary":"LangChain 重新设计了 Deep Agents 的基准评测方式，并在 Harbor 平台上运行覆盖编码、对话和检索三类任务的评估流程。该评测方案用于指导 Deep Agents 的版本迭代与变更发布。","category":"技巧观点","source":"langchain.com","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangChain 发布 Deep Agents 新基准评测方案 - Aioga AI资讯","description":"LangChain 重新设计了 Deep Agents 的基准评测方式，并在 Harbor 平台上运行覆盖编码、对话和检索三类任务的评估流程。该评测方案用于指导 Deep Agents 的版本迭代与变更发布。","url":"https://www.aioga.com/news/cms3dpwfe0apero3flnym5kgh/"},"en":{"title":"LangChain releases new benchmark evaluation scheme for Deep Agents","summary":"LangChain has redesigned the benchmark evaluation method for Deep Agents and runs an assessment process covering encoding, dialogue, and retrieval tasks on the Harbor platform. This evaluation scheme is used to guide the version iteration and change release of Deep Agents.","category":"Insights","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangChain releases new benchmark evaluation scheme for Deep Agents - Aioga AI News","description":"LangChain has redesigned the benchmark evaluation method for Deep Agents and runs an assessment process covering encoding, dialogue, and retrieval tasks on the Harbor platform. Thi...","url":"https://www.aioga.com/en/news/cms3dpwfe0apero3flnym5kgh/","contentTranslated":true,"sourceHash":"9cffaedc8a232801","translatedAt":"2026-07-27T16:03:54.159Z"},"ja":{"title":"LangChain、Deep Agentsの新しいベンチマーク評価プランを発表","summary":"LangChain は Deep Agents のベンチマーク評価方法を再設計し、Harbor プラットフォーム上でコード、対話、検索の三種類のタスクをカバーする評価プロセスを実行しました。この評価プランは、Deep Agents のバージョンの反復更新および変更のリリースを指導するために使用されます。","category":"ヒントと視点","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangChain、Deep Agentsの新しいベンチマーク評価プランを発表 - Aioga AIニュース","description":"LangChain は Deep Agents のベンチマーク評価方法を再設計し、Harbor プラットフォーム上でコード、対話、検索の三種類のタスクをカバーする評価プロセスを実行しました。この評価プランは、Deep Agents のバージョンの反復更新および変更のリリースを指導するために使用されます。","url":"https://www.aioga.com/ja/news/cms3dpwfe0apero3flnym5kgh/","contentTranslated":true,"sourceHash":"9cffaedc8a232801","translatedAt":"2026-07-27T16:04:10.360Z"},"ko":{"title":"LangChain, Deep Agents 새로운 벤치마크 평가 계획 발표","summary":"LangChain은 Deep Agents의 기준 평가 방식을 재설계했으며, Harbor 플랫폼에서 코드 커버리지, 대화 및 검색의 세 가지 유형의 작업을 포함하는 평가 프로세스를 실행합니다. 이 평가 방안은 Deep Agents의 버전 반복 및 변경 출시를 안내하는 데 사용됩니다.","category":"인사이트","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangChain, Deep Agents 새로운 벤치마크 평가 계획 발표 - Aioga AI 뉴스","description":"LangChain은 Deep Agents의 기준 평가 방식을 재설계했으며, Harbor 플랫폼에서 코드 커버리지, 대화 및 검색의 세 가지 유형의 작업을 포함하는 평가 프로세스를 실행합니다. 이 평가 방안은 Deep Agents의 버전 반복 및 변경 출시를 안내하는 데 사용됩니다.","url":"https://www.aioga.com/ko/news/cms3dpwfe0apero3flnym5kgh/","contentTranslated":true,"sourceHash":"9cffaedc8a232801","translatedAt":"2026-07-27T16:05:05.425Z"},"es":{"title":"LangChain lanza un nuevo esquema de evaluación de referencia para Agentes Profundos","summary":"LangChain rediseñó la forma de evaluación de referencia de Deep Agents y ejecuta en la plataforma Harbor procesos de evaluación que cubren tres tipos de tareas: codificación, diálogo y recuperación. Este esquema de evaluación se utiliza para guiar la iteración de versiones y la publicación de cambios de Deep Agents.","category":"Ideas","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangChain lanza un nuevo esquema de evaluación de referencia para Agentes Profundos - Aioga Noticias de IA","description":"LangChain rediseñó la forma de evaluación de referencia de Deep Agents y ejecuta en la plataforma Harbor procesos de evaluación que cubren tres tipos de tareas: codificación, diálo...","url":"https://www.aioga.com/es/news/cms3dpwfe0apero3flnym5kgh/","contentTranslated":true,"sourceHash":"9cffaedc8a232801","translatedAt":"2026-07-27T16:04:49.411Z"},"fr":{"title":"LangChain publie un nouveau protocole d'évaluation de référence pour les Deep Agents","summary":"LangChain a repensé la méthode d'évaluation de référence des Deep Agents et a exécuté sur la plateforme Harbor un processus d'évaluation couvrant trois types de tâches : le codage, la conversation et la récupération. Ce schéma d'évaluation est utilisé pour guider les itérations de versions et les publications de modifications des Deep Agents.","category":"Analyses","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangChain publie un nouveau protocole d'évaluation de référence pour les Deep Agents - Aioga Actualités IA","description":"LangChain a repensé la méthode d'évaluation de référence des Deep Agents et a exécuté sur la plateforme Harbor un processus d'évaluation couvrant trois types de tâches : le codage,...","url":"https://www.aioga.com/fr/news/cms3dpwfe0apero3flnym5kgh/","contentTranslated":true,"sourceHash":"9cffaedc8a232801","translatedAt":"2026-07-27T16:05:52.724Z"},"de":{"title":"LangChain veröffentlicht neuen Benchmark-Testplan für Deep Agents","summary":"LangChain hat die Benchmark-Methode für Deep Agents neu gestaltet und den Bewertungsprozess, der die drei Arten von Aufgaben Abdeckung von Codierung, Dialog und Abruf umfasst, auf der Harbor-Plattform durchgeführt. Dieses Bewertungsschema wird verwendet, um die Versionsiteration und die Veröffentlichung von Änderungen für Deep Agents zu leiten.","category":"技巧观点","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangChain veröffentlicht neuen Benchmark-Testplan für Deep Agents - Aioga KI-News","description":"LangChain hat die Benchmark-Methode für Deep Agents neu gestaltet und den Bewertungsprozess, der die drei Arten von Aufgaben Abdeckung von Codierung, Dialog und Abruf umfasst, auf...","url":"https://www.aioga.com/de/news/cms3dpwfe0apero3flnym5kgh/","contentTranslated":true,"sourceHash":"9cffaedc8a232801","translatedAt":"2026-07-27T16:05:50.117Z"},"pt-BR":{"title":"LangChain lança novo método de avaliação de referência para Agentes Profundos","summary":"LangChain redesenhou a forma de avaliação de referência do Deep Agents e executa no plataforma Harbor um fluxo de avaliação que cobre três tipos de tarefas: codificação, diálogo e recuperação. Esse plano de avaliação é usado para orientar a iteração de versões e o lançamento de alterações do Deep Agents.","category":"技巧观点","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangChain lança novo método de avaliação de referência para Agentes Profundos - Aioga Notícias de IA","description":"LangChain redesenhou a forma de avaliação de referência do Deep Agents e executa no plataforma Harbor um fluxo de avaliação que cobre três tipos de tarefas: codificação, diálogo e...","url":"https://www.aioga.com/pt-BR/news/cms3dpwfe0apero3flnym5kgh/","contentTranslated":true,"sourceHash":"9cffaedc8a232801","translatedAt":"2026-07-27T16:06:30.613Z"},"ru":{"title":"LangChain выпустила новую базовую схему оценки Deep Agents","summary":"LangChain переработала базовый метод оценки Deep Agents и запустила процесс оценки на платформе Harbor, охватывающий три типа задач: кодирование, диалог и поиск. Эта схема оценки используется для руководства по итерации версий и выпуску изменений Deep Agents.","category":"技巧观点","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangChain выпустила новую базовую схему оценки Deep Agents - Aioga Новости ИИ","description":"LangChain переработала базовый метод оценки Deep Agents и запустила процесс оценки на платформе Harbor, охватывающий три типа задач: кодирование, диалог и поиск. Эта схема оценки и...","url":"https://www.aioga.com/ru/news/cms3dpwfe0apero3flnym5kgh/","contentTranslated":true,"sourceHash":"9cffaedc8a232801","translatedAt":"2026-07-27T16:06:36.506Z"},"ar":{"title":"أطلقت LangChain برنامج تقييم معياري جديد للوكلاء الأذكياء","summary":"قام LangChain بإعادة تصميم طريقة التقييم الأساسية لوكلاء Deep، وشغّل على منصة Harbor عملية تقييم تغطي ثلاث فئات من المهام: الترميز، الحوار، والاسترجاع. تُستخدم خطة التقييم هذه لتوجيه تحديثات الإصدارات وإطلاق التغييرات لوكلاء Deep.","category":"技巧观点","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"أطلقت LangChain برنامج تقييم معياري جديد للوكلاء الأذكياء - Aioga أخبار الذكاء الاصطناعي","description":"قام LangChain بإعادة تصميم طريقة التقييم الأساسية لوكلاء Deep، وشغّل على منصة Harbor عملية تقييم تغطي ثلاث فئات من المهام: الترميز، الحوار، والاسترجاع. تُستخدم خطة التقييم هذه لتوج...","url":"https://www.aioga.com/ar/news/cms3dpwfe0apero3flnym5kgh/","contentTranslated":true,"sourceHash":"9cffaedc8a232801","translatedAt":"2026-07-27T16:07:24.208Z"},"hi":{"title":"LangChain ने Deep Agents के लिए नया बेंचमार्क मूल्यांकन योजना जारी की","summary":"LangChain ने Deep Agents के मानक मूल्यांकन तरीके को फिर से डिज़ाइन किया है, और Harbor प्लेटफ़ॉर्म पर कोडिंग, संवाद और पुनःप्राप्ति के तीन प्रकार के कार्यों को कवर करने वाली मूल्यांकन प्रक्रिया चलाई है। यह मूल्यांकन योजना Deep Agents के संस्करण पुनरावृत्ति और परिवर्तन रिलीज़ को मार्गदर्शन करने के लिए उपयोग की जाती है।","category":"技巧观点","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangChain ने Deep Agents के लिए नया बेंचमार्क मूल्यांकन योजना जारी की - Aioga AI समाचार","description":"LangChain ने Deep Agents के मानक मूल्यांकन तरीके को फिर से डिज़ाइन किया है, और Harbor प्लेटफ़ॉर्म पर कोडिंग, संवाद और पुनःप्राप्ति के तीन प्रकार के कार्यों को कवर करने वाली मूल्यां...","url":"https://www.aioga.com/hi/news/cms3dpwfe0apero3flnym5kgh/","contentTranslated":true,"sourceHash":"9cffaedc8a232801","translatedAt":"2026-07-27T16:07:33.854Z"},"it":{"title":"LangChain lancia un nuovo programma di valutazione standard per Deep Agents","summary":"LangChain ha riprogettato il metodo di valutazione di riferimento dei Deep Agents e ha eseguito sul piattaforma Harbor il processo di valutazione che copre tre tipi di compiti: codifica, conversazione e ricerca. Questo schema di valutazione viene utilizzato per guidare l'iterazione delle versioni e il rilascio delle modifiche dei Deep Agents.","category":"技巧观点","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangChain lancia un nuovo programma di valutazione standard per Deep Agents - Aioga Notizie IA","description":"LangChain ha riprogettato il metodo di valutazione di riferimento dei Deep Agents e ha eseguito sul piattaforma Harbor il processo di valutazione che copre tre tipi di compiti: cod...","url":"https://www.aioga.com/it/news/cms3dpwfe0apero3flnym5kgh/","contentTranslated":true,"sourceHash":"9cffaedc8a232801","translatedAt":"2026-07-27T16:08:14.283Z"},"nl":{"title":"LangChain lanceert nieuw benchmark evaluatiesysteem voor Deep Agents","summary":"LangChain heeft de benchmark-evaluatiemethode van Deep Agents opnieuw ontworpen en voert op het Harbor-platform evaluatieprocessen uit die drie soorten taken omvatten: codering, dialoog en informatieopvraging. Deze evaluatiemethode wordt gebruikt om de versie-updates en wijzigingsreleases van Deep Agents te begeleiden.","category":"技巧观点","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangChain lanceert nieuw benchmark evaluatiesysteem voor Deep Agents - Aioga AI-nieuws","description":"LangChain heeft de benchmark-evaluatiemethode van Deep Agents opnieuw ontworpen en voert op het Harbor-platform evaluatieprocessen uit die drie soorten taken omvatten: codering, di...","url":"https://www.aioga.com/nl/news/cms3dpwfe0apero3flnym5kgh/","contentTranslated":true,"sourceHash":"9cffaedc8a232801","translatedAt":"2026-07-27T16:08:09.669Z"},"tr":{"title":"LangChain, Deep Agents için yeni bir kıyaslama değerlendirme planı yayınladı","summary":"LangChain, Deep Agents için standart değerlendirme yöntemini yeniden tasarladı ve Harbor platformunda kodlama, diyalog ve arama olmak üzere üç tür görevi kapsayan değerlendirme sürecini yürüttü. Bu değerlendirme planı, Deep Agents'ın sürüm iterasyonu ve değişiklik yayınlarını yönlendirmek için kullanılır.","category":"技巧观点","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangChain, Deep Agents için yeni bir kıyaslama değerlendirme planı yayınladı - Aioga AI Haberleri","description":"LangChain, Deep Agents için standart değerlendirme yöntemini yeniden tasarladı ve Harbor platformunda kodlama, diyalog ve arama olmak üzere üç tür görevi kapsayan değerlendirme sür...","url":"https://www.aioga.com/tr/news/cms3dpwfe0apero3flnym5kgh/","contentTranslated":true,"sourceHash":"9cffaedc8a232801","translatedAt":"2026-07-27T16:09:02.998Z"},"vi":{"title":"LangChain ra mắt phương án đánh giá chuẩn mới Deep Agents","summary":"LangChain đã thiết kế lại phương pháp đánh giá chuẩn của Deep Agents và triển khai quy trình đánh giá bao quát ba loại nhiệm vụ là mã hóa, đối thoại và truy xuất trên nền tảng Harbor. Phương án đánh giá này được dùng để hướng dẫn việc lặp phiên bản và phát hành thay đổi của Deep Agents.","category":"技巧观点","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangChain ra mắt phương án đánh giá chuẩn mới Deep Agents - Tin tức AI Aioga","description":"LangChain đã thiết kế lại phương pháp đánh giá chuẩn của Deep Agents và triển khai quy trình đánh giá bao quát ba loại nhiệm vụ là mã hóa, đối thoại và truy xuất trên nền tảng Harb...","url":"https://www.aioga.com/vi/news/cms3dpwfe0apero3flnym5kgh/","contentTranslated":true,"sourceHash":"9cffaedc8a232801","translatedAt":"2026-07-27T16:08:53.149Z"},"id":{"title":"LangChain merilis skema penilaian benchmark baru Deep Agents","summary":"LangChain merancang ulang metode evaluasi standar untuk Deep Agents, dan menjalankan proses evaluasi yang mencakup tiga jenis tugas: pengkodean, dialog, dan pengambilan informasi di platform Harbor. Skema evaluasi ini digunakan untuk memandu iterasi versi dan rilis perubahan Deep Agents.","category":"技巧观点","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangChain merilis skema penilaian benchmark baru Deep Agents - Berita AI Aioga","description":"LangChain merancang ulang metode evaluasi standar untuk Deep Agents, dan menjalankan proses evaluasi yang mencakup tiga jenis tugas: pengkodean, dialog, dan pengambilan informasi d...","url":"https://www.aioga.com/id/news/cms3dpwfe0apero3flnym5kgh/","contentTranslated":true,"sourceHash":"9cffaedc8a232801","translatedAt":"2026-07-27T16:09:40.635Z"},"th":{"title":"LangChain เปิดตัวแผนการประเมินมาตรฐานใหม่สำหรับ Deep Agents","summary":"LangChain ได้ออกแบบวิธีการประเมินมาตรฐานของ Deep Agents ใหม่ และรันกระบวนการประเมินครอบคลุมงานสามประเภท ได้แก่ การเข้ารหัส การสนทนา และการค้นหา บนแพลตฟอร์ม Harbor แผนการประเมินนี้ใช้สำหรับชี้แนะแก่การอัปเดตเวอร์ชันและการเผยแพร่การเปลี่ยนแปลงของ Deep Agents","category":"技巧观点","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangChain เปิดตัวแผนการประเมินมาตรฐานใหม่สำหรับ Deep Agents - ข่าว AI Aioga","description":"LangChain ได้ออกแบบวิธีการประเมินมาตรฐานของ Deep Agents ใหม่ และรันกระบวนการประเมินครอบคลุมงานสามประเภท ได้แก่ การเข้ารหัส การสนทนา และการค้นหา บนแพลตฟอร์ม Harbor แผนการประเมินนี้ใ...","url":"https://www.aioga.com/th/news/cms3dpwfe0apero3flnym5kgh/","contentTranslated":true,"sourceHash":"9cffaedc8a232801","translatedAt":"2026-07-27T16:09:50.672Z"},"pl":{"title":"LangChain wprowadza nowy standard testowy Deep Agents","summary":"LangChain przeprojektował sposób przeprowadzania standardowej oceny Deep Agents i uruchamia na platformie Harbor proces oceny obejmujący trzy rodzaje zadań: kodowanie, dialog i wyszukiwanie. Ten plan oceny służy do kierowania iteracjami wersji i wprowadzaniem zmian w Deep Agents.","category":"技巧观点","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangChain wprowadza nowy standard testowy Deep Agents - Aioga Wiadomości AI","description":"LangChain przeprojektował sposób przeprowadzania standardowej oceny Deep Agents i uruchamia na platformie Harbor proces oceny obejmujący trzy rodzaje zadań: kodowanie, dialog i wys...","url":"https://www.aioga.com/pl/news/cms3dpwfe0apero3flnym5kgh/","contentTranslated":true,"sourceHash":"9cffaedc8a232801","translatedAt":"2026-07-27T16:10:42.085Z"}}}}