我们构建了 LangSmith Engine:https://www.langchain.com/langsmith/engine,以加快修复代理的速度,这样开发者可以花更多时间构建新功能,而不是花时间处理漏洞。Engine 是一个平台内的代理,可以在代理开发生命周期的每个步骤自动处理工作:https://www.langchain.com/blog/the-agent-development-lifecycle。代理非常适合这项工作。它们擅长分析大型数据集和编写代码,因此在发现细微问题、编写修复方案以及监控回归方面表现出色。自五月推出以来,Engine 已分析超过 7000 万条跟踪记录,并诊断出数万个问题。
随着今天 v2 版本的推出,引擎现在承担更多工作,帮助用户更快地发布修复,并在问题出现在生产环境之前就进行检测。Engine v2:
引擎扫描您的生产追踪记录以寻找改进空间,从错误到未满足的用户请求。一旦识别出问题,引擎将其分类,并将相关追踪分组为单个问题。每条记录都包含用户需要采取行动的所有信息:根本原因、建议修复、用于评估数据集的真实示例以及持续监控。
引擎利用 LangSmith 收集的数据并付诸实践,帮助您更快地完成代理开发生命周期(ADLC):https://www.langchain.com/blog/the-agent-development-lifecycle,从而提升您的代理性能。
在 Engine v2 中,我们重点推动两项核心任务的改进:识别问题和提出修复方案。
在 Engine v2 中,我们引入了 Red Teaming,这是我们的主动故障排除工具,可在问题影响生产用户之前捕捉诸如幻觉和系统提示违规等问题。Red Teaming 分析您的代理生产追踪和仓库,以了解其目的和行为,然后利用这些上下文测试弱点并识别尚未在生产中显现的问题。您将获得一份相关且经过验证的问题列表,以提前解决潜在失败,从而防患于未然。
Red Teaming 目前已在私有测试版中向现有 LangSmith 部署用户开放。您可以在此申请加入测试版:https://www.langchain.com/langsmith-engine-v2-new-feature-access。
Engine v2 现在能够在生产追踪中发现那些更难检测的问题,这些问题通常在影响客户和大语言模型账单之前会逃过人工审核。即使你的代理看起来按预期工作,延迟的细微下降也可能让用户感到沮丧,而低效的代理轨迹会浪费令牌。Engine v2 会标记这些问题,以便你在它们影响业务之前解决。
每个问题都会被记录在同一个队列中,因此你有一个地方来跟踪和解决所有代理问题。
Engine 已经为它检测到的问题提供了建议的修复方案。但以往,用户需要选择将修复部署到生产环境并在实际流量上监控性能,或者进行手动审核、离线测试,然后再部署修复。未经测试直接上线有可能带来糟糕的客户体验。手动测试则会拖慢产品开发。
对于使用 LangSmith Deployment 的代理:https://www.langchain.com/langsmith/deployment,Engine 现在会在将建议的修复放入你的队列供审核之前进行验证。当 Engine v2 发现问题时,它会首先通过在 Deployment 中对代理运行出错的输入来重现故障。然后,它开始修复开发流程:提出修改方案,用相同的输入进行测试,评估结果,再调整修复方案。每次测试都会让修复方案更完善。一旦 Engine 确认修复解决了问题,它会将修复提供给用户在 LangSmith 中,用户可以点击打开 PR 以快速部署。
除了新功能外,我们还在不断提升 Engine 在核心任务上的性能。当 Engine 发现更高影响的问题并生成更有效的修复时,用户可以更快地解决代理中的问题。
我们定期改进 Engine 的框架、提示和底层模型,以提升在内部和外部基准测试上的性能。正如我们在八月份分享的:https://www.langchain.com/blog/new-in-langsmith-engine-2x-better-issue-detection,Engine 现在在检测问题上的表现提高了超过 2 倍(通过 IssueBench 测量:https://www.langchain.com/blog/issuebench-how-we-evaluate-engine),并提供的修复方案效果提升了 25%(通过 Terminal-Bench 测量:https://www.tbench.ai/)。
对于在自托管环境中运行 LangSmith、并有严格数据政策的团队,我们的下一版本自托管 LangSmith 将支持引入自有密钥(BYOK)功能用于 Engine。这使团队能够提供自己的模型 API 密钥为 Engine 提供推理支持,同时跟踪数据仍保留在您的 VPC 中。我们将在下一次自托管版本中分享更多关于 Engine BYOK 的详细信息。
Engine v2 当前在所有 LangSmith Plus 和 Enterprise 计划的 SaaS 部署中可用。自托管计划对 Engine v2 的支持即将推出。定价详情请查看:https://www.langchain.com/pricing。
如果您使用 Plus 计划,只需几次点击即可启用 Engine: https://docs.langchain.com/langsmith/engine#set-up-engine。
如果您使用 Enterprise 计划,可联系您的客户经理以开始使用。
朗史密斯新手?今天就开始吧:https://smith.langchain.com/。
LangSmith 是我们的智能体工程平台,帮助开发者调试每个智能体决策、评估更改,并一键部署。
We built LangSmith Engine:https://www.langchain.com/langsmith/engine to make fixing agents faster, so developers can spend more time building new capabilities and less time fighting bugs. Engine is an in-platform agent that automates work at each step of the agent development lifecycle:https://www.langchain.com/blog/the-agent-development-lifecycle. Agents are well-suited for this work. They excel at analyzing large datasets and writing code, making them great at spotting subtle issues, writing fixes, and monitoring for regressions. Since launching in May, Engine has analyzed more than 70M traces and diagnosed tens of thousands of issues.
With today’s v2 launch, Engine now takes on more work, helping users ship fixes faster and detect issues before they appear in production. Engine v2:
Engine scans your production traces for areas of improvement, from errors to unmet user requests. Once it has identified an issue, Engine classifies it and groups related traces into a single issue. Each record comes with all the information a user needs to act on it: a root cause, a proposed fix, ground truth examples for your evaluation datasets, and ongoing monitoring.
Engine takes the data that LangSmith collects and puts it into action, helping you move through the agent development lifecycle (ADLC):https://www.langchain.com/blog/the-agent-development-lifecycle faster to improve your agent.
With Engine v2, we focused on driving improvements in two of these core tasks: identifying issues and proposing fixes.
With Engine v2, we’re introducing Red Teaming, our proactive troubleshooting tool to catch issues like hallucinations and violations of system prompts before they affect users in production. Red Teaming analyzes your agent’s production traces and repos to understand its purpose and behavior, then uses that context to test for weaknesses and identify issues that haven’t yet surfaced in production. You get a list of relevant, verified issues to resolve that anticipate future failures, so you can head them off.
Red Teaming is available today in Private Beta to existing LangSmith Deployment users. You can apply to join the Beta here:https://www.langchain.com/langsmith-engine-v2-new-feature-access.
Engine v2 can now find harder-to-detect issues in production traces that often escape human review until they impact customers and LLM bills. Even when your agent appears to function as expected, subtle degradations in latency can frustrate users, and inefficient agent trajectories waste tokens. Engine v2 flags these issues so you can resolve them before they impact your business.
Every issue is tracked in the same queue, so you have one place to track and resolve all your agent issues.
Engine already provided proposed fixes for the issues it detects. But previously, users were presented with a choice to deploy the fix into production and monitor performance on live traffic, or perform a manual review, test offline, and then deploy the fix. Shipping without testing risks a bad customer experience. Manual tests slow down product development.
For agents using LangSmith Deployment:https://www.langchain.com/langsmith/deployment, Engine now validates its proposed fixes before putting them in your queue for review. When Engine v2 spots an issue, it first reproduces the failure by running the offending inputs against your agent in Deployment. Then, it starts the fix development process: proposing a change, testing it against the same inputs, evaluating the results, then adjusting the fix. With each test, the fix gets stronger. Once Engine confirms the fix resolves the issue, it provides the fix to the user in LangSmith, where they can open a PR with a click for quick deployment.
In addition to new capabilities, we’re continually improving Engine’s performance against its core tasks. When Engine finds higher-impact issues and produces more effective fixes, users can resolve problems in their agents faster.
We make regular improvements to Engine’s harness, prompts, and underlying models to raise performance against internal and external benchmarks. As we shared in August:https://www.langchain.com/blog/new-in-langsmith-engine-2x-better-issue-detection, Engine is now more than 2x better at detecting issues (as measured by IssueBench:https://www.langchain.com/blog/issuebench-how-we-evaluate-engine) and provides fixes that are 25% more effective (as measured by Terminal-Bench:https://www.tbench.ai/).
For teams with strict data policies that operate LangSmith in self-hosted environments, our next release of self-hosted LangSmith will support bring-your-own-key (BYOK) for Engine. This enables teams to provide their own model API key to power inference for Engine, while trace data stays within your VPC. We will share more details on BYOK for Engine with our next self-hosted release.
Engine v2 is available in SaaS deployments for all LangSmith Plus and Enterprise plans. Support for Engine v2 in Self-Hosted plans will be coming shortly. Pricing details can be found here:https://www.langchain.com/pricing.
If you’re on a Plus plan, you can enable Engine:https://docs.langchain.com/langsmith/engine#set-up-engine with just a few clicks.
If you’re on an Enterprise plan, you can reach out to your account team to get started.
New to LangSmith? Get started today:https://smith.langchain.com/.
LangSmith, our agent engineering platform, helps developers debug every agent decision, eval changes, and deploy in one click.