今天,我们在 LangSmith 推出了自定义应用程序:https://www.langchain.com/langsmith-platform,这使您可以在 LangSmith 数据之上构建、发布和运行自定义界面,并直接在 LangSmith 内运行它们。
数千个 AI 工程团队使用 LangSmith 来调试、评估和改进他们的智能代理。LangSmith 为这些工作流程提供了默认界面,基于我们看到的最常见用户需求,包括延迟、成本和错误率的预构建仪表板,以及突出显示实验回归的比较视图。但我们也发现,团队常常需要以符合自身质量标准和审核流程的方式与数据交互。
一些团队已经通过在 LangSmith 数据周围构建自定义前端来解决这个问题。他们使用编码代理、LangSmith API 和内部工具来创建数据的自定义界面。但这些应用程序也像任何其他软件一样需要持续维护,包括托管、身份验证、权限和共享。
现在,自定义应用程序让您可以用您的 LangSmith 数据构建所需的界面,将其发布到您的工作区,并跳过托管、身份验证和权限工作。
尝试自定义应用:https://docs.langchain.com/langsmith/custom-apps
使用自定义应用程序,您可以从聊天提示开始或通过编码构建。在 Plus 和 Enterprise 计划中,告诉 LangSmith Chat 您想查看哪些数据以及希望如何可视化,Chat 会将其转换为可运行的应用程序。如果您想自己构建,可以从提供的模板开始,并使用编码代理基于 LangSmith API 构建。
一旦发布,应用程序将成为工作区内的共享界面。您无需导出数据或即兴重建相同的静态视图,可以在重复的工作流程中重用应用程序,并与同事在他们已经查看追踪、实验、注释和评估结果的同一位置共享。
当您需要重复相同的审核流程时,自定义应用程序特别有用。以下是一些常见的使用方式。
人工反馈是改进 AI 系统最重要的输入之一。自动评估可以帮助您发现模式,但您仍然需要人来审核输出、判断行为,并解释对您的应用来说什么是好的表现。
那种审核通常取决于谁在执行工作。工程师可能需要完整的跟踪记录,包括工具调用、元数据和中间步骤。主题专家(SME)可能只需要用户请求、模型响应、相关上下文以及明确的评分标准。
使用自定义应用程序,您可以创建与审核者工作流程匹配的注释界面。您可以向审核者提供他们所需的上下文,隐藏不必要的细节,并引导他们给出您希望收集的特定反馈。
实验有助于评估您对代理进行的更改是否达到了预期效果。在发布更改之前,您通常需要审核对应用程序至关重要的单个结果、分段或失败模式。
这可能意味着比较不同提示版本的输出,检查模型更换是否引入了回归,或按客户细分、失败类别、模型提供者或内部评估标准对结果进行切片。
使用自定义应用程序,您可以围绕这些决策创建实验审核界面。这意味着您可以突出显示最重要的切片,展示代表性示例,对比输出,或将结果打包成发布审核期间使用的格式。
当实验审核是一个重复过程时,这尤其有用。您无需每次都重新构建相同的静态图表或自定义分析,而是可以创建一个与 LangSmith 数据保持连接的共享应用,并在每个审核周期重复使用。
您的跟踪审核过程通常取决于您运行的代理类型。对于客户支持代理,您可能关注语气以及代理是否解决了问题,这只能通过阅读交流来判断。另一方面,研究代理可能根据其来源和所需的工具调用次数来评判。
当审核按计划进行时,自定义应用程序可以让您围绕最常提出的问题创建更集中的用户界面,例如行为变化的位置、哪个步骤导致了失败,或应用程序是否遵循了预期流程。
自定义应用程序适用于 Plus 和企业版计划:
有关一般使用和计费信息,请参阅我们的定价页面:https://www.langchain.com/pricing。
您可以今天尝试自定义应用程序,通过登录或注册:https://smith.langchain.com/ 访问 LangSmith,并查看文档:https://docs.langchain.com/langsmith/custom-apps#custom-apps 了解更多细节。
LangSmith,我们的代理工程平台,帮助开发者调试每一个代理决策,评估变化,并一键部署。
Today, we’re launching Custom Apps in LangSmith:https://www.langchain.com/langsmith-platform, which lets you build, publish, and run custom UIs on top of LangSmith data and run them directly inside LangSmith.
Thousands of AI engineering teams use LangSmith to debug, evaluate, and improve their agents. LangSmith provides default UIs for those workflows based on the most common user requirements we've seen, including prebuilt dashboards for latency, cost, and error rates, and a comparison view that highlights regressions across experiments. But we've seen that teams often need to interact with data in ways that align with their own quality criteria and review processes.
Some teams already solve this by building custom frontends around LangSmith data. They use coding agents, LangSmith APIs, and internal tools to create custom interfaces of their data. But these applications also require ongoing maintenance just like any other piece of software, which includes hosting, authentication, permissions, and sharing.
Now, Custom Apps lets you build the interface you want with your LangSmith data, publish it into your workspace, and skip the hosting, auth, and permissions work.
Try Custom Apps:https://docs.langchain.com/langsmith/custom-apps
With Custom Apps, you can start with a chat prompt or build in code. On Plus and Enterprise plans, tell LangSmith Chat what data you want to see and how you want it visualized, and Chat turns that into a working app. If you'd rather build it yourself, start with provided templates and build against the LangSmith API with your coding agent.
Once published, the app becomes a shared interface inside the workspace. Instead of exporting data or rebuilding the same static view ad hoc, you can reuse the app for recurring workflows and share it with teammates in the same place they already review traces, experiments, annotations, and evaluation results.
Custom Apps are especially useful when you need to repeat the same review process. Here are a few common ways to use them.
Human feedback is one of the most important inputs for improving AI systems. Automated evals can help you find patterns, but you still need people to review outputs, judge behavior, and explain what good looks like for your application.
That review often depends on who is doing the work. An engineer may need the full trace, including tool calls, metadata, and intermediate steps. A subject matter expert (SME) may only need the user request, the model response, relevant context, and a clear rubric.
With Custom Apps, you can create annotation interfaces that match the reviewer’s workflows. You can give reviewers the context they need, hide the details they don’t, and guide them through the specific feedback you want to collect.
Experiments help you assess whether changes to your agent achieved the intended outcome. Before shipping a change, you usually need to review individual results, segments, or failure patterns that matter to your application.
That might mean comparing outputs across prompt versions, checking whether a model swap introduced regressions, or slicing results by customer segment, failure category, model provider, or internal evaluation rubric.
With Custom Apps, you can create experiment review interfaces around those decisions. That means you can highlight the slices that matter most, show representative examples, compare outputs side by side, or package results into the format you use during release reviews.
This is especially useful when experiment review is a recurring process. Instead of rebuilding the same static chart or custom analysis each time, you can create a shared app that stays connected to LangSmith data and reuse it for every review cycle.
Your trace review process typically follows the type of agent you're running. For a customer support agent, you might focus on tone and whether the agent resolved the issue, which you can only judge by reading the exchange. On the other hand, a research agent may be judged on its sources and the number of tool calls required.
When that review happens on a schedule, Custom Apps let you create a more focused UI around the questions you ask most often, such as where behavior changed, which step caused a failure, or whether the application followed the expected process.
Custom Apps are available for Plus and Enterprise plans:
For general usage and billing information, see our pricing page:https://www.langchain.com/pricing.
You can try Custom Apps today by logging in or signing up:https://smith.langchain.com/ for LangSmith, and visit the docs:https://docs.langchain.com/langsmith/custom-apps#custom-apps for more detail.
LangSmith, our agent engineering platform, helps developers debug every agent decision, eval changes, and deploy in one click.