调试代理通常从一个简单的问题开始:代理实际上做了什么?
对于短期的单轮工作流,答案通常很容易找到。对于运行时间较长的代理来说,难度会很快提升。单次会话可能涵盖多次用户轮流、工具调用、重试和子代理交接。完整跟踪仍然是检查执行细节的正确位置,但当你试图理解代理走过的路径时,细节可能会过于繁复。
今天,我们将在 LangSmith 推出“轨迹”:https://www.langchain.com/langsmith-platform,这是一个按时间顺序、对话式的代理会话视图。轨迹汇总了来自人类、人工智能和工具的消息,跨越主代理及所有子代理,然后按首次出现的顺序显示。
你现在可以将线程视为轨迹,使用在线评估器评分轨迹,并将其路由到注释队列或数据集。
在 LangSmith 中尝试 Trajectories:https://smith.langchain.com/
在LangSmith中,代理执行的每一个工作单元都被记录为一次运行。单次操作的运行形成一条轨迹。多回合会话的痕迹被链接成线程。
这种结构为工程师提供了完整的执行树,包括嵌套运行、时序、重试、输入、输出和元数据。当你需要准确理解某事是如何执行的时,追踪就是真相的来源。
但并非所有工作流程都需要完整的执行树,这正是轨迹发挥作用的地方。
轨迹是对线程中痕迹的投影。它去除嵌套的运行结构,保留解释代理行为的消息和动作。每个消息按顺序出现一次,因此你可以将会话解读为代理所走的路径。
对于从 LangChain:https://www.langchain.com/langchain、LangGraph:https://www.langchain.com/langgraph 和 Deep Agents:https://info.langchain.com/deep-agent、OpenAI 和 Claude 等代理 SDK 以及 Codex、Claude Code 和 Cursor 等编码代理发送的轨迹,轨迹可以开箱即用。
大多数代理调试都是从重建对话或工作流开始的。用户报告说代理给出了过时的答案。你打开会话,需要找出代理偏离其预期行为的地方。代理是否误解了用户?是否重用了陈旧的上下文?是否调用了错误的工具?子代理是否返回了不良结果?最终响应是否忽略了工具输出?
答案通常在完整的跟踪记录中某个地方,但它可能埋藏在嵌套的执行细节中。一旦你知道要查看哪个运行,那些细节就是你想要的,但首先找到那个运行本身才是耗时的部分。
使用轨迹功能,你可以从会话的有序路径开始。例如,一个支持代理执行了九轮对话和六十条消息,而客户抱怨答案错误。使用轨迹视图,你可以按顺序扫描对话、工具调用和代理操作,并找到代理重用了旧工具结果而没有获取最新数据的那一轮。
从那里,他们可以跳转到底层跟踪记录查看完整的运行时细节,包括准确的工具输入、输出、时间、重试和嵌套运行结构。不要从每个执行细节开始,先找到关键的行为步骤,然后检查该步骤的跟踪记录。
负责审查代理工作的人员可能不是构建该代理的人。
医疗专家知道临床接待代理是否提出了正确的后续问题。金融服务审查员知道合规代理是否正确处理了案例。支持主管知道升级工作流是否符合政策。这些审查员需要查看查询和代理输出,但不需要解析嵌套运行、重试或执行元数据。
轨迹使这些审查工作流程更容易扩展。将轨迹路由到标注队列,让主题专家(SME)使用可读的会话视图评分行为、标记问题并提供反馈。这些反馈随后可以用于代理改进的其余工作流程;它帮助团队优化提示、更新代理逻辑、改进评估器并构建更好的数据集。
阅读一个轨迹可以告诉你在一次会话中发生了什么,但它不能告诉你相同行为是否在整个生产流量中发生。LangSmith 在线评估器现在可以对轨迹进行评分,为判断代理在一次会话中的行为提供更好的输入。
没有轨迹时,使用运行级评估器对长会话进行评分通常会随着对话逐轮累积而包含重复的上下文。这会使评估器的输入比实际需要的更大、更嘈杂。轨迹会按顺序保留每条消息一次,因此评估器可以专注于代理所采取的路径。
在线评估然后可以突出显示值得仔细查看的轨迹。筛选低分会话,将它们送交人工审核,或保存到数据集中以供将来测试。
对于从事后训练的团队来说,轨迹使生产会话更容易转化为可用示例。它们包含表示代理行为所需的信息,包括系统提示、用户消息、助手响应、工具调用和工具输出。
这使轨迹在复查最终答案之外也非常有用。一个良好的轨迹可以展示代理何时请求澄清、调用了哪些工具、如何使用结果,以及在整个会话中如何调整行为。将高质量轨迹保存到数据集中,并导出用于监督微调工作流,以生产行为作为他们希望模型复现的示例。
随着更多团队使用生产数据来改进代理,轨迹成为一个重要的原始工具。它们不仅捕捉代理的言语,还捕捉产生这些言语的行为。
轨迹现已在美国的所有套餐中可用。
今天就通过登录或注册使用轨迹:https://smith.langchain.com/ 使用 LangSmith,并查看文档:https://docs.langchain.com/langsmith/observability-concepts#trajectories 获取更多详情。
LangSmith 是我们的代理工程平台,帮助开发者调试每一个代理决策、评估变更,并一键部署。
Debugging agents often starts with a simple question: what did the agent actually do?
For short, single-turn workflows, the answer is usually easy to find. For longer-running agents, it gets harder fast. A single session can span many user turns, tool calls, retries, and subagent handoffs. The full trace is still the right place to inspect execution details, but it can be too much detail when you’re trying to understand the path the agent took.
Today, we’re launching Trajectories in LangSmith:https://www.langchain.com/langsmith-platform, a chronological, conversational view of an agent session. A Trajectory aggregates messages from humans, AI, and tools across the main agent and any subagents, then shows them in the order they first appeared.
You can now view threads as trajectories, score trajectories with online evaluators, and route them to annotation queues or datasets.
Try Trajectories in LangSmith:https://smith.langchain.com/
In LangSmith, every unit of work an agent performs is recorded as a run. Runs for a single operation form a trace. Traces from a multi-turn session are linked into a thread.
That structure gives engineers the full execution tree, including nested runs, timing, retries, inputs, outputs, and metadata. When you need to understand exactly how something executed, the trace is the source of truth.
But not every workflow needs the full execution tree, and that's where trajectories come in.
A trajectory is a projection over the traces in a thread. It removes the nested run structure and keeps the messages and actions that explain the agent’s behavior. Each message appears once, in order, so you can read the session as the path the agent took.
Trajectories work out of the box for traces sent from LangChain:https://www.langchain.com/langchain, LangGraph:https://www.langchain.com/langgraph, and Deep Agents:https://info.langchain.com/deep-agent, from agent SDKs like OpenAI and Claude, and from coding agents like Codex, Claude Code, and Cursor.
Most agent debugging starts with reconstructing the conversation or workflow. A user reports that an agent gave an outdated answer. You open the session and need to figure out where the agent deviated from its expected behavior. Did the agent misunderstand the user? Did it reuse stale context? Did it call the wrong tool? Did a subagent return a bad result? Did the final response ignore a tool output?
The answer is usually somewhere in the full trace, but it can be buried inside nested execution detail. That detail is exactly what you want once you know which run to look at, but finding that run in the first place is what takes the time.
With Trajectories, you can start from the ordered path of the session. For example, a support agent runs nine turns and sixty messages and a customer complains the answer was wrong. Using the trajectory view, the you can scan the conversation, tool calls, and agent actions in order, and finds the turn where the agent reused an old tool result instead of fetching current data.
From there, they can jump into the underlying trace for the full runtime detail, including the exact tool input, output, timing, retries, and nested run structure. Instead of starting with every execution detail, first find the behavioral step that matters, then inspect the trace for that step.
The people responsible for reviewing an agent’s work and behavior may not be the people who built the agent.
A healthcare expert knows whether a clinical intake agent asked the right follow-up question. A financial services reviewer knows whether a compliance agent handled a case correctly. A support lead knows whether an escalation workflow matched policy. These reviewers need visibility into queries and an agent’s output, but don’t need to parse nested runs, retries or execution metadata.
Trajectories make those review workflows easier to scale. Route trajectories to annotation queues so subject-matter experts (SMEs) can score behavior, flag issues, and provide feedback using a readable view of the session. That feedback can then feed the rest of the agent improvement workflow; it helps teams refine prompts, update agent logic, improve evaluators, and build better datasets.
Reading one trajectory tells you what happened in one session, but it doesn't tell you whether the same behavior is happening across production traffic. LangSmith online evaluators can now score trajectories, giving them better input for judging agent behavior across a session.
Without trajectories, scoring long sessions with run-level evaluators often includes repeated context as the conversation accumulates turn-over-turn. That can make evaluator inputs larger and noisier than they need to be. A trajectory keeps each message once, in order, so the evaluator can focus on the path the agent took.
Online evals can then surface the trajectories worth a closer look. Filter for low-scoring sessions, route them to human review, or save them into datasets for future testing.
For teams working on post-training, trajectories make production sessions easier to turn into usable examples. They include the information needed to represent agent behavior, including system prompts, user messages, assistant responses, tool calls, and tool outputs.
That makes them useful for more than reviewing final answers. A good trajectory can show when the agent asked for clarification, which tools it called, how it used the results, and how it adapted across a session. Save high-quality trajectories to datasets and export them for supervised fine-tuning workflows, using production behavior as examples of what they want the model to reproduce.
As more teams use production data to improve agents, trajectories become an important primitive. They capture not just what the agent said, but the behavior that produced it.
Trajectories are available now on all plans in the US.
Try Trajectories today by logging in or signing up:https://smith.langchain.com/ for LangSmith, and visit the docs:https://docs.langchain.com/langsmith/observability-concepts#trajectories for more detail.
LangSmith, our agent engineering platform, helps developers debug every agent decision, eval changes, and deploy in one click.