LangChain 发布博客,介绍 deepagents 中的 context modes 如何帮助子智能体分叉 supervisor 的上下文或以隔离方式启动,从而让多智能体工作更快、更省、更聚焦。
大多数框架支持一个子代理功能(subagents):https://www.langchain.com/blog/choosing-the-right-multi-agent-architecture,用于从监督代理生成新任务。子代理能够实现并行推理和上下文隔离:https://docs.langchain.com/oss/python/langchain/multi-agent,从而允许监督代理在不污染上下文窗口的情况下委派工作。
监督代理指定任务,子代理通常在新的上下文窗口中完成任务。这可能导致浪费:子代理可能会重新执行上下文收集操作,例如文件读取,而这些操作已经由监督代理完成。
对于子代理可以受益于监督代理上下文的情况,我们构建了分叉子代理(forked subagents):https://docs.langchain.com/oss/python/deepagents/subagents#forked-subagents。分叉子代理继承监督者的完整对话,而不是从新开始。与隔离子代理相比,分叉可能更快且更经济,因为重用监督者的对话可以利用提示缓存并减少重复工作。
将任务委派给子代理是代理管理自身上下文的一种有效方式。子代理提供上下文隔离:https://docs.langchain.com/oss/python/langchain/multi-agent,因此可以将各个任务的细节从监督代理的上下文窗口中保留。如果你想了解更多,我们已经详细写了:https://www.langchain.com/blog/choosing-the-right-multi-agent-architecture关于不同的多代理架构!
监督者:https://docs.langchain.com/oss/python/deepagents/subagents 是最通用的模式之一,大多数编码框架都采用了它。在这里,监督者维护一个计划并将工作委派给专门的子代理。例如,
监督代理通常只从子代理接收任务的结果;它们的中间推理不会被纳入其上下文窗口。然而,子代理应从监督者那里接收哪些上下文,取决于子代理的用途。
为帮助指定这一点,我们在最新版本的 deepagents 中引入了上下文模式(context modes)。上下文模式规定子代理可以从监督者接收哪些上下文。支持的值为“isolated”和“fork”。
这是 Deep Agents 中子代理的默认和预先存在的行为:https://docs.langchain.com/oss/python/deepagents/overview。子代理在新的上下文窗口中生成,只接收监督者指定的任务描述。
设置 "mode": "fork",监督者的当前状态会传播到子代理,而不是从空状态开始。这实际上是当前线程的一个分叉延续——附加了监督者编写的指令——最后展开为由监督者读取的单个工具结果。
尽管分叉子代理比独立子代理拥有更多上下文,但根据设计会尊重提示缓存。在子代理需要详细上下文以正确执行任务的情况下,分叉可以节省重复的工具调用和上下文收集。
上下文模式的正确选择取决于子代理与工作的关系。一种有用的思考方式是两种常见模式:继续监督者工作的“工人”以及独立评估的“验证者”。
工人在监督者已经收集了上下文或做出决策之后执行一项工作。例如,监督者可能检查一个错误,将其追踪到特定函数,然后委派修复的实现和测试。在孤立状态下启动工人将迫使其重新发现证据。使用分叉时,它会收到监督者的历史记录,并可以从调查中断处继续。监督者在需要完成某项工作但不必关心其达到结论过程中的中间步骤时会调用此方法。
监督者可能会用如下任务调用它:
验证者根据某些标准审查另一个代理的工作——例如,检查差异的正确性、向后兼容性和测试覆盖率。
在这种情况下,继承监督者的推理可能适得其反。验证者应该自行评估工作,而不是受监督者诊断或期望的影响。孤立模式为其提供任务和相关审查材料,而不包含先前的对话。
主管可以通过以下方式调用它:
我们之前写过关于 RubricMiddleware 的文章:https://www.langchain.com/blog/introducing-rubrics-for-deepagents,这又是一个使用独立验证器的例子!
除了工具和中间件之外,上下文模式也是你可以用来将子代理专门化到某个任务的杠杆之一。以下是一些我们认为专门化的子代理及其与上下文模式的关系:
研究人员调查一个问题并向主管返回简明答案。例如,主管可能会委派有关不熟悉的库、竞争对手或技术决策历史的不同问题。
当问题可以独立存在时,研究人员不需要主管的对话。使用 isolated 可以让其上下文集中于手头的问题。这在多个研究人员并行运行时尤其有用:为每个研究人员分支会重复主管的历史,尽管每个研究人员只需要分配给它的问题。
我们可以为子代理提供它自己的功能(如 search_engine 工具)来帮助它完成任务。
记忆代理从交互中识别应该在以后可用的信息——例如,用户偏好、架构决策或在对话中建立的约束。
这里,对话是代理需要分析的材料。使用 fork,记忆代理接收完整的交互内容,并可以决定哪些内容值得保留,而无需主管在任务中重新陈述。
因为我们希望限制记忆代理可以编辑的内容,可以通过设置它在工作中可以编辑的文件限制来专门化子代理。
deepagents 是我们正在构建的一个框架,它总结了我们与成千上万不同团队在部署代理过程中学到的经验。你可以尝试子代理上下文模式(请参见文档:https://docs.langchain.com/oss/python/deepagents/subagents#forked-subagents),以及通过安装实现更多功能:
请通过 GitHub issues:https://github.com/langchain-ai/deepagents/issues,论坛:https://forum.langchain.com/,或在 X / LinkedIn 上告诉我们你的想法!
LangSmith,我们的代理工程平台,帮助开发者调试每个代理决策、评估变化,并一键部署。
Most harnesses support a subagents:https://www.langchain.com/blog/choosing-the-right-multi-agent-architecture feature to spawn new tasks from a supervisor agent. Subagents enable parallel reasoning and context isolation:https://docs.langchain.com/oss/python/langchain/multi-agent, allowing a supervisor agent to delegate work without polluting the context window.
Supervisor agents specify the task, and subagents typically complete the task in a fresh context window. This can lead to waste: subagents may redo context-gathering operations, like file reads, already done by the supervisor.
For cases where subagents can benefit from the supervisor agent’s context, we built forked subagents:https://docs.langchain.com/oss/python/deepagents/subagents#forked-subagents. Forked subagents inherit the supervisor’s full conversation instead of starting fresh. Forking can be faster and cheaper than isolated subagents, since reusing the supervisor’s conversation takes advantage of prompt caching and reduces repeated work.
Delegating tasks to subagents is one effective way that an agent can manage its own context. Subagents provide context isolation:https://docs.langchain.com/oss/python/langchain/multi-agent, so that the details of individual tasks can be withheld from a supervisor agent’s context window. If you’re curious to learn more, we’ve written at length:https://www.langchain.com/blog/choosing-the-right-multi-agent-architecture about different multi-agent architectures!
The supervisor:https://docs.langchain.com/oss/python/deepagents/subagents is one of the most generalizable patterns, and most coding harnesses have adopted it. Here, a supervisor maintains a plan and delegates work to specialized subagents. For example,
The supervisor agent typically receives just the outcome of a task from the subagents; their intermediate reasoning is withheld from its context window. However, what context subagents should receive from the supervisor is dependent on what the subagent is used for.
To help specify this, we introduced context modes in the latest version of deepagents . Context modes specify what context subagents can receive from the supervisor. Supported values are "isolated" and "fork" .
This is the default and pre-existing behavior for subagents in Deep Agents:https://docs.langchain.com/oss/python/deepagents/overview. Subagents spawn with a fresh context window, receiving only the task description specified by the supervisor.
Set "mode": "fork" and the supervisor’s current state propagates to the subagents instead of starting it empty. This is effectively a forked continuation of the current thread— with an added directive written by the supervisor— that is finally unwound into a single tool result read by the supervisor.
Although forked subagents are seeded with more context than isolated subagents, prompt caching is respected by design. In cases where subagents require detailed context to correctly perform their tasks, forking can save repeated tool calls and context-gathering.
The right choice of context mode depends on the subagent’s relationship to the work. A useful way to think about them is with two common patterns: workers that continue the supervisor’s work, and verifiers that evaluate it independently.
A worker carries out a piece of work after the supervisor has already gathered context or made a decision. For example, the supervisor might inspect an error, trace it to a particular function, and then delegate the implementation and testing of a fix. Starting the worker in isolation would force it to rediscover its evidence. With fork , it receives the supervisor’s history and can pick up where the investigation left off. A supervisor calls this when some work needs to be done, but doesn’t necessarily care about the intermediate steps it takes to arrive to a conclusion.
The supervisor might invoke it with a task like:
A verifier reviews another agent’s work against some criteria- for example, checking a diff for correctness, backwards compatibility, and test coverage.
In this case, inheriting the supervisor’s reasoning can be counterproductive. The verifier should evaluate the work itself rather than being anchored by the supervisor’s diagnosis or expectations. isolated mode gives it the task and relevant review materials without the preceding conversation.
The supervisor might invoke it with:
We’ve previously written about RubricMiddleware:https://www.langchain.com/blog/introducing-rubrics-for-deepagents which is another instance of using an independent verifier!
Alongside things like tools and middleware, context modes are one of the levers you can use to specialize subagents to a task. Here's a few subagents we consider specialized and their relationship to context modes:
A researcher investigates a question and returns a condensed answer to the supervisor. For example, a supervisor might delegate separate questions about an unfamiliar library, a competitor, or the history of a technical decision.
When the question can stand on its own, the researcher does not need the supervisor’s conversation. Using isolated keeps its context focused on the question at hand. This is especially useful when several researchers run in parallel: forking each one would duplicate the supervisor’s history even though each researcher only needs its assigned question.
We can give the subagent its own capabilities (like a search_engine tool) to help it complete its task.
A memory agent identifies information from an interaction that should be available later- for example, a user preference, an architectural decision, or a constraint established during the conversation.
Here, the conversation is the material the agent needs to analyze. With fork , the memory agent receives the full interaction and can decide what is worth preserving without requiring the supervisor to restate it in the task.
Because we want to limit exactly what the memorizer agent can edit, we can specialize the subagent by setting restrictions on what files it can edit while it works.
deepagents is a framework we’re building that takes our lessons learned from working with thousands of different teams shipping agents. You can try subagent context modes (see the docs here:https://docs.langchain.com/oss/python/deepagents/subagents#forked-subagents), and much more by installing:
Let us know what you think via GitHub issues:https://github.com/langchain-ai/deepagents/issues, the forum:https://forum.langchain.com/, or on X / LinkedIn!
LangSmith, our agent engineering platform, helps developers debug every agent decision, eval changes, and deploy in one click.