Poolside 发布 Laguna S 2.1,一个 118B 总参数、8B 激活参数的 MoE 模型,支持 1M token 上下文窗口,从训练到发布用时不到九周。
在 Terminal-Bench 2.1 上以 70.2% 的得分超越多数同尺寸模型,在 DeepSWE 上得分 40.4%。 模型权重已开源,完整评估轨迹可在 trajectories.poolside.ai 获取。
今天我们发布了 Laguna S 2.1,这是我们在开发能够执行长远任务并有效利用推理的模型方面迈出的重要一步。
Laguna S 2.1 是一个拥有 118B 总参数的专家混合(Mixture-of-Experts, MoE)模型,每个 token 激活 8B 参数,并且在思考模式和非思考模式下支持最长 1M token 的上下文窗口。从训练开始到上线不到九周,在长远编码基准测试中,其表现可以与体积大数倍的模型相媲美。对于今天发布的每个基准分数,我们都在 trajectories.poolside.ai 发布了最终评估集每次试验的完整轨迹:http://trajectories.poolside.ai/。
截至目前可测量的结果,Laguna S 2.1 在其重量级别中是最有能力的自主编码模型,相差悬殊。
S 2.1 在我们的代理测试环境中启用思考功能时,在 Terminal-Bench 2.1 上得分为 70.2%。其紧凑的体积使其特别适合在本地机器上执行复杂工作。
Terminal-Bench 2.1 报告分数的排名比较。
Terminal-Bench 2.1 评估了一组广泛、高质量的长远任务,其中代理模型通过终端与其环境连接。Laguna S 2.1 在该基准测试的同类模型中表现出众。
在对数坐标轴上展示的已披露总参数数量与 Terminal-Bench 2.1 分数的散点图。Laguna S 2.1 以三角形突出显示。
Today we’re releasing Laguna S 2.1, a significant step forward in our development of models that pursue longer horizon work and make effective use of reasoning.
Laguna S 2.1 is a 118B total parameter Mixture-of-Experts (MoE) model with 8B activated parameters per token and supports a context window of up to 1M tokens in thinking and no-thinking modes. It went from the start of training to launch in under nine weeks, and on long-horizon coding benchmarks it holds its own against models many times its size. For every benchmark score we publish today, we are releasing full trajectories for every trial in the final evaluation set at trajectories.poolside.ai :http://trajectories.poolside.ai/ .
Laguna S 2.1 is, as far as we can measure, the most capable agentic coding model in its weight class by a wide margin .
S 2.1 scores 70.2% on Terminal-Bench 2.1 in our agent harness with thinking enabled. Its compact size makes it uniquely suitable for complex work on local machines.
A ranked comparison of reported Terminal-Bench 2.1 scores.
Terminal-Bench 2.1 evaluates a wide, high-quality set of long-horizon tasks where an agent model is connected to its environment through a terminal. Laguna S 2.1 is a standout model in its size category on this benchmark.
Scatter plot of disclosed total parameter count on a logarithmic axis against Terminal-Bench 2.1 score. Laguna S 2.1 is highlighted as a triangle.