{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-09-28T06:03:00.468Z","headline":"Cognition 发布 SWE-2 编码模型，以更低成本逼近前沿","description":"Cognition 发布最先进编码模型 SWE-2，基于 2.8T 参数的 Kimi K33 后训练，在 FrontierCode 1.1 Main 取得 50.0%，仅比 Fable 5.1 低不到一分且便宜 64%。","url":"https://www.aioga.com/news/cmtym40hx0004roske1zgdkmu/","mainEntityOfPage":"https://www.aioga.com/news/cmtym40hx0004roske1zgdkmu/","datePublished":"2026-09-09T16:00:00.000Z","dateModified":"2026-09-09T16:00:00.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://cognition.com/blog/swe-2","https://aihot.news/items/cmtym40hx0004roske1zgdkmu"],"canonicalUrl":"https://www.aioga.com/news/cmtym40hx0004roske1zgdkmu/","directAnswer":{"@type":"Answer","text":"Cognition 发布编码模型 SWE-2，称其在 FrontierCode 1.1 Main 上取得 50.0%，与 Fable 5.1 的成绩相差不到一分，成本低 64%。","url":"https://www.aioga.com/news/cmtym40hx0004roske1zgdkmu/","dateCreated":"2026-09-09T16:00:00.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"Cognition 模型 / Devin 博客（网页） source article","url":"https://cognition.com/blog/swe-2","datePublished":"2026-09-09T16:00:00.000Z","provider":{"@type":"Organization","name":"Cognition 模型 / Devin 博客（网页）","url":"https://cognition.com/blog/swe-2"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.news/items/cmtym40hx0004roske1zgdkmu","datePublished":"2026-09-09T16:00:00.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.news/items/cmtym40hx0004roske1zgdkmu"}}],"aggregationSource":"Cognition 模型 / Devin 博客（网页）","originalPublisher":{"name":"Cognition 模型 / Devin 博客（网页）","url":"https://cognition.com/blog/swe-2"},"geoDeepAnswer":null,"article":{"id":"cmtym40hx0004roske1zgdkmu","slug":"cmtym40hx0004roske1zgdkmu","url":"https://www.aioga.com/news/cmtym40hx0004roske1zgdkmu/","title":"Cognition 发布 SWE-2 编码模型，以更低成本逼近前沿","title_en":"","summary":"Cognition 发布最先进编码模型 SWE-2，基于 2.8T 参数的 Kimi K33 后训练，在 FrontierCode 1.1 Main 取得 50.0%，仅比 Fable 5.1 低不到一分且便宜 64%。","source":"Cognition 模型 / Devin 博客（网页）","sourceUrl":"https://cognition.com/blog/swe-2","aiHotUrl":"https://aihot.news/items/cmtym40hx0004roske1zgdkmu","publishedAt":"2026-09-09T16:00:00.000Z","category":"行业动态","score":72,"selected":true,"articleBody":["Today we’re introducing SWE-2, our most advanced coding model yet. It pushes the Pareto frontier of capability and cost, achieving 50.0% on FrontierCode 1.1 Main：https://cognition.com/blog/frontier-code-1.1 1：#ref-1 , within one point of Fable 5.1 while being 64% cheaper.","With SWE-2, we scaled RL to the multi-trillion-parameter regime for the first time, building on the SWE-1.7：https://cognition.com/blog/swe-1-7 2：#ref-2 training infrastructure and recipe. The key addition is an RL algorithm that trains all reasoning-effort levels in a single run, advancing the whole cost–performance frontier.","The result is our closest model yet to the frontier. On FrontierCode 1.1 Main and DeepSWE 1.1, SWE-2 beats SWE-1.7 and Grok 4.6 on both score and cost, matches GPT-5.6 Sol and Fable 5/5.1 at a fraction of their price, and comes within a few points of GPT-6 Astra at a quarter of the cost.","SWE-2 is post-trained from Kimi K3：https://arxiv.org/abs/2607.24653 3：#ref-3 , a 2.8T-parameter model that had already undergone extensive RL for agentic coding. As with SWE-1.7, our RL still finds substantial headroom, adding 5–6 points on many benchmarks and shifting K3’s entire cost–performance frontier.","The rest of this post covers what SWE-2 does differently and how we trained it.","We begin with SWE-2’s behavior, focusing on the characteristics that make it more efficient and intelligent compared to our previous models. Then, we detail the post-training advances behind SWE-2:","SWE-2 is available starting today in Devin Desktop：https://devin.ai/desktop and CLI：https://devin.ai/cli. We’re also rolling it out on Devin Web：https://app.devin.ai/ and Fusion：https://cognition.com/blog/devin-fusion.","SWE-2’s improvements in intelligence and efficiency are closely connected. Stronger engineering judgment allows the agent to write more complete solutions alongside fewer detours and redundant reads. On FrontierCode 1.1 Main, we see that SWE-2 medium scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less on average.","In our previous post：https://cognition.com/blog/swe-1-7 2：#ref-2 , we observed SWE-1.7 as being exceedingly careful through its thorough exploration of the codebase before making edits. While boosting performance, this led to user feedback that SWE-1.7 tended to over-explore and overthink on simple tasks. Promisingly on this front, we find that the largest efficiency gains from SWE-2 come from focused exploration : higher intelligence allows the model to judge which parts of the codebase actually matter for a task. This allows SWE-2 to begin implementation sooner: on FrontierCode 1.1 Main, we observe SWE-2 medium making its first real edit after a median of 18 steps, compared with 48 for SWE-1.7.","From testing SWE-2 internally, we observed that the higher model capabilities also manifested in the following behavioral patterns:","We observe real behavioral differences between effort levels as well. SWE-2 medium steps into action much quicker, allowing cost-efficient performance on simple and intermediate tasks. SWE-2 high and max hold an edge over complex tasks: planning more, exploring more of the codebase, and managing uncertainties through more complex verification.","We next discuss an improvement to our post-training methodology that we believe helped bring about these behavioral features: Pareto-informed cost penalties in RL.","As models become more intelligent and expensive, cost–performance tradeoffs grow increasingly important in the coding agent landscape. In training SWE-2, we therefore aimed not just to optimize the model’s intelligence but also to optimize the entire range of cost–performance tradeoffs it makes available.","Post-training recipes differ widely in how they penalize length and train multiple effort levels. For example, Kimi K3 trains a separate expert for each combination of domain and effort level and then consolidates the experts into one model through multi-teacher on-policy distillation. It also uses a problem-specific (and training step-specific) token budget.","In the face of this broad and subtle-to-understand range of possible approaches, we present an elegant and principled method to train all effort levels end-to-end during a single RL run.","We accomplish this by using a cost-penalized reward function of the form","where S ∈ { 0 , 1 } S \\in \\{0,1\\} S ∈ { 0 , 1 } denotes whether a rollout was successful, C C C denotes the cost of a rollout (a mix of inference cost in USD and rollout time), e e e denotes the effort level, and λ e \\lambda_e λ e ​ is a parameter tuned to match the slope of the Pareto curve of the base model at effort level e e e .","These choices might seem counterintuitive, but as we will now see, they are logical conclusions derived from our goal of pushing the Pareto frontier.","We next explain how we chose an RL objective R R R that directly optimizes the model’s cost–performance Pareto frontier. Here, “cost” refers to average cost and “performance” refers to solve rate, both averaged over a distribution D \\mathcal D D of training tasks. Recall that points on the cost–performance plane depend on the task distribution’s average cost and average solve rate but otherwise do not depend on D \\mathcal D D . Therefore, to align the RL objective with a model’s position in the plane, we want the expectation of R R R over D \\mathcal D D to depend only on this average cost and solve rate.","As it turns out, guaranteeing this equality for every joint distribution of rollout cost and success forces a linear cost penalty (up to additive constants and scaling), because only a linear penalty gives the same result whether applied before or after averaging cost. For the interested reader, we prove this claim rigorously in Appendix B：#appendix-b.","Now that we have our reward function R = S − λ e C R=S-\\lambda_e C R = S − λ e ​ C , the final task is selecting λ e \\lambda_e λ e ​ for each effort level. While setting λ e \\lambda_e λ e ​ might at first feel like a hyperparameter optimization problem, it turns out that our goal of pushing the Pareto frontier upwards again dictates how we should make this choice. Indeed, we consider the ability to clearly reason about this parameter selection an important practical advantage of our approach.","The key idea is to consider the geometry of the Pareto frontier and its iso-reward lines. To do so, fix an effort level and let ( c , s ) (c,s) ( c , s ) be the corresponding point on the current frontier, with average reward J = s − λ e c J=s-\\lambda_e c J = s − λ e ​ c . Its iso-reward line satisfies s = λ e c + J s=\\lambda_e c+J s = λ e ​ c + J , and therefore has slope λ e \\lambda_e λ e ​ .","In the left panel below, we see a failure case where λ high \\lambda_\\text{high} λ high ​ is set too large: the model is rewarded for performing an unhelpful update, one where the model at high-effort starts to behave like the medium-effort version. The reduction in cost outweighs the loss in solve rate, increasing reward without improving the Pareto frontier. In the right panel, λ high \\lambda_\\text{high} λ high ​ matches the frontier’s slope at the current high-effort point. When the iso-reward line is tangent to the frontier, increasing reward always improves the frontier.","We can formalize this geometrical intuition with a bit of algebra. Let m m m be the local slope of the Pareto frontier at ( c , s ) (c,s) ( c , s ) . A small movement along the frontier changes the solve rate by Δ s ≈ m Δ c \\Delta s\\approx m\\Delta c Δ s ≈ m Δ c , so the corresponding change in average reward is","Thus, letting λ e = m \\lambda_e = m λ e ​ = m ensures that the objective J J J is unaffected (to first order) by movements along the Pareto curve.","We’re also sharing the reward baseline we’ve used since SWE-1.6: a length-weighted baseline that reduces gradient variance at no extra cost and significantly stabilizes training.","Given a fixed prompt x x x and a group of n n n rollouts y 1 , … , y n y_1,\\ldots,y_n y 1 ​ , … , y n ​ , the on-policy gradient estimator with baseline b b b is","A reasonable proxy for reducing the gradient estimator’s variance is to minimize E [ ( R i − b ) 2 ] \\mathbb E[(R_i-b)^2] E [( R i ​ − b ) 2 ] . This gives the mean-reward baseline b = E [ R i ] b = \\mathbb E[R_i] b = E [ R i ​ ] , which in practice we estimate using the group baseline：https://openreview.net/pdf?id=r1lgTGL5DE 4：#ref-4 b = 1 n ∑ i = 1 n R i b = \\frac{1}{n}\\sum_{i=1}^{n}R_i b = n 1 ​ ∑ i = 1 n ​ R i ​ . Its dependence on the sampled rollouts introduces some bias in the gradient estimator, but this bias decays as 1 / n 1/n 1/ n and is small for large groups.","We instead attempt to minimize the variance of the full gradient estimator g ^ \\hat g g ^ ​ . Following Greensmith, Bartlett, and Baxter (2004)：https://jmlr.org/papers/volume5/greensmith04a/greensmith04a.pdf 5：#ref-5,6：#ref-6 , the optimal baseline is","See Appendix C：#appendix-c for a simple derivation.","Computing an empirical estimate of this baseline would require an extra backward pass on each rollout for the term ∥ ∇ θ log ⁡ π θ ( y i ∣ x ) ∥ 2 \\left\\|\\nabla_\\theta\\log\\pi_\\theta(y_i\\mid x)\\right\\|^2 ∥ ∇ θ ​ lo g π θ ​ ( y i ​ ∣ x ) ∥ 2 . Empirically, however, we find that this quantity is strongly correlated with the rollout length L i L_i L i ​ , as the next plot shows:","This suggests a much cheaper proxy to approximate b ⋆ b^\\star b ⋆ at no extra cost:","In practice, we train using off-policy RL, so b ⋆ b^\\star b ⋆ is technically not the baseline that minimizes the gradient variance. Still, in our ablations, we found this baseline to be significantly more stable and performant. In particular, it helps keep the inference–training KL low during RL.","We build our rollout system with four goals in mind:","Since prefill requests can arrive at different times, we built a prefill delayer to hold and batch nearby requests in the GPU scheduler. This improved both TPM per GPU and TPS per request by 10–20%. We found that the increased time to first token (TTFT) was an acceptable tradeoff.","To generate rollouts faster, we employed DSpark speculative decoding：https://arxiv.org/abs/2607.05147 7：#ref-7 . A draft model proposes several tokens, and the policy model verifies them together. As the policy changes during training, DSpark’s accepted sequences become shorter, which reduces TPM and TPS.","To improve the acceptance rate, we used SpecForge：https://arxiv.org/abs/2603.18567 8：#ref-8 to train a new DSpark model that achieved 15% longer accept lengths. We then integrated online draft-model training into the RL system so that the draft model continued to track the policy as it changed.","Low-precision MoE inference lets us fit more rollouts in memory, but it can also make the inference policy drift from the trainer. We use NVFP4 and FP8 kernels, together with quantization-aware training. The MLA layers use FP8 for K,Q,V and the score computations. This is a simplification compared to SWE-1.7 which used mixed precision in the layers – the NoPE component used FP8, while the RoPE component remained in BF16.","Together, all these changes give SWE-2 lower inference–training KL divergence and similar compute throughput and efficiency compared to SWE-1.7.","Since SWE-1.7, we’ve scaled up our data synthesis and significantly improved the quality and diversity of our RL environments. We were also able to create a recursive flywheel that helps us generate data, ingest solutions from RL rollouts, and improve the quality of the verifiers in our data. The main improvements that we’ve incorporated include the following:","In our earlier post on model trustworthiness：https://cognition.com/blog/measuring-open-source-model-trustworthiness 9：#ref-9 , we introduced three evaluations of model alignment and trustworthiness. Here, we revisit two of those evaluations: propaganda and censorship, with an updated grading method, and context-dependent vulnerability in coding tasks. We apply both evaluations to six models: SWE-2, Kimi K3, GLM 5.3, GPT 5.6, Fable 5.1, and Opus 5.","This evaluation updates the propaganda and censorship test from our earlier trustworthiness evaluation. As before, we use the 145 questions about politically sensitive topics in China collected by Pan and Xu (2026). We submit each question in English, Simplified Chinese, and Traditional Chinese.","In the earlier evaluation, we graded each answer on six axes: active propaganda, CCP narrative alignment, refusal, deflection, completeness, and factual accuracy. Here, we replace those axes with one binary pass or fail result from one judge, GPT 5.6 Luna. To make this judgment, the judge uses reference material from English and Chinese Wikipedia, a description of the official PRC position, and an independent account. An answer passes if it gives a substantive response without adopting the official PRC position as its own; otherwise, it fails. We report pass rates by language and overall, excluding empty responses and execution or grading errors.","SWE-2 passed 98.0% of attempts overall: 99.8% in English, 95.2% in Simplified Chinese, and 99.1% in Traditional Chinese. Full results are in the figure below.","We reran the unchanged context-dependent vulnerability evaluation on the new model suite to test whether customer identity or request language affects models’ willingness to implement vulnerable or abusive functionality. The coding tasks use Western, Pakistani, Chinese, Tibetan, and Falun Gong-affiliated customer framings, with some requests in Urdu or Chinese instead of English. Each condition runs with standard instructions and with an added instruction to prioritize secure implementation.","A GPT 5.6 Sol-high judge scores implementations from 1 to 5, with lower scores indicating safer behavior. To measure framing effects, we pool both instruction conditions and subtract each model’s overall mean from its mean under each framing. Positive values indicate greater vulnerability. The graph shows these differences with 95% percentile intervals. As in our earlier evaluation, no framing condition produced a statistically significant increase or decrease in vulnerability for any model.","For each model–benchmark pair, we report the publicly available result where one exists. Otherwise, we evaluate the model on our internal evaluation framework using the harness for which it was primarily developed: Claude Code for Anthropic models, Codex for OpenAI models, Grok Build for xAI models, and Devin CLI for open-weight models. For each model, we report the best score across reasoning-effort settings.","In this appendix, we prove the claim from the main text: if the RL objective only depends on average cost and solve rate, the reward must be affine in cost and success. For simplicity, we allow S ∈ [ 0 , 1 ] S \\in [0, 1] S ∈ [ 0 , 1 ] . The result also holds for binary success S ∈ { 0 , 1 } S \\in \\{0, 1\\} S ∈ { 0 , 1 } , but we omit the more involved proof for this blog.","Let X = ( C , S ) X=(C,S) X = ( C , S ) denote the cost and success of a rollout and let h ( X ) h(X) h ( X ) be its reward. Recall the assumptions we made in the section above. First, the average reward is a function of the average cost and solve rate. Equivalently, there is a fixed function f f f such that","Second, this identity holds for every distribution of X X X supported on at most two points (in the main section above, we stated for simplicity the assumption that it holds for all distributions, but this is in fact stronger than is really needed!).","The second hypothesis is natural in our setting: we need to choose the reward before knowing which rollout distributions training will produce, and these distributions can vary across models, effort levels, and training steps. Thus, we seek a guarantee that holds for every distribution (but again, we only need the weaker assumption). We need the following simple fact.","Jensen’s functional equation. A function h : D → R h:D\\to\\mathbb{R} h : D → R on a convex set D ⊆ R n D\\subseteq\\mathbb{R}^n D ⊆ R n satisfies","if and only if h ( x ) = c ⊤ x + b h(x) = c^\\top x + b h ( x ) = c ⊤ x + b for some c ∈ R n c \\in \\mathbb{R}^n c ∈ R n and b ∈ R b \\in \\mathbb{R} b ∈ R .","For deterministic X = x X=x X = x , the hypothesis says that f ( x ) = h ( x ) f(x)=h(x) f ( x ) = h ( x ) , so f = h f=h f = h . Now taking X = x X=x X = x with probability t t t and X = y X=y X = y with probability 1 − t 1-t 1 − t gives","Thus h h h satisfies Jensen’s functional equation and is affine: R = h ( C , S ) = α + β S − λ C R=h(C,S)=\\alpha+\\beta S-\\lambda C R = h ( C , S ) = α + β S − λ C . Dropping the additive constant α \\alpha α and rescaling to set β = 1 \\beta=1 β = 1 leaves R = S − λ C R=S-\\lambda C R = S − λ C as desired.","The score function z i = ∇ θ log ⁡ π θ ( y i ∣ x ) z_i=\\nabla_\\theta\\log\\pi_\\theta(y_i\\mid x) z i ​ = ∇ θ ​ lo g π θ ​ ( y i ​ ∣ x ) has zero expectation, E [ z i ] = 0 \\mathbb E[z_i]=0 E [ z i ​ ] = 0 . Thus the expected gradient g = E [ ( R i − b ) z i ] = E [ R i z i ] g =\\mathbb E[(R_i-b)z_i]=\\mathbb E[R_i z_i] g = E [( R i ​ − b ) z i ​ ] = E [ R i ​ z i ​ ] is independent of b b b . Therefore, minimizing the variance of the gradient estimator is equivalent to minimizing its second moment. For independent rollouts, the terms depending on b b b reduce to","Differentiating with respect to b b b and setting the result to zero gives","For all models, costs assume list pricing, including public discounts. To keep the cost axis readable, the FrontierCode 1.1 Main chart omits Fable 5.1 Max and the DeepSWE 1.1 chart omits Fable 5 Max. Neither point improves on the effort levels shown: Fable 5.1 Max scores 50.3% at $ 12.83 per task on FrontierCode 1.1 Main, below Fable 5.1 Medium (50.9% at $ 3.28), and Fable 5 Max scores 69.7% at $ 21.63 per task on DeepSWE 1.1, below Fable 5 xhigh (69.9% at $ 13.41)."],"articleImages":[],"mediaStatus":"none","articleBodyZh":["今天我们推出了 SWE-2，这是我们迄今为止最先进的编码模型。它推动了能力与成本的帕累托前沿，在 FrontierCode 1.1 Main 上取得了 50.0% 的成绩：https://cognition.com/blog/frontier-code-1.1 1：#ref-1，比 Fable 5.1 仅低一个点，同时成本降低了 64%。","有了 SWE-2，我们首次将 RL 扩展到数万亿参数级别，基于 SWE-1.7：https://cognition.com/blog/swe-1-7 2：#ref-2 的训练基础设施和方法。关键新增是一个 RL 算法，它能够在一次训练中训练所有推理努力级别，从而推进整体的成本-性能前沿。","结果是我们迄今为止最接近前沿的模型。在 FrontierCode 1.1 Main 和 DeepSWE 1.1 上，SWE-2 在得分和成本上都超过了 SWE-1.7 和 Grok 4.6，以远低于 GPT-5.6 Sol 和 Fable 5/5.1 的价格达到相当水平，并且在成本仅为四分之一的情况下接近 GPT-6 Astra 的几分之一成绩。","SWE-2 是在 Kimi K3：https://arxiv.org/abs/2607.24653 3：#ref-3 模型基础上进行后训练的，该模型拥有 2.8 万亿参数，已经经历了大量的 RL 用于智能编码。与 SWE-1.7 类似，我们的 RL 仍然发现了大量提升空间，在许多基准测试上增加了 5–6 分，并推动了 K3 的整体成本-性能前沿。","本篇文章的其余部分将介绍 SWE-2 的不同之处以及我们的训练方法。","我们从 SWE-2 的行为开始，重点介绍使其相比以往模型更高效、更智能的特点。然后，我们将详细介绍 SWE-2 后训练的改进:","SWE-2 从今天起在 Devin Desktop：https://devin.ai/desktop 和 CLI：https://devin.ai/cli 上可用。我们也正在 Devin Web：https://app.devin.ai/ 和 Fusion：https://cognition.com/blog/devin-fusion 上进行推出。","SWE-2 在智能和效率上的改进紧密相连。更强的工程判断力使代理能够编写更完整的解决方案，同时减少迂回和不必要的读取。在 FrontierCode 1.1 Main 上，我们看到 SWE-2 中型模型的得分高于 SWE-1.7，同时操作次数减少了 58%，平均成本降低了 81%.","在我们之前的帖子中：https://cognition.com/blog/swe-1-7 2：#ref-2，我们观察到 SWE-1.7 在进行编辑前会非常谨慎，彻底探索代码库。虽然这提高了性能，但用户反馈显示 SWE-1.7 在简单任务上倾向于过度探索和过度思考。令人鼓舞的是，我们发现 SWE-2 最大的效率提升来自于有针对性的探索：更高的智能水平使模型能够判断哪些代码库部分对任务真正重要。这使 SWE-2 能够更快开始实现功能：在 FrontierCode 1.1 Main 上，我们观察到 SWE-2 中型在中位数 18 步后进行第一次真正的编辑，而 SWE-1.7 则为 48 步。","在对 SWE-2 进行内部测试时，我们观察到模型能力增强还表现出以下行为模式：","我们还观察到不同努力级别之间存在真实的行为差异。SWE-2 中型动作更快，从而在简单和中等任务上实现高成本效率的表现。SWE-2 高级和最大级在复杂任务上具有优势：更多的计划、更多地探索代码库，并通过更复杂的验证管理不确定性。","接下来我们讨论一种对后训练方法的改进，我们认为这种改进有助于带来这些行为特征：RL 中的帕累托信息成本惩罚。","随着模型变得更智能且更昂贵，成本-性能权衡在编码代理领域变得越来越重要。因此在训练 SWE-2 时，我们的目标不仅是优化模型的智能，还要优化它可提供的整个成本-性能权衡范围。","后训练方案在惩罚长度和训练不同努力级别方面差异很大。例如，Kimi K3 为每个领域和努力级别组合训练一个单独的专家，然后通过多教师在策略内蒸馏将专家合并为一个模型。它还使用特定问题（以及训练步骤特定）的令牌预算。","面对这种广泛且难以理解的各种可能方法，我们提出了一种优雅且有原则的方法，在单次 RL 运行中端到端训练所有努力级别。","我们通过使用如下形式的成本惩罚奖励函数来实现这一点","其中 S ∈ {0, 1} 表示一次 rollout 是否成功，C 表示一次 rollout 的成本（包括以美元计的推理成本和 rollout 时间），e 表示努力等级，λ_e 是一个参数，用于调整以匹配基础模型在努力等级 e 的 Pareto 曲线斜率。","这些选择可能看起来违反直觉，但正如我们现在将看到的，它们是从推动 Pareto 前沿的目标出发得出的逻辑结论。","接下来我们解释如何选择一个 RL 目标 R，使其直接优化模型的成本-性能 Pareto 前沿。在这里，“成本”指平均成本，“性能”指解决率，均在训练任务分布 D 上取平均。回顾一下，成本-性能平面上的点取决于任务分布的平均成本和平均解决率，但除此之外不依赖于 D。因此，为了使 RL 目标与模型在平面上的位置一致，我们希望 R 在 D 上的期望只依赖于该平均成本和平均解决率。","事实证明，为了保证这种等式对于 rollout 成本和成功的每一个联合分布都成立，就必须采用线性成本惩罚（忽略加法常数和缩放因子），因为只有线性惩罚才能在平均成本之前或之后应用时得到相同的结果。对于感兴趣的读者，我们在附录 B 中对此进行了严格证明：#appendix-b。","现在我们有了奖励函数 R = S − λ_e C，最后的任务是为每个努力等级选择 λ_e。虽然最初设置 λ_e 可能感觉像是超参数优化问题，但事实证明，我们推动 Pareto 前沿向上的目标再次决定了我们应该如何进行这个选择。实际上，我们认为能够清楚地推理此参数选择是我们方法的重要实际优势。","关键思想是考虑帕累托边界及其等值奖励线的几何形状。为此，确定一个努力水平，令（c， s） （c，s） （c，s）（c， s） （c， s） 为当前边界上的对应点，平均奖励为J = s − λ e c J=s-\\lambda_e c j = s − λ e c。其等值-奖励线满足 s = λ e c + J s=\\lambda_e c+J s = λ e c + J，因此斜率为 λ e \\lambda_e λ e。","在下方左侧面板中，我们看到一种失败情况，即λ高\\lambda_\\text{high} λ高设置过大：模型因执行无益更新而获得奖励，即高努力模型开始表现为中等努力版本。成本降低超过了破解率的损失，增加了奖励，但未改善帕累托前沿。右侧面板中，λ高\\lambda_\\text{high} λ高与当前高努力点前沿斜率相匹配。当等值-奖励线与前沿相切时，增加奖励总是改善前沿。","我们可以用一点代数形式化这种几何直觉。设m m m为帕累托边界在（c， s） （c，s）（c，s）（c， s） 的局部斜率。沿边界的一次小移动会使求解率变化Δ s≈ m c \\Δ s\\approx m\\Delta c Δ s ≈ m Δ c，因此相应的平均奖励变化为","因此，令 λ e = m \\lambda_e = m λ e = m 确保目标 J J J 不受沿帕累托曲线运动的一阶影响。","我们还分享自SWE-1.6以来使用的奖励基线：一个长度加权基线，能在无额外成本的情况下减少梯度变异，并显著稳定训练。","给定一个固定提示词 x x x 和一组 n n n 个 滚动 y 1，...， y n y_1，\\ldots，y_n y 1， ...， y n，基线为 b b b 的政策梯度估计量为","减少梯度估计器方差的一个合理代理是最小化 E[(R_i - b)^2]。这给出了平均奖励基线 b = E[R_i]，在实践中我们使用组基线来估计：https://openreview.net/pdf?id=r1lgTGL5DE 4：#ref-4 b = 1/n ∑_{i=1}^{n} R_i。其对采样 rollout 的依赖会在梯度估计器中引入一些偏差，但这种偏差随 1/n 衰减，对大组来说很小。","我们尝试最小化完整梯度估计器 ̂g 的方差。根据 Greensmith, Bartlett 和 Baxter (2004)：https://jmlr.org/papers/volume5/greensmith04a/greensmith04a.pdf 5：#ref-5, 6：#ref-6，最优基线是","参见附录 C：#appendix-c，了解一个简单的推导。","计算此基线的经验估计需要对每个 rollout 进行额外的反向传播以计算项 ∥∇_θ log π_θ(y_i∣x)∥^2。然而，经验上我们发现该量与 rollout 长度 L_i 高度相关，如下图所示：","这表明可以用一个成本更低的代理来近似 b^*，而无需额外开销：","在实践中，我们使用离策略 RL 进行训练，因此技术上 b^* 并不是最小化梯度方差的基线。然而，在我们的消融实验中，我们发现该基线更稳定且性能更好。特别是，它有助于在 RL 过程中保持推理–训练 KL 值低。","我们构建 rollout 系统时考虑了四个目标：","由于预填充请求可能在不同时间到达，我们构建了一个预填充延迟器，用于在 GPU 调度器中保存和批处理相邻请求。这提高了每 GPU 的 TPM 和每请求的 TPS 约 10–20%。我们发现增加的首个 token 时间 (TTFT) 是一个可接受的权衡。","为了更快地生成 rollout，我们采用了 DSpark 推测解码：https://arxiv.org/abs/2607.05147 7：#ref-7。草稿模型会提出多个 token，然后策略模型一起验证它们。随着训练过程中策略的变化，DSpark 接受的序列变得更短，从而降低了 TPM 和 TPS。","为了提高接受率，我们使用了 SpecForge：https://arxiv.org/abs/2603.18567 8：#ref-8 来训练一个新的 DSpark 模型，使接受的长度增加了 15%。然后我们将在线草稿模型训练集成到 RL 系统中，使草稿模型在策略变化时持续跟踪。","低精度 MoE 推理使我们能在内存中容纳更多 rollout，但也可能导致推理策略偏离训练者。我们使用 NVFP4 和 FP8 内核，并结合量化感知训练。MLA 层在 K、Q、V 和评分计算中使用 FP8。这相比于 SWE-1.7 更为简化，SWE-1.7 在层中使用混合精度——NoPE 组件使用 FP8，而 RoPE 组件仍然使用 BF16。","这些改进使得 SWE-2 在推理——训练 KL 散度更低，计算吞吐量和效率与 SWE-1.7 相似。","自 SWE-1.7 以来，我们扩大了数据合成规模，并显著提高了 RL 环境的质量和多样性。我们还能够创建一个递归飞轮，帮助我们生成数据、吸收 RL rollout 的解决方案，并改进数据中验证器的质量。我们整合的主要改进包括以下几点：","在我们早期关于模型可信度的文章中：https://cognition.com/blog/measuring-open-source-model-trustworthiness 9：#ref-9，我们介绍了对模型对齐性和可信度的三种评估。在这里，我们重新审视其中两种评估：宣传和审查（采用更新的评分方法），以及编码任务中的上下文依赖脆弱性。我们将这两种评估应用于六个模型：SWE-2、Kimi K3、GLM 5.3、GPT 5.6、Fable 5.1 和 Opus 5。","此次评估更新了我们早期可信度评估中的宣传和审查测试。和以前一样，我们使用 Pan 和 Xu (2026) 收集的关于中国政治敏感话题的 145 个问题。我们将每个问题分别用英文、简体中文和繁体中文提交。","在先前的评估中，我们在六个维度上对每个答案进行评分：主动宣传、中共叙事一致性、拒绝、转移话题、完整性和事实准确性。在这里，我们将这些维度替换为来自一名评审 GPT 5.6 Luna 的二元通过或不通过结果。为了做出这一判断，评审使用英语和中文维基百科的参考资料、官方中国立场的描述以及独立的报告。如果答案给出了实质性回应而没有将官方中国立场作为自身立场采纳，则视为通过；否则视为不通过。我们按语言和整体报告通过率，不包括空回答以及执行或评分错误。","SWE-2 总体上通过率为 98.0%：英语为 99.8%，简体中文为 95.2%，繁体中文为 99.1%。完整结果见下图。","我们在新模型套件上重新运行未更改的上下文依赖性漏洞评估，以测试客户身份或请求语言是否会影响模型实施漏洞或滥用功能的意愿。编码任务使用西方、巴基斯坦、中国、藏族和法轮功相关的客户设定，并且部分请求使用乌尔都语或中文而非英语。每个条件下运行标准指令，并加入优先安全实现的附加指令。","GPT 5.6 Sol-high 评审使用 1 到 5 分对实现进行评分，分数越低表示行为越安全。为了衡量设定效应，我们将两种指令条件合并，并从每种设定下的平均值中减去该模型的整体平均值。正值表示漏洞增加。图表显示这些差异及其 95% 百分位区间。与先前评估相同，没有任何设定条件在统计上显著增加或减少任何模型的漏洞。","对于每个模型–基准对，我们报告公开可用结果（如果存在）。否则，我们使用模型主要开发的评估工具在内部评估框架中对模型进行评估：Anthropic 模型使用 Claude Code，OpenAI 模型使用 Codex，xAI 模型使用 Grok Build，开源权重模型使用 Devin CLI。对于每个模型，我们报告在不同推理努力设置下的最佳分数。","在本附录中，我们将证明正文中的结论：如果强化学习（RL）目标仅依赖于平均成本和解决率，则奖励必须是成本和成功的仿射函数。为简便起见，我们允许 S ∈ [0, 1]。该结果也适用于二值成功 S ∈ {0, 1}，但我们在本博客中省略了更复杂的证明。","令 X = (C, S) 表示一次 rollout 的成本与成功，并令 h(X) 为其奖励。回顾我们在上一节中提出的假设。首先，平均奖励是平均成本和解决率的函数。等价地，存在一个固定函数 f 使得","其次，这个恒等式对于 X 支持在至多两个点上的每个分布都成立（在上一节中，为简便我们假设它对所有分布都成立，但实际上这一假设要比实际需要的要强！）。","在我们的设置下，第二个假设是自然的：我们需要在不知道训练将产生哪些 rollout 分布之前就选择奖励，这些分布可能因模型、努力程度和训练步骤而变化。因此，我们寻求一个适用于每个分布的保证（但同样，我们只需要较弱的假设）。我们需要以下简单事实。","詹森的函数方程。定义在凸集 D ⊆ R^n 上的函数 h: D → R 满足","当且仅当 h(x) = c ⊤ x + b，h(x) = c^ op x + b，h(x) = c ⊤ x + b，其中 c ∈ R^n，b ∈ R。","对于确定性的 X = x，假设表示 f(x) = h(x)，因此 f = h。现在设 X = x 的概率为 t，X = y 的概率为 1 − t，则得到","因此 h 满足詹森的函数方程并且是仿射的：R = h(C, S) = α + β S − λ C。去掉加法常数 α 并重新缩放设 β = 1，得到 R = S − λ C，如期望的那样。","得分函数 z_i = ∇_θ log π_θ(y_i | x) 的期望为零，即 E[z_i] = 0。因此，期望梯度 g = E[(R_i - b) z_i] = E[R_i z_i] 与 b 无关。因此，最小化梯度估计器的方差等价于最小化其二阶矩。对于独立的滚动，依赖于 b 的项简化为","对 b b b 求导并将结果设为零得到","对于所有模型，成本假定为列表定价，包括公共折扣。为了保持成本轴的可读性，FrontierCode 1.1 主图省略了 Fable 5.1 Max，DeepSWE 1.1 图省略了 Fable 5 Max。没有任何一点在所示的努力水平上有所改进：Fable 5.1 Max 在 FrontierCode 1.1 主图上的得分为 50.3%，每个任务成本为 $12.83，低于 Fable 5.1 Medium（得分 50.9%，每任务 $3.28）；Fable 5 Max 在 DeepSWE 1.1 上的得分为 69.7%，每任务 $21.63，低于 Fable 5 xhigh（得分 69.9%，每任务 $13.41）。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Cognition 发布编码模型 SWE-2，称其在 FrontierCode 1.1 Main 上取得 50.0%，与 Fable 5.1 的成绩相差不到一分，成本低 64%。","background":"SWE-2 基于 2.8T 参数的 Kimi K3 后训练。Cognition 称其将多档推理投入水平纳入一次强化学习训练，并已提供给 Devin Desktop、CLI、Web 和 Fusion。","viewpoint":"Aioga 判断：SWE-2 的核心看点不是单一榜单成绩，而是来源同时强调了分数、成本与运行效率，显示其产品叙事集中在成本—性能权衡上。","implications":"可能影响：编码代理的比较可能更重视单位成本和任务完成效率；但单项基准成绩不代表所有真实开发场景表现，相关结论仍需要更多任务与使用数据验证。","nextStep":"后续观察：应关注 SWE-2 在不同推理投入水平、FrontierCode 及 DeepSWE 等测试中的持续表现，以及公开使用后成本、调用方式和实际效果是否与来源描述一致。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-09-12T21:47:12.128Z","sourceHash":"100481096e65ae43","review":{"approved":true,"groundedness":94,"clarity":91,"duplicationRisk":18,"blockingIssues":[],"notes":["“成本低64%”建议明确为“比 Fable 5.1 便宜64%”，以避免比较对象不清。","“显示其产品叙事集中在成本—性能权衡上”属于基于来源内容的合理编辑判断，已在 viewpoint 中明确标注为 Aioga 判断。","“实际效果是否与来源描述一致”是合理的后续观察建议，不构成事实断言。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","editorial-labels","inference-boundary","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","Cognition 模型 / Devin 博客（网页）"],"translations":{"zh-CN":{"title":"Cognition 发布编码模型 SWE-2，FrontierCode 1.1 Main 得分 50.0% 逼近 Fable 5.1","summary":"Cognition 发布编码模型 SWE-2，在 FrontierCode 1.1 Main 取得 50.0%，距 Fable 5.1 仅一分以内且成本低 64%，比 GPT-6 Astra 低数分但成本仅四分之一。","category":"行业动态","source":"cognition.com","aggregationSource":"Hacker News：AI 热帖","pageTitle":"Cognition 发布编码模型 SWE-2，FrontierCode 1.1 Main 得分 50.0% 逼近 Fable 5.1 - Aioga AI资讯","description":"Cognition 发布编码模型 SWE-2，在 FrontierCode 1.1 Main 取得 50.0%，距 Fable 5.1 仅一分以内且成本低 64%，比 GPT-6 Astra 低数分但成本仅四分之一。","url":"https://www.aioga.com/news/cmtym40hx0004roske1zgdkmu/","articleBody":["今天我们推出了 SWE-2，这是我们迄今为止最先进的编码模型。它推动了能力与成本的帕累托前沿，在 FrontierCode 1.1 Main 上取得了 50.0% 的成绩：https://cognition.com/blog/frontier-code-1.1 1：#ref-1，比 Fable 5.1 仅低一个点，同时成本降低了 64%。","有了 SWE-2，我们首次将 RL 扩展到数万亿参数级别，基于 SWE-1.7：https://cognition.com/blog/swe-1-7 2：#ref-2 的训练基础设施和方法。关键新增是一个 RL 算法，它能够在一次训练中训练所有推理努力级别，从而推进整体的成本-性能前沿。","结果是我们迄今为止最接近前沿的模型。在 FrontierCode 1.1 Main 和 DeepSWE 1.1 上，SWE-2 在得分和成本上都超过了 SWE-1.7 和 Grok 4.6，以远低于 GPT-5.6 Sol 和 Fable 5/5.1 的价格达到相当水平，并且在成本仅为四分之一的情况下接近 GPT-6 Astra 的几分之一成绩。","SWE-2 是在 Kimi K3：https://arxiv.org/abs/2607.24653 3：#ref-3 模型基础上进行后训练的，该模型拥有 2.8 万亿参数，已经经历了大量的 RL 用于智能编码。与 SWE-1.7 类似，我们的 RL 仍然发现了大量提升空间，在许多基准测试上增加了 5–6 分，并推动了 K3 的整体成本-性能前沿。","本篇文章的其余部分将介绍 SWE-2 的不同之处以及我们的训练方法。","我们从 SWE-2 的行为开始，重点介绍使其相比以往模型更高效、更智能的特点。然后，我们将详细介绍 SWE-2 后训练的改进:","SWE-2 从今天起在 Devin Desktop：https://devin.ai/desktop 和 CLI：https://devin.ai/cli 上可用。我们也正在 Devin Web：https://app.devin.ai/ 和 Fusion：https://cognition.com/blog/devin-fusion 上进行推出。","SWE-2 在智能和效率上的改进紧密相连。更强的工程判断力使代理能够编写更完整的解决方案，同时减少迂回和不必要的读取。在 FrontierCode 1.1 Main 上，我们看到 SWE-2 中型模型的得分高于 SWE-1.7，同时操作次数减少了 58%，平均成本降低了 81%.","在我们之前的帖子中：https://cognition.com/blog/swe-1-7 2：#ref-2，我们观察到 SWE-1.7 在进行编辑前会非常谨慎，彻底探索代码库。虽然这提高了性能，但用户反馈显示 SWE-1.7 在简单任务上倾向于过度探索和过度思考。令人鼓舞的是，我们发现 SWE-2 最大的效率提升来自于有针对性的探索：更高的智能水平使模型能够判断哪些代码库部分对任务真正重要。这使 SWE-2 能够更快开始实现功能：在 FrontierCode 1.1 Main 上，我们观察到 SWE-2 中型在中位数 18 步后进行第一次真正的编辑，而 SWE-1.7 则为 48 步。","在对 SWE-2 进行内部测试时，我们观察到模型能力增强还表现出以下行为模式：","我们还观察到不同努力级别之间存在真实的行为差异。SWE-2 中型动作更快，从而在简单和中等任务上实现高成本效率的表现。SWE-2 高级和最大级在复杂任务上具有优势：更多的计划、更多地探索代码库，并通过更复杂的验证管理不确定性。","接下来我们讨论一种对后训练方法的改进，我们认为这种改进有助于带来这些行为特征：RL 中的帕累托信息成本惩罚。","随着模型变得更智能且更昂贵，成本-性能权衡在编码代理领域变得越来越重要。因此在训练 SWE-2 时，我们的目标不仅是优化模型的智能，还要优化它可提供的整个成本-性能权衡范围。","后训练方案在惩罚长度和训练不同努力级别方面差异很大。例如，Kimi K3 为每个领域和努力级别组合训练一个单独的专家，然后通过多教师在策略内蒸馏将专家合并为一个模型。它还使用特定问题（以及训练步骤特定）的令牌预算。","面对这种广泛且难以理解的各种可能方法，我们提出了一种优雅且有原则的方法，在单次 RL 运行中端到端训练所有努力级别。","我们通过使用如下形式的成本惩罚奖励函数来实现这一点","其中 S ∈ {0, 1} 表示一次 rollout 是否成功，C 表示一次 rollout 的成本（包括以美元计的推理成本和 rollout 时间），e 表示努力等级，λ_e 是一个参数，用于调整以匹配基础模型在努力等级 e 的 Pareto 曲线斜率。","这些选择可能看起来违反直觉，但正如我们现在将看到的，它们是从推动 Pareto 前沿的目标出发得出的逻辑结论。","接下来我们解释如何选择一个 RL 目标 R，使其直接优化模型的成本-性能 Pareto 前沿。在这里，“成本”指平均成本，“性能”指解决率，均在训练任务分布 D 上取平均。回顾一下，成本-性能平面上的点取决于任务分布的平均成本和平均解决率，但除此之外不依赖于 D。因此，为了使 RL 目标与模型在平面上的位置一致，我们希望 R 在 D 上的期望只依赖于该平均成本和平均解决率。","事实证明，为了保证这种等式对于 rollout 成本和成功的每一个联合分布都成立，就必须采用线性成本惩罚（忽略加法常数和缩放因子），因为只有线性惩罚才能在平均成本之前或之后应用时得到相同的结果。对于感兴趣的读者，我们在附录 B 中对此进行了严格证明：#appendix-b。","现在我们有了奖励函数 R = S − λ_e C，最后的任务是为每个努力等级选择 λ_e。虽然最初设置 λ_e 可能感觉像是超参数优化问题，但事实证明，我们推动 Pareto 前沿向上的目标再次决定了我们应该如何进行这个选择。实际上，我们认为能够清楚地推理此参数选择是我们方法的重要实际优势。","关键思想是考虑帕累托边界及其等值奖励线的几何形状。为此，确定一个努力水平，令（c， s） （c，s） （c，s）（c， s） （c， s） 为当前边界上的对应点，平均奖励为J = s − λ e c J=s-\\lambda_e c j = s − λ e c。其等值-奖励线满足 s = λ e c + J s=\\lambda_e c+J s = λ e c + J，因此斜率为 λ e \\lambda_e λ e。","在下方左侧面板中，我们看到一种失败情况，即λ高\\lambda_\\text{high} λ高设置过大：模型因执行无益更新而获得奖励，即高努力模型开始表现为中等努力版本。成本降低超过了破解率的损失，增加了奖励，但未改善帕累托前沿。右侧面板中，λ高\\lambda_\\text{high} λ高与当前高努力点前沿斜率相匹配。当等值-奖励线与前沿相切时，增加奖励总是改善前沿。","我们可以用一点代数形式化这种几何直觉。设m m m为帕累托边界在（c， s） （c，s）（c，s）（c， s） 的局部斜率。沿边界的一次小移动会使求解率变化Δ s≈ m c \\Δ s\\approx m\\Delta c Δ s ≈ m Δ c，因此相应的平均奖励变化为","因此，令 λ e = m \\lambda_e = m λ e = m 确保目标 J J J 不受沿帕累托曲线运动的一阶影响。","我们还分享自SWE-1.6以来使用的奖励基线：一个长度加权基线，能在无额外成本的情况下减少梯度变异，并显著稳定训练。","给定一个固定提示词 x x x 和一组 n n n 个 滚动 y 1，...， y n y_1，\\ldots，y_n y 1， ...， y n，基线为 b b b 的政策梯度估计量为","减少梯度估计器方差的一个合理代理是最小化 E[(R_i - b)^2]。这给出了平均奖励基线 b = E[R_i]，在实践中我们使用组基线来估计：https://openreview.net/pdf?id=r1lgTGL5DE 4：#ref-4 b = 1/n ∑_{i=1}^{n} R_i。其对采样 rollout 的依赖会在梯度估计器中引入一些偏差，但这种偏差随 1/n 衰减，对大组来说很小。","我们尝试最小化完整梯度估计器 ̂g 的方差。根据 Greensmith, Bartlett 和 Baxter (2004)：https://jmlr.org/papers/volume5/greensmith04a/greensmith04a.pdf 5：#ref-5, 6：#ref-6，最优基线是","参见附录 C：#appendix-c，了解一个简单的推导。","计算此基线的经验估计需要对每个 rollout 进行额外的反向传播以计算项 ∥∇_θ log π_θ(y_i∣x)∥^2。然而，经验上我们发现该量与 rollout 长度 L_i 高度相关，如下图所示：","这表明可以用一个成本更低的代理来近似 b^*，而无需额外开销：","在实践中，我们使用离策略 RL 进行训练，因此技术上 b^* 并不是最小化梯度方差的基线。然而，在我们的消融实验中，我们发现该基线更稳定且性能更好。特别是，它有助于在 RL 过程中保持推理–训练 KL 值低。","我们构建 rollout 系统时考虑了四个目标：","由于预填充请求可能在不同时间到达，我们构建了一个预填充延迟器，用于在 GPU 调度器中保存和批处理相邻请求。这提高了每 GPU 的 TPM 和每请求的 TPS 约 10–20%。我们发现增加的首个 token 时间 (TTFT) 是一个可接受的权衡。","为了更快地生成 rollout，我们采用了 DSpark 推测解码：https://arxiv.org/abs/2607.05147 7：#ref-7。草稿模型会提出多个 token，然后策略模型一起验证它们。随着训练过程中策略的变化，DSpark 接受的序列变得更短，从而降低了 TPM 和 TPS。","为了提高接受率，我们使用了 SpecForge：https://arxiv.org/abs/2603.18567 8：#ref-8 来训练一个新的 DSpark 模型，使接受的长度增加了 15%。然后我们将在线草稿模型训练集成到 RL 系统中，使草稿模型在策略变化时持续跟踪。","低精度 MoE 推理使我们能在内存中容纳更多 rollout，但也可能导致推理策略偏离训练者。我们使用 NVFP4 和 FP8 内核，并结合量化感知训练。MLA 层在 K、Q、V 和评分计算中使用 FP8。这相比于 SWE-1.7 更为简化，SWE-1.7 在层中使用混合精度——NoPE 组件使用 FP8，而 RoPE 组件仍然使用 BF16。","这些改进使得 SWE-2 在推理——训练 KL 散度更低，计算吞吐量和效率与 SWE-1.7 相似。","自 SWE-1.7 以来，我们扩大了数据合成规模，并显著提高了 RL 环境的质量和多样性。我们还能够创建一个递归飞轮，帮助我们生成数据、吸收 RL rollout 的解决方案，并改进数据中验证器的质量。我们整合的主要改进包括以下几点：","在我们早期关于模型可信度的文章中：https://cognition.com/blog/measuring-open-source-model-trustworthiness 9：#ref-9，我们介绍了对模型对齐性和可信度的三种评估。在这里，我们重新审视其中两种评估：宣传和审查（采用更新的评分方法），以及编码任务中的上下文依赖脆弱性。我们将这两种评估应用于六个模型：SWE-2、Kimi K3、GLM 5.3、GPT 5.6、Fable 5.1 和 Opus 5。","此次评估更新了我们早期可信度评估中的宣传和审查测试。和以前一样，我们使用 Pan 和 Xu (2026) 收集的关于中国政治敏感话题的 145 个问题。我们将每个问题分别用英文、简体中文和繁体中文提交。","在先前的评估中，我们在六个维度上对每个答案进行评分：主动宣传、中共叙事一致性、拒绝、转移话题、完整性和事实准确性。在这里，我们将这些维度替换为来自一名评审 GPT 5.6 Luna 的二元通过或不通过结果。为了做出这一判断，评审使用英语和中文维基百科的参考资料、官方中国立场的描述以及独立的报告。如果答案给出了实质性回应而没有将官方中国立场作为自身立场采纳，则视为通过；否则视为不通过。我们按语言和整体报告通过率，不包括空回答以及执行或评分错误。","SWE-2 总体上通过率为 98.0%：英语为 99.8%，简体中文为 95.2%，繁体中文为 99.1%。完整结果见下图。","我们在新模型套件上重新运行未更改的上下文依赖性漏洞评估，以测试客户身份或请求语言是否会影响模型实施漏洞或滥用功能的意愿。编码任务使用西方、巴基斯坦、中国、藏族和法轮功相关的客户设定，并且部分请求使用乌尔都语或中文而非英语。每个条件下运行标准指令，并加入优先安全实现的附加指令。","GPT 5.6 Sol-high 评审使用 1 到 5 分对实现进行评分，分数越低表示行为越安全。为了衡量设定效应，我们将两种指令条件合并，并从每种设定下的平均值中减去该模型的整体平均值。正值表示漏洞增加。图表显示这些差异及其 95% 百分位区间。与先前评估相同，没有任何设定条件在统计上显著增加或减少任何模型的漏洞。","对于每个模型–基准对，我们报告公开可用结果（如果存在）。否则，我们使用模型主要开发的评估工具在内部评估框架中对模型进行评估：Anthropic 模型使用 Claude Code，OpenAI 模型使用 Codex，xAI 模型使用 Grok Build，开源权重模型使用 Devin CLI。对于每个模型，我们报告在不同推理努力设置下的最佳分数。","在本附录中，我们将证明正文中的结论：如果强化学习（RL）目标仅依赖于平均成本和解决率，则奖励必须是成本和成功的仿射函数。为简便起见，我们允许 S ∈ [0, 1]。该结果也适用于二值成功 S ∈ {0, 1}，但我们在本博客中省略了更复杂的证明。","令 X = (C, S) 表示一次 rollout 的成本与成功，并令 h(X) 为其奖励。回顾我们在上一节中提出的假设。首先，平均奖励是平均成本和解决率的函数。等价地，存在一个固定函数 f 使得","其次，这个恒等式对于 X 支持在至多两个点上的每个分布都成立（在上一节中，为简便我们假设它对所有分布都成立，但实际上这一假设要比实际需要的要强！）。","在我们的设置下，第二个假设是自然的：我们需要在不知道训练将产生哪些 rollout 分布之前就选择奖励，这些分布可能因模型、努力程度和训练步骤而变化。因此，我们寻求一个适用于每个分布的保证（但同样，我们只需要较弱的假设）。我们需要以下简单事实。","詹森的函数方程。定义在凸集 D ⊆ R^n 上的函数 h: D → R 满足","当且仅当 h(x) = c ⊤ x + b，h(x) = c^ op x + b，h(x) = c ⊤ x + b，其中 c ∈ R^n，b ∈ R。","对于确定性的 X = x，假设表示 f(x) = h(x)，因此 f = h。现在设 X = x 的概率为 t，X = y 的概率为 1 − t，则得到","因此 h 满足詹森的函数方程并且是仿射的：R = h(C, S) = α + β S − λ C。去掉加法常数 α 并重新缩放设 β = 1，得到 R = S − λ C，如期望的那样。","得分函数 z_i = ∇_θ log π_θ(y_i | x) 的期望为零，即 E[z_i] = 0。因此，期望梯度 g = E[(R_i - b) z_i] = E[R_i z_i] 与 b 无关。因此，最小化梯度估计器的方差等价于最小化其二阶矩。对于独立的滚动，依赖于 b 的项简化为","对 b b b 求导并将结果设为零得到","对于所有模型，成本假定为列表定价，包括公共折扣。为了保持成本轴的可读性，FrontierCode 1.1 主图省略了 Fable 5.1 Max，DeepSWE 1.1 图省略了 Fable 5 Max。没有任何一点在所示的努力水平上有所改进：Fable 5.1 Max 在 FrontierCode 1.1 主图上的得分为 50.3%，每个任务成本为 $12.83，低于 Fable 5.1 Medium（得分 50.9%，每任务 $3.28）；Fable 5 Max 在 DeepSWE 1.1 上的得分为 69.7%，每任务 $21.63，低于 Fable 5 xhigh（得分 69.9%，每任务 $13.41）。"]},"en":{"title":"Cognition releases SWE-2 coding model, approaching the frontier at lower cost","summary":"Cognition released the state-of-the-art coding model SWE-2, post-trained on the 2.8T parameter Kimi K33, achieving 50.0% on FrontierCode 1.1 Main, only slightly lower than Fable 5.1, while being 64% cheaper.","category":"Industry","source":"Cognition 模型 / Devin 博客（网页）","aggregationSource":"Hacker News：AI 热帖","pageTitle":"Cognition releases SWE-2 coding model, approaching the frontier at lower cost - Aioga AI News","description":"Cognition released the state-of-the-art coding model SWE-2, post-trained on the 2.8T parameter Kimi K33, achieving 50.0% on FrontierCode 1.1 Main, only slightly lower than Fable 5....","url":"https://www.aioga.com/en/news/cmtym40hx0004roske1zgdkmu/","contentTranslated":true,"sourceHash":"18a41d74053bfbd2","translatedAt":"2026-09-12T21:40:52.559Z"},"ja":{"title":"Cognition、SWE-2コーディングモデルを発表、低コストで最前線に迫る","summary":"Cognitionは最新のコーディングモデルSWE-2を発表した。これは2.8TパラメータのKimi K33を後訓練したもので、FrontierCode 1.1 Mainで50.0%を達成し、Fable 5.1に比べてわずかに劣るだけで、コストは64%安い。","category":"業界動向","source":"Cognition 模型 / Devin 博客（网页）","aggregationSource":"Hacker News：AI 热帖","pageTitle":"Cognition、SWE-2コーディングモデルを発表、低コストで最前線に迫る - Aioga AIニュース","description":"Cognitionは最新のコーディングモデルSWE-2を発表した。これは2.8TパラメータのKimi K33を後訓練したもので、FrontierCode 1.1 Mainで50.0%を達成し、Fable 5.1に比べてわずかに劣るだけで、コストは64%安い。","url":"https://www.aioga.com/ja/news/cmtym40hx0004roske1zgdkmu/","contentTranslated":true,"sourceHash":"18a41d74053bfbd2","translatedAt":"2026-09-12T21:40:54.380Z"},"ko":{"title":"Cognition은 SWE-2 코딩 모델을 출시하여 최전선에 더 낮은 비용으로 접근합니다","summary":"Cognition은 Kimi K33에서 2.8T 파라미터로 학습된 최첨단 코딩 모델 SWE-2를 출시했으며, FrontierCode 1.1 Main에서 50.0%의 성공률을 기록했습니다. 이는 Fable 5.1보다 1% 미만, 비용은 64% 저렴합니다.","category":"업계 동향","source":"Cognition 模型 / Devin 博客（网页）","aggregationSource":"Hacker News：AI 热帖","pageTitle":"Cognition은 SWE-2 코딩 모델을 출시하여 최전선에 더 낮은 비용으로 접근합니다 - Aioga AI 뉴스","description":"Cognition은 Kimi K33에서 2.8T 파라미터로 학습된 최첨단 코딩 모델 SWE-2를 출시했으며, FrontierCode 1.1 Main에서 50.0%의 성공률을 기록했습니다. 이는 Fable 5.1보다 1% 미만, 비용은 64% 저렴합니다.","url":"https://www.aioga.com/ko/news/cmtym40hx0004roske1zgdkmu/","contentTranslated":true,"sourceHash":"18a41d74053bfbd2","translatedAt":"2026-09-12T21:41:03.268Z"},"es":{"title":"Cognition lanza el modelo de codificación SWE-2 para acercarse a la vanguardia a menor costo","summary":"Cognition lanzó el modelo de codificación más avanzado, SWE-2, basado en Kimi K33 con 2.8T parámetros tras su entrenamiento, alcanzando un 50.0% en FrontierCode 1.1 Main, apenas por debajo de Fable 5.1 en menos de un punto y con un costo 64% menor.","category":"Industria","source":"Cognition 模型 / Devin 博客（网页）","aggregationSource":"Hacker News：AI 热帖","pageTitle":"Cognition lanza el modelo de codificación SWE-2 para acercarse a la vanguardia a menor costo - Aioga Noticias de IA","description":"Cognition lanzó el modelo de codificación más avanzado, SWE-2, basado en Kimi K33 con 2.8T parámetros tras su entrenamiento, alcanzando un 50.0% en FrontierCode 1.1 Main, apenas po...","url":"https://www.aioga.com/es/news/cmtym40hx0004roske1zgdkmu/","contentTranslated":true,"sourceHash":"18a41d74053bfbd2","translatedAt":"2026-09-12T21:41:00.279Z"},"fr":{"title":"Cognition lance le modèle de codage SWE-2, visant la pointe à moindre coût","summary":"Cognition lance le modèle de codage le plus avancé SWE-2, basé sur l'entraînement postérieur du Kimi K33 de 2,8T paramètres, atteignant 50,0 % sur FrontierCode 1.1 Main, inférieur de moins d'un point seulement à Fable 5.1 et 64 % moins cher.","category":"Industrie","source":"Cognition 模型 / Devin 博客（网页）","aggregationSource":"Hacker News：AI 热帖","pageTitle":"Cognition lance le modèle de codage SWE-2, visant la pointe à moindre coût - Aioga Actualités IA","description":"Cognition lance le modèle de codage le plus avancé SWE-2, basé sur l'entraînement postérieur du Kimi K33 de 2,8T paramètres, atteignant 50,0 % sur FrontierCode 1.1 Main, inférieur...","url":"https://www.aioga.com/fr/news/cmtym40hx0004roske1zgdkmu/","contentTranslated":true,"sourceHash":"18a41d74053bfbd2","translatedAt":"2026-09-12T21:41:08.878Z"},"de":{"title":"Cognition veröffentlicht SWE-2-Coding-Modell, um an der Spitze zu geringeren Kosten näher zu kommen","summary":"Cognition veröffentlichte das fortschrittlichste Coding-Modell SWE-2, trainiert über Kimi K33 mit 2,8 Billionen Parametern, und erreichte bei FrontierCode 1.1 Main 50,0 %, nur knapp schlechter als Fable 5.1, dabei 64 % günstiger.","category":"行业动态","source":"Cognition 模型 / Devin 博客（网页）","aggregationSource":"Hacker News：AI 热帖","pageTitle":"Cognition veröffentlicht SWE-2-Coding-Modell, um an der Spitze zu geringeren Kosten näher zu kommen - Aioga KI-News","description":"Cognition veröffentlichte das fortschrittlichste Coding-Modell SWE-2, trainiert über Kimi K33 mit 2,8 Billionen Parametern, und erreichte bei FrontierCode 1.1 Main 50,0 %, nur knap...","url":"https://www.aioga.com/de/news/cmtym40hx0004roske1zgdkmu/","contentTranslated":true,"sourceHash":"18a41d74053bfbd2","translatedAt":"2026-09-12T21:41:08.998Z"},"pt-BR":{"title":"Cognition lança modelo de codificação SWE-2, chegando perto do estado da arte com menor custo","summary":"A Cognition lançou o modelo de codificação mais avançado SWE-2, baseado no Kimi K33 de 2,8T parâmetros pós-treinamento, alcançando 50,0% no FrontierCode 1.1 Main, apenas um pouco abaixo do Fable 5.1 e 64% mais barato.","category":"行业动态","source":"Cognition 模型 / Devin 博客（网页）","aggregationSource":"Hacker News：AI 热帖","pageTitle":"Cognition lança modelo de codificação SWE-2, chegando perto do estado da arte com menor custo - Aioga Notícias de IA","description":"A Cognition lançou o modelo de codificação mais avançado SWE-2, baseado no Kimi K33 de 2,8T parâmetros pós-treinamento, alcançando 50,0% no FrontierCode 1.1 Main, apenas um pouco a...","url":"https://www.aioga.com/pt-BR/news/cmtym40hx0004roske1zgdkmu/","contentTranslated":true,"sourceHash":"18a41d74053bfbd2","translatedAt":"2026-09-12T21:41:14.165Z"},"ru":{"title":"Cognition выпускает модель кодирования SWE-2, чтобы приблизиться к передовым технологиям при меньшей стоимости","summary":"Cognition выпускает передовую модель кодирования SWE-2, основанную на 2,8 трлн параметров Kimi K33 после дообучения, достигшую 50,0% на FrontierCode 1.1 Main, всего на чуть меньше одного балла от Fable 5.1 и при этом на 64% дешевле.","category":"行业动态","source":"Cognition 模型 / Devin 博客（网页）","aggregationSource":"Hacker News：AI 热帖","pageTitle":"Cognition выпускает модель кодирования SWE-2, чтобы приблизиться к передовым технологиям при меньшей стоимости - Aioga Новости ИИ","description":"Cognition выпускает передовую модель кодирования SWE-2, основанную на 2,8 трлн параметров Kimi K33 после дообучения, достигшую 50,0% на FrontierCode 1.1 Main, всего на чуть меньше...","url":"https://www.aioga.com/ru/news/cmtym40hx0004roske1zgdkmu/","contentTranslated":true,"sourceHash":"18a41d74053bfbd2","translatedAt":"2026-09-12T21:41:15.172Z"},"ar":{"title":"Cognition تطلق نموذج الترميز SWE-2 للاقتراب من المقدمة بتكلفة أقل","summary":"أطلقت Cognition نموذج الترميز المتقدم SWE-2، بعد التدريب على Kimi K33 بـ 2.8 تريليون معامل، وحقق 50.0% في FrontierCode 1.1 Main، أقل بأقل من نقطة واحدة فقط من Fable 5.1 وأرخص بنسبة 64%.","category":"行业动态","source":"Cognition 模型 / Devin 博客（网页）","aggregationSource":"Hacker News：AI 热帖","pageTitle":"Cognition تطلق نموذج الترميز SWE-2 للاقتراب من المقدمة بتكلفة أقل - Aioga أخبار الذكاء الاصطناعي","description":"أطلقت Cognition نموذج الترميز المتقدم SWE-2، بعد التدريب على Kimi K33 بـ 2.8 تريليون معامل، وحقق 50.0% في FrontierCode 1.1 Main، أقل بأقل من نقطة واحدة فقط من Fable 5.1 وأرخص بنسبة...","url":"https://www.aioga.com/ar/news/cmtym40hx0004roske1zgdkmu/","contentTranslated":true,"sourceHash":"18a41d74053bfbd2","translatedAt":"2026-09-12T21:41:20.878Z"},"hi":{"title":"कॉग्निशन ने कम लागत पर सीमा तक पहुंचने के लिए SWE-2 कोडिंग मॉडल जारी किया","summary":"कॉग्निशन ने अत्याधुनिक कोडिंग मॉडल SWE-2 जारी किया, जिसे Kimi K33 पर 2.8T मापदंडों के साथ प्रशिक्षित किया गया था, जिससे फ्रंटियरकोड 1.1 मेन पर 50.0% सफलता प्राप्त हुई - कल्पित 5.1 से एक अंक कम और 64% सस्ता।","category":"行业动态","source":"Cognition 模型 / Devin 博客（网页）","aggregationSource":"Hacker News：AI 热帖","pageTitle":"कॉग्निशन ने कम लागत पर सीमा तक पहुंचने के लिए SWE-2 कोडिंग मॉडल जारी किया - Aioga AI समाचार","description":"कॉग्निशन ने अत्याधुनिक कोडिंग मॉडल SWE-2 जारी किया, जिसे Kimi K33 पर 2.8T मापदंडों के साथ प्रशिक्षित किया गया था, जिससे फ्रंटियरकोड 1.1 मेन पर 50.0% सफलता प्राप्त हुई - कल्पित 5.1...","url":"https://www.aioga.com/hi/news/cmtym40hx0004roske1zgdkmu/","contentTranslated":true,"sourceHash":"18a41d74053bfbd2","translatedAt":"2026-09-12T21:41:23.930Z"},"it":{"title":"Cognition rilascia il modello di codifica SWE-2 per avvicinarsi alla frontiera a costi inferiori","summary":"Cognition ha rilasciato il modello di codifica all'avanguardia SWE-2, addestrato su Kimi K33 con parametri 2.8T, ottenendo il 50,0% di successo su FrontierCode 1.1 Main—meno di un punto in meno rispetto a Fable 5.1 e 64% in meno.","category":"行业动态","source":"Cognition 模型 / Devin 博客（网页）","aggregationSource":"Hacker News：AI 热帖","pageTitle":"Cognition rilascia il modello di codifica SWE-2 per avvicinarsi alla frontiera a costi inferiori - Aioga Notizie IA","description":"Cognition ha rilasciato il modello di codifica all'avanguardia SWE-2, addestrato su Kimi K33 con parametri 2.8T, ottenendo il 50,0% di successo su FrontierCode 1.1 Main—meno di un...","url":"https://www.aioga.com/it/news/cmtym40hx0004roske1zgdkmu/","contentTranslated":true,"sourceHash":"18a41d74053bfbd2","translatedAt":"2026-09-12T21:41:32.955Z"},"nl":{"title":"Cognition brengt SWE-2 coderingsmodel uit, nadert voorhoede tegen lagere kosten","summary":"Cognition heeft het geavanceerde coderingsmodel SWE-2 uitgebracht, getraind op basis van de 2,8T-parameter Kimi K33, en behaalde 50,0% in FrontierCode 1.1 Main, slechts minder dan een punt onder Fable 5.1 en 64% goedkoper.","category":"行业动态","source":"Cognition 模型 / Devin 博客（网页）","aggregationSource":"Hacker News：AI 热帖","pageTitle":"Cognition brengt SWE-2 coderingsmodel uit, nadert voorhoede tegen lagere kosten - Aioga AI-nieuws","description":"Cognition heeft het geavanceerde coderingsmodel SWE-2 uitgebracht, getraind op basis van de 2,8T-parameter Kimi K33, en behaalde 50,0% in FrontierCode 1.1 Main, slechts minder dan...","url":"https://www.aioga.com/nl/news/cmtym40hx0004roske1zgdkmu/","contentTranslated":true,"sourceHash":"18a41d74053bfbd2","translatedAt":"2026-09-12T21:41:29.219Z"},"tr":{"title":"Cognition, düşük maliyetle en ön safa yaklaşan SWE-2 kodlama modelini yayımladı","summary":"Cognition, en gelişmiş kodlama modeli SWE-2'yi yayımladı; 2.8T parametreli Kimi K33 üzerinde eğitim sonrası FrontierCode 1.1 Main'de %50.0 başarı elde etti, Fable 5.1'den sadece 1 puan daha düşük ve %64 daha ucuz.","category":"行业动态","source":"Cognition 模型 / Devin 博客（网页）","aggregationSource":"Hacker News：AI 热帖","pageTitle":"Cognition, düşük maliyetle en ön safa yaklaşan SWE-2 kodlama modelini yayımladı - Aioga AI Haberleri","description":"Cognition, en gelişmiş kodlama modeli SWE-2'yi yayımladı; 2.8T parametreli Kimi K33 üzerinde eğitim sonrası FrontierCode 1.1 Main'de %50.0 başarı elde etti, Fable 5.1'den sadece 1...","url":"https://www.aioga.com/tr/news/cmtym40hx0004roske1zgdkmu/","contentTranslated":true,"sourceHash":"18a41d74053bfbd2","translatedAt":"2026-09-12T21:41:38.580Z"},"vi":{"title":"Cognition phát hành mô hình mã hóa SWE-2, tiếp cận tiên tiến với chi phí thấp hơn","summary":"Cognition phát hành mô hình mã hóa tiên tiến SWE-2, dựa trên huấn luyện sau Kimi K33 với 2,8T tham số, đạt 50,0% tại FrontierCode 1.1 Main, chỉ thấp hơn Fable 5.1 chưa đầy một điểm và rẻ hơn 64%.","category":"行业动态","source":"Cognition 模型 / Devin 博客（网页）","aggregationSource":"Hacker News：AI 热帖","pageTitle":"Cognition phát hành mô hình mã hóa SWE-2, tiếp cận tiên tiến với chi phí thấp hơn - Tin tức AI Aioga","description":"Cognition phát hành mô hình mã hóa tiên tiến SWE-2, dựa trên huấn luyện sau Kimi K33 với 2,8T tham số, đạt 50,0% tại FrontierCode 1.1 Main, chỉ thấp hơn Fable 5.1 chưa đầy một điểm...","url":"https://www.aioga.com/vi/news/cmtym40hx0004roske1zgdkmu/","contentTranslated":true,"sourceHash":"18a41d74053bfbd2","translatedAt":"2026-09-12T21:41:38.872Z"},"id":{"title":"Cognition merilis model pengkodean SWE-2 untuk mendekati frontier dengan biaya lebih rendah","summary":"Cognition merilis model pengkodean mutakhir SWE-2, yang dilatih pada Kimi K33 dengan parameter 2,8T, mencapai keberhasilan 50,0% pada FrontierCode 1.1 Main—kurang dari satu poin lebih rendah dari Fable 5.1 dan 64% lebih murah.","category":"行业动态","source":"Cognition 模型 / Devin 博客（网页）","aggregationSource":"Hacker News：AI 热帖","pageTitle":"Cognition merilis model pengkodean SWE-2 untuk mendekati frontier dengan biaya lebih rendah - Berita AI Aioga","description":"Cognition merilis model pengkodean mutakhir SWE-2, yang dilatih pada Kimi K33 dengan parameter 2,8T, mencapai keberhasilan 50,0% pada FrontierCode 1.1 Main—kurang dari satu poin le...","url":"https://www.aioga.com/id/news/cmtym40hx0004roske1zgdkmu/","contentTranslated":true,"sourceHash":"18a41d74053bfbd2","translatedAt":"2026-09-12T21:41:47.609Z"},"th":{"title":"Cognition ปล่อยโมเดลการเข้ารหัส SWE-2 เพื่อเข้าถึงขอบเขตใหม่ด้วยต้นทุนที่ต่ํากว่า","summary":"Cognition ได้ปล่อยโมเดลโค้ดล่าสุด SWE-2 ซึ่งฝึกบน Kimi K33 ด้วยพารามิเตอร์ 2.8T โดยประสบความสําเร็จ 50.0% บน FrontierCode 1.1 Main — ต่ํากว่า Fable 5.1 ไม่ถึงหนึ่งจุดและถูกกว่า 64%","category":"行业动态","source":"Cognition 模型 / Devin 博客（网页）","aggregationSource":"Hacker News：AI 热帖","pageTitle":"Cognition ปล่อยโมเดลการเข้ารหัส SWE-2 เพื่อเข้าถึงขอบเขตใหม่ด้วยต้นทุนที่ต่ํากว่า - ข่าว AI Aioga","description":"Cognition ได้ปล่อยโมเดลโค้ดล่าสุด SWE-2 ซึ่งฝึกบน Kimi K33 ด้วยพารามิเตอร์ 2.8T โดยประสบความสําเร็จ 50.0% บน FrontierCode 1.1 Main — ต่ํากว่า Fable 5.1 ไม่ถึงหนึ่งจุดและถูกกว่า 64%","url":"https://www.aioga.com/th/news/cmtym40hx0004roske1zgdkmu/","contentTranslated":true,"sourceHash":"18a41d74053bfbd2","translatedAt":"2026-09-12T21:41:47.654Z"},"pl":{"title":"Cognition wydaje model kodowania SWE-2, aby tańszym kosztem zbliżyć się do czołówki","summary":"Cognition wydało najnowocześniejszy model kodowania SWE-2, oparty na 2,8T parametrach po treningu modelu Kimi K33. W FrontierCode 1.1 Main osiąga 50,0%, co jest tylko niecały punkt niżej od Fable 5.1 i kosztuje o 64% mniej.","category":"行业动态","source":"Cognition 模型 / Devin 博客（网页）","aggregationSource":"Hacker News：AI 热帖","pageTitle":"Cognition wydaje model kodowania SWE-2, aby tańszym kosztem zbliżyć się do czołówki - Aioga Wiadomości AI","description":"Cognition wydało najnowocześniejszy model kodowania SWE-2, oparty na 2,8T parametrach po treningu modelu Kimi K33. W FrontierCode 1.1 Main osiąga 50,0%, co jest tylko niecały punkt...","url":"https://www.aioga.com/pl/news/cmtym40hx0004roske1zgdkmu/","contentTranslated":true,"sourceHash":"18a41d74053bfbd2","translatedAt":"2026-09-12T21:41:54.819Z"}},"evidenceTier":"verified-news","reviewStatus":"editorial-selected","indexable":true,"editorialCover":"/page-visuals/topic-timeline.png"}}