{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-09-28T06:03:00.468Z","headline":"Fireworks Research 发布 Ember-1：以少 40% 的 token 达到 Kimi K3 同等质量","description":"Fireworks Research 发布专门化模型 Ember-1，基于 Kimi K3 训练，在保持同等质量的同时减少约 40% 的 token 消耗。","url":"https://www.aioga.com/news/cmuk7v5sc1el9ro9hcldb15xp/","mainEntityOfPage":"https://www.aioga.com/news/cmuk7v5sc1el9ro9hcldb15xp/","datePublished":"2026-09-27T19:09:18.000Z","dateModified":"2026-09-27T19:09:18.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://fireworks.ai/blog/ember-1","https://aihot.news/items/cmuk7v5sc1el9ro9hcldb15xp"],"canonicalUrl":"https://www.aioga.com/news/cmuk7v5sc1el9ro9hcldb15xp/","directAnswer":{"@type":"Answer","text":"Fireworks Research 发布专门化模型 Ember-1。来源称，该模型基于 Kimi K3 训练，在保持同等质量的同时减少约 40% 的 token 消耗，现已提供，并被称为 Fireworks Research 系列模型的首款产品。","url":"https://www.aioga.com/news/cmuk7v5sc1el9ro9hcldb15xp/","dateCreated":"2026-09-27T19:09:18.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"fireworks.ai source article","url":"https://fireworks.ai/blog/ember-1","datePublished":"2026-09-27T19:09:18.000Z","provider":{"@type":"Organization","name":"fireworks.ai","url":"https://fireworks.ai/blog/ember-1"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.news/items/cmuk7v5sc1el9ro9hcldb15xp","datePublished":"2026-09-27T19:09:18.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.news/items/cmuk7v5sc1el9ro9hcldb15xp"}}],"aggregationSource":"Hacker News 热门（buzzing.cc 中文翻译）","originalPublisher":{"name":"fireworks.ai","url":"https://fireworks.ai/blog/ember-1"},"geoDeepAnswer":null,"article":{"id":"cmuk7v5sc1el9ro9hcldb15xp","slug":"cmuk7v5sc1el9ro9hcldb15xp","url":"https://www.aioga.com/news/cmuk7v5sc1el9ro9hcldb15xp/","title":"Fireworks Research 发布 Ember-1：以少 40% 的 token 达到 Kimi K3 同等质量","title_en":"","summary":"Fireworks Research 发布专门化模型 Ember-1，基于 Kimi K3 训练，在保持同等质量的同时减少约 40% 的 token 消耗。","source":"Hacker News 热门（buzzing.cc 中文翻译）","sourceUrl":"https://fireworks.ai/blog/ember-1","aiHotUrl":"https://aihot.news/items/cmuk7v5sc1el9ro9hcldb15xp","publishedAt":"2026-09-27T19:09:18.000Z","category":"行业动态","score":58,"selected":false,"articleBody":["Join us for our inaugural conference, Forge 2026","Ember-1：/models/fireworks/ember-1 is a new specialized model from Fireworks Research that delivers Kimi K3’s quality with 40% fewer tokens. Built on Kimi K3, it learned to cut unnecessary reasoning while keeping the thinking that matters. We tested it on external benchmarks, in live customer A/B tests, and on our own coding and agent workloads, and quality held up in every setting. Available today, Ember-1 kicks off an ongoing series of specialized models by Fireworks, shaped by what developers want next. Ember is just the start of what you could build with the Fireworks Training platform.","We heard from users that they needed K3’s coding capabilities at a lower cost, because its long reasoning traces made automated coding expensive at scale. Turning down K3's reasoning effort didn't solve this. Lower effort settings gave up too much quality. To keep the quality and cut the tokens, the model had to learn to reason more efficiently, and that meant training it.","Getting there took serious research. Our team ran more than 50 training experiments and over 200 evaluations, and developed new training algorithms along the way to shorten reasoning without losing accuracy. We did it all on Fireworks Serverless Training：/training#training-api:~:text=RUN%20THE%20LOOP-,Training%20API,-For%20ML%20researchers. Because we didn’t have to provision or manage GPUs, we could launch experiments as soon as we had an idea, pay only for what we ran, and move from research to launch in a fraction of the usual time and cost.","We trained across a broad set of tasks so the token savings would carry over to many workloads. We then evaluated Ember-1 on the Specialized Intelligence Index：/specialized-intelligence-index/, public benchmarks, and live production traffic to confirm it used fewer tokens with no drop in quality. Ember-1 is Fireworks own model and the first in a series of models from Fireworks Research.","Reasoning models like Kimi K3 spend the majority of their generated tokens, sometimes more than 90%, on internal reasoning rather than the answer itself. This thinking structure is expensive on a single request, but it gets much worse in multi-turn agentic workloads. Every turn replays all prior reasoning back to the model, so context grows roughly quadratically with the number of turns. Long reasoning traces from early turns get re-read (and re-billed) on every subsequent call.","Is all that reasoning actually necessary? Our experiments said no. The reasoning Kimi K3 emits is far longer than the task requires, and the excess can be removed without touching the answer. This was how we created Ember-1, an economic version of Kimi K3 built from specialized intelligence.","Not all of K3's reasoning is wasted. Some of it is self-reflection: revisiting an assumption, responding to feedback, or tracing an outcome back to an earlier decision can help the model recover from mistakes. The opportunity is to preserve this ability while reducing unnecessary reasoning and escaping unproductive loops. We believe that learning from tasks and environment feedback can teach the model to reason more efficiently while maintaining its capabilities.","For agentic tasks, this learning extends across the interaction. The model explores possible actions, incorporates new observations, and refines its reasoning as it progresses. Feedback connects decisions to their consequences, encouraging useful reflection throughout the task.","We carried these insights into a training collection spanning mathematics, coding, instruction following, conversation, search, tool use, and software engineering, covering both standalone problems and extended interactions to enforce adaptation to observations and outcome. Task feedback guides on-policy planning and learning, with an emphasis on preserving capability across this range of settings.","Results on public benchmarks and live A/B tests support this direction: across seven benchmarks and two customer production traffics, Kimi K3’s reasoning could be shortened by 35–50% without sacrificing accuracy. The internalized behavior also shows restrained token use on unsuccessful attempts, reducing prolonged, unproductive reasoning.","Earlier this week, we introduced the Specialized Intelligence Index (SII)：/specialized-intelligence-index/ to benchmark open, closed, and specialized models against real-world tasks created by industry experts.","We evaluated Ember-1 on Doximity’s Bedside Bench, a physician-validated benchmark spanning 500 clinical cases across 10 specialized categories.","The result? Ember-1 set a new pareto frontier for Bedside Bench across both open and closed models including GPT.5 Sol, GPT-6 Astra, and Claude Opus 5 on cost/task.","We also evaluated Ember-1 on the quality-vs-cost frontier across some other industry benchmarks. We computed per-benchmark cost using the public Kimi K3 API pricing (uncached input $3/M tokens, cached input $0.30/M, output $15/M) and plotted it against pass rate for three arms: K3 at reasoning effort low , K3 at reasoning effort high, K3 at reasoning effort max (default) , and Ember-1. Across every benchmark with more than 50 test samples, Ember-1 sits on or near the Pareto frontier, matching K3-max quality at a fraction of the cost, and strictly dominating K3-low. We also analyzed GPT-6 Astra, Claude Opus-5, GPT 5.6 Sol and GLM 5.3, and found that Ember-1 was a leader on the Pareto Frontier.","We took a double-click on the results directly comparing Ember-1 to the original K3, and found the following results:","The most cost optimized way to run K3 is no longer to make it think less, but it’s to run Ember-1 the model that learned to think efficiently.","Benchmarks only tell you so much. Like what we found in the Specialized Intelligence Index results, we wanted to test the model on more real workloads, and to test the model using production traffic. The real test is often whether the model holds up on production traffic, in products users depend on.","We ran live A/B tests with two customers on their production coding workloads. In both cases, Ember-1 delivered impressive token savings, approximately 35% fewer tokens per task at comparable quality . Most of the downstream product metrics held or improved including task completion, success scores, and failure rates all moving in the right direction at substantially lower token cost. Following the A/B tests, one customer is now running Ember-1 in live production, with plans to scale it up to replace the base model entirely.","A large part of Fireworks internal coding/cowork traffic is powered by our own inference service. Before any customer saw the model, we put Ember-1 into Fireworks' internal and let our own developers use it for everyday coding work including things like vibe testing at scale on real tasks.","The outcome we're proudest of: no news. No news is good news. Developers carried on their coding workloads without noticing the switch, while consuming substantially fewer tokens. For a model whose entire value proposition is \"same answers, fewer tokens,\" an invisible rollout on internal traffic is the strongest possible signal.","Ember-1 is rolling out as a serving option alongside the base Kimi K3 model as a Research Preview release on Serverless.To support the rapidly growing open-source ecosystem, we're introducing research releases to give developers two-week serverless access to new research models, making them permanent based on community demand. For agentic coding and other workloads where reasoning tokens account for most of the cost, it delivers the same quality at roughly half the token cost.","Fireworks Research will continue to push the frontier of model efficiency by bringing specialized intelligence to more Ember models to enable you to deploy the most economic models, and reduce your token spend.Token efficiency is becoming a theme of Fireworks.","Looking to take Ember-1 one step further, and optimize it for your use case? We are also launching training support for Ember-1, enabling enterprises to build customized, token-efficient models tailored to their needs with their own data. The future of open-models is specialized models trained on your specific workload.","Trying Ember-1：/models/fireworks/ember-1 out on your workloads? We'd love to hear about your experience, so tag us on X (@FireworksAI_HQ) and let us know what you're building!"],"articleImages":[{"sourceUrl":"https://fireworks.ai/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2Fpv37i0yn%2Fproduction%2F1ae9d150288b660e94ce10c3e6d230f97742f54e-5000x2813.png%3Fauto%3Dformat&w=3840&q=75","alt":"Graph showing 40% fewer tokens than Kimi K3","afterParagraph":0,"url":"/media/articles/cmuemnn9m07yproynbzmmumd3/3b715432e5b27ebf.webp"},{"sourceUrl":"https://fireworks.ai/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2Fpv37i0yn%2Fproduction%2Fdf5a6f5563b24d981bd4c119739a18be88e17146-1810x1208.png%3Fauto%3Dformat&w=3840&q=75","alt":"Figure of Cost/Task SII","afterParagraph":13,"url":"/media/articles/cmuemnn9m07yproynbzmmumd3/661431d7db66d314.webp"},{"sourceUrl":"https://fireworks.ai/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2Fpv37i0yn%2Fproduction%2F7dc4ddef2dc6aadfc3590fc9d4c45f499f743200-1564x1048.png%3Fauto%3Dformat&w=3840&q=75","alt":"Figure 2: Score vs. Duration Chart on Bedside Bench SII","afterParagraph":13,"url":"/media/articles/cmuemnn9m07yproynbzmmumd3/ca996afc2f3d95fb.webp"},{"sourceUrl":"https://fireworks.ai/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2Fpv37i0yn%2Fproduction%2F7fd5a909f34a9100146cb676368a371c34d388fe-2757x1777.png%3Fauto%3Dformat&w=3840&q=75","alt":"Figure 3: Average of 5 Industry Benchmarks on Open and Closed Model Cost/Task","afterParagraph":14,"url":"/media/articles/cmuemnn9m07yproynbzmmumd3/cab4710c3c240284.webp"},{"sourceUrl":"https://fireworks.ai/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Ffireworks-bw.709eb7d4.png&w=3840&q=75&dpl=dpl_FHvX2CCx7sCUwQTx47CmeoPjkzEC","alt":"","afterParagraph":24,"url":"/media/articles/cmuemnn9m07yproynbzmmumd3/9c322a125982a10e.webp"}],"mediaStatus":"ok","articleBodyZh":["加入我们，参加我们的首届大会 Forge 2026","Ember-1：/models/fireworks/ember-1 是 Fireworks Research 推出的新专用模型，它在使用的 token 数量减少 40% 的情况下，实现了 Kimi K3 的质量。基于 Kimi K3 构建，它学会了削减不必要的推理，同时保留关键思考。我们在外部基准测试、实时客户 A/B 测试以及我们自己的编码和代理工作负载中进行了测试，质量在每种环境下都保持稳定。Ember-1 今天即可使用，它开启了 Fireworks 持续推出的一系列专用模型，这些模型由开发者的下一步需求塑造。Ember 只是您可以在 Fireworks Training 平台上构建的开始。","我们听到用户反馈，他们需要 K3 的编码能力但成本更低，因为其长时间的推理路径使得大规模自动化编码变得昂贵。降低 K3 的推理强度并不能解决这个问题。降低推理强度会损失过多质量。为了保持质量并减少 token，模型必须学会更高效地推理，这就意味着必须进行训练。","实现这一目标需要严谨的研究。我们的团队进行了超过 50 项训练实验和 200 多次评估，并在此过程中开发了新的训练算法，以缩短推理时间而不损失准确性。我们完全在 Fireworks Serverless Training：/training#training-api:~:text=RUN%20THE%20LOOP-,Training%20API,-For%20ML%20researchers 上完成。因为无需配置或管理 GPU，我们可以在有想法的第一时间启动实验，仅为实际运行付费，并在比通常时间和成本少得多的情况下从研究进入发布阶段。","我们在广泛的任务集合上进行了训练，以确保 token 节省能够应用于多种工作负载。随后我们在 Specialized Intelligence Index：/specialized-intelligence-index/、公开基准测试以及实时生产流量中评估 Ember-1，以确认其在使用更少 token 的情况下没有质量下降。Ember-1 是 Fireworks 自有模型，也是 Fireworks Research 系列模型中的第一个。","像 Kimi K3 这样的推理模型，将其生成的大部分标记，有时超过 90%，用于内部推理，而不是直接给出答案。这种思考结构在单次请求中代价高昂，但在多轮的自主任务中情况会更糟。每一轮都会将之前的所有推理传回模型，因此上下文随着轮数的增加大致呈二次增长。早期轮次的长推理轨迹会在每次后续调用中被重新读取（并重新计费）。","所有这些推理真的有必要吗？我们的实验表明没有。Kimi K3 输出的推理远比任务需求长，多余部分可以在不触及答案的情况下去除。这就是我们如何创建 Ember-1 的方法，它是一个经济版的 Kimi K3，由专门的智能构建而成。","并非所有 K3 的推理都是浪费。有些是自我反思：重新审视假设、回应反馈或将结果追溯到早期决策可以帮助模型从错误中恢复。机会在于保留这种能力，同时减少不必要的推理并摆脱无效循环。我们认为，从任务和环境反馈中学习可以教会模型更高效地推理，同时保持其能力。","对于自主任务，这种学习延伸到整个交互过程。模型探索可能的行动，整合新的观察，并在执行过程中不断优化推理。反馈将决策与其后果连接起来，激励在整个任务过程中进行有用的反思。","我们将这些见解应用到一个涵盖数学、编码、指令执行、对话、搜索、工具使用和软件工程的训练集合中，覆盖独立问题和延伸交互，以强化对观察和结果的适应性。任务反馈指导策略性规划和学习，强调在各种环境中保持能力。","公共基准测试和实时 A/B 测试的结果支持这一方向：在七个基准和两个客户生产流量中，Kimi K3 的推理可以缩短 35–50% 而不影响准确性。内化行为还显示，在失败尝试中控制标记使用，减少了冗长且无效的推理。","本周早些时候，我们介绍了专门智能指数（Specialized Intelligence Index, SII）：/specialized-intelligence-index/，用于在行业专家创建的真实任务中基准测试开放模型、封闭模型和专门模型。","我们在Doximity的床边基准（Bedside Bench）上评估了Ember-1，这是一个由医生验证的基准，涵盖了10个专业类别中的500个临床案例。","结果如何？Ember-1在床边基准上创造了新的帕累托前沿，无论是开放模型还是封闭模型，包括GPT-5 Sol、GPT-6 Astra和Claude Opus 5，在成本/任务方面均表现出色。","我们还在一些其他行业基准上评估了Ember-1的质量与成本前沿。我们使用公开的Kimi K3 API定价（未缓存输入 $3/百万 token，缓存输入 $0.30/百万，输出 $15/百万）计算每个基准的成本，并将其与三种方案的通过率绘制对比：K3在低推理强度下、K3在高推理强度下、K3在最大推理强度下（默认）以及Ember-1。在每个测试样本超过50的基准中，Ember-1都位于或接近帕累托前沿，以极低的成本匹配K3-max的质量，并严格优于K3-low。我们还分析了GPT-6 Astra、Claude Opus-5、GPT 5.6 Sol和GLM 5.3，发现Ember-1在帕累托前沿上处于领先地位。","我们对直接比较Ember-1与原始K3的结果进行了深入分析，并发现了以下结果：","运行K3成本最优的方法不再是让其减少思考，而是运行Ember-1，这是一种学会高效思考的模型。","基准测试只能告诉你那么多。正如我们在专门智能指数结果中发现的那样，我们希望在更多真实工作负载上测试模型，并使用生产流量测试模型。真正的考验通常是模型在生产流量中是否稳健，以及用户依赖的产品中表现如何。","我们对两位客户的生产编码工作负载进行了实时 A/B 测试。在这两种情况下，Ember-1 都展示了令人印象深刻的令牌节省，大约每个任务使用的令牌减少了 35%，同时质量相当。大部分下游产品指标保持不变或有所改善，包括任务完成率、成功评分和失败率，所有指标都朝着正确方向发展，同时令牌消耗大幅降低。在 A/B 测试之后，其中一位客户现在正在生产环境中使用 Ember-1，并计划扩大规模以完全替代基础模型。","Fireworks 内部的大部分编码/协作流量由我们自己的推理服务提供支持。在任何客户看到模型之前，我们就将 Ember-1 部署到 Fireworks 内部，并让我们自己的开发人员在日常编码工作中使用它，包括在真实任务上进行大规模的 Vibe 测试等。","我们最自豪的结果：没有任何新闻。没有新闻就是好消息。开发人员在不注意到切换的情况下继续进行他们的编码工作，同时消耗的令牌显著减少。对于一个其全部价值主张是“相同答案，更少令牌”的模型来说，在内部流量上的隐形部署是最有力的信号。","Ember-1 正作为服务选项与基础 Kimi K3 模型一起在 Serverless 上以研究预览版的形式推出。为了支持快速增长的开源生态系统，我们引入了研究发布，使开发人员能够在两周内无服务器访问新的研究模型，并根据社区需求将其转为永久使用。对于推理令牌占大部分成本的智能编码和其他工作负载，它以大约一半的令牌成本提供相同的质量。","Fireworks Research 将继续推动模型效率的前沿，通过将专用智能引入更多 Ember 模型，使您能够部署最经济的模型并减少令牌开支。令牌效率正在成为 Fireworks 的一个主题。","想将 Ember-1 推进到下一步，并针对您的用例进行优化吗？我们也正在启动 Ember-1 的训练支持，使企业能够使用自己的数据构建定制化、令牌高效的模型，以满足其需求。开放模型的未来是基于您特定工作负载训练的专用模型。","在您的工作负载上尝试 Ember-1：/models/fireworks/ember-1 吗？我们很想听听您的体验，所以请在 X 上标记我们 (@FireworksAI_HQ) 并告诉我们您正在构建什么！"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Fireworks Research 发布专门化模型 Ember-1。来源称，该模型基于 Kimi K3 训练，在保持同等质量的同时减少约 40% 的 token 消耗，现已提供，并被称为 Fireworks Research 系列模型的首款产品。","background":"Fireworks 表示，用户希望以更低成本使用 Kimi K3 的编码能力，而降低推理力度会损失过多质量。团队进行了超过 50 次训练实验和超过 200 次评估，并在公开基准、生产流量及多类任务中测试 Ember-1。","viewpoint":"Aioga 判断：该发布的核心价值主张是通过专门化训练提升推理效率，而非简单降低推理力度。来源明确说明部分自我反思可能有助于纠错，但仅表示希望通过训练保留这类能力，不能据此断言 Ember-1 已被证实具备该能力。","implications":"可能影响：若来源所述结果能在独立场景中复现，Ember-1 可能降低部分长推理及多轮代理任务的 token 开销。现有材料不足以证明其在所有任务中都与 Kimi K3 等质，也不代表成本和效果已获外部统一验证。","nextStep":"后续观察：应关注独立评测是否复核约 40% 的 token 节省及质量表现，并核实这些结果在公开基准、实际生产流量、编码任务和代理任务中的适用范围，同时观察后续专门化模型是否发布。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-09-27T20:12:50.396Z","sourceHash":"b9680d0e97cbbc04","review":{"approved":true,"groundedness":96,"clarity":93,"duplicationRisk":12,"blockingIssues":[],"notes":[]},"validation":{"passed":true,"mode":"ai-auto","revisions":1,"checks":["schema","length","source-attribution","editorial-labels","inference-boundary","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","Hacker News 热门（buzzing.cc 中文翻译）"],"translations":{"zh-CN":{"title":"Fireworks Research 发布 Ember-1，以更少推理 token 保持 Kimi K3 质量","summary":"Fireworks Research 发布基于 Kimi K3 的专用模型 Ember-1，以约少 40% 的 token 达到与 Kimi K3 相当的质量，今日以 Research Preview 形式在 Serverless 上线。","category":"行业动态","source":"Fireworks AI（网页）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks Research 发布 Ember-1，以更少推理 token 保持 Kimi K3 质量 - Aioga AI资讯","description":"Fireworks Research 发布基于 Kimi K3 的专用模型 Ember-1，以约少 40% 的 token 达到与 Kimi K3 相当的质量，今日以 Research Preview 形式在 Serverless 上线。","url":"https://www.aioga.com/news/cmuk7v5sc1el9ro9hcldb15xp/","articleBody":["加入我们，参加我们的首届大会 Forge 2026","Ember-1：/models/fireworks/ember-1 是 Fireworks Research 推出的新专用模型，它在使用的 token 数量减少 40% 的情况下，实现了 Kimi K3 的质量。基于 Kimi K3 构建，它学会了削减不必要的推理，同时保留关键思考。我们在外部基准测试、实时客户 A/B 测试以及我们自己的编码和代理工作负载中进行了测试，质量在每种环境下都保持稳定。Ember-1 今天即可使用，它开启了 Fireworks 持续推出的一系列专用模型，这些模型由开发者的下一步需求塑造。Ember 只是您可以在 Fireworks Training 平台上构建的开始。","我们听到用户反馈，他们需要 K3 的编码能力但成本更低，因为其长时间的推理路径使得大规模自动化编码变得昂贵。降低 K3 的推理强度并不能解决这个问题。降低推理强度会损失过多质量。为了保持质量并减少 token，模型必须学会更高效地推理，这就意味着必须进行训练。","实现这一目标需要严谨的研究。我们的团队进行了超过 50 项训练实验和 200 多次评估，并在此过程中开发了新的训练算法，以缩短推理时间而不损失准确性。我们完全在 Fireworks Serverless Training：/training#training-api:~:text=RUN%20THE%20LOOP-,Training%20API,-For%20ML%20researchers 上完成。因为无需配置或管理 GPU，我们可以在有想法的第一时间启动实验，仅为实际运行付费，并在比通常时间和成本少得多的情况下从研究进入发布阶段。","我们在广泛的任务集合上进行了训练，以确保 token 节省能够应用于多种工作负载。随后我们在 Specialized Intelligence Index：/specialized-intelligence-index/、公开基准测试以及实时生产流量中评估 Ember-1，以确认其在使用更少 token 的情况下没有质量下降。Ember-1 是 Fireworks 自有模型，也是 Fireworks Research 系列模型中的第一个。","像 Kimi K3 这样的推理模型，将其生成的大部分标记，有时超过 90%，用于内部推理，而不是直接给出答案。这种思考结构在单次请求中代价高昂，但在多轮的自主任务中情况会更糟。每一轮都会将之前的所有推理传回模型，因此上下文随着轮数的增加大致呈二次增长。早期轮次的长推理轨迹会在每次后续调用中被重新读取（并重新计费）。","所有这些推理真的有必要吗？我们的实验表明没有。Kimi K3 输出的推理远比任务需求长，多余部分可以在不触及答案的情况下去除。这就是我们如何创建 Ember-1 的方法，它是一个经济版的 Kimi K3，由专门的智能构建而成。","并非所有 K3 的推理都是浪费。有些是自我反思：重新审视假设、回应反馈或将结果追溯到早期决策可以帮助模型从错误中恢复。机会在于保留这种能力，同时减少不必要的推理并摆脱无效循环。我们认为，从任务和环境反馈中学习可以教会模型更高效地推理，同时保持其能力。","对于自主任务，这种学习延伸到整个交互过程。模型探索可能的行动，整合新的观察，并在执行过程中不断优化推理。反馈将决策与其后果连接起来，激励在整个任务过程中进行有用的反思。","我们将这些见解应用到一个涵盖数学、编码、指令执行、对话、搜索、工具使用和软件工程的训练集合中，覆盖独立问题和延伸交互，以强化对观察和结果的适应性。任务反馈指导策略性规划和学习，强调在各种环境中保持能力。","公共基准测试和实时 A/B 测试的结果支持这一方向：在七个基准和两个客户生产流量中，Kimi K3 的推理可以缩短 35–50% 而不影响准确性。内化行为还显示，在失败尝试中控制标记使用，减少了冗长且无效的推理。","本周早些时候，我们介绍了专门智能指数（Specialized Intelligence Index, SII）：/specialized-intelligence-index/，用于在行业专家创建的真实任务中基准测试开放模型、封闭模型和专门模型。","我们在Doximity的床边基准（Bedside Bench）上评估了Ember-1，这是一个由医生验证的基准，涵盖了10个专业类别中的500个临床案例。","结果如何？Ember-1在床边基准上创造了新的帕累托前沿，无论是开放模型还是封闭模型，包括GPT-5 Sol、GPT-6 Astra和Claude Opus 5，在成本/任务方面均表现出色。","我们还在一些其他行业基准上评估了Ember-1的质量与成本前沿。我们使用公开的Kimi K3 API定价（未缓存输入 $3/百万 token，缓存输入 $0.30/百万，输出 $15/百万）计算每个基准的成本，并将其与三种方案的通过率绘制对比：K3在低推理强度下、K3在高推理强度下、K3在最大推理强度下（默认）以及Ember-1。在每个测试样本超过50的基准中，Ember-1都位于或接近帕累托前沿，以极低的成本匹配K3-max的质量，并严格优于K3-low。我们还分析了GPT-6 Astra、Claude Opus-5、GPT 5.6 Sol和GLM 5.3，发现Ember-1在帕累托前沿上处于领先地位。","我们对直接比较Ember-1与原始K3的结果进行了深入分析，并发现了以下结果：","运行K3成本最优的方法不再是让其减少思考，而是运行Ember-1，这是一种学会高效思考的模型。","基准测试只能告诉你那么多。正如我们在专门智能指数结果中发现的那样，我们希望在更多真实工作负载上测试模型，并使用生产流量测试模型。真正的考验通常是模型在生产流量中是否稳健，以及用户依赖的产品中表现如何。","我们对两位客户的生产编码工作负载进行了实时 A/B 测试。在这两种情况下，Ember-1 都展示了令人印象深刻的令牌节省，大约每个任务使用的令牌减少了 35%，同时质量相当。大部分下游产品指标保持不变或有所改善，包括任务完成率、成功评分和失败率，所有指标都朝着正确方向发展，同时令牌消耗大幅降低。在 A/B 测试之后，其中一位客户现在正在生产环境中使用 Ember-1，并计划扩大规模以完全替代基础模型。","Fireworks 内部的大部分编码/协作流量由我们自己的推理服务提供支持。在任何客户看到模型之前，我们就将 Ember-1 部署到 Fireworks 内部，并让我们自己的开发人员在日常编码工作中使用它，包括在真实任务上进行大规模的 Vibe 测试等。","我们最自豪的结果：没有任何新闻。没有新闻就是好消息。开发人员在不注意到切换的情况下继续进行他们的编码工作，同时消耗的令牌显著减少。对于一个其全部价值主张是“相同答案，更少令牌”的模型来说，在内部流量上的隐形部署是最有力的信号。","Ember-1 正作为服务选项与基础 Kimi K3 模型一起在 Serverless 上以研究预览版的形式推出。为了支持快速增长的开源生态系统，我们引入了研究发布，使开发人员能够在两周内无服务器访问新的研究模型，并根据社区需求将其转为永久使用。对于推理令牌占大部分成本的智能编码和其他工作负载，它以大约一半的令牌成本提供相同的质量。","Fireworks Research 将继续推动模型效率的前沿，通过将专用智能引入更多 Ember 模型，使您能够部署最经济的模型并减少令牌开支。令牌效率正在成为 Fireworks 的一个主题。","想将 Ember-1 推进到下一步，并针对您的用例进行优化吗？我们也正在启动 Ember-1 的训练支持，使企业能够使用自己的数据构建定制化、令牌高效的模型，以满足其需求。开放模型的未来是基于您特定工作负载训练的专用模型。","在您的工作负载上尝试 Ember-1：/models/fireworks/ember-1 吗？我们很想听听您的体验，所以请在 X 上标记我们 (@FireworksAI_HQ) 并告诉我们您正在构建什么！"]},"en":{"title":"Fireworks Research releases Ember-1: achieves Kimi K3 quality with 40% fewer tokens","summary":"Fireworks Research releases the specialized model Ember-1, trained based on Kimi K3, reducing token consumption by about 40% while maintaining the same quality.","category":"Industry","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks Research releases Ember-1: achieves Kimi K3 quality with 40% fewer tokens - Aioga AI News","description":"Fireworks Research releases the specialized model Ember-1, trained based on Kimi K3, reducing token consumption by about 40% while maintaining the same quality.","url":"https://www.aioga.com/en/news/cmuk7v5sc1el9ro9hcldb15xp/","contentTranslated":true,"sourceHash":"6f7c197946e35ca6","translatedAt":"2026-09-27T20:02:16.515Z"},"ja":{"title":"Fireworks ResearchがEmber-1をリリース:トークン数を40%減らしながらKimi K3レベルの品質を達成する","summary":"Fireworks ResearchはKimi K3をベースにした特殊モデルEmber-1をリリースし、トークン消費を約40%削減しつつ品質も維持します。","category":"業界動向","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks ResearchがEmber-1をリリース:トークン数を40%減らしながらKimi K3レベルの品質を達成する - Aioga AIニュース","description":"Fireworks ResearchはKimi K3をベースにした特殊モデルEmber-1をリリースし、トークン消費を約40%削減しつつ品質も維持します。","url":"https://www.aioga.com/ja/news/cmuk7v5sc1el9ro9hcldb15xp/","contentTranslated":true,"sourceHash":"6f7c197946e35ca6","translatedAt":"2026-09-27T20:02:19.626Z"},"ko":{"title":"Fireworks Research가 Ember-1을 출시: 40% 적은 토큰으로 Kimi K3급 품질 달성","summary":"Fireworks Research는 Kimi K3를 기반으로 훈련된 특수 모델인 Ember-1을 출시했으며, 이는 동일한 품질을 유지하면서 토큰 소비를 약 40% 줄여줍니다.","category":"업계 동향","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks Research가 Ember-1을 출시: 40% 적은 토큰으로 Kimi K3급 품질 달성 - Aioga AI 뉴스","description":"Fireworks Research는 Kimi K3를 기반으로 훈련된 특수 모델인 Ember-1을 출시했으며, 이는 동일한 품질을 유지하면서 토큰 소비를 약 40% 줄여줍니다.","url":"https://www.aioga.com/ko/news/cmuk7v5sc1el9ro9hcldb15xp/","contentTranslated":true,"sourceHash":"6f7c197946e35ca6","translatedAt":"2026-09-27T20:02:28.469Z"},"es":{"title":"Fireworks Research lanza Ember-1: misma calidad que Kimi K3 usando 40% menos de tokens","summary":"Fireworks Research lanza el modelo especializado Ember-1, entrenado basado en Kimi K3, logrando la misma calidad con alrededor de un 40% menos de consumo de tokens.","category":"Industria","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks Research lanza Ember-1: misma calidad que Kimi K3 usando 40% menos de tokens - Aioga Noticias de IA","description":"Fireworks Research lanza el modelo especializado Ember-1, entrenado basado en Kimi K3, logrando la misma calidad con alrededor de un 40% menos de consumo de tokens.","url":"https://www.aioga.com/es/news/cmuk7v5sc1el9ro9hcldb15xp/","contentTranslated":true,"sourceHash":"6f7c197946e35ca6","translatedAt":"2026-09-27T20:02:27.635Z"},"fr":{"title":"Fireworks Research lance Ember-1 : atteindre une qualité de niveau Kimi K3 avec 40 % de jetons en moins","summary":"Fireworks Research a lancé le modèle spécialisé Ember-1, entraîné sur Kimi K3, qui réduit la consommation de jetons d’environ 40 % tout en conservant la même qualité.","category":"Industrie","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks Research lance Ember-1 : atteindre une qualité de niveau Kimi K3 avec 40 % de jetons en moins - Aioga Actualités IA","description":"Fireworks Research a lancé le modèle spécialisé Ember-1, entraîné sur Kimi K3, qui réduit la consommation de jetons d’environ 40 % tout en conservant la même qualité.","url":"https://www.aioga.com/fr/news/cmuk7v5sc1el9ro9hcldb15xp/","contentTranslated":true,"sourceHash":"6f7c197946e35ca6","translatedAt":"2026-09-27T20:02:37.285Z"},"de":{"title":"Fireworks Research veröffentlicht Ember-1: gleiche Qualität wie Kimi K3 bei 40% weniger Tokens","summary":"Fireworks Research veröffentlicht das spezialisierte Modell Ember-1, trainiert auf Kimi K3, das bei gleicher Qualität etwa 40% weniger Token verbraucht.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks Research veröffentlicht Ember-1: gleiche Qualität wie Kimi K3 bei 40% weniger Tokens - Aioga KI-News","description":"Fireworks Research veröffentlicht das spezialisierte Modell Ember-1, trainiert auf Kimi K3, das bei gleicher Qualität etwa 40% weniger Token verbraucht.","url":"https://www.aioga.com/de/news/cmuk7v5sc1el9ro9hcldb15xp/","contentTranslated":true,"sourceHash":"6f7c197946e35ca6","translatedAt":"2026-09-27T20:02:36.359Z"},"pt-BR":{"title":"Fireworks Research lança Ember-1: alcançando qualidade nível Kimi K3 com 40% menos tokens","summary":"A Fireworks Research lançou o modelo especializado Ember-1, treinado no Kimi K3, que reduz o consumo de tokens em cerca de 40%, mantendo a mesma qualidade.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks Research lança Ember-1: alcançando qualidade nível Kimi K3 com 40% menos tokens - Aioga Notícias de IA","description":"A Fireworks Research lançou o modelo especializado Ember-1, treinado no Kimi K3, que reduz o consumo de tokens em cerca de 40%, mantendo a mesma qualidade.","url":"https://www.aioga.com/pt-BR/news/cmuk7v5sc1el9ro9hcldb15xp/","contentTranslated":true,"sourceHash":"6f7c197946e35ca6","translatedAt":"2026-09-27T20:02:46.021Z"},"ru":{"title":"Fireworks Research выпускает Ember-1: достижение качества уровня Kimi K3 с 40% меньшим количеством токенов","summary":"Fireworks Research выпустила специализированную модель Ember-1, обученную на Kimi K3, которая снижает потребление токенов примерно на 40% при сохранении такого же качества.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks Research выпускает Ember-1: достижение качества уровня Kimi K3 с 40% меньшим количеством токенов - Aioga Новости ИИ","description":"Fireworks Research выпустила специализированную модель Ember-1, обученную на Kimi K3, которая снижает потребление токенов примерно на 40% при сохранении такого же качества.","url":"https://www.aioga.com/ru/news/cmuk7v5sc1el9ro9hcldb15xp/","contentTranslated":true,"sourceHash":"6f7c197946e35ca6","translatedAt":"2026-09-27T20:02:46.197Z"},"ar":{"title":"أطلقت أبحاث الألعاب النارية إمبر-1: تحقيق جودة بمستوى كيمي K3 مع 40٪ أقل من الرموز","summary":"أصدرت فايروركس ريسيرش النموذج المتخصص إمبر-1، المدرب على Kimi K3، والذي يقلل من استهلاك الرموز بنسبة حوالي 40٪ مع الحفاظ على نفس الجودة.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Fireworks AI（网页）","pageTitle":"أطلقت أبحاث الألعاب النارية إمبر-1: تحقيق جودة بمستوى كيمي K3 مع 40٪ أقل من الرموز - Aioga أخبار الذكاء الاصطناعي","description":"أصدرت فايروركس ريسيرش النموذج المتخصص إمبر-1، المدرب على Kimi K3، والذي يقلل من استهلاك الرموز بنسبة حوالي 40٪ مع الحفاظ على نفس الجودة.","url":"https://www.aioga.com/ar/news/cmuk7v5sc1el9ro9hcldb15xp/","contentTranslated":true,"sourceHash":"6f7c197946e35ca6","translatedAt":"2026-09-27T20:02:55.031Z"},"hi":{"title":"आतिशबाजी अनुसंधान ने एम्बर-1 जारी किया: 3% कम टोकन के साथ Kimi K40-स्तरीय गुणवत्ता प्राप्त करना","summary":"आतिशबाजी अनुसंधान ने Kimi K3 पर प्रशिक्षित विशेष मॉडल Ember-1 जारी किया, जो समान गुणवत्ता बनाए रखते हुए टोकन की खपत को लगभग 40% तक कम कर देता है।","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Fireworks AI（网页）","pageTitle":"आतिशबाजी अनुसंधान ने एम्बर-1 जारी किया: 3% कम टोकन के साथ Kimi K40-स्तरीय गुणवत्ता प्राप्त करना - Aioga AI समाचार","description":"आतिशबाजी अनुसंधान ने Kimi K3 पर प्रशिक्षित विशेष मॉडल Ember-1 जारी किया, जो समान गुणवत्ता बनाए रखते हुए टोकन की खपत को लगभग 40% तक कम कर देता है।","url":"https://www.aioga.com/hi/news/cmuk7v5sc1el9ro9hcldb15xp/","contentTranslated":true,"sourceHash":"6f7c197946e35ca6","translatedAt":"2026-09-27T20:02:55.020Z"},"it":{"title":"Fireworks Research rilascia Ember-1: raggiungendo la qualità di livello Kimi K3 con il 40% di token in meno","summary":"Fireworks Research ha rilasciato il modello specializzato Ember-1, addestrato su Kimi K3, che riduce il consumo di token di circa il 40% mantenendo la stessa qualità.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks Research rilascia Ember-1: raggiungendo la qualità di livello Kimi K3 con il 40% di token in meno - Aioga Notizie IA","description":"Fireworks Research ha rilasciato il modello specializzato Ember-1, addestrato su Kimi K3, che riduce il consumo di token di circa il 40% mantenendo la stessa qualità.","url":"https://www.aioga.com/it/news/cmuk7v5sc1el9ro9hcldb15xp/","contentTranslated":true,"sourceHash":"6f7c197946e35ca6","translatedAt":"2026-09-27T20:03:03.639Z"},"nl":{"title":"Fireworks Research brengt Ember-1 uit: bereikt Kimi K3-niveau kwaliteit met 40% minder tokens","summary":"Fireworks Research bracht het gespecialiseerde model Ember-1 uit, getraind op Kimi K3, dat het tokenverbruik met ongeveer 40% vermindert terwijl de kwaliteit behouden blijft.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks Research brengt Ember-1 uit: bereikt Kimi K3-niveau kwaliteit met 40% minder tokens - Aioga AI-nieuws","description":"Fireworks Research bracht het gespecialiseerde model Ember-1 uit, getraind op Kimi K3, dat het tokenverbruik met ongeveer 40% vermindert terwijl de kwaliteit behouden blijft.","url":"https://www.aioga.com/nl/news/cmuk7v5sc1el9ro9hcldb15xp/","contentTranslated":true,"sourceHash":"6f7c197946e35ca6","translatedAt":"2026-09-27T20:03:03.716Z"},"tr":{"title":"Fireworks Research, Ember-1'i yayınladı: %40 daha az jeton ile Kimi K3 seviyesi kalitesine ulaşmak","summary":"Fireworks Research, Kimi K3 üzerinde eğitilen özel Ember-1 modelini piyasaya sürdü; bu model, token tüketimini yaklaşık %40 azaltırken aynı kaliteyi koruyor.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks Research, Ember-1'i yayınladı: %40 daha az jeton ile Kimi K3 seviyesi kalitesine ulaşmak - Aioga AI Haberleri","description":"Fireworks Research, Kimi K3 üzerinde eğitilen özel Ember-1 modelini piyasaya sürdü; bu model, token tüketimini yaklaşık %40 azaltırken aynı kaliteyi koruyor.","url":"https://www.aioga.com/tr/news/cmuk7v5sc1el9ro9hcldb15xp/","contentTranslated":true,"sourceHash":"6f7c197946e35ca6","translatedAt":"2026-09-27T20:03:12.402Z"},"vi":{"title":"Fireworks Research phát hành Ember-1: đạt chất lượng Kimi K3 với lượng token ít hơn 40%","summary":"Fireworks Research đã phát hành mô hình chuyên biệt Ember-1, được huấn luyện trên Kimi K3, giúp giảm khoảng 40% lượng token tiêu thụ trong khi vẫn giữ nguyên chất lượng.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks Research phát hành Ember-1: đạt chất lượng Kimi K3 với lượng token ít hơn 40% - Tin tức AI Aioga","description":"Fireworks Research đã phát hành mô hình chuyên biệt Ember-1, được huấn luyện trên Kimi K3, giúp giảm khoảng 40% lượng token tiêu thụ trong khi vẫn giữ nguyên chất lượng.","url":"https://www.aioga.com/vi/news/cmuk7v5sc1el9ro9hcldb15xp/","contentTranslated":true,"sourceHash":"6f7c197946e35ca6","translatedAt":"2026-09-27T20:03:12.572Z"},"id":{"title":"Fireworks Research merilis Ember-1: mencapai kualitas level Kimi K3 dengan token 40% lebih sedikit","summary":"Fireworks Research merilis model khusus Ember-1, yang dilatih menggunakan Kimi K3, yang mengurangi konsumsi token sekitar 40% sambil mempertahankan kualitas yang sama.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks Research merilis Ember-1: mencapai kualitas level Kimi K3 dengan token 40% lebih sedikit - Berita AI Aioga","description":"Fireworks Research merilis model khusus Ember-1, yang dilatih menggunakan Kimi K3, yang mengurangi konsumsi token sekitar 40% sambil mempertahankan kualitas yang sama.","url":"https://www.aioga.com/id/news/cmuk7v5sc1el9ro9hcldb15xp/","contentTranslated":true,"sourceHash":"6f7c197946e35ca6","translatedAt":"2026-09-27T20:03:21.252Z"},"th":{"title":"Fireworks Research เปิดตัว Ember-1: บรรลุคุณภาพระดับ Kimi K3 ด้วยโทเคนน้อยลง 40%","summary":"Fireworks Research ได้ปล่อยรุ่นพิเศษ Ember-1 ซึ่งฝึกบน Kimi K3 ซึ่งช่วยลดการใช้โทเค็นได้ประมาณ 40% ในขณะที่ยังคงคุณภาพเดิม","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks Research เปิดตัว Ember-1: บรรลุคุณภาพระดับ Kimi K3 ด้วยโทเคนน้อยลง 40% - ข่าว AI Aioga","description":"Fireworks Research ได้ปล่อยรุ่นพิเศษ Ember-1 ซึ่งฝึกบน Kimi K3 ซึ่งช่วยลดการใช้โทเค็นได้ประมาณ 40% ในขณะที่ยังคงคุณภาพเดิม","url":"https://www.aioga.com/th/news/cmuk7v5sc1el9ro9hcldb15xp/","contentTranslated":true,"sourceHash":"6f7c197946e35ca6","translatedAt":"2026-09-27T20:03:21.231Z"},"pl":{"title":"Fireworks Research wydaje Ember-1: o 40% mniej tokenów dla jakości równej Kimi K3","summary":"Fireworks Research wydaje wyspecjalizowany model Ember-1, trenowany na bazie Kimi K3, zachowując równą jakość przy około 40% mniejszym zużyciu tokenów.","category":"行业动态","source":"Hacker News 热门（buzzing.cc 中文翻译）","aggregationSource":"Fireworks AI（网页）","pageTitle":"Fireworks Research wydaje Ember-1: o 40% mniej tokenów dla jakości równej Kimi K3 - Aioga Wiadomości AI","description":"Fireworks Research wydaje wyspecjalizowany model Ember-1, trenowany na bazie Kimi K3, zachowując równą jakość przy około 40% mniejszym zużyciu tokenów.","url":"https://www.aioga.com/pl/news/cmuk7v5sc1el9ro9hcldb15xp/","contentTranslated":true,"sourceHash":"6f7c197946e35ca6","translatedAt":"2026-09-27T20:03:29.264Z"}},"evidenceTier":"verified-news","reviewStatus":"automated-ingest","indexable":true,"editorialCover":""}}