{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-08-11T09:21:12.743Z","headline":"Making Knowledge Distillation Cheap Enough to Run at Scale","description":"🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/news/cmsn4c69701x9ron5wn1hfxxw/","mainEntityOfPage":"https://www.aioga.com/news/cmsn4c69701x9ron5wn1hfxxw/","datePublished":"2026-08-10T10:05:36.000Z","dateModified":"2026-08-10T10:05:36.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://huggingface.co/blog/MultiverseComputingCAI/efficient-knowledge-distillation","https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw"],"canonicalUrl":"https://www.aioga.com/news/cmsn4c69701x9ron5wn1hfxxw/","directAnswer":{"@type":"Answer","text":"Hugging Face 博客介绍一种面向大语言模型知识蒸馏的系统方案：预先缓存教师模型每个位置最可能的前100个词元，并使用分块KL损失，减少训练阶段的显存占用与教师模型重复计算。","url":"https://www.aioga.com/news/cmsn4c69701x9ron5wn1hfxxw/","dateCreated":"2026-08-10T10:05:36.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"huggingface.co source article","url":"https://huggingface.co/blog/MultiverseComputingCAI/efficient-knowledge-distillation","datePublished":"2026-08-10T10:05:36.000Z","provider":{"@type":"Organization","name":"huggingface.co","url":"https://huggingface.co/blog/MultiverseComputingCAI/efficient-knowledge-distillation"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","datePublished":"2026-08-10T10:05:36.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw"}}],"aggregationSource":"Hugging Face：Blog（RSS）","originalPublisher":{"name":"huggingface.co","url":"https://huggingface.co/blog/MultiverseComputingCAI/efficient-knowledge-distillation"},"geoDeepAnswer":null,"article":{"id":"cmsn4c69701x9ron5wn1hfxxw","slug":"cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/news/cmsn4c69701x9ron5wn1hfxxw/","title":"Making Knowledge Distillation Cheap Enough to Run at Scale","title_en":"","summary":"🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","source":"Hugging Face：Blog（RSS）","sourceUrl":"https://huggingface.co/blog/MultiverseComputingCAI/efficient-knowledge-distillation","aiHotUrl":"https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","publishedAt":"2026-08-10T10:05:36.000Z","category":"行业动态","score":58,"selected":false,"articleBody":["Why distillation recovery is expensive ：#why-distillation-recovery-is-expensive Two systems changes ：#two-systems-changes What this changes in practice ：#what-this-changes-in-practice Scaling to long context lengths ：#scaling-to-long-context-lengths The resulting student ：#the-resulting-student ：https://multiversecomputing.com ：https://github.com/CompactifAI/Full-Chunked-KL-Loss/tree/main ：https://arxiv.org/abs/2608.03796","：https://cdn-uploads.huggingface.co/production/uploads/614a1ebb8f82f1df64d55126/jsYblMP3R5futGF9Y8Kl1.png","Knowledge distillation：https://arxiv.org/abs/1503.02531, training a smaller student model to match the performance of a larger teacher, is a well-known technique in Machine Learning. With the recent wave of open-source Large Language Models, such as gpt-oss：https://huggingface.co/openai/gpt-oss-120b, Qwen：https://huggingface.co/collections/Qwen/qwen35, GLM：https://huggingface.co/zai-org/GLM-5.2, or Kimi：https://huggingface.co/moonshotai/Kimi-K3, it has become a mainstream research topic again. Deploying these very large models is expensive: the recent Kimi-K3 model：https://huggingface.co/moonshotai/Kimi-K3 has 2.8 trillion parameters and needs roughly 3TB of VRAM just to load. Compressing them into smaller models and recovering the original capabilities through knowledge distillation has therefore become standard practice, with companies like Nvidia (Nemotron 3 Puzzle 75B：https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4) or Multiverse Computing (Hypernova 60B：https://huggingface.co/MultiverseComputingCAI/Hypernova-60B-2605) recently releasing high-quality compressed models.","The distillation step is what decides most of the final quality, but it's also usually the most expensive part of the pipeline. Keeping both the teacher and student loaded, and producing a probability distribution over the entire vocabulary for every token, requires enormous amounts of VRAM, typically feasible only with hundreds of GPUs and careful tensor-parallelism strategies. Our latest paper, Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss：https://huggingface.co/papers/2608.03796, tackles this with two systems changes: caching the teacher's top-K logits once so the teacher never has to sit in memory alongside the student, and a new, memory-efficient KL-divergence loss that avoids ever materializing the full vocabulary-size × sequence-length matrix, cutting VRAM use far below what the default implementations in libraries like PyTorch：https://pytorch.org or NVIDIA Megatron-Bridge：https://github.com/NVIDIA-NeMo/Megatron-Bridge achieve. Together, these two changes cut training cost enough to make long-context healing possible on a single GPU, and cheap enough to make large-scale experimentation practical.","The standard setup, online distillation using the Kullback-Leibler divergence loss：https://docs.pytorch.org/docs/2.13/generated/torch.nn.KLDivLoss.html (KL loss), keeps both the teacher and the student loaded at the same time. At every training step, the teacher runs a full forward pass to produce its output distribution, and the student is trained to match it. This is the most expressive setup, since the full teacher distribution is available, but it is also the most memory- and compute-intensive: two full-vocabulary tensors have to be held per token position, and the teacher has to be recomputed on every single step even though its behavior does not change across a training run.","As a practical example, gpt-oss-120b has a vocabulary of 201,088 tokens. At a sequence length of 32K and batch size 4, the teacher-probability tensor alone has shape 4 × 201,088 × 32,768 ; in bfloat16, that's already about 50GB of VRAM for a single tensor. Add gradients, activations, model weights, and optimizer states, and a single training iteration of distillation can peak at roughly 250GB of VRAM, more than even an H200 or B200 GPU can provide. In this post, we show that reformulating the KL loss to process the data in chunks reduces this cost to almost nothing.","：https://cdn-uploads.huggingface.co/production/uploads/668e37fd9c9aa124a3c867e8/-whmAUit96bo1Yy0vLPis.png Dense KL spikes to roughly 250GB, above a single H200's 141GB capacity. The fused chunked loss never forms that spike and peaks at about 128GB. Source: paper Figure 1.","Offline distillation. Instead of recomputing the teacher at every step, we compute its output once, cache the top-100 most likely tokens per position, and train the student against that cache. The teacher never has to sit in memory during training and does not need to be run again once the cache exists, so the same cache can be reused across many ablations.","A fused, chunked KL loss. To see why the loss itself is expensive, picture what it actually builds: for every token position in a sequence and every word in the vocabulary, the loss needs a number describing how much the student's prediction disagrees with the teacher's. Laid out as a grid, that's one row per vocabulary entry and one column per sequence position, for a vocabulary of 100K+ words and a long sequence, that grid is enormous, and the default way of computing a KL loss builds the whole thing before it can produce a single number.","We compare three ways of computing this same loss, all mathematically equivalent:","The GIF below shows the difference between the dense and fused-chunked approaches: one builds the whole comparison grid and holds onto all of it, the other builds and discards one slice at a time, so memory never grows beyond a single chunk.","We have open-sourced the chunked-loss implementation: github.com/CompactifAI/Full-Chunked-KL-Loss：https://github.com/CompactifAI/Full-Chunked-KL-Loss","：https://cdn-uploads.huggingface.co/production/uploads/614a1ebb8f82f1df64d55126/k4K6L0g7crbOk4ZqJuFuy.gif","The table below puts all four setups head to head: online distillation, and the three offline loss implementations just described. Comparing them on a single H200 GPU with Llama 3.1 8B Instruct：https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct as teacher and a 3.2B Llama model as student at an 8K token context, all four reach near-identical training loss, even though the offline runs train against only the cached top-100 logits per token.","：https://cdn-uploads.huggingface.co/production/uploads/668e37fd9c9aa124a3c867e8/hV9Hu8he7f1kZJy_SS9X6.png","The loss curves overlap almost exactly across all four methods, confirming offline distillation with top-100 cached logits is lossless relative to online distillation. Source: paper Figure 2. At this sequence length, the fused chunked loss is not yet the fastest option, its extra backward-pass projection costs a bit of speed, but its real advantage only shows up as context length grows, which the next section demonstrates.","To see the scaling pattern more starkly, we ran an isolated benchmark on a toy output-projection network (no transformer body, just the loss kernel). At 32K tokens, peak memory falls from 85.2 GiB with the dense loss to 5.45 GiB with the fully chunked version, a 15.6× reduction, and the dense loss fails outright from 64K tokens onward. At 256K tokens, the fully chunked loss uses 11.6 GiB against 134.2 GiB for the next-best chunked variant, and is about 3.3× faster per iteration at that length.","Distilling a GPT-OSS 20B model at a 32,768-token context, the memory freed by the fused loss let the setup shrink from four GPU nodes down to one. Step time fell from 57.0 to 12.23 seconds, about 5× faster, and throughput per GPU rose from 74.2 to 345.7 TFLOP/s.","The efficient offline setup is what made a large-scale distillation campaign affordable in the first place. The resulting compact student, distilled from Llama 3.1 8B Instruct down to about 3.2B parameters, retains most of the teacher's accuracy on BoolQ and HellaSwag, stays within about nine points of it on MMLU, at less than half the parameter count.","：https://cdn-uploads.huggingface.co/production/uploads/614a1ebb8f82f1df64d55126/e0VvKmFktQCyoOhhOTD4n.png The student retains most of the teacher's short-context accuracy at less than half the size. Source: paper Figure 6.","This work is part of Multiverse Computing's：https://multiversecomputing.com ongoing research into making distillation and healing practical to run at scale：https://multiversecomputing.com/compactifai, not just as a one-off recipe, but as something teams can iterate on cheaply. The paper also covers additional ablations, such as how the choice of loss function and sequence packing affect recovery quality.","Want the full technical details, including the closed-form gradient behind the fused chunked loss and the complete training configuration? Read the full paper：https://arxiv.org/abs/2608.03796, or get in touch with our team to talk about applying this to your own distillation pipelines.","We have also open-sourced the chunked-loss implementation: github.com/CompactifAI/Full-Chunked-KL-Loss：https://github.com/CompactifAI/Full-Chunked-KL-Loss"],"articleImages":[{"sourceUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1632247447995-noauth.jpeg","alt":"","afterParagraph":0,"url":"/media/articles/cmsn4c69701x9ron5wn1hfxxw/f607df82a17eab06.webp"},{"sourceUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/624c3e54a19f20b19776366d/k5W_QQ63EJHZhIHckzwsi.jpeg","alt":"","afterParagraph":0,"url":"/media/articles/cmsn4c69701x9ron5wn1hfxxw/73d3c8282923458d.webp"}],"mediaStatus":"ok","articleBodyZh":["为什么蒸馏恢复成本高：#why-distillation-recovery-is-expensive 两个系统的变化：#two-systems-changes 这些变化在实践中意味着什么：#what-this-changes-in-practice 扩展到长上下文长度：#scaling-to-long-context-lengths 由此产生的学生模型：#the-resulting-student：https://multiversecomputing.com：https://github.com/CompactifAI/Full-Chunked-KL-Loss/tree/main：https://arxiv.org/abs/2608.03796","：https://cdn-uploads.huggingface.co/production/uploads/614a1ebb8f82f1df64d55126/jsYblMP3R5futGF9Y8Kl1.png","知识蒸馏：https://arxiv.org/abs/1503.02531，训练一个较小的学生模型以匹配较大教师模型的性能，是机器学习中一个著名的技术。随着近期开源大型语言模型的浪潮，比如 gpt-oss：https://huggingface.co/openai/gpt-oss-120b, Qwen：https://huggingface.co/collections/Qwen/qwen35, GLM：https://huggingface.co/zai-org/GLM-5.2, 或 Kimi：https://huggingface.co/moonshotai/Kimi-K3，这一研究主题再次成为主流。部署这些非常大的模型成本很高：近期的 Kimi-K3 模型：https://huggingface.co/moonshotai/Kimi-K3 拥有 2.8 万亿参数，仅加载就需要大约 3TB 的显存。因此，将它们压缩成较小的模型并通过知识蒸馏恢复原有能力，已经成为标准做法，如 Nvidia (Nemotron 3 Puzzle 75B：https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4) 或 Multiverse Computing (Hypernova 60B：https://huggingface.co/MultiverseComputingCAI/Hypernova-60B-2605) 最近发布了高质量的压缩模型。","蒸馏步骤决定了大部分最终质量，但它通常也是流程中最昂贵的部分。保持教师模型和学生模型同时加载，并为每个标记生成整个词汇表的概率分布，需要巨量的显存，通常只能在数百个GPU和谨慎的张量并行策略下实现。我们最新的论文《面向大型语言模型的高效知识蒸馏：离线 Top-K Logits 与融合分块 KL 损失》：https://huggingface.co/papers/2608.03796，通过两个系统方面的改动来解决这个问题：一次性缓存教师模型的 Top-K logits，这样教师模型无需与学生模型同时占用内存；以及一种新的、节省内存的 KL 散度损失，它避免了生成整个词汇表大小 × 序列长度的矩阵，大幅度降低了显存使用量，远低于 PyTorch：https://pytorch.org 或 NVIDIA Megatron-Bridge：https://github.com/NVIDIA-NeMo/Megatron-Bridge 等库的默认实现。结合这两项改动，训练成本下降到足以在单个 GPU 上实现长上下文蒸馏，并且足够低廉，使大规模实验成为可能。","标准配置是使用 Kullback-Leibler 散度损失 (KL 损失)进行在线蒸馏：https://docs.pytorch.org/docs/2.13/generated/torch.nn.KLDivLoss.html，这要求教师模型和学生模型同时加载。在每个训练步骤中，教师模型进行完整的前向传播以生成输出分布，学生模型训练以匹配它。这是最具表现力的配置，因为可以使用完整的教师分布，但它也是最耗内存和计算的：每个标记位置需要保存两个完整词汇表的张量，并且教师模型必须在每一步都重新计算，即使其行为在整个训练过程中不变。","作为一个实际的例子，gpt-oss-120b 的词汇量为 201,088 个标记。在序列长度为 32K、批量大小为 4 的情况下，仅教师概率张量的形状就是 4 × 201,088 × 32,768；以 bfloat16 存储，这单个张量就占用了大约 50GB 的显存。再加上梯度、激活、模型权重和优化器状态，一次蒸馏训练迭代的显存峰值大约可达 250GB，甚至超过 H200 或 B200 GPU 的显存容量。在这篇文章中，我们展示了将 KL 损失重构为分块处理数据可以将成本几乎降为零。","：https://cdn-uploads.huggingface.co/production/uploads/668e37fd9c9aa124a3c867e8/-whmAUit96bo1Yy0vLPis.png 密集 KL 峰值约为 250GB，高于单个 H200 的 141GB 容量。融合分块损失永远不会形成这样的峰值，最大值约为 128GB。来源：论文图 1。","离线蒸馏。我们不是在每一步都重新计算教师模型，而是计算一次其输出，缓存每个位置最可能的前 100 个标记，并让学生模型针对该缓存进行训练。在训练期间教师模型无需驻留在内存中，并且一旦缓存存在就无需再次运行，因此同一个缓存可以在多次消融实验中重复使用。","融合的分块 KL 损失。为了理解为什么损失本身很昂贵，可以想象它实际构建的内容：对于序列中的每个标记位置和词汇表中的每个单词，损失需要一个数值来描述学生的预测与教师的不一致程度。排列成网格，就是每个词汇表条目一行，每个序列位置一列，对于 10 万+ 词的词汇表和长序列，这个网格非常巨大，而默认的 KL 损失计算方式是在生成单个数值之前会先构建整个网格。","我们比较了三种计算同一损失的方法，它们在数学上是等价的：","下方的 GIF 展示了密集与融合分块方法的区别：一种构建整个比较网格并保留所有内容，另一种每次构建并丢弃一片网格，因此内存不会超过单个块。","我们已经开源了分块损失的实现：github.com/CompactifAI/Full-Chunked-KL-Loss：https://github.com/CompactifAI/Full-Chunked-KL-Loss","：https://cdn-uploads.huggingface.co/production/uploads/614a1ebb8f82f1df64d55126/k4K6L0g7crbOk4ZqJuFuy.gif","下表将这四种配置对比：在线蒸馏和刚才描述的三种离线损耗实现。在单个H200 GPU上，使用Llama 3.1 8B Instruct：https：//huggingface.co/meta-llama/Llama-3.1-8B-Instruct 作为教师，以及3.2B Llama模型作为学生，8K令牌上下文下，四者训练损失几乎相同，尽管离线运行只针对每个令牌缓存的前100日志进行训练。","：https://cdn-uploads.huggingface.co/production/uploads/668e37fd9c9aa124a3c867e8/hV9Hu8he7f1kZJy_SS9X6.png","损失曲线在四种方法中几乎完全重叠，证实了使用前100缓存日志的离线蒸馏相对于在线蒸馏是无损的。来源：论文，图2。在这个序列长度下，融合分块丢失还不是最快的选择，其额外的向后传投影会消耗一些速度，但其真正的优势只有在上下文长度增加时才显现，下一节将演示这一点。","为了更清楚地看到缩放模式，我们在一个玩具输出-投影网络上运行了一个孤立的基准测试（没有变压器本体，只有损耗核）。在32K代币时，高峰内存从带稠密丢失的85.2 GiB降至全分块版本的5.45 GiB，减少了15.6×，而从64K代币开始，密集损失则完全失效。在256K代币中，完全分块损失消耗11.6 GiB，而下一个最佳分块版本为134.2 GiB，且每次迭代速度约快3.3×。","在32,768令牌上下文中提取出GPT-OSS 20B模型，融合丢失释放的内存使得配置从四个GPU节点缩减到一个。步进时间从57.0秒降至12.23秒，快约5×，GPU吞吐量从74.2提升至345.7 TFLOP/s。","高效的离线设置正是大规模蒸馏活动之所以能负担得起的原因。最终的紧凑学生，从Llama 3.1 8B Instruct中精炼到约3.2B参数，在BoolQ和HellaSwag上保持了教师的大部分准确度，在MMLU上也保持在大约9分以内，且参数数不到一半。","：https://cdn-uploads.huggingface.co/production/uploads/614a1ebb8f82f1df64d55126/e0VvKmFktQCyoOhhOTD4n.png 学生模型在不到一半的规模下，保持了大部分教师模型的短上下文准确性。来源：论文图6。","这项工作是Multiverse Computing's：https://multiversecomputing.com 持续研究的一部分，旨在使蒸馏和修复在大规模下可实际运行：https://multiversecomputing.com/compactifai，不仅仅作为一次性方案，而是团队可以低成本迭代的内容。论文还涵盖了其他消融实验，例如损失函数选择和序列打包如何影响恢复质量。","想要完整的技术细节吗，包括融合分块损失背后的闭式梯度和完整的训练配置？阅读完整论文：https://arxiv.org/abs/2608.03796，或联系我们的团队讨论将其应用到您自己的蒸馏管道中。","我们还开源了分块损失的实现：github.com/CompactifAI/Full-Chunked-KL-Loss：https://github.com/CompactifAI/Full-Chunked-KL-Loss"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Hugging Face 博客介绍一种面向大语言模型知识蒸馏的系统方案：预先缓存教师模型每个位置最可能的前100个词元，并使用分块KL损失，减少训练阶段的显存占用与教师模型重复计算。","background":"知识蒸馏通过训练较小的学生模型匹配较大教师模型的表现。文章指出，在线蒸馏需同时加载教师和学生，并为每个词元处理完整词表分布，因此在长上下文场景下显存与计算成本较高。","viewpoint":"Aioga 判断，该方案的关键价值在于把教师推理从反复在线计算改为可复用缓存，并避免生成完整的词表大小乘序列长度矩阵。其实际收益仍可能取决于前100个词元缓存对蒸馏质量的影响。","implications":"文章称，分块损失可避免约250GB的显存峰值，在示例中峰值约为128GB；结合离线缓存，长上下文蒸馏可能在单GPU上进行，也可能降低大规模实验的门槛。","nextStep":"值得关注后续公开材料是否进一步说明不同缓存规模、上下文长度和模型组合下的效果，以及压缩后的学生模型在能力恢复、训练成本和部署需求之间的具体权衡。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-08-10T12:22:02.666Z","sourceHash":"4f281a3b889abd04","review":{"approved":true,"groundedness":94,"clarity":91,"duplicationRisk":18,"blockingIssues":[],"notes":["“避免生成完整的词表大小乘序列长度矩阵”可更精确地表述为“避免在内存中完整实例化该矩阵”，以贴近原文的“avoids ever materializing”。","“实际收益仍可能取决于前100个词元缓存对蒸馏质量的影响”属于合理的审慎判断，且已明确归于观点，不构成事实性问题。","“避免约250GB的显存峰值”容易被轻微误读为节省约250GB；若需更严谨，可改为“避免密集KL实现约250GB的峰值，并在示例中将峰值控制在约128GB”。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","Hugging Face：Blog（RSS）"],"translations":{"zh-CN":{"title":"Making Knowledge Distillation Cheap Enough to Run at Scale","summary":"🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"行业动态","source":"huggingface.co","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Making Knowledge Distillation Cheap Enough to Run at Scale - Aioga AI资讯","description":"🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/news/cmsn4c69701x9ron5wn1hfxxw/","articleBody":["为什么蒸馏恢复成本高：#why-distillation-recovery-is-expensive 两个系统的变化：#two-systems-changes 这些变化在实践中意味着什么：#what-this-changes-in-practice 扩展到长上下文长度：#scaling-to-long-context-lengths 由此产生的学生模型：#the-resulting-student：https://multiversecomputing.com：https://github.com/CompactifAI/Full-Chunked-KL-Loss/tree/main：https://arxiv.org/abs/2608.03796","：https://cdn-uploads.huggingface.co/production/uploads/614a1ebb8f82f1df64d55126/jsYblMP3R5futGF9Y8Kl1.png","知识蒸馏：https://arxiv.org/abs/1503.02531，训练一个较小的学生模型以匹配较大教师模型的性能，是机器学习中一个著名的技术。随着近期开源大型语言模型的浪潮，比如 gpt-oss：https://huggingface.co/openai/gpt-oss-120b, Qwen：https://huggingface.co/collections/Qwen/qwen35, GLM：https://huggingface.co/zai-org/GLM-5.2, 或 Kimi：https://huggingface.co/moonshotai/Kimi-K3，这一研究主题再次成为主流。部署这些非常大的模型成本很高：近期的 Kimi-K3 模型：https://huggingface.co/moonshotai/Kimi-K3 拥有 2.8 万亿参数，仅加载就需要大约 3TB 的显存。因此，将它们压缩成较小的模型并通过知识蒸馏恢复原有能力，已经成为标准做法，如 Nvidia (Nemotron 3 Puzzle 75B：https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4) 或 Multiverse Computing (Hypernova 60B：https://huggingface.co/MultiverseComputingCAI/Hypernova-60B-2605) 最近发布了高质量的压缩模型。","蒸馏步骤决定了大部分最终质量，但它通常也是流程中最昂贵的部分。保持教师模型和学生模型同时加载，并为每个标记生成整个词汇表的概率分布，需要巨量的显存，通常只能在数百个GPU和谨慎的张量并行策略下实现。我们最新的论文《面向大型语言模型的高效知识蒸馏：离线 Top-K Logits 与融合分块 KL 损失》：https://huggingface.co/papers/2608.03796，通过两个系统方面的改动来解决这个问题：一次性缓存教师模型的 Top-K logits，这样教师模型无需与学生模型同时占用内存；以及一种新的、节省内存的 KL 散度损失，它避免了生成整个词汇表大小 × 序列长度的矩阵，大幅度降低了显存使用量，远低于 PyTorch：https://pytorch.org 或 NVIDIA Megatron-Bridge：https://github.com/NVIDIA-NeMo/Megatron-Bridge 等库的默认实现。结合这两项改动，训练成本下降到足以在单个 GPU 上实现长上下文蒸馏，并且足够低廉，使大规模实验成为可能。","标准配置是使用 Kullback-Leibler 散度损失 (KL 损失)进行在线蒸馏：https://docs.pytorch.org/docs/2.13/generated/torch.nn.KLDivLoss.html，这要求教师模型和学生模型同时加载。在每个训练步骤中，教师模型进行完整的前向传播以生成输出分布，学生模型训练以匹配它。这是最具表现力的配置，因为可以使用完整的教师分布，但它也是最耗内存和计算的：每个标记位置需要保存两个完整词汇表的张量，并且教师模型必须在每一步都重新计算，即使其行为在整个训练过程中不变。","作为一个实际的例子，gpt-oss-120b 的词汇量为 201,088 个标记。在序列长度为 32K、批量大小为 4 的情况下，仅教师概率张量的形状就是 4 × 201,088 × 32,768；以 bfloat16 存储，这单个张量就占用了大约 50GB 的显存。再加上梯度、激活、模型权重和优化器状态，一次蒸馏训练迭代的显存峰值大约可达 250GB，甚至超过 H200 或 B200 GPU 的显存容量。在这篇文章中，我们展示了将 KL 损失重构为分块处理数据可以将成本几乎降为零。","：https://cdn-uploads.huggingface.co/production/uploads/668e37fd9c9aa124a3c867e8/-whmAUit96bo1Yy0vLPis.png 密集 KL 峰值约为 250GB，高于单个 H200 的 141GB 容量。融合分块损失永远不会形成这样的峰值，最大值约为 128GB。来源：论文图 1。","离线蒸馏。我们不是在每一步都重新计算教师模型，而是计算一次其输出，缓存每个位置最可能的前 100 个标记，并让学生模型针对该缓存进行训练。在训练期间教师模型无需驻留在内存中，并且一旦缓存存在就无需再次运行，因此同一个缓存可以在多次消融实验中重复使用。","融合的分块 KL 损失。为了理解为什么损失本身很昂贵，可以想象它实际构建的内容：对于序列中的每个标记位置和词汇表中的每个单词，损失需要一个数值来描述学生的预测与教师的不一致程度。排列成网格，就是每个词汇表条目一行，每个序列位置一列，对于 10 万+ 词的词汇表和长序列，这个网格非常巨大，而默认的 KL 损失计算方式是在生成单个数值之前会先构建整个网格。","我们比较了三种计算同一损失的方法，它们在数学上是等价的：","下方的 GIF 展示了密集与融合分块方法的区别：一种构建整个比较网格并保留所有内容，另一种每次构建并丢弃一片网格，因此内存不会超过单个块。","我们已经开源了分块损失的实现：github.com/CompactifAI/Full-Chunked-KL-Loss：https://github.com/CompactifAI/Full-Chunked-KL-Loss","：https://cdn-uploads.huggingface.co/production/uploads/614a1ebb8f82f1df64d55126/k4K6L0g7crbOk4ZqJuFuy.gif","下表将这四种配置对比：在线蒸馏和刚才描述的三种离线损耗实现。在单个H200 GPU上，使用Llama 3.1 8B Instruct：https：//huggingface.co/meta-llama/Llama-3.1-8B-Instruct 作为教师，以及3.2B Llama模型作为学生，8K令牌上下文下，四者训练损失几乎相同，尽管离线运行只针对每个令牌缓存的前100日志进行训练。","：https://cdn-uploads.huggingface.co/production/uploads/668e37fd9c9aa124a3c867e8/hV9Hu8he7f1kZJy_SS9X6.png","损失曲线在四种方法中几乎完全重叠，证实了使用前100缓存日志的离线蒸馏相对于在线蒸馏是无损的。来源：论文，图2。在这个序列长度下，融合分块丢失还不是最快的选择，其额外的向后传投影会消耗一些速度，但其真正的优势只有在上下文长度增加时才显现，下一节将演示这一点。","为了更清楚地看到缩放模式，我们在一个玩具输出-投影网络上运行了一个孤立的基准测试（没有变压器本体，只有损耗核）。在32K代币时，高峰内存从带稠密丢失的85.2 GiB降至全分块版本的5.45 GiB，减少了15.6×，而从64K代币开始，密集损失则完全失效。在256K代币中，完全分块损失消耗11.6 GiB，而下一个最佳分块版本为134.2 GiB，且每次迭代速度约快3.3×。","在32,768令牌上下文中提取出GPT-OSS 20B模型，融合丢失释放的内存使得配置从四个GPU节点缩减到一个。步进时间从57.0秒降至12.23秒，快约5×，GPU吞吐量从74.2提升至345.7 TFLOP/s。","高效的离线设置正是大规模蒸馏活动之所以能负担得起的原因。最终的紧凑学生，从Llama 3.1 8B Instruct中精炼到约3.2B参数，在BoolQ和HellaSwag上保持了教师的大部分准确度，在MMLU上也保持在大约9分以内，且参数数不到一半。","：https://cdn-uploads.huggingface.co/production/uploads/614a1ebb8f82f1df64d55126/e0VvKmFktQCyoOhhOTD4n.png 学生模型在不到一半的规模下，保持了大部分教师模型的短上下文准确性。来源：论文图6。","这项工作是Multiverse Computing's：https://multiversecomputing.com 持续研究的一部分，旨在使蒸馏和修复在大规模下可实际运行：https://multiversecomputing.com/compactifai，不仅仅作为一次性方案，而是团队可以低成本迭代的内容。论文还涵盖了其他消融实验，例如损失函数选择和序列打包如何影响恢复质量。","想要完整的技术细节吗，包括融合分块损失背后的闭式梯度和完整的训练配置？阅读完整论文：https://arxiv.org/abs/2608.03796，或联系我们的团队讨论将其应用到您自己的蒸馏管道中。","我们还开源了分块损失的实现：github.com/CompactifAI/Full-Chunked-KL-Loss：https://github.com/CompactifAI/Full-Chunked-KL-Loss"]},"en":{"title":"Making Knowledge Distillation Cheap Enough to Run at Scale","summary":"🔗 Read the original article via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"Industry","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Making Knowledge Distillation Cheap Enough to Run at Scale - Aioga AI News","description":"🔗 Read the original article via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/en/news/cmsn4c69701x9ron5wn1hfxxw/","contentTranslated":true,"sourceHash":"7a527d6593e28aee","translatedAt":"2026-08-10T11:42:11.823Z"},"ja":{"title":"知識蒸留を大規模に運用できるほど安価にすること","summary":"🔗 原文記事はAIHOTより読むことができます。 https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"業界動向","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"知識蒸留を大規模に運用できるほど安価にすること - Aioga AIニュース","description":"🔗 原文記事はAIHOTより読むことができます。 https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/ja/news/cmsn4c69701x9ron5wn1hfxxw/","contentTranslated":true,"sourceHash":"7a527d6593e28aee","translatedAt":"2026-08-10T11:42:12.822Z"},"ko":{"title":"대규모로 실행할 수 있을 만큼 지식 증류 비용 절감","summary":"🔗 원문 읽기 via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"업계 동향","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"대규모로 실행할 수 있을 만큼 지식 증류 비용 절감 - Aioga AI 뉴스","description":"🔗 원문 읽기 via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/ko/news/cmsn4c69701x9ron5wn1hfxxw/","contentTranslated":true,"sourceHash":"7a527d6593e28aee","translatedAt":"2026-08-10T11:42:19.256Z"},"es":{"title":"Hacer que la destilación de conocimiento sea lo suficientemente barata como para ejecutarse a escala","summary":"🔗 Leer original via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"Industria","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Hacer que la destilación de conocimiento sea lo suficientemente barata como para ejecutarse a escala - Aioga Noticias de IA","description":"🔗 Leer original via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/es/news/cmsn4c69701x9ron5wn1hfxxw/","contentTranslated":true,"sourceHash":"7a527d6593e28aee","translatedAt":"2026-08-10T11:42:19.781Z"},"fr":{"title":"Rendre la distillation des connaissances assez bon marché pour être exécutée à grande échelle","summary":"🔗 Lire l'article original via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"Industrie","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Rendre la distillation des connaissances assez bon marché pour être exécutée à grande échelle - Aioga Actualités IA","description":"🔗 Lire l'article original via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/fr/news/cmsn4c69701x9ron5wn1hfxxw/","contentTranslated":true,"sourceHash":"7a527d6593e28aee","translatedAt":"2026-08-10T11:42:25.720Z"},"de":{"title":"Making Knowledge Distillation Cheap Enough to Run at Scale","summary":"🔗 Originaltext lesen via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Making Knowledge Distillation Cheap Enough to Run at Scale - Aioga KI-News","description":"🔗 Originaltext lesen via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/de/news/cmsn4c69701x9ron5wn1hfxxw/","contentTranslated":true,"sourceHash":"7a527d6593e28aee","translatedAt":"2026-08-10T11:42:25.514Z"},"pt-BR":{"title":"Tornando a Destilação do Conhecimento Barata o Suficiente para Operar em Escala","summary":"🔗 Leia o artigo original via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Tornando a Destilação do Conhecimento Barata o Suficiente para Operar em Escala - Aioga Notícias de IA","description":"🔗 Leia o artigo original via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/pt-BR/news/cmsn4c69701x9ron5wn1hfxxw/","contentTranslated":true,"sourceHash":"7a527d6593e28aee","translatedAt":"2026-08-10T11:42:34.256Z"},"ru":{"title":"Сделать дистилляцию знаний достаточно дешёвой для масштабной работы","summary":"🔗 Прочитайте оригинальную статью на сайте AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Сделать дистилляцию знаний достаточно дешёвой для масштабной работы - Aioga Новости ИИ","description":"🔗 Прочитайте оригинальную статью на сайте AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/ru/news/cmsn4c69701x9ron5wn1hfxxw/","contentTranslated":true,"sourceHash":"7a527d6593e28aee","translatedAt":"2026-08-10T11:42:34.316Z"},"ar":{"title":"جعل تقطير المعرفة رخيصًا بما يكفي للتشغيل على نطاق واسع","summary":"🔗 اقرأ المقال الأصلي عبر AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"جعل تقطير المعرفة رخيصًا بما يكفي للتشغيل على نطاق واسع - Aioga أخبار الذكاء الاصطناعي","description":"🔗 اقرأ المقال الأصلي عبر AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/ar/news/cmsn4c69701x9ron5wn1hfxxw/","contentTranslated":true,"sourceHash":"7a527d6593e28aee","translatedAt":"2026-08-10T11:42:41.253Z"},"hi":{"title":"ज्ञान डिस्टिलेशन को बड़े पैमाने पर चलाने के लिए पर्याप्त सस्ता बनाना","summary":"🔗 मूल लेख पढ़ें via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"ज्ञान डिस्टिलेशन को बड़े पैमाने पर चलाने के लिए पर्याप्त सस्ता बनाना - Aioga AI समाचार","description":"🔗 मूल लेख पढ़ें via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/hi/news/cmsn4c69701x9ron5wn1hfxxw/","contentTranslated":true,"sourceHash":"7a527d6593e28aee","translatedAt":"2026-08-10T11:42:42.288Z"},"it":{"title":"Rendere la distillazione della conoscenza abbastanza economica da essere eseguita su larga scala","summary":"🔗 Leggi l'articolo originale via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Rendere la distillazione della conoscenza abbastanza economica da essere eseguita su larga scala - Aioga Notizie IA","description":"🔗 Leggi l'articolo originale via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/it/news/cmsn4c69701x9ron5wn1hfxxw/","contentTranslated":true,"sourceHash":"7a527d6593e28aee","translatedAt":"2026-08-10T11:42:48.221Z"},"nl":{"title":"Kennisdestillatie goedkoop genoeg maken om op grote schaal te draaien","summary":"🔗 Lees het originele artikel via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Kennisdestillatie goedkoop genoeg maken om op grote schaal te draaien - Aioga AI-nieuws","description":"🔗 Lees het originele artikel via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/nl/news/cmsn4c69701x9ron5wn1hfxxw/","contentTranslated":true,"sourceHash":"7a527d6593e28aee","translatedAt":"2026-08-10T11:42:50.896Z"},"tr":{"title":"Bilgi damıtmayı ölçekli çalışacak kadar ucuz hale getirmek","summary":"🔗 Orijinal makaleyi AIHOT üzerinden okuyun · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Bilgi damıtmayı ölçekli çalışacak kadar ucuz hale getirmek - Aioga AI Haberleri","description":"🔗 Orijinal makaleyi AIHOT üzerinden okuyun · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/tr/news/cmsn4c69701x9ron5wn1hfxxw/","contentTranslated":true,"sourceHash":"7a527d6593e28aee","translatedAt":"2026-08-10T11:42:59.561Z"},"vi":{"title":"Làm cho chiết xuất kiến thức đủ rẻ để vận hành ở quy mô lớn","summary":"🔗 Đọc bài gốc qua AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Làm cho chiết xuất kiến thức đủ rẻ để vận hành ở quy mô lớn - Tin tức AI Aioga","description":"🔗 Đọc bài gốc qua AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/vi/news/cmsn4c69701x9ron5wn1hfxxw/","contentTranslated":true,"sourceHash":"7a527d6593e28aee","translatedAt":"2026-08-10T11:42:57.528Z"},"id":{"title":"Membuat Knowledge Distillation Murah Cukup Untuk Dijalankan Secara Massal","summary":"🔗 Baca aslinya via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Membuat Knowledge Distillation Murah Cukup Untuk Dijalankan Secara Massal - Berita AI Aioga","description":"🔗 Baca aslinya via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/id/news/cmsn4c69701x9ron5wn1hfxxw/","contentTranslated":true,"sourceHash":"7a527d6593e28aee","translatedAt":"2026-08-10T11:43:07.590Z"},"th":{"title":"ทําให้การกลั่นกรองความรู้มีราคาถูกพอที่จะดําเนินการในระดับใหญ่","summary":"🔗 อ่านบทความต้นฉบับผ่าน AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"ทําให้การกลั่นกรองความรู้มีราคาถูกพอที่จะดําเนินการในระดับใหญ่ - ข่าว AI Aioga","description":"🔗 อ่านบทความต้นฉบับผ่าน AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/th/news/cmsn4c69701x9ron5wn1hfxxw/","contentTranslated":true,"sourceHash":"7a527d6593e28aee","translatedAt":"2026-08-10T11:43:08.196Z"},"pl":{"title":"Making Knowledge Distillation Cheap Enough to Run at Scale","summary":"🔗 Przeczytaj oryginał via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Making Knowledge Distillation Cheap Enough to Run at Scale - Aioga Wiadomości AI","description":"🔗 Przeczytaj oryginał via AIHOT · https://aihot.virxact.com/items/cmsn4c69701x9ron5wn1hfxxw","url":"https://www.aioga.com/pl/news/cmsn4c69701x9ron5wn1hfxxw/","contentTranslated":true,"sourceHash":"7a527d6593e28aee","translatedAt":"2026-08-10T11:43:15.754Z"}}}}