{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-09-28T06:03:00.468Z","headline":"安全为谁而设？边界感知自蒸馏实现话题内精准拒绝","description":"Hugging Face 新论文提出边界感知自蒸馏方法，解决模型安全对齐中\"话题级拒绝\"过于粗糙的问题。该方法通过升级重试策略将训练数据覆盖率从丢失 19.88% 提示词降至 0.20% 残差失败。","url":"https://www.aioga.com/news/cmtssojsn03txrokagwlns0fs/","mainEntityOfPage":"https://www.aioga.com/news/cmtssojsn03txrokagwlns0fs/","datePublished":"2026-09-08T14:23:07.000Z","dateModified":"2026-09-08T14:23:07.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://huggingface.co/blog/MultiverseComputingCAI/safety-for-whom","https://aihot.news/items/cmtssojsn03txrokagwlns0fs"],"canonicalUrl":"https://www.aioga.com/news/cmtssojsn03txrokagwlns0fs/","directAnswer":{"@type":"Answer","text":"Hugging Face 博客介绍一篇关于边界感知自蒸馏的论文，讨论模型如何在同一主题内区分应拒绝的有害请求与应回答的良性请求，减少话题级安全拒绝过于粗糙的问题。","url":"https://www.aioga.com/news/cmtssojsn03txrokagwlns0fs/","dateCreated":"2026-09-08T14:23:07.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"huggingface.co source article","url":"https://huggingface.co/blog/MultiverseComputingCAI/safety-for-whom","datePublished":"2026-09-08T14:23:07.000Z","provider":{"@type":"Organization","name":"huggingface.co","url":"https://huggingface.co/blog/MultiverseComputingCAI/safety-for-whom"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.news/items/cmtssojsn03txrokagwlns0fs","datePublished":"2026-09-08T14:23:07.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.news/items/cmtssojsn03txrokagwlns0fs"}}],"aggregationSource":"Hugging Face：Blog（RSS）","originalPublisher":{"name":"huggingface.co","url":"https://huggingface.co/blog/MultiverseComputingCAI/safety-for-whom"},"geoDeepAnswer":null,"article":{"id":"cmtssojsn03txrokagwlns0fs","slug":"cmtssojsn03txrokagwlns0fs","url":"https://www.aioga.com/news/cmtssojsn03txrokagwlns0fs/","title":"安全为谁而设？边界感知自蒸馏实现话题内精准拒绝","title_en":"","summary":"Hugging Face 新论文提出边界感知自蒸馏方法，解决模型安全对齐中\"话题级拒绝\"过于粗糙的问题。该方法通过升级重试策略将训练数据覆盖率从丢失 19.88% 提示词降至 0.20% 残差失败。","source":"Hugging Face：Blog（RSS）","sourceUrl":"https://huggingface.co/blog/MultiverseComputingCAI/safety-for-whom","aiHotUrl":"https://aihot.news/items/cmtssojsn03txrokagwlns0fs","publishedAt":"2026-09-08T14:23:07.000Z","category":"行业动态","score":58,"selected":false,"articleBody":["Narrow-boundary safety ：#narrow-boundary-safety Where self-generated safety tuning breaks ：#where-self-generated-safety-tuning-breaks The trade-off, and a trap it hides ：#the-trade-off-and-a-trap-it-hides What this changes ：#what-this-changes Most safety alignment work treats harm as a property of a topic. A prompt is unsafe because it falls into a general category such as weapons, fraud, or self-harm, and guard models like LlamaGuard-3：https://huggingface.co/meta-llama/Llama-Guard-3-8B encode exactly this kind of topic-level taxonomy. Benchmarks like XSTest：https://arxiv.org/abs/2308.01263 and OR-Bench：https://arxiv.org/abs/2405.20947 then probe the failure mode this creates, models that refuse safe prompts because they contain a dangerous-looking word, and refusal-calibration work tries to pull that number back down.","Real deployments rarely fit the topic-level picture. The same base model may be adapted for a general assistant, an educational product, an enterprise system, or a public-sector service, and each setting needs different boundaries within the same topic. A civics tutor and a public-sector assistant can share a model yet require opposite behaviour on politics: both should answer factual questions about an election, but only one may need to refuse a request to write targeted political manipulation. A topic-level guard cannot express that split. LlamaGuard-3, for example, covers elections only as \"factually incorrect information about electoral systems and processes,\" which excludes persuasion and manipulation and, at the same time, excludes the factual prompts a deployment must keep answering.","Our latest paper, Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal ：https://huggingface.co/papers/2609.04482, studies this narrower problem directly. The question is not whether an entire topic should be refused, but which subset of that topic is incompatible with a given deployment policy, and how to train and measure a model against that boundary.","We formalise the setting as a topic universe, all political prompts in our experiments, that contains a target-harmful subset the deployment wants to refuse. The intended policy is not to refuse all of politics, but to refuse the harmful subset while continuing to answer the benign complement. The ideal behaviour is a sharp step: refuse inside the subset, answer everywhere else in the topic.","：https://cdn-uploads.huggingface.co/production/uploads/668e37fd9c9aa124a3c867e8/g_Xngr05EUhlIBNT-mG8s.png","The narrow-boundary setting. A deployment may need to refuse only the political prompts that ask for manipulation or targeted persuasion, while still answering other political prompts, rather than refusing all of politics. A trained model's refusal is smoother than the ideal split and can spill into benign territory near the boundary. Source: paper Figure 1.","A trained model never learns that sharp step. It learns a refusal probability that only approximates the target, and cross-entropy training that raises refusal inside the harmful subset can also push refusal outward into the benign complement. So the real problem is not only raising refusal on harmful prompts, but shaping the behaviour near the boundary itself. We operationalise that boundary as pairs of prompts that share a topic anchor and differ only in intent, one that should be refused and one that should be answered.","We use political persuasion as the testbed, since manipulative persuasion can cause real harm while factual political information stays legitimate, which is exactly the case where topic-level refusal is too blunt.","The natural way to build training data here is self-generation: take the target model, steer it toward a refusal on each harmful prompt, and keep the traces a guard model verifies as genuine refusals. This is the recipe behind methods like ThinkSafe：https://arxiv.org/abs/2601.23143, and we adopt it as our reference, applied to political prompts and measured component by component. Framing the problem as a boundary rather than a topic exposes three weaknesses in that standard pipeline.","The first is a coverage gap. A single steering attempt does not always produce an accepted refusal, and those prompts are silently dropped from the training set. In our audited pool, single-shot generation drops 19.88% of prompts, 8,009 of them, and these failed prompts may well be the hardest examples. We repair this instead of discarding it: an escalating retry strategy, resampling the same prompt through progressively stronger steering, brings the residual failures down to 0.20%, or 79 prompts. Coverage repair leaves 40,293 harmful training prompts where the naive pipeline would have thrown thousands away.","The second is downside reactions. Safety tuning tends to produce false refusals on benign prompts that look superficially dangerous. To compensate, we build in-distribution benign data, including 11,955 verified surface-dangerous benign prompts across 18 semantic types, so the model sees safe prompts with dangerous-looking wording during training rather than only at evaluation.","The third is that ordinary harmful and benign splits do not measure the shape of the boundary at all. A model can improve its harmful-refusal rate simply by expanding refusal into nearby permissible prompts, and a topic-level metric will call that an improvement. Held-out harmful-benign pairs, 1,539 per side, let us measure both sides of the boundary directly.","Training on political refusal data works in the obvious sense. On Qwen3-8B：https://huggingface.co/Qwen/Qwen3-8B, the escalated-coverage model raises in-distribution political refusal from 9.47% to 84.75%, and it also transfers: the mean unsafe-response rate across three broader harmfulness benchmarks, HarmBench：https://www.harmbench.org, StrongREJECT, and WildJailbreak：https://arxiv.org/abs/2406.18510, scored by LlamaGuard-3, falls from 26.26% to 0.14% in the strongest configuration.","Reported alone, those numbers look like a clean win. They are not. At the same checkpoint, over-refusal on XSTest rises from 2.00% to 74.00%. The configuration with the lowest harmful-response rate is also the one that refuses nearly three quarters of plainly safe prompts. It is a blunt refusal machine, not a safer model, and you cannot see that unless you measure the benign side. This is the central message: data composition decides where a checkpoint sits in the space of safety against over-refusal, so the two axes have to be reported together.","Two of our data components pull the over-refusal number back down without giving up the safety gain. Replacing externally adopted compliance responses with verified responses generated by the target model itself lowers XSTest over-refusal from 15.20% to 5.20% under single-shot generation, at a modest harmfulness cost. And the harmful-benign boundary pairs do the most precise work of all.","：https://cdn-uploads.huggingface.co/production/uploads/68db932961906f42259438b7/6W8oCYwy3TnSfdaBsH74C.png","Left: over-refusal on the comply-worthy side of the held-out boundary, lower is better. Runs with the benign boundary data (PB) fall to 0.03 to 0.08; without it, the number rises toward 0.49. Right: refusal on the harmful side, higher is better, which falls only slightly. Source: paper Figure 6.","Concretely, adding the benign boundary data reduces over-refusal on the comply-worthy side of the held-out pairs from 32.94% to 4.16%. Refusal on the harmful side drops only from 91.88% to 87.72%. In other words, most of the false refusals near the boundary disappear while almost all of the genuine refusals survive. There is a real recall cost, and it is small and measurable, which is the point: you can only trade it off deliberately if you are measuring both sides.","The practical takeaway is that safety tuning should not be assessed by harmful-refusal rate alone. A model that refuses more is not automatically safer, and on a narrow boundary the same move that raises refusal on harmful prompts can quietly make the model useless on the legitimate prompts right next to them. Composition of the training data, coverage repair, in-distribution compensation, and boundary pairs are what control that trade-off, and both sides of the intended boundary have to be evaluated for the numbers to mean anything.","This work is part of Multiverse Computing's：https://multiversecomputing.com research into making model behaviour controllable and measurable at the level real deployments care about, rather than at the level of broad topic categories. The same generation pipeline extends to other topics beyond politics, and the paper reports the full set of data-composition ablations behind the results above.","Want the full technical details, including the coverage-repair strategies, the loss routing that separates harmful cross-entropy from benign forward-KL preservation, and the complete held-out boundary evaluation? Read the full paper, or get in touch with our team to talk about deployment-specific safety for your own models."],"articleImages":[{"sourceUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1632247447995-noauth.jpeg","alt":"","afterParagraph":0,"url":"/media/articles/cmtssojsn03txrokagwlns0fs/f607df82a17eab06.webp"},{"sourceUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ZSvndCicLw5gtysncpXOn.png","alt":"","afterParagraph":0,"url":"/media/articles/cmtssojsn03txrokagwlns0fs/698a9295db50e5ef.webp"},{"sourceUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67712fbc8219c3d8608ffe4b/0LN7kqWd3T7eXqi3-liJU.jpeg","alt":"","afterParagraph":0,"url":"/media/articles/cmtssojsn03txrokagwlns0fs/de326de0f60a4249.webp"},{"sourceUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/dd5CkYFbznEo_I9FFnhon.png","alt":"","afterParagraph":0,"url":"/media/articles/cmtssojsn03txrokagwlns0fs/0d597b42b20e1674.webp"}],"mediaStatus":"ok","articleBodyZh":["狭界安全：#narrow-boundary-safety 自我生成的安全调优破裂的地方：#where-self-generated-safety-tuning-breaks 权衡及其隐藏的陷阱：#the-trade-off-and-a-trap-it-hides 这些改变了什么：#what-this-changes 大多数安全对齐工作将伤害视为主题的属性。一个提示被认为是不安全的，因为它属于一个通用类别，例如武器、欺诈或自残，而像 LlamaGuard-3：https://huggingface.co/meta-llama/Llama-Guard-3-8B 这样的防护模型正是编码了这种主题级分类。像 XSTest：https://arxiv.org/abs/2308.01263 和 OR-Bench：https://arxiv.org/abs/2405.20947 这样的基准测试随后探测这种失败模式，即模型因包含看似危险的词而拒绝安全提示，而拒绝校准工作尝试将这一数字拉回到较低水平。","实际部署很少符合主题级的情况。同一个基础模型可能会适配为通用助手、教育产品、企业系统或公共部门服务，而每种场景在同一主题下需要不同的边界。公民教育导师和公共部门助手可以共享模型，但在政治问题上需要相反的行为：两者都应回答关于选举的事实问题，但只有一个可能需要拒绝写作针对特定政治操纵的请求。主题级防护无法表达这种分裂。例如，LlamaGuard-3 仅将选举覆盖为“有关选举系统和流程的事实错误信息”，这排除了劝说与操控，同时也排除了部署必须继续回答的事实性提示。","我们最新的论文，《为谁安全？用于受控 LLM 安全拒绝的边界感知自蒸馏》：https://huggingface.co/papers/2609.04482，直接研究这一更狭窄的问题。问题不在于是否应该拒绝整个主题，而在于该主题的哪些子集与给定部署策略不兼容，以及如何针对该边界训练和评估模型。","我们将设置形式化为一个主题宇宙，也就是我们实验中的所有政治提示，其中包含部署方希望拒绝的目标有害子集。预期的策略不是拒绝所有政治内容，而是拒绝有害子集，同时继续回答良性子集。理想的行为是一个尖锐的分界：在子集内拒绝，在主题的其他地方回答。","：https://cdn-uploads.huggingface.co/production/uploads/668e37fd9c9aa124a3c867e8/g_Xngr05EUhlIBNT-mG8s.png","窄边界设置。部署可能只需要拒绝那些要求操控或针对性说服的政治提示，同时仍然回答其他政治提示，而不是拒绝所有政治内容。训练后的模型的拒绝比理想划分更平滑，并可能在边界附近溢入良性区域。来源：论文图1。","训练后的模型从未学习到那种尖锐分界。它学习的只是近似目标的拒绝概率，而在有害子集内通过交叉熵训练提高拒绝率也可能推动拒绝扩展到良性子集。因此，真正的问题不仅是提高对有害提示的拒绝，还包括在边界附近调整行为。我们将该边界操作化为具有相同主题锚点但意图不同的提示对，一个应被拒绝，一个应被回答。","我们使用政治说服作为测试场景，因为操控性说服可能造成实际伤害，而事实性的政治信息仍然合法，这正是主题级拒绝过于粗糙的情况。","在这里构建训练数据的自然方式是自我生成：让目标模型在每个有害提示上倾向于拒绝，并保留守护模型验证为真实拒绝的痕迹。这就是像ThinkSafe方法背后的流程：https://arxiv.org/abs/2601.23143，我们采用它作为参考，应用于政治提示，并逐个组件进行测量。将问题框定为边界而非主题，揭示了该标准流程的三个弱点。","第一个是覆盖缺口。一次单独的引导尝试并不总是会产生被接受的拒绝，而那些提示会被静默地从训练集中丢弃。在我们审计的样本池中，单次生成会丢弃19.88%的提示，即8,009个提示，这些失败的提示很可能是最难的例子。我们采用修复而非丢弃的方法：使用逐步升级的重试策略，通过逐渐增强的引导重复采样相同的提示，使剩余失败率降至0.20%，即79个提示。覆盖修复使得有害训练提示仍保留40,293个，而简单的管道程序本会丢弃几千个。","第二个是负面反应。安全调优往往会对表面上看起来危险的无害提示产生错误拒绝。为了补偿，我们构建了分布内的无害数据，包括跨18个语义类型的11,955个经过验证的表面危险但实际无害的提示，使模型在训练期间就能看到带有危险表述的安全提示，而不仅仅是在评估时看到。","第三，普通的有害与无害划分根本无法衡量边界的形态。模型可以通过将拒绝扩展到附近的允许提示来提高其有害拒绝率，而主题级别的指标会将其视为改进。保留的有害-无害对，每边1,539个，让我们可以直接测量边界的两侧。","在政治拒绝数据上训练在明显意义上是有效的。在Qwen3-8B：https://huggingface.co/Qwen/Qwen3-8B上，升级覆盖模型将分布内政治拒绝率从9.47%提高到84.75%，并且具有迁移效果：在三个更广泛的有害性基准测试HarmBench：https://www.harmbench.org、StrongREJECT和WildJailbreak：https://arxiv.org/abs/2406.18510中，按LlamaGuard-3评分的平均不安全响应率在最强配置下从26.26%降至0.14%。","单独报告这些数字，看起来像是干净的胜利。实际上并不是。在同一个检查点上，XSTest 的过度拒绝率从 2.00% 上升到 74.00%。有害响应率最低的配置，恰好也是几乎拒绝四分之三明显安全提示的配置。它是一个粗暴的拒绝机器，而不是更安全的模型，除非测量 benign（无害）方面，否则你无法看出这一点。这是中心信息：数据组成决定了一个检查点在安全性与过度拒绝的空间中的位置，因此必须同时报告两个维度。","我们的两个数据组件在不放弃安全增益的情况下，将过度拒绝率降低。用目标模型自身生成的已验证响应替换外部采用的合规响应，在单次生成下将 XSTest 的过度拒绝率从 15.20% 降低到 5.20%，带来适度的有害性成本。而有害-无害边界对的数据对提供了最精确的调节效果。","：https://cdn-uploads.huggingface.co/production/uploads/68db932961906f42259438b7/6W8oCYwy3TnSfdaBsH74C.png","左图：在保留边界的应遵守一侧的过度拒绝，越低越好。使用无害边界数据（PB）的运行结果下降到 0.03 到 0.08；未使用时，数值上升至约 0.49。右图：在有害一侧的拒绝，越高越好，下降幅度很小。来源：论文图 6。","具体来说，增加无害边界数据将保留对的应遵守一侧的过度拒绝率从 32.94% 降至 4.16%。有害一侧的拒绝率仅从 91.88% 降至 87.72%。换句话说，大部分接近边界的错误拒绝消失，而几乎所有的真正拒绝保留。确实存在召回成本，但它小且可测，这是关键：只有在同时测量两方面时，你才能有意识地进行权衡。","实际的要点是，安全调优不应仅通过有害拒绝率来评估。一个拒绝更多的模型并不自动更安全，而且在狭窄的边界上，同样的操作在提高有害提示的拒绝率的同时，可能默默地让模型在紧邻的合法提示上变得无用。训练数据的组成、覆盖修复、分布内补偿以及边界对决定了这种权衡，并且必须评估目标边界的两侧，数字才能有意义。","这项工作是多重宇宙计算的一部分：https://multiversecomputing.com，该研究旨在使模型行为在实际部署关心的级别（而不是在广泛的主题类别级别）可控和可测量。同一代管道扩展到政治以外的其他主题，该论文报告了上述结果背后的全套数据构成消融。","想要了解完整的技术细节，包括覆盖修复策略、将有害交叉熵与良性前向KL保持分开的损失路由，以及完整的保留边界评估吗？请阅读完整论文，或联系我们的团队，讨论您自身模型的部署特定安全性。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Hugging Face 博客介绍一篇关于边界感知自蒸馏的论文，讨论模型如何在同一主题内区分应拒绝的有害请求与应回答的良性请求，减少话题级安全拒绝过于粗糙的问题。","background":"材料称，现有安全对齐常把风险视为主题属性，可能导致模型因提示词包含危险词而拒绝安全请求。论文以政治说服为测试场景，要求模型拒绝操纵性说服或定向说服请求，同时继续回答事实性政治信息。","viewpoint":"Aioga 判断：该材料的核心价值在于把安全拒答从主题分类推进到意图边界识别。来源展示的是论文提出的问题框架与方法方向，尚不足以据此确认其在更多部署场景中的普遍效果。","implications":"可能影响：若方法在相关评测中有效，模型安全策略可能需要同时关注有害请求的拒答与临近边界的良性请求保留；但材料不足以证明该方法已经适用于所有主题、产品或政策环境。","nextStep":"后续观察：需要关注论文完整实验对边界识别、误拒答和残差失败的评测结果，以及该方法在不同部署政策和其他主题中的表现。现有材料未明确说明实际产品采用情况。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-09-08T15:30:04.950Z","sourceHash":"cfc2fb25a085984e","review":{"approved":true,"groundedness":96,"clarity":94,"duplicationRisk":10,"blockingIssues":[],"notes":["“Aioga 判断”在来源材料中未出现，若非固定署名，建议删除或改为“材料认为”，以避免引入不明主体。","evidenceRefs 可进一步对应到具体段落或引文，但不影响当前内容的事实依据。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","editorial-labels","inference-boundary","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","Hugging Face：Blog（RSS）"],"translations":{"zh-CN":{"title":"安全为谁而设？边界感知自蒸馏实现话题内精准拒绝","summary":"Hugging Face 新论文提出边界感知自蒸馏方法，解决模型安全对齐中\"话题级拒绝\"过于粗糙的问题。该方法通过升级重试策略将训练数据覆盖率从丢失 19.88% 提示词降至 0.20% 残差失败。","category":"行业动态","source":"huggingface.co","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"安全为谁而设？边界感知自蒸馏实现话题内精准拒绝 - Aioga AI资讯","description":"Hugging Face 新论文提出边界感知自蒸馏方法，解决模型安全对齐中\"话题级拒绝\"过于粗糙的问题。该方法通过升级重试策略将训练数据覆盖率从丢失 19.88% 提示词降至 0.20% 残差失败。","url":"https://www.aioga.com/news/cmtssojsn03txrokagwlns0fs/","articleBody":["狭界安全：#narrow-boundary-safety 自我生成的安全调优破裂的地方：#where-self-generated-safety-tuning-breaks 权衡及其隐藏的陷阱：#the-trade-off-and-a-trap-it-hides 这些改变了什么：#what-this-changes 大多数安全对齐工作将伤害视为主题的属性。一个提示被认为是不安全的，因为它属于一个通用类别，例如武器、欺诈或自残，而像 LlamaGuard-3：https://huggingface.co/meta-llama/Llama-Guard-3-8B 这样的防护模型正是编码了这种主题级分类。像 XSTest：https://arxiv.org/abs/2308.01263 和 OR-Bench：https://arxiv.org/abs/2405.20947 这样的基准测试随后探测这种失败模式，即模型因包含看似危险的词而拒绝安全提示，而拒绝校准工作尝试将这一数字拉回到较低水平。","实际部署很少符合主题级的情况。同一个基础模型可能会适配为通用助手、教育产品、企业系统或公共部门服务，而每种场景在同一主题下需要不同的边界。公民教育导师和公共部门助手可以共享模型，但在政治问题上需要相反的行为：两者都应回答关于选举的事实问题，但只有一个可能需要拒绝写作针对特定政治操纵的请求。主题级防护无法表达这种分裂。例如，LlamaGuard-3 仅将选举覆盖为“有关选举系统和流程的事实错误信息”，这排除了劝说与操控，同时也排除了部署必须继续回答的事实性提示。","我们最新的论文，《为谁安全？用于受控 LLM 安全拒绝的边界感知自蒸馏》：https://huggingface.co/papers/2609.04482，直接研究这一更狭窄的问题。问题不在于是否应该拒绝整个主题，而在于该主题的哪些子集与给定部署策略不兼容，以及如何针对该边界训练和评估模型。","我们将设置形式化为一个主题宇宙，也就是我们实验中的所有政治提示，其中包含部署方希望拒绝的目标有害子集。预期的策略不是拒绝所有政治内容，而是拒绝有害子集，同时继续回答良性子集。理想的行为是一个尖锐的分界：在子集内拒绝，在主题的其他地方回答。","：https://cdn-uploads.huggingface.co/production/uploads/668e37fd9c9aa124a3c867e8/g_Xngr05EUhlIBNT-mG8s.png","窄边界设置。部署可能只需要拒绝那些要求操控或针对性说服的政治提示，同时仍然回答其他政治提示，而不是拒绝所有政治内容。训练后的模型的拒绝比理想划分更平滑，并可能在边界附近溢入良性区域。来源：论文图1。","训练后的模型从未学习到那种尖锐分界。它学习的只是近似目标的拒绝概率，而在有害子集内通过交叉熵训练提高拒绝率也可能推动拒绝扩展到良性子集。因此，真正的问题不仅是提高对有害提示的拒绝，还包括在边界附近调整行为。我们将该边界操作化为具有相同主题锚点但意图不同的提示对，一个应被拒绝，一个应被回答。","我们使用政治说服作为测试场景，因为操控性说服可能造成实际伤害，而事实性的政治信息仍然合法，这正是主题级拒绝过于粗糙的情况。","在这里构建训练数据的自然方式是自我生成：让目标模型在每个有害提示上倾向于拒绝，并保留守护模型验证为真实拒绝的痕迹。这就是像ThinkSafe方法背后的流程：https://arxiv.org/abs/2601.23143，我们采用它作为参考，应用于政治提示，并逐个组件进行测量。将问题框定为边界而非主题，揭示了该标准流程的三个弱点。","第一个是覆盖缺口。一次单独的引导尝试并不总是会产生被接受的拒绝，而那些提示会被静默地从训练集中丢弃。在我们审计的样本池中，单次生成会丢弃19.88%的提示，即8,009个提示，这些失败的提示很可能是最难的例子。我们采用修复而非丢弃的方法：使用逐步升级的重试策略，通过逐渐增强的引导重复采样相同的提示，使剩余失败率降至0.20%，即79个提示。覆盖修复使得有害训练提示仍保留40,293个，而简单的管道程序本会丢弃几千个。","第二个是负面反应。安全调优往往会对表面上看起来危险的无害提示产生错误拒绝。为了补偿，我们构建了分布内的无害数据，包括跨18个语义类型的11,955个经过验证的表面危险但实际无害的提示，使模型在训练期间就能看到带有危险表述的安全提示，而不仅仅是在评估时看到。","第三，普通的有害与无害划分根本无法衡量边界的形态。模型可以通过将拒绝扩展到附近的允许提示来提高其有害拒绝率，而主题级别的指标会将其视为改进。保留的有害-无害对，每边1,539个，让我们可以直接测量边界的两侧。","在政治拒绝数据上训练在明显意义上是有效的。在Qwen3-8B：https://huggingface.co/Qwen/Qwen3-8B上，升级覆盖模型将分布内政治拒绝率从9.47%提高到84.75%，并且具有迁移效果：在三个更广泛的有害性基准测试HarmBench：https://www.harmbench.org、StrongREJECT和WildJailbreak：https://arxiv.org/abs/2406.18510中，按LlamaGuard-3评分的平均不安全响应率在最强配置下从26.26%降至0.14%。","单独报告这些数字，看起来像是干净的胜利。实际上并不是。在同一个检查点上，XSTest 的过度拒绝率从 2.00% 上升到 74.00%。有害响应率最低的配置，恰好也是几乎拒绝四分之三明显安全提示的配置。它是一个粗暴的拒绝机器，而不是更安全的模型，除非测量 benign（无害）方面，否则你无法看出这一点。这是中心信息：数据组成决定了一个检查点在安全性与过度拒绝的空间中的位置，因此必须同时报告两个维度。","我们的两个数据组件在不放弃安全增益的情况下，将过度拒绝率降低。用目标模型自身生成的已验证响应替换外部采用的合规响应，在单次生成下将 XSTest 的过度拒绝率从 15.20% 降低到 5.20%，带来适度的有害性成本。而有害-无害边界对的数据对提供了最精确的调节效果。","：https://cdn-uploads.huggingface.co/production/uploads/68db932961906f42259438b7/6W8oCYwy3TnSfdaBsH74C.png","左图：在保留边界的应遵守一侧的过度拒绝，越低越好。使用无害边界数据（PB）的运行结果下降到 0.03 到 0.08；未使用时，数值上升至约 0.49。右图：在有害一侧的拒绝，越高越好，下降幅度很小。来源：论文图 6。","具体来说，增加无害边界数据将保留对的应遵守一侧的过度拒绝率从 32.94% 降至 4.16%。有害一侧的拒绝率仅从 91.88% 降至 87.72%。换句话说，大部分接近边界的错误拒绝消失，而几乎所有的真正拒绝保留。确实存在召回成本，但它小且可测，这是关键：只有在同时测量两方面时，你才能有意识地进行权衡。","实际的要点是，安全调优不应仅通过有害拒绝率来评估。一个拒绝更多的模型并不自动更安全，而且在狭窄的边界上，同样的操作在提高有害提示的拒绝率的同时，可能默默地让模型在紧邻的合法提示上变得无用。训练数据的组成、覆盖修复、分布内补偿以及边界对决定了这种权衡，并且必须评估目标边界的两侧，数字才能有意义。","这项工作是多重宇宙计算的一部分：https://multiversecomputing.com，该研究旨在使模型行为在实际部署关心的级别（而不是在广泛的主题类别级别）可控和可测量。同一代管道扩展到政治以外的其他主题，该论文报告了上述结果背后的全套数据构成消融。","想要了解完整的技术细节，包括覆盖修复策略、将有害交叉熵与良性前向KL保持分开的损失路由，以及完整的保留边界评估吗？请阅读完整论文，或联系我们的团队，讨论您自身模型的部署特定安全性。"]},"en":{"title":"Who is security for? Boundary-aware self-distillation enables precise rejection within topics","summary":"A new paper by Hugging Face proposes a boundary-aware self-distillation method to address the problem of overly rough \"topic-level rejection\" in model secure alignment. This method reduces training data coverage from a loss of 19.88% prompt to 0.20% residual failure by upgrading the retry strategy.","category":"Industry","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Who is security for? Boundary-aware self-distillation enables precise rejection within topics - Aioga AI News","description":"A new paper by Hugging Face proposes a boundary-aware self-distillation method to address the problem of overly rough \"topic-level rejection\" in model secure alignment. This method...","url":"https://www.aioga.com/en/news/cmtssojsn03txrokagwlns0fs/","contentTranslated":true,"sourceHash":"6e7b4d18bf184ebe","translatedAt":"2026-09-08T15:23:51.628Z"},"ja":{"title":"安全は誰のためのものなのでしょうか? 境界認識型自己蒸留は、トピック内で正確なリジェクションを可能にします","summary":"Hajin Fasの新しい論文は、モデル安全性アライメントにおける過度に粗い「トピックレベルリジェクション」問題に対処するために境界認識型自己蒸留法を提案しています。 この方法は、再試行戦略をアップグレードすることで、トレーニングデータのカバレッジを19.88%のプロンプト損失から0.20%の残留失敗に削減します。","category":"業界動向","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"安全は誰のためのものなのでしょうか? 境界認識型自己蒸留は、トピック内で正確なリジェクションを可能にします - Aioga AIニュース","description":"Hajin Fasの新しい論文は、モデル安全性アライメントにおける過度に粗い「トピックレベルリジェクション」問題に対処するために境界認識型自己蒸留法を提案しています。 この方法は、再試行戦略をアップグレードすることで、トレーニングデータのカバレッジを19.88%のプロンプト損失から0.20%の残留失敗に削減します。","url":"https://www.aioga.com/ja/news/cmtssojsn03txrokagwlns0fs/","contentTranslated":true,"sourceHash":"6e7b4d18bf184ebe","translatedAt":"2026-09-08T15:23:52.491Z"},"ko":{"title":"안전은 누구를 위한 것인가요? 경계 인지 자기 증류는 주제 내에서 정밀한 거부를 가능하게 합니다","summary":"Hajin Fas의 새 논문은 모델 안전성 정렬에서 지나치게 거친 '주제 수준 거부' 문제를 해결하기 위해 경계 인지 자기 증류법을 제안합니다. 이 방법은 재시도 전략을 업그레이드하여 학습 데이터 적용률을 19.88% 잃는 프롬프트에서 0.20%의 잔여 실패로 줄입니다.","category":"업계 동향","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"안전은 누구를 위한 것인가요? 경계 인지 자기 증류는 주제 내에서 정밀한 거부를 가능하게 합니다 - Aioga AI 뉴스","description":"Hajin Fas의 새 논문은 모델 안전성 정렬에서 지나치게 거친 '주제 수준 거부' 문제를 해결하기 위해 경계 인지 자기 증류법을 제안합니다. 이 방법은 재시도 전략을 업그레이드하여 학습 데이터 적용률을 19.88% 잃는 프롬프트에서 0.20%의 잔여 실패로 줄입니다.","url":"https://www.aioga.com/ko/news/cmtssojsn03txrokagwlns0fs/","contentTranslated":true,"sourceHash":"6e7b4d18bf184ebe","translatedAt":"2026-09-08T15:24:02.056Z"},"es":{"title":"¿Para quién es la seguridad? La autodestilación consciente de los límites permite un rechazo preciso dentro de los temas","summary":"El nuevo artículo de Hajin Fas propone un método de autodestilación consciente de límites para abordar el problema del \"rechazo a nivel temático\" excesivamente aproximado en la alineación de seguridad del modelo. Este método reduce la cobertura de datos de entrenamiento de perder un 19,88% de prompts a un fallo residual del 0,20% al actualizar la estrategia de reintentos.","category":"Industria","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"¿Para quién es la seguridad? La autodestilación consciente de los límites permite un rechazo preciso dentro de los temas - Aioga Noticias de IA","description":"El nuevo artículo de Hajin Fas propone un método de autodestilación consciente de límites para abordar el problema del \"rechazo a nivel temático\" excesivamente aproximado en la ali...","url":"https://www.aioga.com/es/news/cmtssojsn03txrokagwlns0fs/","contentTranslated":true,"sourceHash":"6e7b4d18bf184ebe","translatedAt":"2026-09-08T15:24:02.245Z"},"fr":{"title":"Pour qui est la sécurité ? L’auto-distillation consciente des frontières permet un rejet précis dans les sujets","summary":"Le nouvel article de Hajin Fas propose une méthode d’auto-distillation consciente des frontières pour résoudre le problème d’un « rejet au niveau thématique » trop approximatif dans l’alignement de sécurité des modèles. Cette méthode réduit la couverture des données d’entraînement, passant de 19,88 % d’invites à 0,20 % d’échec résiduel en mettant à jour la stratégie de réessayage.","category":"Industrie","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Pour qui est la sécurité ? L’auto-distillation consciente des frontières permet un rejet précis dans les sujets - Aioga Actualités IA","description":"Le nouvel article de Hajin Fas propose une méthode d’auto-distillation consciente des frontières pour résoudre le problème d’un « rejet au niveau thématique » trop approximatif dan...","url":"https://www.aioga.com/fr/news/cmtssojsn03txrokagwlns0fs/","contentTranslated":true,"sourceHash":"6e7b4d18bf184ebe","translatedAt":"2026-09-08T15:24:11.502Z"},"de":{"title":"Für wen ist Sicherheit da? Grenzbewusste Selbstdestillation ermöglicht eine präzise Ablehnung innerhalb von Themen","summary":"Hajin Fas' neues Papier schlägt eine grenzenbewusste Selbstdestillationsmethode vor, um das Problem der zu groben \"themenbezogenen Ablehnung\" bei der Modellsicherheit anzugehen. Diese Methode reduziert die Trainingsdatenabdeckung von 19,88 % Eingaben auf 0,20 % Restfehler, indem die Wiederversuchsstrategie aufgerüstet wird.","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Für wen ist Sicherheit da? Grenzbewusste Selbstdestillation ermöglicht eine präzise Ablehnung innerhalb von Themen - Aioga KI-News","description":"Hajin Fas' neues Papier schlägt eine grenzenbewusste Selbstdestillationsmethode vor, um das Problem der zu groben \"themenbezogenen Ablehnung\" bei der Modellsicherheit anzugehen. Di...","url":"https://www.aioga.com/de/news/cmtssojsn03txrokagwlns0fs/","contentTranslated":true,"sourceHash":"6e7b4d18bf184ebe","translatedAt":"2026-09-08T15:24:12.103Z"},"pt-BR":{"title":"Para quem é a segurança? A autodestilação consciente de fronteiras permite uma rejeição precisa dentro dos tópicos","summary":"O novo artigo de Hajin Fas propõe um método de autodestilação consciente de fronteiras para abordar o problema da rejeição excessivamente aproximada de \"rejeição em nível tópico\" no alinhamento de segurança do modelo. Esse método reduz a cobertura dos dados de treinamento de perder 19,88% de prompts para 0,20% de falha residual ao atualizar a estratégia de retentativa.","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Para quem é a segurança? A autodestilação consciente de fronteiras permite uma rejeição precisa dentro dos tópicos - Aioga Notícias de IA","description":"O novo artigo de Hajin Fas propõe um método de autodestilação consciente de fronteiras para abordar o problema da rejeição excessivamente aproximada de \"rejeição em nível tópico\" n...","url":"https://www.aioga.com/pt-BR/news/cmtssojsn03txrokagwlns0fs/","contentTranslated":true,"sourceHash":"6e7b4d18bf184ebe","translatedAt":"2026-09-08T15:24:21.547Z"},"ru":{"title":"Для кого нужна безопасность? Самодистилляция, осознанная границами, позволяет точно отвергать темы","summary":"В новой статье Хаджин Фас предлагается метод самодистилляции с учётом границ, чтобы решить проблему чрезмерно грубого «отклонения на уровне темы» при выравнивании безопасности моделей. Этот метод сокращает покрытие обучающих данных с потери 19,88% запросов до 0,20% остаточных неудач за счёт обновления стратегии повторных попыток.","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Для кого нужна безопасность? Самодистилляция, осознанная границами, позволяет точно отвергать темы - Aioga Новости ИИ","description":"В новой статье Хаджин Фас предлагается метод самодистилляции с учётом границ, чтобы решить проблему чрезмерно грубого «отклонения на уровне темы» при выравнивании безопасности моде...","url":"https://www.aioga.com/ru/news/cmtssojsn03txrokagwlns0fs/","contentTranslated":true,"sourceHash":"6e7b4d18bf184ebe","translatedAt":"2026-09-08T15:24:21.774Z"},"ar":{"title":"لمن السلامة المناسبة؟ التقطير الذاتي الواعي للحدود يتيح الرفض الدقيق داخل المواضيع","summary":"تقترح ورقة هاجين فاس الجديدة طريقة التقطير الذاتي الواعية للحدود لمعالجة مشكلة \"الرفض على مستوى الموضوع\" المفرط في محاذاة سلامة النماذج. تقلل هذه الطريقة من تغطية بيانات التدريب من فقدان 19.88٪ من المحفزات إلى 0.20٪ فشل متبقي من خلال ترقية استراتيجية إعادة المحاولة.","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"لمن السلامة المناسبة؟ التقطير الذاتي الواعي للحدود يتيح الرفض الدقيق داخل المواضيع - Aioga أخبار الذكاء الاصطناعي","description":"تقترح ورقة هاجين فاس الجديدة طريقة التقطير الذاتي الواعية للحدود لمعالجة مشكلة \"الرفض على مستوى الموضوع\" المفرط في محاذاة سلامة النماذج. تقلل هذه الطريقة من تغطية بيانات التدريب من...","url":"https://www.aioga.com/ar/news/cmtssojsn03txrokagwlns0fs/","contentTranslated":true,"sourceHash":"6e7b4d18bf184ebe","translatedAt":"2026-09-08T15:24:31.457Z"},"hi":{"title":"सुरक्षा किसके लिए है? सीमा-जागरूक आत्म-आसवन विषयों के भीतर सटीक अस्वीकृति को सक्षम बनाता है","summary":"हाजिन फास का नया पेपर मॉडल सुरक्षा संरेखण में अत्यधिक खुरदरे \"विषय-स्तरीय अस्वीकृति\" की समस्या का समाधान करने के लिए एक सीमा-जागरूक आत्म-आसवन पद्धति का प्रस्ताव करता है। यह विधि पुन: प्रयास रणनीति को अपग्रेड करके प्रशिक्षण डेटा कवरेज को 19.88% संकेतों को खोने से 0.20% अवशिष्ट विफलता तक कम कर देती है।","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"सुरक्षा किसके लिए है? सीमा-जागरूक आत्म-आसवन विषयों के भीतर सटीक अस्वीकृति को सक्षम बनाता है - Aioga AI समाचार","description":"हाजिन फास का नया पेपर मॉडल सुरक्षा संरेखण में अत्यधिक खुरदरे \"विषय-स्तरीय अस्वीकृति\" की समस्या का समाधान करने के लिए एक सीमा-जागरूक आत्म-आसवन पद्धति का प्रस्ताव करता है। यह विधि पु...","url":"https://www.aioga.com/hi/news/cmtssojsn03txrokagwlns0fs/","contentTranslated":true,"sourceHash":"6e7b4d18bf184ebe","translatedAt":"2026-09-08T15:24:31.516Z"},"it":{"title":"Per chi è la sicurezza? L'auto-distillazione consapevole dei confini consente un rifiuto preciso all'interno degli argomenti","summary":"Il nuovo articolo di Hajin Fas propone un metodo di auto-distillazione consapevole dei confini per affrontare il problema del \"rifiuto a livello topico\" eccessivamente approssimativo nell'allineamento della sicurezza dei modelli. Questo metodo riduce la copertura dei dati di addestramento da una perdita del 19,88% dei prompt allo 0,20% di fallimento residuo aggiornando la strategia di ritento.","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Per chi è la sicurezza? L'auto-distillazione consapevole dei confini consente un rifiuto preciso all'interno degli argomenti - Aioga Notizie IA","description":"Il nuovo articolo di Hajin Fas propone un metodo di auto-distillazione consapevole dei confini per affrontare il problema del \"rifiuto a livello topico\" eccessivamente approssimati...","url":"https://www.aioga.com/it/news/cmtssojsn03txrokagwlns0fs/","contentTranslated":true,"sourceHash":"6e7b4d18bf184ebe","translatedAt":"2026-09-08T15:24:40.938Z"},"nl":{"title":"Voor wie is veiligheid? Grensbewuste zelfdistillatie maakt precieze afwijzing binnen onderwerpen mogelijk","summary":"Het nieuwe artikel van Hajin Fas stelt een grensbewuste zelfdestillatiemethode voor om het probleem van te ruwe \"onderwerpniveau afwijzing\" bij modelveiligheidsuitlijning aan te pakken. Deze methode vermindert de dekking van trainingsdata van het verliezen van 19,88% prompts tot 0,20% residuele falen door de herkansingsstrategie te upgraden.","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Voor wie is veiligheid? Grensbewuste zelfdistillatie maakt precieze afwijzing binnen onderwerpen mogelijk - Aioga AI-nieuws","description":"Het nieuwe artikel van Hajin Fas stelt een grensbewuste zelfdestillatiemethode voor om het probleem van te ruwe \"onderwerpniveau afwijzing\" bij modelveiligheidsuitlijning aan te pa...","url":"https://www.aioga.com/nl/news/cmtssojsn03txrokagwlns0fs/","contentTranslated":true,"sourceHash":"6e7b4d18bf184ebe","translatedAt":"2026-09-08T15:24:40.865Z"},"tr":{"title":"Güvenlik kim için? Sınırların farkında kendi-damıtma, konular içinde kesin reddedilmeyi mümkün kılar","summary":"Hajin Fas'ın yeni makalesi, model güvenliği hizalamasında aşırı kaba \"konu düzeyinde reddedilme\" sorununu ele almak için sınır farkında bir kendi-damıtma yöntemi öneriyor. Bu yöntem, tekrar deneme stratejisini yükselterek eğitim verisi kapsamını %19,88 istem kaybından %0,20 kalıntı başarısızlığa indirir.","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Güvenlik kim için? Sınırların farkında kendi-damıtma, konular içinde kesin reddedilmeyi mümkün kılar - Aioga AI Haberleri","description":"Hajin Fas'ın yeni makalesi, model güvenliği hizalamasında aşırı kaba \"konu düzeyinde reddedilme\" sorununu ele almak için sınır farkında bir kendi-damıtma yöntemi öneriyor. Bu yönte...","url":"https://www.aioga.com/tr/news/cmtssojsn03txrokagwlns0fs/","contentTranslated":true,"sourceHash":"6e7b4d18bf184ebe","translatedAt":"2026-09-08T15:24:50.363Z"},"vi":{"title":"Ai là người an toàn? Tự chắt lọc nhận thức ranh giới cho phép loại bỏ chính xác trong các chủ đề","summary":"Bài báo mới của Hajin Fas đề xuất một phương pháp tự chắt lọc nhận thức ranh giới để giải quyết vấn đề \"từ chối cấp độ chủ đề\" quá thô trong việc căn chỉnh an toàn mô hình. Phương pháp này giảm phạm vi bao phủ dữ liệu huấn luyện từ mất 19,88% lỗi nhắc đến còn 0,20% lỗi còn lại bằng cách nâng cấp chiến lược thử lại.","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Ai là người an toàn? Tự chắt lọc nhận thức ranh giới cho phép loại bỏ chính xác trong các chủ đề - Tin tức AI Aioga","description":"Bài báo mới của Hajin Fas đề xuất một phương pháp tự chắt lọc nhận thức ranh giới để giải quyết vấn đề \"từ chối cấp độ chủ đề\" quá thô trong việc căn chỉnh an toàn mô hình. Phương...","url":"https://www.aioga.com/vi/news/cmtssojsn03txrokagwlns0fs/","contentTranslated":true,"sourceHash":"6e7b4d18bf184ebe","translatedAt":"2026-09-08T15:24:50.295Z"},"id":{"title":"Untuk siapa keselamatan? Distilasi diri yang sadar batas memungkinkan penolakan yang tepat dalam topik","summary":"Makalah baru Hajin Fas mengusulkan metode distilasi diri yang sadar batas untuk mengatasi masalah \"penolakan tingkat topik\" yang terlalu kasar dalam penyelarasan keselamatan model. Metode ini mengurangi cakupan data pelatihan dari kehilangan 19,88% prompt menjadi 0,20% kegagalan residual dengan meningkatkan strategi percobaan ulang.","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Untuk siapa keselamatan? Distilasi diri yang sadar batas memungkinkan penolakan yang tepat dalam topik - Berita AI Aioga","description":"Makalah baru Hajin Fas mengusulkan metode distilasi diri yang sadar batas untuk mengatasi masalah \"penolakan tingkat topik\" yang terlalu kasar dalam penyelarasan keselamatan model....","url":"https://www.aioga.com/id/news/cmtssojsn03txrokagwlns0fs/","contentTranslated":true,"sourceHash":"6e7b4d18bf184ebe","translatedAt":"2026-09-08T15:25:00.004Z"},"th":{"title":"ใครเป็นคนปลอดภัย? การกลั่นตัวเองแบบรับรู้ขอบเขตช่วยให้สามารถปฏิเสธได้อย่างแม่นยําภายในหัวข้อ","summary":"บทความใหม่ของ Hajin Fas เสนอวิธีการกลั่นกรองตนเองแบบตระหนักถึงขอบเขตเพื่อแก้ไขปัญหาการปฏิเสธในระดับหัวข้อที่หยาบเกินไปในการปรับแนวความปลอดภัยของโมเดล วิธีนี้ช่วยลดการครอบคลุมข้อมูลฝึกอบรมจากการสูญเสียพรอมต์ 19.88% เหลือ 0.20% ความล้มเหลวที่เหลืออยู่โดยการอัปเกรดกลยุทธ์การลองใหม่","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"ใครเป็นคนปลอดภัย? การกลั่นตัวเองแบบรับรู้ขอบเขตช่วยให้สามารถปฏิเสธได้อย่างแม่นยําภายในหัวข้อ - ข่าว AI Aioga","description":"บทความใหม่ของ Hajin Fas เสนอวิธีการกลั่นกรองตนเองแบบตระหนักถึงขอบเขตเพื่อแก้ไขปัญหาการปฏิเสธในระดับหัวข้อที่หยาบเกินไปในการปรับแนวความปลอดภัยของโมเดล วิธีนี้ช่วยลดการครอบคลุมข้อมูล...","url":"https://www.aioga.com/th/news/cmtssojsn03txrokagwlns0fs/","contentTranslated":true,"sourceHash":"6e7b4d18bf184ebe","translatedAt":"2026-09-08T15:25:00.165Z"},"pl":{"title":"Dla kogo jest ochrona? Świadoma granic samodestylacja umożliwia precyzyjne odrzucenie w ramach tematów","summary":"Nowy artykuł Hajina Fasa proponuje metodę samodestylacji świadomej granic, aby rozwiązać problem zbyt grubego \"odrzucenia na poziomie tematowym\" w zakresie zgodności z bezpieczeństwem modeli. Ta metoda zmniejsza pokrycie danymi treningowymi z utraty 19,88% promptów do 0,20% pozostałych porażek poprzez aktualizację strategii powtórek.","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Dla kogo jest ochrona? Świadoma granic samodestylacja umożliwia precyzyjne odrzucenie w ramach tematów - Aioga Wiadomości AI","description":"Nowy artykuł Hajina Fasa proponuje metodę samodestylacji świadomej granic, aby rozwiązać problem zbyt grubego \"odrzucenia na poziomie tematowym\" w zakresie zgodności z bezpieczeńst...","url":"https://www.aioga.com/pl/news/cmtssojsn03txrokagwlns0fs/","contentTranslated":true,"sourceHash":"6e7b4d18bf184ebe","translatedAt":"2026-09-08T15:25:09.542Z"}},"evidenceTier":"verified-news","reviewStatus":"automated-ingest","indexable":true,"editorialCover":""}}