{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-08-11T09:21:12.743Z","headline":"Mistral 开源 3B 安全模型 Shieldstral，以更小体积匹敌七倍规模模型","description":"Mistral 发布开源 3B 参数安全模型 Shieldstral，通过自然语言\"是/否\"问题而非固定类别来检查 AI 输入输出的安全违规。该模型在部分基准测试中匹敌七倍于其规模的模型，并支持运行者在运行时自行设定标准，无需依赖第三方分类体系，且可在本地运行。","url":"https://www.aioga.com/news/cmsgbb37t004tro5qisawd1ya/","mainEntityOfPage":"https://www.aioga.com/news/cmsgbb37t004tro5qisawd1ya/","datePublished":"2026-08-05T16:35:07.000Z","dateModified":"2026-08-05T16:35:07.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://the-decoder.com/mistrals-open-model-shieldstral-matches-much-larger-safety-models","https://aihot.virxact.com/items/cmsgbb37t004tro5qisawd1ya"],"canonicalUrl":"https://www.aioga.com/news/cmsgbb37t004tro5qisawd1ya/","directAnswer":{"@type":"Answer","text":"Mistral 发布开源 3B 参数安全模型 Shieldstral，采用运行时可定义的自然语言“是/否”问题检查输入输出，并以回答概率计算安全分数。论文称其在综合文本基准上达到 84.9% F1，与约七倍规模的 GPT-OSS-Safeguard-20B 持平。","url":"https://www.aioga.com/news/cmsgbb37t004tro5qisawd1ya/","dateCreated":"2026-08-05T16:35:07.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"the-decoder.com source article","url":"https://the-decoder.com/mistrals-open-model-shieldstral-matches-much-larger-safety-models","datePublished":"2026-08-05T16:35:07.000Z","provider":{"@type":"Organization","name":"the-decoder.com","url":"https://the-decoder.com/mistrals-open-model-shieldstral-matches-much-larger-safety-models"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmsgbb37t004tro5qisawd1ya","datePublished":"2026-08-05T16:35:07.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmsgbb37t004tro5qisawd1ya"}}],"aggregationSource":"The Decoder：AI News（RSS）","originalPublisher":{"name":"the-decoder.com","url":"https://the-decoder.com/mistrals-open-model-shieldstral-matches-much-larger-safety-models"},"geoDeepAnswer":null,"article":{"id":"cmsgbb37t004tro5qisawd1ya","slug":"cmsgbb37t004tro5qisawd1ya","url":"https://www.aioga.com/news/cmsgbb37t004tro5qisawd1ya/","title":"Mistral 开源 3B 安全模型 Shieldstral，以更小体积匹敌七倍规模模型","title_en":"Mistral's open model Shieldstral matches much larger safety models at a fraction of the size","summary":"Mistral 发布开源 3B 参数安全模型 Shieldstral，通过自然语言\"是/否\"问题而非固定类别来检查 AI 输入输出的安全违规。该模型在部分基准测试中匹敌七倍于其规模的模型，并支持运行者在运行时自行设定标准，无需依赖第三方分类体系，且可在本地运行。","source":"The Decoder：AI News（RSS）","sourceUrl":"https://the-decoder.com/mistrals-open-model-shieldstral-matches-much-larger-safety-models","aiHotUrl":"https://aihot.virxact.com/items/cmsgbb37t004tro5qisawd1ya","publishedAt":"2026-08-05T16:35:07.000Z","category":"模型更新","score":60,"selected":false,"articleBody":["A new paper proposes replacing fixed safety categories with yes or no questions that operators can define at runtime without retraining the classifier.","Shieldstral, a 3-billion-parameter model from French AI company Mistral, matches models three times its size on standard text safety benchmarks, according to the paper：https://arxiv.org/abs/2607.25857. Mistral says the model also sets a new high score for joint text and image classification.","Many guardrail models sort content using fixed taxonomies. The paper's authors, including Mistral co-founder Guillaume Lample, point to two problems with this approach: public safety datasets group risks too differently to support one common taxonomy, and the same rules don't fit every use case. Content suitable for a cybersecurity tool could be harmful on a mental health platform. Ad","Operators tell Shieldstral what to check with plain-language questions such as \"Does this content promote violence?\" The model answers only \"yes\" or \"no,\" and the system uses the probability of each response to calculate a safety score between zero and one. Ad DEC_D_Incontent-1","The researchers combined about 54.1 million examples covering safety, harmful content, and manipulation attempts into one format. They applied strict standards to targeted manipulation, moderate standards to general safety data, and lenient standards to response quality.","To teach Shieldstral finer distinctions, the team used another language model to rewrite safe text into unsafe variants. Each example also included a similar but different category that had to be rejected, which trained the model to separate closely related rules rather than make only a broad safe or unsafe judgment. Ad","The authors created the adaptability test categories separately from the training set, using different names and levels of detail. None of the fine-grained test categories directly matches a training category, though 10 of the 12 broader classes have rough counterparts, they say.","Across the combined text benchmarks, Shieldstral posts an F1 score of 84.9 percent. F1 combines precision and recall into one metric, with 100 percent representing a perfect score. That result ties OpenAI's GPT-OSS-Safeguard-20B, which is about seven times larger, and beats Qwen3Guard-8B at 84.0 percent, Nemotron-3.5-Safety-4B at 83.3 percent, and LlamaGuard-4-12B at 69.1 percent. Ad DEC_D_Incontent-2","On images and image-text combinations, Shieldstral scores 83.8 percent, ahead of OmniGuard-7B at 77.6 percent and LlavaGuard-7B at 71.6 percent. Ad","GPT-OSS-Safeguard-20B leads the adaptability benchmark with 94.1 percent, compared with Shieldstral's 91.3 percent. This test uses rules that differ from the training categories or are entirely new. The authors still consider Shieldstral more practical than GPT-OSS-Safeguard-20B：https://the-decoder.com/openai-releases-gpt-oss-safeguard-open-source-models-for-flexible-ai-safety/ and Nemotron-3.5-Safety, since both generate long intermediate reasoning sequences that raise compute costs, while Shieldstral returns a single word.","Shieldstral is based on Mistral's Ministral-3B：https://the-decoder.com/paris-based-mistral-releases-large-3-a-major-new-open-source-ai-model/ with the Pixtral vision encoder：https://the-decoder.com/french-ai-company-mistral-unveils-pixtral-12b-its-first-multimodal-model/. In a validation test with fine-grained categories, synthetic category data raised the F1 score by 23.3 percentage points, which the researchers say was the main driver of the model's ability to adapt to new rules.","Shieldstral is available as an open-weight model under the Apache 2.0 license：https://huggingface.co/mistralai/Shieldstral-1.0-3B.","Safety classifiers sit on either side of the main language model, screening prompts before processing and responses before they reach users. Operators can update these rules without retraining the main model. But because every request passes through the classifier, its size, speed, and cost add up quickly.","Anthropic's Claude Fable 5：https://the-decoder.com/claude-fable-5-the-first-mythos-model-is-powerful-expensive-and-heavily-filtered/ showed how poorly tuned filters can affect real use. Artificial Analysis found that the system automatically routed eight to nine percent of tasks to a weaker model. One medical physicist called Fable 5 unusable because his work often includes the word \"nuclear.\" Other users reported that the system flagged MRI analysis as bioterrorism.","Anthropic tightened the filter after locating a safety issue and says it has since blocked harmless coding tasks more often：https://the-decoder.com/anthropics-fable-5-is-back-worldwide-after-a-two-week-government-ban-over-a-jailbreak/. Shieldstral gives operators more control over that tradeoff. They can write screening criteria at runtime and tailor the filter to a specific app instead of adopting someone else's categories.","These classifiers already play a growing role across the industry. OpenAI uses them for automatic age detection：https://the-decoder.com/openai-rolls-out-age-prediction-to-apply-teen-safeguards-in-chatgpt/ in ChatGPT and routes emotional requests through a safety filter：https://the-decoder.com/openai-adds-parental-controls-to-chatgpt-for-teens/ to stricter models. Claude Code：https://the-decoder.com/claude-codes-new-auto-mode-tries-to-balance-safety-and-speed/ uses a classifier to block external scripts, production deployments, and force pushes. Anthropic's Fable 5 review also requires the company to store inputs and outputs：https://the-decoder.com/claude-fable-5-anthropic-admits-wrong-tradeoff-after-invisibly-throttling-rival-ai-researchers/ for up to 30 days, or up to two years after rule violations.","Stay in the loop on AI. Clear, useful, no fluff.","Follow The Decoder for AI news, background stories and expert analyses.","The Decoder：https://the-decoder.com/"],"articleImages":[{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/07/mistral_ai-3.png","alt":"Image description","afterParagraph":0,"url":"/media/articles/cmsgbb37t004tro5qisawd1ya/3638cfa289b721ef.png"},{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/08/shieldstral-01-architecture.jpg","alt":"Diagram of the Shieldstral architecture showing a fixed system prompt and customizable Instruct, Query, and Document fields feeding into the model, which outputs logits for the yes and no tokens.","afterParagraph":2,"url":"/media/articles/cmsgbb37t004tro5qisawd1ya/41037309096275eb.jpg"},{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/08/shieldstral-06-training-samples.jpg","alt":"Two training data examples showing a text-only sample with a prompt and response about physical violence on the left, and a multimodal sample with an image and question about explicit content on the right. Both are labeled no.","afterParagraph":4,"url":"/media/articles/cmsgbb37t004tro5qisawd1ya/e390707e2c89f1d6.jpg"},{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/08/shieldstral-03-benchmark-overall.jpg","alt":"Bar chart comparing average F1 scores for ten guardrail models. Shieldstral-3B and GPT-OSS-Safeguard-20B lead at 84.9 percent each, while ShieldGemma-9B trails at 54.7 percent.","afterParagraph":8,"url":"/media/articles/cmsgbb37t004tro5qisawd1ya/28fa428b98821a05.jpg"},{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/08/shieldstral-04-adaptability.jpg","alt":"Three groups of bars comparing precision, recall, and F1 on the adaptability benchmark. GPT-OSS-Safeguard-20B leads in F1 at 94.1 percent, followed by Nemotron-3.5-Safety-4B at 91.8 percent and Shieldstral-3B at 91.3 percent.","afterParagraph":9,"url":"/media/articles/cmsgbb37t004tro5qisawd1ya/3679d30e726d0c94.jpg"}],"mediaStatus":"ok","articleBodyZh":["一篇新论文提出，用是或否的问题取代固定的安全类别，操作人员可以在运行时自行定义，而无需重新训练分类器。","据论文介绍，法国人工智能公司Mistral推出的拥有30亿参数的模型Shieldstral，在标准文本安全基准测试中表现可与三倍其规模的模型相媲美：https://arxiv.org/abs/2607.25857。Mistral表示，该模型在文本和图像联合分类上也创造了新的高分记录。","许多护栏模型使用固定的分类法对内容进行排序。包括Mistral联合创始人Guillaume Lample在内的论文作者指出了这种方法的两个问题：公共安全数据集对风险的分组相差过大，无法支持统一的分类法；同样的规则并不适用于所有用例。适用于网络安全工具的内容，在心理健康平台上可能是有害的。","操作人员通过类似“此内容是否宣传暴力？”的自然语言问题告诉Shieldstral需要检查什么。模型仅回答“是”或“否”，系统再根据每个回答的概率计算出一个介于零到一之间的安全评分。","研究人员将约5410万个涵盖安全、有害内容和操纵尝试的示例整合成统一格式。他们对针对性操纵应用严格标准，对一般安全数据应用中等标准，对响应质量应用宽松标准。","为了让Shieldstral学习更细微的区分，团队使用另一种语言模型将安全文本改写为不安全变体。每个示例还包含一个类似但不同的类别，必须被拒绝，这训练模型区分紧密相关的规则，而不仅仅做出广泛的安全或不安全判断。","作者将适应性测试类别与训练集分开创建，使用不同的名称和详细级别。他们表示，没有一个细分测试类别与训练类别完全匹配，尽管12个较广类别中的10个大致对应一个训练类别。","在综合文本基准测试中，Shieldstral 的 F1 分数为 84.9%。F1 将精确率和召回率合并为一个指标，100%表示满分。该结果与 OpenAI 的 GPT-OSS-Safeguard-20B 并列，后者的规模约大七倍，并且超过了 Qwen3Guard-8B 的 84.0%、Nemotron-3.5-Safety-4B 的 83.3% 以及 LlamaGuard-4-12B 的 69.1%。Ad DEC_D_Incontent-2","在图像及图文组合测试中，Shieldstral 得分为 83.8%，领先于 OmniGuard-7B 的 77.6% 和 LlavaGuard-7B 的 71.6%。Ad","GPT-OSS-Safeguard-20B 在适应性基准测试中以 94.1% 领先，而 Shieldstral 为 91.3%。此测试使用与训练类别不同或完全新的规则。作者仍认为 Shieldstral 比 GPT-OSS-Safeguard-20B(https://the-decoder.com/openai-releases-gpt-oss-safeguard-open-source-models-for-flexible-ai-safety/) 和 Nemotron-3.5-Safety 更实用，因为后两者会生成较长的中间推理序列，从而增加计算成本，而 Shieldstral 只返回一个词。","Shieldstral 基于 Mistral 的 Ministral-3B(https://the-decoder.com/paris-based-mistral-releases-large-3-a-major-new-open-source-ai-model/) 和 Pixtral 视觉编码器(https://the-decoder.com/french-ai-company-mistral-unveils-pixtral-12b-its-first-multimodal-model/)。在一个具有细粒度类别的验证测试中，合成类别数据将 F1 分数提高了 23.3 个百分点，研究人员表示这是模型适应新规则能力的主要驱动因素。","Shieldstral 可作为开放权重模型在 Apache 2.0 许可下获取：https://huggingface.co/mistralai/Shieldstral-1.0-3B。","安全分类器位于主语言模型的两侧，分别在处理前筛查提示词，在到达用户前筛查响应。操作人员可以在不重新训练主模型的情况下更新这些规则。但由于每个请求都经过分类器，其大小、速度和成本会迅速累积。","Anthropic 的 Claude Fable 5：https://the-decoder.com/claude-fable-5-the-first-mythos-model-is-powerful-expensive-and-heavily-filtered/ 展示了调校不佳的过滤器如何影响实际使用。Artificial Analysis 发现该系统会自动将 8% 到 9% 的任务路由到性能较弱的模型。一位医学物理学家称 Fable 5 无法使用，因为他的工作经常包含“核”字。其他用户报告系统将 MRI 分析标记为生物恐怖活动。","Anthropic 在发现安全问题后收紧了过滤器，并表示从那时起已经更频繁地阻止无害的编码任务：https://the-decoder.com/anthropics-fable-5-is-back-worldwide-after-a-two-week-government-ban-over-a-jailbreak/。Shieldstral 让操作员在这种权衡中拥有更多控制权。他们可以在运行时编写筛选标准，并针对特定应用调整过滤器，而不是采用别人的分类。","这些分类器在整个行业中已经发挥着越来越重要的作用。OpenAI 在 ChatGPT 中使用它们进行自动年龄检测：https://the-decoder.com/openai-rolls-out-age-prediction-to-apply-teen-safeguards-in-chatgpt/，并将情绪请求通过安全过滤器：https://the-decoder.com/openai-adds-parental-controls-to-chatgpt-for-teens/ 路由到更严格的模型。Claude Code：https://the-decoder.com/claude-codes-new-auto-mode-tries-to-balance-safety-and-speed/ 使用分类器来阻止外部脚本、生产部署和强制推送。Anthropic 对 Fable 5 的评审还要求公司存储输入和输出：https://the-decoder.com/claude-fable-5-anthropic-admits-wrong-tradeoff-after-invisibly-throttling-rival-ai-researchers/ 最长 30 天，或在违反规则后最长两年。","保持对 AI 的了解。清晰、实用、无废话。","关注 The Decoder 获取 AI 新闻、背景故事和专家分析。","解码器：https://the-decoder.com/"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Mistral 发布开源 3B 参数安全模型 Shieldstral，采用运行时可定义的自然语言“是/否”问题检查输入输出，并以回答概率计算安全分数。论文称其在综合文本基准上达到 84.9% F1，与约七倍规模的 GPT-OSS-Safeguard-20B 持平。","background":"现有护栏模型通常依赖固定安全分类体系，但论文作者认为不同数据集和应用场景对风险划分并不一致。Shieldstral 将检查标准改写为运行者可定义的问题，并通过约 5410 万个安全、有害内容和操纵尝试样本训练模型。","viewpoint":"Aioga 判断，Shieldstral 的核心价值在于把安全审核标准从固定分类表转向可调整的问题配置；这可能降低特定场景适配对重新训练的依赖。但当前材料只说明其在部分基准上的表现，不能据此推断真实部署中的整体安全效果。","implications":"若论文结果能够在更多场景中复现，较小模型结合运行时规则的方式，可能为本地部署和定制化审核提供新的技术路径。值得关注的是，模型对自然语言标准的理解是否稳定，以及不同标准设置会否带来审核结果差异。","nextStep":"后续应核对论文完整实验设置、测试集与训练集是否存在交叉影响，并观察 Shieldstral 在文本与图像联合分类、不同语言及具体业务规则下的表现。部署方还需独立测试误报、漏报和规则变更后的稳定性。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-08-05T17:43:47.068Z","sourceHash":"054ce4c0cb3e0572","review":{"approved":true,"groundedness":98,"clarity":94,"duplicationRisk":12,"blockingIssues":[],"notes":["“Aioga 判断”后的内容属于明确标注的分析观点，且以“可能”“不能据此推断”等限定表述，未冒充来源事实。","nextStep 中关于进一步核对实验设置、独立测试误报漏报等属于合理的后续核查建议，并非对现有论文结果作无来源断言。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","low-source-overlap","no-html","independent-ai-review"]}},"tags":["模型更新","The Decoder：AI News（RSS）"],"translations":{"zh-CN":{"title":"Mistral 开源 3B 安全模型 Shieldstral，以更小体积匹敌七倍规模模型","summary":"Mistral 发布开源 3B 参数安全模型 Shieldstral，通过自然语言\"是/否\"问题而非固定类别来检查 AI 输入输出的安全违规。该模型在部分基准测试中匹敌七倍于其规模的模型，并支持运行者在运行时自行设定标准，无需依赖第三方分类体系，且可在本地运行。","category":"模型更新","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Mistral 开源 3B 安全模型 Shieldstral，以更小体积匹敌七倍规模模型 - Aioga AI资讯","description":"Mistral 发布开源 3B 参数安全模型 Shieldstral，通过自然语言\"是/否\"问题而非固定类别来检查 AI 输入输出的安全违规。该模型在部分基准测试中匹敌七倍于其规模的模型，并支持运行者在运行时自行设定标准，无需依赖第三方分类体系，且可在本地运行。","url":"https://www.aioga.com/news/cmsgbb37t004tro5qisawd1ya/","articleBody":["一篇新论文提出，用是或否的问题取代固定的安全类别，操作人员可以在运行时自行定义，而无需重新训练分类器。","据论文介绍，法国人工智能公司Mistral推出的拥有30亿参数的模型Shieldstral，在标准文本安全基准测试中表现可与三倍其规模的模型相媲美：https://arxiv.org/abs/2607.25857。Mistral表示，该模型在文本和图像联合分类上也创造了新的高分记录。","许多护栏模型使用固定的分类法对内容进行排序。包括Mistral联合创始人Guillaume Lample在内的论文作者指出了这种方法的两个问题：公共安全数据集对风险的分组相差过大，无法支持统一的分类法；同样的规则并不适用于所有用例。适用于网络安全工具的内容，在心理健康平台上可能是有害的。","操作人员通过类似“此内容是否宣传暴力？”的自然语言问题告诉Shieldstral需要检查什么。模型仅回答“是”或“否”，系统再根据每个回答的概率计算出一个介于零到一之间的安全评分。","研究人员将约5410万个涵盖安全、有害内容和操纵尝试的示例整合成统一格式。他们对针对性操纵应用严格标准，对一般安全数据应用中等标准，对响应质量应用宽松标准。","为了让Shieldstral学习更细微的区分，团队使用另一种语言模型将安全文本改写为不安全变体。每个示例还包含一个类似但不同的类别，必须被拒绝，这训练模型区分紧密相关的规则，而不仅仅做出广泛的安全或不安全判断。","作者将适应性测试类别与训练集分开创建，使用不同的名称和详细级别。他们表示，没有一个细分测试类别与训练类别完全匹配，尽管12个较广类别中的10个大致对应一个训练类别。","在综合文本基准测试中，Shieldstral 的 F1 分数为 84.9%。F1 将精确率和召回率合并为一个指标，100%表示满分。该结果与 OpenAI 的 GPT-OSS-Safeguard-20B 并列，后者的规模约大七倍，并且超过了 Qwen3Guard-8B 的 84.0%、Nemotron-3.5-Safety-4B 的 83.3% 以及 LlamaGuard-4-12B 的 69.1%。Ad DEC_D_Incontent-2","在图像及图文组合测试中，Shieldstral 得分为 83.8%，领先于 OmniGuard-7B 的 77.6% 和 LlavaGuard-7B 的 71.6%。Ad","GPT-OSS-Safeguard-20B 在适应性基准测试中以 94.1% 领先，而 Shieldstral 为 91.3%。此测试使用与训练类别不同或完全新的规则。作者仍认为 Shieldstral 比 GPT-OSS-Safeguard-20B(https://the-decoder.com/openai-releases-gpt-oss-safeguard-open-source-models-for-flexible-ai-safety/) 和 Nemotron-3.5-Safety 更实用，因为后两者会生成较长的中间推理序列，从而增加计算成本，而 Shieldstral 只返回一个词。","Shieldstral 基于 Mistral 的 Ministral-3B(https://the-decoder.com/paris-based-mistral-releases-large-3-a-major-new-open-source-ai-model/) 和 Pixtral 视觉编码器(https://the-decoder.com/french-ai-company-mistral-unveils-pixtral-12b-its-first-multimodal-model/)。在一个具有细粒度类别的验证测试中，合成类别数据将 F1 分数提高了 23.3 个百分点，研究人员表示这是模型适应新规则能力的主要驱动因素。","Shieldstral 可作为开放权重模型在 Apache 2.0 许可下获取：https://huggingface.co/mistralai/Shieldstral-1.0-3B。","安全分类器位于主语言模型的两侧，分别在处理前筛查提示词，在到达用户前筛查响应。操作人员可以在不重新训练主模型的情况下更新这些规则。但由于每个请求都经过分类器，其大小、速度和成本会迅速累积。","Anthropic 的 Claude Fable 5：https://the-decoder.com/claude-fable-5-the-first-mythos-model-is-powerful-expensive-and-heavily-filtered/ 展示了调校不佳的过滤器如何影响实际使用。Artificial Analysis 发现该系统会自动将 8% 到 9% 的任务路由到性能较弱的模型。一位医学物理学家称 Fable 5 无法使用，因为他的工作经常包含“核”字。其他用户报告系统将 MRI 分析标记为生物恐怖活动。","Anthropic 在发现安全问题后收紧了过滤器，并表示从那时起已经更频繁地阻止无害的编码任务：https://the-decoder.com/anthropics-fable-5-is-back-worldwide-after-a-two-week-government-ban-over-a-jailbreak/。Shieldstral 让操作员在这种权衡中拥有更多控制权。他们可以在运行时编写筛选标准，并针对特定应用调整过滤器，而不是采用别人的分类。","这些分类器在整个行业中已经发挥着越来越重要的作用。OpenAI 在 ChatGPT 中使用它们进行自动年龄检测：https://the-decoder.com/openai-rolls-out-age-prediction-to-apply-teen-safeguards-in-chatgpt/，并将情绪请求通过安全过滤器：https://the-decoder.com/openai-adds-parental-controls-to-chatgpt-for-teens/ 路由到更严格的模型。Claude Code：https://the-decoder.com/claude-codes-new-auto-mode-tries-to-balance-safety-and-speed/ 使用分类器来阻止外部脚本、生产部署和强制推送。Anthropic 对 Fable 5 的评审还要求公司存储输入和输出：https://the-decoder.com/claude-fable-5-anthropic-admits-wrong-tradeoff-after-invisibly-throttling-rival-ai-researchers/ 最长 30 天，或在违反规则后最长两年。","保持对 AI 的了解。清晰、实用、无废话。","关注 The Decoder 获取 AI 新闻、背景故事和专家分析。","解码器：https://the-decoder.com/"]},"en":{"title":"Mistral open-source 3B safety model Shieldstral, competing with models seven times its size with a smaller scale","summary":"Mistral released the open-source 3B-parameter safety model Shieldstral, which checks AI input and output for safety violations through natural language 'yes/no' questions rather than fixed categories. The model matches the performance of models seven times its size on some benchmark tests, supports users in setting standards at runtime without relying on third-party classification systems, and can run locally.","category":"Models","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Mistral open-source 3B safety model Shieldstral, competing with models seven times its size with a smaller scale - Aioga AI News","description":"Mistral released the open-source 3B-parameter safety model Shieldstral, which checks AI input and output for safety violations through natural language 'yes/no' questions rather th...","url":"https://www.aioga.com/en/news/cmsgbb37t004tro5qisawd1ya/","contentTranslated":true,"sourceHash":"bcc3db372e4a5492","translatedAt":"2026-08-05T17:02:35.173Z"},"ja":{"title":"Mistralはオープンソースの3B安全モデルShieldstralをリリースし、より小さいサイズで7倍の規模のモデルに匹敵","summary":"Mistralはオープンソースの3Bパラメータ安全モデルShieldstralを公開しました。このモデルは、固定されたカテゴリではなく、自然言語による「はい/いいえ」の質問でAIの入力と出力の安全違反をチェックします。このモデルは一部のベンチマークテストで自分の規模の7倍のモデルに匹敵し、実行中にユーザーが基準を設定でき、第三者の分類体系に依存せず、ローカルでも実行可能です。","category":"モデル更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Mistralはオープンソースの3B安全モデルShieldstralをリリースし、より小さいサイズで7倍の規模のモデルに匹敵 - Aioga AIニュース","description":"Mistralはオープンソースの3Bパラメータ安全モデルShieldstralを公開しました。このモデルは、固定されたカテゴリではなく、自然言語による「はい/いいえ」の質問でAIの入力と出力の安全違反をチェックします。このモデルは一部のベンチマークテストで自分の規模の7倍のモデルに匹敵し、実行中にユーザーが基準を設定でき、第三者の分類体系に依存せず、ローカル...","url":"https://www.aioga.com/ja/news/cmsgbb37t004tro5qisawd1ya/","contentTranslated":true,"sourceHash":"bcc3db372e4a5492","translatedAt":"2026-08-05T17:02:36.120Z"},"ko":{"title":"Mistral 오픈소스 3B 안전 모델 Shieldstral, 더 작은 규모로 7배 크기 모델에 필적","summary":"Mistral은 오픈소스 3B 파라미터 안전 모델 Shieldstral을 출시했으며, 고정된 카테고리 대신 자연어 '예/아니오' 질문을 통해 AI 입력출력의 안전 위반을 점검합니다. 이 모델은 일부 벤치마크 테스트에서 자신의 규모보다 일곱 배 큰 모델과 맞먹는 성능을 보이며, 실행자가 실행 시 스스로 기준을 설정할 수 있도록 지원하고, 제3자 분류 체계에 의존할 필요 없이 로컬에서 실행할 수 있습니다.","category":"모델 업데이트","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Mistral 오픈소스 3B 안전 모델 Shieldstral, 더 작은 규모로 7배 크기 모델에 필적 - Aioga AI 뉴스","description":"Mistral은 오픈소스 3B 파라미터 안전 모델 Shieldstral을 출시했으며, 고정된 카테고리 대신 자연어 '예/아니오' 질문을 통해 AI 입력출력의 안전 위반을 점검합니다. 이 모델은 일부 벤치마크 테스트에서 자신의 규모보다 일곱 배 큰 모델과 맞먹는 성능을 보이며, 실행자가 실행 시 스스로 기준을 설정할 수 있...","url":"https://www.aioga.com/ko/news/cmsgbb37t004tro5qisawd1ya/","contentTranslated":true,"sourceHash":"bcc3db372e4a5492","translatedAt":"2026-08-05T17:03:23.637Z"},"es":{"title":"Mistral lanza el modelo de seguridad de código abierto 3B Shieldstral, que compite con modelos de siete veces su tamaño con un volumen más pequeño","summary":"Mistral lanzó el modelo de seguridad Shieldsral de 3B parámetros de código abierto, que verifica las violaciones de seguridad en la entrada y salida de la IA mediante preguntas de \"sí/no\" en lenguaje natural en lugar de categorías fijas. Este modelo iguala en algunas pruebas de referencia a modelos siete veces más grandes, y permite a los usuarios establecer sus propios estándares en tiempo de ejecución, sin depender de sistemas de clasificación de terceros, y puede ejecutarse localmente.","category":"Modelos","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Mistral lanza el modelo de seguridad de código abierto 3B Shieldstral, que compite con modelos de siete veces su tamaño con un volumen más pequeño - Aioga Noticias de IA","description":"Mistral lanzó el modelo de seguridad Shieldsral de 3B parámetros de código abierto, que verifica las violaciones de seguridad en la entrada y salida de la IA mediante preguntas de...","url":"https://www.aioga.com/es/news/cmsgbb37t004tro5qisawd1ya/","contentTranslated":true,"sourceHash":"bcc3db372e4a5492","translatedAt":"2026-08-05T17:03:19.749Z"},"fr":{"title":"Mistral ouvre le modèle de sécurité 3B Shieldstral, capable de rivaliser avec un modèle sept fois plus grand avec un volume plus petit","summary":"Mistral a publié le modèle de sécurité open source à 3 milliards de paramètres Shieldstral, qui vérifie les violations de sécurité des entrées et sorties de l'IA à travers des questions de type « oui/non » plutôt que des catégories fixes. Ce modèle rivalise, sur certains tests de référence, avec des modèles sept fois plus grands et permet aux utilisateurs de définir leurs propres critères lors de l'exécution, sans dépendre d'un système de classification tiers, et peut être exécuté localement.","category":"Modèles","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Mistral ouvre le modèle de sécurité 3B Shieldstral, capable de rivaliser avec un modèle sept fois plus grand avec un volume plus petit - Aioga Actualités IA","description":"Mistral a publié le modèle de sécurité open source à 3 milliards de paramètres Shieldstral, qui vérifie les violations de sécurité des entrées et sorties de l'IA à travers des ques...","url":"https://www.aioga.com/fr/news/cmsgbb37t004tro5qisawd1ya/","contentTranslated":true,"sourceHash":"bcc3db372e4a5492","translatedAt":"2026-08-05T17:04:06.232Z"},"de":{"title":"Mistral veröffentlicht das 3B-Sicherheitsmodell Shieldstral, das mit einem siebenmal größeren Modell bei kleinerem Umfang konkurriert","summary":"Mistral hat das Open-Source-Sicherheitsmodell Shieldstral mit 3 Milliarden Parametern veröffentlicht, das Sicherheitsverletzungen bei AI-Ein- und Ausgaben durch Ja/Nein-Fragen anstelle fester Kategorien überprüft. Das Modell kann in einigen Benchmark-Tests mit Modellen konkurrieren, die das Siebenfache seiner Größe haben, und ermöglicht es den Anwendern, Standards zur Laufzeit selbst festzulegen, ohne auf Klassifikationssysteme Dritter angewiesen zu sein, und kann lokal betrieben werden.","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Mistral veröffentlicht das 3B-Sicherheitsmodell Shieldstral, das mit einem siebenmal größeren Modell bei kleinerem Umfang konkurriert - Aioga KI-News","description":"Mistral hat das Open-Source-Sicherheitsmodell Shieldstral mit 3 Milliarden Parametern veröffentlicht, das Sicherheitsverletzungen bei AI-Ein- und Ausgaben durch Ja/Nein-Fragen anst...","url":"https://www.aioga.com/de/news/cmsgbb37t004tro5qisawd1ya/","contentTranslated":true,"sourceHash":"bcc3db372e4a5492","translatedAt":"2026-08-05T17:04:10.023Z"},"pt-BR":{"title":"Mistral lança modelo de segurança de código aberto 3B Shieldstral, que compete com modelos sete vezes maiores com tamanho menor","summary":"A Mistral lançou o modelo de segurança de 3 bilhões de parâmetros de código aberto Shieldstral, que verifica violações de segurança nas entradas e saídas de IA por meio de perguntas de \"sim/não\" em linguagem natural, em vez de categorias fixas. O modelo, em alguns testes de benchmark, rivaliza com modelos sete vezes maiores que ele, e permite que os usuários definam seus próprios padrões em tempo de execução, sem depender de sistemas de classificação de terceiros, podendo ser executado localmente.","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Mistral lança modelo de segurança de código aberto 3B Shieldstral, que compete com modelos sete vezes maiores com tamanho menor - Aioga Notícias de IA","description":"A Mistral lançou o modelo de segurança de 3 bilhões de parâmetros de código aberto Shieldstral, que verifica violações de segurança nas entradas e saídas de IA por meio de pergunta...","url":"https://www.aioga.com/pt-BR/news/cmsgbb37t004tro5qisawd1ya/","contentTranslated":true,"sourceHash":"bcc3db372e4a5492","translatedAt":"2026-08-05T17:04:43.448Z"},"ru":{"title":"Mistral открывает исходный код безопасной модели 3B Shieldstral, которая по производительности сопоставима с семикратной моделью меньшего размера","summary":"Mistral выпустила открытый безопасный модельный параметр Shieldstral с 3B параметрами, который проверяет нарушения безопасности входных и выходных данных ИИ через вопросы «да/нет», а не фиксированные категории. Эта модель в некоторых бенчмарках сопоставима с моделями в семь раз более крупными по размеру и поддерживает возможность для пользователей задавать стандарты во время работы, не полагаясь на сторонние классификационные системы, а также может работать локально.","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Mistral открывает исходный код безопасной модели 3B Shieldstral, которая по производительности сопоставима с семикратной моделью меньшего размера - Aioga Новости ИИ","description":"Mistral выпустила открытый безопасный модельный параметр Shieldstral с 3B параметрами, который проверяет нарушения безопасности входных и выходных данных ИИ через вопросы «да/нет»,...","url":"https://www.aioga.com/ru/news/cmsgbb37t004tro5qisawd1ya/","contentTranslated":true,"sourceHash":"bcc3db372e4a5492","translatedAt":"2026-08-05T17:04:54.974Z"},"ar":{"title":"أطلقت Mistral نموذج الأمان مفتوح المصدر 3B Shieldstral، بحجم أصغر لمنافسة نموذج بمقياس سبعة أضعاف","summary":"أصدرت Mistral نموذج أمني مفتوح المصدر يحتوي على 3 مليارات بارامتر يسمى Shieldstral، للتحقق من انتهاكات السلامة لإدخال وإخراج الذكاء الاصطناعي من خلال أسئلة نعم/لا بدلاً من الفئات الثابتة. يتفوق هذا النموذج في بعض الاختبارات المعيارية على نماذج تعادل سبعة أضعاف حجمه، ويدعم السماح للمشغلين بتحديد المعايير بأنفسهم أثناء التشغيل، دون الحاجة إلى الاعتماد على أنظمة تصنيف الطرف الثالث، ويمكن تشغيله محليًا.","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"أطلقت Mistral نموذج الأمان مفتوح المصدر 3B Shieldstral، بحجم أصغر لمنافسة نموذج بمقياس سبعة أضعاف - Aioga أخبار الذكاء الاصطناعي","description":"أصدرت Mistral نموذج أمني مفتوح المصدر يحتوي على 3 مليارات بارامتر يسمى Shieldstral، للتحقق من انتهاكات السلامة لإدخال وإخراج الذكاء الاصطناعي من خلال أسئلة نعم/لا بدلاً من الفئات ا...","url":"https://www.aioga.com/ar/news/cmsgbb37t004tro5qisawd1ya/","contentTranslated":true,"sourceHash":"bcc3db372e4a5492","translatedAt":"2026-08-05T17:05:44.017Z"},"hi":{"title":"Mistral ने ओपन-सोर्स 3B सुरक्षा मॉडल Shieldstral जारी किया, जो छोटे आकार में सात गुना बड़े मॉडल के बराबर है","summary":"Mistral ने ओपन-सोर्स 3B पैरामीटर सुरक्षा मॉडल Shieldstral जारी किया, जो AI इनपुट और आउटपुट में सुरक्षा उल्लंघनों की जांच के लिए निश्चित श्रेणी के बजाय प्राकृतिक भाषा 'हाँ/नहीं' प्रश्नों का उपयोग करता है। यह मॉडल कुछ बेंचमार्क परीक्षणों में अपनी आकार की तुलना में सात गुना बड़े मॉडल के बराबर प्रदर्शन करता है, और रनटाइम पर संचालन करने वाले उपयोगकर्ताओं को अपने मानक स्वयं निर्धारित करने की अनुमति देता है, किसी तृतीय-पक्ष वर्गीकरण प्रणाली पर निर्भर नहीं करता, और स्थानीय रूप से चलाया जा सकता है।","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Mistral ने ओपन-सोर्स 3B सुरक्षा मॉडल Shieldstral जारी किया, जो छोटे आकार में सात गुना बड़े मॉडल के बराबर है - Aioga AI समाचार","description":"Mistral ने ओपन-सोर्स 3B पैरामीटर सुरक्षा मॉडल Shieldstral जारी किया, जो AI इनपुट और आउटपुट में सुरक्षा उल्लंघनों की जांच के लिए निश्चित श्रेणी के बजाय प्राकृतिक भाषा 'हाँ/नहीं' प्र...","url":"https://www.aioga.com/hi/news/cmsgbb37t004tro5qisawd1ya/","contentTranslated":true,"sourceHash":"bcc3db372e4a5492","translatedAt":"2026-08-05T17:05:39.250Z"},"it":{"title":"Mistral apre il modello di sicurezza 3B Shieldstral, che con un corpo più piccolo compete con modelli di scala sette volte superiore","summary":"Mistral ha rilasciato il modello di sicurezza open source a 3 miliardi di parametri Shieldstral, che verifica le violazioni della sicurezza negli input e output dell'IA tramite domande in linguaggio naturale 'sì/no' invece di categorie fisse. Il modello in alcuni benchmark eguaglia modelli sette volte più grandi e permette agli operatori di impostare autonomamente gli standard durante l'esecuzione, senza fare affidamento su sistemi di classificazione di terze parti, e può essere eseguito localmente.","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Mistral apre il modello di sicurezza 3B Shieldstral, che con un corpo più piccolo compete con modelli di scala sette volte superiore - Aioga Notizie IA","description":"Mistral ha rilasciato il modello di sicurezza open source a 3 miliardi di parametri Shieldstral, che verifica le violazioni della sicurezza negli input e output dell'IA tramite dom...","url":"https://www.aioga.com/it/news/cmsgbb37t004tro5qisawd1ya/","contentTranslated":true,"sourceHash":"bcc3db372e4a5492","translatedAt":"2026-08-05T17:06:32.929Z"},"nl":{"title":"Mistral open source 3B veilig model Shieldstral, met een kleinere omvang om te concurreren met modellen van zeven keer de grootte","summary":"Mistral heeft het open-source 3B-parameter veiligheidsmodel Shieldstral uitgebracht, dat AI-invoer en -uitvoer op veiligheidsinbreuken controleert via natuurlijke taal 'ja/nee'-vragen in plaats van vaste categorieën. Het model evenaart in sommige benchmarks modellen die zeven keer zo groot zijn en ondersteunt dat gebruikers tijdens het draaien zelf normen kunnen instellen, zonder afhankelijk te zijn van externe classificatiesystemen, en kan lokaal worden uitgevoerd.","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Mistral open source 3B veilig model Shieldstral, met een kleinere omvang om te concurreren met modellen van zeven keer de grootte - Aioga AI-nieuws","description":"Mistral heeft het open-source 3B-parameter veiligheidsmodel Shieldstral uitgebracht, dat AI-invoer en -uitvoer op veiligheidsinbreuken controleert via natuurlijke taal 'ja/nee'-vra...","url":"https://www.aioga.com/nl/news/cmsgbb37t004tro5qisawd1ya/","contentTranslated":true,"sourceHash":"bcc3db372e4a5492","translatedAt":"2026-08-05T17:06:23.400Z"},"tr":{"title":"Mistral açık kaynaklı 3B güvenlik modeli Shieldstral, daha küçük boyutla yedi kat büyük modeli karşılayabiliyor","summary":"Mistral, açık kaynaklı 3 milyar parametreli güvenlik modeli Shieldstral’ı yayınladı ve AI giriş-çıkışlarının güvenlik ihlallerini sabit bir sınıf yerine doğal dil ile 'evet/hayır' sorusu yoluyla kontrol ediyor. Bu model, bazı benchmark testlerinde kendi boyutunun yedek katı olan modellerle rekabet ediyor ve çalıştırıcıya standartları çalıştırma sırasında kendi başına belirleme imkanı tanıyor, üçüncü taraf sınıflandırma sistemine ihtiyaç duymadan ve yerel olarak çalıştırılabiliyor.","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Mistral açık kaynaklı 3B güvenlik modeli Shieldstral, daha küçük boyutla yedi kat büyük modeli karşılayabiliyor - Aioga AI Haberleri","description":"Mistral, açık kaynaklı 3 milyar parametreli güvenlik modeli Shieldstral’ı yayınladı ve AI giriş-çıkışlarının güvenlik ihlallerini sabit bir sınıf yerine doğal dil ile 'evet/hayır'...","url":"https://www.aioga.com/tr/news/cmsgbb37t004tro5qisawd1ya/","contentTranslated":true,"sourceHash":"bcc3db372e4a5492","translatedAt":"2026-08-05T17:07:15.540Z"},"vi":{"title":"Mistral mở mã nguồn mô hình an toàn 3B Shieldstral, với kích thước nhỏ hơn nhưng có thể sánh với mô hình gấp bảy lần khối lượng","summary":"Mistral phát hành mô hình bảo mật mã nguồn mở 3B tham số Shieldstral, kiểm tra vi phạm an toàn của đầu vào và đầu ra AI thông qua câu hỏi 'có/không' bằng ngôn ngữ tự nhiên thay vì các loại cố định. Mô hình này trong một số bài kiểm tra chuẩn sánh ngang với các mô hình gấp bảy lần quy mô của nó, đồng thời hỗ trợ người vận hành tự thiết lập tiêu chuẩn khi chạy, không cần phụ thuộc vào hệ thống phân loại của bên thứ ba, và có thể chạy tại địa phương.","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Mistral mở mã nguồn mô hình an toàn 3B Shieldstral, với kích thước nhỏ hơn nhưng có thể sánh với mô hình gấp bảy lần khối lượng - Tin tức AI Aioga","description":"Mistral phát hành mô hình bảo mật mã nguồn mở 3B tham số Shieldstral, kiểm tra vi phạm an toàn của đầu vào và đầu ra AI thông qua câu hỏi 'có/không' bằng ngôn ngữ tự nhiên thay vì...","url":"https://www.aioga.com/vi/news/cmsgbb37t004tro5qisawd1ya/","contentTranslated":true,"sourceHash":"bcc3db372e4a5492","translatedAt":"2026-08-05T17:07:12.248Z"},"id":{"title":"Mistral meluncurkan model keamanan open-source 3B Shieldstral, dengan ukuran lebih kecil namun menyaingi model dengan skala tujuh kali lipat","summary":"Mistral merilis model keamanan berskala 3B parameter open-source Shieldstral, yang memeriksa pelanggaran keamanan input-output AI melalui pertanyaan 'ya/tidak' alih-alih kategori tetap. Model ini dalam beberapa benchmark menandingi model yang ukurannya tujuh kali lipat, dan mendukung pengguna untuk menetapkan standar sendiri saat dijalankan, tanpa perlu mengandalkan sistem klasifikasi pihak ketiga, serta dapat dijalankan secara lokal.","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Mistral meluncurkan model keamanan open-source 3B Shieldstral, dengan ukuran lebih kecil namun menyaingi model dengan skala tujuh kali lipat - Berita AI Aioga","description":"Mistral merilis model keamanan berskala 3B parameter open-source Shieldstral, yang memeriksa pelanggaran keamanan input-output AI melalui pertanyaan 'ya/tidak' alih-alih kategori t...","url":"https://www.aioga.com/id/news/cmsgbb37t004tro5qisawd1ya/","contentTranslated":true,"sourceHash":"bcc3db372e4a5492","translatedAt":"2026-08-05T17:08:00.880Z"},"th":{"title":"Mistral เปิดตัวโมเดลความปลอดภัยแบบโอเพนซอร์ส 3B Shieldstral ซึ่งมีขนาดเล็กกว่าแต่สามารถเทียบเท่ากับโมเดลขนาดเจ็ดเท่า","summary":"Mistral เปิดตัวโมเดลความปลอดภัยแบบเปิด 3B พารามิเตอร์ Shieldstral ซึ่งตรวจสอบการละเมิดความปลอดภัยของอินพุตและเอาต์พุต AI ผ่านคำถาม \"ใช่/ไม่ใช่\" ในภาษาธรรมชาติ แทนการใช้ประเภทคงที่ โมเดลนี้สามารถเทียบเท่ากับโมเดลที่มีขนาดใหญ่กว่าถึงเจ็ดเท่าบางส่วนในการทดสอบมาตรฐาน และรองรับให้ผู้ใช้งานสามารถตั้งค่ามาตรฐานเองในขณะรันโดยไม่ต้องพึ่งพาระบบการจำแนกของบุคคลที่สาม และสามารถรันได้ในเครื่อง","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Mistral เปิดตัวโมเดลความปลอดภัยแบบโอเพนซอร์ส 3B Shieldstral ซึ่งมีขนาดเล็กกว่าแต่สามารถเทียบเท่ากับโมเดลขนาดเจ็ดเท่า - ข่าว AI Aioga","description":"Mistral เปิดตัวโมเดลความปลอดภัยแบบเปิด 3B พารามิเตอร์ Shieldstral ซึ่งตรวจสอบการละเมิดความปลอดภัยของอินพุตและเอาต์พุต AI ผ่านคำถาม \"ใช่/ไม่ใช่\" ในภาษาธรรมชาติ แทนการใช้ประเภทคงที่...","url":"https://www.aioga.com/th/news/cmsgbb37t004tro5qisawd1ya/","contentTranslated":true,"sourceHash":"bcc3db372e4a5492","translatedAt":"2026-08-05T17:08:03.319Z"},"pl":{"title":"Mistral udostępnia otwarty model bezpieczeństwa 3B Shieldstral, który w mniejszej objętości dorównuje modelowi o siedmiokrotnie większej skali","summary":"Mistral opublikował otwarty model bezpieczeństwa o 3 miliardach parametrów Shieldstral, który sprawdza naruszenia bezpieczeństwa wejść i wyjść AI za pomocą pytań w języku naturalnym „tak/nie”, a nie sztywnych kategorii. Model ten w niektórych testach porównawczych dorównuje modelom siedmiokrotnie większym, a także pozwala użytkownikom samodzielnie ustalać standardy w trakcie działania, bez konieczności polegania na systemach klasyfikacji stron trzecich, i może być uruchamiany lokalnie.","category":"模型更新","source":"The Decoder：AI News（RSS）","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"Mistral udostępnia otwarty model bezpieczeństwa 3B Shieldstral, który w mniejszej objętości dorównuje modelowi o siedmiokrotnie większej skali - Aioga Wiadomości AI","description":"Mistral opublikował otwarty model bezpieczeństwa o 3 miliardach parametrów Shieldstral, który sprawdza naruszenia bezpieczeństwa wejść i wyjść AI za pomocą pytań w języku naturalny...","url":"https://www.aioga.com/pl/news/cmsgbb37t004tro5qisawd1ya/","contentTranslated":true,"sourceHash":"bcc3db372e4a5492","translatedAt":"2026-08-05T17:08:49.822Z"}}}}