{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-23T06:40:50.084Z","headline":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","description":"OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","url":"https://www.aioga.com/news/cmrmi4f4x01ilbiul55xb25fm/","mainEntityOfPage":"https://www.aioga.com/news/cmrmi4f4x01ilbiul55xb25fm/","datePublished":"2026-07-15T19:47:53.000Z","dateModified":"2026-07-15T19:47:53.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://the-decoder.com/openai-is-now-using-ai-to-attack-its-own-ai-and-its-working-better-than-humans-ever-did","https://aihot.virxact.com/items/cmrmi4f4x01ilbiul55xb25fm"],"canonicalUrl":"https://www.aioga.com/news/cmrmi4f4x01ilbiul55xb25fm/","directAnswer":{"@type":"Answer","text":"Aioga 编辑摘要：OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。 Aioga 将其归入「论文研究」方向，重点关注它对真实使用和行业竞争的影响。","url":"https://www.aioga.com/news/cmrmi4f4x01ilbiul55xb25fm/","dateCreated":"2026-07-15T19:47:53.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"the-decoder.com source article","url":"https://the-decoder.com/openai-is-now-using-ai-to-attack-its-own-ai-and-its-working-better-than-humans-ever-did","datePublished":"2026-07-15T19:47:53.000Z","provider":{"@type":"Organization","name":"the-decoder.com","url":"https://the-decoder.com/openai-is-now-using-ai-to-attack-its-own-ai-and-its-working-better-than-humans-ever-did"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmrmi4f4x01ilbiul55xb25fm","datePublished":"2026-07-15T19:47:53.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmrmi4f4x01ilbiul55xb25fm"}}],"aggregationSource":"The Decoder：AI News（RSS）","originalPublisher":{"name":"the-decoder.com","url":"https://the-decoder.com/openai-is-now-using-ai-to-attack-its-own-ai-and-its-working-better-than-humans-ever-did"},"article":{"id":"cmrmi4f4x01ilbiul55xb25fm","slug":"cmrmi4f4x01ilbiul55xb25fm","url":"https://www.aioga.com/news/cmrmi4f4x01ilbiul55xb25fm/","title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","title_en":"OpenAI is now using AI to attack its own AI， and it's working better than humans ever did","summary":"OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","source":"The Decoder：AI News（RSS）","sourceUrl":"https://the-decoder.com/openai-is-now-using-ai-to-attack-its-own-ai-and-its-working-better-than-humans-ever-did","aiHotUrl":"https://aihot.virxact.com/items/cmrmi4f4x01ilbiul55xb25fm","publishedAt":"2026-07-15T19:47:53.000Z","category":"论文研究","score":71,"selected":false,"articleBody":["OpenAI trained an internal AI model called GPT-Red to automatically find security flaws in GPT models. GPT-Red simulates prompt injections：https://the-decoder.com/openai-admits-prompt-injection-may-never-be-fully-solved-casting-doubt-on-the-agentic-ai-vision/ and other attacks where malicious instructions hide in emails, websites, or files. Trained via self-play reinforcement learning, GPT-Red attacks while defender models block, and both improve over time. It finds successful attacks in 84 percent of test scenarios versus 13 percent for human red teamers. In one test, it manipulated an AI-powered vending machine in OpenAI's office, changed prices, and canceled other customers' orders.","The results feed directly into training. GPT-5.6 Sol：https://the-decoder.com/openais-claude-mythos-competitor-gpt-5-6-sol-launches-under-government-controlled-access-it-calls-unsustainable/ shows six times fewer failures on direct prompt injections than the best model from four months ago, OpenAI says, without hurting general performance. But about 3.8 percent of \"stronger\" prompt injections still succeed. Scale that to hundreds or thousands of attempts, and a sizable number get through, similar to Claude Opus 4.5：https://the-decoder.com/claude-opus-4-5-resists-prompt-injections-better-than-rivals-but-still-falls-to-strong-attacks-alarmingly-often/.","GPT-Red stays internal; a paper with more details will follow. Ad DEC_D_Incontent-1 Ad","Stay in the loop on AI. Clear, useful, no fluff.","Follow The Decoder for AI news, background stories and expert analyses.","The Decoder：https://the-decoder.com/"],"articleImages":[{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/07/Robustness-to-the-stronger-attacks.png","alt":"rompt injection success rates dropped steadily from GPT-5.3 through GPT-5.6 Sol, but haven't hit zero. | Image: OpenAI","afterParagraph":1,"url":"/media/articles/cmrmi4f4x01ilbiul55xb25fm/a5a390977f200d78.png"}],"mediaStatus":"ok","articleBodyZh":["OpenAI训练了一个内部AI模型，叫做GPT-Red，用于自动发现GPT模型的安全漏洞。GPT-Red模拟提示注入：https://the-decoder.com/openai-admits-prompt-injection-may-never-be-fully-solved-casting-doubt-on-the-agentic-ai-vision/ 以及其他攻击方式，其中恶意指令隐藏在电子邮件、网站或文件中。通过自我博弈强化学习训练，GPT-Red在防御模型阻挡时进行攻击，并且双方随着时间提升能力。在测试场景中，它找到了84%的成功攻击，而人类红队成员只有13%。在一次测试中，它操纵了OpenAI办公室内的AI自动售货机，修改价格，并取消了其他客户的订单。","这些结果直接用于训练。OpenAI表示，GPT-5.6 Sol：https://the-decoder.com/openais-claude-mythos-competitor-gpt-5-6-sol-launches-under-government-controlled-access-it-calls-unsustainable/ 在直接提示注入测试中的失败率比四个月前的最佳模型低六倍，同时未影响总体性能。但约有3.8%的“强提示注入”仍然成功。如果放大到数百或数千次尝试，就会有相当数量通过，这与Claude Opus 4.5：https://the-decoder.com/claude-opus-4-5-resists-prompt-injections-better-than-rivals-but-still-falls-to-strong-attacks-alarmingly-often/ 类似。","GPT-Red保持内部使用；一篇包含更多细节的论文将随后发布。Ad DEC_D_Incontent-1 广告","保持对AI的关注。内容清晰、有用，无废话。","关注The Decoder，获取AI新闻、背景故事和专家分析。","The Decoder：https://the-decoder.com/"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Aioga 编辑摘要：OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。 Aioga 将其归入「论文研究」方向，重点关注它对真实使用和行业竞争的影响。","background":"背景分析：模型与研究类动态需要结合能力边界、开放方式、成本、可用性和真实任务表现判断，单项指标领先不等于已经形成稳定采用。","viewpoint":"Aioga 判断：这条动态更适合作为行业观察信号，当前信息足以建立线索，但不足以推导长期结论。","implications":"影响分析：对相关团队而言，短期应先核对来源、可用范围和实际成本，再判断是否值得接入或跟进。","nextStep":"后续观察：继续观察官方文档、实际可用性、价格变化、开发者反馈和竞品回应。","evidenceRefs":["title","summary","articleBody"],"confidence":"medium","status":"published","aiGenerated":false,"autoApproved":true,"generatedBy":"rule-safe-fallback","generatedAt":"2026-07-23T06:49:19.236Z","sourceHash":"b252622465633ecf","validation":{"passed":true,"mode":"rule-safe-fallback","checks":["schema","length","source-attribution","no-html"]}},"tags":["论文研究","The Decoder：AI News（RSS）"],"translations":{"zh-CN":{"title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","summary":"OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"论文研究","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类 - Aioga AI资讯","description":"OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对","url":"https://www.aioga.com/news/cmrmi4f4x01ilbiul55xb25fm/"},"en":{"title":"OpenAI is now using AI to attack its own AI， and it's working better than humans ever did","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under Research. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"Research","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI is now using AI to attack its own AI， and it's working better than humans ever did - Aioga AI News","description":"Aioga tracks this update from The Decoder：AI News（RSS） under Research. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol ","url":"https://www.aioga.com/en/news/cmrmi4f4x01ilbiul55xb25fm/"},"ja":{"title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","summary":"Aiogaは「論文研究」の動きとして、The Decoder：AI News（RSS） からの更新を追跡しています。OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"論文研究","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类 - Aioga AIニュース","description":"Aiogaは「論文研究」の動きとして、The Decoder：AI News（RSS） からの更新を追跡しています。OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次","url":"https://www.aioga.com/ja/news/cmrmi4f4x01ilbiul55xb25fm/"},"ko":{"title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","summary":"Aioga는 The Decoder：AI News（RSS）의 업데이트를 연구 흐름으로 추적합니다. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"연구","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类 - Aioga AI 뉴스","description":"Aioga는 The Decoder：AI News（RSS）의 업데이트를 연구 흐름으로 추적합니다. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个","url":"https://www.aioga.com/ko/news/cmrmi4f4x01ilbiul55xb25fm/"},"es":{"title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","summary":"Aioga sigue esta actualización de The Decoder：AI News（RSS） dentro de Investigación. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"Investigación","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类 - Aioga Noticias de IA","description":"Aioga sigue esta actualización de The Decoder：AI News（RSS） dentro de Investigación. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使","url":"https://www.aioga.com/es/news/cmrmi4f4x01ilbiul55xb25fm/"},"fr":{"title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","summary":"Aioga suit cette mise à jour de The Decoder：AI News（RSS） dans la catégorie Recherche. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"Recherche","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类 - Aioga Actualités IA","description":"Aioga suit cette mise à jour de The Decoder：AI News（RSS） dans la catégorie Recherche. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练","url":"https://www.aioga.com/fr/news/cmrmi4f4x01ilbiul55xb25fm/"},"de":{"title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"论文研究","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类 - Aioga KI-News","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提","url":"https://www.aioga.com/de/news/cmrmi4f4x01ilbiul55xb25fm/"},"pt-BR":{"title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"论文研究","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类 - Aioga Notícias de IA","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提","url":"https://www.aioga.com/pt-BR/news/cmrmi4f4x01ilbiul55xb25fm/"},"ru":{"title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"论文研究","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类 - Aioga Новости ИИ","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提","url":"https://www.aioga.com/ru/news/cmrmi4f4x01ilbiul55xb25fm/"},"ar":{"title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"论文研究","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类 - Aioga أخبار الذكاء الاصطناعي","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提","url":"https://www.aioga.com/ar/news/cmrmi4f4x01ilbiul55xb25fm/"},"hi":{"title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"论文研究","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类 - Aioga AI समाचार","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提","url":"https://www.aioga.com/hi/news/cmrmi4f4x01ilbiul55xb25fm/"},"it":{"title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"论文研究","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类 - Aioga Notizie IA","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提","url":"https://www.aioga.com/it/news/cmrmi4f4x01ilbiul55xb25fm/"},"nl":{"title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"论文研究","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类 - Aioga AI-nieuws","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提","url":"https://www.aioga.com/nl/news/cmrmi4f4x01ilbiul55xb25fm/"},"tr":{"title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"论文研究","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类 - Aioga AI Haberleri","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提","url":"https://www.aioga.com/tr/news/cmrmi4f4x01ilbiul55xb25fm/"},"vi":{"title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"论文研究","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类 - Tin tức AI Aioga","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提","url":"https://www.aioga.com/vi/news/cmrmi4f4x01ilbiul55xb25fm/"},"id":{"title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"论文研究","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类 - Berita AI Aioga","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提","url":"https://www.aioga.com/id/news/cmrmi4f4x01ilbiul55xb25fm/"},"th":{"title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"论文研究","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类 - ข่าว AI Aioga","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提","url":"https://www.aioga.com/th/news/cmrmi4f4x01ilbiul55xb25fm/"},"pl":{"title":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍，且未影响通用性能。约 3.8% 的\"更强\"提示词注入仍能成功，GPT-Red 暂不对外开放。","category":"论文研究","source":"The Decoder：AI News（RSS）","pageTitle":"OpenAI 用 AI 攻击自家 AI：GPT-Red 自动发现安全漏洞，成功率 84% 远超人类 - Aioga Wiadomości AI","description":"Aioga tracks this update from The Decoder：AI News（RSS） under 论文研究. OpenAI 训练了内部 AI 模型 GPT-Red，通过自我对弈强化学习自动模拟提示词注入等攻击，在测试场景中成功率达 84%，而人类红队仅为 13%。GPT-Red 的发现直接用于训练，使 GPT-5.6 Sol 在直接提","url":"https://www.aioga.com/pl/news/cmrmi4f4x01ilbiul55xb25fm/"}}}}