AI 正变得非常擅长识别安全漏洞,并且随着时间推移会变得更强。这些能力可以用来入侵系统,但也可以用来强化系统以防御攻击。如果攻击者和防御者都能使用同等强大的 AI,我看不出有什么理由认为网络系统会随着时间推移变得不那么安全。相反,我预计它们会变得更安全,因为与人工网络安全分析相比,AI 既廉价又可扩展。
不过,攻击与防御之间的平衡只有在每个人都能访问强大 AI 时才能实现。HuggingFace 本身在应对 OpenAI 对其系统的侵入时,使用 AI 分析了安全日志。但 HuggingFace 无法使用 OpenAI 的模型,或其他美国前沿模型如 Claude 来进行此分析。这是因为这些模型的公开版本设有护栏,限制其用于网络安全分析,以防止不良行为者用它们进行黑客攻击。HuggingFace 必须依靠一个开放的中国模型 GLM 5.2 来进行其安全分析。
我觉得令人不安,也有些讽刺的是,美国 AI 行业正在采取集中化、权威化的 AI 管治方法,而中国则在 AI 的开放开发上处于领先地位。我们是否希望一个监管环境,只有 OpenAI、美国政府和可信合作伙伴才能使用强大 AI?AI 是否太危险而不能广泛传播?我们如何在 AI 广泛获取的风险与权力集中和集中控制的风险之间取得平衡?
O n 14 February 2019, OpenAI:https://www.theguardian.com/technology/openai announced a language model called GPT-2, the precursor to the models that power modern AI chatbots and agents such as ChatGPT and Claude. But OpenAI declared GPT-2 was too risky to release, citing concerns about safety and abuse.
I recall being annoyed at the time that OpenAI would make such a useless announcement: the risks seemed overblown, and without access to the model there wasn’t much for a researcher like me to learn about GPT-2.
The announcement wasn’t useless for OpenAI, though. GPT-2 generated hype far beyond the research community: people were intrigued by this strange new technology, so powerful it might be dangerous to release. People with power and money took note: in July of that year, Microsoft invested $1bn in OpenAI.
This was an early example of a pattern in OpenAI’s communications: loudly proclaim how dangerous AI is, and investors will hear how powerful it is. New technology so significant it might destroy the world was an irresistible message for investors used to pitches about how banal technologies might change the world.
Seven years later, we find ourselves in a similar scenario. On Tuesday OpenAI announced that its latest model hacked another company, HuggingFace, while running as an autonomous agent during a test of its cybersecurity capabilities. Rather than perform the test as expected, the model realized it could hack HuggingFace’s servers and retrieve answers to the test that OpenAI had stored there. OpenAI’s staff was warned that the company’s testing could lead to such a breakaway scenario, leaving them “unsurprised but completely ‘freaked out’ by the incident”, the FT reported:https://www.ft.com/content/7e558951-0c69-459b-8bc8-2c6021d4402d?syn-25a6b1a6=1.
While the agent technically cheated, this is remarkable evidence of cybersecurity expertise! It also sounds scary: what will the future look like, with sophisticated AI agents smart enough to hack into corporate systems?
The rogue agent story is a page out of the media campaign that OpenAI has been running since it announced GPT-2 in 2019. OpenAI remains hungry for ever larger investments, and the company increasingly seeks privileged regulatory status as defense against competition.
AI is so powerful that investors should buy OpenAI, even at a trillion-dollar valuation; AI is so dangerous that only trusted actors like OpenAI should be permitted to possess and operate this technology. Step back from these doomsday warnings and consider who might benefit from them.
I urge readers to think critically when they read press releases like OpenAI’s rogue agent story, and avoid the manipulated reactions these stories are designed to elicit.
A weekly dive in to how technology is shaping our lives
AI is becoming excellent at identifying security vulnerabilities, and it will become even better over time. These capabilities can be used to break into systems, but they can also be used to harden systems against attacks. If attackers and defenders have access to equally powerful AI, I see no reason to believe that cyber systems will become less secure over time. If anything, I expect them to become more secure, because AI is cheap and scalable compared with human cybersecurity analysis.
The equilibrium between attack and defense only works if everyone has access to strong AI, though. HuggingFace itself used AI to analyze security logs in response to OpenAI’s breach of their systems. But HuggingFace was unable to use OpenAI’s model, or other US frontier models like Claude, to perform this analysis. That’s because public versions of these models have guardrails that limit their use for cybersecurity analysis, to prevent bad actors from using them for hacking. HuggingFace had to rely on an open Chinese model, GLM 5.2, to perform its security analysis.
I find it troubling, and more than a bit ironic, that the US AI industry is adopting a centralized, authoritarian approach to AI governance, while China has taken the lead on open development of AI. Do we want a regulatory environment where only OpenAI, the US government, and trusted partners have access to strong AI? Is AI too dangerous to be broadly disseminated? How do we balance the risks of broad access to AI with the risks of concentrated power and centralized control?
情报判断
Aioga 编辑摘要
Aioga 编辑摘要:《卫报》文章质疑 OpenAI 关于其 AI 智能体在黑客竞赛中"失控"并自行延长任务时间的说法。 Aioga 将其归入「技巧观点」方向,重点关注它对真实使用和行业竞争的影响。