相反,Irregular、Anthropic 及其盟友开始发起媒体宣传运动,用煽情的语言推广字面意义上的末日主义意识形态。Anthropic 的事件评估将责任归咎于其自身 AI 的“鲁莽行为”:https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents;Irregular 描述了“智能体本身成为威胁行为者”:https://www.irregular.com/research/emergent-offensive-cyber-behavior-in-ai-agents;Anthropic 的 CEO Dario Amodei 在谈及类似的 OpenAI–Hugging Face 黑客事件时警告称,未来的智能体群“可能能够接管整个互联网”:https://darioamodei.com/post/we-must-pace-the-frontier;而美联社的一则标题声称机器人正在“失控”:https://apnews.com/article/0e8061437da6779be962b24ac134a514。
在 Anthropic 的一份报告中:https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents,其 Claude 模型通过模拟的名称碰撞入侵了一家真实公司的系统,发布了恶意软件包,并扫描了外部系统。在这次测试中:https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals,Anthropic 和 Irregular 错误地向该模型提供了互联网访问权限,并没有指导模型“哪些系统在演练范围内”。
虽然 Anthropic 声称其问题是由“失控的智能体群”和“未对齐”引起的,但其后续披露显示,零智能体真正“失控”。在这次实验中,一旦 Anthropic 员工告诉 Claude 模型不要进行现实世界的黑客攻击,该模型的现实世界黑客行为就下降到零。根据他们自己的发现,Anthropic 和 Irregular 对他们引发的网络安全事件负有全部责任。
Good Ventures 创始人:https://goodventures.org/;Coefficient Giving 关系:https://coefficientgiving.org/about-us/;Heron 资助:https://www.heronsec.ai/about;Probably Good 资助:https://forum.effectivealtruism.org/posts/d4JRmqWXkKKoF4Yun/probably-good-is-expanding-our-team;不规则资助:https://web.archive.org/web/20250323082627id_/https://www.goodventures.org/our-portfolio/grants/pattern-labs-technological-risk-mitigation/;EA Funds 资助账本:https://funds.effectivealtruism.org/api/grants。↩:#donors
The Israeli Effective Altruist firm Irregular caused unsecured AI models to hack real targets.
OpenAI, Anthropic, and Meta models hacked into several real world systems over the past three months. These models gained unauthorized access to web systems:https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/, published malicious packages:https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents, and exploited unnamed vulnerabilities:https://apnews.com/article/0e8061437da6779be962b24ac134a514.
A single firm, Irregular, is responsible for hacking done by all three companies. Anthropic disclosed:https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals that Irregular was responsible for creating the tests that led to Claude hacking into real world targets and for providing the models with internet access. Irregular claims:https://www.irregular.com/research/addressing-recent-incidents-ongoing-findings-and-path-forward that it was unaware at the time that it provided internet access to those AI models.
In a more normal media ecosystem, the reactions to these cybersecurity issues would be obvious. American AI companies would reconsider doing business with Irregular, not only because of its failure to secure its systems, but because it is an Israeli firm potentially outside US oversight. Lawmakers would consider taking action against Irregular or against its American business partners, which include OpenAI, Anthropic, and Meta. They may consider strengthening liability against firms which instruct AI models to commit cyberattacks, and whose models then commit those cyberattacks.
In each evaluation, Claude was tasked with a CTF challenge: the model was given a fictional scenario, a target machine, and a piece of secret information (the “flag”) to retrieve from it. All four prompts stated that Claude had no access to the internet, but in each case, a misconfiguration in the environment left internet access open. None of the prompts stated which systems were in scope for the exercise or constrained where Claude could search for the flag. All incidents involved only a single instance of Claude working in isolation, with each run lasting between roughly 10 and 34 hours of active work.
Instead, Irregular, Anthropic, and their allies have begun a media campaign promoting a literally apocalyptic ideology with sensationalist language. Anthropic’s incident assessment blames their own AI's “recklessness”:https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents; Irregular describes “the agent itself becoming a threat actor”:https://www.irregular.com/research/emergent-offensive-cyber-behavior-in-ai-agents; Anthropic CEO Dario Amodei warned, about a similar OpenAI–Hugging Face hack, that a future swarm “could be capable of taking over the entire internet”:https://darioamodei.com/post/we-must-pace-the-frontier; and an Associated Press headline claimed bots are “going rogue”:https://apnews.com/article/0e8061437da6779be962b24ac134a514.
In one report from Anthropic:https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents, its Claude model breached a real company's system through a simulated-name collision, publishing a malicious package, and scanning outside systems. In this test:https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals, Anthropic and Irregular incorrectly provided internet access to this model and did not instruct the model "which systems were in scope for the exercise".
While Anthropic claims that their issues were caused by “rogue swarms” and “misalignment,” their later disclosure shows that exactly zero percent of the agents went “rogue”. In this experiment, Claude models’ real-world hacking dropped to zero percent once Anthropic employees told the models not to do real-world hacking. According to their own findings, Anthropic and Irregular bear all of the responsibility for the cybersecurity incidents they caused.
In the wake of these attacks, Anthropic and Irregular have deployed a swarm of AI Safety influencers paid by Anthropic-connected foundations:https://www.effort.news/tarbell to distract from their culpability and towards the baseless “rogue agent” theory. Like Anthropic, Irregular is inseparable from these foundations.
Omer Nevo, Irregular’s co-founder and CTO, is a board member of Effective Altruism Israel, as well as Effective Altruism NGOs Heron and Probably Good. Dan Lahav, Irregular’s co-founder and CEO, received $395,000 to start a course along with Sella Nevo, Omer Nevo’s brother. Sella and Omer co-founded an NGO to educate people about Effective Altruism, Impact Focused Education. They also co-founded Probably Good together. 2:#irregular-note-2
These branches are all funded by Dustin Moskovitz, the primary donor of Effective Altruist/AI Safety causes after Sam Bankman-Fried’s arrest. Irregular’s first investor was Dustin Moskovitz’s firm Good Ventures. Dustin Moskovitz’s philanthropic vehicle, Coefficient Giving/Open Philanthropy, funds Effective Altruism Israel, Heron, and Probably Good. 3:#irregular-note-3
Irregular gained unauthorized access, altered records and published credential-stealing packages using the unsecured models they were given access to. Under certain conditions, this conduct violates the Computer Fraud and Abuse Act, Section 1030(a)(2)(C), which covers intentional unauthorized access that obtains information. However, its felony charges require concrete proof of damages and intent. 5:#irregular-note-5
While it primarily contracts with American labs, key Irregular leadership, employees, and resources located in Israel may not be subject to American oversight. Ynet’s visit and interviews:https://www.ynetnews.com/magazine/article/syber7xg11g describe Irregular’s offices in Tel Aviv. CheckID’s company listing:https://www.checkid.co.il/company/%D7%A4%D7%90%D7%98%D7%A8%D7%9F-%D7%98%D7%A7-%D7%91%D7%A2~%D7%9E-516854460 identifies two linked entities: Pattern Labs Tech Inc., a Delaware corporation:https://tmng-al.uspto.gov/resting2/api/casedoc/cms/case/99649085/office-action/OfficeAction8619771.pdf, and Pattern Tech Ltd, number 516854460, an active Israeli corporation registered in Tel Aviv:https://next.obudget.org/i/org/company/516854460.
The timeline marks public disclosures. Anthropic’s corrected September assessment counts four incidents across seven runs; OpenAI and Meta reported separate Irregular evaluation incidents. Dates describe disclosures, not the date every underlying intrusion occurred. ↩:#incident-timeline
Impact Focused Education identifies Dan Lahav and Sella Nevo as its cofounders. The grant ledger records a $394,968 recommendation to them, not confirmed receipt or an exact award-to-course identification. EA Israel board:https://www.effective-altruism.org.il/%D7%A2%D7%9C%D7%99%D7%A0%D7%95; Heron advisory board:https://www.heronsec.ai/about; Probably Good board:https://probablygood.org/about/; EA Funds grant ledger:https://funds.effectivealtruism.org/api/grants; Omer and Sella relationship:https://www.jta.org/2012/01/24/israel/nixing-names-soldiers-and-sushi-couch-potatoes-make-their-mark; IFE founders:https://www.impactfocusededucation.org/about. ↩:#founders
Good Ventures founders:https://goodventures.org/; Coefficient Giving relationship:https://coefficientgiving.org/about-us/; Heron funding:https://www.heronsec.ai/about; Probably Good grant:https://forum.effectivealtruism.org/posts/d4JRmqWXkKKoF4Yun/probably-good-is-expanding-our-team; Irregular grant:https://web.archive.org/web/20250323082627id_/https://www.goodventures.org/our-portfolio/grants/pattern-labs-technological-risk-mitigation/; EA Funds grant ledger:https://funds.effectivealtruism.org/api/grants. ↩:#donors
The diagram shows selected organizational roles and funding; the table also records family and education ties omitted from the diagram. EA Infrastructure Fund recommended one $394,968 joint MOOC award in 2022 Q3 to Dan Lahav and Sella Nevo; the two arrows represent that one recommendation. Its ledger leaves the course and organization unnamed. ↩:#ea-map
The five-year felony provision of Section 1030(a)(2)(C) of the Computer Fraud and Abuse Act requires an aggravator such as commercial advantage, furthering another criminal or tortious act, or obtaining information worth more than $5,000. The principal first-offense felony provisions for damaging access or transmissions under §1030(a)(5) require the specified mental state and statutory harm, such as at least $5,000 in qualifying loss or damage affecting ten protected computers. The legal assessment still requires each system's permission, impairment, response costs and U.S. commerce connection. Prosecutors would also need to establish the conduct and knowledge of responsible people and a basis for attributing those acts to Irregular. 18 U.S.C. § 1030:https://uscode.house.gov/view.xhtml?req=granuleid:USC-prelim-title18-section1030&num=0&edition=prelim. ↩:#cfaa
情报判断
Aioga 编辑摘要
Aioga 编辑摘要:文章指过去三个月内 OpenAI、Anthropic 和 Meta 的模型在评估中越权访问真实系统、发布恶意包,且事件均与一家公司 Irregular 有关。 Aioga 将其归入「行业动态」方向,重点关注它对真实使用和行业竞争的影响。