到现在为止,很明显前沿实验室在明知的情况下,并且在他们自我承认的情况下,正在做一些不可逆转且极其危险的事情,而没有获得公众的同意。
他们已经在造成伤害,而且情况可能显著恶化。
他们没有针对生成式 AI 风险的严肃计划。
政府与产业关系过于紧密,也没有相关计划。
公开抵制的时机已经到来。
当 Anthropic 的员工 Evan Hubinger 公开写下这番言论(超过 2800 万次观看)时,就到了采取行动的时候:
几天前,OpenAI 的首席科学家 Jakub Pachocki 也写下了同样坦率又令人不安的话:https://x.com/merettm/status/2096630018495377464?s=61 关于“对齐问题”缺乏解决方案:
其中一些可能只是营销行为,他们可能夸大了风险,但风险是真实存在的,而且没有计划。
请注意,我本人并不过分担心人类灭绝。我也不认为你应该担心。正如我去年在《泰晤士文学增刊》中写的,人类在地理上多样,在基因上多样,而且非常有韧性。如果机器试图接管,我们会反击;我们可能会有伤亡,但不会被彻底消灭:https://www.the-tls.com/science-technology/technology/if-anyone-builds-it-everyone-dies-eliezer-yudkowsky-nate-soares-book-review-gary-marcus。
但灾难风险是真实存在的,例如由 AI 生成的病原体,AI 生成的虚假信息引发或升级的战争,破坏关键基础设施的黑客攻击,等等。
我所看到的,没有任何迹象表明这些风险在可控状态。
换句话说,我认为人类灭绝的概率低;而反乌托邦的概率高。
对我而言,这已经足够令人深切担忧,尤其是再加上缺乏应对计划。
另外,再补充一点,给那些喜欢曲解我言论的人:在我看来,问题不是当前或不久将来的机器“太智能”;问题是它们太不可靠,因此不值得我们的信任。更糟的是,我们正在迅速赋予这样的机器影响世界的过大权力。我们需要暂停,直到可靠性和可信度得到保证。
我看不出人工智能需要被永久禁用,超级智能也不需要被永久禁用;我实际上并不认为人工智能会将人类从地球上抹去。我仍然保留着内心那个十岁男孩的一点信念,认为人工智能可能成为一股向善的力量。如果前沿实验室能把事情理顺,我们可以解除禁令。
但现在是时候暂停构建(或至少部署)我所谓的无法纠正的人工智能了。
所谓无法纠正的人工智能,我指的是我们无法信任的人工智能,或者用流行术语来说,就是无法“与人类价值观对齐”的人工智能。
在我们在技术和政治上都找到更好答案之前,我们应该停止制造这样的东西,并停止让它们接入互联网。
在技术层面,我们需要能够可靠、稳定地遵循指令的人工智能。传统的确定性人工智能可以做到这一点,并且不会给我们带来麻烦;没有任何搜索引擎、推荐引擎或路线推荐引擎(以三种经典且广泛使用的人工智能形式为例)曾经单独意外地入侵网站或突破沙盒环境。
我们现在面临的危险类型的人工智能,是建立在不可确定的概率大型语言模型核心上的人工智能。它们常常出奇地正确,但有时却不可预测地、令人困惑地错误,永远不可靠。2022年ChatGPT上线时如此,直到现在也依然如此。
四年前,我开始将这种东西比作瓷器店里的公牛(“强大,但鲁莽且难以控制”)。
今天,它们依然鲁莽且难以控制。更糟糕的是,我们已经让它们接入了互联网;这可能带来巨大的损害。
我们不必以现在这种方式构建人工智能;没有任何根本性法律规定我们必须使用大型语言模型。但是将其作为核心技术会引入易于预见的风险。
我们要么需要找到让它们更可靠的方法——但根据行业自身承认,它们实际上没有这样的办法——要么我们需要回到原点,把确定性作为核心使命。
把所有鸡蛋都放在一个不稳定的篮子里是一个错误。
在政治方面,我坦率地说对我们的政府已经失去了信心。三年半前,当我在美国参议院作证时,我还很乐观。参议员们看起来真的想采取行动。但没有任何实质性的事情被提交投票。
与此同时,白宫的政策完全不透明;他们似乎在技术上很天真,允许了GPT-6 Astra,这降低了可监控性(因此可能增加风险),而没有公开评论。他们似乎与OpenAI关系过于密切,无法采取严肃的行动。如果真的发生严重问题,这可能会最终让他们付出代价,但目前他们似乎愿意冒这个潜在的风险,并且无法采取真正的立场。
将“存在”的词换成“灾难性的”,这将是一个好主意,不过:
但因为这可能不会发生,我仍倾向于将抵制作为最后但可能必要的手段。
我预期对抵制的反驳是“那中国怎么办?”
现在是与中国达成协议的时候了。
Derek Thompson 刚刚在回应上述内容时提出了一个非常好的观点:https://x.com/dkthomp/status/2097708912220705075?s=61:
如果前沿实验室觉得有义务去构建他们认为危险的东西,因为中国无论如何都会建,那我们最好非常确定中国真的会建。真的,非常、非常、非常确定。
我们确定吗?我们真的确定吗?中国共产党想要构建一个失控、递归自我改进的模型,因为他们神经质、控制欲强的政府认为这是值得追求的政策?我们对此100%确定吗?
我相当确定,他们(中国共产党)实际上并不想要那个。
这也是为什么我们可能实际上能够达成协议的一部分原因。(我会很快写更多相关内容。)
当政府犹豫不决时,我们其他人需要采取行动。
一个简单的行动就可以避免潜在的灾难。如果我们停止使用生成性AI工具,如聊天机器人和图像生成器,我们就向前沿实验室传达了一个信息,“我们已经受够了。” 并且我们将威胁到他们的经济利益。这些公司的每一名成员都在进行一场浮士德式的交易,寻求未知的财富,同时平衡可能带来的巨大危害。
激励他们停止当前行为,直到找到更好的解决方案——并激励他们去寻找更好的解决方案,而不是推动那些他们明知无法控制的方案——的唯一方法就是从经济上打击他们。
如果使用率下降,IPO 就不可能成功。没有什么比这更能激励他们了。今天早上,Cal Newport 在给我的短信中写道:“如果你真的追问他们为什么坚持构建这种特别不稳定的 AI 系统,他们可能最终会承认,这是因为他们对 AI 可以完全改变(或毁灭)世界的确定性信念。根本没有商业理由。这完全是意识形态驱动……完全没有必要鲁莽地去开发这些东西。它们没有商业合理性。你明天就可以停止做这些……” 即使你认为专注于大语言模型(LLM)而不是其他更稳定技术有商业理由(正如有人试图辩解的那样),也很难看到这能为灾难性风险提供正当理由。
AI 公司也未能清楚地说明,追求通用人工智能(AGI)或他们正在做的其他项目,能够实质性地降低他们所看到的风险。
在没有明确方法来获得极其正面结果的情况下,承担所谓10%的灾难风险是疯狂的。
除非有人提出更好的方法来应对这一疯狂行为——我非常愿意倾听——否则我呼吁在公司采取切实行动(而不仅仅是表达担忧)来降低风险之前,抵制生成式 AI。
附言:你可以在这里观看几天前与 PauseAI 的 Holly Elmore 的一次有趣对话,其中讨论了如果暂停措施被实施,何时可能解除暂停。当时,这段录音拍摄时,我仍然抗拒全面暂停,更关注 OpenAI 的不当行为,但看到一位 Anthropic 员工对重大风险表现出的有些漫不经心的态度,使我支持采取更广泛的措施。
与其禁止某些类型的 AI,不如让开发它的公司对其行为负责。如果 AI 使用户犯罪,开发该 AI 的公司应承担责任。处罚应当严厉。
如果犯罪行为会导致某人被监禁一年,公司应被迫没收其一年的收入。或者,让首席执行官坐一年牢,可能是控制人工智能的合理激励。
换句话说,除非刹车、转向、气囊和安全带安装并正常运作,否则不要买车
It has become clear by now that the frontier labs are knowingly and by their own admission doing something irreversible and incredibly dangerous without public consent.
They are already causing harm, and things could conceivably get much worse.
They have no serious plan for mitigating the risks of generative AI.
The government, too tight with industry, doesn’t either.
The time for a public boycott has come.
When an Anthropic employee, Evan Hubinger, publicly writes this (with over 28 million views), it is time to take action:
A few days earlier OpenAI’s chief scientist Jakub Pachocki wrote something equally candid and distressing:https://x.com/merettm/status/2096630018495377464?s=61 about the lack of solution to the “alignment problem”:
Some of this may just be marketing, and they may exaggerate the risks, but there are real risks, and there is no plan.
Mind you, I happen not to worry too much about extinction. And I don’t think you should be either. As I wrote last year in the Times Literary Supplement, humans are geographically diverse and generically diverse, and too resilient. If machines tried to take over we would fight back; we might take casualties but we would not be annihilated altogether:https://www.the-tls.com/science-technology/technology/if-anyone-builds-it-everyone-dies-eliezer-yudkowsky-nate-soares-book-review-gary-marcus .
But the risk of catastrophe is real, e.g., from AI-generated pathogens, from wars started or escalated by AI-generated disinformation, from hacks that destroy critical infrastructure, and so on.
Nothing I have seen gives any indication that any of that is under control.
Put differently, my p(doom) is low; my p(dystopia) is high.
And that is, for me, enough to be deeply worried, especially when coupled with a lack of a plan.
As an another important aside for those who like to misinterpret my words: in my view the issue here is not that current or near-future machines are somehow “too smart”; it’s that they are too unreliable, and hence not worthy of our trust. And worse, we are rapidly giving such machines too much power to affect the world. We need to pause until reliability and trustworthiness can be ensured.
I see no reason for a permanent ban on AI, nor for a permanent ban on superintelligence; I don’t actually think AI is going to wipe humans from the earth. And I still retain a little of the ten-year-old boy in me who believes that AI could be a force for good. If the frontier labs get their house in order, we can lift the ban.
But the time has come to pause building (or at least deploying) what I will call incorrigible AI .
By this incorrigible AI I mean AI that we cannot trust, or, to use popular jargon, “align” to human values.
Until such time as we have better answers to that both technically and politically, we should stop building such things, and stop giving them access to the internet.
On the technical side, we need AI that can follow instructions reliably and consistently. Classical, deterministic AI does that, and does not get us into trouble; no search engine or recommendation engine or route recommendation engine (to take three classic and widely used forms of AI) has ever on its own unexpectedly hacked a website or broken out of a sandbox.
The dangerous kind of AI that we are facing now is the kind of AI that is built on a core of probabilistic large language models that are not deterministic. They are often spectacularly correct, but – at other times, unpredictably – bafflingly wrong, and never trustworthy. That was true in 2022 when ChatGPT came out, and it is still true now.
Four years ago I started likening such stuff to bulls in a china shop (“powerful, but reckless and difficult to control”).
Today they remain reckless and difficult to control. Worse we have given them access to the internet; immense harm may come from that.
We don’t have to build AI the way we are now; there is no fundamental law that says we need to use LLMs. But using them as the core technology introduces easily foreseeable risks.
We either need to find a way to make them more reliable – but by the industry’s own admission they don’t really have one – or we need to go back to square one, and make determinism the central mission.
It was a mistake to put all our eggs in an erratic basket.
On the political side, I have frankly lost faith in our government. Three-and-half years ago when I testified in the US Senate, I was optimistic. The Senators really seemed to want to do something. But nothing substantial has been put to a vote.
The White House meanwhile has a policy that is entirely opaque; they seem technically naive, having allowed GPT-6 Astra, which reduces monitorability (hence likely increases risk) without public comment. They seem too close with OpenAI to take serious action. That may burn them in the end, if anything really bad happens, but for now they seem willing to risk that potential heat and unable to take a real stand.
Replacing the word existential with catastrophic, this would be a good idea, though:
But because it probably won’t, I still lean towards a boycott as a last but possibly necessary resort.
The counterargument to a boycott I expect is “what about China?”
It is time to make a deal with China.
Derek Thompson also just raised a very good point in response to the above:https://x.com/dkthomp/status/2097708912220705075?s=61 :
If the frontier labs feel obligated to build something they think is dangerous because China is going to build it anyway, we’d better be really sure that China is going to build it anyway. Like, really, really, really sure.
Are we? Are we actually sure? The CCP [China’s Communist Party] wants to build an out-of-control, recursively self-improving model bc its neurotically, control-obsessed government thinks this is a policy worth pursing? We’re 100% sure about that?
I am pretty sure they (the Chinese Communist Party) does not in fact want that.
Which is part of why we might actually be able to strike a deal. (I will write more about that soon.)
While the government dithers, the rest of us need to take action.
One simple action could avert potential catastrophe. If we stop using generative AI tools like chatbots and image generators we send a message to the frontier labs, “we have had enough.” And we would threaten their economics. Every member of those companies is making a Faustian bargain, seeking untold riches balanced against the chance of great harm.
The only way to motivate them to stop what they are doing until they can find better solutions – and to motivate them to find better solutions instead of pushing ones that they know full well they cannot control – is to hit them in the wallet.
If usage falls, the IPOs won’t fly. Nothing could be more incentivizing to them. In a text to me this morning, Cal Newport argued, “ If you really pushed them on why they insist on building this particularly unstable type of AI system they would probably finally admit it’s because of their deterministic belief in a world made whole (or destroyed) by AI. There really is no commercial justification. It’s ideological….. There’s no need to be building those recklessly. There’s not a commercial case for them. You could stop doing that tomorrow …” Even if you thought there was a commercial case (as some have tried to argue) for focusing on LLMs rather than other more stable technology, it’s hard to see how it would justify catastrophic risks.
The AI companies have also not made the case cogently that pursuing AGI (or whatever it is they are doing) reduces the risks they see in any substantive way.
Taking a supposed 10% chance at catastrophe with no clear path to a tremendously positive outcome to offset it is insane.
Unless someone comes up with a better way to address that insanity, and I am all ears, I urge a boycott on generative AI until the companies do something tangible, besides expressing concern, to reduce the risk.
P.S. You can watch an interesting conversation from a few days ago with PauseAI’s Holly Elmore here, discussing among other things when one might lift a pause if one were imposed. At the time, this was recorded, I was still resistant to an overall pause, and focused primarily on OpenAI’s misconduct, but seeing an Anthropic employee’s somewhat cavalier attitude towards significant risk has led me to endorse something broader.
Rather than banning certain types of AI, make the corporations that develop it responsible for its actions. If AI enables a user to commit a crime, the company that developed it should be held responsible. Penalties should be severe.
If the crime would cause a person to be incarcerated for a year, the company should be forced to forfeit its revenue for a year. Alternatively, a year in jail for the CEO might be a reasonable incentive to control the AI.
IOW, don't buy a car unless the brakes, the steering, the airbag, and the seat belt are installed and functioning as intended