Hot take on OpenAI GPT-6 Astra* 1:#footnote-1 , with a challenge to Greg Brockman’s claims about it being AGI toward the end:
Looks to be pretty impressive. Multiple reports suggest it is a genuine advance.
As someone who has campaigned for nearly a decade for (neuro)symbolic world models, often to exceptional hostility, it is extraordinarily vindicating to see that a product from OpenAI explicitly creates and manipulate symbolic world models in the course of some of its most impressive computations.
What we don’t know is how robust that capability is. That is THE key question.
Success on ARC-AGI is great and impressive, but not —despite the name of the task—proof of AGI; I suspect we will see loads of problems with open-ended real world tasks. As with other recent models I would suspect best performance in verifiable domains.
And as a scientist, it’s disappointing that we don’t (yet?) know much about how the system actually works.
Without a clearer sense of what’s under the hood, I feel less confident about both what it can and can’t do, and what new risks we may encounter. I doubt the world is ready.
As ever, enthusiasts got an advance look; skeptics did not. That’s a sound marketing strategy, but it often turns out to be misleading. What we have often seen is initial enthusiasm that gets tempered over time. I suspect we will see that here as well.
The new system appears to be less:https://x.com/tomekkorbak/status/2095596839886274689?s=61 monitorable:https://x.com/tomekkorbak/status/2095596839886274689?s=61 than prior systems, which is not great from a safety perspective. One really doesn’t want more capability in conjunction with less monitorability. But also more alignable:https://x.com/tomekkorbak/status/2095596839886274689?s=61 , not sure why.
Would be great to see whether Astra can make progress on any of the ten tasks that Miles Brundage and I bet on at the end of 2024. (No AI to date has succeeded on any, AFAIK.)
This hot take is VERY tentative, pending more information about how the systems works and more detailed examination of what its limitations are.
Completely agree Gary. We need access to understand. What's being said and IF TRUE I'm trying to wrap my head around .... Is that the intelligence is similar. Yet the model is said to be more capable.
He see that in humans all the time.
Do we live in interesting times? Or just another PR bait that hasn't been thoroughly vetted.
PS. Curious too. This is the second time that arc prize has been there to pat OpenAI on the back on zero day. Just interesting.
情报判断
Aioga 编辑摘要
Gary Marcus 发文评价 GPT-6 Astra,称多项报告显示其是一次真正的进步,并指出 OpenAI 产品在部分重要计算过程中显式创建和操纵符号世界模型。
背景分析
Marcus 表示,GPT-6 Astra 在 ARC-AGI 上取得成功,但这不足以证明其达到 AGI;他同时指出,外界尚不了解该系统的实际工作方式,以及相关能力究竟有多鲁棒。
Aioga 观点
Aioga 判断:材料显示 Marcus 一方面认可 GPT-6 Astra 的技术进展,另一方面对其开放式现实任务表现、系统机制、能力边界及潜在新风险保持审慎,现有信息不足以支持更强结论。