AIE 纽约门票:https://ai.engineer/nyc/2026 现已开放申请,申请链接:https://ai.engineer/code/2026/apply 仅限受邀的 AIE CODE:https://ai.engineer/code/2026。加入我们:https://x.com/aiDotEngineer/status/2078502554200359344!
我们与今天的嘉宾有一种不同寻常的关系:自从共同撰写 InstructGPT 论文:https://arxiv.org/abs/2203.02155 以来,多年来 Diogo Almeida:https://www.youtube.com/watch?v=cJ0EOzey--o 一直在说,可通过 API 可用的前沿模型正走在错误的道路上,从对齐、拒绝到可靠性等各个方面,我们已经放弃了除自回归聊天调优的大型语言模型之外的所有模式:https://docs.typesafe.ai/introduction/machine-learning-primer#the-problems-with-rlhf,因为 ChatGPT 的巨大成功。
在一段当前观看次数约为 4000 万的视频发布中(相比之下,GPT4o 为 2200 万:https://x.com/OpenAI/status/1790072174117613963?s=20,Fable 5 为 1500 万:https://x.com/AnthropicAI/status/2072163884430229756?s=20,Navier Stokes 为 7400 万:https://www.latent.space/p/ainews-here-are-6-clones-of-jev-in,6 Astra 为 1.37 亿:https://x.com/OpenAI/status/2095595741528125780?s=20 ),Diogo 介绍了 Jev,它立即占据了 AI 时间线——我们将跳过完整的 Jev 解释,因为你最喜欢的 AI 影响者/教育者可能已经做过一次。我们还收集了:
官方模式:https://docs.typesafe.ai/patterns 和教程:https://docs.typesafe.ai/cookbooks/ 你应该先看,来自 Allie:https://x.com/allietheicon
基于速度的游戏和电脑使用
语音 + 电脑使用示例:https://x.com/instantricecook/status/2100814590300889426 我们在 1 小时 34 分钟讨论此例子
语音 + 浏览器控制:https://x.com/moritzkremb/status/2100577979021832365?s=20
不容错过的毁灭游戏演示:https://framerusercontent.com/assets/rlL7ImEbISFoYt3IJEHHfvjpY.mp4
在游戏中驾驶汽车:https://x.com/jpschroeder/status/2100347770867458384?s=20
Excalidraw:https://x.com/jackcheng/status/2100729670991802386?s=20
虚拟试戴:https://x.com/nailthy62/status/2101388186916454439?s=12
短信中的引导回复:https://x.com/yuhasbeentaken/status/2101698498567831740/photo/1
Jev用于编码代理:https://docs.typesafe.ai/introduction/coding-agents 有一份官方指南
jev 用于代码检查:https://x.com/ohansemmanuel/status/2101034822760288452?s=12
压实工具调用:https://x.com/tamarajtran/status/2100694549362553153?s=12
Theo的合理反对:https://x.com/theo/status/2100762304862384257 - Diogo 发布了一篇关于 KV 缓存暴政的笔记:https://docs.google.com/document/d/1G61uUB0FifUnmmrPzFQojZ3KpczYKmXGpgEXDJ2l_Zg/edit?tab=t.0,你应该在Jev + 编码代理的播客之后阅读,作为跟进,因为他相信缓存主宰一切:https://x.com/CompleteSkeptic/status/2097738214589215173
基于 Jev 构建的编程语言:https://x.com/southpolesteve/status/2100767781868150938?s=12(Diogo 的最爱)
Jev 用于分析:https://x.com/tarasshyn/status/2101012033340571952 回放和用户旅程审查:https://x.com/regalstreak/status/2101189571375493239?s=12
实体解析:https://x.com/hrishioa/status/2101362082369470675?s=12
自然语言搜索:https://x.com/venturetwins/status/2101341075684434245?s=12
“智能软件:https://youtu.be/cJ0EOzey--o?si=nlFo1Y2XW5SW9vjs&t=697”
Jev的核心目标之一是“消失在背景中”——例如像正则表达式那样平凡无奇
Jev 作为一名法官:https://x.com/langchain/status/2101454284927959080?s=12
Jev 与 LLM 的能力比较:https://x.com/markjaquith/status/2101341256743813558
混合变压器和分类器:https://x.com/nazo_btw/status/2100955791750476048?s=20
关于 confidence API:https://x.com/JinjingLiang/status/2101547529532059736?s=20
Jev 对 GLiNER:https://x.com/george_onx/status/2100293114808119379?s=12 (注意差异/反驳:https://x.com/mkhordoo/status/2101125110681784676?s=12,同意:https://x.com/irl_danb/status/2100935837470843075?s=12,同意:https://x.com/joelgrus/status/2101437270142099965?s=12,同意:https://x.com/mparakhin/status/2101683565520199887?s=12 )
Jev 关于电车难题:https://x.com/The_Alex/status/2100619644973486252?s=20
Jev Bush:https://x.com/zeddotdev/status/2100390620640526554?s=20
相反,我们将重点关注我们独特能提供的东西——对Jev如何以及为何被创建的更广泛的哲学和使命性理解,以及在TypeSafe(ReasoningJev:https://x.com/lateinteraction/status/2101380477495996699?s=12 ?)未来模型中你应当预期什么,以及你应着手的用例和想法,而不是Jev API的第55个低努力克隆:https://www.latent.space/p/ainews-here-are-6-clones-of-jev-in 或进行通用的JevBench:https://x.com/airesearch12/status/2101311992984199580 基准测试——这是Diogo公开拒绝的:https://x.com/CompleteSkeptic/status/2098463042065572038。
Diogo对RLHF相当了解,因为他曾在OpenAI的团队中参与后训练的开创工作——并将三个分支追溯至Christiano等,2017:https://arxiv.org/abs/1706.03741(机器人后空翻演示)、Stiennon等,2020:https://arxiv.org/abs/2009.01325(学习摘要)以及他的“宝宝”,Ouyang等,2022:https://arxiv.org/abs/2203.02155(InstructGPT)。从那时起,每一个创新,从Function Calling:https://www.latent.space/p/devday-2024?utm_source=publication-search 到 Structured Outputs:https://www.latent.space/p/openai-api-and-o1?utm_source=publication-search,再到 Reasoning:https://www.latent.space/p/karina?utm_source=publication-search 都感觉像是在基于字符串的序列到序列预测范式上附加的黑客手段。正如他在播客中提到的,从2023到2024,由于个人及组织低估,他未能成功训练出一个模型,准确解决他认为将LLM作为软件核心的核心问题:可靠性。
Jev的核心创新是“用于校准决策的强化学习”(Reinforcement Learning for Calibrated Decisions),这是一种新颖、未公开的技术,它优化“在系统一任务上具有认识论诚实概率的答案”,而不是依赖人工评分反馈(RLHF):https://www.latent.space/p/rlhf-201?utm_source=publication-search——这会导致幻觉、谄媚以及对人的永久依赖——或者使用可程序验证输出的评分标准(RLVR):https://www.youtube.com/@LatentSpaceTV/search?query=rl——它能解决Navier Stokes问题,但加剧了参差不齐的智能:https://x.com/karpathy/status/1816531576228053133?lang=en,同时与其他软件的集成效果不佳。
我们之前在播客中讨论过校准问题:https://www.latent.space/p/benchmarks-201?utm_source=publication-search,但可能最能理解为何RLCD变得必要的地方是Diogo的AIE演讲:https://www.youtube.com/watch?v=cJ0EOzey--o,该演讲讨论了为什么为人类训练有用的AI助手这一代,会妨碍他们训练用于可组合、可编程自动化AI的模型:https://typesafe.ai/manifesto。
最后他还透露了他对规模定律的反向观点:https://x.com/CompleteSkeptic/status/2073442518117884197——谈及如何在不需要像大型实验室那样数十亿美元的情况下建立现代新实验室…
我们花了相当多的时间讨论Diogo关于『最苦涩的教训』的文章:https://x.com/CompleteSkeptic/status/2098097767512179135:
他的观点是“你得到的正是你所优化的,而在机器学习中最苦涩的教训是,最重要的部分根本不是机器学习本身。”——选择正确的北极星,例如是为用户偏好点赞还是被集成到工具调用中——会让其他一切顺其自然。
我们很高兴能与刚刚染发的Diogo见面:https://x.com/typesafeai/status/2101451220896682107 讨论:
为什么AI能够解决极其困难的问题,但仍未能自动化基础工作
什么是系统一模型以及为什么Jev是为了软件而非聊天而构建的
RLHF、模式崩溃、校准,以及为了优化人类偏好而隐藏的成本
为什么当AI被嵌入软件依赖中时,拒绝会成为问题
为什么TypeSafe拒绝公开基准,而是优化每美元的智力
“最苦涩的教训”:为什么正确的任务和正确的数据可能比计算能力更重要
为什么TypeSafe认为自己是一个数据实验室而非模型实验室
RLCD与RLHF和RLVR作为AI根本不同的北极星
为什么可靠性和稳健性比简单的确定性更重要
Jev的编程原语,以及智力如何映射到软件控制流
为什么开发者应将AI工作流分解为小而可测量的决策
结构化状态如何取代大型提示和系统消息
为什么 Diogo 认为 AI 最终应该融入软件的背景中
“逆向 SaaS 灾难”以及 AI 如何为现有软件提供强大动力
系统一与系统二智能以及推理模型的局限性
暗数据、计算机使用、实时智能,以及 Jev 最大的一些早期应用场景
为什么 Jev 可能重塑基于单一模型架构的编程代理
为什么 Diogo 说他不会用 10 亿美元进行预训练
导致 TypeSafe 的 OpenAI 旅程,以及他认为许多新型实验室在接近 AI 时的错误方法
超越 KV 缓存、共享状态、子代理的编程代理,以及多代理未来
领英: https://www.linkedin.com/in/diogomda:https://www.linkedin.com/in/diogomda
X: https://x.com/CompleteSkeptic:https://x.com/CompleteSkeptic
TypeSafe AI: https://typesafe.ai/:https://typesafe.ai/
Tickets for AIE NYC:https://ai.engineer/nyc/2026 now open, and apply:https://ai.engineer/code/2026/apply for the invite-only AIE CODE:https://ai.engineer/code/2026 . Join us:https://x.com/aiDotEngineer/status/2078502554200359344 !
We have an unusual relationship with today’s guest: for years since coauthoring the InstructGPT paper:https://arxiv.org/abs/2203.02155 , Diogo Almeida:https://www.youtube.com/watch?v=cJ0EOzey--o had been saying that API-available frontier models have been going down the wrong path, everything from the alignment to refusals to reliability perspectives, that we have dropped every mode:https://docs.typesafe.ai/introduction/machine-learning-primer#the-problems-with-rlhf other than autoregressive chat-tuned LLMs because of the overwhelming success of ChatGPT.
In a launch video now viewed ~40M times (by comparison, GPT4o was 22M:https://x.com/OpenAI/status/1790072174117613963?s=20 , Fable 5 was 15M:https://x.com/AnthropicAI/status/2072163884430229756?s=20 , Navier Stokes was 74M:https://www.latent.space/p/ainews-here-are-6-clones-of-jev-in , and 6 Astra was 137M:https://x.com/OpenAI/status/2095595741528125780?s=20 ), Diogo introduced Jev and it immediately took over the AI timeline — we’ll skip full Jev explainers because your favorite AI influencer/educator has probably already done one. We also collected:
the official patterns:https://docs.typesafe.ai/patterns and cookbooks:https://docs.typesafe.ai/cookbooks/ you should see first, from Allie:https://x.com/allietheicon
speed based - games and computer use
the voice + computer use example:https://x.com/instantricecook/status/2100814590300889426 we discuss at 1h34 mins
voice + browser control:https://x.com/moritzkremb/status/2100577979021832365?s=20
The must not miss Doom demo:https://framerusercontent.com/assets/rlL7ImEbISFoYt3IJEHHfvjpY.mp4
Driving cars:https://x.com/jpschroeder/status/2100347770867458384?s=20 in games
Excalidraw:https://x.com/jackcheng/status/2100729670991802386?s=20
virtual try-ons:https://x.com/nailthy62/status/2101388186916454439?s=12
guided responses in text messages:https://x.com/yuhasbeentaken/status/2101698498567831740/photo/1
Jev for coding agents:https://docs.typesafe.ai/introduction/coding-agents has an official guide
jev for linting:https://x.com/ohansemmanuel/status/2101034822760288452?s=12
compacting tool calls:https://x.com/tamarajtran/status/2100694549362553153?s=12
reasonable pushback from Theo:https://x.com/theo/status/2100762304862384257 - Diogo has published a note on the Tyranny of the KV Cache:https://docs.google.com/document/d/1G61uUB0FifUnmmrPzFQojZ3KpczYKmXGpgEXDJ2l_Zg/edit?tab=t.0 that you should read as a followup after the pod for Jev + coding agents, because of his belief that Cache Rules Everything:https://x.com/CompleteSkeptic/status/2097738214589215173
Programming Languages built atop Jev:https://x.com/southpolesteve/status/2100767781868150938?s=12 (Diogo’s fave)
Jev for analytics:https://x.com/tarasshyn/status/2101012033340571952 replay and user journey review:https://x.com/regalstreak/status/2101189571375493239?s=12
entity resolution:https://x.com/hrishioa/status/2101362082369470675?s=12
natural language search:https://x.com/venturetwins/status/2101341075684434245?s=12
“ smart software:https://youtu.be/cJ0EOzey--o?si=nlFo1Y2XW5SW9vjs&t=697 ”
a core goal of Jev is to “disappear into the background” - eg as unremarkable as regex
Jev as a judge:https://x.com/langchain/status/2101454284927959080?s=12
Jev vs LLM capabiltiies:https://x.com/markjaquith/status/2101341256743813558
blending transformers and classifiers:https://x.com/nazo_btw/status/2100955791750476048?s=20
about the confidence api:https://x.com/JinjingLiang/status/2101547529532059736?s=20
Jev vs GLiNER:https://x.com/george_onx/status/2100293114808119379?s=12 (note difference/pushback:https://x.com/mkhordoo/status/2101125110681784676?s=12 , agreed:https://x.com/irl_danb/status/2100935837470843075?s=12 , agreed:https://x.com/joelgrus/status/2101437270142099965?s=12 , agreed:https://x.com/mparakhin/status/2101683565520199887?s=12 )
Jev on trolley problem:https://x.com/The_Alex/status/2100619644973486252?s=20
Jev Bush:https://x.com/zeddotdev/status/2100390620640526554?s=20
Instead we’ll focus on what we can uniquely offer — a broader philosophical and mission-based understanding of how and why Jev was created , and what you should expect next in terms of future models from TypeSafe ( ReasoningJev:https://x.com/lateinteraction/status/2101380477495996699?s=12 ?) and what usecases and ideas you should work on vs the 55th low effort clone of Jev’s API:https://www.latent.space/p/ainews-here-are-6-clones-of-jev-in or doing a generic JevBench:https://x.com/airesearch12/status/2101311992984199580?s=12 benchmark - something Diogo has rejected publicly:https://x.com/CompleteSkeptic/status/2098463042065572038 .
Diogo knows a good deal about RLHF, given that he was on the team that pioneered post-training at OpenAI — and traces the three branches to Christiano et al 2017:https://arxiv.org/abs/1706.03741 (the robot backflip demo), Stiennon et al 2020:https://arxiv.org/abs/2009.01325 (learning to summarize) and his baby, Ouyang et al 2022:https://arxiv.org/abs/2203.02155 (InstructGPT). From there on, every innovation from Function Calling:https://www.latent.space/p/devday-2024?utm_source=publication-search to Structured Outputs:https://www.latent.space/p/openai-api-and-o1?utm_source=publication-search to Reasoning:https://www.latent.space/p/karina?utm_source=publication-search felt like a hack on top of the string based, sequence to sequence prediction paradigm. As he mentions on the pod, from 2023-2024 he struggled unsuccessfully, due to both personal and organization underestimation, to train a model that accurately addressed what he saw as the core problem with making LLMs the heart of software: reliability .
Jev’s core innovation is " Reinforcement Learning for Calibrated Decisions ”, a novel, unpublished technique that optimizes for “answers with epistemically honest probabilities on System One tasks” rather than human rated feedback (RLHF):https://www.latent.space/p/rlhf-201?utm_source=publication-search — which causes hallucinations, sycophancy, and permanent reliance on humans — or programmatically verifiable outputs with rubrics (RLVR):https://www.youtube.com/@LatentSpaceTV/search?query=rl — which solves Navier Stokes but exacerbates jagged intelligence:https://x.com/karpathy/status/1816531576228053133?lang=en and doesn’t integrate well with other software.
We’ve talked about the calibration problem:https://www.latent.space/p/benchmarks-201?utm_source=publication-search before on the pod, but probably the single best place to understand why RLCD became necessary is Diogo’s AIE talk:https://www.youtube.com/watch?v=cJ0EOzey--o , which discusses why a generation of training helpful AI assistants for humans has impaired them for training models for composable, programmable AI for automation:https://typesafe.ai/manifesto .
At the end he also teases his contrarian opinion on scaling laws:https://x.com/CompleteSkeptic/status/2073442518117884197 - which teases how to build a modern neolab without the billions of dollars the major labs have…
We spend a good amount of time discussing Diogo’s essay on the Bitterest Lesson:https://x.com/CompleteSkeptic/status/2098097767512179135 :
His point is that “You get what you optimize for and the bitterest lesson in ML is that the most important part of it isn’t ML at all.” - and picking the right north star, eg upvoting for user preference vs being integrated into tool calls - makes everything else fall in line.
We’re excited to catch up with a freshly dyed:https://x.com/typesafeai/status/2101451220896682107 Diogo to discuss:
Why AI can solve extraordinarily hard problems but still fail to automate basic work
What System One Models are and why Jev is built for software rather than chat
RLHF, mode collapse , calibration, and the hidden costs of optimizing for human preferences
Why refusals become a problem when AI is buried inside software dependencies
Why TypeSafe rejects public benchmarks and optimizes for intelligence per dollar
The “bitterest lesson”: why the right task and the right data can matter more than compute
Why TypeSafe thinks of itself as a data lab rather than a model lab
RLCD vs. RLHF and RLVR as fundamentally different North Stars for AI
Why reliability and robustness matter more than simple determinism
Jev’s programming primitives and how intelligence maps into software control flow
Why developers should decompose AI workflows into small, measurable decisions
How structured state replaces giant prompts and system messages
Why Diogo thinks AI should eventually disappear into the background of software
The “inverse SaaS-pocalypse” and how AI could supercharge existing software
System One vs. System Two intelligence and the limits of reasoning models
Dark data, computer use, real-time intelligence, and Jev’s biggest early use cases
Why Jev could reshape coding agents built around a single-model architecture
Why Diogo says he wouldn’t pre-train with $1 billion
The OpenAI journey that led to TypeSafe and why he thinks many neo-labs are approaching AI incorrectly
Coding agents beyond the KV cache , shared state, sub-agents, and the multi-agent future
LinkedIn: https://www.linkedin.com/in/diogomda:https://www.linkedin.com/in/diogomda
X: https://x.com/CompleteSkeptic:https://x.com/CompleteSkeptic
TypeSafe AI: https://typesafe.ai/:https://typesafe.ai/