Tickets for AIE NYC:https://ai.engineer/nyc/2026 now open, and apply:https://ai.engineer/code/2026/apply for the invite-only AIE CODE:https://ai.engineer/code/2026 . Join us:https://x.com/aiDotEngineer/status/2078502554200359344 !
We have an unusual relationship with today’s guest: for years since coauthoring the InstructGPT paper:https://arxiv.org/abs/2203.02155 , Diogo Almeida:https://www.youtube.com/watch?v=cJ0EOzey--o had been saying that API-available frontier models have been going down the wrong path, everything from the alignment to refusals to reliability perspectives, that we have dropped every mode:https://docs.typesafe.ai/introduction/machine-learning-primer#the-problems-with-rlhf other than autoregressive chat-tuned LLMs because of the overwhelming success of ChatGPT.
In a launch video now viewed ~40M times (by comparison, GPT4o was 22M:https://x.com/OpenAI/status/1790072174117613963?s=20 , Fable 5 was 15M:https://x.com/AnthropicAI/status/2072163884430229756?s=20 , Navier Stokes was 74M:https://www.latent.space/p/ainews-here-are-6-clones-of-jev-in , and 6 Astra was 137M:https://x.com/OpenAI/status/2095595741528125780?s=20 ), Diogo introduced Jev and it immediately took over the AI timeline — we’ll skip full Jev explainers because your favorite AI influencer/educator has probably already done one. We also collected:
the official patterns:https://docs.typesafe.ai/patterns and cookbooks:https://docs.typesafe.ai/cookbooks/ you should see first, from Allie:https://x.com/allietheicon
speed based - games and computer use
the voice + computer use example:https://x.com/instantricecook/status/2100814590300889426 we discuss at 1h34 mins
voice + browser control:https://x.com/moritzkremb/status/2100577979021832365?s=20
The must not miss Doom demo:https://framerusercontent.com/assets/rlL7ImEbISFoYt3IJEHHfvjpY.mp4
Driving cars:https://x.com/jpschroeder/status/2100347770867458384?s=20 in games
Excalidraw:https://x.com/jackcheng/status/2100729670991802386?s=20
virtual try-ons:https://x.com/nailthy62/status/2101388186916454439?s=12
guided responses in text messages:https://x.com/yuhasbeentaken/status/2101698498567831740/photo/1
Jev for coding agents:https://docs.typesafe.ai/introduction/coding-agents has an official guide
jev for linting:https://x.com/ohansemmanuel/status/2101034822760288452?s=12
compacting tool calls:https://x.com/tamarajtran/status/2100694549362553153?s=12
reasonable pushback from Theo:https://x.com/theo/status/2100762304862384257 - Diogo has published a note on the Tyranny of the KV Cache:https://docs.google.com/document/d/1G61uUB0FifUnmmrPzFQojZ3KpczYKmXGpgEXDJ2l_Zg/edit?tab=t.0 that you should read as a followup after the pod for Jev + coding agents, because of his belief that Cache Rules Everything:https://x.com/CompleteSkeptic/status/2097738214589215173
Programming Languages built atop Jev:https://x.com/southpolesteve/status/2100767781868150938?s=12 (Diogo’s fave)
Jev for analytics:https://x.com/tarasshyn/status/2101012033340571952 replay and user journey review:https://x.com/regalstreak/status/2101189571375493239?s=12
entity resolution:https://x.com/hrishioa/status/2101362082369470675?s=12
natural language search:https://x.com/venturetwins/status/2101341075684434245?s=12
“ smart software:https://youtu.be/cJ0EOzey--o?si=nlFo1Y2XW5SW9vjs&t=697 ”
a core goal of Jev is to “disappear into the background” - eg as unremarkable as regex
Jev as a judge:https://x.com/langchain/status/2101454284927959080?s=12
Jev vs LLM capabiltiies:https://x.com/markjaquith/status/2101341256743813558
blending transformers and classifiers:https://x.com/nazo_btw/status/2100955791750476048?s=20
about the confidence api:https://x.com/JinjingLiang/status/2101547529532059736?s=20
Jev vs GLiNER:https://x.com/george_onx/status/2100293114808119379?s=12 (note difference/pushback:https://x.com/mkhordoo/status/2101125110681784676?s=12 , agreed:https://x.com/irl_danb/status/2100935837470843075?s=12 , agreed:https://x.com/joelgrus/status/2101437270142099965?s=12 , agreed:https://x.com/mparakhin/status/2101683565520199887?s=12 )
Jev on trolley problem:https://x.com/The_Alex/status/2100619644973486252?s=20
Jev Bush:https://x.com/zeddotdev/status/2100390620640526554?s=20
Instead we’ll focus on what we can uniquely offer — a broader philosophical and mission-based understanding of how and why Jev was created , and what you should expect next in terms of future models from TypeSafe ( ReasoningJev:https://x.com/lateinteraction/status/2101380477495996699?s=12 ?) and what usecases and ideas you should work on vs the 55th low effort clone of Jev’s API:https://www.latent.space/p/ainews-here-are-6-clones-of-jev-in or doing a generic JevBench:https://x.com/airesearch12/status/2101311992984199580?s=12 benchmark - something Diogo has rejected publicly:https://x.com/CompleteSkeptic/status/2098463042065572038 .
Diogo knows a good deal about RLHF, given that he was on the team that pioneered post-training at OpenAI — and traces the three branches to Christiano et al 2017:https://arxiv.org/abs/1706.03741 (the robot backflip demo), Stiennon et al 2020:https://arxiv.org/abs/2009.01325 (learning to summarize) and his baby, Ouyang et al 2022:https://arxiv.org/abs/2203.02155 (InstructGPT). From there on, every innovation from Function Calling:https://www.latent.space/p/devday-2024?utm_source=publication-search to Structured Outputs:https://www.latent.space/p/openai-api-and-o1?utm_source=publication-search to Reasoning:https://www.latent.space/p/karina?utm_source=publication-search felt like a hack on top of the string based, sequence to sequence prediction paradigm. As he mentions on the pod, from 2023-2024 he struggled unsuccessfully, due to both personal and organization underestimation, to train a model that accurately addressed what he saw as the core problem with making LLMs the heart of software: reliability .
Jev’s core innovation is " Reinforcement Learning for Calibrated Decisions ”, a novel, unpublished technique that optimizes for “answers with epistemically honest probabilities on System One tasks” rather than human rated feedback (RLHF):https://www.latent.space/p/rlhf-201?utm_source=publication-search — which causes hallucinations, sycophancy, and permanent reliance on humans — or programmatically verifiable outputs with rubrics (RLVR):https://www.youtube.com/@LatentSpaceTV/search?query=rl — which solves Navier Stokes but exacerbates jagged intelligence:https://x.com/karpathy/status/1816531576228053133?lang=en and doesn’t integrate well with other software.
We’ve talked about the calibration problem:https://www.latent.space/p/benchmarks-201?utm_source=publication-search before on the pod, but probably the single best place to understand why RLCD became necessary is Diogo’s AIE talk:https://www.youtube.com/watch?v=cJ0EOzey--o , which discusses why a generation of training helpful AI assistants for humans has impaired them for training models for composable, programmable AI for automation:https://typesafe.ai/manifesto .
At the end he also teases his contrarian opinion on scaling laws:https://x.com/CompleteSkeptic/status/2073442518117884197 - which teases how to build a modern neolab without the billions of dollars the major labs have…
We spend a good amount of time discussing Diogo’s essay on the Bitterest Lesson:https://x.com/CompleteSkeptic/status/2098097767512179135 :
His point is that “You get what you optimize for and the bitterest lesson in ML is that the most important part of it isn’t ML at all.” - and picking the right north star, eg upvoting for user preference vs being integrated into tool calls - makes everything else fall in line.
We’re excited to catch up with a freshly dyed:https://x.com/typesafeai/status/2101451220896682107 Diogo to discuss:
Why AI can solve extraordinarily hard problems but still fail to automate basic work
What System One Models are and why Jev is built for software rather than chat
RLHF, mode collapse , calibration, and the hidden costs of optimizing for human preferences
Why refusals become a problem when AI is buried inside software dependencies
Why TypeSafe rejects public benchmarks and optimizes for intelligence per dollar
The “bitterest lesson”: why the right task and the right data can matter more than compute
Why TypeSafe thinks of itself as a data lab rather than a model lab
RLCD vs. RLHF and RLVR as fundamentally different North Stars for AI
Why reliability and robustness matter more than simple determinism
Jev’s programming primitives and how intelligence maps into software control flow
Why developers should decompose AI workflows into small, measurable decisions
How structured state replaces giant prompts and system messages
Why Diogo thinks AI should eventually disappear into the background of software
The “inverse SaaS-pocalypse” and how AI could supercharge existing software
System One vs. System Two intelligence and the limits of reasoning models
Dark data, computer use, real-time intelligence, and Jev’s biggest early use cases
Why Jev could reshape coding agents built around a single-model architecture
Why Diogo says he wouldn’t pre-train with $1 billion
The OpenAI journey that led to TypeSafe and why he thinks many neo-labs are approaching AI incorrectly
Coding agents beyond the KV cache , shared state, sub-agents, and the multi-agent future
LinkedIn: https://www.linkedin.com/in/diogomda:https://www.linkedin.com/in/diogomda
X: https://x.com/CompleteSkeptic:https://x.com/CompleteSkeptic
TypeSafe AI: https://typesafe.ai/:https://typesafe.ai/
