Nathan and Florian sit down to discuss everything happening with open models. Following the Kimi K3 release last week, it feels like everything is accelerating — geopolitics of US v China, economics of open vs. closed models, security at the frontier of AI, and so on. Chapters: 00:00 Welcome & context 04:38 Living with / using Kimi K3 08:53 GLM 5.2’s continued role 12:47 How are the Chinese models this good? 17:41 Data, environments, and a tour of the Chinese labs 19:47 Roundup of Chinese providers: Qwen, DeepSeek, MiniMax… 24:08 The US open-model ecosystem 30:25 Frontier vs. near-frontier, and the cybersecurity case against bans 34:58 Distillation and the Ben Thompson debate 44:12 Predictions and a frontier tier list 48:36 Wrap-up


Listen on Apple Podcasts:https://podcasts.apple.com/us/podcast/interconnects-audio/id1719552353 , Spotify:https://open.spotify.com/show/6XNzfJULeVxR7SneeesDUs , and where ever you get your podcasts:https://www.interconnects.ai/podcast . For other Interconnects interviews, go here:https://www.interconnects.ai/t/interviews .
Share :https://www.interconnects.ai/p/open-models-recap-more-on-kimi-k3?utm_source=substack&utm_medium=email&utm_content=share&action=share
For more educational post-training videos, see the course:https://rlhfbook.com/course I’m putting together.
00:00:06 Nathan Lambert: Okay, welcome back to Interconnects. We’re doing our quarterly open model roundup, which is mostly us just making fun of or explaining, not making fun, why so many distillation takes are bad and understanding the state of where things stand. I think last Thursday was when Kimi K3 was released. I think we will see much much more in the near future. It seems pretty inevitable. Like over the weekend, Xi gave his speech where he directly committed to openness and open source as a strategy. It wasn’t a detailed layout state of affairs.
Qwen announced their next big model is going to be open weight, which is a big change of things. I think there’s just so much to get into. I think Flo you kind of were already going off on some of the performance gap and distillation takes. So we could probably start there and then as I go I have a little bit a little list and we could always go through the topics and the blog that I wrote which all are very nuanced. So I think we have infinite to talk about. So continue rant kind.
00:01:17 Florian Brand: Yeah, I think, or the biggest thing at every model release at least at every open model release is how much or how many months it is behind the closed frontier. Um and people love to put a definite uh definitive number onto this uh which is really really mudding because we have so many different benchmark providers these days and such uh so many different benchmarks as well that every site and I’m not innocent in that either um pulls up their favorite benchmarks to show that the current model or the newly released model is at the frontier which is then counted by the other side pulling up another benchmark and showing oh it’s actually a year behind or something. Um and it like a lot of it seemingly hinges on that question how many months we open models are behind.
00:02:26 Nathan Lambert: Yeah. So I my provocation is that some of the benchmarks are actually reasonably correlated with what people are doing and this is agentic coding and agentic computer use tasks and some of the benchmarks are correlated with the long tail which is where I think Claude and GPT is so valuable. But if it’s it’s like what is the market for Claude Code and Codex right now and if it is software engineering then like the fact that the models are say a couple months behind on that can be a very, very big deal and then I suspect that this model will be okay disclaimer the model weights aren’t out yet supposedly on 20 July 27th and a lot of the discussion will impinge on the assumption that they come.
But like people could post-train this model to very likely match Opus and GPT in many of these kind of niche domains that people want I think watching I mean we both have different views into the post-training open model industry, but there is a ton ton of excitement in progress on making these models like fine-tuned for specific high-value tasks and this has been historically done on a mix of like Qwen and GLM and GLM 5.2 really accelerated this and I I curious on the first person that puts out a blog post like we fine-tuned Kimi K3 on our task because I bet you could get big gains. I think even the you use Kimi K3 more than I do, but my hunch is that it would be a bit of a um rough edged post-training just by how big of a scale up it is and that normally means there’s a lot of performance that could still be extracted from it. No.
00:04:03 Florian Brand: Yeah. Running, running, running and especially post-training that one will be super hard because you need like one node of B300s to just load the weights which is crazy in terms of scale. So will probably take some time and uh a lot of engineering I’ve heard to actually get this into a state where it’s fine-tunable. Um but people you want to talk about using the model like you actually signed up for the the coding program and used it. So like getting this out there is good context.
00:04:38 Nathan Lambert: Yeah. So I signed up on day after release or so uh for the $200 plan uh which is their biggest one similar to to all the others but they have um like I think $40 and $100 as well. Uh but the biggest plan has uh 1 million context and I think or at least it feels like it has also some priority in terms of the API requests because so many people um online are saying that they hit API errors constantly and so far I’ve been I’ve been uh pretty well off if I’m uh going to say that. Um and in terms of model capabilities, it at some ways aside from front end where it is really good, it in some ways it really shines and it excels.
Um even my um expectations even with things like uh some research tasks like I have or at interconnects we now have over a year of data on on open models um and I ask the frontier models to come up with some interesting analysis which we haven’t done before in uh because we do our own analysis and have this published uh and I asked them all right do something new and um surprise me, basically. And a lot of the models or the frontier models u or basically all models latch onto the things we’ve done redo the data analysis part and then do some weird esoteric parts.
Uh Kimi K3 did some more interesting things um I’ve told it explicitly to scrape Reddit um and then it found uh some subreddits I haven’t even considered and then found out for example that the Reddit discussions are um one or two months in uh more recent or they found they find the interesting models one or two months before the download numbers usually take off like they are all onto Qwen and then the people download more Qwen models like those kind of analysis is groundbreaking um but it is something that Kimi surprised me at compared to to all the other frontier um models.
A simple question like can you use this for most of the core work you do in terms of like the exp you you have a distribution of stuff you tend to do most of them are with Codex I think you’re a Codex person rather than Claude person like what percentage do you think the Venn diagram overlaps where this model would be fine
00:07:24 Florian Brand: uh it’s really depends on how much leeway I give it like, the big thing I have seen with Kimi K3 right now I’m I’m working on u the framework we are doing at uh at Prime Intellect, where I work, and the main thing I found with Kimi is its code is a lot simpler uh which makes it way more readable uh but it misses some things that Codex just or like we’re talking 56, 55 and especially 54 would be on those levels. So, I would say that Kimi K3 is like 54-55 level for these kind of tasks.
But if I like I read the code and I say all right that’s really good code and then I give it a pass over with with Codex and it finds all these niche niche cases where it doesn’t excel but for supervising runs or for running uh some experiments it is actually really usable. Um and for some other niche things like you can just let it run. The one downside is but that’s also because the API is completely swamped in terms of users and it their servers are in China. The wall clock time is significantly significantly higher than GPT. But I would say like if I was to to push it and use it in my daily workflow, I would be slower, but I wouldn’t be slowed down by so much that I would say, “All right, that’s unusable.”
00:08:53 Nathan Lambert: And how does this compare to GLM 5.2? Because GLM 5.2 was still a story unfolding in my opinion where like I would go bop around SF and people are like yeah I genuinely use this for this part of my like agentic coding and/or workflow. Um, how do you like I feel like were you in that camp using GLM at all or
00:09:20 Florian Brand: Yeah. like where do you I also use used and use uh GLM mostly because we have an internal endpoint which is really fast and we have or or before that I I also used an API which had I don’t know 200 or 300 tokens per second. Um and if you can do a lot of task at a good enough level like really fast you just use that model compared to going to Codex then selecting the lesser model then selecting the right reasoning effort then selecting fast like I just use GLM get the same result and uh and it’s uh pretty fine like it it definitely is Sonnet-ish level in terms of capabilities and for a lot of cleanup task for a task that just is grunt work. It really works. Like I I would say you could probably go really far for a lot of the work uh with Kimi K3 as the main agent and GLM for for sub agent work.
00:10:24 Nathan Lambert: Something that’s pretty different with Kimi’s announcement and the scale of models this is. I think it’ll take a bit longer for these open models to really be optimized and available across the inference providers. Like GLM 5.2 is pretty fast, but one, we don’t have the weights yet, and then two, like I don’t think it’s going to be as fast of a roll out on adoption as the like 500B, 700B MoE. like I I there’s going to be more problems there which is a very different regime where in the past the Chinese models would finish their RL run and release the model with open weights within hours to days maybe a week and then like immediately the ecosystem kind of knew how to do this.
I think there’s a lot bigger of an infrastructure kind of uplift on this next scale of open weight models which I think we have to factor in like the closed labs do this behind the scenes before announcing the models. So it’s just like that is kind of manipulating the time gap in a way where it could be like an extra month before people can actually post-train and use Kimi at scale for their workflows. And like we love to say as an open weight fan like oh it’s only when the closed model is available that you could take the time gap but like now there’s similar dynamics in open models where it’s like the Kimi API is totally broken. There’s too much supply. There’s too much demand. There’s not enough supply. So it’s not like this model is immediately diffusing like the I’m just I’m just thinking about this as it relates to the performance time gap
00:11:54 Florian Brand: because that that is true but on the other hand the open ecosystem has professionalized quite a lot in the last few months. uh like during or in your initial roll out they all come with some partners which have the weights beforehand. They have the vLLM patches out days or or even weeks before these days which is completely different from from a year ago where basically weights got dropped and the model makers were like all right you got to figure this out. So I expect like the general availability on day one will be pretty okay and then race starts of all the providers starting to optimize to get even higher and higher speeds because it’s so much prestige.
00:12:47 Nathan Lambert: Yeah. Okay. Two directions to go. Why do we think the Chinese models are able to be this good? I think I’ve wrote about there’s a debate in our Discord with in with JSD at at Epoch :https://epoch.ai/ and I think it’s very good and I had this section in my piece that I’m like coming around to think that the Chinese labs are more capital efficient and you can turn capital into compute data and talent in a way that makes the models better and this is really I think this is super important if it actually is some structural advantage whatever the cause I think the cause could be talent is better trained for whatever their education system was to work on problems that make LLMs better.
It could just be that all the compute and talent and everything cost way less in China somehow. Whether it’s a subsidy, whether it’s just average pay being lower. But this is a very big deal as we turn the crank in the model iterations. And if a next generation model costs $10 billion for Anthropic but only $4 billion for Kimi like this this is like could be very huge but it’s not clear why this is the case. For example, I think Big Eagle the Kimi engineer replied to my tweet on this and was like it helps because we’re not trying to push the frontier. are just trying to catch up, which really could be a mindset thing where how the the goals of the labs are scoped in in China so that it cost them way less money to build these models.
But I in the last year we’ve asked a lot of questions on like will the Chinese models fall off. I have thought that the gap between closed and open bottles would grow due to this kind of capital intensity of training and it seems like it’s going the opposite direction which is just like it’s it’s hard to unpack but like do you agree that the labs are keeping up a bit more than we would have expected as in the Chinese labs and why?
00:14:43 Florian Brand: Well, I I actually looked at our uh predictions for uh for this year based on our last year’s recap and we basically said that the gap will stay with within a few months. Uh so that prediction seems to largely hold. Um luckily for us, we didn’t put a concrete number whether it’s 3 months, 6 months or 9 months. So we are safe on that side. Um but I think like the general thing we both felt when we were in China and talking to these people like they are like the researchers themselves are teams of two or 300 people all mid20s and all just want one model to be really good like they don’t seem to do any side quests.
They don’t seem to do anything that uh deviates from from these things. And um they might or in terms of compute which is a really hard question for for us to answer especially as uh these Chinese uh chips are now coming online. We have I also think chips I think chip smuggling increased substantially in the last like 6 to 9 months or the chips that have been smuggled started to become online.
00:15:58 Nathan Lambert: Smuggling is a general term for getting around export restrictions. If the chips are in Malaysia and they’re using them, I I count that similar and I think that that has massively increased in the last six to nine months, this is the partially the result of that and and you’re saying but I just wanted to put that out there of like I do think that they have a lot more compute though than they did when they were training the previous generation of models.
00:16:24 Florian Brand: Yeah. like we like or just for for context two weeks ago I think LongCat released their model which they uh claim and I we know that it is very likely true uh is trained entirely on uh on Chinese chips. uh they didn’t specify publicly which ones but people speculate that it’s uh that it’s some uh Ascends from Huawei um and as the domestic production ramps up and you can they’re probably used most or they are used for for training but they are especially useful for inference which is a huge part of training as well.
So they probably use some mix of uh of Nvidia and other chips for the training part and then an increasingly larger part for the inference part during which during the stage is is really important. So I think their overall compute is increasing and also they don’t actually have a lot of users. So they don’t need to power 1 billion users like ChatGPT has to do, hundreds or thousands of enterprises like Anthropic has to do because they don’t have that magnitude of uh of of paying customers.
00:17:41 Nathan Lambert: Yeah. And I think even those paying customers also, at least on the enterprise side, there’s just like there is company time and chatter when you’re supporting these things. Even if you’re like not a research, even if it’s not in your job, it like does change the attention of the company. And if SSI comes out with a good model, it’ll be the ultimate validation that distractions are are a problem, but that’s an aside that we we can wait on. I think the there’s also rumblings of the data and environments industry starting to appear there.
Do you remember any specific ones? Because when we were in China, it was kind of shocking how little they seem to utilize external data. So just a few months hearing a whole bunch of a month months after our trip we went in April and then just months later in July, we’re are hearing a few things of like new companies in China and them wanting to buy data and things. And that is uh like a funny timeline of how that changes.
00:18:40 Florian Brand: And I would put error bars on what they actually told us.
00:18:44 Nathan Lambert: And that cuz it’s like so close in time that I don’t know.
00:18:49 Florian Brand: Yeah. That that that might that might be true. Uh but like those things are hard to to pinpoint. I but I would say it it seems like the buying of external data is becoming more of a factor. Um which will help the open models catch up to the closed ones if they just buy the same data maybe at a discount because um they buy the the data environments later. But it is it is a factor. How big of a factor like we don’t know. we don’t have any public insights and I doubt that we will get those insights uh from from anyone b uh really uh so that’s definitely one of the parts why um why we are able to to catch up or improve their their model scores.
00:19:47 Nathan Lambert: Okay, roundup of other Chinese model providers. We’ve talked about Kimi, we talked about Zhipu / GLM. I think there will be more GLM models soon that are very good. They might call it like GLM 5.5. Um Qwen, we talked about their biggest model coming. Qwen’s biggest models I will say have tended to relative to the excellence of their small models not had the same like absolute ranking in performance which is a probably a cost of focus. I think it goes with a cloud companies. It’s it’s almost like it’s if you squint it’s almost like Google.
It’s like Qwen has Alibaba has so much opportunity here and the opportunity of getting developers associated with Alibaba Qwen with these small models is such a huge opportunity for their cloud that I think they’re succeeding wildly. But their big models have always not been as excellent as their small models. So I don’t expect their model to be as breakthrough as Kimi K3 or GLM 5.2. I expect it to be covered in the news as major open quite as the open bottle name in China drops giant bottle but I don’t think it will be as sustained as a um news story um DeepSeek you can go if chime in whatever
00:21:01 Florian Brand: the the interesting thing is don’t know how how much you follow this but they are have or they have an endpoint which you can use for a preview version and they’ve updated this endpoint daily so they have some really fast iteration cycle because we the we progress in all these um Twitter um benchmarks. So a lot of these SVG things and three.js like all these visual generation tasks the model has been improving a lot over the last few days. So they have figured out some kind of fast feedback mechanism um which other companies have as well. Uh we we know this or cursor has a lot of blogs about this how they iterate really fast. Um but they seem to continuously upload new checkpoints and make them available.
00:21:47 Nathan Lambert: Um but I agree. I’m guessing it’s like a time gated within their final RL run. It’s like still slightly improving at the end of their RL run and they’re just like checking the box.
00:22:03 Nathan Lambert: Okay. Qwen DeepSeek V4 is supposed to come out a preview version. Um the thing about DeepSeek V4 I think is that the flash model is actually way more popular which is their smaller which seems to be an absolute workhorse for people. So that I think is the model to watch for them. I don’t expect V4 Pro to be a dramatic breakthrough. This is similar to anything like if Xiaomi were to release a new MiMo Pro model soon. I don’t expect it to be as big of a drop but it would probably be a very solid model. It’s just like it’s hard to know. They’re still a pretty new entrance. MiniMax, I think, is playing a different game. I don’t think MiniMax is chasing this um Kimi/GLM moonshot to AGI type vibe.
00:22:46 Florian Brand: Oh, I would, I would disagree there.
00:22:49 Nathan Lambert: You think, Do you think MiniMax is still in this?
00:22:52 Florian Brand: Yeah, I I I I think they they are seeing the tension especially because they are a public company similar to GLM and if you look at the stock performance RIP those stocks in the last few days um it it it it make it seems to make a huge difference and the interesting part will be uh the license because they’ve changed the license a lot uh to be more and more restrictive and um if there’s now a change of heart again after the Xi, uh, speech.
Uh it will be interesting to see whether MiniMax goes back to completely open licenses. It’s also an interesting thing to see um which license will be the license for for K3 because they have said they will open source it but I don’t think they have done any commitments in terms of the actual license where you put on top.
00:23:45 Nathan Lambert: Yeah. I mean that’s it’s super important is the thing. Yeah, we we’ll see. Um, Ling, Meituan, LongCat kind of similar, very strong models, probably getting a lot of value out of them internally. Aren’t don’t have the same developer breakthrough. Um, so what that’s like seven seven to eight Chinese labs. I might have forgotten some. And we can also talk about US labs. Aside um, Gemini 3.6 flash dropped. It looks fine. It’s like it’s like it’s it’s a tiny bump. It’s faster. It’s less of a yapper, but like doesn’t really matter. We’re going to stop we’ll stop sharing this. Um that’s that’s the amount of mention that Gemini gets for us.
But I do think it’s worth talking about the US ecosystem a bit. I think there are emerging players. Thinking machines released their first model. I’ve talked to some of them. they’re very on board for figuring out this how to make a fine-tunable model with Tinker and I think that’s a research area that I really really recommend for most of the open model builders. I think if you can get mind share there you will get massive adoption because it’s more about being fine-tunable for real tasks than it is about having that be best best numbers. Um, so this was their Inkling model which is a one trillion parameter which has like decent but not frontier scores.
I think kind of like DeepSeek V4 they’re going to they’re planning to release a smaller which is like a quarter of the size in total parameters which has really really good performance and if Inkling small preview comes out in a few weeks I do think that that will be a really used model. It’s a good size for kind of automating tasks and kind of domain specific tasks and might not be a like general agent type thing like Kimi and GLM 5.2 but I think that suits their business really well. Um I know that there’s some other the I would say like the smaller players in the US seem well like Arcee released their models earlier this year still chugging along. Poolside has started releasing some models.
They’ve gotten a few in the last few months and seem poised to release more models on top of that. So they’re really going Reflection is perpetually in the model coming soon camp and it really behooves them to get some models or some code or something out so that they can just start getting the developer flywheel going if they’re really committed to open source. It just takes a lot this it’s hard to get the models out. Like I talked to some people at Thinking Machines and it’s like kind of like oh that’s a lot of it’s a lot of work to actually do this I think. And um Nvidia chugging along. I think they’re at the stable player at this point. They’re keeping to release models. They’ll release more soon. They release a lot of data. I’m bullying them to try to get them to release Qwen style small models, which is like Gemma.
Gemma only has these like Qwen competitor models that are super popular. Um, the Gemma models are a little they’re all over the place in sizes or in architectures for the sizes and things like this, but the Gemma models are really really matching the Qwen models in terms of adoption. Um, I’m not sure they’re as easy to use for research, which could take a while. It could take multiple iterations. Like so much of language model research is now designed around small Qwen models and Qwen-based models that like it takes a while. Like people know how to use these models really well and if with the research results. So I hope Gemma keeps coming and can kind of compete in that niche. I don’t know any anyone that I missed here.
00:27:22 Florian Brand: No, I think both are the big players. Uh it’s, it is becoming broader. Uh in terms of model creators like last year, did we have any release aside from Gemma 3 and um GPT-OSS?
00:27:41 Nathan Lambert: was GPT-OSS 2 would go hard and obviously and obviously Nemotron as well. Um, oh, and I think Llama 4 at the start of the year, but uh, I don’t want that to be forgotten, but we are seeing like more players are are are now joining and turning out models at a really incredible rate.
00:27:59 Florian Brand: like Poolside has been releasing three or four models in the last two or three months. Uh and they seem to have figured out some way to turn out models pretty consistently. Um and that’s also something we are seeing on the open source side as well. we are talking about GLM like I think their iterations uh times for the model releases are now between 1 or 2 months with each new iteration becoming better and better which closely resembles what the closed labs are doing like we get a new GPT we get a new Claude every uh 6 weeks or so these days uh so in terms of having uh good enough pipeline uh to release stronger and stronger models they have to or the open source ecosystem has really figured it out or seemingly figured it out.
00:28:59 Nathan Lambert: Yeah, I agree. It’s it’s promising, but it is also so funny that like the US ecosystem started releasing some models and then then you have like Xi on the mic and these two models. It’s just like it’s so hard to catch up because it takes a lot of institutional expertise to train models that people actually use. And I think this is is what the American companies that are releasing models are now realizing is like these are not just benchmaxxed distilled IP theft models.
These are like genuinely good models that people are comparing to on their internal trading benchmarks and then like seeing how hard it is to beat them on measurable things. And I think that that is like I I’ve I’ve picked this sentiment up from a few people in the US trading models and it is just like there’s some I I think people should innovate on like size and fine-tunability and try to like use this potential market that is really close to home but also the pressures for every company is so high to release a model that you can claim as Frontier. I think investors expect that out of so many of these players that they’re kind of trying to do a a pretty hard thing and it’ll be interesting how the next year unfolds for the US China balance.
00:30:25 Florian Brand: Yeah, I think or in general I and a lot of other people have talked about the general ecosystem and that’s also something you’ve talked about at the very beginning. I think we are seeing more and more of a split between the capabilities of models that is good enough for a lot of tasks like uh for a lot of coding tasks the current frontier models both open and closed are good enough. um improvements feel less and less uh important here.
But if we look at the frontiers frontier, so finding new math proofs, finding uh new uh cures, finding new drugs, and inventing new things, that seems to be a whole different beast and probably will be dominated by the very frontier for quite a long time. The big question then becomes how much does that matter uh in terms of the addressable market and also how much of a focus will this be. I think, or my general base case is that we are seeing the frontier close down more and more. We have seen this with Mythos for cyber security GPT... or for biotech that those models won’t be accessible for everyone um and maybe not even external partners if we consider the reports that Anthropic is now spawning or or creating some internal labs to develop drugs.
