{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-23T03:59:02.042Z","headline":"开源模型季度盘点：Kimi K3、Qwen 3.8、WAIC 演讲、知识蒸馏与开源闭源差距","description":"Nathan Lambert 与 Florian Brand 在播客中盘点开源模型最新动态。Kimi K3 发布后，中美 AI 地缘政治、开源与闭源模型的经济性、前沿安全等问题加速演进。Qwen 宣布下一代大模型将开源权重，中国厂商在开源策略上持续加码。讨论还涉及开源模型与闭源前沿的性能差距、知识蒸馏争议，以及后训练对特定任务的价值。","url":"https://www.aioga.com/news/cmrw7dmtf00ukro8gtnl09yk6/","mainEntityOfPage":"https://www.aioga.com/news/cmrw7dmtf00ukro8gtnl09yk6/","datePublished":"2026-07-22T14:09:04.000Z","dateModified":"2026-07-22T14:09:04.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.interconnects.ai/p/open-models-recap-more-on-kimi-k3","https://aihot.virxact.com/items/cmrw7dmtf00ukro8gtnl09yk6"],"canonicalUrl":"https://www.aioga.com/news/cmrw7dmtf00ukro8gtnl09yk6/","directAnswer":{"@type":"Answer","text":"Nathan Lambert 与 Florian Brand 在季度播客中讨论 Kimi K3 发布后的开放模型生态，议题包括中美模型产业、开源与闭源经济性、前沿安全、性能差距、知识蒸馏及后训练价值。","url":"https://www.aioga.com/news/cmrw7dmtf00ukro8gtnl09yk6/","dateCreated":"2026-07-22T14:09:04.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"interconnects.ai source article","url":"https://www.interconnects.ai/p/open-models-recap-more-on-kimi-k3","datePublished":"2026-07-22T14:09:04.000Z","provider":{"@type":"Organization","name":"interconnects.ai","url":"https://www.interconnects.ai/p/open-models-recap-more-on-kimi-k3"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmrw7dmtf00ukro8gtnl09yk6","datePublished":"2026-07-22T14:09:04.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmrw7dmtf00ukro8gtnl09yk6"}}],"aggregationSource":"Nathan Lambert：Interconnects（RSS）","originalPublisher":{"name":"interconnects.ai","url":"https://www.interconnects.ai/p/open-models-recap-more-on-kimi-k3"},"article":{"id":"cmrw7dmtf00ukro8gtnl09yk6","slug":"cmrw7dmtf00ukro8gtnl09yk6","url":"https://www.aioga.com/news/cmrw7dmtf00ukro8gtnl09yk6/","title":"开源模型季度盘点：Kimi K3、Qwen 3.8、WAIC 演讲、知识蒸馏与开源闭源差距","title_en":"Open models recap： more on Kimi K3， Qwen 3.8， Xi's WAIC speech， distillation， the open-closed gap， and what's next","summary":"Nathan Lambert 与 Florian Brand 在播客中盘点开源模型最新动态。Kimi K3 发布后，中美 AI 地缘政治、开源与闭源模型的经济性、前沿安全等问题加速演进。Qwen 宣布下一代大模型将开源权重，中国厂商在开源策略上持续加码。讨论还涉及开源模型与闭源前沿的性能差距、知识蒸馏争议，以及后训练对特定任务的价值。","source":"Nathan Lambert：Interconnects（RSS）","sourceUrl":"https://www.interconnects.ai/p/open-models-recap-more-on-kimi-k3","aiHotUrl":"https://aihot.virxact.com/items/cmrw7dmtf00ukro8gtnl09yk6","publishedAt":"2026-07-22T14:09:04.000Z","category":"技巧观点","score":62,"selected":true,"articleBody":["Nathan and Florian sit down to discuss everything happening with open models. Following the Kimi K3 release last week, it feels like everything is accelerating — geopolitics of US v China, economics of open vs. closed models, security at the frontier of AI, and so on. Chapters: 00:00 Welcome & context 04:38 Living with / using Kimi K3 08:53 GLM 5.2’s continued role 12:47 How are the Chinese models this good? 17:41 Data, environments, and a tour of the Chinese labs 19:47 Roundup of Chinese providers: Qwen, DeepSeek, MiniMax… 24:08 The US open-model ecosystem 30:25 Frontier vs. near-frontier, and the cybersecurity case against bans 34:58 Distillation and the Ben Thompson debate 44:12 Predictions and a frontier tier list 48:36 Wrap-up","Listen on Apple Podcasts：https://podcasts.apple.com/us/podcast/interconnects-audio/id1719552353 , Spotify：https://open.spotify.com/show/6XNzfJULeVxR7SneeesDUs , and where ever you get your podcasts：https://www.interconnects.ai/podcast . For other Interconnects interviews, go here：https://www.interconnects.ai/t/interviews .","Share ：https://www.interconnects.ai/p/open-models-recap-more-on-kimi-k3?utm_source=substack&utm_medium=email&utm_content=share&action=share","For more educational post-training videos, see the course：https://rlhfbook.com/course I’m putting together.","00:00:06 Nathan Lambert: Okay, welcome back to Interconnects. We’re doing our quarterly open model roundup, which is mostly us just making fun of or explaining, not making fun, why so many distillation takes are bad and understanding the state of where things stand. I think last Thursday was when Kimi K3 was released. I think we will see much much more in the near future. It seems pretty inevitable. Like over the weekend, Xi gave his speech where he directly committed to openness and open source as a strategy. It wasn’t a detailed layout state of affairs.","Qwen announced their next big model is going to be open weight, which is a big change of things. I think there’s just so much to get into. I think Flo you kind of were already going off on some of the performance gap and distillation takes. So we could probably start there and then as I go I have a little bit a little list and we could always go through the topics and the blog that I wrote which all are very nuanced. So I think we have infinite to talk about. So continue rant kind.","00:01:17 Florian Brand: Yeah, I think, or the biggest thing at every model release at least at every open model release is how much or how many months it is behind the closed frontier. Um and people love to put a definite uh definitive number onto this uh which is really really mudding because we have so many different benchmark providers these days and such uh so many different benchmarks as well that every site and I’m not innocent in that either um pulls up their favorite benchmarks to show that the current model or the newly released model is at the frontier which is then counted by the other side pulling up another benchmark and showing oh it’s actually a year behind or something. Um and it like a lot of it seemingly hinges on that question how many months we open models are behind.","00:02:26 Nathan Lambert: Yeah. So I my provocation is that some of the benchmarks are actually reasonably correlated with what people are doing and this is agentic coding and agentic computer use tasks and some of the benchmarks are correlated with the long tail which is where I think Claude and GPT is so valuable. But if it’s it’s like what is the market for Claude Code and Codex right now and if it is software engineering then like the fact that the models are say a couple months behind on that can be a very, very big deal and then I suspect that this model will be okay disclaimer the model weights aren’t out yet supposedly on 20 July 27th and a lot of the discussion will impinge on the assumption that they come.","But like people could post-train this model to very likely match Opus and GPT in many of these kind of niche domains that people want I think watching I mean we both have different views into the post-training open model industry, but there is a ton ton of excitement in progress on making these models like fine-tuned for specific high-value tasks and this has been historically done on a mix of like Qwen and GLM and GLM 5.2 really accelerated this and I I curious on the first person that puts out a blog post like we fine-tuned Kimi K3 on our task because I bet you could get big gains. I think even the you use Kimi K3 more than I do, but my hunch is that it would be a bit of a um rough edged post-training just by how big of a scale up it is and that normally means there’s a lot of performance that could still be extracted from it. No.","00:04:03 Florian Brand: Yeah. Running, running, running and especially post-training that one will be super hard because you need like one node of B300s to just load the weights which is crazy in terms of scale. So will probably take some time and uh a lot of engineering I’ve heard to actually get this into a state where it’s fine-tunable. Um but people you want to talk about using the model like you actually signed up for the the coding program and used it. So like getting this out there is good context.","00:04:38 Nathan Lambert: Yeah. So I signed up on day after release or so uh for the $200 plan uh which is their biggest one similar to to all the others but they have um like I think $40 and $100 as well. Uh but the biggest plan has uh 1 million context and I think or at least it feels like it has also some priority in terms of the API requests because so many people um online are saying that they hit API errors constantly and so far I’ve been I’ve been uh pretty well off if I’m uh going to say that. Um and in terms of model capabilities, it at some ways aside from front end where it is really good, it in some ways it really shines and it excels.","Um even my um expectations even with things like uh some research tasks like I have or at interconnects we now have over a year of data on on open models um and I ask the frontier models to come up with some interesting analysis which we haven’t done before in uh because we do our own analysis and have this published uh and I asked them all right do something new and um surprise me, basically. And a lot of the models or the frontier models u or basically all models latch onto the things we’ve done redo the data analysis part and then do some weird esoteric parts.","Uh Kimi K3 did some more interesting things um I’ve told it explicitly to scrape Reddit um and then it found uh some subreddits I haven’t even considered and then found out for example that the Reddit discussions are um one or two months in uh more recent or they found they find the interesting models one or two months before the download numbers usually take off like they are all onto Qwen and then the people download more Qwen models like those kind of analysis is groundbreaking um but it is something that Kimi surprised me at compared to to all the other frontier um models.","A simple question like can you use this for most of the core work you do in terms of like the exp you you have a distribution of stuff you tend to do most of them are with Codex I think you’re a Codex person rather than Claude person like what percentage do you think the Venn diagram overlaps where this model would be fine","00:07:24 Florian Brand: uh it’s really depends on how much leeway I give it like, the big thing I have seen with Kimi K3 right now I’m I’m working on u the framework we are doing at uh at Prime Intellect, where I work, and the main thing I found with Kimi is its code is a lot simpler uh which makes it way more readable uh but it misses some things that Codex just or like we’re talking 56, 55 and especially 54 would be on those levels. So, I would say that Kimi K3 is like 54-55 level for these kind of tasks.","But if I like I read the code and I say all right that’s really good code and then I give it a pass over with with Codex and it finds all these niche niche cases where it doesn’t excel but for supervising runs or for running uh some experiments it is actually really usable. Um and for some other niche things like you can just let it run. The one downside is but that’s also because the API is completely swamped in terms of users and it their servers are in China. The wall clock time is significantly significantly higher than GPT. But I would say like if I was to to push it and use it in my daily workflow, I would be slower, but I wouldn’t be slowed down by so much that I would say, “All right, that’s unusable.”","00:08:53 Nathan Lambert: And how does this compare to GLM 5.2? Because GLM 5.2 was still a story unfolding in my opinion where like I would go bop around SF and people are like yeah I genuinely use this for this part of my like agentic coding and/or workflow. Um, how do you like I feel like were you in that camp using GLM at all or","00:09:20 Florian Brand: Yeah. like where do you I also use used and use uh GLM mostly because we have an internal endpoint which is really fast and we have or or before that I I also used an API which had I don’t know 200 or 300 tokens per second. Um and if you can do a lot of task at a good enough level like really fast you just use that model compared to going to Codex then selecting the lesser model then selecting the right reasoning effort then selecting fast like I just use GLM get the same result and uh and it’s uh pretty fine like it it definitely is Sonnet-ish level in terms of capabilities and for a lot of cleanup task for a task that just is grunt work. It really works. Like I I would say you could probably go really far for a lot of the work uh with Kimi K3 as the main agent and GLM for for sub agent work.","00:10:24 Nathan Lambert: Something that’s pretty different with Kimi’s announcement and the scale of models this is. I think it’ll take a bit longer for these open models to really be optimized and available across the inference providers. Like GLM 5.2 is pretty fast, but one, we don’t have the weights yet, and then two, like I don’t think it’s going to be as fast of a roll out on adoption as the like 500B, 700B MoE. like I I there’s going to be more problems there which is a very different regime where in the past the Chinese models would finish their RL run and release the model with open weights within hours to days maybe a week and then like immediately the ecosystem kind of knew how to do this.","I think there’s a lot bigger of an infrastructure kind of uplift on this next scale of open weight models which I think we have to factor in like the closed labs do this behind the scenes before announcing the models. So it’s just like that is kind of manipulating the time gap in a way where it could be like an extra month before people can actually post-train and use Kimi at scale for their workflows. And like we love to say as an open weight fan like oh it’s only when the closed model is available that you could take the time gap but like now there’s similar dynamics in open models where it’s like the Kimi API is totally broken. There’s too much supply. There’s too much demand. There’s not enough supply. So it’s not like this model is immediately diffusing like the I’m just I’m just thinking about this as it relates to the performance time gap","00:11:54 Florian Brand: because that that is true but on the other hand the open ecosystem has professionalized quite a lot in the last few months. uh like during or in your initial roll out they all come with some partners which have the weights beforehand. They have the vLLM patches out days or or even weeks before these days which is completely different from from a year ago where basically weights got dropped and the model makers were like all right you got to figure this out. So I expect like the general availability on day one will be pretty okay and then race starts of all the providers starting to optimize to get even higher and higher speeds because it’s so much prestige.","00:12:47 Nathan Lambert: Yeah. Okay. Two directions to go. Why do we think the Chinese models are able to be this good? I think I’ve wrote about there’s a debate in our Discord with in with JSD at at Epoch ：https://epoch.ai/ and I think it’s very good and I had this section in my piece that I’m like coming around to think that the Chinese labs are more capital efficient and you can turn capital into compute data and talent in a way that makes the models better and this is really I think this is super important if it actually is some structural advantage whatever the cause I think the cause could be talent is better trained for whatever their education system was to work on problems that make LLMs better.","It could just be that all the compute and talent and everything cost way less in China somehow. Whether it’s a subsidy, whether it’s just average pay being lower. But this is a very big deal as we turn the crank in the model iterations. And if a next generation model costs $10 billion for Anthropic but only $4 billion for Kimi like this this is like could be very huge but it’s not clear why this is the case. For example, I think Big Eagle the Kimi engineer replied to my tweet on this and was like it helps because we’re not trying to push the frontier. are just trying to catch up, which really could be a mindset thing where how the the goals of the labs are scoped in in China so that it cost them way less money to build these models.","But I in the last year we’ve asked a lot of questions on like will the Chinese models fall off. I have thought that the gap between closed and open bottles would grow due to this kind of capital intensity of training and it seems like it’s going the opposite direction which is just like it’s it’s hard to unpack but like do you agree that the labs are keeping up a bit more than we would have expected as in the Chinese labs and why?","00:14:43 Florian Brand: Well, I I actually looked at our uh predictions for uh for this year based on our last year’s recap and we basically said that the gap will stay with within a few months. Uh so that prediction seems to largely hold. Um luckily for us, we didn’t put a concrete number whether it’s 3 months, 6 months or 9 months. So we are safe on that side. Um but I think like the general thing we both felt when we were in China and talking to these people like they are like the researchers themselves are teams of two or 300 people all mid20s and all just want one model to be really good like they don’t seem to do any side quests.","They don’t seem to do anything that uh deviates from from these things. And um they might or in terms of compute which is a really hard question for for us to answer especially as uh these Chinese uh chips are now coming online. We have I also think chips I think chip smuggling increased substantially in the last like 6 to 9 months or the chips that have been smuggled started to become online.","00:15:58 Nathan Lambert: Smuggling is a general term for getting around export restrictions. If the chips are in Malaysia and they’re using them, I I count that similar and I think that that has massively increased in the last six to nine months, this is the partially the result of that and and you’re saying but I just wanted to put that out there of like I do think that they have a lot more compute though than they did when they were training the previous generation of models.","00:16:24 Florian Brand: Yeah. like we like or just for for context two weeks ago I think LongCat released their model which they uh claim and I we know that it is very likely true uh is trained entirely on uh on Chinese chips. uh they didn’t specify publicly which ones but people speculate that it’s uh that it’s some uh Ascends from Huawei um and as the domestic production ramps up and you can they’re probably used most or they are used for for training but they are especially useful for inference which is a huge part of training as well.","So they probably use some mix of uh of Nvidia and other chips for the training part and then an increasingly larger part for the inference part during which during the stage is is really important. So I think their overall compute is increasing and also they don’t actually have a lot of users. So they don’t need to power 1 billion users like ChatGPT has to do, hundreds or thousands of enterprises like Anthropic has to do because they don’t have that magnitude of uh of of paying customers.","00:17:41 Nathan Lambert: Yeah. And I think even those paying customers also, at least on the enterprise side, there’s just like there is company time and chatter when you’re supporting these things. Even if you’re like not a research, even if it’s not in your job, it like does change the attention of the company. And if SSI comes out with a good model, it’ll be the ultimate validation that distractions are are a problem, but that’s an aside that we we can wait on. I think the there’s also rumblings of the data and environments industry starting to appear there.","Do you remember any specific ones? Because when we were in China, it was kind of shocking how little they seem to utilize external data. So just a few months hearing a whole bunch of a month months after our trip we went in April and then just months later in July, we’re are hearing a few things of like new companies in China and them wanting to buy data and things. And that is uh like a funny timeline of how that changes.","00:18:40 Florian Brand: And I would put error bars on what they actually told us.","00:18:44 Nathan Lambert: And that cuz it’s like so close in time that I don’t know.","00:18:49 Florian Brand: Yeah. That that that might that might be true. Uh but like those things are hard to to pinpoint. I but I would say it it seems like the buying of external data is becoming more of a factor. Um which will help the open models catch up to the closed ones if they just buy the same data maybe at a discount because um they buy the the data environments later. But it is it is a factor. How big of a factor like we don’t know. we don’t have any public insights and I doubt that we will get those insights uh from from anyone b uh really uh so that’s definitely one of the parts why um why we are able to to catch up or improve their their model scores.","00:19:47 Nathan Lambert: Okay, roundup of other Chinese model providers. We’ve talked about Kimi, we talked about Zhipu / GLM. I think there will be more GLM models soon that are very good. They might call it like GLM 5.5. Um Qwen, we talked about their biggest model coming. Qwen’s biggest models I will say have tended to relative to the excellence of their small models not had the same like absolute ranking in performance which is a probably a cost of focus. I think it goes with a cloud companies. It’s it’s almost like it’s if you squint it’s almost like Google.","It’s like Qwen has Alibaba has so much opportunity here and the opportunity of getting developers associated with Alibaba Qwen with these small models is such a huge opportunity for their cloud that I think they’re succeeding wildly. But their big models have always not been as excellent as their small models. So I don’t expect their model to be as breakthrough as Kimi K3 or GLM 5.2. I expect it to be covered in the news as major open quite as the open bottle name in China drops giant bottle but I don’t think it will be as sustained as a um news story um DeepSeek you can go if chime in whatever","00:21:01 Florian Brand: the the interesting thing is don’t know how how much you follow this but they are have or they have an endpoint which you can use for a preview version and they’ve updated this endpoint daily so they have some really fast iteration cycle because we the we progress in all these um Twitter um benchmarks. So a lot of these SVG things and three.js like all these visual generation tasks the model has been improving a lot over the last few days. So they have figured out some kind of fast feedback mechanism um which other companies have as well. Uh we we know this or cursor has a lot of blogs about this how they iterate really fast. Um but they seem to continuously upload new checkpoints and make them available.","00:21:47 Nathan Lambert: Um but I agree. I’m guessing it’s like a time gated within their final RL run. It’s like still slightly improving at the end of their RL run and they’re just like checking the box.","00:22:03 Nathan Lambert: Okay. Qwen DeepSeek V4 is supposed to come out a preview version. Um the thing about DeepSeek V4 I think is that the flash model is actually way more popular which is their smaller which seems to be an absolute workhorse for people. So that I think is the model to watch for them. I don’t expect V4 Pro to be a dramatic breakthrough. This is similar to anything like if Xiaomi were to release a new MiMo Pro model soon. I don’t expect it to be as big of a drop but it would probably be a very solid model. It’s just like it’s hard to know. They’re still a pretty new entrance. MiniMax, I think, is playing a different game. I don’t think MiniMax is chasing this um Kimi/GLM moonshot to AGI type vibe.","00:22:46 Florian Brand: Oh, I would, I would disagree there.","00:22:49 Nathan Lambert: You think, Do you think MiniMax is still in this?","00:22:52 Florian Brand: Yeah, I I I I think they they are seeing the tension especially because they are a public company similar to GLM and if you look at the stock performance RIP those stocks in the last few days um it it it it make it seems to make a huge difference and the interesting part will be uh the license because they’ve changed the license a lot uh to be more and more restrictive and um if there’s now a change of heart again after the Xi, uh, speech.","Uh it will be interesting to see whether MiniMax goes back to completely open licenses. It’s also an interesting thing to see um which license will be the license for for K3 because they have said they will open source it but I don’t think they have done any commitments in terms of the actual license where you put on top.","00:23:45 Nathan Lambert: Yeah. I mean that’s it’s super important is the thing. Yeah, we we’ll see. Um, Ling, Meituan, LongCat kind of similar, very strong models, probably getting a lot of value out of them internally. Aren’t don’t have the same developer breakthrough. Um, so what that’s like seven seven to eight Chinese labs. I might have forgotten some. And we can also talk about US labs. Aside um, Gemini 3.6 flash dropped. It looks fine. It’s like it’s like it’s it’s a tiny bump. It’s faster. It’s less of a yapper, but like doesn’t really matter. We’re going to stop we’ll stop sharing this. Um that’s that’s the amount of mention that Gemini gets for us.","But I do think it’s worth talking about the US ecosystem a bit. I think there are emerging players. Thinking machines released their first model. I’ve talked to some of them. they’re very on board for figuring out this how to make a fine-tunable model with Tinker and I think that’s a research area that I really really recommend for most of the open model builders. I think if you can get mind share there you will get massive adoption because it’s more about being fine-tunable for real tasks than it is about having that be best best numbers. Um, so this was their Inkling model which is a one trillion parameter which has like decent but not frontier scores.","I think kind of like DeepSeek V4 they’re going to they’re planning to release a smaller which is like a quarter of the size in total parameters which has really really good performance and if Inkling small preview comes out in a few weeks I do think that that will be a really used model. It’s a good size for kind of automating tasks and kind of domain specific tasks and might not be a like general agent type thing like Kimi and GLM 5.2 but I think that suits their business really well. Um I know that there’s some other the I would say like the smaller players in the US seem well like Arcee released their models earlier this year still chugging along. Poolside has started releasing some models.","They’ve gotten a few in the last few months and seem poised to release more models on top of that. So they’re really going Reflection is perpetually in the model coming soon camp and it really behooves them to get some models or some code or something out so that they can just start getting the developer flywheel going if they’re really committed to open source. It just takes a lot this it’s hard to get the models out. Like I talked to some people at Thinking Machines and it’s like kind of like oh that’s a lot of it’s a lot of work to actually do this I think. And um Nvidia chugging along. I think they’re at the stable player at this point. They’re keeping to release models. They’ll release more soon. They release a lot of data. I’m bullying them to try to get them to release Qwen style small models, which is like Gemma.","Gemma only has these like Qwen competitor models that are super popular. Um, the Gemma models are a little they’re all over the place in sizes or in architectures for the sizes and things like this, but the Gemma models are really really matching the Qwen models in terms of adoption. Um, I’m not sure they’re as easy to use for research, which could take a while. It could take multiple iterations. Like so much of language model research is now designed around small Qwen models and Qwen-based models that like it takes a while. Like people know how to use these models really well and if with the research results. So I hope Gemma keeps coming and can kind of compete in that niche. I don’t know any anyone that I missed here.","00:27:22 Florian Brand: No, I think both are the big players. Uh it’s, it is becoming broader. Uh in terms of model creators like last year, did we have any release aside from Gemma 3 and um GPT-OSS?","00:27:41 Nathan Lambert: was GPT-OSS 2 would go hard and obviously and obviously Nemotron as well. Um, oh, and I think Llama 4 at the start of the year, but uh, I don’t want that to be forgotten, but we are seeing like more players are are are now joining and turning out models at a really incredible rate.","00:27:59 Florian Brand: like Poolside has been releasing three or four models in the last two or three months. Uh and they seem to have figured out some way to turn out models pretty consistently. Um and that’s also something we are seeing on the open source side as well. we are talking about GLM like I think their iterations uh times for the model releases are now between 1 or 2 months with each new iteration becoming better and better which closely resembles what the closed labs are doing like we get a new GPT we get a new Claude every uh 6 weeks or so these days uh so in terms of having uh good enough pipeline uh to release stronger and stronger models they have to or the open source ecosystem has really figured it out or seemingly figured it out.","00:28:59 Nathan Lambert: Yeah, I agree. It’s it’s promising, but it is also so funny that like the US ecosystem started releasing some models and then then you have like Xi on the mic and these two models. It’s just like it’s so hard to catch up because it takes a lot of institutional expertise to train models that people actually use. And I think this is is what the American companies that are releasing models are now realizing is like these are not just benchmaxxed distilled IP theft models.","These are like genuinely good models that people are comparing to on their internal trading benchmarks and then like seeing how hard it is to beat them on measurable things. And I think that that is like I I’ve I’ve picked this sentiment up from a few people in the US trading models and it is just like there’s some I I think people should innovate on like size and fine-tunability and try to like use this potential market that is really close to home but also the pressures for every company is so high to release a model that you can claim as Frontier. I think investors expect that out of so many of these players that they’re kind of trying to do a a pretty hard thing and it’ll be interesting how the next year unfolds for the US China balance.","00:30:25 Florian Brand: Yeah, I think or in general I and a lot of other people have talked about the general ecosystem and that’s also something you’ve talked about at the very beginning. I think we are seeing more and more of a split between the capabilities of models that is good enough for a lot of tasks like uh for a lot of coding tasks the current frontier models both open and closed are good enough. um improvements feel less and less uh important here.","But if we look at the frontiers frontier, so finding new math proofs, finding uh new uh cures, finding new drugs, and inventing new things, that seems to be a whole different beast and probably will be dominated by the very frontier for quite a long time. The big question then becomes how much does that matter uh in terms of the addressable market and also how much of a focus will this be. I think, or my general base case is that we are seeing the frontier close down more and more. We have seen this with Mythos for cyber security GPT... or for biotech that those models won’t be accessible for everyone um and maybe not even external partners if we consider the reports that Anthropic is now spawning or or creating some internal labs to develop drugs."],"articleImages":[{"sourceUrl":"https://substackcdn.com/image/fetch/$s_!mkoP!,e_trim:10:white/e_trim:10:transparent/h_116,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F858a68f7-2e7e-4dd3-bed1-631b36801ce2_1651x357.png","alt":"Interconnects AI","afterParagraph":0,"url":"/media/articles/cmrw7dmtf00ukro8gtnl09yk6/aeec63fae96dffc7.png"},{"sourceUrl":"https://substackcdn.com/image/fetch/$s_!ZugE!,w_144,h_144,c_fill,f_auto,q_auto:good,fl_progressive:steep,g_auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fff7a51-ae70-45b2-a76e-db9f92faab51_1950x1950.png","alt":"Interconnects AI","afterParagraph":0,"url":"/media/articles/cmrw7dmtf00ukro8gtnl09yk6/176a2c507fd693f8.png"}],"mediaStatus":"ok","articleBodyZh":["Nathan 和 Florian 坐下来讨论开放模型的所有动向。在上周 Kimi K3 发布之后，似乎一切都在加速——美国与中国的地缘政治、开放模型与封闭模型的经济学、AI 前沿的安全性等等。章节：00:00 欢迎与背景 04:38 使用 / 体验 Kimi K3 08:53 GLM 5.2 的持续作用 12:47 中国模型为什么这么强？ 17:41 数据、环境以及中国实验室巡礼 19:47 中国提供商综述：Qwen、DeepSeek、MiniMax… 24:08 美国开放模型生态 30:25 前沿与接近前沿，以及反对禁令的网络安全案例 34:58 蒸馏与 Ben Thompson 辩论 44:12 预测与前沿等级列表 48:36 总结","在 Apple Podcasts 收听：https://podcasts.apple.com/us/podcast/interconnects-audio/id1719552353，Spotify：https://open.spotify.com/show/6XNzfJULeVxR7SneeesDUs，以及你获取播客的任何平台：https://www.interconnects.ai/podcast。其他 Interconnects 采访请访问：https://www.interconnects.ai/t/interviews。","分享：https://www.interconnects.ai/p/open-models-recap-more-on-kimi-k3?utm_source=substack&utm_medium=email&utm_content=share&action=share","更多教育性训练后视频，请见我正在整理的课程：https://rlhfbook.com/course","00:00:06 Nathan Lambert：好的，欢迎回到 Interconnects。我们正在进行季度开放模型回顾，主要是我们调侃或解释，当然不是调侃，为什么这么多蒸馏观点不好，并理解当前的状态。我记得 Kimi K3 是在上周四发布的。我认为在不久的将来我们会看到更多。这似乎几乎是不可避免的。周末，习近平发表了演讲，他直接承诺将开放与开源作为战略。这并不是一份详细的事务布局报告。","Qwen 宣布他们的下一个大模型将是开源权重，这将是一个重大变化。我觉得有很多内容可以讨论。我认为 Flo 你已经开始谈论性能差距和蒸馏的内容了，所以我们可能可以从这里开始，然后我也有一个小清单，我们可以随时浏览我写的博客，里面内容都非常细致。我觉得我们有无穷无尽的话题可以聊。所以请继续你的抱怨式讨论。","00:01:17 Florian Brand: 是的，我认为每次发布模型时，至少在每次开源模型发布时，最大的问题是它落后封闭前沿多少或者几个月。人们喜欢给这个问题一个明确的数字，这实际上很混乱，因为现在我们有这么多不同的基准提供者，也有许多不同的基准，每个网站——我自己也不例外——都会拿出自己喜欢的基准来显示当前模型或新发布的模型已经达到了前沿水平，而另一方则会拿出另一个基准来显示，哦，它实际上落后一年左右。这在很大程度上似乎取决于这个问题：开源模型落后多少个月。","00:02:26 Nathan Lambert: 是的。所以我的挑衅是，有些基准实际上与人们正在做的任务相当相关，比如智能代理编码和智能代理计算机使用任务，而有些基准则与长尾相关，我认为 Claude 和 GPT 在这方面非常有价值。但如果市场是针对 Claude Code 和 Codex 的话，而现在市场是软件工程，那么模型落后几个月可能就是一个非常非常大的问题。我怀疑这个模型会表现得不错，声明一下，模型权重还没发布，据说是在 7 月 27 日，而很多讨论将会依赖于权重已发布的假设。","但人们可以对这个模型进行后续训练，从而在很多特定的利基领域中非常可能匹配 Opus 和 GPT，我认为，关注这个，我的意思是我们在后训练开放模型行业中有不同的看法，但在让这些模型针对特定高价值任务进行微调方面，有大量的兴奋和进展。这在历史上是通过 Qwen 和 GLM 的混合方式完成的，而 GLM 5.2 确实加速了这一进程。我很好奇，第一个发布博客文章的人会说“我们在我们的任务上微调了 Kimi K3”，因为我打赌你可以获得很大的提升。我认为，即使你比我更多使用 Kimi K3，但我的直觉是，这将是一个有点粗糙的后训练过程，仅仅因为它的规模扩展很大，而这通常意味着仍然有很多性能可以被挖掘。不是。","00:04:03 Florian Brand: 是的。运行、运行、运行，尤其是后训练，这将非常困难，因为你需要一节点的 B300 来仅仅加载权重，这在规模上是疯狂的。所以可能需要一些时间，并且我听说要投入大量的工程工作才能真正使其处于可微调的状态。嗯，但人们，你想谈论使用模型的事情，你实际上注册了编码计划并使用了它。所以，把这个拿出来是很好的背景信息。","00:04:38 Nathan Lambert: 是的。所以我在发布后的第二天左右注册了 200 美元的计划，这是他们最大的一个，类似于其他所有计划，但他们还有我认为 40 美元和 100 美元的计划。嗯，但最大的计划有 100 万的上下文，我认为，或者至少感觉上在 API 请求方面也有一些优先权，因为很多人在线上说他们不断遇到 API 错误，而到目前为止，如果要我说的话，我的情况还算不错。嗯，就模型能力而言，除了前端很优秀之外，在某些方面它确实非常出色，而且表现优异。","嗯，即使是我对一些研究任务的期望，比如我在互联公司（Interconnects）进行的一些任务，我们现在有一年多关于开放模型的数据，我让前沿模型做一些有趣的分析，这是我们以前从未在自己分析并发表的工作中做过的。我让它们好吧，做一些新的，然后惊艳我，基本上就是这样。很多模型或前沿模型，或者基本上所有模型，都会盯着我们做过的东西，重新做数据分析部分，然后再做一些奇怪、深奥的部分。","嗯，Kimi K3 做了一些更有趣的事情，我明确告诉它去抓取 Reddit，然后它找到了我甚至没考虑过的一些子版块，然后发现例如 Reddit 的讨论是最近一两个月的，或者它们发现有趣的模型通常在下载量激增前一两个月就被关注，比如大家都关注 Qwen，之后人们就更多下载 Qwen 模型，这种分析是开创性的。但是相比其他前沿模型，Kimi 确实让我感到惊讶。","一个简单的问题是，你是否可以在大多数核心工作中使用它？比如说实验工作，你通常会做的很多都是用 Codex，我认为你是 Codex 用户而不是 Claude 用户，那么你觉得这个模型的适用范围与 Codex 的重合度是多少？","00:07:24 Florian Brand：嗯，这真的取决于我给它多少自由度。比如我目前看到 Kimi K3 的大问题是，我正在工作框架上进行研究，这个框架是在我工作的 Prime Intellect 公司进行的，主要发现是 Kimi 的代码简单很多，这使得它更易读，但它缺失了一些 Codex 或者说我们谈论的 56、55 尤其是 54 级别的东西。因此，我会说 Kimi K3 在这些任务上的水平大约是 54-55 级别。","但是如果我愿意，我可以阅读代码，然后说，好吧，这真的是好代码，然后我用 Codex 稍微检查一遍，它会找到所有这些小众的小案例，在这些地方它表现不算出色，但用于监督性运行或进行一些实验，它实际上是挺可用的。嗯，对于一些其他小众的事情，你可以让它直接运行。唯一的缺点是，这也因为 API 完全被用户淹没，而且它们的服务器在中国。实际墙钟时间明显明显比 GPT 高。但是我会说，如果我要在日常工作流程中推动它并使用它，我会更慢，但不会慢到让我说，“好吧，这完全不能用”。","00:08:53 Nathan Lambert：那和 GLM 5.2 相比如何？因为在我看来，GLM 5.2 仍然是一个正在展开的故事，比如我会在旧金山各处走动，人们会说，是的，我真的在我的某些代理式编码和/或工作流程中使用它。嗯，你觉得呢，我感觉你在使用 GLM 吗，还是","00:09:20 Florian Brand：是的。我也主要使用和曾经使用 GLM，因为我们有一个内部端点，它非常快；或者在那之前，我也使用过一个 API，大概每秒 200 或 300 个 token。嗯，如果你能在足够好的水平下快速完成很多任务，你就会用那个模型，而不是去 Codex 然后选择较低性能的模型，再选择合适的推理努力，再选择快速，我只是用 GLM 得到相同的结果，而且它还挺好的。它绝对在功能上达到了 Sonnet 级别，对于很多清理任务或者仅仅是琐碎工作的任务，它真的很有效。我会说，你可能可以用 Kimi K3 作为主要代理，GLM 作为子代理，完成很多工作。","00:10:24 Nathan Lambert: 关于Kimi的发布和模型规模，有一些非常不同的地方。我认为这些开放模型真正被优化并在各个推理提供商上可用可能还需要一段时间。比如GLM 5.2已经相当快，但首先，我们还没有得到权重；其次，我认为它的采纳推广速度不会像500B、700B的MoE那样快。我觉得那里会遇到更多问题，这是一个非常不同的局面，以前中国的模型会在完成强化学习训练后，在几小时到几天，甚至最多一周内发布开放权重的模型，然后生态系统几乎立即就知道如何使用它。","我认为在下一代开放权重模型上，基础设施的提升要大得多，我们必须考虑到这一点。封闭实验室在宣布模型之前会在幕后完成这些工作。所以这在某种程度上会操控时间差，人们实际上可能需要额外一个月才能在大规模工作流程中后训练并使用Kimi。而作为开放权重的粉丝，我们喜欢说，只有在封闭模型可用时，你才可以考虑时间差；但是现在在开放模型中也存在类似的动态，比如Kimi的API完全崩溃了，供应过多，需求过多，供应不足。所以这个模型并不会立即扩散，我只是在考虑它与性能时间差的关系。","00:11:54 Florian Brand: 这确实如此，但另一方面，过去几个月开放生态系统已经专业化了很多。在你最初的发布期间，它们都会带一些合作伙伴，这些合作伙伴提前拿到权重。他们会在几天甚至几周前就发布vLLM补丁，这与一年前完全不同，那时权重刚开始发布，模型制造商会说，好吧，你们自己摸索吧。所以我预计第一天的普遍可用性会相当不错，然后各种提供商开始竞速优化，以获得更高的速度，因为这带来很高的声望。","00:12:47 Nathan Lambert: 是的。好的，有两个方向可以讨论。为什么我们认为中国的模型能够达到这么好的水平？我写过关于这一点的内容，在我们的 Discord 上和 JSD 在 Epoch：https://epoch.ai/ 上有一个讨论，我觉得非常好。我在我的文章中有一个部分，我开始倾向于认为中国的实验室在资本利用方面更高效，你可以以一种使模型更好的方式，将资本转化为计算资源、数据和人才。我认为这非常重要，如果这确实是一种结构性优势，无论原因如何，我认为原因可能是人才更适合他们的教育体系去解决可以提升大语言模型（LLM）的相关问题。","也有可能是在中国，所有的计算资源和人才等成本都要低得多。不管是因为补贴，还是工资水平普遍较低，但在我们推进模型迭代的过程中，这都是一个非常重要的因素。如果下一代模型对 Anthropic 的成本是 100 亿美元，但对 Kimi 只需要 40 亿美元，这可能会产生非常大的影响，但原因尚不明确。例如，我想 Big Eagle（Kimi 的工程师）在我关于这个的推文下回复说，这有帮助，因为我们并不是在尝试推动前沿，只是在努力追赶，这实际上可能是一种心态问题，即中国实验室在设定目标时的思路，使得他们构建这些模型的成本要低得多。","但在过去一年里，我们问了很多类似“中方模型会不会掉队”的问题。我曾认为，由于训练的资本密集性，封闭模型和开放模型之间的差距会扩大，但看来情况正在朝相反方向发展，这很难解释。那么，你是否同意中国实验室的进展比我们预期的要更快一些？为什么会这样？","00:14:43 Florian Brand: 好吧，我实际上看了我们基于去年的回顾对今年的预测，我们基本上说过差距会保持在几个月之内。所以这个预测似乎大体上成立。幸运的是，我们没有给出一个具体的数字，无论是3个月、6个月还是9个月，所以我们在这方面是安全的。但我认为，像我们在中国与这些人交谈时的总体感觉，他们——研究人员本身——都是两三百人的团队，平均二十多岁，他们只想让一个模型非常好，他们似乎不做任何额外的实验。","他们似乎不做任何偏离这些事情的事情。就计算能力而言，这对我们来说是一个非常难回答的问题，尤其是现在这些中国芯片开始上线。我也认为芯片走私在过去六到九个月大幅增加，或者被走私的芯片开始上线。","00:15:58 Nathan Lambert: ‘走私’是规避出口限制的一个通用说法。如果这些芯片在马来西亚并且他们在使用它们，我也算作类似情况，我认为在过去六到九个月这大幅增加，这是部分原因，而且你说的，但我只是想指出，我确实认为他们拥有的计算能力比训练上一代模型时要多得多。","00:16:24 Florian Brand: 是的，比如说，为了提供一些背景信息，两周前我认为LongCat发布了他们的模型，他们声称——我们知道这很可能是真的——完全使用中国芯片训练。他们没有公开说明具体是哪种芯片，但人们推测是华为的Ascend芯片。随着国内生产的增加，它们可能主要用于训练，但对于推理也特别有用，而推理也是训练中非常重要的一部分。","所以他们可能在训练部分使用了某种混合的Nvidia和其他芯片，然后在推理部分使用的比例越来越大，而在这个阶段这真的很重要。所以我认为他们的整体计算量在增加，而且实际上他们没有很多用户。所以他们不需要像ChatGPT那样为10亿用户提供动力，也不需要像Anthropic那样为数百或数千家企业提供服务，因为他们没有那么多付费客户。","00:17:41 Nathan Lambert: 是的。我认为即便是那些付费客户，至少在企业端，也会占用公司时间和交流，当你在支持这些东西时。即使你不是做研究的，即使这不是你的工作，也会改变公司的关注点。如果SSI推出了一个好的模型，那将是对注意力分散问题的最终验证，不过那是另一个话题，我们可以稍后再谈。我认为数据和环境产业开始出现的迹象也有一些传闻。","你记得具体的例子吗？因为当我们在中国的时候，看到他们几乎不利用外部数据，确实有点令人震惊。所以就在我们四月旅行后的几个月，也就是七月，我们听到了一些关于中国新公司想要购买数据的消息。这也是一个有趣的时间线，反映了这种情况是如何变化的。","00:18:40 Florian Brand: 我会在他们实际告诉我们的内容上加上误差范围。","00:18:44 Nathan Lambert: 因为时间太接近，我不太确定。","00:18:49 Florian Brand: 是的，那可能，那可能是对的。不过这些事情很难确定。我会说，看起来购买外部数据正变得越来越重要。这将帮助开放模型赶上封闭模型，如果它们只是购买相同的数据，也许还能享受折扣，因为它们是在后期购买数据环境的。但这确实是一个因素。这个因素有多大，我们不知道。我们没有任何公开的洞察，我怀疑我们能从任何人那里获得这些洞察，所以这绝对是我们能够赶上或提高他们模型评分的原因之一。","00:19:47 Nathan Lambert: 好的，汇总一下其他中国模型提供商。我们谈到了Kimi，谈到了智谱/GLM。我认为很快会有更多非常优秀的GLM模型。他们可能称之为GLM 5.5。我们谈到了Qwen，他们最大的模型即将推出。我会说，Qwen最大的模型相较于他们的小模型在绝对性能排名上一直没有那么出色，这可能是专注策略的代价。我认为这与云公司有关。这几乎就像，如果你仔细看，几乎就像谷歌一样。","就像Qwen、阿里巴巴这里有这么多机会，并且通过这些小模型将开发者与阿里巴巴Qwen联系起来，对于他们的云来说是巨大的机会，我认为他们非常成功。但他们的大模型一直没有像他们的小模型那样出色。所以我不认为他们的模型会像Kimi K3或GLM 5.2那样具有突破性。我预计它会作为重大开放新闻在媒体上报道，就像中国的‘开放瓶’名称大热一样，但我认为它不会像一则新闻故事那样持续。DeepSeek，你可以插话，随你。","00:21:01 Florian Brand: 有趣的是，不知道你们有多少关注这一点，但他们有一个端点，你可以用来预览版本，他们每天都会更新这个端点，所以他们有非常快的迭代周期，因为我们在所有这些推特的基准测试中都有进展。所以很多这些 SVG 事情和 three.js 这些可视化生成任务，这个模型在过去几天里有很大提升。所以他们想出了一种快速反馈机制，其他公司也有。我们知道这个，Cursor 有很多博客讲述他们如何快速迭代。但他们似乎不断上传新的检查点并提供使用。","00:21:47 Nathan Lambert: 但是我同意。我猜它是在他们最终的 RL 运行中时间受控的，就像在他们的 RL 运行结束时仍在略微改进，他们只是勾选了完成框。","00:22:03 Nathan Lambert: 好的。Qwen DeepSeek V4 预览版本应该会发布。我觉得 DeepSeek V4 的情况是 flash 模型实际上更受欢迎，它们的体积更小，看起来对用户来说是一个真正的工作马。所以我认为这是值得关注的模型。我不指望 V4 Pro 会有戏剧性的突破。这和任何事情类似，比如小米如果很快发布一个新的 MiMo Pro 型号，我不指望它会有多大的创新，但它可能会是一个非常可靠的模型。只是很难知道。他们仍然是一个相当新的参与者。MiniMax，我认为是在玩不同的游戏。我不认为 MiniMax 在追求这种 Kimi/GLM 向 AGI 的突破感觉。","00:22:46 Florian Brand: 哦，我不同意。","00:22:49 Nathan Lambert: 你认为，MiniMax 还在追这个吗？","00:22:52 Florian Brand: 是的，我，我，我觉得他们确实看到了这种紧张，尤其是因为他们是一家上市公司，类似于 GLM，如果你看看股价表现，过去几天这些股票 RIP，这似乎有巨大差别。有趣的部分将是许可证，因为他们改变了很多许可证，越来越严格，如果在习…演讲后又改变主意，那会很有意思。","呃，看看 MiniMax 是否会回到完全开放的许可证，这将会很有趣。同样有趣的是看看 K3 将使用哪种许可证，因为他们说过会开源它，但我认为他们并没有在具体的许可证上做出任何承诺。","00:23:45 Nathan Lambert: 是的，这真的非常重要。我们拭目以待。嗯，Ling、美团、LongCat，有点类似，非常强大的模型，可能从中获得了很多内部价值，但并没有同样的开发者突破。嗯，所以大概是七、八个中国实验室，我可能遗漏了一些。我们也可以谈谈美国的实验室。顺便提一下，Gemini 3.6 闪电版发布了，看起来还可以。就是小幅提升，运行更快，少点唠叨，但其实没什么大不了的。我们会停止分享这个，嗯，这就是我们对 Gemini 的提及量。","但我确实认为值得稍微谈一下美国生态系统。我认为有一些新兴玩家。Thinking Machines 发布了他们的第一个模型，我和他们中的一些人聊过，他们非常支持探索如何用 Tinker 制作可微调模型，我认为这是一个我非常推荐给大多数开源模型开发者的研究领域。我认为如果你能在这里占据心智份额，将会获得大量采用，因为这更多是关于让模型能够真正为实际任务微调，而不是追求最佳指标。嗯，所以这是他们的 Inkling 模型，是一个一万亿参数的模型，评分还不错，但并不是前沿水平。","我觉得有点像 DeepSeek V4，他们计划发布一个规模更小的模型，总参数只有原来的大约四分之一，但性能非常好。如果 Inkling 小型预览版在几周内推出，我确实认为那将会是一个被广泛使用的模型。它的大小适合自动化任务和领域特定任务，但可能不会像 Kimi 和 GLM 5.2 那样成为通用型智能体，但我认为这非常适合他们的业务。我知道还有一些其他的，美国一些小型玩家，比如 Arcee，今年早些时候发布了他们的模型，而且依然在运作中。Poolside 也开始发布一些模型。","过去几个月他们发布了一些模型，看起来还准备在此基础上发布更多。因此，他们真的非常反复地处于“模型即将推出”状态，如果他们真的致力于开源，推出一些模型或代码对他们来说非常有利，这样可以开始启动开发者的循环。不过，这件事非常耗费精力，实际上很难推出模型。我和 Thinking Machines 的一些人聊过，他们说这确实很复杂，做这件事要付出很多努力。Nvidia 也在持续推进，我觉得他们现在是一个稳健的参与者，他们会持续发布模型，很快会发布更多，并且发布了大量数据。我正努力推动他们尝试发布 Qwen 风格的小型模型，比如 Gemma。","Gemma 只有一些像 Qwen 的竞争模型，非常受欢迎。嗯，Gemma 的模型在不同尺寸和架构上有些混乱，但在采用率方面，Gemma 模型确实和 Qwen 模型非常接近。我不确定它们在研究上使用是否那么方便，这可能需要一些时间，可能需要多次迭代。现在很多语言模型研究都是围绕小型 Qwen 模型和 Qwen 基础模型设计的，所以需要一些时间。人们已经非常熟悉如何使用这些模型，并结合研究结果。所以我希望 Gemma 能持续发展，并在这个小众领域竞争。我不确定我是否遗漏了其他人。","00:27:22 Florian Brand：不，我认为两者都是大玩家。呃，它确实在变得更广泛。呃，就模型创建者而言，比如去年，除了 Gemma 3 和 GPT-OSS，我们还有其他版本发布吗？","00:27:41 Nathan Lambert：GPT-OSS 2 进展很快，显然，Nemotron 也是如此。呃，哦，我认为年初还有 Llama 4，但我不想让它被遗忘，但我们确实看到，更多玩家现在加入，并以非常惊人的速度推出模型。","00:27:59 Florian Brand：像 Poolside 在过去两三个月里已经发布了三到四个模型。呃，他们似乎找到了一种比较稳定地推出模型的方法。呃，我们在开源方面也看到类似的情况。我们在谈论 GLM，像我认为他们的模型发布迭代时间现在在 1 到 2 个月之间，每个新迭代都变得越来越好，这非常类似封闭实验室的做法，就像我们现在大约每六周会有一个新的 GPT 或 Claude 一样，呃，所以在有足够好的流程去发布越来越强大的模型方面，开源生态系统已经真正找到了方法，或者看起来已经找到了方法。","00:28:59 Nathan Lambert：是的，我同意。这很有前景，但也很有趣，美国生态系统开始发布一些模型，然后你就看到像 Xi 在麦克风上，这两个模型。就像真的很难赶上，因为训练人们实际会使用的模型需要大量的机构专业知识。我认为这是美国 releasing 模型的公司现在才意识到的：这些不仅仅是经过基准优化的、蒸馏后的、侵犯知识产权的模型。","这些模型是真正不错的模型，人们会拿它们和自己的内部交易基准进行比较，然后看看在可衡量的指标上击败它们有多难。我觉得我从美国一些交易模型的人那里感受到这种情绪，我觉得大家应该在规模和微调能力上进行创新，尝试利用这个潜在市场，这个市场离我们很近，但同时每家公司都有很大的压力去发布一个可以称作前沿的模型。我认为投资者对许多这些玩家都有这样的期望，他们实际上正在尝试做一件相当困难的事情，接下来一年美中平衡如何发展会很有意思。","00:30:25 Florian Brand: 是的，我想我和很多其他人都讨论过整体生态系统，你在一开始也提到过这个。我认为我们越来越明显地看到模型能力上的分化，对于很多任务来说已经足够好了，比如很多编程任务，目前的前沿模型，无论是开放还是封闭的，已经足够好了。在这些方面的改进感觉越来越不那么重要。","但如果我们看真正的前沿，例如寻找新的数学证明、发现新疗法、开发新药物和发明新事物，这似乎是完全不同的挑战，并且可能在相当长的一段时间里被非常前沿的模型主导。接下来的大问题是，这在可触及的市场方面有多重要，以及这会成为多大的关注重点。我认为，或者说我的一般基本判断是，我们看到前沿模型越来越封闭。我们在网络安全方面的Mythos GPT看到过……或者在生物技术方面看到过，这些模型不会对所有人开放，甚至考虑到有关Anthropic现在正在孵化或创建一些内部实验室来开发药物的报道，可能甚至不对外部合作伙伴开放。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Nathan Lambert 与 Florian Brand 在季度播客中讨论 Kimi K3 发布后的开放模型生态，议题包括中美模型产业、开源与闭源经济性、前沿安全、性能差距、知识蒸馏及后训练价值。","background":"材料称，Kimi K3 于播客录制前一周发布，Qwen 宣布下一代大模型将采用开放权重策略。两位讨论者还梳理了多家中国模型提供方及美国开放模型生态，但未提供 Qwen 3.8 或 WAIC 演讲的明确事实。","viewpoint":"Aioga 判断，用单一的“落后几个月”衡量开放模型与闭源前沿的差距并不稳健。讨论者指出，不同机构和参与者会选取不同基准，编码、计算机操作及长尾任务也可能呈现不同结论。","implications":"值得关注的是，开放权重策略可能继续影响模型竞争与部署选择，但材料未给出成本、用户规模或市场份额数据。对企业而言，具体任务表现可能比笼统的前沿排名更有参考价值。","nextStep":"后续应核验 Kimi K3 权重是否按讨论中提到的预期时间发布，并关注 Qwen 下一代模型的正式公告、许可证和评测结果；同时应基于多类任务基准审视性能差距与蒸馏争议。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-07-23T03:31:08.431Z","sourceHash":"89660e651ef1d7d3","review":{"approved":true,"groundedness":94,"clarity":89,"duplicationRisk":18,"blockingIssues":[],"notes":["“Aioga 判断”中的主体未在来源材料中出现；如其代表编辑方，建议改为“编辑判断”或直接表述为分析性观点，以免读者误认为 Aioga 是播客讨论者。","“Kimi K3 于播客录制前一周发布”基本对应原文的“released last week”，但改为“于录制前一周内发布”或“于此前一周发布”会更贴近原意。","关于企业应重视具体任务表现的表述属于合理的编辑推论，当前使用“可能”“更有参考价值”等措辞，未冒充来源中的确定事实。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","low-source-overlap","no-html","independent-ai-review"]}},"tags":["技巧观点","Nathan Lambert：Interconnects（RSS）"],"translations":{"zh-CN":{"title":"开源模型季度盘点：Kimi K3、Qwen 3.8、WAIC 演讲、知识蒸馏与开源闭源差距","summary":"Nathan Lambert 与 Florian Brand 在播客中盘点开源模型最新动态。Kimi K3 发布后，中美 AI 地缘政治、开源与闭源模型的经济性、前沿安全等问题加速演进。Qwen 宣布下一代大模型将开源权重，中国厂商在开源策略上持续加码。讨论还涉及开源模型与闭源前沿的性能差距、知识蒸馏争议，以及后训练对特定任务的价值。","category":"技巧观点","source":"interconnects.ai","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"开源模型季度盘点：Kimi K3、Qwen 3.8、WAIC 演讲、知识蒸馏与开源闭源差距 - Aioga AI资讯","description":"Nathan Lambert 与 Florian Brand 在播客中盘点开源模型最新动态。Kimi K3 发布后，中美 AI 地缘政治、开源与闭源模型的经济性、前沿安全等问题加速演进。Qwen 宣布下一代大模型将开源权重，中国厂商在开源策略上持续加码。讨论还涉及开源模型与闭源前沿的性能差距、知识蒸馏争议，以及后训练对特定任务的价值。","url":"https://www.aioga.com/news/cmrw7dmtf00ukro8gtnl09yk6/"},"en":{"title":"Open Source Model Quarterly Review: Kimi K3, Qwen 3.8, WAIC Talks, Knowledge Distillation and the Gap Between Open Source and Closed Source","summary":"Nathan Lambert and Florian Brand review the latest developments of open-source models on their podcast. After the release of Kimi K3, issues such as AI geopolitics between China and the U.S., the economics of open-source and closed-source models, and cutting-edge security have evolved rapidly. Qwen announced that the next generation of large models will have open-source weights, and Chinese companies continue to increase their efforts in open-source strategies. The discussion also touches on the performance gap between open-source models and cutting-edge closed-source models, controversies over knowledge distillation, and the value of post-training for specific tasks.","category":"Insights","source":"Nathan Lambert：Interconnects（RSS）","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"Open Source Model Quarterly Review: Kimi K3, Qwen 3.8, WAIC Talks, Knowledge Distillation and the Gap Between Open Source and Closed Source - Aioga AI News","description":"Nathan Lambert and Florian Brand review the latest developments of open-source models on their podcast. After the release of Kimi K3, issues such as AI geopolitics between China an...","url":"https://www.aioga.com/en/news/cmrw7dmtf00ukro8gtnl09yk6/","contentTranslated":true,"sourceHash":"690b7c77a9fb0d2a","translatedAt":"2026-07-22T17:25:14.143Z"},"ja":{"title":"オープンソースモデル四半期レビュー：Kimi K3、Qwen 3.8、WAIC 講演、知識蒸留とオープンソース・クローズドソースのギャップ","summary":"Nathan Lambert と Florian Brand はポッドキャストでオープンソースモデルの最新動向を振り返りました。Kimi K3 のリリース後、中国と米国の AI 地政学、オープンソースとクローズドモデルの経済性、最先端の安全性などの問題が加速して進化しています。Qwen は次世代大規模モデルの重みをオープンソース化すると発表し、中国の企業はオープンソース戦略を継続的に強化しています。議論はまた、オープンソースモデルとクローズド最先端モデルの性能差、知識蒸留の論争、そして特定タスクに対する事後トレーニングの価値にも及びました。","category":"ヒントと視点","source":"Nathan Lambert：Interconnects（RSS）","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"オープンソースモデル四半期レビュー：Kimi K3、Qwen 3.8、WAIC 講演、知識蒸留とオープンソース・クローズドソースのギャップ - Aioga AIニュース","description":"Nathan Lambert と Florian Brand はポッドキャストでオープンソースモデルの最新動向を振り返りました。Kimi K3 のリリース後、中国と米国の AI 地政学、オープンソースとクローズドモデルの経済性、最先端の安全性などの問題が加速して進化しています。Qwen は次世代大規模モデルの重みをオープンソース化すると発表し、中国の企業はオ...","url":"https://www.aioga.com/ja/news/cmrw7dmtf00ukro8gtnl09yk6/","contentTranslated":true,"sourceHash":"690b7c77a9fb0d2a","translatedAt":"2026-07-22T17:25:34.176Z"},"ko":{"title":"오픈소스 모델 분기별 점검: Kimi K3, Qwen 3.8, WAIC 강연, 지식 증류 및 오픈소스와 클로즈드소스의 차이","summary":"Nathan Lambert와 Florian Brand는 팟캐스트에서 오픈소스 모델의 최신 동향을 정리했습니다. Kimi K3가 출시된 이후, 중미 AI 지정학, 오픈소스와 폐쇄형 모델의 경제성, 최첨단 보안 등 문제가 급속히 진화하고 있습니다. Qwen은 차세대 대형 모델의 가중치를 오픈소스로 공개할 것이라고 발표했으며, 중국 기업들은 오픈소스 전략에 지속적으로 투자하고 있습니다. 논의는 또한 오픈소스 모델과 폐쇄형 최첨단 모델의 성능 차이, 지식 증류 논쟁, 특정 과제를 위한 사후 학습의 가치에도 포함되었습니다.","category":"인사이트","source":"Nathan Lambert：Interconnects（RSS）","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"오픈소스 모델 분기별 점검: Kimi K3, Qwen 3.8, WAIC 강연, 지식 증류 및 오픈소스와 클로즈드소스의 차이 - Aioga AI 뉴스","description":"Nathan Lambert와 Florian Brand는 팟캐스트에서 오픈소스 모델의 최신 동향을 정리했습니다. Kimi K3가 출시된 이후, 중미 AI 지정학, 오픈소스와 폐쇄형 모델의 경제성, 최첨단 보안 등 문제가 급속히 진화하고 있습니다. Qwen은 차세대 대형 모델의 가중치를 오픈소스로 공개할 것이라고 발표했으며,...","url":"https://www.aioga.com/ko/news/cmrw7dmtf00ukro8gtnl09yk6/","contentTranslated":true,"sourceHash":"690b7c77a9fb0d2a","translatedAt":"2026-07-22T17:26:18.916Z"},"es":{"title":"Resumen trimestral de modelos de código abierto: Kimi K3, Qwen 3.8, conferencia WAIC, destilación de conocimiento y la brecha entre código abierto y cerrado","summary":"Nathan Lambert y Florian Brand repasan las últimas novedades de modelos de código abierto en su podcast. Después del lanzamiento de Kimi K3, las cuestiones de geopolítica de IA entre China y Estados Unidos, la economía de los modelos de código abierto y cerrado, y la seguridad de vanguardia han evolucionado rápidamente. Qwen ha anunciado que la próxima generación de grandes modelos tendrá pesos de código abierto, y los fabricantes chinos continúan aumentando su apuesta en la estrategia de código abierto. La discusión también abarca la diferencia de rendimiento entre modelos de código abierto y de vanguardia cerrados, la controversia sobre la destilación de conocimiento, y el valor del posentrenamiento para tareas específicas.","category":"Ideas","source":"Nathan Lambert：Interconnects（RSS）","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"Resumen trimestral de modelos de código abierto: Kimi K3, Qwen 3.8, conferencia WAIC, destilación de conocimiento y la brecha entre código abierto y cerrado - Aioga Noticias de IA","description":"Nathan Lambert y Florian Brand repasan las últimas novedades de modelos de código abierto en su podcast. Después del lanzamiento de Kimi K3, las cuestiones de geopolítica de IA ent...","url":"https://www.aioga.com/es/news/cmrw7dmtf00ukro8gtnl09yk6/","contentTranslated":true,"sourceHash":"690b7c77a9fb0d2a","translatedAt":"2026-07-22T17:26:15.637Z"},"fr":{"title":"Bilan trimestriel des modèles open source : Kimi K3, Qwen 3.8, conférence WAIC, distillation des connaissances et écart entre open source et propriétaire","summary":"Nathan Lambert et Florian Brand passent en revue les dernières actualités des modèles open source dans leur podcast. Après la sortie de Kimi K3, les questions de géopolitique de l'IA entre la Chine et les États-Unis, l'économie des modèles open source et fermés, ainsi que la sécurité de pointe, ont évolué rapidement. Qwen a annoncé que la prochaine génération de grands modèles rendra les poids open source, et les entreprises chinoises continuent de renforcer leur stratégie open source. La discussion porte également sur l'écart de performance entre les modèles open source et fermés de pointe, les controverses autour de la distillation des connaissances, ainsi que la valeur de l'après-formation pour des tâches spécifiques.","category":"Analyses","source":"Nathan Lambert：Interconnects（RSS）","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"Bilan trimestriel des modèles open source : Kimi K3, Qwen 3.8, conférence WAIC, distillation des connaissances et écart entre open source et propriétaire - Aioga Actualités IA","description":"Nathan Lambert et Florian Brand passent en revue les dernières actualités des modèles open source dans leur podcast. Après la sortie de Kimi K3, les questions de géopolitique de l'...","url":"https://www.aioga.com/fr/news/cmrw7dmtf00ukro8gtnl09yk6/","contentTranslated":true,"sourceHash":"690b7c77a9fb0d2a","translatedAt":"2026-07-22T17:27:08.479Z"},"de":{"title":"Quartalsübersicht der Open-Source-Modelle: Kimi K3, Qwen 3.8, WAIC-Vorträge, Wissensdistillation und Unterschiede zwischen Open Source und Closed Source","summary":"Nathan Lambert und Florian Brand besprechen in einem Podcast die neuesten Entwicklungen bei Open-Source-Modellen. Nach der Veröffentlichung von Kimi K3 beschleunigen sich Fragen zu KI-Geopolitik zwischen China und den USA, zur Wirtschaftlichkeit von offenen und geschlossenen Modellen sowie zur Spitzenforschungssicherheit. Qwen kündigte an, dass das nächste Generationmodell offene Gewichte haben wird, und chinesische Unternehmen verstärken weiterhin ihre Open-Source-Strategien. Die Diskussion umfasst auch die Leistungslücken zwischen Open-Source-Modellen und den führenden Closed-Source-Modellen, die Debatte über Wissensdistillation sowie den Wert von Nachtrainingsverfahren für bestimmte Aufgaben.","category":"技巧观点","source":"Nathan Lambert：Interconnects（RSS）","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"Quartalsübersicht der Open-Source-Modelle: Kimi K3, Qwen 3.8, WAIC-Vorträge, Wissensdistillation und Unterschiede zwischen Open Source und Closed Source - Aioga KI-News","description":"Nathan Lambert und Florian Brand besprechen in einem Podcast die neuesten Entwicklungen bei Open-Source-Modellen. Nach der Veröffentlichung von Kimi K3 beschleunigen sich Fragen zu...","url":"https://www.aioga.com/de/news/cmrw7dmtf00ukro8gtnl09yk6/","contentTranslated":true,"sourceHash":"690b7c77a9fb0d2a","translatedAt":"2026-07-22T17:27:00.487Z"},"pt-BR":{"title":"Panorama trimestral de modelos de código aberto: Kimi K3, Qwen 3.8, palestra WAIC, destilação de conhecimento e diferenças entre código aberto e fechado","summary":"Nathan Lambert e Florian Brand comentam as últimas novidades dos modelos de código aberto em um podcast. Após o lançamento do Kimi K3, questões como a geopolítica de IA entre China e EUA, a economia de modelos abertos e fechados, e a segurança de ponta avançaram rapidamente. A Qwen anunciou que a próxima geração de grandes modelos terá pesos de código aberto, e as empresas chinesas continuam a aumentar suas estratégias de código aberto. A discussão também aborda a diferença de desempenho entre modelos de código aberto e de ponta fechados, controvérsias sobre distilação de conhecimento e o valor do pós-treinamento para tarefas específicas.","category":"技巧观点","source":"Nathan Lambert：Interconnects（RSS）","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"Panorama trimestral de modelos de código aberto: Kimi K3, Qwen 3.8, palestra WAIC, destilação de conhecimento e diferenças entre código aberto e fechado - Aioga Notícias de IA","description":"Nathan Lambert e Florian Brand comentam as últimas novidades dos modelos de código aberto em um podcast. Após o lançamento do Kimi K3, questões como a geopolítica de IA entre China...","url":"https://www.aioga.com/pt-BR/news/cmrw7dmtf00ukro8gtnl09yk6/","contentTranslated":true,"sourceHash":"690b7c77a9fb0d2a","translatedAt":"2026-07-22T17:27:50.340Z"},"ru":{"title":"Ежеквартальный обзор открытых моделей: Kimi K3, Qwen 3.8, выступления на WAIC, дистилляция знаний и разрыв между открытыми и закрытыми моделями","summary":"Nathan Lambert и Florian Brand обсудили в подкасте последние события в области открытых моделей. После выпуска Kimi K3 вопросы геополитики AI между Китаем и США, экономичности открытых и закрытых моделей, а также передовой безопасности стали развиваться быстрее. Qwen объявила, что следующая генерация больших моделей будет иметь открытые веса, китайские компании продолжают усиливать свою стратегию в области открытых решений. Обсуждение также затронуло разрыв в производительности между открытыми и закрытыми передовыми моделями, споры вокруг дистилляции знаний и ценность дополнительного обучения для конкретных задач.","category":"技巧观点","source":"Nathan Lambert：Interconnects（RSS）","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"Ежеквартальный обзор открытых моделей: Kimi K3, Qwen 3.8, выступления на WAIC, дистилляция знаний и разрыв между открытыми и закрытыми моделями - Aioga Новости ИИ","description":"Nathan Lambert и Florian Brand обсудили в подкасте последние события в области открытых моделей. После выпуска Kimi K3 вопросы геополитики AI между Китаем и США, экономичности откр...","url":"https://www.aioga.com/ru/news/cmrw7dmtf00ukro8gtnl09yk6/","contentTranslated":true,"sourceHash":"690b7c77a9fb0d2a","translatedAt":"2026-07-22T17:27:56.106Z"},"ar":{"title":"استعراض ربع سنوي للنماذج مفتوحة المصدر: Kimi K3، Qwen 3.8، محاضرة WAIC، تقطير المعرفة والفجوة بين المصادر المفتوحة والمغلقة","summary":"ناثان لامبرت وفلوريان براند يستعرضان آخر التطورات في نماذج المصادر المفتوحة في البودكاست. بعد إصدار Kimi K3، تسارعت القضايا المتعلقة بالجغرافيا السياسية للذكاء الاصطناعي بين الصين والولايات المتحدة، واقتصادية النماذج المفتوحة والمغلقة، والأمان المتقدم. أعلنت Qwen أن الجيل القادم من النماذج الكبيرة سيكون مفتوح المصدر، فيما تواصل الشركات الصينية تعزيز استراتيجيتها في المصادر المفتوحة. كما يشمل النقاش الفجوة في الأداء بين النماذج المفتوحة ونهج النماذج المتقدمة المغلقة، والخلاف حول تقطير المعرفة، وقيمة التدريب اللاحق لمهام محددة.","category":"技巧观点","source":"Nathan Lambert：Interconnects（RSS）","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"استعراض ربع سنوي للنماذج مفتوحة المصدر: Kimi K3، Qwen 3.8، محاضرة WAIC، تقطير المعرفة والفجوة بين المصادر المفتوحة والمغلقة - Aioga أخبار الذكاء الاصطناعي","description":"ناثان لامبرت وفلوريان براند يستعرضان آخر التطورات في نماذج المصادر المفتوحة في البودكاست. بعد إصدار Kimi K3، تسارعت القضايا المتعلقة بالجغرافيا السياسية للذكاء الاصطناعي بين الصين...","url":"https://www.aioga.com/ar/news/cmrw7dmtf00ukro8gtnl09yk6/","contentTranslated":true,"sourceHash":"690b7c77a9fb0d2a","translatedAt":"2026-07-22T17:28:43.219Z"},"hi":{"title":"ओपन-सोर्स मॉडल तिमाही समीक्षा: Kimi K3, Qwen 3.8, WAIC भाषण, ज्ञान आसवन और ओपन-सोर्स व क्लोज़-सोर्स का अंतर","summary":"Nathan Lambert और Florian Brand ने पॉडकास्ट में ओपन-सोर्स मॉडल की नवीनतम प्रगति पर चर्चा की। Kimi K3 के रिलीज़ के बाद, चीन और अमेरिका के एआई भू-राजनीति, ओपन और क्लोज्ड सोर्स मॉडल की आर्थिकता, और उन्नत सुरक्षा जैसी समस्याएँ तेजी से विकसित हो रही हैं। Qwen ने घोषणा की कि अगली पीढ़ी के बड़े मॉडल के वज़न ओपन-सोर्स होंगे, और चीनी कंपनियां ओपन-सोर्स रणनीति में लगातार बढ़ोतरी कर रही हैं। चर्चा में ओपन-सोर्स मॉडल और क्लोज्ड-सोर्स उन्नत संस्करणों के प्रदर्शन के अंतराल, ज्ञान आसवन विवाद, और विशेष कार्यों के लिए पोस्ट-ट्रेनिंग के मूल्य पर भी बात हुई।","category":"技巧观点","source":"Nathan Lambert：Interconnects（RSS）","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"ओपन-सोर्स मॉडल तिमाही समीक्षा: Kimi K3, Qwen 3.8, WAIC भाषण, ज्ञान आसवन और ओपन-सोर्स व क्लोज़-सोर्स का अंतर - Aioga AI समाचार","description":"Nathan Lambert और Florian Brand ने पॉडकास्ट में ओपन-सोर्स मॉडल की नवीनतम प्रगति पर चर्चा की। Kimi K3 के रिलीज़ के बाद, चीन और अमेरिका के एआई भू-राजनीति, ओपन और क्लोज्ड सोर्स मॉडल क...","url":"https://www.aioga.com/hi/news/cmrw7dmtf00ukro8gtnl09yk6/","contentTranslated":true,"sourceHash":"690b7c77a9fb0d2a","translatedAt":"2026-07-22T17:28:53.531Z"},"it":{"title":"Riepilogo trimestrale dei modelli open source: Kimi K3, Qwen 3.8, conferenza WAIC, distillazione della conoscenza e divario tra open source e closed source","summary":"Nathan Lambert e Florian Brand hanno discusso degli ultimi sviluppi dei modelli open source in un podcast. Dopo il rilascio di Kimi K3, la geopolitica dell'AI tra Cina e Stati Uniti, l'economicità dei modelli open source e closed source, e le questioni di sicurezza all'avanguardia si stanno evolvendo rapidamente. Qwen ha annunciato che la prossima generazione di grandi modelli avrà pesi open source, e le aziende cinesi stanno continuando ad intensificare la loro strategia open source. La discussione ha anche toccato il divario di prestazioni tra modelli open source e closed source all'avanguardia, le controversie sulla distillazione della conoscenza e il valore del post-addestramento per compiti specifici.","category":"技巧观点","source":"Nathan Lambert：Interconnects（RSS）","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"Riepilogo trimestrale dei modelli open source: Kimi K3, Qwen 3.8, conferenza WAIC, distillazione della conoscenza e divario tra open source e closed source - Aioga Notizie IA","description":"Nathan Lambert e Florian Brand hanno discusso degli ultimi sviluppi dei modelli open source in un podcast. Dopo il rilascio di Kimi K3, la geopolitica dell'AI tra Cina e Stati Unit...","url":"https://www.aioga.com/it/news/cmrw7dmtf00ukro8gtnl09yk6/","contentTranslated":true,"sourceHash":"690b7c77a9fb0d2a","translatedAt":"2026-07-22T17:29:48.843Z"},"nl":{"title":"Open source model kwartaaloverzicht: Kimi K3, Qwen 3.8, WAIC-presentaties, kennisdistillatie en het verschil tussen open source en closed source","summary":"Nathan Lambert en Florian Brand bespreken in hun podcast de nieuwste ontwikkelingen van open source-modellen. Na de lancering van Kimi K3 versnellen kwesties zoals de AI-geopolitiek tussen China en de VS, de economische aspecten van open en gesloten modellen, en geavanceerde veiligheidsproblemen. Qwen kondigde aan dat het volgende generatie grote model open source-gewichten zal hebben, en Chinese bedrijven verhogen voortdurend hun inzet op het gebied van open source-strategieën. De discussie gaat ook over de prestatiediversiteit tussen open source-modellen en toonaangevende gesloten modellen, de controverse rond kennisdistillatie, en de waarde van natraining voor specifieke taken.","category":"技巧观点","source":"Nathan Lambert：Interconnects（RSS）","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"Open source model kwartaaloverzicht: Kimi K3, Qwen 3.8, WAIC-presentaties, kennisdistillatie en het verschil tussen open source en closed source - Aioga AI-nieuws","description":"Nathan Lambert en Florian Brand bespreken in hun podcast de nieuwste ontwikkelingen van open source-modellen. Na de lancering van Kimi K3 versnellen kwesties zoals de AI-geopolitie...","url":"https://www.aioga.com/nl/news/cmrw7dmtf00ukro8gtnl09yk6/","contentTranslated":true,"sourceHash":"690b7c77a9fb0d2a","translatedAt":"2026-07-22T17:29:38.268Z"},"tr":{"title":"Açık Kaynak Model Çeyrek Değerlendirmesi: Kimi K3, Qwen 3.8, WAIC Sunumu, Bilgi Damıtma ve Açık-Kapalı Kaynak Farkı","summary":"Nathan Lambert ve Florian Brand podcast'te açık kaynak modellerin en son gelişmelerini gözden geçiriyor. Kimi K3 piyasaya sürüldükten sonra, Çin ve ABD AI jeopolitiği, açık ve kapalı kaynak modellerin ekonomikliği, ileri düzey güvenlik gibi konular hızla evriliyor. Qwen, bir sonraki nesil büyük modelin açık kaynak ağırlıklarının paylaşılacağını açıkladı ve Çinli üreticiler açık kaynak stratejilerini güçlendirmeye devam ediyor. Tartışmalar ayrıca açık kaynak modeller ile kapalı kaynak teknolojiler arasındaki performans farkını, bilgi damıtma tartışmalarını ve belirli görevler için sonradan eğitimin değerini de kapsıyor.","category":"技巧观点","source":"Nathan Lambert：Interconnects（RSS）","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"Açık Kaynak Model Çeyrek Değerlendirmesi: Kimi K3, Qwen 3.8, WAIC Sunumu, Bilgi Damıtma ve Açık-Kapalı Kaynak Farkı - Aioga AI Haberleri","description":"Nathan Lambert ve Florian Brand podcast'te açık kaynak modellerin en son gelişmelerini gözden geçiriyor. Kimi K3 piyasaya sürüldükten sonra, Çin ve ABD AI jeopolitiği, açık ve kapa...","url":"https://www.aioga.com/tr/news/cmrw7dmtf00ukro8gtnl09yk6/","contentTranslated":true,"sourceHash":"690b7c77a9fb0d2a","translatedAt":"2026-07-22T17:30:44.784Z"},"vi":{"title":"Tổng kết quý các mô hình mã nguồn mở: Kimi K3, Qwen 3.8, bài thuyết trình WAIC, chưng cất kiến thức và khoảng cách giữa mã nguồn mở và đóng","summary":"Nathan Lambert và Florian Brand đã tổng kết các động thái mới nhất của các mô hình mã nguồn mở trong podcast. Sau khi Kimi K3 được phát hành, các vấn đề như địa chính trị AI giữa Trung Quốc và Mỹ, kinh tế của mô hình mở và đóng, cũng như an ninh tiên tiến, đang có tốc độ tiến triển nhanh hơn. Qwen công bố rằng thế hệ mô hình lớn tiếp theo sẽ mở nguồn trọng số, các nhà sản xuất Trung Quốc tiếp tục tăng cường chiến lược mã nguồn mở. Cuộc thảo luận cũng liên quan tới khoảng cách hiệu suất giữa mô hình mở và tiên tiến đóng, tranh cãi về chưng cất kiến thức, cũng như giá trị của việc huấn luyện bổ sung cho các nhiệm vụ cụ thể.","category":"技巧观点","source":"Nathan Lambert：Interconnects（RSS）","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"Tổng kết quý các mô hình mã nguồn mở: Kimi K3, Qwen 3.8, bài thuyết trình WAIC, chưng cất kiến thức và khoảng cách giữa mã nguồn mở và đóng - Tin tức AI Aioga","description":"Nathan Lambert và Florian Brand đã tổng kết các động thái mới nhất của các mô hình mã nguồn mở trong podcast. Sau khi Kimi K3 được phát hành, các vấn đề như địa chính trị AI giữa T...","url":"https://www.aioga.com/vi/news/cmrw7dmtf00ukro8gtnl09yk6/","contentTranslated":true,"sourceHash":"690b7c77a9fb0d2a","translatedAt":"2026-07-22T17:30:41.465Z"},"id":{"title":"Tinjauan Kuartalan Model Sumber Terbuka: Kimi K3, Qwen 3.8, Presentasi WAIC, Distilasi Pengetahuan dan Kesenjangan antara Sumber Terbuka dan Sumber Tertutup","summary":"Nathan Lambert dan Florian Brand membahas perkembangan terbaru dari model open source dalam podcast mereka. Setelah peluncuran Kimi K3, isu geopolitik AI antara China dan Amerika Serikat, ekonomi model open source dan closed source, serta keamanan mutakhir berkembang dengan cepat. Qwen mengumumkan bahwa model besar generasi berikutnya akan memiliki bobot open source, dan perusahaan-perusahaan China terus memperkuat strategi open source mereka. Diskusi juga mencakup perbedaan kinerja antara model open source dan cutting-edge closed source, kontroversi dalam knowledge distillation, serta nilai fine-tuning untuk tugas-tugas tertentu.","category":"技巧观点","source":"Nathan Lambert：Interconnects（RSS）","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"Tinjauan Kuartalan Model Sumber Terbuka: Kimi K3, Qwen 3.8, Presentasi WAIC, Distilasi Pengetahuan dan Kesenjangan antara Sumber Terbuka dan Sumber Tertutup - Berita AI Aioga","description":"Nathan Lambert dan Florian Brand membahas perkembangan terbaru dari model open source dalam podcast mereka. Setelah peluncuran Kimi K3, isu geopolitik AI antara China dan Amerika S...","url":"https://www.aioga.com/id/news/cmrw7dmtf00ukro8gtnl09yk6/","contentTranslated":true,"sourceHash":"690b7c77a9fb0d2a","translatedAt":"2026-07-22T17:31:30.130Z"},"th":{"title":"การสรุปผลโมเดลโอเพนซอร์สรายไตรมาส: Kimi K3, Qwen 3.8, การบรรยาย WAIC, การกลั่นความรู้ และช่องว่างระหว่างโอเพนซอร์สกับปิดซอร์ส","summary":"Nathan Lambert และ Florian Brand พูดคุยเกี่ยวกับความเคลื่อนไหวล่าสุดของโมเดลโอเพ่นซอร์สในพอดแคสต์ หลังจากที่ Kimi K3 ถูกปล่อยออกมา เรื่องราวเกี่ยวกับภูมิรัฐศาสตร์ AI ระหว่างจีนและสหรัฐ, เศรษฐกิจของโมเดลโอเพ่นซอร์สและโมเดลปิด, รวมถึงความปลอดภัยขั้นสูง ได้พัฒนาอย่างรวดเร็ว Qwen ประกาศว่าโมเดลขนาดใหญ่รุ่นถัดไปจะเปิดน้ำหนักของโมเดลให้เป็นโอเพ่นซอร์ส บริษัทจีนกำลังเพิ่มกลยุทธ์โอเพ่นซอร์สอย่างต่อเนื่อง การสนทนายังครอบคลุมถึงช่องว่างของประสิทธิภาพระหว่างโมเดลโอเพ่นซอร์สกับโมเดลปิดชั้นนำ การถกเถียงเรื่องการกลั่นกรองความรู้ และคุณค่าของการฝึกฝนหลังเพื่อภารกิจเฉพาะ","category":"技巧观点","source":"Nathan Lambert：Interconnects（RSS）","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"การสรุปผลโมเดลโอเพนซอร์สรายไตรมาส: Kimi K3, Qwen 3.8, การบรรยาย WAIC, การกลั่นความรู้ และช่องว่างระหว่างโอเพนซอร์สกับปิดซอร์ส - ข่าว AI Aioga","description":"Nathan Lambert และ Florian Brand พูดคุยเกี่ยวกับความเคลื่อนไหวล่าสุดของโมเดลโอเพ่นซอร์สในพอดแคสต์ หลังจากที่ Kimi K3 ถูกปล่อยออกมา เรื่องราวเกี่ยวกับภูมิรัฐศาสตร์ AI ระหว่างจีนและส...","url":"https://www.aioga.com/th/news/cmrw7dmtf00ukro8gtnl09yk6/","contentTranslated":true,"sourceHash":"690b7c77a9fb0d2a","translatedAt":"2026-07-22T17:31:37.260Z"},"pl":{"title":"Przegląd kwartalny modeli open source: Kimi K3, Qwen 3.8, wykład WAIC, destylacja wiedzy i różnice między open source a zamkniętym oprogramowaniem","summary":"Nathan Lambert i Florian Brand omawiają najnowsze aktualności dotyczące modeli open source w podcaście. Po wydaniu Kimi K3, kwestie geopolityki AI między Chinami a USA, ekonomiczności modeli open source i zamkniętych oraz bezpieczeństwa na granicy przyspieszyły w rozwoju. Qwen ogłosił, że następna generacja dużych modeli będzie miała udostępnione wagi, a chińscy producenci nadal zwiększają zaangażowanie w strategię open source. Dyskusja obejmowała również różnice wydajności między modelami open source a wiodącymi modelami zamkniętymi, kontrowersje związane z distillacją wiedzy oraz wartość treningu po wstępnym szkoleniu dla konkretnych zadań.","category":"技巧观点","source":"Nathan Lambert：Interconnects（RSS）","aggregationSource":"Nathan Lambert：Interconnects（RSS）","pageTitle":"Przegląd kwartalny modeli open source: Kimi K3, Qwen 3.8, wykład WAIC, destylacja wiedzy i różnice między open source a zamkniętym oprogramowaniem - Aioga Wiadomości AI","description":"Nathan Lambert i Florian Brand omawiają najnowsze aktualności dotyczące modeli open source w podcaście. Po wydaniu Kimi K3, kwestie geopolityki AI między Chinami a USA, ekonomiczno...","url":"https://www.aioga.com/pl/news/cmrw7dmtf00ukro8gtnl09yk6/","contentTranslated":true,"sourceHash":"690b7c77a9fb0d2a","translatedAt":"2026-07-22T17:32:23.427Z"}}}}