{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-09-28T06:03:00.468Z","headline":"Dwarkesh Patel 研究：预训练进步主要来自数据改进","description":"Dwarkesh Patel 发布实验分析，在最高 1e19 FLOPs 的算力预算下训练 2019 至 2025 年各年度代表性模型配方与数据语料，发现数据改进带来 12.0x 算力效率提升，模型改进为 3.7x，数据贡献约为模型的 3.24 倍。","url":"https://www.aioga.com/news/cmtsx0wm4041jrob5liwjypsq/","mainEntityOfPage":"https://www.aioga.com/news/cmtsx0wm4041jrob5liwjypsq/","datePublished":"2026-09-08T16:10:16.000Z","dateModified":"2026-09-08T16:10:16.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.dwarkesh.com/p/pretraining-progress-is-mostly-data","https://aihot.news/items/cmtsx0wm4041jrob5liwjypsq"],"canonicalUrl":"https://www.aioga.com/news/cmtsx0wm4041jrob5liwjypsq/","directAnswer":{"@type":"Answer","text":"研究称，在最高1e19 FLOPs算力预算下，2019至2025年代表性模型配方与数据语料的组合实验显示，数据改进带来12.0倍算力效率提升，模型改进为3.7倍。","url":"https://www.aioga.com/news/cmtsx0wm4041jrob5liwjypsq/","dateCreated":"2026-09-08T16:10:16.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"dwarkesh.com source article","url":"https://www.dwarkesh.com/p/pretraining-progress-is-mostly-data","datePublished":"2026-09-08T16:10:16.000Z","provider":{"@type":"Organization","name":"dwarkesh.com","url":"https://www.dwarkesh.com/p/pretraining-progress-is-mostly-data"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.news/items/cmtsx0wm4041jrob5liwjypsq","datePublished":"2026-09-08T16:10:16.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.news/items/cmtsx0wm4041jrob5liwjypsq"}}],"aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","originalPublisher":{"name":"dwarkesh.com","url":"https://www.dwarkesh.com/p/pretraining-progress-is-mostly-data"},"geoDeepAnswer":null,"article":{"id":"cmtsx0wm4041jrob5liwjypsq","slug":"cmtsx0wm4041jrob5liwjypsq","url":"https://www.aioga.com/news/cmtsx0wm4041jrob5liwjypsq/","title":"Dwarkesh Patel 研究：预训练进步主要来自数据改进","title_en":"","summary":"Dwarkesh Patel 发布实验分析，在最高 1e19 FLOPs 的算力预算下训练 2019 至 2025 年各年度代表性模型配方与数据语料，发现数据改进带来 12.0x 算力效率提升，模型改进为 3.7x，数据贡献约为模型的 3.24 倍。","source":"Dwarkesh Patel：Podcast & Blog（RSS）","sourceUrl":"https://www.dwarkesh.com/p/pretraining-progress-is-mostly-data","aiHotUrl":"https://aihot.news/items/cmtsx0wm4041jrob5liwjypsq","publishedAt":"2026-09-08T16:10:16.000Z","category":"行业动态","score":72,"selected":true,"articleBody":["How much of the rapid progress in AI ：https://epoch.ai/gradient-updates/the-least-understood-driver-of-ai-progress that we’ve seen over the last few years 1：#footnote-1 has come from data versus model improvements? The answer has big implications for the economics of frontier labs and the pace of future progress.","We investigate this question at a relatively small scale, and for pretraining specifically, from 2019 to 2025. During each of those years, a new open model recipe was published which codified that year’s publicly known algorithmic tweaks (for example, improvements in architecture, optimizer, initializations, learning rate schedule, hyperparams, etc). And during each of those years, there was also a new public data corpus (produced by broader scrapes and new curation/extraction/filtering techniques).","We train combinations of these year-representative model recipes and data corpuses across different scales of training compute (up to 1e19 FLOPs) 2：#footnote-2 .","Obviously, we can’t compare these different models by their cross-entropy loss against a fixed dataset, since we’re varying the datasets they’re trained on. So instead we evaluate these models on end capabilities as measured by the OLMES ：https://github.com/allenai/olmes eval (which aggregates 10 different relatively easy benchmarks, mostly multiple choice QA). Unfortunately, evaluating end capability rather than pretraining loss adds some noise to our results, as you’ll see in the graphs below, though we try to get cleaner bounds by running multiple seeds.","We find that from 2019 to 2025, 3.24x more compute efficiency gains have come from data improvements rather than model improvements (12.0x for data and 3.7x for models), at the 1e19 FLOPs compute budget 3：#footnote-3 .","Here is a grid which shows how much better a model we train does on the end capability we're testing it on, relative to the 2019 data + architecture baseline, at 3.16e18 FLOPs 4：#footnote-4 .","We find that the gains from data and model improvements are mostly independent and don’t interact (i.e. realizing the gains from some model improvement doesn’t require a specific training datapile, or vice versa). 88% of the variance in the OLMES score can be explained by additive effects of the model and data improvements (using a linear model).","For context, let’s briefly summarize what changed on both the data and the model side from 2019 to 2025.","On the model side, we went from GPT-2 ：http://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf to OLMo-2 ：https://arxiv.org/abs/2501.00656 , including key innovations in optimizers, positional encodings, normalization, activation functions, initializations, and more 5：#footnote-5 .","On the data side, we started with OpenWebText ：http://huggingface.co/datasets/Skylion007/openwebtext in 2019, which contained just web pages linked from Reddit with enough upvotes and then deduplicated and filtered, and thus amounted to only ~9B tokens (this was mostly what GPT-2 was trained on). By 2025, open source data corpuses like UltraFineWeb ：http://huggingface.co/datasets/openbmb/Ultra-FineWeb not only are far larger (by using scrapes of the whole Internet), but also use much more sophisticated filtering (for example, by training a classifier to predict what data will empirically improve model performance).","A naive interpretation of our result is that most of the AI progress from 2019-2024 (the era of pretraining) was just better data engineering (extraction, curation, etc.), and that all the model work during that period was much less important.","But this is probably the wrong way to think about the value of model improvements. Their main contribution was not necessarily compute efficiency - that is, achieving the same performance with fewer FLOPs. Rather, it was making larger amounts of compute usable in the first place. As the number of parameters, context lengths, run duration, and clusters scale up, all kinds of things are prone to breaking (gradients explode or vanish, memory and bandwidth run out, training becomes infeasibly slow). Much of model research has consisted of removing or pushing back these constraints to scaling ：https://www.beren.io/2026-08-23-Architecture-Research-as-Addressing-Constraints-to-Scaling/ . Many of the most important innovations such as MoEs, sparse attention variants, stability innovations (norm placements, initializations, etc.) and system / kernel-level optimizations like FlashAttention ：https://arxiv.org/abs/2205.14135 fall into this category.","The data improvements we investigated here might matter less for larger models. Small models (like the ones we trained) see significant gains from data quality improvements, because they don’t have that much capacity, and so you have to be really careful about what you stuff into them. Whereas big models have so much excess capacity that maybe you just want to throw in as much stuff as you can, even if it’s mostly garbage, and the magic of stochastic gradient descent will separate out the signal from the noise. If you choose to filter aggressively, you’ll have to do dozens of epochs, which empirically gives worse performance ：https://arxiv.org/pdf/2605.19407 than just having a lower average quality but larger dataset. In fact, aggressive data curation is even more harmful once you take into account that frontier models are up to 100x overtrained relative to Chinchilla optimal, in order to minimize the inference compute used for RL and for deployment.","An analogy might be the difference between a sailboat and a container ship - the container ship doesn’t necessarily go faster, but it can lug thousands of tons of cargo (analogous to hundreds of trillions of tokens of pretraining data), and won’t be toppled by choppy waters (analogous to training stably across hundreds of thousands of GPUs).","Now that we have more capacious and sturdy container ships, we don’t have to fret about exactly what we load on board - we can just fill them up with everything that’s even remotely and plausibly useful. Whereas for the tiny flimsy sailboats of 2019, you’d have to be incredibly careful about only carrying the most valuable cargo.","But to the extent that the nature of pretraining progress is simply loading more cargo into this ship, are we running out of cargo? This is a question about the data wall and about how well synthetic data has helped us leap over it. Synthetic data is obviously being widely used at the labs, and we have not at all investigated whether it can effectively expand a data corpus without hurting model performance. If the gains are limited, then the main driver of pretraining progress will stall, because we’re not generating more internet, and you can only curate a fixed set of data by so much. To be clear, we have no active reason to think this. But given how important data seems to be in driving pretraining progress, this seems like a crucial question to investigate.","Ryan Greenblatt noted ：https://www.dwarkesh.com/p/ryan-greenblatt that many of the historical improvements in pretraining data corpuses look like the kind of progress that automated researchers would be able to just test empirically - for example, run ablations trained on different data and see how the model performs. So it’s totally compatible with our results that the data progress which has propelled pretraining since 2019 might speed up a lot if and when we automate AI R&D.","We want to clarify that whether pretraining progress in isolation will speed up or slow down is not really the most important question for overall AI progress, because so many of the gains over the last two years have come from RL.","These are some directions of future research that we think would be really cool, and important questions to answer:","You could run this experiment at larger scales to see whether the data or model improvements are more dependent on scale (and thus far more impactful at the frontier)","What is the marginal value of novel high-quality data for both pre and post-training, as measured by end capabilities?","We want to know broadly how effectively synthetic data works. One concrete question to investigate is this: if you’ve got a small corpus of high quality data, how much better is it to magnify it via synthetic data generation relative to just training on it for multiple epochs ：https://arxiv.org/pdf/2605.19407 ?","You could figure out the implied value of data through lab spending on data brokers, environment producers, etc., relative to their spending on compute and researchers.","We wanted to investigate what role data has played in driving AI progress. There are lots of other ways one could probe this question, and some may be more clever and informative than ours. And even our experiment was done at an extremely small scale. We definitely think it’s plausible that there is something we missed - we’re eager to hear how others would research this question, and ideally to also see their results!","Thanks especially to Charlie O’Neill ：https://x.com/oneill_c?lang=en for many helpful discussions.","We pre-train these model recipes from scratch on these different data corpuses, at varying compute budgets, with multiple independent seeds 6：#footnote-6 . Our compute budgets are: 1e17, 3.16e17, 1e18, 3.16e18 and 1e19 FLOPs. The compute accounting convention is to use nominal compute C = 6ND (N = number of non-embedding parameters, D = tokens of data).","At each compute budget, we vary the number of parameters (and hence number of tokens trained on), to determine the compute-optimal mix for each training recipe x corpus combination. We use held-out loss on the corpus to determine this compute-optimal point. We can then obtain compute scaling curves of downstream performance of each combination, from which we can finally extract our compute multipliers.","We enforce a shared tokenizer and context length across every run: GPT-2 BPE (tiktoken, 50,257 vocab) and T=2048, batch = 262,144 tokens.","The end capabilities of our training runs are highly dependent on hyperparameters. Obviously, there is no way to sweep over all possible sets of hyperparams (hyperparam tuning is a fine art indeed)! We try to control for this as much as possible, and we consider peak learning rate as the main hyperparameter of significance.","Some algorithm vintages do provide specifications of what peak learning rate should be tuned to (as a function of other relevant variables such as model size, data budget, batch size, etc.). These serve as good priors for what we think the optimal learning rate is.","We first sweep learning rates at 5 anchor points - 3 different model sizes and 2 different D/N ratios. We determine the optimal learning rate of these anchor points, and fit an optimal learning rate parametric form","For all the model recipes except OLMo-2, we fit a common exponent a and b , and a model-specific lr₀. For OLMo-2, we use the prescribed optimal learning rate according to the model recipe. The reason we do this for OLMo-2 is that Ai2 published small-model ladders as part of the recipe which specified optimal hyperparameters at the scale we are investigating. We also verify, at the compute-optimal point for 3.16e18 FLOPs, that our production learning rates are at or near optimal.","Explaining some anomalies in our graph","We observe generally increasing compute efficiency across time for both the model and data axes as expected. Some outliers that we observed:","NeoX performs worse than GPT-2 at 1e19 (although it does better across the 1e17 to 3.16e18 range). This might arise from noise in the OLMES evaluation. We also note that on held-out pretraining loss on the FineWeb-Edu corpus, NeoX performs better than GPT-2.","The Piles seems to do much worse than OpenWebText. This is not surprising since the Pile’s main improvement was data corpus diversity over filtering. It has a curated 22-source mixture including PubMed and arXiv papers, GitHub code, legal opinions, patents, and parliamentary proceedings. The amount of cross-domain transfer to OLMES (which is English web-prose MCQ) might be minimal for many of these tokens, thus resulting in lower compute efficiency. We note that by virtue of its larger size, we expect that the Pile should eventually be better than (the really small) OpenWebText at larger scales.","It is also worth noting that the compute multipliers for NeoX and the Pile are obtained by extrapolation, which introduces further potential error.","How compute multipliers were calculated, as well as their error bars","Every point on the compute scaling curves is computed from multiple independently seeded training runs. The error bars there are the standard deviation of the OLMES eval over those seeds.","Consider some given reference level of performance at some compute level for our reference model or data corpus.","We then calculate the compute multiplier by finding the left-most point of the compute scaling curve of our candidate model or corpus that first attains that reference level of performance. The ratio of the compute required by the reference to the compute required by our candidate is the candidate’s compute multiplier","The error bars on the compute multipliers are obtained from a parametric bootstrap of the entire estimation pipeline, and are 1 standard deviation intervals","We do want to highlight that we expect the actual uncertainty in the compute multipliers of the model recipes to be higher than indicated by our error bars. This is because of additional uncertainty introduced by the limited extent of hyperparameter tuning we did, and end capabilities or held-out loss is probably quite sensitive to the exact choice of peak learning rate / batch size / etc.","It is also important to note that there are many reasons why our ablations do not necessarily capture the full scope of compute efficiency gains. Indeed, from 2019 to 2025, we observe year-over-year compute efficiency gains (CEG) of 1.24x [1.19, 1.29] on the model side and 1.51x [1.45, 1.57] on the data side. Measured jointly, we observe a 1.57x YoY CEG [1.49, 1.65] 7：#footnote-7 . This is indeed much lower than Anson Ho et al.’s mean estimate of 3x YoY ：https://arxiv.org/abs/2403.05812 , for the following reasons:","Many of the gains might be scale dependent ：https://arxiv.org/pdf/2511.21622 or might be especially important at longer context, and we are operating at scales too small to realize many of the gains.","For example, OLMo-2’s layer and QK norms, parallel attention + MLP block in NeoX","Inference efficiency optimizations (such as LLama-3’s GQA ：https://arxiv.org/abs/2305.13245 , which is a KV cache optimization) do not show up as compute multipliers in our study. We are also not investigating tokenizer improvements.","The compute multipliers we obtain are pretty sensitive to our choice of model recipe or data corpus for each year. We have chosen what we believe to be representative model recipes or data corpuses. But by no means do we exhaustively conclude that these are the best of each year.","We are looking at compute multipliers with respect to the OLMES benchmark (which combines 10 different relatively easy task types) rather than compute multipliers in getting to some perplexity metric. We would also have very different looking numbers if we were looking at other benchmarks (say, coding- or problem-solving-specific ones), which would probably reward very different methods of data engineering.","We also want to note that we have not investigated other data-side improvements, such as collecting more high-quality data from new sources, human expert generated data, synthetic data generation methods, etc. Most of the corpuses we have investigated are curations (subsets) of the same Common Crawl, rather than expanding the available set of data. This is clearly consumption of a finite stock - there is only so far we can push this lever.","Independence of gains from model recipe and data corpus","Here is the investigation that we did to determine how independent the gains from model recipe and data corpus are. We looked at the grid of OLMES scores at 3.16e18 FLOPs. A linear regression of OLMES score = mean + model effect + data effect gives an R squared of 0.88, which means 88% of the variance in the OLMES score can be explained by additive effects of the model and data improvements, with only ~12% of the variance accounted for by interaction or higher order terms, and eval noise. This hints that complex model-data interactions (where exploiting some model improvement is contingent on some specific data engineering, or vice versa) are relatively minor.","Anson Ho et al. ：https://arxiv.org/abs/2403.05812 estimated software efficiency improvements (in pretraining) of 3x per year (95% CI: 1.5x to 64x). As Ho mentioned in this blog ：https://epochai.substack.com/p/the-least-understood-driver-of-ai , “most software progress might actually be due to data quality improvements” and “from scaling up just a small handful of scale-dependent algorithmic changes”.","We use the C = 6ND nominal convention for accounting for compute.","The 2019 model recipe was GPT-2, and 2025 model recipe was OLMo-2. The 2019 data corpus was OpenWebText, and the 2025 data corpus was UltraFineWeb.","Our implementation of GPT3 encountered some training instabilities (gradient spikes) on the Pile.","These include: Optimizer improvements, warmup + decay schedules, RoPE replacing learning absolute positions, RMSNorm + SwiGLU gated MLPs, Norm reordering, QK-norm, Z-loss regularization and cleaner inits.","For the compute-scaling plots we use at least 3 seeds each. For the 7x7 grid of combinations of model recipes and data corpus at the 3.16e18 budget, we only used 1 seed each.","The 1.57x YoY multiplier is computed using the joint improvement from 2019 model and corpus to 2025 model and corpus, and not the product of the 1.24x model side improvement and 1.51x data side improvement.","The metric for comparison should have been generative perplexity or an analogue, which is the common way to compare pretrains."],"articleImages":[],"mediaStatus":"none","articleBodyZh":["在过去几年里，我们看到的人工智能快速进展（：https://epoch.ai/gradient-updates/the-least-understood-driver-of-ai-progress ）中，有多少来自数据的改进，有多少来自模型的改进？答案对前沿实验室的经济学以及未来进展的速度有重大影响。","我们在相对较小的规模上，特别是针对预训练，从2019年到2025年调查了这个问题。在每一年，都会发布一个新的公开模型配方，总结了当年公开已知的算法调整（例如，架构、优化器、初始化、学习率调度、超参数等的改进）。在这些年份，每年也都有新的公开数据集（通过更广泛的抓取以及新的策划/提取/过滤技术生成）。","我们在不同的训练计算规模（最多达1e19 FLOPs）下训练这些年份代表性的模型配方和数据集的组合。2：#footnote-2","显然，我们不能通过交叉熵损失在固定数据集上比较这些不同模型，因为我们正在改变它们训练的数据集。因此，我们改为通过OLMES衡量模型的最终能力（https://github.com/allenai/olmes eval）（该评估汇总了10个不同的较简单基准，大多为多项选择问答）。不幸的是，评估最终能力而不是预训练损失会给我们的结果带来一些噪声，正如您在下面的图表中看到的，尽管我们尝试通过运行多个随机种子来获得更清晰的界限。","我们发现，从2019年到2025年，在1e19 FLOPs的计算预算下，数据改进带来的计算效率提升比模型改进高3.24倍（数据提升12.0倍，模型提升3.7倍）。3：#footnote-3","这里有一个表格显示了我们训练的模型在我们测试的最终能力上，相对于2019年的数据+架构基线，在3.16e18 FLOPs下表现得有多好。4：#footnote-4","我们发现，数据和模型改进带来的提升大多是独立的，并且不互相影响（即，实现某些模型改进的收益并不需要特定的训练数据集，反之亦然）。使用线性模型，88%的OLMES分数方差可以由模型和数据改进的加性效应解释。","为了提供背景，让我们简要总结一下从2019年到2025年，数据端和模型端发生的变化。","在模型方面，我们从 GPT-2：http://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf 发展到 OLMo-2：https://arxiv.org/abs/2501.00656，包括优化器、位置编码、归一化、激活函数、初始化等方面的关键创新 5：#footnote-5。","在数据方面，我们从 2019 年的 OpenWebText 开始：http://huggingface.co/datasets/Skylion007/openwebtext，其中只包含来自 Reddit 并获得足够点赞的网页，然后进行了去重和过滤，因此总共大约有~90 亿个标记（这大部分就是 GPT-2 的训练数据）。到 2025 年，像 UltraFineWeb 的开源数据语料库：http://huggingface.co/datasets/openbmb/Ultra-FineWeb 不仅规模大得多（通过抓取整个互联网的数据），而且使用了更复杂的过滤方法（例如，通过训练分类器预测哪些数据在实践中会提高模型性能）。","对我们结果的一种天真的解读是，从2019到2024年（预训练时代）的AI大部分进步只是更好的数据工程（提取、整理等），而在此期间的模型工作则不那么重要。","但这种看法可能是错误的。模型改进的主要贡献不一定是计算效率——即用更少的FLOPs实现相同性能。而是首先使更大量的计算变得可用。随着参数数量、上下文长度、运行时间和集群规模的增加，各种问题很容易出现（梯度爆炸或消失、内存和带宽不足、训练变得不可行地缓慢）。大量模型研究的内容就是消除或推迟这些扩展的约束：https://www.beren.io/2026-08-23-Architecture-Research-as-Addressing-Constraints-to-Scaling/。许多最重要的创新，如MoEs、稀疏注意力变体、稳定性创新（归一化位置、初始化等）以及系统/内核级优化，如FlashAttention：https://arxiv.org/abs/2205.14135，都属于这一类别。","我们在这里研究的数据改进对于更大的模型可能影响不大。小型模型（比如我们训练的那些）从数据质量的提升中获得显著收益，因为它们的容量有限，因此你必须非常小心地选择输入内容。而大型模型有如此多的多余容量，以至于你可能只想尽可能多地输入数据，即使大部分是垃圾，通过随机梯度下降的魔力也能从噪声中分离出信号。如果你选择进行严格筛选，你将不得不进行几十轮训练，这在经验上会比拥有平均质量较低但数据集更大的情况表现更差：https://arxiv.org/pdf/2605.19407。事实上，一旦考虑到前沿模型在训练中相对于Chinchilla最优模型被过度训练了多达100倍，以最小化RL和部署时的推理计算，过度的数据策划甚至更有害。","一个类比可能是帆船和货轮的区别——货轮不一定跑得更快，但它可以运载成千上万吨的货物（类似于数百万亿的预训练数据token），并且不会被波涛汹涌的海面推翻（类似于在成千上万GPU上稳定训练）。","现在我们有了更大容量、更坚固的货轮，我们不必担心到底装什么——我们可以把所有哪怕是稍微可能有用的东西都装上去。而对于2019年的那些娇小脆弱的帆船，你必须非常谨慎，只携带最有价值的货物。","但在预训练进展的本质仅仅是向这艘船装载更多货物的程度上，我们是否快要没有货物可装了？这是关于数据壁垒的问题，也是关于合成数据在帮助我们跨越这一壁垒方面的效果问题。显然，各实验室正在广泛使用合成数据，而我们完全没有研究过它是否能够在不影响模型性能的前提下有效扩展数据语料库。如果增益有限，那么预训练进展的主要驱动力将停滞，因为我们无法生成更多的互联网数据，而且你对固定数据集合的整理能力也是有限的。需要明确的是，我们没有实际理由认为会这样。但考虑到数据在推动预训练进展中似乎非常重要，这似乎是一个至关重要的问题，值得研究。","Ryan Greenblatt 指出：https://www.dwarkesh.com/p/ryan-greenblatt 历史上预训练数据语料库的许多改进看起来就像自动化研究者通过实证测试就能完成的进展——例如，运行不同数据训练的消融实验，看看模型表现如何。因此，我们的结果完全兼容这样的观点，即自2019年以来推动预训练的数据进展如果在人工智能研发自动化时，可能会加速很多。","我们想澄清的是，单独预训练进展的加速或减慢并不是总体人工智能进展中最重要的问题，因为过去两年的许多突破都来自强化学习（RL）。","以下是我们认为非常有趣且值得回答的一些未来研究方向：","你可以在更大规模上进行这个实验，以观察数据改进或模型改进对规模的依赖性（因此在前沿领域影响更大）","对于预训练和后训练来说，高质量新数据的边际价值是多少，以最终能力为衡量标准？","我们想大致了解合成数据的效果有多好。一个具体的研究问题是：如果你拥有一个小型高质量数据语料库，通过合成数据生成来放大它相比于仅对它训练多个周期效果有多大：https://arxiv.org/pdf/2605.19407 ?","你可以通过实验室在数据经纪人、环境生成者等方面的支出，相对于他们在计算资源和研究人员上的支出来推测数据的隐含价值。","我们想要研究数据在推动人工智能进步中所起的作用。还有很多其他方法可以探究这个问题，有些方法可能比我们的方法更巧妙、更有信息量。即使我们的实验是在极小规模下进行的，我们也认为我们可能漏掉了一些东西——我们很想听听别人会如何研究这个问题，并理想情况下也希望看到他们的结果！","特别感谢 Charlie O’Neill：https://x.com/oneill_c?lang=en 提供了许多有益的讨论。","我们从零开始在这些不同的数据语料库上预训练这些模型配方，在不同的计算预算下，使用多组独立的随机种子6：#footnote-6。我们的计算预算是：1e17、3.16e17、1e18、3.16e18 和 1e19 FLOPs。计算核算的惯例是使用名义计算 C = 6ND（N = 非嵌入参数数，D = 数据的 token 数）。","在每个计算预算下，我们改变参数数量（从而改变训练的 token 数量），以确定每个训练配方 × 语料组合的计算最优比例。我们使用语料库上的保留损失来确定这个计算最优点。然后我们可以获得每个组合的下游性能的计算扩展曲线，从中最终提取我们的计算乘数。","我们在每次运行中都强制使用统一的分词器和上下文长度：GPT-2 BPE（tiktoken，词汇量 50,257）和 T=2048，批量 = 262,144 tokens。","我们的训练运行的最终能力高度依赖于超参数。显然，不可能扫描所有可能的超参数组合（超参数调优确实是一门精细的艺术）！我们尽量控制这个因素，并认为峰值学习率是最重要的超参数。","一些算法版本确实提供了峰值学习率的调优规格（作为其他相关变量的函数，例如模型大小、数据预算、批量大小等）。这些作为我们认为的最优学习率的良好先验。","我们首先在5个锚点上扫描学习率——3种不同的模型规模和2种不同的D/N比。我们确定这些锚点的最佳学习率，并拟合一个最佳学习率的参数形式。","对于除OLMo-2之外的所有模型配方，我们拟合一个通用的指数a和b，以及一个特定模型的lr₀。对于OLMo-2，我们根据模型配方使用规定的最佳学习率。我们之所以对OLMo-2这样做，是因为Ai2发布了小型模型梯度作为配方的一部分，这些梯度在我们正在研究的规模上指定了最佳超参数。我们还在3.16e18 FLOPs的计算最优点验证，发现我们的生产学习率在最佳范围内或接近最佳。","解释我们图表中的一些异常情况","正如预期的那样，我们观察到模型和数据轴上的计算效率总体上随时间增加。我们观察到的一些异常值：","NeoX在1e19时的表现比GPT-2差（尽管在1e17到3.16e18范围内它表现更好）。这可能源于OLMES评估中的噪声。我们还注意到，在FineWeb-Edu语料库的保留预训练损失上，NeoX表现优于GPT-2。","Pile的表现似乎比OpenWebText差得多。这并不令人意外，因为Pile的主要改进在于数据语料库的多样性，而非过滤。它包含一个由22个来源精心策划的混合，包括PubMed和arXiv论文、GitHub代码、法律意见、专利和议会记录。对于这些许多标记而言，转移到OLMES（即英文网页散文MCQ）的跨领域效果可能很小，因此导致计算效率较低。我们注意到，由于其规模较大，我们预计Pile在更大规模下最终应会比（真正小的）OpenWebText表现更好。","还值得注意的是，NeoX和Pile的计算乘数是通过外推获得的，这引入了进一步的潜在误差。","计算乘数以及它们的误差范围的方法","计算缩放曲线上的每个点都是基于多个独立初始化训练运行得出的。图上的误差范围是这些初始化的OLMES评估的标准差。","考虑某个参考模型或数据语料库在某个计算水平下的一定参考性能水平。","然后，我们通过找到候选模型或语料库的计算缩放曲线中首次达到该参考性能水平的最左端点，来计算计算倍增器。参考所需的计算量与候选所需计算量的比率就是候选的计算倍增器。","计算倍增器的误差条是通过对整个估计流程进行参数自助法（parametric bootstrap）得到的，表示的是1个标准差区间。","我们确实想强调的是，我们预计模型方案的计算倍增器的实际不确定性要高于误差条所显示的。这是因为我们超参数调优的范围有限所引入的额外不确定性，以及最终的能力或保留损失可能对峰值学习率、批量大小等具体选择非常敏感。","同样重要的是要注意，存在许多原因说明我们的消融实验未必完全捕捉到计算效率增益的全部范围。事实上，从2019年到2025年，我们观察到模型端的年均计算效率增益(CEG)为1.24x [1.19, 1.29]，数据端为1.51x [1.45, 1.57]。联合测量时，我们观察到1.57x的年均CEG [1.49, 1.65] 7：#footnote-7。这确实远低于Anson Ho等人的平均估计值3x年均CEG：https://arxiv.org/abs/2403.05812，原因如下：","许多收益可能依赖于规模：https://arxiv.org/pdf/2511.21622，或者在较长的上下文中尤其重要，而我们目前操作的规模太小，无法实现许多收益。","例如，OLMo-2的层和QK范数，NeoX中的并行注意力+MLP模块","推理效率优化（例如 LLama-3 的 GQA：https://arxiv.org/abs/2305.13245，这是一个 KV 缓存优化）在我们的研究中并未显示为计算倍增器。我们也没有研究分词器的改进。","我们得到的计算倍增器对每年的模型方案或数据语料库的选择非常敏感。我们选择了我们认为具有代表性的模型方案或数据语料库。但我们绝不意味着穷尽地得出这些就是每年的最佳选择。","我们正在观察相对于OLMES基准（该基准结合了10种不同的相对简单的任务类型）的计算乘数，而不是在达到某个困惑度指标时的计算乘数。如果我们观察其他基准（比如专门针对编码或问题解决的基准），这些数字也会完全不同，这可能会奖励非常不同的数据工程方法。","我们还想指出，我们没有研究其他数据方面的改进，例如从新来源收集更多高质量数据、人工专家生成的数据、合成数据生成方法等。我们调查的大多数语料库都是对同一Common Crawl的整理（子集），而不是扩展可用的数据集。这显然是对有限库存的消耗——我们能够推动这一杠杆的空间是有限的。","增益独立于模型配方和数据语料","以下是我们为了确定增益与模型配方和数据语料的独立性而进行的调查。我们查看了3.16e18 FLOPs下OLMES得分的网格。对 OLMES得分=均值+模型效应+数据效应 进行线性回归得到的R平方为0.88，这意味着OLMES得分的88%方差可以通过模型和数据改进的加性效应来解释，只有约12%的方差归因于交互或高阶项及评估噪声。这表明复杂的模型-数据交互（即利用某些模型改进依赖于特定数据工程，反之亦然）相对较小。","Anson Ho等人：https://arxiv.org/abs/2403.05812 估计软件效率（在预训练中）每年提高3倍（95%置信区间：1.5倍至64倍）。正如Ho在这篇博客中提到的：https://epochai.substack.com/p/the-least-understood-driver-of-ai，“大多数软件进展实际上可能归因于数据质量的提升”以及“仅从少量与规模相关的算法变化的扩展中实现”。","我们使用C = 6ND名义惯例来计算计算量。","2019年的模型配方是GPT-2，2025年的模型配方是OLMo-2。2019年的数据语料是OpenWebText，2025年的数据语料是UltraFineWeb。","我们实现的GPT-3在Pile数据集上遇到了一些训练不稳定问题（梯度峰值）。","这些包括：优化器改进、预热 + 衰减调度、RoPE 替代学习绝对位置、RMSNorm + SwiGLU 门控 MLP、规范重排序、QK-规范、Z-loss 正则化以及更干净的初始化。","对于计算扩展图，我们每种情况至少使用 3 个随机种子。对于预算为 3.16e18 的模型配方和数据语料组合的 7x7 网格，我们每种情况只使用 1 个随机种子。","1.57 倍的同比乘数是使用从 2019 年模型和语料到 2025 年模型和语料的联合改进计算得出的，而不是 1.24 倍的模型端改进与 1.51 倍的数据端改进的乘积。","比较的指标本应是生成困惑度或类似指标，这也是比较预训练模型的常用方法。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"研究称，在最高1e19 FLOPs算力预算下，2019至2025年代表性模型配方与数据语料的组合实验显示，数据改进带来12.0倍算力效率提升，模型改进为3.7倍。","background":"研究按年度选取公开模型配方和数据语料，在不同训练算力规模下组合训练，并用OLMES评估终端能力；作者指出该指标含有噪声，并通过多次随机种子尝试获得更稳定的界限。","viewpoint":"Aioga 判断：材料支持将数据改进视为该实验范围内更大的效率增益来源，但这项结论针对预训练和所测OLMES能力，不能直接外推到所有AI进展。","implications":"可能影响：若相关结果在更大规模或其他评测中保持，前沿实验室的算力投入与数据工程优先级可能需要重新权衡；但现有材料不足以证明数据改进将持续占优。","nextStep":"后续观察：需要关注更大算力预算、不同任务和评测下的复现结果，以及数据与模型改进是否仍主要呈独立贡献；当前材料仅报告至1e19 FLOPs的预训练实验。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-09-08T17:27:19.814Z","sourceHash":"c4eb24f2d0c8b849","review":{"approved":true,"groundedness":95,"clarity":94,"duplicationRisk":8,"blockingIssues":[],"notes":["“通过多次随机种子尝试获得更稳定的界限”基本对应原文“try to get cleaner bounds by running multiple seeds”，表述准确。","“不能直接外推到所有AI进展”属于基于研究范围的合理限定，且已在 viewpoint 中明确作为判断而非来源原话。","implications 使用“若”“可能”“但现有材料不足以证明”等条件性措辞，未将推测冒充为已证实事实。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","editorial-labels","inference-boundary","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","Dwarkesh Patel：Podcast & Blog（RSS）"],"translations":{"zh-CN":{"title":"Dwarkesh Patel 研究：预训练进步主要来自数据改进","summary":"Dwarkesh Patel 发布实验分析，在最高 1e19 FLOPs 的算力预算下训练 2019 至 2025 年各年度代表性模型配方与数据语料，发现数据改进带来 12.0x 算力效率提升，模型改进为 3.7x，数据贡献约为模型的 3.24 倍。","category":"行业动态","source":"dwarkesh.com","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"Dwarkesh Patel 研究：预训练进步主要来自数据改进 - Aioga AI资讯","description":"Dwarkesh Patel 发布实验分析，在最高 1e19 FLOPs 的算力预算下训练 2019 至 2025 年各年度代表性模型配方与数据语料，发现数据改进带来 12.0x 算力效率提升，模型改进为 3.7x，数据贡献约为模型的 3.24 倍。","url":"https://www.aioga.com/news/cmtsx0wm4041jrob5liwjypsq/","articleBody":["在过去几年里，我们看到的人工智能快速进展（：https://epoch.ai/gradient-updates/the-least-understood-driver-of-ai-progress ）中，有多少来自数据的改进，有多少来自模型的改进？答案对前沿实验室的经济学以及未来进展的速度有重大影响。","我们在相对较小的规模上，特别是针对预训练，从2019年到2025年调查了这个问题。在每一年，都会发布一个新的公开模型配方，总结了当年公开已知的算法调整（例如，架构、优化器、初始化、学习率调度、超参数等的改进）。在这些年份，每年也都有新的公开数据集（通过更广泛的抓取以及新的策划/提取/过滤技术生成）。","我们在不同的训练计算规模（最多达1e19 FLOPs）下训练这些年份代表性的模型配方和数据集的组合。2：#footnote-2","显然，我们不能通过交叉熵损失在固定数据集上比较这些不同模型，因为我们正在改变它们训练的数据集。因此，我们改为通过OLMES衡量模型的最终能力（https://github.com/allenai/olmes eval）（该评估汇总了10个不同的较简单基准，大多为多项选择问答）。不幸的是，评估最终能力而不是预训练损失会给我们的结果带来一些噪声，正如您在下面的图表中看到的，尽管我们尝试通过运行多个随机种子来获得更清晰的界限。","我们发现，从2019年到2025年，在1e19 FLOPs的计算预算下，数据改进带来的计算效率提升比模型改进高3.24倍（数据提升12.0倍，模型提升3.7倍）。3：#footnote-3","这里有一个表格显示了我们训练的模型在我们测试的最终能力上，相对于2019年的数据+架构基线，在3.16e18 FLOPs下表现得有多好。4：#footnote-4","我们发现，数据和模型改进带来的提升大多是独立的，并且不互相影响（即，实现某些模型改进的收益并不需要特定的训练数据集，反之亦然）。使用线性模型，88%的OLMES分数方差可以由模型和数据改进的加性效应解释。","为了提供背景，让我们简要总结一下从2019年到2025年，数据端和模型端发生的变化。","在模型方面，我们从 GPT-2：http://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf 发展到 OLMo-2：https://arxiv.org/abs/2501.00656，包括优化器、位置编码、归一化、激活函数、初始化等方面的关键创新 5：#footnote-5。","在数据方面，我们从 2019 年的 OpenWebText 开始：http://huggingface.co/datasets/Skylion007/openwebtext，其中只包含来自 Reddit 并获得足够点赞的网页，然后进行了去重和过滤，因此总共大约有~90 亿个标记（这大部分就是 GPT-2 的训练数据）。到 2025 年，像 UltraFineWeb 的开源数据语料库：http://huggingface.co/datasets/openbmb/Ultra-FineWeb 不仅规模大得多（通过抓取整个互联网的数据），而且使用了更复杂的过滤方法（例如，通过训练分类器预测哪些数据在实践中会提高模型性能）。","对我们结果的一种天真的解读是，从2019到2024年（预训练时代）的AI大部分进步只是更好的数据工程（提取、整理等），而在此期间的模型工作则不那么重要。","但这种看法可能是错误的。模型改进的主要贡献不一定是计算效率——即用更少的FLOPs实现相同性能。而是首先使更大量的计算变得可用。随着参数数量、上下文长度、运行时间和集群规模的增加，各种问题很容易出现（梯度爆炸或消失、内存和带宽不足、训练变得不可行地缓慢）。大量模型研究的内容就是消除或推迟这些扩展的约束：https://www.beren.io/2026-08-23-Architecture-Research-as-Addressing-Constraints-to-Scaling/。许多最重要的创新，如MoEs、稀疏注意力变体、稳定性创新（归一化位置、初始化等）以及系统/内核级优化，如FlashAttention：https://arxiv.org/abs/2205.14135，都属于这一类别。","我们在这里研究的数据改进对于更大的模型可能影响不大。小型模型（比如我们训练的那些）从数据质量的提升中获得显著收益，因为它们的容量有限，因此你必须非常小心地选择输入内容。而大型模型有如此多的多余容量，以至于你可能只想尽可能多地输入数据，即使大部分是垃圾，通过随机梯度下降的魔力也能从噪声中分离出信号。如果你选择进行严格筛选，你将不得不进行几十轮训练，这在经验上会比拥有平均质量较低但数据集更大的情况表现更差：https://arxiv.org/pdf/2605.19407。事实上，一旦考虑到前沿模型在训练中相对于Chinchilla最优模型被过度训练了多达100倍，以最小化RL和部署时的推理计算，过度的数据策划甚至更有害。","一个类比可能是帆船和货轮的区别——货轮不一定跑得更快，但它可以运载成千上万吨的货物（类似于数百万亿的预训练数据token），并且不会被波涛汹涌的海面推翻（类似于在成千上万GPU上稳定训练）。","现在我们有了更大容量、更坚固的货轮，我们不必担心到底装什么——我们可以把所有哪怕是稍微可能有用的东西都装上去。而对于2019年的那些娇小脆弱的帆船，你必须非常谨慎，只携带最有价值的货物。","但在预训练进展的本质仅仅是向这艘船装载更多货物的程度上，我们是否快要没有货物可装了？这是关于数据壁垒的问题，也是关于合成数据在帮助我们跨越这一壁垒方面的效果问题。显然，各实验室正在广泛使用合成数据，而我们完全没有研究过它是否能够在不影响模型性能的前提下有效扩展数据语料库。如果增益有限，那么预训练进展的主要驱动力将停滞，因为我们无法生成更多的互联网数据，而且你对固定数据集合的整理能力也是有限的。需要明确的是，我们没有实际理由认为会这样。但考虑到数据在推动预训练进展中似乎非常重要，这似乎是一个至关重要的问题，值得研究。","Ryan Greenblatt 指出：https://www.dwarkesh.com/p/ryan-greenblatt 历史上预训练数据语料库的许多改进看起来就像自动化研究者通过实证测试就能完成的进展——例如，运行不同数据训练的消融实验，看看模型表现如何。因此，我们的结果完全兼容这样的观点，即自2019年以来推动预训练的数据进展如果在人工智能研发自动化时，可能会加速很多。","我们想澄清的是，单独预训练进展的加速或减慢并不是总体人工智能进展中最重要的问题，因为过去两年的许多突破都来自强化学习（RL）。","以下是我们认为非常有趣且值得回答的一些未来研究方向：","你可以在更大规模上进行这个实验，以观察数据改进或模型改进对规模的依赖性（因此在前沿领域影响更大）","对于预训练和后训练来说，高质量新数据的边际价值是多少，以最终能力为衡量标准？","我们想大致了解合成数据的效果有多好。一个具体的研究问题是：如果你拥有一个小型高质量数据语料库，通过合成数据生成来放大它相比于仅对它训练多个周期效果有多大：https://arxiv.org/pdf/2605.19407 ?","你可以通过实验室在数据经纪人、环境生成者等方面的支出，相对于他们在计算资源和研究人员上的支出来推测数据的隐含价值。","我们想要研究数据在推动人工智能进步中所起的作用。还有很多其他方法可以探究这个问题，有些方法可能比我们的方法更巧妙、更有信息量。即使我们的实验是在极小规模下进行的，我们也认为我们可能漏掉了一些东西——我们很想听听别人会如何研究这个问题，并理想情况下也希望看到他们的结果！","特别感谢 Charlie O’Neill：https://x.com/oneill_c?lang=en 提供了许多有益的讨论。","我们从零开始在这些不同的数据语料库上预训练这些模型配方，在不同的计算预算下，使用多组独立的随机种子6：#footnote-6。我们的计算预算是：1e17、3.16e17、1e18、3.16e18 和 1e19 FLOPs。计算核算的惯例是使用名义计算 C = 6ND（N = 非嵌入参数数，D = 数据的 token 数）。","在每个计算预算下，我们改变参数数量（从而改变训练的 token 数量），以确定每个训练配方 × 语料组合的计算最优比例。我们使用语料库上的保留损失来确定这个计算最优点。然后我们可以获得每个组合的下游性能的计算扩展曲线，从中最终提取我们的计算乘数。","我们在每次运行中都强制使用统一的分词器和上下文长度：GPT-2 BPE（tiktoken，词汇量 50,257）和 T=2048，批量 = 262,144 tokens。","我们的训练运行的最终能力高度依赖于超参数。显然，不可能扫描所有可能的超参数组合（超参数调优确实是一门精细的艺术）！我们尽量控制这个因素，并认为峰值学习率是最重要的超参数。","一些算法版本确实提供了峰值学习率的调优规格（作为其他相关变量的函数，例如模型大小、数据预算、批量大小等）。这些作为我们认为的最优学习率的良好先验。","我们首先在5个锚点上扫描学习率——3种不同的模型规模和2种不同的D/N比。我们确定这些锚点的最佳学习率，并拟合一个最佳学习率的参数形式。","对于除OLMo-2之外的所有模型配方，我们拟合一个通用的指数a和b，以及一个特定模型的lr₀。对于OLMo-2，我们根据模型配方使用规定的最佳学习率。我们之所以对OLMo-2这样做，是因为Ai2发布了小型模型梯度作为配方的一部分，这些梯度在我们正在研究的规模上指定了最佳超参数。我们还在3.16e18 FLOPs的计算最优点验证，发现我们的生产学习率在最佳范围内或接近最佳。","解释我们图表中的一些异常情况","正如预期的那样，我们观察到模型和数据轴上的计算效率总体上随时间增加。我们观察到的一些异常值：","NeoX在1e19时的表现比GPT-2差（尽管在1e17到3.16e18范围内它表现更好）。这可能源于OLMES评估中的噪声。我们还注意到，在FineWeb-Edu语料库的保留预训练损失上，NeoX表现优于GPT-2。","Pile的表现似乎比OpenWebText差得多。这并不令人意外，因为Pile的主要改进在于数据语料库的多样性，而非过滤。它包含一个由22个来源精心策划的混合，包括PubMed和arXiv论文、GitHub代码、法律意见、专利和议会记录。对于这些许多标记而言，转移到OLMES（即英文网页散文MCQ）的跨领域效果可能很小，因此导致计算效率较低。我们注意到，由于其规模较大，我们预计Pile在更大规模下最终应会比（真正小的）OpenWebText表现更好。","还值得注意的是，NeoX和Pile的计算乘数是通过外推获得的，这引入了进一步的潜在误差。","计算乘数以及它们的误差范围的方法","计算缩放曲线上的每个点都是基于多个独立初始化训练运行得出的。图上的误差范围是这些初始化的OLMES评估的标准差。","考虑某个参考模型或数据语料库在某个计算水平下的一定参考性能水平。","然后，我们通过找到候选模型或语料库的计算缩放曲线中首次达到该参考性能水平的最左端点，来计算计算倍增器。参考所需的计算量与候选所需计算量的比率就是候选的计算倍增器。","计算倍增器的误差条是通过对整个估计流程进行参数自助法（parametric bootstrap）得到的，表示的是1个标准差区间。","我们确实想强调的是，我们预计模型方案的计算倍增器的实际不确定性要高于误差条所显示的。这是因为我们超参数调优的范围有限所引入的额外不确定性，以及最终的能力或保留损失可能对峰值学习率、批量大小等具体选择非常敏感。","同样重要的是要注意，存在许多原因说明我们的消融实验未必完全捕捉到计算效率增益的全部范围。事实上，从2019年到2025年，我们观察到模型端的年均计算效率增益(CEG)为1.24x [1.19, 1.29]，数据端为1.51x [1.45, 1.57]。联合测量时，我们观察到1.57x的年均CEG [1.49, 1.65] 7：#footnote-7。这确实远低于Anson Ho等人的平均估计值3x年均CEG：https://arxiv.org/abs/2403.05812，原因如下：","许多收益可能依赖于规模：https://arxiv.org/pdf/2511.21622，或者在较长的上下文中尤其重要，而我们目前操作的规模太小，无法实现许多收益。","例如，OLMo-2的层和QK范数，NeoX中的并行注意力+MLP模块","推理效率优化（例如 LLama-3 的 GQA：https://arxiv.org/abs/2305.13245，这是一个 KV 缓存优化）在我们的研究中并未显示为计算倍增器。我们也没有研究分词器的改进。","我们得到的计算倍增器对每年的模型方案或数据语料库的选择非常敏感。我们选择了我们认为具有代表性的模型方案或数据语料库。但我们绝不意味着穷尽地得出这些就是每年的最佳选择。","我们正在观察相对于OLMES基准（该基准结合了10种不同的相对简单的任务类型）的计算乘数，而不是在达到某个困惑度指标时的计算乘数。如果我们观察其他基准（比如专门针对编码或问题解决的基准），这些数字也会完全不同，这可能会奖励非常不同的数据工程方法。","我们还想指出，我们没有研究其他数据方面的改进，例如从新来源收集更多高质量数据、人工专家生成的数据、合成数据生成方法等。我们调查的大多数语料库都是对同一Common Crawl的整理（子集），而不是扩展可用的数据集。这显然是对有限库存的消耗——我们能够推动这一杠杆的空间是有限的。","增益独立于模型配方和数据语料","以下是我们为了确定增益与模型配方和数据语料的独立性而进行的调查。我们查看了3.16e18 FLOPs下OLMES得分的网格。对 OLMES得分=均值+模型效应+数据效应 进行线性回归得到的R平方为0.88，这意味着OLMES得分的88%方差可以通过模型和数据改进的加性效应来解释，只有约12%的方差归因于交互或高阶项及评估噪声。这表明复杂的模型-数据交互（即利用某些模型改进依赖于特定数据工程，反之亦然）相对较小。","Anson Ho等人：https://arxiv.org/abs/2403.05812 估计软件效率（在预训练中）每年提高3倍（95%置信区间：1.5倍至64倍）。正如Ho在这篇博客中提到的：https://epochai.substack.com/p/the-least-understood-driver-of-ai，“大多数软件进展实际上可能归因于数据质量的提升”以及“仅从少量与规模相关的算法变化的扩展中实现”。","我们使用C = 6ND名义惯例来计算计算量。","2019年的模型配方是GPT-2，2025年的模型配方是OLMo-2。2019年的数据语料是OpenWebText，2025年的数据语料是UltraFineWeb。","我们实现的GPT-3在Pile数据集上遇到了一些训练不稳定问题（梯度峰值）。","这些包括：优化器改进、预热 + 衰减调度、RoPE 替代学习绝对位置、RMSNorm + SwiGLU 门控 MLP、规范重排序、QK-规范、Z-loss 正则化以及更干净的初始化。","对于计算扩展图，我们每种情况至少使用 3 个随机种子。对于预算为 3.16e18 的模型配方和数据语料组合的 7x7 网格，我们每种情况只使用 1 个随机种子。","1.57 倍的同比乘数是使用从 2019 年模型和语料到 2025 年模型和语料的联合改进计算得出的，而不是 1.24 倍的模型端改进与 1.51 倍的数据端改进的乘积。","比较的指标本应是生成困惑度或类似指标，这也是比较预训练模型的常用方法。"]},"en":{"title":"Dwarkesh Patel Study: Advances in pretraining mainly come from data enhancements","summary":"Dwarkesh Patel published an experimental analysis training representative model recipes and data corpora from 2019 to 2025 with a maximum computing power budget of 1e19 FLOPs, finding that data improvements led to a 12.0x increase in computing efficiency, with model improvements of 3.7x and data contributions about 3.24 times that of the model.","category":"Industry","source":"Dwarkesh Patel：Podcast & Blog（RSS）","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"Dwarkesh Patel Study: Advances in pretraining mainly come from data enhancements - Aioga AI News","description":"Dwarkesh Patel published an experimental analysis training representative model recipes and data corpora from 2019 to 2025 with a maximum computing power budget of 1e19 FLOPs, find...","url":"https://www.aioga.com/en/news/cmtsx0wm4041jrob5liwjypsq/","contentTranslated":true,"sourceHash":"d49d8c06ac46ffa7","translatedAt":"2026-09-08T17:22:00.951Z"},"ja":{"title":"ドワルケシュ・パテル研究:事前訓練の進歩は主にデータの強化によるものです","summary":"Dwarkesh Patelは、2019年から2025年にかけて最大計算能力1e19 FLOPの計算能力予算で代表的なモデルレシピとデータコーパスを訓練した実験解析を発表し、データ改善により計算効率が12.0倍向上し、モデルの改善は3.7倍、データ寄与量はモデルの約3.24倍に達しました。","category":"業界動向","source":"Dwarkesh Patel：Podcast & Blog（RSS）","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"ドワルケシュ・パテル研究:事前訓練の進歩は主にデータの強化によるものです - Aioga AIニュース","description":"Dwarkesh Patelは、2019年から2025年にかけて最大計算能力1e19 FLOPの計算能力予算で代表的なモデルレシピとデータコーパスを訓練した実験解析を発表し、データ改善により計算効率が12.0倍向上し、モデルの改善は3.7倍、データ寄与量はモデルの約3.24倍に達しました。","url":"https://www.aioga.com/ja/news/cmtsx0wm4041jrob5liwjypsq/","contentTranslated":true,"sourceHash":"d49d8c06ac46ffa7","translatedAt":"2026-09-08T17:22:00.414Z"},"ko":{"title":"드와르케시 파텔 연구: 사전 훈련의 발전은 주로 데이터 향상에서 나옵니다","summary":"Dwarkesh Patel은 2019년부터 2025년까지 최대 컴퓨팅 파워 예산 1e19 FLOP로 대표 모델 레시피와 데이터 코퍼라를 훈련시킨 실험 분석을 발표했으며, 데이터 개선으로 컴퓨팅 효율이 12.0배 증가했고, 모델 개선은 3.7배, 데이터 기여도는 모델의 약 3.24배에 달했습니다.","category":"업계 동향","source":"Dwarkesh Patel：Podcast & Blog（RSS）","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"드와르케시 파텔 연구: 사전 훈련의 발전은 주로 데이터 향상에서 나옵니다 - Aioga AI 뉴스","description":"Dwarkesh Patel은 2019년부터 2025년까지 최대 컴퓨팅 파워 예산 1e19 FLOP로 대표 모델 레시피와 데이터 코퍼라를 훈련시킨 실험 분석을 발표했으며, 데이터 개선으로 컴퓨팅 효율이 12.0배 증가했고, 모델 개선은 3.7배, 데이터 기여도는 모델의 약 3.24배에 달했습니다.","url":"https://www.aioga.com/ko/news/cmtsx0wm4041jrob5liwjypsq/","contentTranslated":true,"sourceHash":"d49d8c06ac46ffa7","translatedAt":"2026-09-08T17:22:09.696Z"},"es":{"title":"Estudio de Dwarkesh Patel: Los avances en la preformación provienen principalmente de mejoras de datos","summary":"Dwarkesh Patel publicó un análisis experimental de entrenamiento con recetas representativas de modelos y corpus de datos entre 2019 y 2025 con un presupuesto máximo de potencia de cálculo de 1e19 FLOPs, descubriendo que las mejoras en los datos conducían a un aumento de 12,0 veces en la eficiencia informática, con mejoras en el modelo de 3,7x y contribuciones de datos aproximadamente 3,24 veces superiores a las del modelo.","category":"Industria","source":"Dwarkesh Patel：Podcast & Blog（RSS）","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"Estudio de Dwarkesh Patel: Los avances en la preformación provienen principalmente de mejoras de datos - Aioga Noticias de IA","description":"Dwarkesh Patel publicó un análisis experimental de entrenamiento con recetas representativas de modelos y corpus de datos entre 2019 y 2025 con un presupuesto máximo de potencia de...","url":"https://www.aioga.com/es/news/cmtsx0wm4041jrob5liwjypsq/","contentTranslated":true,"sourceHash":"d49d8c06ac46ffa7","translatedAt":"2026-09-08T17:22:09.804Z"},"fr":{"title":"Étude de Dwarkesh Patel : Les avancées en pré-entraînement proviennent principalement de l’amélioration des données","summary":"Dwarkesh Patel a publié une analyse expérimentale d’entraînement représentant les recettes de modèles et les corpus de données de 2019 à 2025 avec un budget de puissance de calcul maximal de 1e19 FLOPs, constatant que les améliorations des données entraînaient une augmentation de l’efficacité informatique de 12,0, avec des améliorations du modèle de 3,7x et des contributions de données environ 3,24 fois supérieures à celles du modèle.","category":"Industrie","source":"Dwarkesh Patel：Podcast & Blog（RSS）","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"Étude de Dwarkesh Patel : Les avancées en pré-entraînement proviennent principalement de l’amélioration des données - Aioga Actualités IA","description":"Dwarkesh Patel a publié une analyse expérimentale d’entraînement représentant les recettes de modèles et les corpus de données de 2019 à 2025 avec un budget de puissance de calcul...","url":"https://www.aioga.com/fr/news/cmtsx0wm4041jrob5liwjypsq/","contentTranslated":true,"sourceHash":"d49d8c06ac46ffa7","translatedAt":"2026-09-08T17:22:18.657Z"},"de":{"title":"Dwarkesh Patel Studie: Fortschritte im Pretraining basieren hauptsächlich auf Datenverbesserungen","summary":"Dwarkesh Patel veröffentlichte von 2019 bis 2025 eine experimentelle Analyse, die repräsentative Modellrezepte und Datenkorpora mit einem maximalen Rechenleistungsbudget von 1e19 FLOPs trainierte, und stellte fest, dass Datenverbesserungen zu einer 12,0-fachen Steigerung der Recheneffizienz führten, mit Modellverbesserungen von 3,7-fach und Datenbeiträgen, die etwa 3,24-fach so hoch waren wie das Modell.","category":"行业动态","source":"Dwarkesh Patel：Podcast & Blog（RSS）","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"Dwarkesh Patel Studie: Fortschritte im Pretraining basieren hauptsächlich auf Datenverbesserungen - Aioga KI-News","description":"Dwarkesh Patel veröffentlichte von 2019 bis 2025 eine experimentelle Analyse, die repräsentative Modellrezepte und Datenkorpora mit einem maximalen Rechenleistungsbudget von 1e19 F...","url":"https://www.aioga.com/de/news/cmtsx0wm4041jrob5liwjypsq/","contentTranslated":true,"sourceHash":"d49d8c06ac46ffa7","translatedAt":"2026-09-08T17:22:18.945Z"},"pt-BR":{"title":"Estudo Dwarkesh Patel: Os avanços no pré-treinamento vêm principalmente de aprimoramentos de dados","summary":"Dwarkesh Patel publicou uma análise experimental de treinamento representativo de receitas de modelos e corpora de dados de 2019 a 2025, com um orçamento máximo de poder computacional de 1e19 FLOPs, descobrindo que melhorias nos dados levaram a um aumento de 12,0x na eficiência computacional, com melhorias no modelo de 3,7x e contribuições de dados cerca de 3,24 vezes maiores que o modelo.","category":"行业动态","source":"Dwarkesh Patel：Podcast & Blog（RSS）","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"Estudo Dwarkesh Patel: Os avanços no pré-treinamento vêm principalmente de aprimoramentos de dados - Aioga Notícias de IA","description":"Dwarkesh Patel publicou uma análise experimental de treinamento representativo de receitas de modelos e corpora de dados de 2019 a 2025, com um orçamento máximo de poder computacio...","url":"https://www.aioga.com/pt-BR/news/cmtsx0wm4041jrob5liwjypsq/","contentTranslated":true,"sourceHash":"d49d8c06ac46ffa7","translatedAt":"2026-09-08T17:22:27.813Z"},"ru":{"title":"Исследование Дваркеша Пателя: Прогресс в предварительном обучении в основном связан с улучшением данных","summary":"Дваркеш Патель опубликовал экспериментальный анализ, обучающий репрезентативным рецептам моделей и корпусам данных с 2019 по 2025 год с максимальным вычислительной мощностью в 1e19 FLOP, установив, что улучшения данных привели к увеличению эффективности вычислительной работы в 12,0 раза, при этом улучшения модели составили 3,7 раза и увеличили вклад данных примерно в 3,24 раза больше, чем у модели.","category":"行业动态","source":"Dwarkesh Patel：Podcast & Blog（RSS）","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"Исследование Дваркеша Пателя: Прогресс в предварительном обучении в основном связан с улучшением данных - Aioga Новости ИИ","description":"Дваркеш Патель опубликовал экспериментальный анализ, обучающий репрезентативным рецептам моделей и корпусам данных с 2019 по 2025 год с максимальным вычислительной мощностью в 1e19...","url":"https://www.aioga.com/ru/news/cmtsx0wm4041jrob5liwjypsq/","contentTranslated":true,"sourceHash":"d49d8c06ac46ffa7","translatedAt":"2026-09-08T17:22:28.082Z"},"ar":{"title":"دراسة دواركيش باتيل: تأتي التطورات في التدريب المسبق بشكل رئيسي من تحسينات البيانات","summary":"نشر دواركيش باتيل تحليل تجريبي تدريب على وصفات النماذج التمثيلية ومجموعات البيانات من 2019 إلى 2025 بميزانية قوة حوسبة قصوى تبلغ 1e19 FLOPs، ووجد أن تحسينات البيانات أدت إلى زيادة بنسبة 12.0 مرة في كفاءة الحوسبة، مع تحسينات في النموذج بمقدار 3.7 ضعف ومساهمات البيانات حوالي 3.24 ضعف النموذج.","category":"行业动态","source":"Dwarkesh Patel：Podcast & Blog（RSS）","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"دراسة دواركيش باتيل: تأتي التطورات في التدريب المسبق بشكل رئيسي من تحسينات البيانات - Aioga أخبار الذكاء الاصطناعي","description":"نشر دواركيش باتيل تحليل تجريبي تدريب على وصفات النماذج التمثيلية ومجموعات البيانات من 2019 إلى 2025 بميزانية قوة حوسبة قصوى تبلغ 1e19 FLOPs، ووجد أن تحسينات البيانات أدت إلى زيادة...","url":"https://www.aioga.com/ar/news/cmtsx0wm4041jrob5liwjypsq/","contentTranslated":true,"sourceHash":"d49d8c06ac46ffa7","translatedAt":"2026-09-08T17:22:37.065Z"},"hi":{"title":"द्वारकेश पटेल अध्ययन: प्रीट्रेनिंग में प्रगति मुख्य रूप से डेटा संवर्द्धन से आती है","summary":"द्वारकेश पटेल ने 2019 से 2025 तक 1e19 FLOP के अधिकतम कंप्यूटिंग पावर बजट के साथ एक प्रयोगात्मक विश्लेषण प्रशिक्षण प्रतिनिधि मॉडल व्यंजनों और डेटा कॉर्पोरा को प्रकाशित किया, जिसमें पाया गया कि डेटा सुधार के कारण कंप्यूटिंग दक्षता में 12.0 गुना की वृद्धि हुई, जिसमें मॉडल में 3.7x का सुधार और मॉडल के लगभग 3.24 गुना डेटा योगदान था।","category":"行业动态","source":"Dwarkesh Patel：Podcast & Blog（RSS）","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"द्वारकेश पटेल अध्ययन: प्रीट्रेनिंग में प्रगति मुख्य रूप से डेटा संवर्द्धन से आती है - Aioga AI समाचार","description":"द्वारकेश पटेल ने 2019 से 2025 तक 1e19 FLOP के अधिकतम कंप्यूटिंग पावर बजट के साथ एक प्रयोगात्मक विश्लेषण प्रशिक्षण प्रतिनिधि मॉडल व्यंजनों और डेटा कॉर्पोरा को प्रकाशित किया, जिसमें...","url":"https://www.aioga.com/hi/news/cmtsx0wm4041jrob5liwjypsq/","contentTranslated":true,"sourceHash":"d49d8c06ac46ffa7","translatedAt":"2026-09-08T17:22:37.134Z"},"it":{"title":"Studio di Dwarkesh Patel: I progressi nella pre-formazione derivano principalmente dal miglioramento dei dati","summary":"Dwarkesh Patel ha pubblicato un'analisi sperimentale di addestramento rappresentativo ricette di modelli e corpi dati dal 2019 al 2025 con un budget massimo di potenza di calcolo di 1e19 FLOP, rilevando che i miglioramenti dei dati portavano a un aumento di 12,0 volte dell'efficienza di calcolo, con miglioramenti del modello di 3,7 volte e contributi dati circa 3,24 volte rispetto al modello.","category":"行业动态","source":"Dwarkesh Patel：Podcast & Blog（RSS）","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"Studio di Dwarkesh Patel: I progressi nella pre-formazione derivano principalmente dal miglioramento dei dati - Aioga Notizie IA","description":"Dwarkesh Patel ha pubblicato un'analisi sperimentale di addestramento rappresentativo ricette di modelli e corpi dati dal 2019 al 2025 con un budget massimo di potenza di calcolo d...","url":"https://www.aioga.com/it/news/cmtsx0wm4041jrob5liwjypsq/","contentTranslated":true,"sourceHash":"d49d8c06ac46ffa7","translatedAt":"2026-09-08T17:22:46.187Z"},"nl":{"title":"Dwarkesh Patel Studie: Vooruitgang in pretraining komt voornamelijk door dataverbeteringen","summary":"Dwarkesh Patel publiceerde een experimentele analyse waarin representatieve modelrecepten en datacorpora werden getraind van 2019 tot 2025 met een maximaal rekenkrachtbudget van 1e19 FLOPs, waarbij werd vastgesteld dat dataverbeteringen leidden tot een 12,0x toename van de rekenefficiëntie, met modelverbeteringen van 3,7x en databijdragen ongeveer 3,24 keer die van het model.","category":"行业动态","source":"Dwarkesh Patel：Podcast & Blog（RSS）","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"Dwarkesh Patel Studie: Vooruitgang in pretraining komt voornamelijk door dataverbeteringen - Aioga AI-nieuws","description":"Dwarkesh Patel publiceerde een experimentele analyse waarin representatieve modelrecepten en datacorpora werden getraind van 2019 tot 2025 met een maximaal rekenkrachtbudget van 1e...","url":"https://www.aioga.com/nl/news/cmtsx0wm4041jrob5liwjypsq/","contentTranslated":true,"sourceHash":"d49d8c06ac46ffa7","translatedAt":"2026-09-08T17:22:45.271Z"},"tr":{"title":"Dwarkesh Patel Çalışması: Ön eğitimdeki ilerlemeler esas olarak veri geliştirmelerden kaynaklanır","summary":"Dwarkesh Patel, 2019'dan 2025'e kadar 1e19 FLOP maksimum hesaplama gücü bütçesiyle temsil model tarifleri ve veri korpuslarını eğiten deneysel bir analiz yayımladı; veri iyileştirmelerinin hesaplama verimliliğinde 12,0 kat artış, model iyileştirmelerinin 3,7 kat ve veri katkılarının modelin yaklaşık 3,24 katı olduğunu buldu.","category":"行业动态","source":"Dwarkesh Patel：Podcast & Blog（RSS）","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"Dwarkesh Patel Çalışması: Ön eğitimdeki ilerlemeler esas olarak veri geliştirmelerden kaynaklanır - Aioga AI Haberleri","description":"Dwarkesh Patel, 2019'dan 2025'e kadar 1e19 FLOP maksimum hesaplama gücü bütçesiyle temsil model tarifleri ve veri korpuslarını eğiten deneysel bir analiz yayımladı; veri iyileştirm...","url":"https://www.aioga.com/tr/news/cmtsx0wm4041jrob5liwjypsq/","contentTranslated":true,"sourceHash":"d49d8c06ac46ffa7","translatedAt":"2026-09-08T17:22:55.244Z"},"vi":{"title":"Nghiên cứu Dwarkesh Patel: Những tiến bộ trong đào tạo trước chủ yếu đến từ việc nâng cao dữ liệu","summary":"Dwarkesh Patel đã công bố một công thức mô hình đại diện và kho dữ liệu huấn luyện phân tích thực nghiệm từ năm 2019 đến 2025 với ngân sách sức mạnh tính toán tối đa là 1e19 FLOPs, phát hiện rằng các cải tiến dữ liệu dẫn đến hiệu quả tính toán tăng 12,0 lần, với cải tiến mô hình 3,7 lần và đóng góp dữ liệu khoảng 3,24 lần so với mô hình.","category":"行业动态","source":"Dwarkesh Patel：Podcast & Blog（RSS）","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"Nghiên cứu Dwarkesh Patel: Những tiến bộ trong đào tạo trước chủ yếu đến từ việc nâng cao dữ liệu - Tin tức AI Aioga","description":"Dwarkesh Patel đã công bố một công thức mô hình đại diện và kho dữ liệu huấn luyện phân tích thực nghiệm từ năm 2019 đến 2025 với ngân sách sức mạnh tính toán tối đa là 1e19 FLOPs,...","url":"https://www.aioga.com/vi/news/cmtsx0wm4041jrob5liwjypsq/","contentTranslated":true,"sourceHash":"d49d8c06ac46ffa7","translatedAt":"2026-09-08T17:22:54.999Z"},"id":{"title":"Studi Dwarkesh Patel: Kemajuan dalam pra-pelatihan terutama berasal dari peningkatan data","summary":"Dwarkesh Patel menerbitkan analisis eksperimental pelatihan resep model representatif dan korpus data dari tahun 2019 hingga 2025 dengan anggaran daya komputasi maksimum 1e19 FLOP, menemukan bahwa peningkatan data menyebabkan peningkatan efisiensi komputasi sebesar 12,0 kali lipat, dengan peningkatan model sebesar 3,7x dan kontribusi data sekitar 3,24 kali lipat dari model tersebut.","category":"行业动态","source":"Dwarkesh Patel：Podcast & Blog（RSS）","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"Studi Dwarkesh Patel: Kemajuan dalam pra-pelatihan terutama berasal dari peningkatan data - Berita AI Aioga","description":"Dwarkesh Patel menerbitkan analisis eksperimental pelatihan resep model representatif dan korpus data dari tahun 2019 hingga 2025 dengan anggaran daya komputasi maksimum 1e19 FLOP,...","url":"https://www.aioga.com/id/news/cmtsx0wm4041jrob5liwjypsq/","contentTranslated":true,"sourceHash":"d49d8c06ac46ffa7","translatedAt":"2026-09-08T17:23:03.999Z"},"th":{"title":"การศึกษาของ Dwarkesh Patel: ความก้าวหน้าในการฝึกอบรมก่อนเกิดจากการปรับปรุงข้อมูลเป็นหลัก","summary":"Dwarkesh Patel ได้ตีพิมพ์สูตรจําลองและคลังข้อมูลสําหรับการวิเคราะห์เชิงทดลอง โดยมีงบประมาณกําลังประมวลผลสูงสุด 1e19 FLOPs พบว่าการปรับปรุงข้อมูลทําให้ประสิทธิภาพการประมวลผลเพิ่มขึ้น 12.0 เท่า โดยโมเดลดีขึ้น 3.7 เท่า และข้อมูลเพิ่มขึ้นประมาณ 3.24 เท่าของโมเดล","category":"行业动态","source":"Dwarkesh Patel：Podcast & Blog（RSS）","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"การศึกษาของ Dwarkesh Patel: ความก้าวหน้าในการฝึกอบรมก่อนเกิดจากการปรับปรุงข้อมูลเป็นหลัก - ข่าว AI Aioga","description":"Dwarkesh Patel ได้ตีพิมพ์สูตรจําลองและคลังข้อมูลสําหรับการวิเคราะห์เชิงทดลอง โดยมีงบประมาณกําลังประมวลผลสูงสุด 1e19 FLOPs พบว่าการปรับปรุงข้อมูลทําให้ประสิทธิภาพการประมวลผลเพิ่มขึ้...","url":"https://www.aioga.com/th/news/cmtsx0wm4041jrob5liwjypsq/","contentTranslated":true,"sourceHash":"d49d8c06ac46ffa7","translatedAt":"2026-09-08T17:23:04.079Z"},"pl":{"title":"Badanie Dwarkesha Patela: Postępy w pretreningu wynikają głównie z ulepszeń danych","summary":"Dwarkesh Patel opublikował analizę eksperymentalną, trenującą reprezentatywne wzorce modeli i korpusy danych w latach 2019–2025 z maksymalnym budżetem mocy obliczeniowej 1e19 FLOPs, stwierdzając, że poprawa danych doprowadziła do 12,0-krotnego wzrostu efektywności obliczeniowej, z 3,7-krotną poprawą modelu i wkładem danych około 3,24-krotnie większą niż model.","category":"行业动态","source":"Dwarkesh Patel：Podcast & Blog（RSS）","aggregationSource":"Dwarkesh Patel：Podcast & Blog（RSS）","pageTitle":"Badanie Dwarkesha Patela: Postępy w pretreningu wynikają głównie z ulepszeń danych - Aioga Wiadomości AI","description":"Dwarkesh Patel opublikował analizę eksperymentalną, trenującą reprezentatywne wzorce modeli i korpusy danych w latach 2019–2025 z maksymalnym budżetem mocy obliczeniowej 1e19 FLOPs...","url":"https://www.aioga.com/pl/news/cmtsx0wm4041jrob5liwjypsq/","contentTranslated":true,"sourceHash":"d49d8c06ac46ffa7","translatedAt":"2026-09-08T17:23:12.577Z"}},"evidenceTier":"verified-news","reviewStatus":"editorial-selected","indexable":true,"editorialCover":"/page-visuals/topic-timeline.png"}}