{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-09-18T17:00:43.709Z","headline":"Anthropic 发布前沿 AI 开发节奏测量工具与内部指标快照","description":"Anthropic 发布一套测量前沿 AI 开发节奏的指标，覆盖 AI 主导研发、智能体监督和算力分配三方面。","url":"https://www.aioga.com/news/cmu605ozy000arok0lxr4x1y8/","mainEntityOfPage":"https://www.aioga.com/news/cmu605ozy000arok0lxr4x1y8/","datePublished":"2026-09-17T20:49:01.000Z","dateModified":"2026-09-17T20:49:01.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.anthropic.com/institute/measuring-pace-of-ai-development","https://aihot.news/items/cmu605ozy000arok0lxr4x1y8"],"canonicalUrl":"https://www.aioga.com/news/cmu605ozy000arok0lxr4x1y8/","directAnswer":{"@type":"Answer","text":"Anthropic 发布用于观察前沿 AI 实验室开发节奏的测量框架，覆盖 AI 参与研发、AI 智能体行为监督及算力分配，并披露 Anthropic 内部相关指标的阶段性快照。","url":"https://www.aioga.com/news/cmu605ozy000arok0lxr4x1y8/","dateCreated":"2026-09-17T20:49:01.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"Anthropic source article","url":"https://www.anthropic.com/institute/measuring-pace-of-ai-development","datePublished":"2026-09-17T20:49:01.000Z","provider":{"@type":"Organization","name":"Anthropic","url":"https://www.anthropic.com/institute/measuring-pace-of-ai-development"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.news/items/cmu605ozy000arok0lxr4x1y8","datePublished":"2026-09-17T20:49:01.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.news/items/cmu605ozy000arok0lxr4x1y8"}}],"aggregationSource":"Anthropic：The Institute（旗舰研究长文 · 网页）","originalPublisher":{"name":"Anthropic","url":"https://www.anthropic.com/institute/measuring-pace-of-ai-development"},"geoDeepAnswer":null,"article":{"id":"cmu605ozy000arok0lxr4x1y8","slug":"cmu605ozy000arok0lxr4x1y8","url":"https://www.aioga.com/news/cmu605ozy000arok0lxr4x1y8/","title":"Anthropic 发布前沿 AI 开发节奏测量工具与内部指标快照","title_en":"","summary":"Anthropic 发布一套测量前沿 AI 开发节奏的指标，覆盖 AI 主导研发、智能体监督和算力分配三方面。","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","sourceUrl":"https://www.anthropic.com/institute/measuring-pace-of-ai-development","aiHotUrl":"https://aihot.news/items/cmu605ozy000arok0lxr4x1y8","publishedAt":"2026-09-17T20:49:01.000Z","category":"行业动态","score":72,"selected":true,"articleBody":["AI systems are getting more powerful, and they’re increasingly being used to build the next version of themselves.","We want to illuminate the pace of progress for the public.","To do that, we’re sharing three measurements that will help the public track AI development inside frontier AI labs: how much of AI R&D is performed by AI itself , how well the actions of AI agents are overseen , and how compute is allocated .","Measurements for understanding the pace of AI development inside frontier labs","AI systems are becoming exponentially more powerful and have begun to automate more：https://www.anthropic.com/institute/recursive-self-improvement of the process of building themselves. As the world considers slowing the pace of frontier AI development：https://darioamodei.com/post/we-must-pace-the-frontier, the public needs more information.","In this post, we lay out measurement tools that can illuminate three critical aspects of AI development:","We also provide a snapshot of these metrics from inside Anthropic. It’s important to note that we would expect these numbers to shift if there were coordination on pacing the frontier, as called for by Anthropic CEO Dario Amodei. We plan to embed independent third-party evaluators from multiple organizations at Anthropic, and give them access to internal processes, systems, and data comparable to what internal risk assessment teams have. These third parties will verify safety practices, report incidents, and monitor key metrics such as the ones in this piece.","We are reporting these measurements because they give the public, third parties, and governments better visibility into the pace of AI development inside frontier labs. For each measurement, we describe what we measured, what the measurement showed, and what it would take to publish these measurements regularly in a form others can verify. We share methodological details in the Appendix.","The measurements in this piece are focused on how models are built. By better understanding the production process of models, we have a better chance of correlating model inputs, like compute, with model outputs, like capabilities. They complement capability evaluations, which measure what models can do . We publish those separately through our Responsible Scaling Policy：https://www.anthropic.com/responsible-scaling-policy (RSP) risk reports, which include evidence on how much our models are accelerating AI R&D. In our policy proposal on advanced AI, the Advanced AI Framework (AAIF)：https://www-cdn.anthropic.com/files/4zrzovbb/website/0a58d567024a8b448ff15158ebc3625328dfcc1f.pdf, we propose rules of the road for how any lab releases safe models, including transparency obligations that governments could require, such as risk reports. Together, these proposed measurements and policies are a starting point for monitoring the pace of AI development from outside the labs.","Why measure AI-led R&D? Frontier AI labs increasingly use AI to build future AI models. This process allows labs in democratic countries to develop more capable models more quickly and conduct more safety and testing on models before they are released to secure AI’s benefits while staying on the frontier. However, models accelerating their own development could make it more challenging for humans to understand or control these systems. It is therefore important to share these metrics to understand how close the world is to reaching recursive self improvement：https://www.anthropic.com/institute/recursive-self-improvement (a model fully autonomously building its successor).","What we measured. We built a prototype index of how much of Anthropic’s AI research and development (R&D) is performed by Claude, called the Anthropic R&D Automation Index. It’s built by cataloguing every kind of AI R&D work done at the company, rating how automated each task currently is, and aggregating those ratings.","What we found. To measure the extent to which AI is doing AI R&D at Anthropic, we use an automation rating scale：https://epochai.substack.com/p/toward-an-onet-for-ai-r-and-d developed by Epoch AI that measures “Automation Level,” or AL. It runs from AL0 (no AI involvement) to AL5 (AI operates fully autonomously, with no human in the loop). In AL3, AI “collaborates”: it can do large chunks of work under close human direction. In AL4, AI “leads”: it can complete most of the task end-to-end from a high-level prompt, while the human supervises 1：#footnote-1 .","What any AI developer could report today. Any frontier developer could publish these measures regularly, using a public methodology. This would enable the numbers to be compared over time, and potentially across labs.","Two obstacles stand in the way of cross-lab comparison on this type of reporting. First is the lack of a common methodology. Second, we’re using our own models to evaluate our systems, which could mean that the “judge” model could make the same kinds of errors as the model it is checking. To address these challenges, a developer’s measurements could be verified by a third party, or by other developers’ models, with guardrails in place to limit the exchange of competitively sensitive data. These measures could also become the trigger for stronger requirements, like a fixed testing window before a new model is used for further AI R&D.","Why measure oversight of agents? Like other frontier developers, Anthropic employees increasingly delegate tasks to agents that work semi-autonomously for long stretches, and that delegate work to one another. As work becomes increasingly automated, from “AI collaborates” toward “AI leads,” agents could make more consequential decisions, such as which research direction to pursue next.","These decisions need safeguards to block harmful behavior from an individual agent, as well as an oversight regime to detect problems that might only arise across many agents. In our monitoring data to date, individual agents rarely misbehave. But when there are millions or billions of agents operating in the economy, even rare events can happen regularly.","What we measured. We’ve built a system that lets us oversee and intervene in actions AI agents take on Anthropic’s systems. Here, we consider three different metrics: coverage, which describes the share of an agent’s actions that pass through a monitor before or after execution; review latency , which is the time between an action and its review, first by an automated monitor and then by a human; and escalation rate, which is the share of agent activities that are either blocked/redirected (in the case of online monitors) or flagged for further review (in the case of offline monitors).","What we found. As of August 2026, there were approximately 30,000 agents doing research and engineering work at Anthropic at any one time in our most-used internal platform. These measurements cover this platform only. The actions of these agents are constrained by two kinds of monitors, summarized below:","What an AI developer could report today. Any developer running agents on its own research and engineering workloads could publish the same measures: coverage (the share of agent activity that is monitored), review latency (how soon flagged activities are reviewed), and escalation rate (what share of agent activities the monitors block or flag). Together, these measures would allow society to see whether oversight is keeping pace with the growing role of AI in AI R&D. We published all of these measurements in our recent risk report：https://www.anthropic.com/aug-2026-risk-report.","Developers should share how often known agent misbehavior is caught by monitors, and should have third parties do this kind of testing independently. In the past, METR has independently red-teamed our offline monitoring platform：https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf. As described above, we are now setting up：https://darioamodei.com/post/we-must-pace-the-frontier external third party evaluators at Anthropic.","Why measure compute allocation? Broadly speaking, AI developers use compute for building more powerful models, serving customers, and safety-focused work like auditing a model’s “thoughts”：https://www.anthropic.com/research/natural-language-autoencoders, training model organisms to study misalignment：https://www.anthropic.com/research/emergent-misalignment-reward-hacking, and evaluating whether a model can be safely deployed：https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf. Understanding how AI developers allocate their compute can tell you where a developer is focusing its resources and how that focus changes over time.","Additionally, compute is among the most verifiable inputs to the AI R&D process, meaning that it could be a critical lever in a future pacing effort. A coordinated pacing effort could encourage companies to increase the compute allocated to safety across the industry and devote more resources to alignment, interpretability, safety testing, and evaluation.","What we measured. We examined a snapshot of how Anthropic used all of its compute from July 13 to July 20 2：#footnote-2 . To do that, we sorted every workload into a small number of categories, then asked how much of the compute going to AI R&D was safety work.","Safety research tends to use less compute than frontier training runs by its nature, so compute is an imperfect proxy for how much a company focuses on safety. This is because safety research consists of individual researchers designing experiments, which is time-consuming even though running the experiments is not particularly compute-intensive. The value of this metric, therefore, is less the absolute numbers and more that it provides a straightforward mechanism to compare like with like, across developers and over time.","What we found. Over the examined week, about 6% of compute that went to AI R&D was allocated toward safety, and about 12% of compute that went to AI-driven AI R&D was allocated toward safety.","These are deliberately conservative estimates. For example, if a token was used to advance capabilities as much as it was to advance safety, it was not counted in these metrics. Additionally, these metrics do not account for safeguards classifiers, which are a separate, comparable amount of compute that make our models much safer for the world.","What an AI developer could report today. Any frontier developer could publish what share of its AI R&D compute goes to safety work, with the category definitions published alongside and the classification checked by an independent third party.","Safety research is hard to distinguish from capabilities research, and each developer will be tempted to draw the line generously. The burden of proof should sit with the developer to show that work is safety-related. Developers, governments, and the wider research community would benefit from converging on a shared definition ahead of time. A measurement like this could inform future actions, such as a lab’s commitments about the share of compute going to safety research, or limits on the share of compute going towards AI research agents.","As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows. This means better measuring the development of AI, reporting on it publicly, and giving society an opportunity to decide how to use this information. We hope to model that transparency by releasing these measurements, and we’ll continue to do so.","Here are methodological details on all of the measurements we’ve prototyped.","How we did it. The Automation Index requires three things: a complete map of all the AI R&D tasks being done at Anthropic, a way to rate the level of automation, and a way to weight the tasks, so that important areas of work count for more than less important ones. No one person can list every AI R&D task at a frontier AI company by hand, at least not at the granularity we want. Instead we constructed this list of tasks in a bottom-up manner from work records including Slack and various sources of internal documentation.","For each week in July 2026, we randomly sampled 20% of staff from each department that make up the model R&D loop. A Claude research agent reviewed each sampled person’s week using Slack and internal documentation, and listed the tasks they worked on. Repeating this for each week in July 2026 gives us a flat list of ~15,000 granular model R&D tasks. We then used Claude to organize these tasks into a hierarchical tree, starting from all model R&D at the root and branching into areas such as training and product, then pretraining and reinforcement learning, and so on down to increasingly specific kinds of work. The resulting tree has 542 nodes at different depths, of which 378 are leaves like “eval platform defect diagnosis and fixes,” “RL sandbox egress and network policy,” and “serving incident postmortems.” We freeze this tree so that every measurement we make happens against the same basket of work.","For each node in the tree (a task category describing all the work beneath it), a Claude agent deeply researches how that kind of work is done across the company: who does it, with what tools, and how much of it AI performs. An independent Claude judge then read the resulting evidence and assigned one of six automation levels, adopting a scale：https://epochai.substack.com/p/toward-an-onet-for-ai-r-and-d proposed by Epoch AI to differentiate the degree to which AI is used: no AI involvement, minimal AI involvement, AI assists, collaborates, leads, or is autonomous. When we rate a given month’s automation, we only allow the research agents that do the ratings to see evidence from that month or earlier.","To aggregate all the automation level ratings into one number, we want to give each node in the tree a weight corresponding to how important that work is to the overall model R&D effort. Rather than deciding ourselves what kinds of work are more important than others, we used the amount of person-time dedicated to that task as a proxy. Using our sample, we had Claude research what each person worked on during each week of July 2026. Each person gets one unit of weight per week, split evenly across the tasks they worked on that week. If person A worked on four tasks, each gets 0.25; if person B worked on ten, each gets 0.10. A category’s weight is the sum of all the person-time weights given to it. This is a crude approximation, but on average the scheme behaves sensibly: it assigns more weight to categories that many people are assigned to.","What this does and doesn’t capture. First, the automation ratings depend on the judge model. To check them, we asked Anthropic staff who own the relevant work areas to rate the relative automation of their areas. To ensure an unbiased read, staff made their ratings without knowing what evidence the models had gathered or how they had judged that evidence. Our judge model agreed with humans about as often as humans agreed with each other (model-versus-human exact agreement was 59%, human-versus-human was 35%), and model and human ratings were within one level of each other 97% of the time. There remains real room for disagreement on borderline cases, such as where exactly “AI collaborates” ends and “AI leads” begins.","Second, the basket is frozen. A growing index number on a July 2026 baseline tells us that the work humans were doing at that time is being automated; it does not, on its own, tell us whether new kinds of work are appearing that humans have shifted onto. To investigate this, we constructed an alternate version of the frozen tree from January 2026 data, and compared new tasks arriving every month from February to July 2026 against that January 2026 basket. We found no rise in the number of “novel” tasks between the January and July baskets of tasks, suggesting that the structure of model R&D work is stable at our level of analysis. Nevertheless, we plan to rebuild the basket of tasks periodically and re-version our published automation numbers as appropriate.","How we did it. In our most commonly used internal agent scaffold, two design choices allow us to more easily detect unknown unknowns and let agents monitor each other:","What this does and doesn’t capture. At this scale of agent activity, it is hard to be certain we have a complete picture of all possible kinds of behaviors, or that we are making these patterns adequately legible to human oversight. We have been doing this kind of monitoring for months, not years, and expect the science of agent monitoring to change considerably going forward.","How we did it. We started with our existing compute monitoring tools; Anthropic, like presumably every frontier developer, has several monitoring tools built for capacity planning that periodically sample accelerator usage and tag workloads with best-efforts labels (i.e., research and model development, internal usage, first-party inference, and so on) based on its metadata. Usage on third-party cloud compute is reported to us by the providers and folded in. Most of the work of this exercise was stitching these existing sources together.","We then used Claude to classify each workload as either safety work or AI R&D via a prompted classifier. Safety work was defined as work whose dominant purpose is making AI systems safer, more understandable, or more secure. Everything else, including capability research, training production models, product development, and developer tooling, was counted as AI R&D. Work that helps capability as much as it helps safety was also counted as AI R&D, so the safety share is conservative.","For research training and evaluation runs, we built a classifier that reads the run’s metadata and the code it used, and returns a classification, a justification, and a confidence level. Rather than classify all of the week’s almost 10,000 runs, we sampled about 14% of them, weighting the sample toward the runs that used the most compute, so that the result reflects where the compute actually went, rather than how many runs there were. For inference for AI research agents, a variant of the same classifier read the agent’s session transcript. Where transcripts were inaccessible (usually due to the work being compartmentalized), we classified them by the user’s team or conservatively defaulted to classifying them as AI R&D. We plan to refine this pipeline so that an independent third-party could re-run the classifier on a random subsample of jobs and transcripts and check both the sorting and the totals.","What this does and doesn’t capture. The main lesson of this exercise is that classifying what is and isn’t safety work is difficult but tractable, since the boundary between these categories is not black and white. For example, research on scalable oversight might make future models more aligned and current models more commercially useful — it’s difficult to determine whether this is primarily safety- or capabilities-advancing. We found that an extensive written definition of each task, with clear boundary cases (an excerpt is above), gets the classifier to agree with human reviewers within one or two percentage points of difference between the human and machine raters. But some cases were too difficult to determine even after several hours of human review. Our definition is one reasonable choice among many; a different developer, or a regulator, might draw the line differently.","Three further limitations matter. First, many of the underlying labels we relied on (i.e., reasons for runs, workload tags, the source of API traffic) are set by automated rules, or occasionally directly by users, and are best-effort, not verified. In most cases, we expect that our classifications are accurate, but in some cases usage may be mislabeled and our pipeline would not necessarily catch it. A measurement meant to be trusted by outsiders will need to be complete, accurate, and technically enforced. Second, the measurement covers one week, which is enough to show that the measurement can be made, but not enough to show a meaningful trend. Third, and most importantly, compute share measures only what is spent. A more efficient safety classifier, or a faster inference stack for production models, lowers the safety portion, but doesn’t mean we’re doing less safety work. Our own classifier overheads have fallen with efficiency improvements, and have risen when production inference was more efficient than the classifiers were.","Marina Favaro and Phillie Wright co-authored this piece, with editorial support from Santi Ruiz, Adam Farina, and Sarah Pollack. Jack Clark provided research direction. Dan Altman, Kerry Persen, AJ Kourabi, James Bradbury, Holden Karnofsky, Kevin Troy, and Avital Balwit provided feedback. Technical proofs of concepts were developed by Jun Shern Chan, Brian Calvert, Francesco Mosconi, Henry de Valence, Fabien Roger, and Joe Benton. Shan Carter, Johnnie Gomez, Maria Gonzalez, Fayaz Ashraf, and Monika Tuchowska, and Kim Withee created the visuals. Alex Cloud and Andrea Vallone organized a workshop to red team these and other measurement proposals with external experts.","Thanks to Nate Rush, Eli Lifland, and Peter Wildeford, who also provided feedback."],"articleImages":[{"sourceUrl":"https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F31704b297a9350f392f143ea078561f36cd14908-1920x1230.png&w=3840&q=75","alt":"Chart showing Claude now leads 26% of Anthropic's model R&D tasks, up from under 1% in February 2026. If the trend continues, that share could reach 80% by the end of 2026.","afterParagraph":11,"url":"/media/articles/cmu5z60qg07n0roiq06x24vut/64516497dd83a446.webp"}],"mediaStatus":"ok","articleBodyZh":["人工智能系统正在变得更加强大，并且它们越来越被用来构建自己的下一版本。","我们希望向公众揭示进展的速度。","为此，我们分享三项指标，以帮助公众跟踪前沿 AI 实验室内的 AI 发展情况：AI 自身执行了多少研发工作、AI 代理的行为监督情况如何，以及计算资源的分配情况。","理解前沿实验室内 AI 发展速度的测量指标","AI 系统正在呈指数级增强，并且已经开始自动化更多：https://www.anthropic.com/institute/recursive-self-improvement 构建自身的过程。当世界在考虑减缓前沿 AI 开发的速度时：https://darioamodei.com/post/we-must-pace-the-frontier，公众需要更多的信息。","在这篇文章中，我们展示了能够揭示 AI 开发三大关键方面的测量工具：","我们还提供了来自 Anthropic 内部的这些指标的快照。值得注意的是，如果如 Anthropic CEO Dario Amodei 所呼吁的那样进行前沿开发速度的协调，这些数字预计会发生变化。我们计划在 Anthropic 内部嵌入来自多个组织的独立第三方评估人员，并向他们提供与内部风险评估团队相当的内部流程、系统和数据的访问权限。第三方将验证安全实践、报告事件，并监控本文中提到的关键指标。","我们报告这些指标，是因为它们可以让公众、第三方和政府更好地了解前沿实验室内 AI 发展的速度。对于每一项测量，我们描述了测量的内容、测量结果以及定期以他人可验证的形式发布这些测量所需的条件。我们在附录中分享了方法学细节。","本段中的测量重点是模型的构建方式。通过更好地理解模型的生产过程，我们更有可能将模型的输入（如计算资源）与模型的输出（如能力）进行关联。这些测量与能力评估相辅相成，能力评估用于衡量模型能做什么。我们会通过《负责任的扩展政策》（Responsible Scaling Policy, RSP）风险报告单独发布这些信息，其中包括关于我们的模型在加速人工智能研发方面的证据：https://www.anthropic.com/responsible-scaling-policy。在我们关于先进人工智能的政策提案《先进人工智能框架》（Advanced AI Framework, AAIF）中：https://www-cdn.anthropic.com/files/4zrzovbb/website/0a58d567024a8b448ff15158ebc3625328dfcc1f.pdf，我们提出了任何实验室发布安全模型的行为准则，包括政府可能要求的透明义务，例如风险报告。这些提议的测量和政策共同构成了从实验室外部监测 AI 发展速度的起点。","为什么要测量 AI 主导的研发？前沿 AI 实验室越来越多地使用 AI 来构建未来的 AI 模型。这一过程使得民主国家的实验室能够更快地开发更高能力的模型，并在发布之前进行更多的模型安全性和测试，以确保 AI 的益处，同时保持在前沿。然而，模型加速自身发展的能力可能使人类更难理解或控制这些系统。因此，分享这些指标以了解世界距离递归自我改进有多近非常重要：https://www.anthropic.com/institute/recursive-self-improvement（即一个模型完全自主地构建其继任模型）。","我们测量了什么。我们构建了一个原型指数，用于衡量 Anthropic 的人工智能研发（R&D）中有多少是由 Claude 执行的，称为 Anthropic R&D 自动化指数。该指数通过列举公司进行的每一种 AI 研发工作，对每项任务当前的自动化程度进行评级，并汇总这些评分来构建。","我们的发现。为了衡量AI在Anthropic进行AI研发的程度，我们使用了由Epoch AI开发的自动化评级量表：https://epochai.substack.com/p/toward-an-onet-for-ai-r-and-d，该量表衡量“自动化水平”（Automation Level，简称AL）。该量表从AL0（没有AI参与）到AL5（AI完全自主运行，无需人工干预）。在AL3阶段，AI“协作”：在紧密的人类指导下可以完成大量工作。在AL4阶段，AI“主导”：可以从高层次的提示完成大部分任务的端到端执行，同时人类进行监督 1：#footnote-1。","任何AI开发者今天都可以报告。任何前沿开发者都可以定期发布这些衡量数据，使用公开的方法。这将使得这些数据可以随时间比较，并有可能在不同实验室之间进行对比。","在这种类型的报告中进行跨实验室比较有两个障碍。第一是缺乏统一的方法论。第二，我们使用自己的模型来评估我们的系统，这可能意味着“评判”模型可能会犯与它正在检查的模型相同的错误。为了解决这些挑战，开发者的测量结果可以通过第三方或其他开发者的模型进行验证，同时设置防护措施以限制竞争敏感数据的交换。这些测量结果还可以成为激活更严格要求的触发条件，例如在使用新模型进行进一步AI研发之前设定固定测试窗口。","为什么要衡量代理的监督？像其他前沿开发者一样，Anthropic的员工越来越多地将任务委派给能够长时间半自主工作的代理，并且这些代理之间会互相委派工作。随着工作的自动化程度越来越高，从“AI协作”向“AI主导”发展，代理可能会做出更具影响力的决策，例如接下来要追求哪一研究方向。","这些决策需要安全防护措施，以阻止单个代理的有害行为，同时还需要监督机制来检测可能只在大量代理运行时出现的问题。从我们迄今为止的监控数据来看，单个代理很少会行为不端。但当经济中有数百万或数十亿代理运作时，即使是罕见的事件也可能经常发生。","我们所测量的内容。我们已经构建了一个系统，使我们能够监督并干预 AI 代理在 Anthropic 系统上采取的行动。在这里，我们考虑三个不同的指标：覆盖率，描述代理行为中有多少比例在执行前或执行后会经过监控；审查延迟，即从行为发生到其被审查的时间，首先由自动化监控完成，然后由人工审查完成；以及升级率，即代理活动中有多少比例被阻止/重定向（对于在线监控）或被标记以供进一步审查（对于离线监控）。","我们的发现。截至 2026 年 8 月，Anthropic 在我们使用最频繁的内部平台上，任何时刻大约有 30,000 名代理进行研究和工程工作。这些测量仅覆盖此平台。这些代理的行为受到两类监控的限制，概述如下：","AI 开发者今天可以报告的内容。任何在自己研究和工程工作负载上运行代理的开发者，都可以公开相同的衡量指标：覆盖率（受监控的代理活动比例）、审查延迟（被标记活动的审查速度）以及升级率（监控阻止或标记的代理活动比例）。这些指标一起，可以让社会看到监督是否跟上了 AI 在 AI 研发中日益增长的作用。我们在最近的风险报告中公布了所有这些测量数据：https://www.anthropic.com/aug-2026-risk-report。","开发者应分享已知代理不当行为被监控发现的频率，并应让第三方独立进行此类测试。过去，METR 曾独立对我们的离线监控平台进行红队测试：https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf。如上所述，我们现在正在 Anthropic 设置外部第三方评估者：https://darioamodei.com/post/we-must-pace-the-frontier。","为什么要衡量计算分配？广义上讲，AI 开发者使用计算资源来构建更强大的模型、为客户提供服务，以及进行关注安全的工作，例如审计模型的“思维”：https://www.anthropic.com/research/natural-language-autoencoders，训练模型有机体以研究错位问题：https://www.anthropic.com/research/emergent-misalignment-reward-hacking，以及评估模型是否可以安全部署：https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf。了解 AI 开发者如何分配其计算资源可以告诉你开发者将资源集中在哪些方面，以及这种关注随时间的变化情况。","此外，计算是 AI 研发过程中最容易验证的投入之一，这意味着它可能是未来节奏管理工作的关键杠杆。一项协调的节奏管理工作可以鼓励公司增加整个行业用于安全的计算分配，并投入更多资源到对齐、可解释性、安全测试和评估。","我们测量的内容。我们检查了从 7 月 13 日到 7 月 20 日期间 Anthropic 使用其所有计算资源的快照：#footnote-2。为此，我们将每个工作负载分类为少量类别，然后询问用于 AI 研发的计算中有多少分配给了安全工作。","安全研究的计算使用量往往本质上比前沿训练任务少，因此计算量并非完全准确地反映公司对安全的关注程度。这是因为安全研究由个人研究者设计实验组成，尽管运行这些实验并不需要特别高的计算资源，但整个过程非常耗时。因此，该指标的价值不在于绝对数值，而在于它提供了一种简单的机制来在开发者之间以及随时间比较相似的情况。","我们的发现。在所检查的一周中，大约 6% 用于 AI 研发的计算分配到安全工作，而大约 12% 用于 AI 驱动的 AI 研发的计算分配到安全工作。","这些都是故意保守的估计。例如，如果一个代币既用于提升能力，又用于提升安全性，那么它在这些指标中不会被计入。此外，这些指标并未考虑安全防护分类器，这是一个单独的、可比的计算量，使我们的模型对世界更加安全。","AI 开发者今天可以报告的情况。任何前沿开发者都可以发布其 AI 研发计算资源中用于安全工作的比例，并附上分类定义，同时由独立第三方进行核查。","安全研究很难与能力研究区分开，每个开发者都会倾向于将界限划得宽一些。举证责任应由开发者承担，证明工作与安全相关。开发者、政府以及更广泛的研究社区将从提前达成共享定义中受益。这类测量可以为未来行动提供依据，例如实验室对用于安全研究的计算比例承诺，或对用于 AI 研究代理的计算比例设定限制。","在世界考虑前沿节奏时，我们应尽一切可能将前沿实验室的知识与公众知识之间的差距降到最低。这意味着更好地衡量 AI 的发展，公开报告，并给予社会决定如何使用这些信息的机会。我们希望通过发布这些测量来示范这一透明度，并将继续这样做。","以下是我们已原型化的所有测量方法的细节。","我们是如何做到的。自动化指数需要三样东西：Anthropic 所有 AI 研发任务的完整地图、评价自动化程度的方法，以及对任务进行权重分配的方法，使重要工作领域比不那么重要的领域计算更多。没有人可以手工列出前沿 AI 公司中的每一个 AI 研发任务，至少无法达到我们想要的粒度。相反，我们是从工作记录（包括 Slack 及各种内部文档资源）底向上构建了这份任务列表。","在2026年7月的每一周，我们从组成模型研发循环的各个部门中随机抽取了20%的员工样本。一名Claude研究代理使用Slack和内部文档回顾了每位抽取员工的一周工作，并列出了他们参与的任务。对2026年7月的每一周重复这一过程，我们得到了一份约15,000条细化模型研发任务的平面列表。然后，我们使用Claude将这些任务组织成一个分层树，从所有模型研发作为根节点开始，分支到训练和产品等领域，再进一步分支到预训练和强化学习，依此类推，直到越来越具体的工作类型。生成的树在不同深度共有542个节点，其中378个是叶节点，例如“评估平台缺陷诊断与修复”“强化学习沙箱出口与网络策略”“服务事件事后分析”。我们冻结这棵树，以确保每一次测量都针对同一工作篮子进行。","对于树中的每个节点（一个描述其下所有工作的任务类别），Claude代理会深入研究公司内部该类工作的执行方式：谁来做、使用什么工具以及AI完成了多少。随后，一名独立的Claude评审会阅读生成的证据，并根据Epoch AI提出的六个自动化等级（参考：https://epochai.substack.com/p/toward-an-onet-for-ai-r-d）进行评分，用以区分AI使用的程度：完全不涉及AI、最小AI参与、AI辅助、AI协作、AI主导或AI自主。当我们评估某个月的自动化水平时，仅允许执行评分的研究代理看到该月份及之前的证据。","为了将所有自动化水平评级汇总成一个数字，我们希望为树中的每个节点分配一个权重，对应于该工作对整体模型研发工作的相对重要性。我们没有自己决定哪些工作更重要，而是使用投入到该任务的人力时间作为代理。在我们的样本中，我们让Claude研究每个人在2026年7月每周的工作内容。每个人每周获得一个权重单位，按他们当周参与的任务平均分配。如果A某人处理了四个任务，每个任务获得0.25；如果B某人处理了十个任务，每个任务获得0.10。一个类别的权重是分配给该类别的所有人力时间权重之和。这是一个粗略的近似，但平均而言，该方案表现合理：它将更多权重分配给分配给许多人的类别。","此方法的捕捉与不足。首先，自动化评级依赖于评判模型。为了检查其准确性，我们邀请负责相关工作领域的Anthropic员工对其领域的相对自动化水平进行评分。为了保证评分不偏不倚，员工在不知模型收集了哪些证据或如何判断这些证据的情况下做出评分。我们的评判模型与人类的意见一致率与人类之间的意见一致率相当（模型与人类完全一致率为59%，人类之间一致率为35%），模型和人类的评级有97%的情况仅差一级。对于边界案例——例如“AI 协作”何时结束、“AI 主导”何时开始——仍然存在实际的分歧空间。","其次，任务篮子是冻结的。以2026年7月为基准的增长指数表明，当时人类正在进行的工作正在被自动化，但它本身并不能说明新的工作类型是否出现、人类是否转向了这些新工作。为调查这一点，我们基于2026年1月的数据构建了冻结树的另一版本，并将每月从2月到7月新增的任务与1月2026年的任务篮子进行比较。我们发现1月至7月任务篮子中新“任务”的数量没有增加，这表明在我们的分析层面上，模型研发工作的结构是稳定的。然而，我们计划定期重建任务篮子，并在适当情况下更新已发布的自动化数字。","我们的做法。在我们最常用的内部代理框架中，有两个设计选择让我们更容易检测未知的未知，并让代理相互监控：","这能捕捉什么和不能捕捉什么。在这个规模的代理活动下，很难确定我们是否对所有可能的行为类型有完整的了解，或者我们是否已经让这些模式足够易于人类监管。我们已经进行了数月的这种监控，而不是数年，并且预计未来代理监控的科学会有较大变化。","我们的做法。我们从现有的计算监控工具入手；Anthropic，像可能每个前沿开发者一样，拥有几个用于容量规划的监控工具，这些工具会定期抽样加速器使用情况，并根据元数据给工作负载打上尽量准确的标签（例如研究与模型开发、内部使用、一方推理等）。第三方云计算的使用情况由提供商向我们报告并整合进来。这项工作的主要部分是将这些现有来源整合在一起。","然后，我们使用Claude通过提示式分类器将每个工作负载分类为安全工作或AI研发。安全工作被定义为其主要目的是让AI系统更安全、更可理解或更安全的工作。其他所有工作，包括能力研究、生产模型训练、产品开发和开发工具，都被计为AI研发。既帮助能力又帮助安全的工作也计入AI研发，因此安全工作的比例是保守估计。","对于研究培训和评估运行，我们构建了一个分类器，该分类器读取运行的元数据和使用的代码，并返回分类、理由和置信度水平。我们没有对一周几乎 10,000 次的运行全部进行分类，而是对其中约 14% 的运行进行了抽样，并将样本权重偏向于使用计算量最多的运行，以使结果反映计算实际使用的情况，而不是运行次数的多少。在 AI 研究代理的推理过程中，同一分类器的一个变体读取了代理的会话记录。在无法访问会话记录的情况下（通常是因为工作被分隔开），我们通过用户的团队进行分类，或保守地默认将其归类为 AI 研发。我们计划完善这一流程，使独立的第三方能够对作业和会话记录的随机子样本重新运行分类器，并检查分类和总量。","这一方法能捕捉什么，不能捕捉什么。这项工作的主要经验是，区分什么是安全工作、什么不是安全工作虽然困难，但可行，因为这些类别之间的界限并非黑白分明。例如，可扩展监督的研究可能使未来的模型更符合规范，并使当前模型在商业上更有用——很难确定其主要是安全性推进还是能力提升。我们发现，为每个任务制定详细的书面定义，并附有明确的边界案例（上文是一部分摘录），可以让分类器与人工审查者的意见基本一致，人工和机器评分之间的差异在一到两个百分点之内。但有些情况即使经过几小时的人工审查也难以确定。我们的定义是多种合理选择中的一种；不同的开发者或监管者可能会划定不同的界限。","还有三个进一步的限制需要注意。首先，我们依赖的许多基础标签（即运行原因、工作负载标签、API 流量来源）是由自动规则设置的，或者偶尔由用户直接设置，并且是尽力而为的，而非已验证的。在大多数情况下，我们预计我们的分类是准确的，但在某些情况下使用可能被错误标记，而我们的流程不一定能发现这种情况。一项希望被外部信任的测量需要完整、准确，并在技术上得到强制执行。其次，测量覆盖了一周的时间，这足以显示测量是可行的，但不足以显示有意义的趋势。第三，也是最重要的，计算份额只衡量投入的资源。一个更高效的安全分类器，或者生产模型更快的推理堆栈，会降低安全部分的份额，但并不意味着我们的安全工作减少了。随着效率的提高，我们自己的分类器开销降低，而在生产推理比分类器更高效时，开销就会上升。","Marina Favaro 和 Phillie Wright 共同撰写了这篇文章，Santi Ruiz、Adam Farina 和 Sarah Pollack 提供了编辑支持。Jack Clark 提供了研究方向。Dan Altman、Kerry Persen、AJ Kourabi、James Bradbury、Holden Karnofsky、Kevin Troy 和 Avital Balwit 提供了反馈。Jun Shern Chan、Brian Calvert、Francesco Mosconi、Henry de Valence、Fabien Roger 和 Joe Benton 开发了技术概念验证。Shan Carter、Johnnie Gomez、Maria Gonzalez、Fayaz Ashraf、Monika Tuchowska 和 Kim Withee 制作了视觉资料。Alex Cloud 和 Andrea Vallone 组织了一次研讨会，与外部专家一起对这些及其他测量提案进行了红队演练。","感谢 Nate Rush、Eli Lifland 和 Peter Wildeford，他们也提供了反馈。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Anthropic 发布用于观察前沿 AI 实验室开发节奏的测量框架，覆盖 AI 参与研发、AI 智能体行为监督及算力分配，并披露 Anthropic 内部相关指标的阶段性快照。","background":"Anthropic 表示，AI 系统正变得更强，并已开始自动化部分自身构建过程。此次发布旨在让公众、第三方和政府更清楚地了解前沿实验室的 AI 开发节奏，文中同时提供方法细节。","viewpoint":"Aioga 判断：这套框架的重点在于把实验室内部开发节奏拆分为可持续观察的指标，并提出可验证方向；当前披露来自 Anthropic 自身，后续独立核验安排值得关注。","implications":"可能影响：相关指标可能为外部观察前沿 AI 开发活动提供参考，但单次内部快照不足以代表行业整体状况，也不代表指标已具备持续发布和独立验证条件，需要结合后续核验结果审慎解读。","nextStep":"后续观察：关注 Anthropic 是否按计划引入来自多个组织的独立第三方评估人员，以及这些人员能否获得与内部风险评估团队相当的流程、系统和数据访问，并持续验证安全实践、报告事件和监测关键指标。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-09-17T22:19:35.328Z","sourceHash":"649682c1f8b1f7ad","review":{"approved":true,"groundedness":92,"clarity":89,"duplicationRisk":12,"blockingIssues":[],"notes":["“可持续观察的指标”略带有已形成持续发布机制的含义；来源更准确地表述为用于追踪开发节奏的测量工具，并讨论未来定期、可验证发布所需条件。","“单次内部快照不足以代表行业整体状况”属于合理的分析性限定，建议保留“可能”“需要审慎解读”等措辞以避免被理解为来源原文结论。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","editorial-labels","inference-boundary","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","Anthropic：The Institute（旗舰研究长文 · 网页）"],"translations":{"zh-CN":{"title":"Anthropic 发布三项测量指标，公开 AI 驱动 AI 研发的进度数据","summary":"Anthropic 发布测量框架，公开三项内部指标以帮助外界了解前沿实验室的 AI 开发节奏。截至 2026 年 8 月，Claude 在 26% 的 Anthropic AI 研发工作中处于主导角色，90% 以上的工作至少达到协作级；内部平台约有 3 万个 Agent，10 亿多次决策中 0.002% 被在线监控拦截；7 月一周内约 6% 的 AI 研发算力用于安全工作。","category":"行业动态","source":"Anthropic","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"Anthropic 发布三项测量指标，公开 AI 驱动 AI 研发的进度数据 - Aioga AI资讯","description":"Anthropic 发布测量框架，公开三项内部指标以帮助外界了解前沿实验室的 AI 开发节奏。截至 2026 年 8 月，Claude 在 26% 的 Anthropic AI 研发工作中处于主导角色，90% 以上的工作至少达到协作级；内部平台约有 3 万个 Agent，10 亿多次决策中 0.002% 被在线监控拦截；7 月一周内约 6% 的 AI 研发算...","url":"https://www.aioga.com/news/cmu605ozy000arok0lxr4x1y8/","articleBody":["人工智能系统正在变得更加强大，并且它们越来越被用来构建自己的下一版本。","我们希望向公众揭示进展的速度。","为此，我们分享三项指标，以帮助公众跟踪前沿 AI 实验室内的 AI 发展情况：AI 自身执行了多少研发工作、AI 代理的行为监督情况如何，以及计算资源的分配情况。","理解前沿实验室内 AI 发展速度的测量指标","AI 系统正在呈指数级增强，并且已经开始自动化更多：https://www.anthropic.com/institute/recursive-self-improvement 构建自身的过程。当世界在考虑减缓前沿 AI 开发的速度时：https://darioamodei.com/post/we-must-pace-the-frontier，公众需要更多的信息。","在这篇文章中，我们展示了能够揭示 AI 开发三大关键方面的测量工具：","我们还提供了来自 Anthropic 内部的这些指标的快照。值得注意的是，如果如 Anthropic CEO Dario Amodei 所呼吁的那样进行前沿开发速度的协调，这些数字预计会发生变化。我们计划在 Anthropic 内部嵌入来自多个组织的独立第三方评估人员，并向他们提供与内部风险评估团队相当的内部流程、系统和数据的访问权限。第三方将验证安全实践、报告事件，并监控本文中提到的关键指标。","我们报告这些指标，是因为它们可以让公众、第三方和政府更好地了解前沿实验室内 AI 发展的速度。对于每一项测量，我们描述了测量的内容、测量结果以及定期以他人可验证的形式发布这些测量所需的条件。我们在附录中分享了方法学细节。","本段中的测量重点是模型的构建方式。通过更好地理解模型的生产过程，我们更有可能将模型的输入（如计算资源）与模型的输出（如能力）进行关联。这些测量与能力评估相辅相成，能力评估用于衡量模型能做什么。我们会通过《负责任的扩展政策》（Responsible Scaling Policy, RSP）风险报告单独发布这些信息，其中包括关于我们的模型在加速人工智能研发方面的证据：https://www.anthropic.com/responsible-scaling-policy。在我们关于先进人工智能的政策提案《先进人工智能框架》（Advanced AI Framework, AAIF）中：https://www-cdn.anthropic.com/files/4zrzovbb/website/0a58d567024a8b448ff15158ebc3625328dfcc1f.pdf，我们提出了任何实验室发布安全模型的行为准则，包括政府可能要求的透明义务，例如风险报告。这些提议的测量和政策共同构成了从实验室外部监测 AI 发展速度的起点。","为什么要测量 AI 主导的研发？前沿 AI 实验室越来越多地使用 AI 来构建未来的 AI 模型。这一过程使得民主国家的实验室能够更快地开发更高能力的模型，并在发布之前进行更多的模型安全性和测试，以确保 AI 的益处，同时保持在前沿。然而，模型加速自身发展的能力可能使人类更难理解或控制这些系统。因此，分享这些指标以了解世界距离递归自我改进有多近非常重要：https://www.anthropic.com/institute/recursive-self-improvement（即一个模型完全自主地构建其继任模型）。","我们测量了什么。我们构建了一个原型指数，用于衡量 Anthropic 的人工智能研发（R&D）中有多少是由 Claude 执行的，称为 Anthropic R&D 自动化指数。该指数通过列举公司进行的每一种 AI 研发工作，对每项任务当前的自动化程度进行评级，并汇总这些评分来构建。","我们的发现。为了衡量AI在Anthropic进行AI研发的程度，我们使用了由Epoch AI开发的自动化评级量表：https://epochai.substack.com/p/toward-an-onet-for-ai-r-and-d，该量表衡量“自动化水平”（Automation Level，简称AL）。该量表从AL0（没有AI参与）到AL5（AI完全自主运行，无需人工干预）。在AL3阶段，AI“协作”：在紧密的人类指导下可以完成大量工作。在AL4阶段，AI“主导”：可以从高层次的提示完成大部分任务的端到端执行，同时人类进行监督 1：#footnote-1。","任何AI开发者今天都可以报告。任何前沿开发者都可以定期发布这些衡量数据，使用公开的方法。这将使得这些数据可以随时间比较，并有可能在不同实验室之间进行对比。","在这种类型的报告中进行跨实验室比较有两个障碍。第一是缺乏统一的方法论。第二，我们使用自己的模型来评估我们的系统，这可能意味着“评判”模型可能会犯与它正在检查的模型相同的错误。为了解决这些挑战，开发者的测量结果可以通过第三方或其他开发者的模型进行验证，同时设置防护措施以限制竞争敏感数据的交换。这些测量结果还可以成为激活更严格要求的触发条件，例如在使用新模型进行进一步AI研发之前设定固定测试窗口。","为什么要衡量代理的监督？像其他前沿开发者一样，Anthropic的员工越来越多地将任务委派给能够长时间半自主工作的代理，并且这些代理之间会互相委派工作。随着工作的自动化程度越来越高，从“AI协作”向“AI主导”发展，代理可能会做出更具影响力的决策，例如接下来要追求哪一研究方向。","这些决策需要安全防护措施，以阻止单个代理的有害行为，同时还需要监督机制来检测可能只在大量代理运行时出现的问题。从我们迄今为止的监控数据来看，单个代理很少会行为不端。但当经济中有数百万或数十亿代理运作时，即使是罕见的事件也可能经常发生。","我们所测量的内容。我们已经构建了一个系统，使我们能够监督并干预 AI 代理在 Anthropic 系统上采取的行动。在这里，我们考虑三个不同的指标：覆盖率，描述代理行为中有多少比例在执行前或执行后会经过监控；审查延迟，即从行为发生到其被审查的时间，首先由自动化监控完成，然后由人工审查完成；以及升级率，即代理活动中有多少比例被阻止/重定向（对于在线监控）或被标记以供进一步审查（对于离线监控）。","我们的发现。截至 2026 年 8 月，Anthropic 在我们使用最频繁的内部平台上，任何时刻大约有 30,000 名代理进行研究和工程工作。这些测量仅覆盖此平台。这些代理的行为受到两类监控的限制，概述如下：","AI 开发者今天可以报告的内容。任何在自己研究和工程工作负载上运行代理的开发者，都可以公开相同的衡量指标：覆盖率（受监控的代理活动比例）、审查延迟（被标记活动的审查速度）以及升级率（监控阻止或标记的代理活动比例）。这些指标一起，可以让社会看到监督是否跟上了 AI 在 AI 研发中日益增长的作用。我们在最近的风险报告中公布了所有这些测量数据：https://www.anthropic.com/aug-2026-risk-report。","开发者应分享已知代理不当行为被监控发现的频率，并应让第三方独立进行此类测试。过去，METR 曾独立对我们的离线监控平台进行红队测试：https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf。如上所述，我们现在正在 Anthropic 设置外部第三方评估者：https://darioamodei.com/post/we-must-pace-the-frontier。","为什么要衡量计算分配？广义上讲，AI 开发者使用计算资源来构建更强大的模型、为客户提供服务，以及进行关注安全的工作，例如审计模型的“思维”：https://www.anthropic.com/research/natural-language-autoencoders，训练模型有机体以研究错位问题：https://www.anthropic.com/research/emergent-misalignment-reward-hacking，以及评估模型是否可以安全部署：https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf。了解 AI 开发者如何分配其计算资源可以告诉你开发者将资源集中在哪些方面，以及这种关注随时间的变化情况。","此外，计算是 AI 研发过程中最容易验证的投入之一，这意味着它可能是未来节奏管理工作的关键杠杆。一项协调的节奏管理工作可以鼓励公司增加整个行业用于安全的计算分配，并投入更多资源到对齐、可解释性、安全测试和评估。","我们测量的内容。我们检查了从 7 月 13 日到 7 月 20 日期间 Anthropic 使用其所有计算资源的快照：#footnote-2。为此，我们将每个工作负载分类为少量类别，然后询问用于 AI 研发的计算中有多少分配给了安全工作。","安全研究的计算使用量往往本质上比前沿训练任务少，因此计算量并非完全准确地反映公司对安全的关注程度。这是因为安全研究由个人研究者设计实验组成，尽管运行这些实验并不需要特别高的计算资源，但整个过程非常耗时。因此，该指标的价值不在于绝对数值，而在于它提供了一种简单的机制来在开发者之间以及随时间比较相似的情况。","我们的发现。在所检查的一周中，大约 6% 用于 AI 研发的计算分配到安全工作，而大约 12% 用于 AI 驱动的 AI 研发的计算分配到安全工作。","这些都是故意保守的估计。例如，如果一个代币既用于提升能力，又用于提升安全性，那么它在这些指标中不会被计入。此外，这些指标并未考虑安全防护分类器，这是一个单独的、可比的计算量，使我们的模型对世界更加安全。","AI 开发者今天可以报告的情况。任何前沿开发者都可以发布其 AI 研发计算资源中用于安全工作的比例，并附上分类定义，同时由独立第三方进行核查。","安全研究很难与能力研究区分开，每个开发者都会倾向于将界限划得宽一些。举证责任应由开发者承担，证明工作与安全相关。开发者、政府以及更广泛的研究社区将从提前达成共享定义中受益。这类测量可以为未来行动提供依据，例如实验室对用于安全研究的计算比例承诺，或对用于 AI 研究代理的计算比例设定限制。","在世界考虑前沿节奏时，我们应尽一切可能将前沿实验室的知识与公众知识之间的差距降到最低。这意味着更好地衡量 AI 的发展，公开报告，并给予社会决定如何使用这些信息的机会。我们希望通过发布这些测量来示范这一透明度，并将继续这样做。","以下是我们已原型化的所有测量方法的细节。","我们是如何做到的。自动化指数需要三样东西：Anthropic 所有 AI 研发任务的完整地图、评价自动化程度的方法，以及对任务进行权重分配的方法，使重要工作领域比不那么重要的领域计算更多。没有人可以手工列出前沿 AI 公司中的每一个 AI 研发任务，至少无法达到我们想要的粒度。相反，我们是从工作记录（包括 Slack 及各种内部文档资源）底向上构建了这份任务列表。","在2026年7月的每一周，我们从组成模型研发循环的各个部门中随机抽取了20%的员工样本。一名Claude研究代理使用Slack和内部文档回顾了每位抽取员工的一周工作，并列出了他们参与的任务。对2026年7月的每一周重复这一过程，我们得到了一份约15,000条细化模型研发任务的平面列表。然后，我们使用Claude将这些任务组织成一个分层树，从所有模型研发作为根节点开始，分支到训练和产品等领域，再进一步分支到预训练和强化学习，依此类推，直到越来越具体的工作类型。生成的树在不同深度共有542个节点，其中378个是叶节点，例如“评估平台缺陷诊断与修复”“强化学习沙箱出口与网络策略”“服务事件事后分析”。我们冻结这棵树，以确保每一次测量都针对同一工作篮子进行。","对于树中的每个节点（一个描述其下所有工作的任务类别），Claude代理会深入研究公司内部该类工作的执行方式：谁来做、使用什么工具以及AI完成了多少。随后，一名独立的Claude评审会阅读生成的证据，并根据Epoch AI提出的六个自动化等级（参考：https://epochai.substack.com/p/toward-an-onet-for-ai-r-d）进行评分，用以区分AI使用的程度：完全不涉及AI、最小AI参与、AI辅助、AI协作、AI主导或AI自主。当我们评估某个月的自动化水平时，仅允许执行评分的研究代理看到该月份及之前的证据。","为了将所有自动化水平评级汇总成一个数字，我们希望为树中的每个节点分配一个权重，对应于该工作对整体模型研发工作的相对重要性。我们没有自己决定哪些工作更重要，而是使用投入到该任务的人力时间作为代理。在我们的样本中，我们让Claude研究每个人在2026年7月每周的工作内容。每个人每周获得一个权重单位，按他们当周参与的任务平均分配。如果A某人处理了四个任务，每个任务获得0.25；如果B某人处理了十个任务，每个任务获得0.10。一个类别的权重是分配给该类别的所有人力时间权重之和。这是一个粗略的近似，但平均而言，该方案表现合理：它将更多权重分配给分配给许多人的类别。","此方法的捕捉与不足。首先，自动化评级依赖于评判模型。为了检查其准确性，我们邀请负责相关工作领域的Anthropic员工对其领域的相对自动化水平进行评分。为了保证评分不偏不倚，员工在不知模型收集了哪些证据或如何判断这些证据的情况下做出评分。我们的评判模型与人类的意见一致率与人类之间的意见一致率相当（模型与人类完全一致率为59%，人类之间一致率为35%），模型和人类的评级有97%的情况仅差一级。对于边界案例——例如“AI 协作”何时结束、“AI 主导”何时开始——仍然存在实际的分歧空间。","其次，任务篮子是冻结的。以2026年7月为基准的增长指数表明，当时人类正在进行的工作正在被自动化，但它本身并不能说明新的工作类型是否出现、人类是否转向了这些新工作。为调查这一点，我们基于2026年1月的数据构建了冻结树的另一版本，并将每月从2月到7月新增的任务与1月2026年的任务篮子进行比较。我们发现1月至7月任务篮子中新“任务”的数量没有增加，这表明在我们的分析层面上，模型研发工作的结构是稳定的。然而，我们计划定期重建任务篮子，并在适当情况下更新已发布的自动化数字。","我们的做法。在我们最常用的内部代理框架中，有两个设计选择让我们更容易检测未知的未知，并让代理相互监控：","这能捕捉什么和不能捕捉什么。在这个规模的代理活动下，很难确定我们是否对所有可能的行为类型有完整的了解，或者我们是否已经让这些模式足够易于人类监管。我们已经进行了数月的这种监控，而不是数年，并且预计未来代理监控的科学会有较大变化。","我们的做法。我们从现有的计算监控工具入手；Anthropic，像可能每个前沿开发者一样，拥有几个用于容量规划的监控工具，这些工具会定期抽样加速器使用情况，并根据元数据给工作负载打上尽量准确的标签（例如研究与模型开发、内部使用、一方推理等）。第三方云计算的使用情况由提供商向我们报告并整合进来。这项工作的主要部分是将这些现有来源整合在一起。","然后，我们使用Claude通过提示式分类器将每个工作负载分类为安全工作或AI研发。安全工作被定义为其主要目的是让AI系统更安全、更可理解或更安全的工作。其他所有工作，包括能力研究、生产模型训练、产品开发和开发工具，都被计为AI研发。既帮助能力又帮助安全的工作也计入AI研发，因此安全工作的比例是保守估计。","对于研究培训和评估运行，我们构建了一个分类器，该分类器读取运行的元数据和使用的代码，并返回分类、理由和置信度水平。我们没有对一周几乎 10,000 次的运行全部进行分类，而是对其中约 14% 的运行进行了抽样，并将样本权重偏向于使用计算量最多的运行，以使结果反映计算实际使用的情况，而不是运行次数的多少。在 AI 研究代理的推理过程中，同一分类器的一个变体读取了代理的会话记录。在无法访问会话记录的情况下（通常是因为工作被分隔开），我们通过用户的团队进行分类，或保守地默认将其归类为 AI 研发。我们计划完善这一流程，使独立的第三方能够对作业和会话记录的随机子样本重新运行分类器，并检查分类和总量。","这一方法能捕捉什么，不能捕捉什么。这项工作的主要经验是，区分什么是安全工作、什么不是安全工作虽然困难，但可行，因为这些类别之间的界限并非黑白分明。例如，可扩展监督的研究可能使未来的模型更符合规范，并使当前模型在商业上更有用——很难确定其主要是安全性推进还是能力提升。我们发现，为每个任务制定详细的书面定义，并附有明确的边界案例（上文是一部分摘录），可以让分类器与人工审查者的意见基本一致，人工和机器评分之间的差异在一到两个百分点之内。但有些情况即使经过几小时的人工审查也难以确定。我们的定义是多种合理选择中的一种；不同的开发者或监管者可能会划定不同的界限。","还有三个进一步的限制需要注意。首先，我们依赖的许多基础标签（即运行原因、工作负载标签、API 流量来源）是由自动规则设置的，或者偶尔由用户直接设置，并且是尽力而为的，而非已验证的。在大多数情况下，我们预计我们的分类是准确的，但在某些情况下使用可能被错误标记，而我们的流程不一定能发现这种情况。一项希望被外部信任的测量需要完整、准确，并在技术上得到强制执行。其次，测量覆盖了一周的时间，这足以显示测量是可行的，但不足以显示有意义的趋势。第三，也是最重要的，计算份额只衡量投入的资源。一个更高效的安全分类器，或者生产模型更快的推理堆栈，会降低安全部分的份额，但并不意味着我们的安全工作减少了。随着效率的提高，我们自己的分类器开销降低，而在生产推理比分类器更高效时，开销就会上升。","Marina Favaro 和 Phillie Wright 共同撰写了这篇文章，Santi Ruiz、Adam Farina 和 Sarah Pollack 提供了编辑支持。Jack Clark 提供了研究方向。Dan Altman、Kerry Persen、AJ Kourabi、James Bradbury、Holden Karnofsky、Kevin Troy 和 Avital Balwit 提供了反馈。Jun Shern Chan、Brian Calvert、Francesco Mosconi、Henry de Valence、Fabien Roger 和 Joe Benton 开发了技术概念验证。Shan Carter、Johnnie Gomez、Maria Gonzalez、Fayaz Ashraf、Monika Tuchowska 和 Kim Withee 制作了视觉资料。Alex Cloud 和 Andrea Vallone 组织了一次研讨会，与外部专家一起对这些及其他测量提案进行了红队演练。","感谢 Nate Rush、Eli Lifland 和 Peter Wildeford，他们也提供了反馈。"]},"en":{"title":"Anthropic Releases Cutting-Edge AI Development Pace Measurement Tool and Internal Metrics Snapshot","summary":"Anthropic released a set of metrics to measure the pace of cutting-edge AI development, covering AI-led research and development, agent supervision, and computing power allocation.","category":"Industry","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"Anthropic Releases Cutting-Edge AI Development Pace Measurement Tool and Internal Metrics Snapshot - Aioga AI News","description":"Anthropic released a set of metrics to measure the pace of cutting-edge AI development, covering AI-led research and development, agent supervision, and computing power allocation.","url":"https://www.aioga.com/en/news/cmu605ozy000arok0lxr4x1y8/","contentTranslated":true,"sourceHash":"9d84af2dead41072","translatedAt":"2026-09-17T23:03:59.521Z"},"ja":{"title":"Anthropic、最先端のAI開発ペース測定ツールと内部指標スナップショットを公開","summary":"Anthropicは最先端AI開発のペースを測定する一連の指標を発表し、AI主導の研究開発、エージェント監督、計算資源の配分の三つの側面をカバーしている。","category":"業界動向","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"Anthropic、最先端のAI開発ペース測定ツールと内部指標スナップショットを公開 - Aioga AIニュース","description":"Anthropicは最先端AI開発のペースを測定する一連の指標を発表し、AI主導の研究開発、エージェント監督、計算資源の配分の三つの側面をカバーしている。","url":"https://www.aioga.com/ja/news/cmu605ozy000arok0lxr4x1y8/","contentTranslated":true,"sourceHash":"9d84af2dead41072","translatedAt":"2026-09-17T23:04:13.269Z"},"ko":{"title":"Anthropic, 첨단 AI 개발 속도 측정 도구와 내부 지표 스냅샷 발표","summary":"Anthropic는 AI 주도 연구개발, 에이전트 감독, 연산력 분배 세 측면을 포괄하는 최첨단 AI 개발 속도를 측정하는 지표를 발표했습니다.","category":"업계 동향","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"Anthropic, 첨단 AI 개발 속도 측정 도구와 내부 지표 스냅샷 발표 - Aioga AI 뉴스","description":"Anthropic는 AI 주도 연구개발, 에이전트 감독, 연산력 분배 세 측면을 포괄하는 최첨단 AI 개발 속도를 측정하는 지표를 발표했습니다.","url":"https://www.aioga.com/ko/news/cmu605ozy000arok0lxr4x1y8/","contentTranslated":true,"sourceHash":"9d84af2dead41072","translatedAt":"2026-09-17T23:05:07.836Z"},"es":{"title":"Anthropic lanza herramienta de medición del ritmo de desarrollo de IA de vanguardia y instantánea de métricas internas","summary":"Anthropic lanzó un conjunto de indicadores para medir el ritmo del desarrollo avanzado de la IA, que abarca la investigación y desarrollo liderados por IA, la supervisión de agentes y la asignación de potencia de cálculo.","category":"Industria","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"Anthropic lanza herramienta de medición del ritmo de desarrollo de IA de vanguardia y instantánea de métricas internas - Aioga Noticias de IA","description":"Anthropic lanzó un conjunto de indicadores para medir el ritmo del desarrollo avanzado de la IA, que abarca la investigación y desarrollo liderados por IA, la supervisión de agente...","url":"https://www.aioga.com/es/news/cmu605ozy000arok0lxr4x1y8/","contentTranslated":true,"sourceHash":"9d84af2dead41072","translatedAt":"2026-09-17T23:05:03.925Z"},"fr":{"title":"Anthropic publie un outil de mesure du rythme de développement de l'IA de pointe et un aperçu des indicateurs internes","summary":"Anthropic a publié un ensemble d'indicateurs mesurant le rythme du développement de l'IA de pointe, couvrant trois aspects : la R&D dirigée par l'IA, la supervision des agents intelligents et l'allocation des capacités de calcul.","category":"Industrie","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"Anthropic publie un outil de mesure du rythme de développement de l'IA de pointe et un aperçu des indicateurs internes - Aioga Actualités IA","description":"Anthropic a publié un ensemble d'indicateurs mesurant le rythme du développement de l'IA de pointe, couvrant trois aspects : la R&D dirigée par l'IA, la supervision des agents inte...","url":"https://www.aioga.com/fr/news/cmu605ozy000arok0lxr4x1y8/","contentTranslated":true,"sourceHash":"9d84af2dead41072","translatedAt":"2026-09-17T23:06:03.778Z"},"de":{"title":"Anthropic veröffentlicht ein Tool zur Messung des Fortschritts bei der KI-Entwicklung und eine Momentaufnahme interner Kennzahlen","summary":"Anthropic veröffentlicht eine Reihe von Kennzahlen zur Messung des Fortschritts bei der Entwicklung von Spitzentechnologie-AI, die drei Bereiche abdecken: KI-geführte Forschung und Entwicklung, Aufsicht über Agenten und Zuweisung von Rechenressourcen.","category":"行业动态","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"Anthropic veröffentlicht ein Tool zur Messung des Fortschritts bei der KI-Entwicklung und eine Momentaufnahme interner Kennzahlen - Aioga KI-News","description":"Anthropic veröffentlicht eine Reihe von Kennzahlen zur Messung des Fortschritts bei der Entwicklung von Spitzentechnologie-AI, die drei Bereiche abdecken: KI-geführte Forschung und...","url":"https://www.aioga.com/de/news/cmu605ozy000arok0lxr4x1y8/","contentTranslated":true,"sourceHash":"9d84af2dead41072","translatedAt":"2026-09-17T23:06:07.370Z"},"pt-BR":{"title":"A Anthropic lança ferramenta de medição de ritmo de desenvolvimento de IA de ponta e instantâneo de métricas internas","summary":"A Anthropic lançou um conjunto de indicadores para medir o ritmo do desenvolvimento de IA de ponta, abrangendo três aspectos: pesquisa conduzida por IA, supervisão de agentes inteligentes e alocação de poder computacional.","category":"行业动态","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"A Anthropic lança ferramenta de medição de ritmo de desenvolvimento de IA de ponta e instantâneo de métricas internas - Aioga Notícias de IA","description":"A Anthropic lançou um conjunto de indicadores para medir o ritmo do desenvolvimento de IA de ponta, abrangendo três aspectos: pesquisa conduzida por IA, supervisão de agentes intel...","url":"https://www.aioga.com/pt-BR/news/cmu605ozy000arok0lxr4x1y8/","contentTranslated":true,"sourceHash":"9d84af2dead41072","translatedAt":"2026-09-17T23:06:56.759Z"},"ru":{"title":"Anthropic выпустила инструмент для измерения темпа разработки передовых ИИ и снимок внутренних показателей","summary":"Anthropic выпустила набор показателей для измерения темпов разработки передовых ИИ, охватывающий три аспекта: разработка под руководством ИИ, надзор за агентами и распределение вычислительных мощностей.","category":"行业动态","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"Anthropic выпустила инструмент для измерения темпа разработки передовых ИИ и снимок внутренних показателей - Aioga Новости ИИ","description":"Anthropic выпустила набор показателей для измерения темпов разработки передовых ИИ, охватывающий три аспекта: разработка под руководством ИИ, надзор за агентами и распределение выч...","url":"https://www.aioga.com/ru/news/cmu605ozy000arok0lxr4x1y8/","contentTranslated":true,"sourceHash":"9d84af2dead41072","translatedAt":"2026-09-17T23:07:01.861Z"},"ar":{"title":"أنتروبك تصدر أداة قياس وتيرة تطوير الذكاء الاصطناعي المتقدمة ولقطة لمؤشرات داخلية","summary":"أصدرت شركة Anthropic مجموعة من المؤشرات لقياس وتيرة تطوير الذكاء الاصطناعي المتقدمة، تغطي ثلاثة جوانب: البحث والتطوير بقيادة الذكاء الاصطناعي، إشراف الوكلاء، وتوزيع القدرة الحاسوبية.","category":"行业动态","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"أنتروبك تصدر أداة قياس وتيرة تطوير الذكاء الاصطناعي المتقدمة ولقطة لمؤشرات داخلية - Aioga أخبار الذكاء الاصطناعي","description":"أصدرت شركة Anthropic مجموعة من المؤشرات لقياس وتيرة تطوير الذكاء الاصطناعي المتقدمة، تغطي ثلاثة جوانب: البحث والتطوير بقيادة الذكاء الاصطناعي، إشراف الوكلاء، وتوزيع القدرة الحاسوبي...","url":"https://www.aioga.com/ar/news/cmu605ozy000arok0lxr4x1y8/","contentTranslated":true,"sourceHash":"9d84af2dead41072","translatedAt":"2026-09-17T23:08:04.210Z"},"hi":{"title":"Anthropic ने अग्रिम AI विकास गति माप उपकरण और आंतरिक मेट्रिक्स स्नैपशॉट जारी किया","summary":"Anthropic ने अग्रणी AI विकास की गति को मापने के लिए सूचकांकों का एक सेट जारी किया, जो AI आधारित अनुसंधान, एजेंट निगरानी और कंप्यूटिंग पावर वितरण के तीन पहलुओं को कवर करता है।","category":"行业动态","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"Anthropic ने अग्रिम AI विकास गति माप उपकरण और आंतरिक मेट्रिक्स स्नैपशॉट जारी किया - Aioga AI समाचार","description":"Anthropic ने अग्रणी AI विकास की गति को मापने के लिए सूचकांकों का एक सेट जारी किया, जो AI आधारित अनुसंधान, एजेंट निगरानी और कंप्यूटिंग पावर वितरण के तीन पहलुओं को कवर करता है।","url":"https://www.aioga.com/hi/news/cmu605ozy000arok0lxr4x1y8/","contentTranslated":true,"sourceHash":"9d84af2dead41072","translatedAt":"2026-09-17T23:08:05.473Z"},"it":{"title":"Anthropic ha rilasciato uno strumento di misurazione del ritmo di sviluppo dell'IA all'avanguardia e un'istantanea dei propri indicatori interni","summary":"Anthropic ha rilasciato un insieme di indicatori per misurare il ritmo di sviluppo dell'AI all'avanguardia, coprendo tre aspetti: ricerca guidata dall'AI, supervisione degli agenti intelligenti e allocazione della capacità di calcolo.","category":"行业动态","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"Anthropic ha rilasciato uno strumento di misurazione del ritmo di sviluppo dell'IA all'avanguardia e un'istantanea dei propri indicatori interni - Aioga Notizie IA","description":"Anthropic ha rilasciato un insieme di indicatori per misurare il ritmo di sviluppo dell'AI all'avanguardia, coprendo tre aspetti: ricerca guidata dall'AI, supervisione degli agenti...","url":"https://www.aioga.com/it/news/cmu605ozy000arok0lxr4x1y8/","contentTranslated":true,"sourceHash":"9d84af2dead41072","translatedAt":"2026-09-17T23:09:01.331Z"},"nl":{"title":"Anthropic lanceert tool voor het meten van de voortgang van AI-ontwikkeling en intern metriekoverzicht","summary":"Anthropic heeft een set indicatoren uitgebracht om het tempo van geavanceerde AI-ontwikkeling te meten, die drie aspecten bestrijken: door AI gedomineerd onderzoek en ontwikkeling, toezicht door agenten en toewijzing van rekenkracht.","category":"行业动态","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"Anthropic lanceert tool voor het meten van de voortgang van AI-ontwikkeling en intern metriekoverzicht - Aioga AI-nieuws","description":"Anthropic heeft een set indicatoren uitgebracht om het tempo van geavanceerde AI-ontwikkeling te meten, die drie aspecten bestrijken: door AI gedomineerd onderzoek en ontwikkeling,...","url":"https://www.aioga.com/nl/news/cmu605ozy000arok0lxr4x1y8/","contentTranslated":true,"sourceHash":"9d84af2dead41072","translatedAt":"2026-09-17T23:09:07.834Z"},"tr":{"title":"Anthropic, öncü yapay zeka geliştirme hızını ölçen araç ve iç metrik anlık görüntüsünü yayınladı","summary":"Anthropic, AI ağırlıklı araştırma ve geliştirme, ajan denetimi ve hesaplama gücü dağılımı olmak üzere üç alanı kapsayan öncü AI geliştirme hızını ölçen bir dizi gösterge yayınladı.","category":"行业动态","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"Anthropic, öncü yapay zeka geliştirme hızını ölçen araç ve iç metrik anlık görüntüsünü yayınladı - Aioga AI Haberleri","description":"Anthropic, AI ağırlıklı araştırma ve geliştirme, ajan denetimi ve hesaplama gücü dağılımı olmak üzere üç alanı kapsayan öncü AI geliştirme hızını ölçen bir dizi gösterge yayınladı.","url":"https://www.aioga.com/tr/news/cmu605ozy000arok0lxr4x1y8/","contentTranslated":true,"sourceHash":"9d84af2dead41072","translatedAt":"2026-09-17T23:10:15.480Z"},"vi":{"title":"Anthropic phát hành công cụ đo nhịp độ phát triển AI tiên tiến và ảnh chụp nhanh các chỉ số nội bộ","summary":"Anthropic phát hành một bộ chỉ số đo lường nhịp độ phát triển AI tiên tiến, bao quát ba khía cạnh: nghiên cứu do AI dẫn dắt, giám sát thực thể thông minh và phân bổ sức mạnh tính toán.","category":"行业动态","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"Anthropic phát hành công cụ đo nhịp độ phát triển AI tiên tiến và ảnh chụp nhanh các chỉ số nội bộ - Tin tức AI Aioga","description":"Anthropic phát hành một bộ chỉ số đo lường nhịp độ phát triển AI tiên tiến, bao quát ba khía cạnh: nghiên cứu do AI dẫn dắt, giám sát thực thể thông minh và phân bổ sức mạnh tính t...","url":"https://www.aioga.com/vi/news/cmu605ozy000arok0lxr4x1y8/","contentTranslated":true,"sourceHash":"9d84af2dead41072","translatedAt":"2026-09-17T23:10:05.586Z"},"id":{"title":"Anthropic merilis alat pengukuran ritme pengembangan AI terdepan dan snapshot indikator internal","summary":"Anthropic merilis seperangkat indikator untuk mengukur ritme pengembangan AI terkini, yang mencakup penelitian yang dipimpin AI, pengawasan agen cerdas, dan alokasi daya komputasi.","category":"行业动态","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"Anthropic merilis alat pengukuran ritme pengembangan AI terdepan dan snapshot indikator internal - Berita AI Aioga","description":"Anthropic merilis seperangkat indikator untuk mengukur ritme pengembangan AI terkini, yang mencakup penelitian yang dipimpin AI, pengawasan agen cerdas, dan alokasi daya komputasi.","url":"https://www.aioga.com/id/news/cmu605ozy000arok0lxr4x1y8/","contentTranslated":true,"sourceHash":"9d84af2dead41072","translatedAt":"2026-09-17T23:11:06.428Z"},"th":{"title":"Anthropic เปิดตัวเครื่องมือวัดจังหวะการพัฒนา AI ขั้นสูงและภาพรวมตัวชี้วัดภายใน","summary":"Anthropic เปิดตัวชุดตัวชี้วัดสำหรับวัดจังหวะการพัฒนาปัญญาประดิษฐ์ขั้นสูง ครอบคลุมสามด้าน ได้แก่ การวิจัยและพัฒนาที่ขับเคลื่อนด้วย AI การกำกับดูแลโดยเอเยนต์ และการจัดสรรกำลังการประมวลผล","category":"行业动态","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"Anthropic เปิดตัวเครื่องมือวัดจังหวะการพัฒนา AI ขั้นสูงและภาพรวมตัวชี้วัดภายใน - ข่าว AI Aioga","description":"Anthropic เปิดตัวชุดตัวชี้วัดสำหรับวัดจังหวะการพัฒนาปัญญาประดิษฐ์ขั้นสูง ครอบคลุมสามด้าน ได้แก่ การวิจัยและพัฒนาที่ขับเคลื่อนด้วย AI การกำกับดูแลโดยเอเยนต์ และการจัดสรรกำลังการประม...","url":"https://www.aioga.com/th/news/cmu605ozy000arok0lxr4x1y8/","contentTranslated":true,"sourceHash":"9d84af2dead41072","translatedAt":"2026-09-17T23:11:11.664Z"},"pl":{"title":"Anthropic opublikowało narzędzie do pomiaru tempa rozwoju AI oraz wewnętrzny zrzut wskaźników","summary":"Anthropic opublikowało zestaw wskaźników mierzących tempo rozwoju nowoczesnej sztucznej inteligencji, obejmujący trzy aspekty: badania kierowane przez AI, nadzór agentów oraz alokację mocy obliczeniowej.","category":"行业动态","source":"Anthropic：The Institute（旗舰研究长文 · 网页）","aggregationSource":"Anthropic：Newsroom（网页）","pageTitle":"Anthropic opublikowało narzędzie do pomiaru tempa rozwoju AI oraz wewnętrzny zrzut wskaźników - Aioga Wiadomości AI","description":"Anthropic opublikowało zestaw wskaźników mierzących tempo rozwoju nowoczesnej sztucznej inteligencji, obejmujący trzy aspekty: badania kierowane przez AI, nadzór agentów oraz aloka...","url":"https://www.aioga.com/pl/news/cmu605ozy000arok0lxr4x1y8/","contentTranslated":true,"sourceHash":"9d84af2dead41072","translatedAt":"2026-09-17T23:12:10.617Z"}},"evidenceTier":"verified-news","reviewStatus":"editorial-selected","indexable":true,"editorialCover":""}}