{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-08-10T19:40:29.648Z","headline":"智能体真的会用电脑吗？a16z 用数据给出答案","description":"a16z 数据显示，计算机操作智能体在 OSWorld-Verified 基准上的最佳成绩已从一年前的 42% 升至 85%，超过人类测试者约 72% 的水平，Claude Fable 5 以 85% 领先。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","url":"https://www.aioga.com/news/cmsnbzkyr01qjrohfmjqlmjk5/","mainEntityOfPage":"https://www.aioga.com/news/cmsnbzkyr01qjrohfmjqlmjk5/","datePublished":"2026-08-10T14:00:46.000Z","dateModified":"2026-08-10T14:00:46.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.a16z.news/p/can-agents-use-a-computer-yet-weve","https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5"],"canonicalUrl":"https://www.aioga.com/news/cmsnbzkyr01qjrohfmjqlmjk5/","directAnswer":{"@type":"Answer","text":"a16z称，计算机操作智能体在OSWorld-Verified基准上的最佳成绩已从一年前的42%升至85%，高于人类测试者约72%的成绩。文章同时强调，真实业务部署仍受失败率、成本和流程偏离等问题限制。","url":"https://www.aioga.com/news/cmsnbzkyr01qjrohfmjqlmjk5/","dateCreated":"2026-08-10T14:00:46.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"a16z.news source article","url":"https://www.a16z.news/p/can-agents-use-a-computer-yet-weve","datePublished":"2026-08-10T14:00:46.000Z","provider":{"@type":"Organization","name":"a16z.news","url":"https://www.a16z.news/p/can-agents-use-a-computer-yet-weve"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","datePublished":"2026-08-10T14:00:46.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5"}}],"aggregationSource":"a16z：News（RSS）","originalPublisher":{"name":"a16z.news","url":"https://www.a16z.news/p/can-agents-use-a-computer-yet-weve"},"geoDeepAnswer":null,"article":{"id":"cmsnbzkyr01qjrohfmjqlmjk5","slug":"cmsnbzkyr01qjrohfmjqlmjk5","url":"https://www.aioga.com/news/cmsnbzkyr01qjrohfmjqlmjk5/","title":"智能体真的会用电脑吗？a16z 用数据给出答案","title_en":"","summary":"a16z 数据显示，计算机操作智能体在 OSWorld-Verified 基准上的最佳成绩已从一年前的 42% 升至 85%，超过人类测试者约 72% 的水平，Claude Fable 5 以 85% 领先。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","source":"a16z：News（RSS）","sourceUrl":"https://www.a16z.news/p/can-agents-use-a-computer-yet-weve","aiHotUrl":"https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","publishedAt":"2026-08-10T14:00:46.000Z","category":"行业动态","score":72,"selected":true,"articleBody":["America：https://www.a16z.news/t/america | Tech：https://www.a16z.news/t/technology | Opinion：https://www.a16z.news/t/opinion | Culture：https://www.a16z.news/t/culture | Charts：https://www.a16z.news/t/charts","It seems obvious to say, but if you leave Silicon Valley and go out into the rest of the world and tell them, “there are these things called agents, which are pretty smart, and can do tasks with you, and automate some of the repetitive parts of your work”, chances are the first question you’ll get back is, “Can they use a computer?”","This is a good question! Can they, really? The long horizon of productivity potential, out in the real economy we’re going to go unlock over decades, runs through pretty everyday work: can an agent sit (metaphorically) at a desk 24/7, and be trusted to use a web browser, fill out forms, click the right buttons, and not make mistakes? This is the domain of Business Process Outsourcing (BPO), which historically meant, “can this work be outsourced?” but now has a new agentic frontier. We wrote about this last year ：https://a16z.com/the-rise-of-computer-use-and-agentic-coworkers/ , when the computer-use landscape was still mostly a bunch of demos. A lot has happened since then.","The models have improved faster than almost anyone expected. Computer-using agents are beginning to hold up in production at scale and on narrow, repeatable workflows: updating systems of record, moving data through portals, processing tickets, checking records, and handling the long tail of software where no clean API exists. With the right infrastructure, computer-use capabilities can now be deployed to tackle end-to-end tasks at scale, which before required either human supervision or direct human completion.","Today, workflows leveraging computer-use are far from perfect: agents are brittle when work drifts off the runbook, and for certain use-cases where caching is intractable (more below) they are expensive enough that the math does not work everywhere. But we’re seeing production deployments for standardized back-office work, especially where labor would otherwise be clicking through legacy systems by hand; the cost curve is starting to look compelling, considering that workflows leveraging computer-use offer structural advantages such as 24/7 availability and - most importantly - scalability to meet demand.","The first wave of computer-use infrastructure was about making agents capable: seeing, clicking, typing, recovering from mistakes. The next wave is about making them useful inside actual companies. As raw UI navigation becomes a model-layer commodity, the model is no longer the main bottleneck and the durable advantage moves up the stack: context, permissions, process knowledge, validation, escalation, error handling, caching, and the hard-earned understanding of how work actually gets done inside one specific customer’s organization to map a workflow end-to-end. In other words, the frontier is shifting from “can the agent use a computer?” to “can it reliably do this job?”","A year ago the best computer-using model scored 42% on OSWorld-Verified; today’s best scores 85%, above the ~72% humans manage on the same tasks (this means they successfully completed 85 of 100 tasks). In production these general frontier models run much like they do in the benchmark: the labs expose computer use as an API - the model gets a screenshot, returns clicks and keystrokes, with OpenAI’s CUA also layering in accessibility-tree or DOM data where available - and builders wrap that loop in their own harness: a sandboxed VM or browser, plus the orchestration, verification, and retry logic around it. Notably, almost nobody deploys consumer products (Claude, ChatGPT agent mode) for this - founders and enterprises build on the raw APIs, or buy from vendors who package them. And the capability jump is what made those setups viable - “the models weren’t good enough to use in production on their own until Opus 4.6 in February 2026,” as one founder building in the space put it. Somewhere in the last eighteen months, computer use capabilities crossed from demo to being deployable in the field.","Of course, benchmarks aren’t always the best proxy for the viability of a real-world deployment. OSWorld counts completed tasks, so 85% still means 15 of 100 failed, and a business process only finishes if every step does. Back-office work doesn’t grade on a curve: if a person reviews every output, no labor was saved. (It’s analogous to what’s happening in coding right now: the scarce resource is no longer writing the code, it’s vouching for it.)","We found that the best way to think about what matters is to go beyond the benchmark and focus on the core question: can a business process be reliably automated with computer-use capabilities? Under this lens, what makes the biggest difference is everything around the model - that is: verification, escalation, error handling when a retailer portal changes its layout overnight.","Perhaps the clearest tell is that one operator we spoke with, who runs millions of automated tasks a month, couldn’t tell us which model executes them; he hadn’t needed to find out. His vendor swaps models underneath him the way a cloud provider swaps hardware. But he did trust the computer-using agent to run these tasks. Bottom line: when your heaviest users stop checking the leaderboard, the leaderboard has stopped being the story.","Thus, the chart above explains why production deployments exist in 2026 and didn’t in 2024. From there on, what determines whether they work is everything else - and that’s the rest of this piece.","We had various conversations with teams running workflows leveraging computer-use capabilities in production, and we learned from their experiences that protocol following tasks work best. Unsurprisingly, computer-using agents break on more complex workflows where accuracy is harder to verify. The overall takeaway is that computer-using agents are strongest on standardized, repeatable tasks with a clear, well-defined path. The real unlock is the long tail of software where no clean API exists and a person would otherwise be clicking through a UI by hand. In practice, the work looks like updating records in a CRM, QA, logging into government and insurance portals, pulling data off databases and regulatory pages, retail order processing, contract processing, or IT tickets in ServiceNow.","We believe the voice of the user here tells the story much better than any theory. Some examples: a CPG data platform walked us through how they run ~15-20M automated portal interactions a month, using agents as a self-healing fallback for hand-coded scrapers - when a retailer portal changes its UI, the agent diagnoses the break, fixes the automation, and keeps data flowing before an engineer ever sees the error. Once implemented, they told us, they cut the engineering team dedicated to scraper maintenance in half and re-allocated staff capacity to other workflows. In another case, from a global systems integrator, we learned they have 27 live workflows leveraging computer use agents that process ~1,500-2,100 IT tickets a day, with the ultimate goal of redeploying 20-25% of headcount on low-margin managed-services contracts. And finally an agency walked us through how they automated a recruiting workflow end-to-end to populate data in an applicant tracking platform as soon as a candidate interview was over. To do so they run a cheap non-frontier model because it “does everything we need and does it well.”","For the buyers we spoke with, the model itself is rarely the deciding factor, as “the models today are already good enough.” In practice, they evaluate and pay for everything around the model: the infrastructure to run reliably at scale, pass security review, and prove ROI. Users don’t care whether the solution uses a given frontier model; rather they focus on whether it can actually get the task done at scale and reliably. Period.","As a consequence, failure modes matter more than any benchmark, and design for failure needs to be a first-class concern from the start because a solution that does not handle failures well will never be adopted in production. One example of what this looks like in practice and a pattern we ran into more than once: the agent runs the workflow once, the system caches it as deterministic code, runs execute as cheap repeatable code from then on, and the model comes back only when something breaks - to diagnose, fix, and re-cache. With this approach, cost per run falls over a workflow’s lifetime, and cheaper models just lower the bill. What’s interesting about it is how it handles uncertainty. Where before deterministic code simply failed, or a human had to review and fix every break, here the agent absorbs that uncertainty on its own. It’s one way of designing for failure, and it shows what buyers are actually rewarding and using at scale.","We did not encounter more sophisticated use cases among the users we spoke with, which tells us the market is still chipping away at the low-hanging fruit. That said, there is a long list of workflows that can be automated this way before anyone needs the harder tasks.","For founders, the more important shift is what is becoming commoditized. Building a computer-using agent used to mean wrestling with Selenium or Playwright, or more recently Stagehand, and stitching together DOM or video recordings to capture a workflow. That whole execution layer is getting abstracted away, the same way Claude Code abstracted the scaffolding around coding agents. If clicking the right button is no longer the hard part, it is no longer the moat.","Unsurprisingly, the context and knowledge of the workflow are durable. The hard part is not whether an agent can navigate an SAP screen, it is whether it understands how a particular company actually gets work done: the tribal knowledge, the internal terminology, the preferred formats, who to escalate to and when, how to handle failures, how to verify output reliably. In practice that context lives in runbooks, in access and credentials, in test cases and guardrails for when a workflow goes off-script, and increasingly in a single recorded video of someone doing the job once. None of it is general. All of it is specific to a company, and often to one team. That said, it is exactly the kind of specific, unglamorous problem that focused startups tend to solve better than model providers, which is why we think the next generation of agentic coworkers gets built at the application and context layer, not the model layer.","The buyers we spoke with squarely confirmed this. They picked vendors on whether the product reported hours saved without extra work, and whether a junior engineer could run it. For now, the moat is not the frontier capability—rather, it’s being the vendor an enterprise is allowed, and able, to use at scale in production.","The cost data is also encouraging. Take the numbers above as an order of magnitude, not a precise quote. Running an agent costs roughly $6-8 per hour of inference today, but in practice anywhere between $3 and $15 depending on how the harness is built - how often it screenshots, how much context it carries, how much of the work it can hand off to deterministic code. These figures describe the agent operating the UI screenshot by screenshot with a frontier model - the most expensive mode there is. Well-built harnesses reserve that mode for what actually needs it, and let cheap deterministic code handle the repeatable parts - not every workflow can be optimized this way, but where it can, blended cost drops fast. So read the comparison as the worst case, and even then an agent is roughly break-even against offshore BPO at ~$10/hour fully loaded, and pencils out to a 70-80% gross margin against US back-office labor at ~$30-45/hour. In production, the harness drives real-world cost as much as the model does.","Same caveat on speed. In agentic mode, agents are still slower than people, and it is not close - a task someone finishes in two to three minutes can take an agent eight to ten, and academic benchmarks put the gap even wider. Deterministic runs flip this: code executes faster than any human - but for the agentic work the argument is not speed. It is that an agent runs around the clock, costs a fraction of US labor, and scales without hiring.","This comparison works for BPO buyers and ops teams, but the unit economics look different if you are the one selling agent-hours, because costs are less predictable in the wild. COGS are inference plus retries (i.e. failed runs still burn tokens) and margins compress when context grows or screenshot frequency goes up. Vendors manage this by pricing per task, per hour, or per outcome, each with a different risk profile depending on workflow variance. There are also the practical realities of things like monitoring, maintenance and human escalation, which are being priced in, just as they would be with human workforce. There is no one-size-fits-all business model here yet, and the answer varies by vertical.","And the math only gets better - inference keeps getting cheaper, and open-source models are getting good enough for a growing share of these workflows. For any task an agent can reliably solve, embedding computer-use will likely be way more convenient than human labor. So the real question is no longer whether the economics work - it’s how far the set of tasks that can be solved reliably extends, which is where things are heading next.","Over the past year, labs and a wave of startups have poured hundreds of millions into computer-use RL environments - the sandboxes where a model practices real tasks and gets rewarded for finishing them - with companies like Mechanize, Habitat, Fleet, Chakra, Deeptune, Matrices, and Originator building the training and eval substrate underneath the frontier models. That spend is what shows up as better reasoning, better state tracking, and more tolerance for apps that misbehave. The models still need careful harnessing to hold up in production - run-caching being the clearest example - but the raw capability was bought, deliberately, through this training infrastructure and will only get better over time.","Architecture, though, is a different story. Most computer-use systems deployed in production today are single-agent: one model, one task, one session. As workflows get more complex and latency becomes a constraint, multi-agent architectures start to matter. For example, a planner decomposes the workflow, executor agents handle subtasks in parallel and long-running agents bring their own problems: memory, trust, and failure rates that compound over time. The teams doing interesting work here are all building bespoke orchestration, because no standard framework exists yet. The Claude Code analogy is instructive: when coding agents matured, a scaffolding layer emerged to abstract the orchestration away. The same is likely to happen for workflows leveraging computer-use capabilities, and that abstraction layer is one of the more interesting unsolved infrastructure problems in the space.","From here, future developments run along three lines: accuracy, latency, and cost. Accuracy is most important, and, as we explained above, represents the difference between a cool demo and actually solving the problem - catching anomalies, checking its own work, escalating only when it actually needs to. Latency is the one most likely to surprise people: some teams already cut it today by grounding on the accessibility tree instead of screenshots. Standard Intelligence’s general computer action model is trained on an 11-million-hour video dataset, running at 30 FPS, and is an early signal that the step-by-step screenshot loop slowing today’s agents is a solvable problem, not a permanent tax. Cost keeps falling as inference gets cheaper, and smaller non-frontier models take over the routine clicks. These three vectors in addition to security and governance (e.g. credentials, audit logs, data retention, prompt injection, and accountability, permissioning).","Enterprises can and are benefiting from computer-using agents for narrow workflows that have high volume, repetitive steps, stable business rules with legacy interfaces or missing APIs. For now, they are best suited for tasks with immediate, machine-observable evidence of success, tolerable failure consequences, and clear escalation routes. But with the above developments, improvements are real and rapid, making computer-using agents more viable for more types of work.","Let’s just say - the future for computer-use capabilities is bright!"],"articleImages":[{"sourceUrl":"https://substackcdn.com/image/fetch/$s_!yJJV!,e_trim:10:white/e_trim:10:transparent/h_72,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb31fe13-e316-436e-b482-c771eb463be9_1344x257.png","alt":"a16z","afterParagraph":0,"url":"/media/articles/cmsnbzkyr01qjrohfmjqlmjk5/357fd666d1b85d94.png"}],"mediaStatus":"ok","articleBodyZh":["美国：https://www.a16z.news/t/america | 技术：https://www.a16z.news/t/technology | 观点：https://www.a16z.news/t/opinion | 文化：https://www.a16z.news/t/culture | 图表：https://www.a16z.news/t/charts","说出来似乎很显而易见，但如果你离开硅谷，走到世界其他地方，告诉他们：“有一些叫做代理的东西，它们相当聪明，可以与你一起完成任务，并自动化你工作中的一些重复部分”，那么你很可能首先得到的问题是：“它们会用电脑吗？”","这是一个好问题！它们真的会吗？我们在未来几十年中将在真实经济中解锁的生产力潜力，很大程度上涉及非常日常的工作：一个代理能否（比喻地）24/7坐在办公桌前，并被信任使用网页浏览器、填写表格、点击正确的按钮，并且不出错？这是业务流程外包（BPO）的领域，历史上意味着“这项工作能外包吗？”，但现在有了新的代理前沿。我们在去年写过相关内容：https://a16z.com/the-rise-of-computer-use-and-agentic-coworkers/，当时计算机使用的场景仍主要是一些演示。从那时起已经发生了很多变化。","模型的改进速度几乎超过了所有人的预期。会用电脑的代理开始在大规模、狭窄且可重复的工作流程中表现稳定：更新记录系统、通过门户传输数据、处理工单、检查记录以及处理没有干净API的软件长尾部分。有了合适的基础设施，计算机使用能力现在可以部署来应对大规模的端到端任务，而以前这些任务需要人工监督或直接人工完成。","今天，利用计算机的工作流程还远未完美：当工作偏离操作手册时，代理容易出错，而且对于某些缓存无法处理的用例（稍后详细说明），它们的成本高到无法在任何地方都可行。但是我们已经看到标准化后台工作的生产部署，特别是那些原本需要人工点击遗留系统的工作；考虑到利用计算机的工作流程提供了结构性优势，比如全天候可用性，并且最重要的是，可以按需扩展，成本曲线开始变得有吸引力。","第一波计算机使用基础设施是为了让代理具备能力：观察、点击、输入、从错误中恢复。下一波则是让它们在实际公司中变得有用。随着原始界面导航成为模型层的商品，模型不再是主要瓶颈，持久的优势上移到堆栈之上：上下文、权限、流程知识、验证、升级、错误处理、缓存，以及如何在特定客户组织内实际完成工作的宝贵经验，以实现工作流程的端到端映射。换句话说，前沿正在从“代理能否使用计算机？”转向“它能否可靠地完成这项工作？”","一年前，最擅长使用计算机的模型在 OSWorld-Verified 上的得分为 42%；而今天的最佳模型得分为 85%，高于人类在相同任务中大约达到的 72%（这意味着它们在 100 个任务中成功完成了 85 个）。在实际应用中，这些先进的通用模型的运行方式与基准测试中类似：实验室将计算机使用作为 API 暴露——模型接收屏幕截图，返回点击和按键操作，OpenAI 的 CUA 还会在可用时叠加可访问性树或 DOM 数据——而构建者则将这一循环封装在自己的环境中：一个沙盒的虚拟机或浏览器，以及围绕它的编排、验证和重试逻辑。值得注意的是，几乎没有人将消费级产品（如 Claude、ChatGPT 代理模式）用于此用途——创始人和企业直接使用原始 API 构建，或者从将其打包的供应商处购买。而这种能力的跃升正是使这些设置可行的原因——正如一位在该领域构建产品的创始人所说：“在 2026 年 2 月 Opus 4.6 之前，这些模型的能力还不足以单独在生产中使用。”在过去的十八个月里，计算机使用能力从演示阶段跨越到了可在现场部署的阶段。","当然，基准测试并不总是衡量实际部署可行性的最佳指标。OSWorld 统计完成的任务数量，因此 85% 仍意味着 100 个任务中有 15 个失败，而业务流程只有在每个步骤都完成时才算完成。后台工作不会按照曲线评分：如果每个输出都需要人工审核，那么并没有节省任何人力。（这与当前代码编写的情况类似：稀缺资源不再是写代码，而是为代码的正确性担保。）","我们发现，思考什么最重要的最佳方式是超越基准测试，关注核心问题：业务流程能否可靠地通过计算机使用能力实现自动化？在这个视角下，最大区别在于模型之外的所有环节——也就是说：验证、升级，以及当零售商门户网站一夜之间改变布局时的错误处理。","也许最清楚的迹象是，我们采访的一位运营者，他每月运行数百万个自动化任务，却无法告诉我们是哪种模型在执行这些任务；他并不需要去知道。供应商会像云服务提供商更换硬件那样在他不知情的情况下更换模型。但他确实信任使用计算机的代理来运行这些任务。总结来说：当你的最活跃用户不再查看排行榜时，排行榜已经不再是故事的核心了。","因此，上图解释了为什么2026年会有生产部署，而2024年没有。从那时起，决定它们是否有效的是其他一切——这也是本文余下部分讨论的内容。","我们与在生产中运行利用计算机使用能力的工作流的团队进行了各种对话，并从他们的经验中了解到，遵循协议的任务效果最佳。不出意外，计算机使用代理在更复杂、难以验证准确性的工作流中容易失败。总体结论是，计算机使用代理在标准化、可重复且路径清晰明确定义的任务上最为强大。真正的突破在于长尾的软件场景，在这些场景中没有干净的 API，否则人们需要手动点击 UI 执行操作。实际上，工作内容包括在 CRM 中更新记录、质量检查、登录政府和保险门户、从数据库及监管页面提取数据、零售订单处理、合同处理，或 ServiceNow 中的 IT 工单处理。","我们相信，用户的声音讲述这个故事要比任何理论都更有力。一些例子：一家消费品数据平台向我们展示了他们如何每月运行约1500-2000万次自动化门户交互，使用代理作为手动编码爬虫的自我修复备选方案——当零售商门户更改其用户界面时，代理会诊断故障，修复自动化，并在工程师看到错误之前保持数据流动。他们告诉我们，一旦实施，他们将专门用于爬虫维护的工程团队人数减半，并将员工容量重新分配到其他工作流程。另一个例子，来自一家全球系统集成商，我们了解到他们有27个实时工作流程利用计算机使用代理每天处理约1500-2100个IT工单，最终目标是在低利润的托管服务合同上重新部署20%-25%的员工数量。最后，一家机构向我们展示了他们如何从头到尾自动化招聘工作流程，以便在候选人面试结束后立即在申请者跟踪平台中填充数据。为此，他们运行一个便宜的非前沿模型，因为它“能完成我们所需的一切，并且做得很好”。","对于我们交谈过的买家来说，模型本身很少是决定因素，因为“今天的模型已经足够好了”。在实际操作中，他们评估并为模型周围的一切付费：可靠大规模运行的基础设施、安全审查的通过以及投资回报的证明。用户不关心解决方案是否使用某个前沿模型；他们关注的是解决方案是否真的能够大规模可靠地完成任务。仅此而已。","因此，失败模式比任何基准测试都更重要，从一开始就需要将防故障设计作为首要关注点，因为一个无法良好处理失败的解决方案永远不会被生产环境采用。实际中的一个例子，以及我们多次遇到的模式是：代理运行工作流程一次，系统将其缓存为确定性代码，然后从此以廉价可重复代码执行，并且模型只有在出现问题时才返回——用于诊断、修复和重新缓存。通过这种方法，工作流程的每次运行成本会随着时间下降，而更便宜的模型只会进一步降低费用。有趣的是它如何处理不确定性。过去，确定性代码会直接失败，或者需要人工检查和修复每一次中断，而在这里，代理可以自主吸收这种不确定性。这是一种防故障设计方式，并且展示了买家实际上在奖励和大规模使用的内容。","在我们交谈过的用户中，并未遇到更复杂的用例，这告诉我们市场仍在逐步解决低悬果实的问题。也就是说，在任何人需求更困难任务之前，还有很长的工作流程清单可以通过这种方式实现自动化。","对于创始人而言，更重要的变化是哪些正在被商品化。构建使用计算机的代理过去意味着需要处理 Selenium 或 Playwright，或者最近的 Stagehand，并拼接 DOM 或视频记录以捕获工作流程。整个执行层正在被抽象化，就像 Claude Code 将编码代理周围的脚手架抽象化一样。如果点击正确按钮不再是难点，那么它也不再是护城河。","不足为奇的是，工作流的上下文和知识是持久的。困难的部分不在于一个代理能否在 SAP 界面中导航，而在于它是否理解某个公司实际如何完成工作：部落知识、内部术语、偏好的格式、向谁以及何时升级、如何处理故障、如何可靠地验证输出。实际上，这些上下文存在于操作手册、访问权限和凭证、测试用例以及当工作流偏离脚本时的防护措施中，而且越来越多地存在于一段记录了某人完成该工作的单次视频里。没有任何东西是通用的。所有内容都是针对特定公司的，有时甚至针对某个团队。话虽如此，这正是那种专注型创业公司往往比模型提供者解决得更好的具体、平凡的问题，这也是我们认为下一代智能同事将在应用层和上下文层构建，而不是在模型层构建的原因。","我们交谈过的买家明确证实了这一点。他们选择供应商的标准在于产品是否报告了节省的工时而无需额外工作，以及初级工程师是否能操作它。就目前而言，护城河不在于前沿能力——而在于成为企业可以、能够在生产环境中大规模使用的供应商。","成本数据也令人鼓舞。上面的数字应作为数量级参考，而非精确报价。运行一个代理目前大约花费每小时 $6-8 的推理成本，但实际上根据执行框架的构建方式可能在 $3 到 $15 之间——包括截图频率、所携带的上下文量以及可以交给确定性代码处理的工作量。这些数据描述的是代理用前沿模型逐屏操作 UI 的情况——这是最昂贵的模式。构建良好的执行框架会将这种模式保留给真正需要它的部分，让廉价的确定性代码处理可重复的部分——并不是每个工作流都能这样优化，但在可行的情况下，综合成本会迅速下降。因此，把比较看作最坏情况，即便如此，以完全加载约 $10/小时的离岸 BPO 比较，代理大致能够达到收支平衡，相对于美国后台劳动力（约 $30-45/小时），毛利率可达到 70-80%。在生产中，执行框架对实际成本的影响与模型本身一样重要。","速度方面的警告同样适用。在自主模式下，智能体仍然比人类慢，而且差距很大——一个人能在两到三分钟完成的任务，智能体可能需要八到十分钟，而学术基准测试显示差距甚至更大。确定性运行则相反：代码的执行速度比任何人类都快——但对于自主工作的讨论焦点不是速度。关键在于，智能体全天候运行，成本只是美国人工的一小部分，并且无需招聘即可扩展。","这种比较适用于BPO的买家和运营团队，但如果你是在销售智能体小时数，单位经济性看起来就不同，因为现实中的成本不可预测。销售成本是推理成本加上重试（即失败的运行仍然消耗令牌），随着上下文增加或截图频率提高，利润会被压缩。供应商通过按任务、按小时或按结果定价来管理这些风险，每种方式对应的风险取决于工作流的差异。此外，还有监控、维护和人工升级等实际问题的成本需要计入，这和人工劳动力的定价方式类似。目前还没有通用的商业模式，答案因行业而异。","而且数学上情况只会越来越好——推理成本不断降低，开源模型已经足够好，可以处理越来越多的工作流。对于任何智能体能够可靠解决的任务，嵌入计算机使用很可能要比人力更方便。因此，真正的问题不再是经济性是否可行，而是能可靠解决的任务范围有多广，这正是未来的发展方向。","在过去一年里，各实验室和一波初创公司在计算机使用的强化学习环境上投入了数亿资金——这些沙箱让模型练习实际任务，并在完成任务后获得奖励——像Mechanize、Habitat、Fleet、Chakra、Deeptune、Matrices和Originator这样的公司正在前沿模型下搭建训练和评估的基础设施。这些投入带来了更好的推理能力、更好的状态跟踪，以及更高的容忍度用于那些表现异常的应用。这些模型在生产中仍需小心把控——运行缓存是最明显的例子——但其原始能力是通过这一训练基础设施刻意获得的，并且随着时间推移只会越来越强。","然而，架构是另一回事。今天在生产中部署的大多数计算机使用系统都是单代理的：一个模型、一个任务、一个会话。随着工作流程变得更加复杂且延迟成为限制条件，多代理架构开始显得重要。例如，规划器分解工作流程，执行代理并行处理子任务，而长期运行的代理带来自己的问题：内存、信任和随时间累积的失败率。在这一领域做有趣工作的团队都在构建定制的编排，因为目前还不存在标准框架。Claude Code 的类比很有启发性：当编码代理成熟时，出现了一个支架层来抽象编排。同样的情况很可能会出现在利用计算机使用能力的工作流程中，而这个抽象层是该领域尚未解决的更有趣的基础设施问题之一。","从这里看，未来的发展沿着三条路线进行：准确性、延迟和成本。准确性最为重要，正如我们上面解释的，它代表了炫酷演示与实际解决问题之间的区别——捕捉异常、检查自身工作、仅在真正需要时升级。延迟是最可能让人惊讶的：一些团队今天已经通过基于可访问性树而不是截图来减少延迟。Standard Intelligence 的通用计算机操作模型在一个1100万小时的视频数据集上训练，运行速度为30帧每秒，这是一个早期信号，表明缓慢的逐步截图循环是可以解决的问题，而不是永久的负担。随着推理成本下降，成本持续降低，而较小的非前沿模型接管常规点击任务。这三个向量除了安全性和治理（例如凭据、审计日志、数据保留、提示注入以及问责、权限管理）之外。","企业可以并且确实正在从计算机使用代理中受益，特别是对于高量、重复步骤、业务规则稳定、使用遗留接口或缺少 API 的狭窄工作流程。目前，它们最适合用于有立即可机器观测的成功证据、可容忍的失败后果以及清晰升级路径的任务。但随着上述发展，改进真实且快速，使计算机使用代理对更多类型的工作更为可行。","可以说——计算机使用能力的未来非常光明！"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"a16z称，计算机操作智能体在OSWorld-Verified基准上的最佳成绩已从一年前的42%升至85%，高于人类测试者约72%的成绩。文章同时强调，真实业务部署仍受失败率、成本和流程偏离等问题限制。","background":"文章将计算机操作智能体置于后台办公和业务流程外包场景中，涉及更新系统记录、跨门户传递数据、处理工单及操作缺少标准API的软件。当前常见部署方式是模型API配合浏览器或虚拟机、编排、验证和重试机制。","viewpoint":"Aioga判断，基准成绩的跃升说明智能体已具备处理部分标准化、可重复流程的基础，但不能直接等同于稳定完成复杂岗位。文章显示，竞争重点可能从界面操作能力转向权限、上下文、校验、升级和组织流程知识。","implications":"值得关注的是，OSWorld中85%的完成率仍意味着每100项任务约有15项失败，而实际业务流程通常要求每一步都正确。若必须由人工检查全部结果，自动化带来的劳动节省可能有限；成本和异常处理能力也将影响落地范围。","nextStep":"Aioga建议后续观察企业是否披露可复核的生产数据，包括任务完成率、人工介入比例、异常升级方式和单位成本，并将这些指标与具体流程边界对应起来，避免仅以单一基准分数判断产品成熟度。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-08-10T15:23:09.390Z","sourceHash":"24fdc87787199b9d","review":{"approved":true,"groundedness":96,"clarity":94,"duplicationRisk":24,"blockingIssues":[],"notes":["“Aioga判断”和“Aioga建议”明确标示了分析与建议，没有冒充来源事实。","summary、implications均提及失败率及实际业务限制，存在轻微信息重复，但分别承担概述和影响分析功能，不影响通过。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","a16z：News（RSS）"],"translations":{"zh-CN":{"title":"智能体真的会用电脑吗？a16z 用数据给出答案","summary":"a16z 数据显示，计算机操作智能体在 OSWorld-Verified 基准上的最佳成绩已从一年前的 42% 升至 85%，超过人类测试者约 72% 的水平，Claude Fable 5 以 85% 领先。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"行业动态","source":"a16z.news","aggregationSource":"a16z：News（RSS）","pageTitle":"智能体真的会用电脑吗？a16z 用数据给出答案 - Aioga AI资讯","description":"a16z 数据显示，计算机操作智能体在 OSWorld-Verified 基准上的最佳成绩已从一年前的 42% 升至 85%，超过人类测试者约 72% 的水平，Claude Fable 5 以 85% 领先。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk...","url":"https://www.aioga.com/news/cmsnbzkyr01qjrohfmjqlmjk5/"},"en":{"title":"Do agents really know how to use computers? a16z provides the answer with data","summary":"According to a16z data, the best performance of computer-operating agents on the OSWorld-Verified benchmark has risen from 42% a year ago to 85%, surpassing the level of human testers at about 72%, with Claude Fable 5 leading at 85%. 🔗 Read the original via AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"Industry","source":"a16z：News（RSS）","aggregationSource":"a16z：News（RSS）","pageTitle":"Do agents really know how to use computers? a16z provides the answer with data - Aioga AI News","description":"According to a16z data, the best performance of computer-operating agents on the OSWorld-Verified benchmark has risen from 42% a year ago to 85%, surpassing the level of human test...","url":"https://www.aioga.com/en/news/cmsnbzkyr01qjrohfmjqlmjk5/","contentTranslated":true,"sourceHash":"85f73fda431e78af","translatedAt":"2026-08-10T15:21:36.525Z"},"ja":{"title":"エージェントは本当にコンピュータを使えるのか？a16z がデータで答えを示す","summary":"a16z のデータによると、コンピュータ操作エージェントの OSWorld-Verified ベンチマークでの最高成績は、1年前の 42% から 85% に上昇し、人間テスターのおよそ 72% のレベルを超え、Claude Fable 5 は 85% でリードしている。 🔗 原文を読む via AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"業界動向","source":"a16z：News（RSS）","aggregationSource":"a16z：News（RSS）","pageTitle":"エージェントは本当にコンピュータを使えるのか？a16z がデータで答えを示す - Aioga AIニュース","description":"a16z のデータによると、コンピュータ操作エージェントの OSWorld-Verified ベンチマークでの最高成績は、1年前の 42% から 85% に上昇し、人間テスターのおよそ 72% のレベルを超え、Claude Fable 5 は 85% でリードしている。 🔗 原文を読む via AIHOT · https://aihot.virxact.co...","url":"https://www.aioga.com/ja/news/cmsnbzkyr01qjrohfmjqlmjk5/","contentTranslated":true,"sourceHash":"85f73fda431e78af","translatedAt":"2026-08-10T15:21:37.130Z"},"ko":{"title":"에이전트는 정말로 컴퓨터를 사용할 수 있을까? a16z가 데이터를 통해 답을 제시하다","summary":"a16z 데이터에 따르면, 컴퓨터 조작 에이전트의 OSWorld-Verified 벤치마크 최고 점수가 1년 전 42%에서 85%로 상승했으며, 이는 인간 테스트 참가자의 약 72% 수준을 초과했으며, Claude Fable 5가 85%로 선두를 달리고 있다. 🔗 원문 읽기 via AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"업계 동향","source":"a16z：News（RSS）","aggregationSource":"a16z：News（RSS）","pageTitle":"에이전트는 정말로 컴퓨터를 사용할 수 있을까? a16z가 데이터를 통해 답을 제시하다 - Aioga AI 뉴스","description":"a16z 데이터에 따르면, 컴퓨터 조작 에이전트의 OSWorld-Verified 벤치마크 최고 점수가 1년 전 42%에서 85%로 상승했으며, 이는 인간 테스트 참가자의 약 72% 수준을 초과했으며, Claude Fable 5가 85%로 선두를 달리고 있다. 🔗 원문 읽기 via AIHOT · https://aihot.v...","url":"https://www.aioga.com/ko/news/cmsnbzkyr01qjrohfmjqlmjk5/","contentTranslated":true,"sourceHash":"85f73fda431e78af","translatedAt":"2026-08-10T15:21:39.397Z"},"es":{"title":"¿Realmente los agentes inteligentes pueden usar una computadora? a16z da la respuesta con datos","summary":"Los datos de a16z muestran que la mejor puntuación de los agentes inteligentes que operan computadoras en el benchmark OSWorld-Verified ha aumentado del 42 % de hace un año al 85 %, superando aproximadamente el 72 % del nivel de los evaluadores humanos, liderado por Claude Fable 5 con un 85 %. 🔗 Leer el artículo original vía AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"Industria","source":"a16z：News（RSS）","aggregationSource":"a16z：News（RSS）","pageTitle":"¿Realmente los agentes inteligentes pueden usar una computadora? a16z da la respuesta con datos - Aioga Noticias de IA","description":"Los datos de a16z muestran que la mejor puntuación de los agentes inteligentes que operan computadoras en el benchmark OSWorld-Verified ha aumentado del 42 % de hace un año al 85 %...","url":"https://www.aioga.com/es/news/cmsnbzkyr01qjrohfmjqlmjk5/","contentTranslated":true,"sourceHash":"85f73fda431e78af","translatedAt":"2026-08-10T15:21:39.285Z"},"fr":{"title":"Les agents intelligents savent-ils vraiment utiliser un ordinateur ? a16z donne la réponse avec des données","summary":"Les données d’a16z montrent que le score maximal des agents intelligents en manipulation informatique sur le benchmark OSWorld-Verified est passé de 42 % il y a un an à 85 %, dépassant environ 72 % du niveau des testeurs humains, avec Claude Fable 5 en tête à 85 %. 🔗 Lire l’article original via AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"Industrie","source":"a16z：News（RSS）","aggregationSource":"a16z：News（RSS）","pageTitle":"Les agents intelligents savent-ils vraiment utiliser un ordinateur ? a16z donne la réponse avec des données - Aioga Actualités IA","description":"Les données d’a16z montrent que le score maximal des agents intelligents en manipulation informatique sur le benchmark OSWorld-Verified est passé de 42 % il y a un an à 85 %, dépas...","url":"https://www.aioga.com/fr/news/cmsnbzkyr01qjrohfmjqlmjk5/","contentTranslated":true,"sourceHash":"85f73fda431e78af","translatedAt":"2026-08-10T15:21:41.732Z"},"de":{"title":"Können KI-Agenten wirklich Computer benutzen? a16z liefert die Antwort mit Daten","summary":"Daten von a16z zeigen, dass die besten Ergebnisse von Computer-operierenden KI-Agenten im OSWorld-Verified-Benchmark von 42 % vor einem Jahr auf 85 % gestiegen sind, was etwa 72 % über dem Niveau menschlicher Tester liegt. Claude Fable 5 führt mit 85 %. 🔗 Originaltext lesen via AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"行业动态","source":"a16z：News（RSS）","aggregationSource":"a16z：News（RSS）","pageTitle":"Können KI-Agenten wirklich Computer benutzen? a16z liefert die Antwort mit Daten - Aioga KI-News","description":"Daten von a16z zeigen, dass die besten Ergebnisse von Computer-operierenden KI-Agenten im OSWorld-Verified-Benchmark von 42 % vor einem Jahr auf 85 % gestiegen sind, was etwa 72 %...","url":"https://www.aioga.com/de/news/cmsnbzkyr01qjrohfmjqlmjk5/","contentTranslated":true,"sourceHash":"85f73fda431e78af","translatedAt":"2026-08-10T15:21:41.893Z"},"pt-BR":{"title":"Os agentes inteligentes realmente sabem usar computadores? a16z responde com dados","summary":"Dados da a16z mostram que o desempenho dos agentes inteligentes em operações de computador, na referência OSWorld-Verified, aumentou de 42% há um ano para 85%, superando o nível dos testadores humanos em cerca de 72%. Claude Fable 5 lidera com 85%. 🔗 Leia o artigo completo via AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"行业动态","source":"a16z：News（RSS）","aggregationSource":"a16z：News（RSS）","pageTitle":"Os agentes inteligentes realmente sabem usar computadores? a16z responde com dados - Aioga Notícias de IA","description":"Dados da a16z mostram que o desempenho dos agentes inteligentes em operações de computador, na referência OSWorld-Verified, aumentou de 42% há um ano para 85%, superando o nível do...","url":"https://www.aioga.com/pt-BR/news/cmsnbzkyr01qjrohfmjqlmjk5/","contentTranslated":true,"sourceHash":"85f73fda431e78af","translatedAt":"2026-08-10T15:21:43.989Z"},"ru":{"title":"Действительно ли умные агенты умеют пользоваться компьютерами? A16Z предоставляет ответ с помощью данных","summary":"Данные a16z показывают, что лучший балл для операторов компьютеров на бенчмарке OSWorld-Verified вырос с 42% год назад до 85%, превзойдя около 72% среди тестировщиков, при этом Claude Fable 5 лидирует с 85%. 🔗 Прочитайте оригинальную статью на сайте AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"行业动态","source":"a16z：News（RSS）","aggregationSource":"a16z：News（RSS）","pageTitle":"Действительно ли умные агенты умеют пользоваться компьютерами? A16Z предоставляет ответ с помощью данных - Aioga Новости ИИ","description":"Данные a16z показывают, что лучший балл для операторов компьютеров на бенчмарке OSWorld-Verified вырос с 42% год назад до 85%, превзойдя около 72% среди тестировщиков, при этом Cla...","url":"https://www.aioga.com/ru/news/cmsnbzkyr01qjrohfmjqlmjk5/","contentTranslated":true,"sourceHash":"85f73fda431e78af","translatedAt":"2026-08-10T15:21:50.583Z"},"ar":{"title":"هل يمكن للوكيل الذكي فعلاً استخدام الكمبيوتر؟ a16z تقدم الإجابة بالبيانات","summary":"تُظهر بيانات a16z أن أفضل أداء لوكيل العمليات الحاسوبية في معيار OSWorld-Verified قد ارتفع من 42% قبل عام إلى 85%، متجاوزًا مستوى حوالي 72% من المشاركين البشر في الاختبارات، حيث تقدم Claude Fable 5 بنسبة 85%. 🔗 قراءة النص الأصلي عبر AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"行业动态","source":"a16z：News（RSS）","aggregationSource":"a16z：News（RSS）","pageTitle":"هل يمكن للوكيل الذكي فعلاً استخدام الكمبيوتر؟ a16z تقدم الإجابة بالبيانات - Aioga أخبار الذكاء الاصطناعي","description":"تُظهر بيانات a16z أن أفضل أداء لوكيل العمليات الحاسوبية في معيار OSWorld-Verified قد ارتفع من 42% قبل عام إلى 85%، متجاوزًا مستوى حوالي 72% من المشاركين البشر في الاختبارات، حيث تق...","url":"https://www.aioga.com/ar/news/cmsnbzkyr01qjrohfmjqlmjk5/","contentTranslated":true,"sourceHash":"85f73fda431e78af","translatedAt":"2026-08-10T15:21:53.721Z"},"hi":{"title":"क्या एआई वास्तव में कंप्यूटर का उपयोग कर सकते हैं? a16z ने डेटा के साथ उत्तर दिया","summary":"a16z के डेटा से पता चलता है कि कंप्यूटर संचालित एजेंट की OSWorld-Verified बेंचमार्क पर सर्वोत्तम प्रदर्शन एक साल पहले 42% से बढ़कर 85% हो गया है, जो मानव परीक्षकों के लगभग 72% स्तर को पार कर गया है, Claude Fable 5 85% के साथ अग्रणी है। 🔗 मूल लेख पढ़ें AIHOT के माध्यम से · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"行业动态","source":"a16z：News（RSS）","aggregationSource":"a16z：News（RSS）","pageTitle":"क्या एआई वास्तव में कंप्यूटर का उपयोग कर सकते हैं? a16z ने डेटा के साथ उत्तर दिया - Aioga AI समाचार","description":"a16z के डेटा से पता चलता है कि कंप्यूटर संचालित एजेंट की OSWorld-Verified बेंचमार्क पर सर्वोत्तम प्रदर्शन एक साल पहले 42% से बढ़कर 85% हो गया है, जो मानव परीक्षकों के लगभग 72% स्तर...","url":"https://www.aioga.com/hi/news/cmsnbzkyr01qjrohfmjqlmjk5/","contentTranslated":true,"sourceHash":"85f73fda431e78af","translatedAt":"2026-08-10T15:21:54.794Z"},"it":{"title":"Gli agenti intelligenti sanno davvero usare il computer? a16z fornisce una risposta basata sui dati","summary":"I dati di a16z mostrano che il punteggio migliore degli agenti operanti al computer sul benchmark OSWorld-Verified è salito dal 42% di un anno fa all'85%, superando circa il 72% del livello dei tester umani, con Claude Fable 5 in testa all'85%. 🔗 Leggi l'articolo originale via AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"行业动态","source":"a16z：News（RSS）","aggregationSource":"a16z：News（RSS）","pageTitle":"Gli agenti intelligenti sanno davvero usare il computer? a16z fornisce una risposta basata sui dati - Aioga Notizie IA","description":"I dati di a16z mostrano che il punteggio migliore degli agenti operanti al computer sul benchmark OSWorld-Verified è salito dal 42% di un anno fa all'85%, superando circa il 72% de...","url":"https://www.aioga.com/it/news/cmsnbzkyr01qjrohfmjqlmjk5/","contentTranslated":true,"sourceHash":"85f73fda431e78af","translatedAt":"2026-08-10T15:21:57.996Z"},"nl":{"title":"Kunnen AI-agents echt een computer gebruiken? a16z geeft antwoord met data","summary":"Volgens data van a16z is de beste score van computeropererende AI-agents op de OSWorld-Verified benchmark gestegen van 42% een jaar geleden tot 85%, wat ongeveer 72% hoger ligt dan het niveau van menselijke testers. Claude Fable 5 leidt met 85%. 🔗 Lees het originele artikel via AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"行业动态","source":"a16z：News（RSS）","aggregationSource":"a16z：News（RSS）","pageTitle":"Kunnen AI-agents echt een computer gebruiken? a16z geeft antwoord met data - Aioga AI-nieuws","description":"Volgens data van a16z is de beste score van computeropererende AI-agents op de OSWorld-Verified benchmark gestegen van 42% een jaar geleden tot 85%, wat ongeveer 72% hoger ligt dan...","url":"https://www.aioga.com/nl/news/cmsnbzkyr01qjrohfmjqlmjk5/","contentTranslated":true,"sourceHash":"85f73fda431e78af","translatedAt":"2026-08-10T15:21:56.819Z"},"tr":{"title":"Akıllı ajan gerçekten bilgisayar kullanabiliyor mu? a16z verilerle yanıt verdi","summary":"a16z verileri, bilgisayar kullanan akıllı ajanların OSWorld-Verified benchmark'ındaki en iyi sonucunun bir yıl önceki %42'den %85'e yükseldiğini gösteriyor; bu, insan testçilerin yaklaşık %72 seviyesini aşıyor. Claude Fable 5 %85 ile önde. 🔗 Orijinal yazıyı oku via AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"行业动态","source":"a16z：News（RSS）","aggregationSource":"a16z：News（RSS）","pageTitle":"Akıllı ajan gerçekten bilgisayar kullanabiliyor mu? a16z verilerle yanıt verdi - Aioga AI Haberleri","description":"a16z verileri, bilgisayar kullanan akıllı ajanların OSWorld-Verified benchmark'ındaki en iyi sonucunun bir yıl önceki %42'den %85'e yükseldiğini gösteriyor; bu, insan testçilerin y...","url":"https://www.aioga.com/tr/news/cmsnbzkyr01qjrohfmjqlmjk5/","contentTranslated":true,"sourceHash":"85f73fda431e78af","translatedAt":"2026-08-10T15:22:00.699Z"},"vi":{"title":"Các đại lý trí tuệ thực sự có thể sử dụng máy tính không? a16z đưa ra câu trả lời bằng dữ liệu","summary":"Dữ liệu của a16z cho thấy, các đại lý thao tác máy tính đạt điểm tốt nhất trên chuẩn OSWorld-Verified đã tăng từ 42% một năm trước lên 85%, vượt khoảng 72% so với một người thử nghiệm, Claude Fable 5 dẫn đầu với 85%. 🔗 Đọc nguyên văn qua AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"行业动态","source":"a16z：News（RSS）","aggregationSource":"a16z：News（RSS）","pageTitle":"Các đại lý trí tuệ thực sự có thể sử dụng máy tính không? a16z đưa ra câu trả lời bằng dữ liệu - Tin tức AI Aioga","description":"Dữ liệu của a16z cho thấy, các đại lý thao tác máy tính đạt điểm tốt nhất trên chuẩn OSWorld-Verified đã tăng từ 42% một năm trước lên 85%, vượt khoảng 72% so với một người thử ngh...","url":"https://www.aioga.com/vi/news/cmsnbzkyr01qjrohfmjqlmjk5/","contentTranslated":true,"sourceHash":"85f73fda431e78af","translatedAt":"2026-08-10T15:22:00.838Z"},"id":{"title":"Apakah agen cerdas benar-benar bisa menggunakan komputer? a16z memberikan jawaban dengan data","summary":"Data a16z menunjukkan bahwa skor terbaik agen operasi komputer di tolok ukur OSWorld-Verified telah meningkat dari 42% setahun yang lalu menjadi 85%, melampaui sekitar 72% tingkat penguji manusia, dengan Claude Fable 5 memimpin pada 85%. 🔗 Baca artikel asli via AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"行业动态","source":"a16z：News（RSS）","aggregationSource":"a16z：News（RSS）","pageTitle":"Apakah agen cerdas benar-benar bisa menggunakan komputer? a16z memberikan jawaban dengan data - Berita AI Aioga","description":"Data a16z menunjukkan bahwa skor terbaik agen operasi komputer di tolok ukur OSWorld-Verified telah meningkat dari 42% setahun yang lalu menjadi 85%, melampaui sekitar 72% tingkat...","url":"https://www.aioga.com/id/news/cmsnbzkyr01qjrohfmjqlmjk5/","contentTranslated":true,"sourceHash":"85f73fda431e78af","translatedAt":"2026-08-10T15:22:02.899Z"},"th":{"title":"เอเจนต์อัจฉริยะสามารถใช้คอมพิวเตอร์ได้จริงหรือ? a16z ใช้ข้อมูลให้คำตอบ","summary":"ข้อมูลจาก a16z แสดงให้เห็นว่า ผลงานที่ดีที่สุดของเอเจนต์ในการปฏิบัติการบนคอมพิวเตอร์ในมาตรฐาน OSWorld-Verified เพิ่มขึ้นจาก 42% เมื่อหนึ่งปีก่อน เป็น 85% สูงกว่าระดับของผู้ทดสอบมนุษย์ประมาณ 72% Claude Fable 5 นำอยู่ที่ 85% 🔗 อ่านบทความเต็มที่ AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"行业动态","source":"a16z：News（RSS）","aggregationSource":"a16z：News（RSS）","pageTitle":"เอเจนต์อัจฉริยะสามารถใช้คอมพิวเตอร์ได้จริงหรือ? a16z ใช้ข้อมูลให้คำตอบ - ข่าว AI Aioga","description":"ข้อมูลจาก a16z แสดงให้เห็นว่า ผลงานที่ดีที่สุดของเอเจนต์ในการปฏิบัติการบนคอมพิวเตอร์ในมาตรฐาน OSWorld-Verified เพิ่มขึ้นจาก 42% เมื่อหนึ่งปีก่อน เป็น 85% สูงกว่าระดับของผู้ทดสอบมนุ...","url":"https://www.aioga.com/th/news/cmsnbzkyr01qjrohfmjqlmjk5/","contentTranslated":true,"sourceHash":"85f73fda431e78af","translatedAt":"2026-08-10T15:22:03.253Z"},"pl":{"title":"Czy agenci rzeczywiście potrafią korzystać z komputera? a16z odpowiada danymi","summary":"Dane a16z pokazują, że najlepsze wyniki agentów obsługujących komputer w teście OSWorld-Verified wzrosły w ciągu roku z 42% do 85%, przewyższając poziom przeciętnych ludzkich testerów, który wynosi około 72%, a Claude Fable 5 prowadzi z wynikiem 85%. 🔗 Przeczytaj oryginał via AIHOT · https://aihot.virxact.com/items/cmsnbzkyr01qjrohfmjqlmjk5","category":"行业动态","source":"a16z：News（RSS）","aggregationSource":"a16z：News（RSS）","pageTitle":"Czy agenci rzeczywiście potrafią korzystać z komputera? a16z odpowiada danymi - Aioga Wiadomości AI","description":"Dane a16z pokazują, że najlepsze wyniki agentów obsługujących komputer w teście OSWorld-Verified wzrosły w ciągu roku z 42% do 85%, przewyższając poziom przeciętnych ludzkich teste...","url":"https://www.aioga.com/pl/news/cmsnbzkyr01qjrohfmjqlmjk5/","contentTranslated":true,"sourceHash":"85f73fda431e78af","translatedAt":"2026-08-10T15:22:06.780Z"}}}}