{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-27T21:40:51.174Z","headline":"Anthropic 联合 Andon Labs 发布 Drone-Bench，评估 AI 模型自主操控无人机执行定位追踪任务的能力","description":"Anthropic 与 Andon Labs 合作推出 Drone-Bench，用于测试 AI 模型自主操控四旋翼无人机在室内环境中定位并追踪指定人员的能力。该基准将任务分解为 3D 地图重建、定位、导航、目标检测与跟随五个子任务，并通过软件复现实现快速评估。实验表明，该任务链的难度足以区分不同智能水平的模型，并揭示 AI 在物理世界操控能力上的进步轨迹。","url":"https://www.aioga.com/news/cmrz3eamm01jdroey6e0m62on/","mainEntityOfPage":"https://www.aioga.com/news/cmrz3eamm01jdroey6e0m62on/","datePublished":"2026-07-24T15:25:45.300Z","dateModified":"2026-07-24T15:25:45.300Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.anthropic.com/research/project-pilot","https://aihot.virxact.com/items/cmrz3eamm01jdroey6e0m62on"],"canonicalUrl":"https://www.aioga.com/news/cmrz3eamm01jdroey6e0m62on/","directAnswer":{"@type":"Answer","text":"Anthropic 与 Andon Labs 合作开展无人机演示和评估，并形成 Drone-Bench。该基准测试 AI 代理在室内操控四旋翼无人机定位、识别并跟随指定人员的能力，软件化评估用于复现相关任务。","url":"https://www.aioga.com/news/cmrz3eamm01jdroey6e0m62on/","dateCreated":"2026-07-24T15:25:45.300Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"Anthropic source article","url":"https://www.anthropic.com/research/project-pilot","datePublished":"2026-07-24T15:25:45.300Z","provider":{"@type":"Organization","name":"Anthropic","url":"https://www.anthropic.com/research/project-pilot"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmrz3eamm01jdroey6e0m62on","datePublished":"2026-07-24T15:25:45.300Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmrz3eamm01jdroey6e0m62on"}}],"aggregationSource":"Anthropic：Research（发表成果 · 网页）","originalPublisher":{"name":"Anthropic","url":"https://www.anthropic.com/research/project-pilot"},"article":{"id":"cmrz3eamm01jdroey6e0m62on","slug":"cmrz3eamm01jdroey6e0m62on","url":"https://www.aioga.com/news/cmrz3eamm01jdroey6e0m62on/","title":"Anthropic 联合 Andon Labs 发布 Drone-Bench，评估 AI 模型自主操控无人机执行定位追踪任务的能力","title_en":"Project Pilot： Can AI control a drone？","summary":"Anthropic 与 Andon Labs 合作推出 Drone-Bench，用于测试 AI 模型自主操控四旋翼无人机在室内环境中定位并追踪指定人员的能力。该基准将任务分解为 3D 地图重建、定位、导航、目标检测与跟随五个子任务，并通过软件复现实现快速评估。实验表明，该任务链的难度足以区分不同智能水平的模型，并揭示 AI 在物理世界操控能力上的进步轨迹。","source":"Anthropic：Research（发表成果 · 网页）","sourceUrl":"https://www.anthropic.com/research/project-pilot","aiHotUrl":"https://aihot.virxact.com/items/cmrz3eamm01jdroey6e0m62on","publishedAt":"2026-07-24T15:25:45.300Z","category":"论文研究","score":73,"selected":true,"articleBody":["Several of our research projects over the last year have looked at how frontier models interact with the physical world. In Project：https://www.anthropic.com/research/project-vend-1 Vend：https://www.anthropic.com/research/project-vend-2, AI models ran a small shop; Project Fetch：https://www.anthropic.com/research/project-fetch-robot-dog was an early look at robots as the intermediary between digital models and physical objects. As we recently noted in Project Fetch: Phase two：https://www.anthropic.com/research/project-fetch-phase-two, we’re already seeing improvements in model capability such that their ability to use off-the-shelf robots is on track to approach the ease with which coding agents use software tools.","Working again with our partners at Andon Labs：https://andonlabs.com/, we developed a new series of demonstrations and evaluations that assess AI models’ ability to use a flying drone to autonomously perform a simple locate-and-follow task of the kind used in aerial surveillance, culminating in a new benchmark: Drone-Bench.","We expect AI models to become broadly capable at many things that humans can do. Operating hardware, in particular robots, is one such capability. Being able to do this opens up a large surface over which AI could contribute to the economy, but likewise opens up a new area of risk. A key reason why Anthropic has a Frontier Red Team：https://www.anthropic.com/research/team/frontier-red-team is to measure capabilities like this, giving us situational awareness into how close we are to the world in which AI can autonomously pilot robots—with all the attendant benefits and risks. Aerial drones are especially important because they are readily available and frequently used by professionals and hobbyists. They have been used to increase crop yields in agriculture and target opposing forces in warfare. Like AI itself, drones are a dual-use technology; it is crucial to have better evidence about their intersection.","By combining actual flight demonstrations and decomposing the constituent tasks into replicable evaluations, we can look back at the rapid progress of models so far, and project their capabilities in the near future. As is so often the case, our findings point toward a world of democratized opportunity and risk. Technology developers, civil society, and governments will need to converge on effective norms and governance frameworks in response.","The core task we tested in Project Fetch—getting a robot dog to retrieve a beach ball—was neither especially practical nor especially concerning. In this project, we chose an objective with clearer utility and policy relevance: a simple locate-and-follow task used in aerial surveillance. Capabilities like automated person-detection and tracking can have legitimate purposes such as search and rescue, disaster response, and lawful public safety uses. But this is a class of capabilities that is also subject to abuse, either through overreach of a legitimate authority or by unaccountable private individuals or organizations. The work we report here thus more closely matches the “dual-use” nature of AI models.","In these experiments, we ask the model to control a quad-rotor drone in an indoor office environment in order to locate and follow a person. 1 This requires a number of complex sub-tasks. The AI model needs to develop schema for controlling the aircraft, mapping and navigating the obstacle-laden indoor space, finding the target individual from a reference photo, and following them (plus reacquiring the target if they move out of frame).","Individually, there are known algorithms for accomplishing all of these tasks. What is not trivial is for the AI model to understand the challenges, identify the preexisting resources it can use to solve them, adapt those off-the-shelf solutions to its current situation, and execute the mission in real time. As we will see, the difficulty—both individually and in chaining these tasks together—is sufficient to distinguish between models of varying intelligence and plot the trajectory of capability improvement.","Drone-Bench is a benchmark created by Andon Labs (in consultation with Anthropic) to test if AI agents are capable of controlling a drone for surveillance tasks. Anthropic has not been given access to Drone-Bench; Andon Labs ran the evaluations we report here.","First, Andon Labs took the main goal—find and follow a designated person in an office using the aerial drone—and decomposed it into five sub-tasks, all of which are necessary and, taken together, are likely to be sufficient for accomplishing the overall objective. These sub-tasks are:","Next, each of these real-world tasks was reproduced in software so that we could run the models through them multiple times and far faster than needing to set up the physical demo for each instance (this is an improvement over Project Fetch, for example, which was an entirely physical experiment).","It was also important to establish a meaningful baseline of performance. Human-only baselines increasingly don’t reflect the reality of contemporary software engineering, so Andon worked with coding agents to develop algorithms for each sub-task. Putting all of these algorithms crafted by human-AI teams together allowed them to demonstrate end-to-end success, as shown in the below video.","A task is considered completed if the model meets or exceeds the baseline. Thus, if a model can complete all tasks, we can infer that it has the ability to autonomously control a drone to do at least as well on this surveillance task as the team at Andon Labs did.","For more details, check out Andon Labs’ post：https://andonlabs.com/evals/drone-bench about Drone-Bench.","It is worth underscoring that the evaluation’s baseline is neither the floor of unassisted human capability nor the ceiling of what is possible with concerted human-AI collaboration. Rather, it is indicative of what can be achieved in the present by AI experts (but not full-time roboticists) using a realistic suite of modern tools. The interesting question is if and when models operating essentially autonomously reliably pass this baseline of reasonable and realistic effort, as that is the point at which pressure to reduce human oversight may intensify—making deliberate, use case-specific judgments about the appropriate human role all the more important.","Andon tested 15 models from three developers: GPT-4o, GPT-4o Nov, o1, o3, Claude Opus 4, Gemini 2.5 Pro, GPT-5, Gemini 3.1 Pro, Opus 4.5, GPT-5.2, Opus 4.7, GPT-5.5, Opus 4.8, Fable 5, and GPT-5.6 Sol. The overall trend we observe is that newer models get successively further on all sub-tasks. Of these tasks, models are most successful at detection and following, and least successful at reconstruction and localization.","The best performing model was Claude Fable 5, which brings the frontier past the baseline on all tasks except reconstruction. When we then tested its ability to execute the entire demonstration end-to-end on the real drone, it performed noticeably better than the baseline at detecting and following.","However, due to errors from reconstruction that compounded in localization and navigation, it was unable to autonomously navigate between rooms (as you can see in the first part of the below video).","Clearly, Fable’s failure to accurately reconstruct the room is a huge stumbling block. But given models’ capabilities in the other phases, it really just amounts to the missing piece. Once it’s in place, end-to-end performance will suddenly be within reach. This is an advantage of decomposing the evaluation into constituent tasks: we are better positioned to avoid surprise. What would look like a discontinuous jump is revealed to be gradual progress in several necessary, but not sufficient, sub-tasks.","The sub-task view also surfaces encouraging signs. A trend we're seeing when reading Fable 5's submissions is that the model is doing local analysis before submitting its implementation. In one submission, the model calculated the drone's camera extrinsics by analyzing a video from the simulation, estimating the camera tilt to within four degrees of the true value by using the grout lines on the floor to recover the scene's vanishing point. You can see its process below:","In another run, Fable 5 built a 2D top-down reconstruction of what it thought the Follow task's environment looked like, so it could test and iterate on its implementation locally before burning a submission.","The environment differs from the real environment (seen below), but it helped Fable catch some easy bugs!","It’s important to understand how consistently models reach the reference level of performance, as well as whether they can reach it. Here there is obvious room for improvement. When we run 10 simulations, the models reach the human baseline in at least one simulation for four of five tasks. But even Fable 5, the current frontier model, reaches the human baseline on average for only three of the five tasks—and that level of consistency followed six months after the human baseline was exceeded as a one-off for the first time.","Although the complexity of this experiment is greater than some of our previous work and the operating environment of an actual (or simulated) office is more challenging than a wide-open warehouse, the experiment has important limitations: the drones are moving at slow speeds, we only tested in one office floorplan with a limited number of people, and Andon did not test outdoors in large crowds, among many other factors that would have made this more realistic. We still think this pilot provides a real signal about the direction of model capabilities: this evaluation will provide meaningful information about the underlying performance and reliability of models for autonomous targeting and tracking, even though more realistic and diverse experiments would be needed to assess operational capability.","This experiment highlights the potential of commercial-off-the-shelf (COTS) hardware and AI-tailored software to support useful, but possibly risky, tasks.","It is important to take seriously the parallel between AI models’ use of software in agentic coding and AI models’ control of hardware. In the early days of agentic coding, humans approved nearly every tool call. But after only a few months, models are now much more trusted to execute long-horizon tasks with minimal intervention.","In project Fetch, we examined how humans can use models to get robots to perform complex tasks. Now, we investigate many models on a large variety of different robotics tasks in simulation, to see how good models are at controlling robots themselves.","Get updates on our latest red-teaming research and findings."],"articleImages":[{"sourceUrl":"https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F614042bcff1dd452489fcab83e880ed3f880efaf-1999x1446.png&w=3840&q=75","alt":"Drone-Bench step chart: four of five tasks reach near 100% of baseline by mid-2026, while Reconstruct lags at about 47%.","afterParagraph":14,"url":"/media/articles/cmrz3eamm01jdroey6e0m62on/7cbd8cf06b1a7f16.webp"},{"sourceUrl":"https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2Faf663cbd16ca4072d246f34e2faf44236e832ef4-2064x1372.png&w=3840&q=75","alt":"Four views of the same simulated corridor: floor segmentation, edge detection, line detection, and vanishing-point estimate.","afterParagraph":18,"url":"/media/articles/cmrz3eamm01jdroey6e0m62on/18f71b29847ff643.webp"},{"sourceUrl":"https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F7713a1ff0377744cfd0849e70c1c75a41ef8cd28-1999x1380.png&w=3840&q=75","alt":"Drone-Bench chart: models' average run trails their best run by about six months in progress towards baseline.","afterParagraph":21,"url":"/media/articles/cmrz3eamm01jdroey6e0m62on/e5b105857df38e1a.webp"}],"mediaStatus":"ok","articleBodyZh":["在过去一年中，我们的几个研究项目都关注了前沿模型如何与物理世界互动。在项目：https://www.anthropic.com/research/project-vend-1 Vend：https://www.anthropic.com/research/project-vend-2 中，AI 模型管理了一家小商店；项目 Fetch：https://www.anthropic.com/research/project-fetch-robot-dog 是对机器人作为数字模型与物理对象之间中介的早期探索。正如我们最近在项目 Fetch：第二阶段：https://www.anthropic.com/research/project-fetch-phase-two 中提到的，我们已经看到模型能力的提升，使得它们使用现成机器人（off-the-shelf robots）的能力有望接近编码代理使用软件工具的便利性。","再次与我们的合作伙伴 Andon Labs：https://andonlabs.com/ 合作，我们开发了一系列新的演示和评估，评估 AI 模型使用飞行无人机自主执行简单的定位与跟随任务的能力，这类任务通常用于空中监控，最终形成了一个新的基准测试：Drone-Bench。","我们预计 AI 模型将在许多人类能做到的事情上变得普遍有能力。操作硬件，特别是机器人，就是这样的一种能力。能够做到这一点，为 AI 在经济中发挥作用打开了广阔的空间，但同样也带来了新的风险领域。Anthropic 设立前沿红队（Frontier Red Team）：https://www.anthropic.com/research/team/frontier-red-team 的一个关键原因就是为了测量这类能力，从而让我们对 AI 能否自主操作机器人这一世界有清晰的情境认知——以及随之而来的所有利弊。空中无人机尤其重要，因为它们易于获取，且职业人士和爱好者经常使用。它们已被用于提高农业产量和在战争中锁定敌方目标。像 AI 本身一样，无人机也是一项双用途技术；因此，获得关于它们交叉应用的更好证据至关重要。","通过结合实际飞行演示并将组成任务分解为可复制的评估，我们可以回顾迄今为止模型的快速进展，并预测它们在不久的将来的能力。正如常见情况一样，我们的发现指向一个机会与风险普及化的世界。技术开发者、公民社会和政府将需要在应对时汇聚有效的规范和治理框架。","我们在Project Fetch中测试的核心任务——让机器人狗去取回沙滩球——既不特别实用也不特别令人担忧。在这个项目中，我们选择了一个具有更明确用途和政策相关性的目标：用于空中监控的简单定位与跟随任务。像自动人员检测和跟踪这样的能力可以有合法的用途，如搜救、灾害应对和合法的公共安全使用。但这也是一类可能被滥用的能力，无论是通过合法机构的越权，还是通过不负责的个人或组织。因此，我们在此报告的工作更贴近AI模型的“二重用途”性质。","在这些实验中，我们要求模型控制室内办公室环境中的四旋翼无人机，以便定位并跟随一个人。这需要若干复杂的子任务。AI模型需要开发控制飞行器的方案，映射和导航障碍密布的室内空间，从参考照片中找到目标个体，并进行跟随（如果目标移出画面，还需重新获取目标）。","单独来看，实现所有这些任务都有已知算法。非平凡之处在于让AI模型理解这些挑战，识别可用于解决它们的现有资源，将这些现成解决方案适应当前情况，并实时执行任务。正如我们将看到的，无论是单个任务的难度，还是将这些任务串联起来的难度，都足以区分不同智能水平的模型，并描绘能力改进的轨迹。","Drone-Bench 是由 Andon Labs（在与 Anthropic 协商后）创建的一个基准，用于测试 AI 代理是否能够控制无人机进行监控任务。Anthropic 尚未获得 Drone-Bench 的访问权限；Andon Labs 进行了我们在此报告的评估。","首先，Andon Labs 将主要目标——使用空中无人机在办公室中找到并跟踪指定人员——分解为五个子任务，这些子任务都是必要的，并且综合起来很可能足以完成整体目标。这些子任务包括：","接下来，每个这些现实世界的任务都被在软件中再现，以便我们可以多次运行模型，并比每次设置物理演示快得多（这比 Project Fetch 的方法有所改进，例如，Project Fetch 完全是物理实验）。","建立有意义的性能基线也很重要。仅靠人类的基线越来越不能反映当代软件工程的现实，因此 Andon 与编码代理合作，为每个子任务开发算法。将这些由人类和 AI 团队共同制作的算法组合在一起，使他们能够展示端到端的成功，如下视频所示。","如果模型满足或超过基线，则认为任务完成。因此，如果一个模型能够完成所有任务，我们可以推断它有能力自主控制无人机，在这一监控任务上的表现至少与 Andon Labs 团队一样好。","更多详情，请查看 Andon Labs 关于 Drone-Bench 的文章：https://andonlabs.com/evals/drone-bench。","值得强调的是，该评估的基线既不是无人协助的人类能力下限，也不是通过人类与 AI 合作可以达到的上限。而是表明在使用现实的现代工具的情况下，AI 专家（但不是全职机器人专家）目前可以实现的水平。一个有趣的问题是，当模型基本自主运行时，是否以及何时能够可靠地通过这一合理且现实的努力基线，因为这正是减少人类监督压力可能加大的时刻——从而使得对适当人类角色的审慎、针对特定用例的判断尤为重要。","Andon 测试了来自三家开发商的 15 个模型：GPT-4o、GPT-4o Nov、o1、o3、Claude Opus 4、Gemini 2.5 Pro、GPT-5、Gemini 3.1 Pro、Opus 4.5、GPT-5.2、Opus 4.7、GPT-5.5、Opus 4.8、Fable 5 和 GPT-5.6 Sol。我们观察到的总体趋势是，较新的模型在所有子任务上表现越来越好。在这些任务中，模型在检测和跟随方面最成功，在重建和定位方面最不成功。","表现最好的模型是 Claude Fable 5，它在所有任务上除了重建之外都超过了基线。当我们随后测试它在真实无人机上执行整个示范的端到端能力时，它在检测和跟随方面的表现明显优于基线。","然而，由于重建产生的误差在定位和导航中累积，它无法自主地在各个房间间导航（如您在下面视频的前半部分所见）。","显然，Fable 无法准确重建房间是一个巨大的障碍。但鉴于模型在其他阶段的能力，这实际上只是缺少的环节。一旦这个环节到位，端到端的性能就会突然触手可及。这是将评估分解为各个组成任务的一个优势：我们更有能力避免意外。看似断断续续的飞跃，实际上是在几个必要但不充分的子任务中逐步取得的进展。","子任务视角还显现了令人鼓舞的迹象。当我们阅读 Fable 5 的提交内容时，看到的趋势是模型在提交实现前会进行局部分析。在一次提交中，模型通过分析模拟视频计算无人机相机的外参，通过利用地板上的填缝线恢复场景的消失点，将相机倾斜角估计到真实值的四度以内。您可以在下面看到它的过程：","在另一轮运行中，Fable 5 构建了其认为 Follow 任务环境的二维俯视重建图，以便在提交前可以在本地测试和迭代其实现。","该环境与现实环境有所不同（见下图），但它帮助 Fable 捕捉到一些简单的错误！","了解模型达到参考性能水平的稳定性以及它们是否能达到这一水平非常重要。在这方面，显然仍有改进空间。当我们运行 10 次模拟时，模型在五个任务中有四个任务至少有一次达到了人类基准。但即使是当前的前沿模型 Fable 5，平均也仅在五个任务中的三个任务上达到人类基准——而这种稳定性是在首次超过人类基准六个月后才实现的。","尽管这个实验的复杂性高于我们以前的一些工作，并且实际（或模拟）办公室的操作环境也比开放的仓库更具挑战性，但这个实验有重要的局限性：无人机的移动速度较慢，我们只在一个办公楼层平面图中测试，且参与人数有限，并且 Andon 没有在大规模人群的户外测试，此外还有许多其他因素会使实验更接近现实。尽管如此，我们仍认为这一试点实验提供了关于模型能力方向的真实信号：评估将提供有关模型在自主目标定位和跟踪方面底层性能和可靠性的有意义信息，尽管需要更真实和多样化的实验来评估操作能力。","该实验突出了现成商用（COTS）硬件和针对 AI 的软件在支持有用但可能存在风险的任务方面的潜力。","认真对待 AI 模型在自主编程中使用软件与 AI 模型控制硬件之间的类比非常重要。在自主编程的早期阶段，人类几乎批准了每一个工具调用。但仅仅几个月后，模型在执行长周期任务时已经被认为更值得信任，所需的人类干预极少。","在 Fetch 项目中，我们研究了人类如何使用模型让机器人执行复杂任务。现在，我们在模拟环境中研究许多模型在各种不同机器人任务上的表现，以观察模型自身控制机器人的能力有多强。","获取我们最新红队研究和发现的更新。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Anthropic 与 Andon Labs 合作开展无人机演示和评估，并形成 Drone-Bench。该基准测试 AI 代理在室内操控四旋翼无人机定位、识别并跟随指定人员的能力，软件化评估用于复现相关任务。","background":"研究延续 Anthropic 对前沿模型与物理世界交互的关注。正文称 Drone-Bench 由 Andon Labs 在 Anthropic 咨询下创建，Anthropic 未获得该基准的访问权限，文中报告的评估也由 Andon Labs 执行。","viewpoint":"Aioga 判断，Drone-Bench 的重点不只是单项算法表现，而是模型能否识别任务难点、调用并适配现成资源，再把地图构建、定位、导航、目标检测和跟随等环节串联起来实时执行。","implications":"该研究将无人机定位跟随视为双重用途能力：可用于搜救、灾害响应和合法公共安全，也可能被权力越界或缺乏问责的个人与组织滥用。值得关注的是，能力评估与治理规范需要同步推进。","nextStep":"后续值得关注不同模型在各子任务及完整任务链上的表现变化，并结合真实飞行演示与可复现评估持续观察能力进展。技术开发者、公民社会和政府也需讨论相应规范与治理框架。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-07-26T07:49:10.838Z","sourceHash":"b45bc76e5ea4c8a1","review":{"approved":true,"groundedness":96,"clarity":91,"duplicationRisk":18,"blockingIssues":[],"notes":["“软件化评估用于复现相关任务”可改为“将子任务转化为可复现的软件评估”，与来源表述更贴合。","“能力评估与治理规范需要同步推进”属于基于来源的观点性概括，已置于 implications 字段，不构成事实错误。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","low-source-overlap","no-html","independent-ai-review"]}},"tags":["论文研究","Anthropic：Research（发表成果 · 网页）"],"translations":{"zh-CN":{"title":"Anthropic 联合 Andon Labs 发布 Drone-Bench，评估 AI 模型自主操控无人机执行定位追踪任务的能力","summary":"Anthropic 与 Andon Labs 合作推出 Drone-Bench，用于测试 AI 模型自主操控四旋翼无人机在室内环境中定位并追踪指定人员的能力。该基准将任务分解为 3D 地图重建、定位、导航、目标检测与跟随五个子任务，并通过软件复现实现快速评估。实验表明，该任务链的难度足以区分不同智能水平的模型，并揭示 AI 在物理世界操控能力上的进步轨迹。","category":"论文研究","source":"Anthropic","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"Anthropic 联合 Andon Labs 发布 Drone-Bench，评估 AI 模型自主操控无人机执行定位追踪任务的能力 - Aioga AI资讯","description":"Anthropic 与 Andon Labs 合作推出 Drone-Bench，用于测试 AI 模型自主操控四旋翼无人机在室内环境中定位并追踪指定人员的能力。该基准将任务分解为 3D 地图重建、定位、导航、目标检测与跟随五个子任务，并通过软件复现实现快速评估。实验表明，该任务链的难度足以区分不同智能水平的模型，并揭示 AI 在物理世界操控能力上的进步轨迹。","url":"https://www.aioga.com/news/cmrz3eamm01jdroey6e0m62on/"},"en":{"title":"Anthropic and Andon Labs jointly released Drone-Bench to evaluate AI models' ability to autonomously operate drones to perform positioning and tracking tasks","summary":"Anthropic has partnered with Andon Labs to launch Drone-Bench, which is used to test AI models' ability to autonomously control quadrotor drones to locate and track designated individuals in indoor environments. The benchmark breaks the task down into five subtasks: 3D map reconstruction, localization, navigation, target detection, and following, and enables rapid evaluation through software reproduction. Experiments show that the difficulty of this task chain is sufficient to distinguish models with different levels of intelligence and reveals the progress of AI in physical world manipulation capabilities.","category":"Research","source":"Anthropic：Research（发表成果 · 网页）","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"Anthropic and Andon Labs jointly released Drone-Bench to evaluate AI models' ability to autonomously operate drones to perform positioning and tracking tasks - Aioga AI News","description":"Anthropic has partnered with Andon Labs to launch Drone-Bench, which is used to test AI models' ability to autonomously control quadrotor drones to locate and track designated indi...","url":"https://www.aioga.com/en/news/cmrz3eamm01jdroey6e0m62on/","contentTranslated":true,"sourceHash":"4cf2044d608c0676","translatedAt":"2026-07-26T02:25:11.599Z"},"ja":{"title":"AnthropicはAndon Labsと共同でDrone-Benchを発表し、AIモデルが無人機を自律的に操作して位置追跡タスクを実行する能力を評価します","summary":"AnthropicはAndon Labsと協力してDrone-Benchを立ち上げ、AIモデルが屋内環境でクアッドコプターを自主的に操縦し、指定された人物を位置特定および追跡する能力をテストします。このベンチマークはタスクを3Dマップ再構築、位置特定、ナビゲーション、ターゲット検出と追跡の5つのサブタスクに分解し、ソフトウェアシミュレーションを通じて迅速な評価を可能にします。実験によれば、このタスクチェーンの難易度は異なる知能レベルのモデルを区別するのに十分であり、物理世界でのAIの操作能力の進展の軌跡を明らかにします。","category":"論文研究","source":"Anthropic：Research（发表成果 · 网页）","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"AnthropicはAndon Labsと共同でDrone-Benchを発表し、AIモデルが無人機を自律的に操作して位置追跡タスクを実行する能力を評価します - Aioga AIニュース","description":"AnthropicはAndon Labsと協力してDrone-Benchを立ち上げ、AIモデルが屋内環境でクアッドコプターを自主的に操縦し、指定された人物を位置特定および追跡する能力をテストします。このベンチマークはタスクを3Dマップ再構築、位置特定、ナビゲーション、ターゲット検出と追跡の5つのサブタスクに分解し、ソフトウェアシミュレーションを通じて迅速な評...","url":"https://www.aioga.com/ja/news/cmrz3eamm01jdroey6e0m62on/","contentTranslated":true,"sourceHash":"4cf2044d608c0676","translatedAt":"2026-07-26T02:25:21.980Z"},"ko":{"title":"Anthropic가 Andon Labs와 함께 Drone-Bench를 발표하여 AI 모델이 자율적으로 드론을 조종하여 위치 추적 임무를 수행하는 능력을 평가합니다.","summary":"Anthropic은 Andon Labs와 협력하여 Drone-Bench를 출시했으며, 이는 AI 모델이 실내 환경에서 쿼드콥터 드론을 자율적으로 조종하여 특정 인물을 위치 추적하는 능력을 테스트하는 데 사용됩니다. 이 벤치마크는 작업을 3D 지도 재구성, 위치 확인, 내비게이션, 목표 탐지 및 추적의 다섯 가지 하위 작업으로 분해하고, 소프트웨어 재현을 통해 빠른 평가를 실행합니다. 실험 결과, 이 작업 체인의 난이도는 서로 다른 지능 수준의 모델을 구분할 수 있을 정도이며, AI의 물리적 세계 조작 능력 향상 경로를 보여줍니다.","category":"연구","source":"Anthropic：Research（发表成果 · 网页）","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"Anthropic가 Andon Labs와 함께 Drone-Bench를 발표하여 AI 모델이 자율적으로 드론을 조종하여 위치 추적 임무를 수행하는 능력을 평가합니다. - Aioga AI 뉴스","description":"Anthropic은 Andon Labs와 협력하여 Drone-Bench를 출시했으며, 이는 AI 모델이 실내 환경에서 쿼드콥터 드론을 자율적으로 조종하여 특정 인물을 위치 추적하는 능력을 테스트하는 데 사용됩니다. 이 벤치마크는 작업을 3D 지도 재구성, 위치 확인, 내비게이션, 목표 탐지 및 추적의 다섯 가지 하위 작업...","url":"https://www.aioga.com/ko/news/cmrz3eamm01jdroey6e0m62on/","contentTranslated":true,"sourceHash":"4cf2044d608c0676","translatedAt":"2026-07-26T02:25:58.673Z"},"es":{"title":"Anthropic y Andon Labs lanzan Drone-Bench, evaluando la capacidad de los modelos de IA para controlar de forma autónoma drones en la ejecución de tareas de localización y seguimiento","summary":"Anthropic se asoció con Andon Labs para lanzar Drone-Bench, utilizado para probar la capacidad de los modelos de IA de controlar de forma autónoma drones cuatrirrotor para localizar y rastrear personas específicas en entornos interiores. Este estándar descompone la tarea en cinco sub-tareas: reconstrucción de mapas 3D, localización, navegación, detección de objetivos y seguimiento, y permite una evaluación rápida mediante la reproducción por software. Los experimentos muestran que la dificultad de esta cadena de tareas es suficiente para diferenciar entre modelos con distintos niveles de inteligencia y revela la trayectoria de progreso de la IA en la capacidad de manipulación en el mundo físico.","category":"Investigación","source":"Anthropic：Research（发表成果 · 网页）","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"Anthropic y Andon Labs lanzan Drone-Bench, evaluando la capacidad de los modelos de IA para controlar de forma autónoma drones en la ejecución de tareas de localización y seguimiento - Aioga Noticias de IA","description":"Anthropic se asoció con Andon Labs para lanzar Drone-Bench, utilizado para probar la capacidad de los modelos de IA de controlar de forma autónoma drones cuatrirrotor para localiza...","url":"https://www.aioga.com/es/news/cmrz3eamm01jdroey6e0m62on/","contentTranslated":true,"sourceHash":"4cf2044d608c0676","translatedAt":"2026-07-26T02:25:57.475Z"},"fr":{"title":"Anthropic et Andon Labs ont lancé Drone-Bench, évaluant la capacité des modèles d'IA à contrôler de manière autonome des drones pour accomplir des missions de suivi de localisation.","summary":"Anthropic s'est associé à Andon Labs pour lancer Drone-Bench, destiné à tester la capacité des modèles d'IA à contrôler de manière autonome des drones quadrirotors pour localiser et suivre des personnes spécifiques dans des environnements intérieurs. Cette référence divise la tâche en cinq sous-tâches : reconstruction de carte 3D, localisation, navigation, détection de cible et suivi, et permet une évaluation rapide grâce à une reproduction logicielle. Les expériences montrent que cette chaîne de tâches est suffisamment difficile pour différencier des modèles de niveaux d'intelligence différents et révèle la trajectoire des progrès de l'IA dans la capacité de manipulation du monde physique.","category":"Recherche","source":"Anthropic：Research（发表成果 · 网页）","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"Anthropic et Andon Labs ont lancé Drone-Bench, évaluant la capacité des modèles d'IA à contrôler de manière autonome des drones pour accomplir des missions de suivi de localisation. - Aioga Actualités IA","description":"Anthropic s'est associé à Andon Labs pour lancer Drone-Bench, destiné à tester la capacité des modèles d'IA à contrôler de manière autonome des drones quadrirotors pour localiser e...","url":"https://www.aioga.com/fr/news/cmrz3eamm01jdroey6e0m62on/","contentTranslated":true,"sourceHash":"4cf2044d608c0676","translatedAt":"2026-07-26T02:26:34.246Z"},"de":{"title":"Anthropic und Andon Labs haben Drone-Bench veröffentlicht, um die Fähigkeit von KI-Modellen zu bewerten, Drohnen autonom bei der Durchführung von Positionierungs- und Verfolgungsaufgaben zu steuern.","summary":"Anthropic arbeitet mit Andon Labs zusammen, um Drone-Bench einzuführen, das verwendet wird, um die Fähigkeit von KI-Modellen zu testen, eine Quadrocopter-Drohne autonom in Innenräumen zu lokalisieren und eine bestimmte Person zu verfolgen. Dieses Benchmark unterteilt die Aufgabe in fünf Unteraufgaben: 3D-Kartenrekonstruktion, Lokalisierung, Navigation, Zielerkennung und Verfolgung, und ermöglicht eine schnelle Bewertung durch Software-Reproduktion. Experimente zeigen, dass die Schwierigkeit dieser Aufgabenfolge ausreicht, um Modelle mit unterschiedlichem Intelligenzniveau zu unterscheiden und den Fortschritt der KI bei der Steuerung in der physischen Welt aufzuzeigen.","category":"论文研究","source":"Anthropic：Research（发表成果 · 网页）","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"Anthropic und Andon Labs haben Drone-Bench veröffentlicht, um die Fähigkeit von KI-Modellen zu bewerten, Drohnen autonom bei der Durchführung von Positionierungs- und Verfolgungsaufgaben zu steuern. - Aioga KI-News","description":"Anthropic arbeitet mit Andon Labs zusammen, um Drone-Bench einzuführen, das verwendet wird, um die Fähigkeit von KI-Modellen zu testen, eine Quadrocopter-Drohne autonom in Innenräu...","url":"https://www.aioga.com/de/news/cmrz3eamm01jdroey6e0m62on/","contentTranslated":true,"sourceHash":"4cf2044d608c0676","translatedAt":"2026-07-26T02:26:34.504Z"},"pt-BR":{"title":"A Anthropic, em parceria com a Andon Labs, lançou o Drone-Bench, avaliando a capacidade de modelos de IA de controlar autonomamente drones na execução de tarefas de rastreamento de localização.","summary":"A Anthropic, em parceria com a Andon Labs, lançou o Drone-Bench, usado para testar a capacidade de modelos de IA de controlar autonomamente drones quadricópteros para localizar e rastrear pessoas específicas em ambientes internos. Este benchmark divide a tarefa em cinco subtarefas: reconstrução de mapas 3D, localização, navegação, detecção de alvo e acompanhamento, e permite avaliação rápida através de reprodução por software. Os experimentos mostram que a complexidade dessa cadeia de tarefas é suficiente para diferenciar modelos com diferentes níveis de inteligência e revelar o progresso da IA nas capacidades de controle no mundo físico.","category":"论文研究","source":"Anthropic：Research（发表成果 · 网页）","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"A Anthropic, em parceria com a Andon Labs, lançou o Drone-Bench, avaliando a capacidade de modelos de IA de controlar autonomamente drones na execução de tarefas de rastreamento de localização. - Aioga Notícias de IA","description":"A Anthropic, em parceria com a Andon Labs, lançou o Drone-Bench, usado para testar a capacidade de modelos de IA de controlar autonomamente drones quadricópteros para localizar e r...","url":"https://www.aioga.com/pt-BR/news/cmrz3eamm01jdroey6e0m62on/","contentTranslated":true,"sourceHash":"4cf2044d608c0676","translatedAt":"2026-07-26T02:27:09.612Z"},"ru":{"title":"Anthropic совместно с Andon Labs выпустили Drone-Bench, оценивающий способность моделей ИИ самостоятельно управлять дронами для выполнения задач по позиционному отслеживанию.","summary":"Anthropic совместно с Andon Labs запустили Drone-Bench, предназначенный для тестирования способности моделей ИИ самостоятельно управлять квадрокоптерами в помещении, определять местоположение и отслеживать заданных людей. Этот бенчмарк разбивает задачу на пять подзадач: 3D-реконструкция карты, локализация, навигация, обнаружение цели и сопровождение, и позволяет проводить быструю оценку через программную эмуляцию. Эксперименты показывают, что сложность этой цепочки задач достаточна для различения моделей с разным уровнем интеллекта и выявляет траекторию прогресса ИИ в управлении физическим миром.","category":"论文研究","source":"Anthropic：Research（发表成果 · 网页）","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"Anthropic совместно с Andon Labs выпустили Drone-Bench, оценивающий способность моделей ИИ самостоятельно управлять дронами для выполнения задач по позиционному отслеживанию. - Aioga Новости ИИ","description":"Anthropic совместно с Andon Labs запустили Drone-Bench, предназначенный для тестирования способности моделей ИИ самостоятельно управлять квадрокоптерами в помещении, определять мес...","url":"https://www.aioga.com/ru/news/cmrz3eamm01jdroey6e0m62on/","contentTranslated":true,"sourceHash":"4cf2044d608c0676","translatedAt":"2026-07-26T02:27:17.453Z"},"ar":{"title":"أصدرت شركة Anthropic بالتعاون مع Andon Labs منصة Drone-Bench لتقييم قدرة نماذج الذكاء الاصطناعي على التحكم الذاتي في الطائرات بدون طيار لأداء مهام التتبع وتحديد المواقع","summary":"تعاونت شركة Anthropic مع Andon Labs لإطلاق Drone-Bench، لاختبار قدرة نماذج الذكاء الاصطناعي على التحكم الذاتي في الطائرات الرباعية لتحليق في بيئات داخلية وتحديد وتتبع الأشخاص المحددين. يقسم هذا المعيار المهمة إلى خمس مهام فرعية تشمل إعادة بناء خرائط ثلاثية الأبعاد وتحديد المواقع والملاحة والكشف عن الأهداف والمتابعة، ويتم التقييم السريع من خلال محاكاة البرمجيات. أظهرت التجارب أن سلسلة المهام هذه صعبة بما يكفي لتمييز نماذج الذكاء الاصطناعي ذات مستويات الذكاء المختلفة، وتكشف عن مسار تقدم الذكاء الاصطناعي في القدرة على التحكم في العالم الفيزيائي.","category":"论文研究","source":"Anthropic：Research（发表成果 · 网页）","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"أصدرت شركة Anthropic بالتعاون مع Andon Labs منصة Drone-Bench لتقييم قدرة نماذج الذكاء الاصطناعي على التحكم الذاتي في الطائرات بدون طيار لأداء مهام التتبع وتحديد المواقع - Aioga أخبار الذكاء الاصطناعي","description":"تعاونت شركة Anthropic مع Andon Labs لإطلاق Drone-Bench، لاختبار قدرة نماذج الذكاء الاصطناعي على التحكم الذاتي في الطائرات الرباعية لتحليق في بيئات داخلية وتحديد وتتبع الأشخاص المحد...","url":"https://www.aioga.com/ar/news/cmrz3eamm01jdroey6e0m62on/","contentTranslated":true,"sourceHash":"4cf2044d608c0676","translatedAt":"2026-07-26T02:27:58.536Z"},"hi":{"title":"Anthropic ने Andon Labs के साथ मिलकर Drone-Bench जारी किया, जो AI मॉडल की ड्रोन को स्वतः संचालित करके लोकेशन ट्रैकिंग कार्य करने की क्षमता का मूल्यांकन करता है","summary":"Anthropic ने Andon Labs के साथ मिलकर Drone-Bench लॉन्च किया, जिसका उपयोग AI मॉडल की क्षमता का परीक्षण करने के लिए किया जाता है कि वे इनडोर वातावरण में क्वाडकॉप्टर ड्रोन को स्वायत्त रूप से नियंत्रित करके निर्धारित व्यक्ति का पता लगाने और ट्रैक करने में सक्षम हैं। यह बेंचमार्क कार्य को पांच उप-कार्य में विभाजित करता है: 3D मानचित्र पुनर्निर्माण, पोजिशनिंग, नेविगेशन, ऑब्जेक्ट डिटेक्शन और फॉलो। यह सॉफ़्टवेयर सिमुलेशन के माध्यम से तेज़ मूल्यांकन की अनुमति देता है। प्रयोगों से पता चलता है कि यह कार्य श्रृंखला विभिन्न बुद्धिमता स्तरों वाले मॉडल को अलग करने के लिए पर्याप्त चुनौतीपूर्ण है, और AI की भौतिक दुनिया में नियंत्रण क्षमता में प्रगति के मार्ग का खुलासा करती है।","category":"论文研究","source":"Anthropic：Research（发表成果 · 网页）","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"Anthropic ने Andon Labs के साथ मिलकर Drone-Bench जारी किया, जो AI मॉडल की ड्रोन को स्वतः संचालित करके लोकेशन ट्रैकिंग कार्य करने की क्षमता का मूल्यांकन करता है - Aioga AI समाचार","description":"Anthropic ने Andon Labs के साथ मिलकर Drone-Bench लॉन्च किया, जिसका उपयोग AI मॉडल की क्षमता का परीक्षण करने के लिए किया जाता है कि वे इनडोर वातावरण में क्वाडकॉप्टर ड्रोन को स्वायत्त...","url":"https://www.aioga.com/hi/news/cmrz3eamm01jdroey6e0m62on/","contentTranslated":true,"sourceHash":"4cf2044d608c0676","translatedAt":"2026-07-26T02:28:01.884Z"},"it":{"title":"Anthropic e Andon Labs hanno lanciato Drone-Bench, per valutare la capacità dei modelli AI di controllare autonomamente i droni nell'esecuzione di compiti di localizzazione e tracciamento","summary":"Anthropic ha collaborato con Andon Labs per lanciare Drone-Bench, utilizzato per testare la capacità dei modelli di IA di controllare autonomamente droni quadricotteri per localizzare e seguire persone specifiche in ambienti interni. Questo benchmark scompone il compito in cinque sotto-compiti: ricostruzione di mappe 3D, localizzazione, navigazione, rilevamento e inseguimento degli obiettivi, e realizza una valutazione rapida tramite simulazione software. Gli esperimenti mostrano che questa catena di compiti è sufficientemente difficile da distinguere modelli con diversi livelli di intelligenza e rivela la traiettoria dei progressi dell'IA nella capacità di controllo nel mondo fisico.","category":"论文研究","source":"Anthropic：Research（发表成果 · 网页）","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"Anthropic e Andon Labs hanno lanciato Drone-Bench, per valutare la capacità dei modelli AI di controllare autonomamente i droni nell'esecuzione di compiti di localizzazione e tracciamento - Aioga Notizie IA","description":"Anthropic ha collaborato con Andon Labs per lanciare Drone-Bench, utilizzato per testare la capacità dei modelli di IA di controllare autonomamente droni quadricotteri per localizz...","url":"https://www.aioga.com/it/news/cmrz3eamm01jdroey6e0m62on/","contentTranslated":true,"sourceHash":"4cf2044d608c0676","translatedAt":"2026-07-26T02:28:38.661Z"},"nl":{"title":"Anthropic heeft samen met Andon Labs Drone-Bench uitgebracht om het vermogen van AI-modellen te evalueren om drones autonoom te besturen bij het uitvoeren van positionerings- en volgopdrachten.","summary":"Anthropic werkt samen met Andon Labs om Drone-Bench te lanceren, waarmee het vermogen van AI-modellen wordt getest om een quadcopter zelfstandig te besturen en een specifiek persoon in een binnenomgeving te lokaliseren en te volgen. Deze benchmark verdeelt de taak in vijf subtaken: 3D-kaartreconstructie, lokalisatie, navigatie, doelobjectdetectie en volgen, en maakt snelle evaluatie mogelijk via software-simulatie. Experimenten tonen aan dat de moeilijkheidsgraad van deze takenketen voldoende is om modellen met verschillende niveaus van intelligentie te onderscheiden en het ontwikkelingspad van AI in fysieke besturingscapaciteiten te onthullen.","category":"论文研究","source":"Anthropic：Research（发表成果 · 网页）","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"Anthropic heeft samen met Andon Labs Drone-Bench uitgebracht om het vermogen van AI-modellen te evalueren om drones autonoom te besturen bij het uitvoeren van positionerings- en volgopdrachten. - Aioga AI-nieuws","description":"Anthropic werkt samen met Andon Labs om Drone-Bench te lanceren, waarmee het vermogen van AI-modellen wordt getest om een quadcopter zelfstandig te besturen en een specifiek persoo...","url":"https://www.aioga.com/nl/news/cmrz3eamm01jdroey6e0m62on/","contentTranslated":true,"sourceHash":"4cf2044d608c0676","translatedAt":"2026-07-26T02:28:38.420Z"},"tr":{"title":"Anthropic, Andon Labs ile birlikte Drone-Bench'i yayımladı ve AI modellerinin dronları bağımsız olarak kontrol ederek konum izleme görevlerini gerçekleştirme yeteneğini değerlendirdi.","summary":"Anthropic, Andon Labs ile iş birliği yaparak Drone-Bench'i tanıttı; bu, AI modellerinin iç mekanlarda dört rotorlu insansız hava aracını özerk olarak kontrol etme ve belirli kişileri konumlandırıp takip etme yeteneklerini test etmek için kullanılıyor. Bu kıstas, görevi 3B harita yeniden inşası, konumlandırma, navigasyon, nesne tespiti ve takip olmak üzere beş alt göreve ayırıyor ve hızlı değerlendirme için yazılım simülasyonu yoluyla uygulanıyor. Deneyler, bu görev zincirinin zorluk seviyesinin farklı zeka düzeylerindeki modelleri ayırt edebilecek kadar yüksek olduğunu ve AI'nın fiziksel dünyadaki kontrol yeteneklerindeki ilerleme yolunu ortaya koyduğunu gösteriyor.","category":"论文研究","source":"Anthropic：Research（发表成果 · 网页）","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"Anthropic, Andon Labs ile birlikte Drone-Bench'i yayımladı ve AI modellerinin dronları bağımsız olarak kontrol ederek konum izleme görevlerini gerçekleştirme yeteneğini değerlendirdi. - Aioga AI Haberleri","description":"Anthropic, Andon Labs ile iş birliği yaparak Drone-Bench'i tanıttı; bu, AI modellerinin iç mekanlarda dört rotorlu insansız hava aracını özerk olarak kontrol etme ve belirli kişile...","url":"https://www.aioga.com/tr/news/cmrz3eamm01jdroey6e0m62on/","contentTranslated":true,"sourceHash":"4cf2044d608c0676","translatedAt":"2026-07-26T02:29:17.186Z"},"vi":{"title":"Anthropic hợp tác với Andon Labs phát hành Drone-Bench, đánh giá khả năng của mô hình AI trong việc tự động điều khiển drone thực hiện nhiệm vụ theo dõi định vị","summary":"Anthropic hợp tác với Andon Labs ra mắt Drone-Bench, dùng để thử nghiệm khả năng mô hình AI tự điều khiển máy bay không người lái quadcopter trong môi trường trong nhà để định vị và theo dõi người được chỉ định. Bộ chuẩn này phân tách nhiệm vụ thành năm nhiệm vụ con: tái tạo bản đồ 3D, định vị, điều hướng, phát hiện mục tiêu và theo dõi, đồng thời thực hiện đánh giá nhanh thông qua tái hiện phần mềm. Thí nghiệm cho thấy chuỗi nhiệm vụ này đủ khó để phân biệt các mô hình với mức độ thông minh khác nhau và tiết lộ tiến trình cải thiện khả năng điều khiển trong thế giới vật lý của AI.","category":"论文研究","source":"Anthropic：Research（发表成果 · 网页）","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"Anthropic hợp tác với Andon Labs phát hành Drone-Bench, đánh giá khả năng của mô hình AI trong việc tự động điều khiển drone thực hiện nhiệm vụ theo dõi định vị - Tin tức AI Aioga","description":"Anthropic hợp tác với Andon Labs ra mắt Drone-Bench, dùng để thử nghiệm khả năng mô hình AI tự điều khiển máy bay không người lái quadcopter trong môi trường trong nhà để định vị v...","url":"https://www.aioga.com/vi/news/cmrz3eamm01jdroey6e0m62on/","contentTranslated":true,"sourceHash":"4cf2044d608c0676","translatedAt":"2026-07-26T02:29:16.896Z"},"id":{"title":"Anthropic bekerja sama dengan Andon Labs merilis Drone-Bench, untuk menilai kemampuan model AI dalam mengendalikan drone secara mandiri untuk menjalankan tugas pelacakan posisi","summary":"Anthropic bekerja sama dengan Andon Labs meluncurkan Drone-Bench, yang digunakan untuk menguji kemampuan model AI dalam mengendalikan quadcopter secara mandiri untuk menempatkan dan melacak orang yang ditentukan di lingkungan dalam ruangan. Tolok ukur ini membagi tugas menjadi lima sub-tugas: rekonstruksi peta 3D, penentuan posisi, navigasi, deteksi target, dan pengikutan, serta memungkinkan evaluasi cepat melalui reproduksi perangkat lunak. Eksperimen menunjukkan bahwa rantai tugas ini cukup sulit untuk membedakan model dengan tingkat kecerdasan yang berbeda dan mengungkapkan lintasan kemajuan AI dalam kemampuan mengendalikan dunia fisik.","category":"论文研究","source":"Anthropic：Research（发表成果 · 网页）","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"Anthropic bekerja sama dengan Andon Labs merilis Drone-Bench, untuk menilai kemampuan model AI dalam mengendalikan drone secara mandiri untuk menjalankan tugas pelacakan posisi - Berita AI Aioga","description":"Anthropic bekerja sama dengan Andon Labs meluncurkan Drone-Bench, yang digunakan untuk menguji kemampuan model AI dalam mengendalikan quadcopter secara mandiri untuk menempatkan da...","url":"https://www.aioga.com/id/news/cmrz3eamm01jdroey6e0m62on/","contentTranslated":true,"sourceHash":"4cf2044d608c0676","translatedAt":"2026-07-26T02:29:54.655Z"},"th":{"title":"Anthropic ร่วมกับ Andon Labs เปิดตัว Drone-Bench เพื่อประเมินความสามารถของโมเดล AI ในการควบคุมโดรนอัตโนมัติเพื่อปฏิบัติภารกิจติดตามตำแหน่ง","summary":"Anthropic ร่วมกับ Andon Labs เปิดตัว Drone-Bench สำหรับทดสอบความสามารถของโมเดล AI ในการควบคุมโดรนสี่ใบพัดโดยอัตโนมัติเพื่อระบุตำแหน่งและติดตามบุคคลที่กำหนดในสภาพแวดล้อมภายในอาคาร เกณฑ์มาตรฐานนี้แบ่งงานออกเป็น 5 งานย่อย ได้แก่ การสร้างแผนที่ 3D การระบุตำแหน่ง การนำทาง การตรวจจับเป้าหมาย และการติดตาม และใช้การจำลองผ่านซอฟต์แวร์เพื่อประเมินผลอย่างรวดเร็ว การทดลองแสดงให้เห็นว่าลำดับงานนี้มีความยากเพียงพอที่จะสามารถแยกระดับสติปัญญาที่แตกต่างกันของโมเดลได้ และเผยให้เห็นเส้นทางความก้าวหน้าของ AI ในการควบคุมโลกทางกายภาพ","category":"论文研究","source":"Anthropic：Research（发表成果 · 网页）","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"Anthropic ร่วมกับ Andon Labs เปิดตัว Drone-Bench เพื่อประเมินความสามารถของโมเดล AI ในการควบคุมโดรนอัตโนมัติเพื่อปฏิบัติภารกิจติดตามตำแหน่ง - ข่าว AI Aioga","description":"Anthropic ร่วมกับ Andon Labs เปิดตัว Drone-Bench สำหรับทดสอบความสามารถของโมเดล AI ในการควบคุมโดรนสี่ใบพัดโดยอัตโนมัติเพื่อระบุตำแหน่งและติดตามบุคคลที่กำหนดในสภาพแวดล้อมภายในอาคาร เ...","url":"https://www.aioga.com/th/news/cmrz3eamm01jdroey6e0m62on/","contentTranslated":true,"sourceHash":"4cf2044d608c0676","translatedAt":"2026-07-26T02:30:06.354Z"},"pl":{"title":"Anthropic i Andon Labs wspólnie wydali Drone-Bench, oceniający zdolność modeli AI do autonomicznego sterowania dronami w wykonywaniu zadań lokalizacyjnego śledzenia","summary":"Anthropic we współpracy z Andon Labs uruchomiło Drone-Bench, służące do testowania zdolności modeli AI do samodzielnego sterowania dronem typu quadcopter w celu lokalizacji i śledzenia określonych osób w środowisku wewnętrznym. Benchmark dzieli zadanie na pięć podzadań: odtwarzanie mapy 3D, lokalizację, nawigację, wykrywanie celu i śledzenie, oraz umożliwia szybkie ocenianie dzięki odtworzeniu w oprogramowaniu. Eksperymenty wykazały, że łańcuch tych zadań ma wystarczający poziom trudności, aby odróżnić modele o różnym poziomie inteligencji, i ujawnia ścieżkę postępu AI w zakresie zdolności sterowania w świecie fizycznym.","category":"论文研究","source":"Anthropic：Research（发表成果 · 网页）","aggregationSource":"Anthropic：Research（发表成果 · 网页）","pageTitle":"Anthropic i Andon Labs wspólnie wydali Drone-Bench, oceniający zdolność modeli AI do autonomicznego sterowania dronami w wykonywaniu zadań lokalizacyjnego śledzenia - Aioga Wiadomości AI","description":"Anthropic we współpracy z Andon Labs uruchomiło Drone-Bench, służące do testowania zdolności modeli AI do samodzielnego sterowania dronem typu quadcopter w celu lokalizacji i śledz...","url":"https://www.aioga.com/pl/news/cmrz3eamm01jdroey6e0m62on/","contentTranslated":true,"sourceHash":"4cf2044d608c0676","translatedAt":"2026-07-26T02:30:48.827Z"}}}}