文章列出从手部灵巧性、视觉理解、规划反应、速度、力量、安全、续航散热、泛化边界情况到算力、规模化供应链和社会监管等十余项待克服的挑战,认为物理世界的 demo 与现实差距会比 LLM 基准与实际工作的差距更大。
人工智能的进展正迅速推进,但几乎所有可见的进展都集中在知识工作领域,即可以在计算机内部进行的活动。
在旧金山的人工智能圈子里,人们普遍认为机器人很快也会登场。与开发广泛能力的人工智能的竞赛同时进行的,还有同样激烈的开发广泛能力机器人的竞赛——那些拥有身体智能的人形机器。人工工作者可以做饭和打扫、搬运和运送……以及其他所有事情,包括制造更多的自己,根据许多预测,这将带来以“爆炸式”描述最为恰当的经济增长。
换句话说,人们的想法是,数据中心里的人工智能很快将包揽所有智力劳动,而拥有身体的人工智能很快将包揽所有体力劳动。然而,有一个重要的区别:虽然我们能看到智力领域的进展,但人工智能在物理方面大多局限于测试设施和演示视频。还没有类似于ChatGPT的机器人——没有哪一个你我,甚至是大多数人工智能社区的人可以实际接触到。
因此,我们只能依赖演示视频。不幸的是,它们是评估进展的一个很差的工具。我们可能只是在看到100次尝试中的一次成功。场景可能被精心安排,以避免机器人尚未准备好的挑战。视频可能经过编辑,使机器人看起来动作更快、更可靠,而实际上并非如此。这里有一个非常令人印象深刻的演示:https://x.com/deanwball/status/2052282028775637307 ……但可疑的是有大量的镜头切换。
(我还没有太多机会观看最近世界人形机器人比赛的视频:https://www.nature.com/articles/d41586-026-02713-z。这些比赛的视频有价值,因为它们提供了一个较少受挑选影响的公共平台。我看过的少数视频中,有一些令人印象深刻的壮举,但并没有解决我在下面列出的许多挑战……而且也有很多令人惊叹的失败。)
演示吸引人们注意机器人已经能够完成的事情。那么问题就变成了:缺少的是什么?在今天的文章中,我将列出在实现广泛能力的人工工人道路上必须克服的技术挑战。下一次当你看到机器人做出令人印象深刻的动作时,问问自己:机器人展示了这些能力中的哪些?演示场景可能在避开哪些挑战?
(请注意,如果考虑轮式机器人而不是严格的人形机器人,某些挑战会变得更容易。轮式机器人可以承载更多重量,这意味着力量、耐力和电子设备动力的挑战较小。而且轮式机器人不太可能被撞倒。但它们无法爬楼梯1:#footnote-1、跨越杂物,或者调整角度以伸入橱柜。)
人类的手是工程奇迹——可对掌的拇指,诸如此类的特性。人手大约有二十几个“自由度”(每个关节可以弯曲的不同方向或独立关节),大约有17,000个触觉传感器。我们的脑可以用极其精巧的控制手部,通过触觉、视觉,甚至听觉线索完成各种精密任务,精准而可靠。
当前的机器人“操控器”只是苍白的模仿。一些现有的机器人手在某一个或某几个物理属性上可以达到人类标准。例如,有些有多达27个自由度。然而,没有任何一种能够接近匹配人类手在灵活性、敏感性、力量、可靠性及其他物理特性方面的整体组合。正是这种因素的组合:https://itcanthink.substack.com/p/robot-hands-are-getting-better#:~:text=for%20the%20first%20time%2C%20there%20really%20appear%20to%20be%20roughly%20human%20equivalent%20hands%2C%20in%20terms%20of%20dexterity%20and%20sensing%20if%20not%20manufacturability%2C%20strength%2C%20and%20robustness. 尤其难以匹敌:https://x.com/ErenChenAI/status/2078942723864920533,即便演示越来越令人印象深刻:https://x.com/BerntBornich/status/2075253825494237660。例如,一些公司已经设法将数千个触觉传感器塞入机器人的指尖,但没有一个能够让这些微小传感器经受住高强度使用2:#footnote-2。
控制问题可能与物理构建问题一样具有挑战性。一个称职的机器人必须能够找到抓取复杂物体的正确关节位置;规划折叠衬衫、翻煎蛋或在受限空间拧紧螺栓的动作序列;并处理柔软或松散的材料(这可能需要对突然的移动瞬间反应)。
过去十年或二十年中,计算机视觉取得了令人难以置信的进展(并且促成了导致大型语言模型的深度学习繁荣)。但理解复杂的视觉场景——从拥挤的环境中辨认出物体、理解应该从哪里抓取、确定脚步可以安全落地的位置以及如何避免碰倒东西——仍然不是一个解决的问题。
通用机器人必须能够将任务分解为各个步骤,并将这些步骤与其环境关联。如何移动手臂将螺丝刀插入机械装置?通向架子后部香料瓶的最快捷路径是什么?客厅地板上的物品应以什么顺序拾起?
真正的自主性将要求计划规模更大、复杂度更高的任务:做一顿饭、维修浴室管道、修理发动机。更不用说在意外情况出现时重新规划——卡住的螺栓、腐烂的农产品、冲进厨房的孩子。
当当前的人工智能在知识工作任务中失败时,往往是因为它们没有获得足够的上下文。机器人也需要上下文:物资储存在何处?你喜欢食物如何烹饪?养老院的居民需要多少帮助,其步态中的这个小问题是正常的,还是即将绊倒的迹象?
一旦有了上下文,机器人将需要进行推理、计划,并运用判断力和常识。基于大语言模型(LLM)的系统,如ChatGPT和Claude,在这些方面取得了很大进展,但物理领域带来了额外的挑战3:#footnote-3。LLM的成功得益于可用于训练的大规模现有数据池——包含了所有已写书籍的相当大一部分、网络以及其他大量现有数据。对于物理任务来说,要匹配这种数据的广度和深度将会很困难。没有简单的“直接对应”。高效学习、泛化能力和适应性/在职学习似乎是必需的。
AI代理大多孤立操作,并在静态环境中工作。我们很少让它们面对不断变化的环境,或者要求它们进行协调。当我们这么做时,情况往往会失控:https://theaidigest.org/village/goal/design-run-write-up-human-subjects-experiment。在虚拟世界中,孤立操作更容易安排,可以随意创建私人工作空间,而且没有东西过重无法独自搬动。机器人常常需要与人类或其他机器人合作。
当今的通用机器人通常移动速度远低于人类。挑战包括力量、控制(速度越快,计划和反应的时间越少)以及安全性(快速移动的机器人击中你会更重,也更难闪避)。
对于某些应用来说,慢而稳可能完全可以接受:比如我可能不在意家庭机器人用一整夜整理和叠衣服。但慢动作的机器人在做厨师或养老院助手时用途不大。它在仓库里可能会碍事。而且它完成足够工作的能力较差,难以抵消自身成本。
一些工业机器人异常强大。但人形机器人——或其他高度移动的“通用”机器人——通常不是。将强度与可控重量、大量关节及灵活结构结合起来很困难。强力电机会产生更多热量并更快耗尽电池——这是机器人已经面临的两大难题(见下文)。而强大、沉重的机器人带来了更大的安全挑战。
关于通用机器人适当的外形设计,尤其是下半身,目前仍没有定论。它们应该使用轮子还是腿?两条腿、四条腿还是其他数量?轮子成本更低、更稳定、更可靠;腿更适合越过障碍物和爬楼梯。双足结构更灵活,但一旦出现问题,也更容易摔倒。无论如何,问题是:机器人是否能可靠地在工作环境中移动?
通用机器人的安全问题几乎无穷无尽。一个故障或 malfunction 的机器人可能会撞到人、倒在他们身上、掉落物体砸到他们、溅液体在他们身上、打碎玻璃,或者引发火灾。
基于大型语言模型(LLM)的智能体的安全性部分依赖于对离散行为的审查,例如尝试发送邮件或删除文件。机器人不断移动,很难将少数特定动作单独挑出来作为需要审查的潜在危险行为。
如果自动驾驶汽车遇到无法处理的情况或发生故障,它可以靠边停车,或者在最坏的情况下踩刹车。一个通用机器人突然冻结,可能会在炉子上留下东西、在行走中摔倒,或者绊倒它所辅助的人。
当然,危险也可能由人类行为引发,比如一个孩子突然冲到机器人前面。我更希望我的孩子被一个柔软的人撞一下,而不是被金属机器人;按照现状,我宁愿依赖人类的反应和适应能力来避免绊到这个小调皮鬼。
(机器人越强大、越重、越有能力,风险就越大。)
如今的双足机器人通常可以在需要充电前运行几个小时。我怀疑这不会成为限制因素:如果需要解决方法,我们会找到办法,无论是更换电池组、地面充电网,还是通过线缆连接到附近的大电池移动装置 4:#footnote-4。
然后是可靠性问题。机器人有大量的活动部件,其中许多部件必然很挑剔,因为它们被设计用于在尺寸、重量和性能上突破极限。因此,目前对通用人形机器人的尝试经常出现故障:https://blog.robozaps.com/b/challenges-in-humanoid-robotics。(相比之下,人类的身体能够不断从磨损中恢复,并具有相当的自我修复能力,以及弥补小故障的能力。)
Waymo 自动驾驶汽车在处理边缘情况时很困难。最近有人观察到它们驾驶进入积水道路:https://www.nytimes.com/2026/05/22/us/waymo-taxi-suspended-atlanta.html?unlocked_article_code=1.xFA.j1L_.UCPkIKcQKUkk&smid=url-share,或者驶入燃烧的烟花中:https://www.foxnews.com/us/terrified-passengers-film-waymo-autonomous-vehicle-driving-into-live-fireworks-san-francisco。尽管如此,自动驾驶汽车的研发已经进行了二十多年,特别是 Waymo 汽车已经行驶超过 2.2 亿英里:https://waymo.com/blog/shorts/safetydata-june26 —— 是典型美国人一生驾驶里程的 250 倍。
(对于烟花这一点,我愿意给它们留些余地;250周年纪念可不常见。但我很困惑的是,Waymo 的工程技术能如此稳健,以致取得惊人的安全记录:https://secondthoughts.ai/p/autonomous-vehicles-will-save-lives,却又如此粗心大意地轻易驶入深水中。)
在驾驶过程中会出现各种奇怪的边缘情况——经典例子是一只鸭子被持扫帚的轮椅女士追赶:https://www.youtube.com/watch?v=weXDUc5Osto —— 在家庭和企业中操作的机器人将遇到更多这样的情况。它们将面临更广泛的任务,使用更各种工具(不同的工具;不同的机器人身体去抓取这些工具),在更各种环境中操作。在工厂车间这种可控环境之外实现可靠的操作,可能是最难的挑战。
灵巧、协调、视觉理解、规划、反应、理解、合作——并且要快速、安全、可靠地完成所有这些——将需要大量的计算能力。将必要的计算能力整合到机器人本身会增加成本、消耗电池并导致过热。将机器人的大脑放在云端会减慢反应时间,并引入新的故障模式(Wi-Fi中断 → 机器人瘫痪)。
自从ChatGPT发布以来,LLM能够快速扩展,因为现有硬件(GPU)和制造设施(芯片厂)可以轻松适应支持这一新用例。
先进机器人的大规模供应链尚不存在。即使我们有可行的设计,也可能需要数年时间才能实现大规模制造、部署和维护。这类事情不会一蹴而就;特斯拉从首次商业销售到第一百万辆汽车的年份花了14年:#footnote-5。一项分析:https://epoch.ai/publications/how-fast-could-robot-production-scale-up 发现,一旦起跑信号发出(先进人形机器人变得经济有价值),可能需要数年时间才能扩展到每年生产数百万台机器人。相比之下,ChatGPT在发布两个月内就达到了1亿用户6:#footnote-6。
政治、监管、组织和文化障碍:对工作岗位替代和安全的担忧可能限制机器人使用的地点和方式。许多法规在制定时并未考虑机器人——机器人是否算入最低人员配备要求?企业会迅速采用机器人吗?他们会关心可靠性、安全性、责任和维护问题吗?客户愿意接受机器人服务吗?
网络安全、监控、滥用:先进机器人可能是犯罪分子的梦想工具——一个可靠的爪牙,不会背叛主人7:#footnote-7。防止这种情况可能需要对所有机器人进行持续监控,这会引发各种问题。机器人的网络安全必须做到滴水不漏。
成本:一旦其他障碍得到解决,我怀疑这不会成为限制因素。一个不知疲倦的工人,价格相当于一辆新车,将是一个划算的交易,而类人机器人将比汽车小得多、轻得多,运动部件也更少 8:#footnote-8。可能的是,能够胜任任务的机器人的组件、制造技术和训练过程会使早期型号远比汽车昂贵。但即便是昂贵的机器人,也可能会找到早期使用场景——例如,进行危险工作。
在机器人成为能够、可靠、实用的工人,并能在严格控制的工厂环境之外工作之前,许多障碍需要被清除。我列出的清单肯定不完整;而且我仅仅稍微提到了认知方面——理解、规划、行动和反应。
演示视频提供了机器人在理想情况下能完成任务的一个窗口。它们也可能让我们分心,忽视剩余的局限性。对于知识工作而言,人工智能在基准测试与现实世界工作之间存在着持续的性能差距。在物理世界中,我怀疑演示与现实的差距会更大。玩弄大型语言模型来看看它能做什么一直是可行的、便宜的,并且(大部分情况下!)安全的。而要评估机器人的能力,我们将更多依赖于受控演示和制造商声称的能力。绘制它们能力的参差不齐边界将更加困难。
当我为这篇文章做最后润色时,优秀的 Understanding AI 博客发布了《为什么类人机器人短期内无法赶上人类工人》:https://www.understandingai.org/p/why-humanoid-robots-wont-catch-up。我还没读它,但确信值得一看。
想要全面了解机器人能力,可以参考 Epoch 于 2026 年 2 月发布的报告:https://epoch.ai/publications/where-autonomy-works-evaluating-robot-capabilities-in-2026。
分享:https://secondthoughts.ai/p/14-reasons-robotics-is-hard?utm_source=substack&utm_medium=email&utm_content=share&action=share
感谢阿比·奥尔维拉、阿维·帕拉克和塔伦·斯坦布里克纳-考夫曼。
有一些四足机器人,它们的“脚”是轮子,可以在行走和滚动之间切换,并且能够爬楼梯。但到目前为止,没有迹象表明这种设计会成为主流。
可能会找到一些变通方法。例如,一些机器人设计使用摄像头来部分弥补触觉灵敏度,在手腕或其他部位增加额外的“眼睛”;机器人甚至可以通过 Wi-Fi 使用独立的摄像头。机器人也可以整合非人类感官,例如超声波。
请参阅这篇关于生物实验室工作中默会知识重要性的解释:https://secondthoughts.ai/p/tacit-knowledge-the-missing-factor,作为一个例子,说明机器人需要学习的内容之多。
最近的一次演示:https://interestingengineering.com/ai-robotics/figure-03-humanoid-robot-200-hour-shift 显示了一台机器人连续工作 200 小时进行包裹分类,处理了 250,000 个包裹……使用了三台不同的机器人轮流工作,当其他机器人充电时轮换进行。这既是可靠性又是耐久性的演示。
假设犯罪分子设法避免了将自己与机器人联系起来的纸质记录。
至少,相较于内燃机汽车而言。
1. 我想补充的一点是,我们的手能够自我修复累积的损伤和磨损。机器人手无法做到这一点——它们可能难以在成本可控的情况下保持可接受的功能性,尤其是如果它们极其复杂以尝试匹配人类能力的话。
2. 我不知道过热问题。这太残酷了——如果它产生很大的噪音,人们不会允许这些东西在家里乱跑,就像你不会容忍吸尘器始终在你 10 英尺范围内开着一样。
很棒的帖子。同时,人类能量学、代谢和热散发的高效率不可低估。人类可以起床吃一个百吉饼和咖啡,并拥有几小时操作所需的全部能量,这简直令人惊叹。此外,人类还利用极少的能量进行持续的自我修复。进化在这个领域已优化了数百万年。要复制这一点,可谓困难重重。这并不意味着人类无法从与机械工具/机器人协同工作的关系中受益,但我认为我们距离复制人体生物学的全部功能还有很长的路要走。
AI progress is racing along, but virtually all of the visible progress is in the realm of knowledge work, i.e. activities that can take place inside a computer.
In the San Francisco AI scene, there is a widespread belief that robots will soon enter the picture. In parallel with the race to develop broadly capable AI, there is an equally aggressive race to develop broadly capable robots – humanoid machines imbued with physical intelligence. Artificial workers that can cook and clean, fetch and carry… and do everything else, including building more of themselves, leading (in many forecasts) to economic growth best characterized as an “explosion”.
In other words, the thinking goes, AI in the data center will soon subsume all intellectual labor, and AI in humanoid bodies will soon subsume all physical labor. However, there is an important difference: while we can see progress in the intellectual realm, the physical side of AI is mostly confined to test facilities and demo videos. There is no robot equivalent to ChatGPT – nothing that you or I, or even most people in the AI community, can get our hands on.
So we’re stuck with demo videos. Unfortunately, they are a poor tool for assessing progress. We might be seeing the one successful task achieved in 100 attempts. The scenario might have been carefully arranged to avoid challenges the robot isn’t ready for. The video might be edited to make it look like the robot is acting with more speed and reliability than is actually the case. Here’s one very impressive demo :https://x.com/deanwball/status/2052282028775637307 … with a suspiciously large number of camera cuts.
(I have not yet had much chance to watch videos from the recent World Humanoid Robot Games :https://www.nature.com/articles/d41586-026-02713-z . These are valuable for providing a public platform less amenable to cherry-picking. The handful of videos I’ve watched include some impressive feats, but don’t address many of the challenges I list below… and there are also a lot of spectacular failures.)
Demos draw attention to the things a robot can already do. The question then becomes: what’s missing? In today’s post, I’ll catalog the technical challenges that will have to be overcome along the road to broadly capable artificial workers. The next time you watch a robot doing something impressive, ask yourself: which of these capabilities has the robot demonstrated, and which challenges might the demo scenario be avoiding?
(Note that some challenges get easier if we consider wheeled robots rather than strictly humanoid robots. A wheeled robot can carry more weight, meaning that strength, endurance, and power for electronics are less of a challenge. And wheeled robots are less likely to fall over. But they can’t climb stairs 1:#footnote-1 , step over clutter, or angle themselves to reach into a cupboard.)
The human hand is an engineering miracle – opposable thumbs, and all that. It has roughly two dozen “degrees of freedom” (distinct joints and/or directions in which each joint can bend), and approximately 17,000 tactile sensors. Our brains can control our hands with exquisite grace, using touch, sight, and even auditory cues to carry out all manner of delicate tasks, precisely and reliably.
Current robot “manipulators” are a pale imitation. Some existing robot hands can match the human standard on one or another physical attribute. For example, some have as many as 27 degrees of freedom. However, none come close to matching the overall package of flexibility, sensitivity, strength, reliability, and other physical attributes. It is the combination of factors :https://itcanthink.substack.com/p/robot-hands-are-getting-better#:~:text=for%20the%20first%20time%2C%20there%20really%20appear%20to%20be%20roughly%20human%20equivalent%20hands%2C%20in%20terms%20of%20dexterity%20and%20sensing%20if%20not%20manufacturability%2C%20strength%2C%20and%20robustness. that is especially difficult to match :https://x.com/ErenChenAI/status/2078942723864920533 , even if the demos are getting more impressive :https://x.com/BerntBornich/status/2075253825494237660 . For instance, some companies have managed to cram thousands of tactile sensors into a robotic fingertip, but none have managed to make these tiny sensors able to stand up to heavy use 2:#footnote-2 .
The control problem may be as challenging as the problem of physical construction. A competent robot must be able to find the right set of joint positions to grasp a complicated object; plan out the sequence of motions to fold a shirt, flip an omelette, or tighten a bolt in a constrained space; and handle squishy or floppy materials (which can require reacting instantly to a sudden shift).
Computer vision has made incredible strides over the last decade or two (and is responsible for kicking off the deep learning boom that led to LLMs). But making sense of complicated visual scenes – picking out an object from a crowded environment, understanding where it should be grasped, determining where it’s safe to put your feet and how to avoid knocking something over – is not a solved problem.
A general-purpose robot must be able to break down a task into individual steps, and relate those steps to its environment. How do you maneuver your arm to get a screwdriver into a piece of machinery? What’s the quickest way to clear a path to the spice bottle at the back of the shelf? In what order should you pick up the items on the living room floor?
True autonomy will require planning tasks of greater scale and complexity: cooking a meal, plumbing a bathroom, repairing an engine. Not to mention the need to re -plan in the face of surprises – a stuck bolt, a rotten piece of produce, a child darting into the kitchen.
When current AIs fail at a knowledge work task, it’s often because they weren’t provided with sufficient context. Robots will need context, too: where are supplies kept? How do you like your meals cooked? How much assistance does that nursing home resident need, and is that hitch in their stride normal, or a sign that they’re about to stumble?
Once they have context, robots will need to reason, plan, and exercise judgement and common sense. LLM-based systems like ChatGPT and Claude are making great strides in these areas, but the physical domain brings additional challenges 3:#footnote-3 . The success of LLMs has been greatly assisted by the massive pools of pre-existing data that were available for training – a substantial fraction of all books ever written, the web, and other massive pools of pre-existing data. It will be difficult to match this scale of breadth and depth of data for physical tasks. There’s no straightforward equivalent of “just Efficient learning, generalization, and adaptability / on-the-job learning seem like requirements.
AI agents mostly operate in isolation, and in static environments. We rarely put them in situations where things are changing out from under them, or ask them to coordinate. When we do, things often go haywire :https://theaidigest.org/village/goal/design-run-write-up-human-subjects-experiment . Isolation is easier to arrange in the virtual world, where private workspaces can be created at will, and nothing is too heavy to lift on your own. Robots will often need to cooperate with people, or with one another.
Today’s general-purpose robots often move much more slowly than human beings. Challenges include strength, control (higher speed means less time to plan and react), and safety (a fast-moving robot will whack you harder and is harder to dodge).
For some applications, slow and steady may be perfectly acceptable: I may not care if my household robot takes all night to tidy up and fold the laundry. But a slow-motion robot won’t be much use as a cook or nursing-home aide. It might get in the way at a warehouse. And it will have a harder time getting enough work done to pay for itself.
Some industrial robots are extremely strong. But humanoid robots – or other highly mobile, “general-purpose” robots – usually aren’t. It’s difficult to combine strength with manageable weight, a large number of joints, and a maneuverable frame. Powerful motors generate more heat and deplete batteries faster – two areas where robots already struggle (see below). And a strong, heavy robot poses greater safety challenges.
The jury is still out on the appropriate form factor for general-purpose robots, especially with regard to their lower half. Should they have wheels or legs? Two legs, four, or some other number? Wheels are cheaper, more stable, and more reliable; legs are better for stepping over obstacles and climbing stairs. A bipedal frame is more maneuverable, but also more likely to topple if something goes wrong. In any case, the question is: can the robot reliably get around its work environment?
Safety considerations for general-purpose robots are almost limitless. A glitchy or malfunctioning robot could bump into someone, topple onto them, drop something on them, spill something on them, break a glass, or start a fire.
Safety for LLM-based agents relies in part on review of discrete actions, such as attempts to send an email or delete a file. Robots move constantly, and it’s not so easy to single out a few specific motions as the potentially dangerous ones requiring review.
If a self-driving car finds itself in a situation it can’t handle or suffers a glitch, it can pull over or, in the worst case, just hit the brakes. A general-purpose robot that suddenly freezes might leave something on the stove, topple mid-step, or trip the person it was assisting.
And of course danger can be initiated by human action, such as a child darting in front of a robot. I’d much rather my kid be bumped into by a squishy person than a metal robot; and as things stand today, I’d much rather depend on human reflexes and adaptability to avoid tripping over the little rascal.
(The stronger, heavier, and more capable the robot, the greater the risks.)
Today’s bipedal robots can typically run for a few hours before recharging. I suspect this won’t be a limiting factor: if a workaround is needed, we’ll find one, whether that means swapping battery packs, in-floor charging grids, or a cable running to a nearby big-battery-on-wheels 4:#footnote-4 .
Then there’s the question of reliability. Robots have large numbers of moving parts, many of which are necessarily finicky, because they’re engineered to push the envelope on size, weight, and performance. As a result, current attempts at general-purpose humanoid robots experience frequent breakdowns :https://blog.robozaps.com/b/challenges-in-humanoid-robotics . (Contrast the human body, which is constantly recovering from wear and tear, and has substantial ability to self-repair and to compensate for minor breakdowns.)
Waymos struggle to handle edge cases. They’ve recently been observed driving into flooded roads :https://www.nytimes.com/2026/05/22/us/waymo-taxi-suspended-atlanta.html?unlocked_article_code=1.xFA.j1L_.UCPkIKcQKUkk&smid=url-share or over burning fireworks :https://www.foxnews.com/us/terrified-passengers-film-waymo-autonomous-vehicle-driving-into-live-fireworks-san-francisco . This is despite the fact that self-driving cars have been in development for well over two decades, and Waymo cars in particular have driven over 220 million miles :https://waymo.com/blog/shorts/safetydata-june26 – 250 times as many as a typical American drives in their lifetime.
(I’m willing to cut them some slack on the fireworks thing; 250th anniversaries don’t come along all that often. But I am confused at how Waymo engineering can be so robust as to yield an astonishingly good safety record :https://secondthoughts.ai/p/autonomous-vehicles-will-save-lives , and yet so slapdash as to happily drive into deep water.)
For all of the weird edge cases that arise while driving – the classic example being a duck being chased by a broom-wielding woman in a wheelchair :https://www.youtube.com/watch?v=weXDUc5Osto – robots operating in homes and businesses will encounter far more. They will be faced with a wider variety of tasks, using a wider variety of equipment (different tools; different robot bodies to grasp those tools), in a wider variety of environments. Achieving reliable operation outside of the controlled environment of a factory floor may be the hardest challenge of all.
Dexterity, coordination, visual understanding, planning, reacting, understanding, cooperating – and doing all of these quickly, safely, and reliably – will take a lot of computing power. Incorporating the necessary computing capacity into the robot itself will add cost, drain batteries, and contribute to overheating. Leaving the robot’s brains in the cloud will slow down reaction times and introduce new failure modes (Wi-Fi outage → dead robot).
LLMs were able to scale rapidly from the moment ChatGPT was launched, because existing hardware (GPUs) and manufacturing facilities (chip fabs) were easily adapted to support the new use case.
Large-scale supply chains for advanced robots don’t exist yet. Even once we have workable designs, it may take years before they can be manufactured, deployed, and maintained at scale. This sort of thing doesn’t happen overnight; it took 14 years for Tesla to advance from first commercial sales to their first million-car year 5:#footnote-5 . One analysis :https://epoch.ai/publications/how-fast-could-robot-production-scale-up found that once the starting gun is fired (advanced humanoid robots become economically valuable), it might take several years to scale to producing low-millions of robots per year. ChatGPT, by contrast, reached 100 million users 6:#footnote-6 within two months of launch.
Political, regulatory, organizational, and cultural barriers : concerns over job displacement and safety may limit where and how robots can be used. Many regulations were not written with robots in mind – does a robot count toward minimum staffing requirements? Will businesses leap to adopt robots? Will they have concerns over reliability, security, liability, and maintenance? Will customers want to be served by a robot?
Cybersecurity, surveillance, misuse : an advanced robot could be a criminal’s dream – a dependable henchman that can’t betray its owner 7:#footnote-7 . Preventing this might require continuous monitoring of all robots, which raises all sorts of concerns. And the cybersecurity on robots will need to be airtight.
Cost : once the other hurdles are addressed, I suspect this won’t be much of a limiting factor. A tireless worker at the price of a new car would be a bargain, and humanoid robots will be much smaller and lighter than a car, with fewer moving parts 8:#footnote-8 . It could be that the components, manufacturing techniques, and training processes required for capable robots will make early models much more expensive than a car. But even expensive robots would likely find early use cases – for instance, doing hazardous work.
Many hurdles will need to be cleared before robots become capable, reliable, practical workers outside of carefully controlled factory environments. The list I’ve presented is surely incomplete; and I’ve only briefly glossed over the cognitive side – understanding, planning, acting, and reacting.
Demo videos provide a glimpse into what robots can accomplish under ideal circumstances. They can also serve to distract us from the remaining limitations. For knowledge work, there’s a consistent gap in AI performance between benchmarks and real-world work. In the physical world, I suspect the demo / reality gap will be even larger. Playing around with an LLM to see what it can do has been accessible, cheap, and (mostly!) safe. To assess the capabilities of robots, we’ll be much more reliant on controlled demos and manufacturer’s claims. It will be harder to map the jagged boundary of their capabilities.
As I was putting the final touches on this post, the excellent Understanding AI blog posted Why humanoid robots won’t catch up to human workers any time soon:https://www.understandingai.org/p/why-humanoid-robots-wont-catch-up . I haven’t read it yet but I’m sure it’s worth a look.
For a good broad review of robot capabilities, see Epoch’s report from February 2026:https://epoch.ai/publications/where-autonomy-works-evaluating-robot-capabilities-in-2026 .
Share :https://secondthoughts.ai/p/14-reasons-robotics-is-hard?utm_source=substack&utm_medium=email&utm_content=share&action=share
Thanks to Abi Olvera, Avi Parrack, and Taren Stinebrickner-Kauffman.
There are four-legged robots whose “feet” are wheels, which can alternate between rolling and walking, and which can climb stairs. But so far there is no sign of this design going mainstream.
Workarounds may be found. For instance, some robot designs use cameras to partially compensate for tactile sensitivity, with extra “eyes” in the wrists or elsewhere; a robot could even make use of detached cameras via Wi-Fi. Robots could also incorporate non-human senses, such as ultrasound.
See this explanation of the importance of tacit knowledge in bio lab work :https://secondthoughts.ai/p/tacit-knowledge-the-missing-factor for an example of how much robots will need to learn.
One recent demonstration :https://interestingengineering.com/ai-robotics/figure-03-humanoid-robot-200-hour-shift showed a robot sorting packages for 200 hours straight, processing 250,000 packages… using three different robots, taking turns while the others recharged. It was meant as a demonstration of reliability as well as endurance.
Assuming the criminal managed to avoid a paper trail connecting them to the robot.
At least, compared to an internal combustion car.
1. One other element I'd add about our hands is that they're capable of self-repair from accumulated damage and wear-and-tear. Robot hands can not do that - they might struggle to survive with acceptable functionality at a cost-effective rate, especially if they're extremely complex to try and make them match human capabilities.
2. I didn't know about the overheating thing. That's brutal - if it generates a lot of noise, people aren't going to have those things running around in the house anymore than you'd tolerate a vacuum constantly being on at all times within 10 feet of you.
Great post. Also the sheer efficiency of human energetics, metabolism and heat dissapation can not go understated. The fact humans can get up eat a bagel and coffee and have all the energy they need to operate for several hours is nothing short of amazing. Further humans also use that miniscule energy for ongoing self repair. Evolution has been optimizing in this domain for millions of years. Good luck replicating that. That is not to say that humans will gain from a synergistic relationship with mechanical tools/ie robots but I think we're a long way from replicating all things human biology.