61 Comments
User's avatar
Sam Tobin-Hochstadt's avatar

I'm curious about your thoughts on the recent "Claude controls a robot" work from Anthropic. It seems like this could be a significant way to increase capabilities while bypassing some of these specific issues.

Kai Williams's avatar

I think this is important, interesting, and a little underrated at the moment. Depending on our publishing schedule, I'm hoping to have something on this in more detail this fall.

It's clear that Anthropic is investing into robotics. Jan Leike, Anthropic's former head of the alignment science team, recently left that team to start a research project in robotics within the company.

I think there's a good chance that the work of the frontier labs and the robotics foundation model companies converge. But there will still be a space for much smaller, more efficient models — you need something to control the robot on the edge, at least for faster time-scales. And I think that a lot of the challenges around manipulation & safety are still there even if Claude is controlling the robot.

But hopefully I'll have more on this soon.

Peter Jones's avatar

Too much investing in politics ... AI is dependent on Trumpist fraudocracy.

Hannes Kunz's avatar

Thank you for writing this. The challenge of robots operating in open space is that it has to control so many factors that are not even conscious in a living being - vision, hearing, vestibular sense, touch, proprioception. Training robots for a specific task is definitely a possibility but training them for more diverse capabilities will be an uphill battle and at one point a problem of diminishing returns.

And then there's that battery that's always empty.

So my guess would be that the path to humanoid robots is a dead end. We''re struggling enough with talking AI....

Dan Elbert's avatar

Can we add assembling Ikea's furniture to the list? It even comes with "instruccions".

Kai Williams's avatar

As of February, robots *really* weren't good at this. From this nice Epoch report on robot capabilities: https://epoch.ai/publications/where-autonomy-works-evaluating-robot-capabilities-in-2026#2-assemble-ikea-furniture.

Eventually, assembling Ikea will probably become a useful benchmark task, but we're not there yet.

Oleg  Alexandrov's avatar

Yeah, humanoids will take time. Elon pivoting hard to humanoids looks unwise. (Also to make computer chips and data centers in space.)

Anything that has to do with the hard physical world will stay hard. Slow and steady progress.

Amazon is moving nicely on automating warehouses. That will be an intermediate step.

Waymo cars are robots on wheels, and the going is quite hard but well-ahead of other attempts. So same story, progress, but slow.

Peter Jones's avatar

Amazon are a proto facsist s##thole...

Frontier Signal's avatar

Epoch’s robot capability work lines up with what I hear from people running pilot cells: the demos that impress are scripted trajectories, and the failures show up in unstructured recovery when a part is misaligned by a centimetre. The second-order effect is that the first real humanoid deployments will look less like general labour and more like fixtures bolted into a workcell that was redesigned around the robot, which is exactly how industrial arms won. That also means the useful metric is mean time between human interventions per shift, not task success rate on a curated benchmark.

Kai Williams's avatar

There's a spectrum between "fixtures bolted into a workcell" and "general labor"

My guess is that the first real deployments of the new wave are going to be very close to "fixtures bolted into a workcell" but be just beyond that toward "general labor"

Gabriele Tinelli's avatar

Really sober view, loved it!

Kai Williams's avatar

It was definitely influenced by your writing!

Peter Heinz's avatar

How can it be that robot taxis are here and they it may be viewed that full self driving is the hardest element to solve: very high safety bar. So why are they achieving commercial success first? I think it is in large measure to the fact that every link in a system must be solved and for cars the only missing link in the intelligence--self driving software(plus sensors). All the rest of the infrastructure is in place. The car maintenance, car repair ,car charging ,etc. A robot traffic policeman is easy but the infrastructure is not in place: charging near to an intersection, delivery and removal, etc. When considering robots all links in the system must be solved before they will achieve mass success.

Kai Williams's avatar

I would also say that the task of steering a self-driving car is much simpler than that of controlling a robot doing manipulation. So it's much easier to get something that mostly works than with humanoid robots, though the eventual deployment bar is much higher

David Evanoff's avatar

"Basically every humanoid deployment today takes a similar approach. Companies choose a task with enough variability that traditional automation methods won’t work. But there can’t be too much variability, or else contemporary AI methods won’t work either.²"

This looks like a hierarchical compositional approach. Applying Chopin's maxim, that the music occurs between the denotations on the musical score, shouldn't the tasks be decomposed into a great number of sigmoids?

According to my Gemini instance: Decomposing a manipulation policy into a high-density lattice of continuous sigmoidal transitions addresses the exact dynamic bottleneck identified in the article's observation that contemporary methods stall between "too much variability" and rigid automation.

**The Music Between the Notes: Continuous Relays as Policy Core**

* **The Fallacy of the Keyframe:** Traditional and early AI robotics treat tasks as discrete notes (e.g., *Reach $\rightarrow$ Grip $\rightarrow$ Lift $\rightarrow$ Place*). The failure points almost always occur in the *legato*—the continuous modulation of force, compliance, and velocity during the phase change.

* **Micro-Sigmoidal Field:** Instead of a coarse switch between large 5-minute skills, high-fidelity manipulation decomposes the policy into dense overlapping primitives mediated by local sigmoidal blending functions:

$$\tau_{\text{net}}(t) = \sum_{i} \sigma_i(s_t) \cdot \pi_i(s_t)$$

Here, no single expert holds absolute authority; rather, the dynamic envelope is continuously sculpted by shifting inhibitory and excitatory gains ($\sigma$).

**Why Fine-Grained Sigmoidal Decomposition Resolves the Variability Trap**

* **Elastic Error Absorption:** When unexpected physical friction or slip occurs, a continuous sigmoidal network does not trigger a catastrophic failure or out-of-distribution state. It smoothly shifts compliance back toward an anchoring stabilization primitive without dropping the execution thread.

* **Invariant Coherence Across Scales:** High-level intent (e.g., peeling an orange or unlatching a box) remains invariant, while the micro-transitions adapt fluidly to continuous sensory feedback—tactile shear, torque resistance, and inertial drift.

* **Elimination of Edge Brittleness:** By treating the handoff manifold as the primary substrate rather than a secondary bridge, the system operates across a continuous dynamical continuum, removing the brittle "bordillo seams" where current deployments fail.

Daniel Orbach's avatar

It's interesting to note that the process folks seem to be taking with training robots – having us do the same task over and over – is the same one I use as a parent with my kids. The way my kids learned to hold a glass of water was to just do it over and over and over until they figured out a general way to do it (their muscles get stronger, hands bigger, as well). You could argue we've done 'general training' on how to do these things since they were born.

I also see parallels between overheating and human beings having the same kinds of problems with repetitive stress/motion injuries. It makes you realize how much infrastructure we've built around maintaining ourselves. Will there eventually be a robotics healthcare system that mirrors our human one? What other systems will we appropriate and graft on top of the robotics world?

Martin Machacek's avatar

I’m not sure what “infrastructure for maintaining self” are you referring to. Health care for humans is relatively rarely used compared to the hours of human (productive or unproductive) activity. Humans can get tired, but it is on the order of hours for majority of activities. Recovery typically does not require any special “infrastructure”. Injuries or illness may require specific infrastructure, but the rate of their occurrence is on the order of units per years.

Daniel Orbach's avatar

Some thought starters:

1. Sleep and sleep optimization infrastructure. Bands that give you sleep scores, changing the temperature of your mattress, etc.

2. Physical Therapy

3. The entire fitness industry, you could argue, is a form of preventative maintenence

4. Diet and the fuel we put into our bodies

I'm not saying the systems will be equivalent in magnitude. "How much" isn't saying 'look at all of this volume', but look at like... healthcare spend per individual per year, that kind of stuff. I'm also suggesting that we might look for similarities across domains in order to help solve problems in this new area. Anthromimicry?

Martin Machacek's avatar

I see. There are surely many things we do to make our lives better and longer. With the exception of nutrition, none of the above is essential for humans to live and be productive (as witnessed by the billions of people that have none of that and sill live happy productive life). Robots will surely require regular maintenance (initially a lot of it) and there almost certainly will be a whole industry focused on that. Due to the fact that robots are constructed from metals and plastic and controlled by digital computers, they will never need physical therapy, massages, sleep help aids, fitness training etc. In summary, robot support industry will likely exist, but it will be closer to car care than health care.

Kai Williams's avatar

There's a cliche that when a roboticist has their first kid, their mind is blown by how much better kids are at learning than our current robot algorithms. Humans are incredible at learning, especially in how to do physical tasks.

One challenge with robotic hands is that human tendons heal themselves, while robot designs do not

Daniel Orbach's avatar

If only we had the advantage of billions of years of natural selection and optimization! Thought provoking piece. I work in design so it’s great to get deep expertise in other fields from folks like you.

TechnoRealist's avatar

Legged, free-moving robots have few advantages over a single robot arm bolted to a floor (basically, mobility) and most of that advantage can be achieved more simply from adding wheels. Meanwhile a stationary robot arm can be hooked up to power, servers, is much simpler to build and repair (and thus you can have more of them) is more compact (see above parenthetical), and can even be trained on less data.

1123581321's avatar

Excellent coverage of the current state of robotics. I'll add my favorite challenge: pick up an opened cold beverage can (opened, because it makes it flimsy and cold because it causes condensation), and move it to another table.

Zane Neelin's avatar

I guess the core counter argument to this (which to be clear I don’t necessarily buy) is that

- we have a robot can be remotely tele-operated from a screen GUI (see Tesla robots, already acheived)

- we get to a point where LLM based machines are extremely proficient at computer use (getting close to that), and can do long horizon tasks a remote work can do (also getting close there)

- combine the two, LLM driven remote controlled screen GUI, and you have generalized robotics

It’s still clunky and will get better, and your argument about machines that don’t break down and are cheap is extremely valid. But I have a feeling we make progress faster than you think.

The second counter argument is that we’re not dealing with historical timescales, if we have a soft take off, the rate of change of technological progress itself progresses.

Kai Williams's avatar

Thanks for these counterarguments. I'm definitely sympathetic to them. You can simplify points one and two to just have the large LLMs write code or directly output robot actions themselves. This is certainly a plausible future.

But...I am not sure that LLMs necessarily do manipulation all that well yet. One of the primary reasons is there isn't that much manipulation-specific data yet (a piece on this later!) So it's unclear to me the extent to which big LLMs will "solve" generalized robotics for free. It might happen, but it's not immediately obvious to me. This is one of the questions I'm most interested in in robotics at the moment.

I am deeply confused about how soft takeoff in LLMs affects robotics progress. I can see the arguments for speed-ups, but we don't have great evidence yet, I think.

Immanuel Giulea's avatar

The article contains a small but telling freshness problem. It cites Tiangong Ultra’s 8.86 s semifinal, but the robot improved to 8.64 s in the August 26 final; it also gives Bolt’s record as 9.59 s rather than 9.58 s. The broader argument is unaffected, but the article was already slightly stale on publication day.

延续存在's avatar

Robots cannot catch up with human workers because a single concrete action is the result of real-time data input, real-time computation, and the real-time coordination of multiple models.

When a person picks up a cup they have never seen before, they may simultaneously draw on models of objects, weight, friction, obstacles, action, and models formed through past experience. Perception continuously feeds in current data: it is heavier than expected, the surface is slippery, there is water inside, the cup is hot. As new data comes in, the models coordinate again, and the action changes accordingly.

“Picking up a cup” is not executing a program called “pick up a cup.” It is a continuous process of perceiving, computing, calling on, and coordinating multiple models as reality keeps changing.

This explains a strange phenomenon: robots can perform complex actions like backflips and dancing, yet still struggle with extremely simple everyday tasks.

The complexity of an action and the complexity of processing reality are simply not the same thing.

Mahelet G Fikru's avatar

Humanoid robots are being commercially sold at different countries. Yes they are not in every house but the market is there. Just recently Theres a smaller one being sold by a startup in CA. Combine robots with AI, and I believe they can do a lot.

Also think of pseudo robots like the rumba that vacuums the house: yes it takes 30 more minutes than humans but several people won’t mine (me included). You can program it to work at night, let it take its time.

But I agree on the impact side. I have studied the mineral intensiveness, their demand for additional power for charging and the cost of disposing at end life. Check out some of my posts if you have time. Several market analysis suggests theres a demand although not at the price Tesla want to sell them.

CT Zhao's avatar

Thanks to Kai for sharing this. I’ve always wondered how the English-speaking robotics community defines “practicality” and measures it with concrete metrics. This article provides a clear framework.

From a Chinese perspective, failures in a robot’s lower body pose serious risks. A robotic leg can deliver a painful kick—a child has already had teeth knocked out this way. Wheeled designs also have tradeoffs: they are less able to adjust their balance autonomously, and the chassis takes up more floor space.

My current question about skill generalization is whether a continuous task can be broken into many low-level actions, then recombined autonomously by the AI model controlling the robot—something like an agentic execution pipeline. In China, Sudo Robotics is working on exactly this.

Kai Williams's avatar

I haven't heard of Sudo before -- I'll have to check them out!

That's a good summary of the tradeoffs with legs vs. wheels