Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Physical tasks lagged because the world does not arrive pre-labelled

Machines became fluent with text and images while remaining clumsy at picking objects up, and the reasons have to do with data, feedback and the cost of a mistake rather than with intelligence.

By Imran Sheikh3 min read

A futuristic white robot toy performing on a sleek dark surface.
Photograph by Pavel Danilyuk via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The easy things turned out to be the hard things

Researchers noticed decades ago that tasks demanding years of human training — playing strong chess, solving formal proofs, diagnosing from a checklist — fell to machines earlier than tasks a toddler performs without thought. Walking over uneven ground. Picking up a soft object without crushing it. Recognising that a jar lid is stuck.

The usual explanation is evolutionary. Perception and movement were refined over a vast span and are consequently deep, highly optimised and entirely unavailable to introspection, whereas abstract reasoning is recent, shallow and therefore easier to write down. Whether that account is the whole story is arguable, but the pattern it describes is solid.

It has held up well. Text and image work advanced enormously while general-purpose manipulation stayed difficult, and the gap is not closing at anything like the same rate.

There is no vast archive of grasping

Language models were trained on an accumulated written record that already existed. No comparable record exists for physical interaction. Every example of a robot picking something up has to be produced by a robot picking something up, in real time, on hardware that wears out and occasionally breaks what it is holding.

That makes data collection slow and expensive in a way that has no analogue in text. An hour of robot operation yields an hour of data. Running many machines in parallel helps and multiplies the capital cost, and the resulting data is specific to those machines, those grippers and that lighting.

Efforts to pool data across laboratories and across robot designs have made real progress, and they run into the problem that two robots with different bodies do not straightforwardly share experience. A movement that works for one arm is not a movement for another.

Simulation helps and then leaves a gap of its own

The obvious workaround is to practise in a physics simulator, where a million attempts cost electricity rather than hardware. This works well enough to be standard practice, particularly for locomotion, where the relevant physics is reasonably well captured.

The trouble is that simulators are approximations. Friction, deformable materials, contact between surfaces and the behaviour of cloth or liquid are all modelled imperfectly, so a policy trained to exploit the simulator may be exploiting an artefact rather than learning the task. The mismatch has a name and it is a permanent feature of the approach.

The main mitigation is to randomise the simulation aggressively — vary friction, mass, lighting, delay — so that reality looks like one more variation among many. It transfers considerably better than training in a single clean simulation. It is still not the same as having been there.

Being wrong costs something different here

A language model that produces a bad paragraph has wasted somebody’s time. A robot arm that misjudges force breaks the object, the gripper, or a person’s hand. That asymmetry justifies the caution around deployment and it also slows learning, because the cheapest way to learn is by failing repeatedly.

Industrial robotics solved this historically by removing uncertainty rather than handling it: fixed positions, identical parts, cages keeping people away. Those systems are extremely reliable and they are not learning anything. The moment the environment stops being controlled, the approach stops applying.

Systems intended to work alongside people carry certification requirements around force limits, stopping distance and predictable behaviour. A control policy produced by training is hard to certify against those requirements, because the usual method is to reason about what the system will do, and a learned policy does not offer that kind of account.

What has shifted recently, and what to discount

Learned control has genuinely improved, particularly for legged movement, and using language models as a high-level planner that issues instructions to lower-level skills has produced demonstrations that would have looked implausible not long ago. Shared datasets spanning many robot types are a real development rather than a press release.

The counter-case deserves equal weight. Demonstrations are filmed in arranged environments, often with resets between attempts and a success rate that isn’t always stated. The distance between a compelling video and a machine that works unattended in a warehouse for a year is enormous, and it has been underestimated repeatedly for fifty years.

Anyone forecasting when general-purpose physical competence arrives is speculating, including the people who work on it. What is not speculative is the shape of the difficulty: data is scarce, mistakes are expensive, and the physical world does not hold still.

Common questions

Why can a robot do a factory job but not a household one?

Because a factory removes variation deliberately. Parts arrive in known positions and orientations, lighting is constant, and nothing unexpected walks past. A home contains unfamiliar objects in unpredictable arrangements, and handling that requires generalisation rather than repetition of a programmed path.

Does a language model help a robot at all?

It can supply the high-level structure: breaking an instruction into steps and choosing which learned skill to invoke. It contributes nothing to the physical control itself, which is a separate problem of sensing and force. The division of labour is useful and it does not close the harder gap.

Is training entirely in simulation ever enough?

For some locomotion tasks it has worked with only modest adjustment afterwards, because the physics involved is captured tolerably well. For manipulation involving contact, soft materials or fluids, simulation remains a starting point rather than a substitute, and real-world data is still required to finish the job.

In The Worldroboticsembodimentsimulationcontrol
Imran Sheikh
Editor, AI Worth Knowing

Imran has written about how it works, in the world, limits & risks for most of the last decade and thinks most subjects are more interesting once you know how they work.