Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

These systems have a physical footprint and it is built from concrete and electricity

The abstraction of the cloud hides buildings, transformers, cooling plant and a supply chain for chips, and every one of those is a constraint that software cannot argue with.

By Imran Sheikh3 min read

A scientist controls a robotic arm conducting research on a person lying down.
Photograph by Pavel Danilyuk via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Nothing about this is immaterial

The vocabulary of the industry is relentlessly abstract. Models live in the cloud, capacity is elastic, resources are provisioned. Underneath sit specific buildings in specific places, drawing power from specific grids through equipment that took years to order and install.

Those buildings are unusual as industrial facilities go. A large computing site consumes power continuously at a rate comparable to a small town’s, converts essentially all of it into heat, and must then remove that heat reliably or the machines inside shut down. The engineering problem is less about computation than about thermodynamics.

This is not a footnote to the technology. It is one of the two or three real constraints on how far and how fast any of it can go, and it responds to none of the things that usually accelerate software.

The grid is slower than the software industry

Adding computing capacity is quick by industrial standards. Adding the electricity to run it is not. Transmission lines, substations and generating capacity are planned over years to decades, subject to permitting, local consent and physical construction, and none of that compresses because demand arrived sooner than expected.

The result is that siting decisions are increasingly made on the basis of where power is available rather than where the users are. That has consequences for the communities involved — on their grids, their rates, sometimes their water — and those consequences are being argued out locally in many countries at once, with different conclusions.

It also introduces a lag that markets handle badly. Demand for computation can multiply in a year. The infrastructure serving it cannot, and the mismatch shows up as shortages, price movements and long queues for capacity rather than as a shortage of ideas.

Two different energy stories that are constantly confused

Training a large model is a concentrated burst: a great deal of energy over a bounded period, producing an artefact that is then finished. Running a model is a continuous draw that scales with usage and never stops. Both are real and they behave nothing alike.

Public discussion tends to focus on training, because a single large number is memorable, while the cumulative operational draw of serving many users over years is the larger figure for any widely used system. Which dominates depends on how heavily a model is used and for how long, which is exactly the sort of thing that varies enormously and cannot be summarised in one statistic.

Efficiency has improved substantially — better hardware, compression, cheaper serving techniques — and total consumption has risen anyway, because falling unit costs invite more use. That pattern is old and well documented in other industries, and there is no particular reason to expect this one to be exempt.

The chips are a supply chain, not a commodity

The processors doing this work are made in a small number of facilities using equipment produced by an even smaller number of suppliers. Building a new facility takes years and enormous capital, and the specialised memory these chips depend on is similarly concentrated.

Concentration of that kind makes the whole enterprise sensitive to things that have nothing to do with computer science: export controls, trade policy, natural disaster, regional stability. That is a genuine strategic fact and it is one reason governments have become directly involved in a field that until recently they mostly funded from a distance.

It is worth being careful here. Predictions of imminent chip shortages and imminent gluts have both been made confidently and both have been wrong repeatedly. The structural point — few suppliers, long lead times — is solid. Any specific forecast built on it is not.

What this constraint does to the shape of the field

If capability tracks computation, and computation is limited by physical infrastructure, then progress becomes partly a construction question. That reframes several debates. It suggests efficiency research has strategic value beyond cost saving, and it explains why smaller models that run on ordinary hardware attract effort out of proportion to their headline performance.

There is a counter-case worth stating. Efficiency gains have repeatedly arrived from algorithmic work rather than from more hardware, and the relationship between computation and capability has never been as smooth as the tidiest charts imply. Assuming the constraint binds forever is as speculative as assuming it will not.

What is not speculative is that the industry now has a physical bottleneck it did not have when it ran on ordinary servers, and that bottleneck is made of things that take years to build.

Common questions

Is a single query energy-intensive?

One request is small in absolute terms, comparable to other ordinary online activity, and the aggregate matters far more than the individual. Framing this as a personal consumption choice tends to obscure that the significant decisions are about infrastructure, siting and grid planning rather than about individual use.

Why do data centres need water?

Many cooling designs use evaporation, which is efficient in energy terms and consumes water in the process. Alternative designs use more electricity and less water. It is a genuine trade rather than a failure, and which side of it matters more depends heavily on where the facility sits.

Will more efficient models reduce total energy use?

Historically, cheaper computation has increased total consumption rather than reduced it, because lower cost per unit expands demand. Efficiency is still worth pursuing on every other ground, but treating it as a guaranteed route to lower aggregate energy use is optimistic and not well supported by the pattern in comparable industries.

In The Worldenergydata centreshardwareinfrastructure
Imran Sheikh
Editor, AI Worth Knowing

Imran has written about how it works, in the world, limits & risks for most of the last decade and thinks most subjects are more interesting once you know how they work.