Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

An image generator works by learning to remove noise it was taught to add

The dominant method for producing pictures is built on a destruction process run backwards, and the arrangement explains both the quality of the results and several of the characteristic failures.

By Naina Sethi4 min read

Black woman engineer with crossed arms standing in a server room, smiling confidently.
Photograph by Christina Morillo via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Generating a picture is a harder shape of problem than recognising one

Recognition takes a large input and produces a small output: a photograph goes in, a label comes out. Generation runs the other way, and that reversal is not a symmetry. The system has to invent millions of values that are consistent with each other, and almost every arrangement of those values is not a picture of anything.

Early attempts asked a network to produce the whole image in one shot and the results were smeared and characterless, because a model uncertain between two possibilities tends to output their average. An average of two faces isn’t a face. Producing something specific requires committing, and committing all at once turned out to be very difficult.

The method that eventually worked avoids the commitment by breaking it into many small steps. Each step is easy. The difficulty is spread across hundreds of them, which is why generating an image costs so much more computation than classifying one.

Training means learning to undo a controlled destruction

The training procedure is almost perverse in its simplicity. Take a real image, add a small amount of random noise, and ask the network to predict what noise was added. Repeat with more noise, and more again, until the image is indistinguishable from static. At every level of damage the network learns what the damage looked like.

Nothing in this requires a human to label anything, which matters enormously. The correct answer is known because the trainer generated the noise themselves, so the supervision is free and the quantity of training material is limited only by how many images can be gathered. That is the same self-supervised trick that made language modelling practical.

What the network is really learning across all those noise levels is the shape of the space that real images occupy — what tends to be next to what, at every scale from a whole scene down to the texture of a surface. It is a statistical description of picture-ness, encoded as a denoising habit.

Generation is that process run in reverse

To produce a new image, start with pure noise and ask the network what noise it sees. Subtract a little of the answer. Ask again. Each pass makes the field of static slightly more like the sort of thing the network was trained on, and after many passes a coherent picture has emerged from something that contained no picture at all.

This is why the output is different every time even with an identical request: the starting static is different, and everything downstream follows from it. It is also why the number of steps is a dial. Fewer steps is faster and rougher, more steps is slower and generally cleaner, with diminishing returns that arrive sooner than you would expect.

Most production systems do this in a compressed representation rather than at full resolution, using a separate component to encode and decode. The denoising then happens in a much smaller space, which is a large part of why the technique became affordable enough to be offered as a product.

Text steers the process without controlling it

A description is turned into numbers by a separate model trained to place images and captions near each other in a shared space. Those numbers are fed into the denoising network at every step, tilting each prediction towards the region of picture-space that matches the description. The word for this is conditioning, and it is a nudge rather than an instruction.

Because the steering is applied continuously and statistically, the relationship between what was asked and what appears is loose. Attributes migrate between objects. Things that were mentioned go missing. The system is not parsing a specification and checking its work against it; it is being pulled in a direction while it does something else.

Strengthening the pull is possible and it trades off against variety and naturalness. Push too hard towards the description and images become garish and stereotyped, which is a real dial in these systems and one of the more visible ways their output can be recognised.

What the mechanism explains, and what it does not excuse

The recurring failures make sense from this description. Counting is unreliable because nothing in the process maintains a tally. Text within images has been poor because letterforms demand exactness in a system built on plausibility. Anatomical errors at the joins occur because local coherence is easier to learn than global structure.

Some of these have improved substantially, which is worth saying, and the improvements came from better architectures and better training data rather than from any change to the underlying idea. Predicting which remaining failures will fall next is guesswork, and confident forecasts either way should be treated as such.

The mechanism also does not settle the arguments that surround these systems. Whether training on gathered images is permissible, and what the resulting output owes to the people whose work is in the training set, are questions about provenance and law. Knowing how the denoiser works does not answer them.

Common questions

Does the generator store the images it was trained on?

Not as images. What it holds is a set of parameters describing the statistical structure of a very large collection. Under some conditions, particularly when an image appeared many times in training, output closely resembling a specific training example has been demonstrated, so the absence of stored pictures is not the same as an absence of memorisation.

Why do two runs of the same description give different pictures?

Because the process starts from randomly generated static and every step follows from it. Fixing that starting point makes the output repeatable, which is how these systems offer reproducible results when they do.

Is this the same technique used for video and audio?

The same core idea has been adapted to both, with additional machinery for keeping successive frames or samples consistent with each other. Consistency over time is the hard part and it is where those systems still visibly struggle.

How It Worksimage generationdiffusionarchitecture
Naina Sethi
Features writer, AI Worth Knowing

Naina joined to cover how it works, in the world, limits & risks and stayed for the awkward questions and is unreasonably interested in the detail nobody else checks.