Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

A great many published results here cannot be run again by anybody else

Reproducing a result requires the data, the code, the settings and a system that has not changed underneath you, and in this field at least one of those is usually unavailable.

By Daniel Okonkwo4 min read

Macro shot of a laptop displaying coding and data analysis in progress. Ideal for tech themes.
Photograph by Lewis Kang'ethe Ngugi via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

What reproducing a result would actually require

To repeat an experiment here you need the training data in the state it was used, the code including the parts nobody thought worth mentioning, every setting including the ones chosen without deliberation, the random seeds, the versions of every library involved, and enough hardware. Missing any one of these turns repetition into approximation.

This is a harsher requirement than in most empirical fields, and for an odd reason: the experiment is entirely computational, so in principle it should be perfectly repeatable. That expectation makes the actual situation more surprising than it would be for an experiment involving physical materials and living organisms.

The gap is filled by defaults, guesses and correspondence with the authors, all of which introduce variation. When a repeated experiment produces a different number, it is usually impossible to determine whether the original was wrong or the repetition differed in some unrecorded way.

Several of the ingredients are withheld by design

Training data for the largest models is frequently undisclosed, for reasons that include competitive advantage and unresolved questions about what was in it. Model weights may be unavailable. Training procedures are described at a level of detail that conveys the approach without permitting reconstruction.

This is a change from earlier practice in the field, which had been unusually open about code and data by the standards of most disciplines. The reasons offered for the shift are commercial and, in some cases, related to concerns about misuse, and there is genuine disagreement about how much weight the second reason carries.

Whatever the motivation, the effect on verification is the same. A claim that cannot be checked is a claim that has to be believed or set aside, and a field accumulating those is in a different epistemic position from one accumulating checkable results.

Evaluating a moving target is its own problem

A growing share of published work evaluates systems accessed as services rather than as files. Those systems are updated, sometimes without announcement, and a measurement made in one month may not be repeatable in the next. The object of study is not stable, which is an unusual predicament for an empirical discipline.

This affects comparisons most severely. Two results measured months apart against the same named system may have been measured against materially different systems, and nothing in either paper reveals that. Comparisons across such results are common and are weaker than they look.

Some researchers have responded by recording exact dates and version identifiers where those exist, which helps a reader understand what was measured and does not restore the ability to measure it again.

The cost of confirmation is high and nobody is paid for it

Even where everything needed is published, running a large training experiment costs enough that repeating one is a serious decision. Academic groups generally cannot, and organisations that can have limited incentive to spend that money confirming somebody else’s result rather than producing their own.

Publication incentives make this worse in the ordinary way. Novelty is rewarded, confirmation is not, and a negative reproduction attempt is difficult to publish and awkward to be associated with. This is not unique to this field and it interacts badly with the unusually high cost of an attempt.

The consequence is that a substantial body of results has been accepted on the strength of a single run by the party with the strongest interest in it succeeding. That’s a reasonable working assumption in most cases and it is a weaker foundation than the volume of publication suggests.

Variance between identical runs is larger than reported

A separate problem sits underneath all of this. Training is stochastic, and two runs differing only in random seed can produce measurably different results on the same evaluation. Where reported differences between methods are small, the difference between two runs of the same method may be of comparable size.

Reporting a mean and a spread across several runs addresses this and costs several times as much computation, which is why it is done less often than it should be. Single-number comparisons remain common, and the reader is generally given no information about how stable that number is.

This makes a good deal of reported progress harder to assess than it appears. Some of it is certainly real and large. Some of it is within the range that repeated runs would produce anyway, and the two are not distinguishable from the outside.

What has improved, and what a reader should do

The situation is better than it was. Conferences have introduced reproducibility checklists, code release has become expected for academic work, standardised evaluation harnesses have reduced the variation caused by everybody implementing tests differently, and shared model repositories have made a great deal genuinely available.

The improvements are concentrated where the work is cheap enough to repeat, which means the frontier remains the least verifiable part of the field. That is an uncomfortable inversion, since it is also the part attracting the most attention and the most policy interest.

For a reader the practical response is to weight results by what could be checked. A finding independently reproduced by parties with no stake in it deserves considerably more confidence than a headline number from a single unreplicated run, however impressive the number.

Common questions

Is this worse here than in other sciences?

It is different rather than uniformly worse. Other fields struggle with biological variability and small samples, which this field largely avoids. What it has instead is deliberate withholding of the necessary ingredients and a cost of repetition that excludes most potential replicators.

Does releasing model weights make a result reproducible?

It makes the result usable and checkable in the sense that others can evaluate the same artefact. It does not make the training reproducible, which requires the data and the procedure, and those are the parts most often withheld.

Why does a random seed matter so much?

It determines the starting parameters and the order examples are seen, and those propagate through a long training process. The result is not arbitrary, but it does land in a slightly different place each time, and on close comparisons that difference can exceed the effect being measured.

Limits & Risksreproducibilityresearchvalidation
Daniel Okonkwo
Contributing editor, AI Worth Knowing

Daniel covers how it works, in the world, limits & risks and the questions readers actually send in and prefers a plain explanation to a clever one.