How It Works
A model is a file of numbers, and the file is not the whole system
What gets copied, released or stolen is a set of stored values plus the code that knows how to read them, and separating the artefact from the service it sits inside clears up several arguments at once.
By Manish Trivedi4 min read

The artefact is duller than the word suggests
The word model does duty for at least three things: a research idea, a running service, and a specific object sitting on a disk. The third is the concrete one, and separating it out clears up a surprising number of arguments. It is a file, or a small set of files, holding an enormous quantity of numbers.
Those numbers are the parameters, fixed at the moment training stopped. Each is stored in a compact numerical format, so the size of the file is essentially the parameter count multiplied by the bytes used for each value. A large model is a substantial download. A small one fits on a phone with room to spare.
Open the file in an editor and you would see nothing legible. There is no list of facts, no section marked chemistry, no rules anybody wrote down. Whatever the system appears to know is smeared across those blocks in a form that resists reading, which is a different and considerably harder problem.
Numbers are inert without a description of their shape
A block of numbers cannot be used until something knows how many blocks there are, how large each one is, and in what order they should be applied to an input. That description is the architecture, and it lives in code and in a small configuration file rather than inside the weights themselves.
This is why loading a model requires matching software. Given the weights alone, an engineer can often reconstruct the arrangement if the design was published, because the shapes of the blocks constrain what the architecture must have been. If the design is novel and undocumented, the file is close to inert.
A third component travels alongside: the tokeniser, meaning the vocabulary of fragments and the rules for cutting text into them. Pair a model with the wrong one and the output is confident nonsense. It is a small file, trivial in size next to the weights, and without it the weights are unusable.
A checkpoint carries scaffolding the finished model does not
Training systems save their state periodically so that a hardware failure costs hours rather than weeks. What they save is called a checkpoint, and it holds more than the parameters: the optimiser’s running averages, the position reached in the data, the state of the random number generator, and whatever else a resumed run would need.
That makes a checkpoint several times larger than the model inside it. What eventually gets published is usually a stripped version — parameters only, sometimes stored at reduced precision — which is why a released file can be a fraction of the size of the thing that came off the training cluster.
The distinction matters for anyone continuing to train a released model. They can, but the optimiser starts from nothing, having lost the accumulated averages that were smoothing the updates. The early part of a continued run therefore behaves differently from the end of the original one, and not always benignly.
Copying costs nothing, and that governs everything downstream
Once the file exists, duplicating it is an ordinary copy operation. Nothing degrades, nothing is consumed, and the marginal cost is bandwidth. That is unremarkable for software in general and consequential here, because the entire expense sat in producing the file and none of it is recovered by restricting copies.
So arguments about releasing weights are arguments about a file that behaves like every other file the moment it leaves. A release cannot be recalled. Licences, usage terms and takedown requests operate through law and social pressure rather than through anything technical, and the people drafting them know it.
It also explains why protecting a model artefact resembles protecting a confidential document rather than protecting a service. A service can be rate-limited, logged, watched and switched off tomorrow. A copied file is somewhere else now, and it will still work in a decade.
Several things people assume are inside it are not
The training data is not in there. Statistical residue of it is, unavoidably, and fragments of memorised text can sometimes be coaxed back out, but there is no archive to search and no index of what was read. This frustrates everybody from copyright lawyers to the engineers who built the thing.
Nor is most of the behaviour people associate with a product. Standing instructions, filters on input and output, retrieval, tool access and request routing all live in the service wrapped around the file. Two deployments of identical weights can differ enough that users would swear they were different systems.
That gap is worth holding onto whenever you read that a model was released, tested or compared against another. The file is one component in an assembly, and a good deal of the assembly is ordinary software written by ordinary teams. Interesting behaviour is rarely attributable to a single layer.
Common questions
Can you tell what a model can do by looking at the file?
Barely. The file reveals its size and its architecture, which set an upper bound on capacity and identify the family it belongs to. Everything else — what it was trained on, which behaviours were tuned into it, how well it performs on anything — has to be established by running it.
Is the model the same thing as the product people use?
No, and conflating the two causes confusion in both directions. The product includes instructions the model receives before you say anything, filtering, retrieval, memory features and routing between several models. Changes to any of those alter behaviour without a single parameter moving.
Why do released models sometimes come in several file sizes?
Usually because the same parameters have been stored at different numerical precisions, trading fidelity for size and speed. Sometimes it indicates genuinely different models trained at different scales. Naming conventions are not standardised across the field, so the accompanying documentation is the only reliable way to tell which case you are looking at.
Consumer editor, AI Worth Knowing
Manish has written about how it works, in the world, limits & risks for most of the last decade and prefers a plain explanation to a clever one.





