Jargon
Quantisation is lossy compression applied to a model’s numbers
Storing each parameter with less precision makes a model smaller and faster at some cost in fidelity, and the interesting part is which abilities degrade first.
By Daniel Okonkwo3 min read

Precision is a choice, and it was always generous
Every parameter in a model is a number held in a fixed number of bits, and the count determines how finely values can be distinguished. Training has conventionally used relatively generous precision, because the tiny adjustments made during training need somewhere to accumulate and would otherwise be rounded away to nothing.
Running a finished model has no such requirement. The parameters are no longer changing, so the question becomes how coarsely they can be represented before behaviour suffers. That question turns out to have a surprisingly forgiving answer.
Quantisation is the practice of storing them more coarsely on purpose. Fewer bits per parameter means a smaller file, less memory occupied, and less data moved per generated fragment, which translates directly into speed. The word comes from signal processing, where it has meant much the same thing for decades: mapping a continuous range onto a limited set of discrete levels and accepting the error that introduces.
Why it works better than it has any right to
The intuition that halving precision should badly damage a model is reasonable and mostly wrong. Individual parameter values are not precise quantities encoding specific facts; they are contributors to a large weighted sum, and errors introduced across many of them partially cancel.
The distribution of values helps too. Parameters cluster heavily around small magnitudes with a thin tail of large ones, so a representation that allocates its available detail where the values actually sit loses much less than a uniform scheme would.
The practical result is that substantial reductions in precision often produce differences that are hard to detect in ordinary use. Push further and the degradation becomes obvious. Where exactly that boundary falls varies by model and by task, which is why the honest answer to how far you can go is always empirical.
What degrades is not uniform, and that is the risk
Loss from quantisation does not spread evenly across everything a model does. Fluency and general conversational competence survive well, because they are supported by broad statistical regularities that coarse representations preserve.
What tends to suffer first is the fragile end: multi-step arithmetic, precise recall of infrequent facts, careful adherence to a complicated instruction, and performance in languages that were thinly represented in training to begin with. These depend on finer distinctions, and finer distinctions are what got rounded away.
This creates a specific evaluation hazard. A compressed model can look indistinguishable from the original on casual inspection and on general benchmarks, while being measurably worse on exactly the demanding cases people deploy it for. Testing on what a model does easily will not reveal it.
The variants are more than a technicality
The simplest approach rounds parameters after training, which is fast and requires no additional work. More careful methods first run sample data through the model to observe the range of values each part actually produces, and set the quantisation scheme accordingly, which recovers much of the loss.
A more involved option is to train the model with the eventual coarseness in mind, so it adapts to the constraint rather than suffering it afterwards. This performs better and costs training compute, which puts it out of reach for anyone working from a released model rather than building one.
Mixed approaches are common: keep sensitive parts of the network at higher precision and compress the rest heavily. That is an admission that the loss is not uniform, made operational. Deciding which parts deserve the extra bits is done by measurement rather than by theory, and the answer differs from one model to another, which is a fair summary of the state of the art in this whole area.
Why this matters beyond the engineering
Quantisation is a large part of why capable models run on ordinary hardware at all. It is the practical bridge between a released file and something a person can run on their own machine, which has consequences for privacy, cost and independence from any provider.
It also introduces a version problem that the vocabulary handles badly. Two people running the same named model at different precisions are running measurably different systems, and comparisons between them, including published ones, are frequently made without stating which.
The general lesson is one this field keeps supplying. Compression is available, it is remarkably effective, and what it costs is not visible where most people look for it.
Common questions
Does quantisation make a model worse?
Slightly, and the size of the effect depends on how aggressively it is applied and on what you are measuring. Modest reduction is often imperceptible in general use, while aggressive reduction reliably degrades demanding tasks first. Calling it lossless is inaccurate; calling it ruinous is equally so.
Why does it speed up generation rather than just save memory?
Because generating text is largely limited by moving parameters out of memory rather than by arithmetic. Fewer bits per parameter means less data to move for each fragment produced, so the speed improvement follows directly from the size reduction.
Can it be undone?
No. Precision that has been discarded is gone, and expanding the numbers back to a wider format restores the container without restoring the detail. It is lossy in the same sense as compressing an image, and the original file is the only route back.
Contributing editor, AI Worth Knowing
Daniel covers how it works, in the world, limits & risks and the questions readers actually send in and prefers a plain explanation to a clever one.





