Jargon
An embedding is a position, and the geometry is what does the work
Turning words, images or documents into lists of numbers sounds like a technicality, and it is the step that lets a machine treat meaning as distance.
By Zoya Rahman3 min read

A machine needs numbers before it can need anything else
Arithmetic is all a neural network does, so anything it processes must first become numbers. The naive approach assigns each word an arbitrary identifier, which technically works and is useless, because the identifiers carry no relationship to one another. Nothing about the number for cat is closer to the number for kitten than to the number for saxophone.
An embedding solves this by representing each item as a list of numbers — a point in a space of many dimensions, typically hundreds or thousands. The list is not assigned; it is learned, adjusted during training until items used in similar ways end up in similar places.
That is the entire idea, and it is one of the more elegant things in the field. Similarity of meaning becomes proximity in space, and proximity is something arithmetic can handle.
The training signal is company, not definition
The learning procedure rests on an old observation from linguistics: that a word is characterised by the words it appears alongside. Terms that occur in the same contexts tend to be used for similar purposes, and a system that predicts context from a word, or a word from its context, is forced to encode that similarity.
Notice what this does and does not capture. It captures usage, thoroughly and at a scale no lexicographer could match. It does not capture reference, because the training data contains no link between the word and any object in the world. The system learns how a word behaves among other words.
This is the strongest form of a familiar objection, and it deserves to be stated fairly rather than dismissed. It is also worth noting that the resulting representations turn out to be far more useful than the objection would predict, which is a fact the objection has to accommodate.
Direction turned out to carry meaning too
The finding that made embeddings famous was that the space has structure beyond clustering. Differences between positions appeared to correspond to relationships, so that moving from one word to another traced a direction that could be applied elsewhere and land somewhere sensible. Analogies could be computed by arithmetic on positions.
These demonstrations were genuinely striking and they were also somewhat oversold. Later analysis showed that the standard examples depended on details of how the arithmetic was performed and which candidate answers were excluded, and that the effect is real but weaker and less general than the early presentations implied.
The durable lesson is that the geometry is not arbitrary. Something about the structure of usage gets encoded as structure in the space, which is why embeddings work as well as they do for retrieval and comparison.
Context changed what an embedding is
Early embeddings assigned one fixed position per word, which meant a word with several meanings got one position awkwardly averaged between them. The word bank sat somewhere unhelpful between finance and rivers, useful for neither.
Modern systems produce a position for each occurrence, computed from the surrounding text, so the same word occupies different places depending on how it is being used. This is what the layers of a language model are largely doing: repeatedly refining a set of positions in the light of everything else present.
The word embedding now covers both ideas, and the ambiguity causes regular confusion in discussion. A fixed lookup table and a context-dependent representation are quite different objects sharing a name.
Where they are used, and what leaks
The practical workhorse application is search by meaning rather than by keyword. Documents are embedded once, a query is embedded on arrival, and the nearest documents are returned, which finds material that shares no words with the query at all. This is the retrieval half of nearly every system that answers questions from a document collection.
Two cautions matter. Embeddings absorb the associations present in their training text, including the ones nobody wants, and these have been measured directly in the geometry — the same structure that encodes useful relationships also encodes stereotypical ones. Various correction methods exist and their effectiveness is contested.
The second is that an embedding is not anonymised. It is a lossy transformation, not a one-way one, and work on reconstructing source text from embeddings has had more success than most people assume. Treating a vector database as though it contained no sensitive content is a mistake that gets made routinely.
Common questions
How many dimensions does an embedding have?
It varies with the system, from a few dozen for simple applications to several thousand for large models. More dimensions allow finer distinctions and cost more to store and search. There is no natural right answer, and the choice is an engineering trade rather than a discovery about language.
Can embeddings from two different models be compared?
No. Each model learns its own space with its own arrangement, so a position from one is meaningless in another even if the dimensions match. Anything comparing embeddings must produce them all with the same model, and changing models means recomputing everything you have stored.
Is cosine similarity the right way to compare them?
It is the usual default, measuring the angle between two positions while ignoring their length, and it works well in practice. Whether it is optimal depends on how the embeddings were trained, and there is a modest literature arguing that it is applied more automatically than the theory justifies.
Deputy editor, AI Worth Knowing
Zoya joined to cover how it works, in the world, limits & risks and stayed for the awkward questions and would rather show the working than assert the conclusion.





