Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Multimodal means one system takes more than one kind of input

The term describes an architecture rather than an ability, several quite different arrangements are sold under it, and telling them apart from the outside is harder than it should be.

By Daniel Okonkwo3 min read

An open Bible displaying chapters of Acts in Spanish and English, photographed from above.
Photograph by Jesus Vidal via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The definition, and what it is doing work to exclude

A modality is a kind of signal: text, images, audio, video, and less commonly things like depth readings or sensor traces. A multimodal system handles more than one. Stated that plainly the term is almost too broad to be useful, which is part of the problem with it.

The stricter sense means a single model that processes several signal types within the same computation, so that information from one can influence the handling of another at every stage. That is a meaningful architectural claim and it is not what every product described this way actually does.

The looser and very common arrangement chains separate models together: one transcribes speech to text, a language model handles the text, another turns text back into speech. Each component is single-modality and the assembly is described as multimodal, which is understandable marketing and a different thing entirely.

Getting different signals into the same representation

Everything a network processes must become a list of numbers, so the question is how each signal type gets converted. Text becomes fragments looked up in a table. Images are usually cut into patches, each patch converted into a vector. Audio is commonly turned into a representation of frequency over time and then treated much like an image.

Once converted, all of these are sequences of vectors, and the machinery that relates elements of a sequence to one another does not particularly care where they came from. That indifference is the reason the current architecture absorbed other signal types as readily as it did.

The conversion step is not neutral, though. Cutting an image into patches of a chosen size discards fine detail below that scale, which is why systems handling images often struggle with small text, precise counting and exact spatial relationships. The limitation enters at the door, before the model sees anything.

Where the signals meet decides what the system can do

Combining the signals early, so they are processed jointly throughout, allows one to inform the interpretation of another — the tone of a voice affecting how its words are read, or a caption disambiguating what a picture shows. It is more expensive and requires training data pairing the signals.

Combining them late, by processing each separately and merging conclusions at the end, is cheaper and easier to assemble from existing components. It also loses the interaction, which is the entire justification for the approach in the first place.

A very common middle route attaches a trained image encoder to an existing language model through a small connecting layer, then trains mostly that connection. This is efficient and it means the visual side is grafted onto a system whose competence was formed on text, which shows in what it does well.

The grounding argument, which is not settled

One long-standing position holds that a system trained only on text cannot properly understand words referring to physical things, since it has never had access to the referents. On this view adding vision addresses a real deficiency rather than merely adding a feature.

The counter-position notes that text-only systems handle spatial and physical language better than the argument predicts, apparently by absorbing an enormous quantity of written description, and that the improvements from adding vision have been narrower than the strong form of the grounding claim would suggest.

The evidence points both ways and the disagreement remains live. What is not disputed is that vision helps a great deal with tasks that are actually visual, which is a much weaker claim and the one most product capabilities rest on.

Claims are hard to check from outside

Because several architectures share the label, a stated capability rarely tells you what is underneath. A system that describes an image might be examining it within a joint computation or might be running a captioning model and passing the caption along, and the two behave differently when asked something the caption did not cover.

Evaluation is also weaker here than for text alone, since a test must pair signals and there are fewer established collections. Results are correspondingly noisier and comparisons between systems less reliable, a state the field acknowledges and is working on.

The useful question is not whether a system is multimodal but which signals it takes, whether they interact, and at what resolution each is represented. Those determine behaviour. The label determines nothing.

Common questions

Why do these systems misread text inside an image?

Usually because of resolution. The image is divided into patches and reduced before processing, so small lettering can fall below what the representation preserves. Some systems mitigate this by processing regions at higher detail, which costs more and is not applied universally.

Is a system that generates images multimodal?

By the loose definition yes, since it handles text and pixels. Generation and interpretation are separate capabilities though, and many systems do one considerably better than the other. A product combining both frequently uses two distinct models behind one interface.

Does adding a modality make a model generally smarter?

Not in any broad sense. It adds competence on tasks involving that signal and can slightly affect other behaviour, for better or worse, depending on how the training was balanced. Claims of general improvement from adding vision or audio should be read against what was actually measured.

Jargonmultimodalvisionaudioterminology
Daniel Okonkwo
Contributing editor, AI Worth Knowing

Daniel covers how it works, in the world, limits & risks and the questions readers actually send in and prefers a plain explanation to a clever one.