Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Where the computation happens decides what the system is allowed to be

Running a model on the device in your hand and running it in a distant building are not the same product with different plumbing; the choice constrains privacy, cost, capability and who can switch it off.

By Naina Sethi3 min read

Woman strategizing a chess game against a robot arm, illustrating technology and strategy.
Photograph by Pavel Danilyuk via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Two architectures with different politics

A model can execute on the hardware a person is holding, or the input can travel to a data centre and the answer travel back. Everything else about the product may look identical. Underneath, these are different arrangements of responsibility, cost and control, and the differences are not subtle once you look.

On-device work is bounded by the device. Memory is limited, sustained power draw makes the thing hot, and battery is finite. Server-side work is bounded by what the operator will pay for, which is a much looser bound, and by the network sitting between the person and the machine.

Neither is a strictly better choice. The interesting question is which constraint you would rather be subject to, and that depends on what the system does and who is expected to trust it with what.

The phone is a harder engineering problem than it sounds

A modern phone contains dedicated hardware for the kind of arithmetic these models need, and it is genuinely capable. The binding constraint is usually not raw speed but memory bandwidth: moving the parameters from storage into the processor takes time and energy, and for large models the movement dominates the arithmetic.

That is why on-device deployment leans on the techniques that shrink models — reduced numerical precision, pruning, training smaller models to imitate larger ones. Each of those costs some capability, and how much it costs depends on the task in a way that is hard to predict from benchmarks.

There is also a thermal ceiling nobody can engineer away. A phone that runs a heavy model continuously gets warm and then throttles itself, so sustained workloads behave differently from a demonstration lasting thirty seconds. Reviews and launch demonstrations rarely make this visible.

The privacy argument is strong and it isn’t total

If the input never leaves the device, an entire category of risk disappears: no transmission to intercept, no server-side retention, no employee with access to logs, no jurisdiction question about where the data came to rest. For sensitive material this is a substantial and real advantage.

The qualifications matter though. Many products described as on-device run only part of the work locally and send harder requests onward, sometimes without making the handover visible. Telemetry may travel even when content does not. And a local model can still be updated by the vendor, which means the software making decisions about your data changes without your involvement.

So local processing is a meaningful privacy property rather than a guarantee, and the useful question is which specific data stays put rather than whether the label was applied to the product. Details vary between products and change between versions.

The bill lands on different people

Server-side inference costs the operator money on every request, which creates pressure towards usage limits, tiers, advertising, or routing cheap requests to smaller models. The user pays indirectly and the operator carries a cost that grows with success, which is an unusual shape for a software business.

On-device inference costs the operator almost nothing after distribution. The user pays in battery, storage and the price of hardware capable of running it. That shifts the economics decisively and it also shifts them regressively, since capable devices are expensive and the people with older hardware get the degraded experience.

Availability differs too. A local model works on a plane, in a tunnel and during an outage. A remote one does not, and it can also be withdrawn, repriced or altered by the operator at any moment, which is a governance property dressed as a technical one.

How the balance develops is contested

One camp expects most work to migrate to devices as efficiency improves and small models get better, leaving the data centre for genuinely hard requests. The efficiency gains supporting this argument have been real and repeated, and the privacy and cost incentives all point the same way.

The other camp expects the opposite, on the grounds that the most capable models keep growing and that people reliably prefer capability over locality when the difference is visible. On this view local models remain a fallback and a feature checkbox rather than the main event.

Both are predictions and should be read as such. The hybrid arrangement — a small local model handling routine work and escalating the rest — currently looks like the practical answer, though it inherits the drawbacks of both sides and makes it harder for anyone to say where their data actually went.

Common questions

Is a local model always worse than a remote one?

Usually less capable at open-ended work, because size still buys a great deal. For narrow, well-defined tasks — transcription, wake words, photo categorisation, translation of common phrases — the gap can be small enough that the latency and privacy advantages of running locally outweigh it comfortably.

Why do some features work offline and others do not?

Because the product splits its work. Anything handled by a small local model continues without a connection; anything routed to a server stops. The split is a design decision that vendors do not always document, and it can change with an update without any visible announcement.

Does running locally mean no data is collected?

No. Content processed on the device may stay there while usage statistics, error reports and feature telemetry still travel. Those are separate systems with separate settings, and the presence of local processing says nothing about them either way.

In The Worldon-devicecloudprivacydeployment
Naina Sethi
Features writer, AI Worth Knowing

Naina joined to cover how it works, in the world, limits & risks and stayed for the awkward questions and is unreasonably interested in the detail nobody else checks.