Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Machine translation cleared a threshold and left its hardest problems behind

The jump in quality was real and it changed how much of the world people can read, but the failures that remain are more dangerous than the clumsy ones it replaced.

By Naina Sethi4 min read

Close-up of a woman holding a sleek black prosthetic arm in a studio setting.
Photograph by cottonbro studio via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Three approaches, and only one of them scaled

The first serious attempts encoded grammar directly. Linguists wrote rules describing how a sentence in one language maps to a sentence in another, and engineers built systems that applied them. The output was often correct in a stilted way and often catastrophically wrong, because natural language has more exceptions than rules and the exceptions cannot be enumerated.

The second approach abandoned the rules and counted instead. Given a large body of text that already existed in two languages — parliamentary proceedings, treaty documents, subtitles — a system could learn which phrases tended to correspond to which, and assemble a translation from the pieces. The results were noticeably better and unmistakably mechanical, with the seams between the borrowed phrases visible in every paragraph.

The third approach trained a single network to map a whole sentence to a whole sentence, with no explicit phrase table and no rules at all. That is what produced the change everybody noticed. The seams disappeared, because there were no longer any pieces being joined.

Fluency outran accuracy, and that is the problem

A rule-based system that failed produced something obviously broken, and a reader could tell instantly not to trust it. A neural system that fails produces a fluent, well-formed, entirely plausible sentence that happens to say something the original did not. The error is harder to see precisely because the output is better.

This shows up most sharply where a sentence contains information the target language requires and the source language did not supply. Many languages force a choice about formality, gender or number that the original left open, and the system must pick something. It picks whatever was statistically most common in its training data, and that is how translated text acquires assumptions nobody wrote.

Numbers, names and negation are the other recurring weak points. A dropped negation reverses a meaning entirely while leaving a perfectly grammatical sentence behind, which is the exact failure profile you would design if you wanted mistakes to survive proofreading.

The quality gap between languages is enormous

These systems learn from parallel text, and parallel text exists in wildly unequal quantities. Pairs of widely published languages with long histories of official bilingual documentation have vast resources. Many languages with tens of millions of speakers have comparatively little written down in translated form at all.

The result is that quality varies by language pair far more than most users realise, and the interface gives no indication. The same tool, with the same confident presentation, produces professional-grade output for one pair and something closer to a rough gist for another.

Techniques exist to narrow this — training many languages together so that structure learned from well-resourced pairs transfers, or generating synthetic parallel text. They help, and they do not close the gap. How far the gap can be closed by these methods is an open research question rather than a solved engineering problem.

Deployment moved faster than the caveats

Translation is now embedded in places where the consequences of an error are serious: medical intake, immigration and asylum interviews, contractual documents, safety instructions and courtroom material. In each of those settings the professional discipline of interpreting exists for reasons that predate the technology, and those reasons have not gone away.

A human interpreter can say that a phrase does not translate, ask a clarifying question, notice that the speaker misunderstood, and refuse to guess. A translation model does none of these things, and it has no mechanism for signalling that it is out of its depth. It returns a confident sentence for every input it is given.

That does not make the tools useless in serious settings — a rough understanding available immediately is often better than a perfect one available next week. It means the tool and the profession are answering different questions, and treating them as substitutes rather than a sequence is where the harm has tended to occur.

What it did to the work

The effect on translators is contested and probably varies by segment. One account is straightforward displacement: work that used to be paid for is now done by software, and the remaining paid work is post-editing at lower rates and less pleasant conditions. Practitioners in several markets describe exactly that.

The competing account is that volume expanded. Enormous quantities of material that would simply never have been translated at professional rates now get translated, and some fraction of that generates demand for a human to check the parts that matter. Literary, legal and marketing translation, where the point is not the literal content, appear less affected.

Both things can be true in different corners of the same industry, and anyone offering a single number for the net effect is going beyond what the evidence supports.

Common questions

Why does a translation get worse in longer documents?

Consistency is the usual culprit. A term or a name rendered one way early may be rendered differently later, because each passage is handled with limited regard for choices made elsewhere. Professional translation tools maintain glossaries specifically to prevent this, which tells you it is a known and structural weakness.

Is translating through a third language a real thing?

It has been, historically — routing an uncommon pair through a widely resourced language because direct parallel data was scarce. It compounds errors, since anything lost in the first step cannot be recovered in the second. Direct multilingual training has reduced the practice but the underlying data scarcity has not disappeared.

Can these systems handle dialects?

Variably, and usually by flattening them towards whichever standard form dominated the training text. Speakers of regional varieties often find their input understood but their output returned in a register they would not use, which is a data composition issue rather than a design decision anyone made.

In The Worldtranslationlanguagedeploymenterrors
Naina Sethi
Features writer, AI Worth Knowing

Naina joined to cover how it works, in the world, limits & risks and stayed for the awkward questions and is unreasonably interested in the detail nobody else checks.