Illustration of AI-generated translations streaming in while a reviewer compares source and target documents with a magnifying glass

Translation used to be slow, deliberate work. A translator would read the source text, look things up when something was unclear, weigh terminology choices against dictionaries and reference materials, and build the translation one sentence at a time.

Today, you can paste a document into a generative AI or machine translation system and get a translation back in seconds. Watching work that once took hours finish in an instant, it’s easy to conclude that translators are no longer needed — or that translation should cost next to nothing.

AI has, in fact, dramatically cut the time and cost of generating a translation. If you just need the gist of a document, want to read an email in another language, or are sharing information internally, AI output is often good enough to use as is, with no human review at all.

Delivering a translation to a third party with a defined level of quality is a different matter. Producing the text is only one part of the job; someone still has to confirm that it accurately conveys what the source says.

In this article, I’ll split the translation process into two parts — “generation” and “quality assurance” — and look at what AI has changed and what work remains.

Translation output has genuinely improved

Few people would argue that machine translation reads the way it did a decade ago.

A major turning point came around 2016, when neural machine translation entered widespread use. In its 2016 paper on the Google Neural Machine Translation system (GNMT), Google reported that human evaluators found an average of 60% fewer translation errors compared with the earlier phrase-based system, based on tests using isolated simple sentences.

That evaluation ran under limited conditions, but it showed a real leap over the previous generation of systems. What it did not show is that machine translation had reached human parity for every kind of document.

Progress didn’t stop in 2016, either. Neural machine translation kept spreading, and large language models (LLMs) have since joined the toolkit. Today’s AI can produce grammatical, fluent translations in very little time.

So it would be unfair to judge current AI translation by memories of the clunky machine output of the past. This article takes it as a given that AI translation is already useful for many purposes — and asks what work is still left when accuracy has to be guaranteed.

Generating a translation and assuring its quality are different jobs

Producing a translation means reading the source, choosing terminology, and constructing the target text. AI has sped up much of that generation work.

But when a defined level of quality is on the line, the generated translation still has to be checked. For example:

  • Has any source information been dropped, or has anything been added that isn’t in the source?
  • Have modifiers or causal relationships changed, or has a positive statement become negative or vice versa?
  • Are the numbers, units, and proper names correct?
  • Does the terminology fit the field and the context?
  • Are terms and phrasing consistent across the document?
  • Does the translation actually work for its intended use?
  • Did a correction introduce a new error somewhere else?

Quality assurance costs, as I use the term here, don’t just mean running an automated QA tool at the end. They include the time, expertise, and checking work required to compare the source and target texts and decide whether the translation meets the required standard.

Parts of quality assurance can be automated too, of course. AI and computer-assisted translation (CAT) tools are good at suggesting terminology, retrieving past translations, and flagging numeric mismatches or simple omissions. But automating the generation step is not the same as eliminating the quality assurance process.

Post-editing is not a quick read-through

Having a person fix the output of machine translation or AI translation is generally called post-editing.

If the job were just reading the finished translation and fixing awkward wording or typos, it would look light. But checking accuracy against the source means reading both texts. A translation can sound perfectly natural and still get the source wrong.

A post-editor has to understand the source, work out how the AI interpreted it, judge whether that interpretation holds up, and decide how to fix it when it doesn’t. When the initial output is good, post-editing can be faster than translating from scratch. Depending on the difficulty of the text, the quality of the draft, and the standard required, though, finding and fixing errors can still take real time.

When translators review their own work, they’ve already read the source, researched the subject, and made every terminology and drafting decision themselves — and they generally remember why. The review still matters, but it carries less load than reading both texts cold. A third-party reviewer, on the other hand, has to read and interpret both the source and the translation from scratch, making the work itself similar to post-editing.

An AI-generated draft is, likewise, someone else’s translation as far as the reviewer is concerned. On top of reading the source, the reviewer has to read the AI’s output and reconstruct how it understood the text. Compared with reviewing your own translation, that’s an extra layer of work.

Comparison of human translation and AI/MT translation workflows

So neither blanket claim holds up: post-editing isn’t always harder than translating from scratch, and having a first draft doesn’t always cut the work by some fixed percentage.

Research on post-editing effort looks at it from several angles: time, actual editing operations, and cognitive effort — the mental work of understanding the content and deciding whether it’s right. A 2025 study of English-to-Chinese translation found that post-editing with GPT-4 assistance did not significantly reduce working time compared with conventional post-editing, while the amount of editing activity went up. Its effect on cognitive effort was inconsistent as well.

That study involved 26 postgraduate students, so its results won’t generalize to every language, field, or professional translator. The takeaway isn’t that AI assistance is useless — it’s that adding AI doesn’t automatically make the work lighter.

There’s also a standard for this. ISO 18587:2017 sets out requirements for full human post-editing of machine translation output and for post-editor competence. It treats post-editing as a defined process carried out by qualified people — not a cursory skim of the output.

Fluent output can still get the meaning wrong

The great strength of current AI translation is fluency. From a quality assurance standpoint, that strength cuts both ways.

Awkward output announces its own problems. A translation that is grammatical and plausible, on the other hand, can hide a shift in meaning: part of the source silently dropped, the scope of a limitation changed, a causal link added that the source never made.

Terminology works the same way. Take trigger the function: an AI might render it into Japanese as 機能の発火 — literally, “the firing of the function”. The phrase has the shape of technical Japanese, so to some readers it looks right. In practice, the appropriate Japanese rendering may be 機能をトリガーする (“trigger the feature”), 機能を起動する (“activate the feature”), or 関数を呼び出す (“call the function”), depending on the subject matter and the context.

The point of this example isn’t to grade any particular AI system. It’s that fluency and correct meaning are separate questions, and they have to be checked separately.

The quality you need depends on what the translation is for

If you’re using AI translation to get the gist of a document, you’ve met your goal the moment you understand roughly what it says. For that purpose, today’s AI translation is already highly useful.

Supplying a translation to someone else as a deliverable raises the bar. Now it’s not just about readability — the translation has to be checked against the source. Is the meaning the same? Is the terminology right? Does the document do its job? Could an error cause a real-world problem?

That’s why one person can call a translation “good enough” while another says “this isn’t ready to go out”. The two assessments don’t necessarily conflict — they’re based on different purposes and different quality standards.

The perception gap around AI translation isn’t only a disagreement about how good AI is. Much of it comes down to different definitions of when the translation job is done.

Price the total work, not the fact that a machine was involved

Since machine translation went mainstream, post-editing rates have sometimes been set far below translation rates, on the assumption that the machine translates and the human just looks it over. I’ve actually seen postings offering roughly one yen per source character for post-editing.

A number alone doesn’t tell you whether a rate is fair. The workload depends on the quality of the initial output, how specialized the text is, the standard required, the consequences of an error, and the deadline. If the quality bar is modest and the machine output is consistent, a low rate may still be reasonable in relation to the time required. If the job demands detailed comparison against the source and specialist judgment, a first draft may not save much work at all.

A rational price difference, then, shouldn’t come from applying a fixed discount because “a machine did the translation”. It should come from measuring how much the time to reach the required quality actually dropped. The number that matters isn’t how fast the first draft appears — it’s the total time to make the translation ready to deliver.

The checks that remain in patent translation

In patent translation, accurately reflecting the source takes priority over reading smoothly. Some awkwardness may be acceptable when necessary to preserve the technical meaning or the scope of rights accurately. Quality assurance means examining exactly how the wording affects both, including:

  • the elements recited in a claim and the relationships between them;
  • the handling of singulars, plurals, and alternatives;
  • limiting language that has been added or dropped;
  • consistency between the claims and the description;
  • correspondence with reference signs in the drawings; and
  • the semantic range of each chosen term in the target language.

Even if AI cuts the time to produce a first draft, the time required for these checks doesn’t necessarily fall by the same proportion. At the same time, AI and translation tools can improve efficiency in many parts of this work: terminology checks, searches for related documents, comparisons of phrasing, and checks of numbers and reference signs.

The real question isn’t whether to use AI. It’s which parts of the process can be automated — and which still need human judgment.

What got cheaper, and what didn’t

AI has slashed the cost of generating translated text. For some purposes, you can now get a usable translation almost instantly, with no further work. That’s a real advance in translation technology.

But wherever accuracy, subject-matter appropriateness, consistency, or legal effect has to be assured, the generated text still has to be verified. AI can make that work more efficient too, but the cost of verification doesn’t necessarily fall at the same rate as the cost of generating the initial translation.

When you evaluate what AI is doing for you, don’t look only at how fast a translation appears. Look at the total time and work it takes to reach the quality you need.

Edit distance — a measure of how much a text was changed — is sometimes used as a proxy, but it isn’t the same as total effort. Even a translation that needs almost no correction takes time to confirm as correct. A meaningful assessment of AI’s benefits has to account for checking time as well as the volume of edits.

Keeping generation and quality assurance separate makes it easier to step away from the two extremes — “AI makes translation basically free” and “AI is useless for specialist translation” — and toward practical decisions based on purpose, quality, and risk.

References