Neural machine translation has improved to the point where many language pairs now read as intelligible and often quite good. So why is it still not perfect?
The starting point is how humans produce language, drawing on context, emotion and experience, and why those things are difficult for a machine to replicate.
From there, the focus turns to how translation quality is assessed, including the BLEU, TER and GTM scoring systems, and why there is no single perfect translation to measure against. A survey of 150 professional linguists adds their own view: their awareness of AI, how they expect their work to change, and how often they currently use machine translation.
The conclusion is measured. Automatic evaluation is still less reliable than human judgement, and machines cannot yet replace it. For now, the strongest results come from combining smart technology with human input and decision-making.
Download Whitepaper