← Back to news

How good is GPT-4o at machine translation

2 min read

A quick primer on BLEU scores and a tiered view of all 95 languages it supports What BLEU actually measures BLEU (Bilingual Evaluation Understudy) is the most common automatic metric for translation quality. It compares the model’s output with one or more human reference translations and counts overlapping n-grams (word sequences). Scores range from 0 to 100. In practice anything…

Trad AI AI translation machine translation localization CAT-tools workflow
Performance of GPT-4o across languages based on machine translation evaluations

A quick primer on BLEU scores and a tiered view of all 95 languages it supports

What BLEU actually measures

BLEU (Bilingual Evaluation Understudy) is the most common automatic metric for translation quality.

It compares the model’s output with one or more human reference translations and counts overlapping n-grams (word sequences).

Scores range from 0 to 100. In practice anything above ~50 is close to professional human performance; scores below 20 signal heavy post-editing.

GPT-4o’s indicative BLEU values below come from public WMT and FLORES-200 test suites, plus in-house probes run with temperature 0 and no few-shot help. They are averages—scores vary by domain and register.

Four performance tiers across the full 95-language set

*Montenegrin and Bosnian are not given separate ISO codes in the Whisper tokenizer—they share sr (Serbian)—yet empirical testing shows GPT-4o treats their orthography correctly and reaches the same Tier 2 quality.

How to read the table

BLEU bands are indicative, not absolute. GPT-4o can score higher with careful prompting, glossaries, or few-shot examples—especially in Tiers 2 & 3.

Dialects and scripts matter. Serbian Latin vs Cyrillic yields identical scores; Cantonese (Tier 4) fares worse than Standard Chinese (Tier 1).

Data scarcity drives the tiers. Languages with abundant parallel corpora naturally sit higher; truly low-resource tongues fall to Tier 3/4 despite the same architecture.

Why BLEU alone is not enough

A high BLEU does not guarantee perfect style, factual accuracy or domain terminology. Conversely, low BLEU may still be usable for gisting. Always combine automated metrics with human review—particularly for legal, medical or creative texts.

Bottom line: GPT-4o delivers professional-level output in ~15 high-resource languages and solid draft quality in roughly half of the remaining 80. For Montenegrin, Bosnian and other mid-resource varieties the model already matches classic SMT engines; for the dozen truly low-resource tongues, human-in-the-loop workflows remain essential.

GPT-4o scores > 45 BLEU for 15 high-resource languages, 30-45 for ~50 mid-resource, 15-30 for 22 low-resource and < 15 for 8 experimental tongues—quality drops with data scarcity.

#GPT4o #BLEU #TranslationQuality #AIlanguages #NLP #TradAI

https://www.linkedin.com/pulse/how-good-trad-ai-translation-trad-ai-official-ws5ke

More Trad AI news

Previous article

Trad AI Translation vs. Segment-by-Segment Translation

As artificial intelligence continues to evolve, the field of machine translation is witnessing a profound transformation. Traditional machine translation realised in CAT tools operates by translating individual sentences or segments without considering broader contextual information. However, emerging AI models that incorporate wide context provide significant advantages, resulting in translations that are more coherent, accurate, and aligned with the subtleties of…

Next article

Pick the Right API Key for Trad AI Translations

In today's globalized world, the quality of machine translations can significantly impact businesses. Trad AI, an advanced online translation platform, leverages powerful AI models to deliver accurate and contextually nuanced translations. However, the effectiveness of Trad AI hinges largely on selecting the appropriate API key that provides access to the right AI model. This article explains why opting for API…

How Trad AI fits into your workflow

Use your own OpenAI API key, choose model behaviour, and keep every article translation aligned with the tone and terminology your team expects.

See how it works

Try Trad AI

Open the workspace