How good is GPT-4o at machine translation
2 min read
A quick primer on BLEU scores and a tiered view of all 95 languages it supports What BLEU actually measures BLEU (Bilingual Evaluation Understudy) is the most common automatic metric for translation quality. It compares the model’s output with one or more human reference translations and counts overlapping n-grams (word sequences). Scores range from 0 to 100. In practice anything…

A quick primer on BLEU scores and a tiered view of all 95 languages it supports
What BLEU actually measures
BLEU (Bilingual Evaluation Understudy) is the most common automatic metric for translation quality.
It compares the model’s output with one or more human reference translations and counts overlapping n-grams (word sequences).
Scores range from 0 to 100. In practice anything above ~50 is close to professional human performance; scores below 20 signal heavy post-editing.
GPT-4o’s indicative BLEU values below come from public WMT and FLORES-200 test suites, plus in-house probes run with temperature 0 and no few-shot help. They are averages—scores vary by domain and register.
Four performance tiers across the full 95-language set
*Montenegrin and Bosnian are not given separate ISO codes in the Whisper tokenizer—they share sr (Serbian)—yet empirical testing shows GPT-4o treats their orthography correctly and reaches the same Tier 2 quality.
How to read the table
BLEU bands are indicative, not absolute. GPT-4o can score higher with careful prompting, glossaries, or few-shot examples—especially in Tiers 2 & 3.
Dialects and scripts matter. Serbian Latin vs Cyrillic yields identical scores; Cantonese (Tier 4) fares worse than Standard Chinese (Tier 1).
Data scarcity drives the tiers. Languages with abundant parallel corpora naturally sit higher; truly low-resource tongues fall to Tier 3/4 despite the same architecture.
Why BLEU alone is not enough
A high BLEU does not guarantee perfect style, factual accuracy or domain terminology. Conversely, low BLEU may still be usable for gisting. Always combine automated metrics with human review—particularly for legal, medical or creative texts.
Bottom line: GPT-4o delivers professional-level output in ~15 high-resource languages and solid draft quality in roughly half of the remaining 80. For Montenegrin, Bosnian and other mid-resource varieties the model already matches classic SMT engines; for the dozen truly low-resource tongues, human-in-the-loop workflows remain essential.
GPT-4o scores > 45 BLEU for 15 high-resource languages, 30-45 for ~50 mid-resource, 15-30 for 22 low-resource and < 15 for 8 experimental tongues—quality drops with data scarcity.
#GPT4o #BLEU #TranslationQuality #AIlanguages #NLP #TradAI
https://www.linkedin.com/pulse/how-good-trad-ai-translation-trad-ai-official-ws5ke
More Trad AI news
Previous article
Trad AI Translation vs. Segment-by-Segment Translation
As artificial intelligence continues to evolve, the field of machine translation is witnessing a profound transformation. Traditional machine translation realised in CAT tools operates by translating individual sentences or segments without considering broader contextual information. However, emerging AI models that incorporate wide context provide significant advantages, resulting in translations that are more coherent, accurate, and aligned with the subtleties of…
Next article
Pick the Right API Key for Trad AI Translations
In today's globalized world, the quality of machine translations can significantly impact businesses. Trad AI, an advanced online translation platform, leverages powerful AI models to deliver accurate and contextually nuanced translations. However, the effectiveness of Trad AI hinges largely on selecting the appropriate API key that provides access to the right AI model. This article explains why opting for API…
How Trad AI fits into your workflow
Use your own OpenAI API key, choose model behaviour, and keep every article translation aligned with the tone and terminology your team expects.
See how it works