In this post, we announce the release of the 3.0 version of the FLORES+ dataset for evaluating quality of machine translation in hundreds of languages.
But before, as befits the first post in our blog, we tell about who we are, what FLORES+ is, and why we think it is interesting and important. And we finish with the description of a scientific challenge of enriching open-source multilingual datasets with even more new languages.
Automatic translation of texts between languages (traditionally called “machine translation” by the tech folks) has always been an important application and sometimes even a driving force of computer science. It started with the early 1950s Cold War translation experiments. It culminated in an avalanche of papers in the 2010s which, driven by the purpose of improving translation models, introduced attention mechanism, seq2seq modeling, subword tokenization, and transformers architecture (1, 2, 3, 4) — novel approaches that revolutionized the field of AI, eventually leading to the development of large language models as we have them today. But the task of translation (which motivated all this revolution) remained “solved” only for a few dozen “high-resource” languages spoken by tens of millions people in rich countries. Thousands of other languages from all around the world were left behind this technological wave, and most of them still are, as of today.
There were, of course, people who wanted to challenge this. Most researchers in machine translation were searching for technological improvements under a metaphorical streetlight, experimenting with the language pairs for which there were millions of examples of translated sentences to train models with and well-established benchmarks to measure the progress. But then, among others, a group of scientists from Facebook AI Research started benchmarking translation quality for under-represented languages by introducing in 2019 the first FLoRes dataset, focusing on Nepali-English and Sinhala-English translation. In 2021, they rebuilt this benchmark, extending it to 101 languages from all over the world, facilitating development of AI models intended for translation between any pairs of these languages. In 2022, this work culminated in the release of the No Language Left Behind project (NLLB-200), with the extension of the benchmark to 204 language varieties* and opensourcing of a family of neural models capable of translation between nearly any pair of them. They also released the Seed dataset, a set of 6100 sentences from English Wikipedia translated into 40 under-resourced languages, to enrich the training data for machine translation models.
(*We say “language varieties” instead of “languages”, because the FLORES dataset distinguishes some languages written in several scripts, such as Mandarin written either in Traditional or Simplified characters, and several dialects of some languages, such as North Levantine Arabic and South Levantine Arabic.)
To give a glimpse into the FLORES dataset, we illustrate it with a few translations of a sentence sourced from the Wikivoyage website (the other two dataset sources are Wikinews and Wikibooks):
English: When people don't see moose as potentially dangerous, they may approach too closely and put themselves at risk.
Spanish: Si los individuos no perciben que los alces son potencialmente agresivos, podrían acercárseles más de lo adecuado y ponerse en peligro.
Mandarin Chinese (simplified characters): 有些人不把驼鹿当成潜在的危险,可能会靠得太近,将自己置于危险之中。
Hindi: जब लोग मूस को बहुत ख़तरनाक नहीं मानते हैं, तो वे बहुत पास आ सकते हैं और खुद को जोख़िम में डाल सकते हैं.
Egyptean Arabic: لما الناس ميشوفوش أن حيوان الموظ خطر محتمل، ممكن يقربوا كتير ويعرضوا نفسهم للخطر.
Russian: Если люди не воспринимают лосей как потенциальную опасность, они могут подойти слишком близко, подвергая себя риску.
Ligurian: Quande e persoñe no veddan e alce comme potençialmente peigose, peuan avvexinâse tròppo da arente e mettise in peigo.
Tigrinya: ሰባት ንሙስ ሓደገኛ ጌሮም ኣብ ዘይቆፅሩሉ ግዘ ብጣዕሚ ተፀጊዖምዎ ንዓርሶም ንሓደጋ እንዳጋለፁ እዮም።
Aymara: Jaqinakax alsi tarujanakar jan chhijir purt'ayirjam amuyapki ukjax, anchaw jak'acht'asipxaspa, chhijiruw puripxaspa.
Dzongkha: མི་ཚུ་གིས་ ཀ་ཤ་འདི་ ཉེན་ཁ་མེདཔ་སྦེ་མཐོང་པའི་བསྒང་ལས་ སྦོ་ལོགས་ཁར་སྦྱར་ཏེ་ ཁོང་ར་ ནང་དོག་སར་བཙུགས་ནི་གི་ཉེན་ཁ་ཡོད།
And there are about 2000 sentences like this, translated to each of the two hundred languages (an additional thousand sentences has been sourced by Meta but never publicly released, remaining a hidden dataset split).
The NLLB-200 release was achieved as a result of a centralized, intensive, and time-bound research effort by a dedicated team within a commercial company. But further extending the models and datasets towards all 7000 languages spoken (or signed) around the world is hardly possible with this approach. Maintaining and extending open multilingual datasets in the long term is much better managed by a community of volunteers not constrained by project deadlines and ever changing business priorities.
This is the name (abbreviated to OLDI) taken by a group of enthusiasts who decided to further maintain, develop and extend the open FLORES and Seed datasets. Which means, us 🙂. We call our continuously updated versions of the datasets FLORES+ and OLDI-Seed, to distinguish them from the versions originally released by the Facebook AI Research team. Here is our website: https://oldi.org.
In practice, what we mostly do is supporting and promoting the people who improve or extend the FLORES+ and Seed datasets. Additionally, we occasionally find something to adjust or fix in the structure or the documentation of the datasets.
In 2024, to attract more attention to the extensions to our datasets to new languages, we organized a shared task at WMT24 (the largest scientific conference on machine translation). The participants were encouraged to contribute extensions or improvements to massively multilingual datasets and to publish papers about it. As a result (summarized in this paper), we received ten submissions covering 16 languages:
Abdulmumin et al. corrected FLORES+ in Hausa, Northern Sotho (Sepedi), Xitsonga,and isiZulu.
Ahmed et al. translated Seed into Bangla/Bengali
Ali et al. extended FLORES+ with Emakhuwa
Cols extended Seed with Spanish (Latin American) and developed the Seed-CAT tool for computer-aided translation of this dataset
Gordeev et al. extended FLORES+ with Erzya
Kuzhuget et al. extended FLORES+ with Tuvan
Mamasaidov and Shopulatov extended FLORES+ with Karakalpak
Perez-Ortiz et al. extended FLORES+ with Aragonese, Aranese, and Valencian and corrected its Asturian version
Yu et al. (2024) extended FLORES+ with Wu Chinese.
In fact, there were even more contributions; some of them just were not described in papers. Notably, during the last year, we got FLORES+ translated into Chuvash, Meadow Mari, and Dargwa, indigenous languages of various regions of Russia.
The full list of changes to the datasets since their initial release in 2022 can be found in the changelog files for FLORES+ and Seed.
In 2025, there have already been a few more contributions to FLORES+. Recently, we integrated all of them under the new version of FLORES+, version 3.0. The contributions include:
Adding the Ladin language (Val Badia dialect). It is spoken in the Dolomite Alps and looks kinda similar to Italian with a slight vibe of French and German. And it has the letters öëü 🙃. Many thanks to Samuel Frontull and all the translators!
Updating the spelling for Chuvash and Dargwa (these languages are written in Cyrillic, but the translators of FLORES+ had occasionally used similar Latin letters (for example, Ă and I instead of the similar-looking Cyrillic Ӑ and Ӏ). Many thanks to Alexander Antonov and Murtazali Rabadangadzhiev!
Updated the sentence order for the Aranese dialect (this is a variant of the Occitan language spoken in the Aran valley in the Pyrenees); previously, they had been incorrectly aligned with sentences in other languages. Thanks to Oriane Nédey for the contribution!
With the addition of Ladin, there are now 222 different language variations in the dataset, and you can evaluate the quality of translation between any two of them**!
(** Well, almost any two. The FLORES+ dataset consists of two splits, “dev” and “devtest”, and for a few languages, there is only one of the splits, so they cannot be matched to the languages having only the other split. But most languages have both splits.)
We are extremely grateful to all the contributors, and we are happily trying to spread the word about inclusion of these languages in the benchmark, so that machine translation researchers from all over the world could test (and, hopefully, improve) how their MT systems perform with these languages.
This year, we are again organizing a shared task at WMT25, encouraging any types of contributions to massive open multilingual datasets, or even creation of new such datasets. And we would like you to participate!
The datasets expanded with new languages could be one the ones managed by OLDI (FLORES+ for evaluation and OLDI-Seed for training) or by other organizations (e.g. the training dataset SMOL by Google or the BOUQuET benchmark by Meta).
Based on your contribution, you will be able to write a scientific paper about it (see examples from the last year above) and publish it at the most important conference on machine translation: WMT25, held together with EMNLP 2025 (it will be taking place in Suzhou, China this year). Publishing a paper takes time and, unfortunately, costs money, but it provides useful experience, looks good on a resume, and helps promote the authors of the paper and, most importantly, the language for which they are making the contribution.
To start participating, you should declare us your desire to participate as early as possible (for example, by emailing us at info@oldi.org or in the OLDI discord server). And to get published at WMT25, you will have to write and submit the article by mid–August. Of course, we are ready to help with various technical and methodological difficulties, so please reach out to us if you are considering to participate.
This has been a long post; congratulations on reaching the end of it! If you found this text interesting, please consider subscribing to our newsletter (if not already) or sharing it with your friends. We are going to notify you about updates to our datasets, notable events on machine translation and multilingual NLP, and occasional tutorials on language technologies.
And regardless of which path you take, submitting a whole shared task paper, making a contribution to a dataset without a paper, using FLORES+ or Seed data for your next project, or simply reading about language data and technologies, we would be happy to share with you a part of our journey towards creating open technology for every language!

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.