The robust European language model benchmark
(formerly known as ScandEval)
Maintainer
- Dan Saattrup Smart (@saattrupdan, dan.smart@alexandra.dk)
Installation and usage
See the documentation for more information.
Reproducing the evaluation datasets
All datasets used in this project are generated using the scripts located in the src/scripts/dataset_creation folder. To reproduce a dataset, run the corresponding script with the following command
uv run src/scripts/dataset_creation/<name-of-script>.py
Replace with the specific script you wish to execute, e.g.,
uv run src/scripts/dataset_creation/create_allocine.py
Contributors 🙏
A huge thank you to all the contributors who have helped make this project a success!
Contribute to EuroEval
We welcome contributions to EuroEval! Whether you're fixing bugs, adding features, or contributing new datasets, your help makes this project better for everyone.
- General contributions: Check out our contribution guidelines for information on how to get started.
- Adding datasets: If you're interested in adding a new dataset to EuroEval, we have a dedicated guide with step-by-step instructions.
Special thanks
- Thanks to Google for sponsoring Gemini credits as part of their Google Cloud for Researchers Program.
- Thanks @Mikeriess for evaluating many of the larger models on the leaderboards.
- Thanks to OpenAI for sponsoring OpenAI credits as part of their Researcher Access Program.
- Thanks to UWV and KU Leuven for sponsoring the Azure OpenAI credits used to evaluate GPT-4-turbo in Dutch.
- Thanks to Miðeind for sponsoring the OpenAI credits used to evaluate GPT-4-turbo in Icelandic and Faroese.
- Thanks to CHC for sponsoring the OpenAI credits used to evaluate GPT-4-turbo in German.
Citing EuroEval
If you want to cite the framework then feel free to use this:
@inproceedings{smart2025encoder, title={Encoder vs decoder: Comparative analysis of encoder and decoder language models on multilingual NLU tasks}, author={Smart, Dan Saattrup and Enevoldsen, Kenneth and Schneider-Kamp, Peter}, booktitle={Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025)}, pages={561--572}, year={2025} } @inproceedings{smart2023scandeval, author = {Smart, Dan Saattrup}, booktitle = {Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa)}, month = may, pages = {185--201}, title = {{ScandEval: A Benchmark for Scandinavian Natural Language Processing}}, year = {2023} }
